Speech recognition model training method and device, speech recognition method and device, electronic equipment and computer readable storage medium
Patent Information
- Application Number
- CN202410218041.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-27
- Publication Date
- 2025-08-29
AI Technical Summary
In the prior art, the speech recognition model cannot effectively capture semantic information during training, resulting in poor recognition performance.
The initial speech recognition model extracts the speech sample features, obtains the speech sample features, and performs semantic extraction, determines the first loss value, combines the speech recognition results and the speech sample label, determines the second loss value, trains the initial speech recognition model, and obtains the speech recognition model.
Through the extraction of semantic features and determination of loss values, the semantic understanding performance and recognition accuracy of the speech recognition model are significantly improved, and the recognition performance of the speech recognition model is improved.
Smart Images

Figure CN120564699A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a training method for a speech recognition model, a speech recognition method, a device, an electronic device, and a computer-readable storage medium. Background Art
[0002] Artificial Intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0003] In related technologies, the training of speech recognition models usually involves performing speech recognition on speech samples through an initial speech recognition model to obtain a speech recognition result, directly using the speech recognition result to determine a loss value, and training the initial speech recognition model using the loss value to obtain a speech recognition model. This will result in the trained speech recognition model being unable to effectively capture semantic information due to the broad semantic expression in the speech samples, resulting in poor recognition performance of the speech recognition model. Summary of the Invention
[0004] The embodiments of the present application provide a speech recognition model training method, speech recognition method, device, electronic device and computer-readable storage medium, which can effectively improve the recognition performance of the speech recognition model.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] The present invention provides a method for training a speech recognition model, including:
[0007] The initial speech recognition model extracts features from the speech sample to obtain speech sample features;
[0008] Performing semantic extraction on the speech sample feature to obtain a semantic feature of the speech sample, and determining a first loss value based on the semantic feature;
[0009] The initial speech recognition model performs speech recognition on the speech sample based on the speech sample features to obtain a speech recognition result;
[0010] A second loss value is determined according to the speech recognition result and the speech sample label corresponding to the speech sample, and the initial speech recognition model is trained based on the first loss value and the second loss value to obtain the speech recognition model.
[0011] The present invention provides a speech recognition method, including:
[0012] The speech recognition model extracts features from the speech to be recognized to obtain the features of the speech to be recognized;
[0013] The speech recognition model performs speech recognition on the speech to be recognized based on the speech features to obtain a speech recognition result corresponding to the speech to be recognized;
[0014] The speech recognition model is obtained by training an initial speech recognition model based on the semantic features of the speech samples.
[0015] The present invention provides a training device for a speech recognition model, comprising:
[0016] Feature extraction module, used for the initial speech recognition model to extract features from speech samples and obtain speech sample features;
[0017] a semantic extraction module, configured to perform semantic extraction on the speech sample feature to obtain a semantic feature of the speech sample, and determine a first loss value based on the semantic feature;
[0018] A speech recognition module, configured to perform speech recognition on the speech sample based on the speech sample features of the initial speech recognition model to obtain a speech recognition result;
[0019] A training module is used to determine a second loss value according to the speech recognition result and the speech sample label corresponding to the speech sample, and to train the initial speech recognition model based on the first loss value and the second loss value to obtain the speech recognition model.
[0020] The present invention provides a speech recognition device, comprising:
[0021] Feature extraction module, used for speech recognition model to extract features of speech to be recognized and obtain features of speech to be recognized;
[0022] A speech recognition module is used for a speech recognition model to perform speech recognition on the speech to be recognized based on the speech features to be recognized, and obtain a speech recognition result corresponding to the speech to be recognized;
[0023] The speech recognition model is obtained by training an initial speech recognition model based on the semantic features of the speech samples.
[0024] An embodiment of the present application provides an electronic device, including:
[0025] a memory for storing computer-executable instructions or computer programs;
[0026] The processor is used to implement the training method of the speech recognition model and the speech recognition method provided in the embodiments of the present application when executing the computer-executable instructions or computer programs stored in the memory.
[0027] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for causing a processor to execute instructions to implement the speech recognition model training method and speech recognition method provided in the embodiment of the present application.
[0028] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the speech recognition model training method and speech recognition method described in the present invention.
[0029] The embodiments of the present application have the following beneficial effects:
[0030] The initial speech recognition model extracts features from the speech sample to obtain speech sample features, performs semantic extraction on the speech sample features to obtain semantic features of the speech sample, and determines a first loss value based on the semantic features. A second loss value is determined based on the speech recognition results and the speech sample labels. The initial speech recognition model is trained based on the first and second loss values to obtain a speech recognition model. In this way, by performing semantic extraction on the speech sample features, semantic features that can accurately reflect the semantics of the speech sample are obtained. The first loss value determined based on the semantic features can effectively improve the semantic understanding performance of the speech recognition model. The second loss value determined based on the speech recognition results and the speech sample labels corresponding to the speech samples can effectively improve the recognition accuracy of the speech recognition model. The speech recognition model is trained using the first and second loss values, so that the trained speech recognition model can effectively improve the semantic understanding performance and recognition accuracy, thereby effectively improving the recognition performance of the speech recognition model. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is a schematic diagram of the architecture of the speech recognition system provided in an embodiment of the present application;
[0032] Figure 2Schematic diagram of the structure of an electronic device for training a speech recognition model provided in an embodiment of the present application;
[0033] Figure 3 Schematic diagram of the structure of an electronic device for speech recognition provided by an embodiment of the present application;
[0034] Figure 4 Schematic diagram of the flow of the method for training a speech recognition model provided in an embodiment of the present application;
[0035] Figure 5 Schematic diagram of the flow of the speech recognition method provided in the embodiment of the present application;
[0036] Figure 6 This is a schematic diagram of the principle of the training method of the speech recognition model provided in the embodiment of the present application Figure 1 ;
[0037] Figure 7 This is a schematic diagram of the principle of the training method of the speech recognition model provided in the embodiment of the present application Figure 2 ;
[0038] Figure 8 It is a schematic diagram of the principle of the speech recognition method provided in the embodiment of the present application. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0040] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0041] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0043] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0044] 1) Artificial Intelligence (AI): This is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0045] 2) Convolutional Neural Networks (CNN): These are a type of feedforward neural network (FNN) that incorporates convolutional computations and possesses a deep structure. They are a representative algorithm for deep learning. CNNs possess representation learning capabilities and can perform shift-invariant classification on input images based on their hierarchical structure.
[0046] 3) Speech Recognition: Speech recognition is an interdisciplinary field encompassing signal processing, pattern recognition, probability theory and information theory, vocalization and auditory mechanisms, and artificial intelligence. Speech recognition is a high-tech process that enables machines to convert speech signals into corresponding text or commands through recognition and understanding. Speech recognition primarily encompasses three aspects: feature extraction, pattern matching, and model training.
[0047] 4) Machine Learning: Machine Learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0048] 5) End-to-end speech recognition: The goal of speech recognition is to convert the vocabulary in human speech into text. End-to-end speech recognition uses a pure neural network approach, replacing the traditional hybrid training of alignment models, acoustic models, and language models.
[0049] 6) Transformer: A time series model based on the self-attention mechanism. Its transformer component effectively encodes time series information, significantly outperforming LSTM in processing time series information and delivering high speed. It is widely used in natural language processing, computer vision, machine translation, speech recognition, and other fields.
[0050] 7) Conformer: This combines the Transformer and CNN models. The Transformer model excels at capturing content-based global interactions, while the CNN effectively utilizes local features. This allows the model to better model both long-term global interaction information and local features.
[0051] 8) CTC: This is a loss function used in sequence labeling. Traditional sequence labeling algorithms require perfect alignment of input and output symbols at every moment. CTC, however, expands the label set and adds empty elements. After annotating the sequence with the expanded label set, all predicted sequences that can be converted to true sequences through a mapping function are considered correct. This means that predicted sequences can be obtained without requiring data alignment.
[0052] During the implementation of the embodiments of this application, the applicant discovered that the related technology has the following problems:
[0053] In related technologies, the training of speech recognition models usually involves performing speech recognition on speech samples through an initial speech recognition model to obtain a speech recognition result, directly using the speech recognition result to determine a loss value, and training the initial speech recognition model using the loss value to obtain a speech recognition model. This will result in the trained speech recognition model being unable to effectively capture semantic information due to the broad semantic expression in the speech samples, resulting in poor recognition performance of the speech recognition model.
[0054] The embodiments of the present application provide a training method for a speech recognition model, a speech recognition method, an apparatus, an electronic device, and a computer-readable storage medium, which can effectively improve the recognition performance of the speech recognition model. The following describes an exemplary application of the speech recognition system provided by the embodiments of the present application.
[0055] See also Figure 1 , Figure 13 is a schematic diagram of the architecture of the speech recognition system provided in an embodiment of the present application. The terminal (terminal 400 is shown as an example) is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0056] The terminal 400 is used for the user to use the client 410 and display the speech recognition result on the graphical interface 410-1 (graphic interface 410-1 is shown as an example). The terminal 400 and the server 200 are connected to each other via a wired or wireless network.
[0057] In some embodiments, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart TV, a smart watch, a car terminal, etc., but is not limited to this. The electronic device provided in the embodiment of the present application can be implemented as a terminal or as a server. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiment of the present application.
[0058] In some embodiments, the server 200 extracts features from the speech sample through the initial speech recognition model to obtain speech sample features, performs speech recognition on the speech sample based on the speech sample features through the initial speech recognition model to obtain a speech recognition result, trains the initial speech recognition model based on the speech recognition result and the speech sample features to obtain a speech recognition model, and sends the speech recognition model to the terminal 400.
[0059] In other embodiments, the terminal 400 extracts features from the speech sample through the initial speech recognition model to obtain speech sample features, performs speech recognition on the speech sample based on the speech sample features through the initial speech recognition model to obtain a speech recognition result, trains the initial speech recognition model based on the speech recognition result and the speech sample features to obtain a speech recognition model, and sends the speech recognition model to the server 200.
[0060] In other embodiments, the embodiments of the present application can be implemented with the help of cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network within a wide area network or local area network to realize data calculation, storage, processing, and sharing.
[0061] Cloud technology is a general term for network, information, integration, management platform, and application technologies used in the cloud computing business model. It can form a resource pool that can be used flexibly and conveniently on demand. Cloud computing technology will become a key support. The backend services of technical network systems require a large amount of computing and storage resources.
[0062] See also Figure 2 , Figure 2 is a structural diagram of an electronic device for training a speech recognition model provided by an embodiment of the present application, wherein: Figure 2 The electronic device 500 shown may be Figure 1 The server 200 or the terminal 400 in Figure 2 The electronic device 500 shown includes: at least one processor 430, a memory 450, and at least one network interface 420. The various components in the electronic device 500 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 440 is not described in detail. Figure 2 Various buses are labeled as bus system 440 .
[0063] The processor 430 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0064] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 430.
[0065] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0066] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0067] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0068] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB).
[0069] In some embodiments, the speech recognition model training device provided in the embodiments of the present application can be implemented in software. Figure 2 A speech recognition model training device 455 stored in memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a feature extraction module 4551, a semantic extraction module 4552, a speech recognition module 4553, and a training module 4554. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.
[0070] See also Figure 3 , Figure 3 is a structural diagram of an electronic device for speech recognition provided by an embodiment of the present application, wherein: Figure 3 The electronic device 600 shown may be Figure 1 The server 200 or the terminal 400 in Figure 3 The electronic device 600 shown includes: at least one processor 530, a memory 550, and at least one network interface 520. The various components in the electronic device 600 are coupled together via a bus system 540. It is understood that the bus system 540 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 540 is not described in detail. Figure 3 Various buses are labeled as bus system 540 .
[0071] The processor 530 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0072] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 550 may optionally include one or more storage devices that are physically remote from the processor 530.
[0073] The memory 550 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.
[0074] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0075] Operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0076] The network communication module 552 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB).
[0077] In some embodiments, the speech recognition device provided in the embodiments of the present application can be implemented in software. Figure 3 A speech recognition device 555 stored in memory 550 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a feature extraction module 5551 and a speech recognition module 5552. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.
[0078] In other embodiments, the speech recognition device provided in the embodiments of the present application can be implemented in hardware. As an example, the speech recognition device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the speech recognition method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0079] In some embodiments, the terminal or server can implement the speech recognition method provided in the embodiment of the present application by running a computer program or computer executable instructions. For example, the computer program can be a native program (for example, a dedicated speech recognition program) or a software module in the operating system, for example, a speech recognition module that can be embedded in any program (such as an instant messaging client, a photo album program, an electronic map client, a navigation client); for example, it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run. In short, the above-mentioned computer program can be any form of application, module or plug-in.
[0080] The training method of the speech recognition model provided in the embodiment of the present application will be explained in combination with the exemplary application and implementation of the server or terminal provided in the embodiment of the present application.
[0081] See also Figure 4 , Figure 4 This is a flow chart of the training method of the speech recognition model provided in the embodiment of the present application, which will be combined with Figure 4 Steps 101 to 104 are shown for illustration. The training method of the speech recognition model provided in the embodiment of the present application can be implemented by the server or the terminal alone, or by the server and the terminal in collaboration. The following will be illustrated by taking the server alone as an example.
[0082] In step 101, the initial speech recognition model extracts features from the speech sample to obtain speech sample features.
[0083] In some embodiments, the initial speech recognition model includes a feature extraction layer and a speech recognition layer. The feature extraction layer is used to extract features from speech samples to obtain speech sample features, and the speech recognition layer is used to perform speech recognition on speech samples based on the speech sample features.
[0084] For example, see Figure 6 , Figure 6 This is a schematic diagram of the principle of the training method of the speech recognition model provided in the embodiment of the present application Figure 1 , Figure 6 The initial speech recognition model shown includes a feature extraction layer 1 and a speech recognition layer 2. The feature extraction layer 1 is used to extract features from speech samples to obtain speech sample features, and the speech recognition layer 2 is used to perform speech recognition on the speech samples based on the speech sample features.
[0085] In some embodiments, the feature extraction layer includes a coding layer, and the coding layer is used to encode the speech sample to obtain the speech sample features.
[0086] For example, see Figure 7 , Figure 7 This is a schematic diagram of the principle of the training method of the speech recognition model provided in the embodiment of the present application Figure 2 , the encoding layer 21 encodes the speech sample to obtain an encoding result; the linear layer 22 performs feature conversion on the encoding result to obtain speech sample features.
[0087] In some embodiments, before executing the above step 101, the voice sample can be determined in the following manner: obtain the voice to be processed within the target time period, and perform voice segment recognition on the voice to be processed to obtain the target voice segment in the voice to be processed; delete the target voice segment from the voice to be processed to obtain the voice sample.
[0088] In some embodiments, the speech to be processed includes speech segments corresponding to each target moment within the target time period.
[0089] As an example, the start playback time of the to-be-recognized voice A is 0s, and the end playback time is 10s. The to-be-processed voice A includes voice segments corresponding to each target time within the target time period (0s to 10s).
[0090] In some embodiments, the target speech segment is a speech segment in which no audio or noisy speech exists in the speech to be processed. The noisy speech refers to speech without semantic meaning.
[0091] In some embodiments, the above-mentioned voice segment recognition of the speech to be processed to obtain the target voice segment in the speech to be processed can be achieved as follows: if the audio does not exist in the speech segment, the speech segment is determined to be the target speech segment; if the audio exists in the speech segment, noise recognition is performed on the speech segment to obtain a noise recognition result; if the noise recognition result indicates that the noise exists in the speech segment, the speech segment is determined to be the target speech segment.
[0092] In some embodiments, the following processing may be performed on each voice segment: audio recognition is performed on the voice segment to obtain an audio recognition result, where the audio recognition result is used to indicate whether audio exists in the voice segment.
[0093] As an example, the speech A to be recognized includes speech segments A1, speech segments A2 and speech segments A3. If there is no audio in speech segment A1, then speech segment A1 is the target speech segment. If there is audio in speech segment A2 and there is noise in speech segment A2, then speech segment A2 is the target speech segment. If there is audio in speech segment A3 and there is no noise in speech segment A3, then speech segment A3 is not the target speech segment.
[0094] In this way, by acquiring speech to be processed within a target time period and performing speech segment recognition on the speech to be processed, a target speech segment is obtained from the speech to be processed; the target speech segment is then deleted from the speech to be processed to obtain the speech sample. This ensures that the resulting speech sample is free of noise and contains audio at every moment of the speech sample, effectively compressing the storage space for the speech to be processed, reducing the impact of invalid information on the training efficiency of the speech recognition model, and effectively improving the training efficiency of the speech recognition model.
[0095] In some embodiments, the above step 101 can be implemented as follows: performing speech signal processing on the speech sample to obtain frequency domain information of the speech sample; the initial speech recognition model performs feature extraction on the speech sample based on the frequency domain information to obtain the speech sample feature.
[0096] In some embodiments, the above-mentioned speech signal processing includes signal processing processes such as pre-emphasis, framing, windowing, discrete Fourier transform, and Mel filtering.
[0097] In some embodiments, the original speech signal is first passed through a high-pass filter to enhance the high-frequency portion of the speech, and the spectrum can be calculated using the same signal-to-noise ratio across the entire frequency range from low to high frequencies. The high-pass filter scheme used in the present invention is:
[0098] O(n)=k*x(n)-m*x(n-1) (1)
[0099] Among them, O(n) represents the result after high-pass filtering, n represents the audio sampling point index, x represents the specific audio, k and m represent the filter coefficients, k represents the ability to preserve the original high-frequency information, and m represents the ability to suppress the original high-frequency information. The larger K is, the smaller m is, which means that the ability to suppress the original high-frequency information is weaker. In the improvement task of the embodiment of the present application, the acoustic information needs to be combined with the language information. The high-pass filter designed in the embodiment of the present application performs a specific pre-emphasis operation on the audio, which can effectively control the degree of passage of high-frequency information. The sensitivity of language information is generally more important in high-frequency positions. Therefore, in the design of the improvement of the enhancement of language information on acoustic information, this controllable high-pass filter design can bring better results.
[0100] In some embodiments, framing is crucial for extracting fine-grained features, as language information needs to be used in conjunction with acoustic information. In this embodiment, the number of sampling points N after framing is set to 256, encompassing approximately 20ms of time. To avoid excessive spacing between adjacent frames, which can lead to insufficient feature granularity when jointly modeling speech and emotion recognition, an overlapped region between adjacent frames is created, encompassing 128 audio sampling points. This framing scheme improves the granularity of audio features when enhancing acoustic information with language information.
[0101] In some embodiments, for windowing, after the audio is framed, each frame needs to be windowed to increase the continuity of the left and right ends of the frame, because the feature extraction process is essentially a discrete representation of the continuous audio signal, and this discrete representation should be able to express continuous information as much as possible to reduce the spectrum leakage of the extracted audio features. When the language information and acoustic information are deeply integrated, they are more sensitive to this continuity representation. If spectrum leakage occurs, a large error will occur when judging the language information, which will affect the final representation of the acoustic information. Therefore, this fine-grained judgment task requires continuous information before and after to ensure the effect. In view of this, the embodiment of the present application designs a window function, which is expressed as follows:
[0102]
[0103] Where N represents the total number of sampling points, n represents the current sampling point, and O(n) represents the windowed output. This windowing function effectively correlates the discrete information of each frame. This is because the weighted sin function establishes a functional relationship between each frame through specific weighting. When using language information to enhance acoustic information, it allows for a relatively limited continuous representation of discrete information, improving the final effect.
[0104] In some embodiments, the discrete Fourier transform (DFT) is used to transform the signal only in the time domain, making it difficult to discern the true characteristics of the audio signal. Therefore, after framing, a discrete Fourier transform (DFT) is performed to convert the signal into an energy distribution in the frequency domain, thereby characterizing different audio characteristics. This approach involves multiplication in the time domain and convolution in the frequency domain.
[0105] In some embodiments, for the Mel filter, the role of applying the Mel filter is to map the current spectrum into a Mel nonlinear spectrum that conforms to the human ear's perception, and then convert it to the inverse spectrum. After passing through the Mel filter, the human ear can establish a linear mapping relationship between the real audio feature perception and the discrete signal. In order to improve the acoustic information effect improvement scheme based on the language information designed in the embodiment of the present application, a set of trapezoidal filters are designed, with a center frequency of f(m) = 1, 2, 3...M, and the value of M used in the present invention is 22. The trapezoidal filter can retain the original information as much as possible in the low-amplitude part, and transform the high-amplitude part into an information representation that conforms to the human ear. Thereby, the output of the acoustic model is more in line with the prediction requirements of the language model, conforms to the improvement strategy proposed in the embodiment of the present application, and improves the effect.
[0106] In step 102, semantic extraction is performed on the features of the speech sample to obtain semantic features of the speech sample.
[0107] In some embodiments, the semantic extraction can be implemented by an auxiliary coding layer, which is used to extract semantic features from features. Figure 7 , the auxiliary coding layer 23 performs semantic extraction on the speech sample features to obtain the semantic features of the speech sample.
[0108] In some embodiments, the above-mentioned speech sample features include multiple sub-sample features, and the above-mentioned step 102 can be implemented in the following manner: for each of the sub-sample features, semantic recognition is performed on the sub-sample feature to obtain a semantic recognition result; if the semantic recognition result indicates that the sub-sample feature has semantics, the sub-sample feature is determined as the target sample feature; if the number of the target sample features is one, the target sample feature is determined as the semantic feature; if the number of the target sample features is multiple, the target sample features are fused to obtain the semantic feature.
[0109] As an example, speech sample feature B includes subsample feature B1, subsample feature B2, and subsample feature B3. If the semantic recognition result of subsample feature B1 indicates that the subsample feature has semantics, subsample feature B1 is determined as the target sample feature. If the semantic recognition result of subsample feature B2 indicates that the subsample feature does not have semantics, subsample feature B2 is not determined as the target sample feature. If the semantic recognition result of subsample feature B3 indicates that the subsample feature has semantics, subsample feature B3 is determined as the target sample feature. If there are multiple target sample features, target sample feature B3 is fused with target sample feature B1 to obtain the semantic feature.
[0110] In step 103, a first loss value is determined based on the semantic feature.
[0111] In some embodiments, before executing step 103 above, a fusion feature may be obtained by performing feature fusion on the semantic feature and the speech sample feature to obtain a fusion feature.
[0112] In some embodiments, the above step 103 can be implemented as follows: obtaining the semantic label feature corresponding to the speech sample, and determining the first similarity between the semantic label feature and the semantic feature; determining the first loss value based on the fusion feature and the first similarity.
[0113] In some embodiments, the first similarity is negatively correlated with the feature distance between the semantic tag feature and the semantic feature.
[0114] In some embodiments, the above-mentioned determination of the first loss value based on the fusion feature and the first similarity can be achieved as follows: feature extraction is performed on the speech sample label to obtain a speech label feature, and the speech label feature and the semantic label feature are feature-fused to obtain a target label feature; the second similarity between the target label feature and the fusion feature is determined, and the first loss value is determined based on the first similarity and the second similarity.
[0115] In some embodiments, the second similarity is negatively correlated with the feature distance between the semantic tag feature and the semantic feature.
[0116] In some embodiments, determining the first loss value based on the first similarity and the second similarity can be achieved by performing a weighted summation of the first similarity and the second similarity to obtain the first loss value.
[0117] For example, see Figure 7 , Figure 7The att loss shown in can be the second similarity mentioned above, Figure 7 The auxiliary att loss shown in can be the first similarity mentioned above.
[0118] In some embodiments, the weights of the weighted sums corresponding to the first similarity and the second similarity can be specifically set according to actual conditions.
[0119] In some embodiments, the above-mentioned determination of the first loss value based on semantic features can be achieved in the following manner: the initial speech recognition model performs speech recognition on the speech sample based on the fusion feature to obtain a first recognition result; the initial speech recognition model performs speech recognition on the speech sample based on the semantic feature to obtain a second recognition result; and the first loss value is determined based on the first recognition result and the second recognition result.
[0120] In some embodiments, the above-mentioned determination of the first loss value based on the first recognition result and the second recognition result can be achieved by: determining a third similarity between the first recognition result and the voice sample label, and determining a fourth similarity between the second recognition result and the voice sample label; and determining the first loss value based on the third similarity and the fourth similarity.
[0121] For example, see Figure 7 , Figure 7 The att loss shown in can be the third similarity mentioned above, Figure 7 The auxiliary att loss shown in can be the fourth similarity.
[0122] In some embodiments, determining the first loss value based on the third similarity and the fourth similarity can be achieved by performing a weighted summation of the third similarity and the fourth similarity to obtain the first loss value.
[0123] In this manner, by determining the third similarity between the first recognition result and the speech sample label, and determining the fourth similarity between the second recognition result and the speech sample label, and determining the first loss value based on the third and fourth similarities, the first loss value is effectively determined. The first loss value determined by the first and second recognition results effectively enhances the sample size of the initial speech recognition model trained, thereby effectively improving the recognition performance of the speech recognition model trained based on the first loss value.
[0124] In step 104, the initial speech recognition model performs speech recognition on the speech sample based on the speech sample features to obtain a speech recognition result.
[0125] In some embodiments, the initial speech recognition model includes a speech recognition layer. Step 104 can be implemented as follows: the speech recognition layer of the initial speech recognition model performs speech recognition on the speech sample based on the speech sample features to obtain a speech recognition result.
[0126] For example, see Figure 7 , Figure 7 The decoding layer shown in is the above-mentioned speech recognition layer. The decoding layer 24 of the initial speech recognition model performs speech recognition on the speech sample based on the speech sample features to obtain a speech recognition result.
[0127] In step 105, a second loss value is determined according to the speech recognition result and the speech sample label corresponding to the speech sample.
[0128] In some embodiments, the above-mentioned determination of the second loss value based on the speech recognition result and the speech sample label corresponding to the speech sample can be achieved by: determining the fifth similarity between the speech recognition result and the speech sample label, and performing feature extraction on the speech sample label to obtain a speech label feature; determining the sixth similarity between the voice label feature and the speech sample feature, and determining the second loss value based on the fifth similarity and the sixth similarity.
[0129] In some embodiments, determining the second loss value based on the fifth similarity and the sixth similarity can be achieved by performing a weighted summation of the fifth similarity and the sixth similarity to obtain the second loss value.
[0130] For example, see Figure 7 , the fifth similarity can be Figure 6 The att loss shown in , the sixth similarity can be Figure 7 The ctc loss shown in .
[0131] In this way, the fifth similarity between the speech recognition result and the speech sample label is determined, and feature extraction is performed on the speech sample label to obtain a speech label feature; the sixth similarity between the speech label feature and the speech sample feature is determined, and the second loss value is determined based on the fifth similarity and the sixth similarity. The fifth similarity between the speech recognition result and the speech sample label can accurately reflect the difference between the speech recognition result and the speech sample label, thereby evaluating the overall performance of the initial speech recognition model. The sixth similarity between the speech label feature and the speech sample feature can accurately reflect the difference between the speech sample feature and the speech label feature, thereby evaluating the performance of the feature extraction layer of the initial speech recognition model, so that the second loss value determined by the fifth similarity and the sixth similarity can improve the recognition performance of the initial speech recognition model from the overall performance and the feature extraction performance of the initial speech recognition model.
[0132] In step 106, the initial speech recognition model is trained based on the first loss value and the second loss value to obtain the speech recognition model.
[0133] In some embodiments, the above step 106 can be implemented as follows: obtain a first weight corresponding to the first loss value, and a second weight corresponding to the second loss value; multiply the first loss value and the first weight to obtain a first target loss value, and multiply the second loss value and the second weight to obtain a second target loss value; add the first target loss value and the second target loss value to obtain a target loss value, and based on the target loss value, train the initial semantic recognition model to obtain the speech recognition model.
[0134] As an example, the expression of the above target loss value can be:
[0135] L=α1L1+α2L2 (2)
[0136] Among them, L is used to indicate the target loss value, L1 is used to indicate the first loss value, L2 is used to indicate the second loss value, α1 is used to indicate the first weight corresponding to the first loss value, and L2 is used to indicate the second weight corresponding to the second loss value.
[0137] In this way, the initial speech recognition model performs feature extraction on the speech sample to obtain speech sample features, performs semantic extraction on the speech sample features to obtain semantic features of the speech sample, and determines a first loss value based on the semantic features. A second loss value is determined based on the speech recognition results and the speech sample labels. The initial speech recognition model is trained based on the first loss value and the second loss value to obtain a speech recognition model. In this way, by performing semantic extraction on the speech sample features, semantic features that can accurately reflect the semantics of the speech sample are obtained. The first loss value determined based on the semantic features can effectively improve the semantic understanding performance of the speech recognition model. The second loss value determined based on the speech recognition results and the speech sample labels corresponding to the speech samples can effectively improve the recognition accuracy of the speech recognition model. Therefore, by training the speech recognition model using the first loss value and the second loss value, the speech recognition model obtained by training can effectively improve the semantic understanding performance and recognition accuracy of the speech recognition model, thereby effectively improving the recognition performance of the speech recognition model.
[0138] See also Figure 5 , Figure 5 This is a flow chart of the speech recognition method provided by the embodiment of the present application, which will be combined with Figure 5 Steps 201 to 202 are shown for illustration. The training method of the speech recognition model provided in the embodiment of the present application can be implemented by the server or the terminal alone, or by the server and the terminal in collaboration. The following will be illustrated by taking the implementation by the terminal alone as an example.
[0139] In step 201, the speech recognition model extracts features of the speech to be recognized to obtain features of the speech to be recognized.
[0140] In some embodiments, the speech recognition model includes a feature extraction layer, and step 201 can be implemented as follows: the feature extraction layer of the speech recognition model extracts features of the speech to be recognized to obtain features of the speech to be recognized.
[0141] For example, see Figure 6 , Figure 6 The feature extraction layer 1 of the speech recognition model shown extracts features from the speech to be recognized to obtain features of the speech to be recognized.
[0142] In step 202, the speech recognition model performs speech recognition on the speech to be recognized based on the speech features to be recognized, and obtains a speech recognition result corresponding to the speech to be recognized.
[0143] In some embodiments, the speech recognition model is obtained by training an initial speech recognition model based on semantic features of speech samples.
[0144] In some embodiments, the above-mentioned speech recognition model includes a speech recognition layer, and the above-mentioned step 201 can be implemented in the following manner: the speech recognition layer of the speech recognition model performs speech recognition on the speech to be recognized based on the speech features to be recognized, and obtains a speech recognition result corresponding to the speech to be recognized.
[0145] In this manner, the initial speech recognition model extracts features from the speech sample to obtain speech sample features, performs semantic extraction on the speech sample features to obtain semantic features of the speech sample, and determines a first loss value based on the semantic features. A second loss value is determined based on the speech recognition results and the speech sample labels, and the initial speech recognition model is trained based on the first and second loss values to obtain a speech recognition model. Thus, by performing semantic extraction on the speech sample features, semantic features that accurately reflect the semantics of the speech sample are obtained. The first loss value determined based on the semantic features can effectively improve the semantic understanding performance of the speech recognition model. The second loss value determined based on the speech recognition results and the speech sample labels corresponding to the speech samples can effectively improve the recognition accuracy of the speech recognition model. Thus, by training the speech recognition model based on the first and second loss values, the trained speech recognition model can effectively improve both semantic understanding performance and recognition accuracy, thereby effectively improving the recognition performance of the speech recognition model. The trained speech recognition model performs speech recognition on the speech to be recognized based on the speech features to be recognized, resulting in more accurate speech recognition results for the speech to be recognized.
[0146] Below, an exemplary application of the embodiment of the present application in an actual speech recognition application scenario will be described.
[0147] In existing end-to-end speech recognition, deeper networks often have stronger generalization capabilities. Currently, the optimal end-to-end speech recognition method based on neural networks is the conformer-transformer architecture. The conformer component often uses the CTC loss for sequence alignment and loss calculation. The output of the conformer component is input into the transformer component, which uses sequence loss for modeling. The conformer is responsible for acoustic feature mapping, while the transformer is responsible for modeling semantic information. However, due to its inherent independence assumption, the conformer is unable to model semantic information for acoustic sequences, resulting in lower conformer performance. Although the introduction of the transformer can enable the output of the final layer of the conformer to have a certain degree of semantic information modeling capabilities, due to the excessive depth of the network layers, the intermediate layers of the network have difficulty capturing semantic information. Moreover, the combination of the transformer and the conformer often uses a method of calculating losses separately, which is difficult to effectively improve the semantic relevance of the conformer.
[0148] The specific approach of the embodiment of the present application is to add another auxiliary transformer to the conformer-transformer structure. The auxiliary transformer is not used during prediction, but only during training. On the one hand, the output of the auxiliary transformer is output to the attention loss function. On the other hand, according to the number of output frames of the conformer output, the output of the auxiliary transformer is repeated at the frame level to adapt to the output of the conformer. Then, the output of the conformer and the output of the auxiliary transformer are merged together, and the output dimension is changed from 256 to 512. The vector converted to 512 is passed through a linear layer and then changed back to 256 and input into the CTC loss. In this way, an auxiliary semantic CTC loss function is added. The input of this loss function is the merged output of the conformer and the auxiliary transformer, so that the conformer has more semantic information for deep feature fusion, thereby improving the semantic modeling capability of the conformer model and improving the accuracy of the model.
[0149] The specific approach of the embodiment of the present application is to add another auxiliary transformer to the conformer-transformer structure. The auxiliary transformer is not used during prediction. Only during training, the output of the auxiliary transformer is output to the attention loss function on the one hand. On the other hand, according to the number of output frames of the conformer output, the output of the auxiliary transformer is repeated at the frame level to adapt to the output of the conformer. Then, the output of the conformer and the output of the auxiliary transformer are merged together, and the output dimension is changed from 256 to 512. Then, the vector changed to 512 is passed through a linear layer and then changed back to 256 and input into the CTC loss. In this way, an auxiliary semantic CTC loss function is added. The input of this loss function is the merged output of the conformer and the auxiliary transformer, so that the conformer has more semantic information for deep feature fusion, improves the semantic modeling ability of the conformer model, and improves the accuracy of the model. The embodiment of the present application also improves the CTC loss to match the improvement measures of the improvement of acoustic features by language features.
[0150] In addition, the embodiment of the present application also explores an improvement strategy for the fbank feature, with the goal of enabling the extracted features to better capture fine-grained information that fits the combination of transformer language and conformer acoustics, thereby achieving the effect of improving the final effect.
[0151] In some embodiments, see Figure 7 ,like Figure 7As shown, the specific execution can be divided into the following steps: For speech recognition data processing: 50,000 hours of pure speech recognition data are indexed to obtain the text index corresponding to each speech, and the data is used as speech recognition pre-training data; for 50,000 hours of speech data, first follow up the results of manual annotation, perform audio segmentation, filter out silent audio segments, and mark non-human noise or unclear human noise, and segment the speech. The segmentation time range is between 0.8-2.0s. This time range is because most noises in real life occur within this interval. This is used as the UNK label. During training, let the model also learn this UNK label, so as to avoid the model learning both the noise label and the silent label as blank in CTC. Then, when the model is output, the UNK result is filtered out by regular matching, which can effectively improve the model's noise resistance. In the embodiment of the present application, the conformer part needs to be connected to the output of the transformer, that is, the acoustic information needs to be combined with the language information. When the model is not robust to noise, it is easy to mistakenly recognize noise as text, and the misrecognized text often has confusing semantic information. If this confusing semantic information is combined with the acoustic information, it will cause problems in training. Therefore, this approach is also to match the improvement plan of the conformer module of the auxiliary transformer and improve the noise robustness of the model.
[0152] In some embodiments, for feature extraction: the feature extraction process includes pre-emphasis, framing, windowing, discrete Fourier transform, and Mel filtering.
[0153] In some embodiments, the original speech signal is first passed through a high-pass filter to enhance the high-frequency portion of the speech, and the spectrum can be calculated using the same signal-to-noise ratio across the entire frequency range from low to high frequencies. The high-pass filter scheme used in the present invention is:
[0154] O(n)= k*x(n)-m*x(n-1) (3)
[0155] Among them, O(n) represents the result after high-pass filtering, n represents the audio sampling point index, x represents the specific audio, k and m represent the filter coefficients, K represents the ability to preserve the original high-frequency information, and m represents the ability to suppress the original high-frequency information. The larger K is, the smaller m is, which means that the ability to suppress the original high-frequency information is weaker. In the improvement task of the embodiment of the present application, the acoustic information needs to be combined with the language information. The high-pass filter designed in the embodiment of the present application performs a specific pre-emphasis operation on the audio, which can effectively control the degree of passage of high-frequency information. The sensitivity of language information is generally more important in high-frequency positions. Therefore, in the design of the improvement of the enhancement of language information on acoustic information, this controllable high-pass filter design can bring better results.
[0156] In some embodiments, regarding framing: Given that language information needs to be used in conjunction with acoustic information in the embodiments of the present application, the extracted fine-grained features are very important. In the embodiments of the present application, the value of the sampling point N after framing is set to 256, which includes a time of approximately 20ms. At the same time, in order to avoid a large interval between two adjacent frames, which results in insufficient feature granularity when jointly modeling speech and emotion recognition, an area between two adjacent frames is overlapped, and the overlapping area includes 128 audio sampling points. The above-mentioned framing scheme is used to improve the audio feature granularity when the language information strengthens the acoustic information;
[0157] In some embodiments, regarding windowing: after the audio is framed, each frame needs to be windowed to increase the continuity of the left and right ends of the frame, because the feature extraction process is essentially a discrete representation of the continuous audio signal, and this discrete representation should be able to express continuous information as much as possible to reduce the spectrum leakage of the extracted audio features. When the language information and acoustic information are deeply integrated, they are more sensitive to this continuity representation. If spectrum leakage occurs, a large error will occur when judging the language information, which will affect the final representation of the acoustic information. Therefore, this fine-grained judgment task requires continuous information before and after to ensure the effect. In view of this, the embodiment of the present application designs a window function, which is expressed as follows:
[0158]
[0159] Where N represents the total number of sampling points, n represents the current sampling point, and O(n) represents the windowed output. This windowing function effectively correlates the discrete information of each frame. This is because the weighted sin function establishes a functional relationship between each frame through specific weighting. When using language information to enhance acoustic information, it allows for a relatively limited continuous representation of discrete information, improving the final effect.
[0160] In some embodiments, regarding the discrete Fourier transform: Since the signal is only transformed in the time domain, it is difficult to see the characteristics of the actual audio signal. Therefore, after framing, a discrete Fourier transform is also performed to convert it into an energy distribution in the frequency domain, thereby characterizing different audio characteristics. The Fourier transform is not improved in this embodiment of the application, that is, time domain multiplication and frequency domain convolution.
[0161] In some embodiments, for the Mel filter: the role of applying the Mel filter is to map the current spectrum into the Mel nonlinear spectrum that conforms to the human ear's perception, and then convert it to the inverse spectrum. After passing through the Mel filter, the human ear can establish a linear mapping relationship between the real audio feature perception and the discrete signal. In order to improve the acoustic information effect improvement scheme based on the language information designed in the embodiment of the present application, a set of trapezoidal filters are designed, with a center frequency of f(m) = 1, 2, 3...M, and the value of M used in the present invention is 22. The trapezoidal filter can retain the original information as much as possible in the low-amplitude part, and transform the high-amplitude part into an information representation that conforms to the human ear. Thereby, the output of the acoustic model is more in line with the prediction requirements of the language model, conforms to the improvement strategy proposed in the embodiment of the present application, and improves the effect.
[0162] In the embodiment of the present application, another auxiliary transformer is added to the conformer-transformer structure. The auxiliary transformer is not used during prediction. Only during training, the output of the auxiliary transformer is output to the attention loss function on the one hand. On the other hand, according to the number of output frames of the conformer output, the output of the auxiliary transformer is repeated at the frame level to adapt to the output of the conformer. Then, the output of the conformer and the output of the auxiliary transformer are merged together, and the output dimension is changed from 256 to 512. Then, the vector changed to 512 is passed through a linear layer and then changed back to 256 and input into the CTC loss. In this way, an auxiliary semantic CTC loss function is added. The input of this loss function is the merged output of the conformer and the auxiliary transformer, so that the conformer has more semantic information for deep feature fusion, improves the semantic modeling ability of the conformer model, and improves the accuracy of the model. The embodiment of the present application also improves the CTC loss to match the improvement of acoustic features by language features. And the training process is optimized.
[0163] The embodiment of the present application carries out multiple multi-step training in sequence, ensuring that the model of each step forms a model representation with stable parameters and good performance based on the basic data, thereby laying a good foundation for the final combination of language information and acoustic information.
[0164] First, use the above 60,000 hours of data to pre-train the 12-layer conformer and 6-layer conformer models for speech recognition, and extract the above-mentioned modified fbank features. Since this is the pre-training stage of the speech recognition model, a larger learning rate is used for the experiment. At the same time, the trained optimizer is optimized based on Adam. Since the input fbank features are improved through fine-grained adjustment, a small amount of parameter changes during training can easily cause gradient oscillations. Therefore, the embodiment of the present application modifies the Adam optimizer. The specific implementation method is to add a new L2 regularization term when updating the parameters on the basis of the original Adam, and remove the L2 regularization term added during the gradient update in the original Adam. This can reduce the violent parameter oscillations caused by gradient changes, thereby improving the speed of speech recognition training and the model optimization effect. The loss combines CTC loss and attention loss. The present invention also improves the CTC loss. Specifically, the improvement replaces the idea of sequence path optimization through dynamic programming in CTC with a backtracking algorithm. The reason for the change is that although the dynamic programming scheme can perform sequence search relatively quickly, the speech recognition training process essentially requires obtaining all possible sequences. The backtracking algorithm can effectively obtain all sequence solutions and obtain the optimal result among all sequence solutions. Although this solution increases training time, it improves accuracy.
[0165] In each round of training, the above features are input into the conformer in batches. The conformer here can be a conformer, or a transformer, LSTM, TDNN, RNN-T, etc. The output of the conformer represents the acoustic feature vector, and the dimension is generally 16x the number of frames x 256; this output is first input into the linear layer of the improved CTC, converted into 16x the number of frames x the number of text classification indices, and then input into the improved CTC loss to calculate the improved CTC loss.
[0166] The above output is fed into a transformer, which can use either a transformer or an LSTM. The transformer output is 16 times the number of characters in the speech x 256. The specific training method is to predict the next character based on the previous character. The principle is described here: https: / / blog.csdn.net / weixin_45193103 / article / details / 124002864?ydreferer=aHR0cHM6Ly9jbi5iaW5nLmNvbS8%3D. The transformer output is then fed into a linear and softmax function. The output dimension is 16 times the number of characters in the speech x the number of character classification indices. The Attention-To-Text (ATT) loss is then calculated, typically using the KL divergence loss. The CTC loss from step 4 and the attention loss from step 5 are combined to calculate the total loss. Training is then continued until the loss converges to a stable state, resulting in Model 1.
[0167] Load the generated model 1 and add an auxiliary transformer branch after the conformer. Use randomly initialized parameters for this branch. Freeze the original conformer and transformer structures. Then, input the output of the conformer to the auxiliary transformer. The output of the auxiliary transformer is then fed into the attention loss until the loss converges. Save the model to obtain model 2.
[0168] By loading model 2, the auxiliary transformer output is repeatedly processed at the frame level to adapt it to the output of the conformer. The output of the conformer and the auxiliary transformer are then merged, reducing the output dimension from 256 to 512. This 512-dimensional vector is then passed through a linear layer, back to 256, and input into the improved CTC loss. This creates a new auxiliary semantic improved CTC loss function. This loss function takes the combined output of the conformer and auxiliary transformer as input, providing the conformer with more semantic information for deep feature fusion, improving the model's semantic modeling capabilities and boosting model accuracy. During training, the learning rate is reduced to one-tenth of that in step 5, and the auxiliary semantic improved CTC loss, acoustic improved CTC loss, and auxiliary transformer attention loss are merged. The model is saved until the losses converge. This completes the model training phase.
[0169] The semantically improved CTC loss of the encoding layer output and the auxiliary transformer output is called the semantically improved CTC loss LCTC_LM_improved, and the loss of the auxiliary decoding layer is called the auxiliary decoding label smoothing loss (i.e., attention loss) LATT_AUX. The expression of LCTC_LM is as follows:
[0170] LCTC_LM_improved = -logPCTC(y|W*concat(xenc,xdec_aux_repeat)+b) (4)
[0171] Where y represents the label, xenc represents the output of the conformer, xdec_aux_repeat represents the vector of the auxiliary transformer output after frame-level replication, W and b represent the weight and bias of the linear layer, respectively. P represents the ctc label probability.
[0172] The total loss of the encoder is expressed as follows:
[0173] Lenc_total=(1-q)*LCTC_improved+q*LCTC_LM_improved (5)
[0174] The value of q is 0.2, which can be determined based on experiments. The parameter is not unique. The expression of LATT_AUX is as follows:
[0175] LATT_AUX=-logPATT(y|x) (6)
[0176] Among them, y represents the label, x represents the input, that is, the output of the encoder, and P represents the attention label probability.
[0177] The total loss of the transformer is expressed as follows:
[0178] LATT_total=β*LATT+(1-β)*LATT_middle (7)
[0179] The value of β is 0.7, which can be determined based on experiments. The parameter is not unique.
[0180] During training, the total loss is expressed as follows:
[0181] Ltotal=k*Lenc_total+(1-k)*LATT_total (8)
[0182] In some embodiments, see Figure 8 , Figure 8It is a schematic diagram of the principle of the speech recognition method provided by the embodiment of the present application. In the model application stage, when the above-mentioned speech recognition model is applied, the above-mentioned conformer and transformer model parameters are loaded, and the auxiliary transformer is not loaded. When the speech input comes in, the improved 80-dimensional fbank features are extracted, and after model reasoning, the model search score is obtained, and the recognition result is output. When the algorithm is applied, the speech model with good optimization effect can be directly applied to the scene to be recognized, such as speech quality inspection, voice robot, etc. An auxiliary semantic CTC loss function is newly added. The input of this loss function is the combined output of the conformer and the auxiliary transformer, so that the conformer has more semantic information for deep feature fusion, which improves the semantic modeling ability of the conformer and improves the accuracy of the model. In this way, the accuracy of the model is effectively improved without any impact on efficiency. The embodiment of the present application selects the intelligent outbound call scenario for specific explanation.
[0183] In the embodiment of the present application, another auxiliary transformer is added to the conformer-transformer structure. The auxiliary transformer is not used during prediction. Only during training, the output of the auxiliary transformer is output to the attention loss function. On the other hand, according to the number of output frames of the conformer output, the output of the auxiliary transformer is repeated at the frame level to adapt to the output of the conformer. Then, the output of the conformer and the output of the auxiliary transformer are merged together, and the output dimension is changed from 256 to 512. Then, the vector changed to 512 is passed through a linear layer and then changed back to 256 and input into the CTC loss. In this way, an auxiliary semantic CTC loss function is added. The input of this loss function is the merged output of the conformer and the auxiliary transformer, so that the conformer has more semantic information for deep feature fusion, improves the semantic modeling ability of the conformer model, and improves the accuracy of the model. It effectively alleviates the shortcomings of the CTC in the original encoder that only captures the independence hypothesis information, effectively improves the accuracy of the model without affecting efficiency.
[0184] The embodiment of this application explores an improvement strategy for fbank features, with the goal of enabling the extracted features to better capture fine-grained information that fits the combination of transformer language and conformer acoustics, thereby achieving the effect of improving the final effect.
[0185] In the conformer-transformer structure, another auxiliary transformer is added. The auxiliary transformer is not used during prediction. Only during training, the output of the auxiliary transformer is output to the attention loss function on the one hand. On the other hand, according to the number of output frames of the conformer output, the output of the auxiliary transformer is repeated at the frame level to adapt to the output of the conformer. Then, the output of the conformer and the output of the auxiliary transformer are merged together, and the output dimension is changed from 256 to 512. Then, the vector changed to 512 is passed through a linear layer and then changed back to 256 and input into the CTC loss. In this way, an auxiliary semantic CTC loss function is added. The input of this loss function is the merged output of the conformer and the auxiliary transformer, so that the conformer has more semantic information for deep feature fusion, improves the semantic modeling ability of the model conformer, and improves the accuracy of the model. The embodiment of the present application also improves the CTC loss to match the improvement of acoustic features by language features. And the training process is optimized.
[0186] The embodiment of the present application carries out multiple multi-step training in sequence, ensuring that the model of each step forms a model representation with stable parameters and good performance based on the basic data, thereby laying a good foundation for the final combination of language information and acoustic information.
[0187] It is understandable that in the embodiments of the present application, when data related to the voice to be recognized is involved, when the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0188] The following continues to describe the exemplary structure of the speech recognition model training device 455 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 2As shown, the software modules in the training device 455 of the speech recognition model stored in the memory 450 may include: a feature extraction module, which is used for the initial speech recognition model to extract features from the speech sample to obtain speech sample features; a semantic extraction module, which is used to perform semantic extraction on the speech sample features to obtain semantic features of the speech sample, and determine a first loss value based on the semantic features; a speech recognition module, which is used for the initial speech recognition model to perform speech recognition on the speech sample based on the speech sample features to obtain a speech recognition result; a training module, which is used to determine a second loss value based on the speech recognition result and the speech sample label corresponding to the speech sample, and train the initial speech recognition model based on the first loss value and the second loss value to obtain the speech recognition model.
[0189] In some embodiments, the training device of the above-mentioned speech recognition model further includes: a speech processing module, which is used to obtain the speech to be processed within a target time period, and perform speech segment recognition on the speech to be processed to obtain a target speech segment in the speech to be processed; wherein, the target speech segment is a speech segment in which there is no audio or noisy speech in the speech to be processed; the target speech segment is deleted from the speech to be processed to obtain the speech sample.
[0190] In some embodiments, the above-mentioned speech processing module is also used to determine the speech segment as the target speech segment if the audio does not exist in the speech segment; if the audio exists in the speech segment, perform noise recognition on the speech segment to obtain a noise recognition result; if the noise recognition result indicates that the noise exists in the speech segment, determine the speech segment as the target speech segment.
[0191] In some embodiments, the feature extraction module is further used to perform speech signal processing on the speech sample to obtain frequency domain information of the speech sample; the initial speech recognition model performs feature extraction on the speech sample based on the frequency domain information to obtain the speech sample feature.
[0192] In some embodiments, the above-mentioned speech sample features include multiple sub-sample features, and the above-mentioned semantic extraction module is further used to perform semantic recognition on each sub-sample feature to obtain a semantic recognition result. If the semantic recognition result indicates that the sub-sample feature has semantics, the sub-sample feature is determined as the target sample feature; if the number of the target sample features is one, the target sample feature is determined as the semantic feature; if the number of the target sample features is multiple, the target sample features are fused to obtain the semantic feature.
[0193] In some embodiments, the training device for the speech recognition model further includes: a fusion module for fusing the semantic features and the speech sample features to obtain fused features.
[0194] In some embodiments, the above-mentioned semantic extraction module is also used to obtain the semantic label feature corresponding to the speech sample, and determine the first similarity between the semantic label feature and the semantic feature; and determine the first loss value based on the fusion feature and the first similarity.
[0195] In some embodiments, the above-mentioned semantic extraction module is also used for the initial speech recognition model to perform speech recognition on the speech sample based on the fusion feature to obtain a first recognition result; the initial speech recognition model to perform speech recognition on the speech sample based on the semantic feature to obtain a second recognition result; and determine the first loss value based on the first recognition result and the second recognition result.
[0196] In some embodiments, the semantic extraction module is further used to determine a third similarity between the first recognition result and the voice sample label, and to determine a fourth similarity between the second recognition result and the voice sample label; and to determine the first loss value based on the third similarity and the fourth similarity.
[0197] In some embodiments, the above-mentioned training module is also used to determine the fifth similarity between the speech recognition result and the speech sample label, and perform feature extraction on the speech sample label to obtain a speech label feature; determine the sixth similarity between the speech label feature and the speech sample feature, and determine the second loss value based on the fifth similarity and the sixth similarity.
[0198] In some embodiments, the above-mentioned training module is also used to obtain a first weight corresponding to the first loss value and a second weight corresponding to the second loss value; multiply the first loss value and the first weight to obtain a first target loss value, and multiply the second loss value and the second weight to obtain a second target loss value; add the first target loss value and the second target loss value to obtain a target loss value, and train the initial semantic recognition model based on the target loss value to obtain the speech recognition model.
[0199] The following continues to describe the exemplary structure of the speech recognition device 555 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 4As shown, the software modules stored in the speech recognition device 555 of the memory 550 may include: a feature extraction module, which is used for the speech recognition model to extract features of the speech to be recognized and obtain the features of the speech to be recognized; a speech recognition module, which is used for the speech recognition model to perform speech recognition on the speech to be recognized based on the features of the speech to be recognized and obtain the speech recognition result corresponding to the speech to be recognized; wherein, the speech recognition model is obtained by training the initial speech recognition model based on the semantic features of the speech sample.
[0200] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the speech recognition model training method and speech recognition method described in the present invention.
[0201] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will execute the training method and speech recognition method of the speech recognition model provided in the embodiment of the present application.
[0202] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various electronic devices including one or any combination of the above memories.
[0203] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0204] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0205] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0206] In summary, the embodiments of the present application have the following beneficial effects:
[0207] (1) Extracting features from the speech sample using the initial speech recognition model to obtain speech sample features, extracting semantics from the speech sample features to obtain semantic features of the speech sample, and determining a first loss value based on the semantic features. Determining a second loss value based on the speech recognition results and the speech sample labels. Training the initial speech recognition model based on the first loss value and the second loss value to obtain a speech recognition model. In this way, by extracting semantics from the speech sample features, semantic features that can accurately reflect the semantics of the speech sample are obtained. The first loss value determined by the semantic features can effectively improve the semantic understanding performance of the speech recognition model. The second loss value determined based on the speech recognition results and the speech sample labels corresponding to the speech samples can effectively improve the recognition accuracy of the speech recognition model. The speech recognition model is trained using the first loss value and the second loss value, so that the speech recognition model obtained by training can effectively improve the semantic understanding performance and recognition accuracy, thereby effectively improving the recognition performance of the speech recognition model.
[0208] (2) The initial speech recognition model is used to extract features from the speech sample to obtain speech sample features, and semantic extraction is performed on the speech sample features to obtain semantic features of the speech sample, and a first loss value is determined based on the semantic features. A second loss value is determined based on the speech recognition results and the speech sample labels, and the initial speech recognition model is trained based on the first loss value and the second loss value to obtain a speech recognition model. In this way, by performing semantic extraction on the speech sample features, semantic features that can accurately reflect the semantics of the speech sample are obtained. The first loss value determined by the semantic features can effectively improve the semantic understanding performance of the speech recognition model. The second loss value determined based on the speech recognition results and the speech sample labels corresponding to the speech samples can effectively improve the recognition accuracy of the speech recognition model. Thus, the speech recognition model is trained by the first loss value and the second loss value, so that the speech recognition model obtained by training can effectively improve the semantic understanding performance and recognition accuracy, thereby effectively improving the recognition performance of the speech recognition model. The speech recognition model obtained by training performs speech recognition on the speech to be recognized based on the speech features to be recognized, so that the speech recognition result corresponding to the speech to be recognized is more accurate.
[0209] (3) The embodiment of the present application carries out multiple multi-step training in sequence, ensuring that the model of each step forms a model representation with stable parameters and good performance based on the basic data, thereby laying a good foundation for the final language information and acoustic information.
[0210] (4) The embodiment of the present application explores an improvement strategy for fbank features, with the goal of enabling the extracted features to better capture fine-grained information that fits the transformer language and conformer acoustic basis, thereby achieving the effect of improving the final effect.
[0211] (5) In the embodiment of the present application, another auxiliary transformer is added to the conformer-transformer structure. The auxiliary transformer is not used during prediction. Only during training, the output of the auxiliary transformer is output to the attention loss function. On the other hand, according to the number of output frames of the conformer output, the output of the auxiliary transformer is repeatedly operated at the frame level to adapt to the output of the conformer. Then, the output of the conformer and the output of the auxiliary transformer are merged together, and the output dimension is changed from 256 to 512. Then, the vector changed to 512 is passed through a linear layer and then changed back to 256 and input into the CTC loss. In this way, an auxiliary semantic CTC loss function is added. The input of this loss function is the merged output of the conformer and the auxiliary transformer, so that the conformer has more semantic information for deep feature fusion, improves the semantic modeling ability of the conformer model, and improves the accuracy of the model. It effectively alleviates the shortcomings of the CTC in the original encoder that only captures the independence hypothesis information, effectively improves the accuracy of the model without affecting the efficiency.
[0212] (6) Determine the fifth similarity between the speech recognition result and the speech sample label, and perform feature extraction on the speech sample label to obtain a speech label feature; determine the sixth similarity between the speech label feature and the speech sample feature, and determine the second loss value based on the fifth similarity and the sixth similarity. The fifth similarity between the speech recognition result and the speech sample label can accurately reflect the difference between the speech recognition result and the speech sample label, thereby evaluating the overall performance of the initial speech recognition model. The sixth similarity between the speech label feature and the speech sample feature can accurately reflect the difference between the speech sample feature and the speech label feature, thereby evaluating the performance of the feature extraction layer of the initial speech recognition model, thereby enabling the second loss value determined by the fifth similarity and the sixth similarity to improve the recognition performance of the initial speech recognition model in terms of overall performance and feature extraction performance.
[0213] (7) Determining a third similarity between the first recognition result and the speech sample label, and determining a fourth similarity between the second recognition result and the speech sample label; and determining the first loss value based on the third similarity and the fourth similarity. The first loss value determined by the first recognition result and the second recognition result effectively enhances the sample size of the initial speech recognition model, thereby effectively improving the recognition performance of the speech recognition model trained based on the first loss value.
[0214] (8) Determining a third similarity between the first recognition result and the speech sample label, and determining a fourth similarity between the second recognition result and the speech sample label; and determining the first loss value based on the third similarity and the fourth similarity. The first loss value determined by the first recognition result and the second recognition result effectively enhances the sample size of the initial speech recognition model, thereby effectively improving the recognition performance of the speech recognition model trained based on the first loss value.
[0215] (9) For the Mel filter, the role of applying the Mel filter is to map the current spectrum into the Mel nonlinear spectrum that is consistent with the human ear's perception, and then convert it to the inverse spectrum. After passing through the Mel filter, the human ear can establish a linear mapping relationship between the real audio feature perception and the discrete signal. In order to improve the acoustic information effect improvement scheme designed for the language information in the embodiment of the present application, a set of trapezoidal filters are designed, with a center frequency of f(m) = 1, 2, 3...M, and the value of M used in the present invention is 22. The trapezoidal filter can retain the original information as much as possible in the low-amplitude part, and transform the high-amplitude part into an information representation that is consistent with the human ear. As a result, the output of the acoustic model is more in line with the prediction requirements of the language model, in line with the improvement strategy proposed in the embodiment of the present application, and improves the effect.
[0216] (10) A target speech segment in the speech to be processed is obtained by acquiring speech to be processed within a target time period and performing speech segment recognition on the speech to be processed; the target speech segment is deleted from the speech to be processed to obtain the speech sample. As a result, the obtained speech sample is free of noise and audio exists at every moment of the speech sample, thereby effectively compressing the storage space of the speech to be processed, effectively reducing the impact of invalid information on the training efficiency of the speech recognition model, and effectively improving the training efficiency of the speech recognition model.
[0217] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A method for training a speech recognition model, characterized in that: The method comprises: The initial speech recognition model extracts features from the speech sample to obtain speech sample features; Performing semantic extraction on the speech sample feature to obtain a semantic feature of the speech sample, and determining a first loss value based on the semantic feature; The initial speech recognition model performs speech recognition on the speech sample based on the speech sample features to obtain a speech recognition result; A second loss value is determined according to the speech recognition result and the speech sample label corresponding to the speech sample, and the initial speech recognition model is trained based on the first loss value and the second loss value to obtain the speech recognition model.
2. The method according to claim 1, characterized in that Before the initial speech recognition model extracts features from the speech sample to obtain speech sample features, the method further includes: Acquire the speech to be processed within a target time period, and perform speech segment recognition on the speech to be processed to obtain a target speech segment in the speech to be processed; The target speech segment is a speech segment in which there is no audio or noisy speech in the speech to be processed; The target speech segment is deleted from the speech to be processed to obtain the speech sample.
3. The method according to claim 1, characterized in that The speech to be processed includes speech segments corresponding to each target moment within the target time period; and performing speech segment recognition on the speech to be processed to obtain target speech segments in the speech to be processed includes: If the audio does not exist in the voice segment, determining the voice segment as the target voice segment; If the audio exists in the voice segment, performing noise recognition on the voice segment to obtain a noise recognition result; If the noise recognition result indicates that the noise exists in the speech segment, the speech segment is determined as the target speech segment.
4. The method according to claim 1, wherein The initial speech recognition model extracts features from the speech sample to obtain speech sample features, including: Performing speech signal processing on the speech sample to obtain frequency domain information of the speech sample; The initial speech recognition model performs feature extraction on the speech sample based on the frequency domain information to obtain the speech sample features.
5. The method according to claim 1, wherein The speech sample feature includes a plurality of sub-sample features, and the semantic extraction of the speech sample feature to obtain the semantic feature of the speech sample includes: For each of the sub-sample features, perform semantic recognition on the sub-sample feature to obtain a semantic recognition result, and if the semantic recognition result indicates that the sub-sample feature has semantics, determine the sub-sample feature as a target sample feature; If the number of the target sample feature is one, determining the target sample feature as the semantic feature; If the number of the target sample features is greater than one, the target sample features are fused to obtain the semantic feature.
6. The method according to claim 1, wherein Before determining the first loss value based on the semantic feature, the method further includes: fusing the semantic feature with the speech sample feature to obtain a fused feature.
7. The method according to claim 6, characterized in that The determining of a first loss value based on the semantic feature includes: Obtaining a semantic label feature corresponding to the speech sample, and determining a first similarity between the semantic label feature and the semantic feature; Determine the first loss value according to the fusion feature and the first similarity.
8. The method according to claim 7, characterized in that The determining the first loss value according to the fusion feature and the first similarity includes: Extracting features from the speech sample labels to obtain speech label features, and fusing the speech label features with the semantic label features to obtain target label features; Determine a second similarity between the target label feature and the fusion feature, and determine the first loss value according to the first similarity and the second similarity.
9. The method according to claim 6, characterized in that The determining of a first loss value based on the semantic feature includes: The initial speech recognition model performs speech recognition on the speech sample based on the fusion feature to obtain a first recognition result; The initial speech recognition model performs speech recognition on the speech sample based on the semantic features to obtain a second recognition result; The first loss value is determined according to the first recognition result and the second recognition result.
10. The method according to claim 9, characterized in that The determining the first loss value according to the first recognition result and the second recognition result includes: Determining a third similarity between the first recognition result and the voice sample label, and determining a fourth similarity between the second recognition result and the voice sample label; The first loss value is determined according to the third similarity and the fourth similarity.
11. The method according to claim 1, wherein The determining a second loss value according to the speech recognition result and the speech sample label corresponding to the speech sample includes: Determining a fifth similarity between the speech recognition result and the speech sample label, and performing feature extraction on the speech sample label to obtain a speech label feature; A sixth similarity between the voice tag feature and the voice sample feature is determined, and the second loss value is determined according to the fifth similarity and the sixth similarity.
12. The method according to claim 1, characterized in that The training of the initial semantic recognition model based on the first loss value and the second loss value to obtain the speech recognition model includes: Obtaining a first weight corresponding to the first loss value and a second weight corresponding to the second loss value; Multiplying the first loss value by the first weight to obtain a first target loss value, and multiplying the second loss value by the second weight to obtain a second target loss value; The first target loss value and the second target loss value are added to obtain a target loss value, and the initial semantic recognition model is trained based on the target loss value to obtain the speech recognition model.
13. A speech recognition method, characterized in that: The method comprises: The speech recognition model extracts features from the speech to be recognized to obtain the features of the speech to be recognized; The speech recognition model performs speech recognition on the speech to be recognized based on the speech features to obtain a speech recognition result corresponding to the speech to be recognized; The speech recognition model is obtained by training an initial speech recognition model based on the semantic features of the speech samples.
14. A training device for a speech recognition model, characterized in that: The device comprises: Feature extraction module, used for the initial speech recognition model to extract features from speech samples and obtain speech sample features; a semantic extraction module, configured to perform semantic extraction on the speech sample feature to obtain a semantic feature of the speech sample, and determine a first loss value based on the semantic feature; A speech recognition module, configured to perform speech recognition on the speech sample based on the speech sample features of the initial speech recognition model to obtain a speech recognition result; A training module is used to determine a second loss value based on the speech recognition result and the speech sample label corresponding to the speech sample, and to train the initial speech recognition model based on the first loss value and the second loss value to obtain the speech recognition model.
15. A speech recognition device, characterized in that: The device comprises: Feature extraction module, used for speech recognition model to extract features of speech to be recognized and obtain features of speech to be recognized; A speech recognition module is used for a speech recognition model to perform speech recognition on the speech to be recognized based on the speech features to be recognized, and obtain a speech recognition result corresponding to the speech to be recognized; The speech recognition model is obtained by training an initial speech recognition model based on the semantic features of the speech samples.
16. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; A processor, configured to implement the method according to any one of claims 1 to 13 when executing computer-executable instructions or computer programs stored in the memory.
17. A computer-readable storage medium storing computer-executable instructions, characterized in that: When the computer-executable instructions are executed by a processor, the method according to any one of claims 1 to 13 is implemented.
Citation Information
Cited By
Speech recognition optimization method and system based on adaptive dynamic programming
CN121483231A