Text error correction method and device after speech recognition, equipment and storage medium
By multimodal fusion of the text after speech recognition with high-dimensional speech features and fine-tuning the preset large model, the problem of insufficient accuracy of homophone word recognition is solved, and the overall accuracy and semantic understanding ability of speech recognition are improved.
Patent Information
- Application Number
- CN202510403821.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
AI Technical Summary
The existing speech recognition technology has insufficient accuracy in homophone recognition, resulting in poor semantic recognition effect and affecting user experience.
By converting the original text into a voice signal, using an automatic speech recognition system for recognition, combining high-dimensional speech features with text recognition content for multimodal fusion, fine-tuning the preset large model to correct the speech recognition results.
It improves the recognition accuracy of homophones, the error correction results are more in line with grammatical and semantic logic, reduces the cost of data labeling, and enhances the robustness of the model in noisy environments.
Smart Images

Figure CN120260573A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of text error correction, and particularly to a method, apparatus, device and storage medium for text error correction after speech recognition. Background Art
[0002] With the development of artificial intelligence technology, Automatic Speech Recognition (ASR) technology has also been continuously improving. Due to the interference of many factors, speech recognition results often contain errors, especially in homophones and grammatical structures.
[0003] Therefore, how to improve the accuracy of speech recognition results in terms of homophones, so as to improve the speech recognition effect, has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, the present disclosure provides a method, apparatus, device and storage medium for text error correction after speech recognition, so as to solve the problem of how to improve the accuracy of speech recognition results in terms of homophones, so as to improve the speech recognition effect.
[0005] On the one hand, the present disclosure provides a method for text error correction after speech recognition, the method comprising: performing text-to-speech processing on the original text to determine a speech signal; using an automatic speech recognition system to recognize the speech signal to determine text recognition content; combining the original text with the text recognition content into a training data set, and using the training data set to fine-tune a preset large model, including: using an automatic speech recognition system to determine high-dimensional speech features of the speech signal; performing multimodal fusion on the high-dimensional speech features and the text recognition content in the training data set as input data for the preset large model; fine-tuning the preset large model according to the input data and the original text in the training data set; receiving an input speech from a user, and using the fine-tuned preset large model to perform text error correction after speech recognition on the input speech to determine a text error correction result corresponding to the input speech.
[0006] On the other hand, the present disclosure also provides a text error correction device after speech recognition. The device includes: a first processing module for performing text-to-speech processing on the original text to determine a speech signal; a second processing module for using an automatic speech recognition system to recognize the speech signal to determine text recognition content; a pre-fine-tuning module for combining the original text and the text recognition content into a training data set and using the training data set to fine-tune a preset large model, where: the pre-fine-tuning module is specifically configured to use an automatic speech recognition system to determine high-dimensional speech features of the speech signal; perform multimodal fusion of the high-dimensional speech features and the text recognition content in the training data set as input data for the preset large model; fine-tune the preset large model according to the input data and the original text in the training data set; a text error correction module for receiving the user's input speech and using the fine-tuned preset large model to perform text error correction after speech recognition on the input speech to determine a text error correction result corresponding to the input speech.
[0007] On the other hand, the present disclosure also provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the above-mentioned text error correction method after speech recognition.
[0008] On the other hand, the present disclosure also provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to implement the above-mentioned text error correction method after speech recognition.
[0009] On the other hand, the present disclosure also provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the above-mentioned text error correction method after speech recognition.
[0010] Through the text error correction method, device, equipment and storage medium of the above embodiments of the present disclosure, by fine-tuning the preset large model through multimodal fusion of high-frequency speech features and text recognition content, the large model can be made to more easily distinguish between homophonic and heteromorphic words, thereby improving the recognition accuracy for homophonic and heteromorphic words.
[0011] In addition, by only performing text-to-speech processing and speech recognition processing on the original text, a training data set can be constructed, greatly reducing the data annotation cost. At the same time, the error correction result can also be made more in line with grammar and semantic logic, further improving the recognition accuracy for homophonic and heteromorphic words. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the specific embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0013] Figure 1 It is a flowchart of a method for text error correction after speech recognition provided by an embodiment of the present disclosure.
[0014] Figure 2 It is a specific flow chart of a method for text error correction after speech recognition provided by an embodiment of the present disclosure.
[0015] Figure 3 It is a structural schematic diagram of a text error correction device after speech recognition provided by an embodiment of the present disclosure.
[0016] Figure 4 It is a structural schematic diagram of another text error correction device after speech recognition provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0017] With the rapid development of artificial intelligence technology, ASR technology has been widely used in many fields such as voice assistants, intelligent customer service, and in-vehicle voice interaction. However, in actual use, due to interference from factors such as background noise and the speaker's accent, there are often many errors in the speech recognition results.
[0018] In related technologies, the error correction method for speech recognition results often focuses on the level of individual words, that is, it lacks consideration for sentence structure and semantic coherence. Especially for homophones, related technologies are more prone to recognition errors, resulting in insufficient accuracy of semantic recognition effects. For example, "high speed" may be recognized as "high beam", and then "turn on high speed mode" may be recognized as "turn on high beam mode". In speech recognition results, completely different meanings may be expressed due to the difference of one word. The recognition errors of homophones will seriously affect the user experience.
[0019] To solve the above problems, an embodiment of the present disclosure provides a method for correcting text after speech recognition. The method includes: performing text-to-speech processing on the original text to determine a speech signal; using an automatic speech recognition system to recognize the speech signal to determine text recognition content; combining the original text and the text recognition content into a training data set, and using the training data set to fine-tune a preset large model, including: using the automatic speech recognition system to determine high-dimensional speech features of the speech signal; performing multimodal fusion on the high-dimensional speech features and the text recognition content in the training data set as input data for the preset large model; fine-tuning the preset large model according to the input data and the original text in the training data set; receiving the user's input speech, and using the fine-tuned preset large model to correct the text after speech recognition of the input speech to determine a text correction result corresponding to the input speech.
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0021] In an embodiment of the present disclosure, a method for correcting text after speech recognition is provided. Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for correcting text after speech recognition provided by an embodiment of the present disclosure. The process of this method may include the following steps:
[0022] Step S101: Perform text-to-speech processing on the original text to determine a speech signal.
[0023] In this embodiment, the original text may be a standard text for generating a training data set, that is, the correct result of speech recognition.
[0024] Here, the original text may include various homophones in multiple professional fields and specific terms in each field. For example, the professional fields may include, but are not limited to: medicine, intelligent driving, and law.
[0025] Among them, the selection of the original text can be independently set by developers according to the specific application scenario of the preset large model, and no specific limitation is made here.
[0026] Further, after determining the original text, the following preprocessing operations can be performed on the original text:
[0027] Remove special symbols, uniformly represent the date format, uniformly represent synonyms, and split long and difficult sentences.
[0028] Further, the original text after the preprocessing operation is processed using Text-to-Speech (TTS) technology to determine the voice signal corresponding to the original text.
[0029] Here, after determining the voice signal, post-processing can also be performed on the voice data. For example, post-processing can include: audio noise reduction, data enhancement.
[0030] Step S102: Use an automatic speech recognition system to recognize the voice signal and determine the text recognition content.
[0031] In this embodiment, the automatic speech recognition system (ASR system) can refer to a system that converts the input voice signal into the corresponding text recognition content.
[0032] Among them, the process of recognizing the semantic signal can include steps such as determining audio features, converting word units, language modeling, and decoding. The ASR technology used for recognizing the semantic signal can be existing or future ASR technology, which is not specifically limited here.
[0033] The recognition result of the ASR system for the voice signal has the possibility of recognition errors.
[0034] Exemplarily, if the original text is "Turn off cruise control", the text recognition content may correspond to "Please turn off the loop control"; if the original text is "The patient has hepatitis", the text recognition content may correspond to "The patient has an infection".
[0035] Step S103: Combine the original text and the text recognition content into a training data set, and use the training data set to fine-tune a preset large model.
[0036] In this embodiment, the preset large model can be composed of an encoder in the ASR system and a Large Language Model (LLM). Among them, the LLM model includes a decoder.
[0037] Step S103 specifically includes:
[0038] Step S1031: Use an automatic speech recognition system to determine the high-dimensional voice features of the voice signal.
[0039] Among them, the high-dimensional voice features can be a multi-dimensional vector representation extracted from the voice signal and used to characterize the voice content. The high-dimensional voice features can characterize key contents such as the morpheme, intonation, frequency, and timing information of the voice.
[0040] Exemplarily, the high-dimensional voice features can include but are not limited to: Mel Frequency Cepstral Coefficients, Short-Time Fourier Transform.
[0041] Step S1032: Perform multimodal fusion on the high-dimensional speech features and the text recognition content in the training dataset as the input data for the preset large model.
[0042] Among them, multimodal fusion can refer to jointly modeling data of different modalities (such as speech, text, vision, etc.) so that the model can simultaneously understand information from multiple data sources to improve the performance of the task.
[0043] Specifically, multimodal fusion can be to fuse the high-dimensional speech features of the speech modality and the text recognition content of the text modality, so that the preset large model can combine speech features during text correction for more accurate text correction. For example, text correction can include: distinguishing homophones, supplementing missing words, and identifying repeated words.
[0044] Exemplarily, performing multimodal fusion on the high-dimensional speech features and the text recognition content in the training dataset can be achieved through at least one of the following steps:
[0045] Directly splice the text encoding of the high-dimensional speech features and the text recognition content; or,
[0046] Extract the text features of the high-dimensional speech features and the text recognition content, and map the two features into the same vector space for fusion; or,
[0047] Perform weighted fusion on the high-dimensional speech features and the text recognition content.
[0048] Step S1033: Fine-tune the preset large model according to the input data and the original text in the training dataset.
[0049] Among them, use the original text as the label of the training dataset to represent the correct answer; determine the difference comparison between the correct answer and the possibly incorrect recognition result, and calculate the loss value of the preset large model.
[0050] Fine-tune the preset large model according to the loss value until the preset large model meets the preset conditions, and determine the fine-tuned preset large model.
[0051] Step S104: Receive the input speech of the user, and use the fine-tuned preset large model to perform text correction on the text after speech recognition of the input speech, and determine the text correction result corresponding to the input speech.
[0052] In this embodiment, receive the input speech input by the user through the speech input device. Among them, there may be various errors in the input speech. For example, the input speech may include but is not limited to: homophones, missing words, and wrong words. The speech input device may include but is not limited to: smartphones, smart speakers.
[0053] For example, the user's input speech can be "Turn on high-speed mode", and the speech recognition result of the ASR system is "Turn on high-beam mode".
[0054] The preset large model after fine-tuning receives the user's input speech and the speech recognition result, corrects "high-beam" in the speech recognition result to "high-speed", and inputs the correct text correction result as "Turn on high-speed mode".
[0055] In a possible implementation, the preset large model can be customized and fine-tuned for different professional fields, making the preset large model more suitable for the recognition and correction of proper nouns in professional fields.
[0056] In another possible implementation, visual modality input can be added to the multi-modal input of the preset large model, and the preset large model can be fine-tuned visually, enabling the preset large model to correct visual texts such as subtitle texts.
[0057] In the text correction method, device, equipment, and storage medium after speech recognition in the above embodiments of the present disclosure, the preset large model is fine-tuned through the multi-modal fusion of high-frequency speech features and text recognition content, which can make the large model more easily distinguish homophonic and heteromorphic words, thereby improving the recognition accuracy for homophonic and heteromorphic words. By only performing text-to-speech processing and speech recognition processing on the original text, the training data set can be constructed, greatly reducing the data annotation cost. At the same time, the correction result can be made more compliant with grammar and semantic logic, further improving the recognition accuracy for homophonic and heteromorphic words.
[0058] In a possible implementation of step S103 above, step S103 further includes the following steps:
[0059] Step a1, using the original text as the training label and the text recognition content as the training input.
[0060] Among them, the training label (i.e., the above label) refers to the correct result or ideal output that guides the fine-tuning of the preset large model during the training process. The training input refers to the input value of the preset large model during the fine-tuning process.
[0061] During the fine-tuning process of the preset large model, the preset large model compares the predicted text output with the training label, calculates the loss value according to the loss function, and adjusts its own weights according to the loss value. Here, the loss value is the difference between the predicted text output and the training label.
[0062] Step a2, determining the training data set according to the training label and the training input.
[0063] Among them, the determination of the training dataset can be based on different speakers, speech types, and background noise levels to increase the diversity of training. For example, the speech types can include daily conversations and professional field conversations. The background noise levels can include: quiet level, relatively noisy level, and extremely noisy level.
[0064] In the text error correction method, device, equipment, and storage medium of the above embodiments of the present disclosure, traditional multi-modal models usually need to rely on multiple input sources and need to collect and synchronously process a large amount of data of different types. Different from this, this embodiment realizes multi-modal modeling only by relying on the original text, simplifies the model training process, and reduces the complexity and cost of data processing. By using the original text as the training label and combining the text recognition content as the training input, the preset large model can adjust its own parameters according to the loss value during the fine-tuning process, so that the model can more accurately identify the errors in the text and effectively correct them. Especially in the aspect of homophonic words, the model can infer the correct text based on the context and semantics, thereby improving the overall accuracy of speech recognition.
[0065] In a possible implementation manner of the above step S1031, step S1031 is implemented based on the following steps:
[0066] Step b1, use an automatic speech recognition system to receive the speech signal and convert the speech signal into a Log-Mel spectrogram.
[0067] Among them, perform a fast Fourier transform (Fast Fourier Transform, FFT) on each frame of the speech signal to convert the speech signal from a time-domain signal to a frequency-domain signal, that is, generate a spectrogram according to the FFT result of each frame.
[0068] Furthermore, use a Mel filter to filter the spectrogram to retain the frequency information related to human ear perception, convert the spectrogram from a linear frequency axis to a Mel frequency axis, and determine the Mel spectrogram.
[0069] Here, converting the spectrogram to the Mel frequency axis can capture the distribution characteristics of the sound in the Mel frequency. Among them, the Mel frequency can refer to a method of simulating the auditory characteristics of the human ear.
[0070] Even further, perform a logarithmic operation on the Mel spectrogram to convert it into a Log-Mel spectrogram. Among them, the Log-Mel spectrogram can be represented as a two-dimensional tensor, and its shape can be (T, F), where T represents the number of time frames and F represents the number of Mel frequency channels.
[0071] Here, the logarithmic conversion is to simulate the non-linear characteristics of the human ear in sound intensity perception.
[0072] Step b2: Use a one-dimensional convolutional neural network with a preset number of layers and the GELU activation function to extract local features from the Log-Mel spectrogram.
[0073] Among them, using a one-dimensional convolutional neural network with a preset number of layers can capture short-term local features in the time series.
[0074] The Gaussian Error Linear Unit (GELU) activation function can be a smooth non-linear activation function.
[0075] Step b3: Use a Transformer encoder to extract global context features from the Log-Mel spectrogram.
[0076] Among them, the global context features in the Log-Mel spectrogram can refer to the context information of the entire speech sequence.
[0077] Step b4: Perform a preset concatenation on the local features and the global context features to generate high-dimensional speech features of the speech signal.
[0078] Among them, perform a preset concatenation on the local features and the global context features in the feature dimension. Among them, the preset concatenation can include at least one of the following: direct concatenation, linear transformation fusion, and adaptive weighted fusion.
[0079] Here, the high-dimensional speech features after the preset concatenation may have different feature distributions and can be further normalized.
[0080] In the text error correction method, device, equipment, and storage medium of the above embodiments of the present disclosure, the Log-Mel spectrogram is more in line with human ear perception through Mel transformation and logarithmic operation, enhancing the representation ability of the speech signal, thereby improving the text error correction effect of the preset large model. Combining local and global features to improve the speech recognition error correction ability can better identify the errors of homophonic and heteromorphic words, make more accurate context judgments, and also improve the robustness in noisy environments.
[0081] In a possible implementation manner of the above step S1032, step S1032 is implemented based on the following steps:
[0082] Step c1: Split the text recognition content into multiple word units and map the multiple word units to corresponding embedding vectors.
[0083] In this embodiment, a subword segmentation algorithm is used to perform word segmentation on the text recognition content, and the sentence is split into multiple word units.
[0084] For example, split "Turn on high-speed mode" into "Turn on", "high-speed", and "mode".
[0085] Input multiple word units into the embedding layer, which maps them to corresponding embedding vectors. Among them, the embedding layer can use pre-trained word vectors or trainable vectors initialized randomly.
[0086] Step c2: Concatenate the high-dimensional speech features and the embedding vectors in a preset dimension for multi-modal fusion, and use the fused multi-modal input as the input data of a preset large model.
[0087] In this embodiment, the high-dimensional speech features and the embedding vectors are concatenated in a preset dimension to achieve multi-modal fusion, and the fused multi-modal features are used as the input.
[0088] In the text error correction method, device, equipment, and storage medium of the above embodiments of the present disclosure, by combining high-dimensional speech features and text embeddings, the model not only relies on text information but also can use speech features for error correction, thereby enhancing text semantic understanding, reducing recognition errors of homophones, and improving the accuracy of speech recognition. By combining high-dimensional speech features (such as pitch, speech rate, and emotional information), recognition errors caused by different speakers' accents and sentence breaks can be reduced.
[0089] In one possible implementation manner of the above step S1033, step S1033 is implemented based on the following steps:
[0090] Step d1: Send the input data to the large language model in the preset large model to obtain the predicted text determined by the large language model according to the input data.
[0091] Among them, the large language model uses a multi-layer Transformer encoder to extract the deep semantic information of the input data and generate the predicted text.
[0092] Step d2: Compare the predicted text with the original text to determine the loss value corresponding to the loss function of the preset large model.
[0093] Among them, a loss function can be used to calculate the error between the predicted text and the original text.
[0094] For example, the loss function can be a cross-entropy loss function.
[0095] Step d3: Determine the gradient according to the loss value and update the model parameters of the preset large model according to the gradient.
[0096] Among them, backpropagation is used to calculate the gradient of the loss function with respect to the model parameters and update the model parameters of the large model.
[0097] Step d4, if the loss value of the preset large model does not decrease within a preset number of training cycles, determine the preset large model after fine-tuning.
[0098] Among them, if the loss value of the preset large model no longer decreases in a preset number of iterations, stop fine-tuning the preset large model.
[0099] For example, the preset number can be 3, which is not specifically limited here.
[0100] In the method, apparatus, device, and storage medium for text error correction after speech recognition in the above embodiments of the present disclosure, the large model is fine-tuned by comparing the original text with the recognized text, so that it can better adapt to the speech recognition error pattern and improve the text error correction ability. Combining multi-modal fusion, the high-dimensional features of the speech signal and the text information are jointly input to improve the model's understanding ability of the speech input and further improve the accuracy of text error correction.
[0101] In a possible implementation manner of the above embodiment, the method further includes:
[0102] Combine the fused multi-modal input with a preset prompt, and use the combined preset prompt as the input data of the preset large model, where the preset prompt is used to instruct the preset large model to output the text error correction result corresponding to the text recognition content.
[0103] Among them, the preset prompt can be a text guidance input into the preset large model, which restricts the preset large model to output a text error correction result that meets the expectations.
[0104] Exemplarily, the preset prompt can be as follows: "You are a text error correction expert. Given the audio feature is xxx and the output result of ASR is xxx, only output the corrected text."
[0105] Among them, the known audio feature represents the high-dimensional speech feature, and the output result of ASR represents the text recognition content.
[0106] In a possible implementation manner, the professional field of the speech content can be added to the preset prompt. For example, "Please correct the following text according to medical terms."
[0107] In the method, apparatus, device, and storage medium for text error correction after speech recognition in the above embodiments of the present disclosure, guiding the large model to perform the text error correction task through the preset prompt and clarifying the output requirements of the large model can improve the recognition accuracy of the large model and the generalization ability to adapt to different scenarios.
[0108] In a specific embodiment, please refer to Figure 2 , Figure 2It is a schematic flowchart of a specific process of a text error correction method after speech recognition provided by an embodiment of the present disclosure. This process may include the following steps:
[0109] Step S201, convert the Log-Mel spectrogram.
[0110] Among them, an automatic speech recognition system is used to receive a speech signal and convert the speech signal into a Log-Mel spectrogram.
[0111] Step S202, extract local features from the Log-Mel spectrogram.
[0112] Among them, a one-dimensional convolutional neural network with a preset number of layers and a GELU activation function are used to extract local features from the Log-Mel spectrogram.
[0113] Step S203, the encoding layer extracts global context features from the Log-Mel spectrogram.
[0114] Among them, a Transformer encoder is used to extract global context features from the Log-Mel spectrogram.
[0115] Step S204, determine high-dimensional speech features.
[0116] Among them, the local features and the global context features are subjected to a preset splicing to generate high-dimensional speech features of the speech signal.
[0117] Step S205, receive text recognition content.
[0118] Step S206, multi-modal fusion to determine multi-modal input.
[0119] Among them, the high-dimensional speech features and the embedding vector are spliced and processed in a preset dimension for multi-modal fusion, and the fused multi-modal input is used as the input data of a preset large model.
[0120] Step S207, make a prediction on the multi-modal input.
[0121] Step S208, determine the predicted text.
[0122] In one embodiment, a text error correction device 300 after speech recognition is provided. The text error correction device 300 after speech recognition corresponds one-to-one with the text error correction method after speech recognition in the above embodiment. As Figure 3 shown, the text error correction device 300 after speech recognition includes:
[0123] The first processing module 301 is used to perform text-to-speech processing on the original text to determine a speech signal;
[0124] The second processing module 302 is used to identify the voice signal by using an automatic speech recognition system and determine the text recognition content;
[0125] The pre-fine-tuning module 303 is used to combine the original text and the text recognition content into a training data set and fine-tune a preset large model by using the training data set, where: the pre-fine-tuning module 303 is specifically used to determine the high-dimensional voice features of the voice signal by using an automatic speech recognition system; perform multi-modal fusion on the high-dimensional voice features and the text recognition content in the training data set as the input data of the preset large model; fine-tune the preset large model according to the input data and the original text in the training data set;
[0126] The text error correction module 304 is used to receive the input voice of the user, perform text error correction on the text after voice recognition of the input voice by using the fine-tuned preset large model, and determine the text error correction result corresponding to the input voice.
[0127] In one embodiment, the pre-fine-tuning module 303 is used to use the original text as a training label and the text recognition content as a training input;
[0128] Determine the training data set according to the training label and the training input.
[0129] In one embodiment, the pre-fine-tuning module 303 is used to receive the voice signal by using an automatic speech recognition system and convert the voice signal into a Log-Mel spectrogram;
[0130] Extract local features in the Log-Mel spectrogram by using a one-dimensional convolutional neural network with a preset number of layers and a GELU activation function;
[0131] Extract global context features in the Log-Mel spectrogram by using a Transformer encoder;
[0132] Perform a preset splicing on the local features and the global context features to generate high-dimensional voice features of the voice signal.
[0133] In one embodiment, the pre-fine-tuning module 303 is used to segment the text recognition content into multiple word units and map the multiple word units into corresponding embedding vectors;
[0134] Perform splicing processing on the high-dimensional voice features and the embedding vectors in a preset dimension for multi-modal fusion, and use the fused multi-modal input as the input data of the preset large model.
[0135] In one embodiment, the pre-fine-tuning module 303 is used to send the input data to a large language model in the preset large model and obtain the predicted text determined by the large language model according to the input data;
[0136] Compare the predicted text with the original text to determine the loss value corresponding to the loss function of the preset large model;
[0137] Determine the gradient according to the loss value, and update the model parameters of the preset large model according to the gradient;
[0138] If the loss value of the preset large model does not decrease within a preset number of training cycles, determine the fine-tuned preset large model.
[0139] In one embodiment, the pre-fine-tuning module 303 is further configured to combine the fused multi-modal input with a preset prompt, and use the combined preset prompt as the input data of the preset large model, where the preset prompt is used to instruct the preset large model to output the text correction result corresponding to the text recognition content.
[0140] It should be noted that: when implementing the corresponding text correction method after speech recognition provided by the above embodiment, only the above division of each program module is used for illustration. In practical applications, the above processing can be allocated to different program modules according to needs, that is, the internal structure of the above system is divided into different program modules to complete all or part of the above-described processing. In addition, the system provided by the above embodiment and the corresponding Figure 1 The embodiment of the method shown belongs to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be elaborated here.
[0141] The embodiment of the present disclosure also provides a computer device having the above Figure 3 Text correction device after speech recognition shown.
[0142] Please refer to Figure 4 , Figure 4 is a schematic structural diagram of another text correction device after speech recognition provided by the embodiment of the present disclosure. As Figure 4 shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways according to needs. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as a server array, a set of blade servers, or a multi-processor system). Figure 4 One processor 10 is taken as an example herein.
[0143] The processor 10 may be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 may further include a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device may be a complex programmable logic device, a field-programmable gate array, a generic array logic, or any combination thereof.
[0144] Among them, the memory 20 stores instructions that can be executed by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.
[0145] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may further include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0146] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memories.
[0147] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means, Figure 4 Taking connection through a bus as an example.
[0148] The input device 30 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (for example, an LED), and a tactile feedback device (for example, a vibration motor), etc. The above-mentioned display device includes, but is not limited to, a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.
[0149] The computer device further includes a communication interface for the computer device to communicate with other devices or communication networks.
[0150] Embodiments of the present disclosure also provide a computer-readable storage medium. The methods according to the embodiments of the present disclosure can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the methods described herein can be processed by such software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.
[0151] A part of the present disclosure can be applied as a computer program product, such as computer program instructions. When executed by a computer, through the operation of the computer, the methods and / or technical solutions according to the present disclosure can be called or provided. Those skilled in the art should be able to understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0152] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for correcting text after speech recognition, characterized in that, The method includes: Performing text-to-speech processing on the original text to determine a speech signal; Using an automatic speech recognition system to recognize the speech signal and determine the text recognition content; Combining the original text and the text recognition content into a training data set, and using the training data set to fine-tune a preset large model, including: using the automatic speech recognition system to determine the high-dimensional speech features of the speech signal; performing multimodal fusion on the high-dimensional speech features and the text recognition content in the training data set as the input data of the preset large model; fine-tuning the preset large model according to the input data and the original text in the training data set; Receiving the input speech of the user, and using the fine-tuned preset large model to correct the text after speech recognition of the input speech to determine the text correction result corresponding to the input speech.
2. The method according to claim 1, wherein The combining the original text and the text recognition content into a training data set includes: Using the original text as a training label and the text recognition content as a training input; Determining a training data set according to the training label and the training input.
3. The method according to claim 2, wherein Using the automatic speech recognition system to determine the high-dimensional speech features of the speech signal is implemented based on the following steps: Using the automatic speech recognition system to receive the speech signal and convert the speech signal into a Log-Mel spectrogram; Using a one-dimensional convolutional neural network with a preset number of layers and a GELU activation function to extract local features in the Log-Mel spectrogram; Using a Transformer encoder to extract global context features in the Log-Mel spectrogram; Performing a preset splicing on the local features and the global context features to generate the high-dimensional speech features of the speech signal.
4. The method according to claim 3, wherein Performing multimodal fusion on the high-dimensional speech features and the text recognition content in the training data set as the input data of the preset large model is implemented based on the following steps: Segmenting the text recognition content into multiple word units and mapping the multiple word units into corresponding embedding vectors; Performing a splicing process on the high-dimensional speech features and the embedding vectors in a preset dimension for multimodal fusion, and using the fused multimodal input as the input data of the preset large model.
5. The method according to any one of claims 1-4, characterized in that The fine-tuning the preset large model according to the input data and the original text in the training data set includes: Sending the input data to a large language model in the preset large model to obtain the predicted text determined by the large language model according to the input data; Comparing the predicted text with the original text to determine the loss value corresponding to the loss function of the preset large model; Determining a gradient according to the loss value and updating the model parameters of the preset large model according to the gradient; If the loss value of the preset large model does not decrease within a preset number of training cycles, determining the fine-tuned preset large model.
6. The method according to claim 4, characterized in that, The method further includes: Combine the fused multi-modal input with a preset prompt, and use the combined preset prompt as the input data for a preset large model, where the preset prompt is used to indicate the text correction result corresponding to the text recognition content output by the preset large model.
7. A text error correction device after speech recognition, characterized in that, The device includes: A first processing module, configured to perform text-to-speech processing on the original text to determine a voice signal; A second processing module, configured to use an automatic speech recognition system to recognize the voice signal to determine text recognition content; A pre-fine-tuning module, configured to combine the original text and the text recognition content into a training data set, and use the training data set to fine-tune a preset large model, where: the pre-fine-tuning module is specifically configured to use the automatic speech recognition system to determine high-dimensional voice features of the voice signal; perform multi-modal fusion on the high-dimensional voice features and the text recognition content in the training data set as the input data for the preset large model; and fine-tune the preset large model according to the input data and the original text in the training data set; A text correction module, configured to receive the input voice of a user, and use the fine-tuned preset large model to perform text correction after speech recognition on the input voice to determine the text correction result corresponding to the input voice.
8. A computer device, characterized in that, Includes: A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the text correction method after speech recognition according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause a computer to execute the text correction method after speech recognition according to any one of claims 1 to 6.
10. A computer program product, characterized in that, Includes computer instructions, and the computer instructions are used to cause a computer to execute the text correction method after speech recognition according to any one of claims 1 to 6.
Citation Information
Cited By
Speech recognition method, device and system, electronic equipment and storage medium
CN120913565A
Language recognition system, method and equipment combined with large language model error correction and medium
CN121306104A