Voice error correction method and system based on artificial intelligence
By using a context verification model based on attention mechanism in the speech conversion system, the semantic fit between the pinyin sequence and the Chinese character sequence is scored, which solves the problem of inaccurate speech correction caused by polyphonic characters, improves the accuracy of the speech conversion results, and is suitable for terminal applications such as electronic devices and automobiles, improving the user experience.
Patent Information
- Application Number
- CN202510234170.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-28
Smart Images

Figure CN120146035A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voice interaction technology, and in particular to a voice error correction method and system based on artificial intelligence. Background Art
[0002] Polyphonetic characters refer to a Chinese character that has multiple pronunciations in different semantic scenarios (e.g., “重” is pronounced as zhòng or chóng). Its correct pronunciation and the corresponding Chinese character selection are highly dependent on the context. Traditional methods face the following core challenges when solving polyphonetic character matching:
[0003] 1. Limitations of static vocabulary
[0004] It relies on predefined dictionaries (e.g., "行" in "银行" is fixed as háng), and cannot dynamically adapt to unregistered words or complex contexts (e.g., "行" in "行代码" should be read as háng, but read as xíng in "走").
[0005] 2. Insufficient modeling of short-distance dependencies
[0006] The N-gram language model only captures local word order (such as the first 2 to 3 words) and has difficulty understanding long-distance semantic logic (such as the difference in pronunciation of the two characters “重” in “He re-weighed the goods”).
[0007] 3. The ambiguity of ambiguous scenarios
[0008] When the context is not sufficient to clearly point to a single pronunciation (e.g., “好” in “this person is good at talking” can be pronounced as hǎo or hào), traditional models lack the ability to make deep semantic inferences.
[0009] Therefore, the prior art still has the problem that the speech error correction is not accurate enough due to the problem of polyphones when converting speech to text, resulting in low accuracy of the speech input conversion result. Summary of the invention
[0010] In view of the deficiencies in the prior art, the purpose of the present invention is to provide a speech error correction method and system based on artificial intelligence, aiming to solve the problem in the prior art that speech error correction is not accurate enough due to the problem of polyphonetic characters when converting speech to text, resulting in low accuracy of speech input conversion results.
[0011] A first aspect of the present invention is to provide a speech error correction method based on artificial intelligence, the method comprising:
[0012] Acquire the voice input information input at the current time node, and output the corresponding pinyin sequence according to the voice input information;
[0013] Import the pinyin sequence into a preset word database, and traverse multiple Chinese character sequences corresponding to the pinyin sequence in the word database. The multiple Chinese character sequences are homophonic words;
[0014] Use a context verification model based on the attention mechanism to score the semantic fitness of multiple Chinese character sequences and the pinyin sequence respectively, and output the semantic score value of each Chinese character sequence;
[0015] Determine the Chinese character sequence with the highest semantic score value as the target Chinese character sequence, and obtain the target Chinese character text corresponding to the voice input information.
[0016] According to one aspect of the above technical solution, the step of using a context verification model based on the attention mechanism to score the semantic fitness of multiple Chinese character sequences and the pinyin sequence respectively, and output the semantic score value of each Chinese character sequence includes:
[0017] Convert the Chinese character sequence and the pinyin sequence into embedding vectors respectively, and add position encoding correspondingly;
[0018] Through the attention mechanism of the context verification model, make the Chinese character sequence and the pinyin sequence pay attention to each other, and generate context-related target features;
[0019] Perform feature fusion and pooling processing on the target features, and extract the key information of the target features;
[0020] Map the fused target features through the fully connected layer in the context verification model to obtain the semantic score value of each Chinese character sequence.
[0021] According to one aspect of the above technical solution, the step of making the Chinese character sequence and the pinyin sequence pay attention to each other through the attention mechanism of the context verification model, and generating context-related target features includes:
[0022] Through the context verification model, calculate the cross-attention weights between the Chinese character sequence and the pinyin sequence according to the multi-head attention mechanism, and capture the semantic association between the Chinese character sequence and the pinyin sequence;
[0023] Take the Chinese character sequence as the Query vector, take the pinyin sequence as the Key vector and the Value vector, generate an attention matrix, and use scaled dot attention to measure the local correlation between Chinese characters and pinyin;
[0024] Perform weighted summation on the cross-attention weights and the Value vector of the pinyin vector to generate the context-enhanced representation of the Chinese character sequence, and obtain the target features.
[0025] According to one aspect of the above technical solution, the step of mapping the fused target features through the fully connected layer in the context verification model to obtain the semantic score value of each Chinese character sequence includes:
[0026] According to at least one fully connected layer in the context verification model, map the target features to an intermediate dimension through a weight matrix and a bias term;
[0027] Use the activation function of linear output through the output layer in the context verification model to output the score corresponding to the Chinese character sequence, and obtain the semantic score value corresponding to each Chinese character sequence.
[0028] According to one aspect of the above technical solution, the method further includes:
[0029] Use a pre-trained deep learning model to perform background denoising on the voice input information to enhance the voice clarity of the voice input information;
[0030] Perform endpoint detection on the voice input information, determine the start and end points of the voice segment in the voice input information, and perform cutting according to the start and end points of the voice segment to obtain the target voice segment.
[0031] According to one aspect of the above technical solution, after the step of performing endpoint detection on the voice input information, determining the start and end points of the voice segment in the voice input information, and performing cutting according to the start and end points of the voice segment to obtain the target voice segment, it further includes:
[0032] Output the corresponding pinyin sequence according to the target voice segment.
[0033] The second aspect of the present invention lies in providing an artificial intelligence-based speech error correction system, which is applied to the method described in the above technical solution. The system includes:
[0034] A pinyin sequence output module, configured to obtain the voice input information input at the current time node, and output the corresponding pinyin sequence according to the voice input information;
[0035] A Chinese character sequence traversal module, configured to import the pinyin sequence into a preset word database, and traverse multiple Chinese character sequences corresponding to the pinyin sequence in the word database. The multiple Chinese character sequences are homophonic words;
[0036] A Chinese character sequence scoring module, configured to use a context verification model based on an attention mechanism to score the semantic fitness of multiple Chinese character sequences and the pinyin sequence respectively, and output the semantic score value of each Chinese character sequence;
[0037] A Chinese character text output module, which is used to determine the Chinese character sequence with the highest semantic score value as the target Chinese character sequence, and obtain the target Chinese character text corresponding to the voice input information.
[0038] According to one aspect of the above technical solution, the Chinese character sequence scoring module is specifically used for:
[0039] Convert the Chinese character sequence and the pinyin sequence into embedding vectors respectively, and add position encoding correspondingly;
[0040] Through the attention mechanism of the context verification model, make the Chinese character sequence and the pinyin sequence pay attention to each other, and generate context-related target features;
[0041] Perform feature fusion and pooling processing on the target features, and extract the key information of the target features;
[0042] Map the fused target features through the fully connected layer in the context verification model to obtain the semantic score value of each Chinese character sequence.
[0043] The third aspect of the present invention lies in providing a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in the above technical solution is implemented.
[0044] The fourth aspect of the present invention lies in providing an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in the above technical solution is implemented.
[0045] Compared with the prior art, the beneficial effects of adopting the voice error correction method and system based on artificial intelligence shown in the present invention are as follows:
[0046] The method shown in the present invention obtains the voice input information input at the current time node, and outputs the corresponding pinyin sequence according to the voice input information; imports the pinyin sequence into a preset word database, and traverses multiple Chinese character sequences corresponding to the pinyin sequence in the word database, and the multiple Chinese character sequences are homophonic words; uses a context verification model based on the attention mechanism to score the semantic fitness of the multiple Chinese character sequences and the pinyin sequence respectively, and outputs the semantic score value of each Chinese character sequence; determines the Chinese character sequence with the highest semantic score value as the target Chinese character sequence, and obtains the target Chinese character text corresponding to the voice input information. The present invention determines the target Chinese character text by scoring the Chinese character sequences that may correspond to the pinyin sequence through a context verification model based on the attention mechanism, solves the problem that the voice error correction process is not accurate enough due to the polyphonic character problem when converting voice to text in the prior art, and the accuracy of the voice conversion result is low, and can be effectively used for accurate voice error correction in terminal applications such as electronic devices and automobiles, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The above and / or additional aspects and advantages of the present invention will become apparent and easy to understand from the description of the embodiments in conjunction with the following drawings, wherein:
[0048] Figure 1 is a schematic flowchart of a voice error correction method based on artificial intelligence in an embodiment of the present invention;
[0049] Figure 2 is a structural block diagram of a voice error correction system based on artificial intelligence in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] In order to make the objectives, features and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention is made in conjunction with the drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0052] Embodiment 1
[0053] Please refer to Figure 1, the first embodiment of the present invention provides a speech error correction method based on artificial intelligence, and the method includes steps S10 - S40:
[0054] Step S10, obtain the speech input information input at the current time node, and output the corresponding pinyin sequence according to the speech input information.
[0055] In this embodiment, the current time node is the time node when the user inputs speech through a pick-up microphone, which can be the start time of speech input, or the end time of speech input, or any time point during the speech input process. Through speech input, speech input information is obtained.
[0056] It should be noted that the speech input information can be a passage input through speech in an input method. For example, when chatting with an AI model, input "Do you know why the sun rises in the east and sets in the west?", or it can be a control instruction for controlling terminal devices such as electronic devices or vehicles. For example, input "Set an alarm for 7 o'clock tomorrow morning for me", "Set the air conditioner temperature to 26 degrees for me".
[0057] Specifically, after obtaining the speech input information by inputting speech content at the current time node, the corresponding pinyin sequence will be output according to the speech input information. For example, the pinyin sequence output for "Do you know why the sun rises in the east and sets in the west?" includes "ni, zhidao, taiyang, weishenme, shi, dongshengxiluo, ma", and the output pinyin sequence includes at least one.
[0058] Of course, it should also be supplemented that after obtaining the speech input information, the speech input information can also be preprocessed to make the speech content in the speech input information clearer, which is convenient for the accurate output of the pinyin sequence.
[0059] Step S20, import the pinyin sequence into a preset word database, and traverse multiple Chinese character sequences corresponding to the pinyin sequence in the word database. The multiple Chinese character sequences are homophonic words.
[0060] In this embodiment, after outputting the corresponding pinyin sequence according to the speech input information, it is possible to first clarify what the pinyin of the content pre-input by the user is specifically, and then import the corresponding pinyin sequence into a preset word database. This word database is a database built according to the content of a dictionary, which contains at least common words and their corresponding pinyins, and there is a corresponding relationship between them. Then, in this word database, using the pinyin sequence or subsequence as an index, search for all the words corresponding to it, and combine them to obtain several Chinese character sequences.
[0061] For example, let's take the above "Do you know why the sun rises in the east and sets in the west?" as an example. The Chinese character sequences corresponding to the pinyin sequence "zhidao" include "know", "guide", "straight road", etc. The same is true for other pinyin sequences, which will not be explained in detail here.
[0062] That is to say, the Chinese character sequences corresponding to each pinyin sequence contained in the speech input information are multiple, such as the Chinese character sequences "know", "guide", and "straight road". They are polyphones. In order to accurately obtain the final target Chinese character text, it is necessary to correct at least part of the speech content in the speech input information, that is, the Chinese character sequence. For example, in the speech conversion process, the initially generated "Do you know why the sun rises in the east and sets in the west?" is corrected to "Do you know why the sun rises in the east and sets in the west?". It is specifically implemented by steps S30-S40.
[0063] Step S30, using a context verification model based on an attention mechanism to score the semantic fit between the multiple Chinese character sequences and the pinyin sequence, and outputting the semantic score value of each Chinese character sequence.
[0064] In this embodiment, a context verification model based on an attention mechanism is used to score the semantic fit between the multiple Chinese character sequences and the pinyin sequence, and the step of outputting the semantic score value of each Chinese character sequence includes:
[0065] Convert the Chinese character sequence and the pinyin sequence into embedding vectors respectively, and add position codes accordingly;
[0066] The Chinese character sequence and the pinyin sequence are mutually focused through the attention mechanism of the context verification model to generate context-related target features;
[0067] Performing feature fusion and pooling processing on the target features to extract key information of the target features;
[0068] The fused target feature map is mapped through the fully connected layer in the context verification model to obtain the semantic score value of each Chinese character sequence.
[0069] The step of making the Chinese character sequence and the pinyin sequence pay attention to each other through the attention mechanism of the context verification model to generate context-related target features includes:
[0070] By using the context verification model, the cross attention weight between the Chinese character sequence and the pinyin sequence is calculated according to the multi-head attention mechanism to capture the semantic association between the Chinese character sequence and the pinyin sequence;
[0071] Take the Chinese character sequence as the Query vector, take the pinyin sequence as the Key vector and the Value vector, generate an attention matrix, and use scaled dot-product attention to measure the local correlation between Chinese characters and pinyin;
[0072] Perform weighted summation of the cross-attention weights and the Value vector of the pinyin vector to generate a context-enhanced representation of the Chinese character sequence, obtaining the target feature.
[0073] In addition, the step of mapping the fused target feature to obtain the semantic score value of each Chinese character sequence through the fully connected layer in the context verification model includes:
[0074] According to at least one fully connected layer in the context verification model, map the target feature to an intermediate dimension through a weight matrix and a bias term;
[0075] Use the activation function of linear output through the output layer in the context verification model to output the score corresponding to the Chinese character sequence, obtaining the semantic score value corresponding to each Chinese character sequence.
[0076] Specifically, when scoring the semantic fit of a Chinese character sequence, a context verification model based on the attention mechanism is used for scoring. The Chinese character sequence is converted into a vector representation through an embedding layer, and positional encoding is added to capture sequence order information. Then the pinyin sequence is split into initials, finals and tones, or directly converted into a pinyin embedding vector, and positional encoding is also added to capture sequence order information. Then calculate the cross-attention weights between the Chinese character sequence and the pinyin sequence through the multi-head attention mechanism to capture the phonetic association between the two. Specifically, take the Chinese character sequence as the Query vector, take the pinyin sequence as the Key vector and the Value vector (or vice versa), generate an attention matrix, use scaled dot-product attention to measure the local correlation between Chinese characters and pinyin, and finally perform weighted summation of the attention weights and the Value vector of the pinyin sequence to generate a context-enhanced representation of the Chinese character sequence, obtaining the target feature. Concatenate the original Chinese character embedding and the context representation enhanced by attention, retain the original information and attention features, realize hierarchical feature concatenation, and then perform global average pooling or max pooling on the concatenated features to generate a fixed-length semantic vector. Map the pooled semantic vector to the annotation score value through a fully connected layer. Specifically, use an activation function (such as the Sigmoid function) to limit the output within the range of [0,1] to represent the fit probability, or use a regression method to output an unnormalized score, and compare the fit degree through score ranking.
[0077] Step S40, determine the Chinese character sequence with the highest semantic score value as the target Chinese character sequence, obtaining the target Chinese character text corresponding to the speech input information.
[0078] In this embodiment, the target Chinese character sequence is determined according to the semantic score value. Parallel calculations are performed on multiple pairs of Chinese character sequences - pinyin sequences, and the semantic score value of each sequence pair is output. Sequence pairs with low fitness are filtered according to the score threshold, or the optimal candidate sequence pair is output according to the score ranking to obtain the target Chinese character sequence. Then, according to the target Chinese character sequence, the target Chinese character text corresponding to the speech content in the voice input information is determined, so as to realize speech error correction in the speech conversion process.
[0079] It should be noted here that the context verification model based on the attention mechanism in this embodiment is implemented based on cross-attention, rather than self-attention, so as to ensure the semantic alignment across modalities (Chinese characters and pinyin).
[0080] Compared with the prior art, the beneficial effects of adopting the speech error correction method based on artificial intelligence shown in this embodiment are as follows:
[0081] The method shown in this embodiment obtains the voice input information input at the current time node, and outputs the corresponding pinyin sequence according to the voice input information; imports the pinyin sequence into a preset word database, and traverses multiple Chinese character sequences corresponding to the pinyin sequence in the word database. The multiple Chinese character sequences are homophonic words; a context verification model based on the attention mechanism is used to score the semantic fitness of multiple Chinese character sequences and pinyin sequences respectively, and output the semantic score value of each Chinese character sequence; the Chinese character sequence with the highest semantic score value is determined as the target Chinese character sequence to obtain the target Chinese character text corresponding to the voice input information. By using the context verification model based on the attention mechanism to score the Chinese character sequences that may correspond to the pinyin sequence to determine the target Chinese character text, this embodiment solves the problem that the speech error correction process is not accurate enough due to the polyphone problem when converting speech to text in the prior art, and the accuracy of the speech conversion result is low. It can be effectively used for accurate speech error correction in terminal applications such as electronic devices and automobiles, thereby improving the user experience.
[0082] Embodiment Two
[0083] The second embodiment of the present invention also provides a speech error correction method based on artificial intelligence. The method shown in this embodiment further includes:
[0084] Use a pre-trained deep learning model to perform background denoising on the voice input information to enhance the speech clarity of the voice input information;
[0085] Perform endpoint detection on the voice input information, determine the start and end points of the voice segment in the voice input information, and cut according to the start and end points of the voice segment to obtain the target voice segment.
[0086] After the step of performing endpoint detection on the voice input information, determining the start and end points of the voice segment in the voice input information, and cutting according to the start and end points of the voice segment to obtain the target voice segment, the following steps are further included:
[0087] Output the corresponding pinyin sequence according to the target voice segment.
[0088] Specifically, in this embodiment, after obtaining the voice input information, the background noise of the voice input information will also be removed through a pre-trained deep learning model such as Conv-TasNet, removing background noise and enhancing voice clarity. For example, in a noisy environment, the voiceprint features of the target speaker are extracted through a semantic separation model, and the start and end points of the effective voice segment are detected to avoid silent segments or noise segments from being input into the model, which can effectively reduce the data processing volume of the model and improve the conversion efficiency of voice conversion.
[0089] Embodiment III
[0090] Please refer to Figure 2 , the third embodiment of the present invention provides a voice error correction system based on artificial intelligence, which is applied to the method described in any of the above embodiments. The system includes: a pinyin sequence output module 10, a Chinese character sequence traversal module 20, a Chinese character sequence scoring module 30, and a Chinese character text output module 40.
[0091] The pinyin sequence output module 10 is used to obtain the voice input information input at the current time node and output the corresponding pinyin sequence according to the voice input information;
[0092] The Chinese character sequence traversal module 20 is used to import the pinyin sequence into a preset Chinese character database and traverse multiple Chinese character sequences corresponding to the pinyin sequence in the Chinese character database. The multiple Chinese character sequences are homophonic words;
[0093] The Chinese character sequence scoring module 30 is used to score the semantic fit degrees of the multiple Chinese character sequences and the pinyin sequence respectively by using a context verification model based on the attention mechanism, and output the semantic score value of each Chinese character sequence;
[0094] The Chinese character text output module 40 is used to determine the Chinese character sequence with the highest semantic score value as the target Chinese character sequence and obtain the target Chinese character text corresponding to the voice input information.
[0095] In this embodiment, the Chinese character sequence scoring module 30 is specifically used for:
[0096] Convert the Chinese character sequence and the pinyin sequence into embedding vectors respectively, and add position encoding correspondingly;
[0097] Mutual attention is performed between the Chinese character sequence and the pinyin sequence through the attention mechanism of the context verification model to generate context-related target features;
[0098] Perform feature fusion and pooling processing on the target features to extract key information of the target features;
[0099] Map the fused target features through the fully connected layer in the context verification model to obtain the semantic score value of each Chinese character sequence.
[0100] In this embodiment, the Chinese character sequence scoring module 30 is further configured to:
[0101] Through the context verification model, calculate the cross-attention weights between the Chinese character sequence and the pinyin sequence according to the multi-head attention mechanism, and capture the semantic association between the Chinese character sequence and the pinyin sequence;
[0102] Use the Chinese character sequence as the Query vector, the pinyin sequence as the Key vector and the Value vector to generate an attention matrix, and use scaled dot attention to measure the local correlation between Chinese characters and pinyin;
[0103] Perform weighted summation on the cross-attention weights and the Value vector of the pinyin vector to generate a context-enhanced representation of the Chinese character sequence and obtain target features.
[0104] And, the Chinese character sequence scoring module 30 is further configured to:
[0105] According to at least one fully connected layer in the context verification model, map the target features to an intermediate dimension through a weight matrix and a bias term;
[0106] Use the activation function of linear output through the output layer in the context verification model to output the score corresponding to the Chinese character sequence, and obtain the semantic score value corresponding to each Chinese character sequence.
[0107] Compared with the prior art, the speech error correction system based on artificial intelligence shown in this embodiment has the beneficial effects that:
[0108] The system shown in this embodiment obtains the voice input information input at the current time node, and outputs the corresponding pinyin sequence according to the voice input information; imports the pinyin sequence into a preset word database, traverses multiple Chinese character sequences corresponding to the pinyin sequence in the word database, and the multiple Chinese character sequences are homophonic words; uses a context verification model based on the attention mechanism to score the semantic fit degrees of the multiple Chinese character sequences and the pinyin sequence respectively, and outputs the semantic score value of each Chinese character sequence; determines the Chinese character sequence with the highest semantic score value as the target Chinese character sequence, and obtains the target Chinese character text corresponding to the voice input information. This embodiment determines the target Chinese character text by scoring the Chinese character sequences that may correspond to the pinyin sequence through a context verification model based on the attention mechanism, solves the problem in the prior art that the voice error correction process is not accurate enough due to the polyphonic character problem during voice-to-text conversion, and the accuracy of the voice conversion result is low, and can be effectively used for accurate voice error correction in terminal applications such as electronic devices and automobiles, thereby improving the user experience.
[0109] Embodiment 4
[0110] The fourth embodiment of the present invention provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any of the above embodiments is implemented.
[0111] Embodiment 5
[0112] The fifth embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the method described in any of the above embodiments is implemented.
[0113] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0114] The above-described embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. A speech error correction method based on artificial intelligence, characterized in that: The method comprises: Acquire the voice input information input at the current time node, and output the corresponding pinyin sequence according to the voice input information; Importing the pinyin sequence into a preset word database, traversing a plurality of Chinese character sequences corresponding to the pinyin sequence in the word database, wherein the plurality of Chinese character sequences are homophones; Using a context verification model based on an attention mechanism to score the semantic fit between the multiple Chinese character sequences and the pinyin sequence, and outputting a semantic score value for each Chinese character sequence; The Chinese character sequence with the highest semantic score value is determined as the target Chinese character sequence, and a target Chinese character text corresponding to the speech input information is obtained.
2. The speech error correction method based on artificial intelligence according to claim 1, characterized in that: The step of using a context verification model based on an attention mechanism to score the semantic fit between the plurality of Chinese character sequences and the pinyin sequence, and outputting the semantic score value of each Chinese character sequence, comprises: Convert the Chinese character sequence and the pinyin sequence into embedding vectors respectively, and add position codes accordingly; The Chinese character sequence and the pinyin sequence are mutually focused through the attention mechanism of the context verification model to generate context-related target features; Performing feature fusion and pooling processing on the target features to extract key information of the target features; The fused target feature map is mapped through the fully connected layer in the context verification model to obtain the semantic score value of each Chinese character sequence.
3. The speech error correction method based on artificial intelligence according to claim 2 is characterized in that: The step of making the Chinese character sequence and the pinyin sequence pay attention to each other through the attention mechanism of the context verification model to generate context-related target features includes: By using the context verification model, the cross attention weight between the Chinese character sequence and the pinyin sequence is calculated according to the multi-head attention mechanism to capture the semantic association between the Chinese character sequence and the pinyin sequence; The Chinese character sequence is used as a query vector, the pinyin sequence is used as a key vector and a value vector, an attention matrix is generated, and the local correlation between the Chinese characters and the pinyin is measured using scaled point attention; The cross attention weight is weightedly summed with the Value vector of the pinyin vector to generate a context-enhanced representation of the Chinese character sequence to obtain the target feature.
4. The speech error correction method based on artificial intelligence according to claim 3 is characterized in that: The step of mapping the fused target features through the fully connected layer in the context verification model to obtain the semantic score value of each Chinese character sequence includes: According to at least one fully connected layer in the context verification model, mapping the target feature to an intermediate dimension through a weight matrix and a bias term; The scores corresponding to the Chinese character sequences are output by the output layer in the context verification model using a linear output activation function to obtain the semantic score value corresponding to each of the Chinese character sequences.
5. The speech error correction method based on artificial intelligence according to claim 1, characterized in that: The method further comprises: Using a pre-trained deep learning model to perform background denoising on the voice input information to enhance the voice clarity of the voice input information; Endpoint detection is performed on the voice input information to determine the start and end points of the voice segment in the voice input information, and the voice segment is cut according to the start and end points to obtain a target voice segment.
6. The artificial intelligence-based speech error correction method according to claim 5, characterized in that: After performing endpoint detection on the voice input information, determining the start and end points of the voice segments in the voice input information, and cutting the voice segments according to the start and end points to obtain the target voice segments, the method further includes: According to the target speech segment, a corresponding pinyin sequence is output.
7. A speech error correction system based on artificial intelligence, characterized in that: The method applied to any one of claims 1 to 6, wherein the system comprises: A pinyin sequence output module is used to obtain the voice input information input at the current time node and output the corresponding pinyin sequence according to the voice input information; A Chinese character sequence traversal module, used for importing the pinyin sequence into a preset word database, and traversing a plurality of Chinese character sequences corresponding to the pinyin sequence in the word database, wherein the plurality of Chinese character sequences are homophones; A Chinese character sequence scoring module, used to score the semantic fit between the multiple Chinese character sequences and the pinyin sequence respectively using a context verification model based on an attention mechanism, and output a semantic score value for each Chinese character sequence; The Chinese character text output module is used to determine the Chinese character sequence with the highest semantic score as the target Chinese character sequence, and obtain the target Chinese character text corresponding to the voice input information.
8. The artificial intelligence-based speech error correction system according to claim 7, characterized in that: The Chinese character sequence scoring module is specifically used for: Convert the Chinese character sequence and the pinyin sequence into embedding vectors respectively, and add position codes accordingly; The Chinese character sequence and the pinyin sequence are mutually focused through the attention mechanism of the context verification model to generate context-related target features; Performing feature fusion and pooling processing on the target features to extract key information of the target features; The fused target feature map is mapped through the fully connected layer in the context verification model to obtain the semantic score value of each Chinese character sequence.
9. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and capable of being executed on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Voice recognition error correction method and man-machine conversation system
CN110428822A
Error correction method and device based on language model, equipment and storage medium
CN111428474A
Text error correction method, device and system
CN111523306A
Speech recognition error correction method and device and storage medium
CN112735396A
Polyphone pronunciation labeling method and device, equipment and storage medium
CN113268974A