Audio-related text error detection method and device, equipment and storage medium

By integrating audio emotional modality features and text modality features into the error correction system, the problem of low accuracy in audio text misspelling detection in existing technologies is solved, and more accurate misspelling detection is achieved.

CN115563961BActive Publication Date: 2026-07-24IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2022-10-28
Publication Date
2026-07-24

Smart Images

  • Figure CN115563961B_ABST
    Figure CN115563961B_ABST
Patent Text Reader

Abstract

The application discloses an audio-related text error detection method, device, equipment and storage medium. The application extracts the character modality feature of the text to be detected, the emotion modality feature of the input audio related to the text to be detected, fuses the emotion modality feature and the character modality feature, determines the real text corresponding to the text to be detected based on the fused feature, compares the real text and the text to be detected, and obtains the error detection result. In the error detection, the emotion modality feature of the related audio is further fused on the basis of the character modality feature of the text to be detected, so that the prediction result is more accurate. On this basis, the error detection result is determined by comparing the real text and the text to be detected, and the accuracy of the error detection is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and more specifically, to a method, apparatus, device, and storage medium for detecting typos in audio-related text. Background Technology

[0002] With the development of information technology, an era characterized by diversified information transmission methods has arrived. In daily life and work, people receive text information from an increasing number of sources, such as street advertisements, social media posts, and video subtitles. Due to various reasons, typos may appear in text information. Relying solely on manual proofreading and correction of these documents would consume a significant amount of manpower and time.

[0003] In today's era of booming artificial intelligence, especially thanks to advancements in natural language processing technology, various text error detection and correction systems have emerged to help people efficiently check and correct textual errors. The basic process of existing error correction systems is to receive a text that may contain various errors, such as grammatical and lexical errors, as input; process it; locate and correct any potential errors; and return the location and correction results to the user. Taking video subtitles as an example, existing error correction systems typically identify the video subtitles, perform error correction processing based on the context of the subtitle text information, locate potential errors, and return the results to the user.

[0004] Existing error correction methods only utilize plain text information for error correction, resulting in low accuracy in typo detection. Summary of the Invention

[0005] In view of the above problems, this application is proposed to provide a method, apparatus, device, and storage medium for detecting typos in audio-related text, so as to improve the accuracy of typo detection in audio-related text. The specific solution is as follows:

[0006] Firstly, a method for detecting typos in audio-related text is provided, including:

[0007] Obtain the input audio and the text to be detected related to the input audio;

[0008] Extract the emotional modal features of the input audio, and extract the textual modal features of the text to be detected;

[0009] The emotional modality features and the text modality features are fused to obtain fused features;

[0010] The real text corresponding to the text to be detected is determined based on the fusion features;

[0011] By comparing the real text and the text to be detected, the typo detection results in the text to be detected are obtained.

[0012] Secondly, a typo detection device for audio-related text is provided, including:

[0013] The data acquisition unit is used to acquire input audio and text to be detected related to the input audio;

[0014] The feature extraction unit is used to extract the emotional modal features of the input audio and the text modal features of the text to be detected.

[0015] The feature fusion unit is used to fuse the emotional modality features and the text modality features to obtain fused features;

[0016] The real text determination unit is used to determine the real text corresponding to the text to be detected based on the fusion features;

[0017] The misspelling detection unit is used to compare the real text and the text to be detected to obtain the misspelling detection result in the text to be detected.

[0018] Thirdly, a typo detection device for audio-related text is provided, including: a memory and a processor;

[0019] The memory is used to store programs;

[0020] The processor is used to execute the program to implement the various steps of the above-described method for detecting typos in audio-related text.

[0021] Fourthly, a storage medium is provided on which a computer program is stored, which, when executed by a processor, implements the various steps of the above-described method for detecting typos in audio-related text.

[0022] Using the above technical solution, this application simultaneously acquires the related input audio for the text to be detected. For example, when the text to be detected is a subtitle, the input audio can be multimedia audio matching the subtitle. Then, the emotional modality features of the input audio and the text modality features of the text to be detected are extracted. The emotional modality features and the text modality features are fused, and the real text corresponding to the text to be detected is determined based on the fused features. The real text and the text to be detected are compared to obtain the typo detection result. Therefore, this application, when detecting typos in the text to be detected, not only considers the textual modal features of the text to be detected, but also further integrates the emotional modal features of the input audio related to the text to be detected. That is, it makes full use of the emotional modal features of the audio corresponding to the text to assist in the prediction of the real text. Considering that the expression of the text is related to the emotion of the related audio, for example, if the text to be detected is "We all beat our hands together to celebrate Xiaoming's achievement of first place in the class!", the corresponding audio emotion is "happy". It is understandable that when users are in a "happy" mood, they rarely express words like "fear" that express "fear". Therefore, this application can use emotional modal features to assist in the prediction of the real text. Compared with simply relying on the context of the text to be detected to predict the real text, the prediction results are more accurate. On this basis, by comparing the real text and the text to be detected, the typo detection results are determined, which greatly improves the accuracy of typo detection. Attached Figure Description

[0023] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0024] Figure 1 A flowchart illustrating the method for detecting typos in audio-related text provided in this application embodiment;

[0025] Figure 2 This example illustrates the process of marking typos in the text to be detected.

[0026] Figure 3 This example illustrates the structure of an audio text recognition model.

[0027] Figure 4 An example is a schematic diagram of the structure of an audio processing module;

[0028] Figure 5 This example illustrates the structure of a text processing module;

[0029] Figure 6An example is a schematic diagram of the structure of a multimodal fusion module;

[0030] Figure 7 A schematic diagram illustrating the processing flow of a multimodal fusion module is provided.

[0031] Figure 8 A schematic diagram of a typo detection device in audio-related text provided in this application embodiment;

[0032] Figure 9 This is a schematic diagram of the structure of the typo detection device in audio-related text provided in the embodiments of this application. Detailed Implementation

[0033] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0034] This application provides a method for detecting typos in audio-related text, which can be applied to tasks that detect typos in text based on the text to be detected and its related audio data, such as tasks that detect typos in multimedia subtitle text.

[0035] The proposed solution can be implemented based on a terminal with data processing capabilities, such as a mobile phone, computer, server, or cloud platform.

[0036] Next, combined Figure 1 The method for detecting typos in audio-related text described in this application may include the following steps:

[0037] Step S100: Obtain the input audio and the text to be detected related to the input audio.

[0038] Specifically, the text to be detected is the text information that needs to be checked for typos. The text to be detected can be user-inputted text information, text information recognized from images, etc. The text information contained in the text to be detected can include Chinese characters and non-Chinese characters, such as English letters, special symbols, numbers, etc.

[0039] The input audio is audio information related to the text to be detected. It can be user-recorded or extracted from multimedia data. For example, when the text to be detected is subtitle text, the input audio can be multimedia audio that matches the subtitle.

[0040] For example Figure 2, the text to be detected is subtitle text: Let's clap our hands together to celebrate Xiaoming for achieving the first place in the class!

[0041] Meanwhile, the audio data corresponding to the subtitle text is extracted from the multimedia data with subtitle matching.

[0042] It can be known that "papa" in the subtitle "Let's clap our hands together to celebrate Xiaoming for achieving the first place in the class!" is a misspelling, and the correct one should be "pai pai".

[0043] Step S110: Extract the emotional modality features of the input audio, and extract the text modality features of the text to be detected.

[0044] Specifically, when extracting the text modality features, it can be extracted by using a set text feature extraction algorithm, or by using a pre-trained natural language processing model.

[0045] The input audio contains the emotional information of the user, such as emotional types like joy, excitement, sadness, depression, fear, etc. When the user expresses different text contents, the corresponding emotional information may also be different. For example, when the user is in a sad emotion, it is less likely to express words indicating joy such as "burst out laughing".

[0046] Therefore, in order to assist in detecting misspelled words in the text to be detected, the emotional modality features of the input audio related to the text to be detected are extracted in this step.

[0047] Among them, various different algorithms can be used when extracting the emotional modality features. For example, the emotional modality features of the input audio can be extracted through a pre-trained neural network model, etc.

[0048] Step S120: Fuse the emotional modality features and the text modality features to obtain fused features.

[0049] Specifically, the emotional modality features and the text modality features respectively describe the relevant information from the two perspectives of audio and text. In order to more accurately predict the true text corresponding to the text to be detected, the emotional modality features and the text modality features are fused in this step, and the information of the obtained fused features is more abundant and the expression ability is stronger.

[0050] Step S130: Determine the true text corresponding to the text to be detected based on the fused features.

[0051] Specifically, after obtaining the fused features in the above step, the true text corresponding to the text to be detected can be predicted based on the fused features. In this step, a pre-trained neural network model can be used for the prediction of the true text.

[0052] The actual text predicted by this step is the correct text corresponding to the text to be detected as identified in this application.

[0053] Step S140: Compare the real text and the text to be detected to obtain the typo detection results in the text to be detected.

[0054] Specifically, in this step, real text can be used as a benchmark to compare the text to be detected with the real text, determine whether the text to be detected contains typos, and the specific content of the typos, and obtain the typo detection result in the text to be detected.

[0055] For example, in this step, it is possible to match whether there are characters in the text to be detected that are inconsistent with the real text. If so, the inconsistent characters in the text to be detected are regarded as typos in the text to be detected.

[0056] The method for detecting misspellings in audio-related text provided in this application involves acquiring related input audio for the text to be detected. For example, when the text to be detected is subtitle text, the input audio can be multimedia audio matching the subtitles. Then, the emotional modality features of the input audio and the text modality features of the text to be detected are extracted. The emotional modality features and the text modality features are fused together, and the real text corresponding to the text to be detected is determined based on the fused features. The real text and the text to be detected are compared to obtain the misspelling detection result. Therefore, this application, when detecting typos in the text to be detected, not only considers the textual modal features of the text to be detected, but also further integrates the emotional modal features of the input audio related to the text to be detected. That is, it makes full use of the emotional modal features of the audio corresponding to the text to assist in the prediction of the real text. Considering that the expression of the text is related to the emotion of the related audio, for example, if the text to be detected is "We all beat our hands together to celebrate Xiaoming's achievement of first place in the class!", the corresponding audio emotion is "happy". It is understandable that when users are in a "happy" mood, they rarely express words like "fear" that express "fear". Therefore, this application can use emotional modal features to assist in the prediction of the real text. Compared with simply relying on the context of the text to be detected to predict the real text, the prediction results are more accurate. On this basis, by comparing the real text and the text to be detected, the typo detection results are determined, which greatly improves the accuracy of typo detection.

[0057] Optionally, after obtaining the typo detection result in step S140 above, if it is confirmed that the text to be detected contains typos, the position of the typos in the text to be detected can be further determined, and then the typos can be marked in the text to be detected according to the position, so as to intuitively display the typos in the text to be detected.

[0058] Reference Figure 2 For the misspelled word "papa" identified in the text to be detected, it is marked in the form of a rectangular box.

[0059] Of course, the marking form of misspelled words is not limited to rectangular box marking, and other various marking methods can also be adopted, such as highlighting, underlining, etc.

[0060] In some embodiments of the present application, the process of fusing the emotional modality feature and the text modality feature to obtain a fused feature in the above step S120 is described.

[0061] Optionally, the emotional modality feature and the text modality feature extracted in step S110 can be in vector form. The vector dimensions of the emotional modality feature and the text modality feature can be the same or different. On this basis, when performing feature fusion in this step, two features in vector form can be fused to obtain a fused feature.

[0062] When performing vector fusion, various fusion methods can be adopted. In this embodiment, a gated fusion method is provided to fuse the emotional modality feature and the text modality feature in vector form to obtain a fused feature.

[0063] By adopting the gated fusion method, using the emotional modality feature as a gate, some features in the text modality feature are extracted to obtain a fused feature. That is, from the perspective of the emotional modality feature, the most important part of the text modality feature is extracted as the feature representation of the fusion of the emotional modality and the character modality.

[0064] Optionally, several different forms of gated fusion methods are provided in the embodiments of the present application. Examples may include: a gated fusion method of bitwise multiplication, a gated fusion method of bitwise addition or division, etc. For the sake of easy expression, only the gated fusion method of bitwise multiplication is taken as an example in the following embodiments for illustration.

[0065] Furthermore, in order to avoid the loss of global features at the text language level, in this embodiment, the above fused feature can also be added to the text modality feature to obtain a residual fused feature as the final fused feature.

[0066] In order to enhance the richness of the emotional modality feature representation, before performing feature fusion in step S120, a process of representing offset and non-linear transformation of the emotional modality feature can also be added to obtain a processed emotional modality feature for fusing the processed emotional modality feature and the text modality feature in step S120.

[0067] In some embodiments of the present application, for the steps S110 - S130 introduced in the foregoing embodiments, they can be obtained by processing through a pre-trained audio-text recognition model.

[0068] For audio-text recognition models, they can be configured as follows: extract the emotional modality features of the input audio, extract the textual modality features of the input text to be detected, fuse the emotional modality features and the textual modality features, and predict the internal state representation of the real text corresponding to the text to be detected based on the fused features.

[0069] The input to the audio text recognition model can include the text information to be detected, as well as the input audio corresponding to the text to be detected.

[0070] In this embodiment, by pre-training the audio-text recognition model, the powerful learning ability of the neural network model can be utilized to extract the emotional modality features of the input audio and the textual modality features of the text to be detected. Based on this, the real text is predicted after fusion.

[0071] Next, combined Figure 3 As shown, this embodiment provides an optional component structure for the audio text recognition model.

[0072] An audio-to-text recognition model can include an audio processing module, a text processing module, a multimodal fusion module, and an output module. Among them:

[0073] The audio processing module is used to extract the emotional modality features of the input audio.

[0074] Specifically, the input to the audio processing module can be audio features, such as Fbank features. The audio processing module encodes the audio features to obtain the emotional modality features of the input audio.

[0075] The process of extracting audio features (Fbank) from the input audio can include: 1. Pre-emphasizing, framing, and applying a Hamming window to the audio signal, then performing a Short-Time Fourier Transform (STFT) to obtain the spectrum. 2. Calculating the square of the spectrum. Superimposing the energy within each filter band. 3. Taking the logarithm of the output of each filter to obtain the logarithmic power spectrum (Fbank) of the corresponding band. 4. Normalizing all audio feature Fbanks. Normalization improves the model's generalization ability.

[0076] The text processing module is used to extract the text modal features of the text to be detected.

[0077] The multimodal fusion module is used to fuse the emotional modal features and the text modal features to obtain fused features.

[0078] The output module is used to determine the real text corresponding to the text to be detected based on the fusion features.

[0079] The output module can be trained using the MLM (Masked Language Model) method. Based on the fusion features output by the multimodal fusion module, it predicts the real text corresponding to the text to be detected.

[0080] Next, each of the above modules will be explained in detail.

[0081] 1. Audio processing module

[0082] This embodiment describes an optional structure for the audio processing module, such as... Figure 4 As shown, it may include:

[0083] The pre-trained audio coding module is used to encode the audio features of the input audio using a pre-trained audio coding model to obtain the encoded emotional modality features.

[0084] The audio coding model can employ an audio pre-trained model structure, trained using sentiment classification as the pre-training task. This pre-trained model structure can be an audio transformer model, etc. By using sentiment classification as the pre-training task, the audio pre-trained model gains the ability to extract features related to the sentiment type of the input audio. Based on this, the pre-trained audio coding model can be used to encode the audio features of the input audio to extract sentiment modality features.

[0085] The linear transformation module is used to perform a linear transformation on the dimensions of the emotional modality features to output emotional modality features with the same dimensions as the text modality features.

[0086] Specifically, the number of channels of the emotional modality features extracted by the pre-trained audio encoding module may not be directly matched with the dimension of the text modality features extracted by the text processing module. Therefore, it is necessary to perform a linear transformation on the dimension of the emotional modality features through the linear transformation module to output emotional modality features with the same dimension as the text modality features.

[0087] 2. Text Processing Module

[0088] This embodiment describes an optional structural configuration for the text processing module, such as... Figure 5 As shown, it may include:

[0089] The text preprocessing module is used to edit the text to be detected to a set length by padding with specified characters, and to determine the feature representation of the edited text.

[0090] Specifically, to standardize the length of different texts to be detected, this embodiment uses a text preprocessing module to edit the text to be detected to a set length using padding. For texts shorter than the set length, a set padding character, such as [PAD], can be added to the end of the text to supplement it to the set length. For texts longer than the set length, the set length can be truncated from the first character as a single edited text. If the remaining length still exceeds the set length, the truncating operation is repeated. If the remaining length does not exceed the set length, the remaining portion is used as another edited text.

[0091] For each edited text to be detected, a pre-trained tokenizer can be used to encode the text into a feature representation that the model can recognize. Specifically, the edited text to be detected is segmented into words, and each word is encoded to obtain the token feature representation corresponding to the word.

[0092] Among them, the pre-trained tokenizer can adopt pre-trained model structures such as BERT tokenizer.

[0093] The text modality feature extraction module is used to encode the feature representation of the text to be detected to obtain the text modality features of the text to be detected.

[0094] Specifically, the text modality feature extraction module can use a pre-trained model (such as BERT, Transformer, etc.) to encode the feature representation of the text to be detected after the text preprocessing module, so as to obtain the text modality features of the text to be detected.

[0095] 3. Multimodal fusion module

[0096] This embodiment describes an optional structural composition of the multimodal fusion module, such as... Figure 6 As shown, it may include: a feature editing module, a gating fusion module, and a residual connection module.

[0097] The processing flow of each module is combined Figure 7 Explanation:

[0098] The feature editing module is used to perform representation shifts and nonlinear transformations on the emotional modality features to obtain the processed emotional modality features.

[0099] To enhance the representation of sentiment modality features, representation shift and nonlinear transformation can be applied. Representation shift involves adding a learnable bias parameter to each position of the sentiment modality feature. Nonlinear transformation uses nonlinear function layers, such as ReLU, sigmoid, or tanh layers, to nonlinearly transform the shifted sentiment modality features to a relatively small range near 0. For example, the range after a sigmoid transformation is (0, 1), and the range after a tanh transformation is (-1, 1).

[0100] The gated fusion module is used to fuse the processed emotional modality features and the text modality features using a gated fusion method to obtain fused features.

[0101] Specifically, in this embodiment, a gating fusion module is designed to perform bitwise multiplication, bitwise addition, or bitwise division to fuse the processed emotional modality features and text modality features to obtain fused features.

[0102] Figure 7 Taking the bitwise multiplication gating fusion method as an example, by adopting the bitwise multiplication gating fusion method, the emotional modality features are used as the gate to extract some features from the text modality features to obtain the fused features. That is, from the perspective of emotional modality features, the most important part of the text modality features is extracted as the feature representation of the fusion of emotional modality and character modality.

[0103] After the feature editing module processes the sentiment modality features, they gain additional representational shift and nonlinear transformation compared to text modality features. This maps the sentiment modality features to a relatively small range near 0, such as the sigmoid function's range of (0,1). The range and distribution of text modality features remain unchanged. To put it figuratively, each position in the sentiment modality feature after editing is like a faucet (fully open corresponds to the upper bound of the nonlinear function's range, fully closed corresponds to the lower bound), controlling the information at the corresponding position in the text modality feature. The more the faucet is open in the sentiment modality feature, the more information is retained in the corresponding position in the text modality feature; conversely, less is retained. Clearly, this positional multiplication yields the text modality feature portion controlled by the degree of retention from an emotional perspective. In other words, from the perspective of the sentiment modality feature, the most important part of the text modality feature is extracted as the feature representation for the fusion of sentiment and character modality.

[0104] The residual connection module is used to add the fused feature to the text modal feature to obtain the residual fused feature, which is used as the final fused feature.

[0105] Furthermore, to avoid the loss of global features at the text language level, this embodiment can also add the above-mentioned fused features to the text modal features through the residual connection module to obtain residual fused features, which are used as the final fused features.

[0106] In some embodiments of this application, in order to further improve the accuracy of typo detection, after step S140, comparing the real text and the text to be detected to obtain the typo detection results in the text to be detected, a post-processing operation for typo verification can be further added.

[0107] In this embodiment, the post-processing for misspelling verification can be performed from the perspective of sentence semantic fluency, specifically including:

[0108] S1. Delete the typos identified in the text to be detected to obtain the edited text with the typos removed.

[0109] S2. Using a pre-trained language model, calculate the perplexity of the text to be detected and the edited text after deleting typos.

[0110] Specifically, perplexity is an indicator that measures the semantic fluency of a sentence; the more fluent a sentence is semantically, the lower its perplexity.

[0111] A language model is a probabilistic model used to calculate the probability that a sentence is a semantically correct sentence. Perplexity is a sentence-length-normalized metric related to the probability that a language model predicts a sentence. For a perfectly correct sentence, the lower the perplexity of the language model, the better the language model. Conversely, if a very good language model has been selected, then for a given sentence, if the perplexity of the language model is very low, it means that the sentence is highly likely to be correct.

[0112] In this step, to verify whether the previously identified typos are truly typos, the perplexity of the text to be detected and the edited text after deleting the typos are calculated respectively.

[0113] S3. If the perplexity of the edited text after deleting the typos is less than the perplexity of the text to be detected, and the absolute value of the difference between the two is greater than a set threshold, then the typos are taken as the final typo detection result; otherwise, the typos are removed from the final typo detection result.

[0114] Understandably, if the perplexity of the edited text after deleting the typo is less than that of the text to be detected, and the absolute value of the difference between the two is greater than a set threshold, it means that the semantics of the edited text after deleting the typo is more fluent than the semantics of the identified text before deletion. In other words, the deleted typo was indeed a typo, and therefore the deleted typo can be added to the final typo detection result. Conversely, if the perplexity is less than a set threshold, it indicates that the typo identified in the previous steps is a pseudo-typo, and it can be removed from the final typo detection result, meaning it will not be identified as a typo in the end.

[0115] In this embodiment, the accuracy of misspelling recognition is further improved by adding a post-processing operation that performs secondary verification of misspellings from the perspective of sentence semantic fluency.

[0116] The following describes the typo detection device in audio-related text provided in the embodiments of this application. The typo detection device in audio-related text described below can be referred to in correspondence with the typo detection method in audio-related text described above.

[0117] See Figure 8 , Figure 8 This is a schematic diagram of the structure of an audio-related text misspelling detection device disclosed in an embodiment of this application.

[0118] like Figure 8 As shown, the device may include:

[0119] Data acquisition unit 11 is used to acquire input audio and text to be detected related to the input audio;

[0120] The feature extraction unit 12 is used to extract the emotional modal features of the input audio and the text modal features of the text to be detected.

[0121] The feature fusion unit 13 is used to fuse the emotional modality features and the text modality features to obtain fused features;

[0122] The real text determination unit 14 is used to determine the real text corresponding to the text to be detected based on the fusion features;

[0123] The misspelling detection unit 15 is used to compare the real text and the text to be detected to obtain the misspelling detection result in the text to be detected.

[0124] Optionally, if the emotional modality features and the textual modality features are both in vector form, then the process by which the feature fusion unit fuses the emotional modality features and the textual modality features to obtain fused features may include:

[0125] A gating fusion method is used to fuse the vector-based emotional modality features and textual modality features to obtain fused features.

[0126] Optionally, the above gating fusion methods may include gating fusion methods of bitwise multiplication, bitwise addition or division, etc.

[0127] Optionally, after fusing the vector-based emotional modality features and textual modality features using a gated fusion method, the aforementioned feature fusion unit may further include:

[0128] The fused features are added to the text modal features to obtain residual fused features, which are used as the final fused features.

[0129] Optionally, before fusing the vector-based sentiment modality features and text modality features using a gated fusion method, the aforementioned feature fusion unit may further include:

[0130] The emotional modality features are subjected to representation shift and nonlinear transformation to obtain the processed emotional modality features.

[0131] Optionally, the processing of the above-mentioned feature extraction unit 12, feature fusion unit 13 and real text determination unit 14 can be implemented by a pre-trained audio text recognition model. The audio text recognition model is configured to extract the emotional modal features of the input audio, extract the text modal features of the text to be detected, fuse the emotional modal features and the text modal features, and predict the internal state representation of the real text corresponding to the text to be detected based on the fused features.

[0132] The audio text recognition model may include: an audio processing module, a text processing module, a multimodal fusion module, and an output module;

[0133] The audio processing module is used to extract the emotional modality features of the input audio.

[0134] The text processing module is used to extract the text modal features of the text to be detected;

[0135] A multimodal fusion module is used to fuse the emotional modal features and the textual modal features to obtain fused features;

[0136] The output module is used to determine the real text corresponding to the text to be detected based on the fusion features.

[0137] Optionally, the above-mentioned multimodal fusion module may further include:

[0138] The feature editing module is used to perform representation shift and nonlinear transformation on the emotional modality features to obtain the processed emotional modality features;

[0139] The gated fusion module is used to fuse the processed emotional modality features and the text modality features using a gated fusion method to obtain fused features;

[0140] The residual connection module is used to add the fused feature to the text modal feature to obtain the residual fused feature, which is used as the final fused feature.

[0141] Optionally, the above audio processing module may further include:

[0142] The pre-trained audio coding module is used to encode the audio features of the input audio using a pre-trained audio coding model to obtain the encoded emotional modality features. The audio coding model adopts an audio pre-trained model structure and is trained using emotion classification as the pre-training task.

[0143] The linear transformation module is used to perform a linear transformation on the dimensions of the emotional modality features to output emotional modality features with the same dimensions as the text modality features.

[0144] Optionally, the above text processing module may further include:

[0145] The text preprocessing module is used to edit the text to be detected to a set length by filling in set characters, and to determine the feature representation of the edited text to be detected.

[0146] The text modality feature extraction module is used to encode the feature representation of the text to be detected to obtain the text modality features of the text to be detected.

[0147] Optionally, the process by which the above-mentioned misspelling determination unit compares the real text and the text to be detected to obtain the misspelling detection result in the text to be detected may include:

[0148] The system checks whether there are any characters in the text to be detected that are inconsistent with the actual text. If such characters are found, they are treated as typos.

[0149] Optionally, the apparatus of this application may further include: a misspelling verification unit, configured to: after comparing the real text and the text to be detected to obtain a misspelling detection result in the text to be detected, delete the misspellings identified in the text to be detected to obtain an edited text with the misspellings removed; use a pre-trained language model to calculate the perplexity of the text to be detected and the edited text with the misspellings removed respectively; if the perplexity of the edited text with the misspellings removed is less than the perplexity of the text to be detected, and the absolute value of the difference between the two is greater than a set threshold, then the misspelling is taken as the final misspelling detection result; otherwise, the misspelling is removed from the final misspelling detection result.

[0150] Optionally, the apparatus of this application may further include: a misspelling marking unit, configured to: after comparing the real text and the text to be detected to obtain a misspelling detection result in the text to be detected, determine the position of the misspelling in the text to be detected; and mark the misspelling in the text to be detected according to the position.

[0151] The typo detection device in audio-related text provided in this application embodiment can be applied to typo detection devices in audio-related text, such as terminals: mobile phones, computers, etc. Optionally, Figure 9 The diagram shows the hardware structure of a typo detection device in audio-related text. Figure 9 The hardware structure of the device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;

[0152] In this embodiment of the application, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4;

[0153] Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0154] Memory 3 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;

[0155] The memory stores a program, which the processor can call. The program is used for:

[0156] Obtain the input audio and the text to be detected related to the input audio;

[0157] Extract the emotional modal features of the input audio, and extract the textual modal features of the text to be detected;

[0158] The emotional modality features and the text modality features are fused to obtain fused features;

[0159] The real text corresponding to the text to be detected is determined based on the fusion features;

[0160] By comparing the real text and the text to be detected, the typo detection results in the text to be detected are obtained.

[0161] Optionally, the refined and extended functions of the program can be found in the description above.

[0162] This application embodiment also provides a storage medium that can store a program suitable for execution by a processor, the program being used for:

[0163] Obtain the input audio and the text to be detected related to the input audio;

[0164] Extract the emotional modal features of the input audio, and extract the textual modal features of the text to be detected;

[0165] The emotional modality features and the text modality features are fused to obtain fused features;

[0166] The real text corresponding to the text to be detected is determined based on the fusion features;

[0167] By comparing the real text and the text to be detected, the typo detection results in the text to be detected are obtained.

[0168] Optionally, the refined and extended functions of the program can be found in the description above.

[0169] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0170] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0171] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for detecting typos in audio-related text, characterized in that, include: Obtain the input audio and the text to be detected related to the input audio; Extract the emotional modal features of the input audio, and extract the textual modal features of the text to be detected; The emotional modality features and the text modality features are fused to obtain fused features; The real text corresponding to the text to be detected is determined based on the fusion features; By comparing the real text and the text to be detected, the typo detection results in the text to be detected are obtained.

2. The method according to claim 1, characterized in that, The emotional modality features and the textual modality features are both in vector form; The process of fusing the emotional modality features and the textual modality features to obtain the fused features includes: A gating fusion method is used to fuse the vector-based emotional modality features and textual modality features to obtain fused features.

3. The method according to claim 2, characterized in that, After fusing vector-based sentiment modality features and text modality features using a gated fusion method, the process also includes: The fused features are added to the text modal features to obtain residual fused features, which are used as the final fused features.

4. The method according to claim 2, characterized in that, Before fusing vector-based sentiment modality features and text modality features using a gated fusion method, the following steps are also included: The emotional modality features are subjected to representation shift and nonlinear transformation to obtain the processed emotional modality features.

5. The method according to claim 1, characterized in that, The process of extracting the emotional modality features and text modality features and fusing them, and determining the real text corresponding to the text to be detected based on the fused features, is obtained through a pre-trained audio-text recognition model. The audio-text recognition model is configured to extract the emotional modality features of the input audio, extract the textual modality features of the input text to be detected, fuse the emotional modality features and the textual modality features, and predict the internal state representation of the real text corresponding to the text to be detected based on the fused features.

6. The method according to claim 5, characterized in that, The audio-text recognition model includes: an audio processing module, a text processing module, a multimodal fusion module, and an output module; The audio processing module is used to extract the emotional modality features of the input audio. The text processing module is used to extract the text modal features of the text to be detected; A multimodal fusion module is used to fuse the emotional modal features and the textual modal features to obtain fused features; The output module is used to determine the real text corresponding to the text to be detected based on the fusion features.

7. The method according to claim 6, characterized in that, The multimodal fusion module includes: The feature editing module is used to perform representation shift and nonlinear transformation on the emotional modality features to obtain the processed emotional modality features; The gated fusion module is used to fuse the processed emotional modality features and the text modality features using a gated fusion method to obtain fused features; The residual connection module is used to add the fused feature to the text modal feature to obtain the residual fused feature, which is used as the final fused feature.

8. The method according to claim 6, characterized in that, The audio processing module includes: The pre-trained audio coding module is used to encode the audio features of the input audio using a pre-trained audio coding model to obtain the encoded emotional modality features. The audio coding model adopts an audio pre-trained model structure and is trained using emotion classification as the pre-training task. The linear transformation module is used to perform a linear transformation on the dimensions of the emotional modality features to output emotional modality features with the same dimensions as the text modality features.

9. The method according to claim 6, characterized in that, The text processing module includes: The text preprocessing module is used to edit the text to be detected to a set length by filling in set characters, and to determine the feature representation of the edited text to be detected. The text modality feature extraction module is used to encode the feature representation of the text to be detected to obtain the text modality features of the text to be detected.

10. The method according to any one of claims 1-9, characterized in that, The process of comparing the real text and the text to be detected to obtain the typo detection results in the text to be detected includes: The system checks whether there are any characters in the text to be detected that are inconsistent with the actual text. If such characters are found, they are treated as typos.

11. The method according to any one of claims 1-9, characterized in that, After comparing the real text and the text to be detected to obtain the typo detection results in the text to be detected, the method further includes: The identified typos in the text to be detected are deleted to obtain the edited text after the typos are removed; Using a pre-trained language model, the perplexity of the text to be detected and the edited text after deleting typos are calculated respectively; If the perplexity of the edited text after deleting the typos is less than the perplexity of the text to be detected, and the absolute value of the difference between the two is greater than a set threshold, then the typo is taken as the final typo detection result; otherwise, the typo is removed from the final typo detection result.

12. The method according to any one of claims 1-9, characterized in that, After comparing the real text and the text to be detected to obtain the typo detection results in the text to be detected, the method further includes: Determine the location of the misspelled words in the text to be detected; According to the stated location, the misspelled words are marked in the text to be detected.

13. A typo detection device in audio-related text, characterized in that, include: The data acquisition unit is used to acquire input audio and text to be detected related to the input audio; The feature extraction unit is used to extract the emotional modal features of the input audio and the text modal features of the text to be detected. The feature fusion unit is used to fuse the emotional modality features and the text modality features to obtain fused features; The real text determination unit is used to determine the real text corresponding to the text to be detected based on the fusion features; The misspelling detection unit is used to compare the real text and the text to be detected to obtain the misspelling detection result in the text to be detected.

14. A typo detection device in audio-related text, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement each step of the method for detecting typos in audio-related text as described in any one of claims 1 to 12.

15. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the method for detecting typos in audio-related text as described in any one of claims 1 to 12.