Text recognition method and device

By acquiring candidate word groups from speech frames and calculating confidence based on language confidence, the problem of low accuracy in speech recognition under multi-dialect environments is solved, the recognition accuracy is improved and the user experience is enhanced, and flexible candidate word selection is provided.

CN120895041APending Publication Date: 2025-11-04VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511070018.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing speech recognition technologies have low accuracy when faced with multi-dialect environments, especially due to recognition errors caused by the similarity of dialect pronunciations. They also lack dynamic prediction capabilities and multi-stage adaptive mechanisms, which affects user experience.

Method used

By acquiring candidate word groups from speech frames, calculating confidence based on the language confidence of the candidate words' dialect, determining the target words, and combining the language confidence of the user's commonly used dialects to improve recognition accuracy, candidate word groups are provided for the user to choose from.

Benefits of technology

It improves the accuracy of speech recognition in multilingual scenarios for electronic devices, enhances user experience and system trust, overcomes the shortcomings of traditional single-output, and provides flexible modification and selection space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895041A_ABST
    Figure CN120895041A_ABST
Patent Text Reader

Abstract

The invention discloses a text recognition method and device, and belongs to the field of voice processing. The method comprises the steps that first voice is acquired, the first voice comprises N voice frames, and N is a positive integer; a candidate word group corresponding to each voice frame is obtained, each candidate word group comprises at least one candidate word, and each candidate word corresponds to one dialect language; determining the confidence coefficient of each candidate word based on the language confidence coefficient of the dialect language to which each candidate word in each candidate word group belongs; and determining a target word corresponding to each voice frame based on the confidence of all candidate words in the candidate word group corresponding to each voice frame, and generating a text corresponding to the first voice.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of speech processing, and particularly relates to a text recognition method and device thereof. BACKGROUND

[0002] With the development of automatic speech recognition (ASR), ASR is increasingly applied in electronic devices, such as voice input method, intelligent assistant, voice search and voice control.

[0003] At present, there are a large number of different dialects in many regions. Generally, after an electronic device collects dialect speech data of a user, the electronic device determines a target candidate word corresponding to each speech frame from a standard candidate word library according to speech feature information corresponding to each speech frame in the speech data. However, different dialects may have similar pronunciations, but different corresponding characters. Therefore, the determined target candidate word may not be the same as the character that the user wants to transcribe, thereby causing the electronic device to have a low accuracy in transcribing dialect speech. SUMMARY

[0004] An embodiment of the present application aims to provide a text recognition method and device thereof, which can improve the accuracy of an electronic device in recognizing dialect speech text of a user.

[0005] In a first aspect, an embodiment of the present application provides a text recognition method, which comprises: obtaining first speech, the first speech comprising N speech frames, N being a positive integer; obtaining a candidate word group corresponding to each speech frame, each candidate word group comprising at least one candidate word, and each candidate word corresponding to a dialect language; determining a confidence of each candidate word based on a language confidence of a dialect language to which each candidate word in each candidate word group belongs; determining a target word corresponding to each speech frame based on the confidence of all candidate words in the candidate word group corresponding to each speech frame; and generating a text corresponding to the first speech based on each target word.

[0006] In a second aspect, an embodiment of the present application provides a text recognition device, which comprises: an obtaining module and a processing module; the obtaining module is configured to obtain first speech, the first speech comprising N speech frames, N being a positive integer; the obtaining module is further configured to obtain a candidate word group corresponding to each speech frame, each candidate word group comprising at least one candidate word, and each candidate word corresponding to a dialect language; the processing module is configured to determine a confidence of each candidate word based on a language confidence of a dialect language to which each candidate word in each candidate word group belongs; the processing module is further configured to determine a target word corresponding to each speech frame based on the confidence of all candidate words in the candidate word group corresponding to each speech frame; and the processing module is further configured to generate a text corresponding to the first speech based on each target word.

[0007] In a third aspect, an electronic device is provided, which includes a processor and a memory. The memory stores programs or instructions executable on the processor. The programs or instructions, when executed by the processor, implement the steps of the method according to the first aspect.

[0008] In a fourth aspect, a readable storage medium is provided, which stores programs or instructions. The programs or instructions, when executed by a processor, implement the steps of the method according to the first aspect.

[0009] In a fifth aspect, a chip is provided, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to execute programs or instructions to implement the method according to the first aspect.

[0010] In a sixth aspect, a computer program product is provided, which is stored in a storage medium. The computer program product is executed by at least one processor to implement the method according to the first aspect.

[0011] In the embodiments of the present application, the electronic device obtains a first speech, the first speech includes N speech frames, N is a positive integer; obtains a candidate word group corresponding to each speech frame, each candidate word group includes at least one candidate word, and each candidate word corresponds to a dialect language; determines the confidence of each candidate word based on the language confidence of the dialect language to which each candidate word in each candidate word group belongs; determines a target word corresponding to each speech frame based on the confidence of all candidate words in the candidate word group corresponding to each speech frame; and generates a text corresponding to the first speech based on each target word. In this scheme, the electronic device obtains candidate words corresponding to different dialects in the first speech, calculates the probability that each candidate word is a word corresponding to a speech frame by combining the language confidence of the dialect language commonly used by the user, to determine the target word that best fits the current dialect language environment, thereby improving the accuracy of the electronic device in recognizing the text corresponding to the speech in a multi-dialect language environment. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 is a schematic diagram of a text recognition method provided by an embodiment of the present application;

[0013] Figure 2 is an example schematic diagram of an electronic device displaying a candidate word group provided by an embodiment of the present application;

[0014] Figure 3 is a flow example schematic diagram of a text recognition method provided by an embodiment of the present application;

[0015] Figure 4is a structural schematic diagram of a text recognition device provided by an embodiment of the present application;

[0016] Figure 5 is a structural schematic diagram of a text recognition device provided by an embodiment of the present application;

[0017] Figure 6 is a structural schematic diagram of a text recognition device provided by an embodiment of the present application;

[0018] Figure 7 is one of hardware structural schematic diagrams of an electronic device provided by an embodiment of the present application;

[0019] Figure 8 is one of hardware structural schematic diagrams of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0020] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art belong to the scope of protection of the present application.

[0021] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a category and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification means at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in an "or" relationship.

[0022] The terms "at least one", "at least one of", and the like in the specification and claims of the present application refer to any one of the objects, a combination of any two or more of the objects. For example, at least one of a, b, and c can mean "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more, and has a similar meaning to "at least one".

[0023] The identifier in the present application is a character, symbol, image, etc. used to indicate information, which can use an identifier or other container as a carrier to display information, including but not limited to character identifier, image identifier, symbol identifier, etc.

[0024] It should be noted that the text recognition method provided in the embodiments of the present application can be executed by an electronic device such as a mobile phone, a tablet computer, a notebook computer, a palm computer, and a vehicle-mounted electronic device. In some embodiments of the present application, the text recognition method is executed by an electronic device as an execution subject, and the text recognition method provided in the embodiments of the present application is described.

[0025] Currently, the accuracy of speech recognition technology varies in different use scenarios and populations, especially in multi-lingual and multi-dialect scenarios. Existing speech recognition technology is usually based on deep learning algorithms, and a large amount of speech data sets are used for model training. The speech transcription, i.e., the conversion of speech into text, is completed by an acoustic model, a language model, and a decoding algorithm.

[0026] However, the processing of dialect and regional speech features still faces the following challenges:

[0027] 1. Diverse dialect features: There are a large number of different dialects in different countries and regions. The pronunciation of many dialects is similar, but the corresponding meaning has subtle differences, which brings great difficulty to the model in distinguishing specific speech features in the speech signal. For example, words with similar pronunciation in some dialects may seriously affect the accuracy of recognition.

[0028] To this end, although existing speech recognition systems can recognize different dialects in speech, in general, existing speech recognition systems only support standard Mandarin well, and have weak recognition ability for dialects. In the country, due to the large number of dialects and the similarity of dialect pronunciation to varying degrees, for example, it is difficult to distinguish the pronunciation features between some southern regions, which leads to confusion in the model of the existing speech recognition system when recognizing dialects in speech, and thus affects the accuracy of the final speech recognition.

[0029] 2. Limited data resources: Existing speech data resources are relatively rich in mainstream Mandarin or standard speech, but there is less high-quality annotated data for dialects, which limits the fine modeling and optimization of dialects. Existing dialect recognition technology highly depends on long-term corpus training, i.e., a large amount of existing speech data resources are needed to train the dialect recognition technology. However, due to the limitations of dialect corpus collection, such as the dynamicity, incompleteness, and insufficient coverage of speech data, even after a large amount of training, there is still a significant gap in the accurate definition and discrimination of dialect speech in real use scenarios.

[0030] 3. Dynamic changes in user environment: In the actual use of speech recognition, the user's environment may be complex, such as background noise, unstable network, and the user's pronunciation habit changes dynamically, which further affects the recognition effect of the speech recognition system. Therefore, the current speech recognition system lacks dynamic prediction ability for complex dialect environments. Even if certain language-specific features are introduced in the dialect recognition model, there is often a lack of "dynamic screening" and "multi-stage adaptive" deep processing mechanism, which makes it difficult to intelligently optimize according to the context when facing multiple candidate transcription results.

[0031] Moreover, in both Mandarin and dialect recognition, when there are similar pronunciation and ambiguous phonemes in the speech segment, the system has low confidence in a single recognition result, which can easily lead to incorrect transcription. In the current transcription process, there is no subsequent re-screening and optimization step among multiple candidate results. Since users have low tolerance for speech recognition accuracy, frequent errors in dialect speech recognition not only lead to failed interactions, but also significantly affect user experience and trust in the system. This is a major obstacle to the popularization of speech recognition technology.

[0032] Although there are some methods that attempt to alleviate the problem of diverse dialect pronunciation, such as introducing a multi-lingual model or using an end-to-end model to optimize the speech recognition system, these methods often result in increased computational overhead and increased demand for device resources in real-world scenarios. At the same time, the current speech recognition system relies on a single recognition result strategy when faced with complex dialect contexts, which can easily lead to misrecognition, such as incorrect word selection, which can severely affect user experience.

[0033] To this end, in the text recognition method provided by the present application, an electronic device obtains a first speech, the first speech including N speech frames, N being a positive integer; obtains a candidate word group corresponding to each speech frame, each candidate word group containing at least one candidate word, each candidate word corresponding to a dialect language; determines the confidence of each candidate word based on the language confidence of the dialect language to which each candidate word in each candidate word group belongs; determines a target word corresponding to each speech frame based on the confidence of all candidate words in the candidate word group corresponding to each speech frame; and generates a text corresponding to the first speech based on each target word. In this solution, the electronic device obtains candidate words corresponding to different dialects in the first speech, calculates the probability that each candidate word is the word of the corresponding speech frame based on the language confidence of the dialect language commonly used by the user, and determines the target word that best fits the current dialect language environment, thereby improving the accuracy of the electronic device in recognizing the text corresponding to the speech in a multi-dialect language environment.

[0034] The text recognition method and device provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings and specific examples and application scenarios.

[0035] The text recognition method provided by the embodiments of the present application can be executed by a text recognition device. The text recognition device can be an electronic device or a component of the electronic device, such as an integrated circuit or a chip. The text recognition method provided by the embodiments of the present application will be described below by way of example with reference to an electronic device.

[0036] The text recognition method provided by the embodiments of the present application can be executed by a text recognition device. The text recognition device can be an electronic device or a component of the electronic device, such as an integrated circuit or a chip. The text recognition method provided by the embodiments of the present application will be described below by way of example with reference to an electronic device. Figure 1 A flowchart of a text recognition method provided by the embodiments of the present application is shown. The method can be applied to an electronic device. As shown in Figure 1 The text recognition method provided by the embodiments of the present application can include the following steps 201 to 205.

[0037] Step 201: The electronic device acquires a first voice.

[0038] In some embodiments of the present application, the first voice includes but is not limited to a user voice, a video voice, an AI-generated voice, a downloaded voice on the network, or a call voice.

[0039] By way of example, the first voice is voice data acquired by the electronic device when the user uses a voice input method, a voice assistant, or an application program with a voice recognition function.

[0040] By way of example, the electronic device can use a microphone to collect the first voice input.

[0041] In some embodiments of the present application, the first voice includes N voice frames, and N is a positive integer.

[0042] In some embodiments of the present application, the voice frame is a voice segment obtained by the electronic device by dividing the first voice.

[0043] By way of example, the electronic device can divide the first voice into N voice frames based on a preset time length.

[0044] By way of example, the electronic device divides a segment of the first voice into N voice frames according to a preset time length, that is, the first voice is 10s, and if the preset time length is 1s, the first voice can be divided into 10 voice frames.

[0045] It can be understood that the preset time length is user-defined or set by default by the electronic device, and is usually the time length for the user to speak a word.

[0046] Exemplarily, the electronic device can divide the first speech into N speech frames according to the pronunciation features such as volume, tone, etc. in the first speech.

[0047] In step 202, the electronic device obtains a candidate word group corresponding to each speech frame.

[0048] In some embodiments of the present application, each candidate word group contains at least one candidate word.

[0049] In some embodiments of the present application, each candidate word corresponds to a dialect language.

[0050] In some embodiments of the present application, the electronic device obtains speech feature information corresponding to each speech frame, and determines the candidate word group corresponding to each speech frame according to the speech feature information corresponding to each speech frame.

[0051] In some embodiments of the present application, the electronic device determines the candidate word corresponding to each speech frame according to the feature similarity between the speech feature information corresponding to each speech frame and the speech feature information of the candidate words in the candidate word library of different dialect languages, to constitute the candidate word group corresponding to each speech frame.

[0052] In some embodiments of the present application, the speech feature information corresponding to each speech frame can include at least one speech feature.

[0053] In some embodiments of the present application, the at least one speech feature can include at least one of the following: speech tone information, speech volume information, speech tone information, and speech pitch information.

[0054] It can be understood that the various speech features listed above are exemplary and the embodiments of the present application include but are not limited to the various speech features listed above. In actual implementation, the speech features can also include other possible speech features, which can be determined according to actual use requirements, and the embodiments of the present application are not limited.

[0055] In some embodiments of the present application, the speech feature information corresponding to each speech frame can be a speech vector feature of the speech frame.

[0056] For example, the speech vector feature of a speech frame is: [12.3, -0.8, 1.2, …, 0.4].

[0057] It should be noted that the speech vector feature can be 13-dimensional or higher-dimensional, and the present application is not limited.

[0058] In some embodiments of the present application, before obtaining the candidate word group corresponding to each speech frame, the electronic device will first obtain the speech feature information corresponding to each speech frame.

[0059] In some embodiments of the present application, the electronic device compresses the first speech into an M-dimensional MFCC through a mel coefficient algorithm, generates a time sequence feature matrix based on the speech order of the speech frames in the first speech, and then the electronic device performs a smoothing weighting algorithm on the time sequence feature matrix to obtain a global vectorization matrix. The vectorization matrix includes a group of vector groups of each speech frame.

[0060] Illustratively, the electronic device first pre-processes the first speech through a short-time Fourier transform, adopts a Hamming window frame processing, and obtains a time-frequency spectrum matrix of the first speech through a 512-point Fast Fourier Transform (FFT) to extract the feature information of the audio signal of each speech frame. Then, a 40-channel mel filter bank is used for nonlinear frequency analysis, the log energy is calculated, and then a Discrete Cosine Transform (DCT) is used to extract the first M-dimensional Mel-frequency cepstral coefficients (MFCC) of each speech frame. Then, the electronic device constructs a time sequence feature matrix based on the M-dimensional MFCC of each speech frame according to the time sequence of each speech frame. Finally, a sliding average weighting algorithm is used to vectorize the time sequence feature matrix, and the first speech is compressed into an M-dimensional global feature vector.

[0061] It should be noted that the number of M in the above M-dimensional coefficient can be pre-set or user-defined.

[0062] In some embodiments of the present application, taking one speech frame in N speech frames as an example, the electronic device obtains the speech feature information corresponding to the speech frame, matches the speech feature information of the speech frame with the candidate word library of different dialects, such as cosine similarity calculation, obtains the word with the highest semantic feature information similarity in each candidate word library, and takes these words as the candidate words in the candidate word group of the speech frame.

[0063] In some embodiments of the present application, the above-mentioned candidate word library is a word library of multiple dialects pre-stored in the electronic device. Different dialects correspond to different candidate word libraries.

[0064] In other words, the candidate words in different candidate word libraries belong to different dialects, and the candidate words in the same candidate word library belong to the same dialect.

[0065] For example, the electronic device matches the speech feature information corresponding to the speech frame 1 with the candidate words in the candidate word library of the dialect language of Wuhan dialect, obtains the candidate word A with the highest similarity, then matches the speech feature information corresponding to the speech frame 1 with the candidate words in the candidate word library of the dialect language of Mandarin, obtains the candidate word B with the highest similarity, then matches the speech feature information corresponding to the speech frame 1 with the candidate words in the candidate word library of the dialect language of Cantonese, obtains the candidate word C with the highest similarity, and finally forms the candidate word group corresponding to the speech frame 1: candidate word A, candidate word B, and candidate word C.

[0066] In step 203, the electronic device determines the confidence of each candidate word based on the language confidence of the dialect language to which the candidate word belongs.

[0067] In some embodiments of the present application, the language confidence is used to represent the frequency of use of the dialect language.

[0068] It can be understood that the higher the language confidence is, the higher the frequency of use of the dialect language corresponding to the language confidence is.

[0069] In some embodiments of the present application, one language confidence corresponds to one dialect language.

[0070] In some embodiments of the present application, the language confidence is determined according to the ratio between the total number of historical recognized words stored in the electronic device and the total number of words of each dialect language.

[0071] For example, the historical recognized words stored in the electronic device can be stored in the language word library set by the electronic device.

[0072] It can be understood that when the language word library is updated, for example, when a word is added to the language word library, the language confidence of different dialect languages corresponding to the word will also be updated.

[0073] For example, the candidate word A in a candidate word group belongs to the dialect language of Cantonese, and the Cantonese language confidence value of the candidate word A is 38%; the candidate word B in a candidate word group belongs to the dialect language of Mandarin, and the Mandarin language confidence value of the candidate word B is 17%; and the candidate word C in a candidate word group belongs to the dialect language of Wuhan dialect, and the Wuhan dialect language confidence value of the candidate word C is 55%. Therefore, the electronic device determines that the dialect language of the speech frame is Wuhan dialect.

[0074] For example, the language confidence can be in the form of percentage or score, which is not shown in the present application.

[0075] It can be understood that the above language confidence is related to the user's historical use of dialects, that is, the first voice frequently obtained by the electronic device is a certain dialect, and the language confidence value corresponding to the dialect is higher.

[0076] In some embodiments of the present application, the confidence of each candidate word is used to represent the probability of each candidate word being the word of the corresponding voice frame.

[0077] For example, the higher the confidence of a candidate word, the greater the probability that the candidate word is the word of the corresponding voice frame.

[0078] For example, the confidence of candidate word A of a voice frame is 20%, the confidence of candidate word B is 78%, and the confidence of candidate word C is 2%, which indicates that candidate word B has the greatest probability of being the word of the corresponding voice frame, that is, it is the most likely word of the voice frame.

[0079] For example, the confidence of the above candidate word can be in percentage or score form, which is not shown in the present application.

[0080] In some embodiments of the present application, the confidence of the above candidate word is calculated based on the confidence of the dialect language to which the candidate word belongs, the feature similarity between the voice feature information of the candidate word and the voice feature information of the voice frame.

[0081] For example, the electronic device multiplies the feature similarity between the voice feature information of the candidate word and the voice feature information of the voice frame by the confidence of the dialect language to which the candidate word belongs to obtain the confidence of the candidate word.

[0082] For example, taking a candidate word corresponding to a voice frame as an example, the electronic device obtains the feature similarity value between the candidate word and the voice frame, then takes the language confidence of the dialect language to which the candidate word belongs as a weight value, and then performs weighted calculation on the feature similarity value using the weight value to obtain the confidence of the candidate word.

[0083] It can be understood that the confidence of each candidate word for each voice frame can be calculated in the above manner.

[0084] Step 204, the electronic device determines the target word corresponding to each voice frame based on the confidence of all candidate words in the candidate word group corresponding to each voice frame.

[0085] In some embodiments of the present application, the target word corresponding to each voice frame is determined by the electronic device according to the confidence of each candidate word.

[0086] In some embodiments of the present application, the target word corresponding to each voice frame is determined by the electronic device according to the confidence value, or is selected by the user.

[0087] Exemplarily, taking a candidate word group corresponding to one speech frame as an example, in a case where the confidence of a candidate word in the candidate word group is greater than a preset threshold, the electronic device automatically determines the candidate word as a target word corresponding to the speech frame of the candidate word group; in a case where the confidence of any candidate word in the candidate word group is not greater than the preset threshold, the electronic device determines a candidate word selected by the user as the target word corresponding to the speech frame of the candidate word group.

[0088] It can be understood that the process of determining the target word in the candidate word group corresponding to each speech frame in the embodiment can be determined by referring to the above steps.

[0089] In step 205, the electronic device generates text corresponding to the first speech based on each target word.

[0090] In some embodiments of the present application, the electronic device composes the target words corresponding to each speech frame into text corresponding to the first speech according to the time sequence of each speech frame.

[0091] In some embodiments of the present application, the time sequence of each speech frame is the time sequence of the first speech input.

[0092] In the text recognition method provided in the embodiments of the present application, the electronic device obtains a first speech, the first speech including N speech frames, N being a positive integer; obtains a candidate word group corresponding to each speech frame, each candidate word group including at least one candidate word, each candidate word corresponding to a dialect language; determines the confidence of each candidate word based on the language confidence of the dialect language to which each candidate word in each candidate word group belongs; determines a target word corresponding to each speech frame based on the confidence of all candidate words in the candidate word group corresponding to each speech frame; and generates text corresponding to the first speech based on each target word. In the present scheme, the electronic device determines the target word most suitable for the current dialect language environment by obtaining the candidate words corresponding to different dialects in the first speech, combining the language confidence of the dialect language commonly used by the user to calculate the probability that each candidate word is the word corresponding to the speech frame, so as to improve the accuracy of the electronic device in recognizing the text corresponding to the speech in a multi-dialect language environment.

[0093] Optionally, in some embodiments of the present application, the step 203 "the electronic device determines the confidence of each candidate word based on the language confidence of the dialect language to which each candidate word in each candidate word group belongs" can be implemented by the following step 203a.

[0094] In step 203a, the electronic device determines the confidence of the first candidate word based on the language confidence of the dialect language to which the first candidate word in the first candidate word group corresponding to the first speech frame belongs, and the feature similarity between the speech feature information of the first candidate word and the speech feature information of the first speech frame.

[0095] In some embodiments of the present application, the first speech frame is one of the N speech frames.

[0096] In some embodiments of the present application, the first candidate word is one of the first candidate word group.

[0097] In some embodiments of the present application, the feature similarity is the similarity between the speech feature information of the first speech frame and the speech feature information of the first candidate word.

[0098] In some embodiments of the present application, the feature similarity is used to represent the matching degree between the first speech frame and the first candidate word.

[0099] For example, the electronic device calculates the similarity between the speech feature information of the first speech frame and the speech feature information of the first candidate word by using the cosine similarity algorithm, i.e., the feature similarity.

[0100] In some embodiments of the present application, the electronic device determines the candidate word with the largest feature similarity in the at least one candidate word library as the at least one first candidate word.

[0101] For example, since different candidate word libraries correspond to different dialect languages, the electronic device sequentially performs similarity matching between the speech feature information of the first speech frame and each candidate word in the candidate word library corresponding to the dialect language to which the language confidence belongs, based on the ranking of the current language confidence, i.e., from high to low.

[0102] For example, the ranking of the current language confidence from high to low is: Cantonese, Mandarin, Wuhan dialect, and Ganzhou dialect. The electronic device first performs similarity matching between the speech feature information of the first speech frame and the candidate words in the candidate word library corresponding to Cantonese, to determine a candidate word with Cantonese as the dialect language; then performs similarity matching between the speech feature information of the first speech frame and the candidate words in the candidate word library corresponding to Mandarin, to determine a candidate word with Mandarin as the dialect language; then performs similarity matching between the speech feature information of the first speech frame and the candidate words in the candidate word library corresponding to Wuhan dialect, to determine a candidate word with Wuhan dialect as the dialect language; and finally performs similarity matching between the speech feature information of the first speech frame and the candidate words in the candidate word library corresponding to Ganzhou dialect, to determine a candidate word with Ganzhou dialect as the dialect language.

[0103] In some embodiments of the present application, after the electronic device obtains the feature similarity between the first speech frame and the first candidate word, the electronic device weights the feature similarity value between the first speech frame and the first candidate word based on the language weight value corresponding to the first candidate word to obtain a weighted feature similarity value, and determines the weighted feature similarity value as the confidence of the first candidate word.

[0104] In some embodiments of the present application, the language weight value is used to represent the probability that the dialect language to which the first candidate word belongs is the dialect language of the first speech frame.

[0105] In some embodiments of the present application, the language weight value is determined based on the language confidence of the dialect language to which the first candidate word belongs. That is, the language confidence of a dialect language corresponds to a language weight value.

[0106] For example, in one way, the electronic device can directly use the language confidence as the corresponding language weight value; in another way, the electronic device can calculate a language weight value from the language confidence through conversion calculation or the like.

[0107] In this way, the electronic device can calculate the confidence of each candidate word in combination with the language confidence, so as to ensure that the confidence of the candidate word corresponding to the dialect language commonly used by the user is the highest, thereby improving the accuracy of the electronic device in recognizing the text of the speech.

[0108] Optionally, in some embodiments of the present application, the step 204 "the electronic device determines the target word corresponding to each speech frame based on the confidence of all candidate words in the candidate word group corresponding to each speech frame" can be implemented through the following steps 204a and 204b.

[0109] Step 204a: In the case that the confidence of at least one candidate word in the second candidate word group corresponding to the second speech frame is greater than a predetermined threshold, the electronic device determines the candidate word with the highest confidence among the at least one candidate word as the target word corresponding to the second speech frame.

[0110] Step 204b: In the case that the confidence of all candidate words in the second candidate word group corresponding to the second speech frame is not greater than the predetermined threshold, the electronic device determines the candidate word selected by the user as the target word corresponding to the second speech frame.

[0111] In some embodiments of the present application, the second speech frame is one of the N speech frames.

[0112] In some embodiments of the present application, the predetermined threshold is user-defined or preset by the electronic device, and can be set to 90% or 95%.

[0113] In some embodiments of the present application, the confidence level of at least one candidate word in the above-mentioned second candidate phrase is greater than a predetermined threshold, indicating that there is at least one candidate word in the second candidate phrase that can be directly transcribed. The electronic device can directly use these candidate words as the target words of the second speech frame. At this time, the electronic device determines the candidate word with the highest confidence level as the target word of the second speech frame.

[0114] It can be understood that since the higher the confidence level of a candidate word, the greater the probability that the candidate word is the target word of the corresponding speech frame. Therefore, the electronic device determines the candidate word with the highest confidence level as the target word of the second speech frame.

[0115] For example, taking the first speech "佢份工好辛苦,但揾钱唔少" as an example, among the candidate phrases of the fifth speech frame, that is, "辛苦", "新库", and "心苦", the confidence level of "辛苦" is the highest, so "辛苦" is used as the target word corresponding to the fifth speech frame.

[0116] In some embodiments of the present application, the confidence levels of all candidate words in the above-mentioned second candidate phrase are not greater than a predetermined threshold, indicating that there are no candidate words in the second candidate phrase that the electronic device can directly transcribe. Therefore, it is impossible to directly determine the target word corresponding to the second speech frame. Thus, the electronic device needs to display the candidate phrase for the user to select, and determine the candidate word selected by the user as the target word of the corresponding speech frame.

[0117] For example, taking the first speech "佢份工好辛苦,但揾钱唔少" as an example, among the candidate phrases of the seventh speech frame, that is, "揾", "问", "温", and "稳", the confidence levels of all of them are not greater than the preset threshold. Then, "揾", "问", "温", and "稳" are displayed on the interface for the user to select. When the user selects "揾", the electronic device uses this candidate word as the target candidate word.

[0118] In this way, in view of the situation that multiple semantic interpretations are likely to occur for multi-regional dialects, the concept of candidate phrases is proposed in this embodiment. During the speech recognition process, the system not only gives a single candidate result, but also intelligently generates multiple possible recognition results. The design of candidate phrases overcomes the deficiency of traditional single-time output, providing users with a more flexible space for modification and selection.

[0119] Optionally, in some embodiments of the present application, in combination with the above steps 204a and 204b, when the candidate phrase is displayed in the above-mentioned first interface, the text recognition provided by the embodiments of the present application further includes steps 301 and 302.

[0120] Step 301, in a case that the confidence of each candidate word in the candidate word group corresponding to the at least one speech frame is not greater than the predetermined threshold, and the target words corresponding to the other speech frames except the at least one speech frame are determined, the electronic device displays the target words corresponding to each of the other speech frames and each candidate word in the candidate word group corresponding to the at least one speech frame according to the speech order of the N speech frames.

[0121] In some embodiments of the present application, after determining whether the confidence of each candidate word is greater than the preset threshold, the electronic device displays the target words corresponding to the other speech frames and each candidate word in the candidate word group corresponding to the at least one speech frame according to the speech order of the N speech frames.

[0122] In some embodiments of the present application, the electronic device displays the target words corresponding to the other speech frames and each candidate word in the candidate word group corresponding to the at least one speech frame in the first interface.

[0123] For example, the first interface includes at least one target word and at least one candidate word group.

[0124] For example, as shown in FIG. 2, the electronic device collects 10 speech frames corresponding to the first speech "He works hard, but earns a lot of money", wherein the confidence of each candidate word in the candidate word group corresponding to the first to sixth speech frames and the eighth to tenth speech frames is greater than the preset threshold, so the electronic device directly determines these words as target words, and the confidence of each candidate word in the candidate word group corresponding to the seventh speech frame is not greater than the preset threshold, so the electronic device needs to display all candidate words in the candidate word group 22 in the first interface 21 for the user to select. Figure 2

[0125] Step 302, the electronic device responds to the input of the user to the candidate word in the displayed candidate word group, and takes the candidate word selected by the user as the target word corresponding to the speech frame.

[0126] In some embodiments of the present application, the electronic device receives the input of the user to the candidate word in the displayed candidate word group.

[0127] In some embodiments of the present application, the input is used to determine the target word.

[0128] ​In some embodiments of the present application, when multiple candidate words in a candidate phrase are displayed, the above-mentioned first input includes: the user's selection input for the candidate words in the candidate phrase. The selection input may be at least one of the following: ticking, clicking, swiping, and long pressing.

[0129] In some embodiments of the present application, after receiving the above input, if the selected candidate word stays at a preset position for a preset duration, the electronic device will use the candidate word selected by the user as the target word for the corresponding speech frame.

[0130] Exemplarily, the above preset position is the position where the word of this speech frame in the generated text should be displayed.

[0131] Exemplarily, the above preset duration can be user-defined or preset by the electronic device. For example, 0.5s.

[0132] For example, in combination Figure 2 As shown, the electronic device receives the input that the user swipes the character "揾" to the front area of the character "钱". After 0.5s, the electronic device automatically determines the character "揾" as the target word for this speech frame.

[0133] In some embodiments of the present application, after the electronic device selects a candidate word as the target word in the candidate phrase, it updates the first interface to display the target word corresponding to each speech frame to generate the text corresponding to the first speech.

[0134] In this way, by setting the candidate phrase, the electronic device overcomes the deficiency of traditional single output and provides the user with a more flexible modification and selection space.

[0135] Optionally, in some embodiments of the present application, the language confidence of the above dialect language is calculated based on a language word library; after the above step 205 "the electronic device generates the text corresponding to the first speech based on each target word", the text recognition method provided by the embodiments of the present application further includes step 401 and step 402.

[0136] Step 401: The electronic device adds the target word corresponding to each speech frame to the language word library.

[0137] In some embodiments of the present application, the above language word library is a word library that stores words according to the dialect language standard.

[0138] In some embodiments of the present application, the above language word library includes at least one historical word recognized by the electronic device.

[0139] Exemplarily, the above language word library includes at least one word set, and each word set corresponds to a dialect language.

[0140] For example, the above-mentioned word set includes: at least one dialect language whose historical words belong to the dialect language to which the word set belongs.

[0141] Understandably, when an electronic device recognizes a word, it will add that word to the language lexicon, specifically to the word set corresponding to the dialect.

[0142] Step 402: The electronic device updates the language confidence level for each dialect based on the number of words corresponding to each dialect in the language lexicon and the total number of words in the language lexicon.

[0143] In some embodiments of this application, the language confidence level corresponding to each dialect is the ratio of the number of words corresponding to the dialect to the total number.

[0144] In some embodiments of this application, since the electronic device adds the target word to the word set of the dialect in the language lexicon, the total number of words in the language lexicon and the number of words in the corresponding word set will increase. Therefore, based on the number of words in the language lexicon after the addition, the electronic device calculates the ratio of the number of words corresponding to each dialect to the total number of words in the language lexicon, and obtains the language confidence level corresponding to each dialect.

[0145] For example, if a language lexicon includes 100 historical words, and the Cantonese word set contains 50 of them, then the corresponding language confidence value is 50%; the Mandarin word set contains 20 of them, then the corresponding language confidence value is 20%; the Wuhan dialect word set contains 16 of them, then the corresponding language confidence value is 16%; and the Ganzhou dialect word set contains 14 of them, then the corresponding language confidence value is 14%.

[0146] In this way, electronic devices can dynamically update language confidence based on the identified target words, which not only improves the accuracy of dialect recognition but also significantly optimizes the user experience in complex dialect environments.

[0147] The following specific examples illustrate the text recognition method provided in the embodiments of this application.

[0148] Example 1: As Figure 3 As shown, the text recognition method may include the following steps S1 to S5.

[0149] Step S1: The electronic device performs acoustic feature vectorization analysis on the first speech.

[0150] The following uses the first voice message "His job is very hard, but he earns a lot of money" as an example to explain in detail the specific implementation process of electronic devices acquiring voice feature information, including the following steps A1 to A3.

[0151] Step A1, the electronic device pre-processes and extracts feature information of "He works hard, but earns a lot of money".

[0152] Step A1-1, the electronic device uses Short-Time Fourier Transform (STFT) to pre-process and extract features of the user input voice.

[0153] For example, assuming that the voice duration of "He works hard, but earns a lot of money" is 2 seconds and the sampling rate is 16 kHz, formula (1) can be used to calculate the feature information of each frame of voice.

[0154]

[0155] Wherein, each parameter is defined as shown in Table 1:

[0156]

[0157] Table 1

[0158] Step A1-2, the electronic device uses a Mel filter bank to further filter the feature information of each frame of voice.

[0159] For example, the electronic device needs to pass each frame of voice through 40 Mel filters, filter each frame of voice, and convert each frame of voice into Mel coefficients.

[0160] For example, the formula of the filter bank is shown in formula (2) as follows:

[0161]

[0162] Wherein, each parameter is defined as shown in Table 2:

[0163]

[0164] Table 2

[0165] It should be noted that Hm(k) represents the energy of the mth filter.

[0166] Step A1-3, the electronic device adopts MFCC calculation and takes the first 13 coefficients.

[0167] For example, the electronic device uses formula (3) for MFCC calculation.

[0168]

[0169] Wherein, n represents the nth coefficient, in the embodiment, n is [1, 13], and m also takes the first 13 to obtain 13-dimensional coefficients corresponding to each speech frame.

[0170] Step A2, the electronic device constructs a time sequence feature matrix based on the 13-dimensional coefficients corresponding to each speech frame.

[0171] Exemplarily, assuming that the electronic device divides the above-mentioned speech "He works hard, but earns a lot of money" into 10 frames, and constructs a matrix according to the time sequence of each speech frame, and according to the 13-dimensional coefficients corresponding to each speech frame, a 10*3 matrix is obtained:

[0172] Frame 1: [12.3, -0.8, 1.2,..., 0.4];

[0173] Frame 2: [11.9, -1.1, 0.9,..., 0.3];

[0174] Frame 3: [11.9, -1.1, 0.9,..., 0.3];

[0175] Frame 4: [213, -5.2, 0.9,..., 0.8]; ... ...

[0178] Frame 10: [13.1, 0.2, 1.5,..., 0.6].

[0179] Step A3, the electronic device uses a sliding average weighting algorithm to vectorize the time sequence feature matrix to obtain a vector matrix for representing speech feature information of each speech frame in the first speech.

[0180] Exemplarily, the electronic device uses the following formula (4) to vectorize the time sequence feature matrix.

[0181]

[0182] Wherein, v is the vector matrix, T is the Tth vector in the speech frame, and a is a smoothing coefficient.

[0183] Assuming that the smoothing coefficient a = 0.95, after calculating the above-mentioned time sequence feature matrix, the following is finally obtained:

[0184] Frame 1: [12.3*0.95, -0.8*0.95,..., 0.4*0.95];

[0185] Frame 2: [11.9*0.90, -1.1*0.90,..., 0.3*0.90]; .... ....

[0188] ....。

[0189] Finally, a vector matrix is obtained, that is, the vector matrix corresponding to the speech feature information corresponding to each speech frame.

[0190] And at the same time, the vector v corresponding to the first speech is obtained:

[0191] v = ([12.3 * 0.95 + 11.9 * 0.90 +... + item of frame 10] / 10,

[0192] [-0.8 * 0.95 + (-1.1) * 0.90 +... + item of frame 10] / 10, ... ... ...

[0196] [0.4 * 0.95 + 0.3 * 0.90 +... + item of frame 10] / 10).

[0197] In this way, the electronic device can extract the speech feature information in the first speech through the Mel coefficient algorithm. Since the speech feature information of different dialect languages is different, the electronic device judges the dialect language of each speech frame according to the vector group corresponding to each speech frame, that is, the speech feature information, thereby improving the accuracy of speech-to-text conversion. [[ID=CO]]

[0198] Step S2: The electronic device generates candidate phrases for each speech frame based on the language confidence.

[0199] Exemplarily, the electronic device generates a dynamic candidate list for the candidate words corresponding to the speech frames with a confidence lower than the preset threshold of the candidate words, that is, the above-mentioned candidate phrases.

[0200] Exemplarily, the electronic device uses an ASR model to convert the first speech into corresponding text. Since each text has a fixed vector group. For example, Cantonese: "揾" (wen): In Cantonese, "揾" means "look for, seek". For example, "揾钱" means "make money". Gan language (Nanchang dialect): "稳" (wěn): In Gan language, "稳" means "certain" or "stable", not "look for". Therefore, the electronic device can obtain the vector group corresponding to each text, that is, the above-mentioned speech feature information.

[0201] Exemplarily, the electronic device can obtain a vector group of a speech frame in step S1, assuming that it is: [[6.4, -6.2, 9.9,..., 5.3], and then the electronic device calculates the corresponding vector values: [6.6, 4.1, 5.6,..., 4.1] and [6.4, -6.2, 9.9,..., 5.3] using the texts "weng" and "wen" obtained by the ASR. It can be understood that the selection here can be in the candidate word library in the above embodiment step. At this time, the electronic device calculates the similarity of the vector group of the speech frame and the vector group corresponding to "weng" and "wen" using the cosine theorem to obtain the similarity of the speech frame, and then calculates the confidence of the candidate word based on the language confidence of "weng" corresponding to the Cantonese language to obtain the confidence of 78%, and calculates the confidence of the candidate word based on the language confidence of "wen" corresponding to the Ganzhou dialect to obtain the confidence of 40%. Both are less than the preset confidence threshold of 95%, so the system determines that the two transcription texts are not reliable.

[0202] Exemplarily, the electronic device dynamically displays all candidate words when the confidence of the candidate words of a speech frame is not greater than the preset threshold.

[0203] Step S3, the user selects the candidate word through interaction.

[0204] Exemplarily, when the user encounters ambiguous recognized text, i.e., candidate words, the user can manually slide the candidate list, and the candidate box waits for 0.3 seconds. If the user does not slide again, the default selection is effective.

[0205] Step S4, the electronic device updates the language confidence.

[0206] Exemplarily, the language confidence corresponding to different dialects is dynamically updated according to the target word selected by the user in the history, and the ranking is further updated based on the updated language confidence.

[0207] It can be understood that based on the language confidence ranking, the subsequent transcription text is dynamically predicted to improve the transcription accuracy.

[0208] Exemplarily, as in step S3, assuming that before the text corresponding to the first speech is generated according to the input of the user, the electronic device calculates the language confidence corresponding to different dialects by the model and obtains a language confidence array, and the ranking is as follows:

[0209] [Mandarin, probability: 38%

[0210] Ganzhou dialect, probability: 23%

[0211] Wuhan dialect, probability: 22%

[0212] Cantonese, probability: 17%]

[0213] Then, if the user dynamically selects "揾" in step S3, corresponding to Cantonese, the electronic device will recalculate the language confidence corresponding to the dialect and reset the recognition sorting to:

[0214] [Cantonese, probability: 60%

[0215] Mandarin, probability: 18%

[0216] Ganzhou dialect, probability: 13%

[0217] Wuhan dialect, probability: 9%]

[0218] It can be understood that when the user inputs approximate speech subsequently, the electronic device will dynamically generate a candidate list according to the latest confidence sorting array, that is, the sorting of candidate words in the candidate phrase is based on the latest language confidence sorting.

[0219] Step S5: The electronic device performs incremental model optimization.

[0220] Exemplarily, the electronic device improves the dimension of the vector according to the target words selected by the user historically and the target words recognized by the electronic device to gradually optimize the model. Dynamically fuse the acoustic feature vector, that is, the 13-dimensional MFCC moving average and the language confidence-related features, to form a higher-dimensional hybrid vector, and gradually improve the sensitivity of the model to multiple dialects.

[0221] Taking the vector obtained from the example in step S1, "佢份工好辛苦,但揾钱唔少" as an example:

[0222] [[ID=2​​​​​​​​​​​​​​​​​​​​​​​​​​​​​The 13-dimensional MFCC coefficients obtained in step S1 are expanded into 52-dimensional MFCC coefficients, and are continuously trained to improve the fineness of the speech feature information corresponding to the speech frames. In this way, the higher-dimensional vector is used to train the speech recognition model, and after accumulating a certain amount of data, the speech recognition model can be upgraded, and the 52-dimensional vector is used to recognize the language, which can further improve the speech recognition accuracy.

[0234] In this way, the first embodiment of the present application effectively enhances the expression ability of multi-dialect speech features by dynamically fusing acoustic features and language confidence, optimizes the speech recognition effect in a complex multi-language environment, and provides a good user interaction experience. It has significant application value and practical promotion significance in the field of multi-dialect speech recognition.

[0235] Each of the above method embodiments, or each of the various possible implementations of the method embodiments, can be executed alone or in combination with any two or more of them. The specific implementation can be determined according to actual use requirements, and the embodiments of the present application do not limit this.

[0236] The text recognition method provided by the embodiments of the present application can be executed by an electronic device or a text recognition device. In the embodiments of the present application, the text recognition device is taken as an example to illustrate the text recognition device provided by the embodiments of the present application.

[0237] Figure 4 A possible structure schematic diagram of the text recognition device involved in the embodiments of the present application is shown. As shown in Figure 4 The text recognition device 700 can include an acquisition module 701 and a processing module 702.

[0238] The acquisition module 701 is configured to acquire a first speech, wherein the first speech includes N speech frames, and N is a positive integer.

[0239] The acquisition module 701 is further configured to acquire a candidate word group corresponding to each speech frame, wherein each candidate word group includes at least one candidate word, and each candidate word corresponds to a dialect language.

[0240] The processing module 702 is configured to determine a confidence of each candidate word based on a language confidence of a dialect language to which the candidate word belongs.

[0241] The processing module 702 is further configured to determine a target word corresponding to each speech frame based on the confidence of all candidate words in the candidate word group corresponding to the speech frame.

[0242] The processing module 702 is further configured to generate a text corresponding to the first speech based on each target word.

[0243] Optionally, in some embodiments of the present application, the processing module 702 is specifically configured to determine the confidence of the first candidate word in the first candidate word group corresponding to the first speech frame based on the language confidence of the dialect language to which the first candidate word belongs in the first candidate word group corresponding to the first speech frame, and the feature similarity between the speech feature information of the first candidate word and the speech feature information of the first speech frame; wherein the first speech frame is one of the N speech frames; and the first candidate word is one of the first candidate word group.

[0244] Optionally, in some embodiments of the present application, the processing module 702 is specifically configured to:

[0245] In a case where the confidence of at least one candidate word in the second candidate word group corresponding to the second speech frame is greater than a predetermined threshold, the candidate word with the highest confidence among the at least one candidate word is determined as the target word corresponding to the second speech frame;

[0246] In a case where the confidence of all candidate words in the second candidate word group corresponding to the second speech frame is not greater than the predetermined threshold, the candidate word selected by the user is determined as the target word corresponding to the second speech frame;

[0247] The second speech frame is one of the N speech frames.

[0248] Optionally, in some embodiments of the present application, in combination with Figure 4 As shown in Figure 5 , the apparatus further comprises a display module 703; the display module 703 is configured to, in a case where the confidence of all candidate words in the candidate word group corresponding to at least one speech frame of the N speech frames is not greater than the predetermined threshold, and the target words corresponding to other speech frames of the N speech frames except the at least one speech frame have been determined, display the target words corresponding to each speech frame of the other speech frames and each candidate word in the candidate word group corresponding to the at least one speech frame in the order of the N speech frames; and the processing module 702 is further configured to, in response to the input of the user on the displayed candidate words, take the candidate word selected by the user as the target word corresponding to the speech frame.

[0249] Optionally, in some embodiments of the present application, in combination with Figure 4 As shown in Figure 6 , the apparatus further comprises an updating module 704; the language confidence of the dialect language is calculated based on a language word library;

[0250] The processing module 702 is further configured to, after generating the text corresponding to the first speech based on each target word, add the target word corresponding to each speech frame to the language word library;

[0251] The updating module 704 is configured to update the language confidence of each dialect language based on the number of words corresponding to the dialect language in the language dictionary and the total number of words in the language dictionary.

[0252] The language confidence of each dialect language is a ratio of the number of words corresponding to the dialect language to the total number.

[0253] In the text recognition apparatus provided in the embodiments of the present application, the text recognition apparatus acquires a first speech, the first speech including N speech frames, N being a positive integer; acquires a candidate word group corresponding to each speech frame, each candidate word group including at least one candidate word, each candidate word corresponding to a dialect language; determines a confidence of each candidate word based on a language confidence of the dialect language to which each candidate word in each candidate word group belongs; determines a target word corresponding to each speech frame based on the confidence of all candidate words in the candidate word group corresponding to each speech frame; and generates a text corresponding to the first speech based on each target word. In this solution, the text recognition apparatus determines the target word that best fits the current dialect language environment by acquiring the candidate words corresponding to different dialects in the first speech, calculating the probability that each candidate word is the word of the corresponding speech frame in combination with the language confidence of the dialect language commonly used by the user, so as to improve the accuracy of the electronic device in recognizing the text corresponding to the speech in a multi-dialect language environment.

[0254] The text recognition apparatus in the embodiments of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The electronic device can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application are not limited in this regard.

[0255] The text recognition apparatus in the embodiments of the present application can be an apparatus with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, which are not limited in the embodiments of the present application.

[0256] The text recognition apparatus provided in the embodiments of the present application can implement all the processes of the text recognition method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0257] Optionally, as shown in Figure 7 The embodiments of the present application also provide an electronic device 800, which includes a processor 801 and a memory 802. The memory 802 stores programs or instructions that can run on the processor 801. When the programs or instructions are executed by the processor 801, all the steps of the above text recognition method embodiments are implemented, and the same technical effects are achieved. To avoid repetition, details are not described herein.

[0258] It should be noted that the electronic device in the embodiments of the present application includes the mobile electronic device and the non-mobile electronic device described above.

[0259] Figure 8 To implement the hardware structure of an electronic device in the embodiments of the present application.

[0260] The electronic device 100 includes, but is not limited to, a radio frequency unit 101, a network module 102, an audio output unit 103, an input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, and a processor 110, and the like.

[0261] Those skilled in the art can understand that the electronic device 100 can also include a power supply (such as a battery) for supplying power to each component. The power supply can be logically connected to the processor 110 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 8 The electronic device structure shown in the above figure does not constitute a limitation on the electronic device. The electronic device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components, which are not described herein.

[0262] The processor 110 is configured to acquire a first voice, the first voice including N voice frames, N being a positive integer.

[0263] The processor 110 is further configured to acquire a candidate word group corresponding to each voice frame, each candidate word group including at least one candidate word, and each candidate word corresponding to a dialect language.

[0264] The processor 110 is further configured to determine a confidence of each candidate word based on a language confidence of a dialect language to which the candidate word belongs.

[0265] The processor 110 is further configured to determine a target word corresponding to each speech frame based on the confidence of all candidate words in the candidate word group corresponding to the speech frame.

[0266] The processor 110 is further configured to generate a text corresponding to the first speech based on each target word.

[0267] Optionally, in some embodiments of the present application, the processor 110 is specifically configured to determine the confidence of the first candidate word in the first candidate word group corresponding to the first speech frame based on the language confidence of a dialect language to which the first candidate word belongs and a feature similarity between the speech feature information of the first candidate word and the speech feature information of the first speech frame; the first speech frame is one of the N speech frames; and the first candidate word is one of the first candidate word group.

[0268] Optionally, in some embodiments of the present application, the processor 110 is specifically configured to:

[0269] In a case where the confidence of at least one candidate word in the second candidate word group corresponding to the second speech frame is greater than a predetermined threshold, the candidate word with the highest confidence among the at least one candidate word is determined as the target word corresponding to the second speech frame.

[0270] In a case where the confidence of all candidate words in the second candidate word group corresponding to the second speech frame is not greater than the predetermined threshold, a candidate word selected by a user is determined as the target word corresponding to the second speech frame.

[0271] The second speech frame is one of the N speech frames.

[0272] Optionally, in some embodiments of the present application, the display unit 106 is configured to display, in a case where the confidence of all candidate words in a candidate word group corresponding to at least one speech frame of the N speech frames is not greater than the predetermined threshold and target words corresponding to other speech frames of the N speech frames except the at least one speech frame have been determined, the target words corresponding to the other speech frames in the order of the N speech frames and each candidate word in the candidate word group corresponding to the at least one speech frame; and the processor 110 is further configured to determine, in response to an input of a candidate word selected by a user, the candidate word as a target word corresponding to a speech frame.

[0273] Optionally, in some embodiments of the present application, the language confidence of the above dialect language is calculated based on the language library;

[0274] The processor 110 is further configured to add the target word corresponding to each of the speech frames to the language library after generating the text corresponding to the first speech based on each target word.

[0275] The processor 110 is further configured to update the language confidence of each of the dialect languages based on the number of words corresponding to each of the dialect languages in the language library and the total number of all words in the language library.

[0276] The language confidence of each of the dialect languages is the ratio of the number of words corresponding to the dialect language to the total number.

[0277] In the electronic device provided in the embodiments of the present application, the electronic device obtains a first speech, the first speech includes N speech frames, N is a positive integer; obtains a candidate word group corresponding to each speech frame, each candidate word group includes at least one candidate word, and each candidate word corresponds to a dialect language; determines the confidence of each candidate word based on the language confidence of the dialect language to which each candidate word in each candidate word group belongs; determines a target word corresponding to each speech frame based on the confidence of all candidate words in the candidate word group corresponding to each speech frame; and generates a text corresponding to the first speech based on each target word. In this scheme, the electronic device calculates the probability that each candidate word is the word of the corresponding speech frame by obtaining the candidate words corresponding to different dialects in the first speech and combining the language confidence of the dialect language commonly used by the user, to determine the target word that best fits the current dialect language environment, thereby improving the accuracy of the electronic device in recognizing the text corresponding to the speech in a multi-dialect language environment.

[0278] It should be understood that in the embodiments of the present application, the input unit 104 can include a graphics processing unit (GPU) 1041 and a microphone 1042. The graphics processing unit 1041 processes image data of a still picture or a video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 106 can include a display panel 1061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 can include a touch detection device and a touch controller. The other input devices 1072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), trackballs, mice, joysticks, etc., which will not be described here.

[0279] The memory 109 can be used to store software programs and various data. The memory 109 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, etc.), and the like. In addition, the memory 109 can include a volatile memory or a non-volatile memory, or the memory 109 can include both a volatile memory and a non-volatile memory. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 109 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.

[0280] The processor 110 can include one or more processing units; optionally, the processor 110 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 110.

[0281] The embodiments of the present application also provide a readable storage medium, and the readable storage medium stores programs or instructions, which are executed by a processor to realize various processes of the above-mentioned text recognition method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0282] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0283] The embodiment of the present application further provides a chip, which comprises a processor and a communication interface, the communication interface is coupled with the processor, the processor is used for running programs or instructions to realize the processes of the above text recognition method embodiments and achieve the same technical effects. To avoid repetition, details are not described here.

[0284] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0285] The embodiment of the present application provides a computer program product, which is stored in a storage medium, and is executed by at least one processor to realize the processes of the above text recognition method embodiments and achieve the same technical effects. To avoid repetition, details are not described here.

[0286] It should be noted that in this paper, the term "includes", "contains" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes elements inherent to such process, method, article or device. Without more limitations, the element defined by the statement "includes a" does not exclude the presence of another identical element in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the method and device in the embodiment of the present application is not limited to the order of the functions shown or discussed, but can also include the functions performed in a substantially simultaneous manner or in the opposite order according to the functions involved, for example, the described method can be performed in an order different from the described order, and various steps can also be added, omitted or combined. In addition, the features described with reference to some examples can be combined in other examples.

[0287] Through the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned example methods can be realized by means of software and a necessary general hardware platform, and of course, can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer software product in essence or in the form of a part that contributes to the prior art, which is stored in a storage medium (such as a ROM / RAM, a magnetic disc, an optical disc), and includes a plurality of instructions for causing a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0288] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the above-mentioned specific embodiments, and the above-mentioned specific embodiments are only illustrative and not restrictive. Those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.

Claims

1. A text recognition method, characterized in that, The method includes: Obtain the first audio, which includes N audio frames, where N is a positive integer; Obtain candidate word groups corresponding to each speech frame, each candidate word group containing at least one candidate word, and each candidate word corresponding to a dialect; The confidence level of each candidate word is determined based on the language confidence level of the dialect to which each candidate word belongs; Based on the confidence scores of all candidate words in the candidate word group corresponding to each speech frame, the target word corresponding to each speech frame is determined. Based on each target word, generate the text corresponding to the first speech.

2. The method according to claim 1, characterized in that, The step of determining the confidence level of each candidate word based on the language confidence level of the dialect to which each candidate word belongs includes: The confidence level of the first candidate word is determined based on the language confidence level of the dialect to which the first candidate word in the first candidate word group corresponding to the first speech frame belongs, and the feature similarity between the speech feature information of the first candidate word and the speech feature information of the first speech frame. Wherein, the first speech frame is one of the N speech frames; the first candidate word is one of the candidate words in the first candidate word group.

3. The method according to claim 1, characterized in that, The step of determining the target word corresponding to each speech frame based on the confidence scores of all candidate words in the candidate word group corresponding to each speech frame includes: If the confidence level of at least one candidate word in the second candidate word group corresponding to the second speech frame is greater than a predetermined threshold, the candidate word with the highest confidence level among the at least one candidate words shall be determined as the target word corresponding to the second speech frame. If the confidence level of all candidate words in the second candidate word group corresponding to the second speech frame is not greater than the predetermined threshold, the candidate word selected by the user is determined as the target word corresponding to the second speech frame. The second audio frame is one of the N audio frames.

4. The method according to claim 3, characterized in that, The method further includes: If the confidence level of all candidate words in the candidate word group corresponding to at least one of the N speech frames is not greater than the predetermined threshold, and the target words have been determined in the other speech frames besides the at least one speech frame, then the target words corresponding to each of the other speech frames and each candidate word in the candidate word group corresponding to the at least one speech frame are displayed according to the word order of the N speech frames. Receive user input for the displayed candidate words; In response to the input, the candidate word selected by the user is used as the target word of the corresponding speech frame.

5. The method according to claim 1, characterized in that, The language confidence score of the dialect is calculated based on the language lexicon; After generating the text corresponding to the first speech based on each target word, the method further includes: Add each of the target words to the language lexicon; Based on the number of words corresponding to each dialect in the language lexicon and the total number of all words in the language lexicon, update the language confidence level corresponding to each dialect. The language confidence level for each dialect is the ratio of the number of words corresponding to that dialect to the total number of words.

6. A text recognition device, characterized in that, The text recognition device includes: an acquisition module and a processing module; The acquisition module is used to acquire a first voice, which includes N voice frames, where N is a positive integer; The acquisition module is further configured to acquire candidate word groups corresponding to each speech frame, wherein each candidate word group contains at least one candidate word, and each candidate word corresponds to a dialect. The processing module is used to determine the confidence level of each candidate word based on the language confidence level of the dialect to which each candidate word belongs; The processing module is further configured to determine the target word corresponding to each speech frame based on the confidence level of all candidate words in the candidate word group corresponding to each speech frame; The processing module is further configured to generate text corresponding to the first speech based on each of the target words.

7. The apparatus according to claim 6, characterized in that, The processing module is specifically used to determine the confidence level of the first candidate word based on the language confidence level of the dialect to which the first candidate word in the first candidate word group corresponding to the first speech frame belongs, and the feature similarity between the speech feature information of the first candidate word and the speech feature information of the first speech frame. Wherein, the first speech frame is one of the N speech frames; the first candidate word is one of the candidate words in the first candidate word group.

8. The apparatus according to claim 6, characterized in that, The processing module is specifically used for: If the confidence level of at least one candidate word in the second candidate word group corresponding to the second speech frame is greater than a predetermined threshold, the candidate word with the highest confidence level among the at least one candidate words shall be determined as the target word corresponding to the second speech frame. If the confidence level of all candidate words in the second candidate word group corresponding to the second speech frame is not greater than the predetermined threshold, the candidate word selected by the user is determined as the target word corresponding to the second speech frame. The second audio frame is one of the N audio frames.

9. The apparatus according to claim 8, characterized in that, The device further includes: a display module; The display module is configured to, in accordance with the word order of the N speech frames, display the target word corresponding to each of the other speech frames and each candidate word in the candidate word group corresponding to the at least one speech frame, when the confidence level of all candidate words in the candidate word group corresponding to at least one speech frame in the N speech frames is not greater than the predetermined threshold, and the target word has been determined for the other speech frames in the N speech frames except for the at least one speech frame. The processing module is also configured to respond to the user's input of the displayed candidate words and use the candidate word selected by the user as the target word of the corresponding speech frame.

10. The apparatus according to claim 6, characterized in that, The device further includes: an update module; the language confidence of the dialect is calculated based on a language lexicon; The processing module is further configured to generate the text corresponding to the first speech based on each target word, and then add the target word corresponding to each speech frame to the language dictionary. The update module is used to update the language confidence level corresponding to each dialect based on the number of words corresponding to each dialect in the language lexicon and the total number of all words in the language lexicon; The language confidence level for each dialect is the ratio of the number of words corresponding to that dialect to the total number of words.