Text processing method and device, electronic equipment, storage medium and program product
By generating target prompts and combining them with phoneme data and auxiliary information, and using a large language model to correct speech recognition text, the problem of insufficient accuracy in speech recognition error correction in existing technologies is solved, and the accuracy of text processing is improved.
Patent Information
- Application Number
- CN202410703130.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2025-12-02
AI Technical Summary
Existing AI technologies for speech recognition and text correction have room for improvement and are difficult to effectively enhance the accuracy of error correction.
By determining the initial speech recognition text and phoneme data of the target speech, target prompt information is generated. The initial speech recognition text is corrected based on the phoneme data using a large language model. The corrected target recognition text is then generated by combining the phoneme data and auxiliary information.
It improves the accuracy of text processing, and the accuracy of using phoneme data is higher than that of the initial speech recognition text. It provides more effective information to correct speech recognition text and enhances the error correction capability of large language models.
Smart Images

Figure CN121053989A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of artificial intelligence and large models, and specifically to a text processing method, apparatus, electronic device, storage medium, and program product. Background Technology
[0002] Automatic speech recognition (ASR) is a technology that converts human speech into corresponding speech-recognized text. Since there may be errors in the speech-recognized text, it is usually corrected.
[0003] In this field, artificial intelligence (AI) technology is typically used to correct errors in speech recognition text. However, there is still significant room for improvement in AI technology for error correction in speech recognition text. Therefore, there is an urgent need in this field for a text processing method that can better utilize AI technology for error correction. Summary of the Invention
[0004] This application provides a text processing method, apparatus, electronic device, storage medium, and program product, which helps large language models obtain correct text during processing, thereby improving the accuracy of text processing. The technical solution is as follows: On the one hand, embodiments of this application provide a text processing method, the method comprising: Determine the initial speech recognition text and phoneme data for the target speech; Based on the initial speech recognition text and phoneme data of the target speech, target prompt information is generated. The target prompt information is used to prompt the large language model to correct the initial speech recognition text based on the phoneme data. Based on the target prompt information, the target recognition text after correcting the initial speech recognition text is obtained through the large language model.
[0005] On the other hand, embodiments of this application provide a text processing apparatus, the apparatus comprising: The first determining module is used to determine the initial speech recognition text and phoneme data of the target speech; The generation module is used to generate target prompt information based on the initial speech recognition text and phoneme data of the target speech, and the target prompt information is used to prompt the large language model to correct the initial speech recognition text based on the phoneme data; The correction module is used to obtain the corrected target recognition text based on the target prompt information and through the large language model.
[0006] In one possible implementation, the device further includes: The second determining module is used to determine the auxiliary information of the target speech; The generation module is used for: Based on the initial speech recognition text, phoneme data, and auxiliary information of the target speech, the target prompt information is generated. The target prompt information is used to prompt the correction of the initial speech recognition text based on the auxiliary information and phoneme data.
[0007] In one possible implementation, the second determining module, when determining the auxiliary information of the target speech, is specifically used for any of the following: The scene description information of the target scene to which the target speech belongs is used as the auxiliary information; The text information of the context of the target speech is used as the auxiliary information.
[0008] In one possible implementation, the generation module is configured to: Based on the initial speech recognition text, phoneme data, and phoneme reference information, target prompt information is generated. The target prompt information is used to prompt the initial speech recognition text to be corrected based on the phoneme reference information and phoneme data.
[0009] In one possible implementation, the first determining module, when determining phoneme data, is used to: Phoneme recognition is performed on each syllable in the target speech to obtain initial phoneme data; Based on the target type to which the target speech belongs in at least one pronunciation type, the initial phoneme data of the target speech is corrected to obtain the phoneme data.
[0010] In one possible implementation, the device further includes: The acquisition module is used to acquire phoneme reference information of the target speech; The phoneme reference information includes at least one of the following: The target speech belongs to the target type in at least one pronunciation type; The target type corresponds to the target correction relationship, which includes the correspondence between at least one phoneme pronounced according to the target type and the standard pronunciation phoneme; The phoneme correction relationship for each of the at least one pronunciation type.
[0011] In one possible implementation, the first determining module, when determining the initial speech recognition text, is used to: Based on the target speech, multiple candidate texts of the target speech and the confidence level of each candidate text are obtained by a speech recognition model; The average similarity of each candidate text is determined based on the similarity between each candidate text and each other candidate text. Based on the confidence level and average similarity of each candidate text, at least one initial speech recognition text is obtained from the plurality of candidate texts; Specifically, the first determining module, when determining phoneme data, is used to perform any of the following steps: Based on the target speech, phoneme recognition is performed on each syllable in the target speech using a phoneme recognition model to obtain at least one phoneme data and the confidence level of each phoneme data. Each initial speech recognition text is processed by a phoneme conversion tool to obtain at least one phoneme data of the target speech.
[0012] In one possible implementation, if the at least one phoneme data and the confidence level of each phoneme data are obtained through a phoneme recognition model, then the generation module is used to: Based on each initial speech recognition text and its confidence level, each phoneme data and its confidence level, the target prompt information is generated. The target prompt information is used to prompt correction of each initial speech recognition text based on each initial speech recognition text and its confidence level, each phoneme data and its confidence level.
[0013] In one possible implementation, the apparatus further includes a fine-tuning module for performing at least one fine-tuning process on the large language model: The parameter matrix of the large language model to be trained is decomposed into two low-rank matrices. Based on the initial recognition text and sample phoneme data of each speech sample in the sample set, sample prompt information is generated for each speech sample; Based on the sample prompt information of each speech sample, the predicted text after correcting the initial recognized text of each speech sample is obtained through the parameter matrix of the large language model to be trained. The training loss for each speech sample is calculated based on the similarity between the predicted text and the real speech text of each speech sample. Based on the training loss corresponding to each speech sample, the two low-rank matrices are adjusted, and the parameter matrix of the large language model to be trained is updated based on the adjusted two low-rank matrices.
[0014] On the other hand, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the above-described text processing method.
[0015] On the other hand, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described text processing method.
[0016] On the other hand, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the above-described text processing method.
[0017] The beneficial effects of the technical solutions provided in this application are: The text processing method provided in this application determines the initial speech recognition text and phoneme data of the target speech, generates target prompt information based on the initial speech recognition text and phoneme data, and obtains the target recognition text through a large language model based on the target prompt information. This allows the large language model to correct the initial speech recognition text based on the phoneme data. Since the accuracy of the phoneme data is much higher than that of the initial speech recognition text, combining the phoneme data can provide more and more effective information for reference in the processing of the large language model, which helps the large language model obtain the correct text in the processing process, thereby improving the accuracy of text processing. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.
[0019] Figure 1 This is a schematic diagram of the structure of a text processing system provided in an embodiment of this application; Figure 2 A flowchart illustrating a text processing method provided in an embodiment of this application; Figure 3 A flowchart illustrating a text processing method provided in an embodiment of this application; Figure 4 A flowchart illustrating a text processing method provided in an embodiment of this application; Figure 5 A flowchart illustrating a text processing method provided in an embodiment of this application; Figure 6 A schematic diagram illustrating the input and output of a large language model provided in an embodiment of this application; Figure 7 A schematic diagram illustrating the fine-tuning of a large language model provided in this application embodiment; Figure 8 A flowchart illustrating a text processing procedure provided in an embodiment of this application; Figure 9 This is a schematic diagram of the structure of a text processing device provided in an embodiment of this application; Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0020] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.
[0021] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. The terms “comprising” and “including” as used in the embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, or operation, but do not exclude implementation as other features, information, data, steps, or operations supported by this art.
[0022] It is understood that, in the specific embodiments of this application, any user-related data, such as target speech, initial speech recognition text, target text, and phoneme data, is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. In other words, if any object-related data is involved in the embodiments of this application, this data must be obtained with the user's authorization and consent, and in accordance with the relevant laws, regulations, and standards of the country and region.
[0023] The text processing method provided in this application involves artificial intelligence technology, specifically natural language processing, speech processing, and machine learning. For example, machine learning technology can be used to train large language models; for example, in the fine-tuning stage, model compression and quantization techniques can be used, specifically the low-rank decomposition technique, so that only a small number of samples are used for training during fine-tuning to achieve good training results.
[0024] As can be understood, Artificial Intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI essentially studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0025] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0026] Understandably, Natural Language Processing (NLP) is a crucial area within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP deals with natural language—the language people use in daily life—and is closely related to linguistics; it also involves computer science and mathematics. Pre-trained models, a key technique for model training in artificial intelligence, evolved from large language models in NLP. After fine-tuning, large language models can be widely applied to downstream tasks. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.
[0027] As is understandable, Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory, among others. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and pre-trained learning. Pre-trained models represent the latest development in deep learning, integrating all of these techniques.
[0028] As can be understood, model compression and quantization refer to techniques used to reduce model size and accelerate model inference, thereby lowering the storage and computational costs of the model. Model compression typically includes pruning, low-rank decomposition, and knowledge distillation, while model quantization involves converting floating-point parameters in the model into fixed-point or integer parameters, thereby reducing model size and accelerating model inference.
[0029] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, digital twins, virtual humans, robots, AI-generated content (AIGC), conversational interaction, smart healthcare, smart customer service, and game AI. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.
[0030] Figure 1 This application provides a schematic diagram of the structure of a text processing system for a text processing method. (See attached diagram.) Figure 1 As shown, the text processing system includes a server 101 and a terminal 102.
[0031] This text processing system provides text processing services. These services can be used to correct speech recognition text obtained based on ASR technology, resulting in more accurate text recognition. In one possible scenario, for online meetings or lectures, the system can correct the speech recognition text of the speaker or lecturer. In another possible scenario, for online classrooms, the system can correct the speech recognition text of the teacher.
[0032] In one possible implementation scenario, terminal 102 interacts with server 101 to perform text processing. For example, a user inputs target speech to be recognized through terminal 102, and terminal 102 sends the target speech to server 101.
[0033] For the target speech to be recognized, server 101 determines the initial speech recognition text and phoneme information of the target speech, and generates target prompt information based on the initial speech recognition text and phoneme information. This target prompt information is used to prompt the large language model to correct the initial speech recognition text based on the phoneme data. Based on the target prompt information, server 101 obtains the corrected target recognition text of the initial speech recognition text through the large language model. Specifically, server 101 can call the large language model for processing based on the target prompt information to obtain the target recognition text output by the large language model.
[0034] The server 101 sends the target identification text to the terminal 102. The terminal 102 communicates with the server 101 via a network.
[0035] Among them, the large language model is a deep learning model with a huge scale and many parameters. It can learn language structure, grammatical knowledge and semantics from large-scale datasets, thereby correcting the initial speech recognition text based on the target prompt information input.
[0036] Among them, server 101 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, and big data and artificial intelligence platforms.
[0037] The terminal 102 can be a smartphone, tablet computer, laptop computer, digital radio receiver, desktop computer, in-vehicle terminal (such as in-vehicle navigation terminal, in-vehicle computer, etc.), smart speaker, smartwatch, etc., but is not limited to these.
[0038] Figure 2 This is a flowchart illustrating a text processing method provided in an embodiment of this application. The method can be executed by an electronic device. Figure 2 As shown, the method includes the following steps.
[0039] Step 201: The electronic device determines the initial speech recognition text and phoneme data of the target speech.
[0040] The initial speech recognition text is the speech recognition result obtained using ASR technology.
[0041] In this embodiment, the initial speech recognition text obtained by ASR technology can be corrected by combining phoneme data. The phoneme data includes the phonemes corresponding to each syllable of the target speech. For example, a phoneme can be the smallest unit that constitutes a syllable; it is understood that phoneme data can accurately capture each syllable in the target speech.
[0042] Exemplarily, the target voice can be a voice of any type of language, and phonemes can represent each syllable included in the voice of any type of language. For example, for Chinese, the phoneme data can be the pinyin corresponding to each Chinese character. For example, a phoneme can be an initial or a final; the phoneme data corresponding to a Chinese character can include at least one of the corresponding initial or final. For example, the phoneme data corresponding to "啊" is "a", including 1 phoneme; the phoneme data corresponding to the character "拼" in "拼音" is "pin", including 2 phonemes, namely the initial "p" and the final "in". For English, the phoneme can be the phonetic symbol. This application does not limit the language type of the voice, the phonemes corresponding to different languages, etc.
[0043] Exemplarily, the initial speech recognition text can include one or more texts, which can be obtained by using a speech recognition model for recognition. Exemplarily, in step 201, for the determination method of the initial speech recognition text, it includes the following steps 2011: Step 2011: The electronic device recognizes at least one initial speech recognition text of the target voice based on the target voice through a speech recognition model.
[0044] Exemplarily, at least one initial speech recognition text and the confidence of each initial speech recognition text can also be obtained by using a speech recognition model. The confidence represents the possibility that the initial speech recognition text is the true text of the target voice. The higher the confidence, the higher the possibility that the initial speech recognition text is the true text.
[0045] For example, for a segment of voice, a trained ASR model can be used to recognize multiple possible texts of the segment of voice and the score of each text to obtain the initial speech recognition text; among them, there are local differences between the multiple possible texts, and the score of the text can be the confidence, and the higher the score, the higher the possibility that the text is the true text of the segment of voice.
[0046] In one possible example, the ASR model can be used to output the text with the highest score as the initial speech recognition text, that is, the final recognition result of the ASR model, and the initial speech recognition text can be represented as 1-best.
[0047] In another possible example, the ASR model can be used to output multiple possible initial speech recognition texts; for example, the ASR model outputs the top n texts with the highest scores, and the n texts can be arranged in descending order of scores, and can be correspondingly represented as n-best. For example, the text with the highest score determined by the ASR model may not be the recognition result closest to the true text, and other candidate recognition results can be referred to.
[0048] For example, if the actual text of a speech is "The weather is nice today", the ASR model will identify and output the top 5 texts with the highest scores, i.e., the 5-best: 1. The weather is very nice today. 2. The weather is very early today. 3. The weather is very nice today. 4. It smells very sweet today. 5. The weather is very nice in Jingtian.
[0049] Among the five texts, each text had slight differences, and the ASR model scored the highest in recognizing the first text.
[0050] In one possible approach, the similarity between multiple candidate texts can be used to further filter the candidate texts identified by the speech recognition model to obtain the initial speech recognition text. Accordingly, this initial speech recognition text is obtained by performing the following steps: Based on the target speech, multiple candidate texts and the confidence level of each candidate text are obtained by using a speech recognition model. The average similarity of each candidate text is determined based on the similarity between each candidate text and each other candidate text. Based on the confidence level and average similarity of each candidate text, at least one initial speech recognition text is obtained from the multiple candidate texts.
[0051] For example, the similarity between candidate texts can be combined to select several relatively similar texts from multiple candidate texts as the initial speech recognition texts. For instance, for each candidate text, the similarity between the candidate text and each other candidate text can be calculated to obtain at least one similarity for each candidate text, and the average of the at least one similarity can be calculated to obtain the average similarity of the candidate text.
[0052] For example, the confidence score and average similarity score of each candidate text can be combined to filter out the most likely initial speech recognition texts. For instance, for each candidate text, the confidence score and average similarity score can be fused together to obtain a fusion score; for example, the confidence score and average similarity score can be weighted according to their respective weight coefficients to obtain a fusion score. Further, at least one initial speech recognition text with a score greater than the target score threshold is selected from all candidate texts.
[0053] Based on this, the similarity between multiple candidate texts can be combined to eliminate texts that differ significantly from other texts, thereby further improving the accuracy of the initial speech recognition text and providing more accurate input for the processing of large language models, which helps to improve the accuracy of large language models in processing text.
[0054] In one possible approach, phoneme recognition can be performed on the target speech to obtain phoneme data. For example, the electronic device can identify the phoneme data of the target speech using a phoneme recognition model. In another possible approach, the phoneme data can be directly converted from the initial recognized speech text. For example, the electronic device can use a phoneme conversion tool to perform phoneme conversion processing based on the initial recognized speech text to obtain the phoneme data of the target speech.
[0055] For example, in step 201, the method for determining the phoneme data includes the following step 2012: the phoneme data is obtained by performing either step 2012-1 or step 2012-2: Step 2012-1: Based on the target speech, the electronic device performs phoneme recognition on each syllable in the target speech using a phoneme recognition model to obtain at least one phoneme data and the confidence level of each phoneme data. The phoneme recognition model can be a pre-trained model, or it can be a recognition model based on ASR technology used to identify phonemes and output phoneme data.
[0056] For example, the electronic device can input the target speech into the phoneme recognition model and obtain the phoneme data output by the phoneme recognition model. The phoneme recognition model can output multiple phoneme data points of the target speech and the confidence score of each phoneme data point. Each phoneme data point includes the phonemes corresponding to each syllable in the target speech, and each phoneme data point represents a possible pronunciation of the target speech predicted by the phoneme recognition model; the confidence score indicates the probability that the corresponding phoneme data point is the true speech data of the target speech. The higher the confidence score, the closer the corresponding phoneme data point is to the true speech data.
[0057] Step 2012-2: The electronic device performs phoneme conversion processing on each initial speech recognition text using a phoneme conversion tool to obtain at least one phoneme data of the target speech.
[0058] For example, for Chinese, a pinyin transcription tool can be used to transcribe each Chinese character in each initial speech recognition text based on pre-configured initials and finals.
[0059] In one possible approach, the phoneme data may also include phoneme tagging information, such as, but not limited to, the tone corresponding to the pinyin, the separator of the pinyin, etc. That is, the phoneme data includes the phonemes corresponding to each syllable in the target speech, and includes at least one of the tone or apostrophe of each phoneme. For example, taking Chinese characters as an example, the tone of pinyin may include: first tone (yinping), second tone (yangping), third tone (shangsheng), fourth tone (qusheng), etc.; for example, the tone of "pinyin" is all first tone, and its phoneme data can be represented as "pīn yīn", or it can also be represented as "pin1 yin1".
[0060] For example, you can use a pinyin transcription tool and use the following initials and finals for pinyin transcription: Initial consonants: b, c, ch, d, f, g, h, j, k, l, m, n, p, q, r, s, sh, t, w, x, y, z, zh Finals: a, ai, an, ang, ao, e, ei, en, eng, er, i, ia, ian, iang, iao, ie, in, ing, iong, iu, o, ong, ou, u, ua, uai, uan, uang, ue, ui, un, uo, ü, üe For the 5-best text in step 2011, the 5-best pinyin is transcribed as follows: "1. j in t ian t ian qih en h ao 2. j in t ian t ian qih en z ao 3. j in t ian t ian qih en h ao 4. j in t ian t ian qih en h ao 5. j ing t ian t ian qih en h ao”.
[0061] Comparing 5-best text and 5-best pinyin, it's important to note that while the individual possible texts in the 5-best text set may differ, their pinyin may be the same. For example, the first, third, and fourth texts might be different, but their corresponding pinyin entries might be identical. In this case, homophones may exist between different texts, and the accuracy of these homophones' pinyin is higher than the overall text accuracy. For instance, the correct text might be one of these homophones, or another Chinese character with the same pinyin. Using the pinyin of these homophones helps the large language model combine more accurate pronunciations to find the correct text.
[0062] In a possible example, if the phoneme data also includes tones, for the 5-best text in step 2011, the transcribed 5-best pinyin is as follows: "1. j in1 t ian1 t ian1 q i4 h en3 h ao3 2. j in1 t ian1 t ian1 q i4 h en3 z ao3 3. j in1 t ian1 t ian1 q i2 h en3 h ao3 4. j in1 t ian1 t ian2 q i4 h en3 h ao3 5. j ing3 t ian1 t ian1 q i4 h en3 h ao3".
[0063] Among them, in each of the above pinyins with tones, the phonemes and tones can be used to more accurately represent each syllable in the speech; for texts with the same phonemes, these texts can be further distinguished from the tones. For example, the 1st, 3rd, and 4th pinyins without tones are the same, but after adding tones, in the 1st, 3rd, and 4th pinyins, texts with the same phonemes but different tones can be distinguished by the tones. This enables the large language model to obtain more-dimensional features such as phonemes and tones of the text, increasing the amount of information in the feature expression and improving the accuracy of the large language model in obtaining the correct text.
[0064] It should be noted that the accuracy of phoneme data recognition is much higher than that of words. Taking a Chinese character as an example, if there is one incorrect recognition in the initial or final consonant, the whole character is incorrect; but at the phoneme level, half of it is still correctly recognized. For example, the last character "zao" in the 2nd text is incorrectly recognized (it should be "hao"); in the 2nd pinyin corresponding to the 2nd text, in the corresponding "zao", only the "z" is incorrectly recognized and the "ao" is still correct.
[0065] Therefore, by combining phoneme data to correct the initial speech recognition text, more and more effective information can be provided for the large language model to refer to in the processing process, which helps the large language model find the correct text in the processing process and thus improves the accuracy of text processing.
[0066] In a possible example, the phoneme data output by the phoneme recognition model is independent of the initial speech recognition text output by the speech recognition model. Using the phoneme data output by the phoneme recognition model can accurately capture each syllable in the target speech, that is, capture each pronunciation of the speaker. Based on this, the phoneme data output by the phoneme recognition model includes complementary information of the initial speech recognition text.
[0067] Step 202: The electronic device generates target prompt information based on the initial speech recognition text and phoneme data of the target speech. The target prompt information is used to prompt the large language model to correct the initial speech recognition text based on the phoneme data.
[0068] In this step, the electronic device can generate the target prompt information based on the initial speech recognition text and phoneme data, according to a pre-configured prompt template. For example, at least one initial speech recognition text and phoneme data can be added to the prompt template to obtain the target prompt information.
[0069] In one possible embodiment, such as Figure 3 As shown, the text processing method of this application embodiment may further include the following step A1: Step A1: The electronic device determines the auxiliary information for the target speech; Correspondingly, such as Figure 3 As shown, step 202 may include the following step 202a: Step 202a: The electronic device generates target prompt information based on the initial speech recognition text, phoneme data, and auxiliary information of the target speech. The target prompt information is used to prompt the correction of the initial speech recognition text based on the auxiliary information and phoneme data.
[0070] For example, the auxiliary information is used to assist the large language model in correcting the initial speech recognition text of the target speech. The electronic device can add the auxiliary information, the initial speech recognition text, and the phoneme data to the corresponding position in a pre-configured prompt template to obtain the target prompt information.
[0071] In one possible implementation, the context in which the target speech originates can be considered to assist the large language model in correction; alternatively, the context of the target speech can also be considered for correction. Accordingly, step A1 can be implemented by including any one of the following steps A11 and A12: Step A11: The electronic device uses the scene description information of the target scene to which the target speech belongs as the auxiliary information; Step A12: The electronic device uses the text information of the context speech of the target speech as the auxiliary information.
[0072] For example, in step A11, the target scenario refers to the scenario from which the target speech originates. For instance, the target speech could be a sentence or several sentences spoken by a speaker in a target meeting; or, the target speech could be the voice of a speaker in a target lecture; or, the target speech could be the voice of a teacher in a target online classroom, etc.
[0073] For example, the scene description information may include information about at least one scene element of the target scene. For example, the at least one scene element may include, but is not limited to, at least one of the following: the theme, domain, keywords, summary, proper nouns, to-do items, topics, neologisms, entity words, and names involved in the scene.
[0074] For example, the conference theme could be agricultural development, and the conference keywords could be: sustainable development, organic crops, and ecological niche.
[0075] For example, in step A12, the contextual speech can be speech associated with the target speech in the target scene. The text information of the contextual speech can be the corrected text of the contextual speech; for example, the corrected text of the contextual speech can be obtained through a large language model before correcting the initial speech recognition text of the target speech. For example, the contextual speech can be one or more speech items spoken by a speaker in the target meeting before the target speech, or the contextual speech can also include speech items spoken by the speaker that contain meeting keywords, meeting topics, proper nouns, etc.
[0076] In one possible example, the target prompt message can be referred to as a prompt word. This prompt template can include several prompt words; for example, in an online meeting scenario, the prompt template might look like this: "This meeting is about [ ], the terms involved include [ ], its theme is [ ], the n-best text output by the speech recognition system is [ ], the n-best pinyin is [ ], please output the best recognition result based on the above information."
[0077] Based on this, the large language model can correct the n-best text by using the scenario description information such as the theme, keywords, terms, and domain of the meeting in the target prompt information, as well as the n-best pinyin.
[0078] For example, based on this prompt template, the conference's field is added as agricultural development, and corresponding terms or keywords are added such as sustainable development, organic crops, and ecological niche; and the 5-best text and 5-best pinyin are added as shown in the example above; the target prompt information input into the large language model is as follows: "This is a conference about agricultural development, and the following keywords may be encountered: sustainable development, organic crops, and ecological niche. Based on the provided N-best text candidates and N-best phonetic transcription, the text is corrected and the results are output directly."
[0079] N-best candidate text (which can also be the text with the highest score): 1. The weather is very nice today. 2. The weather is very early today. 3. The weather is very nice today. 4. It smells very sweet today. 5. The weather is very nice. N-best Pinyin transcription: 1. j in t ian t ian qih en h ao 2. j in t ian t ian qih en z ao 3. j in t ian t ian qih en h ao 4. j in t ian t ian qih en h ao 5. j ing t ian t ian qih en h ao” The large language model can output the corrected result: "The weather is nice today".
[0080] In one possible example, the contextual speech can be directly used as auxiliary information. For instance, the time-frequency features of the contextual speech can be extracted using Fourier transform, and these features can be used as auxiliary information corresponding to the target speech.
[0081] In yet another possible embodiment, such as Figure 4 As shown, the text processing method of this application embodiment may further include the following step B1: Step B1: The electronic device acquires the phoneme reference information of the target speech; The phoneme reference information includes at least one of the following: The target speech belongs to at least one pronunciation type; The target correction relationship corresponding to the target type includes the correspondence between at least one phoneme pronounced according to the target type and the standard pronunciation phoneme; The phoneme correction relationship for each of the at least one pronunciation type.
[0082] For example, for the same type of language, the pronunciation of the same phoneme may differ between different regions. Pronunciation type can be a classification based on the different pronunciations of a phoneme; a language may correspond to one or more dialects.
[0083] In one example, different regions may have their own regional accents; for instance, for Chinese, different regions may have multiple dialects, and the pronunciation type can be a dialect type, such as Hubei dialect, Northeastern dialect, etc. For example, for English, the pronunciation type can correspond to American English, British English, etc.
[0084] In this context, standard pronunciation phonemes refer to standard pronunciations that do not have accent variations. For example, the pronunciations of the initial consonants s and sh may be confused in dialects, with sh often being pronounced as s.
[0085] Each pronunciation type can correspond to a phoneme correction relationship, which includes the correspondence between at least one phoneme pronounced according to that pronunciation type and the standard pronunciation phoneme. For example, the phoneme correction relationship in Hubei dialect may include s→sh, n→ l That is, the standard phoneme corresponding to the 's' pronounced in Hubei dialect might be 'sh', and the standard phoneme corresponding to the 'n' pronounced in Hubei dialect might be... l。 It should be noted that a phoneme correction relationship for a dialect type may include multiple correspondences between phonemes and standard pronunciation phonemes. This phoneme correction relationship can be configured as needed, and this application does not limit it.
[0086] For example, if the phoneme reference information includes a target type and a corresponding correction relationship, the step of obtaining the phoneme reference information may include: the electronic device determining the target type to which the target speech belongs in at least one pronunciation type; the electronic device obtaining the target correction relationship corresponding to the target type from the phoneme correction relationships corresponding to at least one pronunciation type; the phoneme correction relationship includes the correspondence between at least one phoneme pronounced according to the corresponding pronunciation type and the standard pronunciation phoneme.
[0087] The electronic device can identify the pronunciation type of the target speech through the pronunciation recognition model. That is, the steps of the electronic device to determine the target type of the target speech may include: the electronic device identifies the pronunciation type of the target speech based on the target speech through a pre-trained pronunciation type recognition model to obtain the target type to which the target speech belongs.
[0088] Accordingly, if the text processing method in this application embodiment further includes step B1, such as Figure 4 As shown, step 202 can be implemented by including the following step 202b: Step 202b: The electronic device generates target prompt information based on the initial speech recognition text, phoneme data, and phoneme reference information. The target prompt information is used to prompt the correction of the initial speech recognition text based on the phoneme reference information and phoneme data.
[0089] In this step, the electronic device can add the initial speech recognition text, phoneme data, and phoneme reference information to the corresponding position in the pre-configured prompt template to obtain the target prompt information.
[0090] In one possible example, the target prompt information may specifically include the initial speech recognition text, phoneme data, and phoneme reference information such as target type and target correction relationship. For example, taking the case where the phoneme reference information includes the target type, the target prompt information input into the large language model would be as follows: "The speech recognition system identifies the dialect, which is Hubei dialect. This dialect does not distinguish between retroflex and alveolar consonants, front and back nasal consonants, and its tones are different from Mandarin. Based on the provided N-best text output and N-best pinyin transcription, the system corrects the text and directly outputs the results."
[0091] N-best text: 1. A shocking event, carried horizontally. 2. It's great to be able to withstand such a shocking blow. 3. Jingtian seems to be balanced or 4. The marks look like they're there today. 5. It's good to carry such a huge burden. N-best Pinyin transcription: 1. j ing t ian k ang qil ai h eng h ao 2. j ing t ian k ang qil ai h en h ao 3. j ing t ian k an qil ai h eng h uo 4. j in t ian k an qil ai h eng h ao 5. j ing t ian k ang qil ai h en h ao” The large language model can output the corrected result: "Today looks good".
[0092] In one possible implementation, the phoneme data of the target speech can be obtained by correcting it according to the target correction relation of the target speech type; correspondingly, the method for determining the phoneme data includes: Phoneme identification is performed on each syllable in the target speech to obtain initial phoneme data; Based on the target type to which the target speech belongs in at least one pronunciation type, the initial phoneme data of the target speech is corrected to obtain the phoneme data.
[0093] For example, an electronic device can correct the initial phoneme data based on the target correction relationship corresponding to the target type to which the target speech belongs. For instance, if the target speech belongs to the Hubei dialect, the initial consonant 's' in the initial phoneme data can be corrected to the corresponding standard pronunciation phoneme 'sh', and the initial consonant 'n' in the initial phoneme data can be corrected to the corresponding standard pronunciation phoneme. l。
[0094] For example, the electronic device can perform phoneme recognition on each syllable in the target speech using a phoneme recognition model, and use the recognition result of the phoneme recognition model as the initial phoneme data; then, it can use the target correction relationship to correct each phoneme in the initial phoneme data to obtain the phoneme data.
[0095] It should be noted that for multiple accented pronunciations of the same phoneme, the recognized phoneme may differ depending on the pronunciation type of the speech. For example, in a dialect, "sh" might be pronounced as "s." If the target speech belongs to that dialect type, the recognized phoneme will differ from the phoneme in the standard pronunciation. In this embodiment, by adding phoneme reference information to the prompts, the large language model can combine phoneme correction relationships corresponding to different pronunciation types, the target speech type, or at least one of these factors for processing. This helps the large language model to correct different types of pronunciations by incorporating phoneme reference information, allowing it to utilize the connections between different pronunciation types to obtain more accurate phoneme data, thereby improving the accuracy of text processing.
[0096] In one possible implementation, in step 201, if the confidence scores of the initially identified text and phoneme data are also obtained, the text confidence scores and phoneme data confidence scores can be used to construct target prompt information. Correspondingly, if the at least one phoneme data and the confidence scores of each phoneme data are obtained through a phoneme recognition model, such as... Figure 5 As shown, step 202 can be implemented by including the following step 202c: Step 202c: The electronic device generates the target prompt information based on each initial speech recognition text and its confidence level, each phoneme data and its confidence level. The target prompt information is used to prompt the correction of each initial speech recognition text based on each initial speech recognition text and its confidence level, each phoneme data and its confidence level.
[0097] For example, the 5-best text output by the speech recognition model and its score, along with the 2-best pinyin output by the phoneme recognition model and its score, can be added to a pre-configured prompt template to obtain the target prompt information. For example, the target prompt information input into the large language model might be as follows: "This is a conference on agricultural development, where the following keywords may be encountered: sustainable development, organic crops, and ecological niche. Based on the provided 5-best candidate text and scores, as well as the 2-best phonetic transcription and scores, the text is corrected and the results are directly output."
[0098] The N-best candidate texts, and the score for each text: 1. The weather is great today, 0.9 2. The weather is very early today, 0.8 3. The weather is very good today, 0.7 4. The sweetness is very good today, 0.7 5. The weather is very good for Sedum, 0.7 N-best Pinyin transcription, and a score for each Pinyin entry: 1. j in t ian t ian qih en h ao , 0.9 2. j ing t ian t ian qih en h ao , 0.7” It should be noted that, in this embodiment of the application, the target prompt information can be constructed based on any one of the above steps 202a, 202b, and 202c.
[0099] Alternatively, at least two of steps 202a, 202b, and 202c above can be combined to construct richer target prompt information. In one possible implementation, the electronic device can generate the target prompt information based on the initial speech recognition text, phoneme data, auxiliary information, phoneme reference information, and the confidence levels of each initial speech recognition text and each phoneme data, according to a pre-configured prompt template. That is, the prompt information is constructed by combining the implementation methods of steps 202a, 202b, and 202c above.
[0100] Other possible implementations may also include combining steps 202a and 202b to construct the prompt message; or combining steps 202b and 202c to construct the prompt message; or combining steps 202a and 202c to construct the prompt message. Regardless of which two or three of steps 202a, 202b, or 202c are combined, the implementation method of combining the steps is the same as the implementation method of the corresponding individual steps described above, and will not be elaborated further here.
[0101] It should be noted that since the initial speech recognition text is obtained by the ASR model based on the ASR dictionary, each text unit (such as each Chinese character) in the initial speech recognition text belongs to the scope of the ASR dictionary. And the words included in the ASR dictionary are limited (for example, some rare characters may not be included in the ASR dictionary). Therefore, the ASR model can only recognize the words within the scope of the ASR dictionary in the speech and cannot recognize the out-of-dictionary words outside the ASR dictionary; that is, if the target speech includes out-of-dictionary words, the ASR model cannot output these out-of-dictionary words.
[0102] In the embodiments of the present application, by using a phoneme recognition model to obtain phoneme data, it is possible to capture the phonemes corresponding to each syllable in the target speech, so that the phoneme data includes complementary information in the initial speech recognition text; for example, the phoneme data may include the phonemes of the pronunciation of out-of-dictionary words that the ASR model did not recognize. And constructing the phoneme data and the initial speech recognition text into the target prompt information enables the large language model to start from the basic unit of pronunciation (that is, each phoneme in the phoneme data) and correct the out-of-dictionary words that are not in the ASR dictionary; for example, through the phonemes corresponding to out-of-dictionary words, rare characters, etc. in the phoneme data, the initial speech recognition text is corrected to the correct text corresponding to the target speech.
[0103] For the case of out-of-dictionary words, for example, for the target speech corresponding to "龙行龘龘", the output results of the ASR model and the large language model are compared as follows: The output result of the ASR model is: "龙行达达"; The result after correction by the large language model output is: "龙行龘龘".
[0104] It can be seen that the rare character "龘" is missing in the ASR dictionary, and the ASR model cannot output the correct text "龘". By inputting the target prompt information including the ASR output result "龙行达达" and the pinyin "longxingdada" into the large language model, after correction by the large language model, the correct text output by the large language model should be "龙行龘龘".
[0105] It should be noted further that the error situations in the recognition results of the ASR model can be attributed to different phonetically similar characters or different homophonic characters. For example, phonetic similarity means that the initials or finals of Chinese are the same, and some phonemes of English are the same or similar. For example, the situations where the pronunciations in Chinese are easily confused include: homophonic characters, polyphonic characters, rhyming (same finals, different initials), alliteration (same initials, different finals), etc.
[0106] Since the accuracy of phoneme data is much higher than that of the initial speech recognition text, combining phoneme data can provide more and more effective information for reference in the processing of large language models, which helps the large language models find the correct text during processing and thus improves the accuracy of text processing.
[0107] Furthermore, the initial speech recognition text obtained by the ASR model may include incorrectly identified words; while the phoneme data includes the phonemes of the correct pronunciations corresponding to the incorrect words; through the embodiments of this application, the large language model can combine phoneme data (especially the phonemes of the correct pronunciations corresponding to the incorrect words) for correction, which helps the large language model to correct the incorrect words into correct words, thereby improving the accuracy of the large language model in text processing.
[0108] Furthermore, in this embodiment, starting with each syllable of pronunciation, it can greatly solve the problem of pronunciation confusion (such as homophones, polyphonic characters, rhymes, alliteration, etc.), and effectively deal with text recognition errors related to dialects and accents in speech recognition technology. It greatly improves the accuracy of correcting various possible errors in the text, thereby more realistically and finely restoring the text units corresponding to each syllable in the speech scene, accurately restoring the correct text corresponding to each syllable in the target speech, and thus improving the accuracy and precision of speech recognition text.
[0109] In related technologies, only the initial speech recognition text to be corrected is input into the relevant model without using phoneme data. The accuracy of the output results of the relevant model is much lower than that of the method in the embodiments of this application.
[0110] The following is a specific example comparing the related technologies and the methods of the embodiments of this application: For example, for speech with accents or heavy accented dialects, the initial recognized text output by the ASR model may contain misrecognized text.
[0111] If the method of this embodiment is used, the target prompt information input into the large language model is as follows: "Based on the provided N-best text output and N-best pinyin transcription, the text is corrected and the results are output directly."
[0112] N-best text output: 1. But considering the facts, Ma Ning was unwilling to be led along. 2. But when faced with the fact that Ma Ning is unwilling to lead the way, you realize that Ma Ning is unwilling to do so. 3. Back then, when faced with the fact that Ma Ning was unwilling to be led by the hand, Ma Ning refused. 4. However, considering the facts, Ma Ning was unwilling to be humble. 5. But when faced with the truth, Ma Ning is unwilling to be humble. N-best Pinyin transcription: 1. d an n ian m ian d ui man ing sh i sh imazebuy uan q ianx ing 2. d an nim ian d ui man ing sh i sh imazebuy uan q ian xing 3. d ang n ian m ian d ui man ing sh i sh imazebuy uan q ianx ing 4. d an n ian m ian d ui man ing sh i sh imazebuy uan q ianx in 5. d an nim ian d ui man ing sh i sh imazebuy uan q ian xin”.
[0113] In this embodiment of the application, the output of the large language model is: "When you stare at a horse, the horse is unwilling to move forward"; If the relevant technology method is used, the information input to the relevant model in the relevant technology only contains the text that needs to be corrected, and does not contain the N-best pinyin. After testing, the output result of the relevant model in the relevant technology is: "But when facing Ma Ning, the fact is that Ma is unwilling to lead the way".
[0114] By comparing related technologies and the methods of the embodiments of this application, it can be concluded that the embodiments of this application can effectively correct the misidentified text caused by heavy accents in speech by combining pinyin, and achieve a high accuracy rate.
[0115] Another point to note is that by flexibly setting various possible auxiliary information, such as scene description information, contextual speech text information, or contextual speech, the relevant phonemes in the phoneme data can be matched with the relevant text in the auxiliary information; for example, the text of a professional technical term in a certain field can be matched with the relevant phonemes in the phoneme data corresponding to that technical term. This helps the large language model to effectively correct errors and enables the embodiments of this application to better cope with some difficult speech recognition scenarios, such as speech recognition scenarios in professional fields containing multiple professional technical terms and specialized vocabulary; thereby improving the accuracy and practicality of the large language model in text processing.
[0116] For uncommon or specialized vocabulary, such as the target speech corresponding to "sand gulls soaring together," the initial speech recognition result output by the ASR model is "sand gulls remember." However, in this embodiment, by adding auxiliary information such as the classroom topic "classical poetry" or the contextual phrase "water birds sometimes fly up," and combining the 2-best pinyin output by the phoneme recognition model and the 1-best text output by the speech recognition model, the target prompt information input to the large language model is obtained as follows: "This is an online lesson about classical Chinese poetry. The previous sentence in the lesson is 'Water birds sometimes take flight.' Based on the provided 1-best text and 2-best pinyin, correct the text and output the result directly."
[0117] The 1-best text is: The seagull remembered; 2-best (pinyin: ) sha ou xiang ji sha ou xiang qi” In this embodiment of the application, the text output by the large language model is: "Sand gulls gather".
[0118] Step 203: Based on the target prompt information, the electronic device obtains the target recognition text after correcting the initial speech recognition text through the large language model.
[0119] In this step, the electronic device can call the large language model to process the target prompt information and obtain the target recognition text.
[0120] For example, this large language model can be an LLM model based on the Transformer architecture. This LLM model can output results in an autoregressive form. The target prompt information can be called a token sequence; the LLM model can predict each token (lexical unit) in the target recognition text one by one. For example, the LLM model can predict the next token based on the token sequence corresponding to the input target prompt information; then, based on the original input (i.e., the target prompt information) and the already predicted tokens, it continues to predict the next token, and so on. Here, a token refers to the basic unit of text; in Chinese, this could be a single Chinese character. Alternatively, the BPE (Byte Pair Encoding) method can be used to train the basic units of the text from the text corpus.
[0121] For example, Figure 6 This is a schematic diagram illustrating the input and output of a large language model provided in an embodiment of this application. For example... Figure 6As shown, the input target prompt information includes the prompt word "prompt" from the error-correcting prompt template, the N-best text, the N-best pinyin, and auxiliary information (where N ≥ 1). For example, auxiliary information may include the meeting topic, meeting keywords, technical terms, etc. Figure 3 As shown, after processing by the LLM model, the optimized text can be obtained, which is the target recognition text after correcting the initial speech recognition text.
[0122] For example, the following shows a possible input and output example of a large language model.
[0123] Input may include the following: "This is a conference about agricultural development, and the following keywords may be encountered: sustainable development, organic crops, and ecological niche. Based on the provided N-best text candidates and N-best phonetic transcription, the text is corrected and the results are output directly."
[0124] N-best candidate text (which can also be the text with the highest score): 1. The weather is very nice today. 2. The weather is very early today. 3. The weather is very nice today. 4. It smells very sweet today. 5. The weather is very nice. N-best Pinyin transcription: 1. j in t ian t ian qih en h ao 2. j in t ian t ian qih en z ao 3. j in t ian t ian qih en h ao 4. j in t ian t ian qih en h ao 5. j ing t ian t ian qih en h ao” Correspondingly, the output could include: "The weather is nice today".
[0125] It should be noted that this electronic device can directly use existing large language models without training them. For example, models with strong LLM base models can be used directly. Alternatively, the electronic device can be trained on top of existing large language models, for example, by fine-tuning them with a small number of samples to enhance their performance.
[0126] In one possible embodiment, the large language model is obtained by performing at least one of the following fine-tuning processes: The parameter matrix of the large language model to be trained is decomposed into two low-rank matrices. Based on the initial recognition text and sample phoneme data of each speech sample in the sample set, sample prompt information is generated for each speech sample; Based on the sample prompt information of each speech sample, the predicted text after correcting the initial recognized text of each speech sample is obtained through the parameter matrix of the large language model to be trained. The training loss for each speech sample is calculated based on the similarity between the predicted text and the real speech text of each speech sample. Based on the training loss corresponding to each speech sample, the two low-rank matrices are adjusted, and the parameter matrix of the large language model to be trained is updated based on the adjusted two low-rank matrices.
[0127] For example, the parameter matrix of a large language model is large, but it can be decomposed into two low-rank matrices, which can greatly reduce the dimensionality of the parameter matrix.
[0128] The size of the product of the two low-rank matrices is the size of the parameter matrix.
[0129] The parameter matrix represents the actual parameters of the large language model. For example, the parameter matrix can be a d×d matrix. It can be decomposed into two low-rank matrices of size d×d: a d×r low-rank matrix and an r×d low-rank matrix, where r is less than d. When the d×r and r×d low-rank matrices are multiplied, the size of the resulting matrix serves as the size of the d×d parameter update matrix to be optimized. During each fine-tuning iteration, the d×r and r×d matrices are fine-tuned using the training loss; the updated parameter matrix is then obtained using the fine-tuned d×r and r×d matrices.
[0130] Figure 7 This is a schematic diagram illustrating a fine-tuning process provided in an embodiment of this application. Figure 7 As shown, the actual parameter matrix W∈R of the large language model d×d The d×d parameter matrix can be decomposed into two low-rank matrices, namely matrix A and matrix B. Matrix A can be initialized using a Gaussian distribution, and matrix B can be initialized as a zero matrix.
[0131] The parameter matrix can be frozen, meaning that the predicted text for each speech sample is predicted using the parameter matrix, but the training loss corresponding to each speech sample is not used to adjust the parameter matrix. Instead, the training loss corresponding to each speech sample is used to adjust matrices A and B. Because the dimensions of matrices A and B are much smaller than those of the parameter matrix W, the backpropagation steps for adjusting matrices A and B are much faster than the process of adjusting the parameter matrix W, and good results can be obtained with only a small number of samples.
[0132] It should be noted that a large language model may include multiple parameter matrices, and each parameter matrix can be updated using the fine-tuning methods described above. Furthermore, one or more fine-tuning operations can be performed until a target condition is met. For example, the target condition may include, but is not limited to: the number of fine-tuning operations exceeding a target threshold, or the training loss corresponding to multiple fine-tuning processes stabilizing and falling below a target loss threshold.
[0133] Figure 8 This is a schematic diagram illustrating a text processing process using a large language model, as provided in an embodiment of this application. Figure 8 As shown, the initial speech recognition text of the target speech can be obtained using an ASR model; phoneme data of the target speech can be obtained through a phoneme recognition model or by directly transcribing the initial speech recognition text; and auxiliary information such as the meeting topic, meeting keywords, and contextual speech text information of the target speech can be obtained, or phoneme reference information can also be obtained. Then, the initial speech recognition text, phoneme data, and at least one of the auxiliary information or phoneme reference information are combined with a prompt template to generate target prompt information. The target prompt information is then input into a large language model to obtain the target recognition text output by the large language model.
[0134] The text processing method provided in this application determines the initial speech recognition text and phoneme data of the target speech, generates target prompt information based on the initial speech recognition text and phoneme data, and obtains the target recognition text through a large language model based on the target prompt information. This allows the large language model to correct the initial speech recognition text based on the phoneme data. Since the accuracy of the phoneme data is much higher than that of the initial speech recognition text, combining the phoneme data can provide more and more effective information for reference in the processing of the large language model, which helps the large language model obtain the correct text in the processing process, thereby improving the accuracy of text processing.
[0135] Furthermore, the initial speech recognition text may contain incorrectly identified words; while the phoneme data includes all or part of the phonemes corresponding to the correct pronunciation of the incorrect words; by prompting the large language model to combine phoneme data for correction, it helps the large language model to correct the incorrect words to the correct words, thereby improving the accuracy of the large language model in text processing.
[0136] Furthermore, by using a phoneme recognition model to obtain phoneme data, it is possible to capture the phonemes corresponding to each syllable in the target speech, so that the phoneme data includes complementary information from the initial speech recognition text; this allows the large language model to start from the basic unit of pronunciation, correct out-of-vocabulary words not found in the ASR dictionary, and correct the initial speech recognition text to the correct text corresponding to the target speech.
[0137] Furthermore, by focusing on each syllable of pronunciation, it can greatly resolve pronunciation confusion (such as homophones, polyphonic characters, rhymes, alliteration, etc.) and effectively address text recognition errors related to dialects and accents in speech recognition technology. In addition, by combining phoneme reference information, it can further improve the accuracy of correcting various possible errors in the text, thereby more realistically and finely reproducing the text units corresponding to each syllable in the speech scene, and thus improving the accuracy and precision of speech recognition text.
[0138] Furthermore, by flexibly setting rich auxiliary information, the large language model can match the relevant phonemes in the phoneme data with the relevant text in the auxiliary information; for example, it can match the text of technical terms with the relevant phonemes corresponding to the specialized technical terms; this helps the large language model to effectively correct errors; and enables the embodiments of this application to better cope with some difficult speech recognition scenarios, thereby improving the accuracy and practicality of the large language model in text processing.
[0139] Figure 9 This is a schematic diagram of the structure of a text processing device provided in an embodiment of this application. Figure 9 As shown, the device includes: a first determining module 901, a generating module 902, and a correcting module 903.
[0140] The first determining module 901 is used to determine the initial speech recognition text and phoneme data of the target speech; The generation module 902 is used to generate target prompt information based on the initial speech recognition text and phoneme data of the target speech. The target prompt information is used to prompt the large language model to correct the initial speech recognition text based on the phoneme data. The correction module 903 is used to obtain the corrected target recognition text of the initial speech recognition text based on the target prompt information and through the large language model.
[0141] In one possible implementation, the device further includes: The second determining module is used to determine the auxiliary information of the target speech; This generation module 902 is used for: Based on the initial speech recognition text, phoneme data, and auxiliary information of the target speech, the target prompt information is generated. This target prompt information is used to prompt the correction of the initial speech recognition text based on the auxiliary information and phoneme data.
[0142] In one possible implementation, the second determining module, when determining the auxiliary information of the target speech, is specifically used for any of the following: The scene description information of the target scene to which the target speech belongs is used as this auxiliary information; The text information of the context of the target speech is used as the auxiliary information.
[0143] In one possible implementation, the generation module 902 is used for: Based on the initial speech recognition text, phoneme data, and phoneme reference information, target prompt information is generated. This target prompt information is used to prompt the initial speech recognition text to be corrected based on the phoneme reference information and phoneme data.
[0144] In one possible implementation, the first determining module 901, when determining phoneme data, is used to: Phoneme identification is performed on each syllable in the target speech to obtain initial phoneme data; Based on the target type to which the target speech belongs in at least one pronunciation type, the initial phoneme data of the target speech is corrected to obtain the phoneme data.
[0145] In one possible implementation, the device further includes: The acquisition module is used to acquire phoneme reference information of the target speech. The phoneme reference information includes at least one of the following: The target speech belongs to at least one pronunciation type; The target correction relationship corresponding to the target type includes the correspondence between at least one phoneme pronounced according to the target type and the standard pronunciation phoneme; The phoneme correction relationship for each of the at least one pronunciation type.
[0146] In one possible implementation, the first determining module 901, when determining the initial speech recognition text, is used to: Based on the target speech, multiple candidate texts and the confidence level of each candidate text are obtained by using a speech recognition model. The average similarity of each candidate text is determined based on the similarity between each candidate text and each other candidate text. Based on the confidence and average similarity of each candidate text, at least one initial speech recognition text is obtained from the multiple candidate texts; Specifically, the first determining module 901, when determining phoneme data, is used to perform any of the following steps: Based on the target speech, phoneme recognition model is used to identify each syllable in the target speech, and at least one phoneme data and the confidence level of each phoneme data are obtained. Each initial speech recognition text is processed by a phoneme conversion tool to obtain at least one phoneme data of the target speech.
[0147] In one possible implementation, if the at least one phoneme data and the confidence level of each phoneme data are obtained through a phoneme recognition model, then the generation module 902 is used to: Based on each initial speech recognition text and its confidence level, each phoneme data and its confidence level, the target prompt information is generated. This target prompt information is used to prompt the correction of each initial speech recognition text based on each initial speech recognition text and its confidence level, and each phoneme data and its confidence level.
[0148] In one possible implementation, the device further includes a fine-tuning module for performing at least one fine-tuning process on the large language model: The parameter matrix of the large language model to be trained is decomposed into two low-rank matrices. Based on the initial recognition text and sample phoneme data of each speech sample in the sample set, sample prompt information is generated for each speech sample; Based on the sample prompt information of each speech sample, the predicted text after correcting the initial recognized text of each speech sample is obtained through the parameter matrix of the large language model to be trained. The training loss for each speech sample is calculated based on the similarity between the predicted text and the real speech text of each speech sample. Based on the training loss corresponding to each speech sample, the two low-rank matrices are adjusted, and the parameter matrix of the large language model to be trained is updated based on the adjusted two low-rank matrices.
[0149] The text processing method provided in this application determines the initial speech recognition text and phoneme data of the target speech, generates target prompt information based on the initial speech recognition text and phoneme data, and obtains the target recognition text through a large language model based on the target prompt information. This allows the large language model to correct the initial speech recognition text based on the phoneme data. Since the accuracy of the phoneme data is much higher than that of the initial speech recognition text, combining the phoneme data can provide more and more effective information for reference in the processing of the large language model, which helps the large language model obtain the correct text in the processing process, thereby improving the accuracy of text processing.
[0150] The apparatus in this application embodiment can execute the method provided in this application embodiment, and the implementation principle is similar. The actions performed by each module in the apparatus of each embodiment of this application correspond to the steps in the method of each embodiment of this application. For detailed functional descriptions of each module of the apparatus, please refer to the descriptions in the corresponding methods shown above, which will not be repeated here.
[0151] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example... Figure 10 As shown, the electronic device includes: a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of a text processing method, which, compared with related technologies, can achieve: The text processing method provided in this application determines the initial speech recognition text and phoneme data of the target speech, generates target prompt information based on the initial speech recognition text and phoneme data, and obtains the target recognition text through a large language model based on the target prompt information. This allows the large language model to correct the initial speech recognition text based on the phoneme data. Since the accuracy of the phoneme data is much higher than that of the initial speech recognition text, combining the phoneme data can provide more and more effective information for reference in the processing of the large language model, which helps the large language model obtain the correct text in the processing process, thereby improving the accuracy of text processing.
[0152] In one alternative embodiment, an electronic device is provided, such as Figure 10 As shown, Figure 10 The illustrated electronic device 1000 includes a processor 1001 and a memory 1003. The processor 1001 and the memory 1003 are connected, for example, via a bus 1002. Optionally, the electronic device 1000 may further include a transceiver 1004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 1004 is not limited to one type, and the structure of the electronic device 1000 does not constitute a limitation on the embodiments of this application.
[0153] Processor 1001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 1001 may also be a combination that implements computing functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0154] Bus 1002 may include a pathway for transmitting information between the aforementioned components. Bus 1002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 1002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 10 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0155] The memory 1003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation herein.
[0156] The memory 1003 is used to store computer programs that execute the embodiments of this application, and the execution is controlled by the processor 1001. The processor 1001 is used to execute the computer programs stored in the memory 1003 to implement the steps shown in the foregoing method embodiments.
[0157] Electronic devices include, but are not limited to, servers, terminals, or cloud computing center equipment.
[0158] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the steps and corresponding content of the aforementioned method embodiments.
[0159] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.
[0160] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0161] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. The terms “comprising” and “including” as used in the embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, or operation, but do not exclude implementation as other features, information, data, steps, or operations supported by this art.
[0162] The terms "first," "second," "third," "fourth," "1," "2," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown in the illustrations or text descriptions.
[0163] It should be understood that although arrows indicate various operation steps in the flowcharts of this application's embodiments, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of this application's embodiments, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all steps in each flowchart, based on the actual implementation scenario, may include multiple sub-steps or multiple stages. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and this application's embodiments do not limit this.
[0164] The above description is only an optional implementation method for some implementation scenarios of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application without departing from the technical concept of this application also fall within the protection scope of the embodiments of this application.
Claims
1. A text processing method, characterized in that, The method includes: Determine the initial speech recognition text and phoneme data for the target speech; Based on the initial speech recognition text and phoneme data of the target speech, target prompt information is generated. The target prompt information is used to prompt the large language model to correct the initial speech recognition text based on the phoneme data. Based on the target prompt information, the target recognition text after correcting the initial speech recognition text is obtained through the large language model.
2. The method according to claim 1, characterized in that, The method further includes: Determine the auxiliary information of the target speech; The generation of target prompt information based on the initial speech recognition text and phoneme data of the target speech includes: Based on the initial speech recognition text, phoneme data, and auxiliary information of the target speech, the target prompt information is generated. The target prompt information is used to prompt the correction of the initial speech recognition text based on the auxiliary information and phoneme data.
3. The method according to claim 2, characterized in that, The auxiliary information for determining the target speech includes any one of the following: The scene description information of the target scene to which the target speech belongs is used as the auxiliary information; The text information of the context of the target speech is used as the auxiliary information.
4. The method according to claim 1, characterized in that, The generation of target prompt information based on the initial speech recognition text and phoneme data of the target speech includes: Based on the initial speech recognition text, phoneme data, and phoneme reference information, target prompt information is generated. The target prompt information is used to prompt the initial speech recognition text to be corrected based on the phoneme reference information and phoneme data.
5. The method according to claim 1, characterized in that, The method for determining the phoneme data includes: Phoneme recognition is performed on each syllable in the target speech to obtain initial phoneme data; Based on the target type to which the target speech belongs in at least one pronunciation type, the initial phoneme data of the target speech is corrected to obtain the phoneme data.
6. The method according to claim 4 or 5, characterized in that, The method further includes: Obtain phoneme reference information of the target speech; The phoneme reference information includes at least one of the following: The target speech belongs to the target type in at least one pronunciation type; The target type corresponds to the target correction relationship, which includes the correspondence between at least one phoneme pronounced according to the target type and the standard pronunciation phoneme; The phoneme correction relationship for each of the at least one pronunciation type.
7. The method according to any one of claims 1-6, characterized in that, The initial speech recognition text is obtained by performing the following steps: Based on the target speech, multiple candidate texts of the target speech and the confidence level of each candidate text are obtained by a speech recognition model; The average similarity of each candidate text is determined based on the similarity between each candidate text and each other candidate text. Based on the confidence and average similarity of each candidate text, at least one initial speech recognition text is obtained from the plurality of candidate texts; The phoneme data is obtained by performing any of the following steps: Based on the target speech, phoneme recognition is performed on each syllable in the target speech using a phoneme recognition model to obtain at least one phoneme data and the confidence level of each phoneme data. Each initial speech recognition text is processed by a phoneme conversion tool to obtain at least one phoneme data of the target speech.
8. The method according to claim 7, characterized in that, If the at least one phoneme data and the confidence level of each phoneme data are obtained through the phoneme recognition model, then the generation of target prompt information based on the initial speech recognition text and phoneme data of the target speech includes: Based on each initial speech recognition text and its confidence level, each phoneme data and its confidence level, the target prompt information is generated. The target prompt information is used to prompt correction of each initial speech recognition text based on each initial speech recognition text and its confidence level, each phoneme data and its confidence level.
9. The method according to claim 1, characterized in that, The large language model was obtained by performing at least one of the following fine-tuning processes: The parameter matrix of the large language model to be trained is decomposed into two low-rank matrices. Based on the initial recognition text and sample phoneme data of each speech sample in the sample set, sample prompt information is generated for each speech sample; Based on the sample prompt information of each speech sample, the predicted text after correcting the initial recognized text of each speech sample is obtained through the parameter matrix of the large language model to be trained. The training loss for each speech sample is calculated based on the similarity between the predicted text and the real speech text of each speech sample. Based on the training loss corresponding to each speech sample, the two low-rank matrices are adjusted, and the parameter matrix of the large language model to be trained is updated based on the adjusted two low-rank matrices.
10. A text processing device, characterized in that, The device includes: The first determining module is used to determine the initial speech recognition text and phoneme data of the target speech; The generation module is used to generate target prompt information based on the initial speech recognition text and phoneme data of the target speech, and the target prompt information is used to prompt the large language model to correct the initial speech recognition text based on the phoneme data; The correction module is used to obtain the corrected target recognition text based on the target prompt information and through the large language model.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the text processing method according to any one of claims 1 to 9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the text processing method according to any one of claims 1 to 9.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the text processing method according to any one of claims 1 to 9.
Citation Information
Cited By
Speech recognition method and device, model training method and device, electronic equipment and medium
CN121331105A
Logistics communication verbal skill and dialect real-time talkback and training system and method based on large language model
CN122050375A
Logistics communication speech and dialect real-time talkback and training system and method based on large language model
CN122050375B