Speech Recognition Method and Related Devices, Equipment and Storage Media

Through the autoregressive decoding method of the large language model, combined with feature encoding and decoding, post-processing is directly implemented in the speech recognition process, solving the problems of long response time and heavy computing burden in traditional speech recognition, and improving output accuracy and adaptability.

CN119993163BActive Publication Date: 2025-07-11IFLYTEK CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510454672.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-11
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

On the premise of ensuring that the recognition text has a display form, traditional speech recognition technology has a long response time and a heavy calculation burden, and there are error accumulation problems in the cascade structure, which affects the output accuracy.

Method used

The autoregressive decoding method based on the large language model is adopted. Through the combination of feature encoding and autoregressive decoding, the prompt instructions include recognition and post-processing instructions, which directly realizes speech recognition and post-processing during the decoding process, avoids cascading structure, improves output accuracy and shortens response time.

Benefits of technology

On the premise of ensuring that the recognition text has a display form, the response time of speech recognition is shortened, the calculation burden is reduced, and the output accuracy is improved, adapting to differentiated processing in different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993163B_ABST
    Figure CN119993163B_ABST
Patent Text Reader

Abstract

The present application discloses a speech recognition method and related devices, equipment, and storage media. Among them, the speech recognition method includes: performing feature encoding on the speech to be recognized to obtain encoded features; performing autoregressive decoding on the encoded features by a large language model according to a prompt instruction to obtain a recognition text in a displayable form for the speech to be recognized; wherein the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform speech recognition with the displayable form as an additional target, and various post-processing instructions correspond to different display forms. The above solution can shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition on the premise of ensuring that the recognition text has a displayable form as much as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of speech processing, and in particular, to a speech recognition method and related devices, equipment, and storage media. Background Art

[0002] Automatic Speech Recognition (ASR) technology can extract the text information in speech and transcribe it into text, which is widely used in many fields such as intelligent customer service, intelligent office, smart home, vehicle control, and speech simultaneous interpretation.

[0003] It should be noted that the transcribed text of traditional speech recognition technology is pure text without presentation forms such as punctuation marks, numbers, upper and lower cases, etc. Since it is not convenient for reading or downstream tasks (such as semantic understanding, etc.), in a speech recognition system, a traditional speech recognition model generally cascades a post-processing model for Natural Language Processing (NLP) to post-process the transcribed text presented as pure text into a recognized text with a presentation form. However, this cascaded structure will undoubtedly prolong the response time, increase the computational burden, and due to the error accumulation problem in the cascaded structure, it is easy to degrade the final recognized text. In view of this, how to shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition while ensuring that the recognized text has a presentation form as much as possible has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a speech recognition method and related devices, equipment, and storage media, which can shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition while ensuring that the recognized text has a presentation form as much as possible.

[0005] To solve the above technical problem, a first aspect of this application provides a speech recognition method, including: performing feature encoding on the speech to be recognized to obtain encoded features; performing autoregressive decoding on the encoded features according to a prompt instruction by a large language model to obtain a recognized text with a presentation form of the speech to be recognized; where the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform speech recognition with a presentation form as an additional target, and various post-processing instructions correspond to different presentation forms.

[0006] In order to solve the above technical problems, the second aspect of the present application provides a speech recognition device, including: a feature encoding module and an autoregressive decoding module, the feature encoding module is used to perform feature encoding based on the speech to be recognized to obtain the encoded features; the autoregressive decoding module is used to perform autoregressive decoding on the encoded features according to the prompt instructions based on the large language model to obtain the recognition text of the speech to be recognized with a display form; wherein the prompt instructions include recognition instructions and at least one post-processing instruction, the recognition instructions are used to instruct the large language model to perform speech recognition, and the post-processing instructions are used to instruct the large language model to perform speech recognition with a display form as an additional goal, and various post-processing instructions correspond to different display forms.

[0007] In order to solve the above technical problems, the third aspect of the present application provides an electronic device, which at least includes a memory and a processor coupled to each other, the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the speech recognition method in the above first aspect.

[0008] In order to solve the above technical problems, the fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be executed by a processor, and the program instructions are used to implement the speech recognition method of the first aspect.

[0009] The above scheme performs feature encoding based on the speech to be recognized to obtain the encoded features, and performs autoregressive decoding on the encoded features based on the large language model according to the prompt instruction to obtain the recognized text with a display form of the speech to be recognized, and the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform speech recognition with the additional goal of having a display form, and various post-processing instructions correspond to different display forms, so on the one hand, by inputting the post-processing instruction together with the recognition instruction, forcing the large language model to perform post-processing together with the speech recognition during the autoregressive decoding process, it is possible to ensure that the recognized text has a display form as much as possible, and on the other hand, because the speech recognition and its post-processing are completed together by the large language model during the autoregressive decoding process, that is, there is no need to successively realize speech transcription and post-processing through a cascade structure, which can shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition. Therefore, it is possible to shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition under the premise of ensuring that the recognized text has a display form as much as possible.

[0010] In addition, since various post-processing instructions correspond to different display forms, different post-processing instructions can be configured specifically according to different application scenarios during the speech recognition process, thereby improving the convenience of differentiated processing of different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is a schematic flowchart of an embodiment of the speech recognition method of the present application;

[0012] Figure 2a It is a schematic diagram of the process of an embodiment of the parameter fine-tuning of the present application;

[0013] Figure 2b It is a schematic diagram of the effect of an embodiment of the target attention mask of the present application;

[0014] Figure 2c It is a schematic diagram of the effect of an embodiment of the mask prediction task of the present application;

[0015] Figure 2d It is a schematic diagram of the effect of another embodiment of the mask prediction task of the present application;

[0016] Figure 3 It is a schematic framework diagram of an embodiment of the speech recognition device of the present application;

[0017] Figure 4 It is a schematic framework diagram of an embodiment of the electronic device of the present application;

[0018] Figure 5 It is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. Detailed Embodiments

[0019] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.

[0020] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0021] The terms "system" and "network" are often used interchangeably in this article. The term " / and / " in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the fragment " / " in this article generally represents an "or" relationship between the front and back associated objects. In addition, "multiple" in this article means two or more than two.

[0022] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the speech recognition method of the present application. Specifically, it may include the following steps:

[0023] Step S11: Perform feature encoding on the speech to be recognized to obtain encoded features.

[0024] In an implementation scenario, the speech to be recognized can be streaming speech, that is, the speech to be recognized can be collected in real time while the speaker is speaking; alternatively, the speech to be recognized can also be non-streaming speech. For example, the speaker can start recording when starting to speak and end recording when ending speaking, and a voice recording can be used as the speech to be recognized. Of course, the above examples are only several possible examples of the speech to be recognized, and the acquisition method of the speech to be recognized is not limited herein, and no further examples will be given one by one.

[0025] In an implementation scenario, as a possible implementation manner, feature encoding can be implemented by an encoder. Exemplarily, the encoder can include, but is not limited to, network structures such as Transformer and Conformer. The network structure of the encoder is not limited herein.

[0026] In another implementation scenario, as another possible implementation manner, feature encoding can also be implemented by an encoder that is compatible with both speech and text modalities, so that not only sample speech data can be used for model training during the training process, but also widely existing sample text data can be used to assist in training. The specific training process can refer to the following relevant descriptions and will not be elaborated herein. Exemplarily, the encoder that is compatible with both speech and text modalities can include, but is not limited to: a Conformer network structure composed of several Conformer blocks. Each Conformer block can include a series of multi-head self-attention, depth convolution, and feed-forward layers. For the technical details of the Conformer network structure, refer to the relevant content and will not be elaborated herein. Exemplarily, the last 8-layer block structure of the Conformer network structure can be shared by speech and text. Of course, the above examples are only one possible example of the encoder, and the specific structure of the encoder that is compatible with both speech and text modalities is not limited herein, and no further examples will be given one by one.

[0027] In yet another implementation scenario, as yet another possible implementation example, the feature dimension after performing feature encoding can be different from the feature dimension of the embedding layer in the subsequent large language model. In this case, before performing subsequent autoregressive decoding using the large language model after performing feature encoding, dimensionality conversion can be first performed through an adapter (that is, dimensionality conversion is performed on the output features after performing feature encoding) to be consistent with the feature dimension of the embedding layer. It should be noted that the adapter can include, but is not limited to, network structures such as convolutional layers and fully connected layers. The network structure of the adapter is not limited herein. In addition, the adapter can be fine-tuned together with the large language model for parameters. For the specific details, refer to the subsequent training-related descriptions and will not be elaborated herein.

[0028] It should be noted that, as a possible example, in order to facilitate the encoder to perform feature encoding, before using the encoder to perform feature encoding on the speech to be recognized, the acoustic features of the speech to be recognized can be extracted first. For example, the acoustic features can include but are not limited to FBank, etc. The specific types of acoustic features are not limited herein. On this basis, the encoder can perform feature encoding on the acoustic features to obtain encoded features.

[0029] Step S12: Based on the large language model, autoregressive decoding is performed on the encoded features according to the prompt instructions to obtain the recognition text in the presentation form of the speech to be recognized.

[0030] In the embodiments of the present disclosure, the prompt instructions may include recognition instructions and at least one post-processing instruction. The recognition instructions can be used to instruct the large language model to perform speech recognition, and the post-processing instructions can be used to instruct the large language model to perform speech recognition with the presentation form as an additional target, and various post-processing instructions correspond to different presentation forms respectively. It should be noted that the large language model can include but are not limited to open-source large models such as Llama, Bloom, Spark 13B, GLM6B, etc.; or, the large language model can also include but are not limited to being fine-tuned based on specific corpora; or, the large language model can also include but are not limited to custom large models. The specific sources of the large language model are not limited herein. In addition, during the autoregressive decoding process, at the first decoding, usually by default, the starting character indicating the start of decoding (such as <sos>etc.), and the predicted characters of the first decoding output can be obtained by decoding through the large language model; when continuing to decode, the predicted characters of the previous decoding output are input together with the model input data such as the encoded features, and the predicted characters of the current decoding output can be obtained by decoding through the large language model, and so on, repeating in a loop until the predicted characters of a certain decoding output are the end characters representing the end of decoding (e.g., <eos>), the predicted characters output by each decoding can be sequentially combined to complete speech recognition. Of course, the above description of autoregressive decoding is only a brief explanation of autoregressive decoding. For specific technical details, please refer to the autoregressive decoding, which will not be elaborated here.

[0031] In one implementation scenario, the recognition instruction can be embodied in the form of a natural language description text to instruct the large language model to perform speech recognition. Exemplarily, the recognition instruction can include, but is not limited to, the following content: "Please perform speech recognition based on the encoding features of the input speech", etc. The specific content of the recognition instruction is not limited here.

[0032] In one implementation scenario, similar to the recognition instruction, the post-processing instruction can also be embodied in the form of a natural language description text and instruct the large language model to perform speech recognition with the presentation form as an additional goal. Taking the presentation form including "case" as an example, the post-processing instruction corresponding to the presentation form "case" can include, but is not limited to: "Please distinguish between uppercase and lowercase during recognition", etc. The specific content of the post-processing instruction corresponding to the presentation form "case" is not limited here; or, taking the presentation form including "number" as an example, the post-processing instruction corresponding to the presentation form "number" can include, but is not limited to: "Please distinguish between Arabic numerals and ordinary text during recognition", etc. The specific content of the post-processing instruction corresponding to the presentation form "number" is not limited here; or, taking the presentation form including "punctuation" as an example, the post-processing instruction corresponding to the presentation form "punctuation" can include, but is not limited to: "Please add appropriate punctuation marks to the text according to the semantics during recognition", etc. The specific content of the post-processing instruction corresponding to the presentation form "punctuation" is not limited here. Of course, the above examples are only several possible examples of the post-processing instruction when the presentation form includes cases such as uppercase and lowercase, numbers, and punctuation. The specific content of other possible contents and the post-processing instruction when the presentation form includes other cases are not limited here. For example, the presentation form can also include, but is not limited to, measurement units (such as currency units like US dollars $), etc., and will not be listed one by one here.

[0033] In one implementation scenario, as a possible implementation, please refer to Table 1. Table 1 is a schematic table of an example of the model input data of a large language model. As shown in Table 1, optionally, during the process of autoregressive decoding, the encoded features can also be concatenated with a reference text, and the reference text can specifically be the dialogue text in the way before the speech to be recognized, so as to further refer to the previous dialogue text during the current speech recognition (for example, the dialogue text can be the recognized text obtained by speech recognition of the historical speech before the speech to be recognized and having a presentation form), which helps to improve the speech recognition performance in the long speech stream scenario. In addition, in the case where the encoded features are concatenated with a reference text, as a possible example, the large language model can also output the predicted characters of the characters related to the presentation form in the reference text, so as to correct the characters related to the presentation form in the reference text according to the predicted characters, obtain the corrected text of the reference text, and use the corrected text to overwrite the reference text. For example, the speech to be recognized is "And you?", and the previous dialogue text can be "I plan to go out and play next week." as the reference text. Then, through autoregressive decoding, the recognized text of the speech to be recognized with a presentation form "And you?" and the predicted character "," of the character related to the presentation form "." in the reference text can be obtained. Then, the reference text can be corrected based on this to obtain the corrected text "I plan to go out and play next week,". After that, the corrected text and the recognized text of the speech to be recognized with a presentation form can be concatenated to obtain the latest recognition result "I plan to go out and play next week, and you?". In this way, by repeating this process, during the speech recognition process (especially in the long speech stream scenario), the dialogue text before each recognition can be corrected, and finally, a recognition result that is as accurate as possible and has a presentation form can be obtained.

[0034] Table 1 Schematic table of an example of the model input data of a large language model

[0035]

[0036] It should be noted that, as shown in Table 1, when the prompt instruction does not contain a post-processing instruction, the speech recognition method of the embodiments of the present disclosure can output a recognized text without a presentation form (i.e., a pure text).

[0037] In an implementation scenario, the large language model can perform parameter fine-tuning using sample data, and the sample data can cover at least one of the following types of data: the first sample speech with a presentation form, the first pronunciation sequence of the first sample text with a presentation form, and the first sample speech can be annotated with the first target text that is consistent in content and has a presentation form. For the sake of easy understanding, taking the first sample speech "Please pay twelve yuan" as an example, the first target text it is annotated with can be "Please pay 12 yuan." (i.e., with the presentation forms of "punctuation" and "numbers"); or, taking the first sample speech "Zhang is the new CEO of the company" as an example, the first target text it is annotated with can be "Zhang is the new CEO of the company." (i.e., with the presentation forms of "uppercase and lowercase" and "punctuation"). Of course, the above examples are only several possible examples of the first sample speech and the first target text it is annotated with, and other possible situations of the first sample speech and the first target text it is annotated with will not be exemplified one by one here. In addition, the first sample text can refer to the relevant examples of the foregoing first target text, and the specific content of the first sample text is not limited here and will not be exemplified one by one either. It should be noted that the first pronunciation sequence of the first sample text can include but is not limited to a phoneme sequence, etc., and the specific type of the pronunciation sequence is not limited here. In the above manner, the large language model can perform parameter fine-tuning using sample data, and the sample data can cover at least one of the following types of data: the first sample speech with a presentation form, the first pronunciation sequence of the first sample text with a presentation form, and the first sample speech can be annotated with the first target text that is consistent in content and has a presentation form. Since the number of texts with a presentation form is relatively large compared to the number of speeches annotated with texts with a presentation form, the deficiency in the number of the first sample speeches can be compensated for by the first sample text, thereby assisting the first sample speech in performing parameter fine-tuning on the large language model, enabling the large language model to fully learn and improving its model performance in completing speech recognition and subsequent processing during the autoregressive decoding process.

[0038] In a specific implementation scenario, whether it is the first sample speech with a presentation form or the first sample text with a presentation form, the specific types of the presentation forms it has can either only involve one type (e.g., only involve "uppercase and lowercase", or only involve "punctuation", or only involve "numbers"), or involve multiple types (e.g., simultaneously involve "uppercase and lowercase" and "punctuation", or simultaneously involve "numbers" and "punctuation", or simultaneously involve "uppercase and lowercase", "numbers" and "punctuation"), and this is not limited here.

[0039] In a specific implementation scenario, to unify the display standards of presentation forms, especially for the presentation form of "punctuation marks", it can be preset that for Chinese, it can include commas, periods, question marks, exclamation marks, and semicolons; while for English, it can include commas, periods, question marks, and exclamation marks. In addition, use the corresponding language punctuation in a pure Chinese or pure English context, otherwise use Chinese punctuation marks.

[0040] In a specific implementation scenario, the first sample text with a presentation form can be extracted from channels such as books, periodicals, newspapers, and the Internet. As a possible implementation example, in order to adapt to the voice scenario, the text extracted from the above channels is usually in a written style, so it can be processed into a spoken style, such as adding spoken words, randomly repeating, replacing with synonyms, etc. to enhance the spoken style, and then used as the first sample text. In addition, when constructing the first pronunciation sequence of the first sample text, a text normalization (TN) tool can be used to convert the numbers in the first sample text into corresponding words. For example, in the first sample text "Please pay 12 yuan.", the number "12" needs to be converted into the word "twelve", and in the first sample text "Please call 88886666 for consultation.", the number "88886666" needs to be converted into the word "eight eight eight eight six six six six". Of course, the above examples are only several possible examples of text normalization, and other possible situations will not be listed one by one here.

[0041] In a specific implementation scenario, in order to adapt to the input style of the speech recognition system, the data can also be segmented. For example, for speech data, audio and annotation segmentation can be performed according to the audio segmentation rules; for text data, especially long text data, data segmentation can also be performed. Considering that people tend to pause at punctuation marks when actually speaking, they may not pause or pause in the middle of a normal sentence, and it can also be randomly segmented at the punctuation marks and in the middle of long texts according to a certain probability distribution. In addition, in addition to the above artificial construction and network extraction methods, weakly supervised data can also be constructed through a speech recognition system or a text post-processing model. For example, whisper can be used to obtain weakly supervised data with punctuation marks, numbers, upper and lower cases in unlabeled audio, or a post-processing text model can be used to process the recognition labels to add information such as punctuation marks, numbers, and upper and lower cases. Of course, the above examples are only several possible examples of obtaining sample data in addition to the aforementioned methods, and other possible situations will not be listed one by one again.

[0042] In a specific implementation scenario, after preparing the sample data, the large language model can be fine-tuned accordingly. It should be noted that for fine-tuning the parameters, fine-tuning techniques such as LORA can be used, and the specific technical methods used for parameter fine-tuning are not limited here. Specifically, please refer to Figure 2a , Figure 2a It is a schematic diagram of the process of a parameter fine-tuning embodiment of this application. As Figure 2a shown, feature encoding can be performed based on sample data to obtain sample encoded data, and autoregressive decoding can be performed on the sample encoded features according to the sample instructions based on a large language model to obtain predicted text, and the sample instructions include recognition instructions and post-processing instructions corresponding to the presentation form possessed by the sample data. It should be noted that for the specific meaning of the sample instructions, reference can be made to the relevant description of the foregoing prompt instructions, which will not be elaborated here. On this basis, based on the difference between the predicted text and the expected text of the sample data, a first loss can be obtained, and based on the first loss, the network parameters of the large language model can be adjusted. It should be noted that when the sample data is the first sample speech, the expected text is the first target text, and when the sample data is the first pronunciation sequence, the expected text is the first sample text to which the first pronunciation sequence belongs. In addition, Figure 2a although it is not shown that the model input data of the large language model further includes sample instructions, it does not mean that the model input data of the large language model does not contain sample instructions, but only Figure 2a it is not shown in

[0043] In a specific implementation scenario, in the foregoing parameter fine-tuning implementation example, please continue to refer to Figure 2a , taking the first pronunciation sequence "chi le ma” of the first sample speech "chi le ma” or the first sample text "chi le ma?” as an example, no matter which one is used as the input data, feature extraction can be performed through an encoder to obtain sample encoded features, and after dimension conversion by an adapter, together with the sample instructions (not shown, and in this example, can include recognition instructions and post-processing instructions corresponding to the presentation form "punctuation”), they are used as the model input data of the large language model to perform autoregressive decoding according to the sample instructions through the large language model to obtain predicted text with a presentation form, such as possible situations like "chi le ma?” and "chi le ma.” Of course, optionally, the model input data can also include a sample reference text, and the sample reference text is text data that occurred before the sample data and has a target presentation form. Exemplarily, in Figure 2a the example, the sample reference text can include but is not limited to: "Good morning.”, etc., to assist in recognition especially in long speech stream scenarios, and the specific content of the sample reference text is not limited here.

[0044] In a specific implementation scenario, in the implementation example of the aforementioned parameter fine-tuning, the difference between the predicted text and the expected text of the sample data can be measured by a loss function such as cross entropy to obtain a first loss. For the specific measurement process, please refer to the technical details of the loss function such as cross entropy, which will not be repeated here.

[0045] In a specific implementation scenario, in the implementation example of the aforementioned parameter fine-tuning, in order to further strengthen the attention of the large language model to the characters related to the presentation form during the autoregressive decoding process, before adjusting the network parameters in the aforementioned implementation example, the characters related to the presentation form in the expected text can be selected as the expected characters, and the characters corresponding to the positions of the expected characters in the predicted text can be selected as the predicted characters. Taking the expected text "Have you eaten?" as an example, its presentation form is "punctuation", so "?" can be selected as the expected character, and the character "." corresponding to the position in the predicted text "Have you eaten." can be selected as the predicted character. Of course, the above example is only a possible example in the actual application process, and the specific content of the expected character and the predicted character is not limited here. In addition, the characters related to the presentation form may not be limited to the characters directly corresponding to the presentation form (such as punctuation marks directly corresponding to the presentation form "punctuation", letters directly corresponding to the presentation form "uppercase and lowercase", and Arabic numerals directly corresponding to the presentation form "numbers"), and may further include its adjacent characters (such as directly corresponding characters and the text before and after them). For example, for the text "No. 1, ups and downs, seven or eight or so", the relevant characters in the form of "number" may include "No.", "1", and "name". On this basis, the second loss can be obtained based on the difference between the expected characters and the actual characters. It should be noted that, similar to the measurement method of the aforementioned first loss, the difference between the expected characters and the predicted characters can also be measured based on a loss function such as cross entropy to obtain the second loss. After obtaining the first loss and the second loss, the network parameters of the large language model can be adjusted based on the first loss and the second loss. Exemplarily, the training loss of the large language model can be obtained by weighting the first loss and the second loss, so as to adjust the network parameters of the large language model based on the training loss of the large language model. For example, the weight factor of the first loss can be set to 1, and the weight factor of the second loss can be set to λ, then the training loss L rec It can be expressed as:

[0046] L rec =L ce +λL aux

[0047] In the above formula, L ce represents the first loss, L aux In addition, the second loss L is obtained by using the cross entropy loss function to measure the difference between the expected character and the predicted character. aux For example, the second loss L aux can be expressed as:

[0048]

[0049]

[0050] In the above formula, represents the relevant characters of the presentation form "punctuation", represents the relevant characters of the presentation form "number", represents the relevant characters of the presentation form "case", represents taking the union, represents the relevant characters, represents at the current network parameters of the large language model the probability value that the predicted character is the i-th expected character Therefore, by minimizing the loss function, it is possible to force the predicted character to be as likely as possible to be the expected character.

[0051] In a specific implementation scenario, in the aforementioned example of parameter fine-tuning, the sample data also covers at least one of the following types of data: the second sample speech without a presentation form, the second pronunciation sequence of the second sample text without a presentation form, and the second sample speech is annotated with a second target text with consistent content. Taking the presentation forms including punctuation, numbers, and case as an example, that is, the second target text annotated by the above-mentioned second sample speech without a presentation form is a pure text without punctuation, numbers, and case, and the above-mentioned second sample text without a presentation form is a pure text without punctuation, numbers, and case. It should be noted that in the process of parameter fine-tuning, for the specific process of model training based on the above-mentioned sample data without a presentation form, reference can be made to the specific process of model training based on the sample data with a presentation form. The main difference is that when training the model based on the above-mentioned sample data without a presentation form, the sample prompt instruction can only contain the recognition instruction and not the post-processing instruction, because the post-processing instruction is used to instruct the large language model to perform speech recognition with the presentation form as an additional target. Since the sample data itself does not have a presentation form, the sample prompt instruction can naturally not contain the post-processing instruction. In addition, when training the model based on the above-mentioned sample data without a presentation form, the aforementioned second loss may not need to be measured, because the expected text does not contain the relevant characters of the presentation form, and naturally the second loss does not need to be measured either. Otherwise, reference can be made to the specific process of model training based on the sample data with a presentation form, which will not be elaborated here.

[0052] In an implementation scenario, in order to further enhance the support ability for target presentation forms such as punctuation marks during the speech recognition process in the long speech stream scenario, the sample data can also be accompanied by a sample reference text. The sample reference text can be text data that occurred before the sample data and has the target presentation form. The target presentation form can at least include punctuation marks. And during the parameter fine-tuning process or after the parameter fine-tuning of the large language model, parameter adjustment is performed by executing a masked prediction task based on the relevant characters of the target presentation form in the sample reference text. That is to say, different from the model input data of the large language model in the general scenario shown in Table 1 above, when the target presentation form is punctuation marks, the model input data of the large language model in the long speech stream scenario can refer to Table 2.

[0053] Table 2 Schematic table of an embodiment of the model input data of the large language model in the long speech stream scenario

[0054]

[0055] As shown in Table 2, when the target presentation form is punctuation marks, the model input data of the large language model can include the sample reference text, the sample encoding features, the sample prompt instructions (including the recognition instructions and the post-processing instruction 1, that is, the post-processing instruction corresponding to the presentation form "punctuation marks"), and the decoded characters output at each time step during the autoregressive decoding process. In addition, as mentioned above, the masked prediction task can be executed together during the aforementioned parameter fine-tuning process. At this time, the training loss can not only include the loss shown in the aforementioned formula and its corresponding text, but can also further include the relevant loss of the masked prediction task (specifically, refer to the following relevant description). Or, the masked prediction task can also be executed separately after the aforementioned parameter fine-tuning, which is not limited here. Of course, if there is no long speech stream scenario in the actual application process of speech recognition, the masked prediction task can also be not executed. That is to say, whether to execute the masked prediction task, and when the masked prediction task needs to be executed, specifically in which stage (during the parameter fine-tuning process or after the parameter fine-tuning) the masked prediction task is executed, can both be set according to the actual application needs and are not limited here.

[0056] In a specific implementation scenario, as a possible implementation method, when performing parameter adjustment based on the masked prediction task, specifically, feature encoding can be performed on the sample data to obtain the sample encoding features, and a target attention mask for attention calculation between several model inputs can be configured. And several model inputs include the sample reference text, the sample encoding features, the sample instructions, and the expected text of the sample data. The target attention mask can be configured such that any element in the sample reference text can attend to any element in the sample encoding features after it. That is to say, the sample reference text can attend to the sample encoding features that occur in the future relative to it. As a specific example, please refer to Figure 2b , Figure 2b It is a schematic diagram of the effect of an embodiment of the target attention mask of the present application. As Figure 2b shown, both the length and width of the target attention mask are the total length of elements of several model inputs, and the target attention mask is configured such that any element in the desired text can only focus on any element located before it. Exemplarily, the filled part of the target attention mask can be configured with a value of 1 (i.e., this part of the field of view is visible), in other words, the sample reference text, the sample encoding features, and the sample prompt instructions can all observe all the information before the sample prompt instructions in the attention calculation. The unfilled part can be configured with a value of 0 (i.e., this part of the field of view is invisible), in other words, for <sos>to <eos>For some characters, only historical information can be focused on, while future information cannot be focused on. With this goal, the attention mask can constrain the attention mechanism in autoregressive decoding. Please refer to Figure 2c , Figure 2c is a schematic diagram of the effect of an embodiment of the mask prediction task of this application. As Figure 2c shown, based on this, autoregressive decoding can be performed on the sample reference text and the sample encoding features through the target attention mask according to the sample instructions by the large language model to obtain the predicted text and the first character masked in the sample reference text, and the sample instructions include an identification instruction and a post-processing instruction corresponding to the target display form. Furthermore, based on the difference between the predicted text and the expected text, and the difference between the first character and the related characters (such as Figure 2c shown by the shaded area in the sample reference text as shown, for example, when the target display form is "punctuation", the punctuation at the end of the sample reference text can be masked, and a certain proportion of characters other than this can also be masked), the network parameters of the large language model can be adjusted. It should be noted that when the sample data is the first sample speech, the expected text is the first target text, and when the sample data is the first pronunciation sequence, the expected text is the first sample text to which the first pronunciation sequence belongs. In addition, the difference between the predicted text and the expected text can refer to the relevant description of the aforementioned parameter fine-tuning and will not be elaborated here. The difference between the first character and the related characters can also be measured by a loss function such as cross-entropy, which is not limited here. Of course, in this process, the measurement method of the second loss in the aforementioned parameter fine-tuning part can also be referred to to measure the loss of the difference between the expected character and the predicted character, so as to adjust the network parameters of the large language model by combining the three aspects of differences. Specifically, it can refer to the relevant description in the aforementioned parameter fine-tuning and will not be elaborated here. In the above manner, by configuring the target attention mask, the large language model can be forced to focus on future sample encoding features for the sample reference text during the autoregressive decoding process, which helps to improve the accuracy of predicting the masked characters in the sample reference text, so as to achieve error correction in the long speech stream scenario. And since the mask prediction task and the first character decoding are completed in the same step of calculation, there is no efficiency loss.

[0057] In a specific implementation scenario, please refer to Figure 2d , Figure 2d is a schematic diagram of the effect of another embodiment of the mask prediction task of this application. As Figure 2d As shown, as another possible implementation, different from the foregoing implementation, when adjusting parameters based on the mask prediction task, specifically, feature encoding can also be performed on the sample data to obtain sample encoded features, and based on the large language model, autoregressive decoding of the sample reference text and the sample encoded features is performed through the autoregressive attention mask according to the sample instruction to obtain the predicted text and the second masked character in the sample reference text, and the sample instruction includes an identification instruction and a post-processing instruction corresponding to the target display form. Then, based on the difference between the predicted text and the expected text of the sample data, and the difference between the second character and the relevant characters (such as Figure 2d as shown in the shaded area of the sample reference text. For example, when the target display form is "punctuation", the punctuation at the end of the sample reference text can be masked, and a certain proportion of characters other than this can be further masked), the network parameters of the large language model are adjusted. It should be noted that when the sample data is the first sample voice, the expected text is the first target text, and when the sample data is the first pronunciation sequence, the expected text is the first sample text to which the first pronunciation sequence belongs. As a special example, Figure 2d in <ask>It can be the predicted character (i.e., the second character) output by the large language model for the obscured punctuation at the end of the sample reference text when the target presentation form is "punctuation". When the performance of the large language model is excellent enough, if the end of the sample reference text is obscured, then <ask>The predicted characters output by the large language model are the relevant characters that are masked. If there are no masked characters at the end of the sample reference text, then <ask>The predicted character output by the large language model is the character itself at the end, and without a sample reference text, <ask>The predicted character output by the large language model is empty. Further, the loss here can also be measured by a loss function such as cross entropy. For example, the second character and the related character can be measured by the cross entropy loss function to obtain the corresponding loss L ask :

[0058]

[0059] In the above formula, Indicates the current network parameters in the large language model The second character that is masked is predicted to be the relevant character , so by minimizing the loss function, the possibility that the second character is a related character can be forced to be as large as possible. For the rest of the training loss, please refer to the aforementioned description related to the loss calculation, which will not be repeated here. In addition, the similarities between this embodiment and the previous embodiment can be referred to the previous embodiment, which will not be repeated here. The main difference is that this embodiment adopts an autoregressive attention mask, while the previous embodiment customizes the target attention mask. Different from the previous embodiment that customizes the target attention mask, the autoregressive attention mask can only pay attention to historical information, such as Figure 2b Taking the mask shown in the figure as an example, the difference between the autoregressive attention mask and the autoregressive attention mask is that only the value of the part below the main diagonal of the autoregressive attention mask is configured as 1 (that is, this part of the field of view is visible), and the rest of the value is configured as (that is, this part of the field of view is invisible). On the one hand, compared with configuring a custom target attention mask, it can avoid conflicts with the attention mask of the large language model as much as possible and reduce modifications to the base model. On the other hand, under the autoregressive attention mask mechanism, although the sample reference text part cannot pay attention to the sample encoding features relative to its future, it can still <sos>So far, the sample reference text and sample encoding features have been obtained, so the semantic conditions for correcting errors in the sample reference text are met.

[0060] The above scheme performs feature encoding based on the speech to be recognized to obtain the encoded features, and performs autoregressive decoding on the encoded features based on the large language model according to the prompt instruction to obtain the recognized text with a display form of the speech to be recognized, and the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform speech recognition with the additional goal of having a display form, and various post-processing instructions correspond to different display forms, so on the one hand, by inputting the post-processing instruction together with the recognition instruction, forcing the large language model to perform post-processing together with the speech recognition during the autoregressive decoding process, it is possible to ensure that the recognized text has a display form as much as possible, and on the other hand, because the speech recognition and its post-processing are completed together by the large language model during the autoregressive decoding process, that is, there is no need to successively realize speech transcription and post-processing through a cascade structure, which can shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition. Therefore, it is possible to shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition under the premise of ensuring that the recognized text has a display form as much as possible. In addition, since various post-processing instructions correspond to different display forms, different post-processing instructions can be configured specifically according to different application scenarios during the speech recognition process, thereby improving the convenience of differentiated processing of different application scenarios.

[0061] See also Figure 3 , Figure 3 It is a schematic diagram of the framework of an embodiment of the speech recognition device of the present application. The speech recognition device 30 includes: a feature encoding module 31 and an autoregressive decoding module 32, wherein the feature encoding module 31 is used to perform feature encoding based on the speech to be recognized to obtain the encoded features; the autoregressive decoding module 32 is used to perform autoregressive decoding on the encoded features according to the prompt instructions based on the large language model to obtain the recognition text of the speech to be recognized in the display form; wherein the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform speech recognition with the display form as an additional target, and various post-processing instructions correspond to different display forms.

[0062] In the above scheme, the speech recognition device 30 performs feature encoding based on the speech to be recognized to obtain the encoding feature, and performs autoregressive decoding on the encoding feature based on the large language model according to the prompt instruction to obtain the recognition text of the speech to be recognized with a display form, and the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform speech recognition with a display form as an additional goal, and various post-processing instructions correspond to different display forms, so on the one hand, by inputting the post-processing instruction together with the recognition instruction, forcing the large language model to perform post-processing together with the speech recognition in the autoregressive decoding process, it is possible to ensure that the recognized text has a display form as much as possible, and on the other hand, because the speech recognition and its post-processing are completed together by the large language model in the autoregressive decoding process, that is, there is no need to successively realize speech transcription and post-processing through a cascade structure, which can shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition. Therefore, it is possible to shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition under the premise of ensuring that the recognized text has a display form as much as possible. In addition, since various post-processing instructions correspond to different display forms, different post-processing instructions can be configured specifically according to different application scenarios during the speech recognition process, thereby improving the convenience of differentiated processing of different application scenarios.

[0063] In some disclosed embodiments, the large language model uses sample data for parameter fine-tuning, and the sample data covers at least one of the following types of data: a first sample speech having a presentation form, a first pronunciation sequence of a first sample text having a presentation form, and the first sample speech is annotated with a first target text having consistent content and a presentation form.

[0064] In some disclosed embodiments, the speech recognition device 30 includes a first sample encoding module, which is used to perform feature encoding based on sample data to obtain sample encoding features; the speech recognition device 30 includes a first sample decoding module, which is used to perform autoregressive decoding on the sample encoding features according to sample instructions based on a large language model to obtain predicted text; wherein the sample instructions include recognition instructions and post-processing instructions corresponding to the presentation form of the sample data; the speech recognition device 30 includes a first parameter adjustment module, which is used to obtain a first loss based on the difference between the predicted text and the expected text of the sample data, and adjust the network parameters of the large language model based on the first loss; wherein, when the sample data is the first sample speech, the expected text is the first target text, and when the sample data is the first pronunciation sequence, the expected text is the first sample text belonging to the first pronunciation sequence.

[0065] In some disclosed embodiments, the speech recognition device 30 includes a character selection module, which is configured to select relevant characters in the expected text in a presented form as expected characters, and select characters in the predicted text corresponding to the positions of the expected characters as predicted characters; the speech recognition device 30 includes a loss metric module, which is configured to obtain a second loss based on the difference between the expected characters and the predicted characters; the first parameter adjustment module is specifically configured to adjust the network parameters of the large language model based on the first loss and the second loss.

[0066] In some disclosed embodiments, the sample data is feature-encoded by an encoder compatible with both speech and text modalities, and during the process of parameter fine-tuning, the network parameters of the encoder are frozen; and / or, the sample data also covers at least one of the following types of data: a second sample speech without a presented form, a second pronunciation sequence of a second sample text without a presented form, and the second sample speech is annotated with a second target text with consistent content.

[0067] In some disclosed embodiments, the sample data is also accompanied by a sample reference text, which is text data that occurred before the sample data and has a target presented form. The target presented form includes at least punctuation marks, and during or after the process of parameter fine-tuning, the large language model performs a masked prediction task based on the relevant characters in the target presented form in the sample reference text to adjust the parameters.

[0068] In some disclosed embodiments, the speech recognition device 30 includes a second sample encoding module, which is configured to perform feature encoding on the sample data to obtain sample encoding features; the speech recognition device 30 includes a mask configuration module, which is configured to configure a target attention mask for attention calculation between several model inputs; wherein, the several model inputs include the sample reference text, the sample encoding features, the sample instruction, and the expected text of the sample data, and the target attention mask is configured such that any element in the sample reference text can attend to any element in the sample encoding features after it; the speech recognition device 30 includes a second sample decoding module, which is configured to perform autoregressive decoding on the sample reference text and the sample encoding features by the large language model according to the sample instruction through the target attention mask to obtain a predicted text and the first character masked in the sample reference text; wherein, the sample instruction includes a recognition instruction and a post-processing instruction corresponding to the target presented form; the speech recognition device 30 includes a second parameter adjustment module, which is configured to adjust the network parameters of the large language model based on the difference between the predicted text and the expected text, and the difference between the first character and the relevant characters; wherein, when the sample data is the first sample speech, the expected text is the first target text, and when the sample data is the first pronunciation sequence, the expected text is the first sample text to which the first pronunciation sequence belongs.

[0069] In some disclosed embodiments, the length and width of the target attention mask are both the total length of the elements of several model inputs, and the target attention mask is configured such that any element in the desired text can only attend to any element before it.

[0070] In some disclosed embodiments, the speech recognition device 30 includes a third sample encoding module for performing feature encoding on the sample data to obtain sample encoding features; the speech recognition device 30 includes a third sample decoding module for performing autoregressive decoding on the sample reference text and the sample encoding features through an autoregressive attention mask based on a large language model according to the sample instructions to obtain the predicted text and the second character masked in the sample reference text; wherein the sample instructions include an identification instruction and a post-processing instruction corresponding to the target presentation form; the speech recognition device 30 includes a third parameter adjustment module for adjusting the network parameters of the large language model based on the difference between the predicted text and the desired text of the sample data, and the difference between the second character and the relevant character; wherein, when the sample data is the first sample speech, the desired text is the first target text, and when the sample data is the first pronunciation sequence, the desired text is the first sample text to which the first pronunciation sequence belongs.

[0071] In some disclosed embodiments, the presentation form includes at least one of: upper and lower case, numbers, punctuation; and / or, when the feature dimension after performing feature encoding is different from the feature dimension of the embedding layer in the large language model, perform dimension conversion through an adapter after performing feature encoding and before performing autoregressive decoding to be consistent with the feature dimension of the embedding layer, and the adapter is fine-tuned for parameters together with the large language model; and / or, during the process of performing autoregressive decoding, the encoded features are also concatenated with a reference text, and the reference text is the dialogue text that occurred before the speech to be recognized.

[0072] Please refer to Figure 4 , Figure 4 is a schematic framework diagram of an embodiment of the electronic device of the present application. The electronic device 40 at least includes a memory 41 and a processor 42 that are coupled to each other. The memory 41 stores at least program instructions, and the processor 42 is configured to execute the program instructions to implement the steps in any of the above-mentioned speech recognition method embodiments. Specifically, reference may be made to the foregoing disclosed embodiments, which will not be elaborated herein. It should be noted that the electronic device 40 may include, but is not limited to, devices such as a translator, a learning machine, an e-reader, a headset, a mouse, a smart large screen, etc. The specific type of the electronic device 40 is not limited herein.

[0073] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above-mentioned speech recognition method embodiments. The processor 42 can also be called a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 42 can be implemented by an integrated circuit chip.

[0074] In the above scheme, the electronic device 40 performs feature encoding based on the speech to be recognized to obtain the encoding feature, and performs autoregressive decoding on the encoding feature based on the large language model according to the prompt instruction to obtain the recognition text of the speech to be recognized with a display form, and the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform speech recognition with a display form as an additional goal, and various post-processing instructions correspond to different display forms, so on the one hand, by inputting the post-processing instruction together with the recognition instruction, forcing the large language model to perform post-processing together with the speech recognition during the autoregressive decoding process, it is possible to ensure that the recognized text has a display form as much as possible, and on the other hand, because the speech recognition and its post-processing are completed together by the large language model during the autoregressive decoding process, that is, there is no need to successively realize speech transcription and post-processing through a cascade structure, which can shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition. Therefore, it is possible to shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition under the premise of ensuring that the recognized text has a display form as much as possible. In addition, since various post-processing instructions correspond to different display forms, different post-processing instructions can be configured specifically according to different application scenarios during the speech recognition process, thereby improving the convenience of differentiated processing of different application scenarios.

[0075] See also Figure 5 , Figure 5 The computer-readable storage medium 50 stores program instructions 51 that can be executed by a processor, and the program instructions 51 are used to implement the steps in any of the above-mentioned speech recognition method embodiments.

[0076] In the above scheme, the computer-readable storage medium 50 performs feature encoding based on the speech to be recognized to obtain the encoding feature, and performs autoregressive decoding on the encoding feature based on the large language model according to the prompt instruction to obtain the recognition text of the speech to be recognized with a display form, and the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform speech recognition with a display form as an additional goal, and various post-processing instructions correspond to different display forms, so on the one hand, by inputting the post-processing instruction together with the recognition instruction, forcing the large language model to perform post-processing together with the speech recognition during the autoregressive decoding process, it is possible to ensure that the recognized text has a display form as much as possible, and on the other hand, because the speech recognition and its post-processing are completed together by the large language model during the autoregressive decoding process, that is, there is no need to successively realize speech transcription and post-processing through a cascade structure, which can shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition. Therefore, it is possible to shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition under the premise of ensuring that the recognized text has a display form as much as possible. In addition, since various post-processing instructions correspond to different display forms, different post-processing instructions can be configured specifically according to different application scenarios during the speech recognition process, thereby improving the convenience of differentiated processing of different application scenarios.

[0077] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0078] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.

[0079] In several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0080] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0081] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0082] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of this application. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks or optical discs and other various media that can store program codes.

[0083] If the technical solution of this application involves personal information, before the product applying the technical solution of this application processes personal information, it has clearly informed the personal information processing rules and obtained the individual's independent consent. If the technical solution of this application involves sensitive personal information, before the product applying the technical solution of this application processes sensitive personal information, it has obtained the individual's separate consent and at the same time meets the requirements of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If an individual voluntarily enters the collection scope, it is regarded as consenting to the collection of their personal information; or on the device for personal information processing, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information by themselves, etc.; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.< / sos> < / ask> < / ask> < / ask> < / ask> < / eos> < / sos> < / eos> < / sos>

Claims

1. A speech recognition method, characterized in that, Including: Performing feature encoding on the speech to be recognized to obtain encoded features; Performing autoregressive decoding on the encoded features based on a large language model according to a prompt instruction to obtain a recognized text in the presentation form of the speech to be recognized; wherein, the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, the post-processing instruction is used to instruct the large language model to perform the speech recognition with the presentation form as an additional target, each of the post-processing instructions corresponds to a different presentation form, the large language model is fine-tuned with sample data attached with a sample reference text, the sample reference text is text data that occurs before the sample data and has the target presentation form, the large language model performs a mask prediction task based on relevant characters of the target presentation form in the sample reference text to adjust parameters, and when the attention calculation between several model inputs is implemented through a target attention mask during the mask prediction task, the target attention mask is configured such that any element in the sample reference text can attend to any element in the sample encoded features after it, and only any element in the expected text of the sample data can attend to any element before it, and the sample encoded features are obtained by performing feature encoding on the sample data.

2. The method according to claim 1, wherein The sample data covers at least one of the following types of data: a first sample speech with the presentation form, a first pronunciation sequence of a first sample text with the presentation form, and the first sample speech is annotated with a first target text that is consistent in content and has the presentation form.

3. The method according to claim 2, wherein The steps of the parameter fine-tuning include: Performing feature encoding on the sample data to obtain sample encoded features; Performing autoregressive decoding on the sample encoded features based on the large language model according to a sample instruction to obtain a predicted text; wherein, the sample instruction includes the recognition instruction and a post-processing instruction corresponding to the presentation form of the sample data; Based on the difference between the predicted text and the expected text of the sample data, obtaining a first loss, and adjusting the network parameters of the large language model based on the first loss; Wherein, when the sample data is the first sample speech, the expected text is the first target text, and when the sample data is the first pronunciation sequence, the expected text is the first sample text to which the first pronunciation sequence belongs.

4. The method according to claim 3, characterized in that, Before adjusting the network parameters of the large language model based on the first loss, the method further includes: Selecting relevant characters of the presentation form in the expected text as expected characters, and selecting characters corresponding to the positions of the expected characters in the predicted text as predicted characters; Based on the difference between the expected characters and the predicted characters, obtaining a second loss; Adjusting the network parameters of the large language model based on the first loss includes: Adjusting the network parameters of the large language model based on the first loss and the second loss.

5. The method according to claim 2, wherein The sample data is feature-encoded by an encoder compatible with both speech and text modalities, and during the parameter fine-tuning process, the network parameters of the encoder are frozen; and / or, the sample data further covers at least one of the following types of data: a second sample speech without the presentation form, a second pronunciation sequence of a second sample text without the presentation form, and the second sample speech is labeled with a second target text with consistent content.

6. The method according to claim 2, characterized in that, The target presentation form includes at least punctuation, and the large language model performs a masked prediction task during or after the parameter fine-tuning process to adjust the parameters.

7. The method according to claim 6, wherein When the attention calculation between several model inputs during the masked prediction task is achieved through a target attention mask, the step of adjusting the parameters based on the masked prediction task includes: Performing feature encoding on the sample data to obtain sample encoded features, and configuring a target attention mask for attention calculation between several model inputs; wherein, the several model inputs include the sample reference text, the sample encoded features, the sample instruction, and the expected text; Performing autoregressive decoding on the sample reference text and the sample encoded features by the large language model according to the sample instruction through the target attention mask to obtain a predicted text and a first character masked in the sample reference text; wherein, the sample instruction includes the recognition instruction and a post-processing instruction corresponding to the target presentation form; Adjusting the network parameters of the large language model based on the difference between the predicted text and the expected text, and the difference between the first character and the relevant character; Wherein, when the sample data is the first sample speech, the expected text is the first target text, and when the sample data is the first pronunciation sequence, the expected text is the first sample text to which the first pronunciation sequence belongs.

8. The method according to claim 7, wherein The length and width of the target attention mask are both the total length of the elements of the several model inputs.

9. The method according to claim 6, characterized in that, When the attention calculation between several model inputs during the masked prediction task is achieved through an autoregressive attention mask, the step of adjusting the parameters based on the masked prediction task includes: Performing feature encoding on the sample data to obtain sample encoded features; Performing autoregressive decoding on the sample reference text and the sample encoded features by the large language model according to the sample instruction through the autoregressive attention mask to obtain a predicted text and a second character masked in the sample reference text; wherein, the sample instruction includes the recognition instruction and a post-processing instruction corresponding to the target presentation form; Adjusting the network parameters of the large language model based on the difference between the predicted text and the expected text of the sample data, and the difference between the second character and the relevant character; Wherein, when the sample data is the first sample speech, the expected text is the first target text, and when the sample data is the first pronunciation sequence, the expected text is the first sample text to which the first pronunciation sequence belongs.

10. The method according to any one of claims 1 to 9, characterized in that, The presentation form includes at least one of: capitalization, numbers, punctuation; And / or, when the feature dimension after performing the feature encoding is different from the feature dimension of the embedding layer in the large language model, perform dimension conversion through an adapter after performing the feature encoding and before performing the autoregressive decoding to be consistent with the feature dimension of the embedding layer, and the adapter is fine-tuned together with the large language model for parameters; And / or, during the process of performing the autoregressive decoding, the encoded feature is also concatenated with a reference text, and the reference text is the dialogue text that occurred before the speech to be recognized.

11. A voice recognition device, characterized in that, Comprising: A feature encoding module, configured to perform feature encoding based on the speech to be recognized to obtain an encoded feature; An autoregressive decoding module, configured to perform autoregressive decoding on the encoded feature according to a prompt instruction based on a large language model to obtain a recognized text in a presentation form of the speech to be recognized; wherein, the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, the post-processing instruction is used to instruct the large language model to perform the speech recognition with the presentation form as an additional target, and each of the post-processing instructions corresponds to a different presentation form. The large language model is fine-tuned for parameters using sample data with a sample reference text, and the sample reference text is text data that occurred before the sample data and has a target presentation form. The large language model performs a mask prediction task based on relevant characters of the target presentation form in the sample reference text to adjust parameters, and when the attention calculation between several model inputs during the mask prediction task is achieved through a target attention mask, the target attention mask is configured as follows: any element in the sample reference text can attend to any element in the sample encoded feature after it, and only any element in the expected text of the sample data can attend to any element before it. The sample encoded feature is obtained by feature encoding the sample data.

12. An electronic device, characterized in that, At least comprising a memory and a processor coupled to each other, at least program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the speech recognition method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, Stored with program instructions that can be run by a processor, and the program instructions are used to implement the speech recognition method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method and device for intention classification based on speech recognition result

    CN111177324A

  • Speech recognition method and system, medium, computer equipment, terminal and application

    CN112712804A

  • Speech recognition method and device, equipment and storage medium

    CN117153152A

  • Interaction processing method, device and equipment, man-machine interaction system and program product

    CN118918889A