Speech recognition method, related device, equipment and storage medium
By using a large language model for feature encoding and autoregressive decoding in speech recognition, and combining recognition and post-processing instructions, post-processing is directly performed in the decoding process, the problems of extended response time, increased computational burden and low output accuracy in traditional speech recognition technology are solved, and a faster and more efficient speech recognition process is achieved.
Patent Information
- Application Number
- CN202510454672.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Traditional speech recognition technology has problems such as extended response time, increased computational burden and accumulated errors in post-processing of transcription texts, resulting in low output accuracy of the recognition text.
The speech recognition method based on the large language model is adopted, and through feature encoding and autoregressive decoding, combined with recognition instructions and post-processing instructions, speech recognition and post-processing are directly performed in the autoregressive decoding process to generate recognition text with presentation form.
On the premise of ensuring that the recognition text has a display form, the response time of speech recognition is shortened, the calculation burden is reduced, and the output accuracy is improved. At the same time, through the configuration of different post-processing instructions, it adapts to different application scenarios, improving the convenience of differentiated processing.
Smart Images

Figure CN119993163A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech processing technology, and in particular to a speech recognition method and related devices, equipment and storage media. Background Art
[0002] Automatic Speech Recognition (ASR) technology can extract text information from speech and transcribe it into text. It is widely used in many fields such as smart customer service, smart office, smart home, in-vehicle control, and simultaneous voice interpretation.
[0003] It should be noted that the transcribed text of traditional speech recognition technology is a pure text result without punctuation, numbers, uppercase and lowercase letters, etc. Since it is not convenient to read or perform downstream tasks (such as semantic understanding, etc.), in the speech recognition system, the traditional speech recognition model is generally cascaded for the post-processing model of Natural Language Process (NLP) to post-process the transcribed text presented as a pure text result into a recognized text with a display form. However, this cascade structure will undoubtedly prolong the response time and increase the computational burden, and because of the error accumulation problem of the cascade structure, it is easy to degrade the final recognized text. In view of this, how to shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition while ensuring that the recognized text has a display form as much as possible has become an urgent problem to be solved. Summary of the invention
[0004] The main technical problem solved by the present application is to provide a speech recognition method and related devices, equipment and storage media, which can shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition while ensuring that the recognized text has a display form as much as possible.
[0005] In order to solve the above technical problems, the first aspect of the present application provides a speech recognition method, including: performing feature encoding based on the speech to be recognized to obtain encoded features; performing autoregressive decoding on the encoded features according to prompt instructions based on a large language model to obtain a recognition text with a display form of the speech to be recognized; wherein the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform speech recognition with a display form as an additional goal, and various post-processing instructions correspond to different display forms.
[0006] In order to solve the above-mentioned technical problems, the second aspect of the present application provides a speech recognition device, including: a feature encoding module and an autoregressive decoding module, the feature encoding module is used to perform feature encoding based on the speech to be recognized to obtain the encoded features; the autoregressive decoding module is used to perform autoregressive decoding on the encoded features according to the prompt instructions based on the large language model to obtain the recognition text of the speech to be recognized with a display form; wherein the prompt instructions include recognition instructions and at least one post-processing instruction, the recognition instructions are used to instruct the large language model to perform speech recognition, and the post-processing instructions are used to instruct the large language model to perform speech recognition with a display form as an additional goal, and various post-processing instructions correspond to different display forms.
[0007] In order to solve the above technical problems, the third aspect of the present application provides an electronic device, which at least includes a memory and a processor coupled to each other, the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the speech recognition method in the above first aspect.
[0008] In order to solve the above technical problems, the fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be executed by a processor, and the program instructions are used to implement the speech recognition method of the first aspect.
[0009] The above scheme performs feature encoding based on the speech to be recognized to obtain the encoded features, and performs autoregressive decoding on the encoded features based on the large language model according to the prompt instruction to obtain the recognized text with a display form of the speech to be recognized, and the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform speech recognition with the additional goal of having a display form, and various post-processing instructions correspond to different display forms, so on the one hand, by inputting the post-processing instruction together with the recognition instruction, forcing the large language model to perform post-processing together with the speech recognition during the autoregressive decoding process, it is possible to ensure that the recognized text has a display form as much as possible, and on the other hand, because the speech recognition and its post-processing are completed together by the large language model during the autoregressive decoding process, that is, there is no need to successively realize speech transcription and post-processing through a cascade structure, which can shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition. Therefore, it is possible to shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition under the premise of ensuring that the recognized text has a display form as much as possible.
[0010] In addition, since various post-processing instructions correspond to different display forms, different post-processing instructions can be configured specifically according to different application scenarios during the speech recognition process, thereby improving the convenience of differentiated processing of different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is a flowchart of an embodiment of the speech recognition method of the present application; Figure 2a This is a schematic diagram of the process of parameter fine-tuning in an embodiment of the present application; Figure 2b This is a schematic diagram of the effect of an embodiment of the target attention mask of the present application; Figure 2c This is a schematic diagram of the effect of an embodiment of the mask prediction task of the present application; Figure 2d This is a schematic diagram of the effect of another embodiment of the mask prediction task of the present application; Figure 3 It is a schematic diagram of the framework of an embodiment of the speech recognition device of the present application; Figure 4 It is a schematic diagram of the framework of an embodiment of the electronic device of the present application; Figure 5 It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0012] The scheme of the embodiment of the present application is described in detail below in conjunction with the drawings of the specification.
[0013] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.
[0014] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the fragment " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. In addition, "many" in this article means two or more than two.
[0015] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of the speech recognition method of the present application. Specifically, it may include the following steps: Step S11: Perform feature encoding based on the speech to be recognized to obtain encoding features.
[0016] In one implementation scenario, the speech to be recognized may be streaming speech, that is, the speech to be recognized may be collected in real time while the speaker is speaking; or, the speech to be recognized may be non-streaming speech, such as recording may be started when the speaker starts speaking and may be stopped when the speaker stops speaking, so that a speech recording may be obtained as the speech to be recognized. Of course, the above examples are only several possible examples of the speech to be recognized, and the method of obtaining the speech to be recognized is not limited here, and examples are not given one by one.
[0017] In one implementation scenario, as a possible implementation method, feature encoding can be implemented by an encoder. Exemplarily, the encoder can include but is not limited to network structures such as Transformer and Conformer, and the network structure of the encoder is not limited here.
[0018] In another implementation scenario, as another possible implementation method, feature coding can also be implemented by an encoder compatible with both speech and text modes, so that in the training process, not only sample speech data can be used for model training, but also widely available sample text data can be used for auxiliary training. The specific training process can refer to the following related description, which will not be repeated here. Exemplarily, an encoder compatible with both speech and text modes may include but is not limited to: a Conformer network structure composed of a number of Conformer blocks, each Conformer block may include a series of multi-head self-attention, deep convolution and feedforward layers, and the technical details of the Conformer network structure can be specifically referred to, and the Conformer network structure will not be repeated here. Exemplarily, speech and text can share the last 8 layers of the Conformer network structure. Of course, the above example is only a possible example of an encoder, and the specific structure of the encoder compatible with both speech and text modes is not limited here, and examples are not given one by one.
[0019] In another implementation scenario, as another possible implementation example, the feature dimension after feature encoding may be different from the feature dimension of the embedding layer in the subsequent large language model. In this case, after feature encoding and before the subsequent autoregressive decoding using the large language model, the adapter may be used to perform dimension conversion (i.e., the output feature after feature encoding is performed is dimensionally converted) to be consistent with the feature dimension of the embedding layer. It should be noted that the adapter may include but is not limited to: network structures such as convolutional layers and fully connected layers, and the network structure of the adapter is not limited here. In addition, the adapter can perform parameter fine-tuning together with the large language model. For details, please refer to the subsequent training-related description, which will not be repeated here.
[0020] It should be noted that, as a possible example, in order to facilitate the encoder to perform feature encoding, before using the encoder to perform feature encoding on the speech to be recognized, the acoustic features of the speech to be recognized can be extracted first, such as acoustic features that can include but are not limited to FBank, etc., and the specific types of acoustic features are not limited here. On this basis, the acoustic features can be feature encoded based on the encoder to obtain encoded features.
[0021] Step S12: Based on the large language model, the encoded features are autoregressively decoded according to the prompt instructions to obtain a recognition text with a display form of the speech to be recognized.
[0022] In the disclosed embodiment, the prompt instruction may include a recognition instruction and at least one post-processing instruction. The recognition instruction may be used to instruct the large language model to perform speech recognition. The post-processing instruction may be used to instruct the large language model to perform speech recognition with a presentation form as an additional goal, and various post-processing instructions correspond to different presentation forms. It should be noted that the large language model may include but is not limited to open source large models such as Llama, Bloom, Spark 13B, GLM6B, etc.; or, the large language model may also include but is not limited to fine-tuning parameters based on specific corpus; or, the large language model may also include but is not limited to a custom large model, and the specific source of the large language model is not limited here. In addition, during the autoregressive decoding process, during the first decoding, it is usually defaulted to input the starting character representing the start of decoding together with the model input data such as the encoding features (such as, <sos>When decoding, the predicted character of the previous decoding output is input together with the model input data such as the encoding features, so that the predicted character of the current decoding output can be obtained through large language decoding, and so on, and the cycle repeats until the predicted character of a certain decoding output is the end character representing the end of decoding (such as, <eos>), the predicted characters output by the previous decoding can be sequentially combined to complete the speech recognition. Of course, the above description of autoregressive decoding is only a brief description of autoregressive decoding. For details, please refer to the technical details of autoregressive decoding, which will not be repeated here.
[0023] In one implementation scenario, the recognition instruction may be in the form of a natural language description text to instruct the large language model to perform speech recognition. For example, the recognition instruction may include but is not limited to the following content: "Please perform speech recognition based on the encoding features of the input speech", etc. The specific content of the recognition instruction is not limited here.
[0024] In an implementation scenario, similar to the recognition instruction, the post-processing instruction can also be embodied in the form of a natural language description text, and instruct the large language model to perform speech recognition with the display form as an additional target. Taking the display form including "uppercase" as an example, the post-processing instruction corresponding to the display form "uppercase" may include but is not limited to: "Please distinguish uppercase and lowercase during the recognition process", etc., and the specific content of the post-processing instruction corresponding to the display form "uppercase" is not limited here; or, taking the display form including "numbers" as an example, the post-processing instruction corresponding to the display form "numbers" may include but is not limited to: "Please distinguish Arabic numerals from ordinary text during the recognition process", etc., and the specific content of the post-processing instruction corresponding to the display form "numbers" is not limited here; or, taking the display form including "punctuation" as an example, the post-processing instruction corresponding to the display form "punctuation" may include but is not limited to: "Please add appropriate punctuation marks in the text according to semantics during the recognition process", etc., and the specific content of the post-processing instruction corresponding to the display form "punctuation" is not limited here. Of course, the above examples are only several possible examples of post-processing instructions when the display form includes uppercase and lowercase letters, numbers, punctuation marks, etc., and other possible contents and specific contents of prompt instructions when the display form includes other circumstances are not limited here. For example, the display form may also include but is not limited to measurement units (such as US dollars and other currency units), etc., and no more examples are given here.
[0025] In an implementation scenario, as a possible implementation method, please refer to Table 1, which is a schematic table of an embodiment of the model input data of a large language model. As shown in Table 1, optionally, in the process of performing autoregressive decoding, the coding features can also be spliced with reference text, and the reference text can specifically be a conversation text before the speech to be recognized, so as to further refer to the previous conversation text during the current speech recognition (for example, the conversation text can be a recognition text with a display form obtained by speech recognition of the historical speech before the speech to be recognized), which helps to improve the speech recognition performance in the long speech stream scenario. In addition, in the case where the coding features are spliced with reference text, as a possible example, the large language model can also output the predicted characters of the display form-related characters in the reference text, so as to correct the relevant characters of the display form in the reference text according to the predicted characters, obtain the corrected text of the reference text, and overwrite the reference text with the corrected text. For example, the speech to be recognized is "what about you", and the previous conversation text can be "I plan to go out and play next week." as the reference text. After autoregressive decoding, the recognition text "what about you?" with the presentation form of the speech to be recognized and the predicted characters "," of the related characters "." in the presentation form of the reference text can be obtained. Based on this, the reference text can be corrected to obtain the corrected text "I plan to go out and play next week,". After that, the corrected text and the recognition text with the presentation form of the speech to be recognized can be spliced to obtain the latest recognition result "I plan to go out and play next week, what about you?". This cycle can be repeated, and in the speech recognition process (especially in long speech flow scenarios), the previous conversation text can be corrected each time the recognition is performed, and finally a recognition result that is as accurate as possible and has a presentation form can be obtained.
[0026] Table 1 Schematic table of an embodiment of model input data of a large language model
[0027] It should be noted that, as shown in Table 1, when the prompt instruction does not include a post-processing instruction, the speech recognition method of the embodiment of the present disclosure can output a recognition text without a display form (ie, a plain text).
[0028] In an implementation scenario, the large language model can use sample data for parameter fine-tuning, and the sample data can cover at least one of the following types of data: a first sample voice with a display form, a first pronunciation sequence of a first sample text with a display form, and the first sample voice can be annotated with a first target text with consistent content and a display form. For ease of understanding, taking the first sample voice "Please pay twelve yuan" as an example, the first target text annotated therewith can be "Please pay 12 yuan." (i.e., with the display form of "punctuation" and "number"); or, taking the first sample voice "Zhang is the new CEO of the company" as an example, the first target text annotated therewith can be "Zhang is the new CEO of the company." (i.e., with the display form of "uppercase" and "punctuation"). Of course, the above examples are only several possible examples of the first sample voice and the first target text annotated therewith. Other possible situations of the first sample voice and the first target text annotated therewith will not be given one by one. In addition, the first sample text can refer to the relevant examples of the first target text mentioned above. The specific content of the first sample text is not limited here, and no examples are given one by one. It should be noted that the first pronunciation sequence of the first sample text may include but is not limited to a phoneme sequence, etc., and the specific type of the pronunciation sequence is not limited here. In the above manner, the large language model can use sample data to fine-tune parameters, and the sample data can cover at least one of the following types of data: a first sample voice with a display form, a first pronunciation sequence of a first sample text with a display form, and the first sample voice can be annotated with a first target text with consistent content and a display form. Since the number of texts with a display form is larger than the number of voices annotated with texts with a display form, the first sample text can be used to make up for the defect of insufficient number of first sample voices, thereby assisting the first sample voice to fine-tune the parameters of the large language model, so that the large language model can be fully learned and improve its model performance in completing speech recognition and post-processing during autoregressive decoding.
[0029] In a specific implementation scenario, whether it is a first sample voice with a presentation form or a first sample text with a presentation form, the specific type of presentation form it has may involve only one type (e.g., only "upper and lower case", or only "punctuation", or only "numbers"), or may involve multiple types (e.g., involving both "upper and lower case" and "punctuation", or both "numbers" and "punctuation", or both "upper and lower case", "numbers" and "punctuation"), and is not limited here.
[0030] In a specific implementation scenario, in order to unify the display standards of the display form, especially the display form "punctuation", it can be pre-set: for Chinese, it can include commas, periods, question marks, exclamation marks, and semicolons; and for English, it can include commas, periods, question marks, and exclamation marks. In addition, in the context of pure Chinese or pure English, the corresponding language punctuation is used, otherwise the Chinese punctuation is used.
[0031] In a specific implementation scenario, the first sample text with a display form can be extracted from books, journals, newspapers, the Internet and other channels. As a possible implementation example, in order to adapt to the voice scene, the text extracted from the above channels is usually a written style text, so it can be processed in a colloquial way, such as adding colloquial words, random repetition, synonym replacement and other processing to enhance the spoken style, and then used as the first sample text. In addition, when constructing the first pronunciation sequence of the first sample text, the text normalization (TN) tool can be used to convert the numbers in the first sample text into corresponding text. For example, the number "12" in the first sample text "Please pay 12 yuan." needs to be converted into the text "twelve", and the number "88886666" in the first sample text "Please call 88886666 for consultation." needs to be converted into the text "eight eight eight eight six six six six". Of course, the above examples are only a few possible examples of text regularization, and other possible situations are not given one by one here.
[0032] In a specific implementation scenario, in order to adapt to the input style of the speech recognition system, the data can also be segmented. For example, for speech data, audio and annotation segmentation can be performed according to the audio segmentation rules; and for text data, especially long text data, data segmentation can also be performed. Considering that people tend to pause at punctuation marks when actually speaking, or they may not pause or pause in normal sentences, they can also be randomly segmented according to a certain probability distribution at the punctuation marks and positions in the sentences of long texts. In addition, in addition to the above-mentioned manual construction and network extraction methods, weakly supervised data can also be constructed through a speech recognition system or a text post-processing model. For example, whisper can be used to obtain weakly supervised data with punctuation marks, numbers, and uppercase and lowercase letters for unlabeled audio, or a post-processing text model can be used to process the recognition labels and add punctuation marks, numbers, uppercase and lowercase letters and other information. Of course, the above examples are only a few possible examples of obtaining sample data outside the aforementioned methods, and other possible situations will not be given one by one again.
[0033] In a specific implementation scenario, after the sample data is prepared, the parameters of the large language model can be fine-tuned accordingly. It should be noted that the parameter fine-tuning can adopt fine-tuning technologies such as LORA, and the specific technical methods adopted for parameter fine-tuning are not limited here. Figure 2a , Figure 2a Schematic diagram of the process of fine-tuning parameters in this application. Figure 2a As shown, feature encoding can be performed based on the sample data to obtain sample encoded data, and the sample encoded features can be autoregressively decoded according to the sample instructions based on the large language model to obtain predicted text, and the sample instructions include recognition instructions and post-processing instructions corresponding to the presentation form of the sample data. It should be noted that the specific meaning of the sample instructions can be found in the relevant description of the aforementioned prompt instructions, which will not be repeated here. On this basis, the first loss can be obtained based on the difference between the predicted text and the expected text of the sample data, and the network parameters of the large language model can be adjusted based on the first loss. It should be noted that when the sample data is the first sample speech, the expected text is the first target text, and when the sample data is the first pronunciation sequence, the expected text is the first sample text belonging to the first pronunciation sequence. In addition, Figure 2a It is not indicated that the model input data of the large language model further includes sample instructions, but it does not mean that the model input data of the large language model does not include sample instructions. Figure 2a The above method measures the loss by the difference between the predicted text output by the large language model and the expected text of the sample data, and adjusts the network parameters accordingly, which can force the large language model to learn both speech transcription and post-processing during the autoregressive decoding process.
[0034] In a specific implementation scenario, please continue to refer to the implementation example of the above parameter fine-tuning. Figure 2a , taking the first sample speech "have you eaten?" or the first pronunciation sequence "chi le ma" of the first sample text "have you eaten?" as an example, no matter which one is used as input data, feature extraction can be performed through the encoder to obtain sample encoding features, and after dimension conversion by the adapter, it is used together with sample instructions (not shown, and in this example, it can include recognition instructions and post-processing instructions corresponding to the display form "punctuation") as the model input data of the large language model, so that the large language model can perform autoregressive decoding according to the sample instructions to obtain predicted text with a display form, such as "have you eaten?", "have you eaten." and other possible situations. Of course, optionally, the model input data may also include sample reference text, which is text data that occurs before the sample data and has the target display form. Exemplarily, in Figure 2a In the example of , the sample reference text may include but is not limited to: "Good morning.", etc., in order to assist recognition in long speech flow scenarios, and the specific content of the sample reference text is not limited here.
[0035] In a specific implementation scenario, in the implementation example of the aforementioned parameter fine-tuning, the difference between the predicted text and the expected text of the sample data can be measured by a loss function such as cross entropy to obtain a first loss. For the specific measurement process, please refer to the technical details of the loss function such as cross entropy, which will not be repeated here.
[0036] In a specific implementation scenario, in the implementation example of the aforementioned parameter fine-tuning, in order to further strengthen the attention of the large language model to the characters related to the presentation form during the autoregressive decoding process, before adjusting the network parameters in the aforementioned implementation example, the characters related to the presentation form in the expected text can be selected as the expected characters, and the characters corresponding to the positions of the expected characters in the predicted text can be selected as the predicted characters. Taking the expected text "Have you eaten?" as an example, its presentation form is "punctuation", so "?" can be selected as the expected character, and the character "." corresponding to the position in the predicted text "Have you eaten." can be selected as the predicted character. Of course, the above example is only a possible example in the actual application process, and the specific content of the expected character and the predicted character is not limited here. In addition, the characters related to the presentation form may not be limited to the characters directly corresponding to the presentation form (such as punctuation marks directly corresponding to the presentation form "punctuation", letters directly corresponding to the presentation form "uppercase and lowercase", and Arabic numerals directly corresponding to the presentation form "numbers"), and may further include its adjacent characters (such as directly corresponding characters and the text before and after them). For example, for the text "No. 1, ups and downs, seven or eight or so", the relevant characters in the form of "number" may include "No.", "1", and "name". On this basis, the second loss can be obtained based on the difference between the expected characters and the actual characters. It should be noted that, similar to the measurement method of the aforementioned first loss, the difference between the expected characters and the predicted characters can also be measured based on a loss function such as cross entropy to obtain the second loss. After obtaining the first loss and the second loss, the network parameters of the large language model can be adjusted based on the first loss and the second loss. Exemplarily, the training loss of the large language model can be obtained by weighting the first loss and the second loss, so as to adjust the network parameters of the large language model based on the training loss of the large language model. For example, the weight factor of the first loss can be set to 1, and the weight factor of the second loss can be set to λ, then the training loss L rec It can be expressed as: L rec =L ce +λL aux In the above formula, L ce represents the first loss, L aux In addition, the second loss L is obtained by using the cross entropy loss function to measure the difference between the expected character and the predicted character. aux For example, the second loss L aux It can be expressed as:
[0037]
[0038] In the above formula, Characters related to the display form "punctuation", Characters representing the display form "number", Characters related to "uppercase and lowercase" display format, represents the union of sets, Indicates the relevant characters, Indicates the current network parameters in the large language model The next predicted character is the i-th expected character Therefore, by minimizing the loss function, we can force the probability of the predicted character to be the expected character as large as possible.
[0039] In a specific implementation scenario, in the implementation example of the above-mentioned parameter fine-tuning, the sample data also covers at least one of the following types of data: a second sample voice without a presentation form, a second pronunciation sequence of a second sample text without a presentation form, and a second sample voice annotated with a second target text with the same content. Take the presentation form including punctuation, numbers, and uppercase and lowercase as an example, that is, the second target text annotated by the above-mentioned second sample voice without a presentation form is a pure text text without punctuation, numbers, and uppercase, and the above-mentioned second sample text without a presentation form is a pure text text without punctuation, numbers, and uppercase. It should be noted that in the process of parameter fine-tuning, the specific process of model training based on the above-mentioned sample data without a presentation form can refer to the specific process of model training based on the above-mentioned sample data with a presentation form. The main difference is that when the model is trained based on the above-mentioned sample data without a presentation form, the sample prompt instruction can only include recognition instructions, but not post-processing instructions, because the post-processing instructions are used to instruct the large language model to perform speech recognition with a presentation form as an additional target, and since the sample data itself does not have a presentation form, the sample prompt instruction naturally does not contain post-processing instructions. In addition, when the model is trained based on the sample data without presentation form, the second loss mentioned above may not need to be measured, because the expected text does not contain relevant characters in presentation form, so the second loss may not need to be measured. For other details, please refer to the specific process of model training based on sample data with presentation form, which will not be described here.
[0040] In one implementation scenario, in order to further improve the support capability for target presentation forms such as punctuation during speech recognition in long speech stream scenarios, the sample data may also be accompanied by sample reference text, which may be text data that occurs before the sample data and has the target presentation form, and the target presentation form may at least include punctuation, and the large language model performs mask prediction tasks based on the relevant characters of the target presentation form in the sample reference text to adjust parameters during or after parameter fine-tuning. In other words, different from the model input data of the large language model in the general scenario shown in Table 1 above, when the target presentation form is punctuation, the model input data of the large language model in the long speech stream scenario can refer to Table 2.
[0041] Table 2 Schematic diagram of an embodiment of model input data of a large language model in a long speech stream scenario
[0042] As shown in Table 2, when the target presentation form is punctuation, the model input data of the large language model may include sample reference text, sample encoding features, sample prompt instructions (including recognition instructions and post-processing instructions 1, i.e., post-processing instructions corresponding to the presentation form "punctuation"), and decoded characters output at each time step during the autoregressive decoding process. In addition, as mentioned above, the mask prediction task can be performed together in the aforementioned parameter fine-tuning process. In this case, the training loss can not only include the loss shown in the aforementioned formula and its corresponding text, but also further include the relevant loss of the mask prediction task (for details, please refer to the following related description), or the mask prediction task can also be performed separately after the aforementioned parameter fine-tuning, which is not limited here. Of course, if there is no long speech flow scenario in the actual application of speech recognition, the mask prediction task may not be performed. In other words, whether to perform the mask prediction task, and at which stage (during parameter fine-tuning or after parameter fine-tuning) the mask prediction task is performed when it is required, can be set according to the actual application needs, and are not limited here.
[0043] In a specific implementation scenario, as a possible implementation method, when adjusting parameters based on the mask prediction task, feature encoding can be performed based on sample data to obtain sample encoding features, and a target attention mask for attention calculation between several model inputs can be configured, and the several model inputs include sample reference text, sample encoding features, sample instructions, and expected text of sample data. The target attention mask can be configured so that any element in the sample reference text can pay attention to any element in the sample encoding features that follow it, that is, the sample reference text can pay attention to the sample encoding features that occur relative to it in the future. As a specific example, please refer to Figure 2b , Figure 2b Schematic diagram of the effect of an embodiment of the target attention mask of the present application. Figure 2b As shown, the length and width of the target attention mask are the total length of the elements input by several models, and the target attention mask is configured so that any element in the expected text can only pay attention to any element before it. Exemplarily, the value of the filled part in the target attention mask can be configured to 1 (that is, this part of the field of view is visible). In other words, the sample reference text, sample encoding features, and sample prompt instructions can all observe all information before the sample prompt instruction in the attention calculation. The value of the unfilled part can be configured to 0 (that is, this part of the field of view is invisible). In other words, for <sos>to <eos>For some characters, only historical information can be paid attention to, but not future information. Therefore, the target attention mask can constrain the attention mechanism in autoregressive decoding. Figure 2c , Figure 2c Schematic diagram of the effect of an embodiment of the mask prediction task of this application. Figure 2c As shown, on this basis, the sample reference text and the sample encoding features can be autoregressively decoded according to the sample instructions through the target attention mask based on the large language model to obtain the predicted text and the masked first character in the sample reference text, and the sample instructions include recognition instructions and post-processing instructions corresponding to the target presentation form, and then based on the difference between the predicted text and the expected text, as well as the first character and related characters (such as Figure 2c As shown in the shaded area in the sample reference text shown, when the target presentation form is "punctuation", the punctuation at the end of the sample reference text can be masked, and a certain proportion of characters in addition can be further masked) to adjust the network parameters of the large language model. It should be noted that when the sample data is the first sample speech, the expected text is the first target text, and when the sample data is the first pronunciation sequence, the expected text is the first sample text belonging to the first pronunciation sequence. In addition, the difference between the predicted text and the expected text can refer to the relevant description of the aforementioned parameter fine-tuning, which will not be repeated here. The difference between the first character and the related characters can also be measured by a loss function such as cross entropy, which is not limited here. Of course, in this process, you can also refer to the measurement method of the second loss in the aforementioned parameter fine-tuning part, and measure the loss of the difference between the expected character and the predicted character, so as to adjust the network parameters of the large language model in combination with the three differences. For details, please refer to the relevant description in the aforementioned parameter fine-tuning, which will not be repeated here. The above method, by configuring the target attention mask, can force the large language model to pay attention to the future sample encoding features for the sample reference text during the autoregressive decoding process, which helps to improve the accuracy of predicting the masked characters in the sample reference text, so as to achieve error correction in the long speech flow scenario. Moreover, since the mask prediction task and the first word decoding are completed in the same calculation step, there is no efficiency loss.
[0044] In a specific implementation scenario, please refer to Figure 2d , Figure 2d FIG. 1 is a schematic diagram showing the effect of another embodiment of the mask prediction task of the present application. Figure 2d As shown, as another possible implementation, different from the aforementioned implementation, when adjusting parameters based on the mask prediction task, feature encoding can be performed based on the sample data to obtain sample encoding features, and the sample reference text and the sample encoding features are autoregressively decoded according to the sample instructions through the autoregressive attention mask based on the large language model to obtain the predicted text and the masked second character in the sample reference text, and the sample instructions include the recognition instructions and the post-processing instructions corresponding to the target presentation form, and then based on the difference between the predicted text and the expected text of the sample data, and the difference between the second character and the related characters (such as Figure 2d As shown in the shaded area in the sample reference text, when the target display form is "punctuation", the punctuation at the end of the sample reference text can be masked, and a certain proportion of characters in addition can be further masked) to adjust the network parameters of the large language model. It should be noted that when the sample data is the first sample speech, the expected text is the first target text, and when the sample data is the first pronunciation sequence, the expected text is the first sample text belonging to the first pronunciation sequence. As a special example, Figure 2d middle <ask>It can be the predicted character (i.e., the second character) output by the large language model when the target presentation form is "punctuation" and the punctuation at the end of the sample reference text is obscured. When the performance of the large language model is good enough, if the end of the sample reference text is obscured, then <ask>The predicted characters output by the large language model are the masked related characters. If the ending of the sample reference text is not masked, then <ask>The predicted character output by the large language model is the character at the end itself, and if there is no sample reference text, then <ask>The predicted character output by the large language model is empty. Further, the loss here can also be measured by a loss function such as cross entropy. For example, the second character and the related character can be measured by the cross entropy loss function to obtain the corresponding loss L ask :
[0045] In the above formula, Indicates the current network parameters in the large language model The second character that is masked is predicted to be the relevant character , so by minimizing the loss function, the possibility that the second character is a related character can be forced to be as large as possible. For the rest of the training loss, please refer to the aforementioned description related to the loss calculation, which will not be repeated here. In addition, the similarities between this embodiment and the previous embodiment can be referred to the previous embodiment, which will not be repeated here. The main difference is that this embodiment adopts an autoregressive attention mask, while the previous embodiment customizes the target attention mask. Different from the previous embodiment that customizes the target attention mask, the autoregressive attention mask can only pay attention to historical information, such as Figure 2b Taking the mask shown in the figure as an example, the difference between the autoregressive attention mask and the autoregressive attention mask is that only the value of the part below the main diagonal of the autoregressive attention mask is configured as 1 (that is, this part of the field of view is visible), and the rest of the value is configured as (that is, this part of the field of view is invisible). On the one hand, compared with configuring a custom target attention mask, it can avoid conflicts with the attention mask of the large language model as much as possible and reduce modifications to the base model. On the other hand, under the autoregressive attention mask mechanism, although the sample reference text part cannot pay attention to the sample encoding features relative to its future, it can still <sos>So far, the sample reference text and sample encoding features have been obtained, so the semantic conditions for correcting errors in the sample reference text are met.
[0046] The above scheme performs feature encoding based on the speech to be recognized to obtain the encoded features, and performs autoregressive decoding on the encoded features based on the large language model according to the prompt instruction to obtain the recognized text with a display form of the speech to be recognized, and the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform speech recognition with the additional goal of having a display form, and various post-processing instructions correspond to different display forms, so on the one hand, by inputting the post-processing instruction together with the recognition instruction, forcing the large language model to perform post-processing together with the speech recognition during the autoregressive decoding process, it is possible to ensure that the recognized text has a display form as much as possible, and on the other hand, because the speech recognition and its post-processing are completed together by the large language model during the autoregressive decoding process, that is, there is no need to successively realize speech transcription and post-processing through a cascade structure, which can shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition. Therefore, it is possible to shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition under the premise of ensuring that the recognized text has a display form as much as possible. In addition, since various post-processing instructions correspond to different display forms, different post-processing instructions can be configured specifically according to different application scenarios during the speech recognition process, thereby improving the convenience of differentiated processing of different application scenarios.
[0047] See also Figure 3 , Figure 3 It is a schematic diagram of the framework of an embodiment of the speech recognition device of the present application. The speech recognition device 30 includes: a feature encoding module 31 and an autoregressive decoding module 32, wherein the feature encoding module 31 is used to perform feature encoding based on the speech to be recognized to obtain the encoded features; the autoregressive decoding module 32 is used to perform autoregressive decoding on the encoded features according to the prompt instructions based on the large language model to obtain the recognition text of the speech to be recognized in the display form; wherein the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform speech recognition with the display form as an additional target, and various post-processing instructions correspond to different display forms.
[0048] In the above scheme, the speech recognition device 30 performs feature encoding based on the speech to be recognized to obtain the encoding feature, and performs autoregressive decoding on the encoding feature based on the large language model according to the prompt instruction to obtain the recognition text of the speech to be recognized with a display form, and the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform speech recognition with a display form as an additional goal, and various post-processing instructions correspond to different display forms, so on the one hand, by inputting the post-processing instruction together with the recognition instruction, forcing the large language model to perform post-processing together with the speech recognition in the autoregressive decoding process, it is possible to ensure that the recognized text has a display form as much as possible, and on the other hand, because the speech recognition and its post-processing are completed together by the large language model in the autoregressive decoding process, that is, there is no need to successively realize speech transcription and post-processing through a cascade structure, which can shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition. Therefore, it is possible to shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition under the premise of ensuring that the recognized text has a display form as much as possible. In addition, since various post-processing instructions correspond to different display forms, different post-processing instructions can be configured specifically according to different application scenarios during the speech recognition process, thereby improving the convenience of differentiated processing of different application scenarios.
[0049] In some disclosed embodiments, the large language model uses sample data for parameter fine-tuning, and the sample data covers at least one of the following types of data: a first sample speech having a presentation form, a first pronunciation sequence of a first sample text having a presentation form, and the first sample speech is annotated with a first target text having consistent content and a presentation form.
[0050] In some disclosed embodiments, the speech recognition device 30 includes a first sample encoding module, which is used to perform feature encoding based on sample data to obtain sample encoding features; the speech recognition device 30 includes a first sample decoding module, which is used to perform autoregressive decoding on the sample encoding features according to sample instructions based on a large language model to obtain predicted text; wherein the sample instructions include recognition instructions and post-processing instructions corresponding to the presentation form of the sample data; the speech recognition device 30 includes a first parameter adjustment module, which is used to obtain a first loss based on the difference between the predicted text and the expected text of the sample data, and adjust the network parameters of the large language model based on the first loss; wherein, when the sample data is the first sample speech, the expected text is the first target text, and when the sample data is the first pronunciation sequence, the expected text is the first sample text belonging to the first pronunciation sequence.
[0051] In some disclosed embodiments, the speech recognition device 30 includes a character selection module for selecting relevant characters presented in the expected text as expected characters, and selecting characters corresponding to the positions of the expected characters in the predicted text as predicted characters; the speech recognition device 30 includes a loss measurement module for obtaining a second loss based on the difference between the expected character and the predicted character; and the first parameter adjustment module is specifically used to adjust the network parameters of the large language model based on the first loss and the second loss.
[0052] In some disclosed embodiments, the sample data is feature encoded by an encoder that is compatible with both speech and text modes, and the network parameters of the encoder are frozen during parameter fine-tuning; and / or the sample data also covers at least one of the following types of data: a second sample speech that has no presentation form, a second pronunciation sequence of a second sample text that has no presentation form, and the second sample speech is annotated with a second target text with consistent content.
[0053] In some disclosed embodiments, the sample data is also accompanied by sample reference text, which is text data that occurs before the sample data and has a target presentation form, the target presentation form at least includes punctuation, and the large language model performs a mask prediction task to adjust parameters based on relevant characters of the target presentation form in the sample reference text during or after parameter fine-tuning.
[0054] In some disclosed embodiments, the speech recognition device 30 includes a second sample encoding module for performing feature encoding based on sample data to obtain sample encoding features, and the speech recognition device 30 includes a mask configuration module for configuring a target attention mask for attention calculation between a number of model inputs; wherein the number of model inputs include sample reference text, sample encoding features, sample instructions, and expected text of the sample data, and the target attention mask is configured so that any element in the sample reference text can pay attention to any element in the sample encoding features located thereafter; the speech recognition device 30 includes a second sample decoding module for performing autoregressive decoding on the sample reference text and the sample encoding features through the target attention mask according to the sample instructions based on the large language model to obtain the predicted text and the masked first character in the sample reference text; wherein the sample instructions include recognition instructions and post-processing instructions corresponding to the target presentation form; the speech recognition device 30 includes a second parameter adjustment module for adjusting the network parameters of the large language model based on the difference between the predicted text and the expected text, and the difference between the first character and the related characters; wherein, when the sample data is the first sample speech, the expected text is the first target text, and when the sample data is the first pronunciation sequence, the expected text is the first sample text belonging to the first pronunciation sequence.
[0055] In some disclosed embodiments, the length and width of the target attention mask are both the total length of elements of several model inputs, and the target attention mask is configured so that any element in the expected text can only focus on any element that precedes it.
[0056] In some disclosed embodiments, the speech recognition device 30 includes a third sample encoding module for performing feature encoding based on sample data to obtain sample encoding features; the speech recognition device 30 includes a third sample decoding module for performing autoregressive decoding on sample reference text and sample encoding features through an autoregressive attention mask according to sample instructions based on a large language model to obtain a predicted text and a masked second character in the sample reference text; wherein the sample instructions include recognition instructions and post-processing instructions corresponding to a target presentation form; the speech recognition device 30 includes a third parameter adjustment module for adjusting network parameters of the large language model based on the difference between the predicted text and the expected text of the sample data, and the difference between the second character and related characters; wherein, when the sample data is a first sample speech, the expected text is a first target text, and when the sample data is a first pronunciation sequence, the expected text is a first sample text belonging to the first pronunciation sequence.
[0057] In some disclosed embodiments, the presentation form includes: at least one of uppercase and lowercase, numbers, and punctuation; and / or, when the feature dimension after feature encoding is different from the feature dimension of the embedding layer in the large language model, after feature encoding and before autoregressive decoding, the dimension is first converted by an adapter to be consistent with the feature dimension of the embedding layer, and the adapter and the large language model perform parameter fine-tuning together; and / or, in the process of performing autoregressive decoding, the encoded features are also spliced with reference text, and the reference text is a conversation text that occurs before the speech to be recognized.
[0058] See also Figure 4 , Figure 4 It is a schematic diagram of the framework of an embodiment of an electronic device of the present application. The electronic device 40 includes at least a memory 41 and a processor 42 coupled to each other, the memory 41 stores at least program instructions, and the processor 42 is used to execute the program instructions to implement the steps in any of the above-mentioned speech recognition method embodiments. For details, please refer to the aforementioned disclosed embodiments, which will not be repeated here. It should be noted that the electronic device 40 may include but is not limited to devices such as a translator, a learning machine, an office book, a headset, a mouse, a smart large screen, etc., and the specific type of the electronic device 40 is not limited here.
[0059] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above-mentioned speech recognition method embodiments. The processor 42 can also be called a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 42 can be implemented by an integrated circuit chip.
[0060] In the above scheme, the electronic device 40 performs feature encoding based on the speech to be recognized to obtain the encoding feature, and performs autoregressive decoding on the encoding feature based on the large language model according to the prompt instruction to obtain the recognition text of the speech to be recognized with a display form, and the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform speech recognition with a display form as an additional goal, and various post-processing instructions correspond to different display forms, so on the one hand, by inputting the post-processing instruction together with the recognition instruction, forcing the large language model to perform post-processing together with the speech recognition during the autoregressive decoding process, it is possible to ensure that the recognized text has a display form as much as possible, and on the other hand, because the speech recognition and its post-processing are completed together by the large language model during the autoregressive decoding process, that is, there is no need to successively realize speech transcription and post-processing through a cascade structure, which can shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition. Therefore, it is possible to shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition under the premise of ensuring that the recognized text has a display form as much as possible. In addition, since various post-processing instructions correspond to different display forms, different post-processing instructions can be configured specifically according to different application scenarios during the speech recognition process, thereby improving the convenience of differentiated processing of different application scenarios.
[0061] See also Figure 5 , Figure 5 The computer-readable storage medium 50 stores program instructions 51 that can be executed by a processor, and the program instructions 51 are used to implement the steps in any of the above-mentioned speech recognition method embodiments.
[0062] In the above scheme, the computer-readable storage medium 50 performs feature encoding based on the speech to be recognized to obtain the encoding feature, and performs autoregressive decoding on the encoding feature based on the large language model according to the prompt instruction to obtain the recognition text of the speech to be recognized with a display form, and the prompt instruction includes a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform speech recognition with a display form as an additional goal, and various post-processing instructions correspond to different display forms, so on the one hand, by inputting the post-processing instruction together with the recognition instruction, forcing the large language model to perform post-processing together with the speech recognition during the autoregressive decoding process, it is possible to ensure that the recognized text has a display form as much as possible, and on the other hand, because the speech recognition and its post-processing are completed together by the large language model during the autoregressive decoding process, that is, there is no need to successively realize speech transcription and post-processing through a cascade structure, which can shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition. Therefore, it is possible to shorten the response time of speech recognition, reduce the computational burden of speech recognition, and improve the output accuracy of speech recognition under the premise of ensuring that the recognized text has a display form as much as possible. In addition, since various post-processing instructions correspond to different display forms, different post-processing instructions can be configured specifically according to different application scenarios during the speech recognition process, thereby improving the convenience of differentiated processing of different application scenarios.
[0063] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0064] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.
[0065] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0066] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0067] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0068] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
[0069] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, set clear and prominent signs to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to collect his or her personal information; or on the device that processes personal information, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.< / sos> < / ask> < / ask> < / ask> < / ask> < / eos> < / sos> < / eos> < / sos>
Claims
1. A speech recognition method, characterized in that: include: Perform feature encoding based on the speech to be recognized to obtain encoding features; Based on the large language model, the encoded features are autoregressively decoded according to the prompt instructions to obtain the recognition text of the speech to be recognized in a presentation form; wherein the prompt instructions include recognition instructions and at least one post-processing instruction, the recognition instructions are used to instruct the large language model to perform speech recognition, and the post-processing instructions are used to instruct the large language model to perform the speech recognition with the presentation form as an additional goal, and various post-processing instructions correspond to different presentation forms.
2. The method according to claim 1, characterized in that The large language model uses sample data for parameter fine-tuning, and the sample data covers at least one of the following types of data: a first sample speech having the presentation form, a first pronunciation sequence of a first sample text having the presentation form, and the first sample speech is annotated with a first target text with consistent content and having the presentation form.
3. The method according to claim 2, characterized in that The step of parameter fine-tuning comprises: Perform feature encoding based on the sample data to obtain sample encoding features; Based on the large language model, the sample encoding features are autoregressively decoded according to the sample instructions to obtain a predicted text; wherein the sample instructions include the recognition instructions and post-processing instructions corresponding to the presentation form of the sample data; Obtaining a first loss based on a difference between the predicted text and the expected text of the sample data, and adjusting a network parameter of the large language model based on the first loss; When the sample data is the first sample speech, the expected text is the first target text; when the sample data is the first pronunciation sequence, the expected text is the first sample text to which the first pronunciation sequence belongs.
4. The method according to claim 3, characterized in that Before adjusting the network parameters of the large language model based on the first loss, the method further includes: Selecting a character related to the presentation form in the expected text as the expected character, and selecting a character corresponding to the position of the expected character in the predicted text as the predicted character; Obtaining a second loss based on a difference between the expected character and the predicted character; The adjusting the network parameters of the large language model based on the first loss includes: Based on the first loss and the second loss, a network parameter of the large language model is adjusted.
5. The method according to claim 2, characterized in that: The sample data is feature encoded by an encoder compatible with both speech and text modes, and during the parameter fine-tuning process, the network parameters of the encoder are frozen; And / or, the sample data also covers at least one of the following types of data: a second sample speech that does not have the presentation form, a second pronunciation sequence of a second sample text that does not have the presentation form, and the second sample speech is annotated with a second target text with consistent content.
6. The method according to claim 2, characterized in that The sample data is also accompanied by sample reference text, which is text data that occurs before the sample data and has a target presentation form, wherein the target presentation form includes at least punctuation marks, and the large language model performs a mask prediction task to adjust parameters based on relevant characters of the target presentation form in the sample reference text during or after the parameter fine-tuning process.
7. The method according to claim 6, characterized in that The step of adjusting parameters based on the mask prediction task includes: Perform feature encoding based on the sample data to obtain sample encoding features, and configure a target attention mask for attention calculation between a plurality of model inputs; wherein the plurality of model inputs include the sample reference text, the sample encoding features, a sample instruction, and an expected text of the sample data, and the target attention mask is configured so that any element in the sample reference text can pay attention to any element in the sample encoding features located thereafter; Based on the large language model, the sample reference text and the sample encoding feature are autoregressively decoded through the target attention mask according to the sample instruction to obtain the predicted text and the masked first character in the sample reference text; wherein the sample instruction includes the recognition instruction and the post-processing instruction corresponding to the target presentation form; Adjusting network parameters of the large language model based on the difference between the predicted text and the expected text, and the difference between the first character and the related character; When the sample data is the first sample speech, the expected text is the first target text; when the sample data is the first pronunciation sequence, the expected text is the first sample text to which the first pronunciation sequence belongs.
8. The method according to claim 7, characterized in that The length and width of the target attention mask are both the total length of the elements input by the several models, and the target attention mask is configured so that any element in the expected text can only focus on any element located before it.
9. The method according to claim 6, characterized in that The step of adjusting parameters based on the mask prediction task includes: Perform feature encoding based on the sample data to obtain sample encoding features; Based on the large language model, the sample reference text and the sample encoding feature are autoregressively decoded by an autoregressive attention mask according to the sample instruction to obtain a predicted text and a masked second character in the sample reference text; wherein the sample instruction includes the recognition instruction and a post-processing instruction corresponding to the target presentation form; Adjusting the network parameters of the large language model based on the difference between the predicted text and the expected text of the sample data, and the difference between the second character and the related character; When the sample data is the first sample speech, the expected text is the first target text; when the sample data is the first pronunciation sequence, the expected text is the first sample text to which the first pronunciation sequence belongs.
10. The method according to any one of claims 1 to 9, characterized in that: The display form includes: at least one of uppercase and lowercase letters, numbers, and punctuation marks; and / or, when the feature dimension after performing the feature encoding is different from the feature dimension of the embedding layer in the large language model, after performing the feature encoding and before performing the autoregressive decoding, a dimension conversion is first performed by an adapter to be consistent with the feature dimension of the embedding layer, and the adapter and the large language model are fine-tuned together; And / or, in the process of performing the autoregressive decoding, the encoding features are also spliced with reference text, and the reference text is a conversation text that occurred before the speech to be recognized.
11. A speech recognition device, characterized in that: include: A feature encoding module, used for performing feature encoding based on the speech to be recognized to obtain encoding features; An autoregressive decoding module is used to perform autoregressive decoding on the encoded features according to the prompt instructions based on the large language model to obtain a recognized text with a presentation form of the speech to be recognized; wherein the prompt instructions include a recognition instruction and at least one post-processing instruction, the recognition instruction is used to instruct the large language model to perform speech recognition, and the post-processing instruction is used to instruct the large language model to perform the speech recognition with the presentation form as an additional goal, and various post-processing instructions correspond to different presentation forms.
12. An electronic device, characterized in that: The invention at least comprises a memory and a processor coupled to each other, wherein the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the speech recognition method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the speech recognition method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Method and device for intention classification based on speech recognition result
CN111177324A
Speech recognition method and system, medium, computer equipment, terminal and application
CN112712804A
Speech recognition method and device, equipment and storage medium
CN117153152A
Speech recognition method and device, equipment and storage medium
CN117636873A
Interaction processing method, device and equipment, man-machine interaction system and program product
CN118918889A