Electric power speech recognition method and device based on general speech recognition large model and electric power error correction model
By automatically generating the audio and text data of power terminology, an error correction model in the power field is constructed, and combined with the speech recognition model, the problem of poor power term recognition in the prior art is solved, and high accuracy and low cost power speech recognition is achieved.
Patent Information
- Application Number
- CN202510342685.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-20
AI Technical Summary
The existing speech recognition model does not perform well in the recognition of power terms, and the existing solutions require manual construction of large amounts of labeled data, which is costly and inefficient in training.
Generate power term audio and corresponding text data through automated means, build an error correction model in the power field, and combine the speech recognition model to form a speech recognition model dedicated to the power field.
It significantly improves the recognition accuracy of power terms, reduces the cost of manually constructing training samples, and improves the accuracy of speech recognition.
Smart Images

Figure CN120183398A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech recognition, and particularly relates to a power speech recognition method and device based on a general speech recognition large model and a power error correction model. Background Art
[0002] The power knowledge Q&A engine is an intelligent system integrating technologies such as natural language processing, knowledge graph, and cognitive large models, aiming to provide users with professional question answers and knowledge services in the power field. With the continuous development of intelligent technologies, the power knowledge Q&A engine usually supports users to input power-related query questions by voice, similar to the function of Siri. Existing better speech recognition models, such as OpenAI Whisper, have good speech recognition effects in the general Chinese field, but perform poorly in the recognition of power terms. This is mainly because of the professionalism and particularity of power terms, making it difficult for existing speech recognition models to accurately recognize them.
[0003] Currently, the method to solve this problem is to construct training data of "power term audio - corresponding power term text" manually and perform supervised training on the speech recognition model to improve the recognition accuracy of power terms. For example, for the OpenAI Whisper speech recognition model, in order to support the speech recognition effect in the general Chinese field, 250,000 labeled Chinese audio data in the general field are constructed manually. However, this method has problems such as high training cost and low efficiency, and it is difficult to meet the requirements of actual applications. Summary of the Invention
[0004] The purpose of the present invention is to provide a power speech recognition method and device based on a power speech recognition large model to solve the technical problems that the existing training method of the power speech recognition large model requires a large amount of manual construction of annotations, high training cost, and low training efficiency. The present invention generates power term audio and corresponding text data in an automated manner, constructs a power field error correction model, and combines it with the speech recognition model to form a dedicated speech recognition model for the power field, so as to significantly improve the recognition accuracy of power terms and reduce the cost of manually constructing training samples.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a power speech recognition method based on a general speech recognition large model and a power error correction model, including: Obtain the audio of the power query question; Input the audio of the power query question into a pre-established general speech recognition large model for speech recognition to obtain the text of the preliminary power query question; Match the text of the preliminary power query problem with a pre - constructed error - correction dictionary in the power field. If an error - correction phrase is matched, input the text of the preliminary power query problem and the error - correction phrase into a pre - established error - correction model in the power field to obtain the text of the finally error - corrected power query problem.
[0006] A further improvement of the present invention lies in that the error - correction dictionary in the power field is constructed through the following steps: Input the Chinese text of the original power query problem into a text - to - audio model to generate the Chinese audio of the power query problem. Input the Chinese audio of the power query problem into an audio - to - text model to identify and output the Chinese text of the final power query problem. Compare the Chinese text of the original power query problem with the Chinese text of the final power query problem. If there are differences between the two texts, extract the correct phrases and incorrect phrases of the differences as example entries in the error - correction dictionary. Collect a number of example entries in the error - correction dictionary and construct an error - correction dictionary in the power field.
[0007] A further improvement of the present invention lies in that in the step of inputting the Chinese text of the original power query problem into a text - to - audio model to generate the Chinese audio of the power query problem, the text - to - audio model uses edge - tts. In the step of inputting the Chinese audio of the power query problem into an audio - to - text model to identify and output the Chinese text of the final power query problem, the audio - to - text model uses Whisper.
[0008] A further improvement of the present invention lies in that the pre - established error - correction model in the power field is obtained through the following steps of training: Collect text data in the power field, including correct power texts and error texts generated by speech recognition; generate positive samples, random negative samples, and difficult negative samples based on the collected text data in the power field to obtain pre - processed text data. Convert the pre - processed text data into vectors to obtain output vectors. Extract the feature representations of the output vectors. Use the first large - language model to encode the feature representations of the positive samples, negative samples, and difficult negative samples of the vocabulary to generate vocabulary embedding vectors; use the second large - language model to encode the feature representations of the positive samples, negative samples, and difficult negative samples of the sentences to generate sentence embedding vectors. Use the vocabulary embedding vectors to train the first large - language model to obtain a power word embedding model; use the sentence embedding vectors to train the second large - language model to obtain a sentence word embedding model. Solidify the parameters of the power word embedding model, and train a first semantic classifier for evaluating the semantic integrity of words based on the power word embedding model; solidify the parameters of the power sentence embedding model, and train a second semantic classifier for evaluating the semantic integrity of sentences based on the power sentence embedding model.
[0009] A further improvement of the present invention is that in the step of encoding the feature representations of the positive samples, negative samples, and hard negative samples of the vocabulary using the first large language model to generate word embedding vectors, for the vocabulary, the positive samples are correct power words, and the negative samples are randomly selected power words different from the positive samples; the hard negative samples are words with speech recognition errors. In the step of encoding the feature representations of the positive samples, negative samples, and hard negative samples of the sentence using the second large language model to generate sentence embedding vectors, for the sentence, the positive samples are correct power sentences, and the negative samples are randomly selected power sentences different from the positive samples; the hard negative samples are generated from sentences with speech recognition errors.
[0010] A further improvement of the present invention is that the steps of training the first large language model using the word embedding vectors to obtain a power word embedding model, and training the second large language model using the sentence embedding vectors to obtain a sentence word embedding model specifically include: Use the first large language model to encode the feature representations of the positive samples, negative samples, and hard negative samples of the vocabulary to generate word embedding vectors; perform contrastive learning on the word embedding vectors generated by the first large language model, calculate the cosine similarity between the positive and negative sample pairs respectively, and calculate the first contrast loss L contrastive through iteration until the first contrast loss L contrastive converges to obtain a power word embedding model; Use the second large language model to encode the feature representations of the positive samples, negative samples, and hard negative samples of the sentence to generate sentence embedding vectors; perform contrastive learning on the sentence embedding vectors generated by the second large language model, calculate the cosine similarity between the positive and negative sample pairs respectively, and calculate the second contrast loss L contrastive through iteration until the second contrast loss L contrastive converges to obtain a sentence word embedding model.
[0011] A further improvement of the present invention is that the steps of solidifying the parameters of the power word embedding model and training a first semantic classifier for evaluating the semantic integrity of words based on the power word embedding model, and solidifying the parameters of the power sentence embedding model and training a second semantic classifier for evaluating the semantic integrity of sentences based on the power sentence embedding model specifically include: Solidify the parameters of the power word embedding model, and use the power word embedding model to generate the embedding representation of each word in the query statement; add a first linear layer after the power word embedding model to form a first semantic classifier; use the embedding representation of each word to train the first semantic classifier; through the first semantic classifier, the semantic integrity of the word can be obtained, that is, the word-level loss L word ; Solidify the parameters of the power sentence embedding model, and use the power sentence embedding model to generate the embedding representation of the entire query statement; add a second linear layer after the power sentence embedding model to form a second semantic classifier; use the embedding representation of the entire query statement to train the second semantic classifier; through the second semantic classifier, the semantic integrity of the sentence can be obtained, that is, the sentence-level loss L query ; Combine the word-level loss L word and the sentence-level loss L query , to form a comprehensive loss function Loss: Loss = w1L word + w2L query w1 and w2 are weight coefficients; According to the value of the comprehensive loss function Loss, execute the backpropagation algorithm to update the parameters of the first linear layer and the second linear layer until the comprehensive loss function Loss converges, and obtain the trained first semantic classifier and second semantic classifier.
[0012] A further improvement of the present invention lies in: the step of using the pre-constructed error correction dictionary in the power field to match the text of the preliminary power query problem, and if an error correction phrase is matched; inputting the text of the preliminary power query problem and the error correction phrase into the pre-established power field error correction model to obtain the text of the finally error-corrected power query problem specifically includes: Use the pre-constructed error correction dictionary in the power field to process the text query of the preliminary power query problem. For the first word word1 in the text query of the preliminary power query problem, if the second word word2 in the power field error correction dictionary is matched; input the first word word1 into the first semantic classifier to obtain the first word-level semantic integrity L word1 , input the text query of the preliminary power query problem into the second semantic classifier to obtain the first sentence-level semantic integrity L query1 ; if w1L word1 + w2L query1 < the first threshold, then replace the first word word1 in the text query of the preliminary power query problem with the second word word2 to obtain the text query of the error-corrected power query problem; otherwise, output the text query of the preliminary power query problem as the text of the finally error-corrected power query problem; Input the second word word2 into the first semantic classifier to obtain the second word-level semantic integrity L word2 Input the text query of the corrected power query problem into the second semantic classifier to obtain the second sentence-level semantic integrity L query2 ; If w1L word2 + w2L query2 > the second threshold, take the text query of the corrected power query problem as the text of the finally corrected power query problem; otherwise, output the text query of the preliminary power query problem as the text of the finally corrected power query problem.
[0013] In a second aspect, the present invention provides a power speech recognition device based on a large power speech recognition model, including: An acquisition module for acquiring the audio of the power query problem; A speech recognition module for inputting the audio of the power query problem into a pre-established general speech recognition large model for speech recognition to obtain the text of the preliminary power query problem; An error correction module for matching the text of the preliminary power query problem with a pre-constructed power domain error correction dictionary. If an error correction phrase is matched, input the text of the preliminary power query problem and the error correction phrase into a pre-established power domain error correction model to obtain the text of the finally corrected power query problem.
[0014] In a third aspect, the present invention provides an electronic device, including a processor and a memory. The processor is used to execute a computer program stored in the memory to implement the power speech recognition method based on the general speech recognition large model and the power error correction model.
[0015] In a fourth aspect, the present invention provides a computer-readable storage medium. The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, the power speech recognition method is implemented.
[0016] In a fifth aspect, the present invention provides a computer program product. The computer program product includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the power speech recognition method is implemented.
[0017] Compared with the prior art, the present invention has the following unexpected technical effects: The present invention provides a power speech recognition method based on a general speech recognition large model and a power error correction model, including: obtaining the audio of a power query problem; inputting the audio of the power query problem into a pre-established general speech recognition large model for speech recognition to obtain the text of a preliminary power query problem; using a pre-constructed power domain error correction dictionary to match the text of the preliminary power query problem, and if an error correction phrase is matched; inputting the text of the preliminary power query problem and the error correction phrase into a pre-established power domain error correction model to obtain the text of the finally error-corrected power query problem. The present invention innovatively adds a power error correction model after the speech recognition model, which is more suitable for speech recognition in the power domain; although the general speech recognition model can better recognize speech in different domains, the recognition accuracy for specific professional power speech is lower than that in ordinary domains; adding a power error correction model can perform professional correction in the power domain on the recognized text, significantly improving the recognition accuracy of the speech.
[0018] Further, the power domain error correction dictionary in the present invention is constructed through the following steps: S1. Input the Chinese text of the power query problem into a text-to-audio model to generate the Chinese audio of the power query problem; S2. Input the Chinese audio of the power query problem obtained in S1 into an audio-to-text model to identify and output the final Chinese text of the power query problem; S3. Compare the Chinese text of the power query problem input in S1 with the output in S2. If the two texts are different, extract the correct phrases and incorrect phrases of the difference as example entries in the error correction dictionary; S4. Repeat steps S1 to S3 to obtain several example entries in the error correction dictionary, and construct the power domain error correction dictionary. For existing speech recognition models, such as OpenAI Whisper, if they need to be used in the power domain with high recognition accuracy, a large amount of labeled Chinese audio data needs to be manually constructed for training. This method has problems such as high training cost and low efficiency, and it is difficult to meet the requirements of actual applications; the present invention can automatically generate example entries in the error correction dictionary by using the existing Chinese text library of power query problems through the construction of the power domain error correction dictionary, saving the time and effort of manual annotation, and can quickly construct the power domain error correction dictionary; through the power domain error correction dictionary, the errors in the recognition of professional vocabulary in the power domain by the general speech recognition model can be quickly corrected, improving the recognition accuracy.
[0019] Further, the multi-task learning loss in the present invention: A new context-based multi-task learning loss for power text error correction is proposed, which improves the accuracy and robustness of error correction through the combination of Lword and Lquery.
[0020] Further, the present invention realizes negative sample enhancement in contrastive learning, and improves the generalization ability and error correction effect of the model by introducing difficult negative samples.
[0021] Furthermore, the present invention ensures the accuracy of the error correction dictionary in the power field and the semantic integrity of power query statements through multi-task learning and contrastive learning methods. At the same time, the negative sample enhancement strategy effectively improves the robustness and generalization ability of the model, ensuring the practicability and reliability of the technical solution.
[0022] Furthermore, the present invention not only automatically constructs an error correction dictionary in the power field, but also significantly improves the accuracy and robustness of power speech recognition through multi-task learning and contrastive learning methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is a schematic flowchart of constructing an error correction dictionary in the power field according to an embodiment of the present invention; Figure 2 is a schematic flowchart of training a multi-task power error correction model according to an embodiment of the present invention; Figure 3 is a schematic flowchart of a power speech recognition method based on a large power speech recognition model according to an embodiment of the present invention; Figure 4 is a schematic flowchart of a power speech recognition method based on a large power speech recognition model according to another embodiment of the present invention; Figure 5 is a schematic structural diagram of a power speech recognition device based on a large power speech recognition model according to an embodiment of the present invention; Figure 6 is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments. It should be noted that, without conflict, the embodiments and features in the embodiments of the present invention may be combined with each other.
[0025] The following detailed descriptions are all exemplary descriptions, aiming to provide further details of the present invention. Unless otherwise specified, all technical terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the art to which the present invention belongs. The terms used in the present invention are only for describing specific embodiments and are not intended to limit the exemplary embodiments of the present invention.
[0026] An embodiment of the present invention provides a method for constructing an error correction dictionary in the power field, which significantly reduces the manual construction cost.
[0027] The construction process of the error correction dictionary in the power field is as follows Figure 1 shown, including the following steps: S11. Generate power query audio: Input the Chinese text of the power query question, such as "What is the secondary converter of an electronic voltage transformer?", into the text-to-audio model to generate clear Chinese audio of the power query question. In a specific embodiment, the text-to-audio model can adopt edge-tts, or MaryTTS, eSpeak, etc.
[0028] S12. Speech recognition: Input the Chinese audio of the power query question obtained in S1 into the audio-to-text model, and recognize and output the power query question text - "What is the secondary converter of an electronic voltage protection rod device", and the error recognition result of the speech recognition model for power terms can be obtained, where "mutual inductance" is recognized as "protection rod". In a specific embodiment, the audio-to-text model can adopt OpenAI Whisper, etc.
[0029] Whisper is a pre-trained model for automatic speech recognition, released by Alec Radford et al. from OpenAI in September 2022. Different from previous works such as Wav2Vec 2.0, previous models were pre-trained on unlabeled audio data, while Whisper was pre-trained on a large amount of labeled audio transcription data. The duration of the labeled audio used for its training is as high as 680,000 hours, which is more than an order of magnitude more than the unlabeled training data (60,000 hours) used by Wav2Vec 2.0. In addition, the pre-training data also contains 117,000 hours of multilingual data, and the checkpoint trained by Whisper can be applied to more than 96 languages including Chinese.
[0030] The Whisper model structure is an encoder-decoder transformer, with residual connections before and after the middle attention module. The encoder uses self-attention, and the decoder uses cross-attention. First, the original audio input is transformed into a log-Mel spectrogram through a feature extractor. Then, the transformer encoder encodes the spectrogram to generate a series of encoder hidden states. Finally, the decoder autoregressively predicts the next output word based on the previously output words and the encoder hidden states.
[0031] S13. Construct an error correction dictionary in the power field: Compare the Chinese text of the power query problem between the input of S11 and the output of S12. If there are differences between the two texts, extract the correct phrases and incorrect phrases of the differences as example entries in the error correction dictionary. In a specific embodiment, an example entry in the error correction dictionary (word1, word2) is ('voltage protection rod device', 'voltage transformer').
[0032] Repeat steps S11 to S13 to obtain several example entries in the error correction dictionary, and construct an error correction dictionary for the power field.
[0033] An embodiment of the present invention proposes an error correction model for the power field based on multi-task and contrastive learning.
[0034] If the constructed error correction dictionary for the power field is directly used for error correction of power texts, the error correction error rate is relatively high. For example, for the example entry in the error correction dictionary ('compound', 'load'), it is necessary to consider the semantic integrity of the context of the power query text to determine whether to replace the word 'compound' in the power query text with 'load'.
[0035] Therefore, to further optimize the error correction dictionary for the power field, it is necessary to construct a multi-task large language model. This model can not only learn the professional terms in the power field but also understand the usage of these terms in different contexts, so as to perform error correction more accurately.
[0036] An embodiment of the present invention designs a multi-task learning framework containing two sub-tasks, and its loss function is defined as: Loss = w1L word + w2L query where L word is the semantic integrity loss for a single word, and L query is the semantic integrity loss for the entire query statement. The weights w1 and w2 are pre-configured according to the importance of the specific task. Preferably, w1 + w2 = 1.
[0037] L word is used to evaluate whether each word in the power error correction dictionary has a complete semantic meaning. For example, for the entry in the error correction dictionary “('voltage protection rod device', 'voltage transformer')”, the model needs to determine whether 'voltage protection rod device' is a reasonable power term. If this word does not conform to the common usage in the power field, it is considered that its semantics is incomplete and the loss value will increase.
[0038]
[0039] To achieve this, embodiments of the present invention use the contrastive learning method to train a power vocabulary embedding model; to train the power vocabulary embedding model, positive samples and negative samples need to be prepared. Positive samples are correct power vocabulary, and negative samples are randomly selected power vocabulary different from the positive samples. In addition, hard negative samples are words with speech recognition errors. For example, "voltage transformer" misrecognized as "voltage protection rod device" can be used as a hard negative sample. The specific approach is as follows: Positive samples: Correct power words, such as "voltage transformer".
[0040] Random negative samples: Randomly selected power words different from the positive samples, such as "current transformer".
[0041] Hard negative samples: Generated from words with speech recognition errors, such as "voltage protection rod device".
[0042] L query Used to evaluate whether the entire query statement is semantically complete when putting the words in the error correction dictionary into the original power query problem. For example, after replacing "voltage protection rod device" back with "voltage transformer", the model needs to determine whether the entire query statement "What is the secondary converter of an electronic voltage transformer?" is semantically correct.
[0043]
[0044] To achieve this, the present invention uses the contrastive learning method to retrain a power sentence embedding model. To train the power sentence embedding model, positive samples and negative samples need to be prepared. Positive samples are correct power sentences, and negative samples are randomly selected power sentences different from the positive samples. In addition, hard negative samples are generated from sentences with speech recognition errors. For example, a sentence with "voltage transformer" misrecognized as "voltage protection rod device" can be used as a hard negative sample. The specific approach is as follows: Positive samples: Correct power sentences, such as "What is the secondary converter of an electronic voltage transformer?".
[0045] Random negative samples: Randomly selected power sentences different from the positive samples, such as "What is the secondary converter of an electronic current transformer?".
[0046] Hard negative samples: Generated from sentences with speech recognition errors, such as "What is the secondary converter of an electronic voltage protection rod device?".
[0047] Contrastive loss function:
[0048] y is the label (y = 1 corresponds to the positive sample pair, y = 0 corresponds to the negative sample pair).
[0049] d is the distance between the two samples.
[0050] m is a preset boundary value used to ensure that the distance of negative sample pairs is at least m.
[0051] Train the power vocabulary embedding model and the power sentence embedding model respectively based on the contrast loss function until the contrast loss function converges, and obtain the well-trained power vocabulary embedding model and power sentence embedding model through contrastive learning.
[0052] Multi-task learning loss: The present invention proposes a new context-based multi-task learning loss for power text error correction, which improves the accuracy and robustness of error correction through the combination of Lword and Lquery.
[0053] L word is the semantic integrity loss for a single word; L query is the semantic integrity loss for the entire query sentence;
[0054]
[0055] The loss function of multi-task learning is: Loss = w1L word + w2L query Negative sample enhancement: The present invention realizes negative sample enhancement in contrastive learning, and improves the generalization ability and error correction effect of the model by introducing hard negative samples.
[0056] Through the methods of multi-task learning and contrastive learning, the present invention ensures the accuracy of the error correction dictionary in the power field and the semantic integrity of power query sentences. At the same time, the negative sample enhancement strategy effectively improves the robustness and generalization ability of the model, ensuring the practicability and reliability of the technical solution.
[0057] Please refer to Figure 2 As shown, the embodiment of the present invention provides a multi-task power field error correction model training method, which specifically includes: S21. Data preprocessing: Collect text data in the power field, including correct power texts and error texts generated by speech recognition; generate positive samples, random negative samples and hard negative samples based on the collected text data in the power field to obtain preprocessed text data.
[0058] For vocabulary, the positive sample is the correct power vocabulary, and the negative sample is a randomly selected power vocabulary different from the positive sample; the hard negative sample is the vocabulary with speech recognition errors. For example, if "voltage transformer" is misrecognized as "voltage protection rod device", it can be used as a hard negative sample. The specific method is as follows: Positive samples: Correct power-related terms, such as "voltage transformer".
[0059] Random negative samples: Randomly selected power-related terms different from the positive samples, such as "current transformer".
[0060] Difficult negative samples: Generated from words with speech recognition errors, such as "voltage protection rod device".
[0061] For sentences, the positive samples are correct power-related sentences, and the negative samples are randomly selected power-related sentences different from the positive samples; the difficult negative samples are generated from sentences with speech recognition errors. For example, a sentence with the incorrect recognition of "voltage transformer" as "voltage protection rod device" can be used as a difficult negative sample. The specific method is as follows: Positive samples: Correct power-related sentences, such as "What is the secondary converter of an electronic voltage transformer?"
[0062] Random negative samples: Randomly selected power-related sentences different from the positive samples, such as "What is the secondary converter of an electronic current transformer?"
[0063] Difficult negative samples: Generated from sentences with speech recognition errors, such as "What is the secondary converter of an electronic voltage protection rod device?"
[0064] S22. Input layer reception: Input the preprocessed text data into the input layer of the multi-task power domain error correction model; S23. Embedding layer processing: Convert the preprocessed text data received by the input layer into vectors to obtain output vectors so that the model can process them; S24. Encoder layer processing: Use the multi-head attention mechanism, feed-forward neural network, and layer normalization to extract the feature representations of the output vectors of the embedding layer; S25. Contrastive learning layer processing: LLM encoding: Use the first Large Language Model (LLM for short, the first LLM model) to encode the feature representations of the positive, negative, and difficult negative samples of vocabulary to generate vocabulary embedding vectors; use the second Large Language Model (LLM for short, the second LLM model) to encode the feature representations of the positive, negative, and difficult negative samples of sentences to generate sentence embedding vectors; Similarity calculation: Perform contrastive learning on the vocabulary embedding vectors generated by the first LLM model, and calculate the cosine similarity (or other similarity metrics) of the positive and negative sample pairs respectively. Contrastive loss calculation: Based on the results of the similarity calculation, calculate the contrastive loss L contrastive, ensure that the distance between positive sample pairs is maximized and the distance between negative sample pairs is minimized; iteratively train and continuously update the first LLM model until the contrastive loss L contrastive converges to obtain an electric power word embedding model.
[0065] Similarity calculation: For the sentence embedding vectors generated by the second LLM model, perform contrastive learning and calculate the cosine similarity (or other similarity metrics) of positive and negative sample pairs respectively. Contrastive loss calculation: Based on the results of similarity calculation, calculate the contrastive loss L contrastive , ensure that the distance between positive sample pairs is maximized and the distance between negative sample pairs is minimized; iteratively train and continuously update the second LLM model until the contrastive loss L contrastive converges to obtain an electric power sentence embedding model.
[0066] S26. Multi-task learning layer processing: Freeze the parameters of the contrastive learning layer (parameters of the electric power word embedding model and sentence word embedding model), and train the multi-task learning layer.
[0067] 6.1. Task 1: Word-level semantic integrity judgment Word embedding: Fix the parameters of the electric power word embedding model, and use the electric power word embedding model (obtained by training the above contrastive learning layer) to generate the embedding representation of each word in the query sentence.
[0068] Semantic classifier: Add a first linear layer after the electric power word embedding model to form a first semantic classifier; use the embedding representation of each word to train the first semantic classifier; through the first semantic classifier, the semantic integrity of the word can be obtained, that is, the word-level loss L word , which is used to judge whether the word has a complete semantic meaning.
[0069] Word-level loss L word Calculation: Calculate the semantic integrity loss at the word level to supervise the learning process of the model; .
[0070] 6.2. Task 2: Sentence-level semantic integrity judgment Sentence embedding: Fix the parameters of the electric power sentence embedding model, and use the electric power sentence embedding model (obtained by training the above contrastive learning layer) to generate the embedding representation of the entire query sentence.
[0071] Semantic classifier: Add a second linear layer after the electric power sentence embedding model to form a second semantic classifier; use the embedding representation of the entire query sentence to train the second semantic classifier; through the second semantic classifier, the semantic integrity of the sentence can be obtained, that is, the sentence-level loss L query , which is used to judge whether the entire query sentence is semantically complete.
[0072] Sentence-level loss L query Calculation: Calculate the semantic integrity loss at the sentence level to further optimize the model; 。
[0073] 6.3. Multi-task learning loss Combine the word-level loss L word and the sentence-level loss L query to form a comprehensive loss function Loss, ensuring that the error correction model in the power domain performs well on multiple tasks; Loss = w1L word + w2L query 7. Backpropagation According to the value of the comprehensive loss function Loss, execute the backpropagation algorithm to update the parameters of the first linear layer and the second linear layer to minimize the loss and improve the generalization ability of the error correction model in the power domain.
[0074] Please refer to Figure 3 As shown, an embodiment of the present invention provides a power speech recognition method, including: Obtain the audio of the power query problem; Input the audio of the power query problem into a pre-established general speech recognition large model for speech recognition to obtain the text query of the preliminary power query problem; Process the text query of the preliminary power query problem by using a pre-constructed error correction dictionary in the power domain, and match the first word word1 in the text query of the preliminary power query problem to the second word word2 in the error correction dictionary in the power domain; Input the text query of the preliminary power query problem, the first word word1, and the second word word2 into a pre-established error correction model in the power domain. The first word word1 is input into the first semantic classifier to obtain the first word-level semantic integrity L word1 , and the text query of the preliminary power query problem is input into the second semantic classifier to obtain the first sentence-level semantic integrity L query1 ; If w1L word1 + w2L query1 < the first threshold, then replace the first word word1 in the text query of the preliminary power query problem with the second word word2 to obtain the text query of the error-corrected power query problem; otherwise, output the text query of the preliminary power query problem as the text of the finally error-corrected power query problem; Input the text query of the error-corrected power query problem and the second word word2 into a pre-established error-correction model in the power field. Input the second word word2 into the first semantic classifier to obtain the second word-level semantic integrity L word2 Input the text query of the error-corrected power query problem into the second semantic classifier to obtain the second sentence-level semantic integrity L query2 ; If w1L word2 + w2L query2 > the second threshold, take the text query of the error-corrected power query problem as the text of the finally error-corrected power query problem; otherwise, output the text query of the preliminary power query problem as the text of the finally error-corrected power query problem.
[0075] Please refer to Figure 4 as shown, an embodiment of the present invention provides a power speech recognition method, including: S100. Obtain the audio of the power query problem; S200. Input the audio of the power query problem into a pre-established general speech recognition large model for speech recognition to obtain the text of the preliminary power query problem; S300. Use a pre-constructed error-correction dictionary in the power field to match the text of the preliminary power query problem. If an error-correction phrase is matched; input the text of the preliminary power query problem and the error-correction phrase into a pre-established error-correction model in the power field to obtain the text of the finally error-corrected power query problem.
[0076] The present invention innovatively adds a power error-correction model after the speech recognition model, which is more suitable for speech recognition in the power field; although the general speech recognition model can better recognize speech in different fields, the correct recognition rate for specific professional power speech is lower than that in ordinary fields; adding a power error-correction model can perform professional correction in the power field for the recognized text, significantly improving the correct recognition rate of speech.
[0077] In a specific embodiment, the error-correction dictionary in the power field is constructed through the following steps: Input the Chinese text of the original power query problem into a text-to-audio model to generate the Chinese audio of the power query problem; Input the Chinese audio of the power query problem into an audio-to-text model to recognize and output the final Chinese text of the power query problem; Compare the Chinese text of the original power query problem with the final Chinese text of the power query problem. If the two texts are different, extract the correct phrases and incorrect phrases of the differences as example entries in the error-correction dictionary; Collect several example entries in the error-correction dictionary to construct an error-correction dictionary in the power field.
[0078] In a specific embodiment, in the step of inputting the Chinese text of the original power query question into the text-to-audio model to generate the Chinese audio of the power query question, the edge-tts is adopted for the text-to-audio model; In the step of inputting the Chinese audio of the power query question into the audio-to-text model to recognize and output the Chinese text of the final power query question, the Whisper is adopted for the audio-to-text model.
[0079] In a specific embodiment, the pre-established power domain error correction model is obtained through the following training steps: Collect the text data in the power domain, including correct power texts and error texts generated by speech recognition; generate positive samples, random negative samples, and difficult negative samples based on the collected text data in the power domain to obtain the preprocessed text data; Convert the preprocessed text data into vectors to obtain output vectors; Extract the feature representations of the output vectors; Use the first large language model to encode the feature representations of the positive samples, negative samples, and difficult negative samples of the vocabulary to generate vocabulary embedding vectors; use the second large language model to encode the feature representations of the positive samples, negative samples, and difficult negative samples of the sentences to generate sentence embedding vectors; Use the vocabulary embedding vectors to train the first large language model to obtain a power word embedding model; use the sentence embedding vectors to train the second large language model to obtain a sentence word embedding model; Solidify the parameters of the power word embedding model, and train a first semantic classifier for evaluating the semantic integrity of words based on the power word embedding model; solidify the parameters of the power sentence embedding model, and train a second semantic classifier for evaluating the semantic integrity of sentences based on the power sentence embedding model.
[0080] In a specific embodiment, in the step of using the first large language model to encode the feature representations of the positive samples, negative samples, and difficult negative samples of the vocabulary to generate vocabulary embedding vectors, for the vocabulary, the positive samples are correct power vocabulary, and the negative samples are randomly selected power vocabulary different from the positive samples; the difficult negative samples are the vocabulary with speech recognition errors; In the step of using the second large language model to encode the feature representations of the positive samples, negative samples, and difficult negative samples of the sentences to generate sentence embedding vectors, for the sentences, the positive samples are correct power sentences, and the negative samples are randomly selected power sentences different from the positive samples; the difficult negative samples are generated from the sentences with speech recognition errors.
[0081] In a specific embodiment, the steps of training a first large language model using lexical embedding vectors to obtain a power word embedding model, and training a second large language model using sentence embedding vectors to obtain a sentence word embedding model specifically include: Encode the feature representations of the positive samples, negative samples, and hard negative samples of the vocabulary using the first large language model to generate lexical embedding vectors; perform contrastive learning on the lexical embedding vectors generated by the first large language model, calculate the cosine similarity between the positive and negative sample pairs respectively, and calculate the first contrast loss L based on the result of the similarity calculation contrastive , iteratively train until the first contrast loss L contrastive converges to obtain a power word embedding model; Encode the feature representations of the positive samples, negative samples, and hard negative samples of the sentence using the second large language model to generate sentence embedding vectors; perform contrastive learning on the sentence embedding vectors generated by the second large language model, calculate the cosine similarity between the positive and negative sample pairs respectively, and calculate the second contrast loss L based on the result of the similarity calculation contrastive , iteratively train until the second contrast loss L contrastive converges to obtain a sentence word embedding model.
[0082] In a specific embodiment, the steps of solidifying the parameters of the power word embedding model and training a first semantic classifier for evaluating the semantic integrity of words based on the power word embedding model, and solidifying the parameters of the power sentence embedding model and training a second semantic classifier for evaluating the semantic integrity of sentences based on the power sentence embedding model specifically include: Solidify the parameters of the power word embedding model, and use the power word embedding model to generate the embedding representation of each word in the query sentence; add a first linear layer after the power word embedding model to form a first semantic classifier; use the embedding representation of each word to train the first semantic classifier; the semantic integrity of the word can be obtained through the first semantic classifier, that is, the word-level loss L word ; Solidify the parameters of the power sentence embedding model, and use the power sentence embedding model to generate the embedding representation of the entire query sentence; add a second linear layer after the power sentence embedding model to form a second semantic classifier; use the embedding representation of the entire query sentence to train the second semantic classifier; the semantic integrity of the sentence can be obtained through the second semantic classifier, that is, the sentence-level loss L query ; Combine the word-level loss L word and the sentence-level loss L query to form a comprehensive loss function Loss: Loss = w1L word + w2L query w1 and w2 are weight coefficients; According to the value of the comprehensive loss function Loss, perform the backpropagation algorithm to update the parameters of the first linear layer and the second linear layer until the comprehensive loss function Loss converges, and obtain the trained first semantic classifier and the second semantic classifier.
[0083] In a specific embodiment, the step of using the pre-constructed error correction dictionary in the power field to match the text of the preliminary power query problem, and if an error correction phrase is matched; inputting the text of the preliminary power query problem and the error correction phrase into the pre-established error correction model in the power field to obtain the text of the finally error-corrected power query problem specifically includes: Use the pre-constructed error correction dictionary in the power field to process the text query of the preliminary power query problem. For the first word word1 in the text query of the preliminary power query problem, if the second word word2 is matched in the error correction dictionary in the power field; input the first word word1 into the first semantic classifier to obtain the first word-level semantic integrity L word1 , input the text query of the preliminary power query problem into the second semantic classifier to obtain the first sentence-level semantic integrity L query1 ; if w1L word1 + w2L query1 <the first threshold, then replace the first word word1 in the text query of the preliminary power query problem with the second word word2 to obtain the text query of the error-corrected power query problem; otherwise, output the text query of the preliminary power query problem as the text of the finally error-corrected power query problem; Input the second word word2 into the first semantic classifier to obtain the second word-level semantic integrity L word2 , input the text query of the error-corrected power query problem into the second semantic classifier to obtain the second sentence-level semantic integrity L query2 ; if w1L word2 + w2L query2 >the second threshold, take the text query of the error-corrected power query problem as the text of the finally error-corrected power query problem; otherwise, output the text query of the preliminary power query problem as the text of the finally error-corrected power query problem.
[0084] Please refer to Figure 5 As shown, the embodiment of the present invention provides a power speech recognition device based on a general speech recognition large model and a power error correction model, including: An acquisition module for acquiring the audio of the power query problem; A speech recognition module for inputting the audio of the power query problem into the pre-established general speech recognition large model for speech recognition to obtain the text of the preliminary power query problem; An error correction module is configured to match the text of the preliminary power query problem with a pre-constructed error correction dictionary in the power field. If an error correction phrase is matched, the text of the preliminary power query problem and the error correction phrase are input into a pre-established error correction model in the power field to obtain the text of the finally error-corrected power query problem.
[0085] Please refer to Figure 6 As shown, an electronic device 100 for implementing a power voice recognition method based on a general speech recognition large model and a power error correction model is provided in an embodiment of the present invention. The electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.
[0086] The memory 101 can be used to store the computer program 103. The processor 102 realizes the steps of the power voice recognition method by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device 100 (such as audio data), etc. In addition, the memory 101 may include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0087] The at least one processor 102 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or the processor 102 may also be any conventional processor, etc. The processor 102 is the control center of the electronic device 100 and connects various parts of the entire electronic device 100 through various interfaces and lines.
[0088] The memory 101 in the electronic device 100 stores a plurality of instructions to implement a power voice recognition method based on a general voice recognition large model and a power error correction model. The processor 102 can execute the plurality of instructions to implement: Obtain the audio of the power query problem; Input the audio of the power query problem into a pre-established general voice recognition large model for voice recognition to obtain the text of the preliminary power query problem; Use a pre-constructed power domain error correction dictionary to match the text of the preliminary power query problem. If an error correction phrase is matched, input the text of the preliminary power query problem and the error correction phrase into a pre-established power domain error correction model to obtain the text of the finally error-corrected power query problem.
[0089] If the module / unit integrated in the electronic device 100 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, and read-only memory (ROM, Read-Only Memory).
[0090] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0091] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0092] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0093] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A power speech recognition method based on a general speech recognition large model and a power error correction model, characterized in that: include: Get audio of power query questions; Input the audio of the power inquiry question into a pre-established general speech recognition large model to perform speech recognition to obtain a text of the preliminary power inquiry question; Use the pre-built error correction dictionary in the power field to match the text of the preliminary power query question, if the error correction phrase is matched; The text of the preliminary power query question and the error correction phrase are input into a pre-established power field error correction model to obtain the text of the final power query question after error correction.
2. The electric power speech recognition method based on the general speech recognition large model and the electric power error correction model according to claim 1 is characterized in that: The power field error correction dictionary is constructed by the following steps: To the text-to-audio model, input the Chinese text of the original power query question and generate the Chinese audio of the power query question; Inputting the Chinese audio of the power inquiry question into the audio-to-text model to recognize and output the final Chinese text of the power inquiry question; Compare the Chinese text of the original power query question with the Chinese text of the final power query question, and if there are differences between the two texts, extract the correct phrases and incorrect phrases of the differences as sample entries of the error correction dictionary; Collect several error correction dictionary example entries and construct an error correction dictionary in the power field.
3. The electric power speech recognition method based on the general speech recognition large model and the electric power error correction model according to claim 2 is characterized in that: In the step of inputting the Chinese text of the original power query question into the text-to-audio model and generating the Chinese audio of the power query question, the text-to-audio model adopts edge-tts; In the step of inputting the Chinese audio of the power inquiry question into the audio-to-text model to recognize and output the final Chinese text of the power inquiry question, the audio-to-text model adopts Whisper.
4. The electric power speech recognition method based on the general speech recognition large model and the electric power error correction model according to claim 1 is characterized in that: The pre-established power field error correction model is trained by the following steps: Collect text data in the power field, including correct power text and erroneous text generated by speech recognition; Based on the collected text data in the power field, positive samples, random negative samples and difficult negative samples are generated to obtain preprocessed text data; Convert the preprocessed text data into a vector to obtain an output vector; Extracting feature representation of the output vector; The first large language model is used to encode the feature representations of the positive samples, negative samples, and difficult negative samples of the vocabulary to generate a vocabulary embedding vector; the second large language model is used to encode the feature representations of the positive samples, negative samples, and difficult negative samples of the sentence to generate a sentence embedding vector; The first large language model is trained using vocabulary embedding vectors to obtain an electric power word embedding model; Using sentence embedding vectors to train a second large language model to obtain a sentence word embedding model; Solidify the parameters of the electric power word embedding model, and obtain a first semantic classifier for evaluating the semantic integrity of words based on the training of the electric power word embedding model; The parameters of the electric power sentence embedding model are solidified, and a second semantic classifier for evaluating the semantic completeness of a sentence is obtained based on the training of the electric power sentence embedding model.
5. The electric power speech recognition method based on the general speech recognition large model and the electric power error correction model according to claim 4 is characterized in that: In the step of encoding the feature representations of the positive samples, negative samples and difficult negative samples of the vocabulary using the first large language model to generate the vocabulary embedding vector, for the vocabulary, the positive samples are correct power vocabulary, and the negative samples are randomly selected power vocabulary different from the positive samples; the difficult negative samples are vocabulary with speech recognition errors; In the step of encoding the feature representations of the positive samples, negative samples and difficult negative samples of the sentences using the second large language model to generate sentence embedding vectors, for the sentences, the positive samples are correct electricity sentences, and the negative samples are randomly selected electricity sentences different from the positive samples; Hard negative samples are generated from sentences with speech recognition errors.
6. The electric power speech recognition method based on the general speech recognition large model and the electric power error correction model according to claim 4 is characterized in that: The first large language model is trained by using the vocabulary embedding vector to obtain an electric power word embedding model; The steps of training the second large language model using the sentence embedding vector to obtain the sentence word embedding model specifically include: Encode the feature representations of the positive samples, negative samples, and difficult negative samples of the vocabulary using the first large language model to generate a vocabulary embedding vector; For the vocabulary embedding vector generated by the first large language model, perform contrastive learning, calculate the cosine similarity between the positive and negative sample pairs, and calculate the first contrast loss L based on the similarity calculation results. contrastive , iterative training to the first contrast loss L contrastive Converge and obtain the power word embedding model; The second large language model is used to encode the feature representations of the positive samples, negative samples, and difficult negative samples of the sentence to generate a sentence embedding vector; the sentence embedding vector generated by the second large language model is subjected to comparative learning, and the cosine similarity between the positive and negative sample pairs is calculated respectively. Based on the results of the similarity calculation, the second comparative loss L is calculated. contrastive , iterative training to the second contrast loss L contrastive Converge and obtain the sentence word embedding model.
7. The electric power speech recognition method based on the general speech recognition large model and the electric power error correction model according to claim 4 is characterized in that: The solidified electric power word embedding model parameters are used to obtain a first semantic classifier for evaluating the semantic integrity of a word based on the electric power word embedding model training; The steps of solidifying the electric power sentence embedding model parameters and obtaining a second semantic classifier for evaluating the semantic integrity of a sentence based on the electric power sentence embedding model training specifically include: Solidify the parameters of the electric power word embedding model, use the electric power word embedding model to generate the embedding representation of each word in the query sentence; add the first linear layer after the electric power word embedding model to form the first semantic classifier; use the embedding representation of each word to train the first semantic classifier; the semantic completeness of the word can be obtained through the first semantic classifier, that is, the word-level loss L word ; The parameters of the electric power sentence embedding model are fixed, and the electric power sentence embedding model is used to generate the embedding representation of the entire query sentence; a second linear layer is added after the electric power sentence embedding model to form a second semantic classifier; the second semantic classifier is trained using the embedding representation of the entire query sentence; the semantic completeness of the sentence, i.e., the sentence-level loss L, can be obtained through the second semantic classifier. query ; Combined with word-level loss L word and sentence-level loss L query , forming a comprehensive loss function Loss: Loss=w1L word + w2L query w1 and w2 are weight coefficients; According to the value of the comprehensive loss function Loss, the back propagation algorithm is executed to update the parameters of the first linear layer and the second linear layer until the comprehensive loss function Loss converges to obtain the trained first semantic classifier and the second semantic classifier.
8. The electric power speech recognition method based on the general speech recognition large model and the electric power error correction model according to claim 4 is characterized in that: The pre-built error correction dictionary in the electric power field is used to match the text of the preliminary electric power query question, and if an error correction phrase is matched; The step of inputting the text of the preliminary power query question and the error correction phrase into a pre-established power field error correction model to obtain the final error-corrected power query question text specifically includes: The text query of the preliminary power query question is processed using the pre-built power field error correction dictionary. If the first word word1 in the text query of the preliminary power query question matches the second word word2 in the power field error correction dictionary; the first word word1 is input into the first semantic classifier to obtain the first word level semantic completeness L word1 , the text query of the preliminary power query question is input into the second semantic classifier to obtain the first sentence level semantic completeness L query1 ; if w1L word1 + w2L query1 < the first threshold, then the first word word1 in the text query of the preliminary power query question is replaced with the second word word2 to obtain the text query of the corrected power query question; otherwise, the text query of the preliminary power query question is output as the text of the final corrected power query question; Input the second word word2 into the first semantic classifier to obtain the second word-level semantic completeness L word2 , input the text query of the error correction power query problem into the second semantic classifier to obtain the second sentence level semantic completeness L query2 ; if w1L word2 + w2L query2 >The second threshold, the text query of the error-corrected power query question is used as the text of the final error-corrected power query question; otherwise, the text query of the preliminary power query question is output as the text of the final error-corrected power query question.
9. An electric power speech recognition device based on a general speech recognition large model and an electric power error correction model, characterized in that: include: An acquisition module is used to obtain audio of power query questions; A speech recognition module, used to input the audio of the power inquiry question into a pre-established general speech recognition large model, perform speech recognition, and obtain the text of the preliminary power inquiry question; The error correction module is used to match the text of the preliminary power query question with a pre-built error correction dictionary in the power field, and if an error correction phrase is matched; The text of the preliminary power query question and the error correction phrase are input into a pre-established power field error correction model to obtain the text of the final power query question after error correction.
10. The electric power speech recognition device based on the general speech recognition large model and the electric power error correction model according to claim 9, characterized in that: The power field error correction dictionary is constructed by the following steps: To the text-to-audio model, input the Chinese text of the original power query question and generate the Chinese audio of the power query question; Inputting the Chinese audio of the power inquiry question into the audio-to-text model to recognize and output the final Chinese text of the power inquiry question; Compare the Chinese text of the original power query question with the Chinese text of the final power query question, and if there are differences between the two texts, extract the correct phrases and incorrect phrases of the differences as sample entries of the error correction dictionary; Collect several error correction dictionary example entries and construct an error correction dictionary in the power field.
11. The power speech recognition device based on the general speech recognition large model and the power error correction model according to claim 10, characterized in that: In the step of inputting the Chinese text of the original power query question into the text-to-audio model and generating the Chinese audio of the power query question, the text-to-audio model adopts edge-tts; In the step of inputting the Chinese audio of the power inquiry question into the audio-to-text model to recognize and output the final Chinese text of the power inquiry question, the audio-to-text model adopts Whisper.
12. The electric power speech recognition device based on the general speech recognition large model and the electric power error correction model according to claim 9, characterized in that: The pre-established power field error correction model is trained by the following steps: Collect text data in the power field, including correct power text and erroneous text generated by speech recognition; Based on the collected text data in the power field, positive samples, random negative samples and difficult negative samples are generated to obtain preprocessed text data; Convert the preprocessed text data into a vector to obtain an output vector; Extracting feature representation of the output vector; The first large language model is used to encode the feature representations of the positive samples, negative samples, and difficult negative samples of the vocabulary to generate a vocabulary embedding vector; the second large language model is used to encode the feature representations of the positive samples, negative samples, and difficult negative samples of the sentence to generate a sentence embedding vector; The first large language model is trained using vocabulary embedding vectors to obtain an electric power word embedding model; Using sentence embedding vectors to train a second large language model to obtain a sentence word embedding model; Solidify the parameters of the electric power word embedding model, and obtain a first semantic classifier for evaluating the semantic integrity of words based on the training of the electric power word embedding model; The parameters of the electric power sentence embedding model are solidified, and a second semantic classifier for evaluating the semantic completeness of a sentence is obtained based on the training of the electric power sentence embedding model.
13. The electric power speech recognition device based on the general speech recognition large model and the electric power error correction model according to claim 12, characterized in that: In the step of encoding the feature representations of the positive samples, negative samples and difficult negative samples of the vocabulary using the first large language model to generate the vocabulary embedding vector, for the vocabulary, the positive samples are correct power vocabulary, and the negative samples are randomly selected power vocabulary different from the positive samples; the difficult negative samples are vocabulary with speech recognition errors; In the step of encoding the feature representations of the positive samples, negative samples and difficult negative samples of the sentences using the second large language model to generate sentence embedding vectors, for the sentences, the positive samples are correct electricity sentences, and the negative samples are randomly selected electricity sentences different from the positive samples; Hard negative samples are generated from sentences with speech recognition errors.
14. The electric power speech recognition device based on the general speech recognition large model and the electric power error correction model according to claim 12, characterized in that: The first large language model is trained by using the vocabulary embedding vector to obtain an electric power word embedding model; The steps of training the second large language model using the sentence embedding vector to obtain the sentence word embedding model specifically include: Encode the feature representations of the positive samples, negative samples, and difficult negative samples of the vocabulary using the first large language model to generate a vocabulary embedding vector; For the vocabulary embedding vector generated by the first large language model, perform contrastive learning, calculate the cosine similarity between the positive and negative sample pairs, and calculate the first contrast loss L based on the similarity calculation results. contrastive , iterative training to the first contrast loss L contrastive Converge and obtain the power word embedding model; The second large language model is used to encode the feature representations of the positive samples, negative samples, and difficult negative samples of the sentence to generate a sentence embedding vector; the sentence embedding vector generated by the second large language model is subjected to comparative learning, and the cosine similarity between the positive and negative sample pairs is calculated respectively. Based on the results of the similarity calculation, the second comparative loss L is calculated. contrastive , iterative training to the second contrast loss L contrastive Converge and obtain the sentence word embedding model.
15. The electric power speech recognition device based on the general speech recognition large model and the electric power error correction model according to claim 12, characterized in that: The solidified electric power word embedding model parameters are used to obtain a first semantic classifier for evaluating the semantic integrity of a word based on the electric power word embedding model training; The steps of solidifying the electric power sentence embedding model parameters and obtaining a second semantic classifier for evaluating the semantic integrity of a sentence based on the electric power sentence embedding model training specifically include: Solidify the parameters of the electric power word embedding model, use the electric power word embedding model to generate the embedding representation of each word in the query sentence; add the first linear layer after the electric power word embedding model to form the first semantic classifier; use the embedding representation of each word to train the first semantic classifier; the semantic completeness of the word can be obtained through the first semantic classifier, that is, the word-level loss L word ; The parameters of the electric power sentence embedding model are fixed, and the electric power sentence embedding model is used to generate the embedding representation of the entire query sentence; a second linear layer is added after the electric power sentence embedding model to form a second semantic classifier; the second semantic classifier is trained using the embedding representation of the entire query sentence; the semantic completeness of the sentence, i.e., the sentence-level loss L, can be obtained through the second semantic classifier. query ; Combined with word-level loss L word and sentence-level loss L query , forming a comprehensive loss function Loss: Loss=w1L word + w2L query w1 and w2 are weight coefficients; According to the value of the comprehensive loss function Loss, the back propagation algorithm is executed to update the parameters of the first linear layer and the second linear layer until the comprehensive loss function Loss converges to obtain the trained first semantic classifier and the second semantic classifier.
16. The electric power speech recognition device based on the general speech recognition large model and the electric power error correction model according to claim 12, characterized in that: The pre-built error correction dictionary in the electric power field is used to match the text of the preliminary electric power query question, and if an error correction phrase is matched; The step of inputting the text of the preliminary power query question and the error correction phrase into a pre-established power field error correction model to obtain the final error-corrected power query question text specifically includes: The text query of the preliminary power query question is processed using the pre-built power field error correction dictionary. If the first word word1 in the text query of the preliminary power query question matches the second word word2 in the power field error correction dictionary; the first word word1 is input into the first semantic classifier to obtain the first word level semantic completeness L word1 , the text query of the preliminary power query question is input into the second semantic classifier to obtain the first sentence level semantic completeness L query1 ; if w1L word1 + w2L query1 < the first threshold, then the first word word1 in the text query of the preliminary power query question is replaced with the second word word2 to obtain the text query of the corrected power query question; otherwise, the text query of the preliminary power query question is output as the text of the final corrected power query question; Input the second word word2 into the first semantic classifier to obtain the second word-level semantic completeness L word2 , input the text query of the error correction power query problem into the second semantic classifier to obtain the second sentence level semantic completeness L query2 ; if w1L word2 + w2L query2 >The second threshold, the text query of the error-corrected power query question is used as the text of the final error-corrected power query question; otherwise, the text query of the preliminary power query question is output as the text of the final error-corrected power query question.
17. An electronic device, characterized in that: It comprises a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the power speech recognition method based on a general speech recognition large model and a power error correction model as described in any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by the processor, it implements the power speech recognition method based on the general speech recognition large model and the power error correction model as described in any one of claims 1 to 8.
19. A computer program product, characterized in that The computer program product includes a computer program / instruction, and when the computer program / instruction is executed by a processor, the power speech recognition method based on a general speech recognition large model and a power error correction model as described in any one of claims 1 to 8 is implemented.