Voice Processing Method, Apparatus, Electronic Device, and Storage Medium
By using the inspection voice acoustic model and target decoding diagram in the power inspection scenario, the problem of low recognition accuracy of the general speech recognition model in the power inspection is solved, and the accurate mapping from frame-level text to words and sentences is achieved, which improves the accuracy and efficiency of the power inspection voice recognition.
Patent Information
- Application Number
- CN202211083092.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-09-06
AI Technical Summary
The existing general speech recognition model has low recognition accuracy in power inspection scenarios and cannot meet the recognition requirements of power professional terms. Moreover, the difficulty in collecting power speech data leads to low recognition accuracy.
Using a method based on the patrol speech acoustic model and target decoding diagram, by obtaining power patrol speech information, using the text mapping network, word mapping network and word sentence mapping network, the frame-level text recognition result processing is realized, mapped to words and sentences, and improving the recognition accuracy.
It improves the accuracy of voice recognition, meets the recognition needs of power inspections, and improves the efficiency and quality of inspection operations.
Smart Images

Figure CN115440224B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and particularly to a method, apparatus, electronic device, and storage medium for speech processing and determination. Background Art
[0002] With the rapid progress of artificial intelligence technology and computer performance, speech recognition technology has received increasing attention from more scholars and engineering technicians. An efficient speech recognition technology can provide effective and convenient support for speech recognition applications in different scenarios.
[0003] Currently, the speech recognition method usually uses a general speech recognition model to recognize the features of speech and determine its recognition result. However, the general speech recognition model is usually trained using speech that conforms to the daily speaking habits of the public and does not meet the speech recognition requirements in specific scenarios. Especially in the power inspection scenario, due to the existence of a large number of power professional terms, using this method will cause the problem of low speech recognition accuracy. Even if a speech recognition model is trained using power speech, due to the difficulty in collecting power speech, the trained speech recognition model will also have the problem of low recognition accuracy. Summary of the Invention
[0004] The present invention provides a speech processing method, apparatus, electronic device, and storage medium to improve the accuracy of speech recognition, meet the speech recognition requirements of power inspection, and ensure the inspection efficiency.
[0005] According to one aspect of the present invention, a speech processing method is provided. The method includes:
[0006] Obtain the inspection speech information to be recognized during the power inspection process;
[0007] Process the inspection speech information based on the inspection speech acoustic model to obtain the text recognition result corresponding to each audio frame in the inspection speech information; wherein, the inspection speech acoustic model is trained based on the pre-trained general speech model using past inspection speech data;
[0008] Process each text recognition result based on the target decoding graph to obtain the target inspection text corresponding to the inspection speech information;
[0009] Wherein, the target decoding graph is determined based on a character mapping network, a word mapping network, and a sentence mapping network. The character mapping network includes the mapping relationship between frame-level characters and characters, and the frame-level characters correspond to the text recognition results of audio frames. The word mapping network includes the mapping relationship between characters and words, and the sentence mapping network includes the mapping relationship between words and sentences.
[0010] According to another aspect of the present invention, there is provided a voice processing device, which includes:
[0011] An inspection voice information acquisition module, configured to acquire inspection voice information to be recognized during power inspection;
[0012] A text recognition result determination module, configured to process the inspection voice information based on an inspection voice acoustic model to obtain a text recognition result corresponding to each audio frame in the inspection voice information; wherein, the inspection voice acoustic model is obtained by training a pre-trained general voice model based on inspection voice data;
[0013] A target inspection text determination module, configured to process each text recognition result based on a target decoding graph to obtain a target inspection text corresponding to the inspection voice information;
[0014] Wherein, the target decoding graph is determined based on a character mapping network, a word mapping network, and a sentence mapping network. The character mapping network includes a mapping relationship between frame-level characters and characters, the frame-level characters correspond to the text recognition results of audio frames, the word mapping network includes a mapping relationship between characters and words, and the sentence mapping network includes a mapping relationship between words and sentences.
[0015] According to another aspect of the present invention, there is provided an electronic device, which includes:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores a computer program executable by the at least one processor. When the computer program is executed by the at least one processor, the at least one processor is enabled to execute the voice processing method according to any embodiment of the present invention.
[0019] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the voice processing method according to any embodiment of the present invention when executed.
[0020] The technical solution of the embodiment of the present invention obtains the inspection voice information to be recognized during the power inspection process; processes the inspection voice information based on the inspection voice acoustic model to obtain the text recognition result corresponding to each audio frame in the inspection voice information; the inspection voice acoustic model is trained based on the past inspection voice data for the pre-trained general speech model; processes each text recognition result based on the target decoding graph to obtain the target inspection text corresponding to the inspection voice information, solves the problem of low recognition accuracy in the prior art for inspecting voice recognition based on the general speech recognition model, realizes the recognition of the inspection voice information by using the inspection voice acoustic model to obtain the text recognition result corresponding to each audio frame, and processes the frame-level text recognition result by the target decoding graph, realizes the mapping from the frame-level characters to text, then from text to words, and further from words to sentences, and determines the target inspection text based on the output sentence, achieving the technical effects of improving the accuracy of speech recognition, ensuring the efficiency of inspection, meeting the requirements of power inspection voice recognition, and assisting in improving the quality and efficiency of inspection operations.
[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 is a flowchart of a speech processing method provided in Embodiment 1 of the present invention;
[0024] Figure 2 is a flowchart of a speech processing method provided in Embodiment 2 of the present invention;
[0025] Figure 3 is a flowchart of a speech processing method provided in Embodiment 3 of the present invention;
[0026] Figure 4 is a schematic diagram of a word mapping network provided in Embodiment 3 of the present invention;
[0027] Figure 5 is a schematic diagram of a sentence mapping network provided in Embodiment 3 of the present invention;
[0028] Figure 6It is a schematic diagram of a text mapping network provided in Embodiment 3 of the present invention;
[0029] Figure 7 It is a schematic diagram of a speech recognition method provided in Embodiment 4 of the present invention;
[0030] Figure 8 It is a schematic diagram of a speech recognition method provided in Embodiment 4 of the present invention;
[0031] Figure 9 It is a schematic structural diagram of a speech processing device provided in Embodiment 5 of the present invention;
[0032] Figure 10 It is a schematic structural diagram of an electronic device for implementing the speech processing method of the embodiments of the present invention. Detailed implementation manners
[0033] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] Embodiment 1
[0036] Figure 1 It is a flowchart of a speech processing method provided in Embodiment 1 of the present invention. This embodiment is applicable to speech recognition scenarios. This method can be executed by a speech processing device, which can be implemented in the form of hardware and / or software, and the speech processing device can be configured in a computing device. As Figure 1 shown, the method includes:
[0037] S110. Obtain the inspection voice information to be recognized during the power inspection process.
[0038] In this embodiment, the inspection voice information to be recognized refers to the voice information that needs to be recognized. The inspection voice information can be pre-collected through a sound collection device or can be real-time collected voice. For example, in power inspection operation scenarios such as transmission line inspection and substation inspection, in order to improve the information interaction ability in the operation scenario, relevant voice information in the operation scenario can be real-time collected as the inspection voice information to recognize the inspection voice information and obtain the corresponding text content.
[0039] Specifically, in the process of obtaining the inspection voice information to be recognized, the inspection voice information to be processed can be collected based on a microphone array; the inspection voice information to be recognized can be obtained by performing noise reduction processing on the inspection voice information to be processed.
[0040] Among them, the microphone array can be understood as a sound collection system, which can use multiple microphones to collect sounds from different spatial directions.
[0041] In practical applications, the inspection voice information during the power inspection process can be real-time collected using a microphone array. At this time, there may be noise in the inspection voice information. In order to reduce the interference of environmental noise or channel noise on feature recognition, noise reduction technology can be used to perform noise reduction processing on the inspection voice information to obtain more accurate voice information after eliminating useless information as the inspection voice information to be recognized, so as to improve the accuracy of voice recognition when recognizing the inspection voice information to be recognized.
[0042] S120. Process the inspection voice information based on the inspection voice acoustic model to obtain the text recognition result corresponding to each audio frame in the inspection voice information.
[0043] Among them, the inspection voice acoustic model is trained based on the pre-trained general voice model using past inspection voice data. The inspection voice data can be the voice data collected in the power inspection operation scenario. There are power-specific voices and inspection professional terms in the inspection voice data. These data are closely related to attributes such as power equipment and geographical location, and have the characteristics of small data stock, poor generality, and difficult collection. The general voice model can be trained based on general voice data. The general voice data can be the voice data collected in users' daily life or communication. These data have a large stock, strong generality, and are easy to obtain. Therefore, they are different from the inspection voice data.
[0044] In this embodiment, the inspection voice information can be input into an inspection voice acoustic model. Based on the inspection voice acoustic model, feature processing is performed on the voice information corresponding to each audio frame in the inspection voice information to obtain the character recognition result corresponding to each audio frame. Optionally, the processing method of the inspection voice acoustic model for the inspection voice information can be: based on the inspection voice acoustic model, feature extraction is performed on the inspection voice information to obtain the voice feature corresponding to each audio frame; for the voice feature corresponding to each audio frame, based on the voice feature corresponding to the current audio frame, the character recognition result of the current audio frame is determined.
[0045] Among them, the voice feature can include voiceprint feature, timbre feature, and so on. The character recognition result can be Chinese characters, or numbers or symbols. The method for determining the character recognition result of each audio frame is the same. Any one of the audio frames is used as the current audio frame to illustrate the determination of the character recognition result.
[0046] Specifically, a feature extraction algorithm can be used to extract the voice feature corresponding to each audio frame in the inspection voice information, and then based on the voice feature corresponding to the current audio frame, the character recognition result corresponding to the audio frame is determined. For example, the character recognition results corresponding to audio frames 1, 2, 3, 4, and 5 can be "shu", "shu", "shu", "dian", and "dian" respectively.
[0047] S130. Based on the target decoding graph, process each character recognition result to obtain the target inspection text corresponding to the inspection voice information.
[0048] Among them, the target decoding graph is determined based on a character mapping network, a word mapping network, and a sentence mapping network. The character mapping network includes the mapping relationship between frame-level characters and characters, and the frame-level characters correspond to the character recognition results of audio frames. The word mapping network includes the mapping relationship between characters and words. The sentence mapping network includes the mapping relationship between words and sentences. Taking the decoding graph in the form of WFST (Weighted Finaite-State Transducer) as an example, in the decoding graph, there are a finite number of nodes (i.e., states), and the states are connected by directed line segments with arrows. The text on the directed line segment is the input label, and multiple states and directed line segments form a path. Starting from the initial state, searching is performed through the input label to reach the next state. The state reached after completing the last search is the termination state, and the text combinations on multiple directed line segments form a sentence. Correspondingly, the target decoding graph includes the connections between frame-level characters and characters, between characters and characters, and between words and sentences, and there are corresponding weights on each connection.
[0049] In this embodiment, the text recognition results can be input into the target decoding graph, and the beam search algorithm can be used to search in the target decoding graph to obtain at least one set of word sequences corresponding to the inspection voice information; for each word sequence, based on the weight values corresponding to the words to be output in the current word sequence in the target decoding graph, the weight value of the sentence to be output corresponding to the current word sequence is determined; based on the weight values of the sentences to be output, the target inspection text is determined and displayed.
[0050] Among them, there is a hyperparameter in the beam search algorithm, namely the beam width. The number of sentences to be output corresponds to the beam width. For example, if the beam width is 2, the number of sentences to be output obtained by using the beam search algorithm is 2.
[0051] Specifically, the text recognition result corresponding to each audio frame can be used as a speech recognition sequence, and this speech recognition sequence can be used as the input of the target decoding graph. The beam search algorithm is used to search in the target decoding graph. After the search is completed, at least one set of word sequences can be obtained. For example, if the beam width is 2, word sequence A is "inspection", "site", "maintenance", and word sequence B is "inspection", "exhibition point", "maintenance", then the two sentences to be output are inspection site maintenance and inspection exhibition point maintenance. It should be noted that the method for determining the weight value of the sentence to be output corresponding to each set of word sequences is the same. One set of word sequences can be used as the current word sequence for illustration. The weight values corresponding to the words to be output in the current word sequence can be multiplied to obtain a product value, and the product value can be used as the weight value of the sentence to be output composed of the current word sequence. For example, taking word sequence A as an example, the weight of "inspection connecting site" is 0.7, and the weight of "site connecting maintenance" is 0.6, then the weight value of "inspection site maintenance" is recorded as 0.7 * 0.6 = 0.42. The sentence with the largest weight value among the sentences to be output can be used as the target inspection text. Subsequently, the target inspection text can also be displayed on the display screen for the user to view. Or, the target inspection text can be stored to generate an inspection report or used as a training text, and then used as the target inspection text to train the inspection voice acoustic model later.
[0052] The technical solution of this embodiment is to obtain the inspection voice information to be recognized during the power inspection; process the inspection voice information based on the inspection voice acoustic model to obtain the text recognition result corresponding to each audio frame in the inspection voice information; the inspection voice acoustic model is trained based on the inspection voice data from the trained general speech model; process each text recognition result based on the target decoding graph to obtain the target inspection text corresponding to the inspection voice information, which solves the problem of low recognition accuracy in the prior art for inspection voice recognition based on the general speech recognition model, and realizes the recognition of the inspection voice information by using the inspection voice acoustic model to obtain the text recognition result corresponding to each audio frame. The target decoding graph processes the frame-level text recognition results, realizes the mapping from frame-level characters to text, then from text to words, and further from words to sentences, and determines the target inspection text based on the output sentence, achieving the technical effects of improving the accuracy of speech recognition, ensuring the efficiency of inspection, meeting the requirements of power inspection voice recognition, and assisting in improving the quality and efficiency of inspection operations.
[0053] Embodiment 2
[0054] Figure 2 It is a flowchart of a speech processing method provided according to Embodiment 2 of the present invention. On the basis of the foregoing embodiment, the method further includes training to obtain an inspection voice acoustic model, and the specific implementation manner can refer to the technical solution of this embodiment. Among them, the same or corresponding technical terms as those in the above embodiment will not be described in detail here.
[0055] As Figure 2 shown, the method specifically includes the following steps:
[0056] S210. Obtain the first training data set.
[0057] Among them, the first training data set includes general speech data and text annotation data corresponding to the general speech data.
[0058] In this embodiment, in order to improve the accuracy of the model, as much training data as possible can be obtained to extract its acoustic features based on a large amount of general speech data and train the model to obtain a general speech model.
[0059] S220. Input the general speech data into the speech model to be trained to obtain the first output data.
[0060] Among them, the speech model to be trained can be a network structure model of a Shared Encoder (shared encoder) and a CTC Decoder (Connectionist Temporal Classification decoder). For example, a multi-layer transformer or conformer can be used to construct the Shared Encoder, and a fully connected layer and a softmax layer can be used to construct the CTC Decoder, so as to calculate the frame-level CTC loss (loss) based on the output of the shared encoder and the annotation data corresponding to the general speech data, and to correct the model parameters in the speech model to be trained based on the CTC loss.
[0061] Specifically, the general speech data can be used as the input of the speech model to be trained, and the recognition technology can be used to obtain the first output data corresponding to the general speech data. For example, each audio frame can correspond to a recognition result, that is, the output data of the model, so as to compare the first output data with the annotation data in the text and train the speech model to be trained.
[0062] S230. Train the speech model to be trained according to the first output data and the annotation data in the general speech data.
[0063] It should be noted that since the model parameters in the speech model to be trained are not corrected, there are also corresponding differences between the recognition result identifiers of the output audio frames and the annotation data corresponding to the audio frames. Based on the recognition result identifier of the model output, that is, the first output data, and the corresponding annotation data, an error value is determined, and then the model parameters in the speech model to be trained can be corrected based on the error value.
[0064] Specifically, the recognition result identifier of the current output audio frame can be subjected to loss processing with the annotation data, so as to correct the model parameters of the speech model to be trained according to the obtained loss result and train the speech model to be trained.
[0065] S240. Take the convergence of the loss function in the speech model to be trained as the training target to obtain a general speech model.
[0066] In this embodiment, the convergence of the preset loss function can be used as the training objective. When it is determined that the preset loss function of the speech model to be trained converges, it indicates that this model can be used as a general speech model. For example, the training error of the loss function, that is, the loss parameter, can be used as the condition for detecting whether the loss function has reached convergence currently, such as whether the training error is less than the preset error, or whether the error change trend tends to be stable, or whether the current number of iterations is equal to the preset number. If the convergence condition is detected, such as the training error of the loss function reaches less than the preset error or the error change tends to be stable, it indicates that the training of the speech model to be trained is completed, and at this time, the iterative training can be stopped. If it is detected that the current convergence condition is not reached, the first training data can be further obtained to continue training the speech model to be trained until the training error of the loss function is within the preset range. When the training error of the loss function reaches convergence, the speech model to be trained can be used as a general speech model, so that the inspection speech acoustic model can be trained based on the general speech model subsequently.
[0067] S250. Train the general speech model based on the second training dataset to obtain the inspection speech acoustic model.
[0068] Among them, the second training dataset includes inspection speech data and the corresponding audio frame annotation sequences. The inspection speech data includes power-specific speech and inspection professional terms, etc. For example, professional terms such as line names, hidden danger defect types, equipment conditions, and seasonal characteristics, or special pronunciations and special symbols in power data, such as "kV" is marked as "kilovolt", "±" is marked as "positive and negative", and "110" is marked as "one hundred and ten".
[0069] It should be noted that the general speech model at this time has been pre-trained and has good model parameters itself. Therefore, to determine the inspection speech acoustic model based on the general speech model, only a small amount of inspection speech data with text annotations needs to be used as training samples to iteratively optimize the inspection speech acoustic model, which well solves the problem of difficult collection of inspection speech data and low speech recognition efficiency.
[0070] In this embodiment, the method of training the general speech model with the data in the second training dataset can be: input the inspection speech data into the general speech model to obtain the second output data; train the general speech model according to the second output data and the annotation information in the inspection speech data; use the convergence of the loss function in the general speech model as the training objective to obtain the inspection speech acoustic model.
[0071] Specifically, the inspection voice data in the second training dataset can be input into the general voice model, and then the classification result identification of the current output audio frame is subjected to loss processing with the annotation data corresponding to the audio frame. The model parameters of the general voice model are corrected according to the obtained loss result. When the loss function in the general voice model converges, it can be considered that the training of the general voice model is completed, and an inspection voice acoustic model is obtained. The inspection voice information is input into the inspection voice acoustic model to output the text recognition result corresponding to each audio frame in the inspection voice information.
[0072] S260. Obtain the inspection voice information to be recognized during the power inspection process.
[0073] S270. Process the inspection voice information based on the inspection voice acoustic model to obtain the text recognition result corresponding to each audio frame in the inspection voice information.
[0074] S280. Process each text recognition result based on the target decoding graph to obtain the target inspection text corresponding to the inspection voice information.
[0075] It should be noted that the above S220 to S240 can be executed sequentially or in parallel, and the specific execution order is not limited. The above order is only the order for explaining the technical solutions in each step, not the execution order of each step.
[0076] The technical solution of this embodiment trains the voice model to be trained by using the general voice data that is easy to collect to obtain a general voice model, and then trains the general voice model based on the inspection voice data that is not easy to collect to obtain an inspection voice acoustic model, which saves the collection cost and greatly improves the accuracy of the model for inspecting voice recognition.
[0077] Embodiment III
[0078] Figure 3 It is a flowchart of a voice processing method provided according to Embodiment III of the present invention. On the basis of the foregoing embodiments, the method further includes determining a character mapping network, a word mapping network, and a sentence mapping network, and determining a target decoding graph based on the word mapping network, the sentence mapping network, and the character mapping network. The specific implementation manner can refer to the technical solution of this embodiment. Wherein, the same or corresponding technical terms as those in the above embodiments are not described herein again.
[0079] As Figure 3 shown, the method specifically includes the following steps:
[0080] S310. Determine the word mapping network.
[0081] In this embodiment, the implementation manner of determining the word mapping network may be: obtaining a first inspection text; respectively assigning weights to two adjacent characters in at least one group of words to be applied; determining the weight values between each two adjacent characters based on the weights of each two adjacent characters; generating a word mapping network based on the weight values between each two adjacent characters, so that the word mapping network determines the weight values corresponding to each word to be applied.
[0082] Among them, the first inspection text includes at least one group of words to be applied, and each word to be applied includes at least two characters.
[0083] It should be noted that, in order to improve the convenience of speech processing while improving the speech recognition effect of novel words in the inspection scenario, the novelty of the vocabulary data in the first inspection text, that is, the timeliness is relatively high. Specifically, the first inspection text can be determined by technicians according to the actual working conditions. In practical applications, the words to be applied in each group in the first inspection text can be segmented to obtain at least a pair of adjacent characters. For example, "inspection" can be segmented into "inspect" and "check". The implementation manner of determining the weight value between two adjacent characters may be: when segmenting each word to be applied in the first inspection text, based on the attribute information such as the importance, timeliness, accuracy, and novelty of each group of words to be applied, determine the weight value between adjacent characters. For example, the novelty corresponding to "inspection" is relatively high, and the weight value from "inspect" to "check" can be set to 0.4. The novelty corresponding to "Line Ⅰ" is relatively low, and the weight value from "Ⅰ" to "line" can be set to 0.3. It can be considered that the attribute information of each group of words to be applied is different, and the weight values between two adjacent characters in the corresponding words to be applied may be different. The implementation manner of determining the weight value between two adjacent characters may also be: when segmenting each word to be applied in the first inspection text, based on the attribute information such as the importance, timeliness, accuracy, and novelty of each group of words to be applied, assign score weights to the characters in the corresponding words to be applied, and determine the weight value between two adjacent characters based on the score weights corresponding to the two adjacent characters. For example, the novelty corresponding to "inspection" is relatively high, the score weight of "inspect" is assigned 0.4, and the score weight of "check" is assigned 0.5. Then the weight value from "inspect" to "check" can be 0.4 + 0.5 = 0.9. Further, after determining the weight values between all adjacent characters, a connection relationship between adjacent characters can be established based on the weight values between adjacent characters, and a word mapping network can be constructed, so that based on the weight values between the characters in the word mapping network, the weight values corresponding to the corresponding words to be applied can be determined. For example, the weight value between character 1 and character 2 can represent the probability that character 1 is connected to character 2, and this probability can be used as the weight value of the word formed by character 1 and character 2. Exemplarily, reference can be made to Figure 4 , Figure 4 a schematic diagram showing the mapping of "inspect" and "check" to "inspection" in the word mapping network, <eps>Indicates no input or output. The single-line circle 1 represents the initial state, the single-line circles 2 and 3 represent intermediate states, and the double-line circle represents the termination state.
[0084] It should be noted that the mapping data in the word mapping network is relatively small, which is conducive to its rapid update. Optionally, after generating the word mapping network based on the weight values between adjacent two characters, it further includes: obtaining the newly added inspection text, and determining the inspection text to be discarded in the first inspection text; based on the newly added inspection text and the inspection text to be discarded, updating the first inspection text, so as to update the word mapping network based on the updated first inspection text.
[0085] Among them, the timeliness of the newly added inspection text is higher than that of the inspection text to be discarded. For example, the inspection information on the newly added line can be used as the newly added inspection text to meet the recognition requirements for novel words.
[0086] In practical applications, after determining the newly added inspection text, the first inspection text can be traversed. If it is found that the first inspection text does not contain the words in the newly added inspection text, the non-included words can be automatically added to the first inspection text to update the first inspection text, so as to update the word mapping network based on the updated first inspection text, realizing the rapid expansion of out-of-vocabulary words and meeting the requirements for the rapid update of special terms in practical applications. The data with lower timeliness in the first inspection text can also be deleted from the word mapping network to realize the dynamic update of inspection novel words.
[0087] S320. Determine the phrase mapping network.
[0088] In this embodiment, the implementation manner of determining the phrase mapping network may be: obtaining the second inspection text; determining the weights corresponding to each word to be used; generating a phrase mapping network based on the weights corresponding to at least one word to be used in the corresponding sentence to be used, so that the phrase mapping network determines the weight corresponding to the sentence to be used.
[0089] It should be noted that there is relatively more mapping data in the word-sentence mapping network, which is beneficial to the accuracy of speech recognition. Patrol inspection texts that conform to the content and technical standards of power equipment inspection can be collected. For the accuracy of speech recognition, the patrol inspection texts can cover line names, hidden danger defect types, equipment conditions, and seasonal characteristics, as well as information such as habitual expressions, special pronunciations, and special symbols in patrol inspections. For example, the commonly used voltage level of kilovolts, "±" means "positive and negative", and the line number "Line 1" is "Line Ⅰ". Among them, the second patrol inspection text includes at least one statement to be used, and the statement to be used includes at least one word to be used. For example, each statement to be used in the second patrol inspection text can be segmented to obtain at least one vocabulary sequence, and the vocabulary sequence includes at least one word to be used, such as patrol inspection station, patrol inspection, task, start, and so on.
[0090] In practical applications, the conditional probability algorithm can be used to determine the probability value of each word to be used appearing in the second patrol inspection text, and the probability value can be used as the weight. It is also possible to determine the probability (i.e., weight) of each word to be used appearing when the previous preset number of words appear by combining context information. Then, based on the weights corresponding to each word to be used in the corresponding statement to be used, the weight corresponding to the statement to be used is determined. Further, based on the weights corresponding to each word to be used, the weight between every two words to be used can be determined, and then the connection relationship between the two words to be used is established to generate a word-sentence mapping network, so that based on the weights between the words to be used in the word-sentence mapping network, the weight corresponding to the corresponding statement is determined. For example, the weight between word 1 and word 2 can represent the probability of word 2 appearing after word 1, and this weight can be used as the weight of the sentence composed of word 1 and word 2, that is, the probability of appearance. For example, assume that the second patrol inspection text contains n words to be used, which are represented by ω1, ω2, ω3,..., ωn respectively. Further, the second patrol inspection text can be trained using the N-gram (N-ary) language to calculate the weight of the word appearance. If the weight of a certain word appearance is only related to the previous two words, that is, the trigram language model, taking the trigram language model as an example, the formula for calculating the weight of a sentence can be:
[0091]
[0092]
[0093] Among them, s represents the sentence composed of ω1, ω2, ω3,..., ωn, P(s) is the weight of s, and ∏() represents the product. Exemplarily, see Figure 5 , Figure 5 It can represent a schematic diagram of a phrase mapping network, where the weights on the directional line segments represent the probabilities of connecting to the next word given the current word. The probability of patrol mapping to task is 1.0 * 0.6 = 0.6. The weight of patrol, task, start mapping to "patrol task start" is 1.0 * 0.6 * 1.0 = 0.6. The probability of patrol mapping to person is 1.0 * 0.4 = 0.4. The weight of patrol, person, unlocking mapping to "patrol person unlocking" is 1.0 * 0.4 * 1.0 = 0.4.
[0094] It should be noted that this technical solution can use the maximum likelihood estimation method to estimate the weights corresponding to the sentences to be used, approximate the weights by counting frequencies. There may be a problem of data sparsity in that the frequencies of longer sentences may not be counted. To solve the problem of data sparsity, a smoothing estimation algorithm can be used to estimate the weights of the sentences to be used. For example, the smoothing estimation algorithm can be the linear interpolation method or discounting methods. This technical solution does not make a limitation.
[0095] It should also be noted that in the process of determining the weights corresponding to each word to be used, for each word to be used, the first attribute value corresponding to the co-occurrence of the current word to be used and the previous two words to be used, the second attribute value corresponding to the co-occurrence of the current word to be used and the previous word to be used, the occurrence frequency of the current word to be used, and the corresponding weight parameters can be determined to determine the weight corresponding to the current word to be used.
[0096] Among them, the first attribute value is determined based on the co-occurrence frequency of the current word to be used and the previous two words to be used, and the co-occurrence frequency of the previous two words to be used; the second attribute value is determined based on the co-occurrence frequency of the current word to be used and the previous word to be used, and the occurrence frequency of the previous word to be used. The processing methods for determining the first attribute value of each word to be used are the same. Any word to be used is used as an example of the current word to be used for explanation.
[0097] In practical applications, based on the co-occurrence frequency of the current word to be used and the previous two words to be used, as well as the co-occurrence frequency of the previous two words to be used, the probability value corresponding to the occurrence of the current word to be used when the previous two words to be used co-occur can be determined as the first attribute value. For example, assume that the current word to be used is w, and the previous two words to be used are u and v respectively. The first attribute value p(w|u,v) = c(u,v,w) / c(u,v), where c represents frequency, c(u,v,w) is the frequency of the consecutive occurrence of u, v, and w, and c(u,v) is the frequency of the consecutive occurrence of u and v. Based on the co-occurrence frequency of the current word to be used and the previous word to be used, as well as the occurrence frequency of the previous word to be used, the probability value corresponding to the occurrence of the current word to be used when the previous word to be used occurs can be determined, that is, the second attribute value. For example, the second attribute value p(w|v) = c(v,w) / c(v). Further, based on the first attribute value, the second attribute value, the occurrence frequency, and the corresponding weight parameters corresponding to the current word to be used, the weight corresponding to the current word to be used can be determined. For example, the weight p = p(w|u,v)*α + p(w|v)*β + c(w)*γ, where α, β, and γ are the weight parameters of the first attribute value, the second attribute value, and the occurrence frequency respectively, and α, β, and γ satisfy α + β + γ = 1 and α≥0, β≥0, γ≥0.
[0098] S330. Determine the text mapping network.
[0099] In this embodiment, the implementation method of determining the text mapping network may be: obtaining a third inspection text; for each frame-level tag sequence, during the process of mapping the current frame-level tag sequence into at least one text, determining the mapping attributes corresponding to each frame-level tag in the current frame-level tag sequence; generating a text mapping network based on the mapping attributes corresponding to each frame-level tag, so that the text mapping network determines the weight values corresponding to each text.
[0100] Among them, the third inspection text includes at least one group of frame-level tag sequences. The frame-level tag sequence includes at least one frame-level tag, and the frame-level tag includes a space tag and a frame-level word tag. The frame-level tag can be used to represent the uniqueness of the audio frame and can be the audio recognition result corresponding to the audio frame. For example, the frame-level tag corresponding to audio frame A is "inspection", the frame-level tag corresponding to audio frame B is "check", and the frame-level tag corresponding to audio frame C can be empty. "Inspection" and "check" can be understood as frame-level word tags, and empty means that the recognition result of the audio frame is empty, which can be represented by a space. The mapping attribute can represent the probability of mapping the frame-level tag into a text. It should be noted that the processing method for each group of frame-level tag sequences is the same, and any one of the frame-level tag sequences can be used as the current frame-level tag sequence for description.
[0101] In practical applications, the way to obtain the third inspection text can be to process the inspection voice through an inspection voice acoustic model to obtain the text recognition identifier corresponding to each audio frame in the inspection voice, that is, to obtain a frame-level label sequence. Further, the continuously identical frame-level labels in the frame-level label sequence can be mapped to a single character. During the mapping process, corresponding mapping attributes can be assigned to each frame-level label, and then a character mapping network can be generated based on the mapping attributes corresponding to each frame-level label, so that the character mapping network can obtain the weight value of at least one character corresponding to each frame-level label through each frame-level label. For example, the character mapping network can map the frame-level connection temporal classification label sequence to a single character, and the sequence is a frame-level character sequence. For example, after processing 3 audio frames, the characters output by the CTC Decoder may be "Wei Wei Wei", and the character mapping network maps these sequences to the character "Wei", which can be referred to Figure 6 , <blank>Indicates blank.
[0102] It should be noted that the above S310, S320, and S330 can be executed sequentially or in parallel, and the specific execution order is not limited. The above order is only the order for explaining the technical solutions in each step, not the execution order of each step.
[0103] S340. Fuse the word mapping network and the phrase mapping network to obtain a decoding graph to be processed.
[0104] It should be noted that the word mapping network and the phrase mapping network can be fused, that is, the weight values between adjacent two characters, the weight values corresponding to the words to be applied, and the weight values corresponding to the sentences to be used are fused to obtain a decoding graph to be processed containing the mapping relationship between characters, words to be applied, and sentences to be used.
[0105] Specifically, the method for determining the decoding graph to be processed can be: based on the sum of the weights of adjacent two characters in the word mapping network and the weight values of the words to be applied corresponding to the adjacent two characters, as well as the sum of the weights of at least one word to be used and the weight values of the sentences to be used corresponding to the at least one word to be used in the phrase mapping network, generate a decoding graph to be processed, so that the decoding graph to be processed contains the mapping relationship and weight values between characters, words to be applied, and sentences to be used.
[0106] In this embodiment, the word mapping network contains weight values between multiple groups of adjacent two characters. Correspondingly, the weight values corresponding to the words to be applied corresponding to the adjacent two characters can be determined. The phrase mapping network contains the weights of multiple words to be used in each sentence to be used. Correspondingly, the weight values of each sentence to be used can be determined. The information at each level in the word mapping network and the phrase mapping network can be fused, such as the mapping from characters to words in the word mapping network and the mapping from words to sentences in the phrase mapping network. After fusion, the mapping from characters to words and then to sentences can be obtained, and a decoding graph to be processed containing the mapping relationship and weight values between characters, words to be applied, and sentences to be used is generated.
[0107] S350. Perform pruning processing on the decoding graph to be processed to obtain a decoding graph to be used.
[0108] In this embodiment, the decoding graph to be processed can be pruned according to certain pruning rules to obtain the decoding graph to be used, so as to realize the determination and minimization of the decoding graph through pruning, achieving the effect of compressing the search space and accelerating decoding. Here, determination means ensuring that for a given input, the output is unique; minimization means converting the decoding graph into an equivalent decoding graph with fewer state nodes and edges. For example, at least two identical inputs (i.e., inputlabel) can be merged so that no state has two outgoing arcs for the same input. Two different states connected to the same input can also be merged to minimize the number of states and arcs in the decoding graph to be processed.
[0109] S360. Fuse the character mapping network and the decoding graph to be used to obtain the target decoding graph.
[0110] In this embodiment, the character mapping network and the decoding graph to be used can be fused to obtain the target decoding graph. Further, the target decoding graph can be pruned to realize the determination and minimization of the target decoding graph, and all mapping encodings from the frame-level characters recognized from the audio frame to the character sequence can be realized. When using the beam search algorithm to find the word sequence corresponding to the character recognition result of the audio frame in the target decoding graph, the decoding search efficiency can be improved.
[0111] Specifically, in the process of fusing the character mapping network and the decoding graph to be used to obtain the target decoding graph, it can be as follows: Based on the mapping attributes corresponding to each frame-level label in the character mapping network and the weight values of at least one character corresponding to each frame-level label, as well as the mapping relationships and weight values among the characters, the words to be applied, and the sentences to be used in the decoding graph to be processed, a target decoding graph is generated so that the target decoding graph contains the mapping relationships and weight values among the frame-level labels, characters, words to be applied, and sentences to be used.
[0112] In this embodiment, the character mapping network contains the mapping attributes of frame-level labels. Correspondingly, the weight values of the characters corresponding to multiple frame-level labels can be determined. The decoding graph to be processed contains the mapping relationships and weight values among the characters, the words to be applied, and the sentences to be used. The information at each level in the character mapping network and the decoding graph to be used can be fused. For example, the mapping from the frame-level label to the character in the character mapping network and the mapping from the character to the word and then to the sentence in the decoding graph to be used are fused. After fusion, the mapping from the frame-level label to the character to the word to the sentence can be obtained, and a decoding graph to be processed containing the mapping relationships and weight values among the frame-level labels, characters, words to be applied, and sentences to be used can be obtained.
[0113] S370. Obtain the inspection voice information to be recognized during the power inspection process.
[0114] S380. Process the inspection voice information based on the inspection voice acoustic model to obtain the text recognition result corresponding to each audio frame in the inspection voice information.
[0115] S390. Process each text recognition result based on the target decoding graph to obtain the target inspection text corresponding to the inspection voice information.
[0116] In the technical solution of this embodiment, a word mapping network is generated through the first inspection text with high timeliness, and a word mapping network including the mapping relationship between words and phrases is also generated based on the second inspection text. Finally, by fusing the mapping networks with three levels of information, a target encoding graph is generated. When performing speech recognition, the sentence most suitable for the inspection voice data is searched from the target encoding graph, which not only improves the accuracy and regularity of speech recognition, but also makes the output sentence text conform to a certain language logic, achieving the technical effect of improving the effectiveness of information interaction.
[0117] Embodiment 4
[0118] As an optional embodiment of the above embodiment, Figure 7 It is a schematic diagram of the speech recognition method provided in Embodiment 4 of the present invention. Specifically, the following specific content can be referred to.
[0119] Refer to Figure 7 , the real-time acquisition of inspection voice data can be completed by using a microphone array, and the corresponding audio file can be generated for storage. Furthermore, the collected multi-channel inspection voice information can be input into the speech recognition module in real time, and the inspection voice information can be processed based on the trained inspection voice acoustic model and the target decoding graph, and the recognized target inspection text can be output. The target inspection text and the corresponding inspection voice information are stored in the inspection text library correspondingly, so as to train the model used for inspection voice recognition based on the newly added inspection data in the inspection text library, and improve the speech recognition accuracy of the model.
[0120] It should be noted that with the increase of power grid equipment and the complexity of operation conditions, there may be more and more professional terms such as hidden danger types and entity names, and there may be problems such as inaccurate recognition of these terms, and the model needs to be iteratively updated in a timely manner. Therefore, the dynamic update of the word mapping network and the sentence mapping network is realized through the real-time updated first inspection text, and the accuracy of inspection professional terms is improved. For example, after determining the newly added inspection text, the first inspection text can be traversed. If it is found that the first inspection text does not contain the vocabulary in the newly added inspection text, the unincluded vocabulary can be automatically added to the first inspection text to realize the dynamic update of the inspection novel words, and the word mapping network can be updated based on the updated first inspection text to quickly expand the out-of-vocabulary words and meet the requirements of the rapid update of special terms in practical applications.
[0121] Based on the above solution, refer to Figure 8 , the first training data set (general speech data) can be subjected to feature extraction, and its acoustic features are extracted and input into a network model constructed based on the Shared Encoder and CTC Decoder frameworks, that is, the speech model to be trained, so as to train the speech model to be trained to obtain a general speech model. Furthermore, the second training data set (inspection speech data) can be subjected to feature extraction, and its acoustic features are extracted and input into the general speech model to train the general speech model to obtain an inspection speech acoustic model. A word mapping network can be determined based on the first training text, and a sentence mapping network can be determined based on the second inspection text. When training the general speech model, a character mapping network can be constructed based on the recognition result corresponding to each audio frame. Furthermore, the sentence mapping grid and the word mapping network can be fused, and after fusion, a determinization and minimization process is performed to obtain a decoding graph to be processed. Furthermore, the character mapping network and the decoding graph to be processed are fused to obtain a target encoding graph. In practical applications, the collected inspection speech data can be input into the inspection speech acoustic model to obtain the character recognition result corresponding to each audio frame, and each character recognition result is used as a sequence to input into the target decoding graph, and the beam search algorithm is used to search in the target decoding graph to obtain the target inspection text.
[0122] The technical solution of this embodiment realizes the recognition of inspection speech information by using the inspection speech acoustic model to obtain the character recognition result corresponding to each audio frame. The target decoding graph processes the frame-level character recognition result, realizes the mapping from frame-level characters to words, then from words to phrases, and further from phrases to sentences, and determines the target inspection text based on the output sentence, achieving the technical effects of improving the accuracy of speech recognition, ensuring the efficiency of inspection, meeting the requirements of power inspection speech recognition, and helping to improve the quality and efficiency of inspection operations.
[0123] Embodiment 5
[0124] Figure 9 is a schematic structural diagram of a speech processing device provided according to Embodiment 5 of the present invention. As Figure 9 shown, the device includes: an inspection speech information acquisition module 910, a character recognition result determination module 920, and a target inspection text determination module 930.
[0125] Among them, the inspection voice information acquisition module 910 is used to acquire the inspection voice information to be recognized during the power inspection; the text recognition result determination module 920 is used to process the inspection voice information based on the inspection voice acoustic model to obtain the text recognition result corresponding to each audio frame in the inspection voice information; wherein, the inspection voice acoustic model is obtained by training a pre-trained general voice model based on past inspection voice data; the target inspection text determination module 930 is used to process each text recognition result based on the target decoding graph to obtain the target inspection text corresponding to the inspection voice information; wherein, the target decoding graph is determined based on a character mapping network, a word mapping network, and a sentence mapping network. The character mapping network includes the mapping relationship between the frame-level character and the text, the frame-level character corresponds to the text recognition result of the audio frame, the word mapping network includes the mapping relationship between the text and the word, and the sentence mapping network includes the mapping relationship between the word and the sentence.
[0126] The technical solution of this embodiment obtains the inspection voice information to be recognized during the power inspection; processes the inspection voice information based on the inspection voice acoustic model to obtain the text recognition result corresponding to each audio frame in the inspection voice information; wherein, the inspection voice acoustic model is obtained by training a well-trained general voice model based on the inspection voice data; processes each text recognition result based on the target decoding graph to obtain the target inspection text corresponding to the inspection voice information, solving the problem of low recognition accuracy in the prior art when using a general voice recognition model for inspection voice recognition. It realizes the recognition of the inspection voice information by using the inspection voice acoustic model to obtain the text recognition result corresponding to each audio frame, and the target decoding graph processes the text recognition result at the frame level, realizing the mapping from the frame-level character to the text, then from the text to the word, and further from the word to the sentence, and determining the target inspection text based on the output sentence, achieving the technical effects of improving the accuracy of voice recognition, ensuring the efficiency of inspection, meeting the requirements of power inspection voice recognition, and assisting in improving the quality and efficiency of inspection operations.
[0127] On the basis of the above device, optionally, the inspection voice information acquisition module 910 includes a to-be-processed voice information determination unit and an inspection voice information determination unit.
[0128] The to-be-processed voice information determination unit is used to acquire the to-be-processed inspection voice information based on the microphone array.
[0129] The inspection voice information determination unit is used to obtain the to-be-recognized inspection voice information by performing noise reduction processing on the to-be-processed inspection voice information.
[0130] Based on the above device, optionally, the text recognition result determination module 920 includes a voice feature determination unit and a text recognition result determination unit.
[0131] The voice feature determination unit is configured to extract features from the inspection voice information based on the inspection voice acoustic model to obtain the voice features corresponding to each audio frame.
[0132] The text recognition result determination unit is configured to determine the text recognition result of the current audio frame based on the voice features corresponding to each audio frame and the voice features corresponding to the current audio frame.
[0133] Based on the above device, optionally, the target inspection text determination module 930 includes a word sequence determination unit, an attribute value determination unit, and a target inspection text determination unit.
[0134] The word sequence determination unit is configured to input each text recognition result into the target decoding graph and use the beam search algorithm to search in the target decoding graph to obtain at least one set of word sequences corresponding to the inspection voice information.
[0135] The attribute value determination unit is configured to, for each word sequence, determine the weight value of the to-be-output statement corresponding to the current word sequence based on the weight values corresponding to each to-be-output word in the current word sequence in the target decoding graph.
[0136] The target inspection text determination unit is configured to determine and display the target inspection text based on the attribute values of each to-be-output statement.
[0137] Based on the above device, optionally, the device further includes: an inspection voice acoustic model determination module, and the inspection voice acoustic model determination module includes a first training data set determination unit, a first output data determination unit, a to-be-trained voice model training unit, a general voice model determination unit, and an inspection voice acoustic model determination unit.
[0138] The first training data set determination unit is configured to obtain a first training data set; wherein, the first training data set includes general voice data, and the general voice data is different from the inspection voice data.
[0139] The first output data determination unit is configured to input the general voice data into the to-be-trained voice model to obtain first output data.
[0140] The to-be-trained voice model training unit is configured to train the to-be-trained voice model according to the first output data and the labeled data in the general voice data.
[0141] A general speech model determination unit, which takes the convergence of the loss function in the speech model to be trained as the training objective to obtain a general speech model;
[0142] An inspection speech acoustic model determination unit, which trains the general speech model based on a second training dataset to obtain the inspection speech acoustic model, where the second training dataset includes inspection speech data.
[0143] Based on the above device, optionally, the inspection speech acoustic model determination unit includes a second output data determination subunit, a general speech model training subunit, a general speech model determination subunit, and an inspection speech acoustic model determination subunit.
[0144] The second output data determination subunit is configured to input the inspection speech data into the general speech model to obtain second output data;
[0145] The general speech model training subunit is configured to train the general speech model according to the second output data and the annotation information in the inspection speech data;
[0146] The inspection speech acoustic model determination subunit takes the convergence of the loss function in the general speech model as the training objective to obtain an inspection speech acoustic model.
[0147] Based on the above device, optionally, the device further includes: a target decoding graph determination module, and the target decoding graph determination module includes a mapping network determination unit, a decoding graph to be processed determination unit, a decoding graph to be used determination unit, and a target decoding graph determination unit.
[0148] The mapping network determination unit is configured to determine the word mapping network and determine the sentence mapping network;
[0149] The decoding graph to be processed determination unit is configured to fuse the word mapping network and the sentence mapping network to obtain a decoding graph to be processed;
[0150] The decoding graph to be used determination unit is configured to perform pruning processing on the decoding graph to be processed to obtain a decoding graph to be used;
[0151] The target decoding graph determination unit is configured to fuse the text mapping network and the decoding graph to be used to obtain the target decoding graph.
[0152] Based on the above device, optionally, the mapping network determination unit includes a word mapping network determination subunit, and the word mapping network determination subunit includes a first inspection text determination sub-unit, an adjacent word weight determination sub-unit, a word weight value determination sub-unit, and a word mapping network determination sub-unit.
[0153] The first patrol text determination unit is used to obtain the first patrol text; wherein, the first patrol text includes at least one group of words to be applied, and the words to be applied include at least two characters.
[0154] The adjacent character weight determination unit is used to assign weights to adjacent two characters in at least one group of words to be applied respectively.
[0155] The character weight value determination unit is used to determine the weight value between each adjacent two characters based on the weights of each adjacent two characters.
[0156] The word mapping network determination unit is used to generate the word mapping network based on the weight values between each adjacent two characters, so that the word mapping network determines the weight values corresponding to each word to be applied.
[0157] Based on the above device, optionally, the mapping network determination unit further includes a word mapping network update subunit, and the word mapping network update subunit includes a first patrol text update unit and a word mapping network update unit.
[0158] The first patrol text update unit is used to obtain the new patrol text and determine the patrol text to be discarded in the first patrol text.
[0159] The word mapping network update unit is used to update the first patrol text based on the new patrol text and the patrol text to be discarded, so as to update the word mapping network based on the updated first patrol text.
[0160] Based on the above device, optionally, the mapping network determination unit includes a sentence mapping network determination subunit, and the sentence mapping network determination subunit includes a second patrol text determination unit, a word weight determination unit and a sentence mapping network determination unit.
[0161] The second patrol text determination unit is used to obtain the second patrol text; wherein, the second patrol text includes at least one sentence to be used, and the sentence to be used includes at least one word to be used.
[0162] The word weight determination unit is used to determine the weights corresponding to each word to be used.
[0163] The sentence mapping network determination unit is used to generate the sentence mapping network based on the weights corresponding to at least one word to be used in each sentence to be used, so that the sentence mapping network determines the weight values corresponding to each sentence to be used.
[0164] Based on the above device, optionally, the word weight determination unit is further configured to, for each word to be used, determine the first attribute value corresponding to the co-occurrence of the current word to be used and the previous two words to be used, the second attribute value corresponding to the co-occurrence of the current word to be used and the previous word to be used, the occurrence frequency of the current word to be used, and the corresponding weight parameter, and determine the weight corresponding to the current word to be used; wherein, the first attribute value is determined based on the co-occurrence frequency of the current word to be used and the previous two words to be used, and the co-occurrence frequency of the previous two words to be used; the second attribute value is determined based on the co-occurrence frequency of the current word to be used and the previous word to be used, and the occurrence frequency of the previous word to be used.
[0165] Based on the above device, optionally, the to-be-processed decoding graph determination unit is further configured to generate the to-be-processed decoding graph based on the weights of adjacent two characters in the character mapping network and the weight values of the to-be-applied words corresponding to the adjacent two characters, and the weights of at least one word to be used in the sentence mapping network and the weight values of the to-be-used sentences corresponding to the at least one word to be used, so that the to-be-processed decoding graph includes the mapping relationship and weight values between characters, to-be-applied words and to-be-used sentences.
[0166] Based on the above device, optionally, the device further includes: a character mapping network determination module, and the character mapping network determination module includes a third inspection text determination unit, a mapping attribute determination unit, and a character mapping network determination unit.
[0167] The third inspection text determination unit is configured to obtain a third inspection text; wherein, the third inspection text includes at least one set of frame-level tag sequences, the frame-level tag sequences include at least one frame-level tag, and the frame-level tags include space tags and frame-level character tags.
[0168] The mapping attribute determination unit is configured to, for each frame-level tag sequence, determine the mapping attributes corresponding to the frame-level tags in the current frame-level tag sequence during the process of mapping the current frame-level tag sequence to at least one character.
[0169] The character mapping network determination unit is configured to generate the character mapping network based on the mapping attributes corresponding to the frame-level tags, so that the character mapping network determines the weight values corresponding to the characters.
[0170] Based on the above device, optionally, the target decoding graph determination unit is further configured to generate the target decoding graph based on the mapping attributes corresponding to each frame-level tag in the text mapping network, the weight values of at least one text corresponding to each frame-level tag, and the mapping relationships and weight values among the text, the words to be applied, and the sentences to be used in the decoding graph to be processed, so that the target decoding graph includes the mapping relationships and weight values among the frame-level tags, the text, the words to be applied, and the sentences to be used.
[0171] The speech processing device provided by the embodiments of the present invention can execute the speech processing method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0172] Embodiment Six
[0173] Figure 10 It is a schematic structural diagram of an electronic device for implementing the speech processing method of the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0174] As Figure 10 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0175] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0176] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the speech processing method.
[0177] In some embodiments, the speech processing method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the speech processing method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the speech processing method by any other suitable means (e.g., by means of firmware).
[0178] The various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0179] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable voice processing devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0180] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0181] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0182] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0183] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0184] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0185] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.< / blank> < / eps>
Claims
1. A voice processing method, characterized in that, Including: Obtaining the inspection voice information to be recognized during the power inspection process; Processing the inspection voice information based on the inspection voice acoustic model to obtain the text recognition result corresponding to each audio frame in the inspection voice information; wherein, the inspection voice acoustic model is trained based on the previous inspection voice data for the pre-trained general voice model; Processing each text recognition result based on the target decoding graph to obtain the target inspection text corresponding to the inspection voice information; Wherein, the target decoding graph is determined based on the character mapping network, the word mapping network, and the sentence mapping network. The character mapping network includes the mapping relationship between the frame-level character and the character. The frame-level character corresponds to the text recognition result of the audio frame. The word mapping network includes the mapping relationship between the character and the word. The sentence mapping network includes the mapping relationship between the word and the sentence; Determining the target decoding graph includes: determining the word mapping network and determining the sentence mapping network; fusing the word mapping network and the sentence mapping network to obtain the decoding graph to be processed; performing pruning processing on the decoding graph to be processed to obtain the decoding graph to be used; fusing the character mapping network and the decoding graph to be used to obtain the target decoding graph; The weight value between each adjacent two characters in the word mapping network is determined based on the attribute information in the important attribute, timeliness attribute, accuracy attribute, and novelty attribute of each word to be applied in the first inspection text when performing word segmentation on each word to be applied in the first inspection text.
2. The method according to claim 1, characterized in that The obtaining the inspection voice information to be recognized during the power inspection process includes: Collecting the inspection voice information to be processed based on the microphone array; Obtaining the inspection voice information to be recognized by performing noise reduction processing on the inspection voice information to be processed.
3. The method according to claim 1, wherein The processing the inspection voice information based on the inspection voice acoustic model to obtain the text recognition result corresponding to each audio frame in the inspection voice information includes: Performing feature extraction on the inspection voice information based on the inspection voice acoustic model to obtain the voice feature corresponding to each audio frame; For the voice feature corresponding to each audio frame, determining the text recognition result of the current audio frame based on the voice feature corresponding to the current audio frame.
4. The method according to claim 1, wherein The processing each text recognition result based on the target decoding graph to obtain the target inspection text corresponding to the inspection voice information includes: Inputting each text recognition result into the target decoding graph, and using the beam search algorithm to search in the target decoding graph to obtain at least one group of word sequences corresponding to the inspection voice information; For each word sequence, determining the weight value of the output sentence corresponding to the current word sequence based on the weight value of each word to be output in the current word sequence in the target decoding graph; Determining the target inspection text based on the weight values of each output sentence.
5. The method according to claim 1, characterized in that It further includes: Training to obtain an inspection voice acoustic model; The training to obtain an inspection voice acoustic model includes: Obtain a first training data set; wherein, the first training data set includes general speech data, and the general speech data is different from the inspection speech data; Input the general speech data into the speech model to be trained to obtain first output data; Train the speech model to be trained according to the first output data and the labeled data in the general speech data; Take the convergence of the loss function in the speech model to be trained as the training objective to obtain a general speech model; Train the general speech model based on a second training data set to obtain the inspection speech acoustic model, and the second training data set includes inspection speech data.
6. The method according to claim 5, characterized in that, The training the general speech model based on the second training data set to obtain the inspection speech acoustic model includes: Input the inspection speech data into the general speech model to obtain second output data; Train the general speech model according to the second output data and the labeled information in the inspection speech data; Take the convergence of the loss function in the general speech model as the training objective to obtain the inspection speech acoustic model.
7. The method according to claim 1, characterized in that The determining the word mapping network includes: Obtain a first inspection text; wherein, the first inspection text includes at least one group of words to be applied, and the words to be applied include at least two characters; Assign weights to adjacent two characters in the at least one group of words to be applied respectively; Determine the weight values between each adjacent two characters based on the weights of each adjacent two characters; Generate the word mapping network based on the weight values between each adjacent two characters, so that the word mapping network determines the weight values corresponding to each word to be applied.
8. The method according to claim 7, wherein After the generating the word mapping network based on the weight values between each adjacent two characters, it further includes: Obtain a new inspection text and determine the inspection text to be discarded in the first inspection text; Update the first inspection text based on the new inspection text and the inspection text to be discarded, so as to update the word mapping network based on the updated first inspection text.
9. The method according to claim 1, characterized in that The determining the sentence mapping network includes: Obtain a second inspection text; wherein, the second inspection text includes at least one sentence to be used, and the sentence to be used includes at least one word to be used; Determine the weights corresponding to each word to be used; Generate the sentence mapping network based on the weights corresponding to at least one word to be used in each sentence to be used, so that the sentence mapping network determines the weight values corresponding to each sentence to be used.
10. The method according to claim 9, wherein The determining the weights corresponding to each word to be used includes: For each word to be used, determine the first attribute value corresponding to the co-occurrence of the current word to be used and the previous two words to be used, the second attribute value corresponding to the co-occurrence of the current word to be used and the previous word to be used, the occurrence frequency of the current word to be used, and the corresponding weight parameter, and determine the weight corresponding to the current word to be used; The first attribute value is determined based on the frequency of co-occurrence of the currently to-be-used word and the previous two to-be-used words, as well as the frequency of co-occurrence of the previous two to-be-used words; the second attribute value is determined based on the frequency of co-occurrence of the currently to-be-used word and the previous to-be-used word, as well as the frequency of occurrence of the previous to-be-used word.
11. The method according to claim 1, wherein The fusion of the word mapping network and the sentence mapping network to obtain a to-be-processed decoding graph includes: Generating the to-be-processed decoding graph based on the weight sum of each adjacent pair of characters in the word mapping network and the weight value of the to-be-applied word corresponding to each adjacent pair of characters, as well as the weight sum of at least one to-be-used word corresponding to the sentence mapping network and the weight value of the to-be-used sentence corresponding to the at least one to-be-used word, so that the to-be-processed decoding graph contains the mapping relationship and weight value between characters, to-be-applied words, and to-be-used sentences.
12. The method according to claim 1, wherein The determination of the character mapping network includes: Obtaining a third inspection text; wherein, the third inspection text includes at least one set of frame-level tag sequences, each frame-level tag sequence includes at least one frame-level tag, and each frame-level tag includes a space tag and a frame-level character tag; For each frame-level tag sequence, during the process of mapping the current frame-level tag sequence to at least one character, determining the mapping attribute corresponding to each frame-level tag in the current frame-level tag sequence; Generating the character mapping network based on the mapping attributes corresponding to each frame-level tag, so that the character mapping network determines the weight value corresponding to each character.
13. The method according to claim 1, wherein The fusion of the character mapping network and the to-be-used decoding graph to obtain the target decoding graph includes: Generating the target decoding graph based on the mapping attributes corresponding to each frame-level tag in the character mapping network and the weight value of at least one character corresponding to each frame-level tag, as well as the mapping relationship and weight value between characters, to-be-applied words, and to-be-used sentences in the to-be-processed decoding graph, so that the target decoding graph contains the mapping relationship and weight value between frame-level tags, characters, to-be-applied words, and to-be-used sentences.
14. A voice processing device, characterized in that, including: An inspection voice information acquisition module for acquiring inspection voice information to be recognized during power inspection; A character recognition result determination module for processing the inspection voice information based on an inspection voice acoustic model to obtain the character recognition result corresponding to each audio frame in the inspection voice information; wherein, the inspection voice acoustic model is trained based on past inspection voice data on a pre-trained general voice model; A target inspection text determination module for processing each character recognition result based on the target decoding graph to obtain a target inspection text corresponding to the inspection voice information. Among them, the target decoding graph is determined based on a character mapping network, a word mapping network, and a sentence mapping network. The character mapping network includes the mapping relationship between frame-level characters and characters, and the frame-level characters correspond to the character recognition results of audio frames. The word mapping network includes the mapping relationship between characters and words, and the sentence mapping network includes the mapping relationship between words and sentences; The apparatus further includes: a mapping network determination unit, configured to determine the word mapping network and determine the sentence mapping network; a decoding graph to be processed determination unit, configured to fuse the word mapping network and the sentence mapping network to obtain a decoding graph to be processed; a decoding graph to be used determination unit, configured to perform pruning processing on the decoding graph to be processed to obtain a decoding graph to be used; a target decoding graph determination unit, configured to fuse the character mapping network and the decoding graph to be used to obtain the target decoding graph; The weight value between two adjacent characters in the word mapping network is determined based on the attribute information in the important attribute, timeliness attribute, accuracy attribute, and novelty attribute of each word to be applied in the first inspection text when performing word segmentation on each word to be applied; 15. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the speech processing method according to any one of claims 1-13.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the speech processing method according to any one of claims 1-13 when executed by a processor.
Citation Information
Patent Citations
Model updating method and device, electronic equipment and storage medium
CN111583910A
Speech recognition method and device, equipment and storage medium
CN112599128A
Power equipment voice recognition method and system in inspection scene
CN113851116A