Speech recognition method, device, electronic device and storage medium

By obtaining target conference information, extracting conference conjunctions and determining the weight of hot words according to the ASR engine rules, the lack of hot words in speech recognition engines is solved, and the recognition rate and adaptability are improved.

CN114664307BActive Publication Date: 2025-08-29BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210264521.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-17
Publication Date
2025-08-29
Estimated Expiration
2042-03-17

AI Technical Summary

Technical Problem

The recognition rate of existing speech recognition engines cannot reach 100% in different scenarios, especially the probability of recognition of hot words is insufficient, resulting in frequent misrecognition.

Method used

By obtaining relevant information of the target meeting, extracting the conference conjunctions collection, and determining the rules and weight characteristics based on the hot words of the target ASR engine, determining whether each conference conjunction is a hot word, and inputting the corresponding hot word weight to the ASR engine to improve the recognition rate.

Benefits of technology

It improves the overall recognition rate of speech recognition, adapts to the scene and iterative changes of different ASR engines, reduces misrecognition, and enhances the ability to extract hot words.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114664307B_ABST
    Figure CN114664307B_ABST
Patent Text Reader

Abstract

The present disclosure provides a speech recognition method, device, electronic device, and storage medium. A specific implementation of the method includes: obtaining relevant information of a target meeting, the relevant information of the target meeting including information related to the target meeting; extracting a set of conference-related words from the relevant information; determining whether each conference-related word is a hot word corresponding to the target ASR engine according to a hot word determination rule corresponding to the target ASR engine for automatic speech recognition; inputting the determined hot words into the target ASR engine according to the corresponding hot word weights to realize automatic speech recognition of the speech data of the target meeting. This implementation not only improves the ability to extract hot words, but also can adapt to different ASR engines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of speech recognition technology, and more particularly to a speech recognition method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of speech recognition technology, a large number of speech recognition engines have emerged. Here, an Automatic Speech Recognition (ASR) engine refers to an application used to recognize speech data into text.

[0003] Due to the limitations of existing technology, speech recognition engines still struggle to achieve 100% recognition. To adapt to different scenarios, most ASR engines support hotword input. This allows for improved recognition of hotwords by inputting them, or by inputting them along with their corresponding speech data. Hotwords are an important tool for intervening in ASR recognition results. Summary of the Invention

[0004] Embodiments of the present disclosure provide a speech recognition method, apparatus, electronic device, and storage medium.

[0005] In the first aspect, an embodiment of the present disclosure provides a speech recognition method, including: obtaining relevant information of a target meeting, the relevant information of the target meeting including information related to the target meeting; extracting a set of conference-related words of the relevant information; determining whether each of the conference-related words is a hot word corresponding to the target ASR engine according to a hot word determination rule corresponding to the target ASR engine for automatic speech recognition; and inputting the determined hot words into the target ASR engine according to corresponding hot word weights to realize automatic speech recognition of the speech data of the target meeting.

[0006] In some optional implementations, the relevant information of the target conference includes at least one of the following: content information and participant information of the target conference.

[0007] In some optional embodiments, the hot word determination rule includes a hot word weight determination rule; and the determination of whether each of the conference-associated words is a hot word corresponding to the target ASR engine according to the hot word determination rule corresponding to the target ASR engine for automatic speech recognition includes: determining the hot word weight of each of the conference-associated words in the target ASR engine according to the hot word weight determination rule corresponding to the target ASR engine for automatic speech recognition; and determining whether the conference-associated word is a hot word according to the hot word weight of each of the conference-associated words in the target ASR engine.

[0008] In some optional embodiments, the hot word weight of each of the conference-associated words in the target ASR engine for automatic speech recognition is determined according to the hot word weight determination rule corresponding to the target ASR engine, including: for each of the conference-associated words, determining the hot word weight of the conference-associated word in the target ASR engine according to the weight feature of the conference-associated word, wherein the weight feature includes at least one of the following elements: an entity type label for characterizing the entity type of the conference-associated word, a preset hot word label for characterizing whether a preset hot word dictionary includes the conference-associated word, and a language model probability for characterizing the probability of the conference-associated word appearing in the target ASR engine.

[0009] In some optional embodiments, for each of the conference-associated words, the hot word weight of the conference-associated word in the target ASR engine is determined according to the weight feature of the conference-associated word, including: for each of the conference-associated words, weighted summing the weights corresponding to the elements contained in the weight feature of the conference-associated word to determine the hot word weight of the conference-associated word in the target ASR engine, wherein the weight corresponding to the entity type label is determined based on the correspondence between the entity type label and the weight corresponding to the target ASR engine, the weight corresponding to the preset hot word label is determined based on the correspondence between the hot word label and the weight corresponding to the target ASR engine, and the weight corresponding to the language model probability is determined based on the correspondence between the language model probability and the weight corresponding to the target ASR engine.

[0010] In some optional embodiments, for each of the conference-associated words, the weights corresponding to the elements included in the weight feature of the conference-associated word are weightedly summed to determine the hot word weight of the conference-associated word in the target ASR engine, including: for each of the conference-associated words, according to the first weight coefficient, the second weight coefficient and the third weight coefficient corresponding to the target ASR engine, the weight corresponding to the entity type label of the conference-associated word in the target ASR engine, the weight corresponding to the preset hot word label and the weight corresponding to the language model probability are weighted summed to obtain the hot word weight of the conference-associated word in the target ASR engine.

[0011] In some optional implementations, the target conference is an ongoing audio or video conference.

[0012] In some optional implementations, determining whether the conference-associated words are hot words based on the hot word weights of the conference-associated words in the target ASR engine includes: determining conference-associated words whose hot word weights are greater than a preset hot word weight threshold as hot words.

[0013] In the second aspect, an embodiment of the present disclosure provides a speech recognition device, which includes: an acquisition unit, configured to acquire relevant information of a target meeting, wherein the relevant information of the target meeting includes information related to the target meeting; an extraction unit, configured to extract a set of conference-related words of the relevant information; a hot word determination unit, configured to determine whether each of the conference-related words is a hot word corresponding to the target ASR engine according to a hot word determination rule corresponding to the target ASR engine for automatic speech recognition; and a speech recognition unit, inputting the determined hot words into the target ASR engine according to the corresponding hot word weights to realize automatic speech recognition of the speech data of the target meeting.

[0014] In some optional implementations, the relevant information of the target conference includes at least one of the following: content information and participant information of the target conference.

[0015] In some optional embodiments, the hot word determination rule includes a hot word weight determination rule; and the hot word determination unit is further configured to: determine the hot word weight of each of the conference-associated words in the target ASR engine according to the hot word weight determination rule corresponding to the target ASR engine for automatic speech recognition; and determine whether the conference-associated word is a hot word according to the hot word weight of each of the conference-associated words in the target ASR engine.

[0016] In some optional embodiments, the hot word determination unit is further configured to: for each of the conference-associated words, determine the hot word weight of the conference-associated word in the target ASR engine based on the weight feature of the conference-associated word, wherein the weight feature includes at least one of the following elements: an entity type label for characterizing the entity type of the conference-associated word, a preset hot word label for characterizing whether a preset hot word dictionary includes the conference-associated word, and a language model probability for characterizing the probability of the conference-associated word appearing in the target ASR engine.

[0017] In some optional embodiments, the hot word determination unit is further configured to: for each of the conference-related words, perform weighted summation on the weights corresponding to the elements contained in the weight feature of the conference-related word to determine the hot word weight of the conference-related word in the target ASR engine, wherein the weight corresponding to the entity type label is determined based on the correspondence between the entity type label and the weight corresponding to the target ASR engine, the weight corresponding to the preset hot word label is determined based on the correspondence between the hot word label and the weight corresponding to the target ASR engine, and the weight corresponding to the language model probability is determined based on the correspondence between the language model probability and the weight corresponding to the target ASR engine.

[0018] In some optional embodiments, the hot word determination unit is further configured to: for each of the conference-associated words, according to the first weight coefficient, the second weight coefficient and the third weight coefficient corresponding to the target ASR engine, perform weighted summation on the weight corresponding to the entity type label of the conference-associated word in the target ASR engine, the weight corresponding to the preset hot word label and the weight corresponding to the language model probability to obtain the hot word weight of the conference-associated word in the target ASR engine.

[0019] In some optional implementations, the target conference is an ongoing audio or video conference.

[0020] In some optional implementations, the hot word determination unit is further configured to: determine a conference-related word whose hot word weight is greater than a preset hot word weight threshold as a hot word.

[0021] In a third aspect, an embodiment of the present disclosure provides an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.

[0022] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by one or more processors, implements the method described in any implementation manner in the first aspect.

[0023] To better apply hotword technology in the speech recognition process, the applicant discovered through practical research that the ASR engine itself has different recognition probabilities for different words. For example, some words are easy to recognize and are called "easy-to-recognize words"; conversely, some words are difficult to recognize and are called "difficult-to-recognize words." For easy-to-recognize words, since the ASR engine itself can recognize them well, these words are not very meaningful as hotwords. On the contrary, they may also have side effects, causing other similar-sounding words to be easily misidentified as the hotwords. An effective hotword should be a difficult-to-recognize word that the ASR engine itself has a low recognition probability. In addition, since different ASR engines have different recognition characteristics, and most ASR engines support the input of hotwords with corresponding weights, the same weight has different effects on different engines. Therefore, it is necessary to determine the hotword weights based on the recognition characteristics of different ASR engines and input the hotword weights into the corresponding ASR engines to achieve automatic speech recognition of the speech data of the target meeting.

[0024] The speech recognition method, device, electronic device and storage medium provided by the embodiments of the present disclosure first extract conference-related words related to the target conference from various information sources from the perspective of focusing on "content relevance". Then, from the perspective of focusing on "engine adaptability", the hot word weight of each conference-related word in the target ASR engine is determined, and then based on the hot word weight, whether the conference-related word is a hot word is determined, and the determined hot word and the corresponding hot word weight are input into the target ASR engine, and the speech data of the target conference is automatically recognized based on the hot word weight, thereby improving the recognition rate of the overall speech recognition. In addition, hot word extraction can only determine the situation in which the conference-related words appear in the target conference, without having to pay attention to whether the target ASR engine has already recognized them well. The determination of hot words can adapt to the scenarios of different ASR engines or multiple ASR engines and the impact of ASR engine iteration. In this way, not only the ability of hot word extraction can be improved, but also different ASR engines can be adapted. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Other features, objects, and advantages of the present disclosure will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings. The drawings are for illustration purposes only and are not to be considered as limiting the present invention. In the drawings:

[0026] Figure 1 is an exemplary system architecture diagram in which an embodiment of the present disclosure may be applied;

[0027] Figure 2 is a flow chart of an embodiment of a speech recognition method according to the present disclosure;

[0028] Figure 3 is a structural diagram of an embodiment of a speech recognition device according to the present disclosure;

[0029] Figure 4 It is a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION

[0030] The present disclosure will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the relevant invention are shown in the accompanying drawings.

[0031] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0032] Figure 1An exemplary system architecture 100 is shown to which embodiments of the speech recognition method, apparatus, electronic device, and storage medium of the present disclosure can be applied.

[0033] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is a medium for providing communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0034] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as text processing applications, voice recognition applications, short video social applications, online conferencing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0035] Terminal devices 101, 102, and 103 can be hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with display screens, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the terminal devices listed above. They can be implemented as multiple software or software modules (for example, to provide voice recognition services), or as a single software or software module. No specific limitations are given here.

[0036] In some cases, the speech recognition method provided by the present disclosure may be executed by the terminal devices 101, 102, and 103. Accordingly, the speech recognition apparatus may be provided in the terminal devices 101, 102, and 103. In this case, the system architecture 100 may not include the server 105.

[0037] In some cases, the speech recognition method provided by the present disclosure can be jointly executed by the terminal devices 101, 102, 103 and the server 105, which is not limited by the present disclosure. Accordingly, the speech recognition apparatus can also be respectively provided in the terminal devices 101, 102, 103 and the server 105.

[0038] In some cases, the speech recognition method provided by the present disclosure may be executed by the server 105 , and accordingly, the speech recognition device may also be set in the server 105 . In this case, the system architecture 100 may also not include the terminal devices 101 , 102 , and 103 .

[0039] It should be noted that the server 105 can be hardware or software. When the server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server 105 is software, it can be implemented as multiple software or software modules (for example, to provide distributed services), or it can be implemented as a single software or software module. No specific limitations are given here.

[0040] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0041] Continue to refer Figure 2 , which shows a process 200 of an embodiment of a speech recognition method according to the present disclosure, the speech recognition method includes the following steps:

[0042] Step 201: Obtain relevant information of the target conference.

[0043] In this embodiment, the execution subject of the speech recognition method (eg Figure 1 The terminal devices 101, 102, 103 shown in the figure can obtain relevant information of the target meeting locally or remotely from other electronic devices connected to the execution entity network.

[0044] Here, the relevant information of the target conference may include various information related to the target conference.

[0045] In some optional implementations, the target conference-related information may include at least one of the following: target conference content information and participant information. Specifically, the target conference content information may include at least one of the following: target conference title information, shared content information within the target conference, and audio or subtitles collected during the target conference. The participant information may include participant names and participant identifications.

[0046] In some optional implementations, the target conference may be an ongoing audio or video conference. Accordingly, the speech or subtitles collected during the target conference included in the relevant information of the target conference may also be speech or subtitles that have already been collected during the target conference.

[0047] Step 202: extract a conference-related word set of relevant information.

[0048] In this embodiment, the execution entity can extract the conference-related word set from the relevant information of the target conference in various implementation methods. Here, the conference-related words can be words related to the target conference, that is, words that are likely to appear during the target conference.

[0049] As an example, a machine learning or data mining algorithm may be used to extract a conference-related word set from relevant information, or the conference-related word set may be manually selected from relevant information.

[0050] In some optional implementations, a keyword extraction algorithm can be used to extract a set of conference-related keywords from the relevant information. The keyword extraction algorithm can be, for example, an unsupervised algorithm such as a term frequency-inverse document frequency (TF-IDF) algorithm or topic similarity, or a supervised algorithm such as a statistical machine translation (SMT) model or a sequence tagging model.

[0051] Furthermore, the conference-related word set may be filtered. For example, candidate words with low occurrence frequency, preset stop words, and function words such as prepositions, conjunctions, auxiliary words, and modal particles may be filtered out from the conference-related word set.

[0052] Step 203 : Determine whether each conference-related word is a hot word corresponding to the target ASR engine according to a hot word determination rule corresponding to the target ASR engine for automatic speech recognition.

[0053] Here, the ASR engine may be an application for recognizing speech data as text. The target ASR engine may be an application for recognizing speech data of a target meeting as text. The hotword determination rule may be used to characterize a specific hotword determination strategy. The hotword determination rule may be used to determine whether a conference-related word is a hotword corresponding to the target ASR engine.

[0054] Specifically, the hot word determination rule may include at least one hot word determination item and a corresponding hot word determination condition. The hot word determination item may be related to the target ASR engine feature. The hot word determination item may also be related to the hot word feature. The hot word determination condition may be whether the hot word determination item meets the hot word condition.

[0055] In this embodiment, the execution entity can determine the hot word determination rule corresponding to the ASR engine identifier of the target ASR engine based on the preset correspondence between the ASR engine identifier and the hot word determination rule. That is, corresponding hot word determination rules can be pre-set for different ASR engines. Then, based on the hot word determination rule corresponding to the determined target ASR engine, it is determined whether each hot word determination item of each conference-related word meets the corresponding hot word determination condition. Conference-related words that meet the hot word determination condition are determined as the hot words corresponding to the target ASR engine.

[0056] In some optional implementations, the hot word determination rule may include a hot word weight determination rule. Further, step 203 may include steps 2031 and 2032.

[0057] Step 2031 : Determine the hot word weight of each conference-related word in the target ASR engine according to a hot word weight determination rule corresponding to the target ASR engine for automatic speech recognition.

[0058] For example, the hot word weight determination rule corresponding to the target ASR engine may be, for example, a hot word weight determination correspondence table. The hot word weight determination correspondence table is used to characterize the correspondence between different indicator items including hot words and the weight coefficients of the indicator items. The indicator items may be related to the target ASR engine, for example, the language model probability of the hot word in the target ASR engine. The weight coefficient may represent the degree of influence of the corresponding indicator item on the hot word weight of the hot word in the target ASR engine. The indicator item may also be unrelated to the target ASR engine.

[0059] In some optional implementations, step 2031 may be performed as follows: for each conference-related word, determine the hot word weight of the conference-related word in the target ASR engine according to the weight feature of the conference-related word.

[0060] Here, the weight feature may include at least one of the following elements: an entity type label for characterizing the entity type of the conference-associated word, a preset hot word label for characterizing whether the preset hot word dictionary includes the conference-associated word, and a language model probability for characterizing the probability of the conference-associated word appearing in the target ASR engine.

[0061] In this optional embodiment, the above-mentioned execution entity can directly obtain the entity type label of the conference-related word from the pre-stored correspondence table between conference-related words and entity type labels, or can obtain the entity type label of the conference-related word through, for example, NER (Named Entity Recognition) based on CRF (Conditional random field). A conference-related word can have one or more entity type labels. The entity type label can be used to characterize the entity type of the conference-related word, such as time, name, place name, common word, etc. Common words can be words other than time, name, and place name. For example, the entity type label corresponding to the conference-related word "Beijing" can be "city name". By obtaining the entity type label and then considering the difficulty of the target ASR engine to recognize different entity types, the degree of influence of the entity type label of the conference-related word on the hot word weight of the conference-related word input as a hot word into the target ASR engine can be controlled.

[0062] The preset hot word dictionary can be pre-set based on experience. For example, the preset hot word dictionary can include words with a low recognition rate of the target ASR engine, that is, words that the target ASR can recognize more accurately. The preset hot word label is, for example, "hot word", which indicates that the preset hot word dictionary includes the conference-related words. The preset hot word label is, for example, "non-hot word", which indicates that the preset hot word dictionary does not include the conference-related words. Through the preset hot word label of the conference-related words, the degree of influence of the preset hot word label on the hot word weight of the conference-related words input as hot words into the target ASR engine can be controlled.

[0063] For example, the preset hot word dictionary may include a hot word whitelist and a hot word blacklist. The hot word blacklist may include words with a high recognition rate by the target ASR engine. The hot word whitelist may include words with a low recognition rate by the target ASR engine. If a conference-related word is on the hot word blacklist, it indicates that the target ASR can accurately recognize the conference-related word, and inputting the conference-related word into the target ASR as a hot word is not very meaningful. If a conference-related word is on the hot word blacklist, it indicates that the target ASR has difficulty accurately recognizing the conference-related word. The preset hot word label, for example, is "hot word," indicating that the conference-related word is included in the hot word whitelist. The preset hot word label, for example, is "non-hot word," indicating that the conference-related word is included in the hot word blacklist. The aforementioned execution entity may obtain the language model probability of each conference-related word from the language model of the target ASR engine. The language model probability of a conference-related word in the target ASR engine can be used to represent the probability of the conference-related word appearing. The higher the language model probability of a conference-related word, the easier it is for the target ASR engine's language model to recognize the conference-related word. The target ASR engine may include one or more language models. The language model may be an N-gram language model. The language model probability corresponding to the same conference-related word in different language models can be different. For example, the language model probability corresponding to "Beijing" in language model A may be 1%, while the language model probability corresponding to "Beijing" in language model B may be 1.5%. By obtaining the language model probability of the conference-related word and considering the recognition of the conference-related word in the language model of the target ASR engine, the degree to which the language model probability of the conference-related word in the target ASR engine affects the hot word weight of the conference-related word when it is input as a hot word into the target ASR engine can be controlled.

[0064] Furthermore, step 2031 may also be performed as follows: for each conference-related word, weights corresponding to the elements included in the weight feature of the conference-related word are weighted and summed to determine the hot word weight of the conference-related word in the target ASR engine. The weight corresponding to the entity type tag is determined based on the corresponding relationship between the entity type tag and the weight corresponding to the target ASR engine, the weight corresponding to the preset hot word tag is determined based on the corresponding relationship between the hot word tag and the weight corresponding to the target ASR engine, and the weight corresponding to the language model probability is determined based on the corresponding relationship between the language model probability and the weight corresponding to the target ASR engine.

[0065] On the one hand, for each conference-related word, the execution entity can determine the weight corresponding to the entity type tag of the conference-related word based on a correspondence table corresponding to the target ASR engine, which is used to characterize the correspondence between the entity type tag and the first weight. For example, in the correspondence table, the weight corresponding to the entity type tag that is difficult for the target ASR engine to recognize is larger than the weight corresponding to the entity type tag that is easy for the target ASR engine to recognize. In other words, the greater the weight corresponding to the entity type tag of the conference-related word, the greater the possibility that the conference-related word will be determined as a hot word.

[0066] On the other hand, the execution entity may determine the weight corresponding to the preset hot word label of the conference-related word based on a correspondence table corresponding to the target ASR engine and used to characterize the correspondence between the preset hot word labels and the weights. For example, the weight corresponding to the preset hot word label used to characterize that the preset hot word dictionary includes the conference-related word is larger than the weight corresponding to the preset hot word label used to characterize that the preset hot word dictionary does not include the conference-related word.

[0067] For example, the weight corresponding to the preset hot word tag used to indicate that the hot word whitelist includes the conference-related word is greater than the weight corresponding to the preset hot word tag used to indicate that the hot word blacklist includes the conference-related word. In other words, the greater the weight corresponding to the preset hot word tag of the conference-related word, the greater the possibility that the conference-related word is determined to be a hot word.

[0068] On the other hand, the execution entity may determine the weight corresponding to the language model probability of the conference-related word based on a corresponding relationship corresponding to the target ASR engine, which is used to represent the corresponding relationship between language model probability and weight. For example, the greater the language model probability of the conference-related word, the smaller the weight corresponding to the language model probability of the conference-related word. In other words, the greater the weight corresponding to the language model probability of the conference-related word, the greater the likelihood that the conference-related word will be determined as a hot word.

[0069] It should be noted that when the weight feature includes at least two elements, the weights corresponding to the at least two elements can be obtained successively or simultaneously.

[0070] Finally, for each conference-related word, the execution entity may perform a weighted summation on the weights corresponding to the elements included in the weight feature of the conference-related word to determine the hot word weight of the conference-related word in the target ASR engine.

[0071] Furthermore, for each conference-related word, the weights corresponding to the elements contained in the weight feature of the conference-related word can be weighted and summed to determine the hot word weight of the conference-related word in the target ASR engine. It can also be performed as follows: for each conference-related word, according to the first weight coefficient, the second weight coefficient and the third weight coefficient corresponding to the target ASR engine, the weight corresponding to the entity type label of the conference-related word in the target ASR engine, the weight corresponding to the preset hot word label and the weight corresponding to the language model probability are weighted and summed to obtain the hot word weight of the conference-related word in the target ASR engine.

[0072] Here, the first weight coefficient can represent the importance of the weight corresponding to the entity type label affecting the hot word weight. The second weight coefficient can represent the importance of the weight corresponding to the preset hot word label affecting the hot word weight. The third weight coefficient can represent the importance of the weight corresponding to the language model probability affecting the hot word weight. For example, for example, the hot word weight is S, the weight corresponding to the entity type label is recorded as S1, the weight corresponding to the preset hot word label is recorded as S2, and the weight corresponding to the language model probability is recorded as S3, which can be expressed as S=k1S1+k2S2+k3S3, where k1 is the first weight coefficient corresponding to S1, k2 is the second weight coefficient corresponding to S2, and k3 is the third weight coefficient corresponding to S3. The specific values ​​of k1, k2 and k3 can be set according to the importance of the weight affecting the hot word weight.

[0073] Step 2032: Determine whether the conference-related word is a hot word based on the hot word weight of each conference-related word in the target ASR engine.

[0074] Here, the execution subject may adopt various implementation methods to determine whether the conference-related words are hot words according to the hot word weights of the conference-related words in the target ASR engine.

[0075] In some optional implementations, whether a conference-related word is a hot word may be determined in the following manner: a conference-related word having a hot word weight greater than a preset hot word weight threshold is determined as a hot word.

[0076] In this optional embodiment, the preset hot word weight threshold may be a pre-set smaller value, which may be fixed or customized according to actual conditions. That is, if the hot word weight of a conference-related word is small, it can be considered that the target ASR engine is already able to recognize the conference-related word well, and thus the conference-related word is not very meaningful as a hot word.

[0077] Through this implementation, conference-related words that have been well-recognized by the target ASR engine can be effectively filtered, thereby preventing the well-recognized conference-related words from being input into the target ASR engine as hot words.

[0078] It's important to note that for well-recognized conference-related words, since the ASR engine itself can already recognize them well, using them as hot words is not very meaningful. On the contrary, it may also have the side effect of causing other similar-sounding words to be mistakenly recognized as these conference-related words. An effective hot word should be a difficult word that the target ASR engine has a low probability of recognizing.

[0079] In step 204, each determined hot word is input into the target ASR engine according to the corresponding hot word weight, so as to realize automatic speech recognition of the speech data of the target conference.

[0080] In this embodiment, the execution entity may input the conference-related words determined as hot words in step 203 and their corresponding hot word weights into a target ASR engine, and then use the target ASR engine to perform automatic speech recognition on the speech data of the target meeting, thereby obtaining the text corresponding to the speech data of the target meeting. Here, the target ASR engine may be a speech recognition engine that supports the input of hot words and corresponding hot word weights.

[0081] Specifically, during the decoding process of the target conference's voice data, a determination is made as to whether the decoding path contains a conference-related term identified as a hot word. If so, a weighted incentive is applied to the corresponding decoding path based on the hot word weight corresponding to the conference-related term, thereby improving the accuracy of identifying the conference-related term identified as a hot word and, consequently, the speech recognition rate.

[0082] The speech recognition method provided by the above-mentioned embodiment of the present disclosure first extracts conference-related words related to the target conference from various information sources from the perspective of focusing on "content relevance". Then, from the perspective of focusing on "engine adaptability", the hot word weight of each conference-related word in the target ASR engine is determined, and then based on the hot word weight, whether the conference-related word is a hot word is determined, and the determined hot word and the corresponding hot word weight are input into the target ASR engine, and the speech data of the target conference is automatically recognized based on the hot word weight, thereby improving the recognition rate of the overall speech recognition. In addition, hot word extraction can only determine the situation in which the conference-related words appear in the target conference, without paying attention to whether the target ASR engine has recognized them well. The determination of hot words can adapt to the scenarios of different ASR engines or multiple ASR engines and the impact of ASR engine iteration. In this way, not only the ability of hot word extraction can be improved, but also different ASR engines can be adapted.

[0083] Further references Figure 3 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a speech recognition device. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0084] like Figure 3 As shown, the speech recognition device 300 of this embodiment includes: an acquisition unit 301, an extraction unit 302, a hot word determination unit 303, and a speech recognition unit 304. The acquisition unit 301 is configured to acquire relevant information of a target meeting, wherein the relevant information of the target meeting includes information related to the target meeting; the extraction unit 302 is configured to extract a set of meeting-related words from the relevant information; the hot word determination unit 303 is configured to determine whether each meeting-related word is a hot word corresponding to the target ASR engine according to the hot word determination rules corresponding to the target ASR engine for automatic speech recognition; and the speech recognition unit 304 inputs the determined hot words into the target ASR engine according to the corresponding hot word weights to realize automatic speech recognition of the speech data of the target meeting.

[0085] In this embodiment, the specific processing of the acquisition unit 301, the extraction unit 302, the hot word determination unit 303 and the speech recognition unit 304 of the speech recognition device 300 and the technical effects thereof can be referred to in the respective Figure 2 The relevant descriptions of step 201, step 202, step 203 and step 204 in the corresponding embodiment are not repeated here.

[0086] In some optional implementations, the relevant information of the target conference may include at least one of the following: content information and participant information of the target conference.

[0087] In some optional embodiments, the hot word determination rule may include a hot word weight determination rule; and the hot word determination unit 303 may be further configured to: determine the hot word weight of each conference-associated word in the target ASR engine according to the hot word weight determination rule corresponding to the target ASR engine for automatic speech recognition; and determine whether the conference-associated word is a hot word according to the hot word weight of each conference-associated word in the target ASR engine.

[0088] In some optional embodiments, the hot word determination unit 303 can be further configured to: for each conference-related word, determine the hot word weight of the conference-related word in the target ASR engine based on the weight feature of the conference-related word, wherein the weight feature includes at least one of the following elements: an entity type label for characterizing the entity type of the conference-related word, a preset hot word label for characterizing whether the preset hot word dictionary includes the conference-related word, and a language model probability for characterizing the probability of the conference-related word appearing in the target ASR engine.

[0089] In some optional embodiments, the hot word determination unit 303 can be further configured to: for each conference-related word, perform weighted summation on the weights corresponding to the elements contained in the weight feature of the conference-related word to determine the hot word weight of the conference-related word in the target ASR engine, wherein the weight corresponding to the entity type label is determined based on the correspondence between the entity type label and the weight corresponding to the target ASR engine, the weight corresponding to the preset hot word label is determined based on the correspondence between the hot word label and the weight corresponding to the target ASR engine, and the weight corresponding to the language model probability is determined based on the correspondence between the language model probability and the weight corresponding to the target ASR engine.

[0090] In some optional embodiments, the hot word determination unit 303303 can be further configured to: for each conference-related word, according to the first weight coefficient, the second weight coefficient and the third weight coefficient corresponding to the target ASR engine, the weight corresponding to the entity type label of the conference-related word in the target ASR engine, the weight corresponding to the preset hot word label and the weight corresponding to the language model probability are weightedly summed to obtain the hot word weight of the conference-related word in the target ASR engine.

[0091] In some optional implementations, the target conference may be an ongoing audio or video conference.

[0092] In some optional implementations, the hot word determining unit 303 may be further configured to: determine a conference-related word whose hot word weight is greater than a preset hot word weight threshold as a hot word.

[0093] It should be noted that the implementation details and technical effects of each unit in the speech recognition device provided by the embodiments of the present disclosure can be referred to the description of other embodiments in the present disclosure, and will not be repeated here.

[0094] Reference below Figure 4 , which shows a schematic structural diagram of a computer system 400 suitable for implementing the electronic device of the present disclosure. Figure 4 The computer system 400 shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.

[0095] like Figure 4 As shown, the computer system 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. Various programs and data required for the operation of the computer system 400 are also stored in the RAM 403. The processing device 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0096] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the computer system 400 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 4 The computer system 400 of the electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0097] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0098] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0099] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0100] The computer readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device can realize the following operation: Figure 2 The illustrated embodiment and its alternative implementations illustrate a method for speech recognition.

[0101] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0102] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0103] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. The name of a unit does not necessarily limit the unit itself. For example, an acquisition unit may also be described as a "unit for acquiring relevant information about a target meeting."

[0104] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the present disclosure.

Claims

1. A speech recognition method, comprising: Acquire relevant information of a target conference, wherein the relevant information of the target conference includes information related to the target conference; Extracting a conference-related word set of the relevant information; determining whether each of the conference-related words is a hot word corresponding to the target ASR engine according to a hot word determination rule corresponding to the target ASR engine for automatic speech recognition; as well as The determined hot words are input into the target ASR engine according to the corresponding hot word weights to realize automatic speech recognition of the speech data of the target meeting. The hot word determination rule includes a hot word weight determination rule; and The step of determining whether each of the conference-related words is a hot word corresponding to the target ASR engine according to a hot word determination rule corresponding to the target ASR engine for automatic speech recognition includes: Determining the hot word weight of each of the conference-related words in the target ASR engine according to a hot word weight determination rule corresponding to the target ASR engine for automatic speech recognition; as well as According to the hot word weight of each conference-related word in the target ASR engine, it is determined whether the conference-related word is a hot word.

2. The method according to claim 1, wherein The relevant information of the target conference includes at least one of the following: content information and participant information of the target conference.

3. The method according to claim 1, wherein The step of determining the hot word weight of each conference-related word in the target ASR engine according to a hot word weight determination rule corresponding to the target ASR engine for automatic speech recognition includes: For each of the conference-associated words, the hot word weight of the conference-associated word in the target ASR engine is determined based on the weight feature of the conference-associated word, wherein the weight feature includes at least one of the following elements: an entity type label for characterizing the entity type of the conference-associated word, a preset hot word label for characterizing whether a preset hot word dictionary includes the conference-associated word, and a language model probability for characterizing the probability of the conference-associated word appearing in the target ASR engine.

4. The method according to claim 3, wherein: For each of the conference-related words, determining the hot word weight of the conference-related word in the target ASR engine according to the weight feature of the conference-related word includes: For each of the conference-related words, the weights corresponding to the elements contained in the weight feature of the conference-related word are weighted and summed to determine the hot word weight of the conference-related word in the target ASR engine, wherein the weight corresponding to the entity type label is determined based on the correspondence between the entity type label and the weight corresponding to the target ASR engine, the weight corresponding to the preset hot word label is determined based on the correspondence between the hot word label and the weight corresponding to the target ASR engine, and the weight corresponding to the language model probability is determined based on the correspondence between the language model probability and the weight corresponding to the target ASR engine.

5. The method according to claim 4, wherein For each of the conference-related words, weighted summing of the weights corresponding to the elements included in the weight feature of the conference-related word to determine the hot word weight of the conference-related word in the target ASR engine includes: For each of the conference-related words, according to the first weight coefficient, the second weight coefficient and the third weight coefficient corresponding to the target ASR engine, the weight corresponding to the entity type label of the conference-related word in the target ASR engine, the weight corresponding to the preset hot word label and the weight corresponding to the language model probability are weighted and summed to obtain the hot word weight of the conference-related word in the target ASR engine.

6. The method according to claim 1, wherein The target conference is an ongoing audio or video conference.

7. The method according to claim 1, wherein The determining whether the conference-related words are hot words according to the hot word weights of the conference-related words in the target ASR engine includes: Conference-related words whose hot word weights are greater than a preset hot word weight threshold are determined as hot words.

8. A speech recognition device, comprising: an acquiring unit configured to acquire relevant information of a target conference, wherein the relevant information of the target conference includes information related to the target conference; an extraction unit configured to extract a conference-related word set of the relevant information; as well as The hot word determination unit is configured to determine whether each of the conference-related words is a hot word corresponding to the target ASR engine for automatic speech recognition according to a hot word determination rule corresponding to the target ASR engine, wherein the hot word determination rule includes a hot word weight determination rule. The hot word determination unit is further configured to: determine the hot word weight of each of the conference-related words in the target ASR engine according to a hot word weight determination rule corresponding to the target ASR engine for automatic speech recognition; and determining whether each conference-related word is a hot word according to the hot word weight of each conference-related word in the target ASR engine; The speech recognition unit inputs the determined hot words into the target ASR engine according to the corresponding hot word weights to realize automatic speech recognition of the speech data of the target meeting.

9. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by one or more processors, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Conference audio control method, system and device and computer readable storage medium

    CN110300001A

  • Voice recognition method, apparatus and device and medium

    CN110544477A

  • Speech recognition method and device, electronic equipment and storage medium

    CN112037792A