Speech recognition method and device, electronic equipment and storage medium
By updating the preset hot word set during the speech recognition process and utilizing the association between the proper nouns of the recognized audio and its associated relationship, the problem of untimely hot word set updates is solved, and the accuracy of speech recognition is improved, especially the recognition effect of proper nouns.
Patent Information
- Application Number
- CN202410289259.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-13
- Publication Date
- 2025-09-26
AI Technical Summary
In existing speech recognition methods, the preset hot word set is not updated in a timely manner, resulting in a low recognition accuracy of proper nouns with specific meanings.
By obtaining the proper nouns of the recognized audio, searching for the proper nouns associated with the proper nouns, and updating them to the preset hot word set, the hot word set can be dynamically adjusted for speech recognition.
The accuracy of speech recognition has been improved, especially the recognition of proper nouns with specific meanings.
Smart Images

Figure CN120708603A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a speech recognition method and device, an electronic device, and a storage medium. Background Art
[0002] Speech recognition is a technology that converts speech data into text data. With the continuous development of speech recognition technology, speech recognition has a better recognition effect on commonly used words, but the recognition accuracy of proper nouns with specific meanings is low.
[0003] The existing method of speech recognition is to identify the audio to be recognized input by the user through a preset hot word set. The preset hot word set is configured before speech recognition, and the preset hot word set is fixed during the recognition process. Since users rarely actively update the preset hot word set, the preset hot word set is not updated in a timely manner, which in turn leads to a low recognition accuracy rate for the audio to be recognized. Summary of the Invention
[0004] The present disclosure provides a speech recognition method and apparatus, an electronic device, and a storage medium, the main purpose of which is to solve the problem of low recognition accuracy of audio to be recognized.
[0005] According to a first aspect of the present disclosure, a method for speech recognition is provided, comprising:
[0006] Obtaining audio to be recognized and obtaining a first proper noun in a first recognition result of the recognized audio; the audio to be recognized and the recognized audio are adjacent components of a continuous audio segment;
[0007] Searching a pre-established proper noun database for at least one second proper noun associated with the first proper noun;
[0008] The at least one second proper noun found is updated in the preset hot word set, so as to perform speech recognition on the audio to be recognized based on the updated preset hot word set to obtain a target recognition result.
[0009] Optionally, performing speech recognition on the audio to be recognized based on the updated preset hot word set to obtain a target recognition result includes:
[0010] Performing speech recognition on the audio to be recognized to obtain at least one second recognition result and a first probability corresponding to each of the second recognition results;
[0011] Determining the similarity between each hot word in the updated preset hot word set and the at least one second recognition result;
[0012] If there is a target second recognition result whose similarity exceeds a preset similarity threshold among the at least one second recognition result, assigning a second probability to the target second recognition result;
[0013] Adding the first probability and the second probability corresponding to the second recognition result of the target to obtain an updated first probability of the second recognition result of the target;
[0014] The second recognition result corresponding to the maximum value among all the first probabilities is determined as the target recognition result of the audio to be recognized.
[0015] Optionally, after determining the second recognition result corresponding to the maximum value among all the first probabilities as the target recognition result of the audio to be recognized, the method further includes:
[0016] Matching the target recognition results with all proper nouns in the proper noun library respectively;
[0017] If the matching result is that the target recognition result is identical to the target proper noun in the proper noun library, searching the proper noun library for at least one third proper noun associated with the target proper noun;
[0018] The at least one third proper noun is updated in the preset hot word set, so as to perform speech recognition on the remaining audio to be recognized in the continuous audio based on the updated preset hot word set to obtain a recognition result corresponding to the remaining audio to be recognized.
[0019] Optionally, after updating the at least one third proper noun in the preset hot word set, the method further includes:
[0020] Determining whether the audio to be recognized is the last audio in the continuous audio;
[0021] If it is determined that the audio to be recognized is the last audio in the continuous audio, the at least one second proper noun and the at least one third proper noun in the preset hot word set are deleted.
[0022] Optionally, searching a pre-established proper noun library for at least one second proper noun associated with the first proper noun includes:
[0023] respectively obtaining a correlation score between a fourth proper noun in the proper noun library and the first proper noun; the fourth proper noun being a proper noun having a correlation relationship with the first proper noun;
[0024] If the association relationship score is greater than a preset threshold, the fourth proper noun corresponding to the association relationship score is determined as the second proper noun.
[0025] Optionally, the method for constructing the proper noun library includes:
[0026] In response to the received instruction to establish an association relationship between at least two proper nouns, acquiring the at least two proper nouns;
[0027] Based on the association relationship instruction, the association relationship between the at least two proper nouns is bound to obtain the association relationship between the at least two proper nouns, so as to complete the establishment of the proper noun library.
[0028] According to a second aspect of the present disclosure, there is provided a speech recognition apparatus, comprising:
[0029] an acquisition unit, which acquires the audio to be recognized and the first proper noun in the first recognition result of the recognized audio; the audio to be recognized and the recognized audio are adjacent components of a continuous audio segment;
[0030] a search unit, configured to search a pre-established proper noun library for at least one second proper noun associated with the first proper noun;
[0031] an updating unit, configured to update the at least one second proper noun found in the preset hot word set;
[0032] The recognition unit is used to perform speech recognition on the audio to be recognized based on the updated preset hot word set to obtain a target recognition result.
[0033] Optionally, the identification unit includes:
[0034] a recognition module, configured to perform speech recognition on the audio to be recognized, and obtain at least one second recognition result and a first probability corresponding to each of the second recognition results;
[0035] A first determination module is configured to determine a similarity between each hot word in the updated preset hot word set and the at least one second recognition result;
[0036] An allocating module, configured to allocate a second probability to a target second recognition result when there is a target second recognition result whose similarity exceeds a preset similarity threshold in the at least one second recognition result;
[0037] a calculation module, configured to sum the first probability and the second probability corresponding to the second recognition result of the target to obtain an updated first probability of the second recognition result of the target;
[0038] The second determination module is configured to determine a second recognition result corresponding to a maximum value among all the first probabilities as the target recognition result of the audio to be recognized.
[0039] Optionally, the device further includes:
[0040] a matching unit, configured to, after determining the second recognition result corresponding to the maximum value among all the first probabilities as the target recognition result of the audio to be recognized, match the target recognition result with all the proper nouns in the proper noun library respectively;
[0041] The search unit is further configured to, when a matching result indicates that the target recognition result is identical to a target proper noun in the proper noun library, search the proper noun library for at least one third proper noun associated with the target proper noun;
[0042] The updating unit is further configured to update the at least one third proper noun in the preset hot word set, so as to perform speech recognition on the remaining audio to be recognized in the continuous audio based on the updated preset hot word set to obtain a recognition result corresponding to the remaining audio to be recognized.
[0043] Optionally, the device further includes:
[0044] a judging unit, configured to, after updating the at least one third proper noun in the preset hot word set, judge whether the audio to be recognized is the last audio in the continuous audio;
[0045] A deleting unit is configured to delete the at least one second proper noun and the at least one third proper noun in the preset hot word set when it is determined that the audio to be recognized is the last audio in the continuous audio.
[0046] Optionally, the search unit includes:
[0047] an acquisition module, configured to respectively acquire a correlation score between a fourth proper noun and the first proper noun in the proper noun library; the fourth proper noun being a proper noun having a correlation relationship with the first proper noun;
[0048] The determining module is configured to determine the fourth proper noun corresponding to the association relationship score as the second proper noun when the association relationship score is greater than a preset threshold.
[0049] Optionally, the device further includes:
[0050] a second acquiring unit, configured to acquire the at least two proper nouns in response to a received instruction to establish an association relationship between the at least two proper nouns;
[0051] The binding unit is configured to bind the association relationship between the at least two proper nouns based on the association relationship instruction to obtain the association relationship between the at least two proper nouns, so as to complete the establishment of the proper noun library.
[0052] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0053] at least one processor; and
[0054] a memory communicatively connected to the at least one processor; wherein,
[0055] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect.
[0056] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the first aspect.
[0057] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method as described in the first aspect above.
[0058] The speech recognition method and device, electronic device and storage medium provided by the present disclosure obtain audio to be recognized and a first proper noun in a first recognition result of the recognized audio; the audio to be recognized and the recognized audio are adjacent components of a continuous audio segment; at least one second proper noun that has an association with the first proper noun is searched in a pre-established proper noun library; the at least one second proper noun found is updated in a preset hot word set, so that speech recognition is performed on the audio to be recognized based on the updated preset hot word set to obtain a target recognition result. Compared with the related art, the embodiment of the present disclosure actively updates the preset hot word set by obtaining the proper nouns in the recognition results of the recognized audio, finding at least one second proper noun that is associated with the first proper noun, and updating the at least one second proper noun in the preset hot word set. The audio to be recognized and the recognized audio are adjacent components of a continuous audio segment, and the recognized audio and the audio to be recognized are associated, so the at least one second proper noun and the recognition results of the audio to be recognized are also associated, achieving the effect of hot word enhancement on the recognition results of the audio to be recognized by the updated preset hot word set, thereby improving the accuracy of speech recognition.
[0059] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0061] Figure 1 A flowchart of a speech recognition method provided by an embodiment of the present disclosure;
[0062] Figure 2 A flowchart illustrating a method for performing speech recognition on audio to be recognized based on a preset hot word set provided by an embodiment of the present disclosure;
[0063] Figure 3 A schematic diagram of the structure of a speech recognition device provided by an embodiment of the present disclosure;
[0064] Figure 4 A schematic structural diagram of another speech recognition device provided by an embodiment of the present disclosure;
[0065] Figure 5 A schematic block diagram of an exemplary electronic device provided for an embodiment of the present disclosure. DETAILED DESCRIPTION
[0066] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0067] The following describes the speech recognition method and apparatus, electronic device, and storage medium according to embodiments of the present disclosure with reference to the accompanying drawings.
[0068] Figure 1 A flowchart of a speech recognition method provided by an embodiment of the present disclosure.
[0069] like Figure 1 As shown, the method is applied to a processor, and the method includes the following steps:
[0070] Step 101: obtain audio to be recognized and obtain a first proper noun in a first recognition result of the recognized audio; the audio to be recognized and the recognized audio are adjacent components of a continuous audio segment.
[0071] Proper nouns are words that have special meaning and reference in a specific context. Proper nouns include but are not limited to names of people, places, organizations, brand names, etc.
[0072] During the speech recognition process, speech recognition is performed on a continuous audio segment in frames. For example, a continuous audio segment is divided into three frames, namely a, b, and c. Speech recognition is performed on a first, then on b, and finally on c.
[0073] By obtaining the audio to be recognized and the first proper noun in the first recognition result of the recognized audio, the context and theme of the entire audio can be better understood, thereby guiding the speech recognition process of the audio to be recognized.
[0074] Step 102: Search a pre-established proper noun database for at least one second proper noun associated with the first proper noun.
[0075] There are associations between different proper nouns. For example, if the first proper noun is A, the second proper noun may be B, C, or other tourist attractions of A.
[0076] A second proper noun that is associated with the first proper noun in the recognition results of the recognized audio may appear in the recognition results of the audio to be recognized. By searching for at least one second proper noun that is associated with the first proper noun in a pre-established proper noun library, the understanding of the content of the audio to be recognized can be expanded and deepened, and the accuracy of speech recognition of the audio to be recognized can be improved.
[0077] Step 103 : updating the at least one second proper noun found in the preset hot word set, so as to perform speech recognition on the audio to be recognized based on the updated preset hot word set to obtain a target recognition result.
[0078] In speech recognition, some words or characters are easily misrecognized. In order to improve the accuracy of speech recognition of words or characters that are easily misrecognized, some words or characters that are easily misrecognized can be stored in a preset hot word set. When recognizing the audio to be recognized, if the same recognition result as the hot word in the hot word set appears, the recognition result can be enhanced to improve the recognition accuracy of the hot word, where the hot word is a word or character that is given special attention during the speech recognition process and needs to be accurately recognized.
[0079] By updating the at least one second proper noun found in the preset hot word set, the preset hot word set can be dynamically adjusted. Since the second proper noun is associated with the first proper noun of the recognized audio, the second proper noun is likely to appear in the recognition results of the audio to be recognized. Updating the second proper noun in the preset hot word set can improve the recognition accuracy of the audio to be recognized.
[0080] The present disclosure provides a speech recognition method, which obtains an audio to be recognized and a first proper noun in a first recognition result of the recognized audio; the audio to be recognized and the recognized audio are adjacent components of a continuous audio segment; at least one second proper noun associated with the first proper noun is searched in a pre-established proper noun library; the at least one second proper noun found is updated in a preset hot word set, so that the audio to be recognized is recognized based on the updated preset hot word set to perform speech recognition on the audio to be recognized and obtain a target recognition result. Compared with the related art, the embodiment of the present disclosure actively updates the preset hot word set by obtaining the proper noun in the recognition result of the recognized audio, finding at least one second proper noun associated with the first proper noun, and updating the at least one second proper noun in the preset hot word set. The audio to be recognized and the recognized audio are adjacent components of a continuous audio segment, and the recognized audio and the audio to be recognized are associated, so the recognition result of the at least one second proper noun and the audio to be recognized are also associated, so that the updated preset hot word set has a hot word enhancement effect on the recognition result of the audio to be recognized, thereby improving the accuracy of speech recognition.
[0081] As a refinement of step 103, when performing speech recognition on the audio to be recognized based on the updated preset hot word set to obtain the target recognition result, it can be implemented in the following ways but not limited to: Figure 2 As shown, Figure 2 A flowchart illustrating a method for performing speech recognition on audio to be recognized based on a preset hotword set provided in an embodiment of the present disclosure includes:
[0082] Step 201 : Perform speech recognition on the audio to be recognized to obtain at least one second recognition result and a first probability corresponding to each of the second recognition results.
[0083] The audio to be recognized is input into the speech recognition model, and the speech recognition model will output multiple second recognition results. Each second recognition result has a corresponding first probability, where the first probability includes audio probability and language probability. The audio probability is the probability of the audio to be recognized in different speech units (such as phonemes, characters or words), and the language probability is the probability of the recognition result of the audio to be recognized appearing in natural language.
[0084] Step 202: Determine the similarity between each hot word in the updated preset hot word set and the at least one second recognition result.
[0085] By matching each hot word in the preset hot word set with the at least one second recognition result, it can be determined whether there is a second recognition result similar to a hot word in the preset hot word set among the multiple second recognition results of the audio to be recognized.
[0086] Step 203: If there is a target second recognition result whose similarity exceeds a preset similarity threshold in the at least one second recognition result, a second probability is assigned to the target second recognition result.
[0087] When there is a target second recognition result similar to a hot word in a preset hot word set among multiple second recognition results of the audio to be recognized, it can be determined that the target second recognition result is most likely to be the final recognition result of the audio to be recognized. Therefore, in order to improve the accuracy of speech recognition of the audio to be recognized, it is necessary to assign a second probability to the target second recognition result, where the second probability is an arbitrary value, for example, 10%.
[0088] Step 204 : Add the first probability and the second probability corresponding to the second target recognition result to obtain an updated first probability of the second target recognition result.
[0089] For ease of understanding, an example is provided. Assuming that the first probability is 15% and the second probability is 10%, the updated first probability is 25%.
[0090] Step 205: Determine the second recognition result corresponding to the maximum value among all the first probabilities as the target recognition result of the audio to be recognized.
[0091] Since there are multiple second recognition results for the audio to be recognized, it is necessary to select one second recognition result from the multiple recognition results as the final target recognition result of the audio to be recognized. The second recognition result corresponding to the maximum value of all first probabilities can be determined as the target recognition result of the audio to be recognized.
[0092] For ease of understanding, an example is provided. Suppose that the second recognition results of the audio to be recognized are d, e, and f, respectively. The probability corresponding to d is 14%, the probability corresponding to e is 16%, and the probability corresponding to f is 25%. Then f is the target recognition result of the audio to be recognized.
[0093] In actual applications, after the second recognition result corresponding to the maximum value among all first probabilities is determined as the target recognition result of the audio to be recognized, there may be remaining audio to be recognized behind the audio to be recognized. In order to improve the accuracy of speech recognition of the remaining audio to be recognized, if there is a proper noun in the target recognition result of the audio to be recognized, it is necessary to update other proper nouns that are associated with the proper noun of the target recognition result in the audio to be recognized in the preset hot word library to improve the accuracy of speech recognition of the remaining audio to be recognized. This can be achieved in, but is not limited to, the following manner: matching the target recognition result with all proper nouns in the proper noun library respectively; if the matching result is that the target recognition result is the same as the target proper noun in the proper noun library, searching the proper noun library for at least one third proper noun that is associated with the target proper noun; updating the at least one third proper noun in the preset hot word set, so that speech recognition is performed on the remaining audio to be recognized in the continuous audio based on the updated preset hot word set to obtain the recognition result corresponding to the remaining audio to be recognized.
[0094] To facilitate understanding, an example is provided. Suppose there is a proper noun g in the target recognition result, and there are proper nouns g, h, i, j, k, l, m, n, and o in the proper noun library. There is an association between g and m and n. Since the proper noun g in the target recognition result has the same proper noun in the proper noun library, namely g, m and n that are associated with g will be updated in the preset hot word set.
[0095] In actual application, after the at least one third proper noun is updated in the preset hot word set, since the second proper noun and the third proper noun are only associated with the recognition result of a continuous audio segment, in order to avoid the second proper noun and the third proper noun from affecting the recognition results of other continuous audio segments, when the speech recognition of a continuous audio segment is completed, the second proper noun and the third proper noun updated in the preset hot word set during the speech recognition process of this continuous audio segment need to be deleted. This can be achieved in, but is not limited to, the following manner: determining whether the audio to be recognized is the last audio in the continuous audio; if it is determined that the audio to be recognized is the last audio in the continuous audio, then deleting the at least one second proper noun and the at least one third proper noun in the preset hot word set.
[0096] To facilitate understanding, an example is provided. Suppose a continuous audio segment is divided into three frames, namely a, b, and c. First, speech recognition is performed on a, then on b, and finally on c. After speech recognition is performed on c, the proper nouns updated in the preset hot word library during the recognition process of a, b, and c need to be deleted.
[0097] As a refinement of step 102, when performing the search for at least one second proper noun associated with the first proper noun in the pre-established proper noun library, since the number of second proper nouns associated with the first proper noun may be large, it is impossible to update all second proper nouns in the preset hot word set. Only second proper nouns with a high degree of association with the first proper noun need to be updated in the preset hot word set. This can be achieved in, but not limited to, the following manner: obtaining the association score between a fourth proper noun in the proper noun library and the first proper noun; the fourth proper noun is a proper noun associated with the first proper noun; if the association score is greater than a preset threshold, determining the fourth proper noun corresponding to the association score as the second proper noun; wherein the preset threshold can be any value, and the embodiment of the present disclosure does not limit the specific value of the preset threshold.
[0098] Related to the above embodiment, the proper noun library needs to be constructed in advance to ensure that the association relationship between proper nouns is correct. Regarding the construction of the proper noun library, the following method can be adopted but is not limited to: in response to a received instruction to establish an association relationship between at least two proper nouns, the at least two proper nouns are obtained; based on the association relationship instruction, the association relationship between the at least two proper nouns is bound to obtain the association relationship between the at least two proper nouns to complete the construction of the proper noun library.
[0099] In one possible implementation of the present disclosure, in order to better understand the entire process of speech recognition, an example is provided. A microphone receives the user's speaking audio, that is, a continuous audio segment, and streams the continuous audio segment to a speech recognition model. The speech recognition model includes an acoustic model, that is, a language model. The speech recognition model extracts acoustic features from the received continuous audio segment and sends the acoustic features to the acoustic model to obtain the acoustic probability of the continuous audio segment. One frame of audio corresponds to a set of acoustic probabilities, and a set of acoustic probabilities includes the probabilities of all words corresponding to the current frame. The acoustic probability is beam searched on the language model represented by a weighted finite state converter to obtain the language probability. The first probability is the sum of the acoustic probability and the language probability. The recognition result corresponding to the maximum value of the first probability is determined as the final recognition result of the frame of audio.
[0100] The speech recognition model initializes a frame of audio, and during initialization, merges preset hot words and user-defined hot words to form a complete hot word set; the complete hot word set is segmented, and a hot word matcher is constructed using the segmented hot word set; when the acoustic probability is beam searched through a language model represented by a weighted finite state converter, the hot word matcher is shallowly fused with the language model of the weighted finite state converter. If the recognition result of a frame of audio can match the hot word, a second probability is assigned to the recognition result, and the first probability and the second probability are added to obtain an updated first probability. The recognition result corresponding to the maximum value of the first probability is determined as the final recognition result of the frame of audio.
[0101] In the beam search process, the first maximum probability value of the recognized audio adjacent to the frame audio is obtained, which corresponds to the result; the recognition result is extracted by the proper noun extraction model to extract the proper nouns in the recognition result; according to each proper noun extracted by the proper noun extraction model, the top 20 proper nouns with the highest correlation are searched in the proper noun library; all the proper nouns searched in the proper noun library are segmented, and the original state is not changed in the original hot word set and the representation of this part of the proper nouns is added; then continue to recognize an unrecognized frame of audio, and at this time, the recognition of the subsequent frame of audio is matched with the hot word set representation of the proper noun associated with the proper noun of the currently recognized frame of audio; after the recognition of a piece of audio is completed, the recognition with the maximum first probability is obtained. The recognition result is used as the final speech recognition result, and the proper nouns added to the hot word set are deleted; for example, the user's complete request is "Introduce Sima Yi (yì) in the Rebellion of the Eight Princes", and "Yì" is an uncommon character. The usual speech recognition result is "Introduce Sima Yi in the Rebellion of the Eight Princes". When the user says "Introduce the Rebellion of the Eight Princes", the recognition result of "Introduce the Rebellion of the Eight Princes" is extracted into "The Rebellion of the Eight Princes" through the proper noun extraction model, and then the proper noun library is used to find relevant proper nouns, such as: "Sima Liang", "Sima Wei", "Sima Yi", etc. These relevant proper nouns are added to the current recognition hot word set, which helps the subsequent speech recognition to correct "Sima Yi" to "Sima Yi", and finally obtain the correct recognition result of "Introduce Sima Yi in the Rebellion of the Eight Princes".
[0102] In summary, the embodiments of the present disclosure can achieve the following effects:
[0103] The embodiment of the present disclosure actively updates the preset hot word set by obtaining proper nouns in the recognition results of the recognized audio, finding at least one second proper noun that is associated with the first proper noun, and updating the at least one second proper noun in the preset hot word set. The audio to be recognized and the recognized audio are adjacent components of a continuous audio segment, and the recognized audio and the audio to be recognized are associated, so the at least one second proper noun and the recognition result of the audio to be recognized are also associated, achieving the effect of hot word enhancement on the recognition result of the audio to be recognized by the updated preset hot word set, thereby improving the accuracy of speech recognition.
[0104] Corresponding to the above-mentioned speech recognition method, the present invention also provides a speech recognition device. Since the device embodiment of the present invention corresponds to the above-mentioned method embodiment, details not disclosed in the device embodiment can be referred to the above-mentioned method embodiment and will not be repeated in this invention.
[0105] Figure 3 This is a structural diagram of a speech recognition device provided by an embodiment of the present disclosure, wherein the device is applied to a processor, such as Figure 3 Shown, including:
[0106] An acquisition unit 31 acquires audio to be recognized and a first proper noun in a first recognition result of the recognized audio; the audio to be recognized and the recognized audio are adjacent components of a continuous audio segment;
[0107] a search unit 32 configured to search a pre-established proper noun library for at least one second proper noun associated with the first proper noun;
[0108] An updating unit 33, configured to update the at least one second proper noun found in a preset hot word set;
[0109] The recognition unit 34 is configured to perform speech recognition on the audio to be recognized based on the updated preset hot word set to obtain a target recognition result.
[0110] The speech recognition device provided by the present disclosure obtains an audio to be recognized and obtains a first proper noun in a first recognition result of the recognized audio; the audio to be recognized and the recognized audio are adjacent components of a continuous audio segment; at least one second proper noun associated with the first proper noun is searched in a pre-established proper noun library; the at least one second proper noun found is updated in a preset hot word set, so that the audio to be recognized is recognized based on the updated preset hot word set to perform speech recognition on the audio to be recognized and obtain a target recognition result. Compared with the related art, the embodiment of the present disclosure actively updates the preset hot word set by obtaining the proper noun in the recognition result of the recognized audio, finding at least one second proper noun associated with the first proper noun, and updating the at least one second proper noun in the preset hot word set. The audio to be recognized and the recognized audio are adjacent components of a continuous audio segment, and the recognized audio and the audio to be recognized are associated, so the recognition result of the at least one second proper noun and the audio to be recognized are also associated, so that the updated preset hot word set has a hot word enhancement effect on the recognition result of the audio to be recognized, thereby improving the accuracy of speech recognition.
[0111] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 4 As shown, the identification unit 34 includes:
[0112] The recognition module 341 is configured to perform speech recognition on the audio to be recognized, and obtain at least one second recognition result and a first probability corresponding to each of the second recognition results;
[0113] A first determining module 342 is configured to determine a similarity between each hot word in the updated preset hot word set and the at least one second recognition result;
[0114] An allocating module 343 is configured to allocate a second probability to a target second recognition result when there is a target second recognition result whose similarity exceeds a preset similarity threshold in the at least one second recognition result;
[0115] a calculation module 344 configured to sum the first probability and the second probability corresponding to the second target recognition result to obtain an updated first probability of the second target recognition result;
[0116] The second determination module 345 is configured to determine a second recognition result corresponding to a maximum value among all first probabilities as the target recognition result of the audio to be recognized.
[0117] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 4 As shown, the device also includes:
[0118] a matching unit 35 for, after determining the second recognition result corresponding to the maximum value among all the first probabilities as the target recognition result of the audio to be recognized, matching the target recognition result with all the proper nouns in the proper noun library;
[0119] The search unit 32 is further configured to, when a matching result indicates that the target recognition result is identical to a target proper noun in the proper noun library, search the proper noun library for at least one third proper noun associated with the target proper noun;
[0120] The updating unit 33 is further configured to update the at least one third proper noun in the preset hot word set, so as to perform speech recognition on the remaining audio to be recognized in the continuous audio based on the updated preset hot word set to obtain a recognition result corresponding to the remaining audio to be recognized.
[0121] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 4 As shown, the device also includes:
[0122] a determination unit 36 configured to determine whether the audio to be recognized is the last audio in the continuous audio after updating the at least one third proper noun in the preset hot word set;
[0123] The deleting unit 37 is configured to delete the at least one second proper noun and the at least one third proper noun in the preset hot word set when it is determined that the audio to be recognized is the last audio in the continuous audio.
[0124] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 4 As shown, the search unit 32 includes:
[0125] An acquisition module 321 is configured to respectively acquire a correlation score between a fourth proper noun and the first proper noun in the proper noun library; the fourth proper noun is a proper noun that has a correlation with the first proper noun;
[0126] The determining module 322 is configured to determine the fourth proper noun corresponding to the association relationship score as the second proper noun when the association relationship score is greater than a preset threshold.
[0127] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 4 As shown, the device also includes:
[0128] a second acquiring unit 38 configured to acquire the at least two proper nouns in response to a received instruction to establish an association relationship between the at least two proper nouns;
[0129] The binding unit 39 is configured to bind the association relationship between the at least two proper nouns based on the association relationship instruction to obtain the association relationship between the at least two proper nouns, so as to complete the establishment of the proper noun library.
[0130] It should be noted that the above explanation of the method embodiment is also applicable to the device of the embodiment of the present disclosure, and the principles are the same, which is no longer limited in the embodiment of the present disclosure.
[0131] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0132] Figure 5 A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0133] like Figure 5 As shown, the device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 402 or a computer program loaded from a storage unit 408 into a RAM (Random Access Memory) 403. Various programs and data required for the operation of the device 400 can also be stored in the RAM 403. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An I / O (Input / Output) interface 405 is also connected to the bus 404.
[0134] Various components in device 400 are connected to I / O interface 405, including an input unit 406, such as a keyboard, mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, optical disk, etc.; and a communication unit 409, such as a network card, modem, wireless communication transceiver, etc. Communication unit 409 allows device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0135] The computing unit 401 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as the speech recognition method. For example, in some embodiments, the speech recognition method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the method described above can be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to execute the aforementioned speech recognition method in any other appropriate manner (for example, by means of firmware).
[0136] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0137] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0138] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0139] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0140] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0141] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.
[0142] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.
[0143] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0144] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for speech recognition, characterized in that: include: Obtaining audio to be recognized, and obtaining a first proper noun in a first recognition result of the recognized audio; The audio to be recognized and the recognized audio are adjacent components of a continuous audio segment; Searching a pre-established proper noun database for at least one second proper noun associated with the first proper noun; The at least one second proper noun found is updated in the preset hot word set, so as to perform speech recognition on the audio to be recognized based on the updated preset hot word set to obtain a target recognition result.
2. The method according to claim 1, characterized in that The speech recognition of the audio to be recognized is performed based on the updated preset hot word set to obtain the target recognition result, which includes: Performing speech recognition on the audio to be recognized to obtain at least one second recognition result and a first probability corresponding to each of the second recognition results; Determining the similarity between each hot word in the updated preset hot word set and the at least one second recognition result; If there is a target second recognition result whose similarity exceeds a preset similarity threshold among the at least one second recognition result, assigning a second probability to the target second recognition result; Adding the first probability and the second probability corresponding to the second recognition result of the target to obtain an updated first probability of the second recognition result of the target; The second recognition result corresponding to the maximum value among all the first probabilities is determined as the target recognition result of the audio to be recognized.
3. The method according to claim 2, characterized in that After determining the second recognition result corresponding to the maximum value among all the first probabilities as the target recognition result of the audio to be recognized, the method further includes: Matching the target recognition results with all proper nouns in the proper noun library respectively; If the matching result is that the target recognition result is identical to the target proper noun in the proper noun library, searching the proper noun library for at least one third proper noun associated with the target proper noun; The at least one third proper noun is updated in the preset hot word set, so as to perform speech recognition on the remaining audio to be recognized in the continuous audio based on the updated preset hot word set to obtain a recognition result corresponding to the remaining audio to be recognized.
4. The method according to claim 3, characterized in that After updating the at least one third proper noun in the preset hot word set, the method further includes: Determining whether the audio to be recognized is the last audio in the continuous audio; If it is determined that the audio to be recognized is the last audio in the continuous audio, the at least one second proper noun and the at least one third proper noun in the preset hot word set are deleted.
5. The method according to claim 1, wherein The step of searching a pre-established proper noun library for at least one second proper noun associated with the first proper noun includes: respectively obtaining a correlation score between a fourth proper noun in the proper noun library and the first proper noun; the fourth proper noun being a proper noun having a correlation relationship with the first proper noun; If the association relationship score is greater than a preset threshold, the fourth proper noun corresponding to the association relationship score is determined as the second proper noun.
6. The method according to claim 1, characterized in that The method for constructing the proper noun library includes: In response to the received instruction to establish an association relationship between at least two proper nouns, acquiring the at least two proper nouns; Based on the association relationship instruction, the association relationship between the at least two proper nouns is bound to obtain the association relationship between the at least two proper nouns, so as to complete the establishment of the proper noun library.
7. A speech recognition device, characterized in that: include: an acquiring unit, acquiring the audio to be recognized and acquiring the first proper noun in the first recognition result of the recognized audio; The audio to be recognized and the recognized audio are adjacent components of a continuous audio segment; a search unit, configured to search a pre-established proper noun library for at least one second proper noun associated with the first proper noun; an updating unit, configured to update the at least one second proper noun found in the preset hot word set; The recognition unit is used to perform speech recognition on the audio to be recognized based on the updated preset hot word set to obtain a target recognition result.
8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 6.