Speech recognition method and related device
By identifying the content, intent, and function information of speech in speech recognition, and combining API calls and keyword table matching, the problem of low speech recognition accuracy is solved, achieving efficient recognition of new and specialized vocabulary, and reducing training costs and latency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2025-01-24
- Publication Date
- 2026-07-24
AI Technical Summary
Existing speech recognition technologies have poor accuracy, especially when faced with new words and specialized vocabulary, and the costs of frequent model training and data expansion are high.
By recognizing speech to determine primary content information, intent information, and associated function information, and by using API calls and keyword table matching, combined with fuzzy search and direct concatenation to generate recognition results, the frequency of model training is reduced and recognition accuracy is improved.
It improves the accuracy of speech recognition, especially for new and specialized words, reduces the consumption of computing and human resources, and enhances the user experience.
Smart Images

Figure CN122454976A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a speech recognition method and related equipment. Background Technology
[0002] Speech recognition technology typically converts input speech into text, which can then be further processed based on the recognized text.
[0003] However, the inventors of this disclosure have found that the accuracy of speech recognition in related technologies is unsatisfactory. Summary of the Invention
[0004] This disclosure proposes a speech recognition method and related equipment to solve or partially solve the above-mentioned problems.
[0005] In a first aspect, this disclosure provides a speech recognition method, comprising:
[0006] Receives voice input from the user;
[0007] The speech is identified to determine first content information, intent information, and function information associated with the intent information;
[0008] In response to the mismatch between the first content information obtained from the recognition and the speech, the function corresponding to the function information is invoked to determine the second content information based on the first content information through the application programming interface (API) corresponding to the function;
[0009] Based on the intent information and the second content information, the recognition result of the speech is generated.
[0010] A second aspect of this disclosure provides a voice recognition device, comprising:
[0011] The receiving module is configured to receive user-input voice.
[0012] The recognition module is configured to: recognize the speech to determine first content information, intent information, and function information associated with the intent information corresponding to the speech;
[0013] The calling module is configured to: in response to a mismatch between the first content information obtained from recognition and the speech, call the function corresponding to the function information to determine the second content information based on the first content information through the application programming interface (API) corresponding to the function;
[0014] The generation module is configured to generate the recognition result of the speech based on the intent information and the second content information.
[0015] A third aspect of this disclosure provides a computer device including one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, the programs including instructions for performing the method according to the first aspect.
[0016] A fourth aspect of this disclosure provides a non-volatile computer-readable storage medium containing a computer program that, when executed by one or more processors, causes the processors to perform the method described in the first aspect.
[0017] A fifth aspect of this disclosure provides a computer program product, including a computer program, characterized in that, when executed by a processor, the computer program implements the steps of the method as described in the first aspect.
[0018] The speech recognition method and related device provided in this disclosure identify the speech to determine the first content information, intent information, and function information associated with the intent information. When the first content information obtained by identification does not match the speech, the method calls the function corresponding to the function information to determine the second content information based on the first content information through the application programming interface corresponding to the function. Then, the method generates the speech recognition result based on the intent information and the second content information, thereby obtaining a more accurate recognition result. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A schematic diagram of an exemplary system provided by an embodiment of this disclosure is shown.
[0021] Figure 2 A flowchart illustrating an exemplary method provided by an embodiment of this disclosure is shown.
[0022] Figure 3 A flowchart illustrating another exemplary method provided by an embodiment of this disclosure is shown.
[0023] Figure 4 A schematic diagram of an exemplary apparatus provided by an embodiment of the present disclosure is shown.
[0024] Figure 5A schematic diagram of the hardware structure of an exemplary computer device provided in an embodiment of this disclosure is shown. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0026] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0027] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.
[0028] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.
[0029] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0030] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0031] Figure 1A schematic diagram of an exemplary system 100 provided in an embodiment of this disclosure is shown.
[0032] like Figure 1 As shown, system 100 may include terminal device 102, server 106, and database server 108. A medium (e.g., a network) may be provided between terminal device 102 and server 106 and database server 108 to provide a communication link. This network may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0033] The terminal device 102 may have various applications (APPs) or software installed, such as voice recognition applications, audio applications, collaborative office applications, image processing applications, video conferencing applications, reading applications, video applications, social applications, payment applications, web browsers, and instant messaging tools. In some embodiments, these applications can all be used for voice recognition.
[0034] The terminal device 102 here can be either hardware or software. When the terminal device 102 is hardware, it can be various electronic devices with a display screen, including but not limited to smartphones, tablets, e-book readers, MP3 players, laptops, and desktop computers (PCs). When the terminal device 102 is software, it can be installed in the electronic devices listed above. It can be implemented as multiple software programs or software modules (e.g., to provide distributed services) or as a single software program or software module. No specific limitations are made here.
[0035] Server 106 can be a server providing various services, such as a backend server supporting various applications displayed on terminal device 102. Database server 108 can also be a database server providing various services. It is understood that if server 106 can implement the relevant functions of database server 108, database server 108 may not need to be set up in system 100.
[0036] The server 106 and database server 108 here can be either hardware or software. When they are hardware, they can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When they are software, they can be implemented as multiple software programs or software modules (e.g., used to provide distributed services), or as a single software program or software module. No specific limitations are made here.
[0037] It should be noted that the document editing method provided in this embodiment can be executed by the server 106. It should be understood that... Figure 1The number of terminal devices, users, servers, and database servers shown is merely illustrative. Depending on implementation needs, there can be any number of terminal devices, users, servers, and database servers.
[0038] In one embodiment, a voice recognition application or software may be installed on the terminal device 102, and the user 104 may use the application or software installed on the terminal device 102 to perform voice recognition.
[0039] In a more specific scenario, user 104 can utilize an app with an artificial intelligence (AI) model in terminal device 102 for speech recognition. For example, by inputting a speech signal, user 104 can output a command, causing the AI model to perform a corresponding operation based on that command. Figure 1 As shown, speech 1022 is uploaded to server 106. First, it undergoes speech recognition. Server 106 can then call a speech recognition model to convert speech 1022 into text, and then perform corresponding operations based on the text information. However, as mentioned earlier, the recognition accuracy of speech recognition models in related technologies is unsatisfactory.
[0040] In view of this, embodiments of the present disclosure provide a speech recognition method to solve or partially solve the above-mentioned problems.
[0041] Figure 2 A flowchart illustrating an exemplary method 200 provided in an embodiment of this disclosure is shown. Method 200 can be used for speech recognition. Optionally, method 200 can be... Figure 1 The server-side implementation 106 can also be handled by... Figure 1 The system 100 is implemented.
[0042] like Figure 2 As shown, for example, user 104 can use terminal device 102 to input a voice message 1022 for voice recognition.
[0043] Optionally, the voice 1022 may be a voice recording taken by user 104 through the microphone of terminal device 102, or a voice recording in the environment, or a voice recording selected by user 104 from the local terminal device 102, etc.
[0044] Terminal device 102 can send voice 1022 to server 106 for voice recognition.
[0045] After receiving voice 1022, server 106 can recognize voice 1022. Optionally, the recognition method can be based on artificial intelligence model recognition, voice template matching recognition, or dynamic time warping (DTW) voice recognition method.
[0046] As an optional embodiment, such as Figure 2 As shown, the server 106 can call the speech recognition model 202 to recognize the speech 1022.
[0047] Optionally, the speech recognition model 202 can be a deep learning-based speech recognition model, such as a convolutional neural network (CNN), a recurrent neural network (RNN), or its variant, a long short-term memory network (LSTM); the speech recognition model 202 can also be a speech recognition model based on a hidden Markov model (HMM); the speech recognition model 202 can also be an end-to-end speech recognition model, such as a Transformer or a Conformer.
[0048] It is understandable that for the speech recognition model 202 to achieve accurate recognition, it needs to be trained using various corpora. However, the inventors of this disclosure have discovered that with the development of technology and the barrier-free communication between people brought about by internet technology, new vocabulary is constantly emerging. Especially during voice interaction, some trending words and professional knowledge used by user 104 are not covered in the training corpus of the speech recognition model 202. This makes it more difficult for the speech recognition model 202 to recognize this content not covered by the training corpus, thus affecting the recognition accuracy.
[0049] In related technologies, the recognition performance in this part is improved by adding more and more timely training data or by increasing domain-specific training data. However, the inventors of this disclosure have found that using this method results in high costs for data collection, labeling, and expansion, and the computational and time costs of retraining a large model are enormous. Therefore, constantly expanding training data and constantly fine-tuning the model is not a long-term solution.
[0050] In view of this, such as Figure 2 As shown, the speech recognition model 202 of this embodiment can output first content information 2042, intent information 2044, and function information 2046 associated with the intent information 2044 when recognizing speech 1022. The function information 2046 may include a function that can be used to call the corresponding application programming interface (API) 208 to determine second content information 210 based on the first content information 2042.
[0051] For example, such as Figure 2As shown, the output 204 of the speech recognition model 202 may include intent information 2044 and function information 2046 associated with the intent information 2044, and may also include parameter information 2048 for passing to API 208. The parameter information 2048 further includes first content information 2042, so that after the parameter information 2048 is passed to API 208, API 208 can generate second content information 210 based on the first content information 2042.
[0052] In some embodiments, in order for the speech recognition model 202 to generate content in the format of output 204 based on speech 1022, the speech recognition model 202 can be trained using training data similar to that format.
[0053] For example, some query requests in speech format can be generated in advance, and the content information and intent information in these speech requests can be determined. Then, the corresponding function information can be associated with the intent information to generate a series of training data.
[0054] For example, suppose the speech content is "Play AAAA", where the content information is the song "AAAA" and the intent information is "play". Since the intent information is "play", the function information associated with the intent information can be determined as the function used to call the music playback-related API. Then, this information is used as a label for the speech, and the labeled speech is the training data. It can be understood that this only uses playing music as an example to set up training data. Depending on the intent, there can be training data from various domains, such as novels, live broadcasts, etc. The function information corresponding to the intent information of these different domain training data can be the API of that domain, so that the second content information can be generated based on the first content information of that domain.
[0055] After obtaining the training data, the speech recognition model 202 can be trained using this training data. Alternatively, the speech recognition model can be further trained using the aforementioned training data based on the speech recognition model that has been pre-trained using a general corpus to obtain model 202. This allows the trained model 202 to output the first content information, intent information, and function information associated with the intent information based on the speech.
[0056] In some embodiments, if the speech recognition model 202 is a large language model, in order to make the model 202 output in the format of output 204, a prompt can be set for the model 202. The prompt specifies the output format of the model, so that the model 202 can generate output 204 after recognizing speech 1022, ensuring the standardization of the output format, and thus better realizing API calls.
[0057] Back Figure 2 After receiving output 204, the server 106 can determine whether the first content information 2042 obtained by recognition matches the voice 1022, so as to determine whether the first content information 2042 is correctly recognized.
[0058] In some embodiments, the matching of the first content information 2042 with the speech 1022 can be determined by calculating the confidence level of the first content information 2042. For example, when the confidence level is lower than the confidence level threshold, the first content information 2042 is considered not to match the speech 1022, and when the confidence level is equal to or higher than the confidence level threshold, the first content information 2042 is considered to match the speech 1022.
[0059] However, the inventors of this disclosure have discovered that using the confidence level calculation method to determine whether the first content information 2042 matches the voice 1022 may result in inaccurate calculation results due to the confidence level calculation method, thus failing to accurately determine whether the first content information 2042 matches the voice 1022.
[0060] In view of this, in some embodiments, keywords associated with various intent information can be pre-stored in the database server 108. Then, after obtaining the intent information 2044, multiple keywords associated with the intent information 2044 can be found in the database server 108, and it can be determined whether the first content information 2042 can match these keywords. If the first content information 2042 does not match any of the multiple keywords associated with the intent information 2044, it can be determined that the identified first content information 2042 does not match the voice 1022. If the first content information 2042 matches the target keyword among the multiple keywords associated with the intent information 2044, it can be determined that the identified first content information 2042 matches the voice 1022.
[0061] In this way, it is possible to quickly determine whether the first content information 2042 matches the voice 1022, and the matching result is relatively accurate. It is only necessary to associate the corresponding keywords with specific intent information. Furthermore, it is only necessary to update the keywords as needed in the future, without the need for a time-consuming and labor-intensive model training process.
[0062] Optionally, multiple tables can be pre-stored in the database server 108. Each table includes multiple keywords associated with specific intent information. For example, a table associated with the intent to play music may include keywords such as song title, singer, lyricist, and composer; a table associated with the intent to read a novel may include keywords such as book title and author, and so on. Then, after obtaining the intent information 2044, the table corresponding to the intent information 2044 can be found, and the first content information 2042 can be determined by looking up the table to see if it matches the keywords in the table. If the first content information 2042 does not match any of the keywords in the table, it can be determined that the first content information 2042 does not match the speech 1022. If the first content information 2042 matches the target keyword among the multiple keywords in the table, it can be determined that the first content information 2042 matches the speech 1022.
[0063] In this way, it is possible to quickly determine whether the first content information 2042 matches the voice 1022 by looking up the table, and the matching result is relatively accurate. It is only necessary to configure the keywords in the table corresponding to the specific intent information. Furthermore, updating the table only requires updating the keywords, without the need for a time-consuming and labor-intensive model training process.
[0064] In some embodiments, the multiple keywords include multiple words and / or phrases in the domain associated with the intent information that have been searched more than a preset number of times (e.g., 1000 times, 10000 times, etc.) within a preset time period (e.g., within the last six months, the last three months, the last month, the last week, etc.). In other words, the keywords in the table can be hot words, i.e., words or phrases with high search frequency. For example, popular singers and popular songs in the music playback field. If the first content information 2042 cannot match these hot words, it indicates that the first content information 2042 may belong to obscure vocabulary, and the recognition accuracy of model 202 may be low. The corresponding API can be called to introduce more knowledge to make the recognition of some obscure vocabulary or proper nouns more accurate.
[0065] Furthermore, as an optional embodiment, such as Figure 2As shown, if the first content information 2042 obtained by recognition does not match the speech 1022, the function corresponding to the function information 2046 can be called to determine the second content information 210 based on the first content information 2042 through the API 208 corresponding to the function. Then, based on the intent information 2044 and the second content information 210, the recognition result 212 of the speech 1022 is generated. In this way, when the first content information 2042 obtained by recognition does not match the speech 1022, more knowledge in this field is introduced by calling the API 208 corresponding to the intent information 2044, making the recognition of some obscure words or proper nouns more accurate.
[0066] In some embodiments, such as Figure 2 As shown, when calling API 208, the first content information 2042 can be passed to API 208 as parameter 206, so that API 208 can return the second content information 210 based on parameter 206. Optionally, parameter information 2048 may also include pinyin information 2050 corresponding to the first content information 2042, so that API 208 can return the second content information 210 based on the first content information 2042 and pinyin information 2050 in parameter 206, thereby improving the accuracy of the second content information 210.
[0067] In some embodiments, API 208 can be used to perform a fuzzy search based on the first content information 2042 and / or pinyin information 2050 to determine the second content information 210, thereby correcting the first content information 2042 that has been incorrectly identified.
[0068] Fuzzy search is a search function that allows users to search for relevant results based on the similarity of keywords when they enter them. Unlike traditional exact match search, fuzzy search can handle typos, spelling variations, and partial matches, thereby enhancing user experience and satisfaction. This technology is especially important when dealing with large amounts of data and information, as it can help users quickly find what they need, even if they don't precisely remember the correct spelling or full name of the search term.
[0069] Fuzzy search functionality can include:
[0070] 1. Handling spelling errors: Even if the keywords entered by the user are misspelled, fuzzy search can still find the correct results.
[0071] 2. Partial Match: The system can match complete content even if only part of the keywords entered by the user are matched. For example, entering "aple" can match "apple" and "pineapple".
[0072] 3. Synonym and near - synonym matching: Fuzzy search can identify synonyms and near - synonyms, providing more comprehensive search results. For example, when entering "happy", related content such as "joyful" can be matched.
[0073] 4. Dynamic thesaurus and word segmentation: When processing Chinese data, fuzzy search can achieve more flexible search through pinyin word segmentation and fuzzy matching. For example, when searching for "stir - fried pakchoi", the pinyin "qingchaoxiaobaicai" can be accepted.
[0074] In some embodiments, the fuzzy search can obtain multiple words and / or phrases similar to the first content information 2042. To avoid the infinite expansion of near - synonyms, the parameter information 2048 can also include quantity limit information 2052, which is added to the parameter 206 to limit the number of near - synonyms (candidate content information) returned by the API 208. After the server 106 obtains multiple candidate content information, it can select the information that best matches the speech 1022 by scoring the multiple candidate content information as the second content information 210.
[0075] After obtaining the second content information 210, in some embodiments, the intent information 2044 and the second content information 210 can be input into the speech recognition model 202 to guide the model 202 to output the recognition result 212 of the speech 1022. At this time, the second content information 210 is relatively accurate recognition information, and the re - output of the model 202 will obtain a more accurate recognition result 212.
[0076] However, the inventors of the present disclosure found that using the method of calling the model 202 again requires consuming computing resources again, which may cause a reaction delay of the system 100. Therefore, in some embodiments, the server 106 can directly splice the intent information 2044 and the second content information 210 to generate the recognition result 212 of the speech. At this time, the second content information 210 is relatively accurate recognition information, and a more accurate recognition result 212 can also be obtained by directly splicing the intent information 2044 and the second content information 210 without calling the model 202 again.
[0077] As mentioned above, the model 202 has a high recognition accuracy for some content. Therefore, as another optional embodiment, if the identified first content information 2042 matches the speech 1022, the recognition result 212 of the speech 1022 can be generated based on the intent information 2044 and the first content information 2042, and this recognition result 212 can also be relatively accurate.
[0078] As a more specific example, such as Figure 2 As shown, assuming the content of speech 1022 is "Play AAAA for me", speech recognition model 202 recognizes speech 1022 as "Play BBBB for me" and outputs 204. Server 106 uses "BBBB" to match in the table corresponding to the music playback intent. It finds no matching keyword, indicating that the first content information 2042 recognized by model 202 does not match speech 1022. Then, API 208 corresponding to function information 2046 can be called, and parameters 206 (e.g., "args:BBBB / xxxxxx") (xxxxxx is the pronunciation of "AAAA") can be generated based on parameter information 2048 in output 204 to be passed to API 208. API 208 can perform fuzzy search based on parameter 206 to obtain second content information 210 (e.g., "AAAA"). Finally, server 106 can concatenate intent information 2044 and second content information 210 to form recognition result 212 (e.g., "Play AAAA for me") for output. It is understandable that if model 202 recognizes the speech 1022 as "Play AAAA for me", then server 106 uses "AAAA" to match in the table corresponding to the music playback intent and can find the keyword "AAAA". At this point, there is no need to call API 208 again, but the intent information 2044 and the first content information 2042 can be directly concatenated to form the recognition result 212 for output. In this way, while ensuring the accuracy of the recognition result 212, there is no need to constantly expand the training data or constantly fine-tune the model, saving computing and human resources.
[0079] The inventors of this disclosure have achieved a significant improvement in speech recognition accuracy by applying the above method to speech recognition technology, in the same test set, compared with the speech recognition accuracy using related technologies.
[0080] In some embodiments, after obtaining the recognition result 212, the server 106 can also return the recognition result 212 to the terminal device 102. After receiving the recognition result 212, the terminal device 102 can display it on the screen, and the user 104 can see the recognition result 212. In some embodiments, the terminal device 102 can also perform subsequent operations based on the recognition result 212, such as sending a request to the server 106 to play corresponding music. The server 106 returns the corresponding music to the terminal device 102 for playback, so that the user 104 can achieve automatic playback of the corresponding music by inputting voice 1022 into the terminal device 102.
[0081] As can be seen from the above embodiments, the speech recognition method provided in this disclosure can introduce more knowledge by calling the corresponding API for words that cannot be recognized during the recognition process, making the recognition of some proper nouns and rare words more accurate. This method can correct some domain recognition errors online, especially those involving proper nouns and words with strong time sensitivity, which may not be in the training data of large language models. This method can retrieve such data without training.
[0082] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.
[0083] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0084] This disclosure also provides a speech recognition method. Figure 3 A schematic diagram of another exemplary method 300 provided in an embodiment of this disclosure is shown. Figure 3 As shown, the method 300 may further include the following steps.
[0085] In step 302, the user's voice input is received (e.g., Figure 2 (Voice 1022).
[0086] In step 304, the speech is recognized to determine the first content information corresponding to the speech (e.g., Figure 2 Information 2042), intent information (e.g., Figure 2 Information 2044) and function information associated with the intent information (e.g., Figure 2 Information 2046).
[0087] In step 306, in response to the mismatch between the first content information obtained from the identification and the speech, the function corresponding to the function information is invoked to utilize the application programming interface (API) corresponding to the function (e.g., ...). Figure 2 API 208) determines the second content information based on the first content information (e.g., Figure 2 Information 210).
[0088] In step 308, based on the intent information and the second content information, a speech recognition result is generated (e.g., Figure 2 The identification result 212).
[0089] The speech recognition method provided in this disclosure identifies the speech to determine first content information, intent information, and function information associated with the intent information. When the first content information does not match the speech, the method calls the function corresponding to the function information to determine second content information based on the first content information through the application programming interface corresponding to the function. Then, the method generates the speech recognition result based on the intent information and the second content information, thereby obtaining a more accurate recognition result.
[0090] In some embodiments, the method 300 further includes:
[0091] Obtain multiple keywords associated with the intent information;
[0092] In response to the fact that the first content information obtained by identification does not match any of the plurality of keywords, it is determined that the first content information obtained by identification does not match the speech; or
[0093] In response to the first content information being identified matching a target keyword among the plurality of keywords, it is determined that the first content information being identified matches the speech.
[0094] In this way, it is possible to quickly determine whether the first content information matches the speech, and the matching result is relatively accurate. It is only necessary to configure the keywords in the table corresponding to the specific intent information. Furthermore, updating the table only requires updating the keywords, without the need for a time-consuming and labor-intensive model training process.
[0095] In some embodiments, the plurality of keywords includes multiple words and / or phrases in a domain associated with the intent information that have been searched more than a preset number of times within a preset time period. By adding hot words to the table as keywords, some frequently searched words can be identified first, ensuring processing efficiency. However, for words with lower frequency, the recognition accuracy is usually lower due to their obscurity. Obtaining secondary content information by calling an API can improve the recognition accuracy.
[0096] In some embodiments, recognizing the speech to determine first content information, intent information, and function information associated with the intent information includes: invoking a speech recognition model and outputting the first content information, intent information, and function information associated with the intent information based on the speech. Thus, using a speech recognition model for speech recognition can improve the efficiency and accuracy of initial recognition.
[0097] In some embodiments, generating a speech recognition result based on the intent information and the second content information includes: inputting the intent information and the second content information into the speech recognition model, and outputting the speech recognition result. In this case, the second content information is relatively accurate recognition information, and the model's subsequent output will yield a more accurate recognition result.
[0098] In some embodiments, determining the second content information based on the first content information through the application programming interface (API) corresponding to the function includes: using the first content information as a parameter of the API (e.g., Figure 2 The parameter 206) is passed to the API; the API receives the second content information returned by the API based on the parameter, so that the API can return the second content information corresponding to the first content information based on the received parameter.
[0099] In some embodiments, the API is used to perform a fuzzy search based on the first content information to determine the second content information. Unlike traditional exact match searches, fuzzy search can handle typos, spelling variations, and partial matches, thereby enhancing user experience and satisfaction. This technology is particularly important when dealing with large amounts of data and information, as it can help users quickly find the content they need, even if they do not precisely remember the correct spelling or full name of the search term.
[0100] In some embodiments, generating a speech recognition result based on the intent information and the second content information includes concatenating the intent information and the second content information to generate the speech recognition result. In this case, the second content information is relatively accurate recognition information; directly concatenating the intent information and the second content information can also yield a more accurate recognition result, without needing to call the speech recognition model again, thus reducing latency and improving response speed.
[0101] In some embodiments, the method 300 further includes: in response to the recognition of the first content information matching the speech, generating a speech recognition result based on the intent information and the first content information. In this case, the recognition accuracy of the first content information is relatively high, and the recognition result can also be relatively accurate.
[0102] In some embodiments, the intent information includes intents related to music playback, and the first content information includes at least one of song title, artist, lyricist, and composer. Thus, applying method 300 to a music playback service can improve recall and enhance user experience.
[0103] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.
[0104] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0105] This disclosure also provides a voice recognition device. Figure 4 A schematic diagram of an exemplary apparatus 400 provided in an embodiment of this disclosure is shown. For example... Figure 4 As shown, the device 400 can be used to implement method 200 or method 300, and may further include the following modules.
[0106] The receiving module 402 is configured to receive voice input from the user.
[0107] The recognition module 404 is configured to: recognize the speech to determine first content information, intent information and function information associated with the intent information corresponding to the speech;
[0108] The calling module 406 is configured to: in response to a mismatch between the first content information obtained from recognition and the speech, call the function corresponding to the function information to determine the second content information based on the first content information through the application programming interface (API) corresponding to the function;
[0109] The generation module 408 is configured to generate the recognition result of the speech based on the intent information and the second content information.
[0110] In some embodiments, the identification module 404 is configured to:
[0111] Obtain multiple keywords associated with the intent information;
[0112] In response to the fact that the first content information obtained by identification does not match any of the plurality of keywords, it is determined that the first content information obtained by identification does not match the speech; or
[0113] In response to the first content information being identified matching a target keyword among the plurality of keywords, it is determined that the first content information being identified matches the speech.
[0114] In some embodiments, the plurality of keywords include a plurality of words and / or phrases in a field associated with the intent information that have been searched more than a preset number of times within a preset time period.
[0115] In some embodiments, the recognition module 404 is configured to: invoke a speech recognition model and output first content information, intent information, and function information associated with the intent information based on the speech.
[0116] In some embodiments, the generation module 408 is configured to: input the intent information and the second content information into the speech recognition model, and output the speech recognition result.
[0117] In some embodiments, the calling module 406 is configured to:
[0118] The first content information is passed to the API as a parameter of the API;
[0119] Receive the second content information returned by the API based on the parameters.
[0120] In some embodiments, the API is used to perform a fuzzy search based on the first content information to determine the second content information.
[0121] In some embodiments, the generation module 408 is configured to concatenate the intent information and the second content information to generate the speech recognition result.
[0122] In some embodiments, the generation module 408 is configured to: generate a recognition result of the speech based on the intent information and the first content information in response to the recognition of the first content information matching the speech.
[0123] In some embodiments, the intent information includes an intent related to music playback, and the first content information includes at least one of song title, artist, lyricist, and composer.
[0124] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.
[0125] The apparatus of the above embodiments is used to implement the corresponding method 200 or method 300 in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0126] This disclosure also provides a computer device for implementing the above-described method 200 or method 300. Figure 5 A schematic diagram of the hardware structure of an exemplary computer device 500 provided in an embodiment of this disclosure is shown. The computer device 500 can be used to implement... Figure 1 Server 106 can also be used to implement Figure 1 Terminal device 102. In some scenarios, this computer device 500 can also be used to implement... Figure 1 Database server 108.
[0127] like Figure 5 As shown, the computer device 500 may include: a processor 502, a memory 504, a network module 506, a peripheral interface 508, and a bus 510. The processor 502, memory 504, network module 506, and peripheral interface 508 are interconnected within the computer device 500 via the bus 510.
[0128] Processor 502 may be a central processing unit (CPU), image processor, neural network processor (NPU), microcontroller (MCU), programmable logic device, digital signal processor (DSP), application-specific integrated circuit (ASIC), or one or more integrated circuits. Processor 502 can be used to perform functions related to the techniques described in this disclosure. In some embodiments, processor 502 may also include multiple processors integrated as a single logic component. For example, such as... Figure 5 As shown, processor 502 may include multiple processors 502a, 502b and 502c.
[0129] Memory 504 can be configured to store data (e.g., instructions, computer code, etc.). Figure 5As shown, the data stored in memory 504 may include program instructions (e.g., program instructions for implementing method 200 or method 300 of the embodiments of this disclosure) and data to be processed (e.g., the memory may store configuration files of other modules, etc.). Processor 502 may also access the program instructions and data stored in memory 504 and execute the program instructions to operate on the data to be processed. Memory 504 may include volatile storage devices or non-volatile storage devices. In some embodiments, memory 504 may include random access memory (RAM), read-only memory (ROM), optical disk, magnetic disk, hard disk, solid-state drive (SSD), flash memory, memory stick, etc.
[0130] Network interface 506 can be configured to provide communication with other external devices to computer device 500 via a network. This network can be any wired or wireless network capable of transmitting and receiving data. For example, the network can be a wired network, a local wireless network (e.g., Bluetooth, WiFi, Near Field Communication (NFC), etc.), a cellular network, the Internet, or a combination thereof. It is understood that the type of network is not limited to the specific examples described above.
[0131] The peripheral interface 508 can be configured to connect the computer device 500 to one or more peripheral devices to enable information input and output. For example, peripheral devices may include input devices such as keyboards, mice, touchpads, touch screens, microphones, and various sensors, as well as output devices such as displays, speakers, vibrators, and indicator lights.
[0132] Bus 510 can be configured to transfer information between various components of computer device 500 (e.g., processor 502, memory 504, network interface 506, and peripheral interface 508), such as internal buses (e.g., processor-memory bus), external buses (USB port, PCI-E bus), etc.
[0133] It should be noted that although the architecture of the computer device 500 described above only shows the processor 502, memory 504, network interface 506, peripheral interface 508, and bus 510, in specific implementations, the architecture of the computer device 500 may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the architecture of the computer device 500 described above may only include the components necessary for implementing the embodiments of this disclosure, and does not necessarily include all the components shown in the figures.
[0134] Based on the same inventive concept, corresponding to any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the method 200 or method 300 as described in any of the above embodiments.
[0135] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0136] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to perform the method 200 or method 300 as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0137] Based on the same inventive concept, corresponding to any of the above-described embodiments of method 200 or method 300, this disclosure also provides a computer program product, which includes a computer program. In some embodiments, the computer program is executable by one or more processors to cause the processors to perform the method 200 or method 300. Corresponding to the execution entity for each step in each embodiment of method 200 or method 300, the processor performing the corresponding step may belong to the corresponding execution entity.
[0138] In some embodiments, the computer program product includes program modules for resolving operational conflicts, the program modules being compiled into a stack-based virtual machine binary instruction set (e.g., wasm) and deployed in a terminal device and / or compiled into a static library and deployed in a server.
[0139] The computer program product of the above embodiments is used to cause the processor to execute the method 200 or method 300 as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0140] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.
[0141] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuitry) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0142] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0143] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A speech recognition method, comprising: Receives voice input from the user; The speech is identified to determine first content information, intent information, and function information associated with the intent information; In response to the mismatch between the first content information obtained from the recognition and the speech, the function corresponding to the function information is invoked to determine the second content information based on the first content information through the application programming interface (API) corresponding to the function; Based on the intent information and the second content information, the recognition result of the speech is generated.
2. The method of claim 1, further comprising: Obtain multiple keywords associated with the intent information; In response to the fact that the first content information obtained by identification does not match any of the multiple keywords, it is determined that the first content information obtained by identification does not match the speech; or In response to the first content information being identified matching a target keyword among the plurality of keywords, it is determined that the first content information being identified matches the speech.
3. The method as described in claim 2, wherein, The multiple keywords include multiple words and / or phrases in the field associated with the intent information that have been searched more than a preset number of times within a preset time period.
4. The method of claim 1, wherein, Recognizing the speech to determine first content information, intent information, and function information associated with the intent information, including: The speech recognition model is invoked, and based on the speech, the first content information, intent information, and function information associated with the intent information corresponding to the speech are output.
5. The method of claim 4, wherein, Based on the intent information and the second content information, a speech recognition result is generated, including: The intent information and the second content information are input into the speech recognition model, and the speech recognition result is output.
6. The method of claim 1, wherein, The second content information is determined based on the first content information through the application programming interface (API) corresponding to the function, including: The first content information is passed to the API as a parameter of the API; Receive the second content information returned by the API based on the parameters.
7. The method of claim 6, wherein, The API is used to perform a fuzzy search based on the first content information to determine the second content information.
8. The method of claim 1, wherein, Based on the intent information and the second content information, a speech recognition result is generated, including: The intent information and the second content information are concatenated to generate the speech recognition result.
9. The method of claim 1 or 2, further comprising: In response to the matching of the first content information obtained from the recognition with the speech, a recognition result of the speech is generated based on the intent information and the first content information.
10. The method of claim 1, wherein, The intent information includes intents related to music playback, and the first content information includes at least one of the song title, singer, lyricist, and composer.
11. A voice recognition device, comprising: The receiving module is configured to receive user-input voice. The recognition module is configured to: recognize the speech to determine first content information, intent information, and function information associated with the intent information corresponding to the speech; The calling module is configured to: in response to a mismatch between the first content information obtained from recognition and the speech, call the function corresponding to the function information to determine the second content information based on the first content information through the application programming interface (API) corresponding to the function; The generation module is configured to generate the recognition result of the speech based on the intent information and the second content information.
12. A computer device comprising one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, the programs comprising instructions for performing the method according to any one of claims 1-10.
13. A non-volatile computer-readable storage medium comprising a computer program, which, when executed by one or more processors, causes the processors to perform the method of any one of claims 1-10.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-10.