Speech recognition method, server, speech recognition system and storage medium
By receiving and generating source data of the user data model when the vehicle's business scenario changes, using user characteristic hot words to improve the recognition accuracy of voice requests, solving the problem of voice request recognition in different scenarios, and achieving rapid updates and efficient recognition.
Patent Information
- Application Number
- CN202111509403.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2041-12-10
AI Technical Summary
In the prior art, voice requests issued by users are difficult to be accurately recognized in different scenarios, resulting in recognition errors and user troubles.
By receiving and generating or updating the source data of the user data model when the vehicle's business scenario changes, the user's characteristic hot words improve the recognition accuracy of voice requests, including synchronizing the hot words and interactive hot words when the mobile terminal and the vehicle establish a connection, and loading the corresponding user data model according to the vehicle and user identification.
The accuracy of voice request recognition in different scenarios is improved, and the rapid update and recognition efficiency of user data models are achieved.
Smart Images

Figure CN114360516B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition technology, and in particular to a speech recognition method, a server, a speech recognition system and a storage medium. Background Art
[0002] In related technologies, voice recognition is performed on the user's voice to determine the user's intention, and corresponding operations are performed according to the user's intention, so that the user can perform corresponding control functions through voice interaction and improve recognition accuracy through personalized services. Summary of the Invention
[0003] The present invention provides a speech recognition method, a server, a speech recognition system and a storage medium.
[0004] The speech recognition method of an embodiment of the present invention includes: when the business scenario of a vehicle changes, receiving business scenario information uploaded by the vehicle; generating source data of a user data model based on the business scenario information, the business scenario information including user feature hot words; generating or updating the user data model based on the source data, and the user data model can be used to recognize the voice request sent by the vehicle.
[0005] In the above-mentioned speech recognition method, the user characteristic hot words of the business scenario corresponding to the vehicle are determined through the business scenario information, and the user data model is finally obtained based on the user characteristic hot words. The user data model is used to identify voice requests, which is conducive to improving the accuracy of identifying voice requests in different scenarios.
[0006] The user characteristic hot words are stored in a mobile terminal that is communicatively connected to the vehicle. The speech recognition method includes: determining a change in the business scenario when the mobile terminal and the vehicle are connected. This facilitates rapid updating of the user data model.
[0007] The user feature hotwords include address book hotwords, the business scenario includes a communication scenario, and the vehicle is configured to obtain the address book hotwords sent by the mobile terminal. When the vehicle's business scenario changes, the vehicle receives business scenario information uploaded by the vehicle, including: if the current business scenario is the communication scenario, receiving the address book hotwords sent by the vehicle; and generating source data for the user data model based on the business scenario information, including: synchronizing the address book hotwords with the source data. This facilitates updating information in the address book to the user data model.
[0008] The user characteristic hot words are generated in the vehicle, and the speech recognition method includes: determining the business scenario change when the vehicle generates the user characteristic hot words. In this way, the user data model can be quickly updated.
[0009] The business scenario includes an interaction scenario, and the vehicle is configured to determine the first generated interaction hotword as the user feature hotword. When the vehicle's business scenario changes, the vehicle receives business scenario information uploaded by the vehicle, including: if the current business scenario is the interaction scenario, receiving the interaction hotword sent by the vehicle; and generating source data for the user data model based on the business scenario information, including: synchronizing the interaction hotword with the source data. This facilitates updating relevant information about user-vehicle interactions into the user data model.
[0010] The speech recognition method includes: loading a corresponding user data model based on a vehicle identifier and a user identifier; receiving a speech request sent by the vehicle while the vehicle is traveling; and performing speech recognition on the speech request based on the user data model. In this way, speech recognition can be performed using the corresponding user data model according to different application scenarios.
[0011] The speech recognition method includes: updating the user data model at first preset time intervals based on all received user characteristic hot words; and / or deleting user characteristic hot words that have been stored for a period greater than or equal to a second preset time interval. This ensures efficient updating of the user data model even when business scenarios change.
[0012] The server of an embodiment of the present invention includes a receiving module and a control module. The receiving module is used to: receive business scenario information uploaded by the vehicle when the business scenario of the vehicle changes; the control module is used to: generate source data of the user data model based on the business scenario information, and the business scenario information includes user feature hot words; generate or update the user data model based on the source data, and the user data model can be used to recognize voice requests sent by the vehicle.
[0013] In the above-mentioned server, the user characteristic hot words of the business scenario corresponding to the vehicle are determined through the business scenario information, and the user data model is finally obtained based on the user characteristic hot words. The user data model is used to identify voice requests, which is conducive to improving the accuracy of identifying voice requests in different scenarios.
[0014] A speech recognition system according to an embodiment of the present invention includes a vehicle and a server, wherein the vehicle is used to send business scenario information, wherein the business scenario information includes user feature hot words; the server is used to receive the business scenario information when the business scenario of the vehicle changes; source data of a user data model is generated based on the business scenario information; the user data model is generated or updated based on the source data, and the user data model can be used to recognize voice requests sent by the vehicle.
[0015] In the above-mentioned speech recognition system, the user characteristic hot words of the business scenario corresponding to the vehicle are determined through the business scenario information, and the user data model is finally obtained based on the user characteristic hot words. The user data model is used to identify voice requests, which is conducive to improving the accuracy of identifying voice requests in different scenarios.
[0016] The computer-readable storage medium according to the embodiment of the present invention stores a computer program thereon. When the computer program is executed by a processor, the speech recognition method according to any one of the above embodiments is implemented.
[0017] In the above-mentioned computer-readable storage medium, the user characteristic hot words of the business scenario corresponding to the vehicle are determined through the business scenario information, and the user data model is finally obtained based on the user characteristic hot words, and the voice request is identified through the user data model, which is conducive to improving the accuracy of identifying voice requests in different scenarios.
[0018] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0020] Figure 1 It is a flow chart of the speech recognition method of the present invention;
[0021] Figure 2 is a schematic diagram of a server of the present invention;
[0022] Figure 3 is a schematic diagram of a speech recognition system of the present invention;
[0023] Figures 4 and 5 It is a flow chart of the speech recognition method of the present invention;
[0024] Figure 6 is a schematic diagram of a speech recognition system of the present invention;
[0025] Figures 7 to 9 It is a flow chart of the speech recognition method of the present invention;
[0026] Figure 10 Schematic diagram of a scenario of the speech recognition method of the present invention;
[0027] Figure 11 It is a flow chart of the speech recognition method of the present invention;
[0028] Figure 12 It is a schematic diagram of the connection between the server and the computer-readable storage medium of the present invention.
[0029] Description of main component symbols:
[0030] Server 10 , receiving module 11 , control module 12 , processor 13 , vehicle 20 , mobile terminal 30 , speech recognition system 40 , and computer-readable storage medium 50 . DETAILED DESCRIPTION
[0031] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.
[0032] In related technologies, voice recognition is performed on the user's voice to determine the user's intention, and corresponding operations are performed according to the user's intention, so that the user can perform corresponding control functions through voice interaction and improve recognition accuracy through personalized services.
[0033] Specifically, when a voice request from a user is received, voice recognition can be performed on the voice request to obtain semantic information. When the semantic information is recognizable enough for the machine to discern, the control instruction corresponding to the voice request can be identified based on the semantic information, and the control instruction can be executed to achieve the effect of voice control. However, in actual applications, the voice requests issued by users are often difficult to recognize, which can easily lead to recognition errors and cause trouble for users.
[0034] See also Figure 1 and Figure 2 A speech recognition method according to an embodiment of the present invention can be used in the server 10. The speech recognition method includes:
[0035] 02: When the business scenario of vehicle 20 changes, receive the business scenario information uploaded by vehicle 20;
[0036] 03: Generate source data for the user data model based on business scenario information, which includes user feature hot words;
[0037] 04: Generate or update a user data model based on the source data. The user data model can be used to recognize the voice request sent by the vehicle 20.
[0038] See also Figure 2 、 Figure 3 and Figure 6 The speech recognition method of the embodiment of the present invention can be implemented by the server 10 of the embodiment of the present invention. The server 10 can be connected to the vehicle 20 for communication. The server 10 includes a receiving module 11 and a control module 12. Among them, step 01 can be implemented by the receiving module 11, step 02 can be implemented by the control module 12, and step 03 can be implemented by the control module 12. That is to say, the receiving module 11 can be used to receive the business scenario information uploaded by the vehicle 20 when the business scenario of the vehicle 20 changes. The control module 12 can be used to generate source data of the user data model based on the business scenario information, and the business scenario information includes user feature hot words; generate or update the user data model based on the source data, and the user data model can be used to recognize the voice request sent by the vehicle 20.
[0039] In the above-mentioned speech recognition method and server 10, the user characteristic hot words of the business scenario corresponding to the vehicle 20 are determined through the business scenario information, and the user data model is finally obtained based on the user characteristic hot words, and the voice request is recognized through the user data model, which is conducive to improving the accuracy of recognizing voice requests in different scenarios.
[0040] A business scenario can be a user's actual vehicle usage scenario for vehicle 20. In this actual vehicle usage scenario, the user can interact directly or indirectly with vehicle 20, allowing vehicle 20 to determine the current actual vehicle usage scenario based on the interaction, thereby determining whether the business scenario has changed. Business scenarios may include communication scenarios, navigation scenarios, multimedia scenarios, etc.
[0041] The vehicle 20 will send business scenario information when the current business scenario changes, and the server 10 can receive the business scenario information sent by the vehicle 20. The change of the business scenario can be generated by a new interaction of the user or triggered by the communication connection between the vehicle 20 and other terminal devices.
[0042] The user feature hot words can be text information. The business scenario information includes audio information, and the user feature hot words can be the text result obtained after the audio information is processed by Automatic Speech Recognition (ASR) technology.
[0043] After receiving the business scenario information, the server 10 can generate source data based on the user characteristic hot words included in the business scenario information, and train a user data model based on the generated source data. Since different business scenario information corresponds to different types of business scenarios, the user characteristic hot words included in different business scenario information will also correspond to different types of business scenarios. When the business scenario changes, personal content related to the user will be generated during the change. This personal content often records the user's personal interaction habits and preferences, so corresponding user characteristic hot words can be generated based on the personal content. The user data model obtained through user characteristic hot words training can be used to identify the voice requests issued by the user in the corresponding scenario.
[0044] See also Figure 3 and Figure 4 The user characteristic hot words are stored in the mobile terminal 30, and the mobile terminal 30 is capable of communicating with the vehicle 20. The speech recognition method includes:
[0045] 011: When the mobile terminal 30 and the vehicle 20 establish a connection, a business scenario change is determined.
[0046] See also Figure 2 Step 011 may be implemented by the control module 12. That is, the control module 12 may be configured to determine a change in the service scenario when the mobile terminal 30 and the vehicle 20 establish a connection.
[0047] In this way, the user data model can be quickly updated. Specifically, the mobile terminal 30 can be a smart phone, and the user feature hot words can be the address book stored in the smart phone. The address book can include the name and contact information of the contact. When the mobile terminal 30 and the vehicle 20 establish a connection (such as a Bluetooth connection), the current business scenario of the vehicle 20 can be confirmed, and it can be determined that the current business scenario changes while the connection is established, so that the user feature hot words in the information transmitted by the mobile terminal 30 to the vehicle 20 can be determined. After obtaining the user feature hot words stored in the mobile terminal 30, the vehicle 20 will upload them to the server 10, and the server 10 will use the received user feature hot words to generate source data.
[0048] When the mobile terminal 30 and the vehicle 20 establish a connection, information can be transmitted directly to the vehicle 20. The time when the connection is established between each other is used as a trigger for changes in the business scenario. This can increase the speed of generating the user data model based on the user feature hot words stored in the mobile terminal 30 in the current scenario, thereby facilitating the user data model to be immediately used to identify voice requests associated with the user feature hot words.
[0049] See also Figure 3 and Figure 5The user feature hot words include the address book hot words, the business scenario includes the communication scenario, and the vehicle 20 is used to obtain the address book hot words sent by the mobile terminal 30.
[0050] Step 02 (receiving business scenario information uploaded by vehicle 20 when the business scenario of vehicle 20 changes) includes:
[0051] 021: When the current business scenario is a communication scenario, receiving a hot word in the address book sent by the vehicle 20;
[0052] Step 03 (generating source data for the user data model based on business scenario information) includes:
[0053] 031: Synchronize address book hot words to source data.
[0054] See also Figure 2 Step 021 can be implemented by the receiving module 11, and step 031 can be implemented by the control module 12. That is, the receiving module 11 can be used to receive the address book hotword sent by the vehicle 20 when the current business scenario is the communication scenario. The control module 12 can be used to determine the business scenario change when the mobile terminal 30 and the vehicle 20 establish a connection.
[0055] In this way, the information in the address book can be easily updated to the user data model. Specifically, the address book hot words in the business scenario information can record the address book stored in the mobile terminal 30, and the address book can store the names of contacts. When the vehicle 20 and the mobile terminal 30 establish a connection, it can be determined that the current business scenario is a communication scenario, and it can be determined that the current business scenario has changed, so that the mobile terminal 30 uploads the business scenario information including the address book hot words to the vehicle 20, and then uploads the business scenario information to the server 10 through the vehicle 20, so that the server 10 receives the corresponding address book hot words, and synchronizes the received address book hot words to the source data to update the user data model.
[0056] In one application scenario, the contact names include "Zhang Shan" and "Zhang Shan." If the user sends a voice request for "Call Zhang Shan," server 10 can determine the corresponding user characteristic hotword is "Zhang Shan" based on the corresponding user data model and send the corresponding recognition result to vehicle 20. After receiving the recognition result, vehicle 20 can determine that the user needs to contact "Zhang Shan" and then enter the call scenario with "Zhang Shan" as the contact and make a voice call.
[0057] It will be appreciated that user feature hotwords are stored in the mobile terminal 30, and thus may not be directly associated with the vehicle 20. In this case, when the mobile terminal 30 is in communication with the vehicle 20, the user feature hotwords stored in the mobile terminal 30 are sent to the server 10 via the vehicle 20 to update the user feature hotwords stored in the mobile terminal 30 into the user data model. This allows the recognition of voice requests when the voice request contains content corresponding to the user feature hotwords stored in the mobile terminal 30, thus enabling rapid updating of the user data model without requiring user intervention.
[0058] See also Figure 6 and Figure 7 , user feature hot words are generated in the vehicle 20, and the speech recognition method includes:
[0059] 012: When the vehicle 20 generates a user characteristic hot word, a business scenario change is determined.
[0060] See also Figure 2 Step 012 may be implemented by the control module 12. That is, the control module 12 may be configured to determine a change in a business scenario when the vehicle 20 generates user feature hot words.
[0061] This facilitates rapid updating of the user data model. Specifically, the vehicle 20 can generate corresponding interaction results based on interactions with the user. Based on these interaction results, corresponding user feature hot words can be obtained. This allows the vehicle 20 to identify current business scenario changes when generating the corresponding interaction results and update the user data model based on the user feature hot words in the interaction results.
[0062] It's understandable that users often have personal preferences and habits when interacting. These interactions can be used to identify certain user preferences and habits. Consequently, user-specific hot words can be determined based on business scenario information that reflects these preferences and habits. The user data model generated based on these hot words can then identify the user's corresponding personal preferences and habits during subsequent interactions. This allows subsequent voice requests to be identified based on the hot words in the data model, helping to filter out potential interference and improve recognition rates.
[0063] See also Figure 6 and Figure 8 , the business scenario includes an interaction scenario, and the vehicle 20 is used to determine the interaction hot word generated for the first time as the user feature hot word.
[0064] Step 02 (receiving business scenario information uploaded by vehicle 20 when the business scenario of vehicle 20 changes) includes:
[0065] 022: When the current business scenario is an interactive scenario, receiving an interactive hot word sent by vehicle 20;
[0066] Step 03 (generating source data for the user data model based on business scenario information) includes:
[0067] 032: Synchronize interactive hot words to source data.
[0068] See also Figure 2 Step 022 can be implemented by the receiving module 11, and step 032 can be implemented by the control module 12. That is, the receiving module 11 can be used to receive the interactive hot words sent by the vehicle 20 when the current business scenario is an interactive scenario. The control module 12 can be used to synchronize the interactive hot words with the source data.
[0069] In this way, relevant information about the user's interaction with the vehicle 20 can be easily updated into the user data model.
[0070] The interaction scenario may include a navigation scenario. In a navigation scenario, the user interacts with the vehicle 20 for navigation. The vehicle 20 may determine that the user's destination is "Huancheng Garden" based on the interaction results obtained from the navigation, thereby triggering a change in the business scenario, so that the vehicle 20 can immediately use "Huancheng Garden" as an interaction hot word and upload it to the server 10. The server 10 may synchronize the received interaction hot words to the source data to update the user data model. In this way, when the user's voice request is subsequently received as "Go to Huancheng Garden" (huan indicates pronunciation), the server 10 may identify the corresponding user feature hot word as "Huancheng Garden" based on the user's user data model, thereby determining that the user needs to go to "Huancheng Garden" and sending the recognition result to the vehicle 20. The vehicle 20 may enter the navigation scenario with "Huancheng Garden" as the destination based on the recognition result.
[0071] The interactive scenario may include a music playback scenario. In a music playback scenario, the user interacts with the vehicle 20 to play music. The vehicle 20 may determine that the music the user wants to play is Song A based on the interaction result, and trigger a business scenario change, so that the vehicle 20 can immediately upload the information of Song A as an interactive hot word to the server 10. The server 10 may synchronize the received interactive hot words to the source data to update the user data model. In this way, when it is determined that the user's voice request is to play a piece of music, the server 10 may recognize that the user characteristic hot word corresponding to the music is Song A based on the user's user data model, thereby determining that the user needs to play Song A and sending the recognition result to the vehicle 20. The vehicle 20 may enter the music playback scenario of playing Song A based on the recognition result. The information of Song A may include the name of the music and lyrics.
[0072] See also Figure 9 , the speech recognition method includes:
[0073] 05: Load the corresponding user data model according to the vehicle 20 identification and user identification;
[0074] 06: receiving a voice request sent by the vehicle 20 while the vehicle 20 is traveling;
[0075] 07: Perform voice recognition on voice requests based on the user data model.
[0076] See also Figure 2 Step 06 can be implemented by the receiving module 11, and steps 05 and 07 can be implemented by the control module 12. That is, the receiving module 11 can be used to receive a voice request sent by the vehicle 20 while the vehicle 20 is traveling. The control module 12 can be used to load the corresponding user data model based on the vehicle 20 identifier and the user identifier, and to perform voice recognition on the voice request based on the user data model.
[0077] In this way, speech recognition can be performed using corresponding user data models according to different application scenarios.
[0078] Specifically, in some practical situations, there may be multiple users and multiple vehicles 20 usage scenarios. For the same vehicle 20, it may be driven by different users, and different users may have different needs. Therefore, it is necessary to determine the user currently driving the vehicle 20 based on different user identifiers to further determine the corresponding user data model for voice recognition. For the same user, there may be situations where they drive different vehicles 20, and they may have different voice recognition needs when driving different vehicles 20. For example, when driving different vehicles 20, they may drive one vehicle 20 and often issue work-related voice requests, while when driving other vehicles 20 for daily life, they may drive another vehicle 20 and often issue daily life-related voice requests. As a result, the user has different needs for the two vehicles 20, and it is necessary to determine the vehicle 20 currently driven by the user based on different material identifiers to further determine the corresponding user data model for voice recognition. Each vehicle 20 has a unique vehicle 20 identifier, and each user has a unique user identifier.
[0079] Through the vehicle 20 identifier and the user identifier, the specific vehicle 20 and the specific user in the current application scenario can be determined, so that the corresponding user data model can be loaded according to the vehicle 20 and the user in the current scenario to perform voice recognition on the voice request in the current scenario. In other words, the voice recognition method of the present invention can generate a corresponding user data model according to the vehicle 20 and the user in the business scenario. When the vehicle 20 or the user in the business scenario is different, the corresponding vehicle 20 identifier or the user identifier will also change, so that the server 10 will generate a new user data model according to the different vehicles 20 or users in the business scenario. Since each user data model can correspond to one of the vehicles 20 and one of the users, when a business scenario is formed by one of the vehicles 20 and one of the users, the corresponding user data model is loaded by the corresponding vehicle 20 identifier and the user identifier, which better adapts to the actual application scenario and can improve the recognition rate.
[0080] The process of generating and updating the user data model is decoupled from the process of performing speech recognition based on the user data model, making the two processes independent of each other. When business scenario information and voice requests are generated in a business scenario, the server 10 updates the user data model based on the received business scenario information and can reflect the updated results in the subsequent speech recognition process for the voice request.
[0081] Specifically, see Figure 10 , taking the case where the same vehicle 20 is driven by multiple users as an example. If the user in the current business scenario is determined to be user A based on the user identifier, and the received business scenario information is "go to Huancheng Garden", the vehicle 20 can determine that the current scenario is a navigation scenario, and can determine that the interaction result with the user is "go to Huancheng Garden", so that "Huancheng Garden" can be used as the user feature hot word of user A to update the user data model A. If the user in the current business scenario is determined to be user B based on the user identifier, and the received business scenario information is "go to Huancheng Garden", the vehicle 20 can determine that the current scenario is a navigation scenario, and can determine that the interaction result with the user is "go to Huancheng Garden", so that "Huancheng Garden" can be used as the user feature hot word of user B to update the user data model B.
[0082] Then, in a subsequent application scenario, when the received voice request is "Go to Huancheng Garden", if the user corresponding to the user identifier in the current application scenario is determined to be user A, user data model A will be loaded, and the recognition result obtained will be "Go to Huancheng Garden". If the user corresponding to the user identifier in the current application scenario is determined to be user B, user data model B will be loaded, and the recognition result obtained will be "Go to Huancheng Garden".
[0083] Accordingly, when the same user drives different vehicles 20, the server 10 receives information corresponding to the different business scenarios created by the user and the different vehicles 20, and generates or updates multiple different user data models based on the business scenario information. In this way, when the user drives different vehicles 20, different user data models will be loaded.
[0084] It can be understood that, depending on the actual application situation, based on different vehicles 20 and different users, the corresponding user data model can be comprehensively determined to be loaded based on different voice recognition rates, the storage capacity occupied by the user data model, the real-time recognition effect, and the number of voice assistants.
[0085] Alternatively, if the user ID cannot be determined, a preset voice recognition model may be loaded to perform voice recognition on the received voice request. The user ID cannot be determined because the user has not logged in to the vehicle 20. In this case, the user can still use the preset voice recognition model for basic voice recognition and voice navigation to most locations, but cannot play music or receive voice navigation to the user's relevant locations (such as home or work).
[0086] Please refer to Figure 11 , the speech recognition method includes:
[0087] 081: updating the user data model at a first preset time interval based on all received user feature hot words; or
[0088] 082: Delete user feature hot words whose storage time is greater than or equal to the second preset time interval.
[0089] See also Figure 2 , step 081 can be implemented by the control module 12, or step 082 can be implemented by the control module 12. That is, the control module 12 can be used to update the user data model at a first preset time interval based on all received user characteristic hot words, or can be used to delete user characteristic hot words whose storage time is greater than or equal to a second preset time interval.
[0090] In this way, the efficiency of updating the user data model can be guaranteed when the business scenario changes. In actual applications, due to the different types of business scenarios, in order to promptly determine the current business scenario, as well as the changes in the current business scenario, so as to be able to generate or update the user data model in a timely manner based on the received business scenario information, it is necessary to maintain regular maintenance of the user data model. By updating the user data model based on all user feature hot words at a first preset time interval, it can be beneficial to simplify the structure of the user data model, and avoid the inability to establish new user feature hot word connections in the user data model in a timely manner when the number of user feature hot words is large, which affects the efficiency of updating the user data model when the business scenario changes. The first preset time interval can be 1 day.
[0091] In addition, since some user-characteristic hot words may not be used by the user for a long time, for example, the user may not go to a certain place for a long time, call a certain contact for a long time, or play a certain song for a long time, when the storage time of the corresponding user-characteristic hot word is greater than the second preset time interval, it can be determined that the corresponding user-characteristic hot word is used less frequently and the user may no longer have the need to use it. Therefore, the corresponding user-characteristic hot word is deleted, simplifying the structure of the user data model and facilitating the server 10 to update the new user-characteristic hot word to the user data model. The second preset time interval can be greater than or equal to 30 days.
[0092] Of course, the control module 12 can also be used to update the user data model based on all received user-characteristic hot words at a first preset time interval, and to delete user-characteristic hot words whose storage time is greater than or equal to a second preset time interval. In other words, the control module 12 can implement step 081 and step 082 simultaneously.
[0093] See also Figure 3 and Figure 6 The speech recognition system 40 according to an embodiment of the present invention includes a vehicle 20 and a server 10. The vehicle 20 is configured to transmit business scenario information, which includes user-characteristic hot words. The server 10 is configured to receive the business scenario information when the business scenario of the vehicle 20 changes, generate source data for a user data model based on the business scenario information, and generate or update the user data model based on the source data. The user data model can be used to recognize speech requests transmitted by the vehicle 20.
[0094] In the above-described speech recognition method, the user characteristic hot words corresponding to the business scenario in which the vehicle 20 is located are determined through business scenario information, and a user data model is ultimately obtained based on the user characteristic hot words. The user data model is then used to identify voice requests, which helps improve the accuracy of identifying voice requests in different scenarios. The vehicle 20 and server 10 in the speech recognition system 40 can refer to the vehicle 20 and server 10 in the aforementioned embodiment, so that the speech recognition system 40 of the embodiment of the present invention can achieve the same or similar technical effects. It will not be repeated here.
[0095] See also Figure 12 The computer-readable storage medium 50 of the embodiment of the present invention stores a computer program thereon. When the computer program is executed by the processor 13 of the server 10, the speech recognition method of the above embodiment is implemented.
[0096] For example, when the computer program is executed by the processor 13, the following may be achieved:
[0097] 02: When the business scenario of vehicle 20 changes, receive the business scenario information uploaded by vehicle 20;
[0098] 03: Generate source data for the user data model based on business scenario information, which includes user feature hot words;
[0099] 04: Generate or update a user data model based on the source data. The user data model can be used to recognize the voice request sent by the vehicle 20.
[0100] In the above-mentioned computer-readable storage medium 50, the user characteristic hot words of the business scenario corresponding to the vehicle 20 are determined through the business scenario information, and the user data model is finally obtained based on the user characteristic hot words, and the voice request is identified through the user data model, which is conducive to improving the accuracy of identifying voice requests in different scenarios.
[0101] In the present invention, a computer program includes computer program code. The computer program code may be in source code form, object code form, executable file, or some intermediate form. The computer-readable storage medium 50 may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device. The processor 13 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0102] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0103] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0104] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A speech recognition method, characterized in that: The speech recognition method comprises: When the business scenario of a vehicle changes, receiving business scenario information uploaded by the vehicle; Generating source data of a user data model according to the business scenario information, wherein the business scenario information includes user feature hot words; generating or updating the user data model based on the source data, wherein the user data model can be used to recognize the voice request sent by the vehicle; Loading the corresponding user data model according to the vehicle identification and the user identification; receiving a voice request sent by the vehicle while the vehicle is traveling; Perform speech recognition on the voice request according to the user data model.
2. The speech recognition method according to claim 1, wherein: The user characteristic hot words are stored in a mobile terminal, and the mobile terminal is communicatively connected to the vehicle. The speech recognition method includes: When the mobile terminal and the vehicle establish a connection, the business scenario change is determined.
3. The speech recognition method according to claim 2, wherein: The user feature hot words include address book hot words, the business scenario includes a communication scenario, and the vehicle is used to obtain the address book hot words sent by the mobile terminal; When the business scenario of a vehicle changes, receiving business scenario information uploaded by the vehicle includes: When the current business scenario is the communication scenario, receiving the address book hotword sent by the vehicle; Generate source data for the user data model based on the business scenario information, including: The address book hot words are synchronized to the source data.
4. The speech recognition method according to claim 1, wherein: The user characteristic hot words are generated in the vehicle, and the speech recognition method includes: When the vehicle generates the user characteristic hot word, the business scenario change is determined.
5. The speech recognition method according to claim 4, characterized in that The business scenario includes an interaction scenario, and the vehicle is used to determine the interaction hot word generated for the first time as the user feature hot word; When the business scenario of a vehicle changes, receiving business scenario information uploaded by the vehicle includes: When the current business scenario is the interaction scenario, receiving the interaction hotword sent by the vehicle; Generate source data for the user data model based on the business scenario information, including: The interactive hot words are synchronized to the source data.
6. The speech recognition method according to claim 1, wherein: The speech recognition method comprises: Based on all received user feature hot words, updating the user data model at first preset time intervals; and / or The user characteristic hot words whose storage time is greater than or equal to the second preset time interval are deleted.
7. A server, characterized in that: The server includes a receiving module and a control module. The receiving module is used for: When the business scenario of a vehicle changes, receiving business scenario information uploaded by the vehicle; receiving a voice request sent by the vehicle while the vehicle is traveling; The control module is used for: Generate source data of a user data model based on the business scenario information, wherein the business scenario information includes user feature hot words; generating or updating the user data model based on the source data, wherein the user data model can be used to recognize the voice request sent by the vehicle; Loading the corresponding user data model according to the vehicle identification and the user identification; Perform speech recognition on the voice request according to the user data model.
8. A speech recognition system, characterized in that: include: The vehicle is configured to: send business scenario information, wherein the business scenario information includes user feature hot words; and Server for: When the business scenario of the vehicle changes, receiving the business scenario information; Generate source data of the user data model based on the business scenario information; generating or updating the user data model based on the source data, wherein the user data model can be used to recognize the voice request sent by the vehicle; Loading the corresponding user data model according to the vehicle identification and the user identification; receiving a voice request sent by the vehicle while the vehicle is traveling; Perform speech recognition on the voice request according to the user data model.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the speech recognition method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Recommendation method and recommendation system of hot word
CN104252470A
Speech recognition method, device and apparatus
CN109523991A
Voice interaction method, server and computer readable storage medium
CN112164401A
Speech recognition method and device thereof and storage medium
CN112767917A