A voice recognition method and device, a vehicle terminal, a server and a medium

By combining voiceprint feature recognition on the in-vehicle terminal with a server-generated speech recognition model, the privacy and real-time issues caused by cloud deployment are resolved, achieving efficient speech recognition on the in-vehicle terminal and improving the user experience.

CN116386610BActive Publication Date: 2026-02-06HUIZHOU DESAY SV AUTOMOTIVE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310430336.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-20
Publication Date
2026-02-06
Estimated Expiration
2043-04-20

AI Technical Summary

Technical Problem

Most existing speech recognition methods are deployed in the cloud, which cannot effectively meet users' privacy protection needs, and the real-time performance of speech recognition is poor in offline in-vehicle scenarios, resulting in a reduced user experience.

Method used

User identification is determined by voiceprint feature recognition on the vehicle terminal side, and speech recognition is performed based on the speech recognition model generated by the server to protect user privacy. At the same time, the model is updated and distributed in near real-time through the end-to-cloud collaborative technology framework to improve the real-time performance of speech recognition.

Benefits of technology

While ensuring user privacy protection, the real-time performance and accuracy of voice recognition in the vehicle terminal have been improved, thus enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386610B_ABST
    Figure CN116386610B_ABST
Patent Text Reader

Abstract

The application discloses a voice recognition method and device, a vehicle-mounted terminal, a server and a medium. The method is applied to the vehicle-mounted terminal, performs voiceprint feature recognition on a target voice signal, obtains a target user identifier corresponding to the target voice signal, determines a first voice recognition model of a first user based on the target user identifier and at least one voice recognition model, generates the at least one voice recognition model by the server based on a voice signal sent by the vehicle-mounted terminal, and returns the at least one voice recognition model to the vehicle-mounted terminal, performs voice recognition on the target voice signal according to the first voice recognition model, obtains instruction information corresponding to the target voice signal, and provides a service for the first user based on the instruction information. The method performs voice recognition on the vehicle side, and the at least one voice recognition model required is returned to the vehicle-mounted terminal by the server at a fixed time, rather than being acquired in real time, thereby improving the real-time performance of voice recognition on the basis of meeting the demand of user privacy protection, and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech recognition, and in particular to a speech recognition method and device, a vehicle-mounted terminal, a server and a medium. BACKGROUND

[0002] With the rapid development of technology, the functions of automobiles are becoming more and more rich, and intelligent interaction of a vehicle-mounted speech system has become an important part in the field of automobiles, such as a user can control a vehicle-mounted device through speech.

[0003] Most of the existing speech recognition methods are deployed in the cloud, which cannot better meet the needs of user privacy protection, and in an offline vehicle-mounted scene, the real-time performance of speech recognition is poor, which reduces the user experience. SUMMARY

[0004] The present application provides a speech recognition method and device, a vehicle-mounted terminal, a server and a medium, to improve the real-time performance of speech recognition on the basis of meeting the needs of user privacy protection, and to improve the user experience.

[0005] According to an aspect of the present application, a speech recognition method is provided, applied to a vehicle-mounted terminal, comprising:

[0006] receiving a target speech signal of a first user, and performing voiceprint feature recognition on the target speech signal to obtain a target user identifier corresponding to the target speech signal;

[0007] determining a first speech recognition model of the first user based on the target user identifier and at least one speech recognition model, the at least one speech recognition model being generated by a server based on a speech signal sent by the vehicle-mounted terminal and returned to the vehicle-mounted terminal;

[0008] performing speech recognition on the target speech signal according to the first speech recognition model to obtain instruction information corresponding to the target speech signal, and providing services for the first user based on the instruction information.

[0009] According to another aspect of the present application, a speech recognition method is provided, applied to a server, comprising:

[0010] receiving a speech signal of a second user, the speech signal being sent by a vehicle-mounted terminal;

[0011] determining a vehicle-mounted language model and a vehicle-mounted acoustic model based on the speech signal, the vehicle-mounted language model corresponding to the vehicle-mounted terminal, and the vehicle-mounted acoustic model corresponding to the second user;

[0012] processing the vehicle-mounted language model and the vehicle-mounted acoustic model to obtain a second speech recognition model corresponding to the second user;

[0013] return at least one speech recognition model to the vehicle terminal, the at least one speech recognition model comprising the second speech recognition model.

[0014] According to another aspect of the present application, there is provided a speech recognition apparatus configured in a vehicle terminal, comprising:

[0015] a voiceprint recognition module configured to receive a target speech signal of a first user and perform voiceprint feature recognition on the target speech signal to obtain a target user identifier corresponding to the target speech signal;

[0016] a first determination module configured to determine a first speech recognition model of the first user based on the target user identifier and at least one speech recognition model, the at least one speech recognition model being generated by a server based on a speech signal sent by the vehicle terminal and returned to the vehicle terminal;

[0017] a speech recognition module configured to perform speech recognition on the target speech signal according to the first speech recognition model to obtain instruction information corresponding to the target speech signal, and provide a service for the first user based on the instruction information.

[0018] According to another aspect of the present application, there is provided a speech recognition apparatus configured in a server, comprising:

[0019] a signal receiving module configured to receive a speech signal of a second user, the speech signal being sent by a vehicle terminal;

[0020] a second determination module configured to determine a vehicle language model and a vehicle acoustic model based on the speech signal, the vehicle language model corresponding to the vehicle terminal, and the vehicle acoustic model corresponding to the second user;

[0021] a processing module configured to process the vehicle language model and the vehicle acoustic model to obtain a second speech recognition model corresponding to the second user;

[0022] a returning module configured to return at least one speech recognition model to the vehicle terminal, the at least one speech recognition model comprising the second speech recognition model.

[0023] According to another aspect of the present application, there is provided a vehicle terminal, comprising:

[0024] at least one processor; and

[0025] a memory in communication connection with the at least one processor; wherein

[0026] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the speech recognition method according to the first or second embodiment of the present application.

[0027] According to another aspect of the present application, a server is provided, the server comprising:

[0028] at least one processor; and

[0029] a memory in communication connection with the at least one processor; wherein

[0030] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the speech recognition method according to the third or fourth embodiment of the present application.

[0031] According to another aspect of the present application, a computer readable storage medium is provided, the computer readable storage medium stores computer instructions for enabling a processor to implement the speech recognition method according to any of the embodiments of the present application when executed.

[0032] The embodiments of the present application provide a speech recognition method, device, vehicle-mounted terminal, server and medium, the method is applied to a vehicle-mounted terminal, and the method comprises the following steps: receiving a target speech signal of a first user, performing voiceprint feature recognition on the target speech signal, and obtaining a target user identifier corresponding to the target speech signal; determining a first speech recognition model of the first user based on the target user identifier and at least one speech recognition model, the at least one speech recognition model is generated by a server based on a speech signal sent by the vehicle-mounted terminal and returned to the vehicle-mounted terminal; performing speech recognition on the target speech signal according to the first speech recognition model, obtaining instruction information corresponding to the target speech signal, and providing a service for the first user based on the instruction information. By using the above technical solution, the first speech recognition model is determined based on the at least one speech recognition model on the vehicle side, and speech recognition is performed, so that the user privacy protection requirement can be met, meanwhile, the at least one speech recognition model required is returned to the vehicle-mounted terminal by the server at a time, rather than being acquired in real time, and on the basis of meeting the user privacy protection requirement, the real-time performance of speech recognition is further improved, so that the user experience is improved.

[0033] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0034] In order to make the technical solution in the embodiments of the present application clearer, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some of the embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative effort based on these drawings.

[0035] Figure 1 is a flow chart of a speech recognition method according to the first embodiment of the present application;

[0036] Figure 2 is a flow chart of a speech recognition method according to the second embodiment of the present application;

[0037] Figure 3 is a schematic diagram of determining a vehicle-mounted language model according to the second embodiment of the present application;

[0038] Figure 4 is a structural schematic diagram of implementing a speech recognition method according to the second embodiment of the present application;

[0039] Figure 5 is a flow chart of a speech recognition method according to the second embodiment of the present application;

[0040] Figure 6 is a structural schematic diagram of a speech recognition device according to the third embodiment of the present application;

[0041] Figure 7 is a structural schematic diagram of a speech recognition device according to the fourth embodiment of the present application;

[0042] Figure 8 is a structural schematic diagram of a vehicle-mounted terminal according to the fifth embodiment of the present application;

[0043] Figure 9 is a structural schematic diagram of a server according to the sixth embodiment of the present application. DETAILED DESCRIPTION

[0044] In order to make the technical solution in the embodiments of the present application clearer, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some of the embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative effort based on these drawings.

[0045] It should be noted that the terms "first", "second", and the like in the description and in the claims of the present application and above-described accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular sequential or chronological order. It should be understood that the data thus used can be interchanged under appropriate circumstances so that the embodiments of the application described herein can be implemented in other than the order illustrated or described herein. Furthermore, the terms "comprise" and "have", and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or apparatus that includes a list of steps or units as non-limiting to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to such processes, methods, products, or apparatuses.

[0046] Embodiment one

[0047] Figure 1 It is a flowchart of a speech recognition method according to an embodiment one of the present application. The embodiment can be applicable to the case of recognizing the speech of a user. The method can be executed by a speech recognition device, which can be realized in the form of hardware and / or software. The speech recognition device can be configured in a vehicle terminal.

[0048] In the era of intelligent driving, intelligent interaction of the vehicle-mounted speech system has become an important part and has been widely applied. However, how to apply the off-line vehicle-mounted scene with insufficient hardware computing power to make the end-side speech recognition service achieve the same high recognition accuracy and real-time performance as the cloud model service, and meet the needs of user privacy protection, puts higher requirements on the end-side speech recognition capability.

[0049] It can be considered that the vehicle terminal does not have enough training data matching the scene (such as environment, speaker, and theme) before the deployment of the speech application, and the user usage habits of the vehicle-mounted scene have the following significant features: generally used for home use, and the users of a vehicle are relatively fixed; the speaker style and theme, i.e., usage habits (such as pronunciation, speaking rate, expression habits, commonly used vehicle control instructions, etc.) of each user are relatively fixed, and the above data provide conditions for the pruning of the cloud large model and user adaptive training. Therefore, the current speech recognition model is mainly deployed in the cloud, and the model architecture is generally large.

[0050] Therefore, how to apply and improve the recognition rate at the end side under low resources is the focus of the industry, especially for vehicle-mounted applications, the hardware edge computing power is insufficient, and the demand for timeliness and recognition rate in the vehicle-mounted special field cannot be fully met. Secondly, from the perspective of speech recognition technology, the vehicle terminal does not have enough training data matching the scene before the speech application is deployed, and the mismatch between training and testing can be solved by adaptive technology, so that the model is more suitable for the real deployment environment or the input features of the test are more suitable for the existing model. Therefore, how to use the initial system deployed to obtain speech data in the vehicle-mounted scene, and use the text annotation inferred by the initial system (speaker-independent general system) for adaptive training to further improve the speech recognition accuracy, and how to go out of the research level and enter the industry is a real challenge.

[0051] Based on this, the embodiment of the application provides a speech recognition method, through the technical application framework of the end-side adaptive recognition module and the cloud-side interaction module, the speech recognition demand of the end-side vehicle-mounted field can be realized, the model is updated and issued in quasi-real time through the end-cloud cooperation, so that the speech-driven intelligent vehicle control operation design can be realized, and the speech recognition ability of the vehicle-mounted end side in the offline use of the vehicle control system can be improved, instead of the recognition ability in the general field. As shown in Figure 1 The method comprises the following steps:

[0052] S110, receiving a target speech signal of a first user, and performing voiceprint feature recognition on the target speech signal to obtain a target user identifier corresponding to the target speech signal.

[0053] The first user can be considered as a current user who controls the current vehicle to serve him / her through speech, such as the driver of the current vehicle, or the passenger of the current vehicle, etc. The target speech signal is the speech signal of the first user, which is used to control the current vehicle to serve, such as playing songs, etc. The target user identifier can be understood as the user identifier of the first user recognized.

[0054] In this embodiment, the target speech signal of the first user can be received first, and after receiving the target speech signal, the voiceprint feature recognition is performed on the target speech signal to obtain the target user identifier corresponding to the target speech signal, so as to perform subsequent speech recognition and service, etc. The way of performing voiceprint feature recognition to obtain the target user identifier is not limited, such as directly obtaining the target user identifier corresponding to the target speech signal through the voiceprint feature recognition model, or obtaining the voiceprint feature information by performing voiceprint feature recognition on the target speech signal, and comparing the voiceprint feature information with the pre-stored voiceprint feature information set to determine the target user identifier corresponding to the target speech signal, etc.

[0055] S120, determine the first speech recognition model of the first user based on the target user identifier and at least one speech recognition model, the at least one speech recognition model being generated by the server based on the speech signal sent by the vehicle terminal and returned to the vehicle terminal.

[0056] The at least one speech recognition model can be considered as a speech recognition model suitable for the current vehicle, which can be generated by the server based on the speech signal sent by the vehicle terminal, such as the server can generate a speech recognition model corresponding to each user based on the speech signal of the user using the current vehicle, and return each speech recognition model to the vehicle terminal, so that the vehicle terminal can directly perform speech recognition based on each speech recognition model and perform subsequent services. The process of generating the at least one speech recognition model by the server is not described in detail here. The first speech recognition model can be considered as a speech recognition model corresponding to the first user, which is used for speech recognition of the target speech signal of the first user.

[0057] In one embodiment, the present embodiment can generate a speech recognition model suitable for each user, and each speech recognition model corresponds to a user identifier, so that after the target user identifier is recognized, the first speech recognition model corresponding to the first user can be determined according to the target user identifier.

[0058] S130, perform speech recognition on the target speech signal according to the first speech recognition model, obtain instruction information corresponding to the target speech signal, and provide services for the first user based on the instruction information.

[0059] The instruction information can be instruction information for providing services for the first user, and the specific content of the instruction information is not limited, for example, the instruction information can include specific command words, or can include first user related information, etc.

[0060] Specifically, after the first speech recognition model of the first user is determined, the target speech signal can be speech recognized according to the determined first speech recognition model to obtain the instruction information corresponding to the target speech signal, so as to provide corresponding services for the first user based on the instruction information. For example, when the target speech signal of the first user is "I want to listen to a song", the speech recognition based on the present embodiment can determine that the corresponding instruction information is to play a song and the type of the song is a song or a children's song, etc., so that the current vehicle can play the corresponding song for the first user to respond to the target speech signal of the first user. Or when the target speech signal of the first user is "I want to go to xx destination", the speech recognition based on the present embodiment can determine that the corresponding instruction information is to navigate to the xx destination, and then the current vehicle can provide corresponding navigation services for the first user, etc.

[0061] The embodiment one of the present application provides a voice recognition method, receiving a target voice signal of a first user, and performing voiceprint feature recognition on the target voice signal to obtain a target user identifier corresponding to the target voice signal; determining a first voice recognition model of the first user based on the target user identifier and at least one voice recognition model, the at least one voice recognition model being generated by a server based on a voice signal sent by the vehicle terminal and returned to the vehicle terminal; performing voice recognition on the target voice signal according to the first voice recognition model to obtain instruction information corresponding to the target voice signal, and providing services for the first user based on the instruction information. By using the method, the first voice recognition model is determined based on the at least one voice recognition model on the vehicle side, and voice recognition is performed, which can meet the user privacy protection requirement. Meanwhile, the at least one voice recognition model required is returned to the vehicle terminal by the server at a fixed time, rather than being acquired in real time, which further improves the real-time performance of voice recognition on the basis of meeting the user privacy protection requirement, thereby improving the user experience.

[0062] In one embodiment, the voiceprint feature recognition on the target voice signal to obtain the target user identifier corresponding to the target voice signal comprises:

[0063] Performing voiceprint feature recognition on the target voice signal to obtain first voiceprint feature information, and determining whether the first voiceprint feature information is contained in a voiceprint feature information set stored locally, the voiceprint feature information set being obtained from the server at a fixed time, and the voiceprint feature information set being used to store voiceprint feature information corresponding to each user identifier;

[0064] If yes, the user identifier corresponding to the first voiceprint feature information in the voiceprint feature information set is taken as the target user identifier corresponding to the target voice signal;

[0065] If no, preset voiceprint feature information satisfying a preset feature condition is selected from the voiceprint feature information set, and the user identifier corresponding to the preset voiceprint feature information is taken as the target user identifier corresponding to the target voice signal.

[0066] The first voiceprint feature information can be information obtained by performing voiceprint feature recognition on the target voice signal, and is used to represent the voiceprint feature of the target voice signal. The voiceprint feature information set can be considered as a set obtained from the server at a fixed time, and is used to store voiceprint feature information corresponding to each user identifier. The voiceprint feature information set can be a set of voiceprint feature information configured before the vehicle is offline, or a set of voiceprint feature information constantly stored and updated based on the voice signal of the user in the use process of the current vehicle, which is not limited in the embodiment.

[0067] The preset feature condition can be a pre-set condition for selecting the preset voiceprint feature information. The preset feature condition can be configured according to an experience value, such as a similarity to the first voiceprint feature information being greater than a set threshold, or containing certain feature information at the same time as the first voiceprint feature information, and the like. The preset voiceprint feature information is the voiceprint feature information in the voiceprint feature information set that satisfies the preset feature condition.

[0068] In one embodiment, voiceprint feature recognition can be performed on the target voice signal to obtain first voiceprint feature information. After obtaining the first voiceprint feature information, it is determined whether the voiceprint feature information set stored locally contains the first voiceprint feature information, so as to determine the target user identifier according to the different determination results. For example, if it is determined that the voiceprint feature information set stored locally contains the first voiceprint feature information, it indicates that the voiceprint feature information of the current speaker is stored in the voiceprint feature information set. At this time, the user identifier corresponding to the first voiceprint feature information in the voiceprint feature information set can be taken as the target user identifier corresponding to the target voice signal. In this way, the target user identifier corresponding to the target voice signal can be obtained. If it is determined that the voiceprint feature information set stored locally does not contain the first voiceprint feature information, it indicates that the voiceprint feature information of the current speaker is not stored in the voiceprint feature information set. At this time, the preset voiceprint feature information that satisfies the preset feature condition can be selected from the voiceprint feature information set, and the user identifier corresponding to the preset voiceprint feature information can be taken as the target user identifier corresponding to the target voice signal. On this basis, the determination of the target user identifier is realized, which provides a basis for subsequent determination of the first voice recognition model.

[0069] In one embodiment, after the preset voiceprint feature information that satisfies the preset feature condition is selected from the voiceprint feature information set, the method further includes:

[0070] The first voiceprint feature information is sent to the server, so that the server updates the voiceprint feature information set.

[0071] In one embodiment, if the voiceprint feature information set stored locally does not contain the first voiceprint feature information, that is, the voiceprint feature information of the current speaker is not stored in the voiceprint feature information set, the first voiceprint feature information can be sent to the server, so that the server updates the voiceprint feature information corresponding to each user identifier, so as to obtain the updated voiceprint feature information set at a regular time. On this basis, the voiceprint feature information of different users can be recorded through the update of the voiceprint feature information set, so that more accurate voice recognition can be realized, and the user experience is improved.

[0072] In one embodiment, the method further includes:

[0073] send the target voice signal of the first user to the server, so that the server generates or optimizes the first voice recognition model of the first user based on the target voice signal.

[0074] In one embodiment, the present embodiment can also send the target voice signal of the first user to the server, so that the server can generate or optimize the first voice recognition model of the first user based on the target voice signal. On this basis, the voice recognition model of the user can be enriched or optimized, thereby improving the accuracy of voice recognition and improving user experience.

[0075] For example, if it is judged that the first voiceprint feature information of the first user is not included in the voiceprint feature information set, the first user can be considered as a new user of the current vehicle. At this time, the present embodiment can send the target voice signal of the user (i.e. the first user) to the server, so that the server generates the voice recognition model corresponding to the user based on the target voice signal, thereby providing more accurate service for the user when subsequently performing voice recognition on the voice signal of the user, and realizing the improvement of user experience.

[0076] For example, if the first voiceprint feature information of the first user is already included in the voiceprint feature information set or the user already has a corresponding voice recognition model, the present embodiment can also send the target voice signal of the user to the server according to actual conditions, so as to further optimize the first voice recognition model of the user, thereby more accurately realizing voice recognition.

[0077] Embodiment Two

[0078] Figure 2 is a flowchart of a voice recognition method according to Embodiment Two of the present application. The present embodiment can be applicable to the case of recognizing the voice of a user. The method can be executed by a voice recognition device, which can be realized in the form of hardware and / or software, and can be configured in a server. As shown in Figure 2 The method comprises:

[0079] S210, receiving a voice signal of a second user, the voice signal being sent by a vehicle terminal.

[0080] The second user can be any user using the current vehicle, such as the first user or other users except the first user.

[0081] The present embodiment can receive the voice signal of the second user sent by the vehicle terminal, so as to perform the operation of the subsequent steps based on the voice signal. The present embodiment does not limit the method of receiving the voice signal, as long as the voice signal can be received.

[0082] S220, determine a vehicle language model and a vehicle acoustic model based on the voice signal, the vehicle language model corresponding to the vehicle terminal, and the vehicle acoustic model corresponding to the second user.

[0083] The vehicle language model can be considered as a language model in the vehicle field, and is used for semantic analysis of the voice signal. The vehicle acoustic model can be considered as an acoustic model corresponding to the user, and different users can correspond to different vehicle acoustic models.

[0084] After receiving the voice signal of the second user, the vehicle language model and the vehicle acoustic model can be determined based on the received voice signal. The specific determination process of the model can be different based on the current actual situation, such as distinguishing and determining the process of the vehicle language model and the vehicle acoustic model according to whether the second voice recognition model corresponding to the second user exists or not. The corresponding vehicle language model and vehicle acoustic model can also be directly generated after receiving the voice signal of the second user each time, which is not limited in the embodiment.

[0085] In one embodiment, the vehicle language model and the vehicle acoustic model are determined based on the voice signal, comprising:

[0086] If the second voice recognition model corresponding to the second user does not exist, the vehicle language model and the vehicle acoustic model are generated based on the voice signal.

[0087] In one embodiment, it can be judged whether the second voice recognition model corresponding to the second user exists or not, and the determination of the vehicle language model and the vehicle acoustic model is performed according to the judgment result. For example, if the second voice recognition model corresponding to the second user does not exist, the corresponding vehicle language model and vehicle acoustic model need to be generated based on the voice signal, so as to obtain the second voice recognition model of the second user according to the generated vehicle language model and vehicle acoustic model.

[0088] S230, processing the vehicle language model and the vehicle acoustic model to obtain the second voice recognition model corresponding to the second user.

[0089] After obtaining the vehicle language model and the vehicle acoustic model through the above steps, the vehicle language model and the vehicle acoustic model can be processed to obtain the second voice recognition model corresponding to the second user. The specific means of processing is not limited, and can be determined by the configuration personnel according to the actual situation. For example, the obtained vehicle language model and vehicle acoustic model can be integrated to obtain the second voice recognition model based on the sequence discriminative training system (such as maximum mutual information), and the limited weighted state transition machine (WFST), CTC and other decoder representation technologies.

[0090] S240, return at least one speech recognition model to the vehicle terminal, the at least one speech recognition model comprising the second speech recognition model.

[0091] After obtaining the second speech recognition model corresponding to the second user, at least one speech recognition model containing the second speech recognition model can be returned to the vehicle terminal. The timing of returning is not limited. For example, at least one speech recognition model can be returned to the vehicle terminal after receiving a request from the vehicle terminal. At least one speech recognition model can also be returned to the vehicle terminal at a fixed time. In addition, at least one speech recognition model can be returned according to actual conditions, such as good network conditions or idle time.

[0092] The second embodiment of the present application provides a speech recognition method. The method receives a speech signal of a second user, wherein the speech signal is sent by a vehicle terminal. A vehicle language model and a vehicle acoustic model are determined based on the speech signal. The vehicle language model corresponds to the vehicle terminal, and the vehicle acoustic model corresponds to the second user. The vehicle language model and the vehicle acoustic model are processed to obtain a second speech recognition model corresponding to the second user. At least one speech recognition model is returned to the vehicle terminal, and the at least one speech recognition model includes the second speech recognition model. By receiving the speech signal of the second user, the second speech recognition model of the second user can be determined. At least one speech recognition model containing the second speech recognition model is returned to the vehicle terminal, which meets the demand for speech recognition on the vehicle side, improves the real-time performance of speech recognition, and improves the user experience.

[0093] In one embodiment, the vehicle language model and the vehicle acoustic model are determined based on the speech signal, comprising:

[0094] If there is a second speech recognition model corresponding to the second user, the vehicle language model and the vehicle acoustic model are optimized based on the speech signal.

[0095] The vehicle language model and the vehicle acoustic model are processed to obtain the second speech recognition model corresponding to the second user, comprising:

[0096] The optimized vehicle language model and the optimized vehicle acoustic model are processed to obtain a new speech recognition model corresponding to the second user, and the new speech recognition model is used to replace the second speech recognition model.

[0097] In one embodiment, if the second speech recognition model corresponding to the second user currently exists, it indicates that the in-vehicle language model and the in-vehicle acoustic model corresponding to the second user currently exist, the existing in-vehicle language model and the in-vehicle acoustic model can be optimized based on the speech signal, so as to obtain a new second speech recognition model of the second user according to the optimized in-vehicle language model and the optimized in-vehicle acoustic model, and replace the second speech recognition model with the new speech recognition model. After obtaining the new second speech recognition model of the second user, the existing second speech recognition model can be deleted, and the new second speech recognition model can be saved.

[0098] In one embodiment, the in-vehicle language model and the in-vehicle acoustic model are determined based on the speech signal, comprising:

[0099] The in-vehicle language model is determined according to the general topic model and the in-vehicle topic model; and,

[0100] The in-vehicle acoustic model is obtained by training the original model based on the signal attribute of the speech signal.

[0101] The general topic model can be considered as a topic model commonly used in any scene, and the in-vehicle topic model can be considered as a topic model in the scene of the in-vehicle field. The signal attribute can be used to represent the attribute of the speech signal, such as the voiceprint information, pronunciation characteristics, and speech style of the speech signal.

[0102] It can be considered that in the current field adaptation scene of speech recognition, acoustic model adaptation is more commonly used, and the adjustment and optimization on the language model are mostly concentrated in traditional statistical models, such as interpolation of ngram model through srilm tool package, and deep learning method is less researched. Therefore, in the in-vehicle speech recognition system on the terminal side, the in-field adaptation training of the language model can be performed to further improve the recognition rate.

[0103] In one embodiment, the in-vehicle language model can be determined according to the general topic model and the in-vehicle topic model, so as to improve the accuracy of the in-vehicle language model through the in-vehicle topic model and reduce the error rate of recognition. The specific method for determining the in-vehicle language model is not limited, such as determining the in-vehicle language model according to the type of specific model. For example, for a traditional statistical language model such as ngram, the interpolation method can be used to combine the general topic model and the in-vehicle topic model to obtain an in-vehicle language model which can achieve effective adaptation effect; for a neural network language model, the deep domain applicability technology in transfer learning can be used to align the data distribution of the source domain and the target domain with a deep neural network, so as to solve the distribution difference between the training set (source domain) and the test set (target domain) and overcome the catastrophic forgetting problem of the neural network.

[0104] Figure 3is a schematic diagram of determining a vehicle-mounted language model according to Embodiment Two of the present application, as shown Figure 3 As shown, a general multi-theme scene large model LM0 can be prepared first, such as cross-scene theme corpus 0, theme corpus 1, theme corpus 2, …; vehicle-mounted theme corpus is prepared, such as vehicle-mounted domain theme model car-LM; then pruning is respectively performed, and the pruned corpora are merged and integrated by using an interpolation method, that is, a reference test corpus set can be selected, an optimal interpolation ratio is calculated by using a compute-best-mix method of an Srilm tool package, and interpolation of models under different themes is realized.

[0105] In actual application, training an acoustic model on all scene data and then training a language model under a vehicle-mounted theme can reduce a word error rate of the system from 11% to 3%, thereby improving the accuracy of the vehicle-mounted language model.

[0106] It should be noted that the adaptive technology can compensate for the mismatch of acoustic conditions (different probability distributions) in training data and test data, and further improve the accuracy of speech recognition.

[0107] In one embodiment, the original model can also be trained based on signal properties of the speech signal to obtain a corresponding vehicle-mounted acoustic model.

[0108] For example, a general and robust initial acoustic model (i.e., an original model) can be realized first by using feature adaptation, HMM-GMM / DNN acoustic models, or CTC-based acoustic models.

[0109] Then, a speaker identifier is obtained by using voiceprint recognition, and a speaker-related acoustic model is trained by using an adaptive method, which includes but is not limited to the following methods:

[0110] (1) A speaker subspace method is constructed. Speaker feature representation can be performed based on a traditional speaker vector i-vector and a neural network-based d-vector, and the speaker-related neural network model is trained by splicing the spectral features of the speaker.

[0111] (2) A speaker adaptation method based on a parameter-based framework, such as a base class adaptive training and an intrinsic sound. For example, a plurality of mean vector groups can be used for each Gaussian component first; then a set of interpolation weights unique to each speaker is estimated; and finally, the mean vector base is interpolated to obtain a mean vector unique to the speaker.

[0112] (3) In order to reduce the model overhead, a local adaptive DNN model can also be used, which only adapts part of the model, for example, only the input layer, the output layer, or a specific hidden layer, such as the SVD bottleneck layer adaptive technology, which only adapts the diagonal matrix of the model parameter W after SVD decomposition.

[0113] Finally, in addition to the above-mentioned language model adaptation, speaker adaptation for pronunciation characteristics, speaker adaptation for style and preference topics can also be completed through unsupervised adaptation. Specifically, a (audio, converted text) pair with high confidence can be obtained through a cloud voice recognition initial system; on the basis of the baseline system, a corresponding in-vehicle acoustic model is obtained through adaptive training using multi-task learning techniques such as target function weighting or transfer learning techniques.

[0114] In an embodiment, the method further comprises:

[0115] receiving second voiceprint feature information of a third user, the second voiceprint feature information being sent by the in-vehicle terminal;

[0116] updating the voiceprint feature information set according to the second voiceprint feature information;

[0117] sending the updated voiceprint feature information set to the in-vehicle terminal.

[0118] The second voiceprint feature information can refer to the voiceprint feature information of the third user, and the third user can be any user using the current vehicle, such as the first user or the second user, or any other user except the first user and the second user.

[0119] In an embodiment, the second voiceprint feature information of the third user sent by the in-vehicle terminal can also be received, and the voiceprint feature information set can be updated after receiving the second voiceprint feature information, so as to send the updated voiceprint feature information set to the in-vehicle terminal. The process of updating the voiceprint feature information set is not limited, such as directly adding the second voiceprint feature information to the voiceprint feature information set, or further generating a user identifier corresponding to the second voiceprint feature information to improve the comprehensiveness of updating the voiceprint feature information set.

[0120] In an embodiment, the voiceprint feature information set is updated according to the second voiceprint feature information, comprising:

[0121] generating a user identifier corresponding to the second voiceprint feature information as a user identifier of the third user;

[0122] storing the user identifier of the third user and the second voiceprint feature information in the voiceprint feature information set.

[0123] In one embodiment, a user identifier corresponding to the second voiceprint feature information can be generated, and the generated user identifier is taken as a user identifier of a third user; then the user identifier of the third user is stored in the voiceprint feature information set corresponding to the second voiceprint feature information, so as to realize updating of the voiceprint feature information set.

[0124] Figure 4 is a structural schematic diagram of a voice recognition method according to the second embodiment of the present application, as shown in the figure, the end side adaptive voice recognition can be performed by the following four sub-function modules, the local model of the vehicle end side can be updated through the networking service module, the local model can include voiceprint feature information, voice recognition models of different speaker identifiers; the voiceprint recognition module is used for voiceprint feature recognition on the received target voice signal, to obtain a target user identifier corresponding to the target voice signal; the voice recognition module can perform voice recognition on the target voice signal according to the first voice recognition model of the target user identifier, to obtain instruction information corresponding to the target voice signal, so as to provide subsequent services for the target user. The storage module can store the data generated at the vehicle end side and the local model obtained from the server. Figure 4

[0125] Figure 5 is a flowchart of a voice recognition method according to the second embodiment of the present application, as shown in the figure, first, the local vehicle end side can receive real-time voice of a user (i.e. receiving a target voice signal of a first user), perform voiceprint recognition on the real-time voice, and determine the user ID by judging whether the speaker exists or not, for example, the cloud server can perform clustering by using a pre-trained voiceprint recognition model before the vehicle is offline, to obtain a group of speaker feature basis vectors as a memory unit, and issue the local model, the local model can select the most similar speaker from the memory unit in real time in the early stage, for example, if the speaker exists, the user ID of the user is directly obtained (i.e. performing voiceprint feature recognition on the target voice signal to obtain first voiceprint feature information, and judging whether the voiceprint feature information set stored locally contains the first voiceprint feature information or not; if yes, the user identifier corresponding to the first voiceprint feature information in the voiceprint feature information set is taken as the target user identifier corresponding to the target voice signal), if the speaker does not exist, the new voiceprint information can be stored in the local voiceprint information, and the baseline ID of the most similar voiceprint information is selected as the user ID of the user (i.e. if no, the preset voiceprint feature information satisfying the preset feature condition is selected from the voiceprint feature information set, and the user identifier corresponding to the preset voiceprint feature information is taken as the target user identifier corresponding to the target voice signal). Figure 5

[0126] ​​Then, according to the multi-ID domain adaptive ASR and the user ID, a user adaptive ASR is determined (i.e., determining the first speech recognition model of the first user based on the target user identification and at least one speech recognition model), so as to perform speech recognition, and use the recognized text instruction to provide services for the user by using the product application, for example, to perform vehicle control related function interaction by using voice recognition, semantic understanding (NLU), dialogue management (DM), etc., such as controlling vehicle lights, vehicle windows, air conditioners, aromatherapy, music, radios, making phone calls, navigation, and other applications (i.e., performing voice recognition on the target voice signal according to the first speech recognition model to obtain instruction information corresponding to the target voice signal, and providing services for the target user based on the instruction information).

[0127] In addition, the vehicle side can select whether to upload new voiceprint information to the cloud server, and if so, the new voiceprint information can be uploaded to the cloud server, so as to update the voiceprint feature information, and subsequently pull the updated voiceprint feature information at a regular time (i.e., sending the first voiceprint feature information to the server to update the voiceprint feature information set).

[0128] The vehicle side can select whether to upload the voice signal to the cloud server, and if so, the voice signal can be uploaded to the cloud server, and the cloud server can generate a corresponding user adaptive ASR based on the voice signal, and update the multi-ID domain adaptive ASR, so that the vehicle side can subsequently pull the updated multi-ID domain adaptive ASR at a regular time (i.e., sending the target voice signal to the server to generate or optimize the first speech recognition model of the first user based on the target voice signal). For example, if the voice is selected to be uploaded to the cloud, the cloud will update the acoustic model, language model, and speech recognition model of the speaker through adaptive training to achieve better adaptive effect in the case of connection. Specifically: using the deployed initial system, the voice in the vehicle scene is obtained, and the language model adaptive training is performed using the text label inferred by the initial system (speaker-independent general system); the adaptive training of the acoustic model part is completed by using common technologies such as SVD bottleneck layer self-training and speaker perception training i-vector.

[0129] It can be seen that the speech recognition method provided by the embodiment can realize real-time and high-accuracy speech recognition of the vehicle side under restricted resources through the adaptive strategy of the vehicle domain language model and the speaker acoustic model, and can realize the quasi-real-time updating and distribution of the model through the end-to-cloud collaboration. Therefore, the use of the cloud environment under restricted conditions can be avoided, the intelligent voice interaction of the end side is effectively guaranteed, and the user privacy is protected and the user experience is improved.

[0130] Embodiment Three

[0131] Figure 6 is a structural schematic diagram of a voice recognition device provided according to Embodiment Three of the present application. As shown in the figure, the device comprises: Figure 6

[0132] a voiceprint recognition module 310, configured to receive a target voice signal of a first user, and perform voiceprint feature recognition on the target voice signal to obtain a target user identifier corresponding to the target voice signal;

[0133] a first determination module 320, configured to determine a first voice recognition model of the first user based on the target user identifier and at least one voice recognition model, the at least one voice recognition model being generated by a server based on voice signals sent by the vehicle-mounted terminal and returned to the vehicle-mounted terminal;

[0134] a voice recognition module 330, configured to perform voice recognition on the target voice signal according to the first voice recognition model to obtain instruction information corresponding to the target voice signal, and provide services for the first user based on the instruction information.

[0135] The voice recognition device provided in Embodiment Three of the present application receives a target voice signal of a first user through a voiceprint recognition module, and performs voiceprint feature recognition on the target voice signal to obtain a target user identifier corresponding to the target voice signal. A first determination module determines a first voice recognition model of the first user based on the target user identifier and at least one voice recognition model, the at least one voice recognition model being generated by a server based on voice signals sent by the vehicle-mounted terminal and returned to the vehicle-mounted terminal. A voice recognition module performs voice recognition on the target voice signal according to the first voice recognition model to obtain instruction information corresponding to the target voice signal, and provides services for the first user based on the instruction information. With the device, the first voice recognition model is determined based on at least one voice recognition model on the vehicle side, and voice recognition is performed, which can meet the user privacy protection requirement. Meanwhile, the at least one voice recognition model required is returned to the vehicle-mounted terminal by the server at regular time intervals, rather than being acquired in real time, which further improves the real-time performance of voice recognition on the basis of meeting the user privacy protection requirement, thereby improving the user experience.

[0136] Optionally, the voiceprint recognition module 310 is specifically configured to:

[0137] perform voiceprint feature recognition on the target voice signal to obtain first voiceprint feature information, and determine whether the first voiceprint feature information is contained in a voiceprint feature information set stored locally, the voiceprint feature information set being obtained from the server at regular time intervals, and the voiceprint feature information set being used to store voiceprint feature information corresponding to each user identifier.​

[0138] If yes, a user identifier corresponding to the first voiceprint feature information in the voiceprint feature information set is taken as a target user identifier corresponding to the target voice signal;

[0139] If no, preset voiceprint feature information satisfying a preset feature condition is selected from the voiceprint feature information set, and a user identifier corresponding to the preset voiceprint feature information is taken as a target user identifier corresponding to the target voice signal.

[0140] Optionally, the voice recognition device provided in Embodiment Three of the present application further comprises:

[0141] a first sending module configured to send the first voiceprint feature information to the server after the preset voiceprint feature information satisfying the preset feature condition is selected from the voiceprint feature information set, so that the server updates the voiceprint feature information set.

[0142]

[0143] Optionally, the voice recognition device provided in Embodiment Three of the present application further comprises:

[0144] a second sending module configured to send the target voice signal to the server, so that the server generates or optimizes the first voice recognition model of the first user based on the target voice signal.

[0145] The voice recognition device provided in the embodiments of the present application can execute the voice recognition method provided in Embodiment One of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0146] Embodiment Four

[0147] Figure 7 is a structural schematic diagram of a voice recognition device provided in Embodiment Four of the present application. As shown in the figure, the device comprises: Figure 7

[0148] a signal receiving module 410 configured to receive a voice signal of a second user, the voice signal being sent by a vehicle terminal;

[0149] a second determining module 420 configured to determine a vehicle language model and a vehicle acoustic model based on the voice signal, the vehicle language model corresponding to the vehicle terminal, and the vehicle acoustic model corresponding to the second user;

[0150] a processing module 430 configured to process the vehicle language model and the vehicle acoustic model to obtain a second voice recognition model corresponding to the second user;

[0151] ​​The returning module 440 is configured to return at least one speech recognition model to the vehicle terminal, and the at least one speech recognition model includes the second speech recognition model.

[0152] The speech recognition device provided in the fourth embodiment of the present application receives a speech signal of a second user through a signal receiving module, wherein the speech signal is sent by a vehicle terminal; determines a vehicle language model and a vehicle acoustic model based on the speech signal through a second determining module, wherein the vehicle language model corresponds to the vehicle terminal, and the vehicle acoustic model corresponds to the second user; processes the vehicle language model and the vehicle acoustic model through a processing module to obtain a second speech recognition model corresponding to the second user; and returns at least one speech recognition model to the vehicle terminal through a returning module, wherein the at least one speech recognition model includes the second speech recognition model. By using the device, the second speech recognition model of the second user can be determined by receiving the speech signal of the second user, so that at least one speech recognition model containing the second speech recognition model is returned to the vehicle terminal, the requirement of speech recognition at the vehicle terminal is met, the real-time performance of speech recognition is improved, and the user experience is improved.

[0153] Optionally, the second determining module 420 is specifically configured to:

[0154] If there is no second speech recognition model corresponding to the second user, the vehicle language model and the vehicle acoustic model are generated based on the speech signal.

[0155] Optionally, the second determining module 420 is specifically configured to:

[0156] If there is a second speech recognition model corresponding to the second user, the vehicle language model and the vehicle acoustic model are optimized based on the speech signal.

[0157] The processing module 430 is specifically configured to:

[0158] The optimized vehicle language model and the optimized vehicle acoustic model are processed to obtain a new speech recognition model corresponding to the second user, and the new speech recognition model is used to replace the second speech recognition model.

[0159] Optionally, the second determining module 420 is specifically configured to:

[0160] The vehicle language model is determined according to a general topic model and a vehicle topic model; and

[0161] The original model is trained based on the signal attribute of the speech signal to obtain the vehicle acoustic model.

[0162] Optionally, the speech recognition device provided in the third embodiment of the present application further comprises:

[0163] The information receiving module is configured to receive second voiceprint feature information of a third user, the second voiceprint feature information being sent by the vehicle-mounted terminal.

[0164] The updating module is configured to update the voiceprint feature information set according to the second voiceprint feature information.

[0165] The third sending module is configured to send the updated voiceprint feature information set to the vehicle-mounted terminal.

[0166] Optionally, the updating module is specifically configured to:

[0167] generate a user identifier corresponding to the second voiceprint feature information as the user identifier of the third user;

[0168] store the user identifier of the third user and the second voiceprint feature information in the voiceprint feature information set.

[0169] The speech recognition device provided by the embodiment of the present application can execute the speech recognition method provided by the second embodiment of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0170] Embodiment five

[0171] Figure 8 is a structural schematic diagram of a vehicle-mounted terminal according to the fifth embodiment of the present application. The vehicle-mounted terminal is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The vehicle-mounted terminal can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are merely examples, and are not intended to limit the implementation of the present application described and / or claimed herein.

[0172] As Figure 8As shown, the in-vehicle terminal 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, and the like, which is communicatively connected to the at least one processor 11. The memory stores a computer program that is executable by the at least one processor 11, and the processor 11 can perform various appropriate actions and processes in accordance with the computer program stored in the read-only memory (ROM) 12 or loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the in-vehicle terminal 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0173] Various components in the in-vehicle terminal 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, and the like, an output unit 17, such as various types of displays, a speaker, and the like, a storage unit 18, such as a magnetic disk, an optical disk, and the like, and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 19 allows the in-vehicle terminal 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0174] The processor 11 can be various general and / or special purpose processing components having processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The processor 11 performs various methods and processes described above, such as the speech recognition method.

[0175] In some embodiments, the speech recognition method can be implemented as a computer program that is tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the in-vehicle terminal 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the speech recognition method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the speech recognition method by any other appropriate means, such as by means of firmware.

[0176] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0177] Computer programs used to implement the processes of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program, when executed, can cause instructions defined in the flow charts and / or block diagrams to be implemented. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package and partially on a remote machine or entirely on a remote machine or server.

[0178] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store computer programs for use by or in connection with an instruction execution system, apparatus, or device. Computer-readable storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0179] To provide for interaction with a user, the systems and techniques described here can be implemented on a vehicle terminal having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the vehicle terminal. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0180] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0181] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0182] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, as long as the desired results of the present disclosure are achieved, and the present disclosure is not limited herein.

[0183] The above detailed description does not limit the scope of the present disclosure. It is understood that various modifications, combinations, sub-combinations, and alternatives can be made to the detailed disclosure without departing from the spirit and principles of the present disclosure. Any modifications, equivalent substitutions, improvements, and the like that are made within the spirit and principles of the present disclosure are included in the scope of the present disclosure.

[0184] Embodiment Six

[0185] Figure 9 is a structural diagram of a server provided according to Embodiment Six of the present application. The server is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The server can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.

[0186] As shown in Figure 9 The server 20 includes at least one processor 21, and a memory, such as a read-only memory (ROM) 22, a random access memory (RAM) 23, etc., connected to the at least one processor 21 in communication, where the memory stores computer programs executable by the at least one processor. The processor 21 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 22 or loaded into the random access memory (RAM) 23 from the storage unit 28. In the RAM 23, various programs and data required for the operation of the server 20 can also be stored. The processor 21, the ROM 22, and the RAM 23 are connected to each other through a bus 24. An input / output (I / O) interface 25 is also connected to the bus 24.

[0187] Various components in the server 20 are connected to the I / O interface 25, including: an input unit 26, such as a keyboard, a mouse, etc.; an output unit 27, such as various types of displays, speakers, etc.; a storage unit 28, such as a magnetic disk, an optical disk, etc.; and a communication unit 29, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 29 allows the server 20 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.

[0188] The processor 21 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 21 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 21 performs various methods and processes described above, such as the speech recognition method.

[0189] In some embodiments, the speech recognition method can be implemented as a computer program tangibly embodied in a computer readable storage medium, e.g., storage unit 28. In some embodiments, portions or all of the computer program can be loaded and / or installed onto server 20 via, e.g., ROM 22 and / or communication unit 29. When the computer program is loaded onto RAM 23 and executed by processor 21, one or more steps of the speech recognition method described above can be performed. Alternatively, in other embodiments, processor 21 can be configured to perform the speech recognition method by other means, e.g., with the aid of firmware.

[0190] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, specially designed application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0191] Computer programs used to implement the methods of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor of the machine, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0192] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0193] To provide for interaction with a user, the systems and techniques described here can be implemented on a server having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the server. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0194] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0195] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0196] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present application can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which is not limited herein.

[0197] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A voice recognition method, characterized by, The method is applied to a vehicle terminal and comprises the following steps: receiving a target voice signal of a first user and performing voiceprint feature recognition on the target voice signal to obtain a target user identifier corresponding to the target voice signal; determining a first voice recognition model of the first user based on the target user identifier and at least one voice recognition model, wherein the at least one voice recognition model is generated by a server based on a voice signal sent by the vehicle terminal and returned to the vehicle terminal; performing voice recognition on the target voice signal according to the first voice recognition model to obtain instruction information corresponding to the target voice signal, and providing a service for the first user based on the instruction information; wherein the voiceprint feature recognition on the target voice signal to obtain the target user identifier corresponding to the target voice signal comprises the following steps: performing voiceprint feature recognition on the target voice signal to obtain first voiceprint feature information, and determining whether the first voiceprint feature information is included in a set of voiceprint feature information stored locally, wherein the set of voiceprint feature information is obtained from the server at regular intervals, and the set of voiceprint feature information is used to store voiceprint feature information corresponding to each user identifier; if yes, taking a user identifier corresponding to the first voiceprint feature information in the set of voiceprint feature information as the target user identifier corresponding to the target voice signal; if no, selecting preset voiceprint feature information that meets a preset feature condition from the set of voiceprint feature information, sending the first voiceprint feature information to the server to enable the server to update the set of voiceprint feature information, and taking a user identifier corresponding to the preset voiceprint feature information as the target user identifier corresponding to the target voice signal, wherein the preset feature condition is that the similarity to the first voiceprint feature information is greater than a set threshold.

2. The method of claim 1, further comprising: sending the target voice signal to the server to enable the server to generate or optimize the first voice recognition model of the first user based on the target voice signal.

3. A voice recognition method characterized by, The method is applied to a server and comprises the following steps: receiving a voice signal of a second user, wherein the voice signal is sent by a vehicle terminal; determining a vehicle language model and a vehicle acoustic model based on the voice signal, wherein the vehicle language model corresponds to the vehicle terminal, and the vehicle acoustic model corresponds to the second user; processing the vehicle language model and the vehicle acoustic model to obtain a second voice recognition model corresponding to the second user; returning at least one voice recognition model to the vehicle terminal, wherein the at least one voice recognition model includes the second voice recognition model; wherein the determination of the vehicle language model and the vehicle acoustic model based on the voice signal comprises the following steps: if there is no second voice recognition model corresponding to the second user, generating a vehicle language model and a vehicle acoustic model based on the voice signal; if there is a second voice recognition model corresponding to the second user, optimizing a vehicle language model and a vehicle acoustic model based on the voice signal; The processing of the vehicle-mounted language model and the vehicle-mounted acoustic model obtains a second speech recognition model corresponding to the second user, and the second speech recognition model is obtained by processing the optimized vehicle-mounted language model and the optimized vehicle-mounted acoustic model. The processing of the vehicle-mounted language model and the vehicle-mounted acoustic model obtains a second speech recognition model corresponding to the second user, and the second speech recognition model is obtained by processing the optimized vehicle-mounted language model and the optimized vehicle-mounted acoustic model.

4. The method of claim 3, wherein, The vehicle-mounted language model and the vehicle-mounted acoustic model are determined based on the speech signal, and the vehicle-mounted language model is determined based on the general topic model and the vehicle-mounted topic model. The vehicle-mounted language model is determined based on the general topic model and the vehicle-mounted topic model. The vehicle-mounted acoustic model is obtained by training the original model based on the signal attribute of the speech signal.

5. The method of claim 3, wherein, Further comprising: Receiving second voiceprint feature information of a third user, the second voiceprint feature information being sent by a vehicle-mounted terminal; Updating the voiceprint feature information set according to the second voiceprint feature information; Sending the updated voiceprint feature information set to the vehicle-mounted terminal.

6. The method of claim 5, wherein, The updating of the voiceprint feature information set according to the second voiceprint feature information comprises: Generating a user identifier corresponding to the second voiceprint feature information as a user identifier of the third user; Storing the user identifier of the third user and the second voiceprint feature information in the voiceprint feature information set.

7. A speech recognition apparatus characterized by comprising: The device is configured in a vehicle-mounted terminal, and the device comprises: A voiceprint recognition module configured to receive a target speech signal of a first user and perform voiceprint feature recognition on the target speech signal to obtain a target user identifier corresponding to the target speech signal; A first determination module configured to determine a first speech recognition model of the first user based on the target user identifier and at least one speech recognition model, the at least one speech recognition model being generated by a server based on a speech signal sent by the vehicle-mounted terminal and returned to the vehicle-mounted terminal; A speech recognition module configured to perform speech recognition on the target speech signal based on the first speech recognition model to obtain instruction information corresponding to the target speech signal, and provide services for the first user based on the instruction information; The voiceprint recognition module is specifically configured to: perform voiceprint feature recognition on the target speech signal to obtain first voiceprint feature information, and determine whether the first voiceprint feature information is included in a voiceprint feature information set stored locally, the voiceprint feature information set being obtained from the server at regular intervals, and the voiceprint feature information set being used to store voiceprint feature information corresponding to each user identifier; if yes, a user identifier corresponding to the first voiceprint feature information in the voiceprint feature information set is taken as a target user identifier corresponding to the target speech signal; if no, a preset voiceprint feature information satisfying a preset feature condition is selected from the voiceprint feature information set, the first voiceprint feature information is sent to the server, so that the server updates the voiceprint feature information set, and a user identifier corresponding to the preset voiceprint feature information is taken as a target user identifier corresponding to the target speech signal, wherein the preset feature condition is that the similarity to the first voiceprint feature information is greater than a set threshold.

8. A speech recognition apparatus characterized by comprising: The device is configured in a server, and the device comprises: The signal receiving module is configured to receive a voice signal of a second user, the voice signal being sent by the vehicle-mounted terminal. The second determining module is configured to determine a vehicle-mounted language model and a vehicle-mounted acoustic model based on the voice signal, the vehicle-mounted language model corresponding to the vehicle-mounted terminal, and the vehicle-mounted acoustic model corresponding to the second user. The processing module is configured to process the vehicle-mounted language model and the vehicle-mounted acoustic model to obtain a second voice recognition model corresponding to the second user. The returning module is configured to return at least one voice recognition model to the vehicle-mounted terminal, the at least one voice recognition model including the second voice recognition model. The second determining module is specifically configured to: if there is no second voice recognition model corresponding to the second user, generate a vehicle-mounted language model and a vehicle-mounted acoustic model based on the voice signal; and if there is a second voice recognition model corresponding to the second user, optimize a vehicle-mounted language model and a vehicle-mounted acoustic model based on the voice signal. The processing module is specifically configured to: process the optimized vehicle-mounted language model and the optimized vehicle-mounted acoustic model to obtain a new voice recognition model corresponding to the second user, and replace the second voice recognition model with the new voice recognition model.

9. A vehicle terminal, characterized by The vehicle-mounted terminal comprises: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the voice recognition method in any one of claims 1-2.

10. A server, characterized by The server comprises: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the voice recognition method in any one of claims 3-6.

11. A computer readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the voice recognition method in any one of claims 1-6 when executed. The computer readable storage medium stores computer instructions for enabling the processor to implement the voice recognition method in any one of claims 1-6 when executed.

Citation Information

Patent Citations

  • Voice processing method and device for user personalized service

    CN112185362A