Vehicle speech recognition method, device and equipment and storage medium

By using a pre-trained voiceprint model to recognize vehicle voice in a noisy environment, the problem of inaccurate identification in the prior art is solved, and higher recognition accuracy and robustness are achieved, improving user experience and vehicle operation safety.

CN119964577APending Publication Date: 2025-05-09BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510239230.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to accurately recognize vehicle voice in a noisy environment, which affects the accuracy and robustness of vehicle voice recognition.

Method used

A pre-trained voiceprint model is used to generate a target voiceprint vector corresponding to the target voice, and a target operation is performed within a preset threshold range in response to the similarity of the target voiceprint vector and the vehicle registered voiceprint vector.

Benefits of technology

It improves the accuracy and robustness of vehicle voice recognition in noisy environments, and enhances the user's driving experience and vehicle operation safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964577A_ABST
    Figure CN119964577A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle speech recognition method, and relates to the technical field of artificial intelligence, in particular to the technical fields of automatic driving, intelligent traffic, speech recognition, deep learning and the like. The method comprises the steps that a pre-trained voiceprint model is utilized to generate a target voiceprint vector corresponding to target voice, and the voiceprint model is obtained by training voice sample pairs in different noise environments; calculating the similarity between the target voiceprint vector and a registered voiceprint vector corresponding to the vehicle; and in response to the fact that the similarity is determined to be within the preset similarity threshold range, executing a target operation corresponding to the target voice. According to the method, the accuracy and robustness of vehicle speech recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as autonomous driving, intelligent transportation, speech recognition and deep learning, and more particularly to a vehicle speech recognition method, device, equipment and storage medium. Background Art

[0002] With the development of vehicle technology, vehicles have become an indispensable part of people's lives. As people's requirements for vehicles become higher and higher, people also hope to have a good driving experience while ensuring safe driving. Voice interaction technology, as a convenient and fast technical means, can greatly improve driving safety and has been widely used in automobiles. Users can interact with intelligent voice devices through voice and control intelligent voice devices to perform corresponding operations, such as unlocking car doors, logging in to accounts, etc. Summary of the invention

[0003] The present disclosure provides a vehicle speech recognition method, device, equipment and storage medium.

[0004] According to a first aspect of the present disclosure, a vehicle speech recognition method is provided, comprising: generating a target voiceprint vector corresponding to a target speech using a pre-trained voiceprint model, wherein the voiceprint model is trained using speech sample pairs in different noise environments; calculating the similarity between the target voiceprint vector and a registered voiceprint vector corresponding to the vehicle; and in response to determining that the similarity is within a preset similarity threshold range, executing a target operation corresponding to the target speech.

[0005] According to a second aspect of the present disclosure, a method for training a voiceprint model is provided, comprising: obtaining a training sample set, wherein the training sample pairs in the training sample set include positive sample speech and negative sample speech, the positive sample speech is the speech of a user in a quiet environment, and the negative sample speech is the speech of a user in a noisy environment; using the training sample set to train an initial neural network to obtain a voiceprint model.

[0006] According to a third aspect of the present disclosure, a vehicle voice recognition device is provided, comprising: a generation module, configured to generate a target voiceprint vector corresponding to a target voice using a pre-trained voiceprint model, wherein the voiceprint model is trained using voice sample pairs under different noise environments; a calculation module, configured to calculate the similarity between the target voiceprint vector and a registered voiceprint vector corresponding to the vehicle; and an execution module, configured to execute a target operation corresponding to the target voice in response to determining that the similarity is within a preset similarity threshold range.

[0007] According to a fourth aspect of the present disclosure, a vehicle speech recognition device is provided, comprising: an acquisition module, configured to acquire a training sample set, wherein the training sample pairs in the training sample set include positive sample speech and negative sample speech, the positive sample speech is the speech of a user in a quiet environment, and the negative sample speech is the speech of a user in a noisy environment; a training module, configured to train an initial neural network using the training sample set to obtain a voiceprint model.

[0008] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any implementation manner of the first aspect or the second aspect.

[0009] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method described in any implementation manner of the first aspect or the second aspect.

[0010] According to a seventh aspect of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the computer program implements the method described in any implementation manner of the first aspect or the second aspect.

[0011] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0013] Figure 1 is an exemplary system architecture diagram in which the present disclosure may be applied;

[0014] Figure 2 is a flow chart of an embodiment of a vehicle speech recognition method according to the present disclosure;

[0015] Figure 3 is a flow chart of another embodiment of a vehicle speech recognition method according to the present disclosure;

[0016] Figure 4 is a flowchart of the process of generating registered vectors;

[0017] Figure 5 is a flow chart of another embodiment of a vehicle speech recognition method according to the present disclosure;

[0018] Figure 6 is a flow chart of an embodiment of a method for training a voiceprint model according to the present disclosure;

[0019] Figure 7-1 It is an application flow chart of the training process of the voiceprint model;

[0020] Figure 7-2 It is an application flow chart of the voiceprint adaptive verification process;

[0021] Figure 8 is a structural schematic diagram of an embodiment of a vehicle speech recognition device according to the present disclosure;

[0022] Fig. 9 is a structural schematic diagram of an embodiment of a training device for a voiceprint model according to the present disclosure;

[0023] Fig.10 It is a block diagram of an electronic device used to implement the vehicle voice recognition method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0024] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0025] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0026] Figure 1 An exemplary system architecture 100 is shown to which an embodiment of the vehicle speech recognition method, voiceprint model training method or vehicle speech recognition device, voiceprint model training device of the present disclosure can be applied.

[0027] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, 103, 104, a network 105, and a server 106. The network 105 is used to provide a medium for communication links between the terminal devices 101, 102, 103, 104 and the server 106. The network 105 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.

[0028] The user can use the terminal devices 101, 102, 103, 104 to interact with the server 106 via the network 105 to receive or send information, etc. Various client applications can be installed on the terminal devices 101, 102, 103, 104.

[0029] Terminal devices 101, 102, 103, 104 may be hardware or software. When terminal devices 101, 102, 103, 104 are hardware, they may be various electronic devices, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, etc. When terminal devices 101, 102, 103, 104 are software, they may be installed in the above electronic devices. They may be implemented as multiple software or software modules, or as a single software or software module. No specific limitation is made here.

[0030] The server 106 can provide various services. For example, the server 106 can analyze and process the target speech acquired from the terminal devices 101, 102, 103, 104, and generate a processing result (eg, a target operation).

[0031] It should be noted that the server 106 can be hardware or software. When the server 106 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server 106 is software, it can be implemented as multiple software or software modules (for example, for providing distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0032] It should be noted that the vehicle speech recognition method and the voiceprint model training method provided in the embodiments of the present disclosure are generally executed by the server 106 . Accordingly, the vehicle speech recognition device and the voiceprint model training device are generally arranged in the server 106 .

[0033] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to the implementation requirements.

[0034] Continue to refer Figure 2 , which shows a process 200 of an embodiment of a vehicle speech recognition method according to the present disclosure. The vehicle speech recognition method comprises the following steps:

[0035] Step 201: Generate a target voiceprint vector corresponding to the target speech using a pre-trained voiceprint model.

[0036] In this embodiment, the execution subject of the vehicle speech recognition method (for example Figure 1The server 105 shown in the figure will first obtain the target voice, which is the voice issued by the user for operating the vehicle. For example, when the user is outside or inside the vehicle, the user wants to unlock the vehicle, then the target voice is the voice issued by the user for unlocking the vehicle; for another example, when the user is inside the vehicle, the user wants to operate the vehicle equipment (such as logging in), then the target voice is the voice issued by the user for operating the vehicle equipment.

[0037] Then, the execution subject will use the pre-trained voiceprint model to generate the voiceprint vector corresponding to the target voice, that is, the target voiceprint vector. Here, the execution subject will first collect a large amount of voiceprint audio data and train the voiceprint model through a convolutional neural network. The voiceprint model will be used offline to generate the voiceprint vector of the user's voice. The execution subject will also set a default threshold for the voiceprint model, which is the vector similarity threshold, that is, the similarity threshold between the voiceprint vector generated by the voiceprint model and the voiceprint vector pre-registered by the user.

[0038] When training the voiceprint model, for each user, the corpus will be recorded in different noise environments, such as a noise-free environment, that is, a quiet environment (such as a quiet car environment, a quiet office environment), and a noisy environment, such as an environment with external noise outside the car (the user is outside the car and there is noise outside), an environment with internal noise inside the car (the user is inside the car and there is noise inside the car), an environment with internal and external noise inside the car (the user is inside the car and there is noise outside the car), and an environment with internal noise outside the car (the user is outside the car and there is noise inside the car), etc. The recorded voice contains the vehicle voice wake-up words (such as unlocking, logging in, etc.). Then the user's voice in a quiet environment (no noise) is used as a positive sample, and the user's voice in a noisy environment (such as external noise outside the car, internal and external noise inside the car, internal noise inside the car, internal noise outside the car, etc.) is used as a negative sample, and the initial convolutional neural network is trained using a comparative learning method to obtain a trained voiceprint model.

[0039] Step 202: Calculate the similarity between the target voiceprint vector and the registered voiceprint vector corresponding to the vehicle.

[0040] In this embodiment, the above-mentioned execution entity will determine the vehicle that the user's target voice wants to operate. For example, the vehicle that the user wants to operate can be determined based on the distance between the user and the vehicle. Then the above-mentioned execution entity will obtain the registered voiceprint vector corresponding to the vehicle and calculate the similarity between the target voiceprint vector and the registered voiceprint vector.

[0041] Here, the user's audio will be recorded for each vehicle in advance. In order to improve practicality, each vehicle is only allowed to record a voiceprint for one user. Of course, voiceprints can also be recorded for multiple users according to the situation. This embodiment does not specifically limit this. After recording the user's audio, the voiceprint model will be used to generate a voiceprint vector corresponding to the audio, and the voiceprint vector will be bound to the vehicle to generate a registered voiceprint vector corresponding to the vehicle.

[0042] Step 203: In response to determining that the similarity is within a preset similarity threshold range, executing a target operation corresponding to the target voice.

[0043] In this embodiment, the above-mentioned execution subject will determine whether the similarity between the target voiceprint vector and the registered voiceprint vector is within a preset threshold range. If it is determined that the similarity is within the preset similarity threshold range, it is proved that the target voice has passed the verification. After that, the target voice can be recognized and analyzed to determine the operation corresponding to the target voice and execute the operation. For example, assuming that the target voice is "unlock the car door", it can be determined that the target operation is unlocking the car door. At this time, the above-mentioned execution subject will unlock the car door.

[0044] The vehicle voice recognition method provided by the embodiment of the present disclosure first generates a target voiceprint vector corresponding to the target voice using a pre-trained voiceprint model; then calculates the similarity between the target voiceprint vector and the registered voiceprint vector corresponding to the vehicle; and finally, in response to determining that the similarity is within a preset similarity threshold range, executes the target operation corresponding to the target voice. The vehicle voice recognition method in this embodiment, the voiceprint model in the method can accurately recognize voices in various noise environments, thereby improving the accuracy and robustness of vehicle voiceprint recognition.

[0045] In addition, in the technical solutions involved in this disclosure, the acquisition, storage, use, processing, transportation, provision and disclosure of user personal information (such as user voice involved in this disclosure later) are in compliance with the relevant laws and regulations and do not violate public order and good morals.

[0046] Continue to refer Figure 3 , Figure 3 A process 300 of another embodiment of a vehicle speech recognition method according to the present disclosure is shown. The vehicle speech recognition method comprises the following steps:

[0047] Step 301: Perform noise filtering on the initial speech to obtain the target speech.

[0048] In this embodiment, the execution subject of the vehicle speech recognition method (for example Figure 1The server 105 shown in the figure will perform noise filtering preprocessing on the initial voice. The noise filtering here generally filters out slight noises that are close to the user, such as slight noises generated when the vehicle-mounted equipment is running when the user is in the car, or the sounds made by other members in the vehicle. However, it is impossible to filter out complex noises in the background. After certain noise reduction preprocessing, the interference of surrounding noise can be reduced. The noise reduction processing method can be implemented using existing technologies and is not specifically limited here.

[0049] Step 302: Perform speech recognition on the target speech to determine the speech content corresponding to the target speech.

[0050] In this embodiment, the above-mentioned execution subject will perform speech recognition on the target speech, thereby determining the speech content corresponding to the target speech. For example, the target speech is recognized and the corresponding speech content is determined to be "unlock the car door". The speech recognition method can be implemented using existing technologies and is not specifically limited here.

[0051] Step 303: Match the voice content with the preset wake-up word.

[0052] In this embodiment, the execution subject will match the voice content with the preset wake-up word. Here, multiple wake-up words for vehicle operations are preset, such as "unlock", "login", etc. The execution subject will match the voice content with the preset wake-up word to determine whether the voice content contains the preset wake-up word. If it does, step 304 is executed; if it does not, a prompt voice is generated, such as "Sorry, I don't understand", to prompt the user to re-enter the voice.

[0053] First determine whether the voice contains the preset wake-up word, so as to determine whether to carry out the subsequent recognition process, thereby improving the efficiency of voice recognition.

[0054] Step 304, in response to determining that the voice content successfully matches the preset wake-up word, the target voice is input into the voiceprint model, and the target voiceprint vector is output.

[0055] In this embodiment, when the execution subject determines that the voice content matches the preset wake-up word successfully, it will input the target voice into the voiceprint model and output the target voiceprint vector corresponding to the target voice, thereby accurately generating the vector corresponding to the target voice using the voiceprint model.

[0056] Step 305: Calculate the similarity between the target voiceprint vector and the registered voiceprint vector corresponding to the vehicle.

[0057] Step 305 is basically the same as step 202 of the aforementioned embodiment. The specific implementation method can refer to the aforementioned description of step 202, which will not be repeated here.

[0058] In some optional implementations of this embodiment, the registered voiceprint vector is obtained through the following steps: initializing the number of successful registrations to 0, and performing the following iterative process: recording the user's voice in different noise environments to obtain the current voice; in response to determining that the current voice is recorded successfully, registering the current voice into the voiceprint model; in response to determining that the current voice has been successfully registered into the voiceprint model, adding 1 to the number of successful registrations; in response to determining that the number of successful registrations is less than a registration number threshold, performing the next round of iterative process; in response to determining that the number of successful registrations is greater than or equal to a registration number threshold, using the voiceprint model to generate a voiceprint vector corresponding to the successfully registered voice, and obtaining a registered voiceprint vector corresponding to the user.

[0059] For reference Figure 4 , Figure 4 A flowchart of the process of generating a registered vector is shown, which specifically includes the following steps:

[0060] Step 401: Initialize the number of successful registrations N to 0.

[0061] Step 402: Start recording the user's voice in different noise environments.

[0062] During the first recording and registration, the user's voice in a quiet environment (including the specified wake-up word) will be recorded first to obtain the recorded voice.

[0063] Step 403, determine whether the recording is successful.

[0064] Determine whether the recording is successful. If so, execute step 404. Otherwise, return to step 402 and re-record.

[0065] Step 404: register the recorded speech into the voiceprint model.

[0066] Directly input the recorded speech into the voiceprint model for registration.

[0067] Step 405, determine whether the registration is successful.

[0068] If so, the number of successful registrations N is increased by 1, and step 406 is executed; otherwise, the process returns to step 402 and re-records.

[0069] Step 406, determine whether N is less than the registration times threshold 3.

[0070] If yes, then go back to step 402 and re-record; otherwise, go to step 407. The number of registrations here can be set to other values ​​according to actual needs, and this embodiment does not make any specific limitation to this.

[0071] Here, when the user successfully registers for the first time, the next recording will be performed; when the user records for the second time, the recorded voice will be input into the voiceprint model, and the voiceprint model will determine whether the voice and the registered voice correspond to the same user. If so, registration will be performed; if not, registration will fail. Generally, three voices will be recorded and registered for each user.

[0072] Step 407: Generate a voiceprint vector.

[0073] After the three voices are successfully registered, the voiceprint model will generate a voiceprint vector corresponding to the user, that is, the registered voiceprint vector corresponding to the user. The registered voiceprint vector will be used in the voiceprint verification process.

[0074] That is, in this implementation, the user's voice recorded in different noise environments (noise-free, quiet environment and noisy environment) is first obtained, and then the recorded voice is registered into the voiceprint model. The voiceprint model will generate a corresponding voiceprint vector, thereby obtaining the registered voiceprint vector corresponding to the user.

[0075] Here, since the user will record multiple voices in different noise environments, when the user records for the first time, the recorded voice is directly input into the voiceprint model for registration; when the user records for the second time, the recorded voice is input into the voiceprint model, and the voiceprint model will determine whether the voice and the registered voice correspond to the same user. If so, registration is performed; if not, registration fails. Generally, three voices are recorded and registered for each user. After the three voices are successfully registered, the voiceprint model will generate the voiceprint vector corresponding to the user, that is, the registered voiceprint vector corresponding to the user. By registering the user's voice voiceprint, it can be used in the subsequent voiceprint recognition verification process.

[0076] In some optional implementations of this embodiment, the above method also includes: presetting at least one wake-up word; and recording the user's voice in different noise environments, including: recording the user's voice containing the wake-up word in different noise environments.

[0077] In this implementation, the execution subject will first preset multiple wake-up words, which are generally wake-up words for vehicle operations, such as "unlock", "login", etc. The execution subject will record the user's voice containing the preset wake-up words in different noise environments. Since the preset wake-up words are for vehicle operations, subsequent voice verification based on the voice containing the wake-up words can improve the accuracy of verification.

[0078] Step 306 , in response to determining that the similarity is within a preset similarity threshold range, determining a target device on the vehicle for collecting the target voice.

[0079] In this embodiment, the execution subject will determine whether the similarity between the target voiceprint vector and the registered voiceprint vector is within a preset threshold range. If it is determined that the similarity is within the preset similarity threshold range, the execution subject will determine the target device for collecting the target voice from all the collection devices installed on the vehicle. The collection device here is generally a microphone, and a microphone is installed on each door of the vehicle. The execution subject will determine the distance between the user who makes the target voice and each microphone, or the distance between the target voice and each microphone, so as to use the microphone with the smallest distance from the user as the target device, that is, use the target device to collect the voice made by the user.

[0080] In some optional implementations of this embodiment, step 306 includes: determining distance information between the target voice and each collection device on the vehicle; and determining a target device for collecting the target voice from each collection device according to the distance information.

[0081] In this implementation, the above-mentioned execution entity will determine whether the similarity between the target voiceprint vector and the registered voiceprint vector is within a preset threshold range. If it is determined that the similarity is within the preset similarity threshold range, the distance information between the target voice and each collection device on the vehicle will be determined. The collection device here is generally a microphone, and a microphone will be installed on each door of the vehicle. The above-mentioned execution entity will determine the distance between the user who makes the target voice and each microphone, which can also be the distance between the target voice and each microphone.

[0082] Then, the execution subject will determine the collection device closest to the user among the collection devices and determine it as the target device for collecting the target voice. The device for collecting the voice is determined based on the distance information between the voice and the vehicle, thereby achieving more accurate voice recognition.

[0083] Step 307: Execute the target operation corresponding to the target voice in the area where the target device is located.

[0084] In this embodiment, the above-mentioned execution subject will perform the target operation corresponding to the target voice on the area where the target device is located. Here, each acquisition device can be numbered in advance, for example, the microphone on the main driver's door is numbered 1, the microphone on the co-pilot door is numbered 2, the microphone on the main driver's rear seat door is numbered 3, and the microphone on the co-pilot's rear seat door is numbered 4. Assuming that the target device for collecting the target voice is determined to be the device numbered 2, then it can be determined that the area where the target device is located is the co-pilot area. Then the target voice is recognized to determine the corresponding target operation. Assuming that the target voice is "unlock the door", it can be determined that the target operation is unlocking the door. At this time, the above-mentioned execution subject will unlock the co-pilot's door.

[0085] In some optional implementations of this embodiment, the target operation includes: an unlocking operation; and step 307 includes: determining the vehicle-mounted device corresponding to the area where the target device is located, and unlocking the vehicle-mounted device.

[0086] In this implementation, the target operation includes an unlocking operation, and the execution subject can pre-number each acquisition device, for example, the microphone on the main driver's door is numbered 1, the microphone on the co-pilot door is numbered 2, the microphone on the main driver's rear seat door is numbered 3, and the microphone on the co-pilot's rear seat door is numbered 4. Assuming that the target device for collecting the target voice is determined to be the device numbered 2, then it can be determined that the area where the target device is located is the co-pilot area. Then the execution subject will unlock the co-pilot's door or unlock the vehicle display device in the co-pilot area. Thus, the unlocking operation of the vehicle is realized through voice recognition.

[0087] from Figure 3 It can be seen that Figure 2 Compared with the corresponding embodiment, the vehicle voice recognition method in this embodiment first determines whether the user's voice contains the wake-up word. If it does, the voiceprint model is used to generate a target voiceprint vector corresponding to the target voice. When the similarity between the target voiceprint vector and the registered voiceprint vector is within the threshold range, the target device is determined by the distance between the voice and each acquisition device, and finally the vehicle device corresponding to the area where the target device is located is unlocked. This improves the accuracy and robustness of vehicle voiceprint recognition and also enhances the user experience.

[0088] Continue to refer Figure 5 , Figure 5 A process 500 of another embodiment of a vehicle speech recognition method according to the present disclosure is shown. The vehicle speech recognition method comprises the following steps:

[0089] Step 501: Generate a target voiceprint vector corresponding to the target speech using a pre-trained voiceprint model.

[0090] Step 502: Calculate the similarity between the target voiceprint vector and the registered voiceprint vector corresponding to the vehicle.

[0091] Steps 501-502 are basically consistent with steps 201-202 of the aforementioned embodiment. For specific implementation methods, reference may be made to the aforementioned description of steps 201-202, which will not be repeated here.

[0092] Step 503: In response to determining that the similarity is not within the similarity threshold range, generate voice verification failure prompt information.

[0093] In this embodiment, the execution subject of the vehicle speech recognition method (for example Figure 1When the server 105 shown in the figure determines that the similarity is not within the similarity threshold range, it will generate a voice verification failure prompt message, such as "unlock failed".

[0094] Step 504: Use the target speech to update the voiceprint model to obtain an updated voiceprint model.

[0095] In this embodiment, the above-mentioned execution subject updates the voiceprint model using the target voice to obtain an updated voiceprint model. That is, the target voice that failed the verification is used as a training sample, and the voiceprint model is retrained to obtain an updated voiceprint model. Thus, the adaptive update of the voiceprint model based on the target voice is achieved.

[0096] In some optional implementations of this embodiment, step 504 further includes: using the registered voiceprint vector as a positive sample and the target voiceprint vector as a negative sample, training the voiceprint model by contrastive learning, and obtaining an updated voiceprint model.

[0097] The target voiceprint vector is used as a negative sample, and the corresponding registered voiceprint vector of the user is used as a positive sample. The voiceprint model is trained again by contrastive learning to obtain an updated voiceprint model. The voiceprint model is updated using the speech that failed the verification, thereby improving the recognition accuracy of the updated voiceprint model.

[0098] Step 505: Generate a new target voiceprint vector corresponding to the target speech using the updated voiceprint model.

[0099] In this embodiment, the above-mentioned execution subject will use the updated voiceprint model to generate a new voiceprint vector corresponding to the target speech, that is, a new target voiceprint vector.

[0100] Step 506: Generate a target threshold range according to the similarity between the new target voiceprint vector and the registered voiceprint vector.

[0101] In this embodiment, the execution entity calculates the similarity between the new target voiceprint vector and the registered voiceprint vector, and generates a new similarity threshold range, namely, the target threshold range, according to the similarity.

[0102] Step 507: adaptively adjust the similarity threshold range of the updated voiceprint model to the target threshold range.

[0103] In this embodiment, the above-mentioned execution subject will use the target threshold range to update the original familiarity threshold range, that is, the similarity threshold range of the updated voiceprint model is adaptively adjusted to the target threshold range. Thus, the similarity threshold range of the voiceprint model is adaptively adjusted by the similarity between the target voiceprint vector corresponding to the speech that failed the verification and its corresponding registered voiceprint vector, and the adaptive threshold range adjustment also improves the robustness of the voiceprint model.

[0104] from Figure 5 It can be seen that Figure 3 Compared with the corresponding embodiments, the vehicle speech recognition method in this embodiment uses the target speech to adaptively adjust the threshold of the voiceprint model when the similarity is not within the similarity threshold range, so that the similarity threshold range of the voiceprint model can be adapted according to different noise environments, so that voiceprint recognition can be performed more effectively and accurately in various complex environments, and has better robustness.

[0105] Continue to refer Figure 6 , which shows a process 600 of an embodiment of a method for training a voiceprint model according to the present disclosure. The method for training a voiceprint model includes the following steps:

[0106] Step 601: Obtain a training sample set.

[0107] In this embodiment, the execution subject of the vehicle speech recognition method (for example Figure 1 The server 105 shown in the figure will obtain a training sample set, wherein the training sample pairs in the training sample set include positive sample speech and negative sample speech, the positive sample speech is the speech of the user in a quiet environment, and the negative sample speech is the speech of the user in a noisy environment. That is, the training sample set includes multiple training sample pairs, each training sample pair includes positive sample speech and negative sample speech, the positive sample is the speech of the user in a quiet environment (such as a quiet office environment, a quiet car environment), and the negative sample is the speech of the user in a noisy environment (such as internal noise in the car, internal and external noise in the car, external noise outside the car, internal noise outside the car, etc.). That is, the positive sample speech and negative sample speech in each sample pair correspond to the same user, and different sample pairs correspond to different users, so the training sample set includes the speech of different users in different environments.

[0108] Step 602: Use the training sample set to train the initial neural network to obtain a voiceprint model.

[0109] In this embodiment, the execution subject uses a large number of training sample pairs to train the initial neural network, thereby obtaining a trained voiceprint model. Specifically, the execution subject uses a large number of training sample pairs and a comparative learning method to train the initial neural network, thereby obtaining a trained voiceprint model.

[0110] In some optional implementations of this embodiment, the method further includes: setting a similarity threshold range corresponding to the similarity between different voices for the voiceprint model.

[0111] In this implementation, the above-mentioned execution entity will set a similarity threshold range corresponding to the similarity between different voices for the voiceprint model. The similarity threshold range is generally a default numerical range. Therefore, when performing voice verification, it can be determined whether the voice has passed the verification by judging whether the similarity between the voiceprint vector of the target voice and the registered voiceprint vector is within the similarity threshold range.

[0112] The model is trained through user voice samples in different noise environments, so that the trained voiceprint model can accurately and quickly recognize the user's voice, improving the accuracy of speech recognition.

[0113] In a specific application of the present disclosure, the training process of the voiceprint model and the voiceprint adaptive verification process are shown. Figure 7-1 , Figure 7-1 The application flow of the voiceprint model training process is shown, which specifically includes:

[0114] Step 7011: Obtain positive samples of speech voiceprint audio and negative samples of speech voiceprint audio.

[0115] The above-mentioned execution entity will first obtain a training sample set, which contains a large number of training sample pairs. Each training sample pair contains a positive speech voiceprint audio sample and a negative speech voiceprint audio sample. The positive speech voiceprint audio sample is the user's voice in a quiet environment (such as a quiet office environment, a quiet car environment), and the negative speech voiceprint audio sample is the user's voice in a noisy environment (such as interior noise in the car, interior and exterior noise in the car, exterior noise outside the car, interior noise outside the car, etc.).

[0116] Step 7012, training the convolutional neural network.

[0117] Step 7013, generate a voiceprint model.

[0118] The above-mentioned execution entity will use a large number of training samples to train the convolutional neural network, thereby obtaining a trained voiceprint model, which is used offline.

[0119] Step 7014, setting the experience threshold.

[0120] Set a default similarity threshold range for the voiceprint model. The similarity threshold range is the threshold range of the similarity between the user's voiceprint vector and the registered voiceprint vector when performing voice verification. The threshold can be generated based on the statistical analysis of a large amount of data, or it can be set by relevant staff based on actual experience.

[0121] Thus, a voiceprint model is obtained through training through the above steps, and the voiceprint model can accurately recognize speech in various noise environments, thereby improving the efficiency and accuracy of speech recognition.

[0122] Continue to refer Figure 7-2 , Figure 7-2 The application flow of the voiceprint adaptive verification process is shown, which specifically includes:

[0123] Step 7021, collect audio.

[0124] Step 7022, noise reduction preprocessing.

[0125] Collect the user's current voice and perform noise reduction on it to reduce the interference of surrounding noise.

[0126] Step 7023, offline voiceprint model.

[0127] Step 7024, voiceprint vector.

[0128] The denoised speech is input into the voiceprint model, and the voiceprint model is used to generate the corresponding voiceprint vector.

[0129] Step 7025, perform similarity comparison with the recorded registered voiceprint vector.

[0130] Calculate the similarity between the voiceprint vector and the registered voiceprint vector.

[0131] Step 7026, determine whether it is within the threshold range.

[0132] Determine whether the similarity is within a preset threshold range. If so, execute steps 7027-7028; otherwise, execute steps 7029-70212.

[0133] Step 7027, voiceprint verification is successful.

[0134] Step 7028, scene application.

[0135] If the similarity between the voiceprint vector of the current speech and the registered voiceprint vector is within the preset threshold range, the verification is successful, and the scene application is performed according to the voice content, such as unlocking the car or logging in to an account.

[0136] Step 7029, adaptive threshold strategy module.

[0137] If the similarity between the voiceprint vector of the current speech and the registered voiceprint vector is not within the preset threshold range, the verification fails, the failure result is fed back, and the current speech is sent to the adaptive threshold strategy module.

[0138] Step 70210, send the current speech as a negative sample to the neural network for training.

[0139] Step 70211, generate a new voiceprint model and a new threshold.

[0140] Step 70212, update the voiceprint model threshold.

[0141] The current voice is sent as a negative sample to the voiceprint model to generate a new voiceprint model, and a new similarity threshold range is generated based on the current voice and the registered voice. Finally, the voiceprint model and the new similarity threshold are updated for use in the next voiceprint recognition.

[0142] Further references Figure 8 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a vehicle speech recognition device, and the device embodiment is Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0143] like Figure 8 As shown, the vehicle speech recognition device 800 of this embodiment includes: a generation module 801, a calculation module 802 and an execution module 803. The generation module 801 is configured to generate a target voiceprint vector corresponding to the target speech using a pre-trained voiceprint model, wherein the voiceprint model is trained using speech sample pairs in different noise environments; the calculation module 802 is configured to calculate the similarity between the target voiceprint vector and the registered voiceprint vector corresponding to the vehicle; the execution module 803 is configured to execute the target operation corresponding to the target speech in response to determining that the similarity is within a preset similarity threshold range.

[0144] In this embodiment, in the vehicle speech recognition device 800, the specific processing of the generation module 801, the calculation module 802 and the execution module 803 and the technical effects thereof can be referred to in Figure 2 The relevant descriptions of steps 201 - 203 in the corresponding embodiment are not repeated here.

[0145] In some optional implementations of this embodiment, the above-mentioned vehicle voice recognition device 800 also includes: a prompt module, configured to generate a voice verification failure prompt message in response to determining that the similarity is not within the similarity threshold range; an update module, configured to update the voiceprint model using the target voice to obtain an updated voiceprint model.

[0146] In some optional implementations of this embodiment, the update module is further configured to: use the registered voiceprint vector as a positive sample and the target voiceprint vector as a negative sample, and use contrastive learning to train the voiceprint model to obtain an updated voiceprint model.

[0147] In some optional implementations of this embodiment, the above-mentioned vehicle voice recognition device 800 also includes: a vector generation module, configured to generate a new target voiceprint vector corresponding to the target voice using the updated voiceprint model; a threshold generation module, configured to generate a target threshold range based on the similarity between the new target voiceprint vector and the registered voiceprint vector; a threshold update module, configured to adaptively adjust the similarity threshold range of the updated voiceprint model to the target threshold range.

[0148] In some optional implementations of this embodiment, the above-mentioned vehicle voice recognition device 800 also includes: a registration module for obtaining a registered voiceprint vector, and the registration module is configured to: initialize the number of successful registrations to 0, and perform the following iterative process: record the user's voice in different noise environments to obtain the current voice; in response to determining that the current voice recording is successful, register the current voice into the voiceprint model; in response to determining that the current voice has been successfully registered into the voiceprint model, add 1 to the number of successful registrations; in response to determining that the number of successful registrations is less than the registration number threshold, perform the next round of iterative process; in response to determining that the number of successful registrations is greater than or equal to the registration number threshold, use the voiceprint model to generate a voiceprint vector corresponding to the successfully registered voice, and obtain the registered voiceprint vector corresponding to the user.

[0149] In some optional implementations of this embodiment, the above-mentioned vehicle voice recognition device 800 also includes: a wake-up word preset module, configured to preset at least one wake-up word; and the registration module is further configured to: record the user's voice containing the wake-up word in different noise environments.

[0150] In some optional implementations of this embodiment, the above-mentioned vehicle voice recognition device 800 also includes: a noise reduction module, configured to perform noise filtering on the initial voice to obtain a target voice; a recognition module, configured to perform voice recognition on the target voice and determine the voice content corresponding to the target voice; a matching module, configured to match the voice content with a preset wake-up word; and a generation module is further configured to: in response to determining that the voice content successfully matches the preset wake-up word, input the target voice into the voiceprint model, and output a target voiceprint vector.

[0151] In some optional implementations of this embodiment, the execution module includes: a device determination submodule, configured to determine the target device for collecting the target voice on the vehicle in response to determining that the similarity is within a preset similarity threshold range; and an execution submodule, configured to perform a target operation corresponding to the target voice on the area where the target device is located.

[0152] In some optional implementations of this embodiment, the device determination submodule is further configured to: determine distance information between the target voice and each collection device on the vehicle; and determine a target device for collecting the target voice from each collection device based on the distance information.

[0153] In some optional implementations of this embodiment, the target operation includes: an unlocking operation; and the execution submodule is further configured to: determine the vehicle-mounted device corresponding to the area where the target device is located, and unlock the vehicle-mounted device.

[0154] Further references Fig. 9 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a training device for a voiceprint model. Figure 6 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0155] like Fig. 9 As shown, the training device 900 of the voiceprint model of this embodiment includes: an acquisition module 901 and a training module 902. The acquisition module 901 is configured to acquire a training sample set, wherein the training sample pairs in the training sample set include positive sample speech and negative sample speech, the positive sample speech is the speech of the user in a quiet environment, and the negative sample speech is the speech of the user in a noisy environment; the training module 902 is configured to train the initial neural network using the training sample set to obtain a voiceprint model.

[0156] In the voiceprint model training device 900, the specific processing of the acquisition module 901 and the training module 902 and the technical effects thereof can be referred to in Figure 6 The relevant descriptions of steps 601-602 in the corresponding embodiment are not repeated here.

[0157] In some optional implementations of this embodiment, the above-mentioned voiceprint model training device 900 further includes: a threshold setting module, configured to set a similarity threshold range corresponding to the similarity between different voices for the voiceprint model.

[0158] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0159] Fig.10A schematic block diagram of an example electronic device 1000 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0160] like Fig.10 As shown, the device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0161] Multiple components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0162] The computing unit 1001 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1001 performs the various methods and processes described above, such as a vehicle voice recognition method or a training method for a voiceprint model. For example, in some embodiments, the vehicle voice recognition method or the training method for a voiceprint model may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the vehicle voice recognition method or the training method for a voiceprint model described above may be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to execute the vehicle speech recognition method or the voiceprint model training method in any other appropriate manner (for example, by means of firmware).

[0163] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0164] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0165] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0166] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0167] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0168] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0169] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0170] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A vehicle speech recognition method, comprising: Generate a target voiceprint vector corresponding to the target speech using a pre-trained voiceprint model, wherein the voiceprint model is trained using speech sample pairs in different noise environments; Calculating the similarity between the target voiceprint vector and a registered voiceprint vector corresponding to the vehicle; In response to determining that the similarity is within a preset similarity threshold range, a target operation corresponding to the target voice is performed.

2. The method according to claim 1, further comprising: In response to determining that the similarity is not within the similarity threshold range, generating voice verification failure prompt information; The voiceprint model is updated using the target speech to obtain an updated voiceprint model.

3. The method according to claim 2, wherein: The step of updating the voiceprint model by using the target voice to obtain an updated voiceprint model includes: The registered voiceprint vector is used as a positive sample, and the target voiceprint vector is used as a negative sample. The voiceprint model is trained by contrastive learning to obtain an updated voiceprint model.

4. The method according to claim 2, further comprising: Generating a new target voiceprint vector corresponding to the target speech using the updated voiceprint model; generating a target threshold range according to the similarity between the new target voiceprint vector and the registered voiceprint vector; The similarity threshold range of the updated voiceprint model is adaptively adjusted to the target threshold range.

5. The method according to claim 1, wherein: The registered voiceprint vector is obtained by the following steps: Initialize the number of successful registrations to 0, and perform the following iterative process: record the user's voice in different noise environments to obtain the current voice; in response to determining that the current voice is successfully recorded, register the current voice into the voiceprint model; In response to determining that the current voice has been successfully registered into the voiceprint model, adding 1 to the number of successful registrations; In response to determining that the number of successful registrations is less than the registration number threshold, performing a next round of iteration process; In response to determining that the number of successful registrations is greater than or equal to the registration number threshold, the voiceprint model is used to generate a voiceprint vector corresponding to the successfully registered voice to obtain a registered voiceprint vector corresponding to the user.

6. The method according to claim 5, further comprising: Preset at least one wake-up word; as well as The recording of the user's voice in different noise environments includes: Record the user's voice containing the wake-up word in different noise environments.

7. The method according to claim 1, further comprising: Performing noise filtering on the initial speech to obtain the target speech; Performing speech recognition on the target speech to determine the speech content corresponding to the target speech; Matching the voice content with a preset wake-up word; as well as The method of using the pre-trained voiceprint model to generate a target voiceprint vector corresponding to the target speech includes: In response to determining that the voice content successfully matches the preset wake-up word, the target voice is input into the voiceprint model, and the target voiceprint vector is output.

8. The method according to claim 1, wherein: In response to determining that the similarity is within a preset similarity threshold range, executing a target operation corresponding to the target voice includes: In response to determining that the similarity is within a preset similarity threshold range, determining a target device on the vehicle that collects the target voice; The target operation corresponding to the target voice is performed on the area where the target device is located.

9. The method according to claim 8, wherein: The determining of a target device on the vehicle for collecting the target voice comprises: Determine distance information between the target voice and each acquisition device on the vehicle; A target device for collecting the target voice is determined from the various collection devices according to the distance information.

10. The method according to claim 8, wherein: The target operation includes: an unlocking operation; and The performing the target operation corresponding to the target voice on the area where the target device is located includes: Determine the vehicle-mounted device corresponding to the area where the target device is located, and unlock the vehicle-mounted device.

11. A method for training a voiceprint model, comprising: Acquire a training sample set, wherein the training sample pairs in the training sample set include positive sample speech and negative sample speech, the positive sample speech is the speech of the user in a quiet environment, and the negative sample speech is the speech of the user in a noisy environment; The initial neural network is trained using the training sample set to obtain the voiceprint model.

12. The method according to claim 11, further comprising: A similarity threshold range corresponding to the similarity between different voices is set for the voiceprint model.

13. A vehicle voice recognition device, comprising: A generating module is configured to generate a target voiceprint vector corresponding to the target speech using a pre-trained voiceprint model, wherein the voiceprint model is trained using speech sample pairs in different noise environments; a calculation module configured to calculate the similarity between the target voiceprint vector and a registered voiceprint vector corresponding to the vehicle; The execution module is configured to execute the target operation corresponding to the target speech in response to determining that the similarity is within a preset similarity threshold range.

14. A voiceprint model training device, comprising: An acquisition module is configured to acquire a training sample set, wherein the training sample pairs in the training sample set include positive sample speech and negative sample speech, the positive sample speech is the speech of the user in a quiet environment, and the negative sample speech is the speech of the user in a noisy environment; The training module is configured to train the initial neural network using the training sample set to obtain the voiceprint model.

15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-10 or 11-12.

16. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the method of any one of claims 1-10 or 11-12.

17. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1-10 or 11-12.