Target voice response method and device, medium and electronic equipment
By acquiring the target voice in the vehicle and screening the user voiceprint that adapts to the current state, and using noise reduction and dynamic adjustment of the voiceprint recognition layer, the problem of insufficient voiceprint recognition accuracy caused by changes in user voice noise in the vehicle is solved, and the response efficiency and robustness are improved.
Patent Information
- Application Number
- CN202510840574.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-19
AI Technical Summary
In vehicles, changes in user voice noise lead to insufficient accuracy and robustness of voiceprint recognition, affecting the efficiency of target voice response.
The target voice is acquired in the vehicle, and the user voiceprint that is adapted to the current vehicle status is filtered out through the voiceprint model. After the response, the user voiceprint in the current status is actively registered. The noise reduction layer is used to improve the signal-to-noise ratio, and the voiceprint recognition layer is converted into voiceprint features. The voiceprint recognition is dynamically adjusted in combination with the vehicle status.
It improves the target voice response efficiency, enhances the accuracy and robustness of voiceprint recognition in multi-noise scenarios, and adapts to user voiceprint recognition in different vehicle states.
Smart Images

Figure CN120673767A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of speech processing technology, and in particular relates to a method, device, medium and electronic device for responding to a target speech. Background Art
[0002] Currently, voiceprint recognition is typically performed on users registered in vehicles using a single voiceprint. However, the noise level of user voices varies depending on the vehicle's driving state. This makes it difficult to use a voiceprint registered in a quiet environment to recognize a noisy user's voice, and it's also difficult to use a voiceprint registered in a noisy environment to recognize a quiet user's voice. This results in insufficient voiceprint recognition accuracy and robustness, which in turn reduces the efficiency of responding to the target voice.
[0003] Based on this, the low response efficiency of the target voice is a technical problem that needs to be solved urgently. Summary of the Invention
[0004] The embodiments of the present application provide a method, apparatus, medium, and electronic device for responding to a target voice, thereby improving the efficiency of responding to the target voice at least to a certain extent.
[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.
[0006] According to a first aspect of an embodiment of the present application, a method for responding to a target speech is provided, the method comprising:
[0007] Acquire target voice in target vehicle;
[0008] Inputting the target speech into a voiceprint model to obtain a first voiceprint;
[0009] selecting a third voiceprint from a set of registered voiceprints on the target vehicle based on the first voiceprint, wherein the set of voiceprints includes at least one second voiceprint, and the second voiceprint is a user voiceprint registered based on a set of voices in a corresponding vehicle state;
[0010] Responding to the target voice based on the third voiceprint, obtaining a target user corresponding to the target voice;
[0011] After responding to the target voice based on the third voiceprint, registering the user voiceprint in the current vehicle state based on the target voice.
[0012] In some embodiments of the present application, based on the aforementioned scheme, the voiceprint model includes a noise reduction layer and a voiceprint recognition layer, and inputting the target speech into the voiceprint model to obtain the first voiceprint includes: inputting the target speech into the noise reduction layer to obtain an updated speech, wherein the noise reduction layer is used to improve the signal-to-noise ratio of the target speech; inputting the updated speech into the voiceprint recognition layer to obtain the first voiceprint, wherein the voiceprint recognition layer is used to convert the speech into a corresponding voiceprint.
[0013] In some embodiments of the present application, based on the aforementioned scheme, the registering of the user voiceprint in the current vehicle state based on the target voice includes: obtaining a reference vehicle state corresponding to the third voiceprint, wherein the reference vehicle state is used to indicate a reference vehicle speed and a reference media volume corresponding to the voice set used to register the third voiceprint; obtaining the current vehicle state of the target vehicle, wherein the current vehicle state is used to indicate the current vehicle speed and the current media volume of the target vehicle; if the current vehicle speed is different from the reference vehicle speed, or the current media volume is different from the reference media volume, it is determined that the current vehicle state is different from the reference vehicle state, and the user voiceprint in the current vehicle state is registered based on the target voice.
[0014] In some embodiments of the present application, based on the aforementioned scheme, registering the user voiceprint in the current vehicle state based on the target voice includes: obtaining the updated voice obtained by passing the target voice through the noise reduction layer; adding the updated voice to the initial voice set of the target user in the current vehicle state to obtain the target voice set; if the voice length of the target voice set meets the registration duration, registering the user voiceprint in the current vehicle state based on the target voice set.
[0015] In some embodiments of the present application, based on the aforementioned scheme, registering the user voiceprint in the current vehicle state based on the target voice set includes: inputting the target voice set into the voiceprint model to obtain a fourth voiceprint; and determining the fourth voiceprint as the user voiceprint of the target user in the current vehicle state.
[0016] In some embodiments of the present application, based on the aforementioned scheme, the screening of the third voiceprint from the voiceprint set registered on the target vehicle based on the first voiceprint includes: calculating the voiceprint similarity between the first voiceprint and each of the second voiceprints in the voiceprint set; determining whether there is a voiceprint similarity greater than or equal to a similarity threshold; if there is a voiceprint similarity greater than or equal to the similarity threshold, screening the second voiceprint with the highest voiceprint similarity as the third voiceprint from the voiceprint similarities greater than or equal to the similarity threshold.
[0017] In some embodiments of the present application, based on the aforementioned scheme, responding to the target voice based on the third voiceprint to obtain the target user corresponding to the target voice includes: obtaining the user corresponding to the third voiceprint; and determining the user corresponding to the third voiceprint as the target user corresponding to the target voice.
[0018] According to a second aspect of an embodiment of the present application, a device for responding to a target speech is provided, the device comprising:
[0019] An acquisition module, used to acquire target voice in a target vehicle;
[0020] An input module, configured to input the target speech into a voiceprint model to obtain a first voiceprint;
[0021] a screening module, configured to screen a third voiceprint from a set of voiceprints registered on the target vehicle based on the first voiceprint, wherein the set of voiceprints includes at least one second voiceprint, and the second voiceprint is a user voiceprint registered based on a set of voices in a corresponding vehicle state;
[0022] A response module, configured to respond to the target voice based on the third voiceprint, and obtain a target user corresponding to the target voice;
[0023] A registration module is configured to register the user voiceprint in the current vehicle state based on the target voice after responding to the target voice based on the third voiceprint.
[0024] According to a third aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which at least one computer program instruction is stored. The at least one computer program instruction is loaded and executed by a processor to implement the operations performed by the method described in any one of the first aspects above.
[0025] According to a fourth aspect of an embodiment of the present application, an electronic device is provided, comprising one or more processors and one or more memories, wherein at least one computer program instruction is stored in the one or more memories, and the at least one computer program instruction is loaded and executed by the one or more processors to implement the method described in any embodiment of the first aspect above.
[0026] In the present application, a target voice in a target vehicle is obtained; the target voice is input into a voiceprint model to obtain a first voiceprint; a third voiceprint is screened from a registered voiceprint set on the target vehicle based on the first voiceprint, wherein the voiceprint set includes at least one second voiceprint, and the second voiceprint is a user voiceprint registered based on a voice set in a corresponding vehicle state; the target voice is responded to based on the third voiceprint to obtain a target user corresponding to the target voice; after responding to the target voice based on the third voiceprint, the user voiceprint in the current vehicle state is registered based on the target voice. In other words, during the operation of the target vehicle, not only the currently registered voiceprint set is used to respond to the target voice, but also the user voiceprint in the current vehicle state is actively registered for the target user after the response, so that the target user has a corresponding user voiceprint in a multi-noise scenario. When responding to the next user voice of the target user in the current vehicle state, the user voiceprint in the current vehicle state can be used, thereby improving the response efficiency of the target voice.
[0027] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, explaining the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort. In the drawings:
[0029] Figure 1 A flow chart showing a method for responding to a target voice in an embodiment of the present application is shown;
[0030] Figure 2 A schematic diagram showing a user voiceprint registration process to which an embodiment of the present application can be applied;
[0031] Figure 3 A block diagram of a device for responding to target speech in an embodiment of the present application is shown;
[0032] Figure 4 A schematic structural diagram of an electronic device in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0033] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0034] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0035] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0036] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0037] It should be noted that the term "plurality" as used herein refers to two or more. The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described.
[0038] Figure 1 A flow chart showing a method for responding to a target voice in an embodiment of the present application is shown. The method for responding to a target voice can be executed by a device having a computing and processing function in a target vehicle. Figure 1 As shown, the target speech response method includes:
[0039] Step 101: Acquire target speech in a target vehicle;
[0040] Step 102: Input the target speech into a voiceprint model to obtain a first voiceprint;
[0041] Step 103: Filtering a third voiceprint from a set of registered voiceprints on the target vehicle based on the first voiceprint, wherein the voiceprint set includes at least one second voiceprint, and the second voiceprint is a user voiceprint registered based on a voice set in a corresponding vehicle state;
[0042] Step 104: Responding to the target voice based on the third voiceprint to obtain a target user corresponding to the target voice;
[0043] Step 105: After responding to the target voice based on the third voiceprint, register the user voiceprint in the current vehicle state based on the target voice.
[0044] Through the above steps, the target voice in the target vehicle is obtained; the target voice is input into the voiceprint model to obtain a first voiceprint; based on the first voiceprint, a third voiceprint is screened from the registered voiceprint set on the target vehicle, wherein the voiceprint set includes at least one second voiceprint, and the second voiceprint is a user voiceprint registered based on the voice set in the corresponding vehicle state; the target voice is responded to based on the third voiceprint to obtain the target user corresponding to the target voice; after responding to the target voice based on the third voiceprint, the user voiceprint in the current vehicle state is registered based on the target voice. In other words, during the operation of the target vehicle, not only the currently registered voiceprint set is used to respond to the target voice, but also the user voiceprint in the current vehicle state is actively registered for the target user after the response, so that the target user has the corresponding user voiceprint in a multi-noise scenario. When responding to the next user voice of the target user in the current vehicle state, the user voiceprint in the current vehicle state can be used, thereby improving the response efficiency of the target voice.
[0045] In the embodiment provided in step 101, the target voice is a voice signal for voiceprint recognition emanating from a user riding in a target vehicle. The target vehicle is the vehicle in which the user emanating the target voice is riding. The target voice may be acquired, but is not limited to, via a microphone array or similar voice acquisition device configured in the target vehicle.
[0046] Optionally, in this embodiment, the user who makes the target speech is an individual who needs to use a voiceprint model to verify the speaker's identity, such as the vehicle owner, driver, passenger, etc. Different users may need to verify the speaker's identity in, but are not limited to, various scenarios. For example, for a driver, he needs to verify his identity in various scenarios (such as controlling the doors, windows, wipers, setting the destination, controlling the media volume, etc.) to obtain control of the entire vehicle; for a passenger, he needs to be accurately identified in various scenarios (such as controlling the media volume, adjusting the air conditioner, setting the destination, etc.) to obtain control of some software in the vehicle.
[0047] In the embodiment provided in step 102, the target speech is input into the voiceprint model to obtain a first voiceprint output by the voiceprint model for representing the voiceprint characteristics of the user who utters the target speech.
[0048] In one embodiment of the present application, the voiceprint model includes a noise reduction layer and a voiceprint recognition layer. The target speech can be input into the voiceprint model to obtain the first voiceprint in the following manner, but is not limited to: inputting the target speech into the noise reduction layer to obtain an updated speech, wherein the noise reduction layer is used to improve the signal-to-noise ratio of the target speech; inputting the updated speech into the voiceprint recognition layer to obtain the first voiceprint, wherein the voiceprint recognition layer is used to convert the speech into a corresponding voiceprint.
[0049] Optionally, in this embodiment, the noise reduction layer and voiceprint recognition layer may be, but are not limited to, a noise reduction model and a voiceprint recognition model respectively (i.e., the noise reduction model is deployed before the voiceprint recognition model), or the noise reduction layer may be a functional layer obtained after training the voiceprint recognition model.
[0050] Furthermore, the aforementioned voiceprint recognition model is a machine learning model that identifies or verifies a speaker's identity by analyzing speech features. Its core principle is to convert the personalized acoustic features of speech into mathematical representations, enabling "recognition by sound." In other words, the voiceprint recognition model is content-independent, focusing only on the biometric characteristics of speech (such as the spectrum and formants) rather than the specific content.
[0051] Since the different signal-to-noise ratios of vehicles at different vehicle speeds and media volumes will affect the verification performance of the voiceprint recognition model, in the target voice response method proposed in this application, a noise reduction layer is added to the voiceprint recognition layer. The noise reduction layer is used to first reduce the noise of the acquired user voice, and then the voiceprint is converted through the voiceprint recognition layer, thereby improving the accuracy of voiceprint recognition.
[0052] In the embodiment provided in step 103, it is possible but not limited to obtaining one or more voices recorded by the target user in a scene without human voice and environmental noise in the first use scenario of the target vehicle, modeling is performed through the voiceprint algorithm to obtain the first second voiceprint, and the second voiceprint is bound to the target user (for example: establishing a corresponding relationship, etc.).
[0053] Furthermore, during the operation of the vehicle, for the same target user, the first third voiceprint registered in the noise-free scene is used to respond to his user voice. In the next usage scenario, when the user who makes the user voice is determined to be the target user based on the third voiceprint, the user voice is retained to register the user voiceprint of the target user in the current noise scene (different from the noise-free scene), and after the registration is completed, it is added to the voiceprint set as the second voiceprint, thereby realizing the imperceptible registration of the target user's voiceprint.
[0054] If the vehicle state corresponding to the third voiceprint is different from the current vehicle state, the method proposed in this application uses the registered third voiceprint to determine that the user who makes the user voice is the target user (that is, the third voiceprint is used to respond to the user voice), then the vehicle state corresponding to the third voiceprint can be considered to be similar to the current vehicle state. Therefore, the recognition result of the third voiceprint registered in a similar vehicle state (that is, determining the target user) can be used as the recognition result under the current vehicle state. However, the purpose of the above process is to progressively and imperceptibly register the user voiceprint of the target user in multiple vehicle states, thereby reducing the impact of noise scenes on the voiceprint recognition process.
[0055] Furthermore, in actual application scenarios, the actual storage capacity of the target vehicle is taken into consideration and the recognition interval is set based on the frequency of use of the user's voiceprint, and the user voiceprint with high usage frequency is retained in a certain scenario interval.
[0056] For each target user, each user needs to register the first second voiceprint through user interaction at the beginning of use, and then use the second voiceprint to register their user voiceprints in different vehicle states without perception. In other words, the voiceprint set can include, but is not limited to, at least one second voiceprint corresponding to at least one target user.
[0057] It is understood that for the same target vehicle, it is allowed to register a second voiceprint for one or more users. For each second voiceprint, it has a corresponding target user and vehicle status. For different second voiceprints, the target user or vehicle status they correspond to may be different.
[0058] In one embodiment of the present application, a third voiceprint can be screened from the voiceprint set registered on the target vehicle based on the first voiceprint, but is not limited to, in the following manner: calculating the voiceprint similarity between the first voiceprint and each of the second voiceprints in the voiceprint set; determining whether there is a voiceprint similarity greater than or equal to a similarity threshold; if there is a voiceprint similarity greater than or equal to the similarity threshold, screening the second voiceprint with the highest voiceprint similarity as the third voiceprint from the voiceprint similarities greater than or equal to the similarity threshold.
[0059] Optionally, in this embodiment, if the voiceprint similarity does not exist that is greater than or equal to the similarity threshold, it is considered that the user voiceprint of the user who uttered the target voice is not registered in the target vehicle, and a prompt "User not registered" is displayed and a registration window pops up.
[0060] Optionally, in this embodiment, the similarity threshold may be a preset fixed value, or a variable value determined according to the vehicle speed and the media volume.
[0061] In the embodiment provided in step 104, the user corresponding to the third voiceprint may be searched from users and voiceprints having corresponding relationships, but is not limited to the search, as the target user.
[0062] In one embodiment of the present application, the target user corresponding to the target voice can be obtained by responding to the target voice based on the third voiceprint, but is not limited to the following methods: obtaining the user corresponding to the third voiceprint; and determining the user corresponding to the third voiceprint as the target user corresponding to the target voice.
[0063] In the embodiment provided in step 105, the current vehicle state can be determined based on, but not limited to, the current vehicle speed and the current media volume. For example, when the vehicle speed is 0-50 km / h (kilometers per hour), the media volume is 0-30% and the corresponding driving scene is a low-speed and low-volume scene; when the vehicle speed is 50-80 km / h, the media volume is 0-30% and the corresponding driving scene is a medium-speed and low-volume scene; when the vehicle speed is above 80 km / h, the media volume is 0-30% and the corresponding driving scene is a high-speed and low-volume scene, etc.
[0064] In one embodiment of the present application, the user voiceprint in the current vehicle state can be registered based on the target voice in the following manner, but is not limited to: obtaining a reference vehicle state corresponding to the third voiceprint, wherein the reference vehicle state is used to indicate a reference vehicle speed and a reference media volume corresponding to the voice set used to register the third voiceprint; obtaining the current vehicle state of the target vehicle, wherein the current vehicle state is used to indicate the current vehicle speed and the current media volume of the target vehicle; if the current vehicle speed is different from the reference vehicle speed, or the current media volume is different from the reference media volume, it is determined that the current vehicle state is different from the reference vehicle state, and the user voiceprint in the current vehicle state is registered based on the target voice.
[0065] Optionally, in this embodiment, the current speed of the target vehicle can be obtained in a variety of ways, including but not limited to: directly reading the current speed through a CAN (Controller Area Network) bus, reading the current speed through an OBD-II (On Board Diagnostics II) interface, reading the current speed through an in-vehicle infotainment system API (Application Programming Interface), and calculating the current speed with the assistance of a GPS (Global Positioning System).
[0066] Optionally, in this embodiment, the current media volume of the target vehicle can be obtained in a variety of ways, including but not limited to: directly accessing the current media volume through the audio system, monitoring the current media volume through the automotive audio bus (Automotive Audio Bus, A2B), and obtaining the current media volume using the infotainment system protocol.
[0067] In one embodiment of the present application, the user voiceprint in the current vehicle state can be registered based on the target voice in the following manner, but is not limited to: obtaining the updated voice obtained by passing the target voice through the noise reduction layer; adding the updated voice to the initial voice set of the target user in the current vehicle state to obtain the target voice set; if the voice length of the target voice set meets the registration time length, registering the user voiceprint in the current vehicle state based on the target voice set.
[0068] Optionally, in this embodiment, the initial voice set includes multiple historical user voices for the target user and the current vehicle state, stored during the historical voiceprint recognition process. If the initial voice set for the target user and the current vehicle state cannot be found, the updated voices are used to construct the initial voice set. The voice set can be generated by clustering the voices based on, but not limited to, the user and vehicle state.
[0069] Optionally, in this embodiment, the registration duration is a preset voice length that can be used to register a user's voiceprint.
[0070] In one embodiment of the present application, the user voiceprint in the current vehicle state can be registered based on the target voice set in the following manner, but is not limited to: inputting the target voice set into the voiceprint model to obtain a fourth voiceprint; and determining the fourth voiceprint as the user voiceprint of the target user in the current vehicle state.
[0071] Optionally, in this embodiment, the application process of voiceprint models in related technologies is first described: the voiceprint model includes a training phase, a registration phase, and a verification phase. The training phase trains a voiceprint encoder model using a large number of speech samples. The voiceprint encoder model converts the input speech into a voiceprint feature that represents the speaker's voice characteristics, providing a basis for subsequent verification.
[0072] During the registration phase, users are asked to record a voice sample. This sample is then used to train a speaker model that represents the user's voiceprint characteristics for subsequent voiceprint comparisons. Generally speaking, each user has a unique speaker model that identifies the user.
[0073] The verification phase is where voiceprint recognition is actually performed. During the verification phase, a test voice (i.e., user voice) is read in and converted into voiceprint features using the voiceprint encoder model from the training phase. This voiceprint feature is then compared with the speaker model constructed during the registration phase to determine whether the voice matches a specific speaker, thereby confirming the speaker's identity.
[0074] Differently, in the technical solution proposed in this application, Figure 2 The following is a schematic diagram showing the process of registering a user's voiceprint that can be applied to the embodiment of the present application. Figure 2 As shown, since the target user has been determined through the response process of the target voice, when using the target voice to register the user voiceprint in the current vehicle state, it is only necessary to input the target voice set into the voiceprint model to obtain the fourth voiceprint, and bind the fourth voiceprint to the target user to complete the registration of the user voiceprint of the target user in the current vehicle state.
[0075] This application provides a method for responding to target voices. Based on the registration and verification stages of related technical solutions, an adaptive modeling stage is added. The user voice is classified according to the target user and vehicle status (such as vehicle speed and media volume) identified in the verification stage. Then, a voice set under the same category is used to register the user voiceprint corresponding to the vehicle status for the target user, thus achieving non-perceptual adaptive registration of user voiceprints in multiple scenarios. In terms of technical effects, this method improves the robustness of the voiceprint model in vehicle-mounted scenarios, adapts to different noise reduction algorithms and dynamically changing vehicle status, bridges the gap between theoretical models and practical applications, and enhances recognition accuracy and adaptability in complex environments.
[0076] The following describes an embodiment of the device of the present application, which can be used to execute the target speech response method in the above embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the target speech response method in the above embodiment of the present application.
[0077] See also Figure 3 , shows a block diagram of a target speech response device in an embodiment of the present application.
[0078] like Figure 3 As shown, the target speech response device (300) according to an embodiment of the present application includes:
[0079] Acquisition module 301 , input module 302 , screening module 303 , response module 304 , and registration module 305 .
[0080] Among them, the acquisition module is used to acquire the target voice in the target vehicle; the input module is used to input the target voice into the voiceprint model to obtain a first voiceprint; the screening module is used to screen a third voiceprint from the voiceprint set registered on the target vehicle based on the first voiceprint, wherein the voiceprint set includes at least one second voiceprint, and the second voiceprint is a user voiceprint registered based on the voice set in the corresponding vehicle state; the response module is used to respond to the target voice based on the third voiceprint to obtain the target user corresponding to the target voice; the registration module is used to register the user voiceprint in the current vehicle state based on the target voice after responding to the target voice based on the third voiceprint.
[0081] In some embodiments of the present application, based on the aforementioned solution, the voiceprint model includes a noise reduction layer and a voiceprint recognition layer, and the input module includes:
[0082] a first input unit, configured to input the target speech into the noise reduction layer to obtain an updated speech, wherein the noise reduction layer is configured to improve a signal-to-noise ratio of the target speech;
[0083] The second input unit is configured to input the updated voice into the voiceprint recognition layer to obtain the first voiceprint, wherein the voiceprint recognition layer is configured to convert the voice into a corresponding voiceprint.
[0084] In some embodiments of the present application, based on the aforementioned solution, the registration module includes:
[0085] a first acquiring unit, configured to acquire a reference vehicle state corresponding to the third voiceprint, wherein the reference vehicle state is used to indicate a reference vehicle speed and a reference media volume corresponding to a voice set used to register the third voiceprint;
[0086] a second acquiring unit, configured to acquire the current vehicle state of the target vehicle, wherein the current vehicle state is used to indicate a current vehicle speed and a current media volume of the target vehicle;
[0087] A processing unit is configured to determine, if the current vehicle speed is different from the reference vehicle speed, or the current media volume is different from the reference media volume, that the current vehicle state is different from the reference vehicle state, and to register the user voiceprint in the current vehicle state based on the target voice.
[0088] In some embodiments of the present application, based on the aforementioned scheme, the processing unit is also used to: obtain the updated voice obtained by passing the target voice through the noise reduction layer; add the updated voice to the initial voice set of the target user in the current vehicle state to obtain a target voice set; if the voice length of the target voice set meets the registration time length, register the user voiceprint in the current vehicle state based on the target voice set.
[0089] In some embodiments of the present application, based on the aforementioned scheme, the processing unit is further used to: input the target speech set into the voiceprint model to obtain a fourth voiceprint; and determine the fourth voiceprint as the user voiceprint of the target user in the current vehicle state.
[0090] In some embodiments of the present application, based on the above scheme, the screening module includes:
[0091] a calculation unit, configured to calculate a voiceprint similarity between the first voiceprint and each of the second voiceprints in the voiceprint set;
[0092] A first determining unit, configured to determine whether there is a voiceprint similarity greater than or equal to a similarity threshold;
[0093] The screening unit is configured to screen the second voiceprint with the highest voiceprint similarity from the voiceprint similarities greater than or equal to the similarity threshold as the third voiceprint if there is a voiceprint similarity greater than or equal to the similarity threshold.
[0094] In some embodiments of the present application, based on the aforementioned solution, the response module includes:
[0095] A third acquiring unit, configured to acquire a user corresponding to the third voiceprint;
[0096] The second determining unit is configured to determine the user corresponding to the third voiceprint as the target user corresponding to the target voice.
[0097] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which stores at least one computer program instruction. The at least one computer program instruction is loaded and executed by a processor to implement the operations performed by the method described above.
[0098] Based on the same inventive concept, the present application also provides an electronic device, referring to Figure 4 , shows a schematic structural diagram of an electronic device in an embodiment of the present application, wherein the electronic device includes one or more memories 404, one or more processors 402, and at least one computer program (computer program instruction) stored in the memory 404 and executable on the processor 402, and the method described above is implemented when the processor 402 executes the computer program.
[0099] Among them, Figure 4In the embodiment of the present invention, a bus architecture (represented by bus 400) is shown. Bus 400 may include any number of interconnected buses and bridges, and bus 400 links together various circuits including one or more processors represented by processor 402 and memory represented by memory 404. Bus 400 may also link together various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 405 provides an interface between bus 400 and receiver 401 and transmitter 403. Receiver 401 and transmitter 403 may be the same component, namely a transceiver, which provides a unit for communicating with various other devices over a transmission medium. Processor 402 is responsible for managing bus 400 and general processing, while memory 404 may be used to store data used by processor 402 when performing operations.
[0100] The functions described herein may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored as one or more instructions or codes on or transmitted via a computer-readable medium. Other examples and implementations are within the scope and spirit of this application and the appended claims. For example, due to the nature of software, the functions described above may be implemented using software executed by a processor, hardware, firmware, hardwiring, or a combination of any of these. Furthermore, the functional units may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.
[0101] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0102] The units described as separate components may or may not be physically separate, and the components of the control device may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0103] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store computer program instructions.
[0104] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for responding to a target speech, characterized in that: The method comprises: Acquire target voice in target vehicle; Inputting the target speech into a voiceprint model to obtain a first voiceprint; selecting a third voiceprint from a set of registered voiceprints on the target vehicle based on the first voiceprint, wherein the set of voiceprints includes at least one second voiceprint, and the second voiceprint is a user voiceprint registered based on a set of voices in a corresponding vehicle state; Responding to the target voice based on the third voiceprint, obtaining a target user corresponding to the target voice; After responding to the target voice based on the third voiceprint, registering the user voiceprint in the current vehicle state based on the target voice.
2. The method according to claim 1, characterized in that The voiceprint model includes a noise reduction layer and a voiceprint recognition layer, and inputting the target speech into the voiceprint model to obtain the first voiceprint includes: Inputting the target speech into the noise reduction layer to obtain an updated speech, wherein the noise reduction layer is used to improve the signal-to-noise ratio of the target speech; The updated voice is input into the voiceprint recognition layer to obtain the first voiceprint, wherein the voiceprint recognition layer is used to convert the voice into a corresponding voiceprint.
3. The method according to claim 1, characterized in that The registering of the user voiceprint in the current vehicle state based on the target voice includes: Obtaining a reference vehicle state corresponding to the third voiceprint, wherein the reference vehicle state is used to indicate a reference vehicle speed and a reference media volume corresponding to the voice set used to register the third voiceprint; Acquiring the current vehicle state of the target vehicle, wherein the current vehicle state is used to indicate the current vehicle speed and current media volume of the target vehicle; If the current vehicle speed is different from the reference vehicle speed, or the current media volume is different from the reference media volume, it is determined that the current vehicle state is different from the reference vehicle state, and the user voiceprint in the current vehicle state is registered based on the target voice.
4. The method according to claim 2, characterized in that The registering of the user voiceprint in the current vehicle state based on the target voice includes: Obtaining the updated speech obtained by passing the target speech through the noise reduction layer; Adding the updated voice to the initial voice set of the target user in the current vehicle state to obtain a target voice set; If the speech length of the target speech set meets the registration time, the user voiceprint in the current vehicle state is registered based on the target speech set.
5. The method according to claim 4, characterized in that The registering of the user voiceprint in the current vehicle state based on the target voice set includes: Inputting the target speech set into the voiceprint model to obtain a fourth voiceprint; The fourth voiceprint is determined as the user voiceprint of the target user in the current vehicle state.
6. The method according to claim 1, characterized in that The selecting a third voiceprint from the voiceprint set registered on the target vehicle based on the first voiceprint includes: Calculating a voiceprint similarity between the first voiceprint and each second voiceprint in the voiceprint set; Determining whether there is a voiceprint similarity greater than or equal to a similarity threshold; If the voiceprint similarities are greater than or equal to the similarity threshold, the second voiceprint with the highest voiceprint similarity is selected from the voiceprint similarities greater than or equal to the similarity threshold as the third voiceprint.
7. The method according to claim 1, characterized in that The step of responding to the target voice based on the third voiceprint to obtain a target user corresponding to the target voice includes: Obtaining the user corresponding to the third voiceprint; The user corresponding to the third voiceprint is determined as the target user corresponding to the target voice.
8. A target speech response device, characterized in that: The device comprises: An acquisition module, used to acquire target voice in a target vehicle; An input module, configured to input the target speech into a voiceprint model to obtain a first voiceprint; a screening module, configured to screen a third voiceprint from a set of voiceprints registered on the target vehicle based on the first voiceprint, wherein the set of voiceprints includes at least one second voiceprint, and the second voiceprint is a user voiceprint registered based on a set of voices in a corresponding vehicle state; A response module, configured to respond to the target voice based on the third voiceprint, and obtain a target user corresponding to the target voice; A registration module is configured to register the user voiceprint in the current vehicle state based on the target voice after responding to the target voice based on the third voiceprint.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, which are loaded and executed by a processor to implement the operations performed by the method according to any one of claims 1 to 7.
10. An electronic device comprising a processor and a memory, characterized in that: The memory stores computer program instructions that can be executed by the processor, and when the processor executes the computer program instructions, the processor implements the method according to any one of claims 1 to 7.