A vehicle voice interaction method, device, storage medium and equipment
By acquiring the target user's wake-up voice in the vehicle cabin and combining it with the wake-up area and model judgment, the problem of low accuracy in suppressing wake-up outside the cabin is solved, thus improving the user interaction experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING CO WHEELS TECH CO LTD
- Filing Date
- 2024-11-25
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, during voice interaction in the vehicle cabin, external wake-up suppression processing causes normal wake-up commands from in-vehicle users to be incorrectly suppressed, reducing the accuracy of external wake-up suppression and the user interaction experience.
By acquiring the wake-up voice of the target user, the wake-up area is determined, and by using the external wake-up suppression model and the occupancy information of the wake-up location, combined with acoustic features and frequency band energy analysis, the system can determine from multiple dimensions whether the voice is from outside the vehicle cabin, thereby improving the accuracy of wake-up suppression.
It improves the accuracy of external wake-up suppression, enhances the user's interactive experience, and reduces the impact of external wake-up suppression on the user's daily interactions inside the vehicle.
Smart Images

Figure CN122090835A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle technology, and in particular to a vehicle voice interaction method, device, storage medium and equipment. Background Technology
[0002] With the improvement of people's living standards and the rapid development of the social economy, the usage rate of cars is gradually increasing, and people's requirements for vehicle functions are also getting higher and higher. For example, for the vehicle cabin, it is usually necessary to provide users with interactive functions such as turning on navigation, air conditioning, sunroof, and music through voice control, so as to improve the user's riding experience.
[0003] Currently, during voice interaction in the vehicle cabin, to prevent external sounds (such as shouts or noise) from mistakenly activating the in-vehicle voice assistant system and to avoid potential safety risks or misoperations, external wake-up suppression is typically implemented. This prevents external voice commands from incorrectly controlling the vehicle to perform operations such as opening windows or the trunk, thus improving vehicle safety. However, current external wake-up suppression techniques may inadvertently suppress some normally used wake-up commands from the user inside the vehicle cabin, thereby reducing the accuracy of external wake-up suppression and affecting the user's interactive experience. Summary of the Invention
[0004] In view of the above problems, this application provides a vehicle voice interaction method, device, storage medium and equipment, which can improve the accuracy of suppressing wake-up from outside the vehicle cabin, thereby enhancing the user's interactive experience.
[0005] This application provides a vehicle voice interaction method, including:
[0006] Acquire the target wake-up voice sent by the target user; and determine the wake-up area where the target user is located based on the target wake-up voice;
[0007] The target wake-up voice is input into the external wake-up suppression model to obtain the first determination result;
[0008] Based on the wake-up area of the target user, the characteristic parameters of the target wake-up voice are determined; and based on the characteristic parameters, a second determination result is determined.
[0009] Determine the placeholder information for the wake-up location;
[0010] Based on the first determination result, the second determination result, and the placeholder information, the wake-up result corresponding to the target wake-up voice is determined.
[0011] In one possible implementation, the wake-up region where the target user is located is one of the preset wake-up regions; the preset wake-up region includes a first region, a second region, a third region, and a fourth region; the preset wake-up region is constructed as follows:
[0012] A coordinate system is established with the center of the ground where the front of the vehicle is located as the origin. The center line of the vehicle along the length of the vehicle body is taken as the central axis. Starting from the central axis, the first region, the second region, the third region, and the fourth region are symmetrically divided to both sides along the width of the vehicle body.
[0013] In one possible implementation, determining the wake-up area of the target user based on the target wake-up voice includes:
[0014] The target wake-up voice is localized to determine the coordinate information of the target user based on the localization result; and the wake-up area of the target user is determined based on the coordinate information of the target user.
[0015] In one possible implementation, the step of inputting the target wake-up voice into the external wake-up suppression model to obtain a first determination result includes:
[0016] The target wake-up voice is input into the external wake-up suppression model to obtain the acoustic features of the target wake-up voice;
[0017] The acoustic features of the target wake-up speech are processed by convolution and pooling to obtain the probability value that the target wake-up speech belongs to the speech outside the vehicle cabin.
[0018] The probability value is compared with a preset probability threshold, and a first determination result is determined based on the comparison result.
[0019] In one possible implementation, the extravehicular wake-up suppression model is constructed as follows:
[0020] Acquire sample wake-up voice;
[0021] Using the sample wake-up speech and the target loss function, the initial external wake-up suppression model is trained to obtain the external wake-up suppression model.
[0022] In one possible implementation, the initial external wake-up suppression model is a convolutional neural network (CNN); the initial external wake-up suppression model includes a feature extraction layer, a convolutional layer, and a pooling layer.
[0023] In one possible implementation, the method further includes:
[0024] Obtain the verification wake-up voice;
[0025] The verification wake-up voice is input into the external wake-up suppression model to obtain a prediction result of whether the verification wake-up voice belongs to the external voice of the vehicle cabin;
[0026] When the predicted result of whether the verification wake-up voice belongs to the vehicle cabin external voice is inconsistent with the actual result corresponding to the verification wake-up voice, the verification wake-up voice is used again as the sample wake-up voice to update the cabin external wake-up suppression model.
[0027] In one possible implementation, determining the feature parameters of the target wake-up voice based on the wake-up region where the target user is located, and determining the second determination result based on the feature parameters, includes:
[0028] Based on the wake-up area where the target user is located, determine the preset frequency band corresponding to the target wake-up voice;
[0029] Calculate the energy value of the preset frequency band corresponding to the target wake-up voice; and determine the second determination result based on the comparison result of the energy value and the preset energy threshold.
[0030] In one possible implementation, the placeholder information for determining the wake-up position includes:
[0031] Acquire vehicle visual information and seat pressure information; and determine the occupancy information of the wake-up position based on the vehicle visual information and seat pressure information.
[0032] In one possible implementation, determining the wake-up result corresponding to the target wake-up voice based on the first determination result, the second determination result, and the placeholder information includes:
[0033] When the wake-up area where the target user is located is the first area, the first determination result and the second determination result are both determined to be that the target wake-up voice belongs to the voice outside the vehicle cabin, and the content of the occupancy information is no one, the wake-up result corresponding to the target wake-up voice is determined to be suppression.
[0034] Alternatively, when the wake-up area where the target user is located is the second area, both the first and second determination results indicate that the target wake-up voice belongs to the external voice of the vehicle cabin, and the content of the placeholder information is no one or unavailable, the wake-up result corresponding to the target wake-up voice is determined to be suppression; and when the wake-up area where the target user is located is the second area, if the content of the placeholder information is no one, and one of the first and second determination results indicates that the target wake-up voice belongs to the external voice of the vehicle cabin, and the other determination result indicates that the target wake-up voice belongs to the internal voice of the vehicle cabin, then the wake-up result corresponding to the target wake-up voice is determined to be suppression.
[0035] Alternatively, when the wake-up area where the target user is located is the third area, both the first and second determination results indicate that the target wake-up voice belongs to the external voice of the vehicle cabin, and the content of the placeholder information indicates that someone is present, no one is present, or it is unavailable, the wake-up result corresponding to the target wake-up voice is determined to be suppression; and when the wake-up area where the target user is located is the third area, if the content of the placeholder information indicates that no one is present, and one of the first and second determination results indicates that the target wake-up voice belongs to the external voice of the vehicle cabin, and the other determination result indicates that the target wake-up voice belongs to the internal voice of the vehicle cabin, then the wake-up result corresponding to the target wake-up voice is determined to be suppression.
[0036] Alternatively, when the wake-up area where the target user is located is the fourth area, both the first and second determination results indicate that the target wake-up voice belongs to the external voice of the vehicle cabin, and the content of the placeholder information indicates that someone is present, no one is present, or it is unavailable, the wake-up result corresponding to the target wake-up voice is determined to be suppression; and when the wake-up area where the target user is located is the fourth area, if the content of the placeholder information indicates that no one is present, and one of the first and second determination results indicates that the target wake-up voice belongs to the external voice of the vehicle cabin, and the other determination result indicates that the target wake-up voice belongs to the internal voice of the vehicle cabin, then the wake-up result corresponding to the target wake-up voice is determined to be suppression.
[0037] In one possible implementation, the target wake-up speech is speech containing a wake-up word, wake-up-free speech, or speech containing a preset domain word in the recognition result.
[0038] This application also provides a vehicle voice interaction device, including:
[0039] The first acquisition unit is used to acquire the target wake-up voice sent by the target user; and determine the wake-up area where the target user is located based on the target wake-up voice;
[0040] The input unit is used to input the target wake-up voice into the external wake-up suppression model to obtain a first determination result;
[0041] The first determining unit is configured to determine the feature parameters of the target wake-up voice based on the wake-up area where the target user is located; and determine the second determination result based on the feature parameters.
[0042] The second determining unit is used to determine the placeholder information of the wake-up position;
[0043] The third determining unit is used to determine the wake-up result corresponding to the target wake-up voice based on the first determination result, the second determination result, and the placeholder information.
[0044] In one possible implementation, the wake-up area where the target user is located is one of the preset wake-up areas; the preset wake-up areas include a first area, a second area, a third area, and a fourth area; the device further includes:
[0045] The division unit is used to establish a coordinate system with the center of the ground where the front of the vehicle is located as the origin, take the center line of the vehicle in the length direction as the central axis, and symmetrically divide the first region, the second region, the third region and the fourth region on both sides from the central axis along the width direction of the vehicle.
[0046] In one possible implementation, the first acquisition unit is specifically used for:
[0047] The target wake-up voice is localized to determine the coordinate information of the target user based on the localization result; and the wake-up area of the target user is determined based on the coordinate information of the target user.
[0048] In one possible implementation, the input unit includes:
[0049] The input subunit is used to input the target wake-up speech into the external wake-up suppression model to obtain the acoustic features of the target wake-up speech;
[0050] The processing subunit is used to perform convolution and pooling processing on the acoustic features of the target wake-up speech to obtain the probability value that the target wake-up speech belongs to the speech outside the vehicle cabin.
[0051] The comparison subunit is used to compare the probability value with a preset probability threshold and determine the first judgment result based on the comparison result.
[0052] In one possible implementation, the device further includes:
[0053] The second acquisition unit is used to acquire sample wake-up speech;
[0054] The training unit is used to train the initial external wake-up suppression model using the sample wake-up speech and the target loss function to obtain the external wake-up suppression model.
[0055] In one possible implementation, the initial external wake-up suppression model is a convolutional neural network (CNN); the initial external wake-up suppression model includes a feature extraction layer, a convolutional layer, and a pooling layer.
[0056] In one possible implementation, the device further includes:
[0057] The third acquisition unit is used to acquire the verification wake-up voice;
[0058] The obtaining unit is used to input the verification wake-up voice into the external wake-up suppression model to obtain a prediction result of whether the verification wake-up voice belongs to the external voice of the vehicle cabin;
[0059] The update unit is used to update the external wake-up suppression model when the predicted judgment result of whether the verification wake-up voice belongs to the vehicle cabin external voice is inconsistent with the actual judgment result corresponding to the verification wake-up voice, and re-use the verification wake-up voice as the sample wake-up voice.
[0060] In one possible implementation, the first determining unit is specifically used for:
[0061] A determining subunit is used to determine the preset frequency band corresponding to the target wake-up voice based on the wake-up area where the target user is located;
[0062] The calculation subunit is used to calculate the energy value of the preset frequency band of the target wake-up voice; and to determine the second determination result based on the comparison result of the energy value and the preset energy threshold.
[0063] In one possible implementation, the second determining unit is specifically used for:
[0064] Acquire vehicle visual information and seat pressure information; and determine the occupancy information of the wake-up position based on the vehicle visual information and seat pressure information.
[0065] In one possible implementation, the third determining unit is specifically used for:
[0066] When the wake-up area where the target user is located is the first area, the first determination result and the second determination result are both determined to be that the target wake-up voice belongs to the voice outside the vehicle cabin, and the content of the occupancy information is no one, the wake-up result corresponding to the target wake-up voice is determined to be suppression.
[0067] Alternatively, when the wake-up area where the target user is located is the second area, both the first and second determination results indicate that the target wake-up voice belongs to the external voice of the vehicle cabin, and the content of the placeholder information is no one or unavailable, the wake-up result corresponding to the target wake-up voice is determined to be suppression; and when the wake-up area where the target user is located is the second area, if the content of the placeholder information is no one, and one of the first and second determination results indicates that the target wake-up voice belongs to the external voice of the vehicle cabin, and the other determination result indicates that the target wake-up voice belongs to the internal voice of the vehicle cabin, then the wake-up result corresponding to the target wake-up voice is determined to be suppression.
[0068] Alternatively, when the wake-up area where the target user is located is the third area, both the first and second determination results indicate that the target wake-up voice belongs to the external voice of the vehicle cabin, and the content of the placeholder information indicates that someone is present, no one is present, or it is unavailable, the wake-up result corresponding to the target wake-up voice is determined to be suppression; and when the wake-up area where the target user is located is the third area, if the content of the placeholder information indicates that no one is present, and one of the first and second determination results indicates that the target wake-up voice belongs to the external voice of the vehicle cabin, and the other determination result indicates that the target wake-up voice belongs to the internal voice of the vehicle cabin, then the wake-up result corresponding to the target wake-up voice is determined to be suppression.
[0069] Alternatively, when the wake-up area where the target user is located is the fourth area, both the first and second determination results indicate that the target wake-up voice belongs to the external voice of the vehicle cabin, and the content of the placeholder information indicates that someone is present, no one is present, or it is unavailable, the wake-up result corresponding to the target wake-up voice is determined to be suppression; and when the wake-up area where the target user is located is the fourth area, if the content of the placeholder information indicates that no one is present, and one of the first and second determination results indicates that the target wake-up voice belongs to the external voice of the vehicle cabin, and the other determination result indicates that the target wake-up voice belongs to the internal voice of the vehicle cabin, then the wake-up result corresponding to the target wake-up voice is determined to be suppression.
[0070] In one possible implementation, the target wake-up speech is speech containing a wake-up word, wake-up-free speech, or speech containing a preset domain word in the recognition result.
[0071] This application also provides a vehicle voice interaction device, including: a processor, a memory, and a system bus;
[0072] The processor and the memory are connected via the system bus;
[0073] The memory is used to store one or more programs, the one or more programs including instructions, which, when executed by the processor, cause the processor to perform any of the above-described implementations of the vehicle voice interaction method.
[0074] This application also provides a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to perform any of the above-described implementations of the vehicle voice interaction method.
[0075] This application also provides a computer program product, which, when run on a terminal device, causes the terminal device to execute any of the above-described vehicle voice interaction methods.
[0076] This application provides a vehicle voice interaction method, apparatus, storage medium, and device. First, it acquires a target wake-up voice issued by a target user; then, based on the target wake-up voice, it determines the wake-up area where the target user is located; next, it inputs the target wake-up voice into an external wake-up suppression model to obtain a first determination result; and based on the wake-up area where the target user is located, it determines feature parameters of the target wake-up voice; and based on these feature parameters, it determines a second determination result; then, it determines the occupancy information of the wake-up position; and finally, based on the first determination result, the second determination result, and the occupancy information, it determines the wake-up result corresponding to the target wake-up voice.
[0077] As can be seen, this application utilizes multi-dimensional judgment results such as the external wake-up suppression model, the calculation results of the characteristic parameters of the target wake-up voice, and the occupancy information of the wake-up position to more accurately determine whether the target wake-up voice issued by the target user is an external voice from outside the vehicle cabin, and then performs the corresponding wake-up or suppression interactive operation, thereby improving the accuracy of external wake-up suppression and thus enhancing the interactive experience of the target user. Attached Figure Description
[0078] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0079] Figure 1 A flowchart illustrating a vehicle voice interaction method provided in an embodiment of this application;
[0080] Figure 2 A schematic diagram showing the locations of the weak suppression region, the normal suppression region, the strict suppression region, and the wake-up restriction region provided in the embodiments of this application;
[0081] Figure 3 This is a schematic diagram illustrating the composition of a vehicle voice interaction device provided in an embodiment of this application. Detailed Implementation
[0082] In the automotive field, external wake-up suppression refers to preventing external sounds (such as shouts or noises) from mistakenly activating the vehicle's voice assistant or control systems, thereby avoiding potential security risks or misoperations. This technology is crucial because it relates to vehicle security and privacy. For example, if an outsider could control the vehicle (such as opening windows or the trunk) with voice commands while it is locked, this would pose a security risk.
[0083] However, current methods for suppressing external wake-up calls do not consider the speaker's location information. This may lead to the incorrect suppression of some normally used wake-up commands from users inside the vehicle cabin, thus reducing the accuracy of external wake-up suppression and affecting the user's interactive experience. For example, when a user rests their head on the rear window, their voice requesting to open the passenger window may be recognized as an external wake-up command and suppressed.
[0084] Therefore, improving the accuracy of external wake-up suppression to reduce the loss of user experience in daily interactions within the vehicle due to external wake-up suppression is an urgent problem to be solved.
[0085] To address the aforementioned deficiencies, this application provides a vehicle voice interaction method. First, a target wake-up voice issued by a target user is acquired. Then, based on the target wake-up voice, the wake-up area where the target user is located is determined. Next, the target wake-up voice is input into an external wake-up suppression model to obtain a first determination result. Based on the wake-up area where the target user is located, feature parameters of the target wake-up voice are determined. Based on these feature parameters, a second determination result is determined. Then, the occupancy information of the wake-up position is determined. Finally, based on the first determination result, the second determination result, and the occupancy information, the wake-up result corresponding to the target wake-up voice is determined.
[0086] As can be seen, this application utilizes multi-dimensional judgment results such as the external wake-up suppression model, the calculation results of the characteristic parameters of the target wake-up voice, and the occupancy information of the wake-up position to more accurately determine whether the target wake-up voice issued by the target user is an external voice from outside the vehicle cabin, and then performs the corresponding wake-up or suppression interactive operation, thereby improving the accuracy of external wake-up suppression and thus enhancing the interactive experience of the target user.
[0087] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0088] First Embodiment
[0089] See Figure 1 This is a flowchart illustrating a vehicle voice interaction method provided in this embodiment. The method includes the following steps:
[0090] S101: Obtain the target wake-up voice sent by the target user; and determine the wake-up area where the target user is located based on the target wake-up voice.
[0091] In this embodiment, any voice data requiring determination of whether it belongs to the external voice of the vehicle cabin is defined as the target wake-up voice, and the user issuing the target wake-up voice is defined as the target user. It should be noted that this embodiment does not limit the language type of the target wake-up voice. For example, the target wake-up voice can be spoken by the target user in Mandarin or a dialect, and the target wake-up voice can be composed of Chinese or English. Furthermore, this embodiment does not limit the length of the target wake-up voice; for example, the target wake-up voice can be a sentence or a paragraph.
[0092] Understandably, the target wake-up voice can be obtained as needed, such as by collecting audio data through the vehicle's microphone. After acquiring the target wake-up voice, existing or future voice keyword detection methods are used to detect it. If no preset wake-up word is detected, interaction with the corresponding interactive device in the vehicle's cabin, such as playing music, can be directly achieved based on the target wake-up voice. Conversely, if no preset wake-up word is detected, it is necessary to further determine the wake-up area of the target user to execute subsequent steps S102-S105, determining whether the target wake-up voice belongs to an external voice within the vehicle's cabin.
[0093] In this application, the specific content of the preset wake-up word is not limited. It can be set according to the actual situation and experience value. For example, the preset wake-up word can be set to a fixed wake-up word "ideal classmate", or it can be set to a fixed vehicle control command "open the car window", or it can be a request / command of vehicle control intent that appears in the recognition result after voice recognition, such as "open the car window for me".
[0094] Specifically, one possible implementation is that, after acquiring the target wake-up voice, existing or future sound source localization methods can be used to locate the sound source of the target wake-up voice, and the coordinate information of the target user can be determined based on the localization results. Figure 2 The estimated position coordinates [x, y, z] of the target user are shown in the three-dimensional coordinate system. Then, the wake-up area of the target user can be determined based on the target user's coordinate information.
[0095] The wake-up area where the target user is located is usually one of the preset wake-up areas. This application does not limit the content and division method of the preset wake-up area, and it can be set according to the actual situation and experience. One optional implementation is that the preset wake-up area can include, but is not limited to, a first area, a second area, a third area, and a fourth area. The construction method of the preset wake-up area can be, but is not limited to: establishing a coordinate system with the center ground where the front of the vehicle is located as the origin, taking the vehicle's centerline along the length of the vehicle as the central axis, and symmetrically dividing the first area, second area, third area, and fourth area to both sides from the central axis along the width of the vehicle. For example, taking the first area, second area, third area, and fourth area as a weak suppression area, a normal suppression area, a strict suppression area, and a restricted wake-up area, respectively, as an example... Figure 2 As shown. The preset wake-up area can be constructed in the following ways, but is not limited to: establishing a coordinate system with the center of the ground where the front of the car is located as the origin, and taking the center line of the vehicle along the length of the car body as the central axis (e.g., Figure 2 The vehicle's centerline (where the X-axis is located) is used as a guideline. Starting from this centerline, the vehicle's width is used to symmetrically divide the area into four regions: weak suppression region, normal suppression region, strict suppression region, and restricted wake-up region.
[0096] In this implementation, it should be noted that the division values of the four different regions in the preset wake-up area can be different for different car models. However, it is generally necessary to ensure that most regions (e.g., over 95%) are distributed in the first region (e.g., the weak suppression region) and the second region (e.g., the normal suppression region), while the third region (e.g., the strict suppression region) must include the car door. For example, on a certain car model, the horizontal axis of the division point of each region (e.g., Figure 2 The Y-axis coordinates can be {0, 150, 500, 650}. Thus, after determining the target user's coordinate information [x, y, z], the user can be assigned to a corresponding wake-up region (such as a weak suppression region, a normal suppression region, a strict suppression region, or a restricted wake-up region) based on the value of y in the coordinate information.
[0097] S102: Input the target wake-up voice into the external wake-up suppression model to obtain the first judgment result.
[0098] In this embodiment, after obtaining the target wake-up voice issued by the target user in step S101, in order to improve the accuracy of judging whether the target wake-up voice is external voice, the target wake-up voice can be further input into a pre-constructed external wake-up suppression model to predict a first determination result, which is then used to execute the subsequent step S105. The first determination result is used to determine whether the target wake-up voice belongs to external voices in the vehicle cabin.
[0099] Specifically, one possible implementation is that after acquiring the target wake-up voice from the target user, the target wake-up voice can first be input into a pre-constructed external wake-up suppression model to obtain the acoustic features of the target wake-up voice. These acoustic features refer to feature data used to characterize the voiceprint information of the corresponding speech frame in the target wake-up voice. For example, they could be Mel-scale Frequency Cepstral Coefficients (MFCC) features or Log Mel-filterbank (FBANK) features. It should be noted that this application does not limit the method for extracting the acoustic features of the target wake-up voice, nor does it limit the specific extraction process. Appropriate extraction methods and corresponding feature extraction operations can be selected according to the actual situation. For example, the target wake-up voice can first be segmented into frames to obtain the corresponding speech frame sequence, and then the segmented speech frame sequence can be pre-emphasized; subsequently, the acoustic features of each speech frame can be extracted sequentially.
[0100] Then, the acoustic features (such as FBANK) of the target wake-up speech can be convolved and pooled to obtain the probability value that the target wake-up speech belongs to the speech outside the vehicle cabin (i.e., the probability value that needs to be suppressed outside the field). Next, this probability value (usually ranging from 0 to 1) can be compared with a preset probability threshold, and based on the comparison result, the determination result of whether the target wake-up speech belongs to the speech outside the vehicle cabin (i.e., the first determination result) can be made.
[0101] In this application, the specific value of the preset probability threshold is not limited. It can be set according to the actual situation and experience. For example, the preset probability threshold can be set to 0.4. In this way, when the probability value of the target wake-up voice belonging to the voice outside the vehicle cabin is 0.8, it is greater than the preset probability threshold of 0.4. Therefore, the first determination result is that the target wake-up voice belongs to the voice outside the vehicle cabin.
[0102] It should be noted that this application does not limit the specific structure and processing procedure of the external wake-up suppression model mentioned in step S102, and can set it according to the actual situation and empirical values. Furthermore, the purpose of the pre-constructed external wake-up suppression model in this application is to quickly and accurately calculate the probability of whether the target wake-up voice belongs to the external voice of the vehicle cabin. Then, by comparing it with a preset probability threshold, the comparison result can determine whether the target wake-up voice belongs to the external voice of the vehicle cabin.
[0103] Next, the implementation process of the pre-constructed extravehicular wake-up suppression model in this application will be introduced. One optional implementation method is that the construction process of the extravehicular wake-up suppression model may specifically include: firstly, obtaining sample wake-up voices issued by sample users, and then using the sample wake-up voices and the target loss function (the specific content is not limited and can be selected according to the actual situation and empirical values, such as the CTC (Connectionist Temporal Classification) loss constraint function, etc.) to train the initial extravehicular wake-up suppression model to obtain the extravehicular wake-up suppression model.
[0104] Specifically, in this implementation, a significant amount of preparatory work is required to construct the external wake-up suppression model. First, a large amount of voice data needs to be collected as sample wake-up voices to form the model training data. For example, a large amount of voice data from in-vehicle users or external users can be collected beforehand, such as "open the window" wake-up voice from an in-vehicle user and "open the trunk" wake-up voice from an external user. These can all be used as sample wake-up voices to form the model training data, and the actual judgment results corresponding to these sample wake-up voices are manually labeled to indicate whether they belong to external voices in the vehicle cabin. Next, based on these sample wake-up voices, the corresponding actual judgment results, and the target loss function (such as the CTC loss constraint function), the initial external wake-up suppression model can be trained, thereby generating the external wake-up suppression model.
[0105] One possible implementation is that the external wake-up suppression model can be (but is not limited to) a Convolutional Neural Network (CNN) model. This model can include feature extraction layers, convolutional layers, and pooling layers, with no limit on the number of each layer. For example, the model could include a feature extraction layer to extract the acoustic features (such as FBank features) of the sample wake-up speech, and then, after three to five convolutional layers and one pooling layer, output a probability value representing the probability that the predicted wake-up speech belongs to external vehicle cabin speech.
[0106] Specifically, during model training, a sample wake-up voice can be extracted from the training data as the model input, and the actual judgment result of whether it belongs to the external voice of the vehicle cabin can be used as the output. Multiple rounds of model training are performed, and the prediction results obtained in each round of training are compared with the corresponding manually labeled results. The model parameters are updated according to the difference between the two until the preset conditions are met, such as the target loss function (such as the CTC loss constraint function) is very small and basically unchanged. Then the update of the model parameters is stopped, the training of the external wake-up suppression model is completed, and a trained external wake-up suppression model is generated.
[0107] Based on this, after training and generating an external wake-up suppression model using sample wake-up voices, the generated external wake-up suppression model can be further validated using verification wake-up voices. The specific validation process may include the following steps (1)-(3):
[0108] Step (1): Obtain the verification wake-up voice sent by the verification user.
[0109] In this embodiment, in order to verify the external wake-up suppression model, it is first necessary to obtain the verification wake-up voice issued by the verification user. For example, with the user's permission, 1,000 wake-up voice data spoken by different users inside and outside the vehicle can be collected as verification wake-up voices. The verification wake-up voices refer to the voice information that can be used to verify the external wake-up suppression model. After obtaining these verification wake-up voices and the actual judgment label of each verification wake-up voice as belonging to the external voice of the vehicle cabin, the subsequent steps (2) can be continued.
[0110] Step (2): Input the verification wake-up voice into the external wake-up suppression model to obtain the prediction result of whether the verification wake-up voice belongs to the external voice of the vehicle cabin.
[0111] After obtaining the verification wake-up voice issued by the verification user in step (1), the verification wake-up voice can be further input into the external wake-up suppression model to obtain the prediction result of whether the verification wake-up voice belongs to the external voice of the vehicle cabin, so as to execute the subsequent step (3).
[0112] Step (3): When the predicted judgment result of whether the verification wake-up voice belongs to the vehicle cabin external voice is inconsistent with the actual judgment result corresponding to the verification wake-up voice, the verification wake-up voice is used as the sample wake-up voice again to update the cabin external wake-up suppression model.
[0113] After obtaining the predicted judgment result of whether the verification wake-up voice belongs to the external voice of the vehicle cabin through step (2), if the predicted judgment result is inconsistent with the actual judgment result of whether the verification wake-up voice belongs to the external voice of the vehicle cabin (such as the actual result of manual annotation), the verification wake-up voice can be used again as the sample wake-up voice to adjust and update the parameters of the external wake-up suppression model in a timely manner. This helps to improve the prediction accuracy and precision of the external wake-up suppression model.
[0114] S103: Determine the feature parameters of the target wake-up voice based on the wake-up area where the target user is located; and determine the second judgment result based on the feature parameters.
[0115] In this embodiment, after obtaining the target wake-up voice emitted by the target user in step S101 and determining the wake-up area where the target user is located based on the target wake-up voice, in order to improve the accuracy of judging whether the target wake-up voice is external voice, it is not only necessary to input the target wake-up voice into a pre-constructed external wake-up suppression model to obtain a first judgment result (used to determine whether the target wake-up voice belongs to the vehicle cabin external voice), but also to extract the feature parameters of the target wake-up voice based on the wake-up area where the target user is located, using existing or future methods for extracting voice features. This application does not limit the specific content of the feature parameters and can set them according to actual conditions and empirical values. For example, the feature parameters can be set as the energy value of the extracted target wake-up voice, in joules (J); or the feature parameters can be set as the loudness value of the extracted target wake-up voice, in decibels (dB). Then, based on the feature parameters, a second judgment result can be determined to execute the subsequent step S105. The second judgment result is used to determine whether the target wake-up voice belongs to the vehicle cabin external voice.
[0116] Specifically, one possible implementation is that, after acquiring the target wake-up voice issued by the target user and determining the wake-up area where the target user is located, the preset frequency band corresponding to the target wake-up voice can be determined based on the wake-up area where the target user is located (the specific value is not limited, and the preset frequency band values corresponding to different wake-up areas can be the same or different). Then, the energy value of the target wake-up voice in the preset frequency band corresponding to the wake-up area is calculated. Then, based on the comparison result of the energy value and the preset energy threshold, the determination result of whether the target wake-up voice belongs to the external voice of the vehicle cabin is more accurate (that is, the second determination result).
[0117] In this implementation, the application does not limit the preset frequency bands corresponding to different wake-up regions, and can dynamically select them according to actual conditions and empirical values. For example, the preset frequency band corresponding to the first region (such as the weak suppression region) can be 2500Hz-8000Hz, the preset frequency band corresponding to the third region (such as the strict suppression region) can be 3500Hz-8000Hz, the preset frequency band corresponding to the second region (such as the normal suppression region) can be 2500Hz-8000Hz, and the preset frequency band corresponding to the fourth region (such as the restricted wake-up region) can be 2500Hz-8000Hz, etc.
[0118] Furthermore, this application does not limit the specific value of the preset energy threshold. It can be set according to the actual situation and experience. For example, the preset energy threshold can be set to 75. In this way, when the calculated energy value of the preset frequency band of the target wake-up voice is 72, it is less than the preset energy threshold of 75. Therefore, the second determination result is that the target wake-up voice belongs to the voice outside the vehicle cabin.
[0119] S104: Determine the placeholder information for the wake-up position.
[0120] In this embodiment, in order to improve the accuracy of determining whether the target wake-up voice is an external voice, it is not only necessary to input the target wake-up voice into the pre-constructed external wake-up suppression model in step S102 to obtain a first determination result (used to determine whether the target wake-up voice belongs to the vehicle's external voice), and to determine the characteristic parameters (such as energy value) of the target wake-up voice in step S103 to determine a second determination result (used to determine whether the target wake-up voice belongs to the vehicle's external voice), but also to obtain the vehicle's visual information and seat pressure information at the wake-up time (i.e., the time when the target user finishes saying the wake-up word) through an in-vehicle vision device (such as a camera) (e.g., seat pressure information can be obtained through a gravity sensor pre-installed on the seat). Then, based on the vehicle's visual information and seat pressure information, the occupancy information of the wake-up position is determined for the subsequent step S105.
[0121] The wake-up location refers to the sound zone of the seat, such as the driver's seat or the passenger seat. Each seat has an independent microphone pre-installed on it. The microphone can determine the location from which the sound is coming from based on the audio energy received by different microphones, and thus determine the location of the target user from which the wake-up voice is sent.
[0122] Specifically, in determining the occupancy information of the wake-up location, this embodiment adopts the following method: First, it is determined whether there is an action of opening or closing a car door (focusing on closing the car door). If so, the occupancy information of N-1 wake-up locations (where N is not limited and can be 10, etc., so N-1 is 9) recorded over a period of time (e.g., 9 seconds) can be reset to "no one". That is, the state at the last moment is retained, and the occupancy information of the wake-up location (i.e., "occupied", "no one" or "unavailable") can be updated with the instantaneous state (referring to the moment of wake-up, i.e., the time when the target user finishes saying the wake-up word).
[0123] Conversely, if no action occurs (specifically, closing the car door), the visual information can be determined first. If it is unavailable, the occupancy information of the wake-up location can be set as "unavailable." If the visual information is available, it can be used to determine whether there is someone at the wake-up location. For example, the video or image taken by a pre-installed camera in the car can be used to determine whether there is someone at the wake-up location. If the visual information indicates that there is someone at the wake-up location, the occupancy information of the wake-up location can be set as "occupied." If the visual information indicates that there is no one at the wake-up location, there may be a missed detection. The visual information is updated once per second, but other judgment frequencies can also be used. This application does not limit this. At this point, seat pressure information can be used for further judgment. For example, seat pressure information can be obtained through a gravity sensor pre-installed on the seat to determine whether the seat's load-bearing capacity exceeds a preset weight (e.g., 30kg, the specific value is not limited). If it does not exceed the preset weight, it can be determined that there is no one at the wake-up location. To improve the accuracy of the judgment, the judgment can be performed for a period of time (e.g., 10s). If both visual information and seat pressure information indicate that there is no one at the wake-up location during this period, the occupancy information of the wake-up location can be determined as "no one". However, if the seat pressure information indicates that the seat's load-bearing capacity exceeds the preset weight (e.g., 30kg), the occupancy information of the wake-up location can be determined as "occupied". This process can be repeated to determine the occupancy information of the wake-up location in real time based on vehicle visual information and seat pressure information (specifically, "occupied", "no one", or "unavailable").
[0124] S105: Based on the first judgment result, the second judgment result, and the placeholder information, determine the wake-up result corresponding to the target wake-up voice.
[0125] In this embodiment, step S102 inputs the target wake-up voice into a pre-built external wake-up suppression model to obtain a first determination result (used to determine whether the target wake-up voice belongs to the vehicle cabin external voice). Step S103 determines the characteristic parameters (such as energy value) of the target wake-up voice to determine a second determination result (used to determine whether the target wake-up voice belongs to the vehicle cabin external voice). Step S104 determines the occupancy information of the wake-up position (specifically, "occupied", "unoccupied", or "unavailable") based on vehicle visual information and seat pressure information. Then, based on the first determination result, the second determination result, and the occupancy information, the wake-up result corresponding to the target wake-up voice, i.e., "wake-up" or "suppression", is determined by querying the pre-built Table 1 below. Then, based on the obtained wake-up result, interactive operation with the corresponding interactive device in the vehicle cabin can be realized. That is, when the wake-up result is "wake-up", the corresponding interactive device in the vehicle cabin is allowed to respond to the target wake-up voice signal issued by the target user and enter subsequent interactive operation, such as opening the car window; otherwise, when the wake-up result is "suppression", the target wake-up voice issued by the current target user is blocked and no response is made.
[0126]
[0127] Table 1
[0128] Among them, when the wake-up area where the target user is located is the first area (i.e., the weak suppression area shown in Table 1), and both the first and second judgment results determine that the target wake-up voice belongs to the voice outside the vehicle cabin and the content of the occupant information is "no one", the wake-up result corresponding to the target wake-up voice can be determined to be suppression.
[0129] Alternatively, when the wake-up area where the target user is located is the second area (i.e., the conventional suppression area shown in Table 1), and both the first and second determination results indicate that the target wake-up voice belongs to the external voice of the vehicle cabin, and the content of the placeholder information is "no one" or "unavailable", the wake-up result corresponding to the target wake-up voice can be determined to be suppression; and when the wake-up area where the target user is located is the second area (i.e., the conventional suppression area shown in Table 1), if the content of the placeholder information is "no one", and one of the first and second determination results indicates that the target wake-up voice belongs to the external voice of the vehicle cabin, and the other determination result indicates that the target wake-up voice belongs to the internal voice of the vehicle cabin, then the wake-up result corresponding to the target wake-up voice can be determined to be suppression.
[0130] Alternatively, when the wake-up area where the target user is located is the third area (i.e., the strictly suppressed area shown in Table 1), and both the first and second determination results indicate that the target wake-up voice belongs to the external voice of the vehicle cabin, and the content of the placeholder information is "occupied", "unoccupied", or "unavailable", the wake-up result corresponding to the target wake-up voice can be determined to be suppressed; and when the wake-up area where the target user is located is the third area (i.e., the strictly suppressed area shown in Table 1), if the content of the placeholder information is "unoccupied", and one of the first and second determination results indicates that the target wake-up voice belongs to the external voice of the vehicle cabin, and the other determination result indicates that the target wake-up voice belongs to the internal voice of the vehicle cabin, then the wake-up result corresponding to the target wake-up voice can be determined to be suppressed.
[0131] Alternatively, when the wake-up area where the target user is located is the fourth area (i.e., the restricted wake-up area shown in Table 1), and both the first and second determination results indicate that the target wake-up voice belongs to the external voice of the vehicle cabin, and the content of the placeholder information is "occupied", "unoccupied", or "unavailable", the wake-up result corresponding to the target wake-up voice can be determined to be suppression; and when the wake-up area where the target user is located is the fourth area (i.e., the restricted wake-up area shown in Table 1), if the content of the placeholder information is "unoccupied", and one of the first and second determination results indicates that the target wake-up voice belongs to the external voice of the vehicle cabin, and the other determination result indicates that the target wake-up voice belongs to the internal voice of the vehicle cabin, then the wake-up result corresponding to the target wake-up voice can be determined to be suppression.
[0132] It should be noted that the correspondences shown in Table 1 above are all based on actual situations and experience values, and are statistically derived from existing real data. Adaptive adjustments and updates can also be made, and this application does not limit this.
[0133] In this way, the first judgment result (used to determine whether the target wake-up voice belongs to the vehicle's external voice) is obtained through the external wake-up suppression model. The second judgment result (used to determine whether the target wake-up voice belongs to the vehicle's external voice) is determined by extracting the feature parameters of the target wake-up voice (such as energy value). After determining the occupancy information of the wake-up position based on vehicle visual information and seat pressure information, the judgment information obtained from these multiple dimensions can more accurately determine whether the wake-up or control command (i.e., the target wake-up voice) comes from the occupants of the vehicle. This can prevent external personnel from controlling the doors, windows and other facilities, reduce the safety hazards brought by external personnel, and also reduce the loss of the daily use experience of the users inside the vehicle due to external suppression.
[0134] In addition, it should be noted that, on the one hand, the target wake-up voice mentioned in this application can be the Mandarin wake-up voice or dialect wake-up voice issued by the target user collected by the vehicle voice wake-up module, such as "Li Auto". If it contains a preset wake-up word, it needs to be judged and processed by executing the method described in steps S101-S105 above. If it does not contain a wake-up word, it can directly respond to the target wake-up voice signal and enter subsequent interactive operations, such as playing music. On the other hand, the target wake-up voice mentioned in this application can also be a wake-up-free voice obtained by the vehicle voice recognition module or a voice containing preset domain words in the recognition result. For example, the automatic speech recognition (ASR) module can identify that the voice request or / or command contains vehicle control intent, such as "open the car window for me", or identify a wake-up-free voice with a fixed format, such as "open the car window" (note that if there are other words before and after, it may not be considered a wake-up-free command). If preset domain words are involved (the specific content is not limited, but usually it is words related to vehicle safety such as car windows, doors, tailgate, etc.), they need to be judged and processed by the method described in steps S101-S105 above; otherwise, the target wake-up voice signal can be directly responded to and subsequent interactive operations can be entered, such as playing music.
[0135] In summary, the vehicle voice interaction method provided in this embodiment first acquires the target wake-up voice issued by the target user; then, based on the target wake-up voice, determines the wake-up area where the target user is located; then, inputs the target wake-up voice into a pre-constructed external wake-up suppression model to obtain a first determination result; and based on the wake-up area where the target user is located, determines the feature parameters of the target wake-up voice; and based on the feature parameters, determines a second determination result; next, determines the occupancy information of the wake-up position; and then, based on the first determination result, the second determination result, and the occupancy information, determines the wake-up result corresponding to the target wake-up voice.
[0136] As can be seen, this application utilizes a pre-built external wake-up suppression model, the calculation results of the characteristic parameters of the target wake-up voice, and the occupancy information of the wake-up position, among other multi-dimensional judgment results, to more accurately determine whether the target wake-up voice issued by the target user is an external voice from outside the vehicle cabin, and then performs the corresponding wake-up or suppression interactive operation, thereby improving the accuracy of external wake-up suppression and thus enhancing the interactive experience of the target user.
[0137] Second Embodiment
[0138] This embodiment will introduce a vehicle voice interaction device; please refer to the above method embodiment for related content.
[0139] See Figure 3This is a schematic diagram of the composition of a vehicle voice interaction device provided in this embodiment. The device 300 includes:
[0140] The first acquisition unit 301 is used to acquire the target wake-up voice sent by the target user; and determine the wake-up area where the target user is located based on the target wake-up voice;
[0141] Input unit 302 is used to input the target wake-up voice into the external wake-up suppression model to obtain a first determination result;
[0142] The first determining unit 303 is used to determine the feature parameters of the target wake-up voice based on the wake-up area where the target user is located; and to determine the second determination result based on the energy value.
[0143] The second determining unit 304 is used to acquire vehicle visual information and seat pressure information; and to determine the occupancy information of the wake-up position based on the vehicle visual information and seat pressure information.
[0144] The third determining unit 305 is used to determine the wake-up result corresponding to the target wake-up voice based on the first determination result, the second determination result and the placeholder information.
[0145] In one implementation of this embodiment, the wake-up area where the target user is located is one of the preset wake-up areas; the preset wake-up areas include a first area, a second area, a third area, and a fourth area; the device further includes:
[0146] The division unit is used to establish a coordinate system with the center of the ground where the front of the vehicle is located as the origin, take the center line of the vehicle in the length direction as the central axis, and symmetrically divide the first region, the second region, the third region and the fourth region on both sides from the central axis along the width direction of the vehicle.
[0147] In one implementation of this embodiment, the first acquisition unit 301 is specifically used for:
[0148] The target wake-up voice is localized to determine the coordinate information of the target user based on the localization result; and the wake-up area of the target user is determined based on the coordinate information of the target user.
[0149] In one implementation of this embodiment, the input unit 302 includes:
[0150] The input subunit is used to input the target wake-up speech into the external wake-up suppression model to obtain the acoustic features of the target wake-up speech;
[0151] The processing subunit is used to perform convolution and pooling processing on the acoustic features of the target wake-up speech to obtain the probability value that the target wake-up speech belongs to the speech outside the vehicle cabin.
[0152] The comparison subunit is used to compare the probability value with a preset probability threshold and determine the first judgment result based on the comparison result.
[0153] In one implementation of this embodiment, the apparatus further includes:
[0154] The second acquisition unit is used to acquire sample wake-up speech;
[0155] The training unit is used to train the initial external wake-up suppression model using the sample wake-up speech and the target loss function to obtain the external wake-up suppression model.
[0156] In one implementation of this embodiment, the initial external wake-up suppression model is a convolutional neural network (CNN); the initial external wake-up suppression model includes a feature extraction layer, a convolutional layer, and a pooling layer.
[0157] In one implementation of this embodiment, the apparatus further includes:
[0158] The third acquisition unit is used to acquire the verification wake-up voice;
[0159] The obtaining unit is used to input the verification wake-up voice into the external wake-up suppression model to obtain a prediction result of whether the verification wake-up voice belongs to the external voice of the vehicle cabin;
[0160] The update unit is used to update the external wake-up suppression model when the predicted judgment result of whether the verification wake-up voice belongs to the vehicle cabin external voice is inconsistent with the actual judgment result corresponding to the verification wake-up voice, and re-use the verification wake-up voice as the sample wake-up voice.
[0161] In one implementation of this embodiment, the first determining unit 303 is specifically used for:
[0162] A determining subunit is used to determine the preset frequency band corresponding to the target wake-up voice based on the wake-up area where the target user is located;
[0163] The calculation subunit is used to calculate the energy value of the preset frequency band of the target wake-up voice; and to determine the second determination result based on the comparison result of the energy value and the preset energy threshold.
[0164] In one implementation of this embodiment, the second determining unit 304 is specifically used for:
[0165] Acquire vehicle visual information and seat pressure information; and determine the occupancy information of the wake-up position based on the vehicle visual information and seat pressure information.
[0166] In one implementation of this embodiment, the third determining unit 305 is specifically used for:
[0167] When the wake-up area where the target user is located is the first area, the first determination result and the second determination result are both determined to be that the target wake-up voice belongs to the voice outside the vehicle cabin, and the content of the occupancy information is no one, the wake-up result corresponding to the target wake-up voice is determined to be suppression.
[0168] Alternatively, when the wake-up area where the target user is located is the second area, both the first and second determination results indicate that the target wake-up voice belongs to the external voice of the vehicle cabin, and the content of the placeholder information is no one or unavailable, the wake-up result corresponding to the target wake-up voice is determined to be suppression; and when the wake-up area where the target user is located is the second area, if the content of the placeholder information is no one, and one of the first and second determination results indicates that the target wake-up voice belongs to the external voice of the vehicle cabin, and the other determination result indicates that the target wake-up voice belongs to the internal voice of the vehicle cabin, then the wake-up result corresponding to the target wake-up voice is determined to be suppression.
[0169] Alternatively, when the wake-up area where the target user is located is the third area, both the first and second determination results indicate that the target wake-up voice belongs to the external voice of the vehicle cabin, and the content of the placeholder information indicates that someone is present, no one is present, or it is unavailable, the wake-up result corresponding to the target wake-up voice is determined to be suppression; and when the wake-up area where the target user is located is the third area, if the content of the placeholder information indicates that no one is present, and one of the first and second determination results indicates that the target wake-up voice belongs to the external voice of the vehicle cabin, and the other determination result indicates that the target wake-up voice belongs to the internal voice of the vehicle cabin, then the wake-up result corresponding to the target wake-up voice is determined to be suppression.
[0170] Alternatively, when the wake-up area where the target user is located is the fourth area, both the first and second determination results indicate that the target wake-up voice belongs to the external voice of the vehicle cabin, and the content of the placeholder information indicates that someone is present, no one is present, or it is unavailable, the wake-up result corresponding to the target wake-up voice is determined to be suppression; and when the wake-up area where the target user is located is the fourth area, if the content of the placeholder information indicates that no one is present, and one of the first and second determination results indicates that the target wake-up voice belongs to the external voice of the vehicle cabin, and the other determination result indicates that the target wake-up voice belongs to the internal voice of the vehicle cabin, then the wake-up result corresponding to the target wake-up voice is determined to be suppression.
[0171] In one implementation of this embodiment, the target wake-up speech is speech containing a wake-up word, wake-up-free speech, or speech containing a preset domain word in the recognition result.
[0172] Furthermore, embodiments of this application also provide a vehicle voice interaction device, including: a processor, a memory, and a system bus;
[0173] The processor and the memory are connected via the system bus;
[0174] The memory is used to store one or more programs, the one or more programs including instructions, which, when executed by the processor, cause the processor to perform any of the above-described implementations of the vehicle voice interaction method.
[0175] Furthermore, embodiments of this application also provide a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to perform any of the above-described implementations of the vehicle voice interaction method.
[0176] Furthermore, this application also provides a computer program product, which, when run on a terminal device, causes the terminal device to execute any of the above-described implementation methods of the vehicle voice interaction method.
[0177] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps in the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0178] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0179] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0180] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A vehicle voice interaction method, characterized in that, include: Acquire the target wake-up voice sent by the target user; and determine the wake-up area where the target user is located based on the target wake-up voice; The target wake-up voice is input into the external wake-up suppression model to obtain the first determination result; Based on the wake-up area of the target user, the characteristic parameters of the target wake-up voice are determined; and based on the characteristic parameters, a second determination result is determined. Determine the placeholder information for the wake-up location; Based on the first determination result, the second determination result, and the placeholder information, the wake-up result corresponding to the target wake-up voice is determined.
2. The method according to claim 1, characterized in that, The wake-up area where the target user is located is one of the preset wake-up areas; the preset wake-up areas include a first area, a second area, a third area, and a fourth area; the preset wake-up areas are constructed as follows: A coordinate system is established with the center of the ground where the front of the vehicle is located as the origin. The center line of the vehicle along the length of the vehicle body is taken as the central axis. Starting from the central axis, the first region, the second region, the third region, and the fourth region are symmetrically divided to both sides along the width of the vehicle body.
3. The method according to claim 1, characterized in that, Determining the wake-up area of the target user based on the target wake-up voice includes: The target wake-up voice is localized to determine the coordinate information of the target user based on the localization result; and the wake-up area of the target user is determined based on the coordinate information of the target user.
4. The method according to claim 1, characterized in that, The step of inputting the target wake-up voice into the external wake-up suppression model to obtain a first determination result includes: The target wake-up voice is input into the external wake-up suppression model to obtain the acoustic features of the target wake-up voice; The acoustic features of the target wake-up speech are processed by convolution and pooling to obtain the probability value that the target wake-up speech belongs to the speech outside the vehicle cabin. The probability value is compared with a preset probability threshold, and a first determination result is determined based on the comparison result.
5. The method according to claim 1, characterized in that, The extravehicular wake-up suppression model is constructed as follows: Acquire sample wake-up voice; Using the sample wake-up speech and the target loss function, the initial external wake-up suppression model is trained to obtain the external wake-up suppression model.
6. The method according to claim 5, characterized in that, The method further includes: Obtain the verification wake-up voice; The verification wake-up voice is input into the external wake-up suppression model to obtain a prediction result of whether the verification wake-up voice belongs to the external voice of the vehicle cabin; When the predicted result of whether the verification wake-up voice belongs to the vehicle cabin external voice is inconsistent with the actual result corresponding to the verification wake-up voice, the verification wake-up voice is used again as the sample wake-up voice to update the cabin external wake-up suppression model.
7. The method according to claim 1, characterized in that, The step of determining the feature parameters of the target wake-up voice based on the wake-up region of the target user, and determining the second determination result based on the feature parameters, includes: Based on the wake-up area where the target user is located, determine the preset frequency band corresponding to the target wake-up voice; Calculate the energy value of the preset frequency band corresponding to the target wake-up voice; and determine the second determination result based on the comparison result of the energy value and the preset energy threshold.
8. The method according to claim 1, characterized in that, The placeholder information for determining the wake-up position includes: Acquire vehicle visual information and seat pressure information; and determine the occupancy information of the wake-up position based on the vehicle visual information and seat pressure information.
9. The method according to claim 2, characterized in that, The step of determining the wake-up result corresponding to the target wake-up voice based on the first determination result, the second determination result, and the placeholder information includes: When the wake-up area where the target user is located is the first area, the first determination result and the second determination result are both determined to be that the target wake-up voice belongs to the voice outside the vehicle cabin, and the content of the occupancy information is no one, the wake-up result corresponding to the target wake-up voice is determined to be suppression. Alternatively, when the wake-up area where the target user is located is the second area, both the first and second determination results indicate that the target wake-up voice belongs to the external voice of the vehicle cabin, and the content of the placeholder information is no one or unavailable, the wake-up result corresponding to the target wake-up voice is determined to be suppression; and when the wake-up area where the target user is located is the second area, if the content of the placeholder information is no one, and one of the first and second determination results indicates that the target wake-up voice belongs to the external voice of the vehicle cabin, and the other determination result indicates that the target wake-up voice belongs to the internal voice of the vehicle cabin, then the wake-up result corresponding to the target wake-up voice is determined to be suppression. Alternatively, when the wake-up area where the target user is located is the third area, both the first and second determination results indicate that the target wake-up voice belongs to the external voice of the vehicle cabin, and the content of the placeholder information indicates that someone is present, no one is present, or it is unavailable, the wake-up result corresponding to the target wake-up voice is determined to be suppression; and when the wake-up area where the target user is located is the third area, if the content of the placeholder information indicates that no one is present, and one of the first and second determination results indicates that the target wake-up voice belongs to the external voice of the vehicle cabin, and the other determination result indicates that the target wake-up voice belongs to the internal voice of the vehicle cabin, then the wake-up result corresponding to the target wake-up voice is determined to be suppression. Alternatively, when the wake-up area where the target user is located is the fourth area, both the first and second determination results indicate that the target wake-up voice belongs to the external voice of the vehicle cabin, and the content of the placeholder information indicates that someone is present, no one is present, or it is unavailable, the wake-up result corresponding to the target wake-up voice is determined to be suppression; and when the wake-up area where the target user is located is the fourth area, if the content of the placeholder information indicates that no one is present, and one of the first and second determination results indicates that the target wake-up voice belongs to the external voice of the vehicle cabin, and the other determination result indicates that the target wake-up voice belongs to the internal voice of the vehicle cabin, then the wake-up result corresponding to the target wake-up voice is determined to be suppression.
10. The method according to any one of claims 1-9, characterized in that, The target wake-up speech is speech containing a wake-up word, wake-up-free speech, or speech containing a preset domain word in the recognition result.
11. A vehicle voice interaction device, characterized in that, include: The first acquisition unit is used to acquire the target wake-up voice sent by the target user; and determine the wake-up area where the target user is located based on the target wake-up voice; The input unit is used to input the target wake-up voice into the external wake-up suppression model to obtain a first determination result; The first determining unit is used to determine the feature parameters of the target wake-up voice based on the wake-up area where the target user is located; And based on the aforementioned feature parameters, a second determination result is determined; The second determining unit is used to determine the placeholder information of the wake-up position; The third determining unit is used to determine the wake-up result corresponding to the target wake-up voice based on the first determination result, the second determination result, and the placeholder information.
12. A vehicle voice interaction device, characterized in that, include: Processor, memory, system bus; The processor and the memory are connected via the system bus; The memory is used to store one or more programs, the one or more programs including instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a terminal device, cause the terminal device to perform the method described in any one of claims 1-10.