Wakeup-word-free threshold adjustment method and apparatus, vehicle, and readable storage medium

By adjusting the wake-word threshold through real-time audio data matching and reverse operation, the problem of false triggering in intelligent in-vehicle voice interaction systems has been solved, improving user experience and system adaptability.

CN116052677BActive Publication Date: 2025-11-04SHANGHAI PATEO ELECTRONIC EQUIPMENT MANUFACTURING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211684825.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-27
Publication Date
2025-11-04
Estimated Expiration
2042-12-27

AI Technical Summary

Technical Problem

In intelligent in-vehicle voice interaction systems, voice content from conversations among passengers or in other scenarios may accidentally trigger wake-up-free phrases, leading to a poor user experience.

Method used

By acquiring real-time audio data and matching it with a preset comparison model, a similarity score is calculated. When the score reaches a threshold, a wake-word-free instruction is triggered. After detecting a reverse operation, the threshold is adjusted to avoid false wake-ups.

Benefits of technology

Dynamically adjusting the wake-word threshold improves user experience, avoids false triggers, and enhances the adaptability and user satisfaction of the in-vehicle voice interaction system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116052677B_ABST
    Figure CN116052677B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a wake-up word threshold adjustment method, device, vehicle and readable storage medium. The wake-up word threshold adjustment method comprises: obtaining voice data, the voice data being real-time audio data detected in the vehicle; matching the voice data with a preset comparison model to obtain a first similarity score; in a case where the first similarity score is greater than or equal to a first threshold corresponding to a first wake-up word, triggering a first instruction corresponding to the first wake-up word; and in a case where a reverse operation for the first instruction is detected, adjusting the first threshold corresponding to the first wake-up word. The dynamic adjustment of the pre-set wake-up word threshold in the vehicle is realized, and in the process of continuous dynamic adjustment of the wake-up word threshold, the wake-up word threshold is kept in line with the current use habit of the user, thereby improving the experience of the user using the vehicle voice interaction system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of vehicle control, and particularly relates to a wake-up word threshold adjustment method and device, a vehicle and a readable storage medium. BACKGROUND

[0002] With the continuous development of intelligent vehicle technology, an intelligent vehicle voice interaction system can execute an operation corresponding to a voice instruction input by a user, such as "turn on music" and "turn on air conditioner", without the need for the user to manually set the above-mentioned instruction operation.

[0003] Meanwhile, in order to improve user experience, a plurality of fixed wake-up word items are generally set in advance, and when it is recognized that the user inputs any item in the wake-up word items through voice, the intelligent vehicle voice interaction system can execute an operation instruction corresponding to the item in a wake-up state.

[0004] However, in the intelligent vehicle voice interaction system, due to the fact that the voice content in the conversation of the people in the vehicle or other scenes often contains some wake-up word items, when the user inadvertently utters voice content containing a wake-up word item, the vehicle will be directly triggered to execute the instruction, resulting in a poor user experience. SUMMARY

[0005] The present application provides a wake-up word threshold adjustment method, device, vehicle and readable storage medium to solve the problem of false triggering of wake-up word items.

[0006] To solve the above technical problems, the present application is implemented as follows:

[0007] In a first aspect, the present application provides a wake-up word threshold adjustment method, which comprises:

[0008] Obtaining voice data, the voice data being real-time audio data detected in the vehicle;

[0009] Matching the voice data with a preset comparison model to obtain a first similarity score;

[0010] In the case where the first similarity score is greater than or equal to a first threshold corresponding to a first wake-up word, triggering a first instruction corresponding to the first wake-up word;

[0011] In the case where a reverse operation for the first instruction is detected, adjusting the first threshold corresponding to the first wake-up word.

[0012] In a second aspect, the present application provides a wake-up word threshold adjustment device, which comprises:

[0013] acquire voice data, the voice data being real-time audio data detected in the vehicle;

[0014] match the voice data with a preset comparison model to obtain a first similarity score;

[0015] trigger a first instruction corresponding to the first wake-up word in a case where the first similarity score is greater than or equal to a first threshold value corresponding to the first wake-up word;

[0016] adjust the first threshold value corresponding to the first wake-up word in a case where a reverse operation of the first instruction is detected

[0017] In a third aspect, the present application provides a vehicle, the vehicle comprising an electronic device, the electronic device comprising a memory and a processor, the memory being configured to store a computer program, and the processor being configured to invoke and run the computer program stored in the memory, and the processor implements the wake-up word threshold adjustment method when executing the computer program.

[0018] In a fourth aspect, the present application provides a readable storage medium, the readable storage medium storing a computer program, and the computer program implements the wake-up word threshold adjustment method when executed by a processor.

[0019] In the embodiments of the present application, the real-time audio data detected in the vehicle is matched with a preset comparison model to obtain a first similarity score, and a first instruction corresponding to the first wake-up word is triggered in a case where the first similarity score is greater than or equal to a first threshold value corresponding to the first wake-up word. In a case where a reverse operation of the first instruction is detected during execution of the first instruction corresponding to the first wake-up word, the first threshold value corresponding to the first wake-up word can be adjusted, the dynamic adjustment of the pre-set wake-up word threshold value in the vehicle is realized, and the wake-up word threshold value is kept in line with the current use habit of the user during the continuous dynamic adjustment of the wake-up word threshold value, thereby improving the user experience of using the vehicle voice interaction system. Moreover, the technical solution provided in the embodiments of the present application can automatically adjust the wake-up word threshold value according to the specific situation without manual operation by the user during use in the pre-set condition, so as to make the wake-up word threshold value more in line with the use habit of the user, thereby further improving the user experience of using the vehicle voice interaction system. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to make the technical solutions in the embodiments of the present application or the prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those of ordinary skill in the art without any creative effort.

[0021] Figure 1 is a step flow chart of a threshold adjustment method for a wake-up word free provided by an embodiment of the present application.

[0022] Figure 2 is a step flow chart of another threshold adjustment method for a wake-up word free provided by an embodiment of the present application.

[0023] Figure 3 is a logic block diagram of a threshold adjustment device for a wake-up word free provided by an embodiment of the present application.

[0024] Figure 4 is a structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without any creative effort fall within the scope of the present application.

[0026] Figure 1 is a step flow chart of a threshold adjustment method for a wake-up word free provided by an embodiment of the present application, as shown in Figure 1 , the method can include:

[0027] Step S110, acquiring voice data.

[0028] It should be noted that the voice data is real-time audio data detected in the vehicle, which can specifically include audio data formed by the speaking of a person in the vehicle.

[0029] In the embodiments of the present application, the voice data can be acquired through a vehicle-mounted microphone, and the voice data is sent to a central control voice system of the vehicle machine.

[0030] Step S120, matching the voice data with a preset comparison model to obtain a first similarity score.

[0031] It should be noted that the preset contrast model can be a recognition engine trained based on a large amount of voice data. The contrast model corresponding to the wake-up-free word can be set according to the wake-up-free word set in the current vehicle machine.

[0032] In the embodiment of the application, after the voice data is acquired by the central control voice system of the vehicle machine through the microphone, the contrast model corresponding to the voice data is acquired through recognition of the voice data, and then the voice data is matched with the preset contrast model to obtain the matching degree score of the current acquired voice data and the contrast model, that is, the first similarity score.

[0033] In step S130, if the first similarity score is greater than or equal to the first threshold value corresponding to the first wake-up-free word, a first instruction corresponding to the first wake-up-free word is triggered.

[0034] In the embodiment of the application, the central control voice system of the vehicle machine is pre-provided with a plurality of wake-up-free words, and the plurality of wake-up-free words correspond to different threshold values respectively. The threshold value is the minimum score reference for triggering the wake-up-free word. When the voice data is matched with the preset contrast model to obtain the first similarity score greater than or equal to the first threshold value corresponding to the first wake-up-free word, a first instruction corresponding to the first wake-up-free word is triggered. When the voice data is matched with the preset contrast model to obtain the first similarity score less than the first threshold value corresponding to the first wake-up-free word, a first instruction corresponding to the first wake-up-free word is not triggered, that is, the threshold value condition for triggering the first instruction corresponding to the first wake-up-free word is not met.

[0035] It should be noted that a plurality of wake-up-free words are pre-provided before the vehicle is shipped, and initial threshold values are set for the plurality of wake-up-free words respectively. Specifically, the initial threshold value of the wake-up-free word is generally set to be between 0.6 and 0.7, so as to ensure that the plurality of wake-up-free words can be normally woken up to execute the instructions corresponding to the wake-up-free words when the user normally uses the vehicle. However, since the power amplifier configurations of the vehicles and the noise reduction effects are different, the setting of the initial threshold value also needs to be based on the specific debugging results. For example, the mute effect of the A vehicle is very good, and the signal-to-noise ratio of the microphone to the human voice can reach a good range. In this case, the initial threshold value score corresponding to the A vehicle can be relatively high, for example, it can be set to 0.7. For the B vehicle, the overall condition of the vehicle is not very good, and the mute effect is relatively poor. In this case, the initial threshold value score corresponding to the B vehicle is relatively low, for example, it is 0.60.

[0036] In the embodiment of the present application, the first wake-up word and the first instruction are in one-to-one correspondence, for example, the wake-up word "turn on music" corresponds to the instruction of turning on the music playing function in the car machine.

[0037] In step S140, the first threshold value corresponding to the first wake-up word is adjusted in the case of detecting reverse operation of the first instruction.

[0038] It should be noted that the reverse operation of the first instruction is the operation of the user actively turning off the first instruction within a certain time. The certain time can be 2s to 3s (unit: second), for example, when the user actively turns off the music function within 2s to 3s after triggering the instruction corresponding to "turn on music", it is determined as the reverse operation of the first instruction; when the user actively turns off the music function after 2min (unit: minute) after triggering the instruction corresponding to "turn on music", it is not determined as the reverse operation of the first instruction.

[0039] In the embodiment of the present application, the reverse operation can include the operation of the user manually turning off the first instruction and the operation of the user turning off the first instruction through voice.

[0040] In the embodiment of the present application, in the case of detecting the reverse operation of the first instruction, it is explained that the step of triggering the first instruction corresponding to the first wake-up word at this time is the false wake-up of the first wake-up word, which can be that the first wake-up word is mentioned in the conversation of the people in the car, and the user does not want to trigger the first instruction corresponding to the first wake-up word. Therefore, in the case of detecting the reverse operation of the first instruction, the false wake-up situation can be avoided from frequently occurring by adjusting the first threshold value corresponding to the first wake-up word in time.

[0041] Optionally, before the step of adjusting the first threshold value corresponding to the first wake-up word in the case of detecting the reverse operation of the first instruction, the following sub-steps can be further included:

[0042] In sub-step S1401, the first number of times of detecting the reverse operation of the first instruction is counted in the case of detecting the reverse operation of the first instruction.

[0043] In sub-step S1402, the step of adjusting the first threshold value corresponding to the first wake-up word is executed in the case that the first number of times is greater than a preset reverse operation number threshold value.

[0044] In the embodiment of the present application, in the case of detecting the reverse operation of the first instruction, before adjusting the first threshold value corresponding to the first wake-up word, the first number of times of detecting the reverse operation of the first instruction can be counted by recording, and in the case that the first number of times accumulates to be greater than a preset reverse operation number threshold, the step of adjusting the first threshold value corresponding to the first wake-up word is executed again, which can improve the pertinence and accuracy of threshold adjustment and ensure that each time the first threshold value corresponding to the first wake-up word is adjusted in the direction closer to the user's usage habit.

[0045] It should be noted that the preset reverse operation number threshold is an upper limit value of the number of times of detecting the reverse operation of the same instruction by the user; when the first number of times of detecting the reverse operation of the first instruction is greater than the preset reverse operation number threshold, the step of adjusting the first threshold value corresponding to the first wake-up word is executed. At the same time, after the step of adjusting the first threshold value corresponding to the first wake-up word is executed, the first number of times can be cleared, and the next round of counting the first number of times of detecting the reverse operation of the first instruction is started again, and in the case that the first number of times of detecting the reverse operation of the first instruction in the next round is greater than the preset reverse operation number threshold, the step of adjusting the first threshold value corresponding to the first wake-up word is executed again. The dynamic adjustment of the wake-up word threshold is performed in this way.

[0046] Optionally, the adjusting the first threshold value corresponding to the first wake-up word can include the following sub-steps:

[0047] The sub-step S1411 acquires the original score of the first threshold value corresponding to the first wake-up word.

[0048] It should be noted that the original score of the first threshold value is the score value corresponding to the first threshold value before the step of adjusting the first threshold value corresponding to the first wake-up word is executed this time, which can be the score value after the last time of adjusting the first threshold value corresponding to the first wake-up word, or can be the initial score value corresponding to the initial threshold value set when the vehicle is tested out of the factory.

[0049] The sub-step S1412 adjusts the first preset adjustment score on the basis of the original score to obtain the adjusted first threshold value corresponding to the first wake-up word.

[0050] In the embodiment of the present application, in the case of detecting the reverse operation of the first instruction, it is explained that the step of triggering the first instruction corresponding to the first wake-up free word at this time is the false wake-up of the first wake-up free word, and the user does not want to trigger the first instruction corresponding to the first wake-up free word. Therefore, in order to avoid the frequent occurrence of such false wake-up, the first preset adjustment score can be adjusted on the basis of the original score of the first threshold, the score of the first threshold corresponding to the first wake-up free word is adjusted, and the adjusted first threshold corresponding to the first wake-up free word is obtained. In the case of adjusting the score of the first threshold, the first similarity score required for triggering the first instruction also needs to be correspondingly improved in order to trigger the first instruction corresponding to the first wake-up free word, so that the frequency of false wake-up is reduced.

[0051] It should be noted that the first preset adjustment score is a single execution of adjusting the first preset adjustment score on the basis of the original score, and the specific score of adjusting the first threshold operation corresponding to the first wake-up free word on the basis of the original score is obtained.

[0052] Optionally, the reverse operation of the first instruction includes manual cancellation operation and voice cancellation operation.

[0053] In the case of detecting the reverse operation of the first instruction, adjusting the first threshold corresponding to the first wake-up free word can include adjusting the first threshold corresponding to the first wake-up free word in the case of detecting the manual cancellation operation and / or voice cancellation operation of the first instruction.

[0054] It should be noted that the reverse operation of the first instruction includes manual cancellation of the first instruction and voice cancellation of the first instruction. In the case of detecting the manual cancellation operation and / or voice cancellation operation of the first instruction by the user, the step of adjusting the first threshold corresponding to the first wake-up free word can also be performed.

[0055] For example, during the conversation of the people in the car, the conversation content involves A: "I don't know if the road in Beijing is blocked today", B: "You can open the navigation to check the road condition", which mentions "open the navigation". After obtaining the voice data of the current conversation, "open the navigation" in the voice data can trigger the instruction corresponding to the wake-up free word "open the navigation", but based on the specific conversation content, it can be known that triggering the instruction corresponding to the wake-up free word "open the navigation" is not the user's intention, which is the false wake-up of the wake-up free word "open the navigation". After the above-mentioned triggering of the instruction corresponding to "open the navigation", the user manually closes the navigation interface, at this time, the reverse operation of the user on the instruction "open the navigation" can be detected, and at this time, the threshold corresponding to the wake-up free word "open the navigation" needs to be adjusted to avoid similar situations from occurring again.

[0056] Optionally, after triggering the first instruction corresponding to the first wake-up word in the case that the first similarity score is greater than or equal to the first wake-up word threshold, the method further includes the following sub-steps:

[0057] In the sub-step S1311, the specific content of the voice data is determined.

[0058] It should be noted that the specific content of the voice data can be determined as whether the first wake-up word is included in the voice data, and in the case that the first wake-up word is included in the voice data, it is further determined whether other non-wake-up words exist in the voice data except the first wake-up word; of course, in the case that the wake-up word is not included in the voice data, the subsequent sub-steps S1312 and S1313 do not need to be executed.

[0059] In the sub-step S1312, in the case that the specific content of the voice data includes the first wake-up word and other non-wake-up words, the operation of adjusting the first threshold corresponding to the first wake-up word in the case that the reverse operation of the first instruction is detected is performed.

[0060] In the embodiment of the application, in the case that the specific content of the voice data includes the first wake-up word and other non-wake-up words, it is indicated that the first wake-up word can appear in the conversation of the people in the vehicle, and therefore in the case that the reverse operation of the first instruction is detected, it is confirmed that the first wake-up word is miswoken, and the first threshold corresponding to the first wake-up word needs to be adjusted.

[0061] In the sub-step S1313, in the case that the specific content of the voice data only includes the first wake-up word, the first instruction corresponding to the first wake-up word is normally executed.

[0062] In the embodiment of the application, in the case that the specific content of the voice data only includes the first wake-up word, it is indicated that the first wake-up word appears alone, and when the first wake-up word appears alone, it is considered that the first wake-up word in the voice data is the first instruction that the user wants to trigger corresponding to the first wake-up word. Therefore, in the case that the specific content of the voice data only includes the first wake-up word, the first instruction corresponding to the first wake-up word can be directly executed normally, and step S140 does not need to be executed.

[0063] Optionally, after the voice data is matched with the preset comparison model to obtain the first similarity score, the following sub-steps can be further included:

[0064] The sub-step S1211 comprises: in a case where the first similarity score is less than a first wake-up word threshold, not triggering a first instruction corresponding to the first wake-up word.

[0065] In the embodiment of the present application, in a case where the first similarity score is less than the first wake-up word threshold, it is indicated that the matching degree of the current voice data and the preset comparison model does not meet the triggering of the first instruction corresponding to the first wake-up word. Therefore, in the case where the first similarity score is less than the first wake-up word threshold, the first instruction corresponding to the first wake-up word is not triggered.

[0066] Exemplarily, if the first similarity score of the user's saying “turn on the air conditioner” is 0.65, and the score of the first threshold corresponding to the first wake-up word “turn on the air conditioner” is 0.72, at this time, since 0.65 is less than 0.72, the first similarity score of the user's saying “turn on the air conditioner” does not reach the benchmark of the first threshold 0.72, and the first instruction this time is discarded; if the first similarity score of the user's saying “turn on the air conditioner” this time is 0.85, at this time, since 0.85 is greater than 0.72, the first similarity score of the user's saying “turn on the air conditioner” reaches the benchmark of the first threshold 0.72, and the first instruction corresponding to the first wake-up word “turn on the air conditioner” is executed.

[0067] The sub-step S1212 comprises: determining the specific content of the voice data.

[0068] In the embodiment of the present application, after the first instruction corresponding to the first wake-up word is not triggered in the case where the first similarity score is less than the first wake-up word threshold, the specific content of the voice data can be determined.

[0069] The sub-step S1213 comprises: in a case where the specific content of the voice data only includes the first wake-up word, recording a second number of times of not triggering the first instruction corresponding to the first wake-up word.

[0070] In the embodiment of the present application, in the case that the first similarity score is less than the first wake-up word threshold, the specific content of the voice data can be determined after the first instruction corresponding to the first wake-up word is not triggered. In the case that the specific content of the voice data only includes the first wake-up word, it is indicated that the first wake-up word appears alone, and the first wake-up word in the voice data can be considered as the first instruction corresponding to the first wake-up word that the user wants to trigger. When the user says the first wake-up word to trigger the first instruction corresponding to the first wake-up word, the voice data including only the first wake-up word said by the user is not triggered due to the first similarity score being less than the first wake-up word threshold, which indicates that the score of the first threshold corresponding to the first wake-up word is too high. Therefore, the second number of times of not triggering the first instruction corresponding to the first wake-up word needs to be recorded.

[0071] In the case that the second number of times is greater than the preset number of times of not triggering, the first threshold corresponding to the first wake-up word is adjusted.

[0072] In the embodiment of the present application, when the user says the first wake-up word to trigger the first instruction corresponding to the first wake-up word, the voice data including only the first wake-up word said by the user is not triggered due to the first similarity score being less than the first wake-up word threshold, which indicates that the score of the first threshold corresponding to the first wake-up word is too high. Therefore, the second number of times of not triggering the first instruction corresponding to the first wake-up word needs to be recorded, and in the case that the second number of times is greater than the preset number of times of not triggering, the first threshold corresponding to the first wake-up word is adjusted.

[0073] It should be noted that the preset number of times of not triggering is an upper limit value of the number of times that the first instruction corresponding to the first wake-up word input by the user through voice data is not triggered. In the case that the second number of times of not triggering the first instruction corresponding to the first wake-up word is greater than the preset number of times of not triggering, the step of adjusting the first threshold corresponding to the first wake-up word is executed. After the step of adjusting the first threshold corresponding to the first wake-up word is executed, the second number of times can be cleared, and the second number of times of not triggering the first instruction corresponding to the first wake-up word is recorded again. In the case that the second number of times of not triggering the first instruction corresponding to the first wake-up word recorded in the next round is greater than the preset number of times of not triggering, the step of adjusting the first threshold corresponding to the first wake-up word is executed again. The dynamic adjustment of the wake-up word threshold is performed in this way.

[0074] Optionally, the adjusting the first threshold value corresponding to the first wake-up word in the case that the second number of times is greater than the preset non-triggering number of times threshold value can include the following sub-steps:

[0075] The sub-step S12141 is acquiring an original score of the first threshold value corresponding to the first wake-up word.

[0076] It should be noted that the original score of the first threshold value is a score value corresponding to the first threshold value before this time of performing the step of adjusting the first threshold value corresponding to the first wake-up word, and the score value can be a score value after the last time of performing the step of adjusting the first threshold value corresponding to the first wake-up word, or can be an initial score value corresponding to an initial threshold value set at the time of factory testing of the vehicle.

[0077] The sub-step S12142 is down-regulating a second preset adjustment score on the basis of the original score to obtain an adjusted first threshold value corresponding to the first wake-up word.

[0078] In the embodiment of the application, in the case that the second number of times is greater than the preset non-triggering number of times threshold value, it is indicated that the user wants to trigger the first instruction corresponding to the first wake-up word by inputting the first wake-up word through voice data, but the first instruction is not triggered normally because the first similarity score of the voice data input by the user is less than the first wake-up word threshold value. Therefore, in order to avoid the frequent occurrence of the situation that the instruction that the user wants to trigger is not triggered multiple times, the second preset adjustment score can be down-regulated on the basis of the original score of the first threshold value, the score of the first threshold value corresponding to the first wake-up word is lowered, and an adjusted first threshold value corresponding to the first wake-up word is obtained. In the case that the score of the first threshold value is lowered, the first similarity score required for triggering the first instruction is also lowered accordingly, so that it becomes easy to trigger the first instruction corresponding to the first wake-up word, thereby avoiding the frequent occurrence of the situation that the instruction that the user wants to trigger is not triggered multiple times.

[0079] It should be noted that the second preset adjustment score is a specific score that needs to be down-regulated on the basis of the original score for a single time of performing the operation of down-regulating the second preset adjustment score on the basis of the original score to obtain an adjusted first threshold value corresponding to the first wake-up word, and the second preset adjustment score can be the same as the first preset adjustment score, or can be different. The embodiment of the application does not limit this.

[0080] Figure 2 is a step flowchart of another wake-up word threshold value adjustment method provided by the embodiment of the application, as shown in the figure, the method can include: Figure 2 as shown in the figure, the method can include:

[0081] Step S201, the vehicle is started, and the car machine system is started.

[0082] In the embodiment of the present application, the vehicle start can be a vehicle power-on start state, and the vehicle machine system is started.

[0083] In step S202, the central control voice system is started.

[0084] In the embodiment of the present application, after the user gets on the vehicle and controls the vehicle to start, the central control voice system in the vehicle machine is automatically started, at this time, the microphone of the central control voice system starts continuous recording work, and the voice interaction system works normally.

[0085] In the embodiment of the present application, step S202 can include the following sub-steps:

[0086] In sub-step S2021, voice data is recorded through a microphone.

[0087] In the embodiment of the present application, after the central control voice system is started, the microphone starts continuous recording work, records the real-time audio data detected in the vehicle as voice data, and uploads the voice data to the central control voice system.

[0088] In sub-step S2022, the voice data is matched with a preset comparison model.

[0089] In the embodiment of the present application, after the central control voice system obtains voice data through a microphone, the central control voice system matches the voice data with a preset comparison model, obtains a first similarity score, and in the case that the first similarity score is greater than or equal to a first threshold value corresponding to a first wake-up word, step S203 is executed.

[0090] In sub-step S2023, the voice data is stored in a voice data buffer pool.

[0091] In the embodiment of the present application, the central control voice system of the vehicle can also include a voice data buffer pool for storing the obtained voice data. After the central control voice system matches the voice data with a preset comparison model and obtains a first similarity score, the voice data is stored in the data buffer pool for subsequent use.

[0092] In step S203, a command corresponding to a wake-up word is triggered.

[0093] In the embodiment of the present application, in the case that the first similarity score is greater than or equal to a first threshold value corresponding to a first wake-up word, a first command corresponding to the first wake-up word is triggered.

[0094] It should be noted that the first wake-up word can be any one of a plurality of wake-up words pre-set in the machine, and the first command is an operation command corresponding to the first wake-up word.

[0095] Step S204, taking out the voice data to determine whether there is other data before and after the wake-up word.

[0096] In the embodiment of the present application, the original voice data including the first wake-up word can be taken out from the data buffer pool again, and whether there is other data before and after the first wake-up word in the voice data is determined. In the case where it is determined that there is other data before and after the first wake-up word, step S205 is executed; in the case where it is determined that there is no other data before and after the first wake-up word, step S206 is executed.

[0097] Step S205, whether the user has a reverse operation for the current instruction.

[0098] In the embodiment of the present application, in the case where it is determined that there is other data before and after the first wake-up word, it is detected whether the user has a reverse operation for the current instruction. Specifically, in the case where it is determined that there is other data before and after the first wake-up word, it indicates that the first wake-up word may appear in the sentence of the conversation of the person in the vehicle, and therefore it is necessary to determine whether the user has a reverse operation for the current instruction.

[0099] In the embodiment of the present application, in the case where it is detected that the user has a reverse operation for the current instruction, step S207 is executed; in the case where it is not detected that the user has a reverse operation for the current instruction, step S206 is executed.

[0100] Step S206, normally executing the instruction corresponding to the wake-up word.

[0101] In the embodiment of the present application, in the case where it is determined that there is no other data before and after the first wake-up word and in the case where it is not detected that the user has a reverse operation for the current instruction, the first instruction corresponding to the first wake-up word can be normally executed.

[0102] Step S207, determining whether the number of reverse operations for the current instruction is greater than three.

[0103] In the embodiment of the present application, in the case where it is detected that the user has a reverse operation for the current instruction, it is necessary to determine whether the number of reverse operations for the current instruction is greater than three, and in the case where it is greater than three, step S209 is executed; in the case where it is less than or equal to three, step S208 is executed.

[0104] It should be noted that the three times herein corresponds to the above-mentioned preset reverse operation number threshold, which can be set to three, and of course can be set to other numbers according to actual needs, and the embodiment of the present application does not limit this.

[0105] Step S208, recording the current operation, number + 1.

[0106] In the embodiment of the present application, in the case that the number of reverse operations for the current instruction is determined to be less than or equal to three, the current operation needs to be recorded, and the original number is increased by 1 and saved.

[0107] In step S209, the wake-up word threshold is increased by 0.02.

[0108] In the embodiment of the present application, in the case that the number of reverse operations for the current instruction is determined to be greater than three, the operation of adjusting the first threshold corresponding to the first wake-up word needs to be performed, and specifically, the score of the first threshold corresponding to the first wake-up word can be increased by 0.02.

[0109] It should be noted that 0.02 here corresponds to the first preset adjustment score, which can be set to 0.02, and of course can be set according to actual needs, and the embodiment of the present application does not limit this.

[0110] In the embodiment of the present application, by matching the real-time audio data detected in the vehicle with the preset comparison model, a first similarity score is obtained, and in the case that the first similarity score is greater than or equal to the first threshold corresponding to the first wake-up word, the first instruction corresponding to the first wake-up word is triggered, and in the case that the reverse operation for the first instruction is detected in the process of executing the first instruction corresponding to the first wake-up word, the first threshold corresponding to the first wake-up word can be adjusted, realizing the dynamic adjustment of the pre-set wake-up word threshold in the vehicle, and in the process of continuously dynamically adjusting the wake-up word threshold, the wake-up word threshold is followed to adapt to the current use habit of the user, improving the experience of the user using the vehicle voice interaction system. And the technical solution provided in the embodiment of the present application can realize the self-adjustment of the wake-up word threshold according to the specific situation without manual operation of the user in the process of using the pre-set situation, so as to make the wake-up word threshold more consistent with the user's use habit, and further improve the user's experience of using the vehicle voice interaction system.

[0111] Figure 3 is a logic block diagram of a wake-up word threshold adjustment device provided by the embodiment of the present application. The device 300 can include:

[0112] The acquisition module 310 is configured to acquire voice data, and the voice data is real-time audio data detected in the vehicle.

[0113] The matching module 320 is configured to match the voice data with a preset comparison model to obtain a first similarity score.

[0114] The instruction triggering module 330 is configured to trigger a first instruction corresponding to the first wake-up word if the first similarity score is greater than or equal to a first threshold corresponding to the first wake-up word.

[0115] The first adjusting module 340 is configured to adjust the first threshold corresponding to the first wake-up word if a reverse operation for the first instruction is detected.

[0116] Optionally, the wake-up word threshold adjusting apparatus 300 can further include:

[0117] The statistical module is configured to count a first number of times of detecting the reverse operation for the first instruction if the reverse operation for the first instruction is detected.

[0118] The first execution module is configured to execute the step of adjusting the first threshold corresponding to the first wake-up word if the first number of times is greater than a preset reverse operation number threshold.

[0119] Optionally, the adjusting module 340 can include:

[0120] The first acquisition sub-module is configured to acquire an original score of the first threshold corresponding to the first wake-up word.

[0121] The first adjusting sub-module is configured to increase the first preset adjusting score on the basis of the original score to obtain an adjusted first threshold corresponding to the first wake-up word.

[0122] Optionally, the adjusting module 340 can include:

[0123] The second adjusting sub-module is configured to adjust the first threshold corresponding to the first wake-up word if a manual cancel operation and / or a voice cancel operation for the first instruction is detected.

[0124] Optionally, the wake-up word threshold adjusting apparatus 300 can further include:

[0125] The determining module is configured to determine a specific content of the voice data.

[0126] The second adjusting module is configured to execute the operation of adjusting the first threshold corresponding to the first wake-up word if the reverse operation for the first instruction is detected if the specific content of the voice data includes the first wake-up word and other non-wake-up words.

[0127] The second execution module is configured to normally execute the first instruction corresponding to the first wake-up word if the specific content of the voice data only includes the first wake-up word.

[0128] Optionally, the wake-word threshold adjustment apparatus 300 can further include:

[0129] an instruction inhibition module configured to not trigger a first instruction corresponding to the first wake-word in a case that the first similarity score is less than a first wake-word threshold value.

[0130] a determination module configured to determine a specific content of the voice data.

[0131] a recording module configured to record a second number of times of not triggering the first instruction corresponding to the first wake-word in a case that the specific content of the voice data only includes the first wake-word.

[0132] a third adjustment module configured to adjust the first threshold value corresponding to the first wake-word in a case that the second number of times is greater than a preset number of times of not triggering threshold value.

[0133] Optionally, the adjustment module 300 can include:

[0134] a second acquisition sub-module configured to acquire an original score of the first threshold value corresponding to the first wake-word.

[0135] a third adjustment sub-module configured to down-regulate a second preset adjustment score on the basis of the original score to obtain an adjusted first threshold value corresponding to the first wake-word.

[0136] The present application also provides a vehicle, the vehicle comprising an electronic device, referring to Figure 4 , the electronic device comprising a processor 401, a memory 402, and a computer program 4021 stored in the memory and executable on the processor, the processor implementing the wake-word threshold adjustment method of the foregoing embodiments when executing the program.

[0137] The present application also provides a readable storage medium, the readable storage medium storing a computer program, the computer program being executable on a processor to implement the wake-word threshold adjustment method of the foregoing embodiments.

[0138] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts refer to the part of the method embodiment.

[0139] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other apparatus. Structural requirements of such systems to construct them to practice the operation described above are apparent in light of the above description. In addition, the present application is not directed to any particular programming language. It will be appreciated that a variety of programming languages can be used to implement the present application described herein, and the descriptions above for particular languages are provided for disclosure of the best mode of the present application.

[0140] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the application can be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been described in detail in order not to obscure the understanding of this description.

[0141] Similarly, it is to be understood that the mechanical details of the application sometimes are grouped into single embodiments, figures or descriptions of related embodiments in the above description of example embodiments of the application for the sake of brevity and clarity. However, the disclosure is not to be interpreted as reflecting an intention that the application requires more features than are explicitly recited in each claim. Rather, claim(s) reflect the minimum scope of the present application to the exclusion of any additional features. Accordingly, the claims are hereby expressly incorporated into this detailed description of the illustrated embodiments of the application, with each claim read independently of any other claim(s) hereto.

[0142] Those skilled in the art will appreciate that the modules in the apparatus of the embodiments can be adapted and placed in one or more apparatuses other than the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and further can be divided into more sub-modules or sub-units or sub-components. Any combination of all the features disclosed in the specification (including the accompanying claims, abstract and drawings), and any method or apparatus so disclosed, can be used in any combination, except that at least some of such features and / or processes or units are mutually exclusive. Unless explicitly stated otherwise, each feature disclosed in the specification (including the accompanying claims, abstract and drawings) can be replaced by alternative features that serve the same, equivalent or similar purpose.

[0143] Embodiments of the various components of the application can be implemented in hardware, or as software modules running in one or more processors, or combinations thereof. Those skilled in the art will appreciate that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all of the functionality of some or all of the components in the sequencing apparatus according to the present application. The present application can also be implemented as a program for executing part or all of the methods described herein on a device or apparatus. Such a program can be stored on a computer readable medium or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier medium, or in any other form.

[0144] It should be noted that the above-mentioned embodiments illustrate rather than limit the application, and that one skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word 'comprising' does not exclude the presence of elements or steps other than those listed in a claim. The word 'a' or 'an' preceding an element does not exclude the presence of a plurality of such elements. The application can be implemented by means of both hardware and software, and any combination thereof. In a unit claim, several devices can be listed with a comma. The use of the term 'about' followed by a value and / or a term 'approximately' preceding a value means that the value can vary from the stated value by 10%. The use of any of the following terms in the claims is neither meant to limit the scope nor to introduce a non-combination limitation. The terms 'comprise', 'include', and 'contain' are not used in their exclusive sense. The use of the term 'first','second', and 'third' does not connote any order, quantity, creation or importance, but rather are used to denote one element from another. The use of these terms is interchangeable under appropriate circumstances.

[0145] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.

[0146] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0147] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0148] It should be noted that the above-mentioned embodiments illustrate rather than limit the application, and that one skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word 'comprising' does not exclude the presence of elements or steps other than those listed in a claim. The word 'a' or 'an' preceding an element does not exclude the presence of a plurality of such elements. The application can be implemented by means of both hardware and software, and any combination thereof. In a unit claim, several devices can be listed with a comma. The use of the term 'about' followed by a value and / or a term 'approximately' preceding a value means that the value can vary from the stated value by 10%. The use of any of the following terms in the claims is neither meant to limit the scope nor to introduce a non-combination limitation. The terms 'comprise', 'include', and 'contain' are not used in their exclusive sense. The use of the term 'first','second', and 'third' does not connote any order, quantity, creation or importance, but rather are used to denote one element from another. The use of these terms is interchangeable under appropriate circumstances.

Claims

1. A method for threshold adjustment without a wake-up word, characterized in that, The method comprises the following steps: acquiring voice data, the voice data being real-time audio data detected in a vehicle; matching the voice data with a preset comparison model to obtain a first similarity score; in a case where the first similarity score is greater than or equal to a first threshold value corresponding to a first wake-up word, triggering a first instruction corresponding to the first wake-up word; in a case where a reverse operation of the first instruction is detected, adjusting the first threshold value corresponding to the first wake-up word.

2. The method of claim 1, wherein, Before the step of adjusting the first threshold value corresponding to the first wake-up word in a case where a reverse operation of the first instruction is detected, the method further comprises the following steps: in the case where a reverse operation of the first instruction is detected, counting a first number of times that a reverse operation of the first instruction is detected; in a case where the first number of times is greater than a preset reverse operation number threshold value, performing the step of adjusting the first threshold value corresponding to the first wake-up word.

3. The method of claim 1, wherein, The step of adjusting the first threshold value corresponding to the first wake-up word comprises the following steps: acquiring an original score of the first threshold value corresponding to the first wake-up word; on the basis of the original score, increasing a first preset adjustment score to obtain an adjusted first threshold value corresponding to the first wake-up word.

4. The method of claim 1, wherein, The reverse operation of the first instruction comprises a manual cancellation operation and a voice cancellation operation. The step of adjusting the first threshold value corresponding to the first wake-up word in a case where a reverse operation of the first instruction is detected comprises the following steps: in a case where a manual cancellation operation and / or a voice cancellation operation of the first instruction is detected, adjusting the first threshold value corresponding to the first wake-up word.

5. The method of claim 1, wherein, After the step of triggering a first instruction corresponding to the first wake-up word in a case where the first similarity score is greater than or equal to a first wake-up word threshold value, the method further comprises the following steps: determining specific content of the voice data; in a case where the specific content of the voice data comprises the first wake-up word and other non-wake-up words, performing the operation of adjusting the first threshold value corresponding to the first wake-up word in a case where a reverse operation of the first instruction is detected; in a case where the specific content of the voice data only comprises the first wake-up word, normally executing the first instruction corresponding to the first wake-up word.

6. The method of claim 5, wherein, After the step of matching the voice data with a preset comparison model to obtain a first similarity score, the method further comprises the following steps: in a case where the first similarity score is less than a first wake-up word threshold value, not triggering a first instruction corresponding to the first wake-up word; determining specific content of the voice data; in a case where the specific content of the voice data only comprises the first wake-up word, recording a second number of times that the first instruction corresponding to the first wake-up word is not triggered; in a case where the second number of times is greater than a preset non-triggering number threshold value, adjusting the first threshold value corresponding to the first wake-up word.

7. The method of claim 6, wherein, The step of adjusting the first threshold value corresponding to the first wake-up word in a case where the second number of times is greater than a preset non-triggering number threshold value comprises the following steps: acquiring an original score of the first threshold value corresponding to the first wake-up word; Down-regulate a second preset adjustment score on the basis of the original score, to obtain an adjusted first threshold corresponding to the first wake-up keyword.

8. An apparatus for adjusting a threshold value for an awaking-free word, characterized by, Comprise: An acquisition module is used to acquire voice data, the voice data is real-time audio data detected in the vehicle; A matching module is used to match the voice data with a preset comparison model to obtain a first similarity score; An instruction triggering module is used to trigger a first instruction corresponding to the first wake-up keyword in the case that the first similarity score is greater than or equal to a first threshold corresponding to the first wake-up keyword; A first adjustment module is used to adjust the first threshold corresponding to the first wake-up keyword in the case that a reverse operation for the first instruction is detected.

9. A vehicle characterized by comprising: The vehicle comprises an electronic device, the electronic device comprises a memory and a processor, the memory is used to store a computer program, the processor is used to call and run the computer program stored in the memory, and the processor implements the wake-up keyword threshold adjustment method in any one of claims 1-7 when executing the computer program.

10. A readable storage medium, the readable storage medium storing a computer program, characterized in that, The computer program is executed by the processor to implement the wake-up keyword threshold adjustment method in any one of claims 1-7.

Citation Information

Patent Citations

  • Voice wake-up method, device, terminal and storage medium

    CN107134279A

  • Wake word evaluation

    US9275637B1