Voice wake-up data processing method and device and computer readable storage medium

By determining the interval time and confidence of the audio data in the voice interaction process, and obtaining the target audio data as a sample for voice wake-up training, the problem of missing sample data is solved and the accuracy of user wake-up is improved.

CN119920240APending Publication Date: 2025-05-02MIDEA GROUP WUHAN REFRIGERATION EQUIPMENT CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311442155.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-10-30
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

In the prior art, sample data for voice wake-up training is missing, resulting in low user wake-up accuracy.

Method used

By determining the interval time between the voice interaction flow, the confidence of the reference audio data in the interval time is determined. If the difference between the confidence and the preset reliability threshold is within the preset range, the target audio data of the voice interaction flow of a preset number of times adjacent to the reference audio data is obtained and used as sample data for voice wake-up training.

Benefits of technology

Effectively collect truly effective wake-up audio data, improve the accuracy of user wake-up and meet user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920240A_ABST
    Figure CN119920240A_ABST
Patent Text Reader

Abstract

The invention discloses a voice wake-up data processing method and device and a computer readable storage medium, and the method comprises the steps: determining the interval time between voice interaction processes, and the voice interaction processes are the interaction processes of successfully waking up the device through voice; determining the confidence of the reference audio data in the interval time; if a difference value between the confidence coefficient and a preset confidence coefficient threshold value is within a preset range, obtaining target audio data of the voice interaction process of a preset number of times adjacent to the reference audio data; and taking the target audio data as sample data, wherein the sample data is used for voice wake-up training. According to the invention, real and effective wake-up audio data can be efficiently collected, the audio data is used for later training, the wake-up accuracy of the user is improved, and the user demand is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio processing technology, and in particular to a voice wake-up data processing method, device and computer-readable storage medium. Background Art

[0002] After the terminal device with voice recognition wake-up function is turned on, it is in a dormant state in voice interaction. When the user says the specified wake-up word, the device will be awakened, switch from dormant state to working state, and wait for the user's next instruction. When the device is in working state and idle for a predetermined length of time, the device can go from working state to dormant state again until the user wakes up the device again by voice. Therefore, with the voice wake-up mechanism, the device does not need to be in working state all the time, thus saving energy consumption.

[0003] However, due to factors such as the user's accent, there may be audio data that cannot wake up the device or wakes up the device by mistake. A common solution is to collect sample data by setting a wake-up threshold, for example, collecting sample data when the wake-up threshold is exceeded. If the wake-up threshold is set too high, more valuable audio data will be lost. If the wake-up threshold is set too low, a large amount of useless sample data will exist, which will eventually lead to the lack of sample data for voice wake-up training. Summary of the invention

[0004] The main purpose of the present invention is to provide a voice wake-up data processing method, device and computer-readable storage medium, aiming to improve the problem of missing sample data for voice wake-up training.

[0005] To achieve the above object, the present invention provides a method for processing voice wake-up data, which comprises the following steps:

[0006] Determine the interval time between voice interaction processes, where the voice interaction process is an interaction process in which the voice successfully wakes up the device;

[0007] Determining the confidence level of the reference audio data in the interval time;

[0008] If the difference between the confidence and a preset confidence threshold is within a preset range, obtaining target audio data of the voice interaction process that is adjacent to the reference audio data for a preset number of times;

[0009] The target audio data is used as sample data, and the sample data is used for voice wake-up training.

[0010] Optionally, after the step of determining the interval time between voice interaction processes, the method further includes:

[0011] When the duration of the interval time is greater than a preset steady-state interval value and less than a preset interval maximum value, the step of determining the confidence of the reference audio data in the interval time is performed.

[0012] Optionally, after the step of using the target audio data as sample data, the method further includes:

[0013] Model training is performed based on the target audio data and the corresponding interaction success time to obtain a prediction model, and the prediction model is used to determine the prediction time period of the reference audio data.

[0014] Optionally, the step of performing model training based on the target audio data and the corresponding interaction success time to obtain a prediction model includes:

[0015] Determining the region information corresponding to the target audio data;

[0016] Classifying the sample data according to the regional information;

[0017] The classified sample data and the corresponding interaction success time are used for model training to obtain a prediction model corresponding to each of the regional information.

[0018] Optionally, after the step of performing model training based on the target audio data and the corresponding interaction success time to obtain a prediction model, the step further includes:

[0019] Get real-time audio data;

[0020] Predicting the real-time audio data based on the prediction model to obtain a prediction time period corresponding to the reference audio data;

[0021] According to the audio data within the prediction time period and with a confidence level greater than a preset confidence threshold, as the reference audio data, the target audio data of the voice interaction process adjacent to the reference audio data for a preset number of times is obtained.

[0022] Optionally, the step of determining the confidence of the audio data in the interval time comprises:

[0023] Extracting audio features of the reference audio data;

[0024] Matching the reference audio features with preset keywords to obtain a matching result of the reference audio data;

[0025] The confidence level of the reference audio data is determined according to the matching result.

[0026] Optionally, the step of matching the audio feature with a preset keyword to obtain a matching result of the reference audio data includes:

[0027] Obtaining a wake-up word and a command word in the reference audio data;

[0028] respectively determining a first matching result between the wake-up word and a preset wake-up keyword, and a second matching result between the instruction word and a preset instruction keyword;

[0029] A matching result of the reference audio data is determined according to the first matching result and the second matching result.

[0030] Optionally, the method further comprises:

[0031] A voice wake-up model is trained based on the sample data, and the voice wake-up model is used to determine whether to wake up the device.

[0032] To achieve the above-mentioned purpose, the present invention also provides a voice wake-up data processing device, which includes a memory, a processor, and a voice wake-up data processing program stored in the memory and executable on the processor. When the voice wake-up data processing program is executed by the processor, the various steps of the voice wake-up data processing method described above are implemented.

[0033] To achieve the above objectives, the present invention also provides a computer-readable storage medium, which stores a voice wake-up data processing program. When the voice wake-up data processing program is executed by a processor, the various steps of the voice wake-up data processing method described above are implemented.

[0034] The present invention provides a method, device and computer-readable storage medium for processing voice wake-up data, which determine the interval time between voice interaction processes; determine the confidence of reference audio data in the interval time; if the difference between the confidence and the preset confidence threshold is within a preset range, obtain the target audio data of the voice interaction process adjacent to the reference audio data for a preset number of times; use the target audio data as sample data, and use the sample data for voice wake-up training. This is conducive to efficiently collecting truly effective wake-up audio data, using the audio data for later training, improving the accuracy of user wake-up, and meeting user needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A schematic diagram of the hardware structure of a voice wake-up data processing device according to an embodiment of the present invention;

[0036] Figure 2 This is a flow chart of a first embodiment of a method for processing voice wake-up data according to the present invention;

[0037] Figure 3 It is a flowchart of a second embodiment of the voice wake-up data processing method of the present invention;

[0038] Figure 4 It is a flowchart of a third embodiment of the voice wake-up data processing method of the present invention;

[0039] Figure 5 4 is a flow chart of a fourth embodiment of a method for processing voice wake-up data according to the present invention.

[0040] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0041] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0042] The main solution of the embodiment of the present invention is: determine the interval time between voice interaction processes; determine the confidence of the reference audio data in the interval time; if the difference between the confidence and the preset confidence threshold is within the preset range, obtain the target audio data of the voice interaction process with a preset number of adjacent reference audio data; use the target audio data as sample data, and use the sample data for voice wake-up training. This is conducive to efficiently collecting truly effective wake-up audio data, using the audio data for later training, improving the accuracy of user wake-up, and meeting user needs.

[0043] As an implementation scheme, the voice wake-up data processing device can be as follows Figure 1 shown.

[0044] The embodiment of the present invention relates to a voice wake-up data processing device, which includes: a processor 101, such as a CPU, a memory 102, and a communication bus 103. The communication bus 103 is used to realize connection and communication between these components.

[0045] The memory 102 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Figure 1 As shown, the memory 102 as a computer-readable storage medium may include a voice wake-up data processing program; and the processor 101 may be used to call the voice wake-up data processing program stored in the memory 102 and perform the following operations:

[0046] Determine the interval time between voice interaction processes, where the voice interaction process is an interaction process in which the voice successfully wakes up the device;

[0047] Determining the confidence level of the reference audio data in the interval time;

[0048] If the difference between the confidence and a preset confidence threshold is within a preset range, obtaining target audio data of the voice interaction process that is adjacent to the reference audio data for a preset number of times;

[0049] The target audio data is used as sample data, and the sample data is used for voice wake-up training.

[0050] Optionally, the processor 101 may be used to call a voice wake-up data processing program stored in the memory 102, and perform the following operations:

[0051] When the duration of the interval time is greater than a preset steady-state interval value and less than a preset interval maximum value, the step of determining the confidence of the reference audio data in the interval time is performed.

[0052] Optionally, the processor 101 may be used to call a voice wake-up data processing program stored in the memory 102, and perform the following operations:

[0053] Model training is performed based on the target audio data and the corresponding interaction success time to obtain a prediction model, and the prediction model is used to determine the prediction time period of the reference audio data.

[0054] Optionally, the processor 101 may be used to call a voice wake-up data processing program stored in the memory 102, and perform the following operations:

[0055] Determining the region information corresponding to the target audio data;

[0056] Classifying the sample data according to the regional information;

[0057] The classified sample data and the corresponding interaction success time are used for model training to obtain a prediction model corresponding to each of the regional information.

[0058] Optionally, the processor 101 may be used to call a voice wake-up data processing program stored in the memory 102, and perform the following operations:

[0059] Get real-time audio data;

[0060] Predicting the real-time audio data based on the prediction model to obtain a prediction time period corresponding to the reference audio data;

[0061] According to the audio data within the prediction time period and with a confidence level greater than a preset confidence threshold, as the reference audio data, the target audio data of the voice interaction process adjacent to the reference audio data for a preset number of times is obtained.

[0062] Optionally, the processor 101 may be used to call a voice wake-up data processing program stored in the memory 102, and perform the following operations:

[0063] Extracting audio features of the reference audio data;

[0064] Matching the reference audio features with preset keywords to obtain a matching result of the reference audio data;

[0065] The confidence level of the reference audio data is determined according to the matching result.

[0066] Optionally, the processor 101 may be used to call a voice wake-up data processing program stored in the memory 102, and perform the following operations:

[0067] Obtaining a wake-up word and a command word in the reference audio data;

[0068] respectively determining a first matching result between the wake-up word and a preset wake-up keyword, and a second matching result between the instruction word and a preset instruction keyword;

[0069] A matching result of the reference audio data is determined according to the first matching result and the second matching result.

[0070] Optionally, the processor 101 may be used to call a voice wake-up data processing program stored in the memory 102, and perform the following operations:

[0071] A voice wake-up model is trained based on the sample data, and the voice wake-up model is used to determine whether to wake up the device.

[0072] Based on the hardware architecture of the above-mentioned voice wake-up data processing device, an embodiment of the voice wake-up data processing method of the present invention is proposed.

[0073] Reference Figure 2 , Figure 2 This is a first embodiment of the voice wake-up data processing method of the present invention, and the voice wake-up data processing method includes the following steps:

[0074] Step S10, determining the interval time between voice interaction processes, wherein the voice interaction process is an interaction process in which the voice successfully wakes up the device.

[0075] The method of this embodiment is applied to cloud servers or edge computing devices, etc.

[0076] Optionally, the voice interaction process is an interaction process in which the user's voice successfully wakes up the device. Optionally, the voice interaction process includes successfully waking up the device so that the device responds and successfully sending a control instruction to the device so that the device performs a preset action based on the instruction.

[0077] Different audio data have corresponding confidence levels, and the confidence level of each audio data during local wake-up reception. For example, the confidence level of the audio data collected in the T time period is M(T), the confidence level of the audio data collected in the T-1 time period is M(T-1), the confidence level of the audio data collected in the T-2 time period is M(T-2), and the confidence level of the audio data collected in the Tn time period is M(Tn).

[0078] Optionally, the confidence level of the audio in the voice interaction process is greater than a preset confidence threshold, so the device can be successfully woken up.

[0079] Optionally, the interval time is a time period. Optionally, the interval time Time between two adjacent voice interaction processes = the start time of the latter voice interaction process - the end time of the previous voice interaction process. Exemplarily, voice interaction process 1 and voice interaction process 2 are two adjacent voice interaction processes, wherein voice interaction process 1 occurs before voice interaction process 2, and the interval time Time between voice interaction process 1 and voice interaction process 2 = the start time of voice interaction process 2 - the end time of voice interaction process 1.

[0080] Optionally, when there are multiple consecutive interactions, the interval time Time between groups in the continuous process can be considered to be relatively stable, and when the interval time Time suddenly becomes longer, it can be considered that valuable audio data may be generated.

[0081] Optionally, after step S10, the method further includes: when the duration of the interval time is greater than a preset steady-state interval value and less than a preset interval maximum value, executing step S20.

[0082] The preset steady-state interval value refers to the interval between two consecutive voice interaction processes when the voice interaction process reaches a stable state, reflecting the stability and operation state of the voice interaction process, that is, the preset steady-state interval value is the minimum duration value between two voice interaction processes. The preset maximum interval value is the maximum duration value between two voice interaction processes.

[0083] Optionally, when the duration of the interval time is less than or equal to a preset steady-state interval value, or greater than or equal to a preset interval maximum value, the confidence of the audio data in the interval time is not calculated.

[0084] Step S20, determining the confidence level of the reference audio data in the interval time.

[0085] Optionally, reference audio data collected during the interval between voice interaction processes is obtained, and the confidence level of the reference audio data is determined to be M(Time).

[0086] Optionally, step S20 includes: extracting audio features of the reference audio data; matching the audio features with preset keywords to obtain matching results; and determining the confidence of the reference audio data according to the matching results.

[0087] Optionally, the reference audio feature is matched with a preset keyword according to a preset keyword matching model to obtain a matching result. The preset keyword matching model is trained based on the reference audio feature and the preset keyword. Optionally, the matching result includes a matching degree between the reference audio feature and the preset keyword, which can be expressed as a score.

[0088] Optionally, the wake-up word and the command word in the reference audio data are obtained; for example, the wake-up word is Xiaomei, and the command word is to turn on the air conditioner. A first matching result between the wake-up word and a preset wake-up keyword, and a second matching result between the command word and a preset command keyword are determined respectively; and a matching result of the reference audio data is determined based on the first matching result and the second matching result.

[0089] Optionally, a voice pre-wake-up model is trained based on reference audio data. The voice pre-wake-up model is used to determine whether to pre-wake the device so that when the voice device is in normal working state, the sound in the environment is continuously monitored, and when a specific sound pattern or sound within a specific frequency range is detected, the device is triggered to enter the voice recognition mode. The purpose of pre-wake-up is to prepare the device for voice recognition in advance so as to respond to user commands more quickly. Pre-wake-up is usually performed before wake-up to improve the response speed of the system.

[0090] Optionally, a voice pre-wake-up model is trained based on reference audio data whose difference between the confidence level and a preset confidence threshold is within a preset range, and the voice pre-wake-up model is used to determine whether to pre-wake up the device.

[0091] Step S30: If the difference between the confidence and a preset confidence threshold is within a preset range, then obtaining target audio data of the voice interaction process that is adjacent to the reference audio data for a preset number of times.

[0092] Optionally, the preset confidence threshold is a confidence threshold for effective voice wake-up, wherein when the confidence of the audio data is greater than the preset confidence threshold, the audio data can successfully wake up the device and implement the voice interaction process.

[0093] Optionally, when the confidence is less than a preset confidence threshold, and the difference between the confidence and the preset confidence threshold is within a preset range, the target audio data of the voice interaction process of a preset number of times adjacent to the reference audio data is obtained. Optionally, the preset number is 3 times, voice interaction process 1, voice interaction process 2, voice interaction process 3, interval time Time, voice interaction process 4, voice interaction process 5 and voice interaction process 6, and the target audio data is the audio data corresponding to voice interaction processes 1 to 6.

[0094] Step S40: using the target audio data as sample data, wherein the sample data is used for voice wake-up training.

[0095] Optionally, the target audio data and the corresponding time record are associated as sample data, wherein the sample data is used for voice wake-up training of the device to improve the accuracy of the wake-up of the device.

[0096] The benefit of using the pre-wake-up audio for training after uploading the pre-wake-up audio is that the wake-up accuracy and speed of the device voice assistant can be improved. Pre-wake-up audio uploading can collect more user voice data, which can be used to train the wake-up model of the voice assistant, thereby improving the wake-up accuracy.

[0097] At the same time, pre-wake-up audio upload can also improve the wake-up speed, because the voice assistant can recognize the user's wake-up word and respond more quickly. By using the wake-up uploaded audio for training, the voice assistant can better adapt to the user's voice characteristics and usage habits, improving the performance and user experience of the entire voice interaction system.

[0098] Optionally, a voice wake-up model is trained based on sample data, and the voice wake-up model is used to determine whether to wake up the device, so that when the voice device is in standby or dormant state, the device is activated by a specific keyword or phrase to enter the voice recognition mode. Once the device is awakened, it starts to listen to the user's voice input and transmits it to the background for voice recognition and processing.

[0099] In the technical solution of this embodiment, the interval time between voice interaction processes is determined; the confidence of the reference audio data in the interval time is determined; if the difference between the confidence and the preset confidence threshold is within a preset range, the target audio data of the voice interaction process adjacent to the reference audio data for a preset number of times is obtained; the target audio data is used as sample data, and the sample data is used for voice wake-up training. This is conducive to efficiently collecting truly effective wake-up audio data, using the audio data for later training, improving the accuracy of user wake-up, and meeting user needs.

[0100] Reference Figure 3 , Figure 3This is a second embodiment of the voice wake-up data processing method of the present invention. Based on the first embodiment, after step S40, the method further includes:

[0101] Step S50, performing model training based on the target audio data and the corresponding interaction success time to obtain a prediction model, wherein the prediction model is used to determine a prediction time period of the reference audio data.

[0102] The reference audio data is audio data that fails to wake up but has upload value, and is used for voice wake-up time period prediction.

[0103] Optionally, obtain the actual confidence value (wake-up success) H1 when waking up the local sound; and the total interaction time Ta of the voice control command, for example, 3 voice commands are issued within 1 minute; the interval between the last interactive user's voice commands △ T; the absolute time of the start time of this multi-round interaction Ts; the accent recorded at the time of successful user interaction, i.e., the target audio data and other audio features are sample data M. Using all the above features as independent variables, the dependent variable is the window period for the next valuable pre-wake-up audio to occur, i.e., the predicted time period. Prediction time period T = f(x1, x2, x3, x4, x5...); where x1, x2, x3, x4, x5, etc. correspond to the above H1, Ta, △ T, Ts, M, etc.

[0104] Optionally, determine the regional information corresponding to the target audio data; for example, target audio data corresponding to different provinces. Classify the sample data according to the regional information; for example, classify the sample data of province A into one category, and classify the sample data of province B into another category. Perform model training on the classified sample data and the corresponding interaction success time to obtain a prediction model corresponding to each of the regional information. For example, perform model training on the sample data of province A to obtain a prediction model corresponding to province A, and perform model training on the sample data of province B to obtain a prediction model corresponding to province B.

[0105] Optionally, after training to obtain the prediction model, real-time audio data is obtained; the real-time audio data is predicted based on the prediction model to obtain the prediction time period corresponding to the reference audio data; audio data within the prediction time period with a confidence greater than a preset confidence threshold is used as the reference audio data, and the target audio data of the voice interaction process that is adjacent to the reference audio data for a preset number of times is obtained. Optionally, the reference audio data can be used as the pre-wake-up data of the device to train the device to implement the pre-wake-up function.

[0106] In the technical solution of this embodiment, a prediction model is obtained through model training based on the target audio data and the corresponding interaction success time. This is conducive to efficiently collecting truly effective wake-up audio data, using the audio data for later training, improving the accuracy of user wake-up, and meeting user needs.

[0107] In one embodiment, as Figure 4 shown, record the actual confidence values M(T), M(T - 1), M(T - 2), M(T - n) each time during local wake-up voice collection; where one successful wake-up and sending one command is regarded as a group of interactions; when there are multiple consecutive interactions, it can be considered that the interaction interval time Time between groups during the continuous process is relatively stable, and when Time suddenly becomes longer, it can be considered that valuable data may be generated. Ts is the stable interval value; Tm is the maximum value between two effective interactions, and if it exceeds this value, this logical method is not applicable. When in the normal range Ts < Time < Tm, and the error between M(Time) and the effective wake-up confidence is within the expected range, the wake-up audio of the previous and subsequent consecutive N interactions of this audio is uploaded.

[0108] In one embodiment, as Figure 5 shown, obtain the actual confidence value (when wake-up is successful) H1 during local wake-up voice collection; and the total interaction duration Ta of the voice control command, for example, 3 voice commands are issued within 1 minute; the interval △ T between the last interaction and the user's voice command; the start time Ts of the absolute time of this multi-round interaction; the accent of the user interaction success time record, that is, the audio features such as the target audio data, etc., are used as sample data M. Using all the above features as independent variables, the dependent variable is the window period, that is, the prediction time period, when the next valuable pre-wake-up audio occurs. The prediction time period T = f(x1, x2, x3, x4, x5...); where x1, x2, x3, x4, x5, etc. respectively correspond to the above H1, Ta, △ T, Ts, M, etc. According to the time calculated by the model, the upload threshold range of the pre-wake-up audio is sent down, including the on and off times of the audio upload corresponding to the window period when the next valuable pre-wake-up audio occurs. When the audio data reaches the confidence threshold and hits the prediction time period, it is uploaded.

[0109] Exemplarily, when the user is not clear about the pronunciation method of the device's wake-up sound, or there are differences in each reading, at this time, the user interacts 3 times within 1 minute, with both successful and failed wake-ups. Machine learning is performed using the relevant data before and after the successful wake-up to generate a model, and the prediction time period of the next wake-up audio with a failed wake-up but with upload value is calculated, improving the effectiveness of the uploaded data.

[0110] The present invention also provides a voice wake-up data processing device, which includes a memory, a processor, and a voice wake-up data processing program stored in the memory and executable on the processor. When the voice wake-up data processing program is executed by the processor, the various steps of the voice wake-up data processing method described in the above embodiment are implemented.

[0111] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a voice wake-up data processing program, and when the voice wake-up data processing program is executed by a processor, the various steps of the voice wake-up data processing method described in the above embodiment are implemented.

[0112] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0113] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, system, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, system, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, system, article or device including the element.

[0114] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment system can be implemented by means of software plus a necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a computer-readable storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a parking management device, an air conditioner, or a network device, etc.) to execute the system described in each embodiment of the present invention.

[0115] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A method for processing voice wake-up data, characterized in that: The voice wake-up data processing method comprises: Determine the interval time between voice interaction processes, where the voice interaction process is an interaction process in which the voice successfully wakes up the device; Determining the confidence level of the reference audio data in the interval time; If the difference between the confidence and a preset confidence threshold is within a preset range, obtaining target audio data of the voice interaction process that is adjacent to the reference audio data for a preset number of times; The target audio data is used as sample data, and the sample data is used for voice wake-up training.

2. The voice wake-up data processing method according to claim 1, characterized in that: After the step of determining the interval time between the voice interaction processes, the method further includes: When the duration of the interval time is greater than a preset steady-state interval value and less than a preset interval maximum value, the step of determining the confidence of the reference audio data in the interval time is performed.

3. The voice wake-up data processing method according to claim 1, characterized in that: After the step of using the target audio data as sample data, the method further includes: Model training is performed based on the target audio data and the corresponding interaction success time to obtain a prediction model, and the prediction model is used to determine the prediction time period of the reference audio data.

4. The voice wake-up data processing method according to claim 3, characterized in that: The step of performing model training based on the target audio data and the corresponding interaction success time to obtain a prediction model comprises: Determining the region information corresponding to the target audio data; Classifying the sample data according to the regional information; The classified sample data and the corresponding interaction success time are used for model training to obtain a prediction model corresponding to each of the regional information.

5. The voice wake-up data processing method according to claim 3, characterized in that: After the step of performing model training based on the target audio data and the corresponding interaction success time to obtain a prediction model, the method further includes: Get real-time audio data; Predicting the real-time audio data based on the prediction model to obtain a prediction time period corresponding to the reference audio data; According to the audio data within the prediction time period and with a confidence level greater than a preset confidence threshold, as the reference audio data, the target audio data of the voice interaction process adjacent to the reference audio data for a preset number of times is obtained.

6. The voice wake-up data processing method according to claim 1, characterized in that: The step of determining the confidence of the audio data in the interval time comprises: Extracting audio features of the reference audio data; Matching the reference audio features with preset keywords to obtain a matching result of the reference audio data; The confidence level of the reference audio data is determined according to the matching result.

7. The voice wake-up data processing method according to claim 6, characterized in that: The step of matching the audio feature with the preset keyword to obtain the matching result of the reference audio data comprises: Obtaining a wake-up word and a command word in the reference audio data; respectively determining a first matching result between the wake-up word and a preset wake-up keyword, and a second matching result between the instruction word and a preset instruction keyword; A matching result of the reference audio data is determined according to the first matching result and the second matching result.

8. The voice wake-up data processing method according to claim 1, characterized in that: The method further comprises: A voice wake-up model is trained based on the sample data, and the voice wake-up model is used to determine whether to wake up the device.

9. A voice wake-up data processing device, characterized in that: The voice wake-up data processing device includes a memory, a processor, and a voice wake-up data processing program stored in the memory and executable on the processor. When the voice wake-up data processing program is executed by the processor, the various steps of the voice wake-up data processing method according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a voice wake-up data processing program, and when the voice wake-up data processing program is executed by a processor, each step of the voice wake-up data processing method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Wake-up word recognition method based on eigenvector iterative optimization

    CN121214933A