A voice wake-up method and related devices
By setting high threshold and low threshold conditions in the voice stream, combining historical wake-up recognition results and wake-up audio band update model, the problem of poor wake-up effect in existing voice wake-up technologies is solved, improving the wake-up rate and reducing the false wake-up rate.
Patent Information
- Application Number
- CN202111493721.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2041-12-08
AI Technical Summary
The existing voice wake-up technology has the problem of poor wake-up effect, especially when processing audio data that is not easy to wake up, the wake-up rate is low and the false wake-up rate is high.
By obtaining the current voice segment in real time in the voice stream, performing wake-up recognition processing, setting high threshold and low threshold conditions triggering wake-up instructions, and updating the wake-up model with historical wake-up recognition results and wake-up audio bands, improving wake-up recognition performance.
It effectively improves the wake-up rate of audio data that is not easy to wake up, reduces the false wake-up rate, and improves the voice wake-up effect.
Smart Images

Figure CN114141233B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a voice wake-up method and related devices thereof. Background Art
[0002] Voice wake-up technology is a common artificial intelligence technology; and this voice wake-up technology is used to determine whether to wake up a certain service item (for example, TV drama search function, movie search function, navigation, etc.) according to user voice data. For example, when the user says "Little A, navigate to XXX Community", the terminal device can wake up the navigation service based on this user voice data, so that the terminal device can display the corresponding navigation route.
[0003] However, due to the defects of the voice wake-up technology, the voice wake-up effect is relatively poor. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to provide a voice wake-up method and related devices thereof, which can improve the voice wake-up effect.
[0005] The embodiments of this application provide a voice wake-up method, and the method includes:
[0006] Obtain the current voice segment;
[0007] Perform wake-up recognition processing on the current voice segment to obtain a current wake-up recognition result;
[0008] When the current wake-up recognition result meets the high-threshold wake-up condition, trigger a wake-up instruction;
[0009] When the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, trigger a wake-up instruction.
[0010] In a possible implementation manner, the at least one historical wake-up recognition result includes the previous wake-up recognition result; wherein, the previous wake-up recognition result is obtained by performing wake-up recognition processing on the previous voice segment of the current voice segment; the acquisition time of the previous voice segment is earlier than the acquisition time of the current voice segment;
[0011] The step of triggering a wake-up instruction when the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result includes:
[0012] When the current wake-up recognition result meets the low-threshold wake-up condition and the previous wake-up recognition result meets the low-threshold wake-up condition, trigger a wake-up instruction.
[0013] In a possible implementation, when the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, triggering a wake-up instruction includes:
[0014] When the current wake-up recognition result meets the low-threshold wake-up condition, there is a target wake-up recognition result that meets the low-threshold wake-up condition among the at least one historical wake-up recognition result, and the time difference representation data between the target wake-up recognition result and the current wake-up recognition result meets the first time difference condition, a wake-up instruction is triggered.
[0015] In a possible implementation, the at least one historical wake-up recognition result is obtained by performing wake-up recognition processing on at least one historical speech segment;
[0016] The method further includes:
[0017] When the current wake-up recognition result does not meet the low-threshold wake-up condition, there is a target wake-up recognition result that meets the low-threshold wake-up condition among the at least one historical wake-up recognition result, and the current speech segment does not meet the preset message repetition condition, the at least one historical speech segment is updated according to the current speech segment;
[0018] When the current wake-up recognition result indicates that the current speech segment does not meet the low-threshold wake-up condition, there is a target wake-up recognition result that meets the low-threshold wake-up condition among the at least one historical wake-up recognition result, and the current speech segment meets the preset message repetition condition, the current speech segment is discarded.
[0019] In a possible implementation, the method further includes:
[0020] If the message recognition result of the current speech segment is the same as the message recognition result of the speech segment to be referenced, it is determined that the current speech segment meets the preset message repetition condition; wherein, the acquisition time difference between the speech segment to be referenced and the current speech segment meets the second time difference condition;
[0021] If the message recognition result of the current speech segment is different from the message recognition result of the speech segment to be referenced, it is determined that the current speech segment does not meet the preset message repetition condition.
[0022] In a possible implementation, the current wake-up recognition result is determined using a wake-up recognition model;
[0023] The method further includes:
[0024] When the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, determine the training data to be used according to the wake-up audio segment; wherein, the wake-up audio segment includes the current speech segment and the historical speech segment with the target wake-up recognition result; the training data to be used is used to update the wake-up recognition model.
[0025] In a possible implementation manner, the current wake-up recognition result is determined by using a wake-up recognition model;
[0026] The method further includes:
[0027] When the current wake-up recognition result meets the high-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, determine the training data to be used according to the wake-up audio segment; wherein, the wake-up audio segment includes the current speech segment and the historical speech segment with the target wake-up recognition result; the training data to be used is used to update the wake-up recognition model.
[0028] In a possible implementation manner, the current wake-up recognition result is determined by using a wake-up recognition model;
[0029] The method further includes:
[0030] When the current wake-up recognition result meets the low-threshold wake-up condition and the cloud wake-up recognition result of the current speech segment meets the normal wake-up condition, determine the training data to be used according to the wake-up audio segment; wherein, the cloud wake-up recognition result is obtained by performing wake-up recognition processing on the current speech segment by a cloud wake-up model; the wake-up audio segment includes the current speech segment; the training data to be used is used to update the wake-up recognition model and the cloud wake-up model.
[0031] In a possible implementation manner, the current wake-up recognition result is determined by using a wake-up recognition model;
[0032] The method further includes:
[0033] When the current wake-up recognition result meets the non-wake-up condition and the cloud wake-up recognition result of the current speech segment meets the normal wake-up condition, determine the training data to be used according to the wake-up audio segment; wherein, the cloud wake-up recognition result is obtained by performing wake-up recognition processing on the current speech segment by a cloud wake-up model; the wake-up audio segment includes the current speech segment; the training data to be used is used to update the wake-up recognition model and the cloud wake-up model.
[0034] In a possible implementation manner, the current wake-up recognition result is determined by using a wake-up recognition model;
[0035] The method further includes:
[0036] When the current wake-up recognition result meets the low-threshold wake-up condition and the cloud wake-up recognition result of the current speech segment does not meet the normal wake-up condition, determine the training data to be used according to the wake-up audio segment; wherein, the cloud wake-up recognition result is obtained by the cloud wake-up model performing wake-up recognition processing on the current speech segment; the wake-up audio segment includes the current speech segment; the training data to be used is used to update the wake-up recognition model and the cloud wake-up model.
[0037] In a possible implementation manner, the performing wake-up recognition processing on the current speech segment to obtain a current wake-up recognition result includes: [[ID=ll]]
[0038] If in an unawakened state, use the wake-up recognition model to perform wake-up recognition processing on the current speech segment to obtain a current wake-up recognition result;
[0039] The method further includes:
[0040] If in an awakened state, when it is determined that the current speech segment meets the preset information abnormality condition, determine the training data to be used according to the trigger audio segment of the awakened state; wherein, the training data to be used is used to update the wake-up recognition model.
[0041] In a possible implementation manner, the method is applied to a terminal device; the current wake-up recognition result is determined by using a wake-up recognition model;
[0042] The method further includes:
[0043] When the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, send the wake-up audio segment and the attribute information of the wake-up audio segment to a cloud server, so that the cloud server determines the training data to be used according to the wake-up audio segment and the attribute information of the wake-up audio segment; wherein, the wake-up audio segment includes the current speech segment; the training data to be used is used to update the wake-up recognition model.
[0044] An embodiment of the present application further provides a voice wake-up device, including:
[0045] A voice acquisition unit, configured to acquire a current speech segment;
[0046] A wake-up recognition unit, configured to perform wake-up recognition processing on the current speech segment to obtain a current wake-up recognition result;
[0047] A first trigger unit, configured to trigger a wake-up instruction when the current wake-up recognition result meets a high-threshold wake-up condition;
[0048] A second trigger unit, configured to trigger a wake-up instruction when the current wake-up recognition result meets a low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result.
[0049] An embodiment of the present application further provides a device, where the device includes: a processor, a memory, and a system bus;
[0050] The processor and the memory are connected through the system bus;
[0051] The memory is configured to store one or more programs, and the one or more programs include instructions that, when executed by the processor, cause the processor to execute any implementation manner of the voice wake-up method provided by the embodiment of the present application.
[0052] An embodiment of the present application further provides a computer-readable storage medium, where instructions are stored in the computer-readable storage medium, and when the instructions are run on a terminal device, the terminal device is caused to execute any implementation manner of the voice wake-up method provided by the embodiment of the present application.
[0053] An embodiment of the present application further provides a computer program product, and when the computer program product runs on a terminal device, the terminal device is caused to execute any implementation manner of the voice wake-up method provided by the embodiment of the present application.
[0054] Based on the above technical solutions, the present application has the following beneficial effects:
[0055] In the technical solution provided by the present application, for a terminal device (for example, a smart TV, a smart phone, an in-vehicle intelligent system, etc.), after obtaining the current voice segment from the voice stream in real time, the current voice segment is subjected to wake-up recognition processing to obtain a current wake-up recognition result; if the current wake-up recognition result meets the high-threshold wake-up condition, a wake-up instruction is triggered to wake up a certain service item in the terminal device; if the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, a wake-up instruction is triggered to wake up a certain service item in the terminal device.
[0056] It can be seen that the voice wake-up method provided by the embodiments of the present application can not only perform wake-up processing on audio data that meets the high-threshold wake-up conditions, but also perform wake-up processing on audio data that, although it never meets the high-threshold wake-up conditions, meets the low-threshold wake-up conditions twice. In this way, the wake-up difficulty of certain audio data (such as difficult-to-wake audio data, etc.) can be effectively reduced, thereby effectively improving the wake-up rate, and further facilitating the improvement of the voice wake-up effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0058] Figure 1 Schematic diagram of a voice stream provided by an embodiment of the present application;
[0059] Figure 2 Flowchart of a voice wake-up method provided by an embodiment of the present application;
[0060] Figure 3 Schematic diagram of a voice wake-up technology provided by an embodiment of the present application;
[0061] Figure 4 Schematic diagram of the attribute information of an audio data provided by an embodiment of the present application;
[0062] Figure 5 Schematic diagram of a wake-up type provided by an embodiment of the present application;
[0063] Figure 6 Schematic diagram of mis-triggered audio data provided by an embodiment of the present application;
[0064] Figure 7 Schematic diagram of the structure of a voice wake-up device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] In their research on voice wake-up technology, the inventors discovered that the technology can be implemented using a pre-built neural network model with wake-up recognition capabilities. Furthermore, the neural network model can perform wake-up recognition processing on a piece of audio data to obtain the wake-up probability of the audio data, so that the wake-up instruction is triggered only after the wake-up probability reaches a preset threshold. This indicates that the wake-up performance of the voice wake-up technology implemented based on the neural network model depends on the preset threshold. If the preset threshold is too high, the wake-up rate of the voice wake-up technology may be low because the wake-up probability of a large amount of audio data cannot exceed the preset threshold. If the preset threshold is too low, the wake-up probability of the majority of audio data may exceed the preset threshold, resulting in a high false alarm rate.
[0066] The inventors also found that the voice wake-up technology is usually aimed at voice streams (such as Figure 1 Moreover, for a voice stream including an audio segment that is not easy to wake up, although the wake-up probability of multiple voice segments in the voice stream cannot reach the above-mentioned "preset threshold", the wake-up probability of these voice collection segments is always relatively high.
[0067] Based on the above findings, in order to overcome the technical problems shown in the background technology part, an embodiment of the present application provides a voice wake-up method, which includes: after obtaining the current voice segment from the voice stream in real time, performing wake-up recognition processing on the current voice segment to obtain a current wake-up recognition result; if the current wake-up recognition result meets the high threshold wake-up condition, triggering a wake-up instruction to wake up a certain service item; if the current wake-up recognition result meets the low threshold wake-up condition, and there is a target wake-up recognition result that meets the low threshold wake-up condition in at least one historical wake-up recognition result, triggering a wake-up instruction to wake up a certain service item.
[0068] It can be seen that the voice wake-up method provided in the embodiment of the present application can not only perform wake-up processing on audio data that meet the high threshold wake-up conditions, but also perform wake-up processing on audio data that does not meet the high threshold wake-up conditions but meets the low threshold wake-up conditions twice. This can effectively reduce the wake-up difficulty of certain audio data (for example, audio data that is not easy to wake up, etc.), thereby effectively improving the wake-up rate, and thus helping to improve the voice wake-up effect.
[0069] In addition, the execution subject of the voice wake-up method in the embodiments of the present application is not limited. For example, the voice wake-up method provided in the embodiments of the present application can be applied to data processing devices such as terminal devices or servers. Among them, the terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be an independent server, a cluster server, or a cloud server. Also, in some cases, the voice wake-up method provided in the embodiments of the present application can also be implemented by the cooperation of the terminal device and the server.
[0070] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0071] Method Embodiment 1
[0072] See Figure 2 , which is a flowchart of a voice wake-up method provided by an embodiment of the present application.
[0073] The voice wake-up method provided by the embodiments of the present application includes S1 - S4:
[0074] S1: Obtain the current voice segment.
[0075] The above-mentioned "current voice segment" is used to represent a voice segment collected in real time from a voice stream. For example, as Figure 1 shown, when the "third voice segment" is collected from the voice stream as Figure 1 shown, the third voice segment can be determined as the current voice segment, so that subsequent data processing for the third voice segment can be realized by means of the data processing process for the current voice segment.
[0076] In addition, the execution subject of S1 in the embodiments of the present application is not limited. For example, it can be a terminal device (such as a smart TV, a smart phone, an in-vehicle intelligent system, etc.).
[0077] Based on the relevant content of S1 above, for a terminal device with a voice wake-up function, the terminal device can perform real-time voice acquisition and processing on the voice stream to obtain the current voice segment, so that it can be determined whether the user has the intention to wake up the voice control function in the terminal device based on the current voice segment.
[0078] S2: Perform wake-up recognition processing on the current voice segment to obtain the current wake-up recognition result.
[0079] The above-mentioned "current wake-up recognition result" refers to the wake-up probability of the current speech segment (that is, the occurrence probability of wake-up generated by the current speech segment), so that the "current wake-up recognition result" can indicate whether the current speech segment can generate wake-up (that is, wake up the voice control function in the terminal device).
[0080] In addition, the embodiment of the present application does not limit the determination process of the above-mentioned "current wake-up recognition result". For example, specifically, it may include: using a pre-constructed wake-up recognition model to perform wake-up recognition processing on the current speech segment to obtain the current wake-up recognition result.
[0081] The above-mentioned "wake-up recognition model" is used to perform wake-up recognition processing on the input data of the wake-up recognition model; and the "wake-up recognition model" is a machine learning model (for example, Recurrent Neural Network (RNN), Convolutional Neural Networks (CNN), etc.).
[0082] It should be noted that the embodiment of the present application does not limit the construction process of the above-mentioned "wake-up recognition model", and any existing or future construction method of the wake-up recognition model can be used for implementation.
[0083] In addition, the embodiment of the present application does not limit the execution subject of S2. For example, in order to improve the real-time performance of voice wake-up, the execution subject of S2 can be a terminal device.
[0084] Based on the relevant content of the above S2, after the current speech segment is obtained in real time from the speech stream, with the help of the wake-up recognition model, wake-up recognition processing is performed on the current speech segment to obtain the current wake-up recognition result, so that the current wake-up recognition result can indicate the occurrence probability of the current speech segment triggering a wake-up instruction, so as to determine whether the current speech segment generates wake-up based on the current wake-up recognition result subsequently.
[0085] S3: When the current wake-up recognition result meets the high-threshold wake-up condition, trigger a wake-up instruction.
[0086] The above-mentioned "high-threshold wake-up condition" can be preset; and the high-threshold wake-up condition may include: reaching a first probability threshold. It should be noted that the embodiment of the present application does not limit the above-mentioned "first probability threshold". For example, it may be a normal threshold value set for a machine learning model with a wake-up recognition function (for example, the "preset threshold" mentioned above).
[0087] The above-mentioned "wake-up instruction" is used to generate a wake-up (i.e., to switch the voice control function in the terminal device from the non-wake-up state to the wake-up state), so that other functional modules in the terminal device can learn about this wake-up state, so that these functional modules can provide corresponding service items to the user based on the subsequent voice data input by the user (for example, "Navigate to XXX Community"), such as providing the user with a navigation route to "Arrive at XXX Community".
[0088] In addition, the embodiment of the present application does not limit the execution subject of S3. For example, in order to improve the real-time performance of voice wake-up, the execution subject of S3 can be the terminal device.
[0089] Based on the relevant content of the above S3, after obtaining the current wake-up recognition result (that is, the wake-up recognition result for the current voice segment), it can be determined whether the current wake-up recognition result meets the high-threshold wake-up condition (for example, whether the wake-up rate of the current voice segment reaches the preset threshold), so that when it is determined that the current wake-up recognition result meets the high-threshold wake-up condition, it can be determined that the current voice segment has generated a wake-up, so the wake-up instruction can be directly triggered to inform other functional modules in the terminal device of the wake-up message, so that these functional modules can provide corresponding service items to the user based on the subsequent voice data input by the user.
[0090] S4: When the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, trigger the wake-up instruction.
[0091] The above-mentioned "low-threshold wake-up condition" can be preset; and the low-threshold wake-up condition can include: reaching the second probability threshold but not reaching the first probability threshold. Among them, the second probability threshold is less than the first probability threshold.
[0092] The above-mentioned "historical wake-up recognition result" refers to the result obtained by performing wake-up recognition processing on a voice segment (also called a historical voice segment) in the voice stream whose acquisition time is earlier than the above-mentioned "current voice segment". For example, for Figure 1 For the shown voice stream, if the above-mentioned "current voice segment" is the third voice segment, then the first voice segment and the second voice segment can both be a historical voice segment corresponding to the above-mentioned "current voice segment".
[0093] In addition, the embodiment of the present application does not limit the determination process of the above-mentioned "at least one historical wake-up recognition result". For example, in order to improve the voice wake-up effect, the determination process of the "at least one historical wake-up recognition result" can specifically include steps 11-step 12:
[0094] Step 11: Determine the historical reference interval according to the preset reference duration and the acquisition time of the current voice segment.
[0095] The above-mentioned "preset reference duration" is used to describe the speech duration of a trigger word; moreover, the embodiments of the present application do not limit the acquisition process of the "preset reference duration". For example, it can be preset in advance, especially according to the application scenario. Another example is that the "preset reference duration" can be mined from a large amount of user speech data.
[0096] The above-mentioned "historical reference interval" can be expressed as [the acquisition time of the current speech segment - the preset reference duration, the acquisition time of the current speech segment].
[0097] Step 12: Perform wake-up recognition processing on at least one speech segment collected from the speech stream within the historical reference interval to obtain at least one historical wake-up recognition result.
[0098] Based on the relevant content of the above-mentioned "at least one historical wake-up recognition result", the determination time of these historical wake-up recognition results is relatively close to the determination time of the above-mentioned "current wake-up recognition result". In this way, it can effectively avoid the interference of the historical speech wake-up process on the current wake-up process, which is beneficial to improving the speech wake-up effect.
[0099] The above-mentioned "target wake-up recognition result" refers to the historical wake-up recognition result that meets the low-threshold wake-up condition among the above-mentioned "at least one historical wake-up recognition result"; and the "target wake-up recognition result" is obtained by performing wake-up recognition processing on the target speech segment. Among them, the acquisition time of the "target speech segment" is earlier than the acquisition time of the above-mentioned "current speech segment".
[0100] In addition, the embodiments of the present application do not limit the execution subject of S4. For example, in order to improve the real-time performance of speech wake-up, the execution subject of S4 can be a terminal device.
[0101] Based on the relevant content of the above-mentioned S1 to S4, for the speech wake-up method provided by the embodiments of the present application, after the current speech segment is obtained in real time from the speech stream, wake-up recognition processing is performed on the current speech segment to obtain the current wake-up recognition result; if the current wake-up recognition result meets the high-threshold wake-up condition, a wake-up instruction is triggered to wake up a certain service item in the terminal device; if the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, a wake-up instruction is triggered to wake up a certain service item in the terminal device.
[0102] It can be seen that the voice wake-up method provided by the embodiments of the present application can not only perform wake-up processing on audio data that meets the high-threshold wake-up conditions, but also perform wake-up processing on audio data that, although not meeting the high-threshold wake-up conditions all the time, meets the low-threshold wake-up conditions twice. In this way, the wake-up difficulty of certain audio data (such as difficult-to-wake audio data, etc.) can be effectively reduced, thereby effectively improving the wake-up rate and further facilitating the improvement of the voice wake-up effect.
[0103] Method Embodiment 2
[0104] Actually, for a voice stream including difficult-to-wake audio segments, the wake-up probabilities of multiple consecutive voice segments or multiple voice segments with relatively close acquisition times (such as Figure 1 the "second voice segment" and "third voice segment" shown) in the voice stream are all between the second probability threshold and the first probability threshold. It can be seen that the above-mentioned "target voice segment" is adjacent to the above-mentioned "current voice segment"; or, although the above-mentioned "target voice segment" is not adjacent to the above-mentioned "current voice segment", the time interval between the acquisition time of the above-mentioned "target voice segment" and the acquisition time of the above-mentioned "current voice segment" is relatively small.
[0105] Based on this, in order to improve the voice wake-up effect, another possible implementation manner of S4 is provided in the embodiments of the present application, which may specifically include at least one of S41-S42:
[0106] S41: When the current wake-up recognition result meets the low-threshold wake-up condition and the previous wake-up recognition result meets the low-threshold wake-up condition, trigger a wake-up instruction.
[0107] The above-mentioned "previous wake-up recognition result" is obtained by performing wake-up recognition processing on the previous voice segment of the current voice segment; and the "previous wake-up recognition result" refers to the wake-up probability of the previous voice segment (that is, the occurrence probability of wake-up for the previous voice segment), so that the "previous wake-up recognition result" can indicate whether the previous voice segment can generate a wake-up.
[0108] It should be noted that the determination process of the above-mentioned "previous wake-up recognition result" is similar to the determination process of the "current wake-up recognition result" above.
[0109] The above-mentioned "previous voice segment" refers to a voice segment that is adjacent to the current voice segment and has an acquisition time earlier than the acquisition time of the current voice segment. For example, for Figure 1 the voice stream shown, if the above-mentioned "current voice segment" is the third voice segment, the above-mentioned "previous voice segment" may be the second voice segment.
[0110] Based on the relevant content of S41 above, after obtaining the current wake-up recognition result (i.e., the wake-up recognition result for the current speech segment), if it is determined that the current wake-up recognition result meets the low-threshold wake-up condition and the previous wake-up recognition result meets the low-threshold wake-up condition, it can be determined that the wake-up recognition results of two consecutive speech segments in the speech stream both meet the low-threshold wake-up condition. Thus, it can be inferred that the wake-up audio segment including these two speech segments belongs to the difficult-to-wake-up audio data. Therefore, the wake-up instruction can be directly triggered to inform other functional modules in the terminal device of this wake-up message, so that these functional modules can provide corresponding service items to the user based on the subsequent speech data input by the user.
[0111] S42: When the current wake-up recognition result meets the low-threshold wake-up condition, there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, and the time difference representation data between the target wake-up recognition result and the current wake-up recognition result meets the first time difference condition, trigger the wake-up instruction.
[0112] The above-mentioned "time difference representation data" is used to describe the time interval between two wake-up recognition results; and the embodiments of the present application do not limit the determination process of the "time difference representation data". For example, when the above-mentioned "target wake-up recognition result" is obtained by performing wake-up recognition processing on the target speech segment, the above-mentioned "time difference representation data between the target wake-up recognition result and the current wake-up recognition result" may include: the time difference between the acquisition time of the current speech segment and the acquisition time of the target speech segment.
[0113] It should be noted that the embodiments of the present application do not limit the relationship between the "current speech segment" and the "target speech segment" in S42. For example, the "current speech segment" and the "target speech segment" may be adjacent speech segments (e.g., the third speech segment and the second speech segment), or may be non-adjacent but speech segments with relatively close acquisition times (e.g., the third speech segment and the first speech segment).
[0114] In addition, the embodiments of the present application do not limit the determination method of the above-mentioned "time difference representation data". For example, it can be determined by means of a time difference calculation formula. Another example is that it can be determined with the help of Figure 3 the timer T shown.
[0115] The above-mentioned "first time difference condition" can be preset. For example, as Figure 3 shown, the "first time difference condition" can specifically be: belonging to the time difference region [1 second, 10 seconds].
[0116] Based on the relevant content of S42 above, after obtaining the current wake-up recognition result (i.e., the wake-up recognition result for the current speech segment), when it is determined that the current wake-up recognition result meets the low-threshold wake-up condition, the existing target wake-up recognition result meets the low-threshold wake-up condition, and the time-difference representation data between the target wake-up recognition result and the current wake-up recognition result meets the first time-difference condition, it can be determined that the wake-up recognition results of two speech segments with relatively close acquisition times in the speech stream both meet the low-threshold wake-up condition. Thus, it can be inferred that the wake-up audio segment including these two speech segments belongs to difficult-to-wake-up audio data. Therefore, the wake-up instruction can be directly triggered to inform other functional modules in the terminal device of this wake-up message, so that these functional modules can provide corresponding service items to the user based on the subsequent voice data input by the user.
[0117] Based on the relevant content of S4 above, for a speech stream, if the wake-up recognition results of two consecutive speech segments in the speech stream both meet the low-threshold wake-up condition, or the wake-up recognition results of two speech segments meet the low-threshold wake-up condition within a preset reference duration, it can be determined that the wake-up audio segment including these two speech segments belongs to difficult-to-wake-up audio data. Therefore, the wake-up instruction can be directly triggered to inform other functional modules in the terminal device of this wake-up message, so that these functional modules can provide corresponding service items to the user based on the subsequent voice data input by the user. In this way, the wake-up difficulty of difficult-to-wake-up audio data can be effectively reduced, thereby effectively improving the wake-up rate, and further facilitating the improvement of the voice wake-up effect.
[0118] Method Embodiment 3
[0119] In addition, in some cases, the speech stream may carry repeated character content (e.g., "Little A classmate, navigate to XXX community"). Therefore, to avoid the adverse effects of repeated character content on voice wake-up, another possible implementation manner of the voice wake-up method is provided in the embodiments of the present application. For ease of understanding, it will be described below with examples.
[0120] As an example, when the above "at least one historical wake-up recognition result" is obtained by performing wake-up recognition processing on at least one historical speech segment, the voice wake-up method may further include S5 - S6 in addition to the above S1 - S4:
[0121] S5: When the current wake-up recognition result does not meet the low-threshold wake-up condition, there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, and the current speech segment does not meet the preset message repetition condition, update at least one historical speech segment according to the current speech segment.
[0122] The above-mentioned "preset message repetition condition" can be preset; moreover, the working principle of the "preset message repetition condition" can specifically be: determining whether the message recognition result of the current speech segment is the same as that of the speech segment to be referenced. If so, it is determined that the current speech segment meets the preset message repetition condition; if not, it is determined that the current speech segment does not meet the preset message repetition condition.
[0123] The above-mentioned "message recognition result of the current speech segment" is used to represent the character information carried by the current speech segment; moreover, the embodiments of the present application do not limit the determination process of the "message recognition result of the current speech segment". For example, the current speech segment can be first subjected to speech recognition processing to obtain at least one candidate word; and then these words are combined to obtain the "message recognition result of the current speech segment".
[0124] The above-mentioned "speech segment to be referenced" refers to the speech segment required for determining whether the current speech segment meets the preset message repetition condition; moreover, the time difference between the acquisition time of the speech segment to be referenced and the current speech segment meets the second time difference condition.
[0125] It should be noted that the above-mentioned "second time difference condition" can be preset. For example, as Figure 3 shown, the time interval between the acquisition time of the speech segment to be referenced and the acquisition time of the current speech segment is less than 1 second; or, the time interval between the start point of acquisition of the speech segment to be referenced and the start point of acquisition of the current speech segment is less than 200 milliseconds.
[0126] The above-mentioned "message recognition result of the speech segment to be referenced" is used to represent the character information carried by the speech segment to be referenced; moreover, the determination process of the "message recognition result of the speech segment to be referenced" is similar to the determination process of the "message recognition result of the current speech segment" above.
[0127] Based on the relevant content of the above "preset message repetition condition", for the process of finding the second low-threshold message from the voice stream (that is, there is a target wake-up recognition result that meets the low-threshold wake-up condition in at least one historical wake-up recognition result), after determining that the current wake-up recognition result does not meet the low-threshold wake-up condition, it can be determined whether the message recognition result of the current voice segment is the same as that of the voice segment to be referenced. If they are the same, it means that the character content carried by the current voice segment is the same as that carried by the voice segment to be referenced. Therefore, it can be determined that the current voice segment meets the preset message repetition condition, and thus the current voice segment and its related content (such as wake-up recognition results, speech recognition results, etc.) can be deleted. This is beneficial to avoiding the interference caused by repeated characters to the process of finding the second low-threshold message. If they are different, it means that the character content carried by the current voice segment is different from that carried by the voice segment to be referenced. Therefore, it can be determined that the current voice segment does not meet the preset message repetition condition, and thus the current voice segment and its related content (such as wake-up recognition results, speech recognition results, etc.) can be retained for subsequent reference to continue the process of finding the second low-threshold message with the current voice segment and its related content.
[0128] It should be noted that the embodiments of the present application do not limit the implementation manner of "updating at least one historical voice segment according to the current voice segment" in S5. For example, specifically, it can be: first, delete the historical voice segment with the earliest acquisition time in at least one historical voice segment to obtain at least one remaining voice segment; then, combine the current voice segment with at least one remaining voice segment to obtain the updated at least one historical voice segment.
[0129] S6: When the current wake-up recognition result indicates that the current voice segment does not meet the low-threshold wake-up condition, there is a target wake-up recognition result that meets the low-threshold wake-up condition in at least one historical wake-up recognition result, and the current voice segment meets the preset message repetition condition, discard the current voice segment.
[0130] It should be noted that the embodiments of the present application do not limit the execution subject of S5 - S6. For example, in order to improve the real-time performance of voice wake-up, the execution subject of S5 - S6 can be a terminal device.
[0131] Based on the relevant content of S5 to S6 above, as Figure 3As shown, after determining that the current wake-up recognition result does not meet the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition in at least one historical wake-up recognition result, it can be determined whether the message recognition result of the current speech segment is the same as that of the speech segment to be referenced. If they are the same, the current speech segment can be directly discarded to avoid interference of the current speech segment on the subsequent wake-up process. If they are different, at least one historical speech segment is updated according to the current speech segment, so that the wake-up recognition results of the updated at least one historical speech segment can participate in the subsequent wake-up process as the "at least one historical wake-up recognition result". In this way, the automatic update process of the "at least one historical wake-up recognition result" can be realized, which is beneficial to improving the speech wake-up effect.
[0132] Method Embodiment 4
[0133] In fact, the above "wake-up recognition processing" is usually implemented by means of a wake-up recognition model; and this wake-up recognition model usually needs to be constructed with the help of a large amount of training data.
[0134] The inventor found in the research on the above training data that these training data can be obtained by combining and playing back some pre-recorded wake-up word audio data and multiple scene noise data. However, due to the large difference between the noise data involved in these training data and the noise data that appears in the actual application scenario, the difference between these training data and the speech data that appears in the actual application scenario is also relatively large. As a result, the wake-up recognition model constructed based on these training data cannot accurately perform wake-up recognition processing on the speech data that appears in these actual application scenarios, which easily leads to poor speech wake-up effect.
[0135] It can be seen that in order to overcome the above problems, the speech data that appears in the actual application scenario (especially, difficult-to-wake-up audio data, false wake-up audio data, etc.) can be used to update and train the wake-up recognition model, so that the updated wake-up recognition model has better wake-up recognition performance.
[0136] Based on this, another possible implementation manner of the speech wake-up method is further provided in the embodiments of the present application. In this implementation manner, the speech wake-up method may not only include all or part of the above steps, but may also include S7:
[0137] S7: When the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition in at least one historical wake-up recognition result, determine the training data to be used according to the wake-up audio segment.
[0138] The above-mentioned "wake-up audio segment" refers to audio data carrying a wake-up word; and the "wake-up audio segment" carries one wake-up word (for example, Little A, play, open, search, etc.).
[0139] In addition, when the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition in at least one historical wake-up recognition result, it can be determined that both the current speech segment and the target speech segment (that is, the historical speech segment with the target wake-up recognition result) carry partial characters of the wake-up word. Therefore, the wake-up audio segment can include the current speech segment and the target speech segment.
[0140] It should be noted that the embodiment of the present application does not limit the acquisition process of the above-mentioned "wake-up audio segment", and any existing or future determination process (for example, it can be implemented by extracting the wake-up audio segment from the speech stream with the speech recognition result performed on the speech stream in real time) can be used for implementation.
[0141] The above-mentioned "training data to be used" refers to the training data required for updating the above-mentioned "wake-up recognition model".
[0142] In addition, the embodiment of the present application does not limit the determination process of the above-mentioned "training data to be used". For example, if the update process of the above-mentioned "wake-up recognition model" is implemented on the cloud server, the determination process of the above-mentioned "training data to be used" can specifically include steps 21 - step 22:
[0143] Step 21: When the terminal device determines that the current wake-up recognition result meets the low-threshold wake-up condition and determines that there is a target wake-up recognition result that meets the low-threshold wake-up condition in at least one historical wake-up recognition result, the terminal device sends the wake-up audio segment and the attribute information of the wake-up audio segment to the cloud server.
[0144] The above-mentioned "attribute information of the wake-up audio segment" is used to describe the wake-up audio segment; and the embodiment of the present application does not limit the "attribute information". For example, as Figure 4 shown, the "attribute information" can include at least one of device information, data type, wake-up word annotation, wake-up type, speech recognition text annotation, time information, and whether it is from cloud wake-up.
[0145] The above-mentioned "wake-up word annotation" refers to the wake-up word annotated by the terminal device for the wake-up audio segment, so that the "wake-up word annotation" is used to represent the character content carried by the wake-up audio segment; and the embodiment of the present application does not limit the determination method of the "wake-up word annotation". For example, it can specifically be: first match the speech recognition text of the wake-up audio segment with at least one candidate wake-up word; then determine the candidate wake-up word with a successful match as the wake-up word annotation of the wake-up audio segment.
[0146] It should be noted that the "at least one candidate wake-up word" refers to the words (such as, classmate Xiao A, play, open, search, etc.) that are pre-stored in the terminal device and can trigger a wake-up instruction, so that these words can be referred to when the terminal device performs wake-up recognition processing on an audio data.
[0147] The above "wake-up type" is used to represent the wake-up method of the wake-up audio segment; and the embodiments of the present application do not limit this "wake-up type". For example, it can be implemented Figure 5 using the "wake-up type" shown. It should be noted that Figure 5 in, "high-threshold wake-up" is used to represent meeting the high-threshold wake-up condition; "low-threshold wake-up" is used to represent meeting the low-threshold wake-up condition.
[0148] Based on the relevant content of step 21 above, for the terminal device, when the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition in at least one historical wake-up recognition result, it can first determine the wake-up audio segment including the current speech segment and the target speech segment; then upload the wake-up audio segment and the attribute information of the wake-up audio segment to the cloud server, so that the cloud server can refer to the wake-up audio segment and the attribute information of the wake-up audio segment to update the training data of the wake-up recognition model.
[0149] Step 22: The cloud server determines the training data to be used according to the wake-up audio segment and the attribute information of the wake-up audio segment.
[0150] It should be noted that the embodiments of the present application do not limit the implementation manner of step 22. For example, it can specifically include: first, determine a difficult-to-wake-up audio sample according to the wake-up audio segment, and determine the label information of the difficult-to-wake-up audio sample according to the wake-up word annotation of the wake-up audio segment (or the wake-up word manually annotated for the wake-up audio segment); then use the difficult-to-wake-up audio sample and the label information of the difficult-to-wake-up audio sample to determine the training data to be used, so that the training data to be used includes the difficult-to-wake-up audio sample and the label information of the difficult-to-wake-up audio sample, so as to be able to perform model update training on the wake-up recognition model according to the difficult-to-wake-up audio sample and the label information of the difficult-to-wake-up audio sample subsequently.
[0151] Based on the relevant content of S7 above, after obtaining the current wake-up recognition result, if it is determined that the current wake-up recognition result meets the low-threshold wake-up condition, and it is determined that there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, then it can be determined that the wake-up audio segment including the current speech segment and the target speech segment belongs to the audio data that is not easily processed by the wake-up recognition model for wake-up recognition. Thus, it can be determined that the wake-up audio segment belongs to the not-easily-wake-up audio data. Therefore, the training data to be used can be determined according to the wake-up audio segment, so that the training data to be used can include the wake-up audio segment, so that the wake-up audio segment can be used as a positive example for model update training of the wake-up recognition model in the future, so that the updated wake-up recognition model can better perform wake-up recognition processing on the wake-up audio segment. This is beneficial to improving the wake-up recognition performance of the wake-up recognition model, and thus beneficial to improving the voice wake-up effect.
[0152] In addition, in order to further improve the collection effect of the not-easily-wake-up audio data, another possible implementation manner of the voice wake-up method is provided in the embodiments of the present application. In this implementation manner, the voice wake-up method may not only include all or part of the above steps, but may also include S8:
[0153] S8: When the current wake-up recognition result meets the high-threshold wake-up condition, and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, determine the training data to be used according to the wake-up audio segment.
[0154] It can be seen that after obtaining the current wake-up recognition result, if it is determined that the current wake-up recognition result meets the high-threshold wake-up condition, and it is determined that there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, then it can be determined that the wake-up audio segment including the current speech segment and the target speech segment belongs to the audio data that is not easily processed by the wake-up recognition model for wake-up recognition. Thus, it can be determined that the wake-up audio segment belongs to the not-easily-wake-up audio data. Therefore, the training data to be used can be determined according to the wake-up audio segment, so that the training data to be used can include the wake-up audio segment. In this way, it is possible to collect and process the audio data with the wake-up type of "low-threshold wake-up → high-threshold wake-up", which is beneficial to improving the comprehensiveness of the not-easily-wake-up audio data, thus beneficial to improving the comprehensiveness of the training data to be used, and further beneficial to improving the voice wake-up effect.
[0155] It should be noted that "determining the training data to be used according to the wake-up audio segment" in S8 is similar to the determination process of "the training data to be used" in S7.
[0156] In addition, to further improve the collection effect of audio data that is not easily awakened, another possible implementation of the voice wake-up method is provided in the embodiments of the present application. In this implementation, the voice wake-up method may not only include all or some of the above steps, but may also include S9:
[0157] S9: When the current wake-up recognition result meets the low-threshold wake-up condition and the cloud wake-up recognition result of the current voice segment meets the normal wake-up condition, determine the training data to be used according to the wake-up audio segment.
[0158] The above "cloud wake-up recognition result" is obtained by the cloud wake-up model performing wake-up recognition processing on the current voice segment; and the "cloud wake-up recognition result" is similar to the above "current wake-up recognition result".
[0159] The above "cloud wake-up model" is used to perform wake-up recognition processing on the input data of the cloud wake-up model; and the "cloud wake-up model" is a machine learning model.
[0160] The embodiments of the present application do not limit the association relationship between the above "cloud wake-up model" and the above "wake-up recognition model". For example, the network structure of the above "cloud wake-up model" is more complex than that of the above "wake-up recognition model"; and the wake-up recognition performance of the "cloud wake-up model" is much higher than that of the "wake-up recognition model".
[0161] The embodiments of the present application do not limit the deployment relationship between the above "cloud wake-up model" and the above "wake-up recognition model". For example, due to the limited computing resources of the terminal device, in order to ensure that the terminal device can better serve the user, the wake-up recognition model is usually deployed on the terminal device so that the terminal device can use the wake-up recognition model to perform wake-up trigger processing; and, due to the relatively rich computing resources of the cloud server, in order to improve the wake-up recognition effect, the cloud wake-up model can be deployed on the cloud server so that the cloud server can use the cloud wake-up model to verify the wake-up recognition results generated in the terminal device. Among them, the terminal device can communicate with the cloud server for data.
[0162] In addition, the embodiments of the present application do not limit the acquisition process of the above "cloud wake-up recognition result". For example, specifically, it may include: after the terminal device obtains the current voice segment, the terminal device may send the current voice segment to the cloud server so that the cloud server can use the cloud wake-up model to perform wake-up recognition processing on the current voice segment to obtain the cloud wake-up recognition result of the current voice segment; then, the cloud server may feedback the cloud wake-up recognition result to the terminal device so that the terminal device can use the cloud wake-up recognition result to determine whether the above "current wake-up recognition result" is correct.
[0163] The above "normal wake-up condition" can be preset. For example, the "normal wake-up condition" may include: reaching a first probability threshold.
[0164] Based on the relevant content of S9 above, after obtaining the current wake-up recognition result, if it is determined that the current wake-up recognition result meets the low-threshold wake-up condition, and it is determined that the cloud wake-up recognition result of the current speech segment meets the normal wake-up condition, it can be determined that the current speech segment should actually be able to generate a wake-up. However, due to the limitations of the wake-up recognition model itself, the wake-up recognition model cannot perform accurate wake-up recognition processing on the current speech segment, so that the wake-up audio segment including the current speech segment belongs to the audio data that is not easily recognized by the wake-up recognition model for wake-up recognition processing. Therefore, it can be determined that the wake-up audio segment belongs to the not easily wake-up audio data. Thus, the training data to be used can be determined according to the wake-up audio segment, so that the training data to be used can include the wake-up audio segment. In this way, it is possible to collect and process audio data with the wake-up type of "local low-threshold wake-up → cloud-generated wake-up", which is beneficial to improving the comprehensiveness of the not easily wake-up audio data, thus beneficial to improving the comprehensiveness of the training data to be used, and further beneficial to improving the voice wake-up effect.
[0165] It should be noted that the process of "determining the training data to be used according to the wake-up audio segment" in S9 is similar to the process of determining the "training data to be used" in S7.
[0166] It should also be noted that for any embodiment in the present application that simultaneously involves the "wake-up recognition model" (also known as the local wake-up model) and the "cloud wake-up model", the above "training data to be used" is used to update the wake-up recognition model and the cloud wake-up model.
[0167] Furthermore, it should be noted that the present application embodiment does not limit the update process of the above "wake-up recognition model and cloud wake-up model". For example, the training data to be used can be used to perform update training on the wake-up recognition model and the cloud wake-up model respectively to obtain an updated wake-up recognition model and an updated cloud wake-up model. Another example is that the training data to be used can be used to update and train the cloud wake-up model first to obtain an updated cloud wake-up model; then, according to the preset model simplification rules, the updated cloud wake-up model is simplified to obtain an updated cloud wake-up model.
[0168] In addition, in order to further improve the collection effect of the not easily wake-up audio data, the present application embodiment also provides another possible implementation manner of the voice wake-up method. In this implementation manner, the voice wake-up method may not only include all or some of the above steps, but may also include S10:
[0169] S10: When the current wake-up recognition result meets the non-wake-up condition and the cloud wake-up recognition result of the current voice segment meets the normal wake-up condition, determine the training data to be used according to the wake-up audio segment.
[0170] The above "non-wake-up condition" can be set in advance; for example, specifically, it can be: lower than the second probability threshold. It should be noted that for the relevant content of the above "second probability threshold", please refer to S4 above.
[0171] Based on the relevant content of S10, it can be known that after obtaining the current wake-up recognition result, if it is determined that the current wake-up recognition result meets the non-wake-up condition and the cloud wake-up recognition result of the current voice segment meets the normal wake-up condition, it can be determined that the current voice segment should actually be able to generate a wake-up, but due to the limitations of the wake-up recognition model itself, it is completely impossible to generate a wake-up for the current voice segment by means of the wake-up recognition model, so that the wake-up audio segment including the current voice segment belongs to the audio data that cannot be processed by the wake-up recognition model for wake-up recognition, so it can be determined that the wake-up audio segment belongs to the audio data that is not easily woken up. Therefore, the training data to be used can be determined according to the wake-up audio segment, so that the training data to be used can include the wake-up audio segment. In this way, it is possible to collect and process the audio data with the wake-up type of "locally unable to wake up → cloud generates wake-up", which is beneficial to improving the comprehensiveness of the audio data that is not easily woken up, thus beneficial to improving the comprehensiveness of the training data to be used, and further beneficial to improving the voice wake-up effect.
[0172] It should be noted that "determining the training data to be used according to the wake-up audio segment" in S10 is similar to the determination process of "the training data to be used" in S7.
[0173] Method Embodiment 5
[0174] In fact, since there are not only abnormal situations where it is not easy to wake up but also abnormal situations of false wake-up in the voice wake-up technology, the voice wake-up performance of this voice wake-up technology can be comprehensively measured by the wake-up rate and the false wake-up rate. Based on this, it can be known that in order to improve the wake-up recognition efficiency of the above "wake-up recognition model", a large number of real false wake-up audio data can be collected during the actual application process of the "wake-up recognition model" so that the "wake-up recognition model" can be updated based on these false wake-up audio data later.
[0175] The inventor also found in the research on voice wake-up technology that since the voice stream usually conducts human-computer interaction according to the mode of "wake-up word + user intention", the false wake-up audio data can be collected by means of the wake-up recognition result and the semantic recognition result of the voice stream. For example, Figure 6As shown, the audio data that appears in the pattern of "audio segment triggering a wake-up message + recognized as empty / modal particle / without valid semantics" usually includes mis-wake-up audio segments.
[0176] Based on this, another possible implementation manner of the voice wake-up method is further provided in the embodiments of the present application. In this implementation manner, the voice wake-up method may not only include all or part of the above steps, but may also include S11 - S12:
[0177] S11: After obtaining the current voice segment, determine whether it is in the wake-up state. If so, execute step S12; if not, execute S2 and its subsequent steps.
[0178] The above "wake-up state" is determined according to whether the "wake-up instruction" above is triggered; and this determination process may specifically include: before the wake-up instruction is not triggered, the voice control function in the terminal device is in the non-wake-up state; after the wake-up instruction is triggered, the voice control function in the terminal device is in the wake-up state.
[0179] It can be seen that for a terminal device with the voice control function in the non-wake-up state, after obtaining the current voice segment from the voice stream in real time, it can perform wake-up recognition processing on the current voice segment to determine whether to switch the voice control function in the terminal device from the non-wake-up state to the wake-up state; however, for a terminal device with the voice control function already switched to the wake-up state, after obtaining the current voice segment from the voice stream in real time, it can perform semantic recognition processing on the current voice segment, so that other functional modules in the terminal device can provide corresponding service items to the user according to the semantic recognition result (for example, provide the user with a navigation route of "arriving at XXX community").
[0180] S12: When it is determined that the current voice segment meets the preset information exception condition, determine the training data to be used according to the wake-up state triggering audio segment.
[0181] The above "preset information exception condition" can be set in advance. For example, it may specifically include: the semantic recognition result of the current voice segment meets the preset semantic condition.
[0182] The above "semantic recognition result of the current voice segment" is used to represent the semantic information carried by the voice stream including the current voice segment.
[0183] The above "preset semantic condition" can be preset. For example, the "preset semantic condition" can be any one of empty semantics, modal particles, and invalid semantics. Among them, "empty semantics" means that the semantic recognition result of the above "current speech segment" does not carry any semantic information. The above "modal particle" means that the semantic recognition result of the above "current speech segment" carries a modal particle. The above "invalid semantics" means that the semantic information carried by the semantic recognition result of the above "current speech segment" does not belong to any of the preset candidate function trigger semantics.
[0184] It should be noted that the above "candidate function trigger semantics" is used to trigger a certain function module in the terminal device to provide corresponding service items to the user; and each "candidate function trigger semantics" can be preset.
[0185] In fact, the speech recognition processing for the speech stream is also carried out in real time, and there is a certain distance between the speech segment that can recognize the complete semantics (such as the "fifth speech segment" shown in Figure 1 and the speech segment that can recognize the complete trigger word (such as the "third speech segment" shown in Figure 1 . Based on this, in order to improve the recognition accuracy of false wake-up data, another possible implementation manner of the above "preset information exception condition" is provided in the embodiments of the present application, which can specifically be: the semantic recognition result of the current speech segment meets the preset semantic condition; and the time interval between the acquisition time of the current speech segment and the acquisition time of the speech segment for switching the wake-up state is greater than the preset semantic duration.
[0186] The above "speech segment for switching the wake-up state" refers to the speech segment that determines the wake-up in the speech stream (that is, the speech segment that triggers the wake-up instruction). For example, for the Figure 1 shown speech stream, the above "speech segment for triggering the wake-up state" is the "third speech segment".
[0187] The above "preset semantic duration" is used to represent the speech duration of a complete semantics; and the embodiments of the present application do not limit this "preset duration". For example, it can be preset by relevant personnel. Another example is that it can also be analyzed from a large amount of speech data.
[0188] Based on the relevant content of another possible implementation of the above "preset information exception condition", for a terminal device that has switched to the wake-up state, after obtaining the current voice segment in real time from the voice stream, it can first determine whether the time interval between the current voice segment and the above "switching voice segment of the wake-up state" is greater than the preset duration; if it is greater, it can be determined that the semantic recognition result of the current voice segment should be able to completely represent a semantic information, so it can be determined whether the semantic recognition result of the current voice segment meets the preset semantic condition, so that when it is determined that the preset semantic condition is met, it can be determined that the current voice segment meets the above "preset information exception condition", and thus it can be inferred that the above "wake-up state" is a mis-trigger, so the trigger audio segment of the wake-up state can be determined as mis-awakened audio data.
[0189] The above "trigger audio segment of the wake-up state" carries the trigger word used when switching to this wake-up state; and this "trigger audio segment of the wake-up state" includes the above "switching voice segment of the wake-up state". For example, the "trigger audio segment of the wake-up state" can include Figure 1 the "first voice segment", "second voice segment", and "third voice segment" shown.
[0190] It should be noted that the embodiment of the present application does not limit the determination process of the "training data to be used" in S12, and the determination process of the "training data to be used" in S12 is similar to the determination process of the "training data to be used" in S7 above.
[0191] Based on the relevant content of the above S11 to S12, for a terminal device whose voice control function is in the non-wake-up state, after obtaining the current voice segment in real time from the voice stream, it can perform wake-up recognition processing on the current voice segment to determine whether to switch the voice control function in the terminal device from the non-wake-up state to the wake-up state; however, for a terminal device whose voice control function has been switched to the wake-up state, after obtaining the current voice segment in real time from the voice stream, if it is determined that the current voice segment meets the preset information exception condition, it can be inferred that this wake-up state is a mis-trigger, and thus it can be determined that the trigger audio segment of this wake-up state belongs to mis-awakened audio data, so the training data to be used can be determined according to the trigger audio segment of this wake-up state, so that the training data to be used can include this trigger audio segment, so that the trigger audio segment can be used as a negative example for model update training of the wake-up recognition model in the future, so that the updated wake-up recognition model can better perform wake-up recognition processing on this trigger audio segment, which is beneficial to improving the wake-up recognition performance of the wake-up recognition model and thus beneficial to improving the voice wake-up effect.
[0192] In addition, to further improve the collection effect of mis-awakened audio data, the embodiments of the present application also provide another possible implementation manner of the voice wake-up method. In this implementation manner, the voice wake-up method may not only include all or part of the above steps, but may also include S13:
[0193] S13: When the current wake-up recognition result meets the low-threshold wake-up condition and the cloud wake-up recognition result of the current voice segment does not meet the normal wake-up condition, determine the training data to be used according to the wake-up audio segment.
[0194] It can be seen that for a terminal device in the unawakened state, after obtaining the current voice segment from the voice stream in real time, if it is determined that the current wake-up recognition result meets the low-threshold wake-up condition and it is determined that the cloud wake-up recognition result of the current voice segment does not meet the normal wake-up condition, it can be known that the "wake-up recognition model" mentioned above believes that the wake-up audio segment including the current voice segment may carry a wake-up word, but the "cloud wake-up model" mentioned above believes that the wake-up audio segment including the current voice segment cannot carry a wake-up word. Therefore, it can be inferred that the "wake-up recognition model" gives an incorrect wake-up recognition result for the current voice segment, and thus the wake-up audio segment including the current voice segment can be determined as the training data to be used, so that the training data to be used can include the wake-up audio segment. In this way, it is possible to collect and process audio data with the wake-up type of "local low-threshold wake-up → unable to wake up in the cloud", which is beneficial to improving the comprehensiveness of mis-awakened audio data, thus beneficial to improving the comprehensiveness of the training data to be used, and further beneficial to improving the voice wake-up effect.
[0195] It should be noted that the process of "determining the training data to be used according to the wake-up audio segment" in S13 is similar to the process of determining the "training data to be used" in S7.
[0196] Method Embodiment 6
[0197] To facilitate understanding of the technical solutions provided by the embodiments of the present application, the following will be described in conjunction with two application scenarios.
[0198] Scenario 1: Only use the terminal device to mine difficult-to-wake-up audio data and mis-awakened audio data from the real-time collected voice data.
[0199] For a terminal device with the voice control function in an unawakened state, after obtaining the current voice segment from the voice stream in real time, first use the wake-up recognition model to perform wake-up recognition processing on the current voice segment to obtain the current wake-up recognition result; then determine whether the current wake-up recognition result reaches the first probability threshold. If it reaches the first probability threshold, it can be determined that the current voice segment generates a wake-up, so it can be determined that the wake-up audio segment including the current voice segment carries a wake-up word, so the wake-up instruction can be directly triggered to switch the voice control function in the terminal device from the unawakened state to the awakened state;
[0200] If it does not reach the first probability threshold, it can continue to determine whether the current wake-up recognition result reaches the second probability threshold; if it reaches the second probability threshold, it can be determined that the current voice segment may generate a wake-up, so it can be traced whether there is a target voice segment that meets the low-threshold wake-up condition before the current voice segment. If there is no target voice segment, it can be determined that the current voice segment belongs to the first voice segment that meets the low-threshold wake-up condition, so it can be inferred that the wake-up audio segment including the current voice segment is very likely to generate a wake-up, so it can continue to collect subsequent voice segments from the voice stream in real time to find out whether there is a voice segment that meets the low-threshold wake-up condition for the second time among these subsequent voice segments;
[0201] If there is a target voice segment, it means that the target voice segment and the current voice segment meet the two low-threshold wake-ups, so it can be inferred that the wake-up audio segment including the target voice segment and the current voice segment belongs to difficult-to-wake-up audio data, so the wake-up instruction can be triggered to switch the voice control function in the terminal device from the unawakened state to the awakened state, and the wake-up audio segment and the attribute information of the wake-up audio segment are uploaded to the cloud server so that the cloud server can refer to the wake-up audio segment and the attribute information of the wake-up audio segment to generate training data to be used, so that the wake-up recognition model can be updated and trained using the training data to be used later, so that the updated wake-up recognition model can accurately recognize that the wake-up audio segment carries a valid wake-up word, which is beneficial to improving the voice wake-up effect.
[0202] In addition, for a terminal device in which the voice control function is in a wake-up state, after the current voice segment is obtained in real time from the voice stream, semantic recognition processing can be performed on the current voice segment to obtain the semantic recognition result of the current voice segment. So that when it is determined that the current voice segment meets the "preset information exception condition" above according to the semantic recognition result of the current voice segment, it can be inferred that the terminal device has a false wake-up phenomenon. Therefore, the trigger audio segment in the wake-up state and the attribute information of the trigger audio segment can be uploaded to the cloud server, so that the cloud server can refer to the trigger audio segment and the attribute information of the trigger audio segment to generate training data to be used, so that the training data to be used can be utilized subsequently to update and train the wake-up recognition model, so that the updated wake-up recognition model can accurately identify that the wake-up audio segment carries a valid wake-up word, which is beneficial to improving the voice wake-up effect.
[0203] Scenario 2: Comprehensively utilize the terminal device + cloud server to mine difficult-to-wake-up audio data and false wake-up audio data from the real-time collected voice data.
[0204] For a terminal device in which the voice control function is in an un-wake-up state, after the current voice segment is obtained in real time from the voice stream, the following two processes can be executed respectively:
[0205] Process 1: First, use the local wake-up model (that is, the "wake-up recognition model" above) to perform wake-up recognition processing on the current voice segment to obtain the local wake-up recognition result (that is, the "current wake-up recognition result" above); then judge whether the local wake-up recognition result reaches the first probability threshold. If it reaches the first probability threshold, it can be determined that the current voice segment generates a wake-up, so it can be determined that the wake-up audio segment including the current voice segment carries a wake-up word, so the wake-up instruction can be directly triggered to switch the voice control function in the terminal device from the un-wake-up state to the wake-up state;
[0206] If it does not reach the first probability threshold, it can be continued to judge whether the local wake-up recognition result reaches the second probability threshold; if it reaches the second probability threshold, it can be determined that the current voice segment may generate a wake-up, so it can be traced whether there is a target voice segment that meets the low-threshold wake-up condition before the current voice segment. If there is no target voice segment, it can be determined that the current voice segment belongs to the first voice segment that meets the low-threshold wake-up condition, so it can be inferred that the wake-up audio segment including the current voice segment is very likely to generate a wake-up, so it can continue to collect subsequent voice segments in real time from the voice stream to find out whether there is a voice segment that meets the low-threshold wake-up condition for the second time among these subsequent voice segments;
[0207] If there is a target voice segment, it means that the target voice segment and the current voice segment meet the two low-threshold awakenings. Thus, it can be inferred that the awakening audio segment including the target voice segment and the current voice segment belongs to the difficult-to-awaken audio data. Therefore, an awakening instruction can be triggered to switch the voice control function in the terminal device from the non-awakened state to the awakened state, and upload the awakening audio segment and the attribute information of the awakening audio segment to the cloud server, so that the cloud server can refer to the awakening audio segment and the attribute information of the awakening audio segment to generate the training data to be used, so that the training data to be used can be utilized subsequently to update and train the awakening recognition model, so that the updated awakening recognition model can accurately recognize that the awakening audio segment carries a valid awakening word. This is beneficial to improving the voice awakening effect.
[0208] Process two: Send the current voice segment to the cloud server, so that the cloud server can use the cloud awakening model to perform awakening recognition processing on the current voice segment, obtain the cloud awakening recognition result, and feedback the cloud awakening recognition result to the terminal device. So that after the terminal device receives the cloud awakening recognition result, if it is determined that the above-mentioned "local awakening recognition result" does not reach the first probability threshold, and the cloud awakening recognition result meets the normal awakening condition, it can be inferred that the awakening audio segment including the current voice segment belongs to the difficult-to-awaken audio data. Therefore, the awakening audio segment and the attribute information of the awakening audio segment can be uploaded to the cloud server, so that the cloud server can refer to the awakening audio segment and the attribute information of the awakening audio segment to generate the training data to be used, so that the training data to be used can be utilized subsequently to update and train the awakening recognition model, so that the updated awakening recognition model can accurately recognize that the awakening audio segment carries a valid awakening word. This is beneficial to improving the voice awakening effect;
[0209] If the above-mentioned "local awakening recognition result" is between the second probability threshold and the first probability threshold, and the cloud awakening recognition result does not meet the normal awakening condition, it can be inferred that the terminal device may be wrongly awakened for the current voice segment. Thus, it can be inferred that the awakening audio segment including the current voice segment belongs to the mis-awakening audio data. Therefore, the awakening audio segment and the attribute information of the awakening audio segment can be uploaded to the cloud server, so that the cloud server can refer to the awakening audio segment and the attribute information of the awakening audio segment to generate the training data to be used, so that the training data to be used can be utilized subsequently to update and train the awakening recognition model, so that the updated awakening recognition model can accurately recognize that the awakening audio segment carries a valid awakening word. This is beneficial to improving the voice awakening effect.
[0210] It should be noted that for the above "terminal device + cloud" scenario, if a low-threshold wake-up occurs locally (for example, the low-threshold wake-up condition is met once or twice), but a wake-up occurs in the cloud. At this time, according to the cloud wake-up logic, a wake-up message can be thrown externally (that is, a wake-up instruction is triggered to cause the terminal device to switch from the non-wake-up state to the wake-up state), and at this time, the dual-threshold strategy should clear all state data and start counting again. In addition, if the dual-threshold strategy is adopted locally for wake-up processing and the cloud does not wake up during the first local low-threshold wake-up, when a second wake-up message is generated in a short time, the data of the first low-threshold wake-up should not be uploaded again (because according to the above logic, it has been uploaded once to prevent duplicate upload).
[0211] In addition, for a terminal device with the voice control function in the wake-up state, after the current voice segment is obtained in real time from the voice stream, semantic recognition processing can be performed on the current voice segment to obtain the semantic recognition result of the current voice segment. So that when it is determined that the current voice segment meets the "preset information exception condition" above according to the semantic recognition result of the current voice segment, it can be inferred that the terminal device has had a false wake-up phenomenon. Therefore, the trigger audio segment in the wake-up state and the attribute information of the trigger audio segment can be uploaded to the cloud server, so that the cloud server can refer to the trigger audio segment and the attribute information of the trigger audio segment to generate training data to be used, so that the wake-up recognition model can be updated and trained using the training data to be used later, so that the updated wake-up recognition model can accurately recognize that the wake-up audio segment carries a valid wake-up word, which is beneficial to improving the voice wake-up effect.
[0212] In addition, for the above two scenarios, the cloud server can process the data uploaded by the terminal device using the data processing strategy shown below.
[0213] (1) Data screening
[0214] a. For the wake-up audio data uploaded by the client, it is connected to the cloud voice recognition interface, and the text identifier of the voice recognition text and the voice recognition text are stored together as data annotation information. When storing, a new column of confidence score information can be added in the Figure 4 table shown, and its value is the edit distance between the voice recognition text and the wake-up word annotation.
[0215] b. The data generated by local wake-up, local recognition, and local semantics are uploaded separately and independently. For each piece of uploaded local wake-up data, the voice recognition result or semantic recognition result information whose generation time point is immediately after is queried, and the queried result is stored together with the wake-up data as recognition annotation information.
[0216] (2) Data storage
[0217] Classify and store data according to the calculated confidence score information.
[0218] Based on the relevant content of the above two scenarios, it can be known that the terminal device can adopt a dual-threshold strategy for wake-up processing, which can realize wake-up analysis for voice streams, not easily wake-up audio, and collection and upload of false wake-up data, so that the cloud server can update and train the existing local wake-up model and cloud wake-up model based on a large amount of not easily wake-up audio and a large amount of false wake-up data provided by the terminal device, so that the updated local wake-up model and cloud wake-up model have better wake-up recognition performance. In this way, the purpose of training the model with a large amount of voice data in the real scenario can be realized, which can better accelerate the rhythm of the core effect iteration and evolution, and thus is beneficial to improving the voice wake-up effect.
[0219] Based on the voice wake-up method provided in the above method embodiment, the embodiment of the present application also provides a voice wake-up device, which will be explained and described below with reference to the drawings.
[0220] Device Embodiment
[0221] The device embodiment introduces the voice wake-up device, and for the relevant content, please refer to the above method embodiment.
[0222] See Figure 7 , which is a schematic structural diagram of a voice wake-up device provided by an embodiment of the present application.
[0223] The voice wake-up device 700 provided by the embodiment of the present application includes:
[0224] A voice acquisition unit 701, configured to acquire the current voice segment;
[0225] A wake-up recognition unit 702, configured to perform wake-up recognition processing on the current voice segment to obtain a current wake-up recognition result;
[0226] A first trigger unit 703, configured to trigger a wake-up instruction when the current wake-up recognition result meets the high-threshold wake-up condition;
[0227] A second trigger unit 704, configured to trigger a wake-up instruction when the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition in at least one previous wake-up recognition result.
[0228] In a possible implementation manner, the at least one previous wake-up recognition result includes the previous wake-up recognition result; wherein, the previous wake-up recognition result is obtained by performing wake-up recognition processing on the previous voice segment of the current voice segment; the acquisition time of the previous voice segment is earlier than the acquisition time of the current voice segment;
[0229] The second trigger unit 704 is specifically configured to: trigger a wake-up instruction when the current wake-up recognition result meets the low-threshold wake-up condition and the previous wake-up recognition result meets the low-threshold wake-up condition.
[0230] In a possible implementation, the second trigger unit 704 is specifically configured to: trigger a wake-up instruction when the current wake-up recognition result meets the low-threshold wake-up condition, there is a target wake-up recognition result that meets the low-threshold wake-up condition among the at least one historical wake-up recognition results, and the time difference representation data between the target wake-up recognition result and the current wake-up recognition result meets the first time difference condition.
[0231] In a possible implementation, the at least one historical wake-up recognition result is obtained by performing wake-up recognition processing on at least one historical speech segment;
[0232] The voice wake-up device 700 further includes:
[0233] A message update unit, configured to update the at least one historical speech segment according to the current speech segment when the current wake-up recognition result does not meet the low-threshold wake-up condition, there is a target wake-up recognition result that meets the low-threshold wake-up condition among the at least one historical wake-up recognition results, and the current speech segment does not meet the preset message repetition condition; discard the current speech segment when the current wake-up recognition result indicates that the current speech segment does not meet the low-threshold wake-up condition, there is a target wake-up recognition result that meets the low-threshold wake-up condition among the at least one historical wake-up recognition results, and the current speech segment meets the preset message repetition condition.
[0234] In a possible implementation, the voice wake-up device 700 further includes:
[0235] A repeated recognition unit, configured to determine that the current speech segment meets the preset message repetition condition if the message recognition result of the current speech segment is the same as the message recognition result of the speech segment to be referenced; wherein, the acquisition time difference between the speech segment to be referenced and the current speech segment meets the second time difference condition;
[0236] If the message recognition result of the current speech segment is different from the message recognition result of the speech segment to be referenced, it is determined that the current speech segment does not meet the preset message repetition condition.
[0237] In a possible implementation, the current wake-up recognition result is determined by using a wake-up recognition model;
[0238] The voice wake-up device 700 further includes:
[0239] A first determination unit, configured to determine training data to be used according to a wake-up audio segment when the current wake-up recognition result meets a low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result; wherein, the wake-up audio segment includes the current speech segment and a historical speech segment with the target wake-up recognition result; the training data to be used is used to update the wake-up recognition model.
[0240] In a possible implementation manner, the current wake-up recognition result is determined by using a wake-up recognition model;
[0241] The voice wake-up device 700 further includes:
[0242] A second determination unit, configured to determine training data to be used according to a wake-up audio segment when the current wake-up recognition result meets a high-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result; wherein, the wake-up audio segment includes the current speech segment and a historical speech segment with the target wake-up recognition result; the training data to be used is used to update the wake-up recognition model.
[0243] In a possible implementation manner, the current wake-up recognition result is determined by using a wake-up recognition model;
[0244] The voice wake-up device 700 further includes:
[0245] A third determination unit, configured to determine training data to be used according to the wake-up audio segment when the current wake-up recognition result meets a low-threshold wake-up condition and the cloud wake-up recognition result of the current speech segment meets a normal wake-up condition; wherein, the cloud wake-up recognition result is obtained by performing a wake-up recognition process on the current speech segment by a cloud wake-up model; the wake-up audio segment includes the current speech segment; the training data to be used is used to update the wake-up recognition model and the cloud wake-up model.
[0246] In a possible implementation manner, the current wake-up recognition result is determined by using a wake-up recognition model;
[0247] The voice wake-up device 700 further includes:
[0248] A fourth determination unit, configured to determine training data to be used according to the wake-up audio segment when the current wake-up recognition result meets the non-wake-up condition and the cloud wake-up recognition result of the current speech segment meets the normal wake-up condition; wherein, the cloud wake-up recognition result is obtained by the cloud wake-up model performing wake-up recognition processing on the current speech segment; the wake-up audio segment includes the current speech segment; and the training data to be used is used to update the wake-up recognition model and the cloud wake-up model.
[0249] In a possible implementation manner, the current wake-up recognition result is determined by using a wake-up recognition model;
[0250] The voice wake-up device 700 further includes:
[0251] A fifth determination unit, configured to determine training data to be used according to the wake-up audio segment when the current wake-up recognition result meets the low-threshold wake-up condition and the cloud wake-up recognition result of the current speech segment does not meet the normal wake-up condition; wherein, the cloud wake-up recognition result is obtained by the cloud wake-up model performing wake-up recognition processing on the current speech segment; the wake-up audio segment includes the current speech segment; and the training data to be used is used to update the wake-up recognition model and the cloud wake-up model.
[0252] In a possible implementation manner, the wake-up recognition unit 702 is specifically configured to: if in an unwake state, use the wake-up recognition model to perform wake-up recognition processing on the current speech segment to obtain a current wake-up recognition result;
[0253] The voice wake-up device 700 further includes:
[0254] A mis-trigger recognition unit, configured to, if in a wake state, when determining that the current speech segment meets a preset information anomaly condition, determine training data to be used according to the trigger audio segment of the wake state; wherein, the training data to be used is used to update the wake-up recognition model.
[0255] In a possible implementation manner, the method is applied to a terminal device; the current wake-up recognition result is determined by using a wake-up recognition model;
[0256] The voice wake-up device 700 further includes:
[0257] A data sending unit, configured to send a wake-up audio segment and attribute information of the wake-up audio segment to a cloud server when the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition in at least one historical wake-up recognition result, so that the cloud server determines training data to be used according to the wake-up audio segment and the attribute information of the wake-up audio segment; wherein, the wake-up audio segment includes the current speech segment; the training data to be used is used to update the wake-up recognition model.
[0258] Further, an embodiment of the present application further provides a device, including: a processor, a memory, and a system bus;
[0259] The processor and the memory are connected through the system bus;
[0260] The memory is used to store one or more programs, and the one or more programs include instructions, and when the instructions are executed by the processor, the processor executes any implementation method of the above-mentioned voice wake-up method.
[0261] Further, an embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a terminal device, the terminal device is enabled to execute any implementation method of the above-mentioned voice wake-up method.
[0262] Further, an embodiment of the present application further provides a computer program product, and when the computer program product runs on a terminal device, the terminal device is enabled to execute any implementation method of the above-mentioned voice wake-up method.
[0263] Through the description of the above embodiments, those skilled in the art can clearly understand that all or part of the steps in the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for enabling a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.
[0264] It should be noted that the various embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same and similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0265] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0266] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A voice wake-up method, characterized in that, The method includes: Obtain the current speech segment; Perform wake-up recognition processing on the current speech segment to obtain a current wake-up recognition result; When the current wake-up recognition result meets the high-threshold wake-up condition, trigger a wake-up instruction; When the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, trigger a wake-up instruction; The at least one historical wake-up recognition result includes the previous wake-up recognition result; wherein, the previous wake-up recognition result is obtained by performing wake-up recognition processing on the previous speech segment of the current speech segment; the acquisition time of the previous speech segment is earlier than the acquisition time of the current speech segment; The step of triggering a wake-up instruction when the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result includes: When the current wake-up recognition result meets the low-threshold wake-up condition and the previous wake-up recognition result meets the low-threshold wake-up condition, trigger a wake-up instruction.
2. The method according to claim 1, wherein The step of triggering a wake-up instruction when the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result includes: When the current wake-up recognition result meets the low-threshold wake-up condition, there is a target wake-up recognition result that meets the low-threshold wake-up condition among the at least one historical wake-up recognition result, and the time difference characterization data between the target wake-up recognition result and the current wake-up recognition result meets the first time difference condition, trigger a wake-up instruction.
3. The method according to claim 1, wherein The at least one historical wake-up recognition result is obtained by performing wake-up recognition processing on at least one historical speech segment; The method further includes: When the current wake-up recognition result does not meet the low-threshold wake-up condition, there is a target wake-up recognition result that meets the low-threshold wake-up condition among the at least one historical wake-up recognition result, and the current speech segment does not meet the preset message repetition condition, update the at least one historical speech segment according to the current speech segment; When the current wake-up recognition result indicates that the current speech segment does not meet the low-threshold wake-up condition, there is a target wake-up recognition result that meets the low-threshold wake-up condition among the at least one historical wake-up recognition result, and the current speech segment meets the preset message repetition condition, discard the current speech segment.
4. The method according to claim 3, wherein The method further includes: If the message recognition result of the current speech segment is the same as the message recognition result of the speech segment to be referenced, it is determined that the current speech segment meets the preset message repetition condition; wherein, the acquisition time difference between the speech segment to be referenced and the current speech segment meets the second time difference condition; If the message recognition result of the current speech segment is different from the message recognition result of the speech segment to be referenced, it is determined that the current speech segment does not meet the preset message repetition condition.
5. The method according to claim 1, wherein The current wake-up recognition result is determined using a wake-up recognition model; The method further includes: When the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, determine the training data to be used according to the wake-up audio segment; wherein, the wake-up audio segment includes the current speech segment and the historical speech segment with the target wake-up recognition result; the training data to be used is used to update the wake-up recognition model.
6. The method according to claim 1, wherein The current wake-up recognition result is determined by using the wake-up recognition model; The method further includes: When the current wake-up recognition result meets the high-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, determine the training data to be used according to the wake-up audio segment; wherein, the wake-up audio segment includes the current speech segment and the historical speech segment with the target wake-up recognition result; the training data to be used is used to update the wake-up recognition model.
7. The method according to claim 1, characterized in that, The current wake-up recognition result is determined by using the wake-up recognition model; The method further includes: When the current wake-up recognition result meets the low-threshold wake-up condition and the cloud wake-up recognition result of the current speech segment meets the normal wake-up condition, determine the training data to be used according to the wake-up audio segment; wherein, the cloud wake-up recognition result is obtained by the cloud wake-up model performing wake-up recognition processing on the current speech segment; the wake-up audio segment includes the current speech segment; the training data to be used is used to update the wake-up recognition model and the cloud wake-up model.
8. The method according to claim 1, wherein The current wake-up recognition result is determined by using the wake-up recognition model; The method further includes: When the current wake-up recognition result meets the non-wake-up condition and the cloud wake-up recognition result of the current speech segment meets the normal wake-up condition, determine the training data to be used according to the wake-up audio segment; wherein, the cloud wake-up recognition result is obtained by the cloud wake-up model performing wake-up recognition processing on the current speech segment; the wake-up audio segment includes the current speech segment; the training data to be used is used to update the wake-up recognition model and the cloud wake-up model.
9. The method according to claim 1, wherein The current wake-up recognition result is determined by using the wake-up recognition model; The method further includes: When the current wake-up recognition result meets the low-threshold wake-up condition and the cloud wake-up recognition result of the current speech segment does not meet the normal wake-up condition, determine the training data to be used according to the wake-up audio segment; wherein, the cloud wake-up recognition result is obtained by the cloud wake-up model performing wake-up recognition processing on the current speech segment; the wake-up audio segment includes the current speech segment; the training data to be used is used to update the wake-up recognition model and the cloud wake-up model.
10. The method according to claim 1, characterized in that, Performing wake-up recognition processing on the current speech segment to obtain the current wake-up recognition result includes: If in the non-wake state, use the wake-up recognition model to perform wake-up recognition processing on the current speech segment to obtain the current wake-up recognition result; The method further includes: If in the wake-up state, when it is determined that the semantic recognition result of the current speech segment meets the preset semantic condition, determine the training data to be used according to the trigger audio segment of the wake-up state; wherein, the training data to be used is used to update the wake-up recognition model.
11. The method according to claim 1, wherein The method is applied to a terminal device; the current wake-up recognition result is determined by using a wake-up recognition model; The method further includes: When the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result, send the wake-up audio segment and the attribute information of the wake-up audio segment to the cloud server, so that the cloud server determines the training data to be used according to the wake-up audio segment and the attribute information of the wake-up audio segment; wherein, the wake-up audio segment includes the current speech segment; the training data to be used is used to update the wake-up recognition model.
12. A voice wake-up device, characterized in that, Includes: A voice acquisition unit for acquiring a current speech segment; A wake-up recognition unit for performing wake-up recognition processing on the current speech segment to obtain a current wake-up recognition result; A first trigger unit for triggering a wake-up instruction when the current wake-up recognition result meets the high-threshold wake-up condition; A second trigger unit for triggering a wake-up instruction when the current wake-up recognition result meets the low-threshold wake-up condition and there is a target wake-up recognition result that meets the low-threshold wake-up condition among at least one historical wake-up recognition result; The at least one historical wake-up recognition result includes the previous wake-up recognition result; wherein, the previous wake-up recognition result is obtained by performing wake-up recognition processing on the previous speech segment of the current speech segment; the acquisition time of the previous speech segment is earlier than the acquisition time of the current speech segment; The second trigger unit is specifically configured to: trigger a wake-up instruction when the current wake-up recognition result meets the low-threshold wake-up condition and the previous wake-up recognition result meets the low-threshold wake-up condition.
13. A device, characterized in that, The device includes: a processor, a memory, and a system bus; The processor and the memory are connected through the system bus; The memory is used to store one or more programs, and the one or more programs include instructions, and when the instructions are executed by the processor, the processor executes the method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when the instructions run on the terminal device, the terminal device executes the method according to any one of claims 1 to 11.
15. A computer program product, characterized in that, When the computer program product runs on the terminal device, the terminal device executes the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Terminal equipment standby wake-up method and device, and computer equipment
CN107610702A
Wake-up method and device for voice equipment
CN111897584A