Method and apparatus for predicting a quiet time window, storage medium, and electronic device

CN119580766BActive Publication Date: 2025-11-04QINGDAO HAIER TECH +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411628864.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-11-04
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

[0004]本申请实施例提供了一种静默时间窗的预测方法和装置、存储介质及电子装置,以至少解决相关技术中,通过选择较长的静默时间窗适应不同用户群体在说话时的不同的断句习惯,导致的增加了整体语音交互时长的问题

Benefits of technology

[0017] In the embodiments of the present application, the historical voice interaction data of a target object with a voice device is obtained, a first time difference between the voice interaction time corresponding to the historical voice interaction data and a current time is determined, and a first punctuation time interval corresponding to the historical voice interaction data is determined, a target weight corresponding to the historical voice interaction data is determined according to the first time difference, and a target silence time window of the target object interacting with the voice device is predicted according to the target weight and the first punctuation interval time. That is, the embodiments of the present application determine the target weight according to the first time difference between the voice interaction time corresponding to the historical voice interaction data and the current time, and then predict the target silence time window according to the target weight and the first punctuation interval time. According to the embodiments of the present application, the problem of increasing the overall voice interaction time caused by selecting a longer silence time window to adapt to different punctuation habits of different user groups when speaking in the related art can be solved. Different silence time windows are determined to adapt to different punctuation habits of different user groups, and the overall voice interaction time is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580766B_ABST
    Figure CN119580766B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for predicting a silence time window, a storage medium and an electronic device, and relates to the technical field of smart homes. The method comprises the following steps: obtaining historical voice interaction data of a target object, wherein the historical voice interaction data is used to indicate interactive voice of the target object and a voice device; determining a first time difference between a voice interaction time corresponding to the historical voice interaction data and a current time, and determining a first sentence breaking interval time corresponding to the historical voice interaction data; determining a target weight corresponding to the historical voice interaction data according to the first time difference; and predicting a target silence time window of the target object and the voice device according to the target weight and the first sentence breaking interval time. Through the above method, the problem of increasing the overall voice interaction time caused by selecting a longer silence time window to adapt to different sentence breaking habits of different user groups when speaking can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of smart home, in particular to a silence time window prediction method and device, a storage medium and an electronic device. BACKGROUND

[0002] In the current trend of social life, taking family as the core, voice interaction control based on artificial intelligence (AI) voice technology has become the mainstream demand leading the intelligent home. People aspire to enjoy a more intelligent life, easily master various intelligent devices and systems through simple voice commands, answer questions and obtain operation guidance. Such high recognition rate and rapid voice interaction response not only enhances the seamless communication between users and smart home, but also brings people an unprecedented convenient and intelligent home experience. The voice interaction function in smart home has become an indispensable part of people's daily life. Smart home greatly simplifies the control of intelligent devices and systems, and also provides convenient question and answer and operation guidance services. However, the speaking speed and sentence breaking habits of different users are different, especially those users with slower speaking speed and longer sentence breaking interval. The traditional voice activity detection (VAD) often needs to set a longer silence time window to ensure the integrity of voice input when processing their voice data. However, this uniform processing method undoubtedly increases the voice interaction response time of all users and reduces the voice interaction experience of users.

[0003] For the problem of increasing the overall voice interaction time caused by selecting a longer silence time window to adapt to the different sentence breaking habits of different user groups when speaking in the related art, an effective solution has not been proposed. SUMMARY

[0004] Embodiments of the present application provide a silence time window prediction method and device, a storage medium and an electronic device to at least solve the problem of increasing the overall voice interaction time caused by selecting a longer silence time window to adapt to the different sentence breaking habits of different user groups when speaking in the related art.

[0005] According to one of the embodiments of the present application, a method for predicting a silence time window is provided, including: obtaining historical voice interaction data of a target object, wherein the historical voice interaction data is used to indicate an interactive voice of the target object with a voice device; determining a first time difference between a voice interaction time corresponding to the historical voice interaction data and a current time, and determining a first pause interval time corresponding to the historical voice interaction data; determining a target weight corresponding to the historical voice interaction data according to the first time difference; and predicting a target silence time window of the target object interacting with the voice device according to the target weight and the first pause interval time.

[0006] In one example embodiment, determining a target weight corresponding to the historical voice interaction data according to the first time difference includes: obtaining a plurality of voice interaction data corresponding to the target object, and determining a second time difference between a voice interaction time corresponding to each voice interaction data and the current time; dividing the plurality of voice interaction data and the second time difference corresponding to each voice interaction data into training set data and test set data, and obtaining an initial decay coefficient; training a target model according to the training set data and the initial decay coefficient to obtain a trained target model, wherein the target model is used to determine a weight corresponding to each voice interaction data; verifying the trained target model according to the test set data; and inputting the first time difference into the trained target model to make the trained target model output the target weight in a case where it is determined that the trained target model passes the verification.

[0007] In one example embodiment, determining a first pause interval time corresponding to the historical voice interaction data includes: determining whether a first voice interruption event exists in the historical voice interaction data, wherein the first voice interruption event includes a voice disappearance event or a voiceless event; determining a first time at which the first voice interruption event starts, and determining whether a voice appearance event exists after the first voice interruption event in a case where it is determined that the first voice interruption event exists in the historical voice interaction data; determining a second time at which the voice appearance event starts, and determining the first pause interval time according to a difference between the second time and the first time in a case where it is determined that the voice appearance event exists after the first voice interruption event; determining a third time at which the historical voice interaction data ends, and determining the first pause interval time according to a difference between the third time and the first time in a case where it is determined that the voice appearance event does not exist after the first voice interruption event.

[0008] In one example embodiment, after determining whether the first speech interruption event exists in the historical speech interaction data, the method further comprises: in a case where it is determined that the first speech interruption event exists in the historical speech interaction data, determining whether the speech occurrence event exists in a target time period after the first time; in a case where it is determined that the speech occurrence event exists in the target time period, determining that the speech interruption event is the speech disappearance event; in a case where it is determined that the speech occurrence event does not exist in the target time period, determining that the speech interruption event is the voiceless event.

[0009] In one example embodiment, after predicting the target silence time window of the target object interacting with the speech device according to the target weight and the first punctuation interval time, the method further comprises: in a case where it is determined that the speech device receives first interaction speech uttered by the target object, determining whether a second speech interruption event exists in the first interaction speech; in a case where it is determined that the second speech interruption event exists in the first interaction speech, determining whether the second speech interruption event is a voiceless event; in a case where it is determined that the second speech interruption event is the voiceless event, determining a sum value of the target silence time window and a preset extension time, and determining a first size relationship between a second punctuation interval time corresponding to the first interaction speech and the sum value; in a case where the first size relationship indicates that the second punctuation interval time is greater than or equal to the sum value, ending the sound pickup of the first interaction speech, and determining a first weight corresponding to the first interaction speech after the sound pickup is ended; predicting a first silence time window of the target object interacting with the speech device according to the first weight and the second punctuation interval time, and updating the target silence time window according to the first silence time window.

[0010] In an example embodiment, after determining the first size relationship between the second pause interval time corresponding to the first interactive speech and the sum value, the method further comprises: in a case where the first size relationship indicates that the second pause interval time is less than the sum value, performing endpoint detection on the first interactive speech to obtain a speech occurrence event after the second pause interval; determining whether the unvoiced event exists in a second interactive speech corresponding to the speech occurrence event; in a case where it is determined that the unvoiced event exists in the second interactive speech, determining a second size relationship between a third pause interval time corresponding to the second interactive speech and the sum value; in a case where the second size relationship indicates that the third pause interval time is greater than or equal to the sum value, ending the sound pickup of the second interactive speech, and merging the first interactive speech and the second interactive speech to update the first interactive speech; determining a second weight corresponding to the updated first interactive speech; predicting a second silence time window of the target object interacting with the speech device according to the second weight, the second pause interval time and the third pause interval time, and updating the target silence time window according to the second silence time window.

[0011] In an example embodiment, predicting a second silence time window of the target object interacting with the speech device according to the second weight, the second pause interval time and the third pause interval time comprises: determining a third size relationship between the second pause interval time and the third pause interval time; determining a maximum pause interval time from the second pause interval time and the third pause interval time according to the third size relationship; and determining the second silence time window according to the second weight and the maximum pause interval time.

[0012] In an example embodiment, after predicting the target silence time window of the target object interacting with the speech device according to the target weight and the first pause interval time, the method further comprises: in a case where the speech device receives a third interactive speech, determining whether a sound production object corresponding to the third interactive speech is the target object; in a case where it is determined that the sound production object is not the target object, determining whether a third silence time window of the sound production object interacting with the speech device exists in a silence time window library, wherein the silence time window library comprises: a target silence time window corresponding to the target object; in a case where it is determined that the third silence time window does not exist in the silence time window library, determining a third weight corresponding to the third interactive speech and a fourth pause interval time corresponding to the third interactive speech; predicting the third silence time window according to the third weight and the fourth pause interval time, and storing the third silence time window corresponding to the sound production object into the silence time window library.

[0013] According to another embodiment of the present application, a device for predicting a silence time window is also provided, comprising: an obtaining module configured to obtain historical voice interaction data of a target object, wherein the historical voice interaction data is used to indicate an interactive voice of the target object with a voice device; a first determining module configured to determine a first time difference between a voice interaction time corresponding to the historical voice interaction data and a current time, and determine a first punctuation interval time corresponding to the historical voice interaction data; a second determining module configured to determine a target weight corresponding to the historical voice interaction data according to the first time difference; and a predicting module configured to predict a target silence time window of the target object interacting with the voice device according to the target weight and the first punctuation interval time.

[0014] According to yet another aspect of the present application, a computer readable storage medium having a computer program stored therein is also provided, wherein the computer program is configured to execute the above-mentioned method for predicting a silence time window when running.

[0015] According to yet another aspect of the present application, an electronic device is also provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-mentioned method for predicting a silence time window through the computer program.

[0016] According to yet another aspect of the present application, a computer program product is also provided, comprising a computer program, wherein the computer program executes the above-mentioned method for predicting a silence time window through a processor.

[0017] In the embodiments of the present application, the historical voice interaction data of a target object with a voice device is obtained, a first time difference between the voice interaction time corresponding to the historical voice interaction data and a current time is determined, and a first punctuation time interval corresponding to the historical voice interaction data is determined, a target weight corresponding to the historical voice interaction data is determined according to the first time difference, and a target silence time window of the target object interacting with the voice device is predicted according to the target weight and the first punctuation interval time. That is, the embodiments of the present application determine the target weight according to the first time difference between the voice interaction time corresponding to the historical voice interaction data and the current time, and then predict the target silence time window according to the target weight and the first punctuation interval time. According to the embodiments of the present application, the problem of increasing the overall voice interaction time caused by selecting a longer silence time window to adapt to different punctuation habits of different user groups when speaking in the related art can be solved. Different silence time windows are determined to adapt to different punctuation habits of different user groups, and the overall voice interaction time is reduced. BRIEF DESCRIPTION OF DRAWINGS

[0018] The accompanying drawings, which are incorporated into and constitute a part of this specification, illustrate embodiments consistent with the present application and serve to explain the principles of the present application together with the specification.

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings required by the embodiments or prior art description will be briefly introduced as follows. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative labor.

[0020] Figure 1 is a hardware environment schematic diagram of a silence time window prediction method according to an embodiment of the present application;

[0021] Figure 2 is a flowchart of a silence time window prediction method according to an embodiment of the present application;

[0022] Figure 3 is a whole framework diagram of user voice interaction with a large model according to an optional embodiment of the present application;

[0023] Figure 4 is a voice interaction process schematic diagram according to an embodiment of the present application;

[0024] Figure 5 is a device schematic diagram corresponding to a method for improving voice recognition success rate and shortening voice interaction time according to an optional embodiment of the present application;

[0025] Figure 6 is a flowchart of a method for improving voice recognition success rate and shortening voice interaction time according to an optional embodiment of the present application;

[0026] Figure 7 is a structural block diagram of a silence time window prediction device according to an embodiment of the present application. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and in the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0029] According to an aspect of an embodiment of the present application, a prediction method of a quiet time window is provided. The prediction method of the quiet time window is widely applied to smart home, smart home, smart home device ecology, intelligence house ecology and other whole-house intelligent digital control application scenarios. Optionally, Figure 1 is a hardware environment schematic diagram of a prediction method of a quiet time window according to an embodiment of the present application. In the present embodiment, the prediction method of the quiet time window can be applied to a module in a household appliance. The module can be applied in a hardware environment composed of a household appliance 102 and a server 104 as shown in Figure 1 . As shown in Figure 1 , the server 104 is connected with the household appliance 102 through a network, which can be used to provide services (such as application services, etc.) for nodes or clients installed on nodes, a database can be set on the server or independently of the server, which can be used to provide data storage services for the server 104, cloud computing and / or edge computing services can be configured on the server or independently of the server, which can be used to provide data operation services for the server 104.

[0030] The network can include, but is not limited to, at least one of the following: a wired network, a wireless network. The wired network can include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, a local area network. The wireless network can include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The home appliance 102 can include, but is not limited to, at least one of the following: a smart air conditioner, a smart oven, a smart refrigerator, a smart oven, a smart oven, a smart washing machine, a smart water heater, a smart washing device, a smart dishwasher, a smart projection device, a smart television, a smart clothesline, a smart curtain, a smart audio and video, a smart socket, a smart sound, a smart sound box, a smart fresh air device, a smart kitchen and bathroom device, a smart bathroom device, a smart sweeping robot, a smart window cleaning robot, a smart mopping robot, a smart air purification device, a smart steamer, a smart microwave oven, a smart kitchen treasure, a smart purifier, a smart water dispenser, a smart door lock, etc.

[0031] In the present embodiment, a prediction method of a silence time window is provided, which is applied to the above module, Figure 2 is a flowchart of the prediction method of the silence time window according to the embodiments of the present application, which includes the following steps:

[0032] In step S202, historical voice interaction data of a target object is obtained, wherein the historical voice interaction data is used to indicate the interactive voice of the target object and the voice device;

[0033] In step S204, a first time difference between the voice interaction time corresponding to the historical voice interaction data and the current time is determined, and a first sentence breaking interval time corresponding to the historical voice interaction data is determined.

[0034] In step S206, a target weight corresponding to the historical voice interaction data is determined according to the first time difference;

[0035] In step S208, a target silence time window of the target object interacting with the voice device is predicted according to the target weight and the first sentence breaking interval time.

[0036] By the above steps, the historical voice interaction data of the target object and the voice device is obtained, the historical voice interaction data and a first time difference between the corresponding voice interaction time and the current time are determined, and a first sentence breaking time interval corresponding to the historical voice interaction data is determined, and a target weight corresponding to the historical voice interaction data is determined according to the first time difference; and a target silence time window of the target object and the voice device is predicted according to the target weight and the first sentence breaking interval time. That is, the embodiment of the present application determines the target weight according to the first time difference between the voice interaction time corresponding to the historical voice interaction data and the current time, and then predicts the target silence time window according to the target weight and the first sentence breaking interval time. According to the embodiment of the present application, the problem of increasing the overall voice interaction time caused by selecting a longer silence time window to adapt to different sentence breaking habits of different user groups when speaking in the related art can be solved. Then, different silence time windows are determined to adapt to different sentence breaking habits of different user groups, and the overall voice interaction time is reduced.

[0037] Optionally, the step S206 of determining the target weight corresponding to the historical voice interaction data according to the first time difference comprises: obtaining a plurality of voice interaction data corresponding to the target object, and determining a second time difference between the voice interaction time corresponding to each voice interaction data and the current time; dividing the plurality of voice interaction data and the second time difference corresponding to each voice interaction data into training set data and test set data, and obtaining an initial decay coefficient; training a target model according to the training set data and the initial decay coefficient to obtain a trained target model, wherein the target model is used to determine the weight corresponding to each voice interaction data; verifying the trained target model according to the test set data; and in a case where it is determined that the trained target model passes the verification, inputting the first time difference into the trained target model, so that the trained target model outputs the target weight.

[0038] It can be understood that the specific steps of determining the target weight include:

[0039] A large amount of historical voice interaction data of a target object (which can be a family member) is collected, including the time points of voice start and end. For each piece of historical voice interaction data, a second time difference between its voice interaction time and the current time is calculated, reflecting the degree of newness or oldness of the voice interaction data. The collected multiple voice interaction data and their corresponding second time differences are divided into training set data and test set data. The training set data is used for model training, while the test set data is used to verify the accuracy of the model. A preliminary decay coefficient is selected as the starting point for model training. The decay coefficient determines the degree of influence of historical data on current prediction. The target model (i.e., the model used to determine the weight) is trained using the training set data and the initial decay coefficient. Machine learning algorithms such as weighted moving average algorithm can be used to adjust model parameters and optimize weight distribution. The target model learns how to assign different weights according to the degree of newness or oldness of the data (i.e., the second time difference), so as to more accurately predict the target silence time window.

[0040] After the target model is trained, the test set data is used to verify the performance of the trained target model and evaluate the prediction accuracy of the trained target model. If the model performs well on the test set data, it means that the trained target model can effectively assign weights according to the degree of newness or oldness of the data, thus accurately predicting the silence time window.

[0041] In the case where the model verification is passed, the first time difference of the current voice interaction (i.e., the time difference from the user stopping speaking to the system starting timing) is input into the trained target model. The target model determines the target weight of the current voice interaction data according to the first time difference.

[0042] Through the above steps, the length of the silence time window can be dynamically adjusted according to the user's (i.e., the target object's) recent interaction pattern (e.g., the first time difference). This ensures that each user's voice interaction at different time points can be optimally processed, thereby improving the success rate of voice recognition, shortening the length of voice interaction, and improving user experience.

[0043] Optionally, the determining of the first pause interval time corresponding to the historical voice interaction data in the step S204 comprises: determining whether there is a first voice interruption event in the historical voice interaction data, wherein the first voice interruption event comprises a voice disappearance event or a voiceless event; in a case where it is determined that there is the first voice interruption event in the historical voice interaction data, determining a first time at which the first voice interruption event starts, and determining whether there is a voice appearance event after the first voice interruption event; in a case where it is determined that there is the voice appearance event after the first voice interruption event, determining a second time at which the voice appearance event starts, and determining the first pause interval time according to a difference between the second time and the first time; in a case where it is determined that there is no voice appearance event after the first voice interruption event, determining a third time at which the historical voice interaction data ends, and determining the first pause interval time according to a difference between the third time and the first time.

[0044] It can be understood that the pause interval time is the time of voice silence corresponding to the target.

[0045] Determining the first time of the first voice interruption event of the target object: Once the first voice interruption event is identified, the system records the exact time point at which this event starts.

[0046] Detecting the voice appearance event: Check whether there is a voice appearance event (i.e. the user starts to speak again) after the first voice interruption event. If the voice appearance event is detected, determine the second time at which the voice appearance event starts.

[0047] Calculating the first pause interval time (with voice appearance event): If the voice appearance event is detected after the first voice interruption event, the difference between the second time and the first time is calculated. This difference represents the duration of a voice pause of the target object, which is an important basis for determining the silence time window.

[0048] Calculating the first pause interval time (without voice appearance event): If no voice appearance event is detected after the first voice interruption event (i.e. the user does not speak again, and the voice interaction ends), the time point at which the historical voice interaction data ends, i.e. the third time, is recorded. Then, the difference between the third time and the first time, i.e. the first pause interval time, is calculated to determine the first pause interval time. In this case, the first pause interval time represents the total duration from the time when the user stops speaking to the end of the voice interaction.

[0049] By the above steps, the punctuation habits of each user in historical voice interactions are determined, capturing the possible punctuation interval times that users may take in different situations. These data are crucial for optimizing the silence time window of VAD, as it allows the system to dynamically adjust the waiting time according to individual differences of users, ensuring the complete collection of voice commands and avoiding unnecessary waiting, thus improving the accuracy and efficiency of speech recognition, shortening the total length of voice interaction, and ultimately enhancing the user's smart home experience. This intelligent optimization method embodies the strong potential of AI technology in personalized services and real-time responses.

[0050] In a possible implementation, after determining whether the first voice interruption event exists in the historical voice interaction data, the method further comprises: in a case where it is determined that the first voice interruption event exists in the historical voice interaction data, determining whether the voice appearance event exists in a target time period after the first time; in a case where it is determined that the voice appearance event exists in the target time period, determining that the voice interruption event is the voice disappearance event; and in a case where it is determined that the voice appearance event does not exist in the target time period, determining that the voice interruption event is the voiceless event.

[0051] It can be understood that, in a case where the first voice interruption event exists, it is necessary to determine whether the first voice interruption event is a voiceless event or a voice disappearance event. Specifically, in a case where the voice appearance event exists in a first time period after the start time of the voice interruption event, it is determined that the first voice interruption event is the voice disappearance event; and in a case where the voice appearance event does not exist in the first time period after the start time of the voice interruption event, it is determined that the first voice interruption event is the voiceless event.

[0052] Optionally, after predicting the target silence time window of the target object interacting with the voice device according to the target weight and the first pause interval time, the method further comprises: in the case that it is determined that the voice device receives the first interaction voice uttered by the target object, determining whether a second speech interruption event exists in the first interaction voice; in the case that it is determined that the second speech interruption event exists in the first interaction voice, determining whether the second speech interruption event is a silent event; in the case that it is determined that the second speech interruption event is the silent event, determining a sum value of the target silence time window and a preset extension time, and determining a first size relationship between a second pause interval time corresponding to the first interaction voice and the sum value; in the case that the first size relationship indicates that the second pause interval time is greater than or equal to the sum value, ending the sound pickup of the first interaction voice, and determining a first weight corresponding to the first interaction voice after the sound pickup is ended; predicting a first silence time window of the target object interacting with the voice device according to the first weight and the second pause interval time, and updating the target silence time window according to the first silence time window.

[0053] It can be understood that, in the case that the second speech interruption event is determined to be a silent event, the method of determining the target silence time window comprises:

[0054] The sum of the target silence time window and a preset extension time (i.e., the sum value) is calculated, and a size relationship between the current second pause interval time (i.e., the time from the user stopping speaking to starting speaking again) and the sum value is determined. If the second pause interval time is greater than or equal to the sum value, it indicates that the user's pause time is relatively long, and the voice interaction may have ended.

[0055] If the second pause interval time is indeed greater than or equal to the preset sum value, the system will end the sound pickup of the first interaction voice, i.e., stop recording. Subsequently, the system determines a first weight corresponding to the first interaction voice data obtained after the sound pickup is ended, and the weight reflects the importance of the current voice data in predicting the silence time window.

[0056] Based on the first weight and the second pause interval time, the optimized model (which may be a weighted moving average algorithm or other machine learning models) is used again to predict a first silence time window of the user interacting with the voice device. This prediction value more accurately reflects the user's current speaking and pausing habits. Finally, according to the predicted first silence time window, the system updates the target silence time window in real time to adapt to the user's immediate changes, ensuring accurate capture and efficient processing of the next voice interaction.

[0057] Through this series of dynamic evaluation and real-time adjustment, the embodiments of the present application can accurately capture the personalized speaking habits of different users, optimize the voice interaction performance of smart home voice devices, ensure the integrity and accuracy of voice instructions, significantly shorten the interaction time, and improve the user's voice interaction experience. This real-time optimization mechanism reflects the flexibility and intelligence of AI technology in the smart voice interaction scenario, which can continuously adapt to the personalized needs of users and provide a more smooth and natural interaction experience.

[0058] In a case where the first size relationship indicates that the second pause interval time is less than the sum value, endpoint detection is performed on the first interactive voice to obtain a voice occurrence event after the second pause interval; it is determined whether the voice occurrence event corresponds to a second interactive voice that has the unvoiced event; in a case where it is determined that the second interactive voice has the unvoiced event, a second size relationship between a third pause interval time corresponding to the second interactive voice and the sum value is determined; in a case where the second size relationship indicates that the third pause interval time is greater than or equal to the sum value, the sound pickup of the second interactive voice is ended, and the first interactive voice and the second interactive voice are merged to update the first interactive voice; a second weight corresponding to the updated first interactive voice is determined; a second silence time window of the target object interacting with the voice device is predicted according to the second weight, the second pause interval time and the third pause interval time, and the target silence time window is updated according to the second silence time window.

[0059] It can be understood that, in a case where the second pause interval time is less than the preset sum value, it means that the user may continue to speak soon after a short pause. At this time, the system performs endpoint detection on the first interactive voice to identify a voice occurrence event, that is, the moment when the user starts to speak again. Specifically:

[0060] Determination of the unvoiced event in the second interactive voice: once the voice occurrence event is detected, the system records the new voice data, that is, the second interactive voice, and checks whether there is also an unvoiced event in the second interactive voice. This continuous unvoiced event detection helps the system to better understand the speaking habits of the user.

[0061] Second size relationship between the third pause interval time and the sum value: if the unvoiced event is encountered again in the second interactive voice, the system calculates the third pause interval time from the unvoiced event to the next voice occurrence event, and compares it with the sum value. This comparison helps to determine whether the user has ended the current voice instruction or there may be a longer pause.

[0062] End of sound pickup and voice data merging: if the third pause interval time is greater than or equal to the sum, the system will end the sound pickup of the second interactive voice, and merge the first interactive voice and the second interactive voice to form a complete and updated voice input. This operation ensures that even if the user has a short pause while speaking, the system can capture his voice command completely.

[0063] Determining a second weight: based on the merged voice data, the system will determine a second weight, which reflects the importance of the current user's voice input in predicting the silence time window.

[0064] The method for predicting the second silence time window of the target object interacting with the voice device according to the second weight, the second pause interval time and the third pause interval time can be to determine the third size relationship of the second pause interval time and the third pause interval time; to determine the maximum pause interval time between the second pause interval time and the third pause interval time according to the third size relationship; to determine the second silence time window according to the second weight and the second pause interval time and the maximum pause interval time.

[0065] Through this prediction, the setting of the silence time window can be dynamically adjusted and optimized to adapt to the speaking mode and rhythm of the user, ensuring that the voice command of the user who speaks quickly or slowly can be captured timely and completely.

[0066] Optionally, after the step S208 of predicting the target silence time window of the target object interacting with the voice device according to the target weight and the first pause interval time, the method further comprises: in the case of determining that the voice device receives a third interactive voice, determining whether the voice object corresponding to the third interactive voice is the target object; in the case of determining that the voice object is not the target object, determining whether there is a third silence time window of the voice object interacting with the voice device in the silence time window library, wherein the silence time window library comprises: the target silence time window corresponding to the target object; in the case of determining that there is no third silence time window in the silence time window library, determining a third weight corresponding to the third interactive voice and a fourth pause interval time corresponding to the third interactive voice; predicting the third silence time window according to the third weight and the fourth pause interval time, and storing the third silence time window corresponding to the voice object into the silence time window library.

[0067] It can be understood that the process after the step S208 focuses on processing voice interaction in a multi-person environment, ensuring that the system can accurately distinguish different users and provide the most optimized silence time window setting for each user. Specifically:

[0068] When the voice device receives the third interactive voice, the system's first task is to identify the speaker of this voice and determine whether it is the target object analyzed before. This is achieved through voiceprint recognition technology, which can distinguish the voice characteristics of different members in the family.

[0069] If the identification result shows that the speaker is not the previous target object, it means that another family member has started to use the voice device for interaction. At this time, the system will check the library of silence time windows to see if a third silence time window has been stored for the newly identified user (i.e., the speaker).

[0070] If the speaker's silence time window has not been stored in the library of silence time windows, the system will process it in the same way as the previous target object. First, it will determine the third weight corresponding to the third interactive voice, which reflects the importance of the new user's voice input in voice recognition and silence time window prediction. Then, the system analyzes the pause intervals in the third interactive voice to obtain the fourth pause interval time. By combining the third weight and the fourth pause interval time, the system can predict a third silence time window suitable for the new user and store it in the library of silence time windows for subsequent use.

[0071] The above steps ensure that the smart home voice device can quickly adapt to the unique speaking habits of each user when multiple family members use it, providing each user with the most optimized voice interaction experience. Especially in a multi-person environment, this quick and accurate user identification and personalized silence time window setting avoids the problem of inaccurate recognition or excessively long interaction time that may occur when using a unified setting, thereby significantly improving the efficiency of voice recognition and the user experience of smart homes.

[0072] To better understand the process of predicting the silence time window described above, the following optional embodiments will further describe the implementation method of predicting the silence time window, but not used to limit the technical solutions of the embodiments of the present application.

[0073] Voice interaction application in smart home is increasingly popular, but different user groups have different habits when speaking. To ensure that the voice of each user can be accurately picked up, the traditional approach often chooses a longer silence time window to adapt to most users, which undoubtedly increases the overall interaction time and causes the user to respond slowly. To solve this problem, the optional embodiment of the present application proposes a method for improving the success rate of speech recognition and shortening the length of voice interaction. The above method can adjust the length of the silence time window in real time according to the speaking habits of each user. In this way, not only the voice intention of those who speak slowly and have long sentence breaks can be completely recognized, but also the overall interaction time of normal users is significantly shortened, thereby greatly improving the experience of intelligent voice interaction in smart home.

[0074] In the smart home scene, the convenient entrance of voice interaction is usually provided by advanced devices such as smart brain screens or smart sound boxes with high-sensitivity recording functions. The whole voice interaction process is efficient and smooth: the user only needs to activate the interaction device through a simple wake-up word, and the device will immediately enter an efficient sound pickup state to continuously capture the user's voice instructions. These audio data are then quickly uploaded to the cloud for high-precision audio recognition and NLP (Natural Language Processing) analysis. With the powerful capabilities of large models, the system can accurately identify and output the user's intention, then execute the corresponding voice instructions, and immediately give the user clear and intuitive feedback. This process not only improves the intelligent level of home life, but also ensures seamless and natural interaction experience between the user and the smart home device. Figure 3 The overall framework diagram of the voice interaction between the user and the large model according to the optional embodiment of the present application is shown in Figure 3 .

[0075] When the user speaks, the user's speech is picked up, and the result of the pickup is denoised. After denoising, the voice is subjected to VAD detection to obtain audio data. The audio data is sent to the cloud for processing. Specifically, the smart cloud service performs automatic speech recognition (ASR) on the audio data, and sends the recognition result to the large model to understand the user's intention and perform text-to-speech (TTS) conversion. The converted audio broadcast data is returned to the smart device for voice broadcast.

[0076] According to the framework diagram in Figure 3 , the calculation rule of the voice interaction time can be determined, specifically:

[0077] Figure 4 The voice interaction process according to the embodiment of the present application is shown inFigure 4 as shown:

[0078] User wakes up the device, for example: Xiao You Xiao You;

[0079] User inputs voice interaction control, for example: turn on the light;

[0080] Voice interaction device carries out Figure 3 noise reduction, VAD detection, ASR speech recognition, large model intent understanding, and TTS conversion steps in the process, and then the interaction device plays back the TTS audio response, for example: OK, it has been turned on.

[0081] That is, the interaction time of this time is the interval from the end of the user saying the sentence "turn on the light" to the start of playing back the TTS sound "OK, it has been turned on".

[0082] The optional embodiments of the present application significantly optimize the voice VAD process in the smart home scene. Specifically, the optional embodiments of the present application use the reported audio data to accurately identify different user members in the family through voiceprint recognition technology, and accordingly customize appropriate silence time windows for each user. Further, by carefully analyzing the sentence break interval time of the user audio, combined with the prediction model of the weighted moving average algorithm, the optional embodiments of the present application can dynamically adjust and output a brand new, more accurate silence time window to adapt to the unique speaking habits of each user. This optimization not only improves the accuracy and efficiency of voice interaction, but also brings a more personalized and smooth voice interaction experience for smart home users. Specifically: Figure 5 is a device schematic diagram corresponding to a method for improving speech recognition success rate and shortening voice interaction time according to the optional embodiments of the present application, as shown, specifically comprising the following modules: Figure 5

[0083] (1) Silence time window selection module.

[0084] The core function of the silence time window selection module is to accurately identify the user's voice, and use advanced voiceprint recognition technology to lock which member in the family the current speaker belongs to. Once the user's identity is identified, the system will quickly retrieve and obtain the personalized silence time window value of the member. The personalized silence time window value is then sent to the terminal device as a key parameter for VAD (Voice Activity Detection) to ensure the accuracy and integrity of voice interaction.

[0085] (2) Sentence break interval data collection module.

[0086] ​The pause interval data collection module focuses on collecting and analyzing all pause interval data of the user during the voice interaction process. The pause interval data collection module records the voice pause interaction interval in each interaction in detail and forms a collection. Through in-depth analysis of the collection, the system can accurately capture the user's pause habits and extract the maximum pause interval time. This key data is then sent to the cloud to provide strong support for subsequent model prediction.

[0087] (3) Adaptive model prediction module.

[0088] The adaptive model prediction module is the core intelligent component. Based on the current user data, the adaptive model prediction module uses the weighted moving average algorithm for in-depth analysis and prediction. Through continuous learning and adaptation, the system can output a personalized silence time window value suitable for the next interaction. In addition, the adaptive model prediction module can also dynamically adjust according to possible changes in the user, such as changes in pause habits, to ensure the real-time and accuracy of voice interaction. This adaptive ability enables the system to continuously provide users with more intelligent and convenient home experiences.

[0089] According to the above three modules, a specific process of a method for improving the success rate of voice recognition and shortening the length of voice interaction can be determined, Figure 6 is a flowchart of a method for improving the success rate of voice recognition and shortening the length of voice interaction according to an optional embodiment of the present application, as Figure 6 shown:

[0090] Step S601, endpoint detection.

[0091] Receive noise-reduced audio data and perform VAD endpoint detection.

[0092] Step S602, determine whether the voice appears, disappears, or is silent.

[0093] In the case of determining that the voice disappears (i.e., a voice disappearance event), step S603-1 is performed; in the case of determining that the voice appears (i.e., a voice appearance event), step S603-2 is performed; in the case of determining that it is silent (a silent event), step S603-3 is performed.

[0094] Step S603-1, record time T2 (i.e., the first time).

[0095] If the event of the user's speech disappearing is detected, it means that the user's speech has a pause or will end the conversation, and the time T2 at this time is recorded. Continue to perform endpoint detection.

[0096] Step S603-2, determine whether there is a voice disappearance event before.

[0097] In the case of determining that there is a speech disappearance event before, step S604-1 is performed, and in the case of determining that there is no speech disappearance event before, step S604-2 is performed.

[0098] Step S604-1, record time T1.

[0099] If it is detected that the user starts to speak, if there is a speech disappearance event before, the difference between the current time T1 and the time T2 when the speech disappearance event occurs is recorded, and the audio data is sent to the cloud.

[0100] Step S605, update the sentence interval time set according to the current sentence interval time.

[0101] Wherein, the current sentence interval time td n=T1-T2. The sentence interval time set TD={td 1, td 2, …, td n}.

[0102] Step S604-2, send data to the cloud.

[0103] Step S606, the voiceprint distinguishes the user.

[0104] Step S607, obtain the silence time window ST.

[0105] After the cloud receives the audio data reported by the smart device, it will perform voiceprint recognition to determine which member of the family the current user is, and obtain the silence time window ST (i.e. the target silence time window) to which the member belongs, and send it to the smart device.

[0106] Step S603-3, record time T3.

[0107] If the speaking is continuously detected, and the interval from the speech disappearance exceeds the silence time window, it is considered that the current user has finished speaking. Stop sending voice data to the cloud, and record the time T3 when the voice data is stopped to the cloud.

[0108] Step S608, calculate the interval ΔT=T3-T2 (i.e. the second sentence interval time).

[0109] Step S609, determine whether ΔT is within the silence time window ST.

[0110] In the case of determining that it is within the silence time window, step S601 is performed, and in the case of determining that it is not within the silence time window, step S610 is performed.

[0111] Step S610, determine whether ΔT<ST+CT is met.

[0112] If ΔT < ST (i.e., target silence time window) + CT (i.e., preset extension time) is met, step S601 is performed, and if ΔT < ST + CT is not met, step S611 is performed.

[0113] After the silence time window is exceeded, the delay time CT is increased for endpoint detection, and if the CT delay time is exceeded, the pickup is ended, and the maximum value of the session pause interval time set is sent to the cloud.

[0114] Step S611, the pickup is ended.

[0115] Step S612, the maximum value of the TD set is sent to the cloud.

[0116] Step S613, the data is updated to the model (i.e., target model).

[0117] Step S614, the weighted moving average algorithm prediction model calculation.

[0118] Step S615, a new silence time window ST1 is output.

[0119] The cloud adds the data to the prediction model, calculates the silence time window suitable for the user through the weighted moving average algorithm, and prepares for next use.

[0120] In the optimization of the silence time window algorithm model, the weighted moving average algorithm can be selected in the optional embodiment of the application. This algorithm fully considers the time value of the data, i.e., the influence degree of data at different time points on the predicted value is different. By assigning different weights to the data at each time point, the silence time window of the user can be more accurately predicted, thereby improving the efficiency and accuracy of voice interaction. The algorithm calculation formula is as follows:

[0121] M t =a1Y1+a2Y2+...+a t Y t 。

[0122] Wherein, M t is the moving average number (i.e., silence time window) of the t period; Y t is the observation data of the t period; t is the moving step; a t is the weight number of the t period. Y1 and Y2 are the observation data of the 1st and 2nd periods; a1 and a2 are the moving average numbers of the 1st and 2nd periods. a1, a2, and a t are the weights corresponding to each voice interaction data;

[0123] Specifically, the core of the weighted moving average algorithm lies in assigning different weights to the data within the moving segment. In the optional embodiment of the present application, data of nearly 100 periods is selected as the moving step, which means that the algorithm considers the pause interval data of the user's last 100 voice interactions. For these 100 periods of data, different moving weights a t are assigned, where t represents the period number of the data. The design of these weights follows a basic principle: the more recent the data, the greater the impact on the predicted value. Therefore, a larger weight is assigned to recent data, and a smaller weight is assigned to data of a longer period.

[0124] Through this method, changes in the user's voice interaction habits can be effectively captured. For example, if the user exhibits longer pause intervals in recent interactions, the weighted moving average algorithm will quickly increase the time, thereby outputting a longer silence time window. Conversely, if the user's pause intervals gradually shorten, the algorithm will adjust the time accordingly, outputting a shorter silence time window.

[0125] This dynamic adjustment of the silence time window not only improves the response speed of voice interaction, but also enhances the user experience. Users can interact more naturally with smart home devices without waiting for a long silence time window to confirm voice commands. At the same time, since the algorithm can learn the user's voice habits in real time, it can adapt to the individual needs of different users, providing more accurate and efficient voice services for users.

[0126] According to the optional embodiment of the present application, first, the optional embodiment of the present application adopts audio data pickup technology to capture the user's voice commands in real time and report the data to the cloud for processing. This step ensures the accuracy and completeness of the voice data, providing a solid foundation for subsequent analysis and recognition. Then, through voiceprint recognition technology, the system can accurately distinguish different user members in the family. This function not only provides personalized service experience, but also enhances the security and privacy protection capabilities of the system. Because different user members may have different voice habits and preferences, voiceprint recognition technology can ensure that the system provides the most suitable silence time window settings for each user member. Then, the system uses advanced endpoint detection technology to detect the user's audio input and calculates the pause interval time. This step is crucial for accurately recognizing the user's voice commands. By continuously collecting and analyzing the user's pause data, the system can gain a deep understanding of the user's voice habits and provide strong support for the subsequent prediction model. Finally, using the weighted moving average algorithm prediction model, the system can dynamically adjust the silence time window settings for each user member based on historical data and current data. This prediction model can capture changes in the user's voice habits in real time and adjust the length of the silence time window accordingly, thereby optimizing the efficiency and accuracy of voice interaction.

[0127] That is, the optional embodiments of the present application can improve the interaction success rate: through voiceprint recognition and endpoint detection technology, the system can more accurately identify the user's voice instructions, reduce misjudgment and missed cases, thereby improving the interaction success rate of the abnormal interaction group. It can also shorten the interaction time: through the weighted moving average algorithm prediction model, the system can output the most suitable silent time window setting for the user, reducing unnecessary waiting time and shortening the overall interaction time of the normal interaction group. It can also provide personalized service experience: the system can provide personalized service experience according to the voice habits and preferences of different user members, meet the needs and expectations of different users.

[0128] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, or network device, etc.) execute the methods of various embodiments of the present application.

[0129] Figure 7 is a structural block diagram of a silent time window prediction device according to an embodiment of the present application; as shown in Figure 7 , comprising:

[0130] The acquisition module 72 is configured to acquire historical voice interaction data of a target object, wherein the historical voice interaction data is used to indicate the interaction voice of the target object with a voice device;

[0131] The first determination module 74 is configured to determine a first time difference between the voice interaction time corresponding to the historical voice interaction data and the current time, and determine a first sentence break interval time corresponding to the historical voice interaction data;

[0132] The second determination module 76 is configured to determine a target weight corresponding to the historical voice interaction data according to the first time difference;

[0133] The prediction module 78 is configured to predict a target silent time window of the target object interacting with the voice device according to the target weight and the first sentence break interval time.

[0134] By the above apparatus, the historical voice interaction data of the target object and the voice device is acquired, the historical voice interaction data and a first time difference between the corresponding voice interaction time and the current time are determined, and a first sentence breaking time interval corresponding to the historical voice interaction data is determined, and a target weight corresponding to the historical voice interaction data is determined according to the first time difference; and a target silence time window of the target object and the voice device is predicted according to the target weight and the first sentence breaking time interval. That is, according to the first time difference between the voice interaction time corresponding to the historical voice interaction data and the current time, the target weight is determined, and then the target silence time window is predicted according to the target weight and the first sentence breaking time interval. According to the embodiment of the present application, the problem of increasing the overall voice interaction time caused by selecting a longer silence time window to adapt to different sentence breaking habits of different user groups when speaking in the related art can be solved. Then, different silence time windows are determined to adapt to different sentence breaking habits of different user groups, and the overall voice interaction time is reduced.

[0135] In one example embodiment, the second determining module 76 is further configured to acquire a plurality of voice interaction data corresponding to the target object, and determine a second time difference between the voice interaction time corresponding to each voice interaction data and the current time; divide the plurality of voice interaction data and the second time difference corresponding to each voice interaction data into training set data and test set data, and acquire an initial attenuation coefficient; train a target model according to the training set data and the initial attenuation coefficient to obtain a trained target model, wherein the target model is used to determine a weight corresponding to each voice interaction data; verify the trained target model according to the test set data; and in a case where it is determined that the trained target model passes the verification, input the first time difference into the trained target model, so that the trained target model outputs the target weight.

[0136] In one example embodiment, the first determining module 74 is further configured to determine whether a first voice interruption event exists in the historical voice interaction data, wherein the first voice interruption event includes a voice disappearance event or a voiceless event; in a case where it is determined that the first voice interruption event exists in the historical voice interaction data, determine a first time at which the first voice interruption event starts, and determine whether a voice appearance event exists after the first voice interruption event; in a case where it is determined that the voice appearance event exists after the first voice interruption event, determine a second time at which the voice appearance event starts, and determine the first sentence breaking time interval according to a difference between the second time and the first time; in a case where it is determined that the voice appearance event does not exist after the first voice interruption event, determine a third time at which the historical voice interaction data ends, and determine the first sentence breaking time interval according to a difference between the third time and the first time.

[0137] In an example embodiment, the first determining module 74 is further configured to, in a case where it is determined that the first voice interruption event exists in the historical voice interaction data, determine whether the voice appearance event exists in a target time period after the first time; in a case where it is determined that the voice appearance event exists in the target time period, determine that the voice interruption event is the voice disappearance event; in a case where it is determined that the voice appearance event does not exist in the target time period, determine that the voice interruption event is the voiceless event.

[0138] In an example embodiment, the prediction module 78 is further configured to, in a case where it is determined that the voice device receives a first interaction voice uttered by the target object, determine whether a second voice interruption event exists in the first interaction voice; in a case where it is determined that the second voice interruption event exists in the first interaction voice, determine whether the second voice interruption event is a voiceless event; in a case where it is determined that the second voice interruption event is the voiceless event, determine a sum value of the target silence time window and a preset extension time, and determine a first size relationship between a second pause interval time corresponding to the first interaction voice and the sum value; in a case where the first size relationship indicates that the second pause interval time is greater than or equal to the sum value, end the sound pickup of the first interaction voice, and determine a first weight corresponding to the first interaction voice after the sound pickup is ended; predict a first silence time window of the target object interacting with the voice device according to the first weight and the second pause interval time, and update the target silence time window according to the first silence time window.

[0139] In an example embodiment, the prediction module 78 is further configured to, in a case where the first size relationship indicates that the second pause interval time is less than the sum value, perform endpoint detection on the first interaction voice to obtain a voice appearance event after the second pause interval; determine whether the voiceless event exists in a second interaction voice corresponding to the voice appearance event; in a case where it is determined that the voiceless event exists in the second interaction voice, determine a second size relationship between a third pause interval time corresponding to the second interaction voice and the sum value; in a case where the second size relationship indicates that the third pause interval time is greater than or equal to the sum value, end the sound pickup of the second interaction voice, and merge the first interaction voice and the second interaction voice to update the first interaction voice; determine a second weight corresponding to the updated first interaction voice; predict a second silence time window of the target object interacting with the voice device according to the second weight, the second pause interval time, and the third pause interval time, and update the target silence time window according to the second silence time window.

[0140] In an example embodiment, the prediction module 78 is further configured to predict a second silence time window in which the target object interacts with the voice device according to the second weight, the second pause interval time and the third pause interval time, including: determining a third size relationship between the second pause interval time and the third pause interval time; determining a maximum pause interval time from the second pause interval time and the third pause interval time according to the third size relationship; and determining the second silence time window according to the second weight and the maximum pause interval time.

[0141] In an example embodiment, the prediction module 78 is further configured to, in a case where it is determined that the voice device receives a third interaction voice, determine whether a voice object corresponding to the third interaction voice is the target object; in a case where it is determined that the voice object is not the target object, determine whether a third silence time window in which the voice object interacts with the voice device exists in a silence time window library, wherein the silence time window library includes a target silence time window corresponding to the target object; in a case where it is determined that the third silence time window does not exist in the silence time window library, determine a third weight corresponding to the third interaction voice and a fourth pause interval time corresponding to the third interaction voice; and predict the third silence time window according to the third weight and the fourth pause interval time, and store the third silence time window corresponding to the voice object in the silence time window library.

[0142] Embodiments of the present application also provide a storage medium including a stored program, wherein the program performs any of the above methods when executed.

[0143] Optionally, in the present embodiment, the storage medium can be configured to store program code for performing the following steps:

[0144] S1, obtaining historical voice interaction data of a target object, wherein the historical voice interaction data is used to indicate an interaction voice of the target object with a voice device;

[0145] S2, determining a first time difference between a voice interaction time corresponding to the historical voice interaction data and a current time, and determining a first pause interval time corresponding to the historical voice interaction data;

[0146] S3, determining a target weight corresponding to the historical voice interaction data according to the first time difference;

[0147] S4, predicting a target silence time window in which the target object interacts with the voice device according to the target weight and the first pause interval time.

[0148] The embodiment of the present application further provides an electronic device, comprising a memory and a processor, the memory storing a computer program, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments.

[0149] Optionally, the electronic device can further comprise a transmission device connected with the processor and an input and output device connected with the processor.

[0150] Optionally, in the embodiment, the processor can be configured to execute the following steps by the computer program:

[0151] S1, obtaining historical voice interaction data of a target object, wherein the historical voice interaction data is used to indicate interactive voice of the target object and a voice device;

[0152] S2, determining a first time difference between a voice interaction time corresponding to the historical voice interaction data and a current time, and determining a first pause interval time corresponding to the historical voice interaction data;

[0153] S3, determining a target weight corresponding to the historical voice interaction data according to the first time difference;

[0154] S4, predicting a target silence time window of the target object interacting with the voice device according to the target weight and the first pause interval time.

[0155] The embodiment of the present application further provides a computer program product, comprising a computer program, and the computer program is executed by a processor to execute the steps in any of the above method embodiments.

[0156] Optionally, in the embodiment, the computer program product can be executed by the processor to execute the following steps:

[0157] S1, obtaining historical voice interaction data of a target object, wherein the historical voice interaction data is used to indicate interactive voice of the target object and a voice device;

[0158] S2, determining a first time difference between a voice interaction time corresponding to the historical voice interaction data and a current time, and determining a first pause interval time corresponding to the historical voice interaction data;

[0159] S3, determining a target weight corresponding to the historical voice interaction data according to the first time difference;

[0160] S4, predicting a target silence time window of the target object interacting with the voice device according to the target weight and the first pause interval time.

[0161] Optionally, in the embodiment, the storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various storage medium capable of storing program codes.

[0162] Optionally, the specific examples in the embodiment can refer to the examples described in the above embodiments and optional implementation manners, and the embodiment will not be described herein.

[0163] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be realized by a general computing device, which can be concentrated on a single computing device or distributed on a network composed of multiple computing devices, and optionally, they can be realized by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described herein can be executed in different order, or they can be manufactured into individual integrated circuit modules, or multiple modules or steps thereof can be manufactured into a single integrated circuit module. Thus, the present application is not limited to any specific combination of hardware and software.

[0164] The above only describes the preferred embodiments of the present application, and it should be noted that those skilled in the art can make several improvements and refinements without departing from the principles of the present application, and these improvements and refinements should be considered as the protection scope of the present application.

Claims

1. A method for predicting silent time windows, characterized in that, include: Acquire historical voice interaction data of the target object, wherein the historical voice interaction data is used to indicate the voice interaction between the target object and the voice device; Determine the first time difference between the voice interaction time corresponding to the historical voice interaction data and the current time, and determine the first sentence break interval time corresponding to the historical voice interaction data; The target weights corresponding to the historical voice interaction data are determined based on the first time difference. Predict the target silence time window for the interaction between the target object and the voice device based on the target weight and the first sentence segmentation interval; The determination of the target weights corresponding to the historical voice interaction data based on the first time difference includes: Acquire multiple voice interaction data corresponding to the target object, and determine a second time difference between the voice interaction time corresponding to each voice interaction data and the current time; The multiple voice interaction data and the second time difference corresponding to each voice interaction data are divided into training set data and test set data, and the initial attenuation coefficient is obtained; The target model is trained based on the training set data and the initial attenuation coefficient to obtain the trained target model, wherein the target model is used to determine the weights corresponding to each voice interaction data. The trained target model is validated based on the test set data. If the target model is verified to be successful after training, the first time difference is input into the target model after training so that the target weights are output by the target model after training.

2. The prediction method for the silent time window according to claim 1, characterized in that, Determining the first sentence interval time corresponding to the historical voice interaction data includes: Determine whether a first voice interruption event exists in the historical voice interaction data, wherein the first voice interruption event includes: a voice disappearance event or a silence event; If it is determined that the first voice interruption event exists in the historical voice interaction data, the first time when the first voice interruption event started is determined, and whether there is a voice occurrence event after the first voice interruption event is determined; If a speech interruption event occurs after the first speech interruption event, a second time when the speech interruption event begins is determined, and the first sentence interruption interval time is determined based on the difference between the second time and the first time. If no voice interruption event occurs after the first voice interruption event, a third time when the historical voice interaction data ends is determined, and the first sentence interruption interval time is determined based on the difference between the third time and the first time.

3. The prediction method for the silent time window according to claim 2, characterized in that, After determining whether a first voice interruption event exists in the historical voice interaction data, the method further includes: If it is determined that the first voice interruption event exists in the historical voice interaction data, it is determined whether the voice occurrence event exists within a target time period after the first time. If the voice occurrence event is determined to occur within the target time period, the voice interruption event is determined to be the voice disappearance event; If it is determined that no voice occurrence event occurs within the target time period, the voice interruption event is determined to be the silent event.

4. The prediction method for the silent time window according to claim 1, characterized in that, After predicting the target silence time window for the interaction between the target object and the voice device based on the target weight and the first sentence segmentation interval, the method further includes: If it is determined that the voice device has received the first interactive voice sent by the target object, it is determined whether there is a second voice interruption event in the first interactive voice; If it is determined that the second voice interruption event exists in the first interactive voice, determine whether the second voice interruption event is a silent event; If the second voice interruption event is determined to be the silent event, the sum of the target silence time window and the preset extension time is determined, and the first size relationship between the second sentence interval time corresponding to the first interactive voice and the sum is determined. If the first size relationship indicates that the second sentence interval time is greater than or equal to the sum value, the sound pickup of the first interactive voice is terminated, and the first weight corresponding to the first interactive voice after the sound pickup is terminated is determined. The first silence time window for the interaction between the target object and the voice device is predicted based on the first weight and the second sentence interval time, and the target silence time window is updated based on the first silence time window.

5. The prediction method for the silent time window according to claim 4, characterized in that, After determining the first relationship between the second sentence interval time corresponding to the first interactive voice and the first magnitude relationship of the sum value, the method further includes: When the first size relationship indicates that the second sentence interval time is less than the sum value, endpoint detection is performed on the first interactive voice to obtain the voice occurrence event after the second sentence interval; Determine whether the silent event exists in the second interactive voice corresponding to the voice occurrence event; If it is determined that the silent event exists in the second interactive voice, determine the second relationship between the third sentence interval time corresponding to the second interactive voice and the second magnitude of the sum value; If the second size relationship indicates that the third sentence interval time is greater than or equal to the sum value, the pickup of the second interactive voice ends, and the first interactive voice and the second interactive voice are merged to update the first interactive voice; Determine the second weight corresponding to the updated first interactive voice; The second silence time window for the interaction between the target object and the voice device is predicted based on the second weight, the second sentence interval time, and the third sentence interval time, and the target silence time window is updated based on the second silence time window.

6. The prediction method for the silent time window according to claim 5, characterized in that, Based on the second weight, the second sentence segmentation interval, and the third sentence segmentation interval, a second silence time window for the interaction between the target object and the voice device is predicted, including: Determine the third relationship between the second sentence interval and the third sentence interval; The maximum sentence interval time between the second sentence interval time and the third sentence interval time is determined based on the third size relationship. The second silent time window is determined based on the second weight and the maximum sentence break interval.

7. The prediction method for the silent time window according to claim 1, characterized in that, After predicting the target silence time window for the interaction between the target object and the voice device based on the target weight and the first sentence segmentation interval, the method further includes: If it is determined that the voice device has received a third interactive voice, it is determined whether the voice source corresponding to the third interactive voice is the target object; If it is determined that the voice-emitting object is not the target object, it is determined whether there is a third silent time window in the silent time window library for the interaction between the voice-emitting object and the voice device, wherein the silent time window library includes: the target silent time window corresponding to the target object; If it is determined that the third silent time window does not exist in the silent time window library, the third weight corresponding to the third interactive voice and the fourth sentence interval time corresponding to the third interactive voice are determined. The third silence time window is predicted based on the third weight and the fourth sentence interval, and the third silence time window corresponding to the vocal object is stored in the silence time window library.

8. A prediction device for a silent time window, characterized in that, include: The acquisition module is used to acquire historical voice interaction data of the target object, wherein the historical voice interaction data is used to indicate the voice interaction between the target object and the voice device; The first determining module is used to determine the first time difference between the voice interaction time corresponding to the historical voice interaction data and the current time, and to determine the first sentence interval time corresponding to the historical voice interaction data. The second determining module is used to determine the target weights corresponding to the historical voice interaction data based on the first time difference. The prediction module is used to predict the target silence time window of the interaction between the target object and the voice device based on the target weight and the first sentence interval time; The second determining module is further configured to: acquire multiple voice interaction data corresponding to the target object; determine a second time difference between the voice interaction time corresponding to each voice interaction data and the current time; divide the multiple voice interaction data and the second time difference corresponding to each voice interaction data into training set data and test set data, and acquire an initial attenuation coefficient; train the target model according to the training set data and the initial attenuation coefficient to obtain a trained target model, wherein the target model is used to determine the weights corresponding to each voice interaction data; verify the trained target model according to the test set data; and, if the trained target model is verified to be successful, input the first time difference into the trained target model so that the trained target model outputs the target weights.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method described in any one of claims 1 to 7.

10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 7 through the computer program.

Citation Information

Patent Citations

  • Speech recognition method, control method, model training method and device

    CN114360531A

  • Voice endpoint judgment method and device, equipment, storage medium and product

    CN114495981A