Voice interaction response duration detection method and device, storage medium and equipment
By sliding the window on the voice interactive audio waveform diagram to detect the amplitude of the audio sample point, the voice interaction response time is automatically calculated, which solves the problem of low efficiency and large error in manual acquisition time, and improves the accuracy and efficiency of the test.
Patent Information
- Application Number
- CN202510443459.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, the method of manually obtaining the response time of voice interaction is inefficient and has large errors, which affects the accuracy of test results and product development efficiency.
By obtaining the voice interactive audio waveform diagram of the user and the smart device, sliding on the audio waveform diagram with a preset duration sliding window, detecting the amplitude change of the audio sampling point, automatically determining the end of the user audio and starting time of the device audio, and calculating the response time.
It realizes automatic acquisition of response time, simplifies the operation process, improves the accuracy and efficiency of response time acquisition, and reduces human error.
Smart Images

Figure CN120279890A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart home / smart family. Specifically, it relates to a method, device, storage medium, and equipment for detecting the voice interaction response duration. Background Art
[0002] With the rapid development of artificial intelligence and Internet of Things technologies, intelligent voice technology has become an important part of modern intelligent products. Among them, devices such as smart speakers, smart TVs, and smart air conditioners provide users with a more convenient and user-friendly usage experience through voice interaction.
[0003] In the development process of intelligent voice devices, ensuring the accuracy and response speed of voice interaction is crucial. To ensure that intelligent voice devices can perform high-quality interactions, a large number of tests are required to detect whether indicators such as the response speed during voice interaction are qualified.
[0004] Currently, the commonly used testing methods mostly involve manually recording the audio of the entire interaction process and then manually obtaining the response duration of the voice device to the user's wake-up statement. However, the method of manually obtaining the response duration has problems such as low efficiency and large errors. Summary of the Invention
[0005] This application provides a method, device, storage medium, and equipment for detecting the voice interaction response duration to solve the technical problems of low efficiency and large errors in the method of manually obtaining the response duration.
[0006] In a first aspect, this application provides a method for detecting the voice interaction response duration, including:
[0007] Obtain the audio waveform diagram of the voice interaction between the user and the intelligent device, and create a window with a preset duration on the audio waveform diagram;
[0008] Control the window to slide from left to right at a preset sliding step length, and detect the first amplitude of each audio sampling point in the window after each slide. Determine the end moment of the user's audio according to the first amplitude of each audio sampling point;
[0009] Control the window to slide from right to left at a preset sliding step length, and detect the second amplitude of each audio sampling point in the window after each slide. Determine the start moment of the intelligent device's audio according to the second amplitude of each audio sampling point;
[0010] Determine the difference between the start moment of the intelligent device's audio and the end moment of the user's audio as the voice interaction response duration of the intelligent device.
[0011] Optionally, the determining the end moment of the user's audio according to the first amplitude of each audio sampling point includes:
[0012] If the first amplitude of each audio sampling point within the window is less than or equal to the preset threshold, the start time of the window is determined as the end time of the user's audio.
[0013] Optionally, determining the end time of the user's audio according to the first amplitude of each audio sampling point includes:
[0014] If in a preset number of consecutive slides, the first amplitude of each audio sampling point within the window after each slide is less than or equal to the preset threshold, the start time of the window during the preset number of consecutive slides is determined as the end time of the user's audio.
[0015] Optionally, determining the start time of the intelligent device's audio according to the second amplitude of each audio sampling point includes:
[0016] If the second amplitude of each audio sampling point within the window is less than or equal to the preset threshold, the end time of the window is determined as the start time of the intelligent device's audio.
[0017] Optionally, determining the start time of the intelligent device's audio according to the second amplitude of each audio sampling point includes:
[0018] If in a preset number of consecutive slides, the first amplitude of each audio sampling point within the window after each slide is less than or equal to the preset threshold, the end time of the window during the preset number of consecutive slides is determined as the start time of the intelligent device's audio.
[0019] Optionally, obtaining the audio waveform diagram of the voice interaction between the user and the intelligent device includes:
[0020] Start audio recording and play the voice interaction corpus of the user. After the intelligent device responds to the voice interaction corpus of the user, obtain the audio file of the voice interaction between the user and the intelligent device recorded.
[0021] Load the audio file to generate the audio waveform diagram.
[0022] In a second aspect, the present application provides a detection device for the voice interaction response duration, including:
[0023] An acquisition module, configured to acquire the audio waveform diagram of the voice interaction between the user and the intelligent device.
[0024] A processing module, configured to create a window with a preset duration on the audio waveform diagram.
[0025] The processing module is further configured to control the window to slide from left to right by a preset sliding step length, detect the first amplitude of each audio sampling point within the window after each slide, and determine the audio end moment of the user according to the first amplitude of each audio sampling point.
[0026] The processing module is further configured to control the window to slide from right to left by a preset sliding step length, detect the second amplitude of each audio sampling point within the window after each slide, and determine the audio start moment of the intelligent device according to the second amplitude of each audio sampling point.
[0027] The processing module is further configured to determine the difference between the audio start moment of the intelligent device and the audio end moment of the user as the voice interaction response duration of the intelligent device.
[0028] Optionally, if the first amplitude of each audio sampling point within the window is less than or equal to the preset threshold, the processing module is further configured to determine the time start point of the window as the audio end moment of the user.
[0029] Optionally, if in a continuous preset number of slides, the first amplitude of each audio sampling point within the window after each slide is less than or equal to the preset threshold, the processing module is further configured to determine the time start point of the window in the continuous preset number of slides as the audio end moment of the user.
[0030] Optionally, if the second amplitude of each audio sampling point within the window is less than or equal to the preset threshold, the processing module is further configured to determine the time end point of the window as the audio start moment of the intelligent device.
[0031] Optionally, if in a continuous preset number of slides, the first amplitude of each audio sampling point within the window after each slide is less than or equal to the preset threshold, the processing module is further configured to determine the time end point of the window in the continuous preset number of slides as the audio start moment of the intelligent device.
[0032] Optionally, the processing module is further configured to start audio recording, play the voice interaction corpus of the user, and obtain the audio file of the voice interaction between the user and the intelligent device after the intelligent device responds to the voice interaction corpus of the user.
[0033] The processing module is further configured to load the audio file to generate the audio waveform diagram.
[0034] In a third aspect, the present application provides a computer-readable storage medium, on which a computer execution program is stored. When the computer execution program is executed by a processor, it is used to implement the method for detecting the voice interaction response duration as described in the first aspect above and various possible implementation manners of the first aspect.
[0035] In a fourth aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0036] The memory stores computer execution instructions;
[0037] The processor executes the computer execution instructions stored in the memory to implement the method for detecting the voice interaction response duration as described in the first aspect above and various possible implementation manners of the first aspect.
[0038] In a fifth aspect, the present application provides a program product, including a computer program, which implements the method for detecting the voice interaction response duration as described above when executed by a processor.
[0039] The method, device, storage medium and equipment for detecting the voice interaction response duration provided by the present application obtain an audio waveform diagram of the voice interaction between a user and a smart device, and create a window with a preset duration on the audio waveform diagram; control the window to slide from left to right at a preset sliding step length, and detect the first amplitude of each audio sampling point in the window after each slide. If the first amplitude of each audio sampling point in the window is less than or equal to a preset threshold, determine the start time of the window as the end time of the user's audio; then control the window to slide from right to left at a preset sliding step length, and detect the second amplitude of each audio sampling point in the window after each slide. If the second amplitude of each audio sampling point in the window is less than or equal to a preset threshold, determine the end time of the window as the start time of the smart device's audio, and determine the difference between the start time of the smart device's audio and the end time of the user's audio as the voice interaction response duration of the smart device. This method realizes the automatic acquisition of the response duration, simplifies the extraction process of the response duration, provides a more convenient operation method for users, not only improves the accuracy of obtaining the response duration, but also improves the acquisition efficiency of the response duration, and avoids the human error brought by the traditional acquisition method. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0041] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 Schematic diagram of the hardware environment for the method of detecting the voice interaction response duration provided by the present application;
[0043] Figure 2 Flow chart of the method of detecting the voice interaction response duration provided by the present application Figure 1 ;
[0044] Figure 3 Flow chart of the method of detecting the voice interaction response duration provided by the present application Figure 2 ;
[0045] Figure 4 Schematic diagram of the structure of the device for detecting the voice interaction response duration provided by the present application;
[0046] Figure 5 Schematic diagram of the structure of an electronic device provided by the present application. Detailed implementation manners
[0047] In order to enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0048] It should be noted that the terms "first", "second", etc. in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0049] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0050] According to one aspect of the embodiments of the present application, a method for detecting the voice interaction response duration is provided. The method for detecting the voice interaction response duration is widely applied to whole-house intelligent digital control application scenarios such as Smart Home, smart home, intelligent household device ecosystem, Intelligence House ecosystem, etc. Optionally, in this embodiment, the above method for detecting the voice interaction response duration can be applied to, for example Figure 1 the hardware environment composed of a terminal device 102 and a server 104 as shown. As Figure 1 shown, the server 104 is connected to the terminal device 102 through a network and can be used to provide services (such as application services, etc.) for the terminal or the client installed on the terminal. A database can be set on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data operation services for the server 104.
[0051] The above network can include but is not limited to at least one of the following: wired network, wireless network. The above wired network can include but is not limited to at least one of the following: wide area network, metropolitan area network, local area network. The above wireless network can include but is not limited to at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 is not limited to being a PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing device, smart dishwasher, smart projection device, smart TV, smart clothes hanger, smart curtain, smart audio and video, smart socket, smart speaker, smart sound box, smart fresh air device, smart kitchen and bathroom device, smart bathroom device, smart floor sweeping robot, smart window cleaning robot, smart mopping robot, smart air purification device, smart steam box, smart microwave oven, smart kitchen water heater, smart purifier, smart water dispenser, smart door lock, etc.
[0052] With the rapid development of artificial intelligence and Internet of Things technologies, intelligent voice technology has become an important part of modern consumer electronics products; among them, devices such as smart speakers, smart TVs, and smart air conditioners provide users with a more convenient and user-friendly usage experience through voice interaction. The above-mentioned intelligent devices usually rely on technologies such as natural language processing, speech recognition, and speech synthesis to interact with users.
[0053] During the product development process, it is crucial to ensure the accuracy and response speed of voice interaction. The performance of voice interaction directly affects the user experience, especially in terms of response speed and recognition accuracy. To ensure that intelligent devices can perform high-quality interactions, a large number of tests are required to detect whether indicators such as the response speed during voice interaction are qualified.
[0054] Traditional testing methods mainly rely on manual operations, that is, by recording the audio of the entire interaction process and then manually measuring the response duration of the device to the user's wake-up statement using audio reading software. However, traditional testing methods have the following deficiencies:
[0055] First, manual testing is inefficient, especially in cases where a large number of repeated tests are required, consuming a lot of manpower and time; second, manual measurement is prone to introducing human errors, resulting in inaccurate test results; third, the testing cycle of traditional testing methods is long, which is likely to affect the product's market launch time and increase development costs.
[0056] The method for detecting the response duration of voice interaction provided in this application aims to solve the above technical problems in the prior art. This method converts the audio into two-dimensional digital audio, uses a sliding window of a preset size, and by continuously sliding multiple times, observes the change trend of the amplitude of the sampling points within the window; if there are consecutive preset number of slides in which the amplitude within the window is always less than or equal to the preset threshold, then it is considered that the starting or ending position in this consecutive preset number of slides is a stable time node, thereby achieving the accurate positioning of the start and end time nodes of different corpora in the digital audio, and the automatic calculation of performance indicators such as the delay of the interactive audio. This method simplifies the process of extracting the response duration, provides a more convenient operation method for users, not only improves the acquisition efficiency of the response duration, but also reduces human errors.
[0057] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. These several specific embodiments can be combined with each other, and concepts or processes that are the same or similar may not be repeated in some embodiments. The embodiments of this application will be described below in conjunction with the accompanying drawings.
[0058] Figure 2 Flow schematic of the method for detecting the response duration of voice interaction provided in the embodiments of this applicationFigure 1 。The execution entity of this embodiment can be, for example, an audio processing platform. As Figure 2 shown, the method for detecting the voice interaction response duration provided in this embodiment includes:
[0059] S201: Obtain the audio waveform diagram of the voice interaction between the user and the intelligent device, and create a window with a preset duration on the audio waveform diagram.
[0060] Among them, the audio waveform diagram is used to indicate a graph that visually represents the change of the audio signal over time, and the window is used to obtain the audio signal with a preset duration. The preset duration can be, for example, 20 ms.
[0061] It can be understood that the audio waveform diagram is a two-dimensional digital waveform diagram generated based on the voice audio of the interaction between the user and the intelligent device; the execution entity of this step is the audio processing platform, which can automatically record the interaction audio between the user and the intelligent device, and use the audio processing module set in the audio processing platform to analyze and process the real-time obtained interaction audio; the window refers to a time period applied on the audio waveform diagram, and the size of this window depends on the preset duration, and the preset duration can be set independently according to different analysis requirements or application scenarios.
[0062] The audio processing platform obtains the audio waveform diagram generated based on the voice interaction between the user and the intelligent device, and creates a window with a corresponding time length on this audio waveform diagram according to the currently set preset duration.
[0063] In the embodiments of the present application, the audio processing platform can be set in the intelligent device, or can be set in other intelligent devices in the same Internet of Things as this intelligent device. The present application does not impose special restrictions on the setting of the audio processing platform.
[0064] Exemplarily, the audio processing platform obtains the interaction audio of the user interacting with the intelligent voice device, generates a corresponding audio waveform diagram based on this interaction audio, and creates a window with a corresponding time length of 20 ms on this audio waveform diagram according to the currently set preset duration of 20 ms.
[0065] In the embodiments of the present application, when the user interacts with the intelligent device, the audio processing platform obtains at least one piece of the interaction audio between the user and the intelligent device. During the process of converting the audio waveform diagram, the audio processing platform generates a corresponding audio waveform diagram for any one piece of the interaction audio. That is to say, any one piece of the interaction audio has a unique corresponding audio waveform diagram. Therefore, when determining the voice interaction response duration of the intelligent device, it is also calculated for any one audio waveform diagram.
[0066] S202: Control the window to slide from left to right with a preset sliding step, and detect the first amplitude of each audio sampling point within the window after each slide. Determine the audio end moment of the user according to the first amplitudes of the audio sampling points within the window.
[0067] Among them, the sliding step is used to indicate the size of each slide of the window, the first amplitude is used to indicate the decibel size of the audio sampling points in the user voice data within a preset duration window, and the audio end moment is used to indicate the last sampling point of the audio data stream corresponding to the user voice audio.
[0068] It can be understood that in digital audio processing, the time interval or resolution represented by each sampling point; it is determined by the sampling rate, which refers to the number of times of sampling the continuous audio signal per second. Therefore, the higher the sampling rate, the shorter the time interval represented by each sampling point, and thus the details of the audio signal can be captured and reproduced more precisely; the audio waveform diagram consists of a time horizontal axis and a vertical axis representing the amplitude (such as decibels). The time horizontal axis represents the time process of the audio signal, and each point or scale on the horizontal axis corresponds to a specific time point. The interval between adjacent time points can be determined by the sampling rate. The higher the sampling rate, the higher the resolution on the time axis and the more audio details can be represented. That is to say, each point on the time horizontal axis corresponds to each audio sampling point; in addition, there is an association relationship between the preset sliding step and the audio sampling points. The preset sliding step can be, for example, one audio sampling point or multiple audio sampling points. Exemplarily, the common CD audio sampling rate is 44.1 kHz, which means there are 44,100 sampling points per second. Therefore, the time interval represented by each sampling point is approximately 22.7 μs; in the two-dimensional digital waveform diagram, that is, the audio waveform diagram, the sampling points determine the resolution on the time axis, that is, the time interval between two adjacent minimum units on the waveform diagram. This application does not impose special restrictions on the preset sliding step.
[0069] During the processing of the audio signal, control the window with a preset duration to slide from left to right on the audio waveform diagram with a preset sliding step, and detect the first amplitude of each audio sampling point within the window after each slide. Determine the audio end moment of the user according to the first amplitudes of the audio sampling points within the window.
[0070] In some embodiments, based on the first amplitudes of the audio sampling points obtained within the current window, the envelope of the audio signal within the current window is calculated, and its changing trend is observed; if the envelope remains at a low level within the current window, it indicates the end of the user's audio. At this time, from the audio sampling points within the current window, the earliest occurring moment is screened out and determined as the end moment of the user's audio; if the envelope does not remain at a low level within the current window, it indicates that the user's audio has not ended. At this time, the window of the preset duration is continued to slide, and the envelope of the audio signal within the current window is recalculated to determine whether the user's audio within the current window has ended.
[0071] It can be understood that the envelope is a feature of the audio signal, which describes the overall contour or shape of the signal amplitude changing with time. That is to say, the envelope can be regarded as the "outer shell" of the signal amplitude, which smoothly connects the peaks of the signal waveform, thus providing a macroscopic view of the signal amplitude change.
[0072] S203: Control the window to slide from right to left at a preset sliding step length, and detect the second amplitudes of the audio sampling points within the window after each slide. Determine the start moment of the audio of the intelligent device according to the second amplitudes of the audio sampling points within the window.
[0073] Among them, the second amplitude is used to indicate the decibel size of the audio sampling points in the voice data of the intelligent device within the preset duration window, and the start moment of the audio is used to indicate the first sampling point of the audio data stream corresponding to the voice audio of the intelligent device.
[0074] During the processing of the audio signal, control the window of the preset duration to slide from right to left on the audio waveform diagram at a preset sliding step length, and detect the second amplitudes of the audio sampling points within the window after each slide. Determine the end moment of the user's audio according to the second amplitudes of the audio sampling points within the window.
[0075] S204: Determine the difference between the start moment of the audio of the intelligent device and the end moment of the user's audio as the voice interaction response duration of the intelligent device.
[0076] Among them, the voice interaction response duration is used to indicate the time length elapsed from when the user issues a voice command or question until the voice device or system makes a response.
[0077] According to the currently determined start moment of the audio of the intelligent device and the end moment of the user's audio, calculate the difference between the two, and determine this difference as the voice interaction response duration of the intelligent device.
[0078] Exemplarily, if the end time of the user's audio is 1.1 s and the start time of the audio of the intelligent device is 1.7 s, calculate the time difference between the two, and the result is 0.6 s. Therefore, the voice interaction response duration of the intelligent device is 0.6 s.
[0079] The detection and control method for the voice interaction response duration provided in this embodiment obtains the audio waveform diagram of the voice interaction between the user and the intelligent device, creates a window with a preset duration on the audio waveform diagram, controls the window to slide from left to right at a preset sliding step length, and detects the first amplitudes of each audio sampling point in the window after each slide. Determine the end time of the user's audio according to the first amplitudes of each audio sampling point in the window; then control the window to slide from right to left at a preset sliding step length, and detect the second amplitudes of each audio sampling point in the window after each slide. Determine the start time of the audio of the intelligent device according to the second amplitudes of each audio sampling point in the window; calculate the difference between the start time of the audio of the intelligent device and the end time of the user's audio, so as to obtain the voice interaction response duration of the intelligent device. This method simplifies the extraction process of the response duration, provides a more convenient operation method for users, not only improves the acquisition efficiency of the response duration, but also reduces human errors.
[0080] Figure 3 It is a schematic flow of the detection method for the voice interaction response duration provided in the embodiment of the present application Figure 2 As Figure 3 shown, on the basis of the Figure 2 embodiment, the detection method for the voice interaction response duration is described in detail. The detection method for the voice interaction response duration shown in this embodiment includes:
[0081] S301: Obtain the audio waveform diagram of the voice interaction between the user and the intelligent device, and create a window with a preset duration.
[0082] The audio processing platform obtains the audio waveform diagram generated based on the voice interaction between the user and the intelligent device, and creates a window with a corresponding time length according to the currently set preset duration.
[0083] In some embodiments, the acquisition of the audio waveform diagram of the voice interaction between the user and the intelligent device specifically includes: starting audio recording, playing the voice interaction corpus of the user, and after the intelligent device responds to the voice interaction corpus of the user, obtaining the recorded audio file of the voice interaction between the user and the intelligent device; loading the audio file to generate an audio waveform diagram.
[0084] It can be understood that the voice interaction corpus can be the voice data sent by the user in real time or the voice data set in advance. The present application does not make special restrictions on the voice interaction corpus.
[0085] Exemplarily, the audio processing platform starts audio recording and controls other intelligent devices to play the user's voice interaction corpus. After the intelligent device responds to the user's voice interaction corpus, it continues to record until the response voice of the intelligent device finishes playing, so as to obtain an audio file of the user's voice interaction with the intelligent device; based on this audio file, a corresponding audio waveform diagram is generated.
[0086] S302: Control the window to slide from left to right at a preset sliding step length, and detect the first amplitude of each audio sampling point in the window after each slide.
[0087] S303: If the first amplitude of each audio sampling point in the window is less than or equal to a preset threshold, determine the time start point of the window as the end time of the user's audio.
[0088] Among them, the preset threshold can be, for example, 30dB.
[0089] During the processing of the audio signal, control a window with a preset duration to slide from left to right on the audio waveform diagram at a preset sliding step length, and detect the first amplitude of each audio sampling point in the window after each slide; respectively judge the relationship between the first amplitude of each audio sampling point in the window and the preset threshold. If the first amplitude of each audio sampling point in the window is less than or equal to the preset threshold, it indicates that the user's audio ends. At this time, obtain the time start point of the current window and determine this time start point as the end time of the user's audio.
[0090] Exemplarily, if the audio sampling points and the corresponding first amplitudes obtained within the current 20ms window include: "audio sampling point 1 - 10dB, audio sampling point 2 - 11dB, audio sampling point 3 - 10dB, audio sampling point 4 - 12dB, audio sampling point 5 - 11dB", respectively judge the size relationship between the first amplitude corresponding to each audio sampling point obtained currently and the preset threshold of 30dB. It can be determined that the first amplitude corresponding to each audio sampling point obtained within the current window is less than the preset threshold of 30dB, indicating that the user's voice audio ends. At this time, the time start point corresponding to this window is 1.1s, so the end time of the user's audio can be determined as 1.1s.
[0091] In some embodiments, if in a continuous preset number of slides, the first amplitude of each audio sampling point in the window after each slide is less than or equal to the preset threshold, determine the time start point of the window in the continuous preset number of slides as the end time of the user's audio.
[0092] Among them, the preset number of times can be, for example, 5 times.
[0093] It can be understood that, in order to accurately locate the end moment of the user's audio, a window with a preset duration is slid continuously for multiple times, and each audio sampling point within the window obtained after multiple slides is judged to determine whether the user's interactive audio within the corresponding window has ended, so as to determine the end moment of the user's audio; the strategy of multiple slides and continuous judgments helps to filter out the interference of short-term silence or background noise, thereby improving the accuracy of positioning the voice interaction response duration and enabling the audio processing module to more accurately determine the end moment of the user's audio.
[0094] It should be noted that in the embodiments of the present application, the length of the window used is consistent with the preset duration, and the width of the window is consistent with the preset threshold.
[0095] S304: Control the window to slide from right to left according to a preset sliding step length, and detect the second amplitude of each audio sampling point within the window after each slide.
[0096] S305: If the second amplitude of each audio sampling point within the window is less than or equal to the preset threshold, determine the time end point of the window as the audio start moment of the intelligent device.
[0097] During the processing of the audio signal, control a window with a preset duration to slide from right to left on the audio waveform diagram according to a preset sliding step length, and detect the second amplitude of each audio sampling point within the window after each slide; respectively judge the relationship between the second amplitude of each audio sampling point within the window and the preset threshold. If the second amplitude of each audio sampling point within the window is less than or equal to the preset threshold, it indicates that the audio of the intelligent device has not been played within the current window. At this time, obtain the time end point of the current window and determine this time end point as the audio start moment of the intelligent device.
[0098] Exemplarily, if the audio sampling points and the corresponding second amplitudes obtained within the current 20ms window include: "audio sampling point 6 - 15dB, audio sampling point 7 - 16dB, audio sampling point 8 - 15dB, audio sampling point 9 - 17dB, audio sampling point 10 - 16dB", respectively judge the magnitude relationship between the second amplitude corresponding to each audio sampling point obtained currently and the preset threshold of 30dB. It can be determined that the second amplitude corresponding to each audio sampling point obtained within the current window is less than the preset threshold of 30dB, which indicates that the audio of the intelligent device has not been played within the current window. At this time, the time end point corresponding to this window is 1.7s, so that the end moment of the user's audio can be determined as 1.7s.
[0099] In some embodiments, determining the audio start moment of the intelligent device according to the second amplitude of each audio sampling point within the window includes:
[0100] If, in a continuous preset number of slides, the first amplitude of each audio sample point within the window is less than or equal to a preset threshold after each slide, then determine the end time of the window during the continuous preset number of slides as the start time of the audio of the intelligent device.
[0101] In some embodiments, during the process of a user's voice interaction with an intelligent device, since the interaction may occur in various different environments, different preset thresholds corresponding to different environments are preset accordingly; in addition, due to differences in the decibel levels of the interactive audio between the user and the intelligent device obtained according to the different positions of the audio processing platform, during the process of determining the end time of the user's audio and the start time of the intelligent device's audio, the preset thresholds adopted have different specific values; furthermore, the dynamic adjustment of the preset thresholds can also be optimized based on historical data and machine learning algorithms. The audio processing platform can gradually adjust the threshold settings by continuously learning the user's usage habits and environmental characteristics to make them more personalized and intelligent. For example, in an environment where the user often uses the system, the system can analyze past interaction data to predict and set the most appropriate threshold to improve the efficiency of interaction and the user experience.
[0102] S306: Determine the difference between the start time of the audio of the intelligent device and the end time of the user's audio as the voice interaction response duration of the intelligent device.
[0103] Step S306 is similar to the above-mentioned step S204 and will not be elaborated here.
[0104] The method for detecting the voice interaction response duration provided in this embodiment obtains the audio waveform diagram of the voice interaction between the user and the intelligent device, and creates a window with a preset duration on the audio waveform diagram; controls the window to slide from left to right on the audio waveform diagram according to a preset sliding step length, and detects the first amplitude of each audio sample point within the window after each slide. If the first amplitude of each audio sample point within the window is less than or equal to a preset threshold, then determine the start time of the window as the end time of the user's audio; then control the window to slide from right to left on the audio waveform diagram according to a preset sliding step length, and detect the second amplitude of each audio sample point within the window after each slide. If the second amplitude of each audio sample point within the window is less than or equal to a preset threshold, then determine the end time of the window as the start time of the audio of the intelligent device, and determine the difference between the start time of the audio of the intelligent device and the end time of the user's audio as the voice interaction response duration of the intelligent device. This method realizes the automatic acquisition of the response duration, simplifies the extraction process of the response duration, provides a more convenient operation method for users, not only improves the accuracy of obtaining the response duration, but also improves the acquisition efficiency of the response duration, and avoids the human error brought by the traditional acquisition method.
[0105] Figure 4 This is a schematic structural diagram of the detection device for the voice interaction response duration provided by this application. As Figure 4 shown, this application provides a detection device for the voice interaction response duration. The detection device 400 for the voice interaction response duration includes:
[0106] An acquisition module 401, configured to acquire an audio waveform diagram of the voice interaction between a user and an intelligent device.
[0107] A processing module 402, configured to create a window with a preset duration on the audio waveform diagram.
[0108] The processing module 402 is further configured to control the window to slide from left to right at a preset sliding step length, and detect the first amplitude of each audio sampling point within the window after each slide, and determine the audio end moment of the user according to the first amplitude of each audio sampling point.
[0109] The processing module 402 is further configured to control the window to slide from right to left at a preset sliding step length, and detect the second amplitude of each audio sampling point within the window after each slide, and determine the audio start moment of the intelligent device according to the second amplitude of each audio sampling point.
[0110] The processing module 402 is further configured to determine the difference between the audio start moment of the intelligent device and the audio end moment of the user as the voice interaction response duration of the intelligent device.
[0111] Optionally, if the first amplitude of each audio sampling point within the window is less than or equal to the preset threshold, the processing module 402 is further configured to determine the time start point of the window as the audio end moment of the user.
[0112] Optionally, if in a continuous preset number of slides, the first amplitude of each audio sampling point within the window after each slide is less than or equal to the preset threshold, the processing module 402 is further configured to determine the time start point of the window in the continuous preset number of slides as the audio end moment of the user.
[0113] Optionally, if the second amplitude of each audio sampling point within the window is less than or equal to the preset threshold, the processing module 402 is further configured to determine the time end point of the window as the audio start moment of the intelligent device.
[0114] Optionally, if in a continuous preset number of slides, the first amplitude of each audio sampling point within the window after each slide is less than or equal to the preset threshold, the processing module 402 is further configured to determine the time end point of the window in the continuous preset number of slides as the audio start moment of the intelligent device.
[0115] Optionally, the processing module 402 is further configured to start audio recording and play the user's voice interaction corpus, and after the intelligent device responds to the user's voice interaction corpus, obtain an audio file of the voice interaction between the user and the intelligent device recorded.
[0116] The processing module 402 is further configured to load the audio file to generate the audio waveform diagram.
[0117] Figure 5 It is a schematic structural diagram of an electronic device provided by this application. As Figure 5 shown, this application provides an electronic device, and this electronic device 500 includes: a receiver 501, a transmitter 502, a processor 503, and a memory 504.
[0118] The receiver 501 is configured to receive instructions and data;
[0119] The transmitter 502 is configured to send instructions and data;
[0120] The memory 504 is configured to store computer execution instructions;
[0121] The processor 503 is configured to execute the computer execution instructions stored in the memory 504 to implement each step performed by the method for detecting the voice interaction response duration in the foregoing embodiments. Specifically, reference may be made to the relevant descriptions in the embodiments of the method for detecting the voice interaction response duration described above.
[0122] Optionally, the foregoing memory 504 may be either independent or integrated with the processor 503.
[0123] When the memory 504 is independently provided, the electronic device further includes a bus for connecting the memory 504 and the processor 503.
[0124] The implementation principle and technical effects of the electronic device provided in this embodiment may be referred to the foregoing embodiments, and details are not described herein again.
[0125] This application embodiment also provides a computer-readable storage medium, in which a computer execution program is stored, and when the processor executes the computer execution program, the method described in any one of the foregoing embodiments is implemented.
[0126] This application embodiment also provides a computer program product, including a computer program, and when the computer program is executed by the processor, the method described in any one of the foregoing embodiments is implemented.
[0127] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0128] The integrated modules implemented in the form of software function modules can be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of this application.
[0129] It should be understood that the above processor can be a central processing unit (Central Processing Unit, abbreviated as CPU), and can also be other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application-specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by the execution of the hardware processor, or can be implemented by the combination of hardware and software modules in the processor. The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk or an optical disc, etc.
[0130] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0131] An exemplary storage medium is coupled to a processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master control device.
[0132] It should be noted that in this document, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element.
[0133] The serial numbers of the embodiments of the present application described above are only for description and do not represent the superiority or inferiority of the embodiments.
[0134] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0135] The above are only the preferred embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A method for detecting the response duration of voice interaction, characterized in that, The method includes: Obtain an audio waveform diagram of the voice interaction between the user and the intelligent device, and create a window with a preset duration on the audio waveform diagram; Control the window to slide from left to right at a preset sliding step length, and detect the first amplitude of each audio sampling point in the window after each slide. Determine the audio end time of the user according to the first amplitude of each audio sampling point; Control the window to slide from right to left at a preset sliding step length, and detect the second amplitude of each audio sampling point in the window after each slide. Determine the audio start time of the intelligent device according to the second amplitude of each audio sampling point; Determine the difference between the audio start time of the intelligent device and the audio end time of the user as the voice interaction response duration of the intelligent device.
2. The method according to claim 1, wherein The determining the audio end time of the user according to the first amplitude of each audio sampling point includes: If the first amplitude of each audio sampling point in the window is less than or equal to the preset threshold, determine the time start point of the window as the audio end time of the user.
3. The method according to claim 1, wherein The determining the audio end time of the user according to the first amplitude of each audio sampling point includes: If in a continuous preset number of slides, the first amplitude of each audio sampling point in the window after each slide is less than or equal to the preset threshold, determine the time start point of the window in the continuous preset number of slides as the audio end time of the user.
4. The method according to any one of claims 1 to 3, characterized in that, The determining the audio start time of the intelligent device according to the second amplitude of each audio sampling point includes: If the second amplitude of each audio sampling point in the window is less than or equal to the preset threshold, determine the time end point of the window as the audio start time of the intelligent device.
5. The method according to any one of claims 1 to 3, characterized in that, The determining the audio start time of the intelligent device according to the second amplitude of each audio sampling point includes: If in a continuous preset number of slides, the first amplitude of each audio sampling point in the window after each slide is less than or equal to the preset threshold, determine the time end point of the window in the continuous preset number of slides as the audio start time of the intelligent device.
6. The method according to any one of claims 1 to 3, characterized in that The obtaining the audio waveform diagram of the voice interaction between the user and the intelligent device includes: Start audio recording and play the voice interaction corpus of the user. After the intelligent device responds to the voice interaction corpus of the user, obtain the recorded audio file of the voice interaction between the user and the intelligent device; Load the audio file to generate the audio waveform diagram.
7. A detection device for the response duration of voice interaction, characterized in that, It includes: An obtaining module, configured to obtain an audio waveform diagram of the voice interaction between the user and the intelligent device; A processing module, configured to create a window with a preset duration on the audio waveform diagram; The processing module is further configured to control the window to slide from left to right at a preset sliding step length, and detect the first amplitude of each audio sampling point in the window after each slide. Determine the audio end time of the user according to the first amplitude of each audio sampling point; The processing module is further configured to control the window to slide from right to left at a preset sliding step length, and detect the second amplitude of each audio sampling point in the window after each slide. Determine the audio start time of the intelligent device according to the second amplitude of each audio sampling point; The processing module is further configured to determine the difference between the audio start time of the intelligent device and the audio end time of the user as the voice interaction response duration of the intelligent device.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when running, executes the method according to any one of claims 1 to 6.
9. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 6 through the computer program.
10. A program product, characterized in that, It includes a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 6.
Citation Information
Cited By
Real-time voice conversation method, device and equipment and storage medium
CN121565152A