Method, apparatus, and storage medium for determining voice endpoints

By performing frame-based and Fourier transforming of voice fragments, and using neural network models to judge voice endpoints, the problem of low detection accuracy of voice endpoints in complex noise environments is solved, and a higher detection accuracy is achieved.

CN114005436BActive Publication Date: 2025-06-17JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111436597.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2025-06-17
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

The accuracy of judging voice endpoints in complex noise environments is low.

Method used

By receiving voice clips, performing frame-based operations and performing fast Fourier transforms, inputting neural network models to output judgment scores and start point detection results to determine whether a voice endpoint was detected.

Benefits of technology

Improves the accuracy of detecting voice endpoints in complex noise environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114005436B_ABST
    Figure CN114005436B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, apparatus, and storage medium for determining a voice endpoint. The method includes: receiving a voice segment, performing a framing operation on the voice segment to obtain a plurality of voice segment frames; respectively performing fast Fourier transform processing on the plurality of voice segment frames to obtain a plurality of Fourier spectra corresponding to the plurality of voice segment frames; inputting the plurality of Fourier spectra into a neural network model to output a judgment score and a starting point detection result; and determining whether a voice endpoint is detected through the voice endpoint detection algorithm according to the judgment score and the starting point detection result. By adopting the above technical means, the problem of low accuracy in detecting voice endpoints in a complex noise environment in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of speech recognition, and in particular, to a method, apparatus, and storage medium for determining a speech endpoint. Background Art

[0002] With the development of science and technology, speech recognition is widely used in life. For example, scenarios such as waking up the voice assistant of a client and voice-controlling intelligent robots involve speech recognition. In speech recognition, the detection or determination of speech endpoints is particularly important. A speech endpoint includes the starting point and the ending point of speech. The detection of speech endpoints is to find the starting point of speech from a speech signal containing silence, noise, etc., and start speech recognition. When the ending point of speech is detected, speech recognition ends, thereby realizing multi-round speech interaction. In traditional technologies, the detection of speech endpoints often uses methods based on signal processing statistical metrics to judge speech endpoints, such as energy, zero-crossing rate, etc. Such methods are simple, but lack robustness, especially in complex acoustic scenarios with poor performance. In addition, traditional technologies also use machine models to judge speech endpoints. This method has relatively good robustness in judgment, but the accuracy of judging speech endpoints is low in complex noise environments such as music noise.

[0003] In the process of implementing the concept of the present disclosure, the inventors found that there are at least the following technical problems in the related technologies: the problem of low accuracy in judging speech endpoints in complex noise environments. Summary of the Invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, embodiments of the present disclosure provide a method, apparatus, and storage medium for determining a speech endpoint, so as to at least solve the problem of low accuracy in judging speech endpoints in complex noise environments in the prior art.

[0005] The object of the present disclosure is achieved by the following technical solutions:

[0006] In a first aspect, an embodiment of the present disclosure provides a method for determining a speech endpoint, including: receiving a speech segment, performing a framing operation on the speech segment to obtain a plurality of speech segment frames; respectively performing fast Fourier transform processing on the plurality of speech segment frames to obtain a plurality of Fourier spectra corresponding to the plurality of speech segment frames; inputting the plurality of Fourier spectra into a neural network model, outputting a judgment score and a starting point detection result; and determining whether a speech endpoint is detected according to the judgment score and the starting point detection result through the speech endpoint detection algorithm.

[0007] In an exemplary embodiment, determining whether a voice endpoint is detected by the voice endpoint detection algorithm according to the judgment score and the start point detection result includes: when the start point detection result is that the start point of the voice segment is not detected, determining whether the voice segment is emitted by a target object: if the voice segment is emitted by the target object, marking the voice start point as a true value point and receiving the next voice segment, where the voice start point marked as a true value point is used to indicate that the start point of the voice segment is detected, and the voice endpoint includes the start point of the voice segment; if the voice segment is not emitted by the target object, marking the voice start point as a false value point and receiving the next voice segment, where the voice start point marked as a false value point is used to indicate that the start point of the voice segment is not detected.

[0008] In an exemplary embodiment, determining whether a voice endpoint is detected by the voice endpoint detection algorithm according to the judgment score and the start point detection result includes: when the start point detection result is that the start point of the voice segment is detected, determining whether the voice segment is emitted by a target object: if the voice segment is emitted by the target object, incrementing the number of voiced speech segments by one and receiving the next voice segment, where the voice endpoint includes: the start point of the voice segment and the end point of the voice segment; if the voice segment is not emitted by the target object, incrementing the number of silent voice segments by one, and when the number of silent voice segments is greater than a first preset threshold and the number of voiced speech segments is greater than a second preset threshold, determining that the end point of the voice segment is detected and no longer receiving the next voice segment.

[0009] In an exemplary embodiment, after incrementing the number of silent voice segments by one when the voice segment is not emitted by the target object, the method further includes: when the number of silent voice segments is not greater than the first preset threshold, marking the voice start point as a false value point and receiving the next voice segment, where the voice start point marked as a false value point is used to indicate that the start point of the voice segment is not detected; or when the number of silent voice segments is greater than the first preset threshold but the number of voiced speech segments is not greater than the second preset threshold, marking the voice start point as a false value point and receiving the next voice segment.

[0010] In an exemplary embodiment, performing a frame splitting operation on the voice segment to obtain a plurality of voice segment frames includes: receiving a frame splitting instruction sent by a target object, and determining the frame length and frame shift corresponding to the frame splitting operation according to the frame splitting instruction; performing the frame splitting operation on the voice segment according to the frame length, the frame shift, and the size of the voice segment to obtain the plurality of voice segment frames.

[0011] In an exemplary embodiment, before performing the frame segmentation operation on the voice segment to obtain multiple voice segment frames, the method further includes: continuously receiving voice packets until the size of the received voice packets meets a preset size, and merging the received voice packets into the voice segment.

[0012] In an exemplary embodiment, the neural network model includes: a first preset number of convolutional layers, a second preset number of fully connected layers, and an output layer, where the output layer is composed of the fully connected layer and a softmax layer.

[0013] In an exemplary embodiment, it includes: obtaining ambient noise data and call data; performing phoneme alignment processing on the call data using a speech recognition tool to obtain phoneme alignment data; performing the frame segmentation operation on the ambient noise data and the phoneme alignment data to obtain multiple training data segment frames; respectively performing fast Fourier transform processing on the multiple training data segment frames to obtain multiple training data spectra corresponding to the multiple training data segment frames; performing annotation processing on the multiple training data spectra; and training the neural network model using the multiple training data spectra after the annotation processing.

[0014] In a second aspect, an embodiment of the present disclosure provides a device for determining a voice endpoint, including: a frame segmentation module, configured to receive a voice segment and perform a frame segmentation operation on the voice segment to obtain multiple voice segment frames; a processing module, configured to respectively perform fast Fourier transform processing on the multiple voice segment frames to obtain multiple Fourier spectra corresponding to the multiple voice segment frames; a model module, configured to input the multiple Fourier spectra into a neural network model and output a judgment score and a starting point detection result; and a determination module, configured to determine whether a voice endpoint is detected through the voice endpoint detection algorithm according to the judgment score and the starting point detection result.

[0015] In a third aspect, an embodiment of the present disclosure provides an electronic device. The above electronic device includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; the memory is used to store a computer program; when the processor executes the program stored on the memory, it implements the method for determining a voice endpoint or the method for image processing as described above.

[0016] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium. A computer program is stored on the above computer-readable storage medium, and when the computer program is executed by a processor, it implements the method for determining a voice endpoint or the method for image processing as described above.

[0017] The above technical solutions provided by the embodiments of the present disclosure have at least some or all of the following advantages compared with the prior art: receiving a voice segment, performing a framing operation on the voice segment to obtain a plurality of voice segment frames; respectively performing fast Fourier transform processing on the plurality of voice segment frames to obtain a plurality of Fourier spectra corresponding to the plurality of voice segment frames; inputting the plurality of Fourier spectra into a neural network model, outputting a judgment score and a start point detection result; and determining whether a voice endpoint is detected according to the judgment score and the start point detection result through the voice endpoint detection algorithm. Because the embodiments of the present disclosure sequentially perform a framing operation, fast Fourier transform processing, and input into a neural network model on the voice segment, and determine whether a voice endpoint is detected according to the finally obtained judgment score and the start point detection result through the voice endpoint detection algorithm, therefore, by adopting the above technical means, the problem of low accuracy in detecting voice endpoints in a complex noise environment in the prior art can be solved, and thus the accuracy of detecting voice endpoints can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0020] Figure 1 Schematically shows a hardware structure block diagram of a computer terminal for a method of determining a voice endpoint according to an embodiment of the present disclosure;

[0021] Figure 2 Schematically shows a flowchart of a method of determining a voice endpoint according to an embodiment of the present disclosure;

[0022] Figure 3 Schematically shows an internal network diagram of a neural network model according to an embodiment of the present disclosure;

[0023] Figure 4 Schematically shows a flow schematic diagram of a method of determining a voice endpoint according to an embodiment of the present disclosure;

[0024] Figure 5 Schematically shows a structure block diagram of a device for determining a voice endpoint according to an embodiment of the present disclosure;

[0025] Figure 6 Schematically shows a structure block diagram of an electronic device provided by an embodiment of the present disclosure. Detailed Implementation Manner

[0026] The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments may be combined with each other.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence.

[0028] The method embodiments provided by the embodiments of the present disclosure can be executed on a computer terminal or a similar computing device. Taking running on a computer terminal as an example, Figure 1 A schematic hardware structure block diagram of a computer terminal for a method of determining a voice endpoint according to an embodiment of the present disclosure is shown. As Figure 1 shown, the computer terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a microprocessor (abbreviation: MPU) or a programmable logic device (abbreviation: PLD), etc.), a processing device, and a memory 104 for storing data. Optionally, the above-mentioned computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned computer terminal. For example, the computer terminal may further include more or fewer components than those shown in Figure 1 the figure, or have the same functions as those shown in Figure 1 the figure or different configurations with more functions than those shown in Figure 1 the figure.

[0029] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method of determining a voice endpoint in the embodiments of the present disclosure. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided with respect to the processor 102, and these remote memories may be connected to the computer terminal through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0030] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of a computer terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0031] In an embodiment of the present disclosure, a method for determining a voice endpoint is provided. Figure 2 Schematically shown is a flowchart of a method for determining a voice endpoint according to an embodiment of the present disclosure, as Figure 2 shown. The process includes the following steps:

[0032] Step S202: Receive a voice segment and perform a framing operation on the voice segment to obtain a plurality of voice segment frames;

[0033] Step S204: Respectively perform fast Fourier transform processing on the plurality of voice segment frames to obtain a plurality of Fourier spectra corresponding to the plurality of voice segment frames;

[0034] Step S206: Input the plurality of Fourier spectra into a neural network model and output a judgment score and a start point detection result;

[0035] Step S208: According to the judgment score and the start point detection result, determine whether a voice endpoint is detected through the voice endpoint detection algorithm.

[0036] The embodiment of the present disclosure can be used in a dialogue scenario between an intelligent voice robot and a user. At this time, the target object is a person. The execution subject of the embodiment of the present disclosure is an intelligent voice robot.

[0037] Through the present disclosure, a voice segment is received, and frame division operation is performed on the voice segment to obtain a plurality of voice segment frames; fast Fourier transform processing is respectively performed on the plurality of voice segment frames to obtain a plurality of Fourier spectra corresponding to the plurality of voice segment frames; the plurality of Fourier spectra are input into a neural network model, and a judgment score and a starting point detection result are output; according to the judgment score and the starting point detection result, it is determined whether a voice endpoint is detected through the voice endpoint detection algorithm. Because in the embodiments of the present disclosure, frame division operation, fast Fourier transform processing, and input into the neural network model are sequentially performed on the voice segment, and according to the finally obtained judgment score and the starting point detection result, it is determined whether a voice endpoint is detected through the voice endpoint detection algorithm. Therefore, by adopting the above technical means, the problem of low accuracy in detecting voice endpoints in a complex noise environment in the prior art can be solved, and thus the accuracy of detecting voice endpoints is improved.

[0038] In step S208, determining whether a voice endpoint is detected through the voice endpoint detection algorithm according to the judgment score and the starting point detection result includes: in the case where the starting point detection result is that the starting point of the voice segment is not detected, determining whether the voice segment is emitted by a target object: in the case where the voice segment is emitted by the target object, marking the voice starting point as a true value point, and receiving the next voice segment, where the voice starting point marked as a true value point is used to indicate that the starting point of the voice segment is detected, and the voice endpoint includes the starting point of the voice segment; in the case where the voice segment is not emitted by the target object, marking the voice starting point as a false value point, and receiving the next voice segment, where the voice starting point marked as a false value point is used to indicate that the starting point of the voice segment is not detected.

[0039] For example, using 1 to represent true and 0 to represent false, marking the voice starting point as a false value point means that the label of the voice starting point is recorded as 0, and marking the voice starting point as a true value point means that the label of the voice starting point is recorded as 1.

[0040] The target object can be a person, an animal, or other objects that can emit sounds. If the target object is a person, then determining whether the speech segment is emitted by the target object is determining whether the speech segment is human voice. When the speech segment is received for the first time, the starting point detection result is defaulted to not detecting the starting point of the speech segment. In the case of receiving the speech segment for the first time, it is determined whether the speech segment is emitted by the target object. If the speech segment is emitted by the target object, then the speech starting point is marked as a true value point; if the speech segment is not emitted by the target object, then the speech starting point is marked as a false value point. When the speech segment is received for the first time and the speech starting point is marked as a false value point, it means that the speech segment received for the first time is environmental noise and does not contain the sound emitted by the target object. Then, when the speech segment is received for the second time, the starting point detection result is defaulted to not detecting the starting point of the speech segment. Among them, all sounds that are not emitted by the target object belong to environmental noise.

[0041] It is determined whether it is emitted by the target object according to the judgment score, and then the speech starting point is marked as a true value point or a false value point. Because the marking of the speech starting point of the speech segment received in the previous time as a true value point or a false value point is related to the starting point detection result, it can be said that the starting point detection result is related to the judgment score.

[0042] It can be understood that the speech segment inherits the state of the previous speech segment in sequence. For example, when the speech segment is received for the first time and the speech starting point is marked as a false value point, then when the speech segment is received for the second time, because when the speech segment was received for the first time, the speech starting point was marked as a false value point, the starting point detection result of the speech segment received for the second time is defaulted to not detecting the starting point of the speech segment. When the speech segment is received for the first time and the speech starting point is marked as a true value point, then when the speech segment is received for the second time, because when the speech segment was received for the first time, the speech starting point was marked as a true value point, the starting point detection result of the speech segment received for the second time is defaulted to detecting the starting point of the speech segment.

[0043] In addition, the state of the speech segment includes: detecting the speech starting point, detecting the speech ending point, the number of audible speech segments, and the number of silent speech segments. When receiving each speech segment and determining the speech endpoints, the state of the speech segment is updated.

[0044] In step S208, based on the judgment score and the start point detection result, it is determined whether a voice endpoint is detected through the voice endpoint detection algorithm, including: when the start point detection result is that the start point of the voice segment is detected, it is judged whether the voice segment is emitted by the target object; when the voice segment is emitted by the target object, the number of audible voice segments is incremented by one, and the next voice segment is received. The voice endpoint includes the start point and the end point of the voice segment; when the voice segment is not emitted by the target object, the number of silent voice segments is incremented by one. When the number of silent voice segments is greater than a first preset threshold and the number of audible voice segments is greater than a second preset threshold, it is determined that the end point of the voice segment is detected, and the next voice segment is no longer received.

[0045] The situation where the start point detection result is that the start point of the voice segment is detected must not be the first time a voice segment is received. Because the default start point detection result for the first time a voice segment is received is that the start point of the voice segment is not detected. The start point detection result being that the start point of the voice segment is detected indicates that when the previous voice segment was received during the current determination of the voice endpoint, the voice start point was marked as a true value point. When the start point detection result is that the start point of the voice segment is detected, if it is judged that the voice segment is emitted by the target object, the number of audible voice segments is incremented by one, and the next voice segment is received. When the start point detection result is that the start point of the voice segment is detected, if it is judged that the voice segment is not emitted by the target object, the number of silent voice segments is incremented by one. That is, when each voice segment is received and the voice endpoint is determined, the state of the voice segment is updated. When the number of silent voice segments is greater than a first preset threshold and the number of audible voice segments is greater than a second preset threshold, it is determined that the end point of the voice segment is detected, and the next voice segment is no longer received. When the end point of the voice segment is detected, it indicates that one voice transmission where the current voice segment is located is completed. The first preset threshold and the second preset threshold are set according to specific scenarios.

[0046] In step S208, when the voice segment is not emitted by the target object, after incrementing the number of silent voice segments by one, the method further includes: when the number of silent voice segments is not greater than the first preset threshold, marking the voice start point as a false value point and receiving the next voice segment, where the voice start point marked as a false value point is used to indicate that the start point of the voice segment is not detected; or when the number of silent voice segments is greater than the first preset threshold but the number of audible voice segments is not greater than the second preset threshold, marking the voice start point as a false value point and receiving the next voice segment. When the start point detection result is that the start point of the voice segment is detected and the voice segment is not emitted by the target object, if the number of silent voice segments is not greater than the first preset threshold, or the number of silent voice segments is greater than the first preset threshold but the number of audible voice segments is not greater than the second preset threshold, marking the voice end point as a false value point and receiving the next voice segment. Marking the voice start point as a false value point actually means that the model detects an audible voice segment followed by all non-human voice segments and the number of human voice segments is less than the threshold, so it is considered a false voice; thus, marking this voice start point as a false value point means it is a false voice start point. A voice robot is an intelligent dialogue robot with multiple functions such as automatic phone dialing, multi-round voice interaction, and intelligent intention judgment. It can be applied to various business scenarios such as automatic phone sales, product service promotion, information review, voice notification, and phone collection.

[0047] In the embodiments of the present disclosure, through the above technical means, a complete sentence of voice emitted by the target object can be fully received to avoid omission.

[0048] In step S202, performing a framing operation on the voice segment to obtain multiple voice segment frames, including: receiving a framing instruction sent by the target object, and determining the frame length and frame shift corresponding to the framing operation according to the framing instruction; performing the framing operation on the voice segment according to the frame length, the frame shift, and the size of the voice segment to obtain the multiple voice segment frames.

[0049] The framing operation in the embodiments of the present disclosure may be a method of windowing the voice segment. It should be noted that the framing operation in the embodiments of the present disclosure may be any method in the prior art voice framing operations. It should be noted that in addition to determining the frame length and frame shift corresponding to the framing operation according to the framing instruction, it may also be determined according to the default settings in the system. The system is the system corresponding to the voice end point determination device. For example, with a frame length of 25 ms and a frame shift of 10 ms, a 1-second audio can be divided into 97 frames.

[0050] Before performing step S202, that is, before performing a framing operation on the voice segment to obtain multiple voice segment frames, the method further includes: continuously receiving voice packets until the size of the received voice packets meets a preset size, and combining the received voice packets into the voice segment.

[0051] It should be noted that the embodiments of the present disclosure can directly receive a voice segment and determine the voice endpoints of the received voice segment. It can also continuously receive voice packets until the sizes of the received multiple voice packets meet a preset size, combine the received multiple voice packets into the voice segment, and then determine the voice endpoints of the combined voice segment.

[0052] In an alternative embodiment, the neural network model includes: a first preset number of convolutional layers, a second preset number of fully connected layers, and an output layer, where the output layer is composed of the fully connected layer and a softmax layer.

[0053] It should be noted that the last convolutional layer of the first preset number of convolutional layers is connected to the first fully connected layer of the second preset number of fully connected layers, and the last fully connected layer of the second preset number of fully connected layers is connected to the output layer.

[0054] Optionally, the neural network model includes: three convolutional layers, one fully connected layer, and one output layer, where the output layer is composed of the fully connected layer and a softmax layer.

[0055] In an alternative embodiment, obtain ambient noise data and call data; use a speech recognition tool to perform phoneme alignment processing on the call data to obtain phoneme alignment data; perform the framing operation on the ambient noise data and the phoneme alignment data to obtain multiple training data segment frames; perform the fast Fourier transform processing on the multiple training data segment frames respectively to obtain multiple training data spectra corresponding to the multiple training data segment frames; perform annotation processing on the multiple training data spectra; and train the neural network model using the multiple training data spectra after the annotation processing.

[0056] Performing annotation processing on the multiple training data spectra is to label the multiple training data spectra with labels, where the labels are the judgment scores corresponding to the multiple training data spectra. It should be noted that because the start point detection result is related to the judgment score, labeling the multiple training data spectra will actually label the judgment scores and start point detection results corresponding to the multiple training data spectra at the same time. That is to say, the labels for labeling the multiple training data spectra include the judgment scores and start point detection results corresponding to the multiple training data spectra.

[0057] The training method for the neural network model can be any existing training method in machine learning.

[0058] The environmental noise data can include noises such as cars, wind, factories, and rain. The call data can be telephone channel data with a sampling rate of 8k. The speech recognition tool can be Kaldi. The phoneme-level alignment result of the call data can be generated through the Kaldi tool. For example, with a frame length of 25ms and a frame shift of 10ms, 1 second of audio can be divided into 97 frames, that is, there are 97 results, such as "aabb silence silence cddde...". Then, the non-silent parts are extracted according to the alignment result as phoneme alignment data.

[0059] Through the above technical means, the neural network model in the embodiment of the present disclosure has a model size of 126kb and the number of model parameters is 26,000.

[0060] To better understand the above technical solution, the embodiment of the present disclosure also provides an optional embodiment for explaining the above technical solution.

[0061] Figure 3 Schematically shows the internal network diagram of a neural network model in the embodiment of the present disclosure, as Figure 3 shown:

[0062] The neural network model includes: three convolutional layers, one fully connected layer, and one output layer. Among them, the output layer is composed of the fully connected layer and the softmax layer. The last convolutional layer of the first preset number of convolutional layers is connected to the first fully connected layer of the second preset number of fully connected layers. The last fully connected layer of the second preset number of fully connected layers is connected to the output layer.

[0063] Figure 4 Schematically shows the flowchart of a method for determining a voice endpoint in the embodiment of the present disclosure, as Figure 4 shown:

[0064] S402: Receive a voice segment;

[0065] S404: Perform a framing operation on the voice segment to obtain multiple voice segment frames;

[0066] S406: Perform fast Fourier transform processing on the multiple voice segment frames respectively to obtain multiple Fourier spectra corresponding to the multiple voice segment frames;

[0067] S408: Input the multiple Fourier spectra into the neural network model, and output a judgment score and a start point detection result;

[0068] S410: Determine whether the start point detection result is the start point of the detected voice segment;

[0069] S412: If the start point detection result is that the start point of the speech segment is not detected, determine whether the speech segment is emitted by the target object;

[0070] S414: If the speech segment is emitted by the target object, mark the speech start point as a true value point and receive the next speech segment;

[0071] S416: If the speech segment is not emitted by the target object, mark the speech start point as a false value point and receive the next speech segment;

[0072] S418: If the start point detection result is that the start point of the speech segment is detected, determine whether the speech segment is emitted by the target object;

[0073] S420: If the speech segment is emitted by the target object, increment the number of audible speech segments by one and receive the next speech segment;

[0074] S422: If the speech segment is not emitted by the target object, increment the number of silent speech segments by one. When the number of silent speech segments is greater than the first preset threshold and the number of audible speech segments is greater than the second preset threshold, determine that the end point of the speech segment is detected and stop receiving the next speech segment;

[0075] S424: If the speech segment is not emitted by the target object, increment the number of silent speech segments by one. When the number of silent speech segments is not greater than the first preset threshold, or when the number of silent speech segments is greater than the first preset threshold but the number of audible speech segments is not greater than the second preset threshold, mark the speech end point as a false value point and receive the next speech segment.

[0076] Through the present disclosure, a voice segment is received, and the voice segment is framed to obtain a plurality of voice segment frames; the plurality of voice segment frames are respectively subjected to fast Fourier transform processing to obtain a plurality of Fourier spectra corresponding to the plurality of voice segment frames; the plurality of Fourier spectra are input into a neural network model, and a judgment score and a start point detection result are output; according to the judgment score and the start point detection result, it is determined whether a voice endpoint is detected through the voice endpoint detection algorithm. Because, in the embodiments of the present disclosure, the voice segment is sequentially framed, subjected to fast Fourier transform processing, and input into the neural network model, and according to the finally obtained judgment score and the start point detection result, it is determined whether a voice endpoint is detected through the voice endpoint detection algorithm. Therefore, by adopting the above technical means, the problem of low accuracy in detecting voice endpoints in a complex noise environment in the prior art can be solved, and thus the accuracy of detecting voice endpoints can be improved.

[0077] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk), and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, a component server, or a network device, etc.) to execute the methods of the various embodiments of the present disclosure.

[0078] In this embodiment, a device for determining a voice endpoint is further provided. The device for determining a voice endpoint is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" may be a combination of software and / or hardware that can implement a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0079] Figure 5 A structural block diagram of a device for determining a voice endpoint according to an optional embodiment of the present disclosure is schematically shown, as Figure 5 shown, the device includes:

[0080] A framing module 502, configured to receive a voice segment and perform a framing operation on the voice segment to obtain a plurality of voice segment frames;

[0081] A processing module 504, configured to perform fast Fourier transform processing on the multiple voice segment frames respectively to obtain multiple Fourier spectra corresponding to the multiple voice segment frames;

[0082] A model module 506, configured to input the multiple Fourier spectra into a neural network model and output a judgment score and a start point detection result;

[0083] A determination module 508, configured to determine whether a voice endpoint is detected by the voice endpoint detection algorithm according to the judgment score and the start point detection result.

[0084] The embodiments of the present disclosure can be used in a conversation scenario between an intelligent voice robot and a user. At this time, the target object is a person. The execution subject of the embodiments of the present disclosure is an intelligent voice robot.

[0085] Through the present disclosure, a voice segment is received, and the voice segment is framed to obtain multiple voice segment frames; fast Fourier transform processing is respectively performed on the multiple voice segment frames to obtain multiple Fourier spectra corresponding to the multiple voice segment frames; the multiple Fourier spectra are input into a neural network model to output a judgment score and a start point detection result; whether a voice endpoint is detected is determined by the voice endpoint detection algorithm according to the judgment score and the start point detection result. Because the embodiments of the present disclosure sequentially perform a framing operation, fast Fourier transform processing, and input into a neural network model on the voice segment, and determine whether a voice endpoint is detected by the voice endpoint detection algorithm according to the finally obtained judgment score and start point detection result, therefore, by adopting the above technical means, the problem of low accuracy in detecting voice endpoints in a complex noise environment in the prior art can be solved, and thus the accuracy of detecting voice endpoints can be improved.

[0086] Optionally, the determination module 508 is further configured to, when the start point detection result is that the start point of the voice segment is not detected, determine whether the voice segment is issued by the target object: when the voice segment is issued by the target object, mark the voice start point as a true value point and receive the next voice segment, where the voice start point marked as a true value point is used to indicate that the start point of the voice segment is detected, and the voice endpoint includes the start point of the voice segment; when the voice segment is not issued by the target object, mark the voice start point as a false value point and receive the next voice segment, where the voice start point marked as a false value point is used to indicate that the start point of the voice segment is not detected.

[0087] For example, using 1 to represent true and 0 to represent false, marking the voice start point as a false value point means that the label of the voice start point is recorded as 0, and marking the voice start point as a true value point means that the label of the voice start point is recorded as 1.

[0088] The target object can be a person, an animal, or other objects that can emit sounds. If the target object is a person, then determining whether the speech segment is emitted by the target object is determining whether the speech segment is human voice. When the speech segment is received for the first time, the starting point detection result is defaulted to not detecting the starting point of the speech segment. In the case of receiving the speech segment for the first time, it is determined whether the speech segment is emitted by the target object. If the speech segment is emitted by the target object, then the speech starting point is marked as a true value point. If the speech segment is not emitted by the target object, then the speech starting point is marked as a false value point. When the speech segment is received for the first time and the speech starting point is marked as a false value point, it means that the speech segment received for the first time is environmental noise and does not contain the sound emitted by the target object. Then, when the speech segment is received for the second time, the starting point detection result is defaulted to still not detecting the starting point of the speech segment. Among them, all sounds that are not emitted by the target object belong to environmental noise.

[0089] It can be understood that the speech segment inherits the state of the previous speech segment in sequence. For example, when the speech segment is received for the first time and the speech starting point is marked as a false value point, then when the speech segment is received for the second time, because when the speech segment was received for the first time, the speech starting point was marked as a false value point, so the starting point detection result of the speech segment received for the second time is defaulted to not detecting the starting point of the speech segment. When the speech segment is received for the first time and the speech starting point is marked as a true value point, then when the speech segment is received for the second time, because when the speech segment was received for the first time, the speech starting point was marked as a true value point, so the starting point detection result of the speech segment received for the second time is defaulted to detecting the starting point of the speech segment.

[0090] In addition, the state of the speech segment includes: detecting the speech starting point, detecting the speech ending point, the number of audible speech segments, and the number of silent speech segments. When determining the speech endpoints each time a speech segment is received, the state of the speech segment is updated.

[0091] Optionally, the determining module 508 is further configured to, when the starting point detection result is detecting the starting point of the speech segment, determine whether the speech segment is emitted by the target object: in the case where the speech segment is emitted by the target object, increment the number of audible speech segments by one, and receive the next speech segment. The speech endpoints include: the starting point of the speech segment and the ending point of the speech segment; in the case where the speech segment is not emitted by the target object, increment the number of silent speech segments by one. When the number of silent speech segments is greater than a first preset threshold and the number of audible speech segments is greater than a second preset threshold, it is determined that the ending point of the speech segment is detected, and the next speech segment is no longer received.

[0092] When the starting point detection result is that the starting point of the voice segment is detected, it must not be the first time to receive the voice segment. Because, by default, the starting point detection result of the first time to receive the voice segment is that the starting point of the voice segment is not detected. The fact that the starting point detection result is that the starting point of the voice segment is detected indicates that when the previous voice segment was received to determine the voice endpoint currently, the voice starting point was marked as a true value point. In the case where the starting point detection result is that the starting point of the voice segment is detected, if it is determined that the voice segment is emitted by the target object, the number of audible voice segments is incremented by one, and the next voice segment is received. In the case where the starting point detection result is that the starting point of the voice segment is detected, if it is determined that the voice segment is not emitted by the target object, the number of silent voice segments is incremented by one. That is to say, when receiving each voice segment and determining the voice endpoint, the state of the voice segment is updated. In the case where the number of silent voice segments is greater than a first preset threshold and the number of audible voice segments is greater than a second preset threshold, it is determined that the end point of the voice segment is detected, and the next voice segment is no longer received. When the end point of the voice segment is detected, it indicates that a voice transmission where the current voice segment is located is completed. The first preset threshold and the second preset threshold are set according to specific scenarios.

[0093] Optionally, the determining module 508 is further configured to, when the number of silent voice segments is not greater than the first preset threshold, mark the voice starting point as a false value point and receive the next voice segment, where marking the voice starting point as a false value point is used to indicate that the starting point of the voice segment is not detected; or when the number of silent voice segments is greater than the first preset threshold but the number of audible voice segments is not greater than the second preset threshold, mark the voice starting point as a false value point and receive the next voice segment.

[0094] When the starting point detection result is that the starting point of the voice segment is detected and the voice segment is not emitted by the target object, if the number of silent voice segments is not greater than the first preset threshold, or if the number of silent voice segments is greater than the first preset threshold but the number of audible voice segments is not greater than the second preset threshold, mark the end point of the voice as a false value point and receive the next voice segment. Mark the starting point of the voice as a false value point. In fact, this situation is that the model detects an audible voice segment, and then all are non-human voice segments, and the number of human voice segments is less than the threshold, so it is considered a false speaking voice; therefore, mark the starting point of this voice as a false value point, that is to say, this is a false voice starting point. A voice robot is an intelligent dialogue robot with multiple functions such as automatic phone dialing, multi-round voice interaction, and intelligent intention judgment. It can be applied to various business scenarios such as automatic phone sales, product service promotion, information review, voice notification, and phone collection.

[0095] In the embodiments of the present disclosure, through the above technical means, a complete sentence of voice emitted by the target object can be completely received, avoiding omission.

[0096] Optionally, the framing module 502 is further configured to receive a framing instruction sent by the target object, determine the frame length and frame shift corresponding to the framing operation according to the framing instruction; perform the framing operation on the voice segment according to the frame length, the frame shift, and the size of the voice segment to obtain the multiple voice segment frames.

[0097] The framing operation in the embodiments of the present disclosure may be a method of windowing the voice segment. It should be noted that the framing operation in the embodiments of the present disclosure may be any method in the prior art voice framing operations. It should be noted that in addition to determining the frame length and frame shift corresponding to the framing operation according to the framing instruction, it may also be to determine the frame length and frame shift corresponding to the framing operation according to the default settings in the system. The system is the system corresponding to the voice endpoint determination device. For example, with a frame length of 25 ms and a frame shift of 10 ms, a 1-second audio can be divided into 97 frames.

[0098] Optionally, the framing module 502 is further configured to continuously receive voice packets until the size of the received voice packets meets a preset size, and merge the received voice packets into the voice segment.

[0099] It should be noted that the embodiments of the present disclosure may directly receive a voice segment and determine the voice endpoint of the received voice segment, or may also continuously receive voice packets, wait until the sizes of the received multiple voice packets meet a preset size, merge the received multiple voice packets into the voice segment, and then determine the voice endpoint of the merged voice segment.

[0100] In an alternative embodiment, the neural network model includes: a first preset number of convolutional layers, a second preset number of fully connected layers, and an output layer, where the output layer is composed of the fully connected layer and a softmax layer.

[0101] It should be noted that the last convolutional layer of the first preset number of convolutional layers is connected to the first fully connected layer of the second preset number of fully connected layers, and the last fully connected layer of the second preset number of fully connected layers is connected to the output layer.

[0102] Optionally, the neural network model includes: three convolutional layers, one fully connected layer, and one output layer, where the output layer is composed of the fully connected layer and a softmax layer.

[0103] Optionally, the model module 506 is further configured to obtain ambient noise data and call data; perform phoneme alignment processing on the call data using a speech recognition tool to obtain phoneme-aligned data; perform the frame segmentation operation on the ambient noise data and the phoneme-aligned data to obtain a plurality of training data segment frames; perform the fast Fourier transform processing on the plurality of training data segment frames respectively to obtain a plurality of training data spectra corresponding to the plurality of training data segment frames; perform annotation processing on the plurality of training data spectra; and train the neural network model using the plurality of training data spectra after the annotation processing.

[0104] The ambient noise data may include noises such as cars, wind, factories, and rain. The call data may be telephone channel data with a sampling rate of 8k. The speech recognition tool may be Kaldi, and a phoneme-level alignment result of the call data is generated through the Kaldi tool. For example: with a frame length of 25ms and a frame shift of 10ms, 1 second of audio can be divided into 97 frames, that is, there are 97 results, such as "aabb silence silence cddde...". Then, the non-silent parts are extracted according to the alignment result as the phoneme-aligned data.

[0105] Performing annotation processing on the plurality of training data spectra is to label the plurality of training data spectra with labels, where the labels are the judgment scores corresponding to the plurality of training data spectra. It should be noted that the start point detection result is related to the judgment score. When labeling the plurality of training data spectra, the judgment scores and start point detection results corresponding to the plurality of training data spectra are actually labeled simultaneously. That is to say, the labels for labeling the plurality of training data spectra include the judgment scores and start point detection results corresponding to the plurality of training data spectra.

[0106] The training method for the neural network model can be any existing training method in machine learning.

[0107] In the embodiments of the present disclosure, through the above technical means, a neural network model with a model size of 126 kb and 26,000 model parameters.

[0108] It should be noted that the above-mentioned modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: the above-mentioned modules are all located in the same processor; or, the above-mentioned modules are respectively located in different processors in any combination form.

[0109] Embodiments of the present disclosure provide an electronic device.

[0110] Figure 6 A schematic block diagram of an electronic device provided by an embodiment of the present disclosure is shown.

[0111] Refer to Figure 6 As shown, the electronic device 600 provided by the embodiment of the present disclosure includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604. Among them, the processor 601, the communication interface 602, and the memory 603 complete mutual communication through the communication bus 604; the memory 603 is used to store a computer program; when the processor 601 executes the program stored on the memory, it implements the steps in any one of the above method embodiments.

[0112] Optionally, the above electronic device may further include a transmission device and an input / output device, wherein the input / output device is connected to the above processor.

[0113] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:

[0114] S1, receive a voice segment, and perform a framing operation on the voice segment to obtain a plurality of voice segment frames;

[0115] S2, respectively perform fast Fourier transform processing on the plurality of voice segment frames to obtain a plurality of Fourier spectra corresponding to the plurality of voice segment frames;

[0116] S3, input the plurality of Fourier spectra into a neural network model, and output a judgment score and a start point detection result;

[0117] S4, according to the judgment score and the start point detection result, determine whether a voice endpoint is detected through the voice endpoint detection algorithm.

[0118] Embodiments of the present disclosure also provide a computer-readable storage medium. A computer program is stored on the above computer-readable storage medium, and when the computer program is executed by a processor, it implements the steps in any one of the above method embodiments.

[0119] Optionally, in this embodiment, the above storage medium may be configured to store a computer program for performing the following steps:

[0120] S1. Receive a voice segment, and perform a framing operation on the voice segment to obtain a plurality of voice segment frames;

[0121] S2. Perform fast Fourier transform processing on the plurality of voice segment frames respectively to obtain a plurality of Fourier spectra corresponding to the plurality of voice segment frames;

[0122] S3. Input the plurality of Fourier spectra into a neural network model, and output a judgment score and a starting point detection result;

[0123] S4. According to the judgment score and the starting point detection result, determine whether a voice endpoint is detected through the voice endpoint detection algorithm.

[0124] The computer-readable storage medium may be included in the device / apparatus described in the above embodiment; or it may exist separately without being assembled into the device / apparatus. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.

[0125] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, device, or device.

[0126] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiment and optional implementation manners, and will not be elaborated herein.

[0127] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present disclosure can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present disclosure is not limited to any specific combination of hardware and software.

[0128] The above are only the preferred embodiments of the present disclosure and are not used to limit the present disclosure. For those skilled in the art, the present disclosure can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present disclosure should be included in the protection scope of the present disclosure.

Claims

1. A method for determining the voice endpoint, characterized in that, including: receiving a voice segment and performing a framing operation on the voice segment to obtain a plurality of voice segment frames; respectively performing fast Fourier transform processing on the plurality of voice segment frames to obtain a plurality of Fourier spectra corresponding to the plurality of voice segment frames; inputting the plurality of Fourier spectra into a pre-trained neural network model to output a judgment score and a start point detection result; determining whether a voice endpoint is detected through a voice endpoint detection algorithm according to the judgment score and the start point detection result, and judging whether the voice segment is emitted by a target object according to the judgment score; wherein the start point detection result includes not detecting the start point of the voice segment and detecting the start point of the voice segment; wherein determining whether a voice endpoint is detected through the voice endpoint detection algorithm according to the judgment score and the start point detection result includes: when the start point detection result is not detecting the start point of the voice segment, judging whether the voice segment is emitted by a target object: when the voice segment is emitted by the target object, marking the voice start point as a true value point and receiving the next voice segment, where the voice start point marked as a true value point is used to indicate that the start point of the voice segment is detected, and the voice endpoint includes the start point of the voice segment; when the voice segment is not emitted by the target object, marking the voice start point as a false value point and receiving the next voice segment, where the voice start point marked as a false value point is used to indicate that the start point of the voice segment is not detected; when the start point detection result is detecting the start point of the voice segment, judging whether the voice segment is emitted by a target object: when the voice segment is emitted by the target object, incrementing the number of voiced speech segments by one and receiving the next voice segment, and the voice endpoint includes: the start point of the voice segment and the end point of the voice segment; when the voice segment is not emitted by the target object, incrementing the number of silent voice segments by one, and when the number of silent voice segments is greater than a first preset threshold and the number of voiced speech segments is greater than a second preset threshold, determining that the end point of the voice segment is detected and no longer receiving the next voice segment.

2. The method according to claim 1, characterized in that, after incrementing the number of silent voice segments by one when the voice segment is not emitted by the target object, the method further includes: when the number of silent voice segments is not greater than the first preset threshold, marking the voice start point as a false value point and receiving the next voice segment, where the voice start point marked as a false value point is used to indicate that the start point of the voice segment is not detected; or when the number of silent voice segments is greater than the first preset threshold but the number of voiced speech segments is not greater than the second preset threshold, marking the voice start point as a false value point and receiving the next voice segment.

3. The method according to claim 1, characterized in that, performing the framing operation on the voice segment to obtain a plurality of voice segment frames includes: Receive the framed instruction sent by the target object, and determine the frame length and frame shift corresponding to the framing operation according to the framed instruction; Perform the framing operation on the voice segment according to the frame length, the frame shift, and the size of the voice segment to obtain the multiple voice segment frames.

4. The method according to claim 1, characterized in that, Before performing the framing operation on the voice segment to obtain multiple voice segment frames, the method further includes: Continuously receive voice packets until the size of the received voice packet meets a preset size, and merge the received voice packets into the voice segment.

5. The method according to claim 1, characterized in that, The neural network model includes: a first preset number of convolutional layers, a second preset number of fully connected layers, and an output layer, where the output layer is composed of the fully connected layer and a softmax layer.

6. The method according to claim 1, characterized in that, Including: Obtain environmental noise data and call data; Use a speech recognition tool to perform phoneme alignment processing on the call data to obtain phoneme alignment data; Perform the framing operation on the environmental noise data and the phoneme alignment data to obtain multiple training data segment frames; Perform fast Fourier transform processing on the multiple training data segment frames respectively to obtain multiple training data spectra corresponding to the multiple training data segment frames; Perform annotation processing on the multiple training data spectra; Train the neural network model using the multiple training data spectra after the annotation processing.

7. A device for determining the voice endpoint, characterized in that, Including: A framing module, configured to receive a voice segment and perform a framing operation on the voice segment to obtain multiple voice segment frames; A processing module, configured to perform fast Fourier transform processing on the multiple voice segment frames respectively to obtain multiple Fourier spectra corresponding to the multiple voice segment frames; A model module, configured to input the multiple Fourier spectra into a neural network model and output a judgment score and a start point detection result; A determination module, configured to determine whether a voice endpoint is detected through a voice activity detection algorithm according to the judgment score and the start point detection result, where the start point detection result includes not detecting the start point of the voice segment and detecting the start point of the voice segment, judge whether the voice segment is sent by the target object according to the judgment score, where determining whether a voice endpoint is detected through the voice activity detection algorithm according to the judgment score and the start point detection result includes: In the case where the start point detection result is not detecting the start point of the voice segment, judge whether the voice segment is sent by the target object: In the case where the voice segment is sent by the target object, mark the voice start point as a true value point, and receive the next voice segment, where the voice start point marked as a true value point is used to indicate that the start point of the voice segment is detected, and the voice endpoint includes the start point of the voice segment; In the case where the voice segment is not sent by the target object, mark the voice start point as a false value point, and receive the next voice segment, where the voice start point marked as a false value point is used to indicate that the start point of the voice segment is not detected, When the detection result of the starting point is that the starting point of the speech segment is detected, determine whether the speech segment is emitted by the target object: When the speech segment is emitted by the target object, increment the number of audible speech segments by one and receive the next speech segment. The speech endpoints include: the starting point of the speech segment and the ending point of the speech segment; When the speech segment is not emitted by the target object, increment the number of silent speech segments by one. When the number of silent speech segments is greater than a first preset threshold and the number of audible speech segments is greater than a second preset threshold, determine that the ending point of the speech segment is detected and stop receiving the next speech segment.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method for quickly identifying gender and device thereof, and method for generating algorithm model for identifying gender

    CN111105803A

  • Voice detection method and device, computer equipment and storage medium

    CN112802498A

  • Apparatus and computer program stored in computer-readable medium for improving of voice recognition performance

    KR1020170100705A

  • Speech endpointing based on voice profile

    US8843369B1