Voice endpoint detection method, device, equipment and computer-readable storage medium

By fusing the characteristic vectors of audio signals, video data and reflected wave signals, the problem of low accuracy of voice endpoint detection in noise and complex environments in the prior art is solved, and higher detection accuracy and noise immunity are achieved.

CN112634940BActive Publication Date: 2025-05-06PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011453437.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-11
Publication Date
2025-05-06
Estimated Expiration
2040-12-11

AI Technical Summary

Technical Problem

Existing voice endpoint detection technology is difficult to ensure detection accuracy in noise environments and complex environments, especially when the speaker's pronunciation is distorted, pronunciation speed and tone change, it is prone to recognition errors.

Method used

By obtaining audio signals, video data and reflected wave signals, the spectrum characteristic vector of the audio signal, video characteristic vector of the video data, and reflected wave vector of the reflected wave signal are determined, and fused into the target feature vector, and the pre-trained speech endpoint detection model is input to detect speech endpoints.

Benefits of technology

Through the fusion of multimodal data, this method significantly improves the noise immunity and detection accuracy of voice endpoint detection, especially in complex environments, which can effectively distinguish voice signals from non-voice signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112634940B_ABST
    Figure CN112634940B_ABST
Patent Text Reader

Abstract

The present application relates to artificial intelligence, and provides a method, device, equipment and computer-readable storage medium for detecting speech endpoints, the method comprising: obtaining an audio signal to be detected and video data and a reflected wave signal collected when collecting the audio signal; determining the spectrum data of the audio signal, and determining the spectrum feature vector of the audio signal according to the spectrum data; extracting the image area where the lips are located in the video data, and determining the video feature vector of the video data according to each image area; determining the phase difference between the reflected wave signal and the preset transmitted wave signal, and determining the reflected wave vector of the reflected wave signal according to the phase difference; fusing the spectrum feature vector, the video feature vector and the reflected wave vector to obtain a target feature vector; inputting the target feature vector into a pre-trained speech endpoint detection model to obtain multiple speech endpoints of the audio signal. The present application can improve the accuracy of speech endpoint detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a voice endpoint detection method, apparatus, device and computer-readable storage medium. Background Art

[0002] Voice endpoint detection (VAD), also known as voice activity detection, refers to the technology of locating the start and end points of speech from an audio signal, thereby distinguishing the speech part from the non-speech part of the audio signal. Studies have shown that in a noisy environment or when the speaker's pronunciation is distorted, the pronunciation speed and pitch change, the Lombard / Loud effect will occur, and the application of voice endpoint detection is prone to recognition errors. At present, researchers have also tried to extract sound features through machine learning or deep learning to perform voice endpoint detection. However, the background noise in audio signals in real life is more complex. For example, the sound features of others often interfere with the audio signal, and the detection accuracy is difficult to guarantee. Therefore, how to improve the accuracy of voice endpoint detection has become an urgent problem to be solved. Summary of the invention

[0003] The main purpose of this application is to provide a voice endpoint detection method, device, equipment and computer-readable storage medium, aiming to improve the accuracy of voice endpoint detection.

[0004] In a first aspect, the present application provides a voice endpoint detection method, comprising:

[0005] Acquire an audio signal to be detected and video data and a reflected wave signal collected when collecting the audio signal, wherein the reflected wave signal is collected by performing sound wave detection on the lips of the user;

[0006] Determining frequency spectrum data of the audio signal, and determining a frequency spectrum feature vector of the audio signal according to the frequency spectrum data;

[0007] Extracting the image region where the lips are located in the video data, and determining the video feature vector of the video data according to each of the image regions;

[0008] Determining a phase difference between the reflected wave signal and a preset transmitted wave signal, and determining a reflected wave vector of the reflected wave signal according to the phase difference;

[0009] fusing the frequency spectrum feature vector, the video feature vector and the reflected wave vector to obtain a target feature vector;

[0010] The target feature vector is input into a pre-trained speech endpoint detection model to obtain multiple speech endpoints of the audio signal.

[0011] In a second aspect, the present application further provides a speech endpoint detection device, the speech endpoint detection device comprising:

[0012] An acquisition module, used to acquire an audio signal to be detected and video data and a reflected wave signal collected when the audio signal is collected, wherein the reflected wave signal is collected by performing sound wave detection on the lips of the user;

[0013] A first determining module, configured to determine frequency spectrum data of the audio signal, and determine a frequency spectrum feature vector of the audio signal according to the frequency spectrum data;

[0014] A second determination module is used to extract the image area where the lips are located in the video data, and determine the video feature vector of the video data according to each of the image areas;

[0015] a third determining module, configured to determine a phase difference between the reflected wave signal and a preset transmitted wave signal, and determine a reflected wave vector of the reflected wave signal according to the phase difference;

[0016] A fusion module, used for fusing the spectrum feature vector, the video feature vector and the reflected wave vector to obtain a target feature vector;

[0017] The detection module is used to input the target feature vector into a pre-trained speech endpoint detection model to obtain multiple speech endpoints of the audio signal.

[0018] In a third aspect, the present application also provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the steps of the voice endpoint detection method as described above are implemented.

[0019] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the voice endpoint detection method as described above are implemented.

[0020] The present application provides a voice endpoint detection method, apparatus, device and computer-readable storage medium. The present application obtains an audio signal to be detected and video data and a reflected wave signal collected when collecting the audio signal, then determines the spectrum data of the audio signal, and determines the spectrum feature vector of the audio signal based on the spectrum data, and extracts the image area where the lips are located in the video data, and determines the video feature vector of the video data based on each image area, and determines the phase difference between the reflected wave signal and the preset transmitted wave signal, and determines the reflected wave vector of the reflected wave signal based on the phase difference, and then fuses the spectrum feature vector, the video feature vector and the reflected wave vector to obtain a target feature vector, and then inputs the target feature vector into a pre-trained voice endpoint detection model to obtain multiple voice endpoints of the audio signal. The present application assists in voice endpoint detection through three different modal data, namely the spectrum data of the audio signal and the synchronously collected video data and reflected wave signal, which greatly improves the noise resistance of voice endpoint detection and can effectively improve the detection accuracy of voice endpoint detection in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1 A schematic diagram of the steps of a voice endpoint detection method provided in an embodiment of the present application;

[0023] Figure 2 for Figure 1 A schematic diagram of the sub-step flow of the speech endpoint detection method in FIG.

[0024] Figure 3 A schematic diagram of a scenario for implementing the voice endpoint detection method provided in this embodiment;

[0025] Figure 4 A schematic block diagram of a speech endpoint detection device provided in an embodiment of the present application;

[0026] Figure 5 for Figure 4 A schematic block diagram of a submodule of a speech endpoint detection device in;

[0027] Figure 6 A schematic block diagram of the structure of a computer device provided in an embodiment of the present application.

[0028] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0030] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor do they have to be executed in the order described. For example, some operations / steps can also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions. In addition, although the functional modules are divided in the device schematic, in some cases, the module division can be different from that in the device schematic.

[0031] The embodiments of the present application provide a method, apparatus, device and computer-readable storage medium for voice endpoint detection. The voice endpoint detection method can be applied to a terminal device or a server, and the terminal device can be an electronic device such as a mobile phone, a tablet computer, a laptop computer, a desktop computer, a personal digital assistant and a wearable device; the server can be a single server or a server cluster composed of multiple servers. The following explanation is given by taking the application of the voice endpoint detection method to a server as an example.

[0032] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0033] Please refer to Figure 1 , Figure 1 A schematic flow chart of the steps of a voice endpoint detection method provided in an embodiment of the present application.

[0034] like Figure 1 As shown, the voice endpoint detection method includes steps S101 to S106.

[0035] Step S101: Acquire an audio signal to be detected and video data and a reflected wave signal collected when collecting the audio signal.

[0036] Among them, the reflected wave signal is collected by performing sound wave detection on the user's lips, and the video data is collected by performing video recording on the user's lips. In the research field of multimodal speech recognition, the accuracy of lip reading recognition is very low, but the movement of the lips can often accurately reflect whether a person is speaking, which has a good application prospect in the field of speech endpoint detection. Therefore, by jointly assisting the audio signal to be detected in speech endpoint detection by using the reflected wave signal collected by performing sound wave detection on the user's lips and the video data collected by performing video recording on the user's lips, the accuracy of speech endpoint detection can be greatly improved, and it can effectively cope with noise interference in various environments and of various intensities, and improve the detection accuracy of speech endpoints of audio signals recorded in complex environments.

[0037] In one embodiment, when a user triggers a recording instruction of an audio signal through a smart device, a recorder is turned on to record the audio signal of the surrounding environment; at the same time, a camera of the smart device is turned on to record video, the camera is, for example, a front camera or a rear camera of a smartphone, to record video data including the lips or face of the user; at the same time, sound waves are emitted through a speaker of the smart device, and the sound waves are, for example, high-frequency sound waves that are inaudible to human ears. The sound waves emitted by the speaker will be reflected back to the microphone by the lips, so that the movement of the speaker's lips when speaking can be identified through the change in the phase of the sound waves.

[0038] In one embodiment, the server initiates a data acquisition request to the smart device, requesting to acquire the audio signal to be detected and the video data and reflected wave signal collected when the audio signal is collected. After receiving the data acquisition request sent by the server, the smart device returns the audio signal and the corresponding video data and reflected wave signal to the server according to the data acquisition request.

[0039] In some embodiments, after the smart device collects the audio signal and the video data and reflected wave signal collected when collecting the audio signal, the collected audio signal, video data and reflected wave signal are stored in a cloud database, so that the server can obtain the audio signal, video data and reflected wave signal collected by the smart device through the cloud database.

[0040] It should be noted that in order to further ensure the privacy and security of the above-mentioned audio signals, video data, reflected wave signals and other related information collected synchronously, the above-mentioned audio signals, video data, reflected wave signals and other related information can also be stored in a node of a blockchain. The technical solution of the present application can also be applied to adding other data files stored on the blockchain. The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, point-to-point transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.

[0041] Step S102: Determine frequency spectrum data of the audio signal, and determine a frequency spectrum feature vector of the audio signal according to the frequency spectrum data.

[0042] By performing short-time Fourier transform (STFT) on the audio signal, the spectral data of the audio signal can be determined. By performing feature extraction on the spectral data of the audio signal, the spectral feature vector of the audio signal can be obtained. The spectral feature vector of the audio signal can assist in speech endpoint detection and improve the accuracy of speech endpoint detection.

[0043] In one embodiment, if Figure 2 As shown, step S102 includes: sub-steps S1021 to S1023.

[0044] Sub-step S1021: determining, according to the multiple first timestamps of the spectrum data, a feature vector corresponding to each of the first timestamps.

[0045] It should be noted that the spectrum data includes multiple first timestamps, each first timestamp corresponds to a segment of spectrum data, and according to each first timestamp corresponding to a segment of spectrum data, a feature vector of a segment of spectrum data corresponding to each first timestamp can be determined.

[0046] In one embodiment, based on multiple first timestamps of spectral data, a feature vector corresponding to each first timestamp is determined, including: determining multiple first timestamps of spectral data, and determining multiple frames of spectral data corresponding to each first timestamp; and determining a feature vector corresponding to each first timestamp based on characteristic parameters of the multiple frames of spectral data corresponding to each first timestamp.

[0047] It should be noted that each first timestamp corresponds to a continuous multi-frame spectrum data, and the corresponding relationship between the first timestamp and the multi-frame spectrum data can be flexibly set by the user. For example, the first timestamp is set to correspond to the multi-frame spectrum data within the range of 2 frames before and after the current moment, that is, the first timestamp at time t corresponds to the spectrum data within the range of [t-2, t+2]. In order to facilitate time alignment with subsequent video feature vectors and reflection wave vectors, the corresponding relationship between the timestamp of each modal data (such as the first timestamp) and the sub-modal data (such as multi-frame spectrum data) can be consistent. The feature vector is composed of the feature parameter data of a continuous multi-frame spectrum, and the feature vector is a one-dimensional feature vector.

[0048] Sub-step S1022: performing convolution pooling processing on the feature vectors corresponding to each of the first timestamps.

[0049] The feature vector corresponding to each first timestamp is input into a convolution layer and / or a pooling layer to perform convolution and pooling processing on the multiple feature vectors.

[0050] Sub-step S1023: concatenate the feature vectors corresponding to each of the first timestamps processed by convolution pooling to obtain a spectrum feature vector corresponding to the spectrum data.

[0051] By splicing multiple feature vectors processed by convolution pooling according to the time sequence of each first timestamp, the spectrum feature vector corresponding to the spectrum data can be accurately obtained, and the spectrum feature vector can assist in voice endpoint detection.

[0052] Exemplarily, a range of two frames [t-2, t+2] before and after the first timestamp of the current time t is selected as a segment of spectrum data corresponding to the first timestamp, and characteristic parameters of the segment of spectrum data are extracted to obtain a characteristic vector Xa. The characteristic vector Xa is convolved once, pooled, and convolved twice to obtain a sub-spectrum characteristic vector a of the first timestamp at time t. t =conv a 3(Pool a 2(conv a 1(X a ))), by concatenating the sub-spectrum feature vectors corresponding to each first timestamp of the spectrum data, a spectrum feature vector can be obtained.

[0053] Step S103: extracting the image region where the lips are located in the video data, and determining the video feature vector of the video data according to each of the image regions.

[0054] The target detection algorithm can locate the image area where the lips are located in the video data, and then extract the image area where the lips are located in multiple frames from the video data. According to the time information corresponding to the image area where the lips are located in each frame, the feature extraction is performed on each image area to determine the video feature vector of the video data. The video feature vector of the video data can assist in voice endpoint detection and improve the accuracy of voice endpoint detection.

[0055] In one embodiment, based on multiple third timestamps of the video data, the feature vector corresponding to each third timestamp is determined; the feature vector corresponding to each third timestamp is subjected to convolution pooling processing; the feature vectors corresponding to each third timestamp subjected to convolution pooling processing are spliced ​​to obtain a video feature vector of the video data. It should be noted that the above-mentioned convolution pooling processing includes convolution processing and / or pooling processing, and the video data includes multiple third timestamps, each third timestamp corresponds to a segment of video data, and a segment of video data may include multiple image areas, or may be empty (excluding image areas). By generating a feature vector corresponding to the third timestamp based on the multiple image areas corresponding to the third timestamp, and performing convolution pooling processing and sequential splicing on the feature vector corresponding to the third timestamp, the video feature vector of the video data can be accurately obtained, and the video feature vector can assist in voice endpoint detection.

[0056] In one embodiment, multiple third timestamps of video data are determined, and multiple image regions corresponding to each third timestamp are determined; and feature vectors corresponding to each third timestamp are determined according to feature parameters of the multiple image regions corresponding to each third timestamp. The feature parameters of the image region may include one or more of the number of horizontal axis pixels, the number of vertical axis pixels, the number of channels, and the frame rate. For example, the number of horizontal axis pixels and the number of vertical axis pixels may be 800*1000, the number of RGB channels is 3, and the frame rate is the frequency of continuous appearance of the image region in units of frames.

[0057] For example, the third timestamp is set to correspond to multiple image regions within the range of 2 frames before and after the current moment, that is, the third timestamp at time t corresponds to the image region within the range of [t-2, t+2], and the feature parameters of the image region within the range of [t-2, t+2] corresponding to the third timestamp are extracted to obtain the sub-video feature vector v t =conv v 3(conv v 2(conv v 1(X v ))), by concatenating the sub-video feature vectors corresponding to each third timestamp of the video data, a video feature vector can be obtained.

[0058] Step S104: determine a phase difference between the reflected wave signal and a preset transmitted wave signal, and determine a reflected wave vector of the reflected wave signal according to the phase difference.

[0059] The preset transmission wave signal of the sound wave emitted by the speaker and the reflected wave signal reflected by the lips received by the microphone are obtained. The phase difference between the reflected wave signal and the preset transmission wave signal can be determined through signal processing. The movement of the speaker's lips when speaking can be identified through this phase difference. The phase difference between the reflected wave signal and the preset transmission wave signal is extracted to determine the reflected wave vector of the reflected wave signal. The reflected wave vector of the reflected wave signal can assist in voice endpoint detection and improve the accuracy of voice endpoint detection.

[0060] In one embodiment, the limit of ordinary people's hearing is 17kHz, and most smart devices can emit sound waves at a maximum of 23kHz, so the preset transmission wave signal can be located in the range of 17-23kHz, and the phase difference between the reflected wave signal and the preset transmission wave signal is detected by an acoustic wave coupling detector.

[0061] For example, the preset transmission wave signal is Acos(2πft) and the reflection wave signal is R p =A p cos(2πft-θ p ), where A p is the reflected wave amplitude, θ p is the phase difference, f is the frequency, A and Ap can be consistent, but after reflection they will be affected by the environment and produce slight changes.

[0062] In one embodiment, the reflected wave signal is coupled by a preset parameter factor; the coupled reflected wave signal is filtered by a low-pass filter; and the phase difference between the reflected wave signal and the preset transmitted wave signal is calculated according to the filtered reflected wave signal. It should be noted that the reflected wave signal is coupled by a preset parameter factor, for example, the preset parameter factor is cos(2πft), then the coupled reflected wave signal:

[0063]

[0064] The coupled reflected wave signal is filtered through a high-frequency blocking low-pass filter, for example, the reflected wave signal after filtering:

[0065]

[0066] From the formula of the reflected wave signal after filtering, it can be known that the phase difference θ between the sound source wave signal and the reflected wave signal can be directly calculated based on the reflected wave signal after filtering. p .

[0067] In one embodiment, a reflected wave vector of a reflected wave signal is determined according to a phase difference, including: determining a phase difference corresponding to each second timestamp according to a plurality of second timestamps of the reflected wave signal; determining a difference between phase differences corresponding to each two adjacent second timestamps according to a phase difference corresponding to each second timestamp to obtain a plurality of phase difference values; determining a phase difference vector corresponding to the reflected wave signal according to the plurality of phase difference values; and inputting the phase difference vector corresponding to the reflected wave signal into a first preset neural network to obtain a reflected wave vector corresponding to the reflected wave signal.

[0068] It should be noted that the reflected wave signal includes multiple second timestamps, each second timestamp corresponds to a phase difference, and the phase difference value between the phase differences corresponding to each two adjacent second timestamps is calculated. The phase difference vector corresponding to the reflected wave signal can be obtained by connecting the multiple phase difference values, and then the phase difference vector corresponding to the reflected wave signal is input into the first preset neural network, which is, for example, a convolutional network and an LSTM network, to obtain the reflected wave vector corresponding to the reflected wave signal, which can assist in voice endpoint detection.

[0069] Exemplarily, the phase difference θ corresponding to the second timestamp at the current time t is calculated p , the phase difference corresponding to the second timestamp of the previous frame is θ p-1 , using the phase difference θ of the current frame p Subtract the phase difference θ of the previous frame p-1 , and get the phase difference ΔΦ p , find the phase difference ΔΦ p The k-th derivative ΔΦ p ∈R k As the phase difference vector corresponding to the second timestamp, ΔΦ p ∈R k Input a convolutional network and LSTM network to obtain the reflected wave vector r corresponding to the second timestamp of the current time t t =Conv r 2(Conv r 1(ΔΦ p )), the reflected wave vector can be obtained by splicing the sub-reflected wave vectors corresponding to each second timestamp of the reflected wave signal.

[0070] Step S105: Fusing the frequency spectrum feature vector, the video feature vector and the reflected wave vector to obtain a target feature vector.

[0071] The spectrum feature vector, video feature vector and reflection wave vector are fused to obtain a higher-dimensional target feature vector. By fusing the target feature vectors of the three modal data, the voice endpoint of the audio signal can be detected more accurately.

[0072] In one embodiment, the spectrum feature vector, the video feature vector and the reflection wave vector are combined to obtain a combined feature vector; the combined feature vector is input into a second preset neural network to obtain a target feature vector. i 、Video feature vector v i and the reflected wave vector r i Merge to get the merged feature vector h i =concat(a i ,v i ,r i ), the second preset neural network is, for example, a recursive neural network LSTM network and several fully connected networks, then the merged feature vector h i Input to the second preset neural network to obtain the target feature vector Z = FC (LSTM (h i )). By combining the three modal data and integrating the time series features through the second preset neural network, the features of the previous and next time ranges can be considered, thereby increasing the stability and accuracy of speech endpoint recognition.

[0073] Step S106: input the target feature vector into a pre-trained speech endpoint detection model to obtain multiple speech endpoints of the audio signal.

[0074] The pre-trained speech endpoint detection model is used to detect the speech endpoint of the target feature vector, and multiple speech endpoints of the audio signal can be output. The speech signal and non-speech signal in the audio signal can be distinguished through multiple speech endpoints. The anti-noise ability of speech endpoint detection is greatly improved, and the detection accuracy of speech endpoint detection in complex environments can be effectively improved.

[0075] In one embodiment, a first voice endpoint detection model is trained by multiple labeled audio signals to initialize the parameters of a second voice endpoint detection model; a second voice endpoint detection model is trained by multiple labeled video data to initialize the parameters of the second voice endpoint detection model; a third voice endpoint detection model is trained by multiple labeled reflection wave signals to initialize the parameters of the third voice endpoint detection model; the initialized first voice endpoint detection model, the second voice endpoint detection model and the third voice endpoint detection model are fused to obtain a target voice endpoint detection model; classification cross entropy is used as the objective function, and the target voice endpoint detection model is optimized using a back propagation algorithm to obtain a trained multimodal model.

[0076] Among them, the pre-trained voice endpoint detection model can be a regression classifier, a Bayesian network, a rule-based classifier or a neural network, etc., and the embodiments of the present application do not make specific limitations.

[0077] In one embodiment, the audio signal is smoothed and filtered by a preset smoothing layer to obtain a filtered audio signal. It should be noted that there are often many mutations and discontinuities in the voice endpoint detection process. By smoothing and filtering the audio signal, such as adjusting the offset (FEC) at the front of the segment of the audio signal to be detected and the tail (OVER) at the end, eliminating the segment (MSC) detected as a non-voice signal in the continuous voice signal, and the segment (NDS) recognized as a voice signal in the non-voice signal, the endpoint detection effect of the audio signal can be improved, and the user experience can be improved.

[0078] Please refer to Figure 3 , Figure 3 A schematic diagram of a scenario for implementing the voice endpoint detection method provided in this embodiment.

[0079] like Figure 3 As shown, the server obtains the audio signal to be detected and the video data and reflected wave signal collected when collecting the audio signal, then determines the spectrum data of the audio signal, and determines the spectrum feature vector of the audio signal based on the spectrum data, and extracts the image area where the lips in the video data are located, and determines the video feature vector of the video data based on each image area, and determines the phase difference between the reflected wave signal and the preset transmitted wave signal, and determines the reflected wave vector of the reflected wave signal based on the phase difference, and then fuses the spectrum feature vector, the video feature vector and the reflected wave vector to obtain the target feature vector, and then inputs the target feature vector into a pre-trained voice endpoint detection model to obtain multiple voice endpoints of the audio signal.

[0080] The voice endpoint detection method provided in the above embodiment obtains the audio signal to be detected and the video data and reflected wave signal collected when collecting the audio signal, then determines the spectrum data of the audio signal, and determines the spectrum feature vector of the audio signal based on the spectrum data, and at the same time extracts the image area where the lips are located in the video data, and determines the video feature vector of the video data according to each image area, and determines the phase difference between the reflected wave signal and the preset transmitted wave signal, and determines the reflected wave vector of the reflected wave signal based on the phase difference, and then fuses the spectrum feature vector, the video feature vector and the reflected wave vector to obtain the target feature vector, and then inputs the target feature vector into a pre-trained voice endpoint detection model to obtain multiple voice endpoints of the audio signal. The present application assists in voice endpoint detection through the three different modal data of the spectrum data of the audio signal and the synchronously collected video data and reflected wave signal, which greatly improves the noise resistance of voice endpoint detection and can effectively improve the detection accuracy of voice endpoint detection in complex environments.

[0081] Please refer to Figure 4 , Figure 4 A schematic block diagram of a speech endpoint detection device provided in an embodiment of the present application.

[0082] like Figure 4 As shown, the speech endpoint detection device 200 includes: an acquisition module 201, a first determination module 202, a second determination module 203, a third determination module 204, a fusion module 205 and a detection module 206.

[0083] An acquisition module 201 is used to acquire an audio signal to be detected and video data and a reflected wave signal collected when the audio signal is collected, wherein the reflected wave signal is collected by performing sound wave detection on the lips of the user;

[0084] A first determination module 202, configured to determine frequency spectrum data of the audio signal, and determine a frequency spectrum feature vector of the audio signal according to the frequency spectrum data;

[0085] A second determination module 203 is used to extract the image region where the lips are located in the video data, and determine the video feature vector of the video data according to each of the image regions;

[0086] A third determining module 204 is used to determine a phase difference between the reflected wave signal and a preset transmitted wave signal, and determine a reflected wave vector of the reflected wave signal according to the phase difference;

[0087] A fusion module 205 is used to fuse the spectrum feature vector, the video feature vector and the reflection wave vector to obtain a target feature vector;

[0088] The detection module 206 is used to input the target feature vector into a pre-trained speech endpoint detection model to obtain multiple speech endpoints of the audio signal.

[0089] In one embodiment, Figure 5 As shown, the first determination module 202 includes:

[0090] A determination submodule 2021, configured to determine, according to a plurality of first timestamps of the spectrum data, a feature vector corresponding to each of the first timestamps;

[0091] A processing submodule 2022, configured to perform convolution pooling processing on the feature vector corresponding to each of the first timestamps;

[0092] The splicing submodule 2023 is used to splice the feature vectors corresponding to each of the first timestamps processed by convolution pooling to obtain a spectrum feature vector corresponding to the spectrum data.

[0093] In one embodiment, the determination submodule 2021 is further used for:

[0094] Determine a plurality of first timestamps of the spectrum data, and determine a plurality of frames of spectrum data corresponding to each of the first timestamps;

[0095] A feature vector corresponding to each of the first timestamps is determined according to the feature parameters of the multi-frame spectrum data corresponding to each of the first timestamps.

[0096] In one embodiment, the third determining module 204 is further configured to:

[0097] coupling the reflected wave signal through a preset parameter factor;

[0098] The coupled reflected wave signal is filtered through a low-pass filter;

[0099] According to the reflected wave signal after filtering, the phase difference between the reflected wave signal and the preset transmitted wave signal is calculated.

[0100] In one embodiment, the third determining module 204 is further configured to:

[0101] Determine, according to the plurality of second timestamps of the reflected wave signal, a phase difference corresponding to each of the second timestamps;

[0102] Determine, according to the phase difference corresponding to each of the second timestamps, a difference between the phase differences corresponding to each of two adjacent second timestamps, to obtain a plurality of phase difference values;

[0103] Determining a phase difference vector corresponding to the reflected wave signal according to the multiple phase difference values;

[0104] The phase difference vector corresponding to the reflected wave signal is input into a first preset neural network to obtain a reflected wave vector corresponding to the reflected wave signal.

[0105] In one embodiment, the fusion module 205 is further configured to:

[0106] Merging the frequency spectrum feature vector, the video feature vector and the reflected wave vector to obtain a merged feature vector;

[0107] The combined feature vector is input into a second preset neural network to obtain a target feature vector.

[0108] In one embodiment, the speech endpoint detection device 200 is further used for:

[0109] The audio signal is smoothed and filtered by using a preset smoothing layer to obtain a filtered audio signal.

[0110] It should be noted that technical personnel in the relevant field can clearly understand that for the convenience and conciseness of description, the specific working process of the above-described device and each module and unit can refer to the corresponding process in the aforementioned voice endpoint detection method embodiment, and will not be repeated here.

[0111] The apparatus provided in the above embodiment may be implemented in the form of a computer program. The computer program may be Figure 6 Runs on the computer device shown.

[0112] See also Figure 6 , Figure 6 A schematic block diagram of the structure of a computer device provided in an embodiment of the present application. The computer device may be a server or a terminal device.

[0113] like Figure 6 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.

[0114] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any voice endpoint detection method.

[0115] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.

[0116] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any voice endpoint detection method.

[0117] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0118] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0119] In one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:

[0120] Acquire an audio signal to be detected and video data and a reflected wave signal collected when collecting the audio signal, wherein the reflected wave signal is collected by performing sound wave detection on the lips of the user;

[0121] Determining frequency spectrum data of the audio signal, and determining a frequency spectrum feature vector of the audio signal according to the frequency spectrum data;

[0122] Extracting the image region where the lips are located in the video data, and determining the video feature vector of the video data according to each of the image regions;

[0123] Determining a phase difference between the reflected wave signal and a preset transmitted wave signal, and determining a reflected wave vector of the reflected wave signal according to the phase difference;

[0124] fusing the frequency spectrum feature vector, the video feature vector and the reflected wave vector to obtain a target feature vector;

[0125] The target feature vector is input into a pre-trained speech endpoint detection model to obtain multiple speech endpoints of the audio signal.

[0126] In one embodiment, when the processor implements the step of determining the frequency spectrum feature vector of the audio signal according to the frequency spectrum data, the processor is configured to implement:

[0127] Determine, according to the multiple first timestamps of the spectrum data, a feature vector corresponding to each of the first timestamps;

[0128] Performing convolution pooling processing on the feature vectors corresponding to each of the first timestamps;

[0129] The feature vectors corresponding to each of the first timestamps processed by convolution pooling are concatenated to obtain a spectrum feature vector corresponding to the spectrum data.

[0130] In one embodiment, when the processor implements the step of determining, according to the multiple first timestamps of the spectrum data, the feature vector corresponding to each of the first timestamps, the processor is configured to implement:

[0131] Determine a plurality of first timestamps of the spectrum data, and determine a plurality of frames of spectrum data corresponding to each of the first timestamps;

[0132] A feature vector corresponding to each of the first timestamps is determined according to the feature parameters of the multi-frame spectrum data corresponding to each of the first timestamps.

[0133] In one embodiment, when the processor implements the determining of the phase difference between the reflected wave signal and the preset transmitted wave signal, it is used to implement:

[0134] coupling the reflected wave signal through a preset parameter factor;

[0135] The coupled reflected wave signal is filtered through a low-pass filter;

[0136] According to the reflected wave signal after filtering, the phase difference between the reflected wave signal and the preset transmitted wave signal is calculated.

[0137] In one embodiment, when the processor implements the step of determining the reflected wave vector of the reflected wave signal according to the phase difference, the processor is configured to implement:

[0138] Determine, according to the plurality of second timestamps of the reflected wave signal, a phase difference corresponding to each of the second timestamps;

[0139] Determine, according to the phase difference corresponding to each of the second timestamps, a difference between the phase differences corresponding to each of two adjacent second timestamps, to obtain a plurality of phase difference values;

[0140] Determining a phase difference vector corresponding to the reflected wave signal according to the multiple phase difference values;

[0141] The phase difference vector corresponding to the reflected wave signal is input into a first preset neural network to obtain a reflected wave vector corresponding to the reflected wave signal.

[0142] In one embodiment, when the processor implements the fusion of the spectrum feature vector, the video feature vector and the reflected wave vector to obtain the target feature vector, it is used to implement:

[0143] Merging the frequency spectrum feature vector, the video feature vector and the reflected wave vector to obtain a merged feature vector;

[0144] The combined feature vector is input into a second preset neural network to obtain a target feature vector.

[0145] In one embodiment, the processor is further configured to implement: performing smoothing filtering on the audio signal through a preset smoothing layer to obtain a filtered audio signal.

[0146] It should be noted that technical personnel in the relevant field can clearly understand that for the convenience and brevity of description, the specific working process of the computer device described above can refer to the corresponding process in the aforementioned voice endpoint detection method embodiment, and will not be repeated here.

[0147] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions. The method implemented when the program instructions are executed can refer to the various embodiments of the voice endpoint detection method of the present application.

[0148] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc., equipped on the computer device.

[0149] It should be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0150] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations. It should be noted that, in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "including a..." does not exclude the presence of other identical elements in the process, method, article or system including the element.

[0151] The serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments. The above description is only a specific implementation mode of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A speech endpoint detection method, characterized in that: include: Acquire an audio signal to be detected and video data and a reflected wave signal collected when collecting the audio signal, wherein the reflected wave signal is collected by performing sound wave detection on the lips of the user; Determining frequency spectrum data of the audio signal, and determining a frequency spectrum feature vector of the audio signal according to the frequency spectrum data; Extracting the image region where the lips are located in the video data, and determining the video feature vector of the video data according to each of the image regions; Determining a phase difference between the reflected wave signal and a preset transmitted wave signal, and determining a reflected wave vector of the reflected wave signal according to the phase difference; fusing the frequency spectrum feature vector, the video feature vector and the reflected wave vector to obtain a target feature vector; The target feature vector is input into a pre-trained speech endpoint detection model to obtain multiple speech endpoints of the audio signal.

2. The speech endpoint detection method according to claim 1, characterized in that: The determining the frequency spectrum feature vector of the audio signal according to the frequency spectrum data comprises: Determine, according to the multiple first timestamps of the spectrum data, a feature vector corresponding to each of the first timestamps; Performing convolution pooling processing on the feature vectors corresponding to each of the first timestamps; The feature vectors corresponding to each of the first timestamps processed by convolution pooling are concatenated to obtain a spectrum feature vector corresponding to the spectrum data.

3. The speech endpoint detection method according to claim 2, characterized in that: The determining, according to the plurality of first time stamps of the spectrum data, a feature vector corresponding to each of the first time stamps, comprises: Determine a plurality of first timestamps of the spectrum data, and determine a plurality of frames of spectrum data corresponding to each of the first timestamps; A feature vector corresponding to each of the first timestamps is determined according to the feature parameters of the multi-frame spectrum data corresponding to each of the first timestamps.

4. The speech endpoint detection method according to claim 1, characterized in that: The determining of the phase difference between the reflected wave signal and the preset transmitted wave signal comprises: coupling the reflected wave signal through a preset parameter factor; The coupled reflected wave signal is filtered through a low-pass filter; According to the reflected wave signal after filtering, the phase difference between the reflected wave signal and the preset transmitted wave signal is calculated.

5. The method for voice endpoint detection according to any one of claims 1 to 4, characterized in that: The determining the reflected wave vector of the reflected wave signal according to the phase difference comprises: Determine, according to the plurality of second timestamps of the reflected wave signal, a phase difference corresponding to each of the second timestamps; Determine, according to the phase difference corresponding to each of the second timestamps, a difference between the phase differences corresponding to each of two adjacent second timestamps, to obtain a plurality of phase difference values; Determining a phase difference vector corresponding to the reflected wave signal according to the multiple phase difference values; The phase difference vector corresponding to the reflected wave signal is input into a first preset neural network to obtain a reflected wave vector corresponding to the reflected wave signal.

6. The method for detecting speech endpoints according to any one of claims 1 to 4, characterized in that: The fusing the spectrum feature vector, the video feature vector and the reflected wave vector to obtain a target feature vector includes: Merging the frequency spectrum feature vector, the video feature vector and the reflected wave vector to obtain a merged feature vector; The combined feature vector is input into a second preset neural network to obtain a target feature vector.

7. The method for voice endpoint detection according to any one of claims 1 to 4, characterized in that: The method further comprises: The audio signal is smoothed and filtered by using a preset smoothing layer to obtain a filtered audio signal.

8. A speech endpoint detection device, characterized in that: The speech endpoint detection device comprises: An acquisition module, used to acquire an audio signal to be detected and video data and a reflected wave signal collected when the audio signal is collected, wherein the reflected wave signal is collected by performing sound wave detection on the lips of the user; A first determining module, configured to determine frequency spectrum data of the audio signal, and determine a frequency spectrum feature vector of the audio signal according to the frequency spectrum data; A second determination module is used to extract the image area where the lips are located in the video data, and determine the video feature vector of the video data according to each of the image areas; a third determining module, configured to determine a phase difference between the reflected wave signal and a preset transmitted wave signal, and determine a reflected wave vector of the reflected wave signal according to the phase difference; A fusion module, used for fusing the spectrum feature vector, the video feature vector and the reflected wave vector to obtain a target feature vector; The detection module is used to input the target feature vector into a pre-trained speech endpoint detection model to obtain multiple speech endpoints of the audio signal.

9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the steps of the voice endpoint detection method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the voice endpoint detection method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Voice signal processing method, device, system and facility, and storage medium

    CN110875060A

  • Multi-mode voice endpoint detection method and device

    CN111768760A