Voice endpoint detection method and apparatus, electronic device, and storage medium
By performing beamforming and energy entropy ratio curve analysis on the raw signals of the sensor array, combined with cross-correlation processing, the problem of low efficiency and accuracy of voice endpoint detection in existing technologies is solved, achieving more efficient and accurate voice endpoint detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAOMI EV TECH CO LTD
- Filing Date
- 2022-12-20
- Publication Date
- 2026-04-10
AI Technical Summary
Existing voice endpoint detection methods are inefficient and inaccurate.
By acquiring raw signals from multiple sensors in a sensor array, beamforming and voice endpoint detection are performed. The initial and target times of the voice signal are determined using the energy entropy ratio curve and cross-correlation processing.
It improves the efficiency and accuracy of voice endpoint detection, reduces the number of detections, and enhances the signal-to-noise ratio.
Smart Images

Figure CN116246653B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of speech processing, and particularly relates to a speech endpoint detection method and device, electronic equipment and storage medium. BACKGROUND
[0002] At present, with the development of artificial intelligence, natural language processing and other technologies, speech processing has been widely applied in the fields of intelligent household appliances, robot voice interaction, vehicle-mounted terminals and the like. Speech endpoint detection can include identifying whether there is a speech signal from the collected original signal, and the start time, end time and the like of the speech signal, which is crucial to speech processing. However, the speech endpoint detection method in the related art has the problems of low efficiency and low accuracy. SUMMARY
[0003] The present disclosure provides a speech endpoint detection method, device, electronic equipment, computer readable storage medium and computer program product to at least solve the problem of low efficiency and low accuracy of the speech endpoint detection method in the related art. The technical solutions of the present disclosure are as follows:
[0004] According to a first aspect of an embodiment of the present disclosure, a speech endpoint detection method is provided, comprising: acquiring original signals collected by a plurality of sensors in a sensor array; performing beamforming on a plurality of the original signals to obtain a beam signal; performing speech endpoint detection on the beam signal to obtain an initial time of a speech signal; and obtaining a target time of the speech signal corresponding to the plurality of sensors based on the initial time.
[0005] In an embodiment of the present disclosure, the performing beamforming on a plurality of the original signals to obtain a beam signal comprises: performing orientation estimation on a sound source to obtain an orientation angle between the sound source and the sensor array; and performing beamforming on a plurality of the original signals based on the orientation angle to obtain the beam signal.
[0006] In an embodiment of the present disclosure, the performing beamforming on a plurality of the original signals to obtain a beam signal comprises: in response to the original signal being a wideband signal, performing beamforming on a plurality of the original signals in the frequency domain to obtain a frequency domain beam signal; or, in response to the original signal being a single frequency signal, performing beamforming on a plurality of the original signals in the time domain to obtain a time domain beam signal.
[0007] In an embodiment of the present disclosure, the performing speech endpoint detection on the beam signal to obtain an initial time of a speech signal comprises: performing frame processing on the beam signal to obtain a plurality of frames of beam signals; acquiring an energy-entropy ratio of each frame of beam signal; obtaining the initial time based on the energy-entropy ratio; and wherein the initial time comprises an initial start time and / or an initial end time.
[0008] In one embodiment of the present disclosure, the initial time is obtained based on the energy-entropy ratio, including: obtaining an energy-entropy ratio curve based on the energy-entropy ratio of the multi-frame beam signals, wherein the horizontal coordinate of the energy-entropy ratio curve is time, and the vertical coordinate is the energy-entropy ratio; obtaining a first intersection point and a second intersection point between the energy-entropy ratio curve and a first reference line, and obtaining a third intersection point and a fourth intersection point between the energy-entropy ratio curve and a second reference line, wherein the horizontal coordinate of the first intersection point is less than the horizontal coordinate of the second intersection point; in response to the horizontal coordinate of the third intersection point being less than the horizontal coordinate of the first intersection point, determining the horizontal coordinate of the third intersection point as the initial start time of the speech signal; and / or, in response to the horizontal coordinate of the fourth intersection point being greater than the horizontal coordinate of the second intersection point, determining the horizontal coordinate of the fourth intersection point as the initial end time of the speech signal; wherein the initial time includes the initial start time and / or the initial end time.
[0009] In one embodiment of the present disclosure, the first reference line and the second reference line are both parallel to the horizontal axis, the vertical coordinates of the points on the first reference line are all first thresholds, and the vertical coordinates of the points on the second reference line are all second thresholds, the second threshold being less than the first threshold.
[0010] In one embodiment of the present disclosure, the target time of the speech signal corresponding to the sensor is obtained based on the initial time, including: obtaining a time difference corresponding to the sensor; and delaying the initial time by the time difference to obtain the target time of the speech signal corresponding to the sensor.
[0011] In one embodiment of the present disclosure, the time difference corresponding to the sensor is obtained, including: performing cross-correlation processing on the original signal collected by the sensor and the beam signal to obtain a cross-correlation curve; and determining the time corresponding to the peak value of the cross-correlation curve as the time difference corresponding to the sensor.
[0012] In one embodiment of the present disclosure, the time difference corresponding to the sensor is obtained, including: performing direction estimation on the sound source to obtain an azimuth angle between the sound source and the sensor array; and obtaining the time difference corresponding to the sensor based on the azimuth angle and the distance between the sensor and a reference sensor.
[0013] According to a second aspect of the embodiments of the present disclosure, a device for detecting a voice endpoint is provided, comprising: a collection module configured to perform acquisition of original signals collected by a plurality of sensors in a sensor array; a processing module configured to perform beamforming on the plurality of original signals to obtain a beam signal; a detection module configured to perform voice endpoint detection on the beam signal to obtain an initial time of a voice signal; and an acquisition module configured to perform acquisition of a target time of the voice signal corresponding to the plurality of sensors based on the initial time.
[0014] In an embodiment of the present disclosure, the processing module is further configured to perform: azimuth estimation on a sound source to obtain an azimuth angle between the sound source and the sensor array; and beamforming on the plurality of original signals based on the azimuth angle to obtain the beam signal.
[0015] In an embodiment of the present disclosure, the processing module is further configured to perform: beamforming on the plurality of original signals in a frequency domain to obtain a frequency-domain beam signal in response to the original signals being wideband signals; or beamforming on the plurality of original signals in a time domain to obtain a time-domain beam signal in response to the original signals being single-frequency signals.
[0016] In an embodiment of the present disclosure, the detection module is further configured to perform: frame processing on the beam signal to obtain a plurality of frames of beam signals; acquisition of an energy-entropy ratio of each frame of beam signals; and acquisition of the initial time based on the energy-entropy ratio, wherein the initial time comprises an initial start time and / or an initial end time.
[0017] In an embodiment of the present disclosure, the detection module is further configured to perform: acquisition of an energy-entropy ratio curve based on the energy-entropy ratios of the plurality of frames of beam signals, wherein the energy-entropy ratio curve has a time as an abscissa and an energy-entropy ratio as an ordinate; acquisition of a first intersection and a second intersection between the energy-entropy ratio curve and a first reference line, and acquisition of a third intersection and a fourth intersection between the energy-entropy ratio curve and a second reference line, wherein the abscissa of the first intersection is less than the abscissa of the second intersection; determination of the abscissa of the third intersection as the initial start time of the voice signal in response to the abscissa of the third intersection being less than the abscissa of the first intersection; and / or determination of the abscissa of the fourth intersection as the initial end time of the voice signal in response to the abscissa of the fourth intersection being greater than the abscissa of the second intersection; wherein the initial time comprises the initial start time and / or the initial end time.
[0018] In an embodiment of the present disclosure, the first reference line and the second reference line are both parallel to a horizontal axis, and the vertical coordinates of points on the first reference line are all a first threshold, and the vertical coordinates of points on the second reference line are all a second threshold, and the second threshold is less than the first threshold.
[0019] In an embodiment of the present disclosure, the acquisition module is further configured to perform: acquiring a time difference corresponding to the sensor; and delaying the initial time by the time difference to obtain a target time of the voice signal corresponding to the sensor.
[0020] In an embodiment of the present disclosure, the acquisition module is further configured to perform: performing cross-correlation processing on the original signal collected by the sensor and the beam signal to obtain a cross-correlation curve; and determining a time corresponding to a peak value of the cross-correlation curve as the time difference corresponding to the sensor.
[0021] In an embodiment of the present disclosure, the acquisition module is further configured to perform: performing direction estimation on the sound source to obtain a direction angle between the sound source and the sensor array; and obtaining the time difference corresponding to the sensor based on the direction angle and a distance between the sensor and a reference sensor.
[0022] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, including a processor; a memory for storing processor-executable instructions; and wherein the processor is configured to implement the steps of the method according to the first aspect of the present disclosure.
[0023] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, having stored thereon computer program instructions, which, when executed by a processor, implement the steps of the method according to the first aspect of the present disclosure.
[0024] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, wherein the computer program is executed by a processor of an electronic device to implement the steps of the method according to the first aspect of the present disclosure.
[0025] The technical solution provided by the embodiments of the present disclosure at least brings the following beneficial effects: the original signals collected by multiple sensors can be beamformed to obtain a beam signal, and the beam signal can be subjected to voice endpoint detection to obtain an initial time, so as to obtain target times corresponding to the multiple sensors. Compared with the related art in which the original signals collected by each sensor are individually subjected to voice endpoint detection, in the present solution, only the beam signal needs to be subjected to voice endpoint detection, thereby greatly reducing the number of times of voice endpoint detection, which helps to improve the efficiency of voice endpoint detection of the sensor array, and the signal-to-noise ratio of the beam signal is high, which helps to improve the accuracy of voice endpoint detection.
[0026] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory and are not restrictive of the disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0027] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and serve to explain the principles of the present disclosure, and, do not limit the present disclosure.
[0028] Figure 1 is a flow chart of a voice endpoint detection method according to an exemplary embodiment.
[0029] Figure 2 is a schematic diagram of a voice endpoint detection method according to an exemplary embodiment.
[0030] Figure 3 is a flow chart of a voice endpoint detection method according to another exemplary embodiment.
[0031] Figure 4 is a schematic diagram of an energy entropy ratio curve, a first reference line, and a second reference line in a voice endpoint detection method according to an exemplary embodiment.
[0032] Figure 5 is a flow chart of a voice endpoint detection method according to another exemplary embodiment.
[0033] Figure 6 is a block diagram of a voice endpoint detection system according to an exemplary embodiment.
[0034] Figure 7 is a block diagram of a voice endpoint detection apparatus according to an exemplary embodiment.
[0035] Figure 8 is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0036] In order for those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings.
[0037] It should be noted that the terms "first", "second", and the like in the description and claims of the present disclosure and the foregoing drawings are used only to distinguish similar objects and do not necessarily have a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0038] The acquisition, storage, use, processing, etc. of data in the technical solutions of the present disclosure comply with relevant provisions of national laws and regulations.
[0039] Figure 1 is a flowchart of a voice endpoint detection method according to an exemplary embodiment, as shown in Figure 1 The voice endpoint detection method of the present embodiment comprises the following steps.
[0040] S101, acquiring original signals collected by multiple sensors in a sensor array.
[0041] It should be noted that the execution subject of the voice endpoint detection method of the present embodiment is an electronic device, which includes a mobile phone, a notebook, a desktop computer, a vehicle-mounted terminal, a smart home appliance, etc. The voice endpoint detection method of the present embodiment can be executed by the voice endpoint detection apparatus of the present embodiment, which can be configured in any electronic device to execute the voice endpoint detection method of the present embodiment.
[0042] In the present embodiment, the sensor array comprises multiple sensors. The distribution of the multiple sensors in the sensor array is not limited, for example, multiple sensors can be arranged in at least one set direction, and the interval between two adjacent sensors is set to a certain distance. The set direction is not limited, for example, in the xy two-dimensional coordinate system, multiple sensors can be arranged in the x-axis and y-axis directions, respectively.
[0043] In some examples, the sensor array comprises sensors 1 to 20, which are arranged in the x-axis direction, and the interval between two adjacent sensors is 20 centimeters.
[0044] It should be noted that the sensor is used to collect sound. The sensor is not limited, for example, it can include a microphone (MIC), a sonar, a radar, etc.
[0045] It should be noted that the original signal includes speech signals, and due to environmental factors such as vibration, the original signal also includes noise signals. No strict limitations are placed on the SNR (Signal Noise Ratio) of the original signal; for example, it can be -3dB.
[0046] It is understandable that different sensors will collect different raw signals.
[0047] In one implementation, the controllable sensor acquires the raw signal at a set sampling rate. It should be noted that the set sampling rate is not limited in many ways; for example, the set sampling rate could be 8 kHz (kilohertz).
[0048] In some examples, the sensor array includes M sensors, and each sensor collects raw signals consisting of N sampling points. Therefore, the raw signals collected by the M sensors consist of M*N sampling points. It should be noted that there is a one-to-one correspondence between the sampling points and the sampling time, where M and N are both positive integers.
[0049] It is understandable that the sampling time may include t1, t2 to t N The raw signal collected by the i-th sensor includes S i1 S i2 To S iN , among which, S i1 S i2 To S iN All are sampling point signals, S i1 S i2 To S iN The corresponding sampling times are t1, t2 to t3 respectively. N 1≤i≤M, where i is a positive integer.
[0050] S102 performs beamforming on multiple raw signals to obtain a beam signal.
[0051] In the embodiments of this disclosure, beamforming can be performed on the raw signals collected by multiple sensors to obtain beam signals. It should be noted that the specific method of beamforming is not limited in many ways; for example, any beamforming algorithm in related technologies can be used.
[0052] In some examples, the sensor array includes M sensors, each sensor acquiring raw signals comprising N sampling points. Beamforming is performed on multiple raw signals to obtain a beam signal, which involves weighted summation of the M sampling points at any given sampling time. Therefore, the beam signal comprises beam signals at N sampling times.
[0053] For example, taking M=5 and N=5 as an example, the sampling time can include t1, t2, t5, and the original signal collected by the i th sensor includes S i1 , S i2 to S i5 , wherein S i1 , S i2 to S i5 are sampling point signals, S i1 , S i2 to S i5 correspond to sampling times t1, t2, t5, respectively. 1≤i≤5, i is a positive integer.
[0054] Taking sampling time t1 as an example, the sampling point signals S 11 , S 21 , S 31 , S 41 , S 51 at sampling time t1 can be weighted and summed to obtain the beam signal at sampling time t1. Wherein S 11 is the sampling point signal collected by the first sensor at sampling time t1, S 21 is the sampling point signal collected by the second sensor at sampling time t1, S 31 is the sampling point signal collected by the third sensor at sampling time t1, S 41 is the sampling point signal collected by the fourth sensor at sampling time t1, and S 51 is the sampling point signal collected by the fifth sensor at sampling time t1.
[0055] It should be noted that the acquisition process of the beam signals at sampling times t2 to t5 can refer to the acquisition process of the beam signal at sampling time t1, which will not be described here.
[0056] In an embodiment, the plurality of original signals are beamformed to obtain a beam signal, including estimating the direction of the sound source to obtain the azimuth angle between the sound source and the sensor array, and beamforming the plurality of original signals based on the azimuth angle to obtain the beam signal. Thus, the azimuth angle between the sound source and the sensor array can be considered in the method to beamform the plurality of original signals to obtain the beam signal.
[0057] It should be noted that the azimuth angle can be any angle or an angle interval, which is not limited here. For example, as shown in FIG. 1, the azimuth angle θ is the included angle between the direction of the sound source and the vertical direction. Figure 2
[0058] It should be noted that the specific manner of azimuth estimation of the sound source is not limited too much, for example, CBF (Conventional Beamforming), MVDR (Minimum Variance Distortionless Response), MUSIC (Multiple Signal Classification), CS (Compressed Sensing) and other azimuth estimation algorithms can be used to realize the azimuth estimation.
[0059] In an embodiment, the plurality of original signals are beamformed to obtain a beam signal, including, in response to the original signals being wideband signals, the plurality of original signals are beamformed in the frequency domain to obtain a frequency domain beam signal, or, in response to the original signals being single frequency signals, the plurality of original signals are beamformed in the time domain to obtain a time domain beam signal. Thus, the method can consider that the original signals are wideband signals or single frequency signals, and the plurality of original signals are beamformed in the frequency domain or the time domain, thereby improving the flexibility of beamforming.
[0060] It should be noted that the specific manner of beamforming in the frequency domain or the time domain is not limited too much, for example, the beamforming in the frequency domain can be realized by using any frequency domain beamforming algorithm in the related art, and the beamforming in the time domain can be realized by using any time domain beamforming algorithm in the related art.
[0061] In some examples, the plurality of original signals are beamformed in the frequency domain to obtain a frequency domain beam signal, including, based on the azimuth angle, the plurality of original signals are beamformed in the frequency domain to obtain a frequency domain beam signal.
[0062] In some examples, the plurality of original signals are beamformed in the time domain to obtain a time domain beam signal, including, based on the azimuth angle, the plurality of original signals are beamformed in the time domain to obtain a time domain beam signal.
[0063] S103, speech endpoint detection is performed on the beam signal to obtain an initial time of the speech signal.
[0064] It should be noted that the specific manner of speech endpoint detection on the beam signal is not limited too much, for example, any speech endpoint detection algorithm in the related art can be used to realize the speech endpoint detection.
[0065] It can be understood that the beam signal includes a speech signal and a noise signal, the speech endpoint detection is performed on the beam signal to obtain an initial time of the speech signal, and the initial time of the speech signal refers to the time of the speech signal in the beam signal, which can include the initial start time and / or the initial end time of the speech signal in the beam signal.
[0066] In an embodiment, the beam signal is subjected to voice endpoint detection to obtain an initial time of the voice signal, including subjecting the beam signal to frame processing to obtain a plurality of frames of beam signals, obtaining a target parameter of each frame of beam signals, and obtaining the initial time based on the target parameter. The initial time includes an initial start time and / or an initial end time. It should be noted that the target parameter is not limited, for example, the target parameter can include volume, zero-crossing rate, spectral entropy value, energy entropy ratio, etc.
[0067] It should be noted that the specific manner of frame processing is not limited, for example, any frame processing algorithm in the related art can be used to implement.
[0068] In some examples, obtaining the initial time based on the target parameter includes obtaining a target curve based on the target parameter of the plurality of frames of beam signals, where the abscissa of the target curve is time and the ordinate is the target parameter, obtaining a fifth intersection point and a sixth intersection point between the target curve and a set reference line, where the abscissa of the fifth intersection point is less than the abscissa of the sixth intersection point, the ordinate of the points on the set reference line are all set thresholds, determining the abscissa of the fifth intersection point as the initial start time of the voice signal, and determining the abscissa of the sixth intersection point as the initial end time of the voice signal. It should be noted that the set reference line is parallel to the abscissa, and the set threshold is not limited.
[0069] In S104, a target time of the voice signal corresponding to each of the plurality of sensors is obtained based on the initial time.
[0070] It can be understood that the target time of the voice signal corresponding to each of the plurality of sensors refers to the time of the voice signal in the original signal collected by the sensor, which can include a target start time and / or a target end time of the voice signal in the original signal.
[0071] It can be understood that the target time of the voice signal corresponding to each of the plurality of sensors can be different.
[0072] In an embodiment, obtaining the target time of the voice signal corresponding to each of the plurality of sensors based on the initial time includes obtaining a target start time of the voice signal corresponding to each of the plurality of sensors based on the initial start time, and / or obtaining a target end time of the voice signal corresponding to each of the plurality of sensors based on the initial end time.
[0073] In an embodiment, obtaining the target time of the voice signal corresponding to each of the plurality of sensors based on the initial time includes obtaining a time difference corresponding to each of the plurality of sensors, and delaying the initial time by the time difference to obtain the target time of the voice signal corresponding to each of the plurality of sensors. It can be understood that different sensors can correspond to different time differences, and the time difference refers to the time difference between the target time of the voice signal corresponding to each of the plurality of sensors and the initial time.
[0074] For example, as shown in Figure 2 the sensor array includes sensors 1 to 4, and the time differences corresponding to the sensors 1 to 4 are 0 seconds, 1 second, 2 seconds, and 3 seconds, respectively. If the initial time includes an initial start time t1 seconds and an initial end time t2 seconds, the target start time of the speech signal corresponding to the sensor 1 is t1 seconds, and the target end time is t2 seconds, the target start time of the speech signal corresponding to the sensor 2 is (t1 + 1) seconds, and the target end time is (t2 + 1) seconds, the target start time of the speech signal corresponding to the sensor 3 is (t1 + 2) seconds, and the target end time is (t2 + 2) seconds, and the target start time of the speech signal corresponding to the sensor 4 is (t1 + 3) seconds, and the target end time is (t2 + 3) seconds.
[0075] In some examples, the time difference corresponding to the sensor is obtained by performing direction estimation on the sound source to obtain an azimuth angle between the sound source and the sensor array, and obtaining the time difference corresponding to the sensor based on the azimuth angle and the distance between the sensor and a reference sensor.
[0076] For example, continuing with the example of Figure 2 the reference sensor can be the sensor 1 closest to the sound source, and if the time difference corresponding to the sensor 1 is 0 seconds, the time difference corresponding to the sensor 2 is Δt2 = (d * sinθ) / c, the time difference corresponding to the sensor 3 is Δt3 = (2d * sinθ) / c, and the time difference corresponding to the sensor 4 is Δt3 = (3d * sinθ) / c. Wherein d is the distance between two adjacent sensors, θ is the azimuth angle, and c is the sound speed of the medium.
[0077] The speech endpoint detection method provided by the embodiments of the present disclosure includes obtaining the original signals collected by the plurality of sensors in the sensor array, performing beamforming on the plurality of original signals to obtain a beam signal, performing speech endpoint detection on the beam signal to obtain an initial time of the speech signal, and obtaining the target time of the speech signal corresponding to the plurality of sensors based on the initial time. Thus, the original signals collected by the plurality of sensors can be beamformed to obtain a beam signal, and the beam signal can be subjected to speech endpoint detection to obtain an initial time, so as to obtain the target time corresponding to the plurality of sensors. Compared with the related art in which the original signals collected by each sensor are mostly subjected to speech endpoint detection individually, the present solution only needs to perform speech endpoint detection on the beam signal, greatly reducing the number of times of speech endpoint detection, which helps to improve the efficiency of speech endpoint detection of the sensor array, and the signal-to-noise ratio of the beam signal is high, which helps to improve the accuracy of speech endpoint detection.
[0078] Figure 3 is a flowchart of a speech endpoint detection method according to another exemplary embodiment, as shown in Figure 3As shown, the method for detecting a voice endpoint according to the embodiments of the present disclosure comprises the following steps.
[0079] S301, obtaining original signals collected by a plurality of sensors in a sensor array.
[0080] S302, performing beamforming on the plurality of original signals to obtain a beam signal.
[0081] S303, performing frame processing on the beam signal to obtain a plurality of frames of beam signals.
[0082] The related content of steps S301-S303 can be referred to the above embodiments, which will not be repeated here.
[0083] S304, obtaining an energy-entropy ratio of each frame of beam signals.
[0084] It should be noted that the specific manner of obtaining the energy-entropy ratio is not limited too much, for example, any energy-entropy ratio calculation algorithm in the related art can be used to implement it.
[0085] In one embodiment, obtaining the energy-entropy ratio of each frame of beam signals comprises obtaining a short-time energy and a short-time spectral entropy of any frame of beam signals, and obtaining the energy-entropy ratio of the any frame of beam signals based on the short-time energy and the short-time spectral entropy.
[0086] In some examples, obtaining the energy-entropy ratio of each frame of beam signals can be implemented by the following formula:
[0087]
[0088] LE q = log(1 + E q a)
[0089]
[0090]
[0091]
[0092]
[0093] wherein S q (q, k) is a signal component of the qth frame of beam signals at the kth frequency point, E q is a short-time energy of the qth frame of beam signals, Y q (q, k) is an energy of the qth frame of beam signals at the kth frequency point, p q (q, k) is a normalized spectral probability density function of the qth frame of beam signals at the kth frequency point, H q is a short-time spectral entropy of the qth frame of beam signals, and Q qEnergy-entropy ratio of the qth frame beam signal.
[0094] Wherein, N / 2 represents only taking the positive frequency part, and a is a set coefficient.
[0095] S305, obtaining an energy-entropy ratio curve based on the energy-entropy ratios of the plurality of frame beam signals, wherein the abscissa of the energy-entropy ratio curve is time, and the ordinate is the energy-entropy ratio.
[0096] In some examples, obtaining the energy-entropy ratio curve based on the energy-entropy ratios of the plurality of frame beam signals includes performing curve fitting based on the energy-entropy ratios of the plurality of frame beam signals to obtain the energy-entropy ratio curve. It should be noted that the specific manner of curve fitting is not limited too much, such as linear fitting, least squares method, etc.
[0097] S306, obtaining a first intersection point and a second intersection point between the energy-entropy ratio curve and a first reference line, and obtaining a third intersection point and a fourth intersection point between the energy-entropy ratio curve and a second reference line, wherein the abscissa of the first intersection point is less than the abscissa of the second intersection point.
[0098] S307, in response to the abscissa of the third intersection point being less than the abscissa of the first intersection point, determining the abscissa of the third intersection point as the initial start time of the speech signal.
[0099] S308, in response to the abscissa of the fourth intersection point being greater than the abscissa of the second intersection point, determining the abscissa of the fourth intersection point as the initial end time of the speech signal.
[0100] It should be noted that the first reference line and the second reference line are not limited too much.
[0101] In an embodiment, as shown in FIG. 1, Figure 4 the first reference line L1 and the second reference line L2 are both parallel to the horizontal axis, the ordinate of the points on the first reference line L1 are all the first threshold Th1, the ordinate of the points on the second reference line L2 are all the second threshold Th2, and the second threshold Th2 is less than the first threshold Th1. That is, the first reference line L1 is located above the second reference line L2, and the first threshold Th1 and the second threshold Th2 are not limited too much. For example, the first threshold Th1 can be 1.5 times the noise signal energy, and the second threshold Th2 can be the noise signal energy.
[0102] In some examples, the first threshold Th1 and the second threshold Th2 are as follows:
[0103] Th1 = α1D + σ n
[0104] Th2 = α2D + σ n
[0105] α1 < α2
[0106] Where D is the energy difference between the speech signal and the noise signal, σ n For the pre-acquired noise signal energy, or σ n It can be the average energy of a silent frame, D, σ n It can be preset or updated in real time; we will not impose too many restrictions here.
[0107] Where α1 and α2 are set coefficients.
[0108] In some examples, such as Figure 4 As shown, the first and second intersection points between the energy entropy ratio curve L3 and the first reference line L1 are A and B, respectively. The third and fourth intersection points between the energy entropy ratio curve L3 and the second reference line L2 are C and D, respectively. Points A, B, C, and D are sorted in ascending order of their abscissas, resulting in points C, A, B, and D. The abscissa of point C can be determined as the initial start time of the speech signal, and the abscissa of point D can be determined as the initial end time of the speech signal.
[0109] The voice endpoint detection method provided in the embodiments of this disclosure performs frame-segmentation processing on the beam signal to obtain multiple frames of beam signal, obtains the energy entropy ratio of each frame of beam signal, obtains an energy entropy ratio curve based on the energy entropy ratio of the multiple frames of beam signal, and obtains the initial time of the voice signal by comprehensively considering a first threshold and a second threshold.
[0110] Figure 5 This is a flowchart illustrating a method for detecting a voice endpoint according to another exemplary embodiment, such as... Figure 5 As shown, the voice endpoint detection method of this disclosure includes the following steps.
[0111] S501 acquires the raw signals collected by multiple sensors in the sensor array.
[0112] S502 performs beamforming on multiple raw signals to obtain a beam signal.
[0113] S503 performs voice endpoint detection on the beam signal to obtain the initial time of the voice signal.
[0114] For details regarding steps S501-S503, please refer to the above embodiments; they will not be repeated here.
[0115] S504 performs cross-correlation processing on the raw signal and beam signal acquired by the sensor to obtain the cross-correlation curve.
[0116] S505 determines the time corresponding to the peak value of the cross-correlation curve as the time difference corresponding to the sensor.
[0117] It should be noted that the specific manner of cross-correlation processing is not limited too much, for example, any cross-correlation algorithm in the related art can be used to implement.
[0118] It should be noted that the abscissa of the cross-correlation curve is time, and the ordinate is the correlation parameter, which is used to represent the correlation of the original signal and the beam signal at a certain time. If the correlation parameter is large, it indicates that the correlation of the original signal and the beam signal at a certain time is strong. If the correlation parameter is large, it indicates that the correlation of the original signal and the beam signal at a certain time is high. Conversely, if the correlation parameter is small, it indicates that the correlation of the original signal and the beam signal at a certain time is weak.
[0119] Continuing with the example of Figure 2 , the original signal and the beam signal collected by the sensor 1 can be cross-correlated to obtain a cross-correlation curve E1. The time corresponding to the peak value of the cross-correlation curve E1 is determined as the time difference corresponding to the sensor 1.
[0120] The original signal and the beam signal collected by the sensor 2 can be cross-correlated to obtain a cross-correlation curve E2. The time corresponding to the peak value of the cross-correlation curve E2 is determined as the time difference corresponding to the sensor 2.
[0121] The original signal and the beam signal collected by the sensor 3 can be cross-correlated to obtain a cross-correlation curve E3. The time corresponding to the peak value of the cross-correlation curve E3 is determined as the time difference corresponding to the sensor 3.
[0122] The original signal and the beam signal collected by the sensor 4 can be cross-correlated to obtain a cross-correlation curve E4. The time corresponding to the peak value of the cross-correlation curve E4 is determined as the time difference corresponding to the sensor 4.
[0123] It should be noted that the cross-correlation curves E1 to E4 are not shown in Figure 2 .
[0124] S506, delaying the initial time by the time difference to obtain the target time of the speech signal corresponding to the sensor.
[0125] The related content of step S506 can be referred to the above-mentioned embodiments, which will not be described here.
[0126] The speech endpoint detection method provided by the embodiments of the present disclosure can cross-correlate the original signal and the beam signal collected by the sensor to obtain a cross-correlation curve, and determine the time corresponding to the peak value of the cross-correlation curve as the time difference corresponding to the sensor, so as to realize the acquisition of the time difference of the sensor.
[0127] On the basis of any of the above embodiments, for example, Figure 6As shown, the voice endpoint detection system 100 includes a sensor array 110, an orientation estimation module 120, a beamforming module 130, an endpoint detection module 140, a time delay estimation module 150, and a time delay processing module 160.
[0128] The sensor array 110 includes a plurality of sensors for collecting raw signals, the orientation estimation module 120 is configured to estimate the orientation angle between a sound source and the sensor array 110, the beamforming module 130 is configured to perform beamforming on the raw signals collected by the plurality of sensors in the sensor array 110 to obtain a beam signal, the endpoint detection module 140 is configured to perform voice endpoint detection on the beam signal to obtain an initial time of a voice signal, the time delay estimation module 150 is configured to obtain a time difference corresponding to a sensor, and the time delay processing module 160 is configured to delay the initial time by the time difference to obtain a target time of the voice signal corresponding to the sensor.
[0129] In some examples, the time delay estimation module 150 is further configured to perform cross-correlation processing on the raw signals collected by the sensors and the beam signal to obtain a cross-correlation curve, and determine a time corresponding to a peak value of the cross-correlation curve as the time difference corresponding to the sensor.
[0130] In some examples, the time delay estimation module 150 is further configured to obtain the time difference corresponding to the sensor based on the orientation angle and a distance between the sensor and a reference sensor.
[0131] Figure 7 FIG. 1 is a block diagram of a voice endpoint detection device according to an example embodiment. Referring to FIG. 1, the voice endpoint detection device 100 includes a sensor array 110, an orientation estimation module 120, a beamforming module 130, an endpoint detection module 140, a time delay estimation module 150, and a time delay processing module 160. Figure 7 The voice endpoint detection device 200 of the present disclosure includes a collection module 210, a processing module 220, a detection module 230, and an acquisition module 240.
[0132] The collection module 210 is configured to perform collection of raw signals collected by a plurality of sensors in a sensor array.
[0133] The processing module 220 is configured to perform beamforming on the plurality of raw signals to obtain a beam signal.
[0134] The detection module 230 is configured to perform voice endpoint detection on the beam signal to obtain an initial time of a voice signal.
[0135] The acquisition module 240 is configured to obtain a target time of a voice signal corresponding to a plurality of sensors based on the initial time.
[0136] In an embodiment of the present disclosure, the processing module 220 is further configured to perform: azimuth estimation on the sound source to obtain an azimuth angle between the sound source and the sensor array; and beamforming on the plurality of original signals based on the azimuth angle to obtain the beam signal.
[0137] In an embodiment of the present disclosure, the processing module 220 is further configured to perform: in response to the original signal being a wideband signal, beamforming on the plurality of original signals in a frequency domain to obtain a frequency domain beam signal; or, in response to the original signal being a single frequency signal, beamforming on the plurality of original signals in a time domain to obtain a time domain beam signal.
[0138] In an embodiment of the present disclosure, the detection module 230 is further configured to perform: frame processing on the beam signal to obtain a plurality of frames of beam signals; obtaining an energy-entropy ratio of each frame of beam signal; and obtaining the initial time based on the energy-entropy ratio. The initial time includes an initial start time and / or an initial end time.
[0139] In an embodiment of the present disclosure, the detection module 230 is further configured to perform: obtaining an energy-entropy ratio curve based on the energy-entropy ratio of the plurality of frames of beam signals, wherein the abscissa of the energy-entropy ratio curve is time and the ordinate is the energy-entropy ratio; obtaining a first intersection and a second intersection between the energy-entropy ratio curve and a first reference line, and obtaining a third intersection and a fourth intersection between the energy-entropy ratio curve and a second reference line, wherein the abscissa of the first intersection is less than the abscissa of the second intersection; in response to the abscissa of the third intersection being less than the abscissa of the first intersection, determining the abscissa of the third intersection as the initial start time of the voice signal; and / or, in response to the abscissa of the fourth intersection being greater than the abscissa of the second intersection, determining the abscissa of the fourth intersection as the initial end time of the voice signal; wherein the initial time includes the initial start time and / or the initial end time.
[0140] In an embodiment of the present disclosure, the first reference line and the second reference line are parallel to the abscissa, the ordinate of the points on the first reference line are all a first threshold, the ordinate of the points on the second reference line are all a second threshold, and the second threshold is less than the first threshold.
[0141] In an embodiment of the present disclosure, the obtaining module 240 is further configured to perform: obtaining a time difference corresponding to the sensor; and delaying the initial time by the time difference to obtain a target time of the voice signal corresponding to the sensor.
[0142] In an embodiment of the present disclosure, the acquisition module 240 is further configured to perform: performing cross-correlation processing on the original signals collected by the sensors and the beam signals to obtain a cross-correlation curve; determining a time corresponding to a peak of the cross-correlation curve as the time difference corresponding to the sensor.
[0143] In an embodiment of the present disclosure, the acquisition module 240 is further configured to perform: performing direction estimation on the sound source to obtain a direction angle between the sound source and the sensor array; and obtaining the time difference corresponding to the sensor based on the direction angle and a distance between the sensor and a reference sensor.
[0144] As to the apparatus in the above-mentioned embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and thus will not be described here in detail.
[0145] The apparatus for detecting a voice endpoint provided by the embodiments of the present disclosure acquires original signals collected by a plurality of sensors in a sensor array, performs beam forming on the plurality of original signals to obtain beam signals, performs voice endpoint detection on the beam signals to obtain an initial time of a voice signal, and obtains target times of the voice signal corresponding to the plurality of sensors based on the initial time. In this way, the original signals collected by the plurality of sensors can be subjected to beam forming to obtain beam signals, and the beam signals can be subjected to voice endpoint detection to obtain an initial time, so as to obtain target times corresponding to the plurality of sensors. Compared with the related art in which original signals collected by each sensor are mostly subjected to voice endpoint detection individually, the present solution only needs to perform voice endpoint detection on the beam signals, greatly reducing the number of times of voice endpoint detection, which helps to improve the efficiency of voice endpoint detection of the sensor array, and the signal-to-noise ratio of the beam signals is high, which helps to improve the accuracy of voice endpoint detection.
[0146] Figure 8 is a block diagram of an electronic device 300 according to an exemplary embodiment.
[0147] As shown in Figure 8 the above-mentioned electronic device 300 includes:
[0148] a memory 310 and a processor 320, a bus 330 connecting different components including the memory 310 and the processor 320, and the memory 310 stores a computer program which, when executed by the processor 320, implements the method for detecting a voice endpoint described in the embodiments of the present disclosure.
[0149] Bus 330 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration bus, a processor or local bus using any of a variety of bus architectures. By way of example, these architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0150] Electronic device 300 typically includes a variety of electronic device readable media. These media can be any available media that is located either internally or externally to electronic device 300, including both volatile and nonvolatile media, removable and non-removable media.
[0151] Memory 310 also can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 340 and / or cache memory 350. Electronic device 300 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 360 can be provided for reading from and writing to a non-removable, non-volatile magnetic media (not shown and typically called a "hard drive"). Figure 8 Although not shown, a magnetic hard disk drive can also be used for reading from and writing to non-removable, non-volatile magnetic media (e.g., a "hard drive"). Although not shown, a magnetic hard disk drive can also be used for reading from and writing to non-removable, non-volatile magnetic media (e.g., a "hard drive"). Figure 8 Although not shown, a magnetic hard disk drive can also be used for reading from and writing to non-removable, non-volatile magnetic media (e.g., a "hard drive"). Although not shown, a magnetic hard disk drive can also be used for reading from and writing to non-removable, non-volatile magnetic media (e.g., a "hard drive").
[0152] Program / utility 380 having a set (at least one) of program modules 370 can be stored in, for example, memory 310 by way of example, such program modules 370 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or a combination can include implementation of a network environment. Program modules 370 generally carry out the functions and / or methodologies of embodiments of the disclosure as described herein.
[0153] The electronic device 300 can also communicate with one or more external devices 390 such as a keyboard or pointing devices, a display 391, etc.; one or more devices that enable a user to interact with the electronic device 300; and / or one or more devices (e.g., net cards, modems, etc.) that enable the electronic device 300 to communicate with one or more other computing devices. Such communication can occur via an input / output (I / O) interface 392. Still yet, such communication can occur electronically over a network 393 such as a local area network (LAN) and / or a wide area network (WAN) such as the Internet. As an example, the network 393 can be enabled by a modem, network adapter, or other means for attaching the electronic device 300 to the network 393. The modem, network adapter or other means can be connected to the bus 330 via the I / O interface 392 and / or other hardware (not shown). It will be appreciated that the network 393 can also be implemented as a private network; and / or an electronic, electrical, and / or optical (e.g., infrared) connection. In some embodiments, the network 393 can be implemented using a wireless communications protocol or technology. Figure 8 As shown, the network adapter 393 communicates with the other components of the electronic device 300 via bus 330. It should be appreciated that although not shown, other hardware and / or software components could be used in conjunction with the electronic device 300. These include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.
[0154] The processor 320 performs various function applications and data processing by running programs stored in the memory 310.
[0155] It should be noted that the implementation process and technical principles of the electronic device of the present embodiment are described above in the description of the voice endpoint detection method of the present disclosure, which will not be repeated here.
[0156] The electronic device provided by the present embodiment can execute the voice endpoint detection method as described above, obtain the original signals collected by the plurality of sensors in the sensor array, perform beamforming on the plurality of original signals to obtain a beam signal, perform voice endpoint detection on the beam signal to obtain an initial time of the voice signal, and obtain a target time of the voice signal corresponding to the plurality of sensors based on the initial time. Thus, the original signals collected by the plurality of sensors can be beamformed to obtain a beam signal, and the beam signal can be subjected to voice endpoint detection to obtain an initial time, so as to obtain a target time corresponding to the plurality of sensors. Compared with the related art in which the original signals collected by each sensor are mostly subjected to voice endpoint detection individually, the present solution only needs to perform voice endpoint detection on the beam signal, thereby greatly reducing the number of times of voice endpoint detection, helping to improve the efficiency of voice endpoint detection of the sensor array, and the signal-to-noise ratio of the beam signal is high, which helps to improve the accuracy of voice endpoint detection.
[0157] In order to implement the above-mentioned embodiments, the present disclosure further provides a computer readable storage medium having stored thereon computer program instructions which, when executed by a processor, implement the steps of the voice endpoint detection method provided by the present disclosure.
[0158] Optionally, the computer readable storage medium can be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disc, and optical data storage device, etc.
[0159] To achieve the above-mentioned embodiments, the present disclosure further provides a computer program product comprising a computer program, characterized by, when the computer program is executed by a processor of an electronic device, realizing the voice endpoint detection method as described above.
[0160] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the present disclosure disclosed herein. It is intended that the present disclosure cover any and all variations of the present disclosure including those variations contained within the scope of the present disclosure, as well as those adaptations implementing features of the present disclosure that are within its skill of the art. It is intended that the specification and examples be considered exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0161] It should be understood that the present disclosure is not limited to the precise structures as herein described and illustrated in the drawings, and that various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the claims that follow.
Claims
1. A method of detecting a voice endpoint, characterized by, The method comprises: obtaining original signals collected by a plurality of sensors in a sensor array; performing beamforming on the plurality of original signals to obtain a beam signal; performing voice endpoint detection on the beam signal to obtain an initial time of a voice signal; obtaining a target time of the voice signal corresponding to the plurality of sensors based on the initial time; the voice endpoint detection on the beam signal to obtain the initial time of the voice signal comprises: performing frame processing on the beam signal to obtain a plurality of frames of beam signals; obtaining an energy-entropy ratio of each frame of beam signals; obtaining an energy-entropy ratio curve based on the energy-entropy ratios of the plurality of frames of beam signals, wherein the abscissa of the energy-entropy ratio curve is time and the ordinate is the energy-entropy ratio; obtaining a first intersection point and a second intersection point between the energy-entropy ratio curve and a first reference line, and obtaining a third intersection point and a fourth intersection point between the energy-entropy ratio curve and a second reference line, wherein the abscissa of the first intersection point is less than the abscissa of the second intersection point; in response to the abscissa of the third intersection point being less than the abscissa of the first intersection point, determining the abscissa of the third intersection point as an initial start time of the voice signal; and / or in response to the abscissa of the fourth intersection point being greater than the abscissa of the second intersection point, determining the abscissa of the fourth intersection point as an initial end time of the voice signal.
2. The method of claim 1, wherein, the beamforming on the plurality of original signals to obtain the beam signal comprises: performing orientation estimation on a sound source to obtain an orientation angle between the sound source and the sensor array; performing beamforming on the plurality of original signals based on the orientation angle to obtain the beam signal.
3. The method of claim 1, wherein, the beamforming on the plurality of original signals to obtain the beam signal comprises: in response to the original signals being wideband signals, performing beamforming on the plurality of original signals in the frequency domain to obtain a frequency domain beam signal; or in response to the original signals being single frequency signals, performing beamforming on the plurality of original signals in the time domain to obtain a time domain beam signal.
4. The method of claim 1, wherein, The first reference line and the second reference line are parallel to the horizontal axis, the ordinate of the points on the first reference line is a first threshold, the ordinate of the points on the second reference line is a second threshold, and the second threshold is less than the first threshold.
5. The method according to any one of claims 1-4, characterized in that, the target time of the voice signal corresponding to the plurality of sensors based on the initial time comprises: obtaining a time difference corresponding to the sensor; delaying the initial time by the time difference to obtain the target time of the voice signal corresponding to the sensor.
6. The method of claim 5, wherein, the time difference corresponding to the sensor comprises: performing cross-correlation processing on the original signal collected by the sensor and the beam signal to obtain a cross-correlation curve; determining the time corresponding to the peak value of the cross-correlation curve as the time difference corresponding to the sensor.
7. The method of claim 5, wherein, the time difference corresponding to the sensor comprises: performing orientation estimation on a sound source to obtain an orientation angle between the sound source and the sensor array; obtaining the time difference corresponding to the sensor based on the orientation angle and the distance between the sensor and a reference sensor.
8. A device for detecting a voice endpoint, characterized in that The method comprises: a collection module configured to perform obtaining original signals collected by a plurality of sensors in a sensor array; a processing module configured to perform beamforming on the plurality of original signals to obtain a beam signal; a detection module configured to perform voice endpoint detection on the beam signal to obtain an initial time of a voice signal; an obtaining module configured to obtain a target time of the voice signal corresponding to a sensor based on the initial time; the detection module is further configured to perform: frame processing on the beam signal to obtain a plurality of frames of beam signals; obtain an energy-entropy ratio of each frame of beam signals; obtain the initial time based on the energy-entropy ratio; wherein the initial time includes an initial start time and / or an initial end time; the detection module is further configured to perform: obtain an energy-entropy ratio curve based on the energy-entropy ratios of the plurality of frames of beam signals, wherein the abscissa of the energy-entropy ratio curve is time and the ordinate is the energy-entropy ratio; obtain a first intersection point and a second intersection point between the energy-entropy ratio curve and a first reference line, and obtain a third intersection point and a fourth intersection point between the energy-entropy ratio curve and a second reference line, wherein the abscissa of the first intersection point is less than the abscissa of the second intersection point; in response to the abscissa of the third intersection point being less than the abscissa of the first intersection point, determine the abscissa of the third intersection point as the initial start time of the voice signal; and / or in response to the abscissa of the fourth intersection point being greater than the abscissa of the second intersection point, determine the abscissa of the fourth intersection point as the initial end time of the voice signal.
9. The apparatus of claim 8, wherein, the processing module is further configured to perform: azimuth estimation of a sound source to obtain an azimuth angle between the sound source and the sensor array; perform beamforming on the plurality of original signals based on the azimuth angle to obtain the beam signal.
10. The apparatus of claim 8, wherein, the processing module is further configured to perform: in response to the original signal being a wideband signal, perform beamforming on the plurality of original signals in the frequency domain to obtain a frequency domain beam signal; or in response to the original signal being a single frequency signal, perform beamforming on the plurality of original signals in the time domain to obtain a time domain beam signal.
11. The apparatus of claim 8, wherein, The first reference line and the second reference line are parallel to the horizontal axis, the ordinate of the points on the first reference line is a first threshold, the ordinate of the points on the second reference line is a second threshold, and the second threshold is less than the first threshold.
12. The apparatus of any one of claims 8-11, wherein, the obtaining module is further configured to perform: obtain a time difference corresponding to the sensor; delay the initial time by the time difference to obtain the target time of the voice signal corresponding to the sensor.
13. The apparatus of claim 12, wherein, the obtaining module is further configured to perform: perform cross-correlation processing on the original signal collected by the sensor and the beam signal to obtain a cross-correlation curve; determine the time corresponding to the peak value of the cross-correlation curve as the time difference corresponding to the sensor.
14. The apparatus of claim 12, wherein, the obtaining module is further configured to perform: azimuth estimation of a sound source to obtain an azimuth angle between the sound source and the sensor array; obtain the time difference corresponding to the sensor based on the azimuth angle and the distance between the sensor and a reference sensor.
15. An electronic device, comprising: comprise: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to: Steps to implement the method of any of claims 1-7.
16. A computer-readable storage medium having stored thereon computer program instructions, wherein, The program instructions, when executed by a processor, implement steps of the method of any of claims 1-7.
Citation Information
Patent Citations
Apparatus and method for beamforming to obtain voice and noise signals
CN105532017A
Depression detection method based on microphone array
CN112349297A
Systems and methods for adaptive beamforming
US20220086564A1