Sound processing system, sound processing device, sound processing method, and program
The acoustic processing system enhances speech position detection in vehicles by analyzing spatial and acoustic features with a mathematical model, addressing noise interference and reducing hardware complexity, thereby improving estimation accuracy and stability.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- HONDA MOTOR CO LTD
- Filing Date
- 2022-03-29
- Publication Date
- 2026-04-17
AI Technical Summary
Existing voice recognition systems in vehicles face challenges in accurately detecting the speaker's position due to noise contamination and the dispersion of speech position over time, which complicates the estimation of speech position and requires significant hardware resources for complex calculations.
An acoustic processing system that analyzes the spatial spectrum of sound sources using multiple channels, employing a mathematical model to estimate the speaker's position based on spatial and acoustic features, and utilizes a random forest for high-speed estimation.
The system provides stable and accurate estimation of speech position, even in noisy environments, reducing computational load and improving estimation accuracy by using a mathematical model and acoustic features, while minimizing hardware requirements.
Smart Images

Figure 0007847460000004 
Figure 0007847460000005 
Figure 0007847460000006
Abstract
Description
Technical Field
[0001] The present invention relates to an acoustic processing system, an acoustic processing apparatus, an acoustic processing method, and a program.
Background Art
[0002] Voice recognition systems that extract voice commands from uttered voices in vehicles using voice recognition technology have become widespread. In such voice recognition systems, various devices and functions can be operated according to the extracted voice commands. For example, Patent Document 1 describes a voice recognition system that separates the voice of a speaker from voices input by a plurality of microphones installed in a vehicle. The voice recognition system includes a storage device that stores preset information indicating the sound source position of the speaker's voice, and a voice recognition unit that refers to the preset information of the speaker stored in the storage device, separates the speaker's voice from the voices input from the microphones, performs voice recognition, and recognizes voice commands.
[0003] Further, the voice recognition system described in Patent Document 1 further includes a sensor that detects the position of the speaker's seat, and the storage device stores preset information for each position of the speaker's seat, obtains the position of the speaker's seat from the sensor, searches for the preset information from the storage device based on the obtained position of the seat, and outputs it to the voice recognition unit. Based on the recognized voice commands, various navigation processes are performed.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] Some of the devices and functions to be controlled are related to the speaker's position. For example, when opening and closing a window, the target of the control is typically the window closest to the speaker. For voice commands indicating such operations, detection of the speaker's position, which is the sound source, is required. On the other hand, various types of noise are picked up by microphones installed in the vehicle's interior. In addition to engine noise and other driving sounds, music emitted by sound equipment and spoken voices can also be considered noise. Noise contamination can make it difficult to accurately detect the speaker's position, and the detected speech position tends to become dispersed and unstable over time. To improve the accuracy of speech position detection, it is conceivable to improve the estimation performance of speech position by using more microphones. However, performing complex calculations on multi-channel audio signals requires more hardware resources.
[0006] One of the objectives of the present invention is to economically provide an acoustic processing system, an acoustic processing device, an acoustic processing method, and a program that can stably detect the position of speech. [Means for solving the problem]
[0007] (1) The present invention has been made to solve the above problems, and one aspect of the present invention is a spatial spectrum analysis unit that analyzes the spatial spectrum of a sound source frame by frame over a predetermined period based on multiple channel acoustic signals, and an utterance position estimation unit that estimates the utterance position based on the analyzed spatial spectrum using at least a mathematical model that shows the relationship between the spatial spectrum of the sound source and the utterance position, which is the position of the speaker making the utterance. An acoustic feature analysis unit analyzes acoustic feature quantities for each frame based on the acoustic signals of the multiple channels and separate channels, Equipped with The mathematical model shows the relationship between acoustic features, spatial spectra, and speech position. The speech position estimation unit estimates the speech position based on the analyzed acoustic features and spatial spectrum. It is an acoustic processing system.
[0009] (2) Other aspects of the present invention include: (1) The acoustic processing system may include an acoustic feature analysis unit that analyzes the frequency characteristics of the acoustic signals of the separate channels and the frequency of zero intersections as the acoustic feature quantities.
[0010] (3) Other aspects of the present invention include:(2) The acoustic processing system may include an acoustic feature analysis unit that analyzes the power spectral density as the frequency characteristics of the acoustic signals of the separate channels.
[0011] (4) Other aspects of the present invention include: (3) The sound processing system wherein the spatial spectral analysis unit analyzes each of the predetermined candidate speech positions which are candidates for the speech position. A vector containing the element values The aforementioned spatial spectrum as The speech position estimation unit may calculate the confidence level that the speaker making the utterance is located for each candidate speech position.
[0012] (5) Other aspects of the present invention include: (4) The acoustic processing system is such that the mathematical model may be a random forest.
[0013] (6) Another aspect of the present invention includes a spatial spectrum analysis unit that analyzes the spatial spectrum of a sound source frame by frame over a predetermined period based on multiple channel acoustic signals, and a speech position estimation unit that estimates the speech position based on the analyzed spatial spectrum using at least a mathematical model that shows the relationship between the spatial spectrum of the sound source and the spatial distribution of the sound source. An acoustic feature analysis unit analyzes acoustic feature quantities for each frame based on the acoustic signals of the multiple channels and separate channels, Equipped with The mathematical model shows the relationship between acoustic features, spatial spectra, and speech position. The speech position estimation unit estimates the speech position based on the analyzed acoustic features and spatial spectrum. It is an acoustic processing device.
[0014] (7) Another aspect of the present invention relates to a computer. (6) It could also be a program to function as an acoustic processing device.
[0015] (8)Another aspect of the present invention is an acoustic processing method in an acoustic processing system, comprising: a spatial spectrum analysis step in which a spatial spectrum analysis unit analyzes the spatial spectrum of a sound source frame by frame over a predetermined period based on acoustic signals from multiple channels; and a speech position estimation step in which an utterance position estimation unit estimates the utterance position based on the analyzed spatial spectrum, using at least a mathematical model that shows the relationship between the spatial spectrum of the sound source and the utterance position, which is the position of the speaker making the utterance. An acoustic feature analysis step that analyzes acoustic features for each frame based on the acoustic signals of the multiple channels and separate channels, Execute The mathematical model shows the relationship between acoustic features, spatial spectra, and utterance location, and the utterance location estimation step estimates the utterance location based on the analyzed acoustic features and spatial spectra. It may also be an acoustic processing method. [Effects of the Invention]
[0016] According to the present invention, the position of speech can be reliably detected. (1) of the present invention, (6), (7) or (8) According to this method, the speech position corresponding to the spatial spectrum derived frame by frame from the acquired acoustic signal is estimated based on the relationship between known spatial spectra and speech positions. Therefore, a more stable speech position can be estimated than when the speech position is directly determined from the spatial spectrum. For example, the phenomenon of frames in which the speech position is determined appearing intermittently within a speech interval is mitigated. Furthermore, the speech position is estimated using acoustic features analyzed from acoustic signals of channels separate from the multiple channels used to estimate the spatial spectrum. Since the speech position is estimated by referring to the speech position dependence of the acoustic features of the spoken speech, the accuracy of speech position estimation can be improved even in noisy environments. In addition, by using a mathematical model, estimation accuracy can be ensured even when the acoustic signals of multiple channels and the acoustic signals of separate channels are not synchronized.
[0018] (2) According to this method, by using the frequency characteristics of the acoustic signal and the frequency of zero intersections as acoustic features, the characteristics of the spoken speech can be grasped more reliably. Therefore, the accuracy of estimating the speech position can be further improved.
[0019] (3) According to this method, the intensity for each frequency is analyzed by a simple calculation based on the discrete Fourier transform. Therefore, the accuracy of utterance location estimation can be improved economically.
[0020] (4)In this configuration, the confidence level of the speaker's location is calculated for each predetermined candidate utterance location, based on the spatial spectrum. Therefore, the computational load is reduced compared to calculating the spatial spectrum for every sound source location. Furthermore, unlike directly deriving the utterance location from the spatial spectrum, the influence of its absolute value on the likelihood of determining the utterance location is reduced. As a result, stable estimation of the utterance location becomes possible.
[0021] (5) In this configuration, the operations performed by the individual decision trees constituting the random forest are parallel, allowing for high-speed estimation of utterance positions. Because there is little dependence on specific input values used as explanatory variables, utterance positions can be estimated stably. [Brief explanation of the drawing]
[0022] [Figure 1] This is a schematic block diagram showing an example configuration of the sound processing system according to this embodiment. [Figure 2] This figure shows a first example of microphone placement. [Figure 3] This figure shows a second example of microphone placement. [Figure 4] This table illustrates voice operations in response to voice commands. [Figure 5] This flowchart shows an example of voice control processing according to this embodiment. [Figure 6] This is an explanatory diagram showing the first example of speech location detection. [Figure 7] This figure shows the first example of an estimated utterance location. [Figure 8] This is an explanatory diagram showing a second example of speech location detection. [Figure 9] This figure shows a second example of the estimated utterance position. [Figure 10] This flowchart shows a first example of conversion to an utterance-based utterance position according to this embodiment. [Figure 11] This is an explanatory diagram showing an example of the execution of the first example of conversion to utterance-based utterance position. [Figure 12]This flowchart shows a second example of conversion to an utterance-based utterance position according to this embodiment. [Figure 13] This is an explanatory diagram showing an example of determining the utterance position for each sub-segment. [Figure 14] This is an explanatory diagram showing an example of simultaneous speech detection. [Figure 15] This diagram illustrates patterns of simultaneous speech. [Figure 16] This figure shows examples of speech detection rates and speech location detection rates for each speech detection method. [Figure 17] This figure shows the first example of speech location detection rate and speech detection rate for different in-car environments. [Figure 18] This figure shows a second example of the speech location detection rate and speech detection rate for different in-car environments. [Figure 19] This figure shows the speech location detection rate and examples of speech detection rates for each operating state of the vehicle. [Figure 20] The figure shows examples of simultaneous speech detection rate, simultaneous speech detection accuracy, and single speech detection accuracy for each operating state of the vehicle. [Modes for carrying out the invention]
[0023] Embodiments of the present invention will be described below with reference to the drawings. First, an example of the configuration of the sound processing system S1 according to this embodiment will be described. Figure 1 is a schematic block diagram showing an example of the configuration of the sound processing system S1 according to this embodiment. The sound processing system S1 acquires sound signals from multiple channels and estimates the speaker's position based on the acquired sound signals. The sound processing system S1 identifies a voice command (instruction) as the content of the utterance transmitted by the acquired voice signal. In this application, the position of the speaker making the utterance may be referred to as the "utterance position". If the identified voice command is related to the utterance position, the sound processing system S1 executes the processing instructed according to the identified voice command based on the estimated utterance position.
[0024] The instructed processing includes controlling the operation of equipment connected to the sound processing system S1. The controlled equipment may include components of the sound processing system S1, or it may include other equipment not belonging to the sound processing system S1. In the following description, the case where the sound processing system S1 is configured as part of an in-vehicle system installed in a vehicle and has the function of operating various equipment installed in the vehicle via voice commands. In this application, "vehicle-mounted" means not only being actually installed in a vehicle, but also being primarily intended for use within a vehicle, or being suitable for such use. In other words, "vehicle-mounted" is not intended to limit the embodiment to use when actually installed in a vehicle, nor to exclude use when not installed in a vehicle.
[0025] The acoustic processing system S1 is composed of one or more devices. In the example shown in Figure 1, the acoustic processing system S1 comprises two acoustic processing devices 10 and 20, three microphones 30, and a controlled device 40. The three microphones 30 are distinguished by sub-numbers 30-1 to 30-3. The acoustic processing device 10 and the acoustic processing device 20, and the acoustic processing device 10 and each controlled device 40 are connected to enable the transmission and reception of various types of data wirelessly or via wired connections. The acoustic processing devices 10 and 20 and the controlled device 40 can be connected, for example, using a CAN (Controller Area Network).
[0026] The acoustic processing device 10 acquires acoustic signals from microphones 30-1 and 30-2, respectively. The acoustic processing device 10 analyzes the spatial spectrum of the sound source for each frame over a predetermined period from the acquired two-channel acoustic signals. The acoustic processing device 10 uses a mathematical model that shows the relationship between the spatial spectrum, the pair of acoustic features, and the speaker's position, and estimates the speaker's position based on the analyzed spatial spectrum and acoustic features. The acoustic processing device 10 acquires speech segment information indicating the speech segment in which the speech was uttered and speech information indicating the content of the utterance from the acquired acoustic signals. The acoustic signals used to acquire the speech information may be acoustic signals obtained by performing spatial filtering (described later) on the two-channel acoustic signals.
[0027] The sound processing device 10 determines the utterance position within a speech segment based on the length of the speech segment in which the utterance position was estimated, for each estimated utterance position. When the acquired speech information indicates a voice command related to the utterance position, the sound processing device 10 executes processing according to the voice command instructed by the utterance information, based on the determined utterance position. The voice command may instruct the operation control of the controlled device 40. The sound processing device 10 outputs a control signal indicating the operation mode to be instructed to the controlled device 40. As will be described later, in relation to the utterance position, the necessity of operation control may differ depending on the utterance position, depending on the voice command or the controlled device 40. Depending on the utterance position, the processing instructed by the voice command or the manner of operation control instructed to the controlled device 40 may differ.
[0028] The acoustic processing unit 20 acquires an acoustic signal from the microphone 30-3 and analyzes acoustic features from the acquired acoustic signal frame by frame. The acoustic processing unit 20 notifies the acoustic processing unit 10 of the acoustic features obtained through the analysis. The acoustic processing unit 20 and the microphone 30-3 may be primarily intended for acquiring acoustic signals that include noise components such as engine noise. The acoustic processing unit 20 outputs the acquired acoustic signal to the acoustic processing unit 10 as a reference signal.
[0029] Microphones 30-1 to 30-3 each capture sound arriving at the device and are equipped with an electroacoustic transducer (actuator) that converts the sound pressure of the captured sound, which is an electrical signal indicating its intensity, into an acoustic signal. Microphones 30-1 and 30-2 output their converted acoustic signals to the acoustic processing device 10 and the acoustic processing device 20, respectively. The acoustic signal output to the acoustic processing device 20 is used to remove noise components. Microphone 30-3 outputs its converted acoustic signal to the acoustic processing device 20.
[0030] Next, we will describe examples of the placement of microphones 30-1 to 30-3. In the examples in Figures 2 and 3, microphones 30-1 and 30-2 are positioned symmetrically in front of the driver's seat and passenger seat, with the middle of the space between them in the vehicle interior. In the illustrated examples, the front corresponds to the left side of the diagram. The distance between microphones 30-1 and 30-2 is, for example, 5 to 10 cm. The acoustic signal acquired at this position contains a relatively large amount of the voice of the driver seated in the driver's seat or the voice of the passenger seated in the passenger seat. In the example in Figure 2, microphone 30-3 is positioned at the rear right end of the rear seat. In the example in Figure 3, it is positioned at the rear center of the rear seat. The acoustic signal acquired at these positions contains a relatively large amount of noise components such as engine noise and friction noise between the road surface and the wheels.
[0031] The controlled device 40 is a device that is subject to operation control based on voice commands. The controlled device 40 operates according to the operation mode instructed by the control signal input from the sound processing device 10. In the example in Figure 1, the controlled device 40 includes an audio device 42 (e.g., a car audio system), an air conditioner 44 (e.g., an air conditioner), a window opener 46 (e.g., a power window), and a steering heater 48 (e.g., a steering heater).
[0032] Next, an example of the configuration of the sound processing device 10 according to this embodiment will be described. The sound processing device 10 includes an A / D conversion unit 112, a communication unit 114, and a control unit 120. The sound processing device 10 also includes an input / output unit (not shown) for inputting and outputting various types of data wirelessly or via wired connection using a predetermined input / output method with the sound processing device 20 and the controlled device 40.
[0033] The A / D (Analog-to-Digital) conversion unit 112 samples the analog audio signals input from microphones 30-1 and 30-2, respectively, at a predetermined sampling frequency and converts them into digital audio signals. The A / D conversion unit 112 outputs the converted audio signals to the control unit 120. The A / D conversion unit 112 is configured, for example, to include an A / D converter.
[0034] The communication unit 114 wirelessly connects to the communication network NW and communicates with other devices via the communication network NW. A voice recognition server may be designated as the recipient device. The designated voice recognition server may form a cloud 50 together with other server devices. The communication unit 114 is configured to include a communication interface that enables communication using a predetermined communication method. The communication method is, for example, 5G(5 th Any of the following may be used: General Mobile Communication System (5th generation mobile communication system), LTE-A (Long Term Evolution - Advanced), IEEE 802.11, etc.
[0035] The control unit 120 performs various calculations to realize and control the functions of the sound processing device 10. The control unit 120 may be realized by dedicated components, or it may be realized as a computer equipped with a processor and storage media such as ROM (Read Only Memory) and RAM (Random Access Memory). The control unit 120 may be configured as, for example, an ECU (Engine Control Unit). The processor reads a predetermined program stored in ROM beforehand, loads the read program into RAM, and uses the RAM storage area as a working area. The processor realizes the functions of the control unit 120 by executing the processes instructed by various instructions written in the read program. The realized functions may include the functions of each part described later. In the following description, the execution of processes instructed by instructions written in a program may be referred to as "executing a program" or "program execution." The processor is, for example, a CPU (Central Processing Unit).
[0036] The control unit 120 is comprised of a spatial spectral analysis unit 122, a speech position estimation unit 124, a speech information acquisition unit 126, a speech position processing unit 128, a command processing unit 130, and a spatial filtering unit 134. The spatial spectrum analysis unit 122 calculates a spatial spectrum for each frame of a predetermined time length from the acoustic signals input for each channel from the A / D conversion unit 112. The spatial spectrum is a spatial feature that shows the distribution of intensity depending on the sound source location. The spatial spectrum analysis unit 122 outputs the calculated spatial spectrum to the speech location estimation unit 124. In this embodiment, element values that make up the spatial spectrum are calculated for each candidate sound source location (hereinafter referred to as "candidate location"). The spatial spectrum is represented as a vector containing the element values for each candidate location. The positions of individual seats in a vehicle may be used as candidate locations. For example, if there are 5 seats in the vehicle, the spatial spectrum is represented as a 5-dimensional vector.
[0037] The spatial spectrum analysis unit 122 calculates, for example, a MUSIC (Multiple Signal Classification) spectrum as a spatial spectrum. The MUSIC spectrum can be calculated using the following procedure. The spatial spectrum analysis unit 122 performs a discrete Fourier transform for each frame and calculates the transformation coefficients converted to the frequency domain. The spatial spectrum analysis unit 122 generates an input vector for each frequency that includes the transformation coefficients for each channel as elements. The spatial spectrum analysis unit 122 calculates the spectral correlation matrix R as the expected value of the matrix obtained by transposing the generated input vector and the product of that transposed vector and the input vector. sp It is calculated as follows.
[0038]
number
[0039] In equation (1), * represents the complex conjugate transpose operator. E(...) represents the expected value of ... The spatial spectral analysis unit 122 calculates the spectral correlation matrix R sp Solve the eigenvalue problem of λ i and eigenvector e i The eigenvalue λ is calculated. i and eigenvector e i The number of sets corresponds to the number of channels. The spatial spectral analysis unit 122 uses, for example, (2) to determine the number of sound sources that can be detected (hereinafter referred to as the "number of detectable sound sources"), a pre-set transfer function vector d(θ), and an eigenvector e i The element values P(θ) of the spatial spectrum (hereinafter referred to as "frequency-specific spatial spectrum") are calculated for each frequency using this method. In this embodiment, the number of sound sources to be detectable may be set to 1. The transfer function vector d(θ) is a vector whose element values are the transfer functions from the candidate position θ to the positions of each microphone 30-1, 30-2 (hereinafter referred to as "receiving positions").
[0040]
number
[0041] In Equation (2), |…| represents the absolute value of …. M is a preset positive integer value less than N, indicating the number of detectable sound sources. K is the number of eigenvectors e i held by the sound source localization unit 121. M is a positive integer not exceeding N. N is a positive integer corresponding to the number of channels (in the example of FIG. 1, N = 2). The spatial spectrum analysis unit 122 calculates the signal-to-noise ratio (S / N ratio) for each frequency band based on the acoustic signals of each channel, and selects a frequency band k in which the calculated S / N ratio is higher than a preset threshold value. The spatial spectrum analysis unit 122 calculates the eigenvalues λ i for each frequency in the selected frequency band k, and uses the square root of the maximum eigenvalue λ max (k) as the element value P k (θ) of the spatial spectrum, and performs weighted addition among the frequency bands k to calculate the element value P ext (θ) of the extended spatial spectrum shown in Equation (3).
[0042] [Number]
[0043] In Equation (3), Ω represents a set of frequency bands. |Ω| represents the number of frequency bands in the set. Therefore, the extended spatial spectrum P ext (θ) has relatively few noise components and reflects the characteristics of the frequency bands in which the value of the frequency band spatial spectrum P k (θ) is large. The spatial spectrum analysis unit 122 adopts the vector including the element value P ext (θ) of this extended spatial spectrum as the above-described spatial spectrum.
[0044] The speech position estimation unit 124 receives the spatial spectrum from the spatial spectrum analysis unit 122 and acoustic features from the acoustic processing unit 20 for each frame. As will be described later, the acoustic features represent the acoustic characteristics of the acoustic signal as vectors. The acoustic features include power spectral density (PSD) and the number of zero crossings (ZC). For example, if the number of samples in the sub-interval of each frame analyzed at once by the acoustic processing unit 20 is 512, the power spectral density is represented as a 256-dimensional vector. The number of zero crossings is represented as a scalar value. In that case, the acoustic features are represented as a 257-dimensional vector. The speech position estimation unit 124 concatenates the spatial spectrum and acoustic features to construct an input vector for a predetermined mathematical model.
[0045] The speech position estimation unit 124 calculates an output vector for each frame based on the input vector, using a mathematical model that shows the relationship between the input vector and the output vector. The calculated output vector represents information about the estimated speech position (which may be referred to as "estimated speech position information" in the following explanation). The input vector and output vector correspond to the explanatory variable and the dependent variable, respectively. The output vector includes a confidence score for each predetermined speaker position as an element value. The confidence score indicates the probability that the speaker is located at that speaker position and is expressed as a real number. The higher the confidence score, the greater the probability that the speaker is located at that position. The range of values for each confidence score is normalized to a predetermined range (for example, between 0 and 1). As an example, if there are 5 seats in the vehicle, the output vector is represented as a 5-dimensional vector. The output vector can also be viewed as information showing the spatial distribution of the probability that the speaker is located in that frame. For example, the speech position with the highest calculated confidence level, which is higher than a predetermined confidence threshold, is estimated as the speech position (frame-based speech position) where the speaker is located in that frame. The confidence threshold only needs to be significantly larger than the expected value of the confidence level selected by chance. The speech position estimation unit 124 outputs the estimated speech position information to the speech information processing unit 128.
[0046] The speech position estimation unit 124 can use, for example, a random forest as a mathematical model. A random forest is a type of ensemble machine learning model that includes multiple decision trees as weak learners. The output from the random forest is the average value of the outputs from the multiple decision trees. A decision tree is a machine learning model that has a tree structure with multiple nodes, one node as the root, multiple branches from each node, and separate nodes continuing from each branch until the end. Each node is assigned an attribute, and each branch is assigned a value for that attribute. For each decision tree, for example, the confidence level of each speech position with respect to the input vector can be obtained as an output value. Alternatively, each decision tree may be configured to obtain an index indicating any of the speech positions as an output value for the input vector. In that case, the index indicating the speech positions from multiple decision trees may be used as the estimated speech position information, which is the output from the random forest.
[0047] The speech information acquisition unit 126 receives an acoustic signal from the spatial filtering unit 134. The acoustic signal input to the speech information acquisition unit 126 is derived based on a two-channel acoustic signal acquired via the A / D conversion unit 112. The speech information acquisition unit 126 transmits the input acoustic signal to the cloud 50 via the communication unit 114 and the communication network NW. The speech information acquisition unit 126 receives from the speech recognition server, which constitutes the cloud 50, for each utterance detected from the transmitted acoustic signal, associating speech section information indicating the utterance section with speech content information indicating the utterance content. The speech section information indicates the start time when a single utterance began and the end time when it ended. The speech content information describes the information recognized from the utterance as text expressed in natural language. The speech information acquisition unit 126 outputs the received speech section information to the speech position processing unit 128 and the speech content information to the command processing unit 130.
[0048] The speech position processing unit 128 receives estimated speech position information from the speech position estimation unit 124 and speech segment information from the speech information acquisition unit 126 for each frame. The speech position processing unit 128 identifies the section in the speech segment information where the speech position is estimated, for each predetermined speech position, as a position-specific speech segment. For example, the speech position processing unit 128 can identify the speech position with the highest confidence level indicated in the input speech segment information and identify the section of that frame as a position-specific speech segment for the identified speech position. When determining the utterance position for each frame, the speech position processing unit 128 may identify a sub-interval that is part of the utterance section and spans multiple frames including that frame, and identify the utterance position with the highest reliability in the entire sub-interval as the utterance position for that frame.
[0049] Furthermore, if the speech segment information is represented by an index of the speech position for each frame, the speech position processing unit 128 can identify that frame as a speech segment by the position of the speech position indicated by that index. Here, when identifying a sub-interval corresponding to each frame, the speech position processing unit 128 can identify the speech position with the highest frequency among the speech positions for each frame in the sub-interval (by majority vote), and then identify the frame corresponding to that sub-interval as a speech segment by the position of the identified speech position.
[0050] The speech position processing unit 128 determines the speech position (speech-based speech position) in a speech segment based on the position-specific speech segments identified within that speech segment. The speech position processing unit 128 can determine the speech position in a speech segment as the speech position where the ratio within the speech segment is greater than or equal to a predetermined ratio and the position-specific speech segment is the longest. The speech position processing unit 128 outputs speech position information indicating the defined speech position to the command processing unit 130.
[0051] There may be multiple utterance locations where the ratio of utterance sections by location within an utterance section exceeds a predetermined ratio. In such cases, the utterance location processing unit 128 may determine it as simultaneous utterance without specifying one utterance location. The predetermined ratio is set in the utterance location processing unit 128 in advance to be greater than or equal to the probability of an utterance location being selected by chance, and more preferably to be significantly greater than that ratio. The reciprocal of the number of seats in the vehicle may be used as the probability of selection by chance.
[0052] The speech position processing unit 128 converts the defined speech position and related information into a predetermined format required by the command processing unit 130. For example, when a voice command is used that requires identification of the driver's seat (front right, FRF) and other locations as speech positions, the speech position processing unit 128 includes driver's seat identification information indicating whether or not it is the driver's seat in the speech position information. When a voice command is used that requires identification of the driver's seat, passenger seat (front left, FL), and other locations as speech positions, the speech position processing unit 128 may include front seat identification information indicating either the driver's seat, passenger seat, or other locations in the speech position information. The speech position processing unit 128 may also include simultaneous utterance identification information indicating either simultaneous utterance or other (single utterance) in the speech position information.
[0053] The command processing unit 130 receives speech content information from the speech information acquisition unit 126 and speech position information from the speech position processing unit 128 for each utterance. The command processing unit 130 refers to a command list (not shown) that has been set up in advance on its own part and determines whether or not the utterance content shown in the utterance content information contains a voice command. The command processing unit 130 may also refer to the command list and determine whether or not the voice command included in the utterance content is a voice command related to the utterance position (which may be referred to as a "utterance position-related voice command" in the following description).
[0054] The command list corresponds to data that, for example, indicates keywords, operation mode information, and target device information for each voice command. The keywords may include one or more words relating to either the target device, the operation mode, or both, indicated by each voice command. The operation mode information may include information such as the operation mode, or its elements such as operation characteristics, the object being operated on, and the state being the control target. For utterance position-related voice commands, an operation mode may be set for each utterance position identified by that voice command. The target device information may include one or more pieces of information that identify or help identify the target device, such as type, name, model number, IP (Internet Protocol) address, and MAC (Media Access Control) address. For utterance position-related voice commands, a separate target device may be set for each utterance position identified by that voice command.
[0055] The command processing unit 130 refers to the command list and searches for a voice command in which the words or phrases contained in the input utterance information match the set keywords. The command processing unit 130 may perform known morphological analysis on the text representing the utterance information to determine the words or phrases and parts of speech expressed in the text. The command processing unit 130 may also compare each independent word among the defined words or phrases with one of the commands described in the command list. Words or phrases whose part of speech is a noun, verb, adjective, or adverb are identified as independent words.
[0056] When the command processing unit 130 detects a voice command that matches a keyword, it reads out the target device and operation mode information related to that voice command. If the detected voice command is a voice command related to speech position and a target device is set for each speech position, the command processing unit 130 identifies the target device corresponding to the speech position indicated by the speech position information. If the detected voice command is a voice command related to speech position and an operation mode is set for each speech position, the command processing unit 130 identifies the operation mode corresponding to the speech position indicated by the speech position information.
[0057] The command processing unit 130 generates control information to instruct the identified target device to operate in a specified operating mode, and transmits the generated control signal to one of the controlled devices 40 as the target device. The controlled device 40 waits for the control signal from the sound processing device 10 and operates according to the operating mode instructed by the input control signal.
[0058] If the utterance position information includes simultaneous utterance identification information indicating simultaneous utterance, the command processing unit 130 may discard the utterance content information and output guidance instructions to the sound device 42 to instruct individual speakers to repeat their utterances. When the sound device 42 receives guidance instructions from the command processing unit 130, it plays guidance audio that conveys guidance information for guiding individual repeat utterances. The guidance audio may convey the following messages as guidance information: for example, "Please repeat your words one by one in turn," or "Please speak again so that two or more people do not overlap," etc.
[0059] The spatial filtering unit 134 receives two channels of acoustic signals from the A / D conversion unit 112 and speech position information from the speech position estimation unit 124. The spatial filtering unit 134 performs spatial filtering on each channel of acoustic signals based on the speech position information to generate a single channel of acoustic signal after filtering. The spatial filtering unit 134 outputs the generated acoustic signal to the speech information acquisition unit 126 as an acoustic signal for acquiring speech information.
[0060] The spatial filtering unit 134 determines, for each channel, a filter coefficient that has a directivity where the gain is higher in the direction of the speech position indicated by the speech position information than in other directions, in spatial filtering. For example, the spatial filtering unit 134 pre-sets filter coefficients for each speech position and identifies the filter coefficient corresponding to the speech position indicated by the speech position information. The spatial filtering unit 134 performs filtering on the acoustic signal of the corresponding channel using the identified filter coefficient, and the added signal obtained by adding the processed acoustic signals between channels is obtained as the filtered acoustic signal.
[0061] The spatial filtering unit 134 can use, for example, known methods for controlling directivity such that the gain in the direction of the identified sound source location is higher than in other directions, such as the delayed sum method or the filter-and-thumb beamformer, as spatial filtering processing. The spatial filtering unit 134 may also use a sound source separation processing method such as the GHDSS (Geometric High-order Decorrelation-based Source Separation) method to separate or extract sound from the direction of the identified sound source location from sound from other directions, and acquire a one-channel acoustic signal for acquiring speech information arriving from the speech location identified from the two-channel acoustic signal.
[0062] Next, an example of the configuration of the sound processing device 20 according to this embodiment will be described. The sound processing device 20 includes an A / D conversion unit 212 and a control unit 220. The sound processing device 20 also includes an input / output unit (not shown) for inputting and outputting various types of data wirelessly or via wired connection using a predetermined input / output method with the sound processing device 10 and the controlled device 40.
[0063] The A / D conversion unit 212 samples the analog audio signals input from microphones 30-1 to 30-3 at a predetermined sampling frequency and converts them into digital audio signals. The A / D conversion unit 212 outputs the converted audio signals for each channel to the control unit 220. The A / D conversion unit 212 is configured, for example, to include an A / D converter. In the following description, the channels corresponding to microphones 30-1 to 30-3 will be referred to as channels 1 to 3.
[0064] The control unit 220 performs various calculations to realize and control the functions of the sound processing device 10. The control unit 220 may be realized by dedicated components, or it may be realized as a computer equipped with a processor and a storage medium. The control unit 220 may be configured as, for example, an ECU (Engine Control Unit). The processor reads a predetermined program stored in ROM beforehand, loads the read program into RAM, and uses the RAM storage area as a working area. The processor executes the processing instructed by various instructions written in the read program to realize the functions of the control unit 220. The realized functions may include the functions of each part described later.
[0065] The control unit 220 includes an acoustic feature analysis unit 222, a low-pass filter 224, and a noise reduction unit 226. The acoustic feature analysis unit 222 analyzes the acoustic features of the acoustic signal of channel 3 input from the A / D conversion unit 212 for each frame of a predetermined length. The power spectrum and zero crossing number are associated with the acoustic features calculated for each frame and output to the speech position estimation unit 124 of the acoustic processing device 10. The acoustic feature analysis unit 222 is composed of a frequency analysis unit 222a and a zero intersection analysis unit 222b.
[0066] The frequency analysis unit 222a analyzes features that show frequency characteristics as acoustic features. The frequency analysis unit 222a calculates power spectral density as an acoustic feature. Power spectral density is power per unit frequency. The frequency analysis unit 222a can perform a discrete Fourier transform on the input acoustic signal for each frame to calculate the transformation coefficients in the frequency domain, and calculate the power spectrum as the absolute value of the square of the obtained transformation coefficients. Power spectral density can serve as a clue to determine whether or not speech is being spoken. The zero intersection analysis unit 222b detects zero intersections in the input acoustic signal for each frame. A zero intersection is the point in time when the signal value of each sample constituting the acoustic signal changes from a positive value to a negative value, or from a negative value to a positive value. The zero intersection analysis unit 222b defines the number of zero intersections detected for each frame as the zero intersection count. The zero intersection count is also a type of acoustic feature.
[0067] The low-pass filter 224 primarily passes low-frequency components below a predetermined cutoff frequency (e.g., 50-200 Hz) from the acoustic signals of channels 1-3 input from the A / D conversion unit 212. These low-frequency components mainly consist of noise and contain almost no spoken voice components. The low-pass filter 224 outputs an acoustic signal indicating the low-frequency components that passed through each channel to the noise reduction unit 226.
[0068] The noise reduction unit 226 receives an acoustic signal representing low-frequency components from the low-pass filter 224. The noise reduction unit 226 uses the acoustic signal from channel 3 as a noise signal to perform noise reduction processing and removes noise components contained in the acoustic signals from channels 1 and 2. As a noise reduction process, the noise reduction unit 226 implements, for example, active noise control (ANC).
[0069] To achieve ANC, the noise reduction unit 226 is connected to a speaker (not shown) for presenting the canceled sound and includes an adaptive filter. The adaptive filter is used to estimate filter coefficients that indicate the transmission path of the canceled sound from the speaker to microphones 30-1 and 30-2, respectively. The adaptive filter determines the filter coefficients so that the intensity of the acoustic signal representing the low-frequency component extracted for channels 1 and 2 approximates (minimizes) to 0. The noise reduction unit 226 generates a canceled sound signal by performing a convolution operation on the noise signal using the filter coefficients determined by the adaptive filter. When determining the filter coefficients, for example, the LMS (Least Mean Square) method can be used. The noise reduction unit 226 supplies the generated canceled sound signal to the speaker. The speaker emits canceled sound based on the canceled sound signal supplied from the noise reduction unit 226. Therefore, in microphones 30-1 and 30-2, the canceling sound arriving from the speaker and the noise components arriving from the noise source cancel each other out, resulting in the acquisition of an acoustic signal with the noise components removed or reduced.
[0070] As described above, the operation of the sound processing devices 10 and 20 is individually controlled by the control units 120 and 220, respectively. Synchronization of operation between the sound processing devices 10 and 20 is not guaranteed. That is, a time difference and fluctuation may occur on a sample-by-sample basis between the acoustic signal used for spatial spectrum analysis in the spatial spectrum analysis unit 122 and the acoustic signal used for acoustic feature analysis in the acoustic feature analysis unit 222. This time difference is difficult to distinguish from the difference in phase difference between channels due to the speech position (sound source position). Therefore, it is not practical to directly use the two-channel acoustic signal acquired by the sound processing device 10 and the one-channel acoustic signal acquired by the sound processing device 20 to calculate the spatial spectrum. Furthermore, even if the spatial spectrum calculated frame by frame from the two-channel acoustic signal and the acoustic feature calculated frame by frame from a different one-channel acoustic signal are combined, it does not necessarily provide clues to improving the accuracy of speech position estimation.
[0071] In this embodiment, estimation accuracy can be improved by using a mathematical model that shows the relationship between the known spatial spectrum and the utterance position as a relationship between explanatory and dependent variables. By further referencing acoustic features as explanatory variables, the utterance position can be estimated using the variation in acoustic features due to the utterance position as a clue, even when synchronization with the spatial spectrum is not guaranteed, thus improving estimation accuracy.
[0072] The speech position estimation unit 124 has a set of parameters for a mathematical model pre-configured. The sound processing device 10 may include a model learning unit (not shown) for calculating the parameter set through learning. The model learning unit has training data pre-configured. The training data consists of a large number of training sets. Each training set includes known input vectors that serve as explanatory variables and output vectors that serve as objective variables, and these are associated with each other. For example, as the objective variable, an output vector representing a certain speech position may be set, where the element value of the dimension corresponding to that speech position is 1, and the element values of the dimension corresponding to all other speech positions are 0.
[0073] The model learning unit recursively updates the parameter set so that the overall difference between the estimated value obtained by performing calculations using a mathematical model on a known input vector and the corresponding output vector is small. The model learning unit can use one of the following as the loss function indicating the magnitude of the difference: for example, the sum of squared errors, cross-entropy, or a linear combination of any two. The model learning unit can use, for example, the gradient method to update the parameter set.
[0074] When a random forest is used as the mathematical model, the model learning unit may perform the following steps during learning: (1) Random sampling is performed from the entire training data using the bootstrap method to classify it into B (B is an integer greater than or equal to 2) subsamples. Each subsample contains multiple training sets. (2) B decision trees are generated using each subsample as training data. (3) For each decision tree, nodes are generated by performing the following steps until a predetermined number of nodes is reached: (3-1) Some of the explanatory variables of the training data are randomly selected. (3-2) The explanatory variable that best classifies the training data from the selected explanatory variables, and the threshold used for that classification, are set as the threshold used for classifying new nodes. By using randomly sampled training data and randomly selected explanatory variables, a group of decision trees with low correlation is generated. This enables fast learning.
[0075] Next, we will explain examples of voice operations performed by the command processing unit 130. Figure 4 is a table illustrating voice operations corresponding to voice commands. In the illustrated example, saying "Play music" instructs music playback regardless of the utterance position. In this case, distinction based on utterance position is not required. As keywords related to the voice command, it is sufficient to detect that the utterance contains phrases such as "music" and "play". For each type of voice operation, we will show examples of utterances, the functions for each utterance position, and the necessary functions. Five utterance positions are listed: driver's seat (FR), passenger seat (FL), rear right (RR), rear middle (RM), and rear left (RR).
[0076] Voice control types include operations independent of the speaking position, operations related to the speaking position, safety-related operations, operations that may distract the driver due to operations by passengers, and simultaneous speaking. An operation independent of the utterance position refers to the execution of a predetermined function regardless of the utterance position, in accordance with the recognized voice command. In the illustrated example, in response to the utterance "Play music," the command processing unit 130 instructs the audio device 42 to play music regardless of the utterance position. In this case, distinction based on the utterance position is not required. The command processing unit 130 can recognize a voice command related to music playback by detecting that the utterance contains keywords such as "music" and "play."
[0077] A speech-location-related operation refers to the implementation of a predetermined function in accordance with a recognized speech-location-related voice command, depending on the speech location. In the illustrated example, in response to the utterance "Lower the air conditioner temperature," the command processing unit 130 instructs the air conditioner 44 to lower the temperature at the speech location. However, if the speech location is the rear seat, this instruction is ignored. In this case, it is necessary to refer to the speech location information and distinguish at least between the driver's seat, the passenger seat, and other locations. The command processing unit 130 can recognize a voice command related to lowering the temperature by detecting from the utterance content that the words, for example, "air conditioner," "temperature," and "lower" are included as keywords.
[0078] Safety-related operations are those that are performed according to voice commands uttered by the driver, and whose execution by voice commands from passengers other than the driver is restricted. In the illustrated example, in response to the utterance "Open the window (other than my seat)" from the driver's seat, the command processing unit 130 instructs the window opener 46 to open the specified window other than the driver's seat. However, if the utterance is made from a seat other than the driver's seat, it is ignored. In this case, it is necessary to refer to the utterance location information and distinguish at least between the driver's seat and other locations. In other words, this voice command is determined to be invalid from other locations. The command processing unit 130 can detect keywords from the utterance content, for example, words related to the location of each seat other than the driver's seat (e.g., "passenger seat," "rear right," "rear left," etc.), "window," and "open," and recognize a voice command related to opening a window at that location. Voice commands related to safety-related operations can also be considered as utterance location-related voice commands.
[0079] An operation that may distract the driver due to an operation by a passenger refers to an operation that is performed according to a voice command uttered by the driver, and whose execution by a voice command uttered by a passenger other than the driver is restricted. In the illustrated example, in response to the utterance "Steering heater, on" from the driver's seat, the command processing unit 130 instructs the steering heater 48 to be heated. However, if the utterance is made from a seat other than the driver's seat, it is ignored. In this case, it is necessary to refer to the utterance location information and distinguish at least between the driver's seat and other locations. In other words, this voice command is determined to be invalid from other locations. The command processing unit 130 can detect from the utterance content that the words, for example, "steering heater" and "on" are included as keywords, and recognize the voice command related to heating the steering heater 48. Voice commands related to operations that may distract the driver due to an operation by a passenger can also be considered utterance location-related voice commands.
[0080] Simultaneous utterance refers to a state in which utterances are made from multiple seats within a single utterance interval. In this case, the command processing unit 130 rejects the detected voice command, even if a voice command can be detected from the utterance content. The command processing unit 130 outputs guidance instructions to the audio device 42 and plays guidance audio to guide individual speakers to repeat their utterances. In this case, it is sufficient to detect simultaneous utterance by referring to the utterance position information.
[0081] Next, an example of voice control processing according to this embodiment will be described. Figure 5 is a flowchart showing an example of voice control processing according to this embodiment. (Step S102) The spatial spectrum analysis unit 122 calculates the spatial spectrum frame by frame based on the two-channel acoustic signals input from microphones 30-1 and 30-2. (Step S104) The frequency analysis unit 222a analyzes the power spectral density frame by frame based on the one-channel acoustic signal input from the microphone 30-3. (Step S106) The zero intersection analysis unit 222b analyzes zero intersections frame by frame based on the one-channel acoustic signal input from the microphone 30-3 and counts the number of zero intersections. (Step S108) The speech position estimation unit 124 constructs an input vector for each frame that includes power spectral density, power spectral density, and zero intersection as elements. The speech position estimation unit 124 uses a mathematical model to calculate an output vector from the constructed input vector that includes the confidence level for each speech position as an element, and estimates the speech position information (frame-based speech position).
[0082] (Step S110) The spatial filtering unit 134 performs spatial filtering on the two-channel acoustic signals input from microphones 30-1 and 30-2, directs its directivity towards the estimated speech position, and acquires a one-channel acoustic signal. (Step S114) The speech information acquisition unit 126 transmits the acquired 1-channel acoustic signal to the cloud 50, and acquires speech segment information and speech content information from the cloud 50 for each utterance. (Step S116) The speech position processing unit 128 determines a position-specific speech segment based on the speech position information estimated for each frame in the speech segment indicated in the acquired speech segment information. The speech information acquisition unit 126 determines the speech position with the most position-specific speech segments as the speech position (speech-based speech position) for that speech segment.
[0083] (Step S118) The command processing unit 130 refers to the command list and determines whether or not it can detect a voice command from the utterance content shown in the acquired utterance content information. If it determines that it can be detected (Step S118 YES), it proceeds to the process in Step S120. If it determines that it cannot be detected (Step S118 NO), it terminates the process shown in Figure 5. (Step S120) The command processing unit 130 determines whether the utterance position information indicates simultaneous utterance. If it determines that it indicates simultaneous utterance (Step S120 YES), the process proceeds to step S122. If it determines that it indicates single utterance (Step S120 NO), the process proceeds to step S124. (Step S122) The command processing unit 130 causes the audio device 42 to present guidance audio as guidance information for instructing individual speakers to re-speak. After that, the process shown in Figure 5 is terminated.
[0084] (Step S124) The command processing unit 130 refers to the command list and determines whether the detected voice command is a voice command related to the speech position. If it is determined to be a voice command related to the speech position (Step S124 YES), the process proceeds to step S128. If it is determined not to be a voice command related to the speech position (Step S124 NO), the process proceeds to step S126. (Step S126) The command processing unit 130 performs operation control on the controlled device 40 instructed by the voice command, in accordance with the detected voice command, regardless of the utterance position. After that, the process shown in Figure 5 is terminated.
[0085] (Step S128) The command processing unit 130 refers to the command list and determines whether the voice command detected at the speech position indicated in the speech position information is valid or not. If it is determined to be valid (Step S128 YES), the process proceeds to step S130. If it is determined to be invalid (Step S128 NO), the detected voice command is rejected and the process in Figure 5 is terminated. (Step S130) The command processing unit 130 performs operation control related to the utterance position in accordance with the detected voice command for the controlled device 40 instructed by the voice command. After that, the process shown in Figure 5 is terminated.
[0086] Next, we will explain in more detail the first example of a frame-based utterance positioning method. Figure 6 is an explanatory diagram showing a first example of utterance position detection. In the illustrated example, since the shift length is shorter than the frame length, utterance position information is acquired at time intervals corresponding to the shift length. In this example, the utterance position with the maximum spatial spectral value for each candidate utterance position, which is greater than or equal to a predetermined value, is detected. As illustrated in Figure 7, the estimated utterance position (estimated utterance position) may not be continuous across multiple frames, but may be acquired intermittently on a frame-by-frame basis. If the speaker is seated, the ground truth of the utterance position for each utterance should be constant for each utterance. Stability of the utterance position is required to determine the utterance position for each utterance.
[0087] Figure 8 is an explanatory diagram showing a second example of utterance position detection. In the illustrated example, the utterance position processing unit 128 sets a sub-interval for each frame, which includes that frame, and determines the utterance position with the highest confidence level among the utterance position information in the set sub-interval, where the confidence level is above a predetermined confidence threshold, as the utterance position for that frame. The frame length, shift length, and sub-interval duration are typically, for example, 20-50ms, 10-20ms, and 300-1000ms. In the example in Figure 8, the sub-interval T for each frame is centered on the target frame to be processed and includes one or more preceding frames that precede the target frame and one or more succeeding frames that follow the target frame. By determining the utterance position for each frame for each corresponding sub-interval T, stable estimation of the utterance position is achieved. As illustrated in Figure 9, the estimated utterance position approximates the true value in that it continues across multiple frames, eliminating the phenomenon of intermittent occurrences. However, even when an utterance is made and its position should be estimated, there are cases where the utterance position cannot be detected for multiple frames.
[0088] The speech position processing unit 128 according to this embodiment can determine the speech base speech position using the method described below. Figure 10 is a flowchart showing a first example of conversion to the speech base speech position according to this embodiment. (Step S202) The speech position processing unit 128 processes the m-th utterance R as the utterance to be processed. m Select this option. (Step S204) The speech position processing unit 128 initializes the cumulative detection time P(K) for each predetermined speech position K to 0.
[0089] (Step S206) The speech position processing unit 128 processes speech R for each speech position K. m In the speech interval, the interval in which the speech position K is estimated is called the position-specific speech interval V. i (i is an integer between 1 and N, and N is an integer indicating the number of detected location-specific speech segments) (Step S208) The speech position processing unit 128 processes the speech interval V for each speech position K. iAnother section length L i The cumulative detection period P(K) is calculated by adding these values together. (Step S210) The speech position processing unit 128 determines the speech position K where the cumulative detection period P(K) is maximum as speech position R m This is set as the utterance base utterance position for that. After that, the process in Figure 10 is terminated.
[0090] Figure 11 is an explanatory diagram showing an example of the execution of the first example of conversion to utterance-based utterance position. In the illustrated example, T s (m), T e (m) represents the utterance R m This indicates the start and end times of the utterance interval. V1 to V3 are detected as position-specific utterance intervals included in the utterance interval. L1 to L3 indicate the interval lengths of the position-specific utterance intervals V1 to V3 detected within the utterance interval, respectively. The utterance positions for position-specific utterance intervals V1, V2, and V3 are estimated to be 0, 0, and 1, respectively. At this time, the cumulative detection period P(0) for utterance position 0 is L1 + L2, and the cumulative detection period P(1) for utterance position 1 is L3. The cumulative detection periods P(2) to P(4) for utterance positions 2 to 4 are all 0. Eutterance positions 0, 1, 2, 3, and 4 represent the driver's seat (FR), passenger seat (FL), rear right (RR), rear middle (RM), and rear left (RR), respectively. At this time, the utterance position processing unit 128 determines that utterance position 0, which has the maximum cumulative detection period, is utterance R m Utterance-based utterance position SSL(R m ) can be defined as follows.
[0091] Figure 12 is a flowchart showing a second example of conversion to utterance-based utterance position according to this embodiment. The example in Figure 12 also includes simultaneous utterance detection. (Step S222) The utterance position processing unit 128 processes the m-th utterance U m Select this option. (Step S224) The speech position processing unit 128 processes the sub-interval S containing the f-th frame in the selected utterance. f Select this option. (Step S226) The speech position processing unit 128 determines the cumulative detection period Q for each speech position K. fInitialize (K) to 0.
[0092] (Step S228) The speech position processing unit 128 processes the small section S f In this context, the interval in which the utterance position K is estimated is called the position-specific utterance interval V. i Identify it as such. (Step S230) The speech position processing unit 128 processes the speech interval V for each speech position K. i Another section length L i Add the cumulative detection period Q f Calculate (K). (Step S232) The speech position processing unit 128 performs cumulative detection during the Q f The speech position K where (K) is maximized is within the sub-interval S. f Sub-segment utterance position K f It shall be defined as follows. (Step S234) The speech position processing unit 128 advances the frame f to be processed to the next frame. (Step S236) The speech position processing unit 128 controls the speech U m The system determines whether or not the next frame to be processed exists within the speech interval. If it is determined that it exists (step S236 YES), the system proceeds to step S240. If it is determined that it does not exist (step S236 NO), the system proceeds to step S238.
[0093] (Step S238) The speech position processing unit 128 processes each speech position K and speaks U m Within the speech interval, the speech position K is set as the sub-interval speech position S. f The number of localities N K It is counted as follows: locality number N K This indicates the time length in frames of the cumulative utterance period (K) in Figure 10. (Step S240) The speech position processing unit 128 counts the localization number N K The system determines whether there are multiple utterance positions whose ratio to the total number of frames within the utterance interval is greater than or equal to a predetermined ratio. If it is determined that there are multiple positions (step S240 YES), the system proceeds to step S242. If it is determined that there is only one position (step S240 NO), the system proceeds to step S244. (Step S242) The speech position processing unit 128 controls the speech U m The speech state in this case is determined to be simultaneous speech. After that, the process shown in Figure 12 is terminated. (Step S244) The speech position processing unit 128 controls the localization number N K The utterance position K that maximizes this is utterance U m This is set as the utterance base utterance position for that. After that, the process in Figure 12 is terminated.
[0094] Figure 13 is an explanatory diagram showing an example of determining the utterance position for each sub-segment. In the illustrated example, utterance U m In the speech interval relating to sub-interval S, f In this case, the intervals in which utterance positions 0 and 1 are estimated are identified as position-specific utterance intervals V1 and V2, respectively. L1 and L2 indicate the interval lengths of position-specific utterance intervals V1 and V2, respectively. At this time, the cumulative detection period Q for utterance position 0. f (0) is L1, and the cumulative detection period P(1) for speech position 1 is L2. Cumulative detection period Q for speech positions 2-4 f (2) ~ Q f (4) is 0 in all cases. At this time, the speech position processing unit 128 determines the speech position 0 that has the maximum cumulative detection period to be sub-interval S f Sub-segment utterance position K f It can be defined as follows.
[0095] Figure 14 is an explanatory diagram showing an example of simultaneous speech detection. In the illustrated example, the utterance U estimated by Cloud 50 n The speech interval (estimated speech interval) related to this contains a single utterance R n This includes the following. Here, the localization count for utterance segment 2 is 22, and the localization counts for the other utterance segments are 0. The utterance position 2, which has the highest localization count, is utterance U. n It is determined as the utterance base utterance position for this. In contrast, utterance U mAs in the case of estimated speech intervals, multiple utterances may actually be included at different times. If each of the multiple utterances contains a voice command, it may be difficult to determine which voice command to use. In this embodiment, the speech position processing unit 128 detects events included in a single speech interval as simultaneous utterances, even if they are multiple utterances at different times. When simultaneous utterances are detected, guidance information is output to prompt the user to make individual utterances again.
[0096] In the example in Figure 14, the utterance U m The estimated speech intervals related to this include 33 sub-intervals and 2 actual speech R m1 , R m2 This includes the utterance R. m1 , R m2 The utterance position for each of these is different. Utterance U m In the first half of the estimated speech interval related to this, the speech position for each sub-interval is approximately 0, and the speech R m1 This corresponds to the speech segment. In the latter half of the speech segment, the sub-segment speech position is approximately 1, and the speech R m2 This corresponds to the speech segment U. m In the estimated utterance related to this, the localization count for utterance position 0 is 12, and the localization count for utterance position 1 is 21. The localization count for utterance positions 2 to 4 is 0 each. The ratio of the localization count for utterance position 0 to the number of sub-intervals within the utterance interval is 0.36 (≒12 / 33), and the ratio of the localization count for utterance position 1 to the number of sub-intervals is 0.63 (≒21 / 33), both exceeding a predetermined ratio (e.g., 0.2). Therefore, the utterance state in this utterance interval is determined to be simultaneous utterance.
[0097] Next, we will describe the evaluation experiments conducted on this embodiment. The evaluation experiments were conducted from the following perspectives: (1) the basic performance of the proposed method, (2) speech position detection under various conditions, and (3) the channel number dependence of simultaneous speech detection under various conditions. In the evaluation experiment, voice data and driving noise (noise data) were individually recorded using microphones 30-1 to 30-3 inside a vehicle equipped with five seats. The voice data and noise data were then mixed and supplied to sound processing devices 10 and 20 installed outside the vehicle. This reproduced speech under noise. However, depending on the experimental conditions, music played by sound equipment 42 was also recorded as music data and further mixed. As sound sources, speech was randomly selected from 200 English speech utterances and 52 Japanese speech utterances.
[0098] First, (1) the basic performance of the proposed method will be explained. Here, the speech detection rate (SDR) and speech localization rate (SLR) were calculated for each experimental condition as indicator values for the experimental results. The speech detection rate is the ratio of the number of correctly detected speech segments to the number of correct speech segments. The speech localization rate is the ratio of the number of speech segments in which the seat was correctly identified as the speech location to the number of detected speech segments. The following four methods were set as experimental conditions: (i) speech location estimation using a mathematical model on a frame-by-frame basis, (ii) speech location estimation for each sub-segment corresponding to each frame, (iii) speech-based speech location estimation based on the speech location estimated using a mathematical model on a frame-by-frame basis, and (iv) speech-based speech location estimation based on the speech location estimated using a mathematical model on a frame-by-frame basis, taking sub-segments into consideration. However, under experimental conditions (i) and (ii), the correctness of the speech segment was evaluated by checking whether the ratio of the number of frames in which the speech position could be detected by the speech position processing unit 128 to the number of frames in the speech segment was 0.5 or greater. For each experimental condition, 250 utterances were used. 250 utterances correspond to 50 utterances per seat. The experiment was conducted with the vehicle windows and sunroof closed and the vehicle stationary.
[0099] Figure 16 shows the speech detection rate and speech location detection rate for each speech detection method. The speech detection rate is lowest for experimental condition (i), and even higher for experimental condition (ii). For experimental conditions (iii) and (iv), the speech detection rate is 100%. This indicates that speech can be detected with almost certainty. Even in experimental condition (i), where the speech detection rate is lowest, it is 74.1%, which is better than when using the spatial spectrum directly without a model. This supports the idea that the phenomenon of intermittent detection due to incorrect detection of speech segments is mitigated. The speech location detection rate is lowest for experimental condition (ii), and increases in the order of experimental conditions (i), (iii), and (iv). The speech location detection rate is 88.5% even in experimental condition (ii), and 95.9%, 99.6%, and 100% for experimental conditions (ii), (iii), and (iv). This indicates that the speech location can be detected with almost certainty.
[0100] Next, (2) Speech position detection under various conditions will be explained. In speech position detection under various conditions, the speech detection rate and speech position detection rate were determined for each of the different combinations of the number of channels in the acoustic signal, the operating state of the vehicle, and the in-vehicle environment as experimental conditions. Two types of acoustic signal channels (2 channels and 3 channels), two or five types of vehicle operating states, and 14 types of in-vehicle environment conditions were set.
[0101] Two-channel refers to the case where the frame-based speech position is detected based on the spatial spectrum derived from the acoustic signals picked up by microphones 30-1 and 30-2. Three-channel refers to the case where the frame-based speech position is detected using not only the spatial spectrum but also the power spectral density and zero crossing number derived from the acoustic signal picked up by microphone 30-3. Two operating states were set: stopped and driving at 45 mph. Five operating states were also set: a: stopped, b: idling, c: driving at 45 mph, d: driving at 65 mph, and e: driving at 72 mph. At 65 mph and 72 mph, the noise level is 5 dB and 10 dB higher, respectively, than at 45 mph.
[0102] The 14 different in-car environments are as follows: A: Basic configuration (all windows closed, sunroof closed), B: HATS (Head and Torso Simulator) with mask, C: HATS facing inward, D: HATS facing outward, E: Sunroof open, F: FL windows open, G: FR windows closed, H: RL windows open, I: RR windows open, J: FLFR windows open, K: RLRR windows open, L: All windows and sunroof open, M: Reclining HATS lying down, N: Reclining HATS standing up. HATS refers to a simulated head with a torso, indicating that the simulated head is seated in the driver's seat. The placement of HATS inside the vehicle simulates the propagation of sound by a person inside, including reflection, diffraction, and absorption. B, C, D, K, and L above represent the differences in the placement of the simulated head in the basic configuration, that is, in an environment with all windows and the sunroof closed. HATS with mask refers to a state where a simulated head wearing a mask over the mouth is seated in the driver's seat facing forward. HATS inward refers to a state where the simulated head is tilted 90 degrees inward towards the vehicle while seated in the driver's seat. HATS outward refers to a state where the simulated head is tilted 90 degrees outward towards the vehicle while seated in the driver's seat.
[0103] Figure 17 shows the speech location detection rate and speech detection rate obtained for 14 different in-vehicle environments while the vehicle was stationary. 5-seat evaluation refers to distinguishing and detecting speech in five different locations: driver's seat, passenger seat, rear right, rear middle, and rear left. 3-seat evaluation refers to distinguishing and detecting speech in three different locations: passenger seat, rear left, and rear middle. (a) shows the speech location detection rate with 3 channels, (b) shows the speech location detection rate with 2 channels, (c) shows the speech detection rate with 3 channels, and (d) shows the speech detection rate with 2 channels. With 3 channels, both the speech location detection rate and speech detection rate exceed 95%. With 2 channels, both the speech location detection rate and speech detection rate exceed 95%, except for the speech location detection rate in in-vehicle environment D. The speech location detection rate in in-vehicle environment D is approximately 92%. Furthermore, no significant difference was observed between the 5-seat evaluation and the 3-seat evaluation in terms of either the speech location detection rate or the speech detection rate.
[0104] Figure 18 shows the speech location detection rate and speech detection rate obtained for 15 different in-vehicle environments while driving at 45 mph. The 15 in-vehicle environments include in-vehicle environments A to O, as well as in-vehicle environment A'. In-vehicle environment A' refers to music playback in basic form A. (a) shows the speech location detection rate with 3 channels, (b) shows the speech location detection rate with 2 channels, (c) shows the speech detection rate with 3 channels, and (d) shows the speech detection rate with 2 channels. With 3 channels, both the speech location detection rate and speech detection rate exceed 95% except for in-vehicle environments J and L. However, with 2 channels, the speech location detection rate drops to 90% or less for in-vehicle environments J, L, as well as in-vehicle environments B, C, D, F, G, and H. This indicates that referencing acoustic features based on a single-channel acoustic signal in speech estimation contributes to improving the speech location detection rate even under noise caused by driving.
[0105] Figure 19 shows the speech location detection rate and speech detection rate obtained for five different vehicle operating states in the in-vehicle environment A. (a) shows the speech location detection rate with 3 channels, (b) shows the speech location detection rate with 2 channels, (c) shows the speech detection rate with 3 channels, and (d) shows the speech detection rate with 2 channels. For all operating states, regardless of the number of channels, both the speech location detection rate and speech detection rate exceed 95%.
[0106] Next, (3) the channel number dependence of simultaneous speech detection under various conditions will be explained. In simultaneous speech detection under various conditions, the simultaneous speech detection rate, simultaneous speech detection accuracy, and single speech detection rate were determined for each different combination of the number of channels in the acoustic signal and the operating state of the vehicle as experimental conditions. However, the in-car environment was set to have all windows closed and the sunroof closed. The simultaneous speech detection rate (SSDR) refers to the ratio of speech segments that were correctly detected as simultaneous speech to the total number of speech segments detected. The simultaneous speech detection accuracy (SSDA) refers to the ratio of the difference between the number of speech segments that were correctly detected as simultaneous speech and the number of speech segments that were incorrectly detected as simultaneous speech to the total number of speech segments detected. The single speech detection rate (SinSDR) refers to the ratio of the number of speech segments that were correctly detected as single speech to the total number of speech segments detected. For example, M speech segments U1~U M Assume that the data contains N1 simultaneous utterance segments and N2 single utterance segments. Let C1 and S1 be the number of simultaneous utterance segments that were correctly detected as simultaneous utterances and the number of single utterance segments that were detected as single utterances, respectively. Let C2 and S2 be the number of single utterance segments that were correctly detected as single utterances and the number of simultaneous utterance segments that were detected as simultaneous utterances, respectively. In this case, the simultaneous utterance detection rate SSDR is C1 / N1. The simultaneous utterance detection accuracy SSDA is (C1-S2) / N1. The single utterance detection rate SinSDR is C2 / N2.
[0107] Two options were set for the number of channels in the acoustic signal: 2 channels and 3 channels. Five conditions, a to e, were set for the operating state of the vehicle. In the evaluation of simultaneous speech, 50 evaluation files (W1-W) were selected from a total of 250 audio data files, combining Japanese and English speech. 50 The following were selected. For each pair of evaluation data with adjacent indices, such as [W1,W2], [W2,W3], etc., five patterns of simultaneous utterance data were generated. Each evaluation data contains one utterance.
[0108] Figure 15 illustrates patterns of simultaneous utterances. (a) No simultaneous utterance, (b) Preceding utterance W m+1 Part of the latter half is the subsequent utterance W m+2 (c) Preceding utterance W m+1 The period is the subsequent utterance W m+2 (d) Subsequent utterance W m+1 Part of the first half is a preceding utterance W m+2 (e) Preceding utterance W m+1 The end of the subsequent utterance W m+2 It matches the beginning of. In pattern (a), the preceding utterance W m+1 From the end of the subsequent utterance W m+2 The period G1 to the beginning of the following utterance was randomly selected from a predetermined range of 100 to 500 ms. In patterns (b) and (d), the period G2 from the beginning of the following utterance to the end of the preceding utterance was randomly selected from a predetermined range of 100 to ml. ml is half the length of the shorter of the preceding or following utterance.
[0109] Figure 20 shows the simultaneous speech detection rate, simultaneous speech detection accuracy, and single-utterance speech detection accuracy obtained for five different vehicle operating states as simultaneous speech evaluations. (a) shows the 3-channel simultaneous speech evaluation, and (b) shows the 2-channel simultaneous speech evaluation. The simultaneous speech detection rate, simultaneous speech detection accuracy, and single-utterance speech detection accuracy were all around 90%. Regarding single-utterance speech detection accuracy, a tendency to decrease with increasing driving speed was observed, but this was not significant at practical driving speeds (below 45 mph).
[0110] (modified version) The above embodiments may be implemented in modified form. Modifications include substituting some configurations with others, combining some configurations with others, and omitting some configurations. For example, the acoustic feature analysis unit 222 can use any type of feature as long as it can represent the characteristics of the spoken speech as an acoustic feature. For example, instead of spectral power density, one of the following may be used: Mel-frequency cepstrum (MFCC), delta cepstrum, linear prediction coefficients (LPC), etc. Also, the derivation of the zero crossing number may be omitted.
[0111] Furthermore, although the example given was the calculation of the MUSIC spectrum by the spatial spectrum analysis unit 122, it is not limited to this. The spatial spectrum analysis unit 122 may also calculate other types of spatial spectra, for example, spatial spectra obtained by beamforming in the direction of individual sound source positions. Beamforming is a technique for controlling directivity by adding or filtering signals with different gains and / or delays for each channel.
[0112] The mathematical model used in the speech position estimation unit 124 is not necessarily limited to random forests; other types of machine learning models may be used. These other types of machine learning models may include convolutional neural networks (CNNs), recurrent neural networks, and others. The process shown in Figure 10 may include a step to determine whether or not simultaneous utterances occur as an utterance state within an utterance interval, similar to the process in Figure 12. That is, the utterance position processing unit 128 determines that simultaneous utterances occur if there are multiple utterance positions K where the ratio of the cumulative detection period P(K) for each utterance position K to the utterance interval is equal to or greater than a predetermined ratio. The utterance position processing unit 128 may determine that a single utterance occurs if there is only one utterance position K where the ratio of the cumulative detection period P(K) for each utterance position K to the utterance interval is equal to or greater than a predetermined ratio, and then proceed to the process in step S210.
[0113] The acoustic processing device 20 does not necessarily have to be a device primarily intended for noise reduction, as long as it acquires an acoustic signal from the microphone 30-3 and is equipped with an acoustic feature analysis unit 222. The spatial filtering unit 134 and the speech information acquisition unit 126 of the sound processing device 10 are provided in separate devices in the sound processing system S1 and may be omitted from the sound processing device 10. The spatial filtering unit 134 may be omitted, and the speech information acquisition unit 126 may receive an acoustic signal from either one channel of the A / D conversion unit 112.
[0114] Part or all of the sound processing device 10 may constitute part of the sound equipment 42. The sound processing devices 10 and 20 may be configured as a single unit. In that case, one of the control units 120 and 220 may have the function of the other, while the other may be omitted. One of the A / D conversion units 112 and 212 may have the function of the other, while the other may be omitted. The number of channels in the acoustic signal used to calculate the spatial spectrum is not limited to two channels, but may be three or more. Furthermore, the parameters related to the above processing can be arbitrarily set as long as the desired effects according to this embodiment are achieved.
[0115] The speech information acquisition unit 126 may perform a known voice activity detection (VAD) on the acoustic signal of any channel to detect a speech segment. More specifically, the speech information acquisition unit 126 may calculate the number of zero crossings and power for each frame, and define a speech segment as the point in time when a state in which the calculated power is equal to or greater than a predetermined power threshold and the number of zero crossings is equal to or greater than a predetermined frequency (e.g., 200 to 500 times per second) continues for a predetermined time (e.g., 0.2 to 0.5 seconds) or longer, and the point in time when that state stops as the end point, with all other segments defined as non-speech segments. Alternatively, the speech information acquisition unit 126 may use a mathematical model to determine whether a frame belongs to a speech segment based on the number of zero crossings and power. The mathematical model may pre-learn the relationship between pairs of zero crossings and power as explanatory variables and the relationship between whether a frame belongs to a speech segment as the objective variable. When the acoustic processing devices 10 and 20 are configured as a single unit, the power derived from the zero crossing number and power spectral density calculated by the acoustic feature analysis unit 222 may be used for the speech detection process.
[0116] Furthermore, the speech information acquisition unit 126 may perform known speech recognition processing in the detected speech segment to determine the speech content. If the speech information acquisition unit 126 itself detects the speech segment and determines the speech content, it does not need to receive speech content information and speech segment information from the cloud 50. The speech information acquisition unit 126 does not need to transmit the acquired acoustic signal to the cloud 50 to request speech segment detection and speech recognition processing.
[0117] As described above, the sound processing system S1 according to this embodiment includes a spatial spectrum analysis unit 122 that analyzes the spatial spectrum of a sound source for each frame over a predetermined period based on acoustic signals from multiple channels, and a speech position estimation unit 124 that estimates the speech position based on the analyzed spatial spectrum, using at least a mathematical model that shows the relationship between the spatial spectrum of the sound source and the speech position, which is the position of the speaker making the utterance. This configuration estimates the speech position corresponding to the spatial spectrum derived frame by frame from the acquired acoustic signal, based on the relationship between known spatial spectra and speech positions. Therefore, it is possible to estimate a more stable speech position than when the speech position is directly determined from the spatial spectrum. For example, the phenomenon of intermittently occurring frames within a speech interval where the speech position is determined is mitigated.
[0118] The acoustic processing system S1 may include an acoustic feature analysis unit 222 that analyzes acoustic features for each frame based on acoustic signals from multiple channels and separate channels. The mathematical model described above shows the relationship between acoustic features, spatial spectra, and speech position, and the speech position estimation unit 124 may estimate the speech position based on the analyzed acoustic features and spatial spectra. In this configuration, the speech position is estimated using acoustic features analyzed from acoustic signals of channels separate from the multiple channels used to estimate the spatial spectrum. Since the speech position is estimated by referring to the speech position dependence of the acoustic features of the spoken speech, the accuracy of speech position estimation can be improved even in noisy environments. Furthermore, by using a mathematical model, estimation accuracy can be ensured even when the acoustic signals of multiple channels and the acoustic signals of separate channels are not synchronized.
[0119] The acoustic feature analysis unit 222 may analyze the frequency characteristics of the acoustic signals of the acoustic feature analysis channels of the separate channels mentioned above, as well as the frequency of zero intersections, as acoustic features. This configuration allows for a more reliable understanding of the characteristics of spoken speech by using the frequency characteristics of the acoustic signal and the frequency of zero intersections as acoustic features. Therefore, the accuracy of estimating the speech location can be improved.
[0120] The acoustic feature analysis unit 222 may analyze the power spectral density of the acoustic signals of the separate channels mentioned above as acoustic features. This configuration allows for the analysis of frequency-specific intensity through a simple calculation based on the Discrete Fourier Transform. Therefore, it is possible to improve the accuracy of utterance location estimation economically.
[0121] The spatial spectrum analysis unit 122 may calculate a spatial spectrum for each predetermined candidate speech location (e.g., a seat), and the speech location estimation unit 124 may calculate the confidence level that the speaker making the utterance is located at each candidate speech location. In this configuration, the confidence level of the speaker's location is calculated for each given candidate utterance location, based on the spatial spectrum. Therefore, the computational load is reduced compared to calculating the spatial spectrum for every sound source location. Also, unlike directly deriving the utterance location from the spatial spectrum, the influence of its absolute value on the likelihood of determining the utterance location is reduced. As a result, stable estimation of the utterance location becomes possible.
[0122] The mathematical model may be a random forest. This configuration allows for fast utterance position estimation because the operations performed by each decision tree constituting the random forest are parallel. Furthermore, because it has little dependence on specific input values used as explanatory variables, it can stably estimate utterance positions.
[0123] Although one embodiment of this invention has been described in detail above with reference to the drawings, the specific configuration is not limited to that described above, and various design changes can be made without departing from the spirit of this invention. [Explanation of Symbols]
[0124] S1…Acoustic processing system, 10, 20…Acoustic processing device, 30 (30-1~30-3)…Microphone, 40…Controlled equipment, 42…Acoustic equipment, 44…Air conditioner, 46…Window opener, 48…Steering heater, 50…Cloud, 112…A / D conversion unit, 114…Communication unit, 120…Control unit, 122…Spatial spectrum analysis unit, 124…Speech position estimation unit, 126…Speech information acquisition unit, 128…Speech position processing unit, 130…Command processing unit, 134…Spatial filtering unit, 212…A / D conversion unit, 222…Acoustic feature analysis unit, 222a…Frequency analysis unit, 222b…Zero intersection analysis unit, 224…Low-pass filter, 226…Noise removal unit
Claims
1. A spatial spectrum analysis unit analyzes the spatial spectrum of a sound source frame by frame over a predetermined period based on multiple channel acoustic signals, A speech position estimation unit that estimates the speech position based on the analyzed spatial spectrum, using at least a mathematical model that shows the relationship between the spatial spectrum of the sound source and the speech position, which is the position of the speaker making the utterance. The system includes an acoustic feature analysis unit that analyzes acoustic feature quantities for each frame based on the acoustic signals of the multiple channels and a separate channel, The mathematical model shows the relationship between acoustic features, spatial spectrum, and speech position, and the speech position estimation unit, The speech position is estimated based on the analyzed acoustic features and spatial spectrum. Acoustic processing system.
2. The acoustic feature analysis unit is, The aforementioned acoustic features include analyzing the frequency characteristics of the acoustic signals of the separate channels and the frequency of zero intersections. The acoustic processing system according to claim 1.
3. The acoustic feature analysis unit is, The power spectral density is analyzed as the frequency characteristic of the acoustic signals of the aforementioned separate channels. The acoustic processing system according to claim 2.
4. The spatial spectrum analysis unit is, A vector containing elemental values for each predetermined candidate utterance position, which is a candidate for the utterance position, is calculated as the spatial spectrum. The aforementioned speech position estimation unit, For each of the aforementioned candidate speech location positions, calculate the confidence level of the speaker who is making the utterance. The acoustic processing system according to claim 1.
5. The mathematical model is a random forest. The acoustic processing system according to claim 4.
6. A spatial spectrum analysis unit analyzes the spatial spectrum of a sound source frame by frame over a predetermined period based on multiple channel acoustic signals, A speech position estimation unit that estimates the speech position based on the analyzed spatial spectrum, using at least a mathematical model that shows the relationship between the spatial spectrum of the sound source and the spatial distribution of the sound source, The system includes an acoustic feature analysis unit that analyzes acoustic feature quantities for each frame based on the acoustic signals of the multiple channels and a separate channel, The mathematical model shows the relationship between acoustic features, spatial spectrum, and speech position, and the speech position estimation unit, The speech position is estimated based on the analyzed acoustic features and spatial spectrum. Acoustic processing device.
7. A program for causing a computer to function as the sound processing device described in claim 6.
8. A method for acoustic processing in an acoustic processing system, The spatial spectrum analysis unit performs a spatial spectrum analysis step in which it analyzes the spatial spectrum of a sound source frame by frame over a predetermined period based on acoustic signals from multiple channels. The speech position estimation unit estimates the speech position based on the analyzed spatial spectrum, using a mathematical model that shows the relationship between the spatial spectrum of the sound source and the speech position, which is the position of the speaker making the utterance. The process involves performing an acoustic feature analysis step that analyzes acoustic features for each frame based on the acoustic signals of the multiple channels and separate channels, The mathematical model shows the relationship between acoustic features, spatial spectrum, and speech position, and the speech position estimation step is, The speech position is estimated based on the analyzed acoustic features and spatial spectrum. Sound processing methods.
Citation Information
Patent Citations
Sound source localization system and sound source localization method
JP2013044950A
Voice recognition system, voice recognizing method and its program
WO2006025106A1
Emergency vehicle detection
WO2021080989A1