Acoustic processing system, acoustic processing device, acoustic processing method, and program

The acoustic processing system addresses noise interference in voice recognition by estimating utterance positions and processing voice commands, improving the accuracy and reliability of device control in vehicles.

JP7814213B2Active Publication Date: 2026-02-16HONDA MOTOR CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022053157
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-29
Publication Date
2026-02-16
Estimated Expiration
2042-03-29

AI Technical Summary

Technical Problem

Existing voice recognition systems in vehicles face challenges in accurately determining the speaker's position due to noise interference and mixed sounds, leading to unstable recognition of voice commands.

Method used

An acoustic processing system that estimates the utterance position for each frame based on spatial spectra of multiple channels, identifies speech sections, and processes voice commands accordingly, incorporating noise reduction and simultaneous speech handling.

Benefits of technology

Stably determines the speaker's position for each utterance, enhancing the accuracy of voice command recognition and enabling reliable device control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007814213000004
    Figure 0007814213000004
  • Figure 0007814213000005
    Figure 0007814213000005
  • Figure 0007814213000006
    Figure 0007814213000006
Patent Text Reader

Abstract

To stably estimate a speech position for each speech production.SOLUTION: A speech position estimation part estimates a speech position for each frame of a prescribed period on the basis of a spatial spectrum of a sound source analyzed from an acoustic signal of a plurality of channels, a speech information acquiring part acquires speech section information indicating a speech section in which speech is produced on the basis of the acoustic signal of at least one of the plurality of channels, a speech position processing part identifies an every position speech section in which the speech position is estimated in the speech section, and determines the speech position on the basis of the amount of the every position speech position.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a sound processing system, a sound processing device, a sound processing method, and a program. [Background technology]

[0002] Speech recognition systems that use speech recognition technology to extract voice commands from speech in vehicles are becoming widespread. These voice recognition systems enable the operation of various devices and functions according to the extracted voice commands. For example, Patent Document 1 describes a voice recognition system that separates a speaker's voice from voices input through multiple microphones installed in a vehicle. The voice recognition system includes a storage device that stores preset information indicating the sound source position of the speaker's voice, and a voice recognition unit that refers to the speaker's preset information stored in the storage device, separates the speaker's voice from the voice input through the microphones, performs voice recognition, and recognizes voice commands.

[0003] The voice recognition system described in Patent Document 1 further includes a sensor for detecting the seat position of the speaker, and the storage device stores preset information for each seat position of the speaker, acquires the seat position of the speaker from the sensor, searches the storage device for the preset information based on the acquired seat position, and outputs it to the voice recognition unit. Various navigation processes are performed based on the recognized voice commands. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] International Publication No. 2006 / 025106 Summary of the Invention [Problem to be solved by the invention]

[0005] Some devices and functions to be operated are related to the speaker's position. For example, when opening or closing a window, the window closest to the speaker is typically the target of operation. A voice command indicating such an operation requires the detection of the speaker's position, which is the sound source. In this regard, Patent Document 1 describes a method for processing the phase information and intensity distribution of a received voice command to analyze the directivity of the voice and identify the speaker's position. Phase information and intensity distribution are typically analyzed in frame units, which are shorter in time than the speech. On the other hand, a voice command is completed in a single utterance and can be uttered by different speakers. However, various noises are mixed into microphones installed in the vehicle cabin. In addition to driving noise such as the engine, music emitted by audio equipment and spoken voices can also be noise. The mixed noise can reduce the recognition rate of the voice command's utterance position. Therefore, a stable estimation of the utterance position for each utterance has been anticipated.

[0006] An object of the present invention is to provide a sound processing system, a sound processing device, a sound processing method, and a program that can stably estimate an utterance position for each utterance. [Means for solving the problem]

[0007] (1) The present invention has been made to solve the above-mentioned problems, and one aspect of the present invention includes an utterance position estimation unit that estimates an utterance position where an utterance is made for each frame of a predetermined period based on a spatial spectrum of a sound source analyzed from acoustic signals of multiple channels; an utterance information acquisition unit that acquires utterance section information indicating an utterance section where a voice is uttered based on acoustic signals of at least one channel of the multiple channels; and a speech information acquisition unit that acquires, for each utterance position, information indicating the utterance section where the utterance position is made in the utterance section. The utterance position estimation unit Identifying the estimated position-specific speech section, The longest utterance position is , an utterance position in the utterance section as and a speech position processing unit that determines the speech position.

[0008] (2) Another aspect of the present invention is the acoustic processing system of (1), wherein the speech position processing unit may identify, for each predetermined speech position, a frame in which the speech position is identified as the position-specific speech section.

[0009] (3) Another aspect of the present invention is the acoustic processing system of (1), wherein the speech position processing unit may identify, for each frame, a small section that is part of the speech section and spans multiple frames including the frame in question, and identify the speech position in the frame based on the speech position related to the small section.

[0010] (4) Another aspect of the present invention is an acoustic processing system as defined in (1), which may include a command processing unit that executes processing based on the speech position-related voice command when the voice command indicated by the speech content in the speech section is a speech position-related voice command related to a specific speech position and the speech position is the specific speech position.

[0011] (5) Another aspect of the present invention is the acoustic processing system of (4), wherein the speech position processing unit may determine that simultaneous speech occurs in the speech section when there are multiple speech positions in which the ratio of the position-specific speech sections in the speech section is equal to or greater than a predetermined ratio.

[0012] (6) Another aspect of the present invention is the acoustic processing system of (5), wherein when simultaneous speech is determined in the speech section, the command processing unit may output guidance information that guides individual re-speech.

[0013] (7) Another aspect of the present invention is a speech position estimation unit that estimates a speech position where an utterance is made for each frame of a predetermined period based on a spatial spectrum of a sound source analyzed from acoustic signals of multiple channels; an utterance information acquisition unit that acquires speech section information indicating a speech section where a voice is uttered based on acoustic signals of at least one channel of the multiple channels; and a speech information acquisition unit that acquires, for each speech position, a speech section where the utterance position is located in the speech section. The utterance position estimation unit Identifying the estimated position-specific speech section, The longest utterance position is, an utterance position in the utterance section as and a speech position processing unit that determines the speech position.

[0014] (8) Another aspect of the present invention is a computer of It may also be a program for causing the sound processing device (7) to function.

[0015] (9) Another aspect of the present invention is a sound processing method for a sound processing system, comprising: an utterance position estimation step in which an utterance position estimation unit estimates an utterance position where an utterance was made for each frame of a predetermined period based on a spatial spectrum of a sound source analyzed from sound signals of multiple channels; an utterance information acquisition step in which an utterance information acquisition unit acquires utterance section information indicating an utterance section where a voice was uttered based on sound signals of at least one channel of the multiple channels; and an utterance position processing step in which, for each utterance position, the utterance position is calculated in the utterance section. The utterance position estimation unit Identifying the estimated position-specific speech section, The longest utterance position is , an utterance position in the utterance section as and an utterance position processing step for determining the utterance position. [Effects of the Invention]

[0016] According to the present invention, it is possible to stably estimate the utterance position for each utterance. According to the aspects (1), (7), (8), or (9) of the present invention, the speech position in the speech section is determined based on the amount of the position-specific speech section estimated for each speech position in the estimated speech section. Even if the speech position estimated for each frame is unstable, the speech position in the speech section is uniquely determined. Therefore, the speech position can be stably determined for each utterance.

[0017] According to the aspect (2), for a given speech position estimated for each frame, that frame is included in the position-specific speech section of that speech position, so that the position-specific speech section used to determine the speech position can be immediately determined.

[0018] According to the aspect (3), the speech position of the corresponding frame is determined by referring to the speech position in a small section spanning multiple frames, and the frame is included in the position-specific speech section related to the speech position. Therefore, the speech position for each frame can be determined more stably than when only the speech positions of individual frames are referred to. As a result, the speech position in the speech section can be determined more stably.

[0019] According to the fourth aspect, a process based on an utterance position-related voice command instructed by a speech content from a specific utterance position is executed, and a process based on the utterance position-related voice command instructed by a speech content from a position other than the specific utterance position is not executed, so that the process for a speaker at a specific utterance position can be more reliably restricted.

[0020] According to the aspect (5), it is possible to estimate the possibility that multiple utterances were made in a speech section, which provides a clue to resolve the situation where it is unclear which utterance content to use to identify the voice command.

[0021] According to the sixth aspect, multiple speakers at different positions are prompted to individually re-speak, which facilitates the resolution of an issue where processing based on a voice command is not executed due to simultaneous speech. [Brief explanation of the drawings]

[0022] [Figure 1] 1 is a schematic block diagram illustrating an example of the configuration of a sound processing system according to an embodiment of the present invention. [Figure 2] FIG. 1 is a diagram illustrating a first example of microphone placement. [Figure 3] FIG. 10 is a diagram illustrating a second example of microphone placement. [Figure 4] 10 is a table illustrating voice operations in response to voice commands. [Figure 5] 10 is a flowchart illustrating an example of a voice operation process according to the present embodiment. [Figure 6] FIG. 10 is an explanatory diagram showing a first detection example of an utterance position. [Figure 7] FIG. 10 is a diagram showing a first example of an estimated utterance position. [Figure 8] FIG. 10 is an explanatory diagram showing a second example of detection of an utterance position. [Figure 9] FIG. 10 is a diagram showing a second example of an estimated utterance position. [Figure 10] 10 is a flowchart showing a first example of conversion to an utterance-based utterance position according to the present embodiment. [Figure 11] FIG. 10 is an explanatory diagram showing an example of execution of a first example of conversion to an utterance-based utterance position. [Figure 12] 10 is a flowchart showing a second example of conversion to an utterance-based utterance position according to the present embodiment. [Figure 13] FIG. 10 is an explanatory diagram showing an example of determining a speech position for each small section. [Figure 14] FIG. 10 is an explanatory diagram showing an example of simultaneous speech determination. [Figure 15] FIG. 10 is a diagram illustrating an example of a simultaneous speech pattern. [Figure 16] 10A and 10B are diagrams illustrating examples of speech detection rates and speech position detection rates for each speech detection method. [Figure 17] FIG. 10 is a diagram showing a first example of an utterance position detection rate and an utterance detection rate for each in-vehicle environment. [Figure 18] FIG. 10 is a diagram showing a second example of the utterance position detection rate and the utterance detection rate for each in-vehicle environment. [Figure 19] 10A and 10B are diagrams illustrating examples of utterance position detection rates and utterance detection rates for each operating state of a vehicle. [Figure 20] 10A to 10C are diagrams illustrating examples of simultaneous utterance detection rates, simultaneous utterance detection accuracies, and single utterance detection accuracies for each operating state of a vehicle. DETAILED DESCRIPTION OF THE INVENTION

[0023] Hereinafter, embodiments of the present invention will be described with reference to the drawings. First, an example configuration of a sound processing system S1 according to this embodiment will be described. FIG. 1 is a schematic block diagram showing an example configuration of the sound processing system S1 according to this embodiment. The sound processing system S1 acquires sound signals from multiple channels and estimates the position of a speaker based on the acquired sound signals. The sound processing system S1 identifies a voice command (instruction) as the speech content transmitted by the acquired voice signal. In this application, the position of the speaker who is speaking may be referred to as the "speech position." If the identified voice command is related to the speech position, the sound processing system S1 executes processing instructed in accordance with the identified voice command based on the estimated speech position.

[0024] The instructed processing includes controlling the operation of devices connected to the sound processing system S1. The devices to be controlled may include components of the sound processing system S1, or may include other devices that do not belong to the sound processing system S1. The following description mainly focuses on the case where the sound processing system S1 is configured as a part of an in-vehicle system installed in a vehicle and has the function of controlling the operation of various devices installed in the vehicle using voice commands. In this application, "on-board" includes the meaning of actually being mounted on a vehicle, as well as the meaning of being primarily intended for use in a vehicle or being suitable for such use. That is, "on-board" is not intended to limit the embodiment to use in a vehicle, nor is it intended to exclude use outside of a vehicle.

[0025] The sound processing system S1 is configured to include one or more devices. In the example of Fig. 1, the sound processing system S1 includes two sound processing devices 10 and 20, three microphones 30, and a control target device 40. The three microphones 30 are distinguished by sub-numbers 30-1 to 30-3. The sound processing device 10 and the sound processing device 20, and the sound processing device 10 and each control target device 40 are connected to each other so that various data can be transmitted and received wirelessly or via a cable. The sound processing devices 10 and 20 and the control target device 40 can be connected using, for example, a CAN (Controller Area Network).

[0026] The sound processing device 10 acquires sound signals from the microphones 30-1 and 30-2, respectively. The sound processing device 10 analyzes the spatial spectrum of the sound source from the acquired two-channel sound signals for each frame over a predetermined period. The sound processing device 10 estimates the speaker position based on the analyzed spatial spectrum and acoustic feature using a mathematical model that indicates the relationship between the speaker position and a set of spatial spectrum and acoustic feature. The sound processing device 10 acquires, from the acquired sound signals, speech section information that indicates the speech section in which speech was spoken and speech information that indicates the content of the speech. The sound signals used to acquire the speech information may be sound signals obtained by performing spatial filtering (described later) on the two-channel sound signals.

[0027] For each estimated speech position, the sound processing device 10 determines the speech position in the speech section based on the length of the position-specific speech section in which the speech position is estimated in the speech section. When the acquired speech information indicates a voice command related to the speech position, the sound processing device 10 executes processing in accordance with the voice command indicated by the speech information based on the determined speech position. The voice command may instruct operation control of the control-target device 40. The sound processing device 10 outputs a control signal indicating an operation mode (operation mode) to be instructed to the control-target device 40. As will be described later, in relation to the speaking position, the necessity of operational control may differ depending on the speaking position, depending on the voice command or the control-target device 40. The processing instructed by the voice command or the mode of operational control instructed to the control-target device 40 may differ depending on the speaking position.

[0028] The sound processing device 20 acquires a sound signal from the microphone 30-3 and analyzes sound features from the acquired sound signal for each frame. The sound processing device 20 notifies the sound processing device 10 of the sound features obtained by the analysis. The sound processing device 20 and the microphone 30-3 may be primarily intended to acquire a sound signal containing noise components such as engine sounds. The sound processing device 20 outputs the acquired sound signal to the sound processing device 10 as a reference signal.

[0029] Each of the microphones 30-1 to 30-3 includes an electroacoustic transducer (actuator) that picks up sound arriving at the microphone and converts the sound pressure of the picked up sound into an electric signal indicating its intensity into an acoustic signal. The microphones 30-1 and 30-2 output the converted acoustic signals to the sound processing devices 10 and 20, respectively. The acoustic signal output to the sound processing device 20 is used to remove noise components. The microphone 30-3 outputs the converted acoustic signal to the sound processing device 20.

[0030] Next, an example of the arrangement of microphones 30-1 to 30-3 will be described. In the examples of FIGS. 2 and 3, microphones 30-1 and 30-2 are arranged symmetrically in front of the driver's seat and passenger seat provided in the vehicle cabin, with the midpoint between them. In the illustrated example, the front corresponds to the left side of the drawing. The distance between microphones 30-1 and 30-2 is, for example, 5 to 10 cm. The acoustic signals acquired at these positions contain relatively large components of the voice of the driver seated in the driver's seat or the voice of the passenger seated in the passenger seat. In the example of FIG. 2, microphone 30-3 is arranged at the right rear end of the rear seat. In the example of FIG. 3, it is arranged at the center rear end of the rear seat. The acoustic signals acquired at these positions contain relatively large components of noise components such as engine noise and friction noise between the road surface and the wheels.

[0031] The control target devices 40 are devices that are subject to operational control based on voice commands. The control target devices 40 operate in accordance with operational modes instructed by control signals input from the sound processing device 10. In the example of Fig. 1, the control target devices 40 include an audio device 42 (e.g., car audio), an air conditioner 44 (e.g., air conditioner), a window opening / closing device 46 (e.g., power window), and a steering heater 48 (e.g., steering heater).

[0032] Next, a configuration example of the sound processing device 10 according to this embodiment will be described. The sound processing device 10 includes an A / D conversion unit 112, a communication unit 114, and a control unit 120. The sound processing device 10 also includes an input / output unit (not shown) for wirelessly or wiredly inputting and outputting various types of data to and from the sound processing device 20 and the control target device 40 using a predetermined input / output method.

[0033] The A / D (Analog-to-Digital) conversion unit 112 samples the analog acoustic signals input from the microphones 30-1 and 30-2 at a predetermined sampling frequency and converts them into digital acoustic signals. The A / D conversion unit 112 outputs the converted acoustic signals to the control unit 120. The A / D conversion unit 112 is configured to include, for example, an A / D converter.

[0034] The communication unit 114 is wirelessly connected to the communication network NW and communicates with other devices via the communication network NW. A voice recognition server may be designated as the device to be the other party. The designated voice recognition server may form a cloud 50 together with other server devices. The communication unit 114 is configured to include a communication interface that enables communication using a predetermined communication method. The communication method may be, for example, 5G (5 th Any of the following may be used: General Mobile Communication System (5th generation mobile communication system), LTE-A (Long Term Evolution - Advanced), IEEE802.11, etc.

[0035] The control unit 120 performs various types of arithmetic processing to realize and control the functions of the sound processing device 10. The control unit 120 may be realized by a dedicated component, or may be realized as a computer including a processor and storage media such as a read-only memory (ROM) and a random access memory (RAM). The control unit 120 may be configured as, for example, an engine control unit (ECU). The processor reads a predetermined program stored in advance in the ROM, deploys the read program in the RAM, and uses the storage area of ​​the RAM as a working area. The processor realizes the functions of the control unit 120 by executing processes instructed by various instructions written in the read program. The realized functions may include the functions of each unit described below. In the following description, executing processes instructed by instructions written in a program may be referred to as "executing a program" or "executing a program." The processor is, for example, a central processing unit (CPU).

[0036] The control unit 120 includes a spatial spectrum analysis unit 122, an utterance position estimation unit 124, an utterance information acquisition unit 126, an utterance position processing unit 128, a command processing unit 130, and a spatial filtering unit 134. The spatial spectrum analysis unit 122 calculates a spatial spectrum for each frame of a predetermined time length from the acoustic signals input for each channel from the A / D conversion unit 112. The spatial spectrum is a spatial feature that indicates the distribution of intensity depending on the sound source position. The spatial spectrum analysis unit 122 outputs the calculated spatial spectrum to the utterance position estimation unit 124. In this embodiment, element values ​​constituting the spatial spectrum for each candidate sound source position (hereinafter referred to as "candidate position") are calculated. The spatial spectrum is expressed as a vector including element values ​​for each candidate position. The positions of each seat in the vehicle may be used as the candidate positions. As an example, if there are five seats in the vehicle, the spatial spectrum is expressed as a five-dimensional vector.

[0037] The spatial spectrum analysis unit 122 calculates, for example, a MUSIC (Multiple Signal Classification) spectrum as the spatial spectrum. The MUSIC spectrum can be calculated using the following procedure. The spatial spectrum analysis unit 122 performs a discrete Fourier transform for each frame to calculate transformation coefficients transformed into the frequency domain. The spatial spectrum analysis unit 122 generates, for each frequency, an input vector containing, as elements, the transformation coefficients for each channel. The spatial spectrum analysis unit 122 calculates the expected value of a matrix that is the product of a transposed vector obtained by transposing the generated input vector and the input vector, as a spectral correlation matrix R. sp It is calculated as follows.

[0038]

number

[0039] In equation (1), * denotes the complex conjugate transpose operator, and E(...) denotes the expectation value of.... The spatial spectrum analysis unit 122 calculates the spectral correlation matrix R sp Solve the eigenvalue problem of i and the eigenvector e i The calculated eigenvalue λ i and the eigenvector e i The number of pairs corresponds to the number of channels. The spatial spectrum analysis unit 122 determines, for example, the number of sound sources that can be detected using (2) (hereinafter referred to as the "number of detectable sound sources"), a preset transfer function vector d(θ), and an eigenvector e i is used to calculate the element value P(θ) of the spatial spectrum for each frequency (hereinafter referred to as "frequency-specific spatial spectrum"). In this embodiment, the number of detectable sound sources may be set to 1. The transfer function vector d(θ) is a vector having, as element values, the transfer functions from the candidate position θ to the positions of the individual microphones 30-1 and 30-2 (hereinafter referred to as "sound receiving positions").

[0040]

number

[0041] In equation (2), |...| indicates the absolute value. M indicates the number of detectable sound sources and is a positive integer value less than N, which is set in advance. K is the eigenvector e i M is the number of channels. M is a positive integer equal to or less than N. N is a positive integer corresponding to the number of channels (N=2 in the example of FIG. 1). The spatial spectrum analysis unit 122 calculates the S / N ratio (Signal-to-Noise ratio) for each frequency band based on the acoustic signal of each channel, and selects a frequency band k in which the calculated S / N ratio is higher than a preset threshold value. The spatial spectrum analysis unit 122 calculates the eigenvalue λ for each frequency in the selected frequency band k. i The largest eigenvalue λ max The square root of (k) is the element value P of the spatial spectrum k (θ) is used as a weighting coefficient, and weighted sum is performed between frequency bands k to obtain the element value P of the extended spatial spectrum shown in equation (3). ext Calculate (θ).

[0042]

number

[0043] In equation (3), Ω denotes a set of frequency bands, and |Ω| denotes the number of frequency bands in the set. Therefore, the extended spatial spectrum P ext (θ) has relatively few noise components and the frequency band spatial spectrum P k The spatial spectrum analysis unit 122 calculates the element value P ext The vector containing (θ) is taken as the spatial spectrum mentioned above.

[0044] The utterance position estimation unit 124 receives the spatial spectrum from the spatial spectrum analysis unit 122 and the acoustic features from the sound processing device 20 for each frame. As will be described later, the acoustic features represent the acoustic features of the sound signal as vectors. The acoustic features include a power spectrum density (PSD) and a number of zero crossings (ZC). As an example, if the number of samples included in the small section per frame analyzed at one time by the sound processing device 20 is 512, the power spectrum density is represented by a 256-dimensional vector. The zero crossings are represented by a scalar value. In that case, the acoustic features are represented by a 257-dimensional vector. The utterance position estimation unit 124 concatenates the spatial spectrum and the acoustic features to form an input vector for a predetermined mathematical model.

[0045] The utterance position estimation unit 124 calculates an output vector for an input vector constructed for each frame using a mathematical model showing the relationship between input vectors and output vectors. The calculated output vector represents information on the estimated utterance position (hereinafter, sometimes referred to as "estimated utterance position information"). The input vector and output vector correspond to explanatory variables and target variables, respectively. The output vector includes the reliability of each predetermined speaker position as an element value. The reliability indicates the possibility that a speaker is located at that speaker position and is expressed as a real number. The larger the reliability value, the more likely the speaker is located. The range of each reliability is normalized within a predetermined range (e.g., between 0 and 1). As an example, if there are five seats in a vehicle, the output vector is expressed as a five-dimensional vector. The output vector can also be seen as information indicating the spatial distribution of the possibility that a speaker is located in that frame. For example, the utterance position with the highest calculated reliability, which is higher than a predetermined reliability threshold, is estimated as the utterance position where the speaker speaking in that frame is located (frame-based utterance position). The reliability threshold may be a value significantly larger than the expected value of reliability selected by chance. The utterance position estimation unit 124 outputs the estimated utterance position information to the utterance information processing unit 128.

[0046] The utterance position estimation unit 124 can use, for example, a random forest as a mathematical model. A random forest is a type of ensemble machine learning model that includes multiple decision trees as weak learners. The output from the random forest is the average value of the outputs from the multiple decision trees. A decision tree is a machine learning model that has multiple nodes, with one node as the root, multiple branches for each node, and a tree structure in which a separate node is connected to each branch until it reaches its end. An attribute is assigned to each node, and a value for the attribute is assigned to each branch. For each decision tree, for example, the reliability of each utterance position for an input vector is configured to be obtained as an output value. Alternatively, the system may be configured to obtain an index indicating one of the utterance positions for an input vector from each decision tree as an output value. In this case, the index indicating the utterance position from the multiple decision trees may be used as estimated utterance position information, which is output from the random forest.

[0047] The utterance information acquisition unit 126 receives an acoustic signal from the spatial filtering unit 134. The acoustic signal input to the utterance information acquisition unit 126 is derived based on two-channel acoustic signals acquired via the A / D conversion unit 112. The utterance information acquisition unit 126 transmits the input acoustic signal to the cloud 50 via the communication unit 114 and the communication network NW. The utterance information acquisition unit 126 receives, for each utterance detected from the transmitted acoustic signal, utterance section information indicating the utterance section and utterance content information indicating the content of the utterance, in association with each other, from the speech recognition server constituting the cloud 50. The utterance section information indicates the start time when one utterance started and the end time when it ended. The utterance content information describes text in which information recognized from the utterance is expressed in a natural language. The utterance information acquisition unit 126 outputs the received utterance section information to the utterance position processing unit 128 and outputs the utterance content information to the command processing unit 130.

[0048] The utterance position processing unit 128 receives the estimated utterance position information from the utterance position estimation unit 124 and the utterance section information from the utterance information acquisition unit 126 for each frame. For each predetermined speech position, the utterance position processing unit 128 identifies a section in the utterance section indicated in the utterance section information where the utterance position is estimated as a position-specific utterance section. For example, the utterance position processing unit 128 can identify the utterance position indicated by the input utterance section information with the highest reliability, and identify the section of that frame as the position-specific utterance section of the identified utterance position. When determining the speech position for each frame, the speech position processing unit 128 may identify a small section that is part of the speech section and spans multiple frames including that frame, and identify the speech position with the highest reliability in the entire small section as the speech position in that frame.

[0049] When the speech section information is expressed by an index of the speech position for each frame, the speech position processing unit 128 can identify the frame as the position-specific speech section of the speech position indicated by the index. Here, when identifying a sub-section corresponding to each frame, the speech position processing unit 128 can identify the most frequent utterance position among the utterance positions for each frame in the sub-section (by majority vote), and identify the frame corresponding to that sub-section as the position-specific speech section of the identified utterance position.

[0050] The utterance position processing unit 128 determines an utterance position in an utterance section (utterance-based utterance position) based on the position-specific utterance section identified in the utterance section. The utterance position processing unit 128 can determine an utterance position whose ratio in the utterance section is equal to or greater than a predetermined ratio and whose position-specific utterance section is the longest as the utterance position in the utterance section. The utterance position processing unit 128 outputs utterance position information indicating the determined utterance position to the command processing unit 130.

[0051] There may be a plurality of utterance positions where the ratio of position-specific utterance periods in an utterance period is equal to or greater than a predetermined ratio. In such a case, the utterance position processor 128 may determine simultaneous utterances without identifying a single utterance position. As the predetermined ratio, a value equal to or greater than the probability that an utterance position is selected by chance, more preferably a value significantly greater than that ratio, is preset in the utterance position processor 128. The probability of selection by chance may be the reciprocal of the number of seats in the vehicle.

[0052] The utterance position processing unit 128 converts the determined utterance position and related information into a predetermined format required by the command processing unit 130. For example, when a voice command is used that requires discrimination between the driver's seat (front right, FRF) and other utterance positions, the utterance position processing unit 128 includes driver's seat identification information indicating whether the utterance position is the driver's seat or not in the utterance position information. When a voice command is used that requires discrimination between the driver's seat, the passenger seat (front left, FL), and other utterance positions, the utterance position processing unit 128 may include front seat identification information indicating either the driver's seat, the passenger seat, or other utterance positions in the utterance position information. Furthermore, the utterance position processing unit 128 may include simultaneous utterance identification information indicating either simultaneous utterance (simultaneous utterance) or other utterance positions (single utterance) in the utterance position information.

[0053] The command processing unit 130 receives utterance content information from the utterance information acquisition unit 126 and utterance position information from the utterance position processing unit 128 for each utterance. The command processing unit 130 refers to a command list (not shown) set in advance in itself and determines whether or not a voice command is included in the utterance content indicated in the utterance content information. The command processing unit 130 may also refer to the command list and determine whether or not a voice command included in the utterance content is a voice command related to the utterance position (hereinafter, may be referred to as an "utterance position related voice command").

[0054] The command list corresponds to data indicating, for example, keywords, operation mode information, and target device information for each voice command. Keywords may include one or more words or phrases related to either or both of the target device and operation mode indicated by each voice command. Operation mode information may include information such as the operation mode, its elemental operation characteristics, the operation target, and the state that is the target of control. For an utterance position-related voice command, an operation mode may be set for each utterance position identified by the voice command. For target device information, one or more pieces of information such as type, name, model number, IP (Internet Protocol) address, and MAC (Media Access Control) address may be set as information that identifies or helps identify the target device. For an utterance position-related voice command, a target device may be set for each utterance position identified by the voice command.

[0055] The command processing unit 130 refers to the command list and searches for a voice command in which a word or phrase included in the input utterance content information matches the set keyword. The command processing unit 130 may perform a known morphological analysis on the text expressing the utterance content information to determine the word or phrase and part of speech expressed in the text. The command processing unit 130 may match each of the determined words or phrases with one of the commands described in the command list for each independent word. Words whose part of speech is a noun, verb, adjective, or adverb are identified as independent words.

[0056] When the command processing unit 130 detects a voice command that matches a keyword, it reads out the target device and operation mode information related to the voice command. If the detected voice command is an utterance position-related voice command and a target device is set for each utterance position, the command processing unit 130 identifies the target device corresponding to the utterance position indicated by the utterance position information. If the detected voice command is an utterance position-related voice command and an operation mode is set for each utterance position, the command processing unit 130 identifies the operation mode corresponding to the utterance position indicated by the utterance position information.

[0057] The command processing unit 130 generates control information for instructing the identified target device to operate in the identified operating mode, and transmits the generated control signal to one of the control target devices 40 as the target device. The control target device 40 waits for a control signal from the sound processing device 10, and operates in accordance with the operating mode instructed by the input control signal.

[0058] If the utterance position information includes simultaneous utterance identification information indicating simultaneous utterances, the command processing unit 130 may reject the utterance content information and output a guidance instruction to the audio device 42 to instruct the audio device 42 to respeak each individual speaker. When the audio device 42 receives a guidance instruction from the command processing unit 130, it plays a guidance voice that conveys guidance information for guiding each individual speaker to respeak. The guidance voice may convey the following messages as the guidance information: for example, "Please speak again one by one in order," "Please speak again so that two or more people do not overlap," etc.

[0059] The spatial filtering unit 134 receives two-channel acoustic signals from the A / D conversion unit 112 and receives utterance position information from the utterance position estimation unit 124. The spatial filtering unit 134 performs spatial filtering on the acoustic signals for each channel based on the utterance position information to generate a filtered one-channel acoustic signal. The spatial filtering unit 134 outputs the generated acoustic signals to the utterance information acquisition unit 126 as acoustic signals for acquiring utterance information.

[0060] In spatial filtering, the spatial filtering unit 134 determines, for each channel, a filter coefficient having directivity such that the gain is higher in the direction of the speech position indicated by the speech position information than in other directions. For example, a filter coefficient is set in advance for each speech position in the spatial filtering unit 134, and the spatial filtering unit 134 identifies a filter coefficient corresponding to the speech position indicated by the speech position information. The spatial filtering unit 134 performs filtering processing on the acoustic signal of the corresponding channel using the identified filter coefficient, and adds the processed acoustic signals between channels to obtain an added signal as the filtered acoustic signal.

[0061] The spatial filtering unit 134 may use, for example, a known method for controlling directivity so that the gain in the direction of the identified sound source position is higher than in other directions, such as a delay-and-sum method or a filter-and-sum beamformer, as the spatial filtering process. The spatial filtering unit 134 may use a sound source separation processing method that can separate or extract sound from the direction of the identified sound source position from sound from other directions, such as a GHDSS (Geometric High-order Decorrelation-based Source Separation) method, to acquire a one-channel acoustic signal for acquiring speech information that arrives from the identified speech position from two-channel acoustic signals.

[0062] Next, a configuration example of the sound processing device 20 according to this embodiment will be described. The sound processing device 20 includes an A / D conversion unit 212 and a control unit 220. The sound processing device 20 includes an input / output unit (not shown) for wirelessly or wiredly inputting and outputting various data to and from the sound processing device 10 and the control target device 40 using a predetermined input / output method.

[0063] A / D conversion unit 212 samples the analog acoustic signals input from microphones 30-1 to 30-3 at a predetermined sampling frequency and converts them into digital acoustic signals. A / D conversion unit 212 outputs the converted acoustic signals of each channel to control unit 220. A / D conversion unit 212 is configured to include, for example, an A / D converter. In the following description, the channels corresponding to microphones 30-1 to 30-3, respectively, will be referred to as channels 1 to 3.

[0064] The control unit 220 performs various types of arithmetic processing to realize and control the functions of the sound processing device 10. The control unit 220 may be realized by a dedicated component, or may be realized as a computer including a processor and a storage medium. The control unit 220 may be configured as, for example, an ECU (Engine Control Unit). The processor reads out a predetermined program stored in advance in a ROM, deploys the read program in a RAM, and uses the storage area of ​​the RAM as a working area. The processor realizes the functions of the control unit 220 by executing processes instructed by various commands written in the read program. The realized functions may include the functions of each unit described below.

[0065] The control unit 220 includes an acoustic feature analysis unit 222 , a low-pass filter 224 , and a noise reduction unit 226 . The acoustic feature analysis unit 222 analyzes the acoustic feature for each frame of a predetermined length for the acoustic signal of channel 3 input from the A / D conversion unit 212. The acoustic feature calculated for each frame is output to the speech position estimation unit 124 of the sound processing device 10 in association with the power spectrum and the number of zero crossings. The acoustic feature analysis unit 222 includes a frequency analysis unit 222a and a zero crossing point analysis unit 222b.

[0066] The frequency analysis unit 222a analyzes features indicating frequency characteristics as acoustic features. The frequency analysis unit 222a calculates power spectral density as the acoustic feature. The power spectral density is the power per unit frequency. The frequency analysis unit 222a performs a discrete Fourier transform on the acoustic signal input for each frame to calculate transform coefficients in the frequency domain, and can calculate the absolute value of the square of the obtained transform coefficients as the power spectrum. The power spectral density can be a clue for determining whether or not speech is being heard. The zero crossing point analysis unit 222b detects the zero crossing points of the input acoustic signal for each frame. A zero crossing point is a point at which the signal value of each sample constituting the acoustic signal changes from a positive value to a negative value, or from a negative value to a positive value. The zero crossing point analysis unit 222b determines the number of zero crossing points detected for each frame as the zero crossing count. The zero crossing count is also a type of acoustic feature.

[0067] Low-pass filter 224 mainly passes low-frequency components below a predetermined cutoff frequency (for example, 50 to 200 Hz) from the acoustic signals of channels 1 to 3 input from A / D conversion unit 212. The low-frequency components mainly consist of noise components and contain almost no spoken audio components. Low-pass filter 224 outputs an acoustic signal indicating the passed low-frequency components for each channel to noise reduction unit 226.

[0068] The noise elimination unit 226 receives an audio signal indicating low-frequency components from the low-pass filter 224. The noise elimination unit 226 performs noise elimination processing using the audio signal of channel 3 as a noise signal, and eliminates noise components contained in the audio signals of channels 1 and 2. The noise elimination unit 226 realizes, for example, active noise control (ANC) as the noise elimination processing.

[0069] To achieve ANC, the noise cancellation unit 226 is connected to a speaker (not shown) for presenting the cancellation sound and includes an adaptive filter. The adaptive filter is used to estimate filter coefficients that indicate the transmission paths of the cancellation sound from the speaker to each of the microphones 30-1 and 30-2. The adaptive filter determines the filter coefficients so that the intensity of the acoustic signals indicating the low-frequency components extracted for channels 1 and 2 approximates (minimizes) zero. The noise cancellation unit 226 generates a cancellation signal by performing a convolution operation on the noise signal using the filter coefficients determined by the adaptive filter. For example, the LMS (Least Mean Square) method can be used to determine the filter coefficients. The noise cancellation unit 226 supplies the generated cancellation signal to the speaker. The speaker emits a cancellation sound based on the cancellation signal supplied from the noise cancellation unit 226. Therefore, the canceling sound coming from the speaker and the noise component coming from the noise source cancel each other out at the microphones 30-1 and 30-2, so that an acoustic signal from which the noise component has been removed or reduced is acquired.

[0070] As described above, the operations of the sound processing devices 10 and 20 are individually controlled by the control units 120 and 220, respectively. Synchronization of operations between the sound processing devices 10 and 20 is not guaranteed. That is, a time difference and fluctuation may occur on a sample-by-sample basis between the sound signal used for spatial spectrum analysis in the spatial spectrum analysis unit 122 and the sound signal used for acoustic feature analysis in the acoustic feature analysis unit 222. This time difference is difficult to distinguish from differences in inter-channel phase differences due to the speech position (sound source position). Therefore, it is not realistic to calculate a spatial spectrum by directly using two-channel sound signals acquired by the sound processing device 10 and one-channel sound signals acquired by the sound processing device 20. Furthermore, even if a spatial spectrum calculated for each frame from two-channel sound signals and an acoustic feature calculated for each frame from a different one-channel sound signal are simultaneously calculated, this does not necessarily lead to an improvement in the accuracy of estimation of the speech position.

[0071] In this embodiment, the estimation accuracy can be improved by using a mathematical model that shows the relationship between the known spatial spectrum and the speech position as a relationship between an explanatory variable and a target variable. By further referring to acoustic features as explanatory variables, the speech position can be estimated using the variation in acoustic features due to the speech position as a clue even when synchronization with the spatial spectrum is not guaranteed, thereby improving the estimation accuracy.

[0072] A parameter set of a mathematical model is set in advance in the speech position estimation unit 124. The sound processing device 10 may be provided with a model learning unit (not shown) for calculating the parameter set by learning. Training data is set in advance in the model learning unit. The training data is configured to include a large number of training sets. One training set includes known input vectors that serve as explanatory variables and output vectors that serve as objective variables, and these are associated with each other. For example, as the objective variable, a vector is set as an output vector representing a certain speech position, in which the value of the element of the dimension corresponding to that speech position is 1 and the value of the element of the dimension corresponding to other speech positions is 0.

[0073] The model training unit recursively updates the parameter set so that the overall difference between an estimated value obtained by performing a calculation using a mathematical model on a known input vector and an output vector corresponding to that input vector becomes small. The model training unit can use, as a loss function indicating the magnitude of the difference, for example, any one of sum of squares error, cross entropy, etc., or a linear combination of any of these. The model training unit can use, for example, a gradient method to update the parameter set.

[0074] When a random forest is used as the mathematical model, the model learning unit may perform the following steps in learning: (1) Randomly sampling the entire training data using the bootstrap method to classify it into B (B is an integer equal to or greater than 2) subsamples. Each subsample includes multiple training sets. (2) Using each subsample as training data, B decision trees are generated. (3) For each decision tree, the following steps are performed to generate nodes until a predetermined number of nodes is reached: (3-1) Randomly select a portion of the explanatory variables of the training data. (3-2) Among the selected explanatory variables, the explanatory variable that best classifies the training data and the threshold used for that classification are determined as the threshold used to classify the new node. By using the randomly sampled training data and the randomly selected explanatory variables, a group of decision trees with low correlation is generated. This enables fast learning.

[0075] Next, examples of voice operations executed by the command processing unit 130 will be described. FIG. 4 is a table illustrating voice operations corresponding to voice commands. In the illustrated example, an utterance of "play music" commands music playback regardless of the utterance position. In this case, there is no need to distinguish between the utterance positions. Keywords related to the voice command can be detected, for example, by detecting that the utterance contains the words "music" and "play." Examples of utterances, functions for each utterance position, and required functions are shown for each type of voice operation. Five utterance positions are listed: driver's seat (FR), passenger seat (FL), rear right (RR), rear middle (RM), and rear left (RR).

[0076] The types of voice operations include operations unrelated to the speaking position, operations related to the speaking position, operations related to safety, operations by passengers that may distract the driver, and simultaneous speech. An operation unrelated to the speaking position refers to realizing a predetermined function according to a recognized voice command regardless of the speaking position. In the illustrated example, in response to the utterance "Play music," the command processing unit 130 instructs the audio device 42 to play music regardless of the speaking position. In this case, there is no need to distinguish based on the speaking position. The command processing unit 130 can detect that the utterance contains keywords, such as "music" and "play," and recognize a voice command related to music playback.

[0077] An operation related to an utterance position refers to realizing a predetermined function depending on the utterance position in accordance with a recognized utterance position-related voice command. In the illustrated example, in response to the utterance "Turn down the air conditioner temperature," the command processing unit 130 instructs the air conditioner 44 to lower the temperature at the utterance position. However, if the utterance position is a rear seat, it is ignored. In this case, it is necessary to refer to the utterance position information and distinguish at least between the driver's seat, the passenger seat, and other positions. The command processing unit 130 can recognize a voice command related to lowering the temperature by detecting, for example, the inclusion of keywords such as "air conditioner," "temperature," and "turn down" from the utterance content.

[0078] An operation related to safety design is performed according to a voice command uttered by the driver, and the execution of such a command by passengers other than the driver is restricted. In the illustrated example, in response to an utterance from the driver's seat of "Open the window (other than my seat)," the command processing unit 130 instructs the window opener 46 to open the specified window other than the driver's seat. However, if the utterance is from a seat other than the driver's seat, the command is ignored. In this case, it is necessary to refer to the utterance position information and distinguish at least between the driver's seat and other positions. In other words, the voice command is determined to be invalid in other positions. The command processing unit 130 can detect keywords from the utterance, such as the inclusion of words related to the positions of seats other than the driver's seat (e.g., "passenger seat," "right rear," "left rear," etc.), the words "window," and "open," and recognize the voice command for opening the window at that position. Note that voice commands related to operations related to safety design can also be considered as utterance position-related voice commands.

[0079] An operation that may distract the driver due to a passenger's operation is implemented according to a voice command uttered by the driver, and implementation by a voice command from a passenger other than the driver is restricted. In the illustrated example, the command processing unit 130 instructs the steering heater 48 to be heated in response to the utterance "Steering heater, on" from the driver's seat. However, if the utterance is from a seat other than the driver's seat, the command is ignored. In this case, it is necessary to refer to the utterance position information and distinguish at least between the driver's seat and other positions. In other words, in other positions, the voice command is determined to be invalid. The command processing unit 130 can detect the inclusion of keywords, such as "steering heater" and "on," from the utterance content and recognize the voice command for heating the steering heater 48. Note that a voice command related to an operation that may distract the driver due to a passenger's operation can also be considered an utterance position-related voice command.

[0080] Simultaneous speech refers to speech occurring at multiple seats in a single speech section. In this case, even if a voice command can be detected from the speech content, the command processing unit 130 rejects the detected voice command. The command processing unit 130 outputs a guidance instruction to the audio device 42 and plays a guidance voice to guide each speaker to resume speaking. In this case, it is sufficient to refer to the speech position information and detect simultaneous speech.

[0081] Next, an example of the voice operation process according to this embodiment will be described. FIG. 5 is a flowchart showing an example of the voice operation process according to this embodiment. (Step S102) The spatial spectrum analysis unit 122 calculates a spatial spectrum for each frame based on the two-channel acoustic signals input from the microphones 30-1 and 30-2. (Step S104) The frequency analysis unit 222a analyzes the power spectrum density for each frame based on the one-channel acoustic signal input from the microphone 30-3. (Step S106) The zero crossing point analysis unit 222b analyzes the zero crossing points for each frame based on the one-channel acoustic signal input from the microphone 30-3, and counts the number of zero crossing points. (Step S108) The utterance position estimation unit 124 constructs an input vector containing the power spectral density, the power spectral density, and the zero crossing point as elements for each frame. Using a mathematical model, the utterance position estimation unit 124 calculates an output vector containing the reliability of each utterance position as elements from the constructed input vector, and estimates utterance position information (frame-based utterance position).

[0082] (Step S110) The spatial filtering unit 134 performs spatial filtering on the two-channel acoustic signals input from the microphones 30-1 and 30-2, and acquires one-channel acoustic signal by directing the directivity to the estimated speaking position. (Step S114) The utterance information acquisition unit 126 transmits the acquired one-channel acoustic signal to the cloud 50, and acquires utterance section information and utterance content information for each utterance from the cloud 50. (Step S116) The utterance position processing unit 128 determines position-specific utterance sections based on the utterance position information estimated for each frame in the utterance section indicated in the acquired utterance section information. The utterance information acquisition unit 126 determines the utterance position with the largest number of position-specific utterance sections as the utterance position for that utterance section (utterance-based utterance position).

[0083] (Step S118) The command processing unit 130 refers to the command list and determines whether or not a voice command can be detected from the utterance content indicated in the acquired utterance content information. If it is determined that the voice command can be detected (step S118 YES), the process proceeds to step S120. If it is determined that the voice command cannot be detected (step S118 NO), the process of FIG. 5 ends. (Step S120) The command processing unit 130 determines whether the utterance position information indicates simultaneous utterances. If it is determined that the utterance position information indicates simultaneous utterances (YES in step S120), the process proceeds to step S122. If it is determined that the utterance position information indicates single utterances (NO in step S120), the process proceeds to step S124. (Step S122) The command processing unit 130 causes the audio device 42 to present a guidance voice as guidance information for instructing re-speaking guidance for each individual speaker. After that, the process of FIG. 5 ends.

[0084] (Step S124) The command processing unit 130 refers to the command list and determines whether the detected voice command is an utterance position related voice command. If it is determined to be an utterance position related voice command (step S124 YES), the process proceeds to step S128. If it is determined not to be an utterance position related voice command (step S124 NO), the process proceeds to step S126. (Step S126) The command processing unit 130 executes operational control, unrelated to the utterance position, on the control-target device 40 instructed by the voice command, in accordance with the detected voice command. Then, the processing of FIG. 5 ends.

[0085] (Step S128) The command processing unit 130 refers to the command list and determines whether the voice command detected at the utterance position indicated in the utterance position information is valid. If it is determined to be valid (YES in step S128), the process proceeds to step S130. If it is determined to be invalid (NO in step S128), the detected voice command is rejected, and the process in FIG. 5 ends. (Step S130) The command processing unit 130 executes operational control related to the utterance position in accordance with the detected voice command for the control-target device 40 instructed by the voice command. Then, the processing of FIG. 5 ends.

[0086] Next, a first example of a method for determining frame-based utterance positions will be described in more detail. FIG. 6 is an explanatory diagram showing a first example of detection of an utterance position. In the illustrated example, since the shift length is shorter than the frame length, utterance position information is acquired at time intervals corresponding to the shift length. In this example, an utterance position where the spatial spectrum value for each candidate utterance position is equal to or greater than a predetermined value and is maximum is detected. As exemplified in FIG. 7, the estimated utterance position (estimated utterance position) may not be continuous across multiple frames, but may be acquired intermittently on a frame-by-frame basis. If the speaker is seated, the ground truth of the utterance position for each utterance should be consistent for each utterance. Determining the utterance position for each utterance requires stability of the utterance position.

[0087] FIG. 8 is an explanatory diagram showing a second example of utterance position detection. In the illustrated example, the utterance position processing unit 128 sets, for each frame, a subsection spanning multiple frames including that frame, and determines the utterance position with the highest reliability in the utterance position information for the set subsection, where the reliability is equal to or greater than a predetermined reliability threshold, as the utterance position for that frame. The frame length, shift length, and time length of the subsection are typically, for example, 20 to 50 ms, 10 to 20 ms, or 300 to 1000 ms. In the example of FIG. 8, the subsection T for each frame is a period centered on the target frame to be processed, including one or more preceding frames that precede the target frame and one or more succeeding frames that follow the target frame. By determining the utterance position for each frame for each corresponding subsection T, stable utterance position estimation is achieved. As illustrated in FIG. 9, the estimated utterance position continues across multiple frames, and the intermittent phenomenon is eliminated, approximating the true value. However, even though an utterance is made and the utterance position should be estimated, a state in which the utterance position cannot be detected may continue over a plurality of frames.

[0088] The utterance position processing unit 128 according to this embodiment can determine the utterance-based utterance position using the method described below. Fig. 10 is a flowchart showing a first example of conversion to the utterance-based utterance position according to this embodiment. (Step S202) The utterance position processing unit 128 selects the m-th utterance R m Select . (Step S204) The utterance position processing unit 128 initializes the accumulated detection time P(K) for each predetermined utterance position K to zero.

[0089] (Step S206) The utterance position processing unit 128 calculates the utterance R for each utterance position K. m The section in which the utterance position K is estimated in the utterance section is defined as the position-specific utterance section V i (i is an integer between 1 and N, and N is an integer indicating the number of detected position-specific speech segments). (Step S208) The utterance position processing unit 128 calculates the position-specific utterance section V iAnother section length L i are added together to calculate the cumulative detection period P(K). (Step S210) The utterance position processing unit 128 calculates the utterance position K at which the cumulative detection period P(K) is maximized as the utterance R. m The process of FIG. 10 then ends.

[0090] FIG. 11 is an explanatory diagram showing an example of the first conversion example to an utterance-based utterance position. In the illustrated example, T s (m), T e (m) is the utterance R m L1 indicates the start time and end time of the utterance section of the position-specific utterance section V1. V1 to V3 are detected as position-specific utterance sections included in the utterance section. L1 to L3 indicate the section lengths of the position-specific utterance sections V1 to V3 detected within the utterance section, respectively. The utterance positions for the position-specific utterance sections V1, V2, and V3 are estimated to be 0, 0, and 1, respectively. In this case, the cumulative detection period P(0) for utterance position 0 is L1+L2, and the cumulative detection period P(1) for utterance position 1 is L3. The cumulative detection periods P(2) to P(4) for utterance positions 2 to 4 are all 0. Utterance positions 0, 1, 2, 3, and 4 indicate the driver's seat (FR), passenger seat (FL), rear right (RR), rear middle (RM), and rear left (RR), respectively. In this case, the utterance position processing unit 128 determines the utterance position 0 with the longest cumulative detection period as the utterance R. m Utterance-based utterance location SSL(R m ) can be defined as

[0091] 12 is a flowchart showing a second example of conversion to an utterance-based utterance position according to this embodiment. The example in FIG. 12 also includes determination of simultaneous utterances. (Step S222) The utterance position processing unit 128 selects the m-th utterance U as the utterance to be processed. m Select . (Step S224) The utterance position processing unit 128 calculates the subsection S including the f-th frame in the selected utterance. f Select . (Step S226) The utterance position processing unit 128 calculates the cumulative detection period Q for each utterance position K. fInitialize (K) to 0.

[0092] (Step S228) The utterance position processing unit 128 calculates the subsection S f The section where the utterance position K is estimated is called the position-specific utterance section V i Identify as: (Step S230) The utterance position processing unit 128 calculates the position-specific utterance section V i Another section length L i Add up the cumulative detection period Q f Calculate (K). (Step S232) The utterance position processing unit 128 calculates the cumulative detection period Q f The utterance position K at which (K) is maximized is the small section S f Subsection-specific utterance position K for f It is defined as follows. (Step S234) The utterance position processing unit 128 advances the frame f to be processed to the next frame. (Step S236) The utterance position processing unit 128 m It is determined whether the next frame to be processed exists in the speech section. If it is determined that the next frame exists (YES in step S236), the process proceeds to step S240. If it is determined that the next frame does not exist (NO in step S236), the process proceeds to step S238.

[0093] (Step S238) The utterance position processing unit 128 calculates the utterance U for each utterance position K. m In the speech section of f The number of positions is N K Count as follows: Number of locations N K indicates the time length in units of the number of frames of the cumulative speech period (K) in FIG. (Step S240) The utterance position processing unit 128 calculates the counted number of localizations N K It is determined whether there are a plurality of utterance positions where the ratio of the number of utterance positions to the total number of frames in the utterance section is equal to or greater than a predetermined ratio. If it is determined that there are a plurality of utterance positions (YES in step S240), the process proceeds to step S242. If it is determined that there is one utterance position (NO in step S240), the process proceeds to step S244. (Step S242) The utterance position processing unit 128 m The speech state at the time is determined to be simultaneous speech. Then, the process shown in FIG. (Step S244) The utterance position processing unit 128 calculates the number of positions N K The utterance position K where is the maximum is called utterance U m The utterance position of the utterance base is then determined as the utterance position of the utterance base. Then, the process of FIG. 12 ends.

[0094] FIG. 13 is an explanatory diagram showing an example of determining the utterance position for each small section. In the illustrated example, utterance U m In the speech section related to f The sections where utterance positions 0 and 1 are estimated are identified as position-specific utterance sections V1 and V2, respectively. L1 and L2 indicate the section lengths of the position-specific utterance sections V1 and V2, respectively. In this case, the cumulative detection period Q for utterance position 0 is f (0) is L1, and the cumulative detection period P(1) for utterance position 1 is L2. The cumulative detection period Q for utterance positions 2 to 4 is f (2)~Q f (4) are both 0. In this case, the utterance position processing unit 128 sets the utterance position 0 where the cumulative detection period is the longest as the small section S f Subsection-specific utterance position K for f It can be defined as:

[0095] 14 is an explanatory diagram showing an example of simultaneous speech determination. In the illustrated example, the speech U estimated by the cloud 50 n The utterance section (estimated utterance section) related to n Here, the number of localizations in utterance section 2 is 22, and the constants for the other utterance sections are all 0. Utterance position 2 with the highest number of localizations is utterance U n The utterance base utterance position is determined as the utterance base utterance position for . mIn reality, multiple utterances may be included at different times, such as the estimated utterance section related to (1). When a voice command is included as the content of each of the multiple utterances, it may be difficult to determine which voice command should be adopted. In this embodiment, the utterance position processing unit 128 detects events included in one utterance section as simultaneous utterances, even if the multiple utterances are temporally different. When simultaneous utterances are detected, guidance information is output to prompt the user to make individual utterances again.

[0096] In the example in Figure 14, utterance U m The estimated utterance section for contains 33 subsections, and two actual utterances R m1 , R m2 In this case, utterance R m1 , R m2 The utterance position for each of the utterances is different. m In the first half of the estimated speech section related to , the subsection-specific speech position is approximately 0, and the utterance R m1 In the latter half of the speech section, the speech position for each subsection is roughly 1, and the speech R m2 corresponds to the utterance section of utterance U m In the estimated utterance related to this, the localization number for speech position 0 is 12, and the localization number for speech position 1 is 21. The localization numbers for speech positions 2 to 4 are each 0. The ratio of the localization number for speech position 0 to the number of small sections in the speech section is 0.36 (≒12 / 33), and the ratio of the localization number for speech position 1 to the number of small sections is 0.63 (≒21 / 33), which exceeds a predetermined ratio (e.g., 0.2). Therefore, the speech state in this speech section is determined to be simultaneous speech.

[0097] Next, an evaluation experiment conducted on this embodiment will be described. The evaluation experiment was conducted from the following perspectives: (1) basic performance of the proposed method, (2) speech position detection under various conditions, and (3) channel number dependency of simultaneous speech detection under various conditions. In the evaluation experiment, voice data and running noise (noise data) were recorded by microphones 30-1 to 30-3, respectively, inside a vehicle with five seats inside the vehicle. The voice data and noise data were mixed and supplied to sound processing devices 10 and 20 installed outside the vehicle. This reproduces speech in a noisy environment. However, depending on the experimental conditions, music played by audio device 42 was also recorded as music data and further mixed. As sound sources, utterances randomly selected from 200 English utterances and 52 Japanese utterances were used.

[0098] First, (1) we explain the basic performance of the proposed method. Here, we calculated the speech detection rate (SDR) and speech localization rate (SLR) as indicators of the experimental results for each experimental condition. The speech detection rate is the ratio of the number of correctly detected speech periods to the number of correct speech periods. The speech localization rate is the ratio of the number of speech periods in which seats were correctly detected as speech positions to the number of detected speech periods. The following four methods were set as experimental conditions: (i) speech localization using a mathematical model for each frame, (ii) speech localization for each small section corresponding to each frame, (iii) speech-based speech localization based on speech positions estimated using a mathematical model for each frame, and (iv) speech-based speech localization based on speech positions estimated using a mathematical model for each frame, taking small sections into account. However, under experimental conditions (i) and (ii), whether the speech section was correctly detected was evaluated based on whether the ratio of the number of frames in which the speech position processing unit 128 could detect the speech position in the speech section to the number of frames in the speech section was 0.5 or more. 250 utterances were used for each experimental condition. 250 utterances corresponds to 50 utterances per seat. The experiment was conducted with the vehicle stopped and the windows and sunroof closed.

[0099] Figure 16 shows the speech detection rate and speech location detection rate for each speech detection method. The speech detection rate is lowest under experimental condition (i), and is even higher under experimental condition (ii). Under experimental conditions (iii) and (iv), the speech detection rate is 100%. This indicates that speech can be detected with almost certainty. Even under experimental condition (i), which has the lowest speech detection rate, the rate is 74.1%, which is better than when the spatial spectrum is used directly without using a model. This confirms that the phenomenon of incorrectly detecting speech segments and resulting in intermittent speech is mitigated. The detection rate of utterance position was lowest for experimental condition (ii), and increased in the order of experimental conditions (i), (iii), and (iv). The detection rate of utterance position was 88.5% for experimental condition (ii), and 95.9%, 99.6%, and 100% for experimental conditions (ii), (iii), and (iv). This indicates that the utterance position can be detected with almost certainty.

[0100] Next, we will explain (2) Speech position detection under various conditions. In speech position detection under various conditions, we calculated the speech detection rate and speech position detection rate for each of the different combinations of the number of acoustic signal channels, vehicle operating state, and in-vehicle environment as experimental conditions. Two conditions were set for the number of acoustic signal channels (2 and 3 channels), two or five vehicle operating states, and 14 in-vehicle environment conditions.

[0101] The two-channel case refers to a case where the frame-based speech position is detected based on the spatial spectrum derived from the acoustic signals picked up by microphones 30-1 and 30-2, while the three-channel case refers to a case where the frame-based speech position is detected using not only the spatial spectrum but also the power spectral density and zero-crossing count derived from the acoustic signal picked up by microphone 30-3. Two operating conditions were set: stopped and cruising at 45 mph. Five operating conditions were set: a: stopped, b: idling, c: cruising at 45 mph, d: cruising at 65 mph, and e: cruising at 72 mph. The noise levels at 65 mph and 72 mph are 5 dB and 10 dB higher, respectively, than at 45 mph.

[0102] The 14 vehicle interior environments were as follows: A: Basic (all windows closed, sunroof closed), B: HATS (Head and Torso Simulator) with mask, C: HATS facing inward, D: HATS facing outward, E: Sunroof open, F: FL windows open, G: FR windows closed, H: RL windows open, I: RR windows open, J: FLFR windows open, K: RLRR windows open, L: All windows open, sunroof open, M: Reclining HATS lying down, N: Reclining HATS standing up. HATS refers to a dummy head with a torso attached, seated in the driver's seat. The HATS was installed in the vehicle to simulate the reflection, diffraction, and absorption of sound propagation caused by a person inside. B, C, D, K, and L above represent the differences in the placement of the dummy head in the basic configuration, i.e., with all windows and the sunroof closed. HATS with mask refers to a dummy head wearing a mask around its mouth and seated in the driver's seat facing forward. HATS inward-facing refers to the state in which the dummy head is tilted 90 degrees inward and seated in the driver's seat. HATS outward-facing refers to the state in which the dummy head is tilted 90 degrees outward and seated in the driver's seat.

[0103] Figure 17 shows the speech location detection rate and speech detection rate obtained for 14 different in-car environments while the vehicle was stopped. The five-seat evaluation refers to distinguishing between five locations: the driver's seat, the passenger seat, the rear right, the middle rear, and the rear left. The three-seat evaluation refers to distinguishing between three locations: the passenger seat, the rear left, and the middle rear. (a) shows the speech location detection rate for three channels, (b) shows the speech location detection rate for two channels, (c) shows the speech detection rate for three channels, and (d) shows the speech detection rate for two channels. For three channels, both the speech location detection rate and speech detection rate exceed 95%. For two channels, both the speech location detection rate and speech detection rate exceed 95% except for in-car environment D. The speech location detection rate for in-car environment D is also approximately 92%. Furthermore, no significant differences were found in either the speech location detection rate or speech detection rate between the five-seat and three-seat evaluations.

[0104] Figure 18 shows the speech localization rate and speech detection rate obtained for 15 different in-car environments while traveling at 45 mph. The 15 in-car environments include in-car environments A through O as well as in-car environment A'. In-car environment A' refers to music playback in basic mode A. (a) shows the speech localization rate for three channels, (b) shows the speech localization rate for two channels, (c) shows the speech localization rate for three channels, and (d) shows the speech localization rate for two channels. For three channels, both the speech localization rate and speech detection rate exceed 95% except for in-car environments J and L. However, for two channels, the speech localization rate drops below 90% for in-car environments J and L as well as B, C, D, F, G, and H. This indicates that referring to acoustic features based on a single-channel acoustic signal in speech estimation contributes to improving the speech localization rate even in noisy driving environments.

[0105] Figure 19 shows the speech position detection rate and speech detection rate obtained for five different vehicle operating states in in-vehicle environment A. (a) shows the speech position detection rate for three channels, (b) shows the speech position detection rate for two channels, (c) shows the speech detection rate for three channels, and (d) shows the speech detection rate for two channels. For all operating states, both the speech position detection rate and speech detection rate exceed 95%, regardless of the number of channels.

[0106] Next, we will explain (3) the channel number dependency of simultaneous speech detection under various conditions. In simultaneous speech detection under various conditions, the simultaneous speech detection rate, simultaneous speech detection accuracy, and single speech detection rate were calculated for each of the experimental conditions, which were different combinations of the number of acoustic signal channels and the vehicle operating state. However, the vehicle interior environment was set to all windows closed and the sunroof closed. The simultaneous speech detection rate (SSDR) refers to the ratio of speech segments that were correctly detected as simultaneous speech to the number of detected speech segments. The simultaneous speech detection accuracy (SSDA) refers to the ratio of the difference between the number of speech segments that were correctly detected as simultaneous speech and the number of speech segments that were incorrectly detected as simultaneous speech to the number of detected speech segments. The single speech detection rate (SinSDR) refers to the ratio of the number of speech segments that were correctly detected as single speech to the number of detected speech segments. For example, for M speech segments U1 to U M Assume that there are N1 simultaneous speech segments and N2 independent speech segments. Of the simultaneous speech segments, the number of segments in which simultaneous speech was correctly detected and the number of segments detected as independent speech are denoted as C1 and S1, respectively. Of the independent speech segments, the number of segments in which independent speech was correctly detected and the number of segments detected as simultaneous speech are denoted as C2 and S2, respectively. In this case, the simultaneous speech detection rate SSDR is C1 / N1. The simultaneous speech detection accuracy SSDA is (C1-S2) / N1. The independent speech detection rate SinSDR is C2 / N2.

[0107] Two types of acoustic signal channels, two channels and three channels, and five types of vehicle operating conditions, a to e, were set. For the simultaneous speech evaluation, 50 evaluation files W1 to W2 were selected from a total of 250 audio data sets, including Japanese and English audio. 50 For each pair of evaluation data with adjacent indices, such as [W1, W2], [W2, W3], etc., five patterns of simultaneous utterance data were generated. Each evaluation data set contains one utterance.

[0108] FIG. 15 shows examples of simultaneous speech patterns. (a) No simultaneous speech, (b) Previous speech W m+1 The latter part of the following utterance W m+2 overlaps with the first half of (c) the preceding utterance W m+1 The period is the subsequent utterance W m+2 (d) the subsequent utterance W m+1 The first half of the preceding utterance W m+2 overlaps with the latter part of (e) the preceding utterance W m+1 The end of the following utterance W m+2 In pattern (a), the preceding utterance W m+1 From the end of the following utterance W m+2 The period G1 from the beginning of the following utterance to the end of the preceding utterance was randomly selected from a predetermined range of 100 to 500 ms. In patterns (b) and (d), the period G2 from the beginning of the following utterance to the end of the preceding utterance was randomly selected from a predetermined range of 100 to ml. ml is half the length of the shorter of the preceding or following utterance.

[0109] Figure 20 shows the simultaneous speech detection rate, simultaneous speech detection accuracy, and single-utterance detection accuracy obtained for five vehicle operating states as simultaneous speech evaluations. (a) shows the three-channel simultaneous speech evaluation, and (b) shows the two-channel simultaneous speech evaluation. The simultaneous speech detection rate, simultaneous speech detection accuracy, and single-utterance detection accuracy were all around 90%. There is a tendency for single-utterance detection accuracy to decrease as the driving speed increases, but this is not significant at practical driving speeds (below 45 mph).

[0110] (Variation) The above-described embodiment may be realized with modifications, including replacing some components with other components, combining some components with other components, and omitting some components. For example, the acoustic feature analysis unit 222 can use any type of feature as long as it can represent the features of the uttered speech as an acoustic feature. For example, instead of the spectral power density, any of Mel-Frequency Cepstrum (MFCC), Delta Cepstrum, Linear Prediction Coefficients (LPC), etc. may be used. Also, the derivation of the zero-crossing number may be omitted.

[0111] Although the example has been described in which the spatial spectrum analysis unit 122 calculates the MUSIC spectrum, the present invention is not limited to this. The spatial spectrum analysis unit 122 may calculate other types of spatial spectra, for example, spatial spectra obtained by beamforming in the direction of each sound source position. Beamforming is a method of controlling directivity by adding a gain and / or delay that differs for each channel, or by filtering and adding the resulting signals.

[0112] The mathematical model used in the utterance position estimation unit 124 is not necessarily limited to a random forest, and other types of machine learning models may be used. The other types of machine learning models may be any of a convolutional neural network (CNN), a recurrent neural network, etc. 10 may include a step of determining whether or not simultaneous speech occurs as a speech state in an utterance section, similar to the process in Fig. 12. That is, when there are multiple utterance positions K where the ratio of the cumulative detection period P(K) for each utterance position K to the utterance section is equal to or greater than a predetermined ratio, the utterance position processing unit 128 determines that simultaneous speech occurs. When there is one utterance position K where the ratio of the cumulative detection period P(K) for each utterance position K to the utterance section is equal to or greater than the predetermined ratio, the utterance position processing unit 128 may determine that single speech occurs, and perform the process of step S210.

[0113] As long as the sound processing device 20 acquires a sound signal from the microphone 30-3 and includes the sound feature analysis unit 222, it does not necessarily have to be a device whose main purpose is to remove noise. The components corresponding to the spatial filtering unit 134 and the speech information acquisition unit 126 of the sound processing device 10 may be provided in separate devices in the sound processing system S1 and may be omitted from the sound processing device 10. The spatial filtering unit 134 may be omitted, and the speech information acquiring unit 126 may receive an acoustic signal of any one channel from the A / D converting unit 112 .

[0114] A part or all of the sound processing device 10 may constitute a part of the acoustic device 42 . The sound processing devices 10 and 20 may be configured as an integrated unit. In this case, one of the control units 120 and 220 may have the function of the other, and the other may be omitted. One of the A / D conversion units 112 and 212 may have the function of the other, and the other may be omitted. The number of channels of the acoustic signal used to calculate the spatial spectrum is not limited to 2, and may be 3 or more. In addition, the parameters related to the above processing can be set arbitrarily as long as the desired effects of this embodiment can be achieved.

[0115] The speech information acquisition unit 126 may detect a speech interval by performing a known voice activity detection (VAD) process on the acoustic signal of one of the channels. More specifically, the speech information acquisition unit 126 may calculate the number of zero crossings and power for each frame, and determine a section starting from the point when the calculated power is equal to or greater than a predetermined power threshold and the number of zero crossings is equal to or greater than a predetermined frequency (e.g., 200 to 500 times per second) for a predetermined period of time (e.g., 0.2 to 0.5 seconds) as a speech interval, and ending from the point when this state ends as an end point. The speech information acquisition unit 126 may also use a mathematical model to determine whether each frame belongs to a speech interval based on the number of zero crossings and power. The mathematical model may be used to learn in advance the relationship between the pair of the number of zero crossings and power, which are explanatory variables, and whether the frame belongs to a speech interval, which is a target variable. When the sound processing devices 10 and 20 are configured as an integrated device, the number of zero crossings calculated by the sound feature analysis unit 222 and the power derived from the power spectral density may be used in the voice detection process.

[0116] Furthermore, the speech information acquisition unit 126 may perform known speech recognition processing on the detected speech section to determine the speech content. When the speech information acquisition unit 126 itself detects the speech section and determines the speech content, it is not necessary to receive speech content information and speech section information from the cloud 50. The speech information acquisition unit 126 does not need to transmit the acquired acoustic signal to the cloud 50 and request speech section detection and speech recognition processing.

[0117] As described above, the acoustic processing system S1 according to this embodiment includes an utterance position estimation unit 124 that estimates the speech position where an utterance was made for each frame of a predetermined period based on the spatial spectrum of the sound source analyzed from acoustic signals of multiple channels; an utterance information acquisition unit 126 that acquires speech section information indicating the speech section where sound was spoken based on acoustic signals of at least one channel of the multiple channels; and an utterance position processing unit 128 that identifies, for each speech position, a position-specific speech section in which the utterance position was estimated in the speech section and determines the speech position in the speech section based on the amount of position-specific speech sections. According to this configuration, the speech position in the speech section is determined based on the amount of the position-specific speech section estimated for each speech position in the estimated speech section. Even if the speech position estimated for each frame is unstable, the speech position in the speech section is uniquely determined. Therefore, the speech position can be stably determined for each utterance.

[0118] The utterance position processing unit 128 may identify, for each predetermined utterance position, a frame in which the utterance position is identified as a position-specific utterance section. According to this configuration, for a given speech position estimated for each frame, that frame is included in the position-specific speech section for that speech position, so that the position-specific speech section used to determine the speech position can be immediately determined.

[0119] The utterance position processing unit 128 may identify, for each frame, a small section that is part of the utterance section and spans multiple frames including the frame, and identify the utterance position in the frame based on the utterance position related to the small section. According to this configuration, the speech position of the corresponding frame is determined by referring to the speech position in a small section spanning multiple frames, and the frame is included in the position-specific speech section related to the speech position. Therefore, the speech position for each frame can be determined more stably than when only the speech positions of individual frames are referred to. As a result, the speech position in the speech section can be determined more stably.

[0120] The voice command indicated by the speech content in the speech section may be a speech position-related voice command related to a specific speech position, and when the speech position is the specific speech position, a command processing unit 130 may be provided which executes processing based on the speech position-related voice command. According to this configuration, a process based on an utterance position-related voice command instructed by a speech content from a specific utterance position is executed, and a process based on the utterance position-related voice command instructed by a speech content from a position other than the specific utterance position is not executed, thereby making it possible to more reliably restrict the process for a speaker located at a specific utterance position.

[0121] The utterance position processing unit 128 may determine that simultaneous utterances occur in an utterance period when there are a plurality of utterance positions in which the ratio of position-specific utterance periods in the utterance period is equal to or greater than a predetermined ratio. According to this configuration, it is possible to estimate the possibility that multiple utterances were made in a speech section, thereby providing a clue to resolve the situation where it is unclear which utterance content to use to identify the voice command.

[0122] When simultaneous speech is determined in the speech section, the command processing unit 130 may output guidance information that guides individual re-speech. According to this configuration, multiple speakers at different positions are prompted to individually re-speak, which facilitates the resolution of an issue where processing based on a voice command is not executed due to simultaneous speech.

[0123] One embodiment of the present invention has been described in detail above with reference to the drawings, but the specific configuration is not limited to that described above, and various design changes and the like are possible within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]

[0124] S1...sound processing system, 10, 20...sound processing device, 30 (30-1 to 30-3)...microphone, 40...control target device, 42...acoustic equipment, 44...air conditioner, 46...window opener, 48...steering heater, 50...cloud, 112...A / D conversion unit, 114...communication unit, 120...control unit, 122...spatial spectrum analysis unit, 124...utterance position estimation unit, 126...utterance information acquisition unit, 128...utterance position processing unit, 130...command processing unit, 134...spatial filtering unit, 212...A / D conversion unit, 222...acoustic feature analysis unit, 222a...frequency analysis unit, 222b...zero crossing point analysis unit, 224...low pass filter, 226...noise elimination unit

Claims

1. an utterance position estimation unit that estimates an utterance position where an utterance is made for each frame of a predetermined period based on a spatial spectrum of a sound source analyzed from acoustic signals of multiple channels; an utterance information acquisition unit that acquires utterance section information indicating an utterance section in which a voice is uttered based on an acoustic signal of at least one channel of the plurality of channels; an utterance position processing unit that specifies, for each utterance position, a position-specific utterance section in which the utterance position is estimated by the utterance position estimation unit in the utterance section, and determines the utterance position in the utterance section where the position-specific utterance section is longest as the utterance position in the utterance section. Sound processing system.

2. The utterance position processing unit For each predetermined speech position, a frame in which the speech position is identified is identified as the position-specific speech section. The sound processing system of claim 1 .

3. The utterance position processing unit identifying, for each frame, a small section that is part of the speech section and spans a plurality of frames including the frame; Identifying an utterance position in the frame based on the utterance position in the small section The sound processing system of claim 1 .

4. a command processing unit configured to execute processing based on an utterance position-related voice command when the voice command indicated by the utterance content in the utterance section is an utterance position-related voice command related to a specific utterance position and the utterance position is the specific utterance position; The sound processing system of claim 1 .

5. The utterance position processing unit When there are a plurality of speech positions in which the ratio of the positional speech periods in the speech period is equal to or greater than a predetermined ratio, the speech period is determined to be simultaneous speech. The sound processing system of claim 4 .

6. When it is determined that simultaneous utterances occur in the utterance section, the command processing unit: Output guidance information for individual recurrences The sound processing system of claim 5 .

7. an utterance position estimation unit that estimates an utterance position where an utterance is made for each frame of a predetermined period based on a spatial spectrum of a sound source analyzed from acoustic signals of multiple channels; an utterance information acquisition unit that acquires utterance section information indicating an utterance section in which a voice is uttered based on an acoustic signal of at least one channel of the plurality of channels; an utterance position processing unit that specifies, for each utterance position, a position-specific utterance section in which the utterance position is estimated by the utterance position estimation unit in the utterance section, and determines the utterance position in the utterance section where the position-specific utterance section is longest as the utterance position in the utterance section. Sound processing equipment.

8. A program for causing a computer to function as the sound processing device according to claim 7.

9. an utterance position estimation step in which an utterance position estimation unit estimates an utterance position where an utterance is made for each frame of a predetermined period based on a spatial spectrum of a sound source analyzed from acoustic signals of multiple channels; an utterance information acquisition step in which an utterance information acquisition unit acquires utterance section information indicating an utterance section in which speech is uttered based on an acoustic signal of at least one channel of the plurality of channels; an utterance position processing step in which an utterance position processing unit specifies, for each utterance position, a position-specific utterance section in which the utterance position is estimated by the utterance position estimation unit in the utterance section, and determines the utterance position in the utterance section with the longest position-specific utterance section as the utterance position in the utterance section. Acoustic processing methods.

Citation Information

Patent Citations

  • Sound signal processor, sound signal processing method, and program

    JP2015155975A

  • Voice processing device, voice processing method and program

    JP2018169473A

  • Voice recognition system, voice recognizing method and its program

    WO2006025106A1