Signal extraction device and program
Patent Information
- Application Number
- PCT/JP2025/012225
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2026-10-01
Smart Images

Figure JP2025012225_01102026_PF_FP_ABST
Abstract
Description
Signal extraction apparatus and program
[0001] The present invention relates to a signal extraction apparatus and a program.
[0002] It is known that when listening to an attention-target sound source (for example, the voice of a conversation partner, an announcement in a train, etc.) in a noisy environment, a user pays attention to the target that the user wants to hear, and extracts an attention-target sound source signal from an audio signal mixed with noise through a perceptual cognition process. As an extension of this auditory ability, application to an apparatus that estimates the direction of attention from brain waves and amplifies an audio signal in the direction to which attention is directed has been proposed.
[0003] Su, Enze, et al. "STAnet: A spatiotemporal attention network for decoding auditory spatial attention from EEG." IEEE Transactions on Biomedical Engineering 69.7 (2022): 2233-2242.
[0004] Conventionally, information obtained from brain waves tends to be limited to rough direction estimation such as "right or left". However, recent studies (for example, auditory cortical entrainment and the phenomenon in which EEG synchronizes with the envelope of an attended sound) suggest that brain waves may reflect not only the direction but also "frequency components of a specific sound source to which the user strongly directs attention".
[0005] For this reason, originally, simply "applying beamforming to the direction where attention is directed" cannot achieve sufficient separation when another sound source or noise existing in the same direction is mixed, leaving the problem that the target sound source cannot be completely extracted.
[0006] The present invention has been made in view of the above circumstances, and an object of the present invention is to provide a signal extraction apparatus and a program that can effectively block unnecessary sound sources and extract an attention-target sound source.
[0007] A signal extraction device according to a first aspect of the present invention comprises: an estimation unit that calculates electroencephalogram (EEG) features using EEG measurement data acquired from an external device and estimates attention direction information using the EEG features; a calculation unit that generates a mixed speech corresponding to the attention direction information using audio signal data acquired from an external device and calculates frequency features by performing a short-time Fourier transform on the mixed speech; and an output unit that generates a frequency mask using a mask generation model that takes the frequency features as input and outputs an enhanced audio signal using the frequency mask and frequency features.
[0008] A program according to a second aspect of the present invention is a program that causes a computer to function as a signal extraction device according to the first aspect.
[0009] According to the present invention, it is possible to provide a signal extraction device and program that can effectively block unwanted sound sources and extract sound sources of interest.
[0010] Figure 1 is a block diagram showing an example of the configuration of a signal extraction device according to an embodiment. Figure 2 is a block diagram showing an example of the functional configuration of a signal extraction device according to an embodiment. Figure 3 is a flowchart showing an example of an overview of the processing of a signal extraction device according to an embodiment. Figure 4 is a flowchart showing an example of the details of the processing of an information processing device according to an embodiment. Figure 5 is a flowchart showing an example of the details of the processing of an information processing device according to an embodiment. Figure 6 is a flowchart showing an example of the details of the processing of an information processing device according to an embodiment. Figure 7 is a block diagram showing an example of the hardware configuration of an information processing device according to an embodiment.
[0011] Each embodiment is described below with reference to the drawings. Each embodiment illustrates an apparatus or method for realizing the technical idea of the invention. The drawings are schematic or conceptual. Hereinafter, components having substantially the same function and configuration are denoted by the same reference numeral. The numbers following the letters that make up the reference numerals are used to distinguish elements that are referred to by reference numerals containing the same letters and have similar configurations. When it is not necessary to distinguish between elements indicated by reference numerals containing the same letters or numbers, these elements are referred to by reference numerals containing only letters or numbers.
[0012] The functional configuration of the signal extraction device 1 according to the embodiment will be described using Figures 1 and 2. Figure 1 is a block diagram showing an example of the configuration of the signal extraction device according to the embodiment. Figure 2 is a block diagram showing an example of the functional configuration of the signal extraction device according to the embodiment.
[0013] The signal extraction device 1 is an electronic device such as a computer, and may be, for example, a television receiver (including internet television), a PC (Personal Computer), a mobile terminal (e.g., a tablet, smartphone, laptop, feature phone, digital music player, e-book reader, smartwatch, etc.), a game console (home game console, portable game console), a VR (Virtual Reality) terminal, an AR (Augmented Reality) terminal, etc.
[0014] The signal extraction device 1 may be connected to an external server or the like via a network, for example. In this case, multiple signal extraction devices 1 can be connected to the external server via the network, enabling them to communicate with each other and share necessary information. The external server may be a storage device or the like.
[0015] As shown in Figure 2, the signal extraction device 1 is connected to external devices 7 and 8 so that data / information can be sent and received from each other. External device 7 includes an electroencephalogram measurement unit 71 and an audio signal measurement unit 72.
[0016] The electroencephalogram (EEG) measurement unit 71 preprocesses the raw EEG signal data to remove artifacts such as data errors and signal distortions from the raw data, and outputs an artifact-free EEG signal. More specifically, the EEG measurement unit 71 uses the raw EEG data of the C channel at time t, represented by Equation 1 below, as input. Here, we model the raw EEG data represented by equation 1 as shown in equation 2. Here, A is the mixture matrix and s(t) is the independent component vector. The independent component vector obtained by removing artifacts such as electromyography and blinking is taken as s'(t) and recomposed as shown in equation 3 below. Baseline correction and other processes are applied to the above equation 3, and a pre-processed EEG signal, represented by the following equation 4, is output after noise reduction, independent component analysis (ICA), artifact removal, etc. The electroencephalogram measurement unit 71 outputs the pre-processed EEG signal to the signal extraction device 1.
[0017] The audio signal measurement unit 72 acquires sound from the environment, for example, using a microphone array, and outputs the acquired sound to the signal extraction device 1. More specifically, the audio signal measurement unit 72 acquires multi-channel sound with a microphone array, takes the analog microphone signal as input, performs A / D conversion, samples it, and outputs an audio signal represented by the following equation 5. t is the sampling index (integer), and M is the number of microphone channels (integer). The audio signal measurement unit 72 outputs the audio signal to the signal extraction device 1.
[0018] The external device 8 includes an audio signal presentation unit 81. The audio signal presentation unit 81 receives the enhanced audio signal represented by the following equation 6, which has been enhanced by the signal extraction device 1. Details of the enhanced audio signal y(t) will be described later. The audio signal presentation unit 81 converts the enhanced audio signal y(t) into an analog signal using D / A conversion or the like. The audio signal presentation unit 81 has the function of presenting the converted audio to the user through a speaker, headphones, earphones, etc.
[0019] The signal extraction device 1 comprises a control unit 2, a data storage unit 3, a program storage unit 4, a communication unit 5, and an input / output unit 6.
[0020] The control unit 2 includes, as processing functions for implementing this embodiment, an electroencephalogram feature calculation unit 21, an attention direction estimation unit 22, an audio signal adjustment unit 23, a frequency feature calculation unit 24, a mask generation unit 25, and an enhanced signal generation unit 26. The control unit 2 comprehensively controls the data storage unit 3, the program storage unit 4, the communication unit 5, and the input / output unit 6.
[0021] The electroencephalogram (EEG) feature calculation unit 21 acquires pre-processed EEG signals from the EEG measurement unit 71 and calculates EEG features. Specifically, the EEG feature calculation unit 21 takes the pre-processed EEG signals represented by equation 4 as input and divides the pre-processed EEG signals into windows τ (width W samples).
[0022] The electroencephalogram feature calculation unit 21 calculates, for each channel c, the signal x after alpha band filtering, for example. α,c Calculate the power of (t) so that it is represented by the following equation 7. These values for each band and channel are combined to output the electroencephalogram (EEG) feature represented by the following equation 8. Note that τ is the window index (integer) when the electroencephalogram (EEG) is divided into time windows of a certain length, and the EEG features are vectors that include various band powers and phase indices.
[0023] The attention direction estimation unit 22 uses electroencephalogram (EEG) features to determine whether the user is directing their attention to the right or left. Note that the user's attention direction is not limited to right or left; it can be in various directions such as forward / backward or up / down. In this embodiment, the explanation will assume the user is directing their attention to the right or left.
[0024] The attention direction estimation unit 22, for example, in the case of a linear model, estimates attention direction information represented by the following equation 9 using the electroencephalogram features represented by equation 8. If d'(τ) is greater than 0, "Right" is estimated as the result, and if d'(τ) is less than 0, "Left" is estimated as the result. In other words, the attention direction estimation unit 22 outputs attention direction information d(τ) according to the value of d'(τ). Furthermore, the attention direction estimation unit 22 uses softmax to estimate the probability p to the right. R (τ) and the probability p on the left L A method for estimating (τ) may also be used. Note that the electroencephalogram feature calculation unit 21 and the attention direction estimation unit 22 are examples of estimation units.
[0025] The audio signal adjustment unit 23 acquires an audio signal from the audio signal measurement unit 72 and obtains the result from the attention direction estimation unit 22. Based on the result from the attention direction estimation unit 22, the audio signal adjustment unit 23 adjusts the audio signal in each direction using beamforming or the like. Specifically, the audio signal adjustment unit 23 prepares in advance a beamformer for the right direction represented by equation 10 below and a beamformer for the left direction represented by equation 11 below. If d'(τ) is greater than 0 and the attention direction information d(τ) is "Right", then the right-direction beamformer shown in equation 10 is applied. On the other hand, if d'(τ) is less than 0 and the attention direction information d(τ) is "Left", then the left-direction beamformer shown in equation 11 is applied.
[0026] Furthermore, if the attention direction information is probabilistic, the beamformer represented by the following equation 12 is applied. The audio signal adjustment unit 23 applies a beamformer according to the attention direction information, then synthesizes the audio using the process represented by the following equation 13, and outputs a single-channel mixed audio. Furthermore, if the attention direction information is in the form of a probability, the audio signal adjustment unit 23 synthesizes the audio using the process represented by equation 14 below and outputs a single-channel mixed audio. The frequency feature calculation unit 24 calculates the mixed audio signal x from the audio signal adjustment unit 23. audio A short-time Fourier transform (STFT) is performed on (t) to obtain frequency features. Mixed audio signal x audio The equation obtained by taking the short-time Fourier transform of (t) is shown in Equation 15 below. Also, the mixed audio signal X obtained by the short-time Fourier transform. audio The formula used to calculate the amplitude of (f, τ) is shown in Equation 16 below. Note that f is the frequency bin and τ is the index of the STFT frame. The frequency feature calculation unit 24 outputs the frequency features calculated in equations 15 and 16 to the mask generation unit 25. Note that the audio signal adjustment unit 23 and the frequency feature calculation unit 24 are examples of calculation units.
[0027] The mask generation unit 25 generates a frequency mask that blocks noise and allows the attention target sound source to pass through, based on the frequency features output by the frequency feature calculation unit 24. Alternatively, the mask generation unit 25 may generate the frequency mask using the brainwave features output by the brainwave feature calculation unit 21 and the frequency features.
[0028] More specifically, the mask generation unit 25 applies the frequency features as input to the mask generation model, as shown in equation 17 below. Furthermore, the mask generation unit 25 applies frequency features and electroencephalogram features as input to the mask generation model, as shown in equation 18 below. The mask generation unit 25 can output a frequency mask that reflects the user's level of attention by using the above equation 17 or equation 18.
[0029] The enhanced signal generation unit 26 applies the frequency mask output by the mask generation unit 25 to reconstruct the audio signal in the time domain, removing unwanted components. Specifically, the enhanced signal generation unit 26 performs spectral synthesis using the equation represented by the following equation 19. The enhanced signal generation unit 26 converts equation 19 back into the time domain by performing an inverse short-time Fourier transform and outputs the enhanced audio signal y(t). The enhanced signal generation unit 26 outputs the enhanced audio signal y(t) to the audio signal presentation unit 81. Note that the mask generation unit 25 and the enhanced signal generation unit 26 are examples of output units.
[0030] Furthermore, the audio signal presentation unit 81 may be configured to present emphasized audio to the ear in the direction of attention and unemphasized audio to the opposite ear, allowing the signal extraction device 1 to determine from the brainwaves whether the direction of attention is in the opposite direction. If the attention is in the opposite direction, adjustments may be made by switching the mask generation model in the mask generation unit 25 to a different model.
[0031] The data storage unit 3 stores data or information required for implementing the embodiment, and data or information generated in the process of executing various processes. The data storage unit 3 records data or information generated by processing of, for example, an electroencephalogram feature calculation unit 21, an attention direction estimation unit 22, an audio signal adjustment unit 23, a frequency feature calculation unit 24, a mask generation unit 25, and an enhanced signal generation unit 26.
[0032] The program storage unit 4 stores programs necessary for executing various controls and processes according to the embodiment. That is, the program storage unit 4 stores an information processing program, a control program, and other application programs for executing processing in the signal extraction device 1 according to the embodiment.
[0033] The communication unit 5 has a function of transmitting and receiving data, information, and information processing programs according to the embodiment to and from an external device or the like via a network.
[0034] The input / output unit 6 transmits data or information input by a user to the control unit 2. The input / output unit 6 also outputs data or information received from the control unit 2 to a user. The input / output unit 6 includes, for example, an input unit 61 and an output unit 62.
[0035] An outline of processing of the signal extraction device 1 according to the embodiment will be described with reference to FIG. 3. FIG. 3 is a flowchart showing an example of an outline of processing of the signal extraction device according to the embodiment.
[0036] As shown in FIG. 3, the processing in the signal extraction device 1 according to the present embodiment is divided into five steps, which are, for example, a data collection step S1, an electroencephalogram analysis step S2, an audio signal processing step S3, an attention target sound source step S4, and an audio signal presentation step S5. Note that the content of processing in the following operation description is an example, and various processes capable of obtaining similar results can be appropriately used.
[0037] Data processing step S1 simultaneously acquires electroencephalogram (EEG) signals from the external device 7 and audio signals from the microphone array. Furthermore, data processing step S1 synchronizes the EEG measurement data and audio measurement data chronologically and records the user's attention direction, EEG features, and audio signal frequency features for use in subsequent processing.
[0038] The electroencephalogram (EEG) analysis step S2 estimates the direction of attention from the acquired EEG data and further calculates EEG features (such as band power and phase synchronization). The processing in EEG analysis step S2 makes it possible to estimate when and in which direction the user is paying attention, or in what frequency band sound sources the user is responding to most strongly.
[0039] In the audio signal processing step S3, beamforming and other techniques are applied to the audio signal to synthesize the audio corresponding to the general direction as a single channel, and then frequency analysis (STFT, etc.) is performed. In the audio signal processing step S3, the frequency characteristics of the audio obtained here are linked to the electroencephalogram (EEG) characteristics and managed.
[0040] The attention-focused sound source step S4 incorporates frequency attention information contained in brainwaves into the mask generation model, even when multiple sound sources are present in the same direction. Specifically, the attention-focused sound source step S4 uses a mask generation model that takes brainwave features and audio signal frequency features as input to estimate a frequency mask that allows only the sound source components that the user is truly paying attention to to pass through. The attention-focused sound source step S4 blocks out unwanted sound sources and noise in the same direction.
[0041] In the audio signal presentation step S5, the audio signal, which has been enhanced by applying the frequency mask estimated in the attention target sound source step S4, is returned to the time domain and presented to the user. As a result of the audio signal presentation step S5, the user can clearly hear the audio (i.e., the attention target sound source) that has been enhanced to reflect the attention information derived from brainwaves.
[0042] Next, a more detailed procedure relating to the processing of the signal extraction device 1 according to the embodiment will be described using Figures 4, 5, and 6. Figure 4 is a flowchart showing an example of the detailed processing of the information processing device according to the embodiment. Figure 5 is a flowchart showing an example of the detailed processing of the information processing device according to the embodiment. Figure 6 is a flowchart showing an example of the detailed processing of the information processing device according to the embodiment. Note that the processing content in the following operation description is just an example, and various processes that can obtain similar results can be used as appropriate.
[0043] For example, the process in Figure 4 is initiated when the signal extraction device 1 outputs a signal to the external device 7 requesting electroencephalogram (EEG) measurement data and audio signal measurement data. Alternatively, the process in Figure 4 may be initiated when the user of the signal extraction device 1 requests an arbitrary start command via the input unit 61, and may be executed at arbitrary timings and periods.
[0044] The electroencephalogram feature calculation unit 21 acquires electroencephalogram measurement data via the electroencephalogram measurement unit 71 of the external device 7 (step S10). For example, the electroencephalogram measurement data includes at least the pre-processed EEG signal and the date and time on which the electroencephalogram measurement unit 71 acquired the raw EEG data.
[0045] The audio signal adjustment unit 23 acquires audio signal measurement data via the audio signal measurement unit 72 of the external device 7 (step S11). For example, the audio signal measurement data includes at least the audio signal and the date and time on which the audio signal measurement unit 72 acquired the audio signal with the microphone array.
[0046] The control unit 2 synchronizes the electroencephalogram (EEG) measurement data acquired by the EEG feature calculation unit 21 and the audio signal data acquired by the audio signal adjustment unit 23 in a time series (step S12). That is, the control unit 2 uses the date and time included in the EEG measurement data and the date and time included in the audio signal data to associate the EEG measurement data and audio signal data at the same time on the same date and store them in the data storage unit 3.
[0047] The electroencephalogram (EEG) feature calculation unit 21 calculates EEG features using the EEG measurement data (step S13). The EEG feature calculation unit 21 stores the calculated EEG measurement data in the data storage unit 3 (step S14).
[0048] The attention direction estimation unit 22 estimates the user's attention direction information using the electroencephalogram (EEG) features calculated by the EEG feature calculation unit 21 (step S15). The attention direction estimation unit 22 stores the estimated user's attention direction information in the data storage unit 3 (step S16).
[0049] The audio signal adjustment unit 23 determines whether the user's attention direction is to the right by acquiring attention direction information (step S17). The audio signal adjustment unit 23 may also be configured to determine whether the user's attention direction is to the left.
[0050] If the audio signal adjustment unit 23 determines that the user's attention direction is to the right, it applies a beamformer for the right direction to the audio signal measurement data (step S18). On the other hand, if the audio signal adjustment unit 23 determines that the user's attention direction is not to the right, i.e., that the user's attention direction is to the left, it applies a beamformer for the left direction to the audio signal measurement data (step S19).
[0051] By performing the processing in step S18 or step S19, the audio signal adjustment unit 23 generates a single-channel mixed audio (step S20). The frequency feature calculation unit 24 performs a short-time Fourier transform on the mixed audio generated by the audio signal adjustment unit 23 (step S21).
[0052] The frequency feature calculation unit 24 calculates the frequency features of the audio signal by performing a short-time Fourier transform on the mixed audio and calculating the amplitude (step S22). The frequency feature calculation unit 24 stores the frequency features in the data storage unit 3.
[0053] The control unit 2 associates frequency features and electroencephalogram (EEG) features for an arbitrary time period and stores them in the data storage unit 3 (step S23). The mask generation unit 25 incorporates the associated frequency features and EEG features into the mask generation model (step S24). The mask generation unit 25 may also incorporate only the frequency features into the mask generation model.
[0054] The mask generation unit 25 outputs a frequency mask based on the frequency features and electroencephalogram features incorporated into the mask generation model (step S25). The enhanced signal generation unit 26 applies the frequency mask output by the mask generation unit 25 to the frequency features and performs spectral synthesis (step S26).
[0055] The enhanced signal generation unit 26 reconstructs the spectrally synthesized audio signal into the time domain by performing an inverse short-time Fourier transform and outputs an enhanced audio signal (step S27). The enhanced signal generation unit 26 outputs the enhanced audio signal to the user via the external device 8 (step S28) and terminates the process.
[0056] According to the signal extraction device 1 of this embodiment, by utilizing frequency attention features derived from brainwaves (such as EEG patterns in which the user is strongly synchronized to a specific speaker or frequency band) in mask generation, it becomes possible to generate a frequency mask that selectively passes only sound source components corresponding to strong attention shown by the brainwaves, even within the same direction. Brainwave features are input to the mask generation model along with frequency features, and based on the correlation between the two, the frequency components of the attention target sound source are emphasized, and unwanted sound source components with low attentional responses on the EEG are suppressed. In other words, by directly incorporating brainwave information (including frequency attention features as well as direction) into speech signal processing, it becomes possible to provide a technology that effectively blocks unwanted sound sources within the same direction that cannot be completely removed by simple direction-based beamforming, and extracts the attention target sound source. As a result, even in noisy environments or multi-speaker environments, the sounds that the user truly wants to hear can be emphasized with higher accuracy.
[0057] Next, an example of the hardware configuration of the signal extraction device 1 according to the embodiment will be described with reference to Figure 7. Figure 7 is a block diagram showing an example of the hardware configuration of the signal extraction device according to the embodiment.
[0058] As shown in Figure 7, the signal extraction device 1 includes, for example, a processor 27, a ROM (Read Only Memory) 28, a RAM (Random Access Memory) 29, a storage 30, a communication interface 31, and an input / output interface 32.
[0059] The control unit 2, which includes the electroencephalogram feature calculation unit 21, attention direction estimation unit 22, audio signal adjustment unit 23, frequency feature calculation unit 24, mask generation unit 25, and enhanced signal generation unit 26 shown in Figure 1, corresponds to, for example, the processor 27 and RAM 29. The data storage unit 3 corresponds to, for example, the ROM 28, RAM 29, and storage 30. The program storage unit 4 corresponds to, for example, the ROM 28 and storage 30. The communication unit 5 corresponds to, for example, the communication interface 31. The input / output unit 6 corresponds to, for example, the input / output interface 32 and input / output device 33.
[0060] The processor 27, ROM 28, RAM 29, storage 30, communication interface 31, and input / output interface 32 are each connected via a bus communication line (BUS) and can send and receive data from each other. Alternatively, the processor 27, ROM 28, RAM 29, storage 30, communication interface 31, and input / output interface 32 may be configured to send and receive data from each other via wireless communication or the like.
[0061] The processor 27 includes at least one processor such as a CPU (Central Process Unit), MPU (microprocessing unit), GPU (Graphics Processing Unit), or FPGA (field-programmable gate array), and can realize various functions of the signal extraction device 1 by executing programs such as system software, application software, or firmware stored in the ROM 28, RAM 29, or storage 30.
[0062] ROM 28 is a read-only non-volatile memory. ROM 28 non-temporarily stores the startup program required when the signal extraction device 1 of this embodiment is started. The signal extraction device 1 is started when the processor 27 executes the program in ROM 28. ROM 28 is, for example, an EPROM (Erasable Programmable Read Only Memory) and stores various startup settings in addition to the startup program.
[0063] RAM 29 is a volatile memory that can be written to and read from. RAM 29 temporarily stores programs necessary for processing by the processor 27 and data necessary for executing those programs. The processor 27, for example, executes a program in RAM 29 to perform calculations on the data in RAM 29 and stores the calculation results in RAM 29.
[0064] The storage 30 is composed of non-volatile memory such as an HDD (Hard Disk Drive) or SSD (Solid State Drive). The storage 30 non-temporarily stores programs executed by the processor 27 and data necessary for program execution. The processor 27 reads the programs and data from the storage 30 into the RAM 29 and executes various functions by running the programs.
[0065] The communication interface 31 is connected to a communication network and enables the reception of information from and transmission of information to external server devices.
[0066] The input / output interface 32 is connected to an input device 34 and an output device 35, respectively. The input / output interface 32 enables the input of information from the input device 34 and the output of information to the output device 35. The input device 34 includes, for example, a keyboard, mouse, touch panel, and disk drive. The input device 34 is not limited to these and may include any other input device. The output device 35 includes, for example, a display and disk drive. The output device 35 is not limited to these and may include audio output means such as a speaker or any other output device. The input device 34 and the output device 35 may be configured as an input / output device 33 that has the functions of both.
[0067] The program for operating the signal extraction device 1 according to the embodiment is provided to the computer, for example, via a computer-readable storage medium 36. This storage medium 36 is a non-temporary computer-readable storage medium. Non-temporary computer-readable storage media include, for example, disks such as flexible disks, optical disks (CD-ROM, CD-R, DVD-ROM, DVD-R, etc.), magneto-optical disks (MO, etc.), semiconductor memory, USB memory, etc.
[0068] Furthermore, the above program may be stored on a server device on a communication network, downloaded from the server device, and temporarily stored in storage 30.
[0069] For example, in response to an input signal instructing the activation of the signal extraction device 1, the processor 27 reads a program from the storage 30 into the program area of the RAM 29, and also reads the data necessary for executing the program from the storage 30 into the data area of the RAM 29. The processor 27 performs calculations on the data in the data area according to the program and writes the calculation results to the data area. Through these operations, the processor 27, RAM 29, storage 30, communication interface 31, and input / output interface 32 cooperate to perform at least some of the functions of the components of the signal extraction device 1, namely the electroencephalogram feature calculation unit 21, the attention direction estimation unit 22, the speech signal adjustment unit 23, the frequency feature calculation unit 24, the mask generation unit 25, and the enhanced signal generation unit 26.
[0070] It should be noted that the present invention is not limited to the embodiments described above, and can be modified in various ways during implementation without departing from its essence. Furthermore, each embodiment may be combined as appropriate, and in that case, the combined effects can be obtained. Moreover, the above embodiments include various inventions, and various inventions can be extracted by selecting combinations from the multiple constituent elements disclosed. For example, if the problem can be solved and effects obtained even if some constituent elements are deleted from all the constituent elements shown in the embodiment, then the configuration with these deleted constituent elements can be extracted as an invention.
[0071] 1...Signal extraction device 2...Control unit 21...Electroencephalogram feature calculation unit 22...Attention direction estimation unit 23...Audio signal adjustment unit 24...Frequency feature calculation unit 25...Mask generation unit 26...Enhanced signal generation unit 3...Data storage unit 4...Program storage unit 5...Communication unit 6...Input / output unit 61...Input unit 62...Output unit 7...External device 71...Electroencephalogram measurement unit 72...Audio signal measurement unit 8...External device 81...Audio signal presentation unit 27...Processor 28...ROM 29...RAM 30...Storage 31...Communication interface 32...Input / output interface 33...Input / output device 34...Input device 35...Output device 36...Storage medium
Claims
1. A signal extraction device comprising: an estimation unit that calculates electroencephalogram (EEG) features using EEG measurement data acquired from an external device and estimates attention direction information using the EEG features; a calculation unit that generates a mixed voice corresponding to the attention direction information using audio signal data acquired from an external device and calculates frequency features by performing a short-time Fourier transform on the mixed voice; and an output unit that generates a frequency mask using a mask generation model that takes the frequency features as input and outputs an enhanced audio signal using the frequency mask and the frequency features.
2. The signal extraction device according to claim 1, wherein the output unit generates the frequency mask using a mask generation model that takes the frequency features and the electroencephalogram features as inputs.
3. The signal extraction device according to claim 1, wherein the estimation unit estimates the attention direction information according to the numerical values of a linear model using the electroencephalogram features.
4. A program that causes a computer to function as a signal extraction device according to any one of claims 1 to 3.