Signal processing device, signal processing method, and signal processing program
The signal processing device uses DNNs to estimate correction filters based on user-specific acoustic characteristics, addressing the challenge of varying head and wearing conditions, thereby improving noise cancellation by up to 15 dB and enhancing usability.
Patent Information
- Application Number
- JP2022530116
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-11
- Filing Date
- 2021-05-26
- Publication Date
- 2025-08-20
- Estimated Expiration
- 2041-05-26
AI Technical Summary
Conventional noise canceling technologies face challenges in achieving optimal noise reduction due to variations in user-specific head and wearing conditions, making it difficult to place microphones at the eardrum position for accurate noise cancellation.
A signal processing device that uses machine learning, specifically Deep Neural Networks (DNNs), to estimate correction filters based on user-specific acoustic characteristics without requiring a microphone at the eardrum, by focusing on device characteristics measured with microphones inside the headphones, thereby optimizing noise cancellation.
The solution significantly improves noise cancellation effectiveness by up to 15 dB compared to default settings, adapting to individual user conditions and environments, enhancing usability and noise reduction performance.
Smart Images

Figure 0007726208000002 
Figure 0007726208000003 
Figure 0007726208000004
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a signal processing device, a signal processing method, a signal processing program, a signal processing model production method, and an acoustic output device. [Background technology]
[0002] In recent years, with the spread of portable audio players, noise reduction systems that provide listeners (users) with a good playback sound field space by reducing noise from the external environment have become popular for audio output devices for portable audio players (e.g., headphones, earphones, etc.).
[0003] In relation to the above technology, a technology that uses a noise canceling (NC) filter to suppress noise at the position of the user's eardrum has become widespread. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-015585 Summary of the Invention [Problem to be solved by the invention]
[0005] However, conventional technology leaves room for further improvements in usability. For example, conventional technology sometimes requires a signal at the eardrum position to maximize the NC effect at the eardrum position, but product specifications make it difficult to place a microphone at the eardrum position.
[0006] Therefore, the present disclosure proposes a new and improved signal processing device, signal processing method, signal processing program, signal processing model production method, and audio output device that can promote further improvements in usability. [Means for solving the problem]
[0007] According to the present disclosure, there is provided a signal processing device including an acquisition unit that acquires acoustic characteristics inside a user's ear that are isolated from the outside world, an NC filter unit that generates sound data that is in antiphase with environmental sound that has leaked into the user's ear, a correction unit that corrects the sound data using a correction filter, and a determination unit that determines filter coefficients of the correction filter based on the acoustic characteristics. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of a configuration for NC optimization according to an embodiment. [Figure 2] FIG. 1 is a diagram illustrating an overview of a function related to NC filter determination according to an embodiment. [Figure 3] 1A to 1C are diagrams illustrating an example of the configuration of an NC filter according to an embodiment when it is designed and used [Figure 4] FIG. 1 is a diagram illustrating an overview of functions for NC optimization during use according to an embodiment. [Figure 5] FIG. 1 is a diagram illustrating an overview of functions for NC optimization during use according to an embodiment. [Figure 6] FIG. 10 is a diagram illustrating an example of an HM characteristic according to the embodiment. [Figure 7A] FIG. 10 is a diagram showing an example of a simulation result of the NC effect according to the embodiment. [Figure 7B] FIG. 10 is a diagram showing an example of a simulation result of the NC effect according to the embodiment. [Figure 8A] FIG. 10 is a diagram showing an example of a simulation result of the NC effect according to the embodiment. [Figure 8B] FIG. 10 is a diagram showing an example of a simulation result of the NC effect according to the embodiment. [Figure 9] 1 is a diagram illustrating an example of the configuration of a signal processing system according to an embodiment. [Figure 10] FIG. 1 is a diagram illustrating an overview of functions for NC optimization according to an embodiment. [Figure 11] FIG. 10 is a diagram illustrating an example of an estimation result of a second DNN according to the embodiment. [Figure 12] FIG. 1 is a diagram illustrating an overview of functions of a signal processing system according to an embodiment. [Figure 13] FIG. 1 is a diagram illustrating an overview of functions of a signal processing system according to an embodiment. [Figure 14] 3 is a flowchart showing a processing flow of the signal processing system according to the embodiment. [Figure 15] 3 is a flowchart showing a processing flow of the signal processing system according to the embodiment. [Figure 16] 3 is a flowchart showing a processing flow of the signal processing system according to the embodiment. [Figure 17] FIG. 10 is a diagram illustrating an overview of a function for storing and referencing a correction filter according to an embodiment. [Figure 18A] FIG. 10 is a diagram showing a flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 18B] FIG. 10 is a diagram showing a flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 18C] FIG. 10 is a diagram showing a flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 19] FIG. 10 is a diagram showing a flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 20] 10 is a flowchart showing the flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 21A] FIG. 10 is a diagram showing a flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 21B] FIG. 10 is a diagram showing a flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 21C] FIG. 10 is a diagram showing a flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 22A] FIG. 10 is a diagram showing a flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 22B] FIG. 10 is a diagram showing a flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 22C]FIG. 10 is a diagram showing a flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 23A] FIG. 10 is a diagram showing a flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 23B] FIG. 10 is a diagram showing a flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 23C] FIG. 10 is a diagram showing a flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 23D] FIG. 10 is a diagram showing a flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 24] 10 is a flowchart showing the flow of a process for storing and referencing a correction filter according to the embodiment. [Figure 25] FIG. 1 is a block diagram of a signal processing system according to an embodiment. [Figure 26] FIG. 2 illustrates an example of a storage unit according to the embodiment. [Figure 27] 3 is a flowchart showing a processing flow in the signal processing device according to the embodiment. [Figure 28] FIG. 10 is a diagram illustrating an example of a display screen that displays a list of correction filters according to the embodiment. [Figure 29] FIG. 10 is a diagram illustrating an example of a display screen that displays a list of correction filters according to the embodiment. [Figure 30] FIG. 10 is a diagram illustrating an overview of a function for updating a correction filter according to an embodiment. [Figure 31] FIG. 10 is a diagram illustrating an overview of a function when adjusting the gain of a correction filter according to an embodiment. [Figure 32] FIG. 10 is a diagram illustrating an overview of a function when adjusting the gain of a correction filter according to an embodiment. [Figure 33] 10 is a flowchart showing a processing flow when adjusting the gain of a correction filter according to the embodiment. [Figure 34] FIG. 1 is a diagram illustrating an example of a hardware configuration of a signal processing device according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0010] The explanation will be given in the following order. 1. One embodiment of the present disclosure Introduction 1.2.NC Personal Optimization 1.3.Signal Processing System Configuration 2. Functions of the signal processing system 2.1. First DNN 2.2. Second DNN 2.3. The Third DNN 2.4. Correction filter estimation process 2.5. The Fourth DNN 2.6.Processing Flow 2.7.Memorizing and referencing correction filters 2.8. The Fifth DNN 2.9. The Sixth DNN 2.10. Functional configuration example 2.11. Signal Processing System 2.12. Processing Variations 3. Hardware configuration example 4. Summary
[0011] <<1. One embodiment of the present disclosure>> <1.1. Introduction> The volume and air density inside headphones, etc. may vary depending on the user's physical characteristics, such as head shape and ear size, and external factors, such as whether or not the user is wearing glasses or a hat. Therefore, the characteristics of the signal when the sound from the signal after the noise reduction signal has been applied reaches the user's ear may vary depending on the volume and air density inside the headphones, etc., and therefore may vary depending on the user. The characteristics of the signal when the sound from the signal after the noise reduction signal has been applied reaches the user's ear may also vary depending on how the headphones, etc. are worn.
[0012] The standard (default) NC filter (hereinafter referred to as "α default") installed in the product may be determined based on the standard head shape and wearing condition at the time of design. Therefore, when used by the user, there may be differences in the head shape and wearing condition compared to the default, and the optimal NC effect may not be achieved. For this reason, there was room for further improvement in usability.
[0013] Therefore, the present disclosure proposes a new and improved signal processing device, signal processing method, and signal processing model creation method that can promote further improvements in usability.
[0014] <1.2.NC Personal Optimization> First, NC personal optimization will be described. FIG. 1 is a diagram showing an example of a configuration for NC personal optimization. Microphone MI11 indicates a microphone for FF (Feed Forward) NC (hereinafter referred to as the "first microphone") arranged inside headphones HP11. Microphone MI12 indicates a microphone for FB (Feed Back) NC (hereinafter referred to as the "second microphone") arranged inside headphones HP11. Microphone MI13 indicates a microphone arranged at the eardrum position (hereinafter referred to as the "third microphone"). Acoustic characteristics F0 indicate the acoustic characteristics (spatial acoustic characteristics) from noise source N to the first microphone. Acoustic characteristics F1 indicate the acoustic characteristics from the first microphone to the third microphone. Note that acoustic characteristics F1 are leakage characteristics that do not pass through the space inside headphones HP11. Device characteristics H1 indicate the acoustic characteristics from the driver (speaker) of headphones HP11 to the third microphone. Device characteristics H2 indicate the acoustic characteristics from the driver of headphones HP11 to the second microphone. Microphone characteristics M1 indicate the microphone characteristics of the first microphone. The microphone characteristic M2 indicates the microphone characteristic of the second microphone, and the microphone characteristic M3 indicates the microphone characteristic of the third microphone.
[0015] Next, we will explain the overview of the function for personal NC optimization. In Figure 2, the NC filter that maximizes the NC effect amount for the standard head and wearing condition at the time of design is determined. This NC filter is the α default that is installed in the product. In Figure 2, the α default is determined based on the device characteristics H1 and acoustic characteristics F1 at the time of design. The following formula (1) shows the calculation formula for determining the α default.
[0016]
number
[0017] The device characteristic H1 and the acoustic characteristic F1 may differ between users. For this reason, focusing on the device characteristic H1, personal optimization can be performed by correcting the H1 default M1 (hereinafter referred to as the "H1M1 characteristic" as appropriate) included in the above formula (1) between users. However, in this case, it is necessary to place a microphone near the eardrum, making it difficult to measure the device characteristic H1 in the user's usage environment. Therefore, in this embodiment, for example, focusing on the device characteristic H2, the device characteristic H1 is estimated based on the similarity between the device characteristic H1 and the device characteristic H2.
[0018] 3(A) and 3(B) are diagrams showing configuration examples at the time of design and at the time of use. Device characteristics H2 default indicates the device characteristics H2 at the time of design. Device characteristics H2 user indicates the device characteristics H2 when personal optimization is performed.
[0019] Next, an overview of the function for personal NC optimization during use will be described using Figures 4 and 5. Note that descriptions similar to those in Figure 2 will be omitted as appropriate. Also, Figure 2 illustrates a case where the device characteristic H1 default is used as the device characteristic H1, while Figures 4 and 5 illustrate a case where the device characteristic H1 user is used. In Figure 4, the standard α default is used for the product's NC filter. However, acoustic characteristics may change depending on the device characteristic H1 user, based on the user's wearing state, etc. Therefore, Figure 5 corrects the acoustic characteristics that may change in Figure 4. For example, Figure 5 focuses on the device characteristic H2 user, and performs correction using a correction filter that cancels the difference between the device characteristic H2 user and the device characteristic H2 default. Note that, for convenience of explanation, Figure 5 illustrates a case where correction is performed immediately after application of the device characteristic H1 user. However, correction may be performed before or after application of the α default, or the α default itself may be corrected. In addition, in actual products, correction may often be performed by narrowing the frequency band to approximately 100 Hz or less to avoid adverse effects.
[0020] Next, the HM characteristics included in the above formula (1) will be explained using Figure 6. Figure 6(A) shows the H1M characteristics measured with a microphone placed at the eardrum position. Figure 6(B) shows the H2M characteristics measured with a microphone for FBNC. Each of Figures 6(A) and 6(B) contains data on HM characteristics measured approximately 440 times while changing the wearing state. Note that all the data shown in Figures 6(A) and 6(B) was measured using a dummy head, so there is no difference due to the shape of the head. The horizontal axis represents frequency (Hz) and the vertical axis represents sound pressure (dB).
[0021] As mentioned above, it is difficult to measure the H1M characteristic data shown in FIG. 6A in a user's operating environment. If the H1M characteristic were measurable, the optimal correction filter coefficient α could be determined by calculation rather than estimation. Furthermore, the correction filter coefficient α is determined based on the H1M characteristic, and cannot be determined based on the H2M characteristic. Therefore, as mentioned above, the α default is corrected to offset the difference based on the H2M characteristic, focusing on the device characteristic H2 user. However, as shown in FIGS. 6A and 6B, the H1M characteristic and the H2M characteristic may differ significantly above approximately 200 Hz. Examples of factors that may cause significant differences in the HM characteristic include the shape of the user's ear canal, ear hair, and the temperature and humidity of the room, but there may also be various other factors. For this reason, it is desirable to perform correction within a frequency band (e.g., approximately 100 Hz) where the H1M characteristic and the H2M characteristic tend to be similar. Specifically, in bands with similar tendencies, appropriate correction was possible by substituting H2M characteristics. However, there were cases where appropriate correction was not possible because similarity could not be guaranteed due to individual differences in head shape and wearing conditions between users.
[0022] Next, a simulation of the NC effect will be explained using Figure 7. Figure 7A shows an example of the simulation results measured with a microphone placed at the eardrum position. Figure 7A includes five graphs. Of these, graph LA1 shows the simulation results for an unexposed state in which the user is not wearing headphones or the like. Graph LA2 shows the simulation results for a case in which the user is wearing headphones or the like and no NC is performed. Graph LA3 shows the simulation results for a case in which NC is performed with the α default. Graph LA4 shows the simulation results for a case in which NC is performed with an optimal NC filter that maximizes the amount of NC effect. Graph LA5 shows the simulation results for a case in which NC is performed with an NC filter (corrected filter) corrected with a correction filter estimated by machine learning. The indicators on the vertical and horizontal axes are the same as those in Figure 6.
[0023] In FIG. 7A, the lower the sound pressure on the vertical axis, the higher the NC effect. Note that the NC effect here also includes the effect of sound insulation. Furthermore, comparing graphs LA3 and LA4, it can be seen that there can be a difference of approximately 15 dB in the bands where the difference is large. Graphs LA3 to LA5 show that applying a correction filter to the α default built into the product can approach the optimal NC filter. The closer graph LA5 is to graph LA4, the closer the NC filter corrected with the correction filter estimated by machine learning will be to the optimal NC filter, thereby improving the NC effect. Furthermore, FIG. 7B shows the frequency characteristics (gain) of the NC filters corresponding to graphs LA3 to LA5 in FIG. 7A.
[0024] FIG. 8 shows an example of a simulation result when the user in FIG. 7 changes the wearing condition by putting on or taking off headphones, etc. Note that the graphs included in FIG. 8 are the same as those in FIG. 7, and therefore their explanations are omitted. Comparing FIG. 7 and FIG. 8, it can be seen that errors in the wearing condition have a significant impact on the NC effect and the characteristics of the NC filter. For example, below 200 Hz, the difference between graph LA4 and graph LA5 is larger in FIG. 8 than in FIG. 7. For example, graph LA3 decreases sharply from around 350 Hz in FIG. 7, whereas it decreases gradually from around 200 Hz in FIG. 8.
[0025] In the following embodiments, a case will be described in which a correction filter is estimated using machine learning such as a DNN (Deep Neural Network). By using machine learning such as a DNN, a correction filter can be appropriately estimated according to the shape of the user's head, the wearing state, external environmental sounds, and the like, without any bandwidth limitations. This allows the signal processing device 10 to achieve NC optimization with a wider bandwidth and with a higher degree of freedom. Note that the DNN used in the embodiments is an example of artificial intelligence.
[0026] In the following embodiments, a DNN (hereinafter referred to as a "correction filter coefficient estimation DNN" or "first DNN" as appropriate) that receives H2M characteristics measured by an FBNC microphone as input and outputs a correction filter coefficient (correction filter coefficient) for optimally correcting a noise canceling signal generated based on measurement data measured by an FFNC microphone will be described. Note that the first DNN is not limited to correcting the noise canceling signal, and may output a correction filter coefficient for optimally correcting a filter that generates a noise canceling signal based on measurement data measured by an FFNC microphone. Also, a DNN (hereinafter referred to as a "correction determination DNN" or "second DNN" as appropriate) that determines whether correction is necessary when the NC effect amount is sufficient when optimization is performed or when the NC effect amount is insufficient even after correction due to large leakage will be described.
[0027] Hereinafter, the correction filter according to the embodiment may be, for example, an FIR (Finite Impulse Response) filter having a finite impulse response.
[0028] Hereinafter, the corrected filter according to the embodiment may be, for example, a filter obtained by applying a correction filter at the time of use or the like to the α default.
[0029] In the following embodiment, a case where the NC effect amount is estimated in an environment set in accordance with the JEITA standard is shown, but the NC effect amount may be estimated in an environment set in accordance with other standards, not limited to the JEITA standard. The signal processing device 10 can estimate the optimization effect by estimating the NC effect amount, and therefore can determine whether or not to perform optimization.
[0030] In the following, the embodiment will be described using headphones 20 as an example of an audio output device.
[0031] <1.3. Signal Processing System Configuration> The configuration of a signal processing system 1 according to an embodiment will be described. FIG. 9 is a diagram illustrating an example of the configuration of the signal processing system 1. As illustrated in FIG. 9, the signal processing system 1 includes a signal processing device 10 and headphones 20. Various devices can be connected to the signal processing device 10. For example, headphones 20 are connected to the signal processing device 10, and information is shared between the devices. The signal processing device 10 and the headphones 20 are connected to an information and communication network via wireless or wired communication so that they can communicate information and data with each other and operate in cooperation. The information and communication network can be configured using the Internet, a home network, an IoT (Internet of Things) network, a P2P (Peer-to-Peer) network, a proximity communication mesh network, or the like. For wireless communication, for example, Wi-Fi, Bluetooth (registered trademark), or a technology based on a mobile communication standard such as 4G or 5G can be used. For wired communication, power line communication technology such as Ethernet (registered trademark) or PLC (Power Line Communications) can be used.
[0032] The signal processing device 10 and the headphones 20 may be provided separately as multiple computer hardware devices on-premise, an edge server, or the cloud, or the functions of any multiple devices among the signal processing device 10 and the headphones 20 may be provided as a single device. For example, the signal processing device 10 and the headphones 20 may function as an integrated device that communicates with an external information processing device. Furthermore, a user can communicate information and data with the signal processing device 10 and the headphones 20 via a user interface (including a Graphical User Interface: GUI) and software (composed of a computer program (hereinafter also referred to as a program)) running on a terminal device (not shown) (a personal device such as a PC (Personal Computer) or a smartphone that includes a display as an information display device, voice input, and keyboard input).
[0033] (1) Signal Processing Device 10 The signal processing device 10 is an information processing device that performs processing to determine coefficients of a correction filter (filter coefficients) for optimal NC for an individual user. Specifically, the signal processing device 10 acquires acoustic characteristics in the user's ear, which is isolated from the outside world. The signal processing device 10 then generates sound data that is out of phase with the environmental sound that has leaked into the user's ear, and corrects the sound using a correction filter. The signal processing device 10 also determines the correction filter coefficients based on the acoustic characteristics. This allows the signal processing device 10 to estimate correction filter coefficients for optimization without requiring a signal at the eardrum position. Furthermore, the signal processing device 10 can perform optimization processing without relying on the experience or discretion of a designer. This leaves room for the signal processing device 10 to further improve usability.
[0034] The signal processing device 10 also has a function of controlling the overall operation of the signal processing system 1. For example, the signal processing device 10 controls the overall operation of the signal processing system 1 based on information shared between the devices. Specifically, the signal processing device 10 determines correction filter coefficients for optimization based on information received from the headphones 20.
[0035] The signal processing device 10 is realized by a PC (Personal Computer), a server, etc. Note that the signal processing device 10 is not limited to a PC, a server, etc. For example, the signal processing device 10 may be a computer hardware device such as a PC or a server in which the functions of the signal processing device 10 are implemented as an application.
[0036] (2) Headphones 20 The headphones 20 are headphones used by a user to listen to sound. The headphones 20 are not limited to headphones and may be any type of audio output device that has a driver and a microphone and can separate the space including the user's eardrum from the outside world. For example, the headphones 20 may be earphones.
[0037] The headphones 20, for example, collect the measurement sound output from the driver with a microphone.
[0038] <<2. Functions of the signal processing system>> This concludes the description of the configuration of the signal processing system 1. Next, we will describe the functions of the signal processing system 1. The functions of the signal processing system 1 include a function of estimating a correction filter coefficient for correcting the α default in order to perform optimal NC for each individual user, and a function of determining whether to perform optimal NC correction for each individual user.
[0039] FIG. 10 is a diagram showing an overview of a function for performing optimal NC for an individual user. The signal processing system 1 measures acoustic characteristics (H2 user M2 characteristics) based on a signal collected by the second microphone. Then, the signal processing system 1 estimates correction filter coefficients using a first DNN that estimates correction filter coefficients based on the measured acoustic characteristics. The signal processing system 1 also estimates the NC effect of the α default based on the measured H2 user M2 characteristics, and applies a correction filter if the correction effect is expected to be sufficient using a second DNN that determines whether the correction effect is expected to be sufficient. The first DNN and the second DNN will be described below.
[0040] <2.1. First DNN> The first DNN receives as input the H2 user M2 characteristics based on the signal picked up by the second microphone and outputs a correction filter coefficient. The first DNN performs optimization using Adam as an example of an optimization method. The first DNN uses a correction filter coefficient based on the H1 user M3 as training data. Here, the first DNN may use, for example, a gradient method to obtain a correction filter coefficient that minimizes the NC simulation results as training data. The first DNN outputs this correction filter coefficient and inputs the H2 user M2 characteristics as training data. The first DNN may use a loss function to convert both the training data and the estimated data into frequency characteristics using FFT (Fast Fourier Transform), and then use a common low-pass filter to calculate the average (mean value) from the sum of the absolute values of the differences in each band.
[0041] <2.2. Second DNN> The second DNN receives as input acoustic characteristics (e.g., impulse response time signal and FFT frequency signal) based on the signal picked up by the second microphone and the corrected filter coefficients, and outputs whether or not to perform correction. The second DNN performs optimization using Adam as an example of an optimization method. The second DNN uses a loss function based on cross-entropy. The second DNN performs an NC simulation using the H2 user M2 characteristics, microphone characteristics M1, microphone characteristics M3, and the corrected filter coefficients. The second DNN then uses as training data labeled with whether or not to perform correction based on whether the NC effect amount, which is the correction effect obtained as a result of the simulation, is equal to or exceeds a predetermined threshold. Here, the NC effect amount refers to the amount of suppression when comparing the sound pressure at the eardrum position under a specified noise source and in a noise environment between a state in which headphones 20 are not worn and a state in which NC is enabled. For example, the signal processing system 1 may perform 1 / 3 octave band analysis for both a normal state in which the headphones 20 are not worn and a state in which NC is enabled, and process the amount of suppression or noise suppression rate of each band as the NC effect amount.
[0042] FIG. 11 shows the estimation result of the second DNN when the noise reduction rate is used as the NC effect amount. Specifically, the estimation result shows the estimation of the noise reduction rate of the correction filter coefficient α using the H2M2 characteristic as an input. Here, if correction is not performed when the noise reduction rate is equal to or greater than a predetermined threshold, the data can be divided into four quadrants as shown in FIG. 11. In FIG. 11, the predetermined threshold is 0.7. Here, the horizontal axis represents the correct data, and the vertical axis represents the estimated data. The signal processing system 1 trains the second DNN in response to the input of the corrected filter coefficient.
[0043] Next, optimization based on the noise reduction rate will be described. Here, the functions of the signal processing system 1 include a function of estimating whether or not noise is reduced by correcting the NC filter. The signal processing system 1 estimates whether or not noise is reduced using a DNN that outputs a noise reduction rate (hereinafter referred to as a "noise reduction rate estimation DNN" or a "third DNN" as appropriate). The third DNN will be described below.
[0044] <2.3. The third DNN> The third DNN takes the H2 user M2 characteristics, H2M2 characteristics, and α default as inputs and outputs the noise suppression rate. The third DNN uses Adam optimization as an example of an optimization method. The third DNN uses a loss function based on the mean square error.
[0045] 2.4. Correction filter estimation process FIG. 12 is a diagram illustrating an overview of the functions of a signal processing system according to an embodiment. FIG. 12 illustrates a case in which a first DNN and a second DNN function together. In FIG. 12, the integrated first DNN and second DNN are collectively referred to as "DNN." In the DNN illustrated in FIG. 12, the H2 user M2 characteristic and the corrected filter are input, and the correction filter coefficient and whether or not to perform correction are output. In addition, in the DNN illustrated in FIG. 12, the corrected filter may be the final output. Note that, although FIG. 12 illustrates a case in which the first DNN and the second DNN are configured to be connected in a fully connected layer so as to be integrated, the first DNN and the second DNN may also be configured to be arranged separately.
[0046] Next, a case will be described in which a correction filter that measures ambient environmental sound and corrects a difference based on the acoustic characteristics of the environmental sound is estimated. Here, the correction filter that corrects an error in the user's wearing state as described above will be referred to as a "wearing error correction filter" or a "first correction filter," as appropriate. Also, a correction filter that corrects a difference based on the acoustic characteristics of the environmental sound will be referred to as an "environmental sound difference correction filter" or a "second correction filter," as appropriate. Here, when estimating the first correction filter, the measurement sound may be buried in noise unless the environment is relatively quiet. When estimating the second correction filter, a relatively loud noise may be desirable because it makes it easier to measure the characteristics of the environmental sound. Therefore, the signal processing system 1 determines whether to estimate the first correction filter or the second correction filter, depending on the noise level of the environmental sound.
[0047] FIG. 13 is a diagram showing an outline of processing using a first correction filter and a second correction filter in addition to the processing of FIG. 12. When estimating the first correction filter, processing is performed based on the same input / output information as in FIG. 12. Here, the functions of the signal processing system 1 include a function of estimating a correction filter coefficient based on the surrounding environmental sound. The signal processing system 1 estimates a corrected filter using a DNN that outputs a second correction filter coefficient (hereinafter referred to as an "environmental sound difference correction filter coefficient estimation DNN" or "fourth DNN" as appropriate). Below, a description will be given of the processing in which the fourth DNN estimates the second correction filter.
[0048] <2.5. The Fourth DNN> The fourth DNN receives the signal collected by the first microphone and the corrected filter at the target time as input and outputs the second correction filter coefficient. The fourth DNN performs optimization using Adam as an example of an optimization method. The fourth DNN uses H1M3 and acoustic characteristics F1 user to measure the surrounding sound field with various environmental sounds. In this case, the signal processing system 1 estimates optimal filter coefficients based on the signal collected by the first microphone and the signal collected by the third microphone. The signal processing system 1 then estimates correction filter coefficients that correct the difference between the α default and the optimal filter coefficients, for example, using a gradient method. The signal processing system 1 then generates training data using the signal collected by the first microphone as input and the estimated correction filter coefficients as output. The fourth DNN may use a loss function to weight both the training data and the estimated data for each frequency band, and then calculate an average from the sum of the amplitude and phase distances for each band. Here, weighting for each frequency band is, for example, weighting based on the exclusion of high frequencies where no NC effect can be expected using a low-pass filter, or the exclusion of low frequencies where frequency resolution is low using a high-pass filter.
[0049] <2.6. Processing flow> Fig. 14 is a flowchart showing the flow of the process related to Fig. 13. The signal processing system 1 determines whether to perform correction based on the first correction filter or the second correction filter depending on the volume of the ambient sound when the optimization function is executed. The flow of the process related to the signal processing device 10 will be described in detail later.
[0050] Fig. 15 is a flowchart showing a process flow for determining whether or not to perform correction based on the second correction filter after determination based on the environmental sound, in addition to the process of Fig. 14. The signal processing system 1 determines whether or not to perform correction based on the second correction filter depending on the magnitude of the estimated second correction filter coefficient.
[0051] Fig. 16 is a modified example of Fig. 15. Fig. 16 is a flowchart showing the flow of processing for comparing the current corrected NC effect estimation result with the new corrected NC effect estimation result to determine whether or not to perform correction. In Fig. 16, it is not necessary to determine whether or not to perform correction based on a comparison with threshold values as shown in Figs. 14 and 15.
[0052] <2.7. Saving and referencing correction filters> FIG. 17 shows an overview of the signal processing system 1's function when it stores (saves) correction filter coefficients and performs processing based on the history of the correction filter coefficients when executing the optimization function. In recent years, products equipped with multiple α defaults, which are preset NC filters, have become popular. While FIG. 17 illustrates a case where optimization processing is performed based on one α default, processing may also be performed based on multiple α defaults. Here, DNN1 in FIG. 17 is the first DNN. DNN2 in FIG. 17 is a DNN that estimates the NC effect when an NC filter with a predetermined filter coefficient is used (hereinafter referred to as the "NC effect estimation DNN" or the "fifth DNN" as appropriate). DNN3 in FIG. 17 is a DNN that estimates the NC effect in an environment set by a predetermined standard (hereinafter referred to as the "NC effect user environment estimation DNN" or the "sixth DNN" as appropriate). Details of the fifth and sixth DNNs will be described later. The NC effect JEITA in FIG. 17 is the NC effect amount in a noise environment according to the JEITA standard. It is possible to use the noise suppression rate as the NC effect amount in a noise environment in the JEITA standard, but in the case of the noise suppression rate, the output will be a single numerical value, which is not sufficient as input for DNN3, so here we will not use the noise suppression rate as the NC effect amount.
[0053] Next, the flow of processing for storing and referencing a correction filter will be described using Figs. 18 to 24. Figs. 18 to 24 will be described using an example of a memory (for example, the storage unit 120) stored by the signal processing device 10. In Figs. 18 to 24, a numerical value is calculated as an index by performing predetermined processing such as weighting and averaging based on the NC effect amount of each band. Note that the predetermined processing is not limited to weighting and averaging based on the NC effect amount of each band, and may be any processing that calculates a numerical value as an index of the NC effect amount. This numerical value is calculated between 0 and 1. In addition, the larger the numerical value, the higher the NC performance will be described. First, the processing for updating a first correction filter stored in memory will be described. Fig. 18 shows a case where optimization processing is not performed.
[0054] 18A shows a state in which nothing is stored in the memory of the correction filter. For example, this is the initial state at the time of purchase. Hereinafter, the state in which the first correction filter is used will be referred to as "N. Standard." Also, hereinafter, the state in which the headphones 20 are worn and optimization has not been performed will be referred to as "O. Unknown."
[0055] FIG. 18B shows a state in which the NC effect amount is stored when the user uses the headphones 20 while traveling on a train without wearing anything, such as glasses, that does not affect the wearing state of the headphones 20. Here, the state of traveling on a train will be referred to as "B. Train" as appropriate. Note that the wearing state is "O. Unknown" because optimization processing has not been performed. Here, an NC effect amount of "0.55" is stored for "B. Train" in the "O. Unknown" state. The signal processing device 10 stores the actual measured value of the NC effect amount for "B. Train" in the "O. Unknown" state. The signal processing device 10 stores the environmental sounds for "B. Train." Note that, for convenience of explanation, the label "B. Train" has been used in the description, but it is assumed that the headphones 20 do not need to recognize that the usage environment at that time is "B. Train."
[0056] FIG. 18C shows a state in which the NC effect amount is stored when the user uses headphones 20 while traveling by bus after "B. Train." Here, the state of traveling by bus is hereinafter referred to as "C. Bus" where appropriate. Here, it is assumed that the NC effect amount in "C. Bus" is greater than the NC effect amount in "B. Train." Here, an NC effect amount of "0.60" is stored for "C. Bus" in the "O. Unknown" state. In FIG. 18C, the signal processing device 10 stores the actual measured value of the NC effect amount in "C. Bus" in the "O. Unknown" state. The signal processing device 10 stores the environmental sound of "C. Bus."
[0057] Next, FIG. 19 shows a case where the user notices the optimization function and executes it in a quiet environment without removing the headphones 20. Here, the state in which the user executes the optimization function without wearing any glasses or other protective gear will be referred to as "P. (wearing) not present" hereinafter. The signal processing device 10 determines that the spatial characteristics of the wear when executing the optimization function in the "P. not present" state are different from those in the "N. standard" state, and estimates the correction filter (p) as the first correction filter. The signal processing device 10 also estimates the NC effect amount for cases where the correction filter (p) is applied and not applied. Here, as the NC effect amount when the correction filter (p) is applied, an NC effect amount of "0.70" is stored in "C. bass" in the "P. not present" state. The signal processing device 10 stores an estimated value of the NC effect amount in "C. bass" in the "P. not present" state. Note that when the correction filter (p) is not applied, an actual measurement value is stored in "O. unknown," and this actual measurement value is used as the NC effect amount. Furthermore, an NC effect amount of "0.74" is stored for "A.JEITA" in the "P. No" state. The signal processing device 10 stores an estimated value of the NC effect amount for "A.JEITA" in the "P. No" state.
[0058] The signal processing device 10 compares the two NC effect amounts, "O. Unknown" and "P. None," for "C. Bus," and updates the first correction filter (S21). Here, the signal processing device 10 compares the NC effect amount of "0.60" for "O. Unknown" with the NC effect amount of "0.70" for "P. None." Since the NC effect amount for "P. None" is greater, the signal processing device 10 updates the first correction filter to the correction filter (p). Next, the signal processing device 10 uses the updated first correction filter to store the NC effect amount when the headphones 20 are used on "C. Bus" while being worn (S22). Here, an NC effect amount of "0.68" is stored for "C. Bus" in the "P. None" state. Next, the signal processing device 10 measures the NC effect amount when the headphones 20 are used on "B. Train" while being worn, and compares it with the NC effect amount when the headphones 20 are used on "C. Bus" (S23). The signal processing device 10 overwrites the NC effect amount because the NC effect amount for "B. Train" is greater. The signal processing device 10 deletes (erases) the memory of "C. Bus" because the environmental sound conditions at the time when the maximum NC effect amount was stored have changed from "C. Bus" to "B. Train."
[0059] After that (for example, at a later date), the signal processing device 10 stores the NC effect amount when the user wears the glasses and uses them in "B. Train" and "C. Bus" without executing the optimization function (S24). Here, an NC effect amount of "0.64" is stored for "B. Train" in the "O. Unknown" state. Next, it is assumed that the user executes the optimization function in a quiet environment without removing the headphones 20. Hereinafter, the state in which optimization is performed while wearing the glasses will be referred to as "Q. Glasses" as appropriate. The signal processing device 10 determines that the characteristics when worn during execution in the "Q. Glasses" state are different from those in "N. Standard" and "P. None," and estimates the correction filter (q) as the first correction filter (S25). The signal processing device 10 also estimates the effect amount for each of "A. JEITA" and "B. Train" in the "Q. Glasses" state. Here, an NC effect amount of "0.70" is stored for "A. JEITA" in the "Q. Glasses" state, and an NC effect amount of "0.71" is stored for "B. Train." Here, since an actual measurement value is stored for "O. Unknown," this actual measurement value is used for the NC effect amount for "B. Train" in the "Q. Glasses" state. If an actual measurement value is not stored for "O. Unknown," the NC effect amount for "A. JEITA" in the "Q. Glasses" state is estimated together with the environmental sound of "B. Train" as an input. Then, the signal processing device 10 compares the two NC effect amounts for "O. Unknown" and "Q. Glasses" in "B. Train" and updates the first correction filter (S26). Here, the signal processing device 10 compares the NC effect amount of "0.64" for "O. Unknown" with the NC effect amount of "0.71" for "Q. Glasses," and since the comparison shows that the NC effect amount of "Q. Glasses" is greater, the signal processing device 10 updates the first correction filter to the correction filter (q).
[0060] FIG. 20 is a flowchart showing the flow of the processing according to FIGS.
[0061] In order to determine whether the correction filters are close to the characteristics of the H2 user M2, the signal processing device 10 may search the list in memory in order of the NC effect amount or the number of times the correction filters have been determined to be close to the characteristics of the H2 user M2, rather than in order of storage or address. This allows the signal processing device 10 to select a correction filter with higher reliability. Here, some users may execute the optimization function less frequently. The headphones 20 may be used multiple times while the optimization function is not executed. For this reason, the signal processing device 10 may store the NC effect amount in the "O. Unknown" state and use it to search for nearby characteristics. The signal processing device 10 may store, for example, (1) the "average value of the NC effect amount in the target wearing state," (2) the "average value of the NC effect amount when the wearing state is unknown," (3) the "number of times the headphones 20 were used when the correction filter was selected in the target wearing state," and (4) the "number of times the headphones 20 were used when the correction filter was selected when the wearing state was unknown," etc., for each correction filter, and use these to search for nearby characteristics.
[0062] Since the number of times in (3) above is likely to depend on the variation in the user's wearing state, the signal processing device 10 may search for nearby characteristics in descending order of the number of times. Here, correction filters with a large number of times in (3) above tend to be in the same wearing state even if the user repeatedly wears and removes the filter multiple times, and therefore may be highly reliable. If correction filters with the same number of times are included, the signal processing device 10 may search in order of the NC effect amount in (1) above. Furthermore, if correction filters with the same NC effect amount are included in (1) above, the signal processing device 10 may search in order of the number of times in (4) above. Then, the signal processing device 10 may search in order of the NC effect amount in (2) above. Note that this search order is an example and is not limited to this search order.
[0063] Next, the process of updating the memory that stores the second correction filter will be described with reference to Figures 21 to 24. Note that the same descriptions as those in Figures 18 to 20 will be omitted as appropriate.
[0064] FIG. 21A shows the memory of the second correction filter at the time of initialization. Hereinafter, the state of the memory of the second correction filter at the time of initialization will be referred to as "A.JEITA (Through)" as appropriate, and the environmental sound at that time will be referred to as "A.JEITA" as appropriate. Also, the state of the memory of the second correction filter after the initial state will be referred to as "n. Standard" as appropriate, and the wearing information at that time will be referred to as "N. Standard" as appropriate. Also, correction filters are represented by a combination of "a" and "n." In FIG. 21A, the signal processing device 10 accesses the memory of the second correction filter at the time of initialization.
[0065] FIG. 21B shows a state in which the NC effect amount when used in "B. Train" is stored without the user wearing anything and without executing the optimization function. Here, an NC effect amount of "0.62" is stored in "NC filter (an) B. Train" in the "O. Unknown" state. In FIG. 21B, the signal processing device 10 stores the actual measured value of the NC effect amount in "NC filter (an) B. Train" in the "O. Unknown" state. The signal processing device 10 stores the environmental sounds of "B. Train."
[0066] FIG. 21C shows a state in which the NC effect amount is stored when the user executes the optimization function for "B. Train." In FIG. 21C, the signal processing device 10 estimates a second correction filter and an NC effect amount. Here, an NC effect amount of "0.72" is stored for "NC filter (bn) B. Train" in the "O. Unknown" state. The signal processing device 10 stores an estimated value of the NC effect amount for "NC filter (bn) B. Train" in the "O. Unknown" state. The process continues to FIG. 22.
[0067] In FIG. 22A, the signal processing device 10 compares the actual measurement value of "NC filter (an) B. train" in the "O. unknown" state with the estimated value of "NC filter (bn) B. train." Specifically, the signal processing device 10 compares the NC effect amount of "0.62," which is the actual measurement value of "NC filter (an) B. train" in the "O. unknown" state, with the NC effect amount of "0.71," which is the estimated value of "NC filter (bn) B. train." Because the newly estimated value of "NC filter (bn) B. train" is larger, the signal processing device 10 determines that this correction filter has higher NC performance and updates the second correction filter. FIG. 22A shows a state in which the actual measurement value of "NC filter (bn) B. train" is stored without the user removing the headphones 20.
[0068] FIG. 22B shows a state in which the NC effect amount is stored when the user uses the headphones 20 in "C. Bass" without removing them and the environmental sound changes. Here, an NC effect amount of "0.66" is stored in "NC filter (bn) C. Bass" in the "O. Unknown" state. In FIG. 22B, the signal processing device 10 stores an estimated value of the NC effect amount in "C. Bass" in the "O. Unknown" state.
[0069] FIG. 22C shows the state in which the NC effect amount is stored when optimization is performed in a quiet environment with the user not wearing any glasses or other items (the "P. None" state) after that (for example, at a later date). In this case, the signal processing device 10 assumes that the headphones 20 have been detached and clears all values corresponding to the "O. Unknown" state. The signal processing device 10 determines that the "P. None" state has characteristics different from those of the "N. Standard" state stored in the memory, and estimates a correction filter (p) corresponding to "P. None." The signal processing device 10 also estimates the NC effect amount of the "NC filter (ap) A.JEITA" and the "NC filter (an) A.JEITA" in the "P. None" state. Here, an NC effect amount of "0.77" is stored in the "NC filter (ap) A.JEITA" in the "P. None" state, and an NC effect amount of "0.68" is stored in the "NC filter (an) A.JEITA." Then, based on the estimation result, the signal processing device 10 updates the second correction filter to the correction filter (p) because the estimated value of the estimated "NC filter (ap) A.JEITA" is larger. Then, the process continues to FIG.
[0070] FIG. 23A shows a state in which the NC effect amount is stored when the user uses the headphones 20 in "B. Train" and "C. Bus" without removing them. In FIG. 23A, the signal processing device 10 stores estimated values of the NC effect amount in "B. Train" and "C. Bus" in the "P. No" state. Here, an NC effect amount of "0.78" is stored for "B. Train" in the "P. No" state, and an NC effect amount of "0.70" is stored for "C. Bus."
[0071] FIG. 23B shows a state in which the NC effect amount is stored when the user wears the glasses afterward (for example, at a later date) and uses them in "C. Bus" and "D. Airplane" without executing the optimization function after wearing them. Here, since the user does not execute the optimization function after wearing the glasses, the amount is stored as "O. Unknown." In FIG. 23(B), the signal processing device 10 stores the actual measured values of the NC effect amount for "C. Bus" and "D. Airplane" in the "O. Unknown" state. Here, an NC effect amount of "0.58" is stored for "C. Bus" in the "O. Unknown" state, and an NC effect amount of "0.62" is stored for "D. Airplane."
[0072] FIG. 23C shows the state in which the NC effect amount is stored when the optimization function is executed in a quiet environment while the user is still wearing the headphones 20. The signal processing device 10 determines that the "Q. Glasses" state has characteristics different from the "N. Standard" and "P. None" states, and estimates a correction filter (q) corresponding to "Q. Glasses." The signal processing device 10 also estimates the NC effect amounts for the "NC filter (ap) A. JEITA," the "NC filter (bp) B. Train," and the "NC filter (bq) B. Train" in the "Q. Glasses" state. Here, an NC effect amount of "0.74" is stored for the "NC filter (ap) A. JEITA" in the "Q. Glasses" state, an NC effect amount of "0.66" is stored for the "NC filter (bp) B. Train," and an NC effect amount of "0.77" is stored for the "NC filter (bq) B. Train."
[0073] In FIG. 23D, the newly estimated value of "NC filter (bq) B. Train" is larger, so the signal processing device 10 determines that this NC performance is high and selects the correction filter (q) as the second correction filter. FIG. 23D shows a state in which the NC effect amount is stored when the user uses the headphones 20 for "C. Bus" and "D. Airplane" while wearing the headphones 20. In FIG. 23D, the signal processing device 10 stores estimated values of the NC effect amount for "C. Bus" and "D. Airplane" while wearing the headphones 20. Here, an NC effect amount of "0.70" is stored for "C. Bus" in the "Q. Glasses" state, and an NC effect amount of "0.78" is stored for "D. Airplane."
[0074] FIG. 24 is a flowchart showing the flow of the processing according to FIGS.
[0075] <2.8. The Fifth DNN> Next, optimization based on the estimation result of the correction filter will be described. Here, the function of the signal processing system 1 includes a function for estimating the NC effect when an NC filter having a predetermined filter coefficient is used. The signal processing system 1 estimates the NC effect using a fifth DNN. The fifth DNN will be described below.
[0076] The fifth DNN takes the H2 user M2 characteristics and corrected filter coefficients as input and outputs the NC effect amount. The fifth DNN may also take the H2M2 characteristics as input in addition to the above. The fifth DNN performs optimization using Adam as an example of an optimization method. The fifth DNN uses a loss function based on the root mean square error. The fifth DNN performs an NC simulation using the training data generated by the first DNN, and the NC effect amount obtained as a result of the simulation is used as training data.
[0077] <2.9. The 6th DNN> Next, optimization based on an environment set by a predetermined standard will be described. Here, the functions of the signal processing system 1 include a function for estimating the NC effect in an environment set by a predetermined standard. The signal processing system 1 estimates the NC effect using a sixth DNN. The sixth DNN will be described below.
[0078] The sixth DNN receives as input the NC effect amount in a noise environment of a predetermined standard, the corrected filter coefficient, and the characteristics of the environmental sound in the user's usage environment, and outputs the NC effect amount in the user's usage environment. The sixth DNN uses a loss function based on the root mean square error. The sixth DNN uses the NC effect amount obtained as a result of the NC simulation as training data. For example, the sixth DNN performs an NC simulation using an NC filter, a correction filter, and data on environmental sound (e.g., sound data on environmental sound measured by the first to third microphones) and characteristics, and uses the NC effect amount obtained as a result of the simulation as training data.
[0079] <2.10. Functional configuration example> FIG. 25 is a block diagram showing an example of the functional configuration of a signal processing system 1 according to an embodiment.
[0080] (1) Signal Processing Device 10 25, the signal processing device 10 includes a communication unit 100, a control unit 110, and a storage unit 120. The signal processing device 10 includes at least the control unit 110.
[0081] (1-1) Communication unit 100 The communication unit 100 has a function of communicating with an external device. For example, in communication with the external device, the communication unit 100 outputs information received from the external device to the control unit 110. Specifically, the communication unit 100 outputs information received from the headphones 20 to the control unit 110. For example, the communication unit 100 outputs a signal picked up by a microphone provided in the headphones 20 to the control unit 110.
[0082] In communication with an external device, the communication unit 100 transmits information input from the control unit 110 to the external device. Specifically, the communication unit 100 transmits information related to the acquisition of a picked-up signal input from the control unit 110 to the headphones 20. The communication unit 100 is configured as a hardware circuit (such as a communications processor), and can be configured to perform processing by a computer program running on the hardware circuit or on another processing device (such as a CPU) that controls the hardware circuit.
[0083] (1-2) Control Unit 110 The control unit 110 has a function of controlling the operation of the signal processing device 10. For example, the control unit 110 performs a process of determining correction filter coefficients for performing optimal NC for an individual user.
[0084] 25, the control unit 110 has an acquisition unit 111, a processing unit 112, and an output unit 113. The control unit 110 may be configured with a processor such as a CPU, and may be configured to read software (computer programs) that realize the functions of the acquisition unit 111, processing unit 112, and output unit 113 from the storage unit 120 and perform processing. Furthermore, one or more of the acquisition unit 111, processing unit 112, and output unit 113 may be configured as a hardware circuit (such as a processor) separate from the control unit 110, and may be configured to be controlled by a computer program that runs on the separate hardware circuit or on the control unit 110.
[0085] ·Acquisition part 111 The acquisition unit 111 has a function of acquiring acoustic characteristics in the user's ear isolated from the outside world. The acquisition unit 111 acquires acoustic characteristics based on, for example, a sound pickup signal obtained by collecting a test sound output into the ear. For example, the acquisition unit 111 acquires acoustic characteristics based on a sound pickup signal collected by a microphone of a sound output device.
[0086] The acquisition unit 111 acquires data stored in the storage unit 120. For example, the acquisition unit 111 acquires information related to correction filter coefficients.
[0087] Processing section 112 The processing unit 112 has a function for controlling the processing of the signal processing device 10. As shown in Fig. 25 , the processing unit 112 has a determination unit 1121, an NC filter unit 1122, a correction unit 1123, a generation unit 1124, and a judgment unit 1125. The determination unit 1121, the NC filter unit 1122, the correction unit 1123, the generation unit 1124, and the judgment unit 1125 of the processing unit 112 may each be configured as an independent computer program module, or multiple functions may be configured as a single integrated computer program module.
[0088] Decision unit 1121 The determination unit 1121 has a function of determining correction filter coefficients based on the acoustic characteristics acquired by the acquisition unit 111 .
[0089] The determination unit 1121 determines the correction filter coefficients using a trained model (for example, a first DNN) that receives acoustic characteristics as input and outputs filter coefficients. For example, the determination unit 1121 determines the correction filter coefficients using a trained model that has trained using acoustic characteristics estimated at the position of the user's eardrum as training data.
[0090] The determination unit 1121 receives acoustic characteristics and sound data as input and determines the correction filter coefficients using a trained model (for example, a second DNN) that outputs whether or not to correct the sound data. For example, the determination unit 1121 determines the correction filter coefficients using a trained model that has learned, as training data, appended information labeled with whether or not to correct based on a noise reduction rate estimated based on the acoustic characteristics and the sound data.
[0091] The determination unit 1121 determines the correction filter coefficients using a trained model (for example, a third DNN) that receives acoustic characteristics, pre-measured acoustic characteristics, and sound data as inputs and outputs a noise reduction rate. For example, the determination unit 1121 determines the correction filter coefficients using a trained model that has trained using, as training data, a noise reduction rate based on acoustic characteristics estimated at the position of the user's eardrum and sound data.
[0092] The determination unit 1121 receives as input a sound signal collected by a microphone different from the microphone that measured the acoustic characteristics and sound data, and determines a correction filter coefficient using a trained model (fourth DNN) that outputs a correction filter coefficient that corrects a difference in filter coefficients based on environmental sounds in the user's environment. For example, the determination unit 1121 determines the correction filter coefficient using a trained model that has trained, as training data, a filter coefficient that corrects a difference in filter coefficients based on acoustic characteristics estimated at the position of the user's eardrum.
[0093] The determination unit 1121 determines the correction filter coefficients using a trained model (for example, a fifth DNN) that receives acoustic characteristics and sound data as input and outputs an NC effect amount. For example, the determination unit 1121 determines the correction filter coefficients using a trained model that has trained, as training data, an effect amount based on acoustic characteristics estimated at the position of the user's eardrum.
[0094] The determination unit 1121 receives the NC effect amount in an environment defined by a predetermined standard, sound data, and the acoustic characteristics of environmental sound in the user environment as inputs, and determines the correction filter coefficients using a trained model (sixth DNN) that outputs the NC effect amount in the user environment. For example, the determination unit 1121 determines the correction filter coefficients using a trained model that has learned the NC effect amount based on the sound data, filter coefficients, and the acoustic characteristics of environmental sound in the user environment as training data.
[0095] NC filter part 1122 The NC filter unit 1122 has a function of generating sound data that is in the opposite phase to the environmental sound that has leaked into the user's ear. For example, the NC filter unit 1122 generates sound data that is in the opposite phase to the acoustic characteristics of the environmental sound acquired by the acquisition unit 111.
[0096] Correction unit 1123 The correction unit 1123 has a function of correcting, using a correction filter, the sound data generated by the NC filter unit 1122. Specifically, the correction unit 1123 performs the correction using the correction filter coefficient determined by the determination unit 1121.
[0097] ·Generation unit 1124 The generation unit 1124 has a function of generating a trained model. The generation unit 1124 generates a trained model that is trained by, for example, inputting input data and output data into a loss function. The determination unit 1121 determines correction filter coefficients estimated using the trained model generated by the generation unit 1124.
[0098] ·Judgment section 1125 The determination unit 1125 has a function of determining whether or not to use a correction filter to correct the sound data generated by the NC filter unit 1122. For example, the determination unit 1125 determines whether or not a sufficient correction effect can be expected by using a correction filter, and if a sufficient correction effect can be expected, determines that correction should be performed using the correction filter.
[0099] The determination unit 1125 determines the noise level of the environmental sound. The determination unit 1125 determines whether to use the first correction filter or the second correction filter, depending on the noise level of the environmental sound.
[0100] Output section 113 The output unit 113 has a function of outputting sound data corrected by the correction unit 1123. The output unit 113 provides the corrected sound data to, for example, the headphones 20 via the communication unit 100. Upon receiving the corrected sound data, the headphones 20 reproduce sound based on the corrected sound data. This allows the user to preview the sound corrected by the correction filter.
[0101] (1-3) Storage section 120 The storage unit 120 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 120 has a function of storing computer programs and data (including one form of program) related to the processing in the signal processing device 10.
[0102] Fig. 26 shows an example of the storage unit 120. As shown in Fig. 26, the storage unit 120 may have items such as "correction filter coefficient ID," "correction filter coefficient," "execution state," "usage environment 1," and "usage environment 2."
[0103] "Correction filter coefficient ID" indicates identification information for identifying a correction filter coefficient. "Correction filter coefficient" indicates a correction filter coefficient. "Execution status" indicates the execution status of the optimization function. In the example shown in FIG. 26, conceptual information such as "Execution status #1" and "Execution status #2" is stored in "Execution status", but in reality, data such as "N. Standard" and "O. Unknown" is stored. "Usage environment 1" and the like indicate the user's usage environment. In the example shown in FIG. 26, conceptual information such as "Usage environment #1" and "Usage environment #2" is stored in "Usage environment 1", but in reality, data such as "B. Train" and "C. Bus" is stored.
[0104] (2) Headphones 20 As shown in FIG. 25, the headphones 20 include a communication unit 200, a control unit 210, and an output unit 220.
[0105] (2-1) Communications Department 200 The communication unit 200 has a function of communicating with an external device. For example, in communication with the external device, the communication unit 200 outputs information received from the external device to the control unit 210. Specifically, the communication unit 200 outputs information received from the signal processing device 10 to the control unit 210. For example, the communication unit 200 outputs information related to the acquisition of sound data corrected by a correction filter to the control unit 210.
[0106] (2-2) Control Unit 210 The control unit 210 has a function of controlling the operation of the headphones 20. For example, the control unit 210 transmits, via the communication unit 200, to the signal processing device 10, acoustic characteristics based on a sound signal collected by a microphone.
[0107] (2-3) Output unit 220 The output unit 220 is realized by a member capable of outputting sound, such as a speaker, etc. The output unit 220 outputs sound based on sound data.
[0108] <2.11. Signal Processing System> The functions of the signal processing system 1 according to the embodiment have been described above. Next, the processing of the signal processing system 1 will be described.
[0109] FIG. 27 is a flowchart showing a processing flow in the signal processing device 10 according to the embodiment. The signal processing device 10 acquires acoustic characteristics in the user's ear isolated from the outside world (S101). Next, the signal processing device 10 determines a correction filter coefficient using a trained model that outputs a correction filter coefficient when the acquired acoustic characteristics are input (S102). Then, the signal processing device 10 generates sound data that is out of phase with the environmental sound that has leaked into the user's ear (S103). Next, the signal processing device 10 determines whether or not to perform correction using a correction filter (S104). If the signal processing device 10 determines to perform correction using a correction filter (S104; YES), it corrects the generated sound data using the determined correction filter coefficient (S105). If the signal processing device 10 determines not to perform correction using a correction filter (S104; NO), it ends the information processing.
[0110] <2.12. Processing Variations> (Selecting a correction filter using the UI) In the above embodiment, the signal processing device 10 determines whether to perform correction using machine learning such as DNN, but the present invention is not limited to this example. For example, the signal processing device 10 may determine whether to perform correction by receiving a selection from a user.
[0111] Whether a higher NC effect amount provides a more comfortable experience for the user may depend on the user's subjective opinion. For example, a higher NC effect amount may result in discomfort for the user when, for example, high-frequency noise that was previously masked by the noise is relatively emphasized due to significant suppression of mid- and low-frequency noise, resulting in harshness. The signal processing device 10 may determine whether or not to perform correction by presenting the NC effect amount using the current filter coefficient, the NC effect amount using estimated correction filter coefficients, the NC effect amount using correction filter coefficients stored in memory, or the like, and accepting a selection from the user. For example, the signal processing device 10 may display a list of correction filters on a mobile terminal such as a smartphone (hereinafter, appropriately referred to as the "terminal device 30") to accept a selection from the user. For example, the signal processing device 10 may display a list of correction filters according to the user's wearing state. This allows the signal processing device 10 to explicitly select a correction filter. Furthermore, the signal processing device 10 may allow the user to check the NC effect amount for any environmental sound.
[0112] FIG. 28 shows an example of a display screen displaying a list of correction filters. In FIG. 28, the list of correction filters includes “Standard,” “Filter 1,” and “Filter 2.” Here, “Standard” is a correction filter estimated by the signal processing device 10 when, for example, the user is not wearing anything. “Filter 1” is a correction filter estimated by the signal processing device 10 when, for example, the user is wearing glasses. “Filter 2” is a correction filter estimated by the signal processing device 10 when, for example, the user is wearing a hat. In FIG. 28, the display screen HG11 displaying the list of correction filters includes a predetermined area SK11 in which a correction filter based on a new measurement is added as an option when the user operates (e.g., clicks or taps) measurement B11. The display screen HG11 also includes a predetermined area SK12 in which the characteristics of a correction filter selected by the user are highlighted. When the user operates preview C11 included in the display screen HG11, the terminal device 30 outputs sound based on the correction filter selected by the user, for example.
[0113] Upon receiving an operation for the preview C11, the signal processing device 10 may perform processing for outputting sound based on a correction filter selected by the user. This allows the user to preview sound based on the selected correction filter. Here, during preview, the signal processing device 10 may select and play a sound (e.g., a song) stored in the terminal device 30 so that the user can easily recognize the difference between the correction filters included in the list. Alternatively, the signal processing device 10 may play any sound selected in advance by the user. This allows the signal processing device 10 to easily compare correction filters in the user's usage environment. The signal processing device 10 may also perform processing for displaying the H2 user M2 characteristics. This allows the signal processing device 10 to visually grasp the H2 user M2 characteristics. The signal processing device 10 may also perform processing for allowing the user to name each correction filter on the UI of the terminal device 30. This allows the signal processing device 10 to easily use different correction filters by allowing the user to name them. In this case, the ease of recognizing and operating the information displayed on the UI may be reduced. For this reason, the signal processing device 10 may perform processing to enable the user to compare the listening experience using only the UI of the headphones 20 using a guide voice or the like. Furthermore, the signal processing device 10 may perform processing to enable the user's terminal device 30 or a server to which the terminal device 30 is connected to perform estimation processing of the correction filter coefficients.
[0114] Next, a case will be described in which correction filters for errors in ambient environmental sounds are managed and operated by the terminal device 30. Here, the display screen of the terminal device 30 is provided with a tab for correcting a wearing error and a tab for correcting a difference in environmental sounds, and the user selects a tab to switch the list of correction filters. FIG. 29 shows an example of a display screen in which a list of first correction filters and a list of second correction filters are managed and selected using tabs. Note that descriptions similar to those in FIG. 28 will be omitted as appropriate. The display screen HG21 includes tabs TB11 and TB12 that the user selects to switch the list of correction filters. When the user selects tab TB11 or tab TB12 included in the display screen HG21, the terminal device 30 displays a list of correction filters corresponding to tab TB11 or tab TB12. Upon receiving the user's selection, the signal processing device 10 may perform processing to switch the list of correction filters corresponding to the tab selected by the user. This allows the user to manage and select correction filters separately depending on the type of correction filter. Furthermore, when the second correction filter tab is selected, the signal processing device 10 may display the acoustic characteristics of the environmental sound targeted by the default NC filter of the product and the acoustic characteristics of the environmental sound in the user's environment, which allows the user to refer to them when making a selection.
[0115] The terminal device 30 according to the embodiment is not limited to a mobile terminal such as a smartphone, but may be any information processing device that can accept an operation related to a correction filter from a user.
[0116] (Processing when environmental sounds change at any time) In the above embodiment, the signal processing device 10 updates the estimated correction filter coefficients triggered by a user operation, but this is not limiting. The signal processing device 10 may also update the estimated correction filter coefficients as needed for environmental sounds that change as needed. As shown in FIG. 30 , the signal processing device 10 may update the correction filter coefficients in accordance with changes in the environmental sounds by crossfading the correction filters. This allows the signal processing device 10 to update the correction filter coefficients without causing sound interruptions or a sense of discomfort. Note that the signal processing device 10 may update the correction filter coefficients based on any process, not limited to crossfading.
[0117] (NC filter estimation) In the above embodiment, the signal processing device 10 has estimated a correction filter coefficient for a difference between environmental sounds, but the signal processing device 10 may estimate a filter coefficient for an NC filter. For example, the signal processing device 10 may estimate a filter coefficient that minimizes the signal picked up by the third microphone, based on a signal picked up by the first microphone and a signal picked up by the third microphone. In the above embodiment, the signal processing device 10 has used correction filter coefficients estimated for various environmental sounds as training data, but the signal processing device 10 may estimate a correction filter coefficient by determining a standard filter coefficient.
[0118] (Processing when adjusting gain) In the above embodiment, the signal processing device 10 determines a correction filter coefficient and performs correction using the determined correction filter coefficient. However, the signal processing device 10 may perform correction by adjusting the filter gain without determining a correction filter coefficient. In this case, the signal processing device 10 may add an offset based on the error between the H2M2 characteristics and the H2 user M2 characteristics. The signal processing device 10 may also adjust this offset and calculate an offset value that minimizes the sum of squares of the error. If the minimum sum of squares error of this offset value is smaller than a predetermined threshold, the signal processing device 10 may perform correction using the offset value as an adjustment value for the gain. The signal processing device 10 may also accept adjustments from the user based on the offset value. This allows the signal processing device 10 to make adjustments according to the user's subjective preferences and hearing condition. If the minimum sum of squares error of the offset value is larger than a predetermined threshold, the signal processing device 10 may also estimate the correction filter coefficient.
[0119] Fig. 31 shows a case where the least square sum error of the offset value is smaller than a predetermined threshold, Fig. 31(A) shows the case before gain adjustment, and Fig. 31(B) shows the case after gain adjustment.
[0120] Fig. 32 shows a case where the least square sum error of the offset value is greater than a predetermined threshold, Fig. 32(A) shows the case before gain adjustment, and Fig. 32(B) shows the case after gain adjustment.
[0121] FIG. 33 is a flowchart showing the flow of processing when adjusting the gain.
[0122] (Error correction) In the above embodiment, the case where errors due to individual differences between users and the wearing state are corrected has been described, but the correction is not limited to these cases. The correction according to the embodiment also includes, for example, the case where errors due to individual differences of the headphones 20, etc. are corrected.
[0123] <<3. Hardware configuration example>> Finally, an example of the hardware configuration of a signal processing device according to an embodiment will be described with reference to Fig. 34. Fig. 34 is a block diagram showing an example of the hardware configuration of a signal processing device according to an embodiment. Note that the signal processing device 900 shown in Fig. 34 can realize, for example, the signal processing device 10 and headphones 20 shown in Fig. 25. Information processing by the signal processing device 10 and headphones 20 according to the embodiment is realized by cooperation between software (composed of a computer program) and hardware described below.
[0124] 34, the signal processing device 900 includes a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, and a RAM (Random Access Memory) 903. The signal processing device 900 also includes a host bus 904a, a bridge 904, an external bus 904b, an interface 905, an input device 906, an output device 907, a storage device 908, a drive 909, a connection port 910, and a communication device 911. Note that the hardware configuration shown here is an example, and some of the components may be omitted. The hardware configuration may also include components other than those shown here.
[0125] The CPU 901 functions as, for example, an arithmetic processing device or a control device, and controls the overall operation or part of the operation of each component based on various computer programs recorded in the ROM 902, the RAM 903, or the storage device 908. The ROM 902 is a means for storing programs loaded into the CPU 901 and data used for calculations. The RAM 903 temporarily or permanently stores, for example, the programs loaded into the CPU 901 and data (parts of the programs) such as various parameters that change as appropriate when the programs are executed. These are connected to each other by a host bus 904a consisting of a CPU bus or the like. The CPU 901, the ROM 902, and the RAM 903 can realize the functions of the control unit 110 and the control unit 210 described with reference to FIG. 25, for example, in cooperation with software.
[0126] The CPU 901, ROM 902, and RAM 903 are interconnected via, for example, a host bus 904a capable of high-speed data transmission. On the other hand, the host bus 904a is connected to an external bus 904b, which has a relatively low data transmission speed, via, for example, a bridge 904. Furthermore, the external bus 904b is connected to various components via an interface 905.
[0127] The input device 906 is realized by a device to which information is input by a listener, such as a mouse, keyboard, touch panel, button, microphone, switch, or lever. The input device 906 may also be, for example, a remote control device using infrared or other radio waves, or an externally connected device such as a mobile phone or PDA that is compatible with the operation of the signal processing device 900. The input device 906 may also include, for example, an input control circuit that generates an input signal based on information input using the above-mentioned input means and outputs the signal to the CPU 901. An administrator of the signal processing device 900 can input various data to the signal processing device 900 and instruct processing operations by operating the input device 906.
[0128] Alternatively, the input device 906 may be formed by a device that detects the position of the user. For example, the input device 906 may include various sensors such as an image sensor (e.g., a camera), a depth sensor (e.g., a stereo camera), an acceleration sensor, a gyro sensor, a geomagnetic sensor, a light sensor, a sound sensor, a distance measurement sensor (e.g., a ToF (Time of Flight) sensor), and a force sensor. The input device 906 may also acquire information about the state of the signal processing device 900 itself, such as the attitude and movement speed of the signal processing device 900, and information about the space around the signal processing device 900, such as brightness and noise around the signal processing device 900. The input device 906 may also include a GNSS module that receives GNSS signals from GNSS (Global Navigation Satellite System) satellites (e.g., GPS signals from GPS (Global Positioning System) satellites) to measure position information including the latitude, longitude, and altitude of the device. Regarding the location information, the input device 906 may detect the location by transmitting and receiving data via Wi-Fi (registered trademark), a mobile phone, a PHS, a smartphone, or by short-range communication. The input device 906 may realize the function of the acquisition unit 111 described with reference to FIG. 25, for example.
[0129] The output device 907 is formed by a device capable of visually or audibly notifying the user of acquired information. Examples of such devices include display devices such as CRT display devices, liquid crystal display devices, plasma display devices, EL display devices, laser projectors, LED projectors, and lamps, audio output devices such as speakers and headphones, and printer devices. The output device 907 outputs, for example, results obtained from various processes performed by the signal processing device 900. Specifically, the display device visually displays the results obtained from various processes performed by the signal processing device 900 in various formats such as text, images, tables, and graphs. On the other hand, the audio output device converts audio signals consisting of reproduced audio data, acoustic data, etc. into analog signals and outputs them audibly. The output device 907 can, for example, implement the functions of the output unit 113 and the output unit 220 described with reference to FIG. 25 .
[0130] The storage device 908 is a data storage device formed as an example of a storage unit of the signal processing device 900. The storage device 908 is realized, for example, by a magnetic storage device such as an HDD, a semiconductor storage device, an optical storage device, or a magneto-optical storage device. The storage device 908 may include a storage medium, a recording device that records data on the storage medium, a reading device that reads data from the storage medium, and a deletion device that deletes data recorded on the storage medium. The storage device 908 stores computer programs executed by the CPU 901, various data, and various data acquired from the outside. The storage device 908 can realize, for example, the function of the storage unit 120 described with reference to FIG. 25 .
[0131] The drive 909 is a reader / writer for a storage medium, and is built into or externally attached to the signal processing device 900. The drive 909 reads information recorded on a removable storage medium such as an attached magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, and outputs the information to the RAM 903. The drive 909 can also write information to the removable storage medium.
[0132] The connection port 910 is a port for connecting an external device, such as a Universal Serial Bus (USB) port, an IEEE1394 port, a Small Computer System Interface (SCSI), an RS-232C port, or an optical audio terminal.
[0133] The communication device 911 is, for example, a communication interface formed by a communication device or the like for connecting to the network 920. The communication device 911 is, for example, a communication card for a wired or wireless LAN (Local Area Network), LTE (Long Term Evolution), Bluetooth (registered trademark), or WUSB (Wireless USB). The communication device 911 may also be a router for optical communication, a router for ADSL (Asymmetric Digital Subscriber Line), or a modem for various communications. The communication device 911 can transmit and receive signals, for example, between the Internet and other communication devices in accordance with a predetermined protocol such as TCP / IP. The communication device 911 can implement, for example, the functions of the communication unit 100 and the communication unit 200 described with reference to FIG. 25 .
[0134] The network 920 is a wired or wireless transmission path for information transmitted from devices connected to the network 920. For example, the network 920 may include public networks such as the Internet, telephone networks, and satellite communication networks, as well as various LANs (Local Area Networks) including Ethernet (registered trademark), and WANs (Wide Area Networks). The network 920 may also include dedicated network such as an IP-VPN (Internet Protocol-Virtual Private Network).
[0135] The above describes an example of a hardware configuration capable of realizing the functions of the signal processing device 900 according to the embodiment. Each of the above components may be realized using general-purpose components, or may be realized by hardware specialized for the function of each component. Therefore, the hardware configuration to be used can be changed as appropriate depending on the technical level at the time of implementing the embodiment.
[0136] <<4. Summary>> As described above, the signal processing device 10 according to the embodiment performs a process of determining correction filter coefficients based on the acoustic characteristics of the inside of a user's ear, which is isolated from the outside world. The signal processing device 10 also performs a process of correcting sound data that is out of phase with environmental sounds that have leaked into the user's ear using a correction filter. This allows the signal processing device 10 to determine correction filter coefficients for optimization without requiring, for example, an acoustic signal at the eardrum position, which is difficult to incorporate into a product. Furthermore, the signal processing device 10 can promote improvement of the NC effect by performing correction using a correction filter.
[0137] Therefore, it is possible to provide a new and improved signal processing device, a signal processing method, a signal processing model manufacturing method, and an audio output device that can promote further improvements in usability.
[0138] Although the preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the technical scope of the present disclosure is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical idea described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.
[0139] For example, each device described in this specification may be realized as a standalone device, or some or all of them may be realized as separate devices. For example, the signal processing device 10 and headphones 20 shown in Fig. 25 may each be realized as a standalone device. Furthermore, for example, they may be realized as a server device connected to the signal processing device 10 and headphones 20 via a network or the like. Furthermore, the function of the control unit 110 of the signal processing device 10 may be configured to be provided by a server device connected via a network or the like.
[0140] Furthermore, the series of processes performed by each device described herein may be realized using software, hardware, or a combination of software and hardware. The computer programs constituting the software are stored in advance, for example, on a recording medium (non-transitory medium) provided inside or outside each device. Then, each program is loaded into RAM when executed by a computer and executed by a processor such as a CPU.
[0141] Furthermore, the processes described herein using flowcharts do not necessarily have to be performed in the order shown. Some process steps may be performed in parallel. Additional process steps may be employed, and some process steps may be omitted.
[0142] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.
[0143] The following configurations also fall within the technical scope of the present disclosure. (1) an acquisition unit that acquires acoustic characteristics in the user's ear that are isolated from the outside world; an NC filter unit that generates sound data in an opposite phase to the environmental sound that has leaked into the user's ear; a correction unit that corrects the sound data using a correction filter; a determination unit that determines a filter coefficient of the correction filter based on the acoustic characteristics; A signal processing device comprising: (2) The acquisition unit The acoustic characteristics are obtained based on a sound pickup signal obtained by collecting the test sound output into the ear. The signal processing device according to (1) above. (3) The determination unit The filter coefficients are determined using a trained model that takes acoustic characteristics as input and outputs filter coefficients. The signal processing device according to (1) or (2). (4) The determination unit The filter coefficients are determined using the trained model that has been trained using the acoustic characteristics estimated at the position of the user's eardrum as training data. The signal processing device according to (3) above. (5) The determination unit The filter coefficients are determined using a trained model that receives acoustic characteristics and sound data as input and outputs whether or not to correct the sound data. The signal processing device according to any one of (1) to (4). (6) The determination unit The filter coefficients are determined using the trained model that has been trained using training data that is attached information labeled with whether or not to perform correction based on a noise suppression rate estimated based on acoustic characteristics and sound data. The signal processing device according to (5) above. (7) The determination unit The filter coefficients are determined using a trained model that takes acoustic characteristics, previously measured acoustic characteristics, and sound data as inputs and outputs a noise suppression rate. The signal processing device according to any one of (1) to (6). (8) The determination unit The filter coefficients are determined using the trained model that has been trained using the acoustic characteristics estimated at the position of the eardrum of the user and the noise reduction rate based on sound data as training data. The signal processing device according to (7) above. (9) The determination unit The filter coefficients are determined using a trained model that receives as input a sound signal collected by a microphone different from the microphone used to measure the acoustic characteristics and sound data, and outputs correction filter coefficients that correct differences in filter coefficients based on environmental sounds in the user's environment. The signal processing device according to any one of (1) to (8). (10) The determination unit The filter coefficients are determined using the trained model, which has been trained using as training data filter coefficients that correct differences in filter coefficients based on acoustic characteristics estimated at the position of the user's eardrum. The signal processing device according to (9) above. (11) The determination unit The filter coefficients are determined using a trained model that takes acoustic characteristics and sound data as inputs and outputs the NC effect amount. The signal processing device according to any one of (1) to (10) above. (12) The determination unit The filter coefficients are determined using the trained model that has been trained using training data of an effect amount based on acoustic characteristics estimated at the position of the user's eardrum. The signal processing device according to (11) above. (13) The determination unit The filter coefficients are determined using a trained model that receives as input the NC effect amount in an environment defined by a predetermined standard, sound data, and the acoustic characteristics of the environmental sound in the user's environment, and outputs the NC effect amount in the user's environment. The signal processing device according to any one of (1) to (12) above. (14) The determination unit The filter coefficients are determined using the trained model that has been trained using the sound data, the filter coefficients, and the NC effect amount based on the acoustic characteristics of the environmental sound in the user's environment as training data. The signal processing device according to (13) above. (15) 1. A computer-implemented signal processing method comprising: an acquisition step of acquiring acoustic characteristics in the user's ear isolated from the outside world; an NC filtering process for generating sound data in an antiphase with the environmental sound leaking into the user's ear; a correction step of correcting the sound data using a correction filter; a determining step of determining filter coefficients of the correction filter based on the acoustic characteristics; A signal processing method comprising: (16) an acquisition step for acquiring acoustic characteristics in the user's ear isolated from the outside world; an NC filter step for generating sound data in an antiphase with the environmental sound leaking into the user's ear; a correction step of correcting the sound data using a correction filter; a determination step of determining filter coefficients of the correction filter based on the acoustic characteristics; A signal processing program that causes a computer to execute the above. (17) A signal processing model manufacturing method for manufacturing a model for optimal noise canceling by learning the acoustic characteristics based on a signal collected in advance by a microphone and the correction filter coefficients for optimal noise canceling as inputs, the method determining whether to correct filter coefficients based on acoustic characteristics based on a signal collected by a microphone, determining filter coefficients for optimal noise canceling, and generating a noise canceling signal based on the determined filter coefficients. (18) An audio output device comprising an output unit that outputs noise-canceled sound based on a signal provided from a signal processing device, wherein the signal processing device determines filter coefficients for optimal noise cancellation based on acoustic characteristics based on a sound signal collected by a microphone of the audio output device, and provides a signal generated based on the determined filter coefficients. [Explanation of symbols]
[0144] 1. Signal Processing System 10. Signal Processing Device 20 headphones 30 Terminal Equipment 100 Communications Department 110 control section 111 Acquisition Department 112 Processing section 1121 Decision Section 1122 NC filter section 1123 Correction Unit 1124 Generation part 1125 Judgment section 113 Output section 200 Communications Department 210 Control Unit 220 Output section
Claims
1. an acquisition unit that acquires acoustic characteristics in the user's ear that are isolated from the outside world; an NC filter unit that generates sound data that is in an opposite phase to the environmental sound that has leaked into the user's ear; a correction unit that corrects the sound data using a correction filter; a determination unit that determines a filter coefficient of the correction filter used in the correction unit according to output information output by a trained model that outputs predetermined information related to the acoustic characteristics and the sound data when the acoustic characteristics and sound data are input, by inputting the sound data generated by the NC filter unit and the acoustic characteristics acquired by the acquisition unit into the trained model; A signal processing device comprising:
2. The acquisition unit The acoustic characteristics are obtained based on a sound pickup signal obtained by collecting the test sound output into the ear. The signal processing device according to claim 1 .
3. The determination unit The filter coefficients are determined using the trained model, which receives acoustic characteristics and sound data as inputs and outputs information for determining the filter coefficients. The signal processing device according to claim 1 .
4. The determination unit The filter coefficients are determined using the trained model that has been trained using the acoustic characteristics estimated at the position of the user's eardrum as training data. The signal processing device according to claim 3 .
5. The determination unit The filter coefficients are determined using the trained model, which receives acoustic characteristics and sound data as input and outputs whether or not to correct the sound data. The signal processing device according to claim 1 .
6. The determination unit The filter coefficients are determined using the trained model that has been trained using training data that is attached information labeled with whether or not to perform correction based on a noise suppression rate estimated based on acoustic characteristics and sound data. The signal processing device according to claim 5 .
7. The determination unit The filter coefficients are determined using the trained model that receives acoustic characteristics, previously measured acoustic characteristics, and sound data as inputs and outputs a noise suppression rate. The signal processing device according to claim 1 .
8. The determination unit The filter coefficients are determined using the trained model that has been trained using the acoustic characteristics estimated at the position of the eardrum of the user and the noise reduction rate based on sound data as training data. The signal processing device according to claim 7 .
9. The determination unit The filter coefficients are determined using the trained model that receives as input a sound signal collected by a microphone different from the microphone used to measure the acoustic characteristics and sound data, and outputs correction filter coefficients that correct differences in filter coefficients based on environmental sounds in the user's environment. The signal processing device according to claim 1 .
10. The determination unit The filter coefficients are determined using the trained model, which has been trained using as training data filter coefficients that correct differences in filter coefficients based on acoustic characteristics estimated at the position of the user's eardrum. The signal processing device according to claim 9 .
11. The determination unit The filter coefficients are determined using the trained model that receives the acoustic characteristics and sound data as inputs and outputs the NC effect amount. The signal processing device according to claim 1 .
12. The determination unit The filter coefficients are determined using the trained model that has been trained using training data of an effect amount based on acoustic characteristics estimated at the position of the user's eardrum. The signal processing device according to claim 11 .
13. The determination unit The filter coefficients are determined using the trained model that receives as input the NC effect amount in an environment defined by a predetermined standard, sound data, and the acoustic characteristics of environmental sounds in the user environment, and outputs the NC effect amount in the user environment. The signal processing device according to claim 1 .
14. The determination unit The filter coefficients are determined using the trained model that has been trained using the sound data, the filter coefficients, and the NC effect amount based on the acoustic characteristics of the environmental sound in the user's environment as training data. The signal processing device according to claim 13 .
15. 1. A computer-implemented signal processing method comprising: an acquisition step of acquiring acoustic characteristics in the user's ear isolated from the outside world; an NC filter process for generating sound data that is in antiphase with the environmental sound that has leaked into the user's ear; a correction step of correcting the sound data using a correction filter; a determination process for determining filter coefficients of the correction filter to be used in the correction process in accordance with output information output by a trained model that outputs predetermined information related to the acoustic characteristics and the sound data when the acoustic characteristics and sound data are input, by inputting the sound data generated by the NC filter process and the acoustic characteristics acquired by the acquisition process; A signal processing method comprising:
16. an acquisition step for acquiring acoustic characteristics in the user's ear isolated from the outside world; an NC filter step for generating sound data in antiphase with the environmental sound leaking into the user's ear; a correction step of correcting the sound data using a correction filter; a determination step of determining filter coefficients of the correction filter used in the correction step according to output information output by a trained model that outputs predetermined information related to the acoustic characteristics and the sound data when the acoustic characteristics and sound data are input, by inputting the sound data generated by the NC filter step and the acoustic characteristics acquired by the acquisition step; A signal processing program that causes a computer to execute the above.
Citation Information
Patent Citations
Signal processing apparatus, sound apparatus, and signal processing method
JP2010259008A
Acoustic correction apparatus, and acoustic correction method
JP2011015080A
Signal processor, signal processing method and computer program
JP2016015585A
Earphone device, headphone device, and method
JP2019054337A
In-ear active noise reduction earphone
US9792893B1