Signal processing apparatus, method, and program
The signal processing device improves user satisfaction and audibility in noisy conditions by determining sound parameters for open-ear hearable devices, addressing the issue of external sound interference.
Patent Information
- Application Number
- PCT/JP2024/005515
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-16
- Publication Date
- 2025-08-21
AI Technical Summary
Open-ear hearable devices struggle to process external sounds effectively when they are loud, leading to reproduced sounds becoming inaudible.
A signal processing device that determines parameters for reproduced sound based on external sound information and user feedback to enhance user satisfaction beyond audibility, using spatial masking cancellation.
Enhances user satisfaction and maintains audibility of reproduced sounds in noisy environments by controlling sound output through unique parameter determination and processing.
Smart Images

Figure JP2024005515_21082025_PF_FP_ABST
Abstract
Description
Signal processing device, method and program
[0001] The disclosed technology relates to a technology for processing a signal.
[0002] It is thought that by listening to virtual sounds, which are sounds played back from a device, while listening to real sounds, which are external sounds, the satisfaction and understanding of real-world experiences can be improved.
[0003] To achieve this, for example, an open-ear hearable device that allows the user to hear external sounds naturally could be used. However, open-ear hearable devices have difficulty processing external sounds, and if the external sounds are loud, the reproduced sounds may become inaudible.
[0004] Therefore, a technique has been proposed that uses spatial masking cancellation to present an easy-to-listen dialogue while maintaining the playback level of external sounds (see, for example, Non-Patent Document 1).
[0005] NHK Science and Technology Research Laboratories, "Audio presentation technology that is easy for the elderly to hear," [online], May 2019, [Retrieved February 1, 2024], Internet <URL: https: / / www.nhk.or.jp / strl / publica / rd / 175 / 4.html>
[0006] However, the above technology only considers ease of listening.
[0007] The disclosed technology aims to provide a signal processing device, method, and program for controlling the playback sound output from a speaker of a wearable device using a method different from conventional methods.
[0008] A signal processing device, which is one aspect of the disclosed technology, includes a determination unit that determines parameters related to the reproduced sound so that an evaluation value related to user satisfaction other than ease of audibility for the reproduced sound in a situation where there is external sound, which is sound other than the reproduced sound output from the speaker of the wearable device that reproduces sound without completely blocking the external auditory canal, is high, and a processing unit that processes the input reproduced sound signal so that the reproduced sound output from the speaker of the wearable device becomes the reproduced sound determined by the parameters.
[0009] According to the disclosed technology, it is possible to control the reproduced sound output from the speaker of a wearable device using a method different from conventional methods.
[0010] FIG. 1 is a diagram illustrating an example of the functional configuration of a signal processing device. FIG. 2 is a diagram illustrating an example of a processing procedure of a signal processing method. FIG. 3 is a diagram illustrating an example of a wearable device. FIG. 4 is a diagram illustrating an example of processing by a determination unit. FIG. 5 is a diagram illustrating an example of processing by a determination unit. FIG. 6 is a diagram illustrating an experimental example. FIG. 7 is a diagram illustrating an experimental example. FIG. 8 is a diagram illustrating an experimental example. FIG. 9 is a diagram illustrating an experimental example. FIG. 10 is a diagram illustrating an experimental example. FIG. 11 is a diagram illustrating an experimental example. FIG. 12 is a diagram illustrating an experimental example. FIG. 13 is a diagram illustrating an experimental example. FIG. 14 is a diagram illustrating an example of the functional configuration of a computer.
[0011] Hereinafter, embodiments of the disclosed technology will be described with reference to the drawings. Note that components having the same functions in the drawings are given the same reference numerals, and redundant description will be omitted.
[0012] 1, the signal processing device includes, for example, an external sound information acquisition unit 1, a determination unit 2, and a processing unit 3. The signal processing unit may further include at least one of a speaker 4 and an input unit 5.
[0013] The signal processing method is realized, for example, by each component of the signal processing device performing the processes from step S1 to step S3 described below and shown in FIG.
[0014] Each component of the signal processing device will be described below.
[0015] <External Sound Information Acquisition Unit 1> External sound, which is sound other than the reproduced sound output from a speaker of a wearable device that reproduces sound without completely blocking the external ear canal, is input to the external sound information acquisition unit 1.
[0016] The external sound information acquisition unit 1 acquires external sound information, which is information about external sounds (step S1). The acquired external sound information is output to the determination unit 2.
[0017] Examples of external sound information include at least one of the sound pressure, direction, localization, and spatial distribution of the external sound. The direction of the external sound is, for example, the direction of the sound source of the external sound when the position of the user wearing the wearable device is taken as the origin. The direction of the external sound is at least one of the azimuth angle and the elevation angle.
[0018] The localization of an external sound is the direction and distance of the source of the external sound. The distance of the external sound is, for example, the distance between the source of the external sound and the user wearing the wearable device.
[0019] Examples of wearable devices that reproduce sound without completely blocking the ear canal include open-ear earphones, shoulder speakers, audio glasses, and smart glasses.
[0020] Note that "without completely blocking the external auditory canal" here means that the sound radiating unit is not placed in region R1, which corresponds to the axis passing through the entrance of the ear canal, as shown by the long dashed line in Figure 3. The sound radiating unit may be placed in peripheral region R2, such as the helix and earlobe, as shown by the dashed line in the same figure. An example of a device in which the sound radiating unit is placed in a peripheral region is audio glasses. Furthermore, the sound radiating unit may be placed in region R3, which is around the triangular fossa, antihelic crus, and navicular fossa, as shown by the dashed line in the same figure. An example of a device in which the sound radiating unit is placed in this region is open-ear earphones.
[0021] <Determining Unit 2> The determining unit 2 receives external sound information.
[0022] The determination unit 2 determines parameters related to the reproduced sound so as to increase an evaluation value related to user satisfaction other than audibility for the reproduced sound in a situation where external sound is present (step S2). The determined parameters are output to the processing unit 3.
[0023] Examples of parameters related to the reproduced sound are at least one of the sound pressure, direction, localization, and spatial distribution of the reproduced sound. The parameters related to the reproduced sound may include at least one of the sound pressure, direction, localization, and spatial distribution of the reproduced sound. The direction of the reproduced sound is the direction of the sound source position of the reproduced sound. The direction is at least one of the azimuth angle and the elevation angle. The localization of the reproduced sound is the direction and distance of the sound source position of the reproduced sound. The distance of the reproduced sound is, for example, the distance between the sound source position of the reproduced sound and the user wearing the wearable device.
[0024] An example of an evaluation value related to the user's satisfaction level other than audibility for reproduced sound in a situation where external sound is present is L2 below.
[0025] L2 = Satisfaction level (external sound information, user information, parameters related to the reproduced sound) In this L2 formula, "satisfaction level" is a function that takes external sound information, user information, and parameters related to the reproduced sound as arguments and outputs a satisfaction level based on the values of these arguments. Examples of user information are information that identifies the user and the user's location.
[0026] The evaluation value related to the user's satisfaction level other than the ease of listening may be L2 as follows.
[0027] L2 = Lack of discomfort (external sound information, user information, parameters related to the reproduced sound) In this L2 formula, "lack of discomfort" is a function that takes external sound information, user information, and parameters related to the reproduced sound as arguments and outputs the lack of discomfort based on the values of these arguments.
[0028] It is assumed that a table is predefined in which evaluation values related to user satisfaction other than audibility are associated with each set of values of at least one parameter related to external sound information, user information, and reproduced sound. In this case, the determination unit 2 selects, based on the predefined table, a set of values of at least one parameter related to external sound information, user information, and reproduced sound that results in the highest evaluation value related to user satisfaction other than audibility.
[0029] For example, assume that evaluation values (1 to 5) related to user satisfaction other than audibility are defined for each pair of values of the elevation angle of the playback sound, which is one of the parameters related to the playback sound, and information (A, B) identifying the user, which is one of the user information, as shown in Figure 4. The larger the evaluation value, the higher the user satisfaction. In this case, when user A is determined, the determination unit 2 determines "elevation angle of the playback sound = 90°" as the parameter related to the playback sound.
[0030] Assume also that evaluation values (1 to 5) related to user satisfaction other than audibility are defined for each pair of values of the elevation angle of the reproduced sound, which is one of the parameters related to the reproduced sound, and the user position (seat P1, seat P2), which is one of the user information, as shown in Fig. 5. In this case, when the user is sitting in seat P1, the determination unit 2 determines "elevation angle of the reproduced sound = 130°" as the parameter related to the reproduced sound.
[0031] As in these examples, when determining an evaluation value related to the user's satisfaction other than ease of audibility for the reproduced sound in a situation where external sound is present, if user information is required, the user information may be input to the determination unit 2.
[0032] Furthermore, as in these examples, when external sound information is not necessary to calculate an evaluation value related to user satisfaction other than audibility for reproduced sound in a situation where external sound is present, the external sound information does not need to be input to the determination unit 2. In this case, the signal processing device does not need to include the external sound information acquisition unit 1.
[0033] Note that the determination unit 2 may use information other than the external sound information, the user information, and the parameters related to the reproduced sound when calculating an evaluation value related to the user's satisfaction other than the ease of audibility with respect to the reproduced sound in a situation where external sound is present. In this case, information other than the external sound information, the user information, and the parameters related to the reproduced sound may be input to the determination unit 2. An example of information other than the external sound information, the user information, and the parameters related to the reproduced sound is information about the model of the wearable device.
[0034] The determination unit 2 may further determine parameters related to the reproduced sound so that an evaluation value indicating the ease of audibility of at least one of the external sound and the reproduced sound is increased.
[0035] An example of an evaluation value that indicates the ease of hearing at least one of the external sound and the reproduced sound is L1 below.
[0036] L1 = audibility (external sound information, user information, parameters related to the reproduced sound) In this L1 formula, "audibility" is a function that takes external sound information, user information, and parameters related to the reproduced sound as arguments and outputs the audibility based on the values of these arguments.
[0037] In this case, the determination unit 2 determines, for example, parameters related to external sound information, user information, and reproduced sound that maximize the output value when L1 and L2 are input into a predetermined non-decreasing function. An example of a non-decreasing function is λ1L1+λ2L2. λ1 and λ2 are coefficients that balance L1 and L2 and are predetermined positive real numbers.
[0038] <Processing Unit 3> The processing unit 3 receives as input the parameters determined by the determination unit 2. The processing unit 3 also receives as input a reproduction sound signal that is a signal of the reproduction sound to be reproduced.
[0039] The processing unit 3 processes the input playback sound signal so that the playback sound output from the speaker of the wearable device becomes the playback sound determined by the parameters (step S3).
[0040] The processing unit 3 generates a filter for converting the playback sound output from the speaker of the wearable device into a playback sound determined by parameters. In this case, the processing unit 3 processes the playback sound signal using the generated filter to generate a processed playback sound signal. The generated processed playback sound signal is output to the speaker 4 of the wearable device.
[0041] The speaker 4 of the wearable device generates sound based on the processed playback sound signal.
[0042] In this way, by determining parameters related to the reproduced sound so that the evaluation value related to user satisfaction other than ease of hearing for the reproduced sound in a situation where external noise is present is high, it is possible to control the reproduced sound output from the speaker of the wearable device using a method different from conventional methods.
[0043] [Modifications] The specific configurations of the embodiments of the disclosed technology are not limited to the configurations described above. The specific configurations of the embodiments of the disclosed technology can be appropriately modified in design, etc., within the scope of the spirit of the embodiments of the disclosed technology.
[0044] For example, the signal processing device may include an input unit 5 indicated by a dashed line in Fig. 1. The input unit 5 is, for example, an input device such as a keyboard, a mouse, or a touch panel. In this case, a user wearing the wearable device may be able to determine parameters related to the reproduced sound by operating the input unit 5. In other words, the determination unit 2 may be able to determine parameters input by the user as parameters related to the reproduced sound.
[0045] Furthermore, the determination unit 2 may determine parameters for the playback sound corresponding to the current time based on parameters for the playback sound determined in the past. For example, suppose that the user inputs information to the determination unit 2 that the sound pressure of the playback sound is high by operating the input unit 5. In this case, the determination unit 2 may determine parameters for the playback sound corresponding to the current time by lowering the value of "sound pressure of the playback sound" among the parameters for the playback sound determined one time ago by a predetermined value. In this way, by gradually changing the parameters for the playback sound determined in the past, the parameters for the playback sound corresponding to the current time can be determined, thereby reducing the processing load of the determination unit 2. This also makes it possible to make the change in the playback sound more gradual, thereby reducing the discomfort that the user may feel.
[0046] If the speakers of the wearable device are composed of a left speaker and a right speaker, the determination unit 2 may determine both parameters related to the reproduced sound corresponding to the left speaker and parameters related to the reproduced sound corresponding to the right speaker. The process of determining the parameters related to the reproduced sound corresponding to the left speaker and the parameters related to the reproduced sound corresponding to the right speaker is the same as the process of determining parameters related to the reproduced sound by the determination unit 2 described above. In this case, the determined parameters related to the reproduced sound corresponding to the left speaker and the determined parameters related to the reproduced sound corresponding to the right speaker are output to the processing unit 3.
[0047] In this case, the processing unit 3 receives the reproduced sound signal corresponding to the left and the reproduced sound signal corresponding to the right.
[0048] The processing unit 3 processes the input playback sound signal corresponding to the left so that the playback sound output from the left speaker of the wearable device becomes the playback sound determined by the playback sound parameters corresponding to the determined left speaker, and generates a processed left playback sound signal. The processed left playback sound signal is output to the left speaker of the wearable device. The left speaker of the wearable device generates sound based on the processed left playback sound signal.
[0049] The processing unit 3 processes the input right-side playback sound signal so that the playback sound output from the right speaker of the wearable device is determined by the playback sound parameters corresponding to the determined right speaker, thereby generating a processed right-side playback sound signal. The processed right-side playback sound signal is output to the right speaker of the wearable device. The right speaker of the wearable device generates sound based on the processed right-side playback sound signal.
[0050] The processing unit 3 can control the direction and positioning of the reproduced sound, for example, by using the difference between the sound pressure of the reproduced sound output from the left speaker of the wearable device and the sound pressure of the reproduced sound output from the right speaker of the wearable device, or by using a positioning filter based on the HRTF (Head Related Transfer Function).
[0051] The various processes described in the embodiments of the disclosed technology may not only be performed chronologically in the order described, but may also be performed in parallel or individually depending on the processing capacity of the device performing the processes or as needed.
[0052] For example, data may be exchanged directly between the components of the signal processing device, or may be exchanged via a storage unit (not shown).
[0053] Furthermore, a device (terminal) for using the device, system, or method of the present invention via a network (telecommunications line) may also be provided. The "device (terminal) for use" may be provided with functions (e.g., control function, decoding function, restoration function, input / output function, etc.) necessary to obtain the effects of implementing the device, system, or method of the present invention.
[0054] It goes without saying that other modifications are possible without departing from the spirit of the present invention.
[0055] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[0056] [Experimental example] In a situation where commentary audio is presented through earphones while listening to classical music (external sound), we conducted a subjective evaluation experiment to verify the relationship between each presentation condition and the impression that contributes to satisfactory listening.
[0057] The presentation conditions were 12 conditions in total, with the sound pressure level of the earphone playback sound being 75 dB or 69 dB for each of the six conditions shown in Table 1 of FIG.
[0058] A speaker for emitting external sound is placed as shown in Figure 7. The sound pressure level of the external sound is assumed to be 75 dB.
[0059] There were nine subjects. Each subject evaluated each of the 12 presentation conditions once. The evaluation items were the nine items shown in Figure 8. Each subject was asked to rate each item on a scale of 1 to 5.
[0060] The correlation between the evaluation values of each item is shown in Figure 9. From Figure 9, it can be seen that the correlation value between evaluation item 1, "Satisfaction with listening to music using commentary audio," and evaluation item 5, "Lack of discomfort with commentary audio," is 0.6, indicating a correlation.
[0061] Figure 10 shows the average value and variance of each evaluation item when the sound pressure level of the sound reproduced by earphones is 75 dB, and the average value and variance of each evaluation item when the sound pressure level of the sound reproduced by earphones is 65 dB. Figure 10 shows that the evaluation values for satisfaction, ease of audibility, and lack of discomfort vary significantly depending on the sound pressure level of the sound reproduced by earphones. These results demonstrate that the sound pressure level of the sound reproduced by earphones affects ease of audibility, satisfaction, and other factors.
[0062] Figures 11 to 13 show the evaluation scores for the ease of audibility and the lack of discomfort for subjects A, B, and C when the sound pressure level of the sound reproduced through earphones was 65 dB. From Figures 11 and 12, it can be seen that subjects A and B had higher evaluation scores when the HRFT elevation / depression angle was 90°. From Figure 13, it can be seen that subject C had a higher evaluation score when the HRFT elevation / depression angle was 135°. This shows that the elevation / depression angle of the sound reproduced through earphones affects the ease of audibility and the lack of discomfort.
[0063] [Program, Recording Medium] The functions realized by the components described in this specification may be implemented in circuitry or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), CPUs (Central Processing Units), conventional circuits, and / or combinations thereof, programmed to realize the described functions. A processor includes transistors and other circuits and is considered to be circuitry or processing circuitry. A processor may also be a programmed processor that executes a program stored in a memory.
[0064] In this specification, a circuitry, unit, or means is hardware that is programmed to realize or performs the described functions, which may be any hardware disclosed herein or any hardware known to be programmed to realize or perform the described functions.
[0065] If the hardware is a processor considered to be a type of circuitry, the circuitry, means, or unit is a combination of the hardware and software used to configure the hardware and / or processor.
[0066] The various processes described above can be implemented by loading a program that executes each step of the above method into the recording unit 2020 of the computer 2000 shown in Figure 14, and operating the control unit 2010, input unit 2030, output unit 2040, display unit 2050, etc.
[0067] The program describing the processing contents can be recorded on a computer-readable recording medium, which may be, for example, a magnetic recording device, an optical disk, a magneto-optical recording medium, a semiconductor memory, or any other suitable recording medium.
[0068] The program may be distributed by, for example, selling, transferring, lending, etc. portable recording media such as DVDs and CD-ROMs on which the program is recorded. Furthermore, the program may be stored in a storage device of a server computer, and then transferred from the server computer to other computers via a network, thereby distributing the program.
[0069] A computer that executes such a program may first temporarily store the program recorded on a portable recording medium or transferred from a server computer in its own storage device. Then, when executing a process, the computer reads the program stored on its own recording medium and executes the process in accordance with the read program. Alternatively, the computer may read the program directly from a portable recording medium and execute the process in accordance with the program. Furthermore, the computer may execute the process in accordance with the program each time a program is transferred from a server computer to the computer. Alternatively, the server computer may not transfer the program to the computer, but may instead execute the process through a so-called ASP (Application Service Provider) service, which realizes the processing function by issuing an execution instruction and obtaining the results. Furthermore, the server computer may execute the process at the terminal using a so-called SaaS (Software as a Service) service, which allows users to use part of a server computer along with the program. In this embodiment, the program includes information used for processing by an electronic computer that is equivalent to a program (such as data that is not a direct instruction to a computer but has properties that dictate computer processing).
[0070] Furthermore, in this embodiment, the device is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.
Claims
1. A signal processing device comprising: a determination unit that determines parameters related to a reproduced sound so that an evaluation value related to user satisfaction other than ease of audibility for a reproduced sound output from a speaker of a wearable device that reproduces sound without completely blocking the external auditory canal is increased in a situation where there is external sound other than the reproduced sound output from the speaker of the wearable device; and a processing unit that processes an input reproduced sound signal so that the reproduced sound output from the speaker of the wearable device becomes a reproduced sound determined by the parameters.
2. A signal processing device according to claim 1, wherein the determination unit further determines parameters relating to the reproduced sound so as to increase an evaluation value representing the ease of audibility of at least one of the external sound and the reproduced sound.
3. A signal processing device according to claim 1, wherein the determination unit determines parameters relating to the reproduced sound that maximize the output value when an evaluation value relating to the user's satisfaction with the reproduced sound in a situation where the external sound is present, other than audibility, and an evaluation value representing the audibility of at least one of the external sound and the reproduced sound, are input into a predetermined non-decreasing function.
4. A signal processing device according to any one of claims 1 to 3, wherein the parameters relating to the reproduced sound include at least one of the sound pressure, direction, localization and spatial distribution of the reproduced sound.
5. A signal processing method including: a step in which a determination unit determines parameters related to a reproduced sound so that an evaluation value related to a user's satisfaction other than ease of audibility for a reproduced sound in a situation where there is external sound other than the reproduced sound output from a speaker of a wearable device that reproduces sound without completely blocking the external auditory canal is increased; and a step in which a processing unit processes an input reproduced sound signal so that the reproduced sound output from the speaker of the wearable device becomes a reproduced sound determined by the parameters.
6. A program for causing a computer to execute each step of the signal processing method of claim 5.
Citation Information
Patent Citations
Neurofuzzy-based device for programmable hearing aids
JP2001520412A
How hearing aids work and hearing aids
JP2015501114A
Sound electronic circuit and method for adjusting sound level thereof
WO2006051586A1
Control processing device, control processing method, and program
WO2019131159A1