Audio processing device and program thereof
The audio processing device adjusts audio signals based on user position and hearing aid characteristics to enhance sound clarity for users with hearing impairments in 6DoF applications, addressing the challenge of acoustic mismatch in existing technologies.
Patent Information
- Application Number
- JP2024059651
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-02
- Publication Date
- 2025-10-15
AI Technical Summary
Existing 6DoF applications do not adequately consider the acoustic characteristics of hearing assistive devices, leading to challenges in making sound reproduction easier for users with hearing impairments, particularly when using binaural playback.
An audio processing device that adjusts acoustic characteristics of audio signals based on the user's position, listening direction, and the specific hearing aid device being used, applying hearing assistance characteristics to enhance sound clarity for users.
The solution ensures that sound reproduction is easier to hear for users with hearing aids by accounting for the acoustic characteristics of their devices, improving immersion and clarity in 6DoF applications.
Smart Images

Figure 2025156905000001_ABST
Abstract
Description
[Technical Field]
[0001] The present application relates to a speech processing device and a program therefor, and to a technology for adding acoustic characteristics of a hearing aid device in, for example, a 6DoF (Degrees of Freedom) application that allows a listener to set their listening position. [Background technology]
[0002] In recent years, audio technology supporting 6DoF applications (see Non-Patent Documents 6 and 7) has been developed by extending object-based audio systems (see Non-Patent Documents 3-5) that can manipulate audio content using audio signals and audio metadata (see Non-Patent Documents 1 and 2). 6DoF refers to the degree of freedom of object movement in six directions in three-dimensional space. A distinctive feature of 6DoF applications is that a listener can set an arbitrary position and orientation as the listening position and listening direction, and the content is simulated as viewed from the set listening position while facing the listening direction. In other words, 6DoF applications differ from conventional audio systems in that the reproduced sound changes through rendering according to the set listening position or listening direction.
[0003] While 6DoF applications have primarily been used in electronic games, their applications are expanding to include work experience and virtual travel. Therefore, system designs designed solely for use by hearing-impaired users are likely unable to fully meet diverse needs. For example, when users requiring hearing assistance use 6DoF applications, it is expected that the playback sound will be adapted to their hearing assistive devices. This is because users using hearing aids with directional sound collection (see Non-Patent Document 8) perceive a sound space where sound is collected in a directional manner, whereas hearing-impaired users without directional sound collection perceive an omnidirectional sound space. The directional sound collection function of hearing assistive devices selectively collects sounds arriving directly in front of the user, making them easier to hear. This sound collection characteristic is based on the typical human behavior of turning the head toward an object of interest. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] ITU-R BS.2076-1, Audio Definition Model, June 2017 [Non-patent document 2] ITU-R BS.2125-0, A serial representation of the Audio Definition Model, January 2019 [Non-patent document 3] ISO / IEC 23008-3:2019, Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3 3D audio, 2019 [Non-patent document 4] ETSI TS 103 190-2, Digital Audio Compression (AC-4) Standard; Part2: immersive and personalized audio, V1.2.1,2018-02 [Non-Patent Document 5] ATSC Standard: A / 342:2021 Part 3, MPEG-H System, 11 March 2021 [Non-patent document 6] MPEG-I Immersive Audio Encoder Input Format, Version 5, April 4, 2023 [Non-Patent Document 7] Report ITU-R BT.2420-5 (09 / 2022), Collection of usage scenarios of advanced immersive sensory media systems (09 / 2022) [Non-patent document 8] Theodore H. Benema, translated by Nakagawa Tatsuo, Introduction to Digital Hearing Aids for Clinicians, Kaibundo Publishing Co., Ltd., September 2008 Summary of the Invention [Problem to be solved by the invention]
[0005] In 6DoF applications, two-channel audio is often provided using binaural playback. Binaural playback is a method of processing the audio signals of audio objects so that sound is heard in a specific direction. The number of audio objects that make up a given piece of content is not necessarily limited to one, but is usually multiple. However, with binaural playback, regardless of the number of audio objects, the components of each audio object are aggregated into two-channel audio signals, but the impact of the acoustic characteristics of hearing aids, such as directional sound collection functions, is not taken into consideration.
[0006] The embodiments of the present application have been made to solve the above-mentioned problems, and one of the objects of the present application is to make the sound reproduced by rendering easier to hear for a user of a hearing assistive device. [Means for solving the problem]
[0007] [1] One aspect of this embodiment is an audio processing device that includes an audio support characteristic setting unit that sets audio support characteristics, which are acoustic characteristics of an audio support device, based on audio support-related information regarding audio support; an audio support characteristic application unit that applies the audio support characteristics to an audio signal; and a rendering unit that adjusts the acoustic characteristics of the audio signal to which the audio support characteristics have been applied, for each audio object, in accordance with a target position, a listening position, and a listening direction. According to the configuration in [1], the acoustic characteristics of the hearing aid device are applied to the audio signal, and the acoustic characteristics are adjusted according to the target position of each audio object and the user's listening position and listening direction. By taking into account the influence of the hearing aid device on the acoustic characteristics, a reproduced sound that is as easy to hear as when using the hearing aid device is realized.
[0008] [2] One aspect of this embodiment is the above-mentioned audio processing device, wherein the hearing assistance characteristic setting unit determines the sensitivity of the hearing assistance device based on the directionality of the sound collection function of the hearing assistance device and the direction of the target position relative to the listening position, and sets the hearing assistance characteristic including the sensitivity. According to the configuration in [2], the directivity of the sound collection function of the hearing assist device according to the target position of the audio object is reflected in the reproduced sound, resulting in reproduced sound that is as easy to hear as when using the sound collection function of the hearing assist device.
[0009] [3] One aspect of this embodiment is the above-mentioned audio processing device, wherein the hearing assistance characteristic setting unit sets the directionality of the sound collection function for each model of the hearing assistance device, and the directionality of the model indicated by the hearing assistance-related information may be specified. Generally, the influence of the directionality of a sound collection function varies depending on the model of hearing assist device, but with this configuration, the directionality of the specified device is reflected in the reproduced sound, resulting in reproduced sound that is as easy to hear as when using the sound collection function of the specified model of hearing assist device.
[0010] [4] One aspect of this embodiment is the above-mentioned audio processing device, wherein the hearing assistance characteristic setting unit may determine whether or not the hearing assistance characteristic needs to be set based on hearing assistance necessity information included in the hearing assistance-related information. According to the configuration of [4], whether or not to include the influence of the hearing assist device on the acoustic characteristics is determined depending on whether or not hearing assistance is required. Therefore, whether or not to provide the same reproduced sound as when using the sound collection function of the hearing assist device can be easily changed depending on whether or not hearing assistance is required.
[0011] [5] One aspect of this embodiment is the above-mentioned audio processing device, wherein the rendering unit may set a smaller gain for a rear object, which is an audio object whose target direction, which is the direction of the target position based on the listening position, is closer to the front direction than to the rear direction, than for a front object, which is an audio object whose target direction is closer to the rear direction than to the front direction. According to the configuration of [5], the level of the component of the front object in the reproduced sound is relatively higher than the level of the component of the rear object, so that the sound of the front object in the reproduced sound is heard more clearly than the sound of the rear object.
[0012] [6] One aspect of this embodiment is the above-mentioned audio processing device, wherein the rendering unit does not need to set a smaller gain for an audio object that is provided in priority over other audio objects than for the forward object when the target direction is closer to the rearward direction than to the forward direction. According to the configuration of [6], even if the target direction of a given audio object is closer to the rear, a gain smaller than that of a forward object is not set for that audio object. The level of the audio object component in the reproduced sound is maintained without being relatively lower than the level of the forward object component. Therefore, the sound of that object is heard without losing clarity.
[0013] [7] One aspect of this embodiment may be a program for causing a computer to function as the above-described voice processing device. According to the configuration in [7], the acoustic characteristics of the hearing aid device are applied to the audio signal, and the acoustic characteristics are adjusted according to the target position, listening position, and listening direction for each audio object. By taking into account the influence of the hearing aid device on the acoustic characteristics, a reproduced sound that is as easy to hear as when using the hearing aid device is realized. [Effects of the Invention]
[0014] According to this embodiment, the sound reproduced by rendering can be provided to a user of a hearing assist device in a way that is easy to hear. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a schematic block diagram illustrating an example of an overview of a voice processing system according to an embodiment of the present invention. [Figure 2] 1 is a schematic block diagram illustrating an example of a functional configuration of a voice processing device according to an embodiment of the present invention. [Figure 3] 4 is a schematic block diagram illustrating an example of the functional configuration of a weighting coefficient calculation unit according to the present embodiment. FIG. [Figure 4] FIG. 1 is a schematic diagram illustrating a 6DoF content space. [Figure 5] 10 is a flowchart illustrating an example of audio processing according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0016] First, an overview of an embodiment of the present application will be described with reference to the drawings. FIG. 1 is a schematic block diagram illustrating an overview of an audio processing system S1 according to this embodiment. The audio processing system S1 includes an audio processing device 10 and a playback device 20. The audio processing system S1 is an example of an object-based audio system. In the following description, an example will be given in which the audio processing system S1 is capable of playing back 6DoF audio content in a three-dimensional 6DoF content space.
[0017] The audio processing device 10 functions as an audio rendering device that simulates sounds emitted from audio objects in a content space and heard by a user. The transmission characteristics of sound waves from an audio object to a user are determined by the target position of the audio object in the content space and the user's listening position. The target position of the audio object is indicated by audio object position information. The user's listening position and listening direction are indicated by user position information. The audio processing device 10 adjusts the acoustic characteristics of the audio signal of the audio object according to the target position, listening position and listening direction, in accordance with the capabilities of the playback device 20, and converts the signal into a playback signal for the audio object.
[0018] The voice processing device 10 according to this embodiment further sets hearing assistance-related information. The hearing assistance-related information may include information regarding the need for hearing assistance. If hearing assistance is required, the voice processing device 10 applies the acoustic characteristics of the hearing assistance device (sometimes referred to herein as "hearing assistance characteristics") to the audio signal before performing rendering. In rendering, the acoustic characteristics of the audio signal are adjusted according to the target direction of the audio object and the user's listening position and listening direction. A typical hearing assistance device is a hearing aid. A hearing assistance device may also be an electronic device that does not primarily provide hearing assistance but primarily performs other functions, such as a multi-function mobile phone (a so-called smartphone), a voice guidance device, a receiver, or a sound collector.
[0019] When hearing aid is not required, the audio processing device 10 performs rendering on the audio signal of the audio object without setting the hearing aid characteristics, and adjusts the acoustic characteristics. The audio processing device 10 outputs to the playback device 20 a playback signal that includes, as a component, an audio signal whose acoustic characteristics have been adjusted for each audio object, regardless of whether or not the setting of hearing assistance characteristics is required.
[0020] The playback device 20 plays back sound based on a playback signal input from the audio processing device 10. The playback device 20 is an acoustic device that has two playback sound sources and is attached or worn by the user with each playback sound source in contact with or close to each ear. The playback device 20 may be any of two-channel headphones, two-channel earphones, a wearable speaker, a neck speaker, a bone conduction receiver, etc.
[0021] When a user requires hearing assistance, the acoustic characteristics of the playback signal include the hearing assistance characteristics in addition to the acoustic characteristics corresponding to the target position, listening position, and listening direction of the audio object. Therefore, the sound of the audio object that the user focuses on is presented as a sound with acoustic characteristics similar to that presented using the hearing assistance device. Because the playback device 20 presents a sound with characteristics familiar to the user through the hearing assistance device, the sound of the audio object is easier to hear, improving the sense of immersion. When the user does not require hearing assistance, the audio processing device 10 causes the playback device 20 to present sounds having acoustic characteristics according to the target position, listening position, and listening direction, without including hearing assistance characteristics.
[0022] The number of audio objects provided in audio content may be one or more. The audio processing device 10 may vary the target position of each of two or more audio objects, or may vary the target position of some audio objects and fix the target positions of the remaining audio objects to predetermined positions. The audio processing device 10 may vary the listening direction of an audio object with a variable target position, such as the direction in front of the user. In the audio processing device 10, audio objects whose audio object position information can be variably set are limited to essential audio objects, and changes to the audio object position information of other audio objects may not be permitted. Essential audio objects are audio objects whose sound is played with priority over other audio objects. Essential audio objects are also called priority audio objects. Examples of essential audio objects include the dialogue of actors in movies and dramas, and commentary in sports broadcasts and news programs. Essential audio objects are identified using an essential flag included in the metadata of the audio content.
[0023] Next, a description will be given of input information to the voice processing device 10 according to this embodiment. As illustrated in Fig. 2, user position information, hearing assistance related information, voice object position information, and a voice signal are input to the voice processing device 10.
[0024] The user position information indicates the user's position in the content space. The user position information may include information indicating the user's orientation. The user's orientation is, for example, represented by the forward direction of the user's head. In this application, the user's position may be referred to as the "listening position," and the user's orientation may be referred to as the "listening direction." Real position information or virtual position information may be acquired as the user position information. The audio content may specify whether the real position information or the virtual position information is to be used. The audio processing device 10 can acquire real user position information of the user to be detected from an input device equipped with a position sensor, such as a head-mounted display (HMD) or a human tracker. The audio processing device 10 may acquire virtual user position information indicating a specified position or orientation from an input device that can specify a position in response to a received operation, such as a mouse, a touch sensor, or a joystick.
[0025] Hearing assistance-related information includes information indicating whether a user needs hearing assistance (sometimes referred to as "hearing assistance necessity information" in this application). Hearing assistance-related information is acquired, for example, from an input device. Hearing assistance device wearing information indicates whether a user is wearing a hearing assistance device. Information indicating whether a hearing assistance device is being worn can also be considered one form of hearing assistance necessity information. Information indicating whether a hearing assistance device is being worn may be represented using a wearing flag. The wearing flag may be expressed as a one-bit value indicating whether a hearing assistance device is being worn. For example, a wearing flag value of 1 indicates that a hearing assistance device is being worn, and a wearing flag value of 0 indicates that a hearing assistance device is not being worn.
[0026] The hearing aid-related information may include hearing aid device information about the hearing aid device worn by the user. The hearing aid device information may include model information indicating the model of the hearing aid device. The model information may be expressed, for example, by the model number of the hearing aid device. The model information may also be expressed by being included in identification information (for example, a serial number) for identifying each device. The weighting coefficient calculation unit 110 (described below) of the voice processing device 10 may set hearing aid device wearing information indicating whether or not a hearing aid device is being worn, depending on whether a wired or wireless connection with an input device is detected. The weighting coefficient calculation unit 110 may receive hearing aid device information from the connected hearing aid device during the connection process with the hearing aid device or after the connection is established.
[0027] The model information may include sound collection function information, or may not include sound collection function information but may be associated and pre-set in the weighting coefficient calculation unit 110. The weighting coefficient calculation unit 110 can identify sound collection function information corresponding to the acquired model information. The sound collection function information is information indicating the characteristics of the sound collection function of the hearing assist device. The sound collection function information includes, for example, directivity information. The directivity information may be expressed as sensitivity for each sound source direction relative to the sound collector, or may express a typical directivity distinction, such as omnidirectional, bidirectional, or unidirectional. The directivity distinction may be expressed as a value. For directivities such as bidirectional and unidirectional, which have higher sensitivity in a specific direction than in other directions, information on the direction of the main axis is included.
[0028] The hearing aid-related information may include hearing aid characteristic acquisition procedure information related to the selection of a procedure for acquiring a hearing aid characteristic. The hearing aid characteristic acquisition procedure information may be expressed, for example, using a procedure selection flag. The procedure selection flag may be expressed as a one-bit value indicating whether a weighting coefficient table or a weighting coefficient algorithm is to be used as the procedure, i.e., processing method, for acquiring the hearing aid characteristic. For example, a procedure selection flag value of 1 indicates a weighting coefficient algorithm, and a procedure selection flag value of 0 indicates a weighting coefficient table. The hearing aid characteristics are expressed using weighting coefficients for each of multiple frequency bands. The multiple frequency bands cover the audible range, including the frequency band (typically 100 Hz to 4 kHz) that comprises the main components of human speech. The multiple frequency bands may be, for example, three frequency bands divided into a low range (typically 250 Hz or less), a mid range (higher than 250 Hz and lower than 2.5 kHz), and a high range (2.5 kHz or more), or may be frequency bands divided into equal widths in the logarithmic domain, such as 1 / 3 octave bands. The weighting coefficients correspond to the gains relative to the amplitude of the components of each frequency band of the collected audio signal.
[0029] The audio object position information is information indicating the position of an audio object in a content space. The audio object position information may be provided in association with the audio signal of the audio object in the audio content, or may be acquired independently by the audio processing device 10. Required audio objects may be indicated in the audio content. Required audio objects may be indicated, for example, using a required flag. The required flag may be provided as part of metadata included in the audio content. The audio processing device 10 may independently acquire audio object position information for required audio objects, but may not need to acquire audio object position information for other audio objects included in the audio content. The audio processing device 10 may acquire audio object position information indicating a specified position from, for example, an input device that allows a position to be specified in response to a received operation.
[0030] An audio signal is a signal that indicates the waveform of the audio emitted from an audio object. The audio signal is provided for each audio object as part of the audio content. The audio content is composed of the audio signal and metadata for each audio object. The metadata may or may not include one or both of the audio object position information and the required flag.
[0031] Next, an example of the functional configuration of the audio processing device 10 will be described. Fig. 2 is a schematic block diagram showing an example of the functional configuration of the audio processing device 10 according to this embodiment. However, Fig. 2 shows an example in which the hearing assistance characteristics of a hearing assistance device worn by a user are expressed using weighting coefficients for each passband. The hearing assistance characteristics correspond to the amount of adjustment of the frequency characteristics for a collected sound signal. The audio processing device 10 includes a weighting coefficient calculation unit 110, a weighting coefficient application unit 120, and a binaural rendering unit .
[0032] The weighting coefficient calculation unit 110 calculates the direction of the audio object position indicated in the audio object position information as the audio object direction in the content space, using the user position indicated in the user position information as a reference. The weighting coefficient calculation unit 110 calculates the direction from the user position to the audio object position for each audio object, using the listening direction indicated in the user position information as a reference.
[0033] The weighting coefficient calculation unit 110 acquires a wearing flag that constitutes the hearing assistance-related information and determines whether or not a hearing assistance device is being worn based on the wearing flag. When determining that a hearing assistance device is being worn, the weighting coefficient calculation unit 110 acquires model information of the hearing assistance device that constitutes the hearing assistance-related information. The weighting coefficient calculation unit 110 identifies sound collection function information corresponding to the acquired model information and identifies the directivity of the sound collection function of the hearing assistance device indicated in the identified sound collection function information. Based on the identified directivity and the audio object direction, the weighting coefficient calculation unit 110 refers to the sound collection function information and calculates a weighting coefficient representing the hearing assistance characteristics in accordance with the hearing assistance characteristic acquisition method indicated in the hearing assistance characteristic acquisition procedure information included in the hearing assistance-related information. The weighting coefficient calculation unit 110 outputs the weighting coefficient calculated for each audio object to the weighting coefficient application unit 120.
[0034] It should be noted that if no hearing assistance related information is input to the weighting coefficient calculation unit 110, or if the hearing assistance related information is input but the wearing flag indicates that no hearing assistance device is being worn, the weighting coefficient calculation unit 110 does not calculate a weighting coefficient. In this case, no weighting coefficient is output from the weighting coefficient calculation unit 110, and the weighting coefficient application unit 120 outputs the audio signal input for each audio object directly to the binaural rendering unit 130 as an object-specific audio signal.
[0035] The weighting coefficient application unit 120 applies the weighting coefficients input from the weighting coefficient calculation unit to the audio signal for each audio object. By applying the weighting coefficients, the frequency characteristics of the hearing assist device are added to the audio signal, resulting in an audio signal with adjusted frequency characteristics. The weighting coefficient application unit 120 performs filtering processing on the input audio signal using different passbands, converting it into pass signals with different frequency bands. For each frequency band, the weighting coefficient application unit 120 multiplies the amplitude of the pass signal by the weighting coefficient corresponding to that frequency band to generate an amplitude-adjusted pass signal. The weighting coefficient application unit 120 outputs the weighted coefficient-applied audio signal, obtained by adding the amplitude-adjusted pass signals between frequency bands, as an object-specific audio signal to the binaural rendering unit 130 for each audio object.
[0036] The binaural rendering unit 130 converts the object-specific audio signals input from the weighting coefficient application unit 120 into object-specific binaural signals, which are audio signals indicating audio waveforms arriving at the left and right ears of the user for each audio object, based on audio object position information and user position information. Here, a head-related transfer function, which indicates the transfer characteristics of sound waves to the left and right ears of the user located at the target position indicated by the audio object and the listening position indicated by the user position information, is applied to the object-specific audio signals input to the binaural rendering unit 130, and audio signals for each set consisting of an audio object and a channel (sometimes referred to as "object-specific channel signals" herein) are obtained for each channel corresponding to the left and right ears (sometimes referred to as "left channel" or "right channel" herein). The object-specific binaural signals include the object-specific channel signals for each of the left and right channels.
[0037] The binaural rendering unit 130 adjusts the amplitude of the object-specific channel signal for each audio object for each of the left and right channels by multiplying the signal value of the object-specific channel signal by a predetermined gain. The binaural rendering unit 130 generates a left-channel, two-channel binaural signal by adding (i.e., mixing) the object-specific channel signals after amplitude adjustment for each of the left and right channels between audio objects. The binaural rendering unit 130 outputs the generated binaural signal to the playback device 20 as a playback signal.
[0038] The binaural rendering unit 130 can derive head-related transfer functions corresponding to the target position, listening position, and listening direction for each audio object using, for example, the following method. Head-related transfer function data indicating head-related transfer functions corresponding to the left and right ears for each audio object direction is set in advance in the binaural rendering unit 130. The binaural rendering unit 130 identifies the direction from the listening position to the target position based on the listening direction (i.e., the front direction of the user) as the audio object direction, and determines the head-related transfer function corresponding to the identified audio object direction by referring to the preset head-related transfer function data. The listening direction is specified by the user position information. A volume model indicating the relationship between the sound source distance and the volume attenuation rate is also set in advance in the binaural rendering unit 130. For example, the inverse square law is used as the volume model for audio objects that can be approximated as a point sound source in the content space. The inverse proportional law is used for audio objects that can be approximated as a line sound source. The binaural rendering unit 130 specifies the distance from the listening position to the target position as the sound source distance, and calculates an attenuation factor corresponding to the specified sound source distance using a volume model preset for each sound object.The binaural rendering unit 130 then multiplies the head-related transfer functions corresponding to the sound object directions for each of the left and right ears by the calculated attenuation factor to calculate head-related transfer functions corresponding to the target position, listening position, and listening direction.
[0039] Prior to applying the head-related transfer functions, the binaural rendering unit 130 may render each audio object to a group of virtual speakers in the content space (e.g., 22.2ch audio, 11.2ch audio, 5.1ch audio, etc.), and then apply head-related transfer functions based on the directions of each virtual speaker. That is, in an audio reproduction system including multiple virtual speakers, the binaural rendering unit 130 pre-sets, for each of the left and right ears, head-related transfer functions corresponding to directions from a predetermined listening position based on a predetermined listening direction to positions where each virtual speaker is located (hereinafter referred to as "virtual speaker directions"). Then, for the audio object directions identified for each audio object, the binaural rendering unit 130 identifies, for each of the left and right ears, head-related transfer functions corresponding to the most approximate virtual speaker directions among multiple pre-set virtual speaker directions, for example. Alternatively, the binaural rendering unit 130 may specify two or more adjacent virtual speaker directions sandwiching an audio object direction specified for each audio object, and may calculate a head-related transfer function corresponding to the audio object direction by interpolating the head-related transfer functions corresponding to each specified virtual speaker direction. In the interpolation of the head-related transfer functions, a panning law such as the sine law or the tangent law may be applied.
[0040] (Weighting coefficient calculation section) Next, an example of the functional configuration of the weighting coefficient calculation unit 110 according to this embodiment will be described. Fig. 3 is a schematic block diagram showing an example of the functional configuration of the weighting coefficient calculation unit 110 according to this embodiment. The weighting coefficient calculation unit 110 includes an audio object direction calculation unit 112 and a weighting coefficient determination unit 114. A weighting coefficient table WT and a weighting coefficient calculation algorithm WA are set in advance in the weighting coefficient calculation unit 110. The weighting coefficient table WT is a data table showing weighting coefficients for each frequency band for each sound source direction. The weighting coefficient calculation algorithm WA is a mathematical model for calculating as output the weighting coefficients for each frequency band corresponding to the sound source direction as input. A specific example of the weighting coefficient calculation algorithm WA will be described later.
[0041] The audio object direction calculation unit 112 uses the listening position indicated in the user position information as a reference and calculates, for each audio object, the direction of the target position indicated in the audio object position information as the audio object direction. The audio object direction may be expressed as an angle relative to the listening direction indicated in the user position information. The audio object direction calculation unit 112 outputs the calculated audio object direction to the weighting coefficient determination unit 114.
[0042] The weighting coefficient determination unit 114 determines whether the hearing assistance related information includes a procedure selection flag. If the procedure selection flag is included, the weighting coefficient determination unit 114 selects whether to use the weighting coefficient table WT or the weighting coefficient calculation algorithm WA when determining the weighting coefficients in accordance with the procedure selection flag. If the procedure selection flag is not included, the weighting coefficient determination unit 114 selects a predetermined one of the weighting coefficient table WT and the weighting coefficient calculation algorithm WA.
[0043] When selecting the weighting coefficient table WT, the weighting coefficient determination unit 114 reads out the weighting coefficients for each frequency band corresponding to the calculated voice object direction from the weighting coefficient table WT. When the weighting coefficient calculation algorithm WA is selected, the weighting coefficient determination unit 114 executes the following process to calculate the weighting coefficient for each frequency band: The weighting coefficient calculation algorithm WA will be described with reference to FIG.
[0044] FIG. 4 is a schematic diagram illustrating a 6DoF content space. In the following description, a three-dimensional global coordinate system with an origin O as the reference point is used. In FIG. 4, u, f u ,s i are vectors indicating the user's listening position, the listening direction from the listening position u, i.e., a position in front of the user at a reference distance r0, and the target direction of the i-th (i is a natural number indicating the audio object) audio object, respectively. Here, r0 = |f u-u|. The position of the ith audio object relative to the listening position is s i Therefore, the angle ξ between the listening direction and the audio object i is i is the unit vector f u -u and vector s indicating the position of audio object i relative to the listening position i -u and its absolute value |f u -u|. However, in equation (1), normalization is performed as r0=1. The audio object direction calculation unit 112 calculates the distance between the listening position u and the position s of the audio object i using equation (1). i Based on the angle ξ i can be calculated as the audio object direction related to audio object i.
[0045]
number
[0046] The weighting coefficient determination unit 114 determines whether the hearing assistance-related information includes directivity information. The directivity information is information indicating the directivity of the sound collection function of the hearing assistance device. If directivity information is present, the weighting coefficient determination unit 114 adopts the directivity indicated in the directivity information. The weighting coefficient determination unit 114 determines whether the hearing assistance-related information includes model number information. The model number information is information indicating the model number of the hearing assistance device. The model number information is an example of model information indicating the model of the hearing assistance device. The weighting coefficient determination unit 114 is configured to preset device data indicating directivity information corresponding to each model number information. If model number information is present, the weighting coefficient determination unit 114 may refer to the device data and adopt the directivity indicated in the directivity information corresponding to the model number information.
[0047] The directivity may be expressed by an index indicating the type of directivity, such as omnidirectional, unidirectional, or bidirectional, or may be expressed by the sensitivity for each sound source direction relative to the listening position. i is p(ξi )s i and sound s provided by sound object i i The directivity of the sound collection function p(ξ i The weighting coefficient determination unit 114 determines the directivity p(ξ i ) within the frequency band to calculate the weighting coefficient. The weighting coefficient determination unit 114 outputs the weighting coefficient calculated for each frequency band to the weighting coefficient application unit 120.
[0048] For example, if the directivity of the sound collection function is omnidirectional, the directivity p o (ξ i ) is the audio object direction ξ i This is typically the case when sound is collected by a hearing aid equipped with an omnidirectional microphone. In this case, the component s^ of the i-th speech object in the collected signal is i O is, s i However, the directivity p(ξ i The sensitivity of the sound collection function is normalized so that ) is 1.
[0049] If the directionality of the sound collection function is bidirectional, the directionality p b (ξ i ) is cosξ i This is typically the case when sound is collected by a hearing aid equipped with a bidirectional microphone. In this case, the component s^ of the i-th speech object in the collected signal is proportional to i b is cosξ i s i However, the directivity p b (ξ i ) is normalized so that the maximum value of the directivity p, i.e., the sensitivity in the direction of the principal axis where the directivity is the highest, is 1. b (ξ i ) being a negative value corresponds to the phase of the audio signal output from the sound collection function being inverted.
[0050] If the directivity of the sound collection function is unidirectional, the directivity p c (ξi ) is (1+cosξ i ) / 2. This is typically the case when sound is collected by a hearing aid equipped with a unidirectional microphone. In this case, the component s^ of the i-th speech object in the collected signal is i c is (1+cosξ i )s i / 2. However, the directivity p b (ξ i The sensitivity of the sound collection function is normalized so that the maximum value of (1), i.e., the sensitivity in the principal axis direction where the sensitivity is highest, is 1.
[0051] Note that, although the above description exemplifies a case in which the weighting coefficient indicates dependency on the directionality of the sound collection function, this is not limiting. The weighting coefficient calculation unit 110 may further add an adjustment amount according to the user's hearing ability to the weighting coefficient for each frequency band obtained by the above process. Furthermore, the weighting coefficient calculation unit 110 may scale the weighting coefficient for each frequency band so that the intensity of the frequency component of the hearing assistance signal for that frequency band falls within the range between the maximum audible level, which is the maximum intensity that the user can hear, and the minimum audible level, which is the minimum intensity. Here, the frequency component of the hearing assistance signal corresponds to the sum of the multiplied values between the audio objects obtained by multiplying the frequency component of the object-specific audio signal by the weighting coefficient for that frequency band.
[0052] Next, an example of audio processing according to this embodiment will be described below with reference to a flowchart of FIG. (Step S102) The weighting coefficient calculation unit 110 acquires user position information indicating the listening position of the user. (Step S104) The weighting coefficient calculation unit 110 acquires voice object position information related to the voice object. (Step S106) The weighting coefficient application unit 120 acquires an audio signal related to an audio object. (Step S108) The weighting coefficient calculation unit 110 acquires hearing assistance information related to the user's use of the hearing assistance device.
[0053] (Step S110) The weighting coefficient calculation unit 110 determines whether the hearing assistance function is required by referring to the wearing flag of the hearing assistance device. If it is determined that the hearing assistance function is required (YES in step S110), the process proceeds to step S112. If it is determined that the hearing assistance function is not required (NO in step S110), the weighting coefficient application unit 120 does not apply a weighting coefficient to the audio signal for each audio object, but outputs the audio signal to the binaural rendering unit 130 as an object-specific audio signal. Then, the process proceeds to step S124. (Step S112) The audio object direction calculation unit 112 calculates the audio object direction as the direction from the listening position indicated in the user position information to the target position indicated in the audio object position information for each audio object, with the listening direction as the reference.
[0054] (Step S114) The weighting coefficient determination unit 114 determines the procedure, i.e., the processing method, for calculating the weighting coefficients by referring to the procedure selection flag. If it is determined that the weighting coefficient calculation algorithm WA is to be used (YES in step S114), the process proceeds to step S116. If it is determined that the weighting coefficient table WT is to be used (NO in step S114), the process proceeds to step S118. (Step S116) The weighting coefficient determination unit 114 reads out the weighting coefficients for each frequency band corresponding to the calculated voice object direction from the weighting coefficient table WT. (Step S118) The weighting coefficient determination unit 114 refers to the directionality of the sound collection function of the hearing assistive device being worn, and calculates weighting coefficients for each frequency band based on the sensitivity corresponding to the calculated sound object direction. (Step S120) The weighting coefficient determination unit 114 outputs the weighting coefficients for each frequency band determined for each voice object to the weighting coefficient application unit 120.
[0055] (Step S122) The weighting coefficient application unit 120 multiplies the frequency band components of the audio signal for each audio object by the weighting coefficient for that frequency band, and mixes the resulting multiplied values between frequency bands to obtain object-specific auditory hearing assistance audio signals. The weighting coefficient application unit 120 outputs the object-specific auditory hearing assistance audio signals for each audio object to the binaural rendering unit 130 as object-specific audio signals.
[0056] (Step S124) The binaural rendering unit 130 applies a head-related transfer function from the target position to each of the left and right ears of a user who is located at the listening position and has the front of his or her head facing the listening direction to the object-specific audio signal for each audio object, thereby generating object-specific channel signals for each set of audio object and each of the left and right ears. The binaural rendering unit 130 mixes the object-specific channel signals between audio objects and synthesizes the channel signals for each of the left and right ears. (Step S126) The binaural rendering unit 130 outputs the binaural signal consisting of the channel signals for the left and right ears as a playback signal to the playback device 20. After that, the processing in FIG. 5 ends.
[0057] The binaural rendering unit 130 may reset the gain for an audio object located behind or to the side of the user (hereinafter may be referred to as a "rear object") so that the level of the object-specific binaural signal is relatively smaller than that for an audio object located in front of or to the side of the user (hereinafter may be referred to as a "forward object"). Here, the binaural rendering unit 130 may (1) lower the level of the object-specific binaural signal for the rear object while maintaining the level of the object-specific binaural signal for the front object unchanged, (2) increase the level of the object-specific binaural signal for the front object and maintain the level of the object-specific binaural signal for the rear object unchanged, or (3) increase the level of the object-specific binaural signal for the front object and lower the level of the object-specific binaural signal for the rear object.
[0058] Here, the binaural rendering unit 130 can determine, as a forward object, an audio object whose target direction based on the listening position is closer to the front direction of the user (i.e., the listening direction) than to the back direction of the user. The binaural rendering unit 130 can determine, as a rearward object, an audio object whose target direction based on the listening position is closer to the back direction of the user than to the front direction of the user.
[0059] For audio objects whose gains have been reset, the binaural rendering unit 130 multiplies the signal value of the object-specific binaural signal by the reset gain instead of the predetermined gain, thereby synthesizing the playback signal to include the object-specific binaural signal whose amplitude has been adjusted. Therefore, the components of the front object are relatively more numerous than the components of the rear object in the playback sound presented based on the playback signal by the playback device 20. This makes the sound of the front object relatively easier for the user to hear than the sound of the rear object.
[0060] In addition, the binaural rendering unit 130 does not need to reset the gain so that the level of the binaural signal for each object is relatively low, regardless of whether the required object corresponds to a rear object located behind the user or to the side of the user. For the binaural signal for each required object for which audio or rendering function is to be provided preferentially, the gain is not reduced even if the required object is located behind the user. Therefore, the audibility of the required object component in the reproduced sound is maintained.
[0061] The binaural rendering unit 130 may display a content space screen indicating the listening position, listening direction, and audio object position for each audio object in the 6DoF space on the display. The audio processing device 10 may specify one or a combination of the listening position, listening direction, and audio object position based on an operation signal input from an input device. In addition, although the above description has been given with reference to an example in which the present invention is applied to rendering of audio objects in a 6DoF content space, the present invention is not limited to this. The present invention may also be applied to rendering of audio objects in a two-dimensional 4DoF content space. In this case, the height direction or elevation angle direction in the three-dimensional space may be omitted.
[0062] The weighting coefficient calculation unit 110 may display a setting screen for setting the hearing assistance-related information on a display (not shown). The weighting coefficient calculation unit 110 may set element information that is an element of the hearing assistance-related information based on an operation signal input from an input device. Settable element information includes whether or not a hearing assistance device is being worn (corresponding to a wearing flag), model number information, or some or all of the type of directivity of the sound collection function. As the type of directivity, any one of predetermined typical directivity candidates (e.g., bidirectional, unidirectional, etc.) may be selected.
[0063] The weighting coefficient calculation unit 110 may also be connected wirelessly or wired to a hearing assistive device worn by the user, and acquire model number information or directional type information in response to a request from the connected hearing assistive device or provided spontaneously by the hearing assistive device. In this case, the weighting coefficient calculation unit 110 may acquire, from the connected hearing assistive device, an adjustment amount for the weighting coefficient according to the hearing ability for each frequency band. The acquired adjustment amount is used to adjust the weighting coefficient. The weighting coefficient calculation unit 110 may also acquire, from the connected hearing assistive device, an intensity range (i.e., maximum and minimum values) that the user can hear for each frequency band. The acquired intensity range is used to scale the weighting coefficient.
[0064] In the above description, a case has been exemplified in which a weighting coefficient based on the directivity of the sound collection function and the audio object direction is set for each frequency band, but this is not limited to this. The weighting coefficient may be set for all frequency bands collectively, regardless of the frequency band. The user of the sound processing device 10 is not necessarily limited to hearing-impaired people, but may also be able-bodied people. For able-bodied people, the wearing flag may be set to indicate that a hearing assistive device is not being worn, or that a hearing assistive device is being worn. By setting it to wear, the influence of the acoustic characteristics of the sound collection function and the audio object direction on the reproduced sound is simulated. In addition, instead of a weighting coefficient for each frequency band as a parameter related to the hearing assistance characteristics, an adjustment amount for the strength of hearing assistance may be used for each frequency width (e.g., discrete frequency) that is more finely divided than a predetermined frequency band.
[0065] As described above, the audio processing device 10 according to this embodiment includes an audio hearing assistance characteristic setting unit (e.g., weighting coefficient calculation unit 110) that sets audio hearing assistance characteristics, which are the acoustic characteristics of an audio hearing assistance device (e.g., a hearing aid) based on audio hearing assistance-related information related to audio hearing assistance, an audio hearing assistance characteristic application unit (e.g., weighting coefficient application unit 120) that applies audio hearing assistance characteristics to audio signals, and a rendering unit (e.g., binaural rendering unit 130) that adjusts the acoustic characteristics of the audio signal to which the audio hearing assistance characteristics have been applied, for each audio object, in accordance with the target position, listening position, and listening direction. With this configuration, the acoustic characteristics of the hearing aid device are applied to the audio signal, and the acoustic characteristics are adjusted according to the target position, listening position, and listening direction for each audio object. By taking into account the effect of the hearing aid device on the acoustic characteristics, reproduced sound that provides ease of listening equivalent to that when the hearing aid device is used is realized.
[0066] In addition, the hearing assistance characteristic setting unit may determine the sensitivity of the hearing assistance device based on the directionality of the sound collection function of the hearing assistance device and the direction of the target position relative to the listening position, and set the hearing assistance characteristics including the determined sensitivity. With this configuration, the directivity of the sound collection function of the hearing assist device according to the target position of the audio object is reflected in the reproduced sound, resulting in reproduced sound that is as easy to hear as when using the sound collection function of the hearing assist device.
[0067] Furthermore, the hearing assistance characteristic setting unit may set the directivity of the sound collection function for each model (for example, model number) of hearing assistance device, and specify the directivity of the model indicated in the hearing assistance related information. Generally, the influence of the directionality of the sound collection function varies depending on the model of the hearing assistive device. With this configuration, the directionality of the specified device is reflected in the reproduced sound. Therefore, the reproduced sound is as easy to hear as when using the sound collection function of the specified model of hearing assistive device.
[0068] The hearing assistance characteristic setting unit may also determine whether or not a hearing assistance characteristic needs to be set based on hearing assistance necessity information (for example, a wearing flag) included in the hearing assistance related information. With this configuration, whether or not to include the effect of the hearing assist device on the acoustic characteristics is determined depending on whether or not hearing assistance is required, so whether or not to provide reproduced sound equivalent to that when using the sound collection function of the hearing assist device can be easily changed depending on whether or not hearing assistance is required.
[0069] In addition, the rendering unit may set a smaller gain for a rear object, which is an audio object whose target direction is closer to the rear direction than to the front direction, than for a front object, which is an audio object whose target direction, which is the direction of the target position based on the listening position, is closer to the front direction than to the rear direction. With this configuration, the level of the component of the front object in the reproduced sound is relatively higher than the level of the component of the rear object, so that the sound of the front object in the reproduced sound is heard more clearly than the sound of the rear object.
[0070] In addition, the rendering unit may not set a smaller gain for an audio object (e.g., a required object) that is provided with priority over other audio objects, than for a forward object, if the target direction is closer to the rearward direction than to the forward direction. With this configuration, even if the target direction of a sound object to be provided preferentially is closer to the rear, a gain smaller than that of a forward object is not set for the sound object. The level of the sound object component in the reproduced sound is maintained without being relatively lower than the level of the forward object component. Therefore, the sound of the sound object is heard without losing clarity.
[0071] Note that all or part of the above-described audio processing device 10, for example, all or part of the weighting coefficient calculation unit 110, the weighting coefficient application unit 120, and the binaural rendering unit 130, may be implemented by a computer. The functions may be realized by recording a program for implementing the functions on a computer-readable recording medium, loading the program recorded on the recording medium into a processor of a computer system, and executing the processing instructed by the instructions written in the program. Note that the term "computer system" as used herein refers to a computer system built into the audio processing device 10, including hardware such as an operating system (OS) and peripheral devices. Furthermore, the term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into a computer system. Furthermore, the term "computer-readable recording medium" may also include devices that dynamically store programs for a short period of time, such as communication lines when transmitting programs via networks such as the Internet or communication lines such as telephone lines, or devices that store programs for a certain period of time, such as volatile memory within the computer systems that serve as the server or client in such cases. Furthermore, the above program may be one that realizes part of the functions described above, or may be one that can realize the functions described above in combination with a program already recorded in the computer system.
[0072] Furthermore, part or all of the audio processing device 10 in the above-described embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the audio processing device 10 may be individually implemented as a processor, or part or all of the blocks may be integrated into a processor. The integrated circuit implementation method is not limited to LSI, and may be implemented using a dedicated circuit or a general-purpose processor. Furthermore, if an integrated circuit implementation technology that can replace LSI emerges due to advances in semiconductor technology, an integrated circuit based on that technology may be used.
[0073] The embodiments of the present invention have been described in detail above with reference to the drawings, but the specific configuration is not limited to that described above, and various design changes can be made within the scope of the gist of the present invention. [Explanation of symbols]
[0074] S1...sound processing system, 10...sound processing device, 20...playback device, 110...weighting coefficient calculation unit, 112...sound object direction calculation unit, 114...weighting coefficient determination unit, 120...weighting coefficient application unit, 130...binaural rendering unit
Claims
1. a hearing assistance characteristic setting unit that sets hearing assistance characteristics, which are acoustic characteristics of the hearing assistance device, based on hearing assistance-related information related to hearing assistance; a hearing aid feature application unit that applies the hearing aid feature to an audio signal; a rendering unit that adjusts the acoustic characteristics of the audio signal to which the hearing assistance characteristics have been applied, according to the target position, listening position, and listening direction for each audio object. Audio processing device.
2. the hearing assistance characteristic setting unit determines the sensitivity of the hearing assistance device based on the directionality of a sound collection function of the hearing assistance device and the direction of the target position relative to the listening position; Set the hearing aid characteristics including the sensitivity The audio processing device according to claim 1 .
3. the hearing assistance characteristic setting unit sets the directionality of a sound collection function for each model of the hearing assistance device, Identifying the directionality of the model indicated in the hearing assistance related information The audio processing device according to claim 2 .
4. The hearing assistance characteristic setting unit determines whether or not the hearing assistance characteristic needs to be set based on hearing assistance necessity information included in the hearing assistance related information. The audio processing device according to claim 1 .
5. The rendering unit A smaller gain is set for a rear object, which is an audio object whose target direction is closer to the rear direction than to the front direction, than for a front object, which is an audio object whose target direction is closer to the front direction than to the rear direction, which is the direction of the target position based on the listening position. The audio processing device according to claim 1 .
6. The rendering unit For a voice object that is provided with priority over other voice objects, if the target direction is closer to the rear direction than the front direction, a gain smaller than that of the forward object is not set. The audio processing device according to claim 5 .
7. To the computer A program for causing the speech processing device according to claim 1 to function.
Citation Information
Patent Citations
ITRBS.2125-0,
ITRBS.2076-1,
IEC23008-3