Sound localization training apparatus and method using personalized head-related transfer function

The sound directionality training device and method address the challenge of impaired sound localization in hearing-impaired individuals by using personalized HRTFs and spatial sound generation, improving spatial awareness and reducing hazard exposure.

WO2025170145A1PCT designated stage Publication Date: 2025-08-14IND ACADEMIC COOP FOUND HALLYM UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/017167
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-05
Filing Date
2024-11-04
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Individuals with hearing loss, particularly those with bilateral or unilateral hearing impairment, face challenges in discerning sound directionality due to impaired interaural difference information, relying heavily on frequency cues which are not adequately addressed by commercial Head-Related Transfer Functions (HRTFs), leading to difficulties in spatial awareness and increased risk in hazardous situations.

Method used

A sound directionality training device and method utilizing personalized Head-Related Transfer Functions (HRTFs) generated from individual upper body scan data and binaural sound recordings, combined with spatial sound generation and training environments, to enhance sound localization skills through virtual or augmented reality training.

Benefits of technology

Improves the ability of hearing-impaired individuals to discern sound directionality by tailoring training sounds to their unique anatomical characteristics, enhancing spatial awareness and reducing exposure to hazards through personalized sound directionality training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024017167_14082025_PF_FP_ABST
    Figure KR2024017167_14082025_PF_FP_ABST
Patent Text Reader

Abstract

A sound localization training apparatus and method are disclosed. The sound localization training apparatus according to the present specification may comprise: a function generation unit for generating a personalized head-related transfer function by using upper body scan data of a trainee or a sound source recorded through microphones mounted on both ears of the trainee; a spatial sound generation unit for generating spatial sound information related to a place captured using optical equipment; a training setting unit for receiving information related to a training environment and a training difficulty; a training unit for performing a sound localization training by outputting, through headphones, according to the information related to the training environment, a first training sound sound-processed through the personalized head-related transfer function or a second training sound sound-processed through the personalized head-related transfer function and the spatial sound information; and a result generation unit for generating result information about the sound localization training.
Need to check novelty before this filing date? Find Prior Art

Description

Sound directionality training device and method using personalized head-related transfer function

[0001] The present invention relates to a sound directionality training device and method, and more particularly, to a sound directionality training device and method using a personalized head-related transfer function.

[0002] This application claims priority to Korean Patent Application No. 10-2024-0017568, filed on February 5, 2024, the entire disclosure of which is incorporated herein by reference.

[0003] The material described in this section merely provides background information on the embodiments described herein and does not necessarily constitute prior art.

[0004] The ability to locate and orient sound sources in everyday life is essential for recognizing one's surroundings. Individuals with bilateral hearing impairment, such as those with hearing loss in both ears, those with unilateral hearing loss, and / or those with similar degrees of hearing loss but a different hearing pattern due to the use of a hearing aid in only one ear, experience impaired ability to locate and orient sound sources. This can lead to difficulties in daily life and increased risk of exposure to hazardous situations. Therefore, sound directionality training is essential to overcome this impaired ability to discern sound directionality.

[0005] In general, interaural difference information, which refers to the physical difference in the sound reaching the two ears, including the interaural time difference (ITD) and interaural intensity difference (ILD), is the most important information used to recognize the location of a sound source. Another cue, the frequency cue (spectral cue), is important sound directional information used to detect the directionality of a sound source. In the case of a person with normal hearing, the frequency cue is used to supplement the interaural difference information. The frequency cue refers to a cue caused by the frequency change of the sound due to the human body structures such as the ears, face, and torso before the sound reaches the ears. Since the shapes of the ears, face, and torso vary from person to person, the frequency cue has large individual differences.

[0006] Patients with different degrees of hearing loss between the two ears have difficulty properly utilizing interaural information to discern the direction of a sound source, as do normal individuals, due to the difference in hearing. Consequently, patients with significant interaural hearing loss rely heavily on frequency cues, which are physical variations in sound quality depending on an individual's anatomical structure, to discern the direction of a sound source, rather than relying on binaural cues. However, patients with hearing loss wear hearing aids. These patients receive sounds through the aid's microphone. However, the aid's microphone may be positioned behind the ear or toward the head, rather than at the entrance to the ear canal. This may result in different characteristics compared to sounds received at the entrance of the external auditory canal.

[0007] Head-Related Transfer Function (HRTF) is a spatial sound generation formula used in digital virtual spaces. The HRTF is created based on data measured using a standardized body model. The inter-incidence information is determined by the size of the human head and does not vary greatly from person to person. Therefore, individuals with normal hearing may have no difficulty experiencing spatial awareness in virtual spaces through sounds generated using commercially available HRTFs. However, patients with hearing loss who rely heavily on frequency cues due to different hearing levels on both sides may have more difficulty experiencing realistic spatial awareness when using commercially available HRTFs. Therefore, providing patients with hearing loss with commercially available HRTFs in digital spaces to conduct sound directionality training may have difficulties in achieving substantial improvements in sound directionality.

[0008] The purpose of this specification is to provide a sound directionality training device and method.

[0009] This specification is not limited to the above-mentioned tasks, and other tasks not mentioned will be clearly understood by those skilled in the art from the description below.

[0010] In order to solve the above-described problem, a sound directionality training device according to the present specification may include a function generation unit that generates a personalized head-related transfer function using a trainee's upper body scan data or a sound source recorded through a microphone mounted on the trainee's binaural area; a spatial sound generation unit that generates spatial sound information related to a location photographed using optical equipment; a training setting unit that receives information related to a training environment and training difficulty; a training unit that outputs a first training sound processed through the personalized head-related transfer function or a second training sound processed through the personalized head-related transfer function and spatial sound information to headphones according to the information related to the training environment to perform the sound directionality training; and a result generation unit that generates result information for the sound directionality training.

[0011] According to one embodiment of the present specification, when the function generation unit generates the personalized head transfer function using the upper body scan data, the function generation unit can generate the upper body scan data by extracting coordinate data of the facial surface and upper body surface contour of the trainee from data obtained by photographing the trainee using optical equipment.

[0012] According to one embodiment of the present specification, the spatial sound generation unit can generate the spatial sound information by generating coordinate information of cloud points related to a plurality of mesh shapes constituting each object located at the photographed location.

[0013] According to one embodiment of the present specification, the training setting unit can set at least one of the number of directions in which sound sources are presented, the angle difference in directions, and the number of volumes according to information related to the difficulty level.

[0014] According to one embodiment of the present specification, the training environment is characterized in that it is a virtual reality or augmented reality-based training environment, and the training unit can output the first training sound when the training environment is virtual reality-based, and can output the second training sound when the training environment is augmented reality-based.

[0015] According to one embodiment of the present specification, the training unit can record information about the location of the sound source selected by the trainee or the time taken for the training sound to be output and the location of the sound source to be selected.

[0016] According to one embodiment of the present specification, the result generation unit can generate the result information using at least one of a correct response rate, a root of the mean square of an angular error, an absolute value mean of an angular error, an error index, or a Localization Bias.

[0017] According to one embodiment of the present specification, the result generation unit can generate information on the recommended difficulty level of sound directionality training based on the result information.

[0018] The sound directionality training device according to the present specification may be a component of a sound directionality training system including a training assistance device that visually provides a virtual reality or augmented reality to a trainee, outputs training sounds through headphones, and links the trainee's response with the virtual reality or augmented reality.

[0019] A sound directionality training method according to the present specification for solving the above-described problem may include: (a) a step in which a processor generates a personalized head-related transfer function using upper body scan data of a trainee or a sound source recorded through a microphone mounted on the trainee's binaural area; (b) a step in which the processor generates spatial acoustic information related to a location photographed using optical equipment; (c) a step in which the processor receives information related to a training environment and a training difficulty; (d) a step in which the processor outputs a first training sound acoustically processed through the personalized head-related transfer function or a second training sound acoustically processed through the personalized head-related transfer function and spatial acoustic information to headphones according to the information related to the training environment, thereby performing the sound directionality training; and (e) a step in which the processor generates result information for the sound directionality training.

[0020] According to one embodiment of the present specification, in step (a), when the processor generates the personalized head transfer function using the upper body scan data, step (a) may further include a step in which the processor extracts coordinate data of the shape of the face surface and upper body surface of the trainee from data obtained by photographing the trainee using optical equipment to generate the upper body scan data.

[0021] According to one embodiment of the present specification, the step (b) may be a step in which the processor generates the spatial sound information by generating coordinate information of cloud points related to a plurality of mesh shapes constituting each object located at the photographed location.

[0022] According to one embodiment of the present specification, the step (c) may be a step in which the processor sets at least one of the number of directions in which sound sources are presented, the angle difference in directions, and the number of volumes according to information related to the difficulty.

[0023] According to one embodiment of the present specification, the training environment is characterized by being a virtual reality or augmented reality-based training environment, and the step (d) may be a step in which the processor outputs the first training sound when the training environment is virtual reality-based, and may be a step in which the processor outputs the second training sound when the training environment is augmented reality-based.

[0024] According to one embodiment of the present specification, the step (d) may be a step in which the processor records information about the location of the sound source selected by the trainer or the time taken for the training sound to be output and the location of the sound source to be selected.

[0025] According to one embodiment of the present specification, the step (e) may be a step in which the processor generates the result information using at least one of a correct response rate, a root of the mean square of an angular error, an absolute value of an angular error, an error index, or a Localization Bias.

[0026] According to one embodiment of the present specification, the step (e) may further include the step of the processor generating information on the recommended difficulty level of sound directionality training based on the result information.

[0027] Other specific details of the present invention are included in the detailed description and drawings.

[0028] According to one aspect of the present disclosure, the ability to discern sound directionality of a hearing-impaired patient can be relatively improved compared to before training by training sound directionality using training sounds processed through a personalized head-related transfer function and / or spatial acoustic information.

[0029] According to another aspect of the present specification, more effective sound directionality training can be conducted by providing information on the next training difficulty level based on the trainer's training results.

[0030] The effects of the present invention are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.

[0031] FIG. 1 is a block diagram of a sound directionality training system according to one embodiment of the present specification.

[0032] Figure 2 is an example image of scanning a trainer to generate a personalized head transfer function.

[0033] Figure 3 is an example image of recording sound using a microphone mounted on the trainer's head to create a personalized head-related transfer function.

[0034] Figure 4 is a flowchart of a sound directionality training method according to one embodiment of the present specification.

[0035] Figure 5 is a flowchart showing an example of the steps for generating a personalized head-to-head transfer function.

[0036] Figure 6 is a flowchart showing an example of training progress according to the training environment.

[0037] Figure 7 is a flowchart of a sound directionality training method according to another embodiment of the present specification.

[0038] The advantages and features of the invention disclosed in this specification, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, this specification is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided only to ensure that the disclosure of this specification is complete and to fully inform those of ordinary skill in the art (hereinafter referred to as "skilled workers") of the scope of this specification, and the scope of rights of this specification is defined only by the scope of the claims.

[0039] The terminology used herein is for the purpose of describing embodiments and is not intended to limit the scope of the present disclosure. In this specification, singular forms also include plural forms, unless specifically stated otherwise. As used herein, the terms "comprises" and / or "comprising" do not exclude the presence or addition of one or more other components in addition to the components mentioned.

[0040] Throughout the specification, the same reference numerals refer to the same elements, and the term "and / or" includes each and every combination of the elements mentioned. Although terms such as "first," "second," etc. are used to describe various elements, these elements are not limited by these terms. These terms are used only to distinguish one element from another. Therefore, it should be understood that a first element mentioned below may also be a second element within the technical scope of the present invention.

[0041] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in their common sense to those skilled in the art to which this specification pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.

[0042] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings.

[0043] FIG. 1 is a block diagram of a sound directionality training system according to one embodiment of the present specification.

[0044] Referring to FIG. 1, a sound directionality training system (1) according to one embodiment of the present specification may include a sound directionality training device (10) and a training assistance device (20). The sound directionality training device (10) may include a function generation unit (100), a spatial sound generation unit (110), a training setting unit (120), a training unit (130), and a result generation unit (140).

[0045] The above function generation unit (100) can generate a personalized head related transfer function (HRTF) using the trainee's upper body scan data and / or a sound source recorded through a microphone mounted on the trainee's abdomen. The personalized HRTF may refer to a head related transfer function that reflects the individual's physical characteristics.

[0046] Figure 2 is an example image of scanning a trainer to generate a personalized head transfer function.

[0047] Referring to FIG. 2, a scanning assembly may be used to obtain upper body scan data of the trainee. The scanning assembly may scan the trainee using an optical device (200). The optical device may correspond to a LiDAR (Light Detection and Ranging), a single camera, and / or a multi-camera, which are examples and are not limited by the optical device. The scanning assembly may be installed so that the entire face and upper body of the trainee are scanned when scanning the trainee.

[0048] For example, the scanning assembly may include a movable rail (210) installed to surround the trainer 360 degrees in a horizontal direction. At this time, the optical equipment (200) may move along the movable rail (210) and scan the trainer.

[0049] The above-mentioned movable rail (210) may be installed on the floor around the trainee. In addition, when the trainee is sitting on a chair, the movable rail may be fixed to the upper part of the backrest of the chair and installed to horizontally surround the trainee 360 ​​degrees.

[0050] The optical equipment (200) may further include a driving unit (not shown) that automatically moves along the moving rail (210) when receiving an input signal. Alternatively, the optical equipment (200) may be manually moved along the moving rail (210) by a training assistant.

[0051] As another example, the scanning assembly may include a stand (not shown) that secures the optical device (200) to a predetermined position and a rotating member (not shown) that rotates the trainer horizontally in place. The optical device (200) can scan the trainer while rotating in place. Additionally, the training assistant can scan while holding the optical device (200) and rotating around the trainer.

[0052] The above function generation unit (100) can convert the scanned data into coordinate data for a three-dimensional space and extract coordinate data of the shape of the trainer's face surface and upper body surface. The function generation unit (100) can extract coordinate data using techniques such as point cloud and / or feature point extraction, which are merely examples and are not limited by the above techniques.

[0053] At this time, the function generation unit (100) can extract data on the angle of the earlobe, the angle of the face, the shape of the curvature of the earlobe, and / or the depth of the earlobe, among the data on the shape of the face surface, at a relatively higher resolution than data on other parts.

[0054] The above function generation unit (100) inputs the extracted upper body scan data into a head transfer function extraction algorithm to calculate the acoustic gain and frequency change for the sound direction, thereby generating a head transfer function personalized for the trainee.

[0055] The function generation unit (100) can calculate the frequency change of sound due to the pinna effect according to the shape of the auricle by using data corresponding to the angle of the auricle, the shape of the auricle curvature, and the depth of the auricle of the trainee. In addition, the function generation unit (100) can calculate the interaural time difference (ITD), which is the difference in time for sound to reach both ears, and the interaural intensity difference, which is the difference in sound intensity between the two ears, by using data on the distance between the ears of the trainee. In addition, the function generation unit (100) can calculate the frequency and acoustic gain change due to the reverberation of sound caused by the head and the upper body by using data on the shapes of the head and the upper body.

[0056] The above function generation unit (100) can integrate the calculated values ​​and formulate the overall change in sound characteristics reaching each ear according to the individual's physical characteristics into three-dimensional data (change in frequency and acoustic gain according to the location of the sound).

[0057] The above function generation unit (100) can generate a personalized head transfer function using the above three-dimensional data.

[0058] Figure 3 is an example image of recording sound using a microphone mounted on the trainer's head to create a personalized head-related transfer function.

[0059] Referring to FIG. 3, a speaker assembly may be used to record sound sources through microphones mounted on the trainer's ears. The trainer may be positioned in a soundproof room in which at least one speaker is installed. The trainer may be positioned at a distance of 1.5 m from the speakers. A frame having a predetermined shape in which speakers are installed may be installed in the soundproof room. At least one frame may be installed. A plurality of speakers may be installed in the frame at a preset angle. A sound source having a preset spectrum may be sequentially reproduced from each speaker, and the sound source may be recorded through the microphone. In addition, one speaker may rotate around the trainer at a preset angle along the frame and reproduce the sound source.

[0060] At this time, the speaker can reproduce the sound source at different angles in the horizontal direction and / or vertical direction with respect to the trainer.

[0061] Figure 3 (a) is an example image of a plan view looking down from the direction of the trainer's head. Using the speaker assembly, the sound source can be reproduced at different horizontal angles relative to the trainer.

[0062] Figure 3 (b) is an example image of the trainer's face viewed from the front. Using the speaker assembly, the sound source can be reproduced at different vertical angles relative to the trainer.

[0063] Additionally, the soundproof room may be equipped with at least one speaker installed on a movable and / or height-adjustable stand. Using this, sound sources can be recorded by playing them from different horizontal and / or vertical angles relative to the trainer. The method for recording the sound sources is merely an example and is not limited to the above method.

[0064] If the trainee wears a hearing aid and / or cochlear implant, the position of the microphone can be adjusted to record sound sources depending on the type of the device.

[0065] The above function generation unit (100) can input recorded sound source data into a head-related transfer function extraction algorithm to generate a personalized head-related transfer function.

[0066] The personalized head-related transfer function can generate an algorithm that reflects changes in the sound characteristics according to the individual's physical characteristics, such as changes in frequency, frequency band, and / or binaural hearing level, when a sound source is input. Through this, the personalized head-related transfer function can acoustically process the input sound source according to the individual's characteristics.

[0067] The function generation unit (100) can generate a code that can identify the trainer. Thereafter, the function generation unit (100) can match the personalized head-to-head transfer function with a code corresponding to the trainer and store it in a database.

[0068] The above spatial sound generation unit (110) can generate spatial sound information related to a location photographed using optical equipment. The spatial sound generation unit (110) can receive data photographing the space where the trainee is located from the training assistance device (20).

[0069] The training assistance device (20) may include a housing of a predetermined shape that is mounted on the head. The housing may include a display device of a predetermined shape that provides virtual reality and / or augmented reality to the trainee. In addition, the training assistance device (20) may include optical equipment such as LiDAR, a single camera, and / or a dual camera that captures the space where the trainee is located. In addition, the training assistance device (20) may include a GPS module that generates location information of the trainee. In addition, the training assistance device (20) may include a microphone that can detect ambient noise, headphones that output training sounds, and an accelerometer, a gyroscope, and / or a magnetometer for detecting the movement of the trainee. In addition, the training assistance device (20) may include a stereoscopic system for implementing a three-dimensional space, an optical tracking system, infrared sensors, a camera system, a time-of-flight sensor, stereo vision, and / or a depth map camera.

[0070] The above-mentioned trainee can wear the above-mentioned training assistance device (20) and take pictures of the space in which he or she is currently located. The above-mentioned training assistance device (20) can record GPS location information, information related to the height of the trainee's head and / or the inclination of the head when taking pictures.

[0071] The above trainer can wear the above training assistance device (20) and move around in the space to take pictures.

[0072] The spatial sound generation unit (110) can generate information related to the space using the captured data. The spatial sound generation unit (110) can generate information related to objects such as walls, ceilings, and furniture in the space.

[0073] The spatial sound generation unit (110) can model the photographed object into a plurality of mesh shapes. For example, the spatial sound generation unit (110) can model the photographed object into a plurality of triangular mesh shapes. The spatial sound generation unit (110) can generate three-dimensional coordinate information for cloud points constituting the mesh. Each mesh can further include content related to color, texture, normal vector, etc.

[0074] The above spatial sound generation unit (110) may use algorithms such as Structure from Motion (SFM), UV Mapping, and / or Normal Mapping to generate information about the space in which the trainer is located, which are examples and are not limited by the above algorithms.

[0075] The above-described spatial sound generation unit (110) can generate spatial sound information of the space using information about the generated space. The spatial sound information may refer to information related to changes in sound due to the characteristics of objects located in the space. The spatial sound generation unit (110) may use ray tracing and / or ambisonic technology to generate the spatial sound information, which is an example and is not limited by the above-described technology.

[0076] The above spatial sound generation unit (110) can structure the position information and spatial sound information of the trainer and store them in a database.

[0077] In this specification, sound directionality training can output training sounds to a trainee wearing the training assistance device (20) through headphones. The trainee can input the location where the training sounds are estimated to have been generated.

[0078] The above training setting unit (120) can receive information related to the training environment and training difficulty.

[0079] The above training setting unit (120) can receive information about a virtual reality-based and / or augmented reality-based training environment. The trainer can provide the virtual reality and / or augmented reality through the display of the training assistance device (20).

[0080] The above training assistance device (20) may include input devices such as a controller, hand tracking, and / or voice recognition device that input information to the sound directionality training device (10). The trainee may use the input devices to input the location where the training sound is estimated to have been generated.

[0081] The above training setting unit (120) can receive information on a plurality of predetermined training difficulty levels. Depending on each training difficulty level, at least one of the following may be included: the number of directions in which sound sources are presented, the angle difference between the directions, and the volume level.

[0082] For example, the training difficulty may range from level 1 to level 5. In level 1, sound sources may be presented to the left and right of the trainee. In level 2, sound sources may be presented in front, left, and right of the trainee. In level 3, sound sources may be presented in front, behind, left, and right of the trainee. In level 4, sound sources may be presented in front, behind, left, right, 45 degrees forward, and 45 degrees backward of the trainee. In level 5, sound sources may be presented in front, behind, left, right, 45 degrees forward, 135 degrees forward, 45 degrees backward, and 135 degrees backward of the trainee. These are examples and are not limited by the difficulty and number of directions, and sound sources may be presented at horizontal and / or vertical locations of the trainee.

[0083] As another example, the training difficulty may range from Level 1 to Level 5. In Level 1, sound sources may be presented at 180-degree intervals relative to the trainee's front. In Level 2, sound sources may be presented at 120-degree intervals relative to the trainee's front. In Level 3, sound sources may be presented at 90-degree intervals relative to the trainee's front. In Level 4, sound sources may be presented at 60-degree intervals relative to the trainee's front. In Level 5, sound sources may be presented at 30-degree intervals relative to the trainee's front. These are examples and are not limited by the difficulty and angular intervals, and may be presented at horizontal and vertical angles relative to the trainee.

[0084] As another example, the training difficulty may range from Level 1 to Level 7. Levels 1 to 5 may have the same difficulty based on the number of directions or angular intervals. In Level 6, two sound sources with different volumes may be presented. In Level 7, four sound sources with different volumes may be presented. These are examples and are not limited by the difficulty or volume levels.

[0085] The above training setting unit (120) sets the sound source presented according to the level of difficulty as an example, and is not limited to the above method, and can set the sound source presented in various ways, such as a moving sound source, an accelerated moving sound source, and a moving sound source whose volume changes. When a moving sound source is presented, the trainee can input the direction in which the sound source moves.

[0086] In addition, the training setting unit (120) may receive information regarding the type of training. The type of training may correspond to sound localization accuracy training, which trains the ability to distinguish the direction from which a sound is generated, and spatial discrimination training, which relates to the range of the ability to distinguish the direction of a sound in space.

[0087] The above training unit (130) can perform the sound directionality training by outputting the first training sound processed through the personalized head-related transfer function or the second training sound processed through the personalized head-related transfer function and spatial sound information to headphones according to information related to the training environment.

[0088] The above training unit (130) can input the trainee's information and retrieve the personalized head transfer function of the trainee from the above database.

[0089] When the above training environment is a virtual reality-based training environment, the training unit (130) can output the first training sound. Through this, the training unit (130) can output the first training sound to the trainee without applying variables such as reverberation in space due to the size of the space, the distance from the sound source, and / or the frequency characteristics of the sound source.

[0090] When the above training environment is an augmented reality-based training environment, the training unit (130) can output the second training sound. The training unit (130) can retrieve the spatial sound information using the trainee's GPS information. Through this, the training unit (130) can output the second training sound to the trainee with variables applied due to reverberation in the space, such as the area of ​​the space, the distance from the sound source, and / or the frequency characteristics of the sound source.

[0091] The training unit (130) can preferentially process the training sound using the spatial acoustic information. The training unit (130) can calculate changes in sound frequency and / or sound reverberation caused by objects within the captured space. Through this, the training sound can be processed to suit the characteristics of the corresponding space. Thereafter, the training unit (130) can input the training sound processed using the spatial acoustic information into a personalized head-related transfer function. The training unit (130) can process the spatially acoustically processed training sound to suit the physical characteristics of the trainee. Through this, the training unit (130) can generate a second training sound that has been processed to suit the physical characteristics of the corresponding space and the trainee.

[0092] For example, a table may be positioned in the filmed space. The sound source may be positioned below the table. The trainee may be positioned to the left of the table. The training unit (130) may calculate the sound that reaches the trainee as the initial training sound is changed by the table, floor, wall, and / or ceiling. Thereafter, the training unit (130) may input the calculated sound into the personalized head-related transfer function to generate the second training sound. This is an example and is not limited by the above-mentioned situation.

[0093] In addition, the training unit (130) can generate the first training sound by processing the training sound using the personalized head-related transfer function. Thereafter, the training unit (130) can generate the second training sound by processing the first training sound using the spatial acoustic information.

[0094] For example, a table may be positioned in the filmed space. A sound source may be positioned below the table. A trainee may be positioned to the left of the table. The training unit (130) may input a training sound into the personalized head-related transfer function to generate the first training sound. Thereafter, the training unit (130) may generate a second training sound that is heard by the trainee by being changed by the table, floor, wall, and / or ceiling when the first training sound is generated below the table. This is an example and is not limited by the above situation.

[0095] When the above training environment is a virtual reality-based training environment, the training unit (130) can provide the virtual training environment to the trainee through the display of the training assistance device (20). The virtual training environment can be a realistic environment such as a cafe, a sidewalk next to a road, a subway, a school, a sports stadium, etc. In addition, the virtual training environment can be an unrealistic environment such as outer space or a spaceship. The trainee can select the virtual training environment. In addition, the training unit (130) can arbitrarily select the virtual training environment and provide it to the trainee.

[0096] When the above training environment is an augmented reality-based training environment, the above training environment may refer to the space where the trainee is located.

[0097] The training unit (130) may output environmental sounds corresponding to the training environment as training sounds. Furthermore, the training unit (130) may output nonsensical words and / or meaningful words of 1 to 3 syllables as training sounds. Furthermore, the training unit (130) may output nonsensical sentences and / or meaningful sentences as training sounds. These are merely examples and are not limited by the training sounds.

[0098] In addition, the training unit (130) can adjust the signal-to-noise ratio between the environmental sound and the target sound source within an interval of 1 dB.

[0099] The training unit (130) can record information about the location of the sound source selected by the trainee and / or the time taken for the training sound to be output and the location of the sound source to be selected. The training unit (130) can provide the location of a plurality of sound sources where the sound source is presented to the trainee through a display. The trainee can select one of the locations of the plurality of sound sources. The training unit (130) can store the location of the sound source selected by the trainee. The location of the sound source can include a coordinate value of the sound source. The training unit (130) can store the time taken for the training sound to be output and the trainee to select the location of the sound source. The training unit (130) can store the location of the sound source and / or the time taken for selection as log data.

[0100] The training unit (130) can provide visual feedback through the display of the training assistance device (20). If the trainee inputs the correct location of the sound source, the training unit (130) can display 'O' and / or 'Correct answer' on the display to provide visual feedback. In addition, if the trainee inputs the incorrect location of the sound source, the training unit (130) can display 'X' and / or 'Incorrect answer' on the display to provide visual feedback. This is an example and is not limited by the symbols and / or sentences.

[0101] The training unit (130) can provide visual feedback in stages according to the time it takes for the trainee to select the location of the sound source. For example, the location of the sound source can be provided with visual feedback in five stages according to the time it takes to select it. If the trainee selects the location of the sound source within 3 seconds, the training unit (130) can display 'Perfect' on the display. If the trainee selects the location of the sound source between 3 and 5 seconds, the training unit (130) can display 'Great' on the display. If the trainee selects the location of the sound source between 5 and 7 seconds, the training unit (130) can display 'Good' on the display. If the trainee selects the location of the sound source between 7 and 9 seconds, the training unit (130) can display 'Poor' on the display. If the trainee selects the location of the sound source for more than 9 seconds, the training unit (130) can display 'Bad' on the display. This is an example and is not limited by the above steps and words.

[0102] The above result generation unit (140) can generate result information for the sound directionality training. The above result generation unit (140) can generate the result information using at least one of the percent correct trial, root mean square error (RMSE), mean absolute error (MAE), error index, and localization bias.

[0103] The above-mentioned positive response rate refers to the accuracy with which a sound generated from a specific direction can be distinguished from a sound generated from another direction, and can be calculated using [Mathematical Formula 1].

[0104]

[0105] The above RMSE is the root mean square of the difference between the direction of the sound estimated by the trainer and the actual direction of the sound. The lower the RMSE, the more accurately the direction of the sound can be distinguished, and can be calculated using [Mathematical Formula 2].

[0106]

[0107] The above MAE is the average of the differences between the direction of the sound guessed by the trainer and the actual direction of the sound. A lower MAE means that the direction of the sound can be distinguished more accurately, and can be calculated using [Mathematical Formula 3].

[0108]

[0109] The above error index is the average of the difference between the direction of the sound estimated by the trainer and the actual direction of the sound, expressed in degrees. A lower error index means that the direction of the sound can be distinguished more accurately, and can be calculated using [Mathematical Formula 4].

[0110]

[0111] The above Localization Bias refers to the bias of the trainer toward sounds in a specific direction, and the lower the Localization Bias, the more accurately the trainer can distinguish sounds from all directions, and can be calculated using [Mathematical Formula 5].

[0112]

[0113] In the above [Mathematical Formula 1] to [Mathematical Formula 5], the response value is the position value of the sound source selected by the trainer, and the actual value may mean the position value of the sound source from which the training sound is presented.

[0114] The above result generation unit (130) can generate the result information using the log data generated during the sound directionality training process.

[0115] The above result generation unit (130) can generate information on the recommended difficulty level of sound directionality training according to the above result information.

[0116] The above result generation unit (130) can generate information on the recommended difficulty level of sound directionality training according to the above result information.

[0117] For example, the result generation unit (130) may generate information on the recommended difficulty level based on the correct response rate. If the trainee's correct response rate for sound directionality training is equal to or greater than a preset value, the result generation unit (130) may generate information related to a difficulty level that is one level higher than the difficulty level of the corresponding training. The result generation unit (130) may provide feedback to the trainee so that the trainee can proceed with sound directionality training at the generated difficulty level. If the trainee proceeds with training at the generated difficulty level, the result generation unit (130) may output information related to the generated difficulty level to the training setting unit (120). In addition, if the correct response rate is equal to or less than a preset value, the result generation unit (130) may generate information related to a difficulty level that is one level lower than the difficulty level of the corresponding training. Alternatively, the result generation unit (130) may generate information to retry training at the difficulty level of the corresponding training. This is an example and is not limited by the above method.

[0118] As another example, the result generation unit (130) may generate information on the recommended difficulty level in stages according to the correct response rate. If the correct response rate for the trainee's sound directionality training is 90% or higher, the result generation unit (130) may generate information related to a difficulty level that is two levels higher than the difficulty level of the corresponding training. If the correct response rate for the trainee's sound directionality training is between 80% and 90%, the result generation unit (130) may generate information related to a difficulty level that is one level higher than the difficulty level of the corresponding training. If the correct response rate for the trainee's sound directionality training is between 70% and 80%, the result generation unit (130) may generate difficulty information to retry at the difficulty level of the corresponding training. If the correct response rate for the trainee's sound directionality training is between 60% and 70%, the result generation unit (130) may generate information related to a difficulty level that is one level lower than the difficulty level of the corresponding training. If the correct response rate for the above-mentioned trainee's sound directionality training is less than 60%, the result generation unit (130) may generate information related to a difficulty level two levels lower than the difficulty level of the training. This is merely an example and is not limited by the range of the correct response rate and the information related to the training difficulty.

[0119] The above result generation unit (130) can generate information related to the recommended difficulty level based on result information generated using not only the correct response rate but also the root of the mean square of the angle error, the mean absolute value of the angle error, the error index, and the Localization Bias.

[0120] Additionally, the information related to the recommended difficulty level may include information regarding the focus of the training. For example, a trainee may have lower accuracy for sounds presented from the right than for sounds presented from the left. In this case, the information related to the recommended difficulty level may include information indicating that the training focus should be on improving discrimination of sounds from the right. This is merely an example and is not limited by the above.

[0121] The above result generation unit (140) can match information related to the trainer's code, personalized head transfer function, and recommended difficulty level and store them in a database.

[0122] The above function generation unit (100), spatial sound generation unit (110), training setting unit (120), training unit (130), and result generation unit (140) may include a processor, ASIC (application-specific integrated circuit), other chipset, logic circuit, register, communication modem, data processing device, etc. known in the technical field to which the present invention belongs in order to execute calculation and various control logic. In addition, when the above-described control logic is implemented in software, the function generation unit (100), spatial sound generation unit (110), training setting unit (120), training unit (130), and result generation unit (140) may be implemented as a set of program modules. At this time, the program modules may be stored in the memory device and executed by the processor.

[0123] Below, a sound directionality training method using a sound directionality training device according to the present specification is described. However, in describing the sound directionality training method according to the present specification, repetitive descriptions of each component are omitted.

[0124] Figure 4 is a flowchart of a sound directionality training method according to one embodiment of the present specification.

[0125] Referring to FIG. 4, in step S10, the processor may generate a personalized head-related transfer function using the trainee's upper body scan data or a sound source recorded through a microphone mounted on the trainee's binaural area. In step S11, the processor may generate spatial acoustic information related to a location photographed using optical equipment. In step S12, the processor may receive information related to a training environment and training difficulty from the trainee. The processor may generate content related to sound directionality training according to the input information. In step S13, the processor may output a first training sound acoustically processed through the personalized head-related transfer function or a second training sound acoustically processed through the personalized head-related transfer function and spatial acoustic information to headphones according to the training environment, thereby performing the sound directionality training. In step S14, the processor may generate result information regarding the sound directionality training.

[0126] Figure 5 is a flowchart showing an example of the steps for generating a personalized head-to-head transfer function.

[0127] Referring to FIG. 5, in step S10-1, the processor may receive input on whether to use the upper body scan data of the trainee. If the upper body scan data is used (step S10-1 Y), in step S10-2, the processor may extract coordinate data of the facial surface and upper body surface shape from the data photographing the trainee to generate the upper body scan data. Thereafter, the upper body scan data may be input into a head transfer function extraction algorithm to generate the personalized head transfer function.

[0128] When the above-mentioned upper body scan data is not used (step S10-1 N), in step S10-2', the processor can input a sound source recorded through a microphone mounted on the trainer's chin into the head-mounted transfer function extraction algorithm to generate the personalized head-mounted transfer function. The processor can match the generated personalized head-mounted transfer function with the trainer's code and store the matched code in a database.

[0129] In step S11, the processor can model each object located at the photographed location into multiple mesh shapes. Thereafter, the processor can generate coordinate information for cloud points associated with each mesh shape. Through this, the processor can generate information about the photographed location. Using the generated information, the processor can generate spatial acoustic information.

[0130] In the above step S12, the processor can set at least one of the number of directions in which sound sources are presented, the angle difference in directions, and the number of volumes according to information related to the difficulty level.

[0131] Figure 6 is a flowchart showing an example of training progress according to the training environment.

[0132] Referring to FIG. 6, in step S12, the processor may receive information about a virtual reality or augmented reality-based training environment. When the training environment is a virtual reality-based training environment (step S13-1 Y), in step S13-2, the processor may output the first training sound to perform sound directionality training. When the training environment is an augmented reality-based training environment (step S13-1 N), in step S13-2', the processor may output the second training sound to perform sound directionality training. In step S13-3, the processor may record information about the location of the sound source selected by the trainee and / or the time taken by the trainee to select the sound source.

[0133] In the above step S14, the processor may generate the result information using the information recorded in the above step S13-3. The processor may generate the result information by calculating at least one of the correct response rate, the root of the mean square of the angular error, the mean absolute value of the angular error, the error index, and / or the localization bias using the recorded information.

[0134] Figure 7 is a flowchart of a sound directionality training method according to another embodiment of the present specification.

[0135] Referring to Fig. 7, steps S20 to S23 are identical to steps S10 to S13, and thus, a repetitive description thereof will be omitted. In step S24, the processor may generate information on the recommended difficulty level of sound directionality training based on the result information. Thereafter, the processor may receive a signal indicating whether sound directionality training is in progress. When a signal indicating sound directionality training progress is received (step S25 Y), the processor may perform sound directionality training set based on the recommended difficulty level.

[0136] The sound directionality training method according to the present specification can be implemented in the form of a computer program written to perform each step on a computer and recorded on a computer-readable recording medium. The aforementioned computer program may include code coded in a computer language, such as C / C++, C#, JAVA, Python, or machine language, that can be read by the processor (CPU) of the computer through the device interface of the computer, so that the computer reads the program and executes the methods implemented as a program. Such code may include functional code related to functions that define the functions necessary to execute the methods, and may include control code related to execution procedures necessary for the processor of the computer to execute the functions according to a predetermined procedure. In addition, such code may further include memory reference-related code regarding which location (address address) in the internal or external memory of the computer should reference additional information or media necessary for the processor of the computer to execute the functions. In addition, if the processor of the computer needs to communicate with any other computer or server located remotely in order to execute the functions, the code may further include communication-related code regarding how to communicate with any other computer or server located remotely using the communication module of the computer, and what information or media to send and receive during communication.

[0137] The above storage medium refers to a medium that stores data semi-permanently and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specifically, examples of the storage medium include, but are not limited to, ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device. That is, the program can be stored in various recording media on various servers that the computer can access or in various recording media on the user's computer. In addition, the medium can be distributed across network-connected computer systems, so that computer-readable code can be stored in a distributed manner.

[0138] While the embodiments of this specification have been described with reference to the attached drawings, those skilled in the art will appreciate that the present invention can be implemented in other specific forms without altering the technical spirit or essential features thereof. Therefore, the embodiments described above should be understood to be illustrative in all respects and not restrictive.

[0139] [Explanation of symbols]

[0140] 1: Sound Directionality Training System

[0141] 10: Sound Directionality Training Device

[0142] 20: Training Aid Device

[0143] 100: Function creation section

[0144] 110: Spatial sound generation unit

[0145] 120: Training Settings

[0146] 130: Training Department

[0147] 140: Result generation section

Claims

1. A function generation unit that generates a personalized head-related transfer function using the upper body scan data of the trainee or a sound source recorded through a microphone mounted on the trainee's abdomen; A spatial sound generation unit that generates spatial sound information related to a location photographed using optical equipment; A training setting section that receives information related to the training environment and training difficulty; A training unit that performs the sound directionality training by outputting a first training sound processed through the personalized head-related transfer function or a second training sound processed through the personalized head-related transfer function and spatial sound information to headphones according to information related to the training environment; and A sound directionality training device, comprising a result generation unit that generates result information for the above sound directionality training.

2. In claim 1, When the above function generation unit generates the personalized head transfer function using the upper body scan data, The above function generation part is, A sound directionality training device that extracts coordinate data of the facial surface and upper body surface contour of a trainee from data taken of the trainee using optical equipment to generate the upper body scan data.

3. In claim 1, The above spatial sound generation unit, A sound directionality training device that generates spatial sound information by generating coordinate information of cloud points related to a plurality of mesh shapes constituting each object located at the above-mentioned photographed location.

4. In claim 1, The above training settings section, A sound directionality training device that sets at least one of the number of directions in which sound sources are presented, the angle difference in direction, and the number of volumes according to information related to the above difficulty level.

5. In claim 1, The above training environment is characterized by being a virtual reality or augmented reality-based training environment. A sound directionality training device, wherein the training unit outputs the first training sound when the training environment is virtual reality-based, and outputs the second training sound when the training environment is augmented reality-based.

6. In claim 1, The above training department, A sound directionality training device that records information about the location of a sound source selected by the above trainer or the time taken for the training sound to be output and the location of the sound source to be selected.

7. In claim 1, The above result generation unit, A sound directionality training device that generates the result information by using at least one of a correct response rate, a root mean square of an angular error, an absolute value of an angular error, an error index, or a localization bias.

8. In claim 1, The above result generation unit, A sound directionality training device that generates information on the recommended difficulty level of sound directionality training based on the above result information.

9. A sound directionality training device according to any one of claims 1 to 8; and A sound directional training system comprising a training assistance device that visually provides a virtual reality or augmented reality to a trainee, outputs training sounds through headphones, and links the trainee's responses with the virtual reality or augmented reality. 10.(a) A step in which a processor generates a personalized head-related transfer function using the upper body scan data of the trainer or a sound source recorded through a microphone mounted on the trainer's abdomen; (b) a step in which the processor generates spatial sound information related to a location photographed using optical equipment; (c) a step in which the processor receives information related to the training environment and training difficulty; (d) a step of the processor performing the sound directionality training by outputting the first training sound processed through the personalized head-related transfer function or the second training sound processed through the personalized head-related transfer function and spatial sound information to headphones according to information related to the training environment; and (e) a step of the processor generating result information for the sound directionality training; a sound directionality training method comprising:

11. In claim 10, In the step (a), when the processor generates the personalized head transfer function using the upper body scan data, Step (a) above, A sound directionality training method, further comprising a step of generating the upper body scan data by extracting coordinate data of the shape of the face surface and upper body surface of the trainee from data captured by the processor using optical equipment.

12. In claim 10, Step (b) above, A sound directionality training method, wherein the processor generates spatial sound information by generating coordinate information of cloud points related to a plurality of mesh shapes constituting each object located at the photographed location.

13. In claim 10, Step (c) above, A sound directionality training method, wherein the processor sets at least one of the number of directions in which sound sources are presented, the angle difference in direction, and the number of volumes according to information related to the difficulty level.

14. In claim 10, The above training environment is characterized by being a virtual reality or augmented reality-based training environment. Step (d) above, When the above training environment is virtual reality-based, the step of the processor outputting the first training sound, A sound directionality training method, wherein the processor outputs the second training sound when the training environment is based on augmented reality.

15. In claim 10, Step (d) above, A sound directionality training method, wherein the processor records information about the location of the sound source selected by the trainer or the time taken for the training sound to be output and the location of the sound source to be selected.

16. In claim 10, Step (e) above, A sound directionality training method, wherein the processor generates the result information by using at least one of a correct response rate, a root mean square of an angular error, an absolute value of an angular error, an error index, or a localization bias.

17. In claim 10, Step (e) above, A method for sound directionality training, further comprising a step in which the processor generates information on a recommended difficulty level of sound directionality training based on the result information.

18. A computer program written to perform each step of the sound directionality training method according to any one of claims 10 to 17 on a computer and recorded on a computer-readable recording medium.

Citation Information

Patent Citations

  • Method for providng virtual-reality based on multi omni-direction camera and microphone, sound signal processing apparatus, and image signal processing apparatus for performin the method

    KR1020180090022A

  • Method and distributed processing system for estimating performance of heterogeneous processing unit

    KR1020210008597A

  • A Novel Fructose C4 epimerases and Preparation Method for producing Tagatose using the same

    KR102558526B1

  • Food packing envelope with excellent ability to exhaust internal pressure

    KR102686469B1

  • Using bluetooth / wireless hearing aids for personalized HRTF creation

    US20230041038A1