Information processing method, information processing system, and program

An information processing device facilitates easy and accurate personalization of HRTFs by user evaluation, addressing the challenges of costly and cumbersome acoustic measurements, enhancing three-dimensional sound reproduction.

WO2025206184A1PCT designated stage Publication Date: 2025-10-02SONY GROUP CORP

Patent Information

Application Number
PCT/JP2025/012468
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-28
Filing Date
2025-03-27
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing methods for determining individual head-related transfer functions (HRTFs) are cumbersome and costly, making it difficult to personalize sound reproduction for multiple users.

Method used

An information processing device that presents multiple sample sounds processed with different HRTFs, allowing users to evaluate and compare them, determining a recommended HRTF based on their preferences, thereby optimizing sound reproduction parameters.

Benefits of technology

Enables easy and accurate personalization of HRTFs without the need for expensive and cumbersome acoustic measurements, improving the three-dimensional sound perception through headphones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025012468_02102025_PF_FP_ABST
    Figure JP2025012468_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present technology relates to an information processing method, an information processing system, and a program that make it possible to easily perform personal optimization of sound reproduction parameters relating to sound reproduction. A determination unit determines a recommendation parameter that is a sound reproduction parameter to be recommended to a user on the basis of the user's evaluation responses to a plurality of test sounds obtained by respectively applying a plurality of sound reproduction parameters related to sound reproduction to a predetermined sound source. The present technology can be applied to a case of, for example, performing personal optimization of sound reproduction parameters relating to sound reproduction such as a head-related transfer function (HRTF).
Need to check novelty before this filing date? Find Prior Art

Description

Information processing method, information processing system, and program

[0001] The present technology relates to an information processing method, an information processing system, and a program, and in particular to an information processing method, an information processing system, and a program that enable easy personal optimization of sound reproduction parameters related to sound reproduction, for example.

[0002] An example of a sound reproduction parameter related to sound reproduction that is applied to sound is the head-related transfer function (HRTF). The HRTF is a transfer function that represents the transfer characteristics from a predetermined sound source position to the ear (eardrum position).

[0003] HRTFs differ from person to person due to individual differences in the shape of the torso, head, pinna, etc. To determine an individual's HRTF, a speaker is placed at the sound source position for which the HRTF is to be obtained, and a microphone is placed near the eardrum of the user (listener) to perform acoustic measurements. This makes it possible to determine the HRTF for that individual user at the sound source position for which the HRTF is to be obtained.

[0004] A technology called binaural reproduction has been proposed, which applies (convolutes) the HRTF of a predetermined sound source position to a sound source and reproduces the resulting sound through headphones, thereby simulating the same situation (sound at the eardrum position) as when a sound actually propagates through space from the predetermined sound source position and reaches the user's ear. Binaural reproduction is described, for example, in Patent Literature 1. Binaural reproduction allows a user to experience sound as if it is coming from a predetermined sound source position, even though the sound is being reproduced through headphones. If a user's individual HRTF can be accurately determined and the HRTF can be appropriately applied to a sound source and reproduced through headphones, it is possible to achieve such accuracy that it is almost impossible to distinguish the difference between sound reproduced from the actual sound source position and sound reproduced through headphones.

[0005] International Publication No. 2021 / 187229

[0006] Acoustic measurements to determine an individual's HRTF require large-scale equipment, and due to issues such as the effort required for users to go directly to the location where the equipment is installed, as well as the time and cost required to install and operate the equipment, it is practically difficult to perform acoustic measurements to determine the individual HRTFs of many users.

[0007] Therefore, there is a demand for a technique that can easily perform personal optimization of HRTF, that is, obtain an individual's HRTF or an HRTF close to the individual's HRTF.

[0008] The same applies to other parameters related to sound reproduction that differ from user to user for physical or preference reasons, in addition to HRTFs.

[0009] The present technology has been made in view of such circumstances, and makes it possible to easily perform personal optimization of sound reproduction parameters related to sound reproduction.

[0010] The information processing method of the present technology is an information processing method that includes determining recommended parameters, which are sound reproduction parameters to be recommended to a user, based on the user's evaluation responses to multiple sample sounds obtained by applying multiple sound reproduction parameters related to sound reproduction to a specified sound source.

[0011] The information processing system or program of the present technology is an information processing system that includes a determination unit that determines recommended parameters, which are sound reproduction parameters to be recommended to a user, based on the user's evaluation responses to multiple sample sounds obtained by applying multiple sound reproduction parameters related to sound reproduction to a specified sound source, or a program for causing a computer to function as such an information processing system.

[0012] In this technology, recommended parameters, which are sound reproduction parameters to be recommended to a user, are determined based on the user's evaluation responses to a plurality of sample sounds obtained by applying a plurality of sound reproduction parameters related to sound reproduction to a predetermined sound source.

[0013] The information processing system may be an independent device, or may be an internal block constituting a single device, or may be composed of multiple independent devices.

[0014] The program can be provided by transmitting it via a transmission medium or by recording it on a recording medium.

[0015] 1 is a block diagram showing an example configuration of an embodiment of an information processing device to which the present technology is applied. FIG. 1 is a flowchart illustrating an overview of the processing of the information processing device 10. FIG. 1 is a flowchart illustrating details of the processing of personal optimization of HRTFs. FIG. 2 is a diagram showing an example of the progress of the processing of personal optimization of HRTFs in a tournament format. FIG. 2 is a flowchart illustrating details of the processing of the preliminary phase of stage Ns in step S12. FIG. 3 is a flowchart illustrating details of the processing of the final phase of stage Ns in step S13. FIG. 4 is a flowchart illustrating details of the processing of the defense phase of stage Ns in step S15. FIG. 5 is a diagram showing an example display of a setting screen (window) as a first example of a UI. FIG. 6 is a diagram showing an example display of an optimization screen (window) as a second example of a UI. FIG. 7 is a diagram showing an example display of an experience screen (window) as a third example of a UI. FIG. 8 is a diagram showing another example progress of the processing of personal optimization of HRTFs in a tournament format. FIG. 9 is a block diagram showing an example configuration of another embodiment of an information processing device to which the present technology is applied. FIG. 10 is a flowchart illustrating an example of the processing of the information processing device 110. FIG. 11 is a flowchart illustrating an example of the listening comparison processing in step S102. FIG. 12 is a flowchart illustrating an example of the adjustment processing in step S107. FIG. 13 is a flowchart illustrating an example of the processing of adjusting the direction of HRTF localization for direction #n performed in the adjustment processing. 1 is a flowchart illustrating an example of the process of adjusting the timbre of an HRTF for direction #n performed in the adjustment process. FIG. 2 is a diagram illustrating an example display of an adjustment screen (window) as a fourth example of a UI. FIG. 3 is a diagram illustrating an example display of an adjustment screen (window) as a fourth example of a UI. FIG. 4 is a diagram illustrating another example display of the adjustment screen. FIG. 5 is a diagram illustrating the accuracy of a recommended HRTF determined by a listening comparison process without presenting a criterion for judgment. FIG. 6 is a diagram illustrating the accuracy of a recommended HRTF after adjustment obtained by an adjustment process performed by presenting a criterion, with a recommended HRTF determined by a listening comparison process with a criterion for judgment set as the target HRTF. FIG. 7 is a diagram illustrating a first method of multi-environment personalization. FIG. 8 is a diagram illustrating a second method of multi-environment personalization. FIG. 9 is a diagram illustrating a third method of multi-environment personalization. FIG. 10 is a diagram illustrating a fourth method of multi-environment personalization.1 is a diagram illustrating an example of a process for generating a sample sound in a listening comparison process and a process for adjusting the direction of localization / timbre in an adjustment process when performing second to fourth methods of multiple environment personalization.

[0016] <Configuration example of one embodiment of information processing device>

[0017] FIG. 1 is a block diagram showing an example of the configuration of an embodiment of an information processing device to which the present technology is applied.

[0018] In recent years, stereophonic playback, which can reproduce three-dimensional sound direction, distance, spread, reverberation, etc., has become widespread in fields such as movies, music, and games. Ideally, stereophonic content would be played from multiple speakers spatially arranged around a user (listener). However, creating such a speaker playback environment with multiple speakers can be difficult for users due to financial or spatial constraints. Therefore, in headphone playback, in which stereophonic content is played back using headphones, binaural playback, which recreates a speaker playback environment or a viewing experience similar to that of a speaker playback environment, has attracted attention as a highly valuable technology.

[0019] Here, headphone playback includes listening to sound using headphones, as well as listening to sound using sound output devices that are used in contact with a person's ear, such as earphones and neck speakers, and sound output devices that are used in close proximity to a person's ear.

[0020] Binaural playback uses HRTF, a type of sound reproduction parameter related to sound reproduction. HRTF is a function that represents the transmission characteristics of sound from the sound source to the human ear (eardrum), and plays an important role in three-dimensional perception of the direction, distance, spread, reverberation, etc. of the sound source. The frequency characteristics of sound emitted from the sound source change depending on the shape of the torso, head, and pinna before it reaches the eardrum of the user listening to the sound. The way in which the frequency characteristics of the sound change varies depending on the direction from which the sound arrives. Humans perceive three-dimensional information about the sound source based on the difference in the frequency characteristics of the sound obtained from the sounds heard by the two ears.

[0021] HRTFs differ from person to person due to individual differences in the shape of the torso, head, and pinna, etc. To determine an individual's HRTF, a speaker is placed at the sound source position for which the HRTF is to be obtained, and a microphone is placed near the user's eardrum to perform acoustic measurements. This makes it possible to determine the HRTF for that individual user at the sound source position for which the HRTF is to be obtained.

[0022] In binaural playback, a sound source is convolved with the HRTF of a specific sound source position, and the resulting sound is played back over headphones, thereby simulating the experience of sound actually propagating through space from the specific sound source position and reaching the user's ears. Binaural playback allows the user to experience sound as if it is coming from the specific sound source position, even though it is played back over headphones. If the user's individual HRTF can be accurately determined and that HRTF is appropriately applied to the sound source for playback over headphones, it is possible to achieve such accuracy that it is almost impossible to distinguish the difference between sound played back at the actual sound source position and sound played back over headphones.

[0023] Acoustic measurements to determine an individual's HRTF require large-scale equipment, and due to issues such as the effort required for users to go directly to the location where the equipment is installed, as well as the time and cost required to install and operate the equipment, it is practically difficult to perform acoustic measurements to determine the individual HRTFs of many users.

[0024] Therefore, the information processing device 10 in FIG. 1 functions as an optimization device that can easily perform personal optimization of HRTF, that is, obtain an individual's HRTF or an HRTF close to an individual's HRTF.

[0025] The information processing device 10 presents to the user a plurality of sample sounds obtained by applying a plurality of HRTFs to a predetermined sound source, and based on the user's evaluation responses to the plurality of sample sounds, the information processing device 10 determines an appropriate HRTF as a recommended HRTF to be recommended to the user, and recommends the HRTF to the user.

[0026] For example, the information processing device 10 selects, for example, two HRTFs as multiple HRTFs and presents to the user sample sounds obtained by applying each of the two HRTFs to a predetermined sound source, i.e., two sample sounds obtained by convolving each of the two HRTFs with the predetermined sound source. Presenting the sample sounds to the user with HRTFs applied is also referred to as a (HRTF) match. The HRTFs in this technology include BRTFs (Binaural Room Transfer Functions) (or the HRTFs in this technology can be interpreted as HRTFs or BRTFs). That is, the HRTF applied to the predetermined sound source may be an HRTF of an arbitrary environment (an HRTF including the acoustic characteristics of an arbitrary environment), or an HRTF obtained by synthesizing (applying) an RTF (room transfer function) of an arbitrary environment to an HRTF of an anechoic chamber. The HRTF of a certain environment, i.e., the BRTF, may be an HRTF obtained by performing measurements in that environment, or a BRTF obtained by synthesizing the RTF of that environment with an HRTF of an anechoic chamber. An anechoic HRTF is an HRTF that does not include the acoustic characteristics of the environment, for example, an HRTF measured in an anechoic chamber.

[0027] The user listens to the two sample sounds and compares them, and then answers their evaluation of which they prefer (better), specifically which they feel is closer to the true three-dimensional sound (direction, distance, spread, reverberation, etc.).

[0028] The information processing device 10 selects two HRTFs to be next compared based on the user's answer (evaluation), and compares the two HRTFs. The information processing device 10 repeats the selection of two HRTFs to be next compared based on the user's answer and the comparison of the two HRTFs as necessary, and determines a personally optimized HRTF (personal HRTF) as a recommended HRTF to be recommended to the user.

[0029] Although it may take skill to evaluate a large number of sample sounds at once, even a novice user can easily evaluate which of two sample sounds is better.

[0030] The information processing device 10 is characterized by various aspects of the process for determining recommended HRTFs, such as the algorithm for selecting two HRTFs to be compared, the UI (user interface) provided to the user for personal optimization, and other aspects.

[0031] In Figure 1, an information processing device 10 as an optimization device has a database 11, a selection unit 12, a sound source storage unit 13, a preview sound generation unit 14, a determination unit 15, an internal state storage unit 16, an image storage unit 17, a UI generation unit 18, a preview sound presentation unit 21, an answer reception unit 22, and an image presentation unit 23.

[0032] The database 11 stores a large number (plurality) of HRTFs. For example, the database 11 stores a large number of HRTFs representing transfer characteristics corresponding to various combinations of human body shapes, such as the shapes of the torso, head, and pinna, for various sound source positions in one or more directions including the front (of the user).

[0033] The selection unit 12 selects two HRTFs to be compared from the HRTFs stored in the database 11, and supplies them to the sample sound generation unit 14. The selection unit 12 selects the two HRTFs to be compared, for example, based on the internal state stored in the internal state storage unit 16. The two HRTFs to be compared are hereinafter also referred to as HRTF-A and HRTF-B.

[0034] The sound source storage unit 13 stores one or more sound sources as personal optimization sound sources, which are sound sources (predetermined sound sources) used for personal optimization.

[0035] The personal optimization sound source may be, for example, a monaural or multi-channel sound source with nearly flat frequency characteristics in the frequency band from approximately 50 Hz to 16 kHz, or an object audio sound source. This is because the shape of the peaks and notches in the HRTF frequency characteristics, particularly in the region from approximately 4 kHz to 16 kHz, is related to the perception of sound source direction. In other words, this is to ensure that the differences in the shape of the peaks and notches for each HRTF applied to the personal optimization sound source are reflected in the sample sound generated by applying the HRTF to the personal optimization sound source.

[0036] An example of a monaural or multi-channel sound source with nearly flat frequency characteristics in the frequency band of approximately 50 Hz to 16 kHz is pink noise. In addition to pink noise, other sounds such as waterfalls, flowing water, wind, fire, and thunder can be used as personal optimization sound sources. Other examples of personal optimization sound sources include the sounds of vehicles such as cars and helicopters, and the sounds of moving animals such as birds and horses.

[0037] The sample sound generation unit 14 generates sample sounds by applying the two HRTFs to be compared, HRTF-A and HRTF-B, from the selection unit 12, to the personal optimization sound source stored in the sound source storage unit 13, and supplies the sample sounds to the sample sound presentation unit 21. In other words, the sample sound generation unit 14 convolves the HRTF-A and HRTF-B into the personal optimization sound source, respectively, to generate sample sounds of HRTF-A and HRTF-B, and supplies the sample sounds to the sample sound presentation unit 21.

[0038] It is desirable for the test sound to properly reflect the differences in HRTF, so the type of sound to use as the test sound must be carefully determined.

[0039] When generating the sample sound, the sound source positions of the HRTFs applied to the personalized sound source (sound source positions where the HRTFs represent transfer characteristics) include positions in one or more directions (one or more directions) including the front. Sounds from the sides are often localized with some degree of accuracy even when an HRTF that is not personalized is applied, but sounds from the front are often localized behind or inside the user's head if a personalized HRTF is not applied. For this reason, it is desirable to configure the sample sound to focus on sounds from the front. Here, "forward" (of the user) means a direction in which both the azimuth and elevation angles (as seen from the user) are within the range of approximately -30 degrees to +30 degrees.

[0040] The preview sound generator 14 can generate, for example, pink noise as a preview sound that is emitted from a sound source position in only one direction, i.e., the front direction, in the forward direction. The preview sound generator 14 can generate, for example, pink noise as a preview sound that is emitted from each of the sound source positions of the left front, the front direction, and the right front, in that order. The preview sound generator 14 can generate, for example, pink noise as a preview sound that is emitted from each of the sound source positions of the lower front, the front direction, and the upper front, in that order. The front direction refers to a direction in which both the azimuth angle and the elevation angle are approximately 0 degrees. The left front refers to a direction in which the azimuth angle is approximately -30 degrees and the elevation angle is approximately 0 degrees, and the right front refers to a direction in which the azimuth angle is approximately 30 degrees and the elevation angle is approximately 0 degrees. The lower front refers to a direction in which the azimuth angle is approximately 0 degrees and the elevation angle is approximately -30 degrees, and the upper front refers to a direction in which the azimuth angle is approximately 0 degrees and the elevation angle is approximately 30 degrees. For the azimuth angle, the left direction is negative and the right direction is positive, and for the elevation angle, the downward direction is negative and the upward direction is positive.

[0041] In addition, the preview sound generator 14 can generate, as preview sounds, sounds such as the sound of a waterfall, flowing water, wind, fire, thunder, etc., which is emitted from a sound source position in one direction ahead, sounds which are emitted in sequence from sound source positions in two or more directions including the front, and sounds which are emitted simultaneously from sound source positions in two or more directions including the front.The preview sound generator 14 can generate, as preview sounds, sounds of vehicles such as cars or helicopters, or animals such as birds or horses moving in front from left to right or right to left.

[0042] The determination unit 15 receives the user's evaluation of the sample sound from the response receiving unit 22. The determination unit 15 makes various decisions based on the user's (evaluation) response. For example, based on the user's response, the determination unit 15 determines the HRTF applied to the sample sound that the user evaluated as good, out of the two HRTFs, HRTF-A and HRTF-B, as the winner of the competition (first sound reproduction parameter). Also, for example, the determination unit 15 determines the HRTF applied to the sample sound that the user evaluated as bad as the loser of the competition. Furthermore, the determination unit 15 determines a recommended HRTF (recommended parameter) to be recommended to the user based on the internal state stored in the internal state storage unit 16. The determination unit 15 updates the internal state stored in the internal state storage unit 16 based on the results of various decisions. The internal state is updated based on the various decisions (the results of the decisions), and various decisions are made based on the user's responses. Therefore, determining the recommended HRTF based on the internal state can be said to be determining the recommended HRTF based on the user's answer.

[0043] The internal state storage unit 16 stores internal states. The internal states are internal information necessary for determining recommended HRTFs. The internal states include, for example, the history of the user's answers, status information for each HRTF stored in the database 11, and information about the current stage and phase in personal optimization. HRTF statuses include winner and loser. Other HRTF statuses include "not presented," which indicates that the applied sample sound is an unpresented HRTF (unpresented parameter) that has not yet been presented to the user. A stage is a processing unit for determining one HRTF as the current champion (third sound reproduction parameter) in personal optimization, and a phase is a processing unit that constitutes a stage. In this embodiment, the stages are composed of three phases: a qualifying phase (first phase), a final phase (second phase), and a defense phase (third phase), as described below.

[0044] The image storage unit 17 stores various images and the like used in the UI generated by the UI generation unit 18 .

[0045] The UI generation unit 18 generates an image as a UI using the internal state stored in the internal state storage unit 16 and the image stored in the image storage unit 17 as necessary, and supplies it to the image presentation unit 23.

[0046] The preview sound presentation unit 21 presents the preview sound from the preview sound generation unit 14 and various other sounds to the user. For example, the preview sound presentation unit 21 presents the preview sound to the user by causing a sound output device (not shown) that outputs sound, such as headphones, to output the preview sound from the preview sound generation unit 14.

[0047] The answer receiving unit 22 receives the user's evaluation answer to the sample sound, and supplies it to the determination unit 15. For example, the answer receiving unit 22 receives the user's (evaluation) answer accepted by an input device (not shown) that accepts user input, such as a keyboard, a mouse, a touch panel, or a microphone, and supplies it to the determination unit 15.

[0048] The image presentation unit 23 presents to the user an image as a UI from the UI generation unit 18 and various other images. For example, the image presentation unit 23 presents the UI to the user by displaying the UI from the UI generation unit 18 on a display device (not shown) that displays images, such as a personal computer (PC), a mobile terminal such as a smartphone, or an xR (AR, VR, MR) device such as a head-mounted display.

[0049] The information processing device 10 can be configured as a single device (housing) such as a game console, or as a server-client system such as a web server and web browser. When the information processing device 10 is configured as a server-client system, at least the components from the database 11 to the image presentation unit 23, which are enclosed by dotted lines in the figure, can be provided on the client, and the remaining blocks can be provided on the server. In the information processing device 10, the preview sound generation unit 14 generates preview sounds by applying HRTFs stored in the database 11 to personal optimization sound sources stored in the sound source storage unit 13. However, the preview sounds can be generated in advance by applying the HRTFs stored in the database 11 to the personal optimization sound sources stored in the sound source storage unit 13. In this case, the information processing device 10 can be configured without the database 11, the sound source storage unit 13, and the preview sound generation unit 14, for example, and the application (program) implementing the information processing device 10 can be run even in an environment with limited processing resources. Of the database 11, the sound source storage unit 13, and the sample sound generation unit 14, the database 11 may not be used in the recommendation process, but may be used when finally recommending the HRTFs resulting from personal optimization to the user. Furthermore, the application only needs to play the sample sound to process the sample sound, making application implementation easier. Furthermore, in the information processing device 10, storing sample sounds generated in advance by applying HRTFs to each of the personal optimization sound sources may require less storage capacity than storing both the personal optimization sound source and the HRTF. In particular, the storage capacity can be reduced by compressing and storing the sample sounds using, for example, Advanced Audio Coding (AAC) or MP3.

[0050] <Overview of processing by information processing device 10>

[0051] FIG. 2 is a flowchart illustrating an outline of the processing of the information processing apparatus 10 of FIG.

[0052] In step S1, the selection unit 12 selects two HRTFs, HRTF-A and HRTF-B, to be compared next from among the unpresented HRTFs (whose internal status is unpresented) stored in the database 11. The selection unit 12 supplies the HRTFs, HRTF-A and HRTF-B, to be compared next to each other to the sample sound generation unit 14, and the process proceeds from step S1 to step S2.

[0053] In step S2, the sample sound generation unit 14 generates sample sounds by applying (convolving) the HRTF-A and HRTF-B from the selection unit 12 to the personal optimization sound source stored in the sound source storage unit 13. The sample sound generation unit 14 supplies the HRTF-A sample sound and the HRTF-B sample sound to which HRTF-A and HRTF-B have been applied, respectively, to the sample sound presentation unit 21, and the process proceeds from step S2 to step S3.

[0054] In step S3, the preview sound presentation unit 21 presents the HRTF-A preview sound and the HRTF-B preview sound from the preview sound generation unit 14 to the user, i.e., pits HRTF-A against HRTF-B, and the process proceeds to step S4. Here, the image presentation unit 23 can present an image related to the preview sound to the user along with the preview sound presented by the preview sound presentation unit 21. For example, the UI generation unit 18 can generate a UI that displays an image representing the preview sound or a moving image synchronized with the preview sound as an image related to the preview sound, and the image presentation unit 23 can present that UI to the user.

[0055] In step S4, the answer receiving unit 22 waits for the user's answer to be input regarding the evaluation of the HRTF-A sample sound and the HRTF-B sample sound, and receives the answer (from the user). The answer receiving unit 22 supplies the received user's answer to the determination unit 15, and the process proceeds from step S4 to step S5.

[0056] In step S5, the determination unit 15 determines the winner and loser, etc., of the HRTF-A and HRTF-B that were played against each other, based on the user's answer received from the answer receiving unit 22. Furthermore, the determination unit 15 updates the internal state stored in the internal state holding unit 16 based on the determination of the winner and loser, etc., and the process proceeds from step S5 to step S6.

[0057] In step S6, the determination unit 15 determines whether to end the personal optimization. If it is determined in step S6 that the personal optimization should not be ended, for example, if the user has not operated the information processing device 10 to end the personal optimization, the process returns to step S1, and the same process is repeated.

[0058] On the other hand, if it is determined in step S6 that personal optimization is to be ended, for example, if the user operates the information processing device 10 to end personal optimization, the processing proceeds to step S7.

[0059] In step S7, the determination unit 15 determines a recommended HRTF to be recommended to the user based on the internal state stored in the internal state storage unit 16. Furthermore, the determination unit 15 updates the status of the HRTF determined as the recommended HRTF in the internal state stored in the internal state storage unit 16 to "recommended," indicating that the HRTF is recommended to the user, and the process proceeds from step S7 to step S8.

[0060] In step S8, the UI generation unit 18 identifies a recommended HRTF based on the internal state stored in the internal state storage unit 16, recommends the recommended HRTF to the user, and then the process ends. In recommending the recommended HRTF, the UI generation unit 18 generates a UI recommending the recommended HRTF and causes the image presentation unit 23 to present it to the user.

[0061] Of the above steps S1 to S9, steps S1 to S8 are the HRTF personal optimization process.

[0062] In the personal HRTF optimization process, the recommended HRTF to be recommended to the user is determined based on the user's evaluation responses to multiple sample sounds obtained by applying two competing HRTFs, HRTF-A and HRTF-B, to the personal optimization sound source. This facilitates personal HRTF optimization. That is, the user's personal HRTF (or an HRTF similar to the user's personal HRTF) can be obtained without, for example, preparing the large-scale equipment required for acoustic measurements to obtain the user's personal HRTF or visiting a location where such equipment is installed.

[0063] <Details of the HRTF personal optimization process>

[0064] FIG. 3 is a flowchart illustrating the details of the process of personal optimization of HRTFs.

[0065] In step S11, the information processing device 10 causes the determination unit 15 to initialize the stage number Ns, which identifies the stage in the internal state stored in the internal state storage unit 16, to 1, and the process proceeds to step S12.

[0066] In step S12, the information processing device 10 performs the preliminary phase of stage Ns, and the processing proceeds to step S13. In the preliminary phase, two HRTFs, HRTF-A and HRTF-B, are selected from the unpresented HRTFs to be pitted against each other, and a winner is determined from the two HRTFs-A and HRTF-B that have been pitted against each other based on the user's evaluation of two sample sounds to which the two HRTFs-A and HRTF-B have been applied, and this process is repeated until the number of winners reaches a predetermined number N.

[0067] In step S13, the information processing device 10 performs the final phase of stage Ns, and the process proceeds to step S 14. In the final phase, a winner (second sound reproduction parameter) is determined from among the winners of the preliminary phase based on the users' evaluation responses to the sample sounds to which the winner (the HRTF) of the preliminary phase has been applied.

[0068] In step S14, the information processing device 10 determines whether the current stage Ns is stage 2 or higher.

[0069] If it is determined in step S14 that the current stage Ns is stage 2 or higher, the process proceeds to step S15, where the information processing device 10 executes the defense phase of stage Ns, and the process proceeds to step S17. In the defense phase, a new current champion is determined based on the user's evaluation responses to the sample sounds to which the winner (the HRTF of the winner) of the final phase and the HRTF of the current champion have been applied.

[0070] On the other hand, if it is determined in step S14 that the current stage Ns is not stage 2 or higher, that is, if the current stage Ns is stage 1, the process proceeds to step S16.

[0071] In step S16, the information processing device 10 determines the winner of the final phase as the current champion. The information processing device 10 updates the HRTF status as the winner of the final phase in the internal state stored in the internal state holding unit 16 to the current champion, which represents the winner of the defense phase, and the process proceeds from step S16 to step S17.

[0072] In step S17, the information processing device 10 increments the stage number Ns by 1, and the process proceeds to step S18.

[0073] In step S18, the information processing device 10 determines whether to end the personal optimization. If it is determined in step S18 that the personal optimization should not be ended, for example, if the user has not operated the information processing device 10 to end the personal optimization, the process returns to step S12, and the same process is repeated.

[0074] On the other hand, if it is determined in step S18 that personal optimization is to be ended, for example, if the user operates the information processing device 10 to end personal optimization, the processing proceeds to step S19.

[0075] In step S19, the determination unit 15 determines the current champion as the personal optimization result, i.e., the recommended HRTF, based on the internal state stored in the internal state storage unit 16, and updates the internal state stored in the internal state storage unit 16 based on that determination, and the processing then ends.

[0076] 3, in step S18, which is the timing after the end of a stage, the user can decide at their own will whether to end personal optimization or continue (advance to the next stage). If personal optimization is to be ended, in step S19, the current champion at that time is determined as the recommended HRTF.

[0077] After Stage 1 is completed, the user can terminate personal optimization if they wish, even if they are still in the middle of the stage. In this case, the current champion at that time will be determined as the recommended HRTF.

[0078] The timing to end the personal optimization can be determined by the information processing device 10 according to some rule, rather than by the user's will. For example, a predetermined number of stages can be set in advance as a predetermined number of stages for ending the personal optimization, and the personal optimization can be ended when the predetermined number of stages have been completed. Furthermore, for example, the personal optimization can be ended at a stage when the average time required for the user to answer a question in that stage exceeds a threshold. Even if the information processing device 10 determines the timing to end the personal optimization according to some rule, the user can proceed to the next stage, i.e., continue the personal optimization, if desired. Furthermore, for example, the personal optimization can be ended when there are no unpresented HRTFs among the HRTFs stored in the database 11.

[0079] FIG. 4 is a diagram showing an example of the progress of the personal HRTF optimization process of FIG. 3 in a tournament format.

[0080] 4 shows the progression from stage 1 to stage 3. The stages are made up of a qualifying phase, a final phase, and a defense phase.

[0081] First, the preliminary phase of stage 1 is carried out. In the preliminary phase, two HRTFs, HRTF-A and HRTF-B, are selected from the unpresented HRTFs to compete against each other. Then, these two HRTFs, A and B, compete against each other. That is, two sample sounds to which HRTF-A and HRTF-B have been applied are presented to the user. After that, a winner is determined from the two HRTFs, A and B, that have been competed against each other, based on the user's response (evaluation) to the two sample sounds to which HRTF-A and HRTF-B have been applied. The above process is repeated until a specified number N of winners is reached.

[0082] In FIG. 4 , the predetermined number N is set to 4, and therefore, the process of determining a winner from the two HRTFs-A and HRTF-B that have been matched is repeated until the predetermined number N=4 is reached. In FIG. 4 , #i represents the i-th match. For example, in the first match #1, HRTF1 and HRTF2 are matched as HRTF-A and HRTF-B, and one of the HRTFs is determined to be the winner. Note that, for simplicity's sake, in FIG. 4 , only one HRTF is always determined as the winner from the two HRTFs-A and HRTF-B that have been matched. Here, since the predetermined number N is 4, the preliminary phase involves matches until the predetermined number N=4 is reached, i.e., four matches.

[0083] In FIG. 4 , the matchups (competitors) for the preliminary phase, i.e., the two HRTFs (HRTF-A and HRTF-B) to be matched in the preliminary phase, appear to be determined in advance. However, the two HRTFs (HRTF-A and HRTF-B) to be matched in the preliminary phase can be selected after the previous match has ended. The next two HRTFs to be matched can be randomly selected from unsubmitted HRTFs, for example. Alternatively, for example, the similarity between each HRTF stored in the database 11 and other HRTFs can be calculated. The next two HRTFs to be matched can be selected from unsubmitted HRTFs that have a high similarity to the HRTF that won the previous match, for example, the previous match. If a higher similarity indicates a higher degree of similarity, an unsubmitted HRTF that has a similarity to the HRTF that won the previous match that is equal to or greater than a threshold value can be selected. The next two HRTFs to be matched can be selected from unsubmitted HRTFs that have a high integrated value (additive value, multiplication value, etc.) of similarity to each of the HRTFs that won the previous matches.

[0084] As described above, when selecting two HRTFs to be competed next from unpresented HRTFs that are highly similar to the HRTF that was the winner in a past competition, as shown in the lower part of Figure 4, as the stage progresses, it is expected that HRTFs that are closer in position in the HRTF parameter space that represents HRTFs with HRTF parameters as the axis will be selected as the two HRTFs to be competed next.

[0085] The similarity between HRTFs can be an index value that represents the auditory similarity between HRTFs. For example, the similarity between HRTFs can be ITD (interaural time difference), ILD (interaural sound pressure difference), the frequency of HRTF peaks or notches, or a value based on one or more of these.

[0086] Once four winning HRTFs, which is the specified number N, have been determined in the preliminary phase of Stage 1, the final phase of Stage 1 begins. As described with reference to FIG. 3 , in the final phase, a winner is determined from the winners of the preliminary phase based on the user's responses to the sample sounds to which the winner (the HRTF) of the preliminary phase has been applied. That is, in the final phase, two HRTFs to be competed against each other are selected from the HRTFs that have been winners in the preliminary phase and have not been defeated in the final phase. Then, a competition between these two HRTFs is held, and a winner is determined based on the user's responses. The above process is repeated until one HRTF remains as the winner. The HRTF that remains as the winner in the final phase is determined as the winner of the final phase.

[0087] Once the winning HRTF is determined in the final phase, the defense phase begins. In the defense phase, the winner (or the HRTF that became the winner) of the final phase competes against the current champion HRTF, and a new current champion is determined based on the user's responses to the sample sounds to which the winner of the final phase and the current champion HRTF have been applied. That is, in the defense phase, the HRTF that became the winner of the final phase and the HRTF that is the current champion are selected as the two HRTFs to compete against each other, and a winner is determined based on the user's responses. The HRTF that won the defense phase is determined as the new current champion, and the stage ends.

[0088] In Stage 1, since there is no current champion, the defense phase does not take place, and the winner of the final phase is determined to be the new current champion, ending Stage 1. Note that Stage 1 can also be interpreted as a defense phase being held in which the non-existent current champion competes against the winner of the final phase, and the winner of the final phase being determined to be the new current champion by default.

[0089] When Stage 1 ends, Stage 2 begins. A preliminary phase, a final phase, and a defense phase of Stage 2 are performed, and thereafter, the stage progresses until the user operates the information processing device 10 to end personal optimization, for example.

[0090] <Qualification Phase>

[0091] FIG. 5 is a flowchart illustrating the details of the process of the preliminary phase of stage Ns in step S12 of FIG.

[0092] In the preliminary phase, in step S31, the selection unit 12 selects two HRTFs, HRTF-A and HRTF-B, from the unpresented HRTFs stored in the database 11 as competitors to be next put into a match.

[0093] The selection of HRTF-A and HRTF-B as competitors can be performed randomly, for example. Furthermore, the selection of HRTF-A and HRTF-B as competitors can be performed according to some rules, for example, based on the user's past answers. One method for selecting HRTF-A and HRTF-B as competitors according to some rules based on the user's past answers is, for example, to select two HRTFs to be competed next from unsubmitted HRTFs that are highly similar to the HRTF that was determined to be the winner based on the user's answers in a past competition, as described with reference to FIG. 3 . According to this method, an HRTF suitable for the user can be determined and recommended as the recommended HRTF with an expectedly fewer number of competitions (which also means fewer responses from the user) than when competitor HRTFs are selected randomly.

[0094] The selection unit 12 supplies the HRTF-A and HRTF-B of the opponent to the trial sound generation unit 14, and the process proceeds from step S31 to step S32.

[0095] In step S32, the sample sound generation unit 14 generates sample sounds by applying the HRTF-A and HRTF-B from the selection unit 12 to the personal optimization sound source stored in the sound source storage unit 13. The sample sound generation unit 14 supplies the HRTF-A sample sound and the HRTF-B sample sound to which HRTF-A and HRTF-B have been applied, respectively, to the sample sound presentation unit 21, and the process proceeds from step S32 to step S33.

[0096] In step S33, the preview sound presentation unit 21 presents the HRTF-A preview sound and the HRTF-B preview sound from the preview sound generation unit 14 to the user, i.e., a match between HRTF-A and HRTF-B. As described in Fig. 2, the image presentation unit 23 can present an image related to the preview sound to the user in addition to the preview sound presented by the preview sound presentation unit 21.

[0097] The user listens to the HRTF-A sample sound and the HRTF-B sample sound, and responds to evaluate the sample sounds by selecting one of four options, for example, "A (HRTF-A) is good," "B (HRTF-B) is good," "both are good," or "both are bad." After waiting for the user's response (evaluation), the process proceeds from step S33 to step S34.

[0098] In step S34, the answer receiving unit 22 receives the user's answer and supplies it to the determining unit 15, and the process proceeds to step S35.

[0099] In step S35 , the determining unit 15 determines the user's answer based on the user's answer received from the answer receiving unit 22 .

[0100] If it is determined in step S35 that the user's answer is "A is good," the process proceeds to step S36. In step S36, the determination unit 15 determines HRTF-A as the winner (advance to the final (phase)) and HRTF-B as the loser (eliminated), and updates the internal state stored in the internal state holding unit 16 based on this determination. The process then proceeds from step S36 to step S40.

[0101] If it is determined in step S35 that the user's answer is "B is better," the process proceeds to step S37. In step S37, the determination unit 15 determines HRTF-A as the loser and HRTF-B as the winner, and updates the internal state stored in the internal state storage unit 16 based on this determination. The process then proceeds from step S37 to step S40.

[0102] If it is determined in step S35 that the user's answer is "both are good," the process proceeds to step S38. In step S38, the determination unit 15 determines both HRTF-A and HRTF-B as winners, and updates the internal state stored in the internal state storage unit 16 based on that determination. The process then proceeds from step S38 to step S40.

[0103] If it is determined in step S35 that the user's answer is "both bad," the process proceeds to step S39. In step S39, the determination unit 15 determines that both HRTF-A and HRTF-B are losers, and updates the internal states stored in the internal state storage unit 16 based on this determination. The process then proceeds from step S39 to step S40.

[0104] An HRTF that has lost will not be selected as a competitor thereafter. However, there is a possibility that an HRTF that should have been the winner (advance to the finals) will be selected as the loser due to a user error in judgment. If the user's subsequent answers indicate a high possibility of the user's error in judgment, for example, if an HRTF highly similar to the losing HRTF is frequently selected as the winner in subsequent competitions, the losing HRTF can be selected as a competitor again. In this case, the determination unit 15 updates (changes) the status of the losing HRTF in its internal state to not presented.

[0105] In step S40, the deciding unit 15 determines whether the number of winners (number of finalists) is equal to or greater than a specified number N.

[0106] If it is determined in step S40 that the number of winners is not equal to or greater than the specified number N, the process returns to step S31, and the same process is repeated.

[0107] If it is determined in step S40 that the number of winners has reached or exceeded the specified number N, the process of the preliminary phase ends (returns from the process of the preliminary phase).

[0108] By setting the prescribed number N to a power of 2, it is possible to prevent so-called seeded winners (extra winners) from occurring in the tournament where winners compete against each other in the final phase.

[0109] In addition, in FIG. 5 , the user's answer is limited to four options: "A is good," "B is good," "both are good," and "both are bad." However, the user's answer can be selected from options other than these four options. For example, the user's answer can be limited to two options: "A is good" and "B is good," or three options: "A is good," "B is good," and "skip." If the user's answer is "skip," the HRTF status as a competitor remains undisclosed. Furthermore, if the user's answer is limited to two options: "A is good" and "B is good," the number of HRTF matches in the preliminary phase, and therefore the number of user answers, will be equal to the specified number N.

[0110] <Final Phase>

[0111] FIG. 6 is a flowchart illustrating the details of the process of the final phase of the stage Ns in step S13 of FIG.

[0112] In the final phase, in step S51, the selection unit 12 selects two HRTFs, HRTF-A and HRTF-B, as competitors to be competed next from among the HRTFs that have become finalists among the HRTFs stored in the database 11. Finalists are HRTFs that have not become losers in the final phase among the winners of the preliminary phase. Immediately after the start of the final phase, all winners of the preliminary phase have the status of finalists.

[0113] As with the preliminary phase, in the final phase, instead of deciding all of the matchups in advance, the next opponents can be selected sequentially after the previous match is over. By selecting the next opponents sequentially, it is possible to deal with cases where the number of winners in the preliminary phase (the number of players advancing to the final) is not a power of 2, for example, due to the prescribed number N not being a power of 2, or the user responding "both good" or "both bad" in the preliminary phase.

[0114] In the final phase, HRTF-A and HRTF-B can be selected as competitors, for example, randomly. Randomly selecting competitors in the final phase is the simplest and most effective method. However, it is desirable to select competitors in the final phase so as not to create a bias in the number of HRTFs competing (the number of times an HRTF is selected as a competitor). Therefore, if the matches after the start of the final phase are considered the first round, competitors for the first round can be randomly selected from the finalists who have never been selected as a competitor in the first round. After all of the finalists have been selected as competitors once, the subsequent matches can be considered the second round, and competitors for the second round can be randomly selected from the finalists who have never been selected as a competitor in the second round. Similar matches can be made, with the third, fourth, and so on.

[0115] Note that, for example, if it is possible to obtain a ranking of the suitability (suitability ranking) of the HRTFs of the finalists for the user, the HRTFs to be used as competitors can be selected based on the suitability ranking. For example, based on the suitability ranking, two HRTFs with the most distant suitability rankings can be selected as the competitors' HRTF-A and HRTF-B. For example, if the winners of the preliminary phase have four HRTFs, HRTF1 through HRTF4, and the suitability rankings of these four HRTFs are first through fourth, respectively, the HRTF1 and HRTF4, which are the most distant in suitability ranking, can be selected as the competitors first. Thereafter, among HRTF2 and HRTF3, which have never been selected by a competitor, the HRTF1 and HRTF4, which are the most distant in suitability ranking, can be selected as the competitors. The suitability rankings can be obtained, for example, by estimation using some method or by having the user input the rankings. The suitability ranking of the HRTFs for the finalists can be estimated based on, for example, the distribution of the HRTFs for the finalists in the HRTF parameter space, the similarity between the HRTFs for the finalists, etc. For example, the center of gravity of the distribution of the HRTFs for the finalists in the HRTF parameter space can be found, and the closer the HRTF is to the center of gravity, the higher the suitability ranking can be estimated.

[0116] The selection unit 12 supplies the HRTF-A and HRTF-B of the opponent to the trial sound generation unit 14, and the process proceeds from step S51 to step S52.

[0117] In step S52, similar to step S32 in the preliminary phase, the trial sound generation unit 14 generates HRTF-A trial sounds and HRTF-B trial sounds by applying the HRTF-A and HRTF-B, respectively, from the selection unit 12. The trial sound generation unit 14 supplies the HRTF-A trial sounds and HRTF-B trial sounds to the trial sound presentation unit 21, and the process proceeds from step S52 to step S53.

[0118] In step S53, similar to step S33 in the preliminary phase, the trial sound presentation unit 21 presents the HRTF-A trial sound and the HRTF-B trial sound from the trial sound generation unit 14 to the user, i.e., a competition between HRTF-A and HRTF-B. As described in Fig. 2, the image presentation unit 23 can present an image related to the trial sound to the user in addition to the trial sound presented by the trial sound presentation unit 21.

[0119] The user listens to the HRTF-A and HRTF-B sample sounds and responds to evaluate the sample sounds by choosing from three options, for example, "A (HRTF-A) is good," "B (HRTF-B) is good," or "Skip." After waiting for the user's response (evaluation), the process proceeds from step S53 to step S54.

[0120] In step S54, the answer receiving unit 22 receives the user's answer and supplies it to the determining unit 15, and the process proceeds to step S55.

[0121] In step S55 , the determining unit 15 determines the user's answer based on the user's answer received from the answer receiving unit 22 .

[0122] If it is determined in step S55 that the user's answer is "A is good," the process proceeds to step S56. In step S56, the determination unit 15 determines HRTF-A as the finalist (remainer) and HRTF-B as the loser (eliminated), and updates the internal state stored in the internal state storage unit 16 based on these determinations. The process then proceeds from step S56 to step S59.

[0123] If it is determined in step S55 that the user's answer is "B is better," the process proceeds to step S57. In step S57, the determination unit 15 determines HRTF-A as the loser and HRTF-B as the finalist, and updates the internal state stored in the internal state holding unit 16 based on this determination. The process then proceeds from step S57 to step S59.

[0124] If it is determined in step S55 that the user's response is "skip," the process proceeds to step S58. In step S58, the determination unit 15 determines both HRTF-A and HRTF-B as finalists, and updates the internal states stored in the internal state holding unit 16 based on that determination. The process then proceeds from step S58 to step S59.

[0125] In step S59, the deciding unit 15 determines whether the number of players remaining in the finals has reached one.

[0126] If it is determined in step S59 that the number of players remaining in the finals is not 1, the process returns to step S51, and the same process is repeated.

[0127] If it is determined in step S59 that the number of finalists has reached one, the determination unit 15 determines the HRTF of that one finalist as the winner of the final phase, and based on that determination, updates the internal state stored in the internal state holding unit 16. Then, the processing of the final phase ends (returns from the processing of the final phase).

[0128] If the user's answer is either "A is good" or "B is good," the number of HRTF matches in the final phase, and therefore the number of answers given by the user, will be the specified number N-1.

[0129] Here, if it is not desired to make the user aware of the flow of the personal optimization process, such as which phase the user is currently in, or if it is not desired to change the UI between the preliminary phase and the final phase, the user's answer in the final phase can be limited to the same four options as in the preliminary phase: "A is good," "B is good," "both are good," and "both are bad." In this case, when the user's answer is "both are good" or "both are bad," the information processing device 10 can internally treat the user's answer of "both are good" or "both are bad" as an answer of "skip." In this case, both of the two HRTFs of the competitors who answered "both are good" or "both are bad" remain as finalists. Note that there is also a method in which both of the two HRTFs of the competitors who answered "both are bad" are deemed losers. However, with this method, if both of the HRTFs of the competitors who answered "both are bad" are deemed losers and only one HRTF remains as a finalist, that HRTF is determined to be the winner of the final phase by a process of elimination. By process of elimination, to prevent a winner from being determined in the final phase, the two HRTFs of the competitors in the match who answered "both bad" should be designated as finalists rather than losers. Furthermore, if the user continues to select "both good" or "both bad" in the match after the number of finalists reaches two, i.e., in the match (final) to determine the winner of the final phase, a situation will arise in which no winner will be determined no matter how much time passes. To prevent this situation from occurring, the user can be prompted to select either "A is good" or "B is good" after several consecutive responses of "both good" or "both bad" in the final phase. Methods of prompting the user include, for example, displaying a text message, visually highlighting (e.g., periodically flashing) the two buttons for inputting the responses "A is good" and "B is good," or temporarily disabling the two buttons for inputting the responses "both good" and "both bad."

[0130] <Defense Phase>

[0131] FIG. 7 is a flowchart illustrating the details of the defense phase processing of stage Ns in step S15 of FIG.

[0132] In the final phase, in step S71, the selection unit 12 selects the winner of the final phase and the current champion from among the HRTFs stored in the database 11 as the two HRTFs, HRTF-A and HRTF-B, to be the next competitors. To prevent the user from knowing which of HRTF-A and HRTF-B is the current champion, it is possible to randomly select which of HRTF-A and HRTF-B will be the current champion. Note that it is also possible to make it clear to the user which of HRTF-A and HRTF-B is the current champion.

[0133] The selection unit 12 supplies the HRTF-A and HRTF-B of the opponent to the trial sound generation unit 14, and the process proceeds from step S71 to step S72.

[0134] In step S72, the trial sound generation unit 14 generates HRTF-A trial sounds and HRTF-B trial sounds by applying the HRTF-A and HRTF-B, respectively, from the selection unit 12, similar to step S32 in the preliminary phase. The trial sound generation unit 14 supplies the HRTF-A trial sounds and HRTF-B trial sounds to the trial sound presentation unit 21, and the process proceeds from step S72 to step S73.

[0135] In step S73, similar to step S33 in the preliminary phase, the trial sound presentation unit 21 presents the HRTF-A trial sound and the HRTF-B trial sound from the trial sound generation unit 14 to the user, i.e., a competition between HRTF-A and HRTF-B. As described in Fig. 2, the image presentation unit 23 can present an image related to the trial sound to the user in addition to the trial sound presented by the trial sound presentation unit 21.

[0136] The user listens to the HRTF-A sample sound and the HRTF-B sample sound, and responds to evaluate the sample sounds by choosing, for example, "A (HRTF-A) is good" or "B (HRTF-B) is good." After waiting for the user's response (evaluation), the process proceeds from step S73 to step S74.

[0137] In step S74, the answer receiving unit 22 receives the user's answer and supplies it to the determining unit 15, and the process proceeds to step S75.

[0138] In step S75 , the determining unit 15 determines the user's answer based on the user's answer received from the answer receiving unit 22 .

[0139] If it is determined in step S75 that the user's answer is "A is good," the process proceeds to step S76. In step S76, the determination unit 15 determines HRTF-A as the new current champion and HRTF-B as the loser (eliminated), and updates the internal state stored in the internal state storage unit 16 based on this determination.

[0140] If it is determined in step S75 that the user's answer is "B is better," the process proceeds to step S77. In step S77, the determination unit 15 determines HRTF-A as the loser and HRTF-B as the new current champion, and updates the internal state stored in the internal state storage unit 16 based on this determination.

[0141] After the processing of steps S76 and S77, the processing of the defense phase ends (returns from the processing of the defense phase).

[0142] If the user's answer is either "A is good" or "B is good," the number of HRTF matches in the defense phase, and therefore the number of user answers, will be one.

[0143] Here, if you do not want the user to be aware of the flow of the personal optimization process, such as which phase they are currently in, or if you do not want to change the UI for each phase, the user's answers in the defense phase can be limited to four options, the same as in the preliminary phase: "A is good," "B is good," "both are good," and "both are bad." In this case, if the user answers "both are good" or "both are bad," the defense phase can be restarted. Also, in the defense phase, as in the battle to determine the winner of the final phase described above, if the user continues to select "both are good" or "both are bad," a situation will occur in which a new reigning champion will never be determined. In the defense phase, as in the final phase described above, if the user repeatedly answers "both are good" or "both are bad," the user can be prompted to select either "A is good" or "B is good," preventing a situation in which a new reigning champion will never be determined.

[0144] As described above, by performing personal optimization in stages consisting of a qualifying phase, a final phase, and a defense phase (for Stage 1, the stage consists of the qualifying phase and the final phase), it is possible to quickly determine an HRTF that is (somewhat) suitable for the user as the recommended HRTF. That is, at the end of Stage 1, the HRTF that is most suitable for the user is determined for the current champion from among the HRTFs that have been applied to the sample sounds presented up to that point. As the stage progresses, an HRTF that is more suitable for the user than in the previous stage is determined for the current champion. Therefore, an HRTF that is somewhat suitable for the user can be determined at an earlier stage than when all of the HRTFs stored in the database 11 are compared.

[0145] Furthermore, by performing personal optimization in units of stages, the information processing device 10 becomes scalable with respect to the number of HRTFs stored in the database 11. In other words, an increase or decrease in the number of HRTFs stored in the database 11 does not affect the stages. A user who does not want to spend time on personal optimization can simply end the personal optimization after a small number of stages. On the other hand, a user who wants to be recommended an HRTF that is more suitable for them, even if it takes time for personal optimization, can simply increase the number of stages until an HRTF that satisfies the user (applied to the sample sound) is obtained.

[0146] According to the information processing device 10, the needs of various users, such as a user who wants to complete personal optimization quickly, or a user who wants to take their time to listen to sample sounds with various HRTFs applied and carefully perform personal optimization, can be flexibly met by increasing or decreasing the number of stages.

[0147] <UI>

[0148] Hereinafter, a display example of a UI generated by the UI generating unit 18 (FIG. 1) on a PC will be described, assuming that the information processing device 10 is realized by executing an application program on the PC.

[0149] FIG. 8 is a diagram showing a display example of a setting screen (window) as a first example of a UI.

[0150] The UI includes a settings screen for making various settings for personal optimization, an optimization screen for performing personal optimization, and an experience screen for experiencing sound to which the recommended HRTF obtained through personal optimization has been applied. Fig. 8 shows an example of the settings screen of the UI.

[0151] The setting screen 50 has an image display area 51 , an environment selection box 52 , an audio setting button 53 , and an optimization start button 54 .

[0152] The image display area 51 is provided in approximately the upper half of the setting screen 50. The image display area 51 displays an image, such as a photograph, of the environment for which a personally optimized HRTF is desired, selected by operating the environment selection box 52. The image displayed in the image display area 51 helps the user visualize what kind of sound will be produced in that environment, based on the size of the environment, such as a listening room, for which the HRTF is desired to be obtained, the materials of the walls, floor, and ceiling, the placement of speakers, the placement of other objects, etc.

[0153] The environment selection box 52 is provided below the image display area 51. Clicking on the environment selection box 52 displays a pull-down menu listing various environments (their names), such as movie theaters, film production studios, concert halls, live music venues, music production studios, listening rooms, and anechoic chambers. The user can select from the environments displayed in the pull-down menu of the environment selection box 52 as environments for listening to the sample sound (environments for personal optimization), i.e., as candidates for environments for listening to the sample sound and obtaining a personally optimized HRTF. The image display area 51 displays an image of the environment selected in the environment selection box 52 from among the images stored in the image storage unit 17.

[0154] The size of the environment from which you want to obtain HRTF can be selected, for example, from "large," "medium," "small," or a specific number such as "10m."You can also select the environment from which you want to obtain HRTF from a variety of different environments, such as actual movie theaters, film production studios, concert halls, live music venues, music production studios, and listening rooms.

[0155] The audio setting button 53 is provided below the environment selection box 52. The audio setting button 53 is operated when making various audio settings such as the audio device used in the PC and the sampling rate.

[0156] The optimization start button 54 is provided below the audio setting button 53. The optimization start button 54 is operated when personal optimization is to be performed.

[0157] FIG. 9 is a diagram showing a display example of an optimized screen (window) as a second example of a UI.

[0158] The optimization screen 60 is displayed in place of the setting screen 50 when the optimization start button 54 on the setting screen 50 (FIG. 8) is operated.

[0159] The optimization screen 60 has an image display area 61, a preview sound play button 62, an answer selection button 63, a button 64 for advancing to the next question, a button 65 for returning to the previous question, a progress bar 66, a preview sound selection box 67, a keyboard shortcut list display button 68, an expand button 69, and a keyboard shortcut assignment display section 70. The optimization screen 60 further has a favorite button 71, a ranking display section 72, a play button 73, an HRTF decision button 74, an optimization cancel button 75, and an optimization end button 76.

[0160] The image display area 61 occupies approximately the top two-thirds of the optimization screen 60. The image display area 61 displays images related to the sample sound to which the HRTF has been applied (hereinafter also referred to as related images). Functionally, the related images are images that help the user recognize how stereoscopically the sample sound should sound. For example, the related images may be images of the environment in which the sample sound will be listened to, selected in the environment selection box 52 on the setting screen 50. Furthermore, the related images may be images representing the sample sound or moving images synchronized with the sample sound.

[0161] For example, for a sample sound in which HRTFs are applied to pink noise and played (output) in turn from the left, front, and right speakers, an animation can be displayed in which the left, front, and right speakers are highlighted in turn in an image of the environment in which the sample sound is to be listened to, in synchronization with the playback of the sample sound. Also, for example, a still image indicating the positions of the left, front, and right speakers can be displayed in the image of the environment in which the sample sound is to be listened to.

[0162] For example, for a preview sound of a vehicle such as a car or helicopter, or an animal such as a bird or horse moving from left to right in front of you, a live-action or simple animated moving image showing the movement can be displayed.

[0163] When a user uses an xR device such as a VR headset as a display device, the related images can be displayed on the VR headset in a manner similar to how they would appear in the real world. When related images are displayed on the display screen of a PC or smartphone, it may be difficult to accurately convey three-dimensional information (such as direction and position) of a sound source to the user. For example, depending on the screen size of the display screen, there may be a difference between the actual direction of an object such as a speaker displayed on the display screen and the correct direction (true direction) of that object. By displaying related images on the VR headset in a manner similar to how they would appear in the real world, three-dimensional information of a sound source can be accurately conveyed to the user.

[0164] The preview sound playback button 62 is provided at the bottom center of the image display area 61 and is operated when preview sound is to be played (presented). The preview sound playback buttons 62 include a "Listen to A" button and a "Listen to B" button. When the "Listen to A" button is operated, preview sound with HRTF-A applied to the personal optimization sound source is played, and when the "Listen to B" button is operated, preview sound with HRTF-B applied to the optimization sound source is played.

[0165] When the "Listen to A" button displays "Listen to A" and the "Listen to B" button displays "Listen to B," it is difficult for the user to recognize the HRTF applied to the sample sound played in response to the operation of the "Listen to A" button and the "Listen to B" button (the HRTF applied to the optimization sound source in generating the sample sound). Therefore, it is possible to eliminate as much as possible preconceived notions about what HRTF has been applied, and perform a so-called blind evaluation of the sample sound (the HRTF applied to the sample sound).

[0166] For the "Listen to A" button, instead of the "A" displayed thereon, it is possible to use, for example, "ID_001" or "ID_002" as an identification number assigned to the HRTF stored in the database 11. The same applies to the "B" displayed on the "Listen to B" button. In this case, the user can recognize the HRTF (the identification number) applied to the sample sound played in response to the operation of the "Listen to A" button and the "Listen to B" button.

[0167] The answer selection button 63 is provided below the preview sound playback button 62 and is operated when inputting an answer to evaluate the preview sound. The answer selection buttons 63 include a "Choose A" button, a "Choose B" button, a "Choose Both" button, and a "Discard Both" button. The "Choose A" button, the "Choose B" button, the "Choose Both" button, and the "Discard Both" button correspond to the four-choice answers "A is good," "B is good," "Both are good," and "Both are bad," respectively, as described in FIG. 5 . The answer selection button 63 can be designed so that it cannot be operated until each of the "Listen to A" button and the "Listen to B" button serving as the preview sound playback button 62 has been operated at least once. In this case, the user can more easily understand how to operate the optimization screen 60, for example, the operating sequence of operating the preview sound playback button 62 and then operating the answer selection button 63.

[0168] The Next Question button 64 is located at the lower right of the answer selection button 63 and is operated to confirm the answer entered by operating the answer selection button 63 and proceed to the next match (question). When the Next Question button 64 is operated, two HRTFs to be played next are selected based on the answer previously entered by operating the answer selection button 63. The Next Question button 64 can be configured so that it can only be operated after operating one of the answer selection buttons 63, namely, the "Choose A" button, the "Choose B" button, the "Choose Both" button, or the "Discard Both" button. In this case, the user can more easily understand how to operate the optimization screen 60, for example, the operating sequence of operating the Next Question button 64 after operating the answer selection button 63. To make the operating sequence easier for the user to understand, the user can be notified by highlighting (periodically flashing) the button to be operated next. For example, when advancing to the next match, the "Listen to A" button may be initially highlighted, and after the user operates "Listen to A" to finish listening to the sample sound with HRTF-A applied, the "Listen to B" button may be highlighted. After the user operates the "Listen to B" button to finish listening to the sample sound with HRTF-B applied, four buttons may be highlighted: the "Choose A" button, the "Choose B" button, the "Choose Both" button, and the "Discard Both" button. After the user operates any of the "Choose A" button, the "Choose B" button, the "Choose Both" button, or the "Discard Both" button, the button 64 for proceeding to the next question may be highlighted.

[0169] The Return to Previous Question button 65 is provided at the lower left of the answer selection button 63 and is operated to return to the previous match (question). When the Return to Previous Question button 65 is operated, the HRTFs to be matched return to the HRTFs matched in the previous match, and the internal state stored in the internal state storage unit 16 is also returned to the internal state at the time of the previous match.

[0170] The progress bar 66 is located below the center of the answer selection button 63 and displays values ​​corresponding to the progress of a stage, the progress of a phase (whether the current phase is the preliminary phase, the final phase, or the defense phase), the progress within each phase, etc., using numbers, bars, etc. The value corresponding to the progress of a stage can be calculated based on a predetermined number of stages set as the number of stages at which personal optimization ends and the current number of stages. For example, for the preliminary phase, the value corresponding to the progress within each phase can be calculated based on the current number of players who have advanced to the finals and a predetermined number N, and for the final phase, the value can be calculated based on the number of players remaining in the finals.

[0171] The sample sound selection box 67 is located to the left of the answer selection button 63 and is operated when selecting a personal optimization sound source. If the user continues to listen to sample sounds that use the same sound source, the brain may adapt and it may become difficult to distinguish between sample sounds to which different HRTFs have been applied. By operating the sample sound selection box 67 and appropriately changing the personal optimization sound source, it is possible to prevent the brain from adapting and becoming unable to distinguish between the sample sounds.

[0172] A keyboard shortcut list display button 68 is provided below the preview sound selection box 67 and is operated to display a list of shortcut key assignments for keyboard shortcuts.

[0173] The expand button 69 is provided in the upper right corner of the image display area 61, and is operated when the image display area 61 is to be displayed (expanded) in a separate window from the optimized screen 60. When the image display area 61 is displayed in a separate window, the separate window can be displayed in full screen.

[0174] The keyboard shortcut assignment display unit 70 is displayed on the button to which the shortcut key is assigned, and displays the assigned shortcut key. By looking at the display on the keyboard shortcut assignment display unit 70 displayed on the button, the user can recognize the shortcut key assigned to that button. By assigning a shortcut key to each button, the user can operate each button using the keyboard instead of the mouse. Therefore, the user can concentrate on evaluating the sample sound without being distracted by operating the mouse.

[0175] The favorite button 71 is provided as part of the "Listen to A" and "Listen to B" buttons that serve as the preview sound playback buttons 62, and is a star-shaped button in FIG. 9 . When the favorite button 71 provided on the "Listen to A" button is operated, the HRTF-A applied to the preview sound played in response to the operation of the "Listen to A" button is registered as a favorite (favorite registration). If the HRTF-A applied to the preview sound played in response to the operation of the "Listen to A" button has been registered as a favorite, the star-shaped favorite button 71 provided on the "Listen to A" button is lit. The lit favorite button 71 allows the user to easily recognize that the HRTF-A applied to the preview sound played in response to the operation of the "Listen to A" button has been registered as a favorite, or that it has already been registered as a favorite. The same applies to the favorite button 71 provided on the "Listen to B" button. On the optimization screen 60, a star-shaped display with the same shape as the favorite button 71 is also displayed in the ranking display section 72 that displays the rankings of HRTFs. For HRTFs that have been registered as favorites, a star is also displayed in the ranking display section 72. Therefore, the user can check the HRTFs that have been registered as favorites using the ranking display section 72.

[0176] The ranking display unit 72 is provided in the lower right corner of the image display area 61 and displays, for example, rankings of HRTFs and the like stored in the database 11. For example, the ranking display unit 72 displays the HRTFs stored in the database 11 and the HRTFs (identification numbers) that have been used in previous battles in order of ranking. For example, the current champion may be ranked at the top of the HRTF ranking, with the more times a player has won a battle, the higher the ranking. Furthermore, HRTFs that are highly similar to the HRTFs ranked higher can be ranked close to the top ranking.

[0177] The play button 73 is displayed to the right of the HRTF (its identification number) displayed in the ranking display area 72, and is operated when playing back a sample sound to which that HRTF has been applied. Therefore, by operating the play button 73 on the optimization screen 60, it is possible to listen to a sample sound to which an HRTF other than the competing HRTF-A and HRTF-B has been applied.

[0178] The HRTF decision button 74 is displayed to the right of the play button 73 displayed in the ranking display section 72, and is operated when the HRTF displayed on the left is to be decided as the recommended HRTF. If the user finds a favorite HRTF, even before the progress bar 66 reaches 100%, the user can decide that HRTF as the recommended HRTF and end personal optimization by operating the HRTF decision button 74 to the right of the HRTF in the ranking display section 72.

[0179] An optimization stop button 75 is provided below the ranking display section 72 and is operated when personal optimization is to be stopped. When the optimization stop button 75 is operated, the screen returns from the optimization screen 60 to the setting screen 50 (FIG. 8).

[0180] The optimization end button 76 is provided below the ranking display section 72 and is operated to end personal optimization. The user can end personal optimization by operating the optimization end button 76 even before the progress bar 66 reaches 100%. In this case, the HRTF of the current champion can be determined as the recommended HRTF at that time. Note that if the optimization end button 76 is operated before the end of Stage 1, for example, the HRTF with the highest appropriateness ranking among the HRTFs remaining at that time can be determined as the recommended HRTF. In other words, if the optimization end button 76 is operated during the preliminary phase of Stage 1, the HRTF with the highest appropriateness ranking among the HRTFs of the winner can be determined as the recommended HRTF. If the optimization end button 76 is operated during the final phase of Stage 1, the HRTF with the highest appropriateness ranking among the HRTFs of the finalists can be determined as the recommended HRTF. Furthermore, the optimization end button 76 can be disabled (inoperable) until the end of Stage 1.

[0181] FIG. 10 is a diagram showing a display example of a trial screen (window) as a third example of a UI.

[0182] The trial screen 90 is displayed in place of the optimization screen 60, for example, when the recommended HRTF is determined on the optimization screen 60 and personal optimization is completed.

[0183] The trial screen 90 has an optimization result display area 91 , a trial button 92 , a selection box 93 , a save button 94 , an end button 95 , and a continue button 96 .

[0184] The optimization result display area 91 is provided in approximately the upper half of the experience screen 90, and displays the recommended HRTF (its identification number).

[0185] The Experience button 92 is located below the optimization result display area 91 and is operated when experiencing stereophonic content, which is sound obtained by applying the recommended HRTF displayed in the optimization result display area 91 to a specific sound content. Operating the Experience button 92 transitions the experience screen 90 to a presentation screen displaying stereophonic content, where the stereophonic content is presented with the recommended HRTF applied to the specific sound content. On the presentation screen, the user can select from multiple sound content candidates the sound content to which the recommended HRTF is to be applied. Also, on the presentation screen, the user can load their own sound content as sound content to which the recommended HRTF is to be applied. Furthermore, on the presentation screen, the user can select, for example, multiple top-ranked HRTFs and receive presentations of stereophonic content obtained by applying each of the multiple HRTFs. In this case, the user can compare multiple HRTFs. Additionally, on the application screen, sound content can be played in stereo without applying an HRTF. In this case, the user can compare the case where the sound content is played in stereo with the case where the sound content is played in stereo without applying an HRTF.

[0186] A selection box 93 is provided below the experience button 92, and is operated when the user wants to select the HRTF that he or she ultimately desires. In the selection box 93, the user can select the HRTF that he or she ultimately desires from the HRTFs displayed in the ranking display section 72 (FIG. 9), for example.

[0187] A save button 94 is provided to the right of the selection box 93 and is operated when the HRTF selected in the selection box 93 is to be saved as a file in a predetermined format. When the information processing device 10 is configured as a server-client system, the HRTF file selected in the selection box 93 is downloaded from the server and saved in the client.

[0188] The end button 95 is located at the bottom of the experience screen 90 and is operated to end the experience screen 90 and return to the setting screen 50 ( FIG. 8 ). The continue button 96 is located to the right of the experience button 92 and is operated to continue personal optimization. When the continue button 96 is operated, the screen returns to the optimization screen 60 of FIG. 9 , and one more stage of personal optimization is performed. When returning to the optimization screen 60, the progress bar 66 returns to 0%, and the progress rate within the additional stage is displayed on the progress bar 66. When the additional stage is completed, the screen transitions from the optimization screen 60 to the experience screen 90 of FIG. 10. When the continue button 96 is operated again on the experience screen 90, the screen returns to the optimization screen 60, and one more stage of personal optimization is performed again. Note that if no unpresented HRTFs remain, the continue button 96 is disabled.

[0189] <Modification>

[0190] FIG. 11 is a diagram showing another example of the progress of the personal HRTF optimization process in a tournament format.

[0191] In FIG. 4 , two HRTFs, HRTF-A and HRTF-B, are pitted against each other in a single match. However, three or more HRTFs can be pitted against each other at the same time. When pitting three or more HRTFs against each other at the same time, the user can select any number of good HRTFs from the three or more HRTFs that have been pitted against each other. Selecting any number of good HRTFs includes selecting only one best HRTF or selecting multiple good HRTFs. Furthermore, selecting all of the pitted HRTFs as good or all of the pitted HRTFs as bad is also possible. Specifically, for example, when pitting three HRTFs, HRTF-A, HRTF-B, and HRTF-C, the user can select from five options: “A (HRTF-A) is good,” “B (HRTF-B) is good,” “C (HRTF-C) is good,” “all good,” or “all bad,” or four options: “A is good,” “B is good,” “C is good,” or “skip.”

[0192] When three or more HRTFs are pitted against each other at once, the HRTFs that receive a "good" response in the preliminary phase can be declared the winner and advance to the final phase. HRTFs that receive a "good" response in the final phase can be declared finalists. If the number of HRTFs pitted against each other at once is M, then in the final phase, if the number of finalists becomes M', which is less than M, the M' finalists can be pitted against each other. In the defense phase, the HRTF of the current champion can be pitted against the HRTF of the winner of the final phase.

[0193] The number of HRTFs that are pitted against each other at one time can be changed for each phase or stage. For example, in the preliminary phase, the number of HRTFs that are pitted against each other at one time can be a predetermined number equal to or greater than three, and in the final and defense phases, the number of HRTFs that are pitted against each other at one time can be set to two. In this way, in the preliminary phase, the user can roughly listen to the trial sounds to which each HRTF has been applied for many HRTFs and eliminate HRTFs that do not suit the user, and then in the final and defense phases, the user can carefully listen to the trial sounds to which each HRTF has been applied for the remaining HRTFs and narrow down the choices, thereby enabling the user to efficiently find an HRTF that suits them.

[0194] FIG. 11 shows a case where three HRTFs compete in one match in the preliminary phase, and two HRTFs compete in one match in the final and defense phases.

[0195] In theory, all of the HRTFs stored in the database 11 can be selected as the HRTFs to be compared at one time in the selection unit 12. However, from the perspective of determining HRTFs that are somewhat suitable for the user at an earlier stage than when all of the HRTFs stored in the database 11 are compared, it is possible to select a number of HRTFs to be compared at one time that is fewer than the number of HRTFs stored in the database 11, for example, two, three, or four. In particular, by selecting two HRTFs to be compared at one time and having the user evaluate two sample sounds to which each of the two HRTFs has been applied, even a novice user can easily make an evaluation, as described with reference to FIG. 1 .

[0196] There are no particular restrictions on the number of directions of the sound source positions of the HRTFs, as long as the HRTFs can be stored in the database 11. For example, it is possible to adopt HRTFs that correspond to widely used speaker layouts such as 2.0, 5.1, 7.1, 5.1.2, 5.1.4, 7.1.2, 7.1.4, 9.1.6, 13.1, and 22.2 channels, or it is also possible to adopt HRTFs that correspond to unique speaker layouts.

[0197] The user's answer may be a continuous value rather than a discrete value selected from options such as "A is good," "B is good," "both are good," or "both are bad." For example, "A is good" may correspond to -1.0 and "B is good" may correspond to 1.0, and the user may answer using a continuous value between -1.0 and 1.0. In this case, the next two HRTFs to be matched may be selected based on a value obtained by applying the continuous value of the user's answer to the similarity with HRTF-A and HRTF-B. For example, a weighted sum may be applied to the similarity with HRTF-A that is inversely proportional to the continuous value's proximity to -1.0, and a weighted sum may be applied to the similarity with HRTF-B that is inversely proportional to the continuous value's proximity to 1.0. The HRTF with the larger weighted sum may be selected as the next HRTF to be matched.

[0198] Even if the user gives the same answer, the next HRTF to be compared may be changed based on the time taken to give the answer, the number of times the user listened to and compared the sample sounds to which HRTF-A and HRTF-B were applied, etc. For example, even if the user answers "A is good," if the user took a long time to give the answer or if the user listened and compared the sample sounds to which HRTF-A and HRTF-B were applied many times, it can be assumed that the user was unsure of their answer (decision), and the degree to which the answer is influenced can be reduced.

[0199] The user's answer can be obtained by having the user directly input which of the sample sounds to which HRTF-A and HRTF-B are applied is better, or by methods other than such direct input. For example, the competition between HRTF-A and HRTF-B can be in the form of a game, such as a game in which the user points in the direction from which a sound is heard, and the user's performance in the game, for example, the error between the correct direction (the direction of the sound source position) and the direction pointed by the user, can be used to indirectly obtain the user's answer as to which of HRTF-A and HRTF-B is more suitable for the user.

[0200] The HRTFs to be compared can be changed based on user information acquired by the user input or by sensing the user, such as by photographing the user. User information includes, for example, the user's attribute information, physical characteristics, and measurement data related to the user. Attribute information includes, for example, gender, age, and race. Physical characteristics include, for example, head width and head circumference. Measurement data related to the user includes, for example, head photographs, ear photographs, HRTFs from a single direction, personal headphone characteristics (characteristics from the headphone speaker unit to the ear canal), and other results of simple acoustic measurements or photography that can be performed at home. For example, the HRTF database 11 can be divided into databases of HRTFs suitable for each race, such as Asians, Blacks, and Whites. Then, based on the user information, an HRTF database suitable for the user's race can be selected as the database from which the HRTFs to be compared are selected.

[0201] Furthermore, based on user information, it is possible to change the number of times sample sounds are compared (such as the number of matches) required for personal optimization, such as the prescribed number N or the prescribed number of stages. For example, statistics on the number of times comparisons are performed in personal optimization can be kept for each race, and the statistics can be used to set the number of times the comparisons are required for personal optimization based on the race.

[0202] The information processing device 10 can improve the algorithm for subsequent personal optimization based on past user responses. For example, there is a high probability that HRTFs that the same user has answered as good are highly similar. Therefore, the similarity between HRTFs can be improved to a more likely value based on the user's responses.

[0203] Any HRTF can be used as the HRTF to be stored in the database 11. For example, HRTFs measured on humans, HRTFs obtained by acoustic simulation, HRTFs generated using certain parameters, etc. can be used. Furthermore, HRTFs to be selected as HRTFs to be matched can be stored in the database 11 in advance or newly generated. For example, by performing principal component analysis on HRTFs previously stored in the database 11 and changing the principal component scores of each principal component, a new HRTF not stored in the database 11 can be generated. HRTFs can be generated in advance or each time a HRTF to be matched is selected. Furthermore, for HRTFs to be selected as HRTFs to be matched, for example, HRTFs for multiple environments can be generated by combining RTFs (room impulse response (RIR) (reverberation characteristics of the environment)) for each of the multiple environments with a single HRTF measured in an anechoic chamber. In this case, it is not necessary to prepare HRTFs for various environments.

[0204] This technology can be applied not only to personal optimization of HRTFs in stereophonic sound reproduction using headphones, but also to personal optimization of HRTFs in transaural reproduction.

[0205] Furthermore, the sound reproduction parameters that are the subject of personal optimization are not limited to HRTFs. For example, parameters related to sound reproduction that differ from person to person due to physical or preference reasons, such as parameters of hearing aids or sound collectors, parameters of noise canceling or noise reduction, and parameters of equalizers, can be used as the sound reproduction parameters that are the subject of personal optimization.

[0206] Furthermore, this technology can be used in situations such as finding the perfect instrument sound in music production, or the perfect sound effect in film or game sound production. For producers, searching through a huge database for the instrument sound or sound effect they have in mind is a time-consuming task. In contrast, this technology allows users to quickly find the instrument sound or sound effect from the database that most closely matches their image by pitting instrument sounds and sound effects against each other.

[0207] <Configuration example of another embodiment of information processing device>

[0208] FIG. 12 is a block diagram showing an example configuration of another embodiment of an information processing device to which the present technology is applied.

[0209] Personal optimization of HRTFs is an important technology that significantly affects the experience of 3D sound reproduction using headphones. The most accurate personal optimization is based on acoustic measurement, but this is time-consuming and costly. For this reason, various methods have been proposed to easily perform personal optimization of HRTFs.

[0210] Here, in this specification, the accuracy of personal optimization means, when a user listens to a sound (virtual sound source) in which HRTFs for various directions have been applied to the sound source, how close the sound sounds when the sound source is played from an actual speaker, from the perspective of various criteria such as the direction of localization, sense of distance, timbre, etc. The accuracy of personal optimization may refer to the accuracy (good or bad) for each HRTF direction or for each criterion, or it may refer to the overall accuracy that summarizes the various directions and criteria.

[0211] The information processing device 10 in FIG. 1 provides a method for easily performing personal optimization of HRTFs, in which the user listens to and compares sample sounds to which HRTFs in the database 11 have been applied, and selects the HRTF that suits them (the way it sounds to their liking).

[0212] The information processing device 10 selects two HRTFs from the database 11 and compares the two HRTFs. That is, two sample sounds to which the two HRTFs have been applied are presented to the user. The user listens to and compares the two sample sounds and responds by rating which HRTF (the sample sound to which the HRTF has been applied) suits them best, for example, which HRTF (the sample sound to which the HRTF) matches the originally intended sound localization direction, distance, timbre, and other characteristics. Based on the user's response, the information processing device 10 uses a unique algorithm to select a new combination of two HRTFs from the database 11 and compares them. The information processing device 10 repeats the selection of two HRTFs based on the user's response and the comparison of the two HRTFs as necessary, ultimately determining a recommended HRTF to be recommended to the user from the HRTFs in the database 11.

[0213] The personal optimization method by the information processing device 10 can directly consider the user's subjective hearing, compared to other simple personal optimization methods, such as methods of estimating HRTFs based on photographs of the user's ears or head, 3D (dimensional) scans, etc. Therefore, the personal optimization method by the information processing device 10 can obtain HRTFs that are subjectively satisfactory to the user. How humans perceive HRTFs (sounds to which HRTFs have been applied) is a complex issue involving the auditory mechanisms of the human ear and brain, and it is difficult to predict how such perceptions will be perceived from photographs of the ears or head, 3D scans, etc. The personal optimization method by the information processing device 10 involves the user listening to and comparing sample sounds to which HRTFs have been applied, and has the advantage of being able to directly consider (access) how humans perceive HRTFs (sounds to which HRTFs have been applied).

[0214] However, it may be difficult to obtain sufficient accuracy in personal optimization even with personal optimization of the information processing device 10. Furthermore, in personal optimization of the information processing device 10, the user is asked to listen to and compare sample sounds to which HRTFs have been applied, and to respond to an answer as to which HRTF they feel suits them best, which may make it difficult to make a judgment when responding.

[0215] One reason why it is difficult to achieve sufficient accuracy in personal optimization is that there are limited HRTFs in the database 11. Therefore, for all directions in 4π space, for example, for each direction expressed by an azimuth angle with an integer value ranging from -180 degrees to +180 degrees and an elevation angle with an integer value ranging from -90 degrees to +90 degrees, the probability that an HRTF that suits the user from the perspective of all criteria such as direction of localization, sense of distance, and tone color is present in the database 11 is low.

[0216] Another reason why it is difficult to achieve sufficient accuracy in personal optimization is that it is time-efficient to have a user listen to and compare HRTFs for all directions (sample sounds to which each is applied) from the perspective of all various criteria. From the perspective of time efficiency, the information processing device 10 recommends, for example, focusing on listening to and comparing HRTFs for the front (HRTFs for front sound source positions) (HRTFs that generate sounds emitted from front sound source positions), which are prone to differences in accuracy. Therefore, accuracy is not guaranteed for directions (HRTFs) that the user has not compared. Furthermore, accuracy is not guaranteed for criteria that the user did not consider when answering the HRTF that they feel suits them.

[0217] On the other hand, having a user listen to and compare HRTFs (sample sounds to which the HRTFs have been applied) from many directions, or having a user consider all of the various criteria when answering which HRTF they feel is best, makes it difficult for the user to make a decision when answering. For example, when comparing two HRTFs (sample sounds to which the HRTFs have been applied), the user may feel that one HRTF is better for a certain direction or certain criteria, but the other HRTF is better for another direction or other criteria, and may perceive each of the two HRTFs to be compared (pitched against each other) as having its own advantages and disadvantages. In this case, it is difficult for the user to determine which is better when answering the question, and the user may be unsure how to answer.

[0218] Therefore, in the information processing device 110 in Fig. 12, a recommended HRTF is determined by the user listening to and comparing two HRTFs, and the recommended HRTF is also adjusted. By adjusting the recommended HRTF, the number of HRTF options can be substantially increased. Furthermore, by adjusting the direction (of the HRTF) or the criteria for judgment that the user has not compared, the accuracy of personal optimization of the HRTF can be improved.

[0219] Furthermore, the information processing device 110 can determine (narrow down) the direction and judgment criteria of the HRTFs to be compared by listening to them, taking into account the details of the adjustment of the recommended HRTF, thereby enabling the user to easily answer questions without hesitation.

[0220] 12, the information processing device 110 includes a database 111, a sound source storage unit 112, a listening comparison unit 113, an adjustment unit 114, a UI generation unit 115, and a UI unit 116. Similar to the information processing device 10 in FIG. 1, the information processing device 110 functions as an optimization device that can easily perform personal optimization of HRTFs.

[0221] The database 111 stores a large number (plurality) of HRTFs, similar to the database 11 in FIG.

[0222] The sound source storage unit 112 stores a personal optimization sound source, similar to the sound source storage unit 13 in FIG.

[0223] The listening comparison unit 113 performs a listening comparison process using the HRTF stored in the database 111 and the personal optimization sound source stored in the sound source storage unit 112 .

[0224] The listening comparison process involves a user listening to and comparing sample sounds obtained by applying a plurality of HRTFs to a given sound source, and determining a (tentative) recommended HRTF based on the user's response to an evaluation of the sample sounds.

[0225] In the listening comparison process, similar to the information processing device 10, for example, a plurality of HRTFs, for example, two HRTFs, are selected from the HRTFs stored in the database 111, and a comparison is performed in which trial sounds obtained by applying the two HRTFs to a personal optimization sound source stored in the sound source storage unit 112 are presented to the user. The user listens to and compares the two trial sounds and responds to evaluate the trial sounds (the HRTFs applied to them). In the listening comparison process, the selection of two HRTFs to be compared next and the comparison of the two HRTFs are repeated as necessary, and a recommended HRTF to be recommended to the user is determined based on the user's responses in each comparison. For example, the process described with reference to FIG. 3 can be used as the listening comparison process.

[0226] In the listening comparison process of the listening comparison unit 113 , the UI unit 116 presents the sample sound to the user, and the user's response is received by the UI unit 116 and supplied to the listening comparison unit 113 .

[0227] The adjustment unit 114 performs adjustment processing using the HRTF stored in the database 111 and the personal optimization sound source stored in the sound source storage unit 112 .

[0228] The adjustment process is a process of adjusting a target HRTF, which is an HRTF for an individual determined in some way, such as a recommended HRTF determined in the listening comparison process, as the target HRTF to be adjusted.

[0229] In the adjustment process, a listening sound obtained by applying the pre-adjustment target HRTF to the personal optimization sound source stored in the sound source storage unit 112 is presented to the user. Furthermore, the target HRTF is adjusted in response to the user's manipulation to adjust the target HRTF, and a listening sound obtained by applying the adjusted target HRTF to the personal optimization sound source is presented to the user. The user listens to the listening sound to which the adjusted target HRTF has been applied, and repeats the adjustment of the target HRTF as necessary.

[0230] In the adjustment process of the adjustment unit 114 , the UI unit 116 presents the sample sound to the user, and operation information indicating the user's operation to adjust the target HRTF is received by the UI unit 116 and supplied to the adjustment unit 114 .

[0231] The UI generation unit 115 generates an image as a UI and supplies it to the UI unit 116 .

[0232] The UI unit 116 is configured in the same manner as the preview sound presentation unit 21, response reception unit 22, and image presentation unit 23 in Figure 1, and presents sounds such as preview sounds to the user and presents images such as UI to the user, as well as receiving input such as responses from the user and HRTF adjustment operations, and supplies them to the necessary blocks.

[0233] <Processing of information processing device 110>

[0234] FIG. 13 is a flowchart illustrating an example of processing by the information processing apparatus 110 in FIG.

[0235] In step S101, the information processing device 110 determines whether or not this is the first time the user has performed personal optimization.

[0236] If it is determined in step S101 that this is the first time the user is performing personal optimization, the process proceeds to step S102.

[0237] In step S102, the listening comparison unit 113 performs a listening comparison process, for example, a process similar to the personal optimization (FIG. 2) of the information processing device 10 in FIG. 1, to determine a provisional recommended HRTF, and the process proceeds to step S103.

[0238] In step S103, the adjustment unit 114 sets the provisional recommended HRTF obtained in this listening comparison process as the target HRTF, and the process proceeds to step S106.

[0239] On the other hand, if it is determined in step S101 that this is not the first time the user has performed personal optimization, the process proceeds to step S104.

[0240] In step S104, the listening comparison unit 113 performs the listening comparison process again and determines whether or not to re-determine a provisional recommended HRTF.

[0241] If it is determined in step S104 that the listening comparison process should be performed again, for example, if the user performs an operation to perform the listening comparison process again, the process proceeds to step S102, and the same process is performed thereafter.

[0242] If it is determined in step S104 that the listening comparison process will not be performed again, for example, if the user performs an operation so as not to perform the listening comparison process, the process proceeds to step S105.

[0243] In step S105, the adjustment unit 114 sets the provisional recommended HRTF obtained in the previous listening comparison process as the target HRTF, and the process proceeds to step S106.

[0244] In step S106, the adjustment unit 114 determines whether or not to perform adjustment processing.

[0245] If it is determined in step S106 that adjustment processing is to be performed, for example, if the user performs an operation to perform adjustment processing, the processing proceeds to step S107.

[0246] In step S107, the adjustment unit 114 performs an adjustment process to adjust the target HRTF, and the process proceeds to step S108.

[0247] In step S108, the adjustment unit 114 determines the adjusted target HRTF as the final recommended HRTF, and the process proceeds to step S110.

[0248] On the other hand, if it is determined in step S106 that the adjustment process is not to be performed, for example, if the user performs an operation so as not to perform the adjustment process, the process proceeds to step S109.

[0249] In step S109, the adjustment unit 114 determines the target HRTF as the final recommended HRTF, and the process proceeds to step S110.

[0250] In step S110, the UI generation unit 115 performs a recommendation process for recommending the recommended HRTF to the user, for example, a process for generating a UI for recommending the recommended HRTF and presenting it to the user, and then the process ends.

[0251] Among the above-described steps S101 to S110, the steps S101 to S109 are the processing for personal optimization of HRTFs in the information processing device 110.

[0252] In the personal HRTF optimization process in the information processing device 110, a listening comparison process is performed in the same manner as in the information processing device 10 in Fig. 1, i.e., a provisional recommended HRTF to be recommended to the user is determined based on the user's evaluation responses to sample sounds applied with two HRTFs, HRTF-A and HRTF-B, that are compared (listened to and compared).The provisional recommended HRTF is then used as a target HRTF, and an adjustment process is performed as necessary to adjust the target HRTF in accordance with the user's adjustment operation.

[0253] In the information processing device 110, the process of determining a provisional recommended HRTF that will serve as a target HRTF is not limited to the same listening comparison process as in the information processing device 10. That is, for example, an HRTF estimated by a method of estimating an HRTF based on a photograph of the user's ears or head, a 3D scan, or the like can be set as the provisional recommended HRTF as the target HRTF.

[0254] <Listening comparison processing>

[0255] FIG. 14 is a flowchart illustrating an example of the listening comparison process in step S102 of FIG.

[0256] In step S121, the listening comparison unit 113, like the information processing device 10, selects two HRTFs to compare, i.e., two HRTFs to be compared, HRTF-A and HRTF-B, from the HRTFs stored in the database 111, and compares them (presents them to the user). Furthermore, in step S121, the listening comparison unit 113 presents the user with a first judgment criterion, which is a judgment criterion for answering which of the two HRTFs, HRTF-A and HRTF-B, the user feels is most suitable, and the process proceeds from step S121 to step S122.

[0257] In step S122, the listening comparison unit 113 waits for the user to listen to and compare the two presented HRTFs, A and B (or the sample sounds to which HRTF-A and HRTF-B have been applied), and give an answer in accordance with the first judgment criterion, and then receives the answer given in accordance with the first judgment criterion (via the UI unit 116), and the processing proceeds to step S123.

[0258] In step S123, the listening comparison unit 113 determines whether or not the condition for ending the listening comparison process has been met.

[0259] If it is determined in step S123 that the conditions for terminating the listening comparison process are not met, for example, if the user has not performed an operation to terminate the listening comparison process, the process returns to step S121 and the same process is repeated.

[0260] Furthermore, if it is determined in step S123 that the condition for ending the listening comparison process is satisfied, for example, if the user performs an operation to end the listening comparison process, the process proceeds to step S124.

[0261] In step S124, the listening comparison unit 113 determines a provisional recommended HRTF based on the user's response and ends the listening comparison process. For example, the listening comparison unit 113 determines the HRTF that is the winner when the end condition for the listening comparison process is satisfied as the provisional recommended HRTF.

[0262] Here, one HRTF stored in the database 111 refers to a (set of) HRTFs for all directions for a specific person who is the subject of a competition in the listening comparison process, and the HRTFs for all directions are assigned the same identification number (ID). The same applies to the database 11 in Figure 1. For example, if the database 111 stores the HRTFs for all directions actually measured for 30 people, the HRTF (for all directions) for the ith person is assigned the identification number #i.

[0263] The specific person may be a real person or a virtual person. For a real person, the HRTFs for all directions of the real person can be obtained, for example, by performing actual measurements or acoustic simulations of the real person. For a virtual person, the HRTFs for all directions of the virtual person can be obtained, for example, by performing acoustic simulations using various information about the virtual person, such as attributes such as gender, age, and race, and physical characteristics such as head size.

[0264] The omnidirectional HRTF for a specific person refers to one or more predetermined directions, including the front (of the user), centered on the specific person in 4π space. The omnidirectional direction can be, for example, a single direction such as the front, or multiple directions such as the directions of each speaker in a specific speaker layout such as 5.1 channels, or directions whose azimuth and elevation angles are integer values ​​or multiples of predetermined integer values.

[0265] An HRTF for a certain direction means an HRTF that makes a sound source (position) appear to be in that direction, and the direction of the HRTF (the direction of the HRTF's sound source position) means the direction in which the HRTF makes the sound source appear to be.

[0266] The (provisional) recommended HRTF is one HRTF to be matched, that is, a set of HRTFs for all directions of a person.

[0267] In the listening comparison process, in the HRTF comparison, it is possible to select HRTFs with different identification numbers for each direction. For example, for the front direction (C) (front of the user), an HRTF with identification number #i=5 is selected, and for the left front (L) and right front (R), an HRTF with identification number #i=10 is selected. However, if HRTFs with different identification numbers are selected for each direction in the HRTF comparison, the consistency of the HRTFs for each direction may be lost, creating a sense of discomfort. Furthermore, the number of HRTF combinations to be selected in the HRTF comparison may become enormous, which may increase the number of times the user must answer, placing a heavy burden on the user. For this reason, in the HRTF comparison, HRTFs with the same identification number, i.e., HRTFs of the same person, are selected for each direction.

[0268] Furthermore, the criteria for judging the accuracy of personal optimization of HRTF, that is, the criteria (elements) for judging the proximity of the sound to which the HRTF has been applied to the sound source when it is played from an actual speaker, include, for example, the direction of localization (azimuth direction (horizontal), elevation direction (vertical)), sense of distance, tone color, etc. (whether they suit the user).

[0269] As a judgment criterion (first judgment criterion) in the listening comparison process, for example, a sense of distance (whether there is a sense of distance or not) (whether the sense of distance is correct) can be adopted from among the direction of localization, sense of distance, and tone color.

[0270] Regarding sound localization direction and timbre, the general relationship between HRTFs and the way in which the HRTFs (or the sounds to which the HRTFs are applied) are heard has become somewhat clear. Therefore, it is relatively easy to determine how to adjust the HRTFs so that the sound localization direction (how it is heard (the direction perceived as the direction of the sound source)) and timbre (how it is heard) change to a certain way of hearing, i.e., to design an algorithm for adjusting the HRTFs. It is also possible to design a process that homogenizes the sound localization direction and timbre of each HRTF stored in advance in the database 111 to a certain extent.

[0271] Homogenizing the direction of HRTF localization means making the deviation of the direction of HRTF localization (the deviation of the direction of HRTF localization in a certain direction relative to another direction (the amount and direction of deviation)) uniform (to the same extent) for each HRTF. Homogenizing the timbre of HRTFs means making the timbre of HRTFs uniform for each HRTF. For example, in a case where HRTFs that produce a dull timbre and HRTFs that produce a sharp timbre are mixed, this means making the timbre of each HRTF uniform to either a dull timbre or a sharp timbre, or a timbre that is neither dull nor sharp. Methods for homogenizing the direction of HRTF localization and timbre include, for example, aligning the interaural time difference (ITD) or interaural sound pressure difference (ILD) of the HRTFs, or setting a constant value for the ratio of low-frequency components to high-frequency components in the frequency characteristics of the HRTFs. When there are more high-frequency components than low-frequency components, the tone tends to be perceived as sharp, and when there are fewer high-frequency components, the tone tends to be perceived as dull.

[0272] On the other hand, in order to clarify the relationship between HRTFs and how they are heard, it is necessary to unravel the mechanisms of hearing in the human ear and brain, and this relationship remains largely unknown.

[0273] Therefore, although it is possible to adjust the HRTF to effect a certain change in the direction of localization and timbre, it is difficult to adjust the HRTF to effect a certain change in the sense of distance.

[0274] Therefore, in the listening comparison process, in order to have the user answer based on the judgment criterion of (there is) a sense of distance that makes it difficult to adjust the HRTF, the sense of distance (is) presented to the user as the first judgment criterion that is used to judge when answering the HRTF.When presenting the first judgment criterion, a message can also be presented to the user that, since small discrepancies in the direction of localization and timbre can be adjusted later, there is no need to worry too much about them at the moment.

[0275] In the listening comparison process, by having the user respond using the sense of distance (or whether there is such a sense of distance) as a criterion for judgment, the HRTF whose sense of distance matches the user can be determined as the provisional recommended HRTF.

[0276] By appropriately combining the criteria for the comparison listening process with the criteria for the adjustment process performed after the comparison listening process, it is possible to achieve higher accuracy in the personal optimization of HRTFs.

[0277] For example, in the adjustment process in which a provisional recommended HRTF is used as the target HRTF, the user can adjust the HRTF using criteria different from those used in the listening comparison process, for example, by having the user adjust the HRTF using the direction of positioning in which the HRTF can be adjusted and / or the timbre (being a match) as criteria, thereby achieving more accurate personal optimization of the HRTF.

[0278] In the adjustment process, one or both of the direction of localization and the tone color can be used as the criteria for judgment when performing adjustment operations (by the user), but in the following, both the direction of localization and the tone color will be used as the criteria for judgment when performing adjustment operations.

[0279] <Adjustment processing>

[0280] FIG. 15 is a flowchart illustrating an example of the adjustment process in step S107 of FIG.

[0281] In step S131, the adjustment unit 114 adjusts the localization direction and timbre of the HRTF for direction #1 in the set of omnidirectional HRTFs assigned a certain ID as the target HRTF in response to a user operation, and the process proceeds to step S132. Adjusting the localization direction and timbre of the HRTF means adjusting the HRTF so as to adjust the localization direction and timbre of the sound to which the HRTF is applied (so that the localization direction and timbre undergo a predetermined change).

[0282] Here, in the adjustment process, one or more N directions (a predetermined plurality of directions) of all directions are predetermined as directions in which HRTFs are adjusted in response to user operation, and the N directions are assigned numbers #n in the order in which they are adjusted. The N directions may be all or some of the directions in all directions. The directions in which HRTFs are adjusted in response to user operation are also referred to as operation adjustment target directions. The operation adjustment target directions and the number (N) of operation adjustment target directions can be arbitrarily set by the user, for example.

[0283] In step S132, the adjustment unit 114 sets the variable n, which indicates the operation adjustment target direction, to 1 as an initial value, and the process proceeds to step S133.

[0284] In step S133, the adjustment unit 114 determines whether the variable n is equal to the total number N of operation adjustment target directions.

[0285] If it is determined in step S133 that the variable n is not equal to the total number N of directions to be adjusted by operation, i.e., if adjustment of the localization direction and timbre of all HRTFs for each of the N directions to be adjusted by operation has not yet been completed, the process proceeds to step S134.

[0286] In step S134, the adjuster 114 increments the variable n by 1, and the process proceeds to step S135.

[0287] In step S135, the adjustment unit 114 sets (calculates) an initial value for the HRTF adjustment for direction #n using the HRTF adjustment values ​​for some or all of directions #1 to #n-1, and the process proceeds to step S136. For example, the adjustment unit 114 sets the HRTF adjustment value for direction #n-1, which was adjusted immediately before, as the initial value for the HRTF adjustment for direction #n.

[0288] In step S136, the adjustment unit 114 adjusts the direction of HRTF localization and tone color for direction #n in the set of omnidirectional HRTFs assigned a certain ID as the target HRTF in response to the user's operation, and the processing returns to step S133.

[0289] Then, in step S133, if it is determined that the variable n is equal to the total number N of directions to be adjusted by operation, that is, if adjustment of the localization direction and timbre of all HRTFs for each of the N directions to be adjusted by operation has been completed, processing proceeds to step S137.

[0290] In step S137, the adjustment unit 114 uses the HRTF adjustment values ​​for some or all of the directions #1 to #N to determine (calculate) the HRTF localization direction and adjustment values ​​for adjusting the tone for each of the directions other than directions #1 to #N (directions excluding directions #1 to #N) out of all directions, and the processing proceeds to step S138.

[0291] In step S138, the adjustment unit 114 adjusts the HRTF for each of the directions other than directions #1 to #N among all directions using the adjustment value determined in step S137 (generating HRTFs with adjusted localization direction and timbre), and then ends the adjustment process.

[0292] Here, the HRTF (localization direction and timbre) for each of directions #1 to #N, which are the directions to be adjusted by operation among all directions, is adjusted in response to user operation, but it is possible to adjust the HRTF for each of all directions in response to user operation. However, adjusting the HRTF for each of all directions in response to user operation takes time and places a heavy burden on the user.

[0293] 15 , the HRTFs for each of the N directions #1 to #N that are to be adjusted by operation out of all directions are adjusted in response to user operation, and in steps S137 and S138, the adjusted HRTF values ​​(or adjusted HRTFs) for each of the directions other than the directions #1 to #N that are to be adjusted by operation out of all directions are calculated using the adjusted HRTF values ​​(or adjusted HRTFs) for some or all of the directions that are to be adjusted by operation. This allows the HRTF adjustment to be performed in a short time and with a reduced burden on the user.

[0294] As a method for calculating the HRTF adjustment values ​​for each of the directions other than the direction to be adjusted by operation using the HRTF adjustment values ​​for some or all of the directions to be adjusted by operation, it is possible to employ interpolation such as linear interpolation or spline interpolation.In addition, it is possible to employ a method for estimating the HRTF adjustment values ​​by using, as learning data, a machine learning model that has performed machine learning such as deep learning using the HRTF adjustment results from multiple users.

[0295] In a method of estimating HRTF adjustment values ​​using a machine learning model, for example, the direction to be adjusted by operation is set to just one direction, the front direction (C), and the HRTF adjustment value for the front direction (C), which is the direction to be adjusted by operation, is provided as input data to the machine learning model, allowing it to estimate HRTF adjustment values ​​for each of the other directions out of all directions. In this case, in addition to the HRTF adjustment value for the front direction (C), which is the direction to be adjusted by operation, information obtained by the listening comparison process, such as the identification number of the (tentative) recommended HRTF that will be the target HRTF, can be provided as input data to the machine learning model.

[0296] FIG. 16 is a flowchart illustrating an example of the process of adjusting the direction of HRTF localization for direction #n, which is performed in the adjustment process of FIG.

[0297] In step S141, the adjustment unit 114 generates (plays) a trial sound to which the HRTF for the current direction #n (the current HRTF for direction #n) is applied, and presents it to the user. It also presents to the user a second judgment criterion, which is the judgment criterion when performing an operation to adjust the HRTF, and the process proceeds to step S142.

[0298] As described above, the judgment criteria (second judgment criteria) for the adjustment process can be, for example, the direction of localization and the tone (whether they match) out of the direction of localization, sense of distance, and tone.

[0299] Furthermore, in the adjustment process, pink noise, for example, can be used as a personal optimization sound source to which HRTFs are applied to generate a sample sound. By using a wideband signal such as pink noise as the personal optimization sound source, it is possible to generate a sample sound that appropriately (easily understandable to the user) reflects differences in HRTFs, and HRTF adjustment can be performed appropriately (accurately).

[0300] In step S142, the adjustment unit 114 waits for the user to listen to the HRTF (or the trial sound to which the HRTF has been applied) for the presented current direction #n and perform an operation to adjust the direction of localization in accordance with the second judgment criterion, and then receives (via the UI unit 116) operation information representing the operation performed in accordance with the second judgment criterion, and processing proceeds to step S143.

[0301] In step S143, the adjustment unit 114 adjusts the HRTF (localization direction of the HRTF) for the current direction #n in accordance with the operation information, and the process proceeds to step S144.

[0302] In step S144, the adjustment unit 114 generates a trial sound to which the adjusted HRTF for direction #n (the adjusted HRTF for direction #n) has been applied, and presents the trial sound to the user, and the process proceeds to step S145.

[0303] In step S145, the adjustment unit 114 determines whether or not to end the adjustment of the HRTF for direction #n.

[0304] If it is determined in step S145 that the HRTF adjustment for direction #n is not to be terminated, for example, if the user feels that the direction of the sample sound localization presented in the previous step S144 does not match direction #n and performs an operation to make adjustments again, the process returns to step S141 and the same process is repeated.

[0305] Also, if it is determined in step S145 that the adjustment of the HRTF for direction #n is to be terminated, for example, if the user feels that the direction of the localization of the sample sound presented in the previous step S144 matches direction #n and performs an operation to terminate the adjustment of the HRTF for direction #n, the process of adjusting the direction of the localization of the HRTF for direction #n is terminated.

[0306] FIG. 17 is a flowchart illustrating an example of the process of adjusting the timbre of the HRTF for direction #n, which is performed in the adjustment process of FIG.

[0307] In step S151, similar to step S141 in FIG. 16, the adjustment unit 114 generates a trial sound to which the HRTF for the current direction #n has been applied, and presents this to the user. The adjustment unit 114 also presents to the user a second judgment criterion, which is a judgment criterion when performing an operation to adjust the HRTF, and the process proceeds to step S152.

[0308] In step S152, the adjustment unit 114 waits for the user to listen to the HRTF for the presented current direction #n and perform an operation to adjust the tone in accordance with the second judgment criterion, then receives operation information representing the operation performed in accordance with the second judgment criterion, and processing proceeds to step S153.

[0309] In step S153, the adjustment unit 114 adjusts the HRTF (timbre) for the current direction #n in accordance with the operation information, and the process proceeds to step S154.

[0310] In step S154, the adjustment unit 114 generates a sample sound to which the adjusted HRTF for direction #n has been applied, and presents the sample sound to the user, and the process proceeds to step S155.

[0311] In step S155, the adjustment unit 114 determines whether or not to end the adjustment of the HRTF for direction #n.

[0312] If it is determined in step S155 that the HRTF adjustment for direction #n is not to be terminated, for example, if the user feels that the tone of the sample sound presented in the previous step S154 does not suit them and performs an operation to make adjustments again, the process returns to step S151 and the same process is repeated.

[0313] Furthermore, if it is determined in step S155 that the adjustment of the HRTF for direction #n is to be terminated, for example, if the user feels that the tone of the sample sound presented in the previous step S154 suits them and performs an operation to terminate the adjustment of the HRTF for direction #n, the process of adjusting the tone of the HRTF for direction #n is terminated.

[0314] <UI>

[0315] 18 and 19 are diagrams showing a display example of an adjustment screen (window) as a fourth example of a UI.

[0316] The adjustment screen is a screen (UI) for adjusting HRTFs in the adjustment process.

[0317] FIG. 18 shows an adjustment screen 150 presented to the user when adjusting the direction of HRTF localization and tone color for the front direction (C) determined as direction #1, for example.

[0318] The adjustment screen 150 has an adjustment instruction message 151 , a localization direction adjustment section 152 , a tone adjustment section 153 , a back button 154 , and a next button 155 .

[0319] On the adjustment screen 150, an adjustment instruction message 151, a localization direction adjustment section 152, and a tone adjustment section 153 are arranged in that order from top to bottom, and below the tone adjustment section 153, a back button 154 and a next button 155 are arranged side by side on the left and right, respectively.

[0320] The adjustment instruction message 151 presents the user with a second criterion for performing an operation to adjust the HRTF, and instructs the user to perform the adjustment operation in accordance with the second criterion. The adjustment instruction message 151 instructs the user to adjust the HRTF for the front direction (C) as direction #1, using the localization direction and timbre (matching (adjusting) them) as criteria for judgment. The adjustment instruction message 151 allows the user to easily understand that the adjustment operation should be performed so that the localization direction and timbre match.

[0321] The localization direction adjustment unit 152 is a rectangular area for adjusting the direction of localization, and a 4π space image, which is an image of 4π space, is displayed on the localization direction adjustment unit 152. In the 4π space image, the depth direction is the front, and (an image of) a listener simulating a user is placed in the center, and (images of) speakers are placed according to a predetermined speaker layout. Furthermore, the localization direction adjustment unit 152 displays a cursor 161, a play button 162, and a reset button 163 superimposed on the 4π space image.

[0322] The cursor 161 is displayed in the direction of the HRTF being adjusted on the adjustment screen 150 as seen by the listener on the 4π spatial image, i.e., in Fig. 18, it is positioned in the front direction (C) as direction #1, and is operated when adjusting the direction of HRTF localization. That is, the user operates the cursor 161 so that the localization direction (position) (the direction from which the sample sound is heard) is the direction in which the cursor 161 is located, i.e., in Fig. 18, it is the front direction (C) as direction #1. For example, if the sample sound is heard from a direction to the left of the front direction (C), the user operates the cursor 161 to move the sample sound (sound source) to the right.

[0323] The play button 162 is operated to play back a sample sound to which the HRTF for the current direction #1 has been applied. The reset button 163 is operated to reset the HRTF localization direction to its initial state (returning it to the state before adjustment).

[0324] The timbre adjustment section 153 is a rectangular area for adjusting the timbre, and displays the frequency characteristics of the HRTF (or the sample sound to which the HRTF is applied) as the timbre of the HRTF for the current direction #1. Furthermore, the timbre adjustment section 153 displays a play button 172 and a reset button 173.

[0325] The play button 172 is operated to play back a sample sound to which the HRTF for the current direction #1 has been applied. The reset button 173 is operated to reset the frequency characteristics of the HRTF for direction #1, and therefore the timbre, to their initial states.

[0326] The back button 154 is operated to return to the previous setting screen, and the next button 155 is operated to proceed to the next setting screen.

[0327] In the adjustment screen 150 configured as described above, the cursor 161 in the localization direction adjustment unit 152 has four arrows, for example, "up," "down," "left," and "right," and the user can adjust the localization direction of the HRTF by operating (clicking, for example) the arrows. When the user operates the "up," "down," "left," or "right" arrow, the adjustment unit 114 adjusts the HRTF so that the localization direction of the HRTF with respect to the front direction (C) as direction #1 moves up, down, left, or right by a predetermined amount (angle), respectively. The operating means operated by the user to adjust the localization direction can be any operating means, such as a touch panel, a mouse, a PC keyboard, or a cross key or joystick on a game controller for a game console.

[0328] The adjustment of the HRTF to move (change) the direction of localization can be performed by any method.

[0329] For example, the localization in the azimuth direction (horizontal direction) (left-right direction) is related to the interaural time difference (ITD) and interaural pressure difference (ILD) of the HRTF. Therefore, by changing the ITD and / or ILD of the HRTF, it is possible to adjust the direction of localization (the direction in which the user perceives the sound source position) to move in the azimuth direction.

[0330] Furthermore, for example, localization in the elevation direction (vertical direction) (up and down direction) is related to the notch frequency present in the high frequency range of HRTF, above 4 kHz. Furthermore, increasing the notch frequency makes it easier to feel localization in the direction of higher elevation angles. Therefore, by raising or lowering the notch frequency of the HRTF, it is possible to adjust the direction of localization to move in the elevation direction.

[0331] It should be noted that the relationship between the ITD, ILD, and notch frequency of an HRTF and the direction of HRTF localization is not a simple one, such that, for example, changing the ITD by a certain amount will shift the direction of localization by a certain amount. Furthermore, when the ITD, ILD, and notch frequency of an HRTF are artificially changed in response to user operation, it cannot be denied that the change may adversely affect elements of the HRTF (or the way the sound to which it is applied is heard) other than the direction of HRTF localization.

[0332] Therefore, HRTFs with dense measurement points (directions) are used as HRTFs (HRTFs for all directions) assigned the same identification number and stored in database 111, and adjustment of the direction of localization of HRTF for direction #n can be performed using HRTFs measured at measurement points in directions near direction #n. As HRTFs with dense measurement points, for example, HRTFs measured at measurement points every 5 degrees in both the azimuth direction and the elevation direction can be used.

[0333] For example, if the user feels that the localization direction of the HRTF for the direction 30 degrees forward and left is slightly shifted to the left and performs an operation to move the localization direction to the right, the HRTF measured at the measurement point immediately to the right of the measurement point for the direction 30 degrees forward and left, in this case the measurement point for the direction 25 degrees forward and left, can be used as the adjusted HRTF. In this case, the HRTF for the direction 30 degrees forward and left is replaced with the HRTF for the direction 25 degrees forward and left, but in this technology, this is also included in the adjustment of the HRTF for the direction 30 degrees forward and left.

[0334] If it is desired to adjust the direction of localization more precisely, an interpolated HRTF can be generated for an arbitrary point other than the measurement point by interpolating using HRTFs measured at multiple measurement points near the arbitrary point, and the localization direction of the HRTF can be adjusted using the interpolated HRTF. For example, the interpolated HRTF can be used as the HRTF after the adjustment of the direction of localization.

[0335] Adjusting the HRTF (localization direction) using measured or interpolated HRTFs as described above may prevent adverse effects on HRTF elements other than the HRTF localization direction, as occurs when adjustments are made that artificially change the HRTF.

[0336] Regarding the tone adjustment section 153, for example, the user can adjust the frequency characteristics of the HRTF, and therefore the tone, by manipulating the frequency characteristics displayed on the tone adjustment section 153.

[0337] The timbre adjustment unit 153 may also include, for example, a one-dimensional slider bar. In this case, the user can adjust the timbre by amplifying a specific frequency band in the HRTF frequency characteristics when the slider bar is moved in one direction, and attenuating the specific frequency band when the slider bar is moved in the other direction. Multiple slider bars can be provided, each responsible for a different frequency band. The slider bar can be configured to have only discrete values ​​for the gain of the frequency band, such as three levels (-1, 0, 1) or five levels (-2, -1, 0, 1, 2), or it can be configured to have continuous values. Configuring the slider bar to have continuous values ​​allows for fine adjustment of the timbre, thereby meeting the needs of advanced users. Configuring the slider bar to have discrete values ​​limits the timbre adjustment, making it easier for beginners to find an appropriate amount of adjustment for the timbre.

[0338] Furthermore, the tone adjustment unit 153 can adjust the center frequency, bandwidth, and gain (parameters) of the frequency band, like a so-called parametric EQ (equalizer). In this case, the frequency characteristics of the HRTF, and therefore the tone, can be adjusted more flexibly.

[0339] Alternatively, the timbre adjustment unit 153 may be provided with a slider bar that performs principal component analysis of the frequency characteristics of the HRTFs stored in the database 111 and changes the scores of principal components, such as the first and second principal components, of the HRTF frequency characteristics. In this case, the user can adjust the frequency characteristics of the HRTFs and, ultimately, the timbre by operating the slider bar to change the principal component scores. The user can use any operating device to adjust the timbre, such as a touch panel, a mouse, a PC keyboard, or a cross key or joystick on a game controller for a game console. For example, if a mouse with a mouse wheel is used as the operating device to adjust the sound localization direction and timbre, the user can click the cursor 161 with the mouse to adjust the sound localization direction and turn the mouse wheel to adjust the timbre. For example, turning the mouse wheel in the + direction sharpens the timbre, and turning the mouse wheel in the - direction dulls the timbre. In this case, the user can smoothly adjust the sound localization direction and the timbre without having to change the operating device.

[0340] Note that, on the adjustment screen 150, the adjustment instruction message 151 presents the second criterion for performing the HRTF adjustment operation in natural language, but the second criterion (as well as the first criterion) can be presented by means other than natural language. For example, the adjustment screen 150 displays a localization direction adjustment unit 152 for adjusting the localization direction and a timbre adjustment unit 153 for adjusting the timbre. The display of the localization direction adjustment unit 152 and the timbre adjustment unit 153 allows the user to recognize that they are performing operations to adjust the localization direction and the timbre, respectively. Therefore, the display of the localization direction adjustment unit 152 and the timbre adjustment unit 153 can also be said to present the second criterion for performing the HRTF adjustment operation.

[0341] When the user has completed adjustment of the HRTF localization direction and timbre for the front direction (C) determined as direction #1 on adjustment screen 150, he or she operates next button 155. When next button 155 is operated, an adjustment screen for adjusting the HRTF localization direction and timbre for the next direction is presented to the user.

[0342] 15, the HRTF localization direction and timbre for each direction are adjusted, but it is also possible to adjust the HRTF localization direction and timbre for multiple directions, such as two directions, at once. For example, an adjustment screen can be presented that adjusts the HRTF localization direction and timbre for one direction, or an adjustment screen can be presented that adjusts the HRTF localization direction and timbre for each of multiple directions at once.

[0343] For example, the directions #1, #2, #3, #4, #5, ... are determined to be the front direction (C), left front (L), right front (R), left side (Ls), right side (Rs), ...

[0344] When the next button 155 is operated on the adjustment screen 150 for adjusting the HRTF localization direction and tone color for the front direction (C) as direction #1, an adjustment screen for adjusting the HRTF localization direction and tone color for the left front (L) as direction #2 next to direction #1 can be presented instead of the adjustment screen 150.

[0345] Furthermore, when the next button 155 is operated on the adjustment screen 150 for adjusting the localization direction and timbre of the HRTF for the front direction (C) as direction #1, an adjustment screen for adjusting the localization direction and timbre of the HRTF for the left front (L) as direction #2 after direction #1 and the right front (R) as direction #3 after direction #2 can be presented instead of the adjustment screen 150. Furthermore, thereafter, an adjustment screen for adjusting the localization direction and timbre of the HRTF for the left side (Ls) as direction #4 after direction #3 and the right side (Rs) as direction #5 after direction #4 can be presented.

[0346] FIG. 19 shows an adjustment screen 180 for adjusting the direction and tone of the HRTF for the left front (L) as direction #2 and the HRTF for the right front (R) as direction #3, and an adjustment screen 210 for adjusting the direction and tone of the HRTF for the left side (Ls) as direction #4 and the right side (Rs) as direction #5.

[0347] The adjustment screen 180 has an adjustment instruction message 181 , a localization direction adjustment section 182 , tone adjustment sections 183L and 183R, a back button 184 , and a next button 185 .

[0348] The adjustment instruction message 181, the localization direction adjustment unit 182, the back button 184, and the next button 185 correspond to the adjustment instruction message 151, the localization direction adjustment unit 152, the back button 154, and the next button 155 in Fig. 18. The timbre adjustment units 183L and 183R correspond to the timbre adjustment unit 153 in Fig. 18.

[0349] On the adjustment screen 180, an adjustment instruction message 181 and a localization direction adjustment section 182 are arranged in that order from top to bottom, and tone adjustment sections 183L and 183R are arranged side by side on the left and right, respectively, below the localization direction adjustment section 182. Furthermore, below the tone adjustment sections 183L and 183R, a back button 184 and a next button 185 are arranged side by side on the left and right, respectively.

[0350] The adjustment instruction message 181 is a message instructing that the HRTF for the left front (L) as direction #2 and the HRTF for the right front (R) as direction #3 be adjusted using the direction of localization and tone color as the criteria for judgment.

[0351] The localization direction adjustment unit 182 displays a 4π spatial image in which listeners and speakers are arranged, similar to the localization direction adjustment unit 152 in Fig. 18. Furthermore, the localization direction adjustment unit 182 displays cursors 191L ​​and 191R, play buttons 192L and 192R, and reset buttons 193L and 193R superimposed on the 4π spatial image.

[0352] The cursor 191L ​​is displayed at a position in the left front (L) as direction #2 as seen by the listener on the 4π spatial image, and is operated when adjusting the direction of HRTF localization for direction #2. The cursor 191R is displayed at a position in the right front (R) as direction #3 as seen by the listener on the 4π spatial image, and is operated when adjusting the direction of HRTF localization for direction #3.

[0353] The play button 192L is operated to play back a sample sound to which the current HRTF for the left front (L) as direction #2 has been applied, and the play button 192R is operated to play back a sample sound to which the current HRTF for the right front (R) as direction #3 has been applied. The reset button 193L is operated to reset the direction of HRTF localization for the left front (L) as direction #2 to its initial state, and the reset button 193R is operated to reset the direction of HRTF localization for the right front (R) as direction #3 to its initial state.

[0354] The timbre adjustment section 183L displays the frequency characteristics of the current HRTF for the left front (L) as direction #2 as the timbre of that HRTF, and also displays a play button 202L and a reset button 203L. The timbre adjustment section 183R displays the frequency characteristics of the current HRTF for the right front (R) as direction #3 as the timbre of that HRTF, and also displays a play button 202R and a reset button 203R.

[0355] The play button 202L is operated to play back a sample sound to which the current HRTF for the left front (L) as direction #2 is applied, and the play button 202R is operated to play back a sample sound to which the current HRTF for the right front (R) as direction #3 is applied. The reset button 203L is operated to reset the frequency characteristics (timbre) of the HRTF for the left front (L) as direction #2 to the initial state, and the reset button 203R is operated to reset the frequency characteristics of the HRTF for the right front (R) as direction #3 to the initial state.

[0356] On adjustment screen 180, it is possible to adjust the HRTF localization direction and timbre for each of two directions #2 and #3 at once (within adjustment screen 180), and the adjustment method is the same as that on adjustment screen 150 of FIG. 18, which allows adjustment of the HRTF localization direction and timbre for one direction #1, so a description thereof will be omitted.

[0357] When the back button 184 is operated on the adjustment screen 180, an adjustment screen for adjusting the HRTF for the forward direction, that is, the adjustment screen 150 in FIG. 18, is displayed instead of the adjustment screen 180.

[0358] Furthermore, when the user has completed adjusting the HRTF localization direction and tone for each of directions #1 and #2 on adjustment screen 180 and operates next button 185, adjustment screen 180 is replaced with an adjustment screen for adjusting the HRTF for the next (or subsequent) direction, i.e., adjustment screen 210.

[0359] The adjustment screen 210 has an adjustment instruction message 211, a localization direction adjustment section 212, tone adjustment sections 213L and 213R, a back button 214, and a next button 215. The adjustment instruction message 211 to the next button 215 correspond to the adjustment instruction message 181 to the next button 185 on the adjustment screen 180, respectively.

[0360] The adjustment screen 210 is similar to the adjustment screen 180, except that instead of adjusting the HRTFs for the two directions #2 and #3, namely the left front (L) and the right front (R), adjustments are made to the direction of HRTF localization and tone for the two directions #4 and #5, namely the left side (Ls) and the right side (Rs), respectively, and therefore a description thereof will be omitted.

[0361] When adjusting HRTFs for a certain direction #n, for example, the HRTF adjustment value for the previous direction #n-1 can be set as the initial value for HRTF adjustment for direction #n. In this case, the user can start adjusting HRTFs for direction #n from the initial value.

[0362] For example, if the adjustment of the HRTF localization direction for direction #1 is performed by adjusting the localization direction two steps upward, the adjustment of the HRTF localization direction for the next direction #2 can be started from a state in which the localization direction has been adjusted two steps upward. If the adjustment of the HRTF localization direction for direction #1 is performed by adjusting the localization direction two steps upward, the adjustment of the HRTF localization direction for the next direction #2 is also likely to be performed by adjusting the localization direction two steps upward. Therefore, by starting the adjustment of the HRTF localization direction for direction #2 from a state in which the localization direction has been adjusted two steps upward, the effort required for the user to perform the investigation can be reduced.

[0363] FIG. 20 is a diagram showing another display example of the adjustment screen.

[0364] In Figures 18 and 19, the HRTF adjustment for each of directions #1 to #N, which are predetermined directions to be adjusted by operation, is performed in the order of predetermined directions #1 to #N, but the user can select any direction as the direction to be adjusted by operation and adjust the HRTF for the selected direction to be adjusted by operation.

[0365] FIG. 20 shows an adjustment screen 250 when the user selects a direction to be adjusted by operation and adjusts the HRTF for the selected direction to be adjusted by operation.

[0366] The adjustment screen 250 has an instruction message 251 , a localization direction adjustment section 252 , a tone adjustment section 253 , and a complete button 254 .

[0367] On the adjustment screen 250, an instruction message 251, a localization direction adjustment section 252, a tone adjustment section 253, and a complete button 254 are arranged in this order from the top.

[0368] The instruction message 251 is a message that instructs the user. Before the user selects the direction to be adjusted by operation, a message instructing the user to select the direction to be adjusted by operation (in FIG. 20, "Please select the direction you want to adjust") is displayed as the instruction message 251.

[0369] The localization direction adjustment section 252 is a rectangular area for adjusting the direction of localization, and a 4π spatial image is displayed on the localization direction adjustment section 252. In the 4π spatial image, as in the cases of Fig. 18 and Fig. 19 , the depth direction is the front, a listener simulating a user is placed in the center, and speakers are arranged according to a predetermined speaker layout.

[0370] The tone adjustment section 253 is a rectangular area for adjusting the tone, and no particular display is provided in the tone adjustment section 253 before the user selects the direction to be adjusted by operation.

[0371] The complete button 254 is operated when the user wishes to complete the adjustment of the HRTF for the selected direction.

[0372] The user can select a direction to be adjusted in the 4π spatial image displayed on the localization direction adjustment unit 252. The user can select any direction in the 4π spatial image as the direction to be adjusted in the operation.

[0373] However, when any direction can be selected as the direction to be adjusted by operation, in the adjustment process of Figure 15, the algorithm used to calculate HRTF adjustment values ​​for all directions other than the direction to be adjusted by operation, using HRTF adjustment values ​​for some or all of the directions to be adjusted by operation, needs to be carefully designed so that it operates appropriately no matter which direction the user selects as the direction to be adjusted by operation, which increases the cost of designing the algorithm.

[0374] Therefore, it is possible to restrict the selection of the operation adjustment target direction so that the operation adjustment target direction can be selected from a plurality of predetermined directions, thereby reducing the cost of designing the algorithm.

[0375] 20 , the localization direction adjustment unit 252 displays a plurality of direction marks superimposed on the 4π spatial image. The direction marks represent (the positions of) directions that can be selected as directions to be adjusted for operation. By selecting a direction mark, the user can select the direction represented by the direction mark as the direction to be adjusted for operation.

[0376] When the user selects the direction to be adjusted by selecting the direction mark, the display contents of the instruction message 251, the localization direction adjustment section 252, and the tone adjustment section 253 are changed (updated). The direction to be adjusted by the user's selection of the direction mark is also referred to as the selected direction.

[0377] The instruction message 251 is changed to a message instructing that the HRTF for the selected direction be adjusted using the direction of localization and timbre as criteria for judgment (in Figure 20, it reads "Please adjust the direction of localization and timbre").

[0378] In the localization direction adjustment unit 252, a cursor 281 is displayed at the position of the selected direction (the position of the direction mark selected by the user) on the 4π spatial image, and a play button 282 and a reset button 283 are also displayed.

[0379] The cursor 281 is operated when adjusting the direction of HRTF localization for the selected direction. That is, by operating the cursor 281, the user can adjust the direction of HRTF localization for the selected direction so that the direction of localization coincides with the direction in which the cursor 281 is positioned (selected direction).

[0380] The play button 282 is operated to play back a sample sound to which the current HRTF for the selected direction has been applied, while the reset button 283 is operated to reset the HRTF localization direction for the selected direction to its initial state (returning it to the state before adjustment).

[0381] The frequency characteristics of the HRTF are displayed as the current HRTF tone for the selected direction in the tone adjustment section 253. Furthermore, a play button 292 and a reset button 293 are displayed in the tone adjustment section 253.

[0382] The play button 292 is operated to play a preview sound to which the current HRTF for the selected direction has been applied, and the reset button 293 is operated to reset the frequency characteristics (timbre) of the HRTF for the selected direction to the initial state.

[0383] <Effects of combining listening comparison processing and adjustment processing>

[0384] Below, we will explain the effect of combining the listening comparison process and the adjustment process, but before that, we will explain the accuracy of recommended HRTFs determined by the listening comparison process without presenting any criteria.

[0385] As described above, for the HRTFs and the way in which they are heard, the relationship between the direction of localization and the timbre is relatively clear, and it is relatively easy to adjust the HRTFs so that the direction of localization and the timbre change to a certain way of hearing. Therefore, the adjustment process adjusts the direction of localization and the timbre, and therefore the direction of localization and the timbre are used as the criteria (second criteria) for the user's adjustment operation.

[0386] Therefore, small discrepancies in the direction of localization and timbre can be adjusted by the adjustment process. Therefore, in the listening comparison process that determines a provisional recommended HRTF that is to be used as the target HRTF in the adjustment process, a judgment criterion different from the judgment criterion used when the user performs the adjustment operation, for example, a sense of distance that is difficult to adjust by the adjustment process, is adopted as the judgment criterion (first judgment criterion) when the user answers which HRTF they feel suits them, thereby making it possible to perform highly accurate personal optimization from the perspectives of various judgment criteria, in this case, the sense of distance, direction of localization, and timbre.

[0387] That is, in the listening comparison process, the HRTF with the best sense of distance (suitable for the user) is determined as the provisional recommended HRTF, and in the adjustment process performed using this provisional recommended HRTF as the target HRTF, the localization direction and timbre of the target HRTF are adjusted to suit the user. As a result, an HRTF whose sense of distance, localization direction, and timbre all suit the user can be obtained as the final recommended HRTF.

[0388] In order to perform highly accurate personal optimization, it is important to appropriately combine the judgment criterion (first judgment criterion) used when answering questions in the listening comparison process with the judgment criterion (second judgment criterion) used when performing adjustment operations in the adjustment process. For example, the judgment criterion used when performing adjustment operations in the adjustment process is preferably an element of the hearing of an HRTF (or a sound to which an HRTF is applied) (sense of distance, direction of localization, timbre, etc.) for which the relationship between the HRTF and how that HRTF is heard is relatively clear, such as direction of localization and timbre, and for which the HRTF can be easily adjusted. Furthermore, for example, the judgment criterion used when answering questions in the listening comparison process is preferably different from the judgment criterion used when performing adjustment operations in the adjustment process, such as an element for which the relationship between the HRTF and how that HRTF is heard is not clear.

[0389] Furthermore, in order to perform highly accurate personal optimization, it is also important to present the user with criteria for judgment.

[0390] FIG. 21 is a diagram illustrating a listening comparison process in which no criteria for answering are presented.

[0391] Although sounds from the sides can often be localized with a certain degree of accuracy even when HRTFs that are not personally optimized are applied, sounds from the front will often be localized behind or inside the head if a personally optimized HRTF is not applied. For this reason, it is desirable to use sounds from the front, such as the front (C), left front (L), or right front (R), as the sample sounds.

[0392] 21, in the listening comparison process, HRTFs for a plurality of directions in front, for example, the left front (L), the front direction (C), and the right front (R), are applied to the personal optimization sound source for HRTF-A and HRTF-B to be compared, which are selected from the HRTFs stored in the database 111. As a result, trial sounds are generated at the sound source positions of the left front (L), the front direction (C), and the right front (R) in that order, and are presented to the user.

[0393] However, in FIG. 21, in the listening comparison process, no criteria are presented to the user when answering which of HRTF-A and HRTF-B suits them best.

[0394] In this case, the user may answer with the direction of sound localization in mind as a criterion for judgment, or with the sense of distance or timbre in mind as a criterion for judgment.In addition, for example, the user may answer with an HRTF that they vaguely feel is good, without being particularly conscious of any criterion for judgment.

[0395] For example, if a user answers using the direction of localization as the criterion for judgment, without being conscious of distance or tone, and a recommended HRTF (in Figure 21, the HRTF (set) with identification number "ID_036") is determined based on that answer, the localization directions of the HRTFs for the left front (L), front direction (C), and right front (R) applied to the sample sounds presented to the user in the listening comparison process for that recommended HRTF will be highly accurate, as shown by the circles in the figure.

[0396] However, the direction of HRTF localization other than the left front (L), front direction (C), and right front (R) for the recommended HRTF, for example, the left side (Ls) and right side (Rs), is not necessarily accurate and remains unknown, as indicated by the question marks in the figure.Furthermore, the accuracy of elements of HRTF hearing other than the direction of HRTF localization for the left front (L), front direction (C), right front (R), left side (Ls), and right side (Rs) for the recommended HRTF, such as sense of distance and timbre, is also unknown, as indicated by the question marks in the figure.

[0397] As described above, if the user is not presented with criteria for answering questions in the listening comparison process, the accuracy of the elements of HRTF hearing (sense of distance, direction of localization, timbre, etc.) for directions other than the direction of the HRTF applied to the sample sounds presented to the user in the listening comparison process will be unclear in terms of the accuracy of personal optimization. Furthermore, the accuracy of elements of HRTF hearing (sense of distance, direction of localization, timbre, etc.) that the user is not aware of as criteria for judgment will also be unclear.

[0398] Furthermore, with regard to the burden on the user in the listening comparison process, the criteria for answering questions are left up to the user, which can make beginners, in particular, have difficulty in answering questions. For example, when two HRTFs, A and B, are compared against each other, one HRTF may be perceived as better when the sense of distance is used as the criterion for judgment, while the other HRTF may be perceived as better when the direction of localization is used as the criterion for judgment, making it difficult to determine which is better. Furthermore, when sample sounds to which HRTFs for multiple directions are applied are presented, one HRTF may be perceived as better for one direction and the other HRTF may be perceived as better for another direction, making it difficult to determine which is better.

[0399] Furthermore, with regard to the user's satisfaction with the recommended HRTF determined in the listening comparison process, when two HRTFs, HRTF-A and HRTF-B, are compared, only sample sounds from a portion of all directions, i.e., sample sounds to which HRTFs for a portion of all directions have been applied, are presented, so experienced users who are particular about quality may not be very satisfied with the recommended HRTF.

[0400] FIG. 22 is a diagram illustrating the accuracy of the adjusted recommended HRTF obtained by the adjustment process performed by presenting the judgment criteria, with the recommended HRTF determined by the listening comparison process using the judgment criteria as the target HRTF.

[0401] FIG. 22 shows a listening comparison process in which criteria for answering are presented, and an adjustment process in which the criteria are presented and which is performed using the recommended HRTF determined by the listening comparison process as the target HRTF.

[0402] In Figure 22, five directions, namely, the front direction (C), the left front (L), the right front (R), the left side (Ls), and the right side (Rs), are the directions to be adjusted for operation (predetermined multiple directions).

[0403] 22, in the listening comparison process, a portion of the HRTFs for each of the five directions to be adjusted, for example, the HRTF for the front direction (C), which is one direction in front, is applied to the personal optimization sound source, for example, HRTFs for the front direction (C) that are selected from the HRTFs stored in the database 111 and used for the HRTF-A and HRTF-B to be compared in the five directions to be adjusted in the listening comparison process. As a result, a sample sound that sounds from the sound source position in the front direction (C) is generated and presented to the user.

[0404] Furthermore, in FIG. 22, in the listening comparison process, the user is presented with the sense of distance (or the presence of distance) as a criterion for deciding which of HRTF-A and HRTF-B suits them best.

[0405] In this case, the user answers using the sense of distance as a criterion. In the listening comparison process, a recommended HRTF (in FIG. 22, the HRTF (set) with the identification number "ID_036") is determined based on the user's answer. As a result, the sense of distance of the HRTF in the front direction (C) applied to the sample sounds presented to the user in the listening comparison process is highly accurate, as indicated by the circle in the figure.

[0406] When there is a sense of distance for listening sounds to which HRTFs for the front direction (C) are applied, it is generally highly likely that the sense of distance for listening sounds to which HRTFs for other directions, such as the front left (L), front right (R), left side (Ls), right side (Rs), etc. are also applied. Therefore, when the sense of distance for HRTFs for the front direction (C) is highly accurate, the sense of distance for HRTFs for each of the other directions, the front left (L), front right (R), left side (Ls), and right side (Rs), will also be highly accurate (highly likely), as indicated by the circles in the figure.

[0407] As described above, the recommended HRTF determined by the listening comparison process has a high degree of accuracy in the sense of distance. However, for elements of HRTF hearing other than the sense of distance, such as the direction of localization and timbre, the accuracy is not necessarily high and the accuracy is unknown, as indicated by the question marks in the figure.

[0408] Therefore, in FIG. 22 , in the adjustment process, the recommended HRTF determined by the listening comparison process is set as the target HRTF, and the localization direction and timbre are presented to the user as criteria for determining whether to adjust the target HRTF.

[0409] In this case, the user performs an operation to adjust the HRTF localization direction and timbre for each of the five target directions for adjustment of the target HRTF (HRTF for all directions) using the localization direction and timbre as criteria. In the adjustment process, the HRTF localization direction and timbre for each of the five target directions for adjustment are adjusted in accordance with the user's operation. Therefore, the accuracy of the adjusted HRTF localization direction and timbre for each of the five target directions for adjustment, i.e., the front direction (C), left front (L), right front (R), left side (Ls), and right side (Rs), is improved, as indicated by the circles in the figure.

[0410] As described above, by using the listening comparison process that presents the sense of distance as a criterion for determining the answer, and the adjustment process that presents the direction of localization and timbre as criteria for determining the answer and that is performed using the recommended HRTF determined by the listening comparison process as the target HRTF, it is possible to improve the accuracy of all of the sense of distance, direction of localization, and timbre of the HRTF for each of the five directions that are subject to adjustment: front (C), left front (L), right front (R), left side (Ls), and right side (Rs), in terms of the accuracy of personal optimization.

[0411] That is, the adjustment process can improve the accuracy of directions and judgment criteria that are not confirmed in the listening comparison process. Specifically, the listening comparison process presents a sense of distance as a judgment criterion for answering, and generates a sample sound to which an HRTF for some of the directions to be adjusted (the front direction (C) in FIG. 22 ) is applied, thereby improving the accuracy of the sense of distance of the HRTF for each of the directions to be adjusted. Furthermore, the adjustment process presents a localization direction and timbre that are different from the judgment criterion for answering as a judgment criterion for the adjustment operation, and adjusts the localization direction and timbre of the HRTF for each of the directions to be adjusted. This improves the accuracy of the localization direction and timbre not only for the HRTF applied to the sample sound presented to the user in the listening comparison process, but also for HRTFs for directions other than the direction of the HRTF.

[0412] Furthermore, regarding the burden on the user in the listening comparison process, the user is presented with criteria for answering, making it easier for the user to answer. For example, even a beginner user can easily answer according to the presented criteria. By using a simple criterion for answering, such as perceived distance, the user can easily determine which of the two HRTFs, A and B, that were used in the comparison, is superior or inferior in terms of perceived distance, significantly reducing the user's confusion in answering. Furthermore, by simplifying the presentation of the sample sounds in the listening comparison process, for example, by presenting sample sounds to which an HRTF for one direction, such as the front direction (C), is applied, the user can significantly reduce the user's confusion in answering compared to presenting sample sounds to which HRTFs for multiple directions are applied.

[0413] Regarding the burden on the user in the adjustment process, the user is presented with criteria for making adjustment operations, which makes it easier for the user to perform the adjustment operations.

[0414] Regarding the user's satisfaction with the recommended HRTF after adjustment in the adjustment process, the adjustment process can adjust HRTFs in directions other than the HRTF applied to the sample sound in the listening comparison process, in addition to the direction of that HRTF, so even experienced users with particular preferences can obtain recommended HRTFs that they are satisfied with.

[0415] Here, it is important to consider what criteria are used as the criteria for determining an answer in the listening comparison process. In the above example, the sense of distance, which is difficult to adjust the HRTF for, was used as the criteria for determining an answer. However, if the direction of localization were used, the recommended HRTF determined in the listening comparison process may be an HRTF with high accuracy for the direction of localization but low accuracy for the sense of distance (no sense of distance). Currently, it is difficult to adjust an HRTF that changes the sense of distance, and therefore it is difficult to adjust the HRTF to create a sense of distance in the adjustment process. Therefore, in this case, as in the case of FIG. 22 , it is difficult to increase the accuracy of all of the sense of distance, direction of localization, and timbre for all of the HRTFs for each direction to be adjusted.

[0416] Furthermore, one method for selecting the two HRTFs to be compared in the listening comparison process is to select the next two HRTFs to be compared from unpresented HRTFs that are highly similar to the winning HRTF in a past comparison based on the user's answers, as described in Figure 3. This method makes it possible to determine the recommended HRTFs with fewer expected comparisons than when the two HRTFs to be compared, HRTF-A and HRTF-B, are selected at random.

[0417] If the user is not presented with criteria for answering questions in the listening comparison process, it is unclear what criteria the user will use to answer. Therefore, it is necessary to adopt a value for HRTF similarity that comprehensively reflects the similarity of various elements of how HRTFs are heard, such as direction of localization, sense of distance, and timbre.

[0418] On the other hand, when the sense of distance is presented to the user as a criterion for answering questions in the listening comparison process, a value that reflects only the similarity in the sense of distance can be adopted as the HRTF similarity.

[0419] Note that, although the adjustment process is performed here using the recommended HRTF determined in the listening comparison process as the target HRTF, the adjustment process can also be performed using an individual's HRTF determined by any method other than the listening comparison process as the target HRTF. For example, the user's HRTF can be estimated from an image of the user's ear, and the adjustment process can be performed using this HRTF as the target HRTF. In this case, by adopting an algorithm that can estimate an HRTF with high accuracy in the sense of distance as the algorithm for estimating the HRTF from the image of the user's ear, it is possible to obtain an HRTF with high accuracy in the sense of distance, direction of localization, and timbre as the HRTF after the adjustment process.

[0420] The localization direction and timbre can be homogenized for the HRTFs stored in the database 111. If it is difficult to adjust the localization direction and timbre of the HRTFs to the user's satisfaction using adjustment processing alone, it is effective to homogenize the localization direction and timbre for the HRTFs stored in the database 111.

[0421] In the adjustment process, the HRTFs for the directions to be adjusted by operation are not adjusted for each direction to be adjusted by operation according to the user's operation, but rather the HRTFs for multiple or all directions to be adjusted by operation can be adjusted together according to the user's operation (of one parameter).

[0422] The adjustment process can adjust the sense of distance in addition to the direction of HRTF localization and timbre. It is difficult to adjust the sense of distance of an HRTF so that increasing the value increases the sense of distance in proportion to the value. On the other hand, it is possible to adjust the HRTF so that, for example, by operating a slider bar, the sense of distance of the HRTF changes little by little in response to the operation of the slider bar. In this case, the user operates the slider bar to find the position where the sense of distance is maximized.

[0423] Adjustment of HRTFs so that the sense of distance of the HRTFs changes in response to slider bar operation can be performed, for example, by performing principal component analysis on the HRTFs stored in database 111 and changing the principal component scores of predetermined principal components, such as the first principal component or the second principal component, of the target HRTF in response to slider bar operation. Alternatively, adjustment of HRTFs that change the sense of distance can be performed by, for example, generating a generative model that generates HRTFs from latent variables by performing machine learning such as deep learning in advance, and outputting the HRTFs generated by the generative model from the latent variables in response to slider bar operation as HRTFs after the adjustment of the sense of distance.

[0424] In the listening comparison process, for example, as explained in Figure 22, when a sample sound is generated by applying an HRTF for the front direction (C) to a personal optimization sound source and presented to the user, if the two HRTFs being compared, HRTF-A and HRTF-B, differ in their respective localization directions or timbres, it may be difficult to compare the sense of distance between HRTF-A and HRTF-B due to the influence of the differences in their localization directions and timbres.

[0425] Therefore, in the listening comparison process, an adjustment process is performed to adjust the localization direction and timbre of each of HRTF-A and HRTF-B in accordance with the user's adjustment operation, and sample sounds can be generated to which the adjusted HRTF-A and HRTF-B are applied. When adjusting the localization direction of each of HRTF-A and HRTF-B, HRTF-A and HRTF-B are adjusted so that the localization direction of one HRTF becomes the same as the localization direction of the other HRTF, or so that each localization direction becomes a predetermined reference direction. The same applies to the adjustment of timbre.

[0426] As described above, by generating sample sounds that apply adjusted HRTF-A and HRTF-B, which have been adjusted so that the direction of localization and timbre are the same for each HRTF-A and HRTF-B, the user can compare the sense of distance for each HRTF-A and HRTF-B without being affected by differences in direction of localization or timbre.

[0427] The adjustment process performed in the listening comparison process is also referred to as pre-adjustment process, and the adjustment process performed after the listening comparison process is also referred to as post-adjustment process.

[0428] In the post-adjustment process, the target HRTF can be adjusted using the adjusted HRTF, which is the result of adjusting the HRTF localization direction and timbre in the pre-adjustment process. That is, in the post-adjustment process, the recommended HRTF adjusted in the pre-adjustment process can be used as the target HRTF (its initial value) to adjust the target HRTF (its localization direction and timbre). In this case, if adjustment values ​​for the localization direction and timbre suited to the user can be estimated from the localization direction and timbre of the recommended HRTF adjusted in the pre-adjustment process, the target HRTF can be adjusted using the estimated values, and the post-adjustment process can be skipped.

[0429] Furthermore, in the listening comparison process, the adjusted HRTFs, which are the results of adjusting the HRTF localization direction and timbre obtained through the pre-adjustment process, can be used to adjust the HRTFs used to generate subsequent sample sounds. That is, the adjusted HRTFs-A and HRTF-B obtained through the pre-adjustment process, which adjusts the localization direction and timbre for each of the HRTFs-A and HRTF-B of a given match, can be used to adjust the localization direction and timbre for each of the HRTFs-A and HRTF-B of the next match in the same way as adjusting the adjusted HRTFs-A and HRTF-B. In this case, it is possible to eliminate the need for the user to perform adjustment operations, or to reduce the degree or frequency of the operations, so that the localization direction and timbre of each of the HRTFs-A and HRTF-B are the same for each match.

[0430] <Personal HRTF optimization for multiple environments>

[0431] Below, we will explain the method for personal optimization of HRTFs for multiple environments.

[0432] HRTFs vary depending on the environment, such as a movie theater or a music production studio. Even within the same movie theater or music production studio, HRTFs vary depending on the size (large, medium, small) and construction. For this reason, users may want to obtain individually optimized HRTFs for multiple environments (multiple sets of HRTFs that include the acoustic (reverberation) characteristics of each environment).

[0433] Hereinafter, personal optimization that obtains personally optimized HRTFs for each of multiple environments will also be referred to as multi-environment personalization.

[0434] FIG. 23 is a diagram illustrating a first method of multi-environment personalization.

[0435] Hereinafter, for example, three environments #1, #2, and #3 are used as the multiple environments, and multiple environment personalization involves obtaining the user's HRTFs for each of the three environments #1 to #3. Furthermore, multiple environment personalization can be performed by determining a recommended HRTF through a listening comparison process, or by determining a recommended HRTF through a listening comparison process and then adjusting the target HRTF through an adjustment process using the recommended HRTF as the target HRTF; however, a description of the adjustment process will be omitted where appropriate.

[0436] The first method of multiple environment personalization uses a database of environments #1 to #3, which stores the HRTFs of multiple people measured in each of three environments #1 to #3. HRTFs that include (reflect) the acoustic characteristics of environment #k, such as the HRTF measured in environment #k (where k = 1, 2, or 3), are also referred to as the HRTF of environment #k. The database of environment #k stores the HRTFs (omnidirectional HRTFs) of environment #k for multiple people.

[0437] In a first method of multi-environment personalization, the user performs a comparison listening process using the HRTFs for environment #1 stored in the database for environment #1 to determine a recommended HRTF for environment #1, then performs a comparison listening process using the HRTFs for environment #2 stored in the database for environment #2 to determine a recommended HRTF for environment #2, and then performs a comparison listening process using the HRTFs for environment #3 stored in the database for environment #3 to determine a recommended HRTF for environment #3.

[0438] The first method of multi-environment personalization requires measuring the HRTFs of multiple people in each environment and preparing a database, and the cost of preparing the database increases in proportion to the number of environments.

[0439] Furthermore, in the first method of multi-environment personalization, the user must perform a listening comparison process for each environment using the HRTFs stored in the database for that environment, which places a heavy burden on the user. That is, in the listening comparison process, the user listens to and compares sample sounds to which two HRTFs selected from the database for that environment are applied, and then answers the question about the HRTF that best suits the user, repeatedly, until the recommended HRTF for that environment is determined. In the first method of multi-environment personalization, the user must repeatedly listen to and compare sample sounds and answer the question for each environment, which is tedious.

[0440] FIG. 24 is a diagram illustrating a second method of multi-environment personalization.

[0441] The second method of multi-environment personalization uses a database that stores anechoic chamber HRTFs for multiple people who have been measured in an anechoic chamber.

[0442] In the second method of multiple environment personalization, the HRTFs for environments #1 to #3 are generated by combining the HRTFs for each of environments #1 to #3 with the HRTF for the anechoic chamber, and a database of environments #1 to #3 is constructed in which the HRTFs for each of environments #1 to #3 are stored.

[0443] Thereafter, for example, similar to the first method of multiple environment personalization, the user performs a listening comparison process using the HRTFs for environment #1 stored in the database for environment #1 to determine a recommended HRTF for environment #1, then performs a listening comparison process using the HRTFs for environment #2 stored in the database for environment #2 to determine a recommended HRTF for environment #2, and then performs a listening comparison process using the HRTFs for environment #3 stored in the database for environment #3 to determine a recommended HRTF for environment #3.

[0444] In the second method of multi-environment personalization, the HRTFs of an environment are generated by combining the HRTFs of the anechoic chamber with the RTFs of the environment. Therefore, it is only necessary to prepare a database storing the HRTFs of the anechoic chamber, and there is no need to prepare databases for environments #1 to #3. Therefore, multi-environment personalization can be performed efficiently.

[0445] However, in the second method of multi-environment personalization, as described above, if the user performs a listening comparison process for each environment using the HRTFs stored in the database for that environment, the burden on the user will be heavy, just as in the first method of multi-environment personalization.

[0446] Therefore, in the second method of multi-environment personalization, the user selects one of environments #1 to #3 as the selected environment, performs a listening comparison process using the HRTFs stored in the database for the selected environment, and determines the recommended HRTF for the selected environment. Then, for an environment other than the selected environment among environments #1 to #3, the listening comparison process is not performed, and the recommended HRTF for that environment is determined to be the HRTF with the same identification number as the recommended HRTF for the selected environment, i.e., the HRTF generated by combining the anechoic chamber HRTF used to generate the recommended HRTF for the selected environment with an RTF. This is because an HRTF measured in another environment for a person whose HRTF suited the user in one environment is measured is likely to suit the user in that other environment.

[0447] As described above, a listening comparison process is performed using the HRTFs stored in the database of the selected environment to determine the recommended HRTF for the selected environment, and for environments other than the selected environment, the listening comparison process is not performed, and the HRTF of that environment with the same identification number (ID) as the recommended HRTF for the selected environment is determined to be the recommended HRTF for that environment.This allows the user to listen to and compare the sample sounds and respond only for the selected environment, thereby minimizing the burden on the user and enabling efficient personalization of multiple environments.

[0448] In addition, for discerning experts, the second method of multiple environment personalization can also be used, as in the first method of multiple environment personalization, to perform listening comparison processing for each environment using the HRTFs stored in the database for that environment.

[0449] FIG. 25 is a diagram illustrating a third method of multi-environment personalization.

[0450] The third method of multi-environment personalization, like the second method, uses a database that stores anechoic chamber HRTFs for multiple people who have been measured in an anechoic chamber.

[0451] In the third method of multi-environment personalization, the user performs a listening comparison process using anechoic chamber HRTFs stored in a database to determine recommended anechoic chamber HRTFs.

[0452] After that, the recommended HRTF for each of the environments #1 to #3 is synthesized with the recommended HRTF for the anechoic chamber to generate the recommended HRTF for each of the environments #1 to #3.

[0453] In the third method of multi-environment personalization, the listening comparison process is performed using HRTFs from an anechoic chamber. Therefore, it is sufficient to prepare only the HRTF database containing the anechoic chamber HRTFs, and there is no need to prepare databases for environments #1 to #3. This allows for efficient multi-environment personalization.

[0454] Furthermore, in a third method of multiple environment personalization, the recommended HRTF for each of environments #1 to #3 is generated by combining the RTFs for each of environments #1 to #3 with the recommended HRTF for the anechoic chamber determined by the listening comparison process. Therefore, the user can easily obtain recommended HRTFs for various environments. Furthermore, because the user only needs to compare the sample sounds and respond to the anechoic chamber HRTFs, the burden on the user is minimized, and multiple environment personalization can be performed efficiently.

[0455] In addition, experienced users who are particular about their listening habits can use the first or second method of multi-environment personalization and perform listening comparison processing for each environment using the HRTFs stored in the database for that environment.

[0456] FIG. 26 is a diagram illustrating a fourth method of multi-environment personalization.

[0457] The fourth method of multi-environment personalization, like the second method, uses a database that stores anechoic chamber HRTFs for multiple people who have been measured in an anechoic chamber.

[0458] In the fourth method of multi-environment personalization, a certain environment, for example, a standard environment for listening to music, is used as the standard environment, and the HRTF of the standard environment is generated by synthesizing the HRTF of the standard environment with the HRTF of an anechoic chamber, and a database of the standard environment is constructed in which the HRTF of the standard environment is stored.

[0459] Then, in the fourth method of multi-environment personalization, the user performs a listening comparison process using HRTFs of standard environments stored in a database of standard environments to determine recommended HRTFs for the standard environment.

[0460] In the fourth method of multiple environment personalization, the anechoic chamber HRTF with the same identification number (ID) as the recommended HRTF for the standard environment, i.e., the anechoic chamber HRTF used to generate the recommended HRTF for the standard environment, is used as the recommended anechoic chamber HRTF, and the recommended HRTF for each of environments #1 to #3 is generated by combining the recommended anechoic chamber HRTF with the RTF for each of environments #1 to #3.

[0461] In the fourth method of multi-environment personalization, the listening comparison process is performed using HRTFs for a standard environment, which are generated by synthesizing HRTFs for a standard environment with HRTFs for an anechoic chamber. Therefore, it is sufficient to prepare only the database containing HRTFs for the anechoic chamber, and there is no need to prepare databases for environments #1 to #3. This allows for efficient multi-environment personalization.

[0462] Furthermore, in a fourth method of multiple environment personalization, the recommended HRTFs for each of environments #1 to #3 are generated by combining the RTFs for each of environments #1 to #3 with the recommended anechoic chamber HRTF used to generate the recommended HRTF for the standard environment, which was determined by a listening comparison process using the HRTF for the standard environment. This allows the user to easily obtain recommended HRTFs for various environments. Furthermore, because the user only needs to compare sample sounds and respond to questions about the HRTFs for the standard environment, the burden on the user is minimized, and multiple environment personalization can be performed efficiently.

[0463] Furthermore, in the fourth method of multi-environment personalization, the listening comparison process is performed using HRTFs from a standard environment, which makes it easier for the user to respond in the listening comparison process. Sample sounds with standard reverberation may be easier to compare than sample sounds without environmental reverberation, and presenting the user with sample sounds that are easy to compare makes it easier for the user to respond with the HRTF that suits them.

[0464] Figure 27 is a diagram illustrating an example of the process of generating a sample sound in the listening comparison process and the process of adjusting the localization direction / tone in the adjustment process when performing the second to fourth methods of multi-environment personalization.

[0465] In generating the sample sound, in step S171, the listening comparison unit 113 determines whether or not to combine an RTF with the HRTF of the anechoic chamber to be compared.

[0466] In step S171, if it is determined that an RTF is to be synthesized with the anechoic chamber HRTF, for example, if an RTF of a specified environment that can be synthesized with the anechoic chamber HRTF can be obtained, processing proceeds to step S172.

[0467] In step S172, the listening comparison unit 113 combines the RTF in the predetermined environment with the HRTF in the anechoic chamber, and the process proceeds to step S173.

[0468] In step S173, the listening comparison unit 113 generates a sample sound to which the HRTF of the anechoic chamber has been synthesized with the RTF of the specified environment, i.e., the HRTF (BRTF) of the specified environment has been applied, and the process of generating the sample sound is then terminated.

[0469] Also, if it is determined in step S171 that an RTF is not to be synthesized with the anechoic chamber HRTF, for example, if an RTF of a specified environment that can be synthesized with the anechoic chamber HRTF cannot be obtained, processing proceeds to step S174.

[0470] In step S174, the listening comparison unit 113 generates a sample sound to which the anechoic chamber HRTF has been applied, and the sample sound generation process ends. When performing the listening comparison process, the RTFs are combined with the anechoic chamber HRTFs. Only the necessary RTFs can be combined with the anechoic chamber HRTFs after the listening comparison process has started, or all RTFs can be combined with the anechoic chamber HRTFs in advance. When performing the listening comparison process, the sample sound to which the HRTFs have been applied can be generated each time a comparison is made after the listening comparison process has started, or sample sounds to which the anechoic chamber HRTFs, into which the RTFs of each environment have been combined, have been generated in advance.

[0471] In adjusting the direction of localization / tone color, in step S181, the adjustment unit 114 adjusts the recommended HRTF for the anechoic chamber in accordance with operation information representing the user's operation to adjust the direction of localization / tone color, and the process of adjusting the direction of localization / tone color is then completed.

[0472] The recommended anechoic chamber HRTF adjusted in step S181 means the anechoic chamber HRTF used to generate the recommended HRTF for a specified environment, if the recommended HRTF for that specified environment is determined from the HRTF for that specified environment generated by combining the anechoic chamber HRTF with the RTF for that specified environment in the listening comparison process.

[0473] If the recommended anechoic chamber HRTF adjusted in step S181 is the anechoic chamber HRTF used to generate the recommended anechoic chamber HRTF for a predetermined environment, the recommended anechoic chamber HRTF after adjustment is combined with the RTF for the predetermined environment to generate the recommended anechoic chamber HRTF after adjustment, as the adjustment of the recommended anechoic chamber HRTF for a predetermined environment. In the adjustment process, each time the recommended anechoic chamber HRTF is adjusted, the RTF for the predetermined environment is combined with the recommended anechoic chamber HRTF after adjustment, and a sample sound is generated to which the recommended anechoic chamber HRTF after adjustment combined with the RTF for the predetermined environment is applied. In the adjustment process, it is not impossible, but unrealistic, to combine the RTFs for each environment with all the adjusted recommended anechoic chamber HRTFs in advance, or to generate a sample sound to which the recommended anechoic chamber HRTF after such combination is applied.

[0474] <Description of a computer to which this technology is applied>

[0475] Next, the above-described series of processes can be performed by hardware or software. When the series of processes is performed by software, the programs that make up the software are installed on a general-purpose computer or the like.

[0476] FIG. 28 is a block diagram showing an example of the configuration of an embodiment of a computer in which a program for executing the above-described series of processes is installed.

[0477] The program can be recorded in advance on the hard disk 905 or ROM 903 as a recording medium built into the computer.

[0478] Alternatively, the program can be stored (recorded) on a removable recording medium 911 driven by the drive 909. Such a removable recording medium 911 can be provided as a so-called package software. Here, examples of the removable recording medium 911 include a flexible disk, a CD-ROM (Compact Disc Read Only Memory), an MO (Magneto Optical) disk, a DVD (Digital Versatile Disc), a magnetic disk, and a semiconductor memory.

[0479] The program can be installed into the computer from the removable recording medium 911 as described above, or can be downloaded to the computer via a communication network or a broadcasting network and installed on the built-in hard disk 905. That is, the program can be transferred to the computer wirelessly from a download site via an artificial satellite for digital satellite broadcasting, or transferred to the computer via a wired network such as a LAN (Local Area Network) or the Internet.

[0480] The computer includes a CPU (Central Processing Unit) 902 , to which an input / output interface 910 is connected via a bus 901 .

[0481] When a user inputs a command via an input / output interface 910 by operating an input unit 907, the CPU 902 executes a program stored in a read-only memory (ROM) 903 in accordance with the command. Alternatively, the CPU 902 loads a program stored on a hard disk 905 into a random access memory (RAM) 904 and executes the program.

[0482] As a result, the CPU 902 performs processing according to the flowchart described above or processing performed by the configuration of the block diagram described above. Then, the CPU 902 outputs the processing results from the output unit 906 via the input / output interface 910, or transmits them from the communication unit 908, or further records them on the hard disk 905, as necessary.

[0483] The input unit 907 is made up of a keyboard, a mouse, a microphone, etc. The output unit 906 is made up of an LCD (Liquid Crystal Display), a speaker, etc.

[0484] In this specification, the processing performed by a computer according to a program does not necessarily have to be performed in chronological order according to the order described in the flowchart. In other words, the processing performed by a computer according to a program also includes processing that is executed in parallel or individually (for example, parallel processing or object-based processing).

[0485] The program may be processed by a single computer (processor), or may be distributed among multiple computers. Furthermore, the program may be transferred to and executed on a remote computer.

[0486] Furthermore, in this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0487] It should be noted that the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the present technology.

[0488] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.

[0489] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.

[0490] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0491] Furthermore, the effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0492] The present technology can have the following configurations.

[0493] <1> An information processing method including: determining recommended parameters, which are sound reproduction parameters to be recommended to a user, based on a user's evaluation responses to a plurality of sample sounds obtained by applying a plurality of sound reproduction parameters related to sound reproduction to a predetermined sound source. <2> The information processing method described in <1>, further including selecting the plurality of sound reproduction parameters. <3> The information processing method described in <2>, wherein the selection of the plurality of sound reproduction parameters includes selecting the plurality of sound reproduction parameters from unpresented parameters, which are sound reproduction parameters that have not been presented to the user to the sample sounds to which the plurality of sound reproduction parameters have been applied. <4> The information processing method described in <2> or <3>, wherein the selection of the plurality of sound reproduction parameters includes selecting the plurality of sound reproduction parameters from a database that stores sound reproduction parameters. <5> The information processing method described in <4>, wherein the selection of the plurality of sound reproduction parameters includes selecting a number of sound reproduction parameters that is smaller than the number of sound reproduction parameters stored in the database. <6> The information processing method described in <5>, wherein the selection of the plurality of sound reproduction parameters includes selecting two sound reproduction parameters. <7> The information processing method described in <3>, wherein the determination of the recommended parameters includes: a first phase in which a first sound reproduction parameter is determined from the plurality of sound reproduction parameters based on the user's evaluation response to the plurality of sample sounds to which each of the plurality of sound reproduction parameters selected from the unpresented parameters has been applied, and this determination is repeated until a specified number of the first sound reproduction parameters is reached; a second phase in which a second sound reproduction parameter is determined from the first sound reproduction parameters of the first phase based on the user's evaluation response to the sample sounds to which the first sound reproduction parameter of the first phase has been applied; and a third phase in which a new third sound reproduction parameter is determined based on the user's evaluation response to the sample sounds to which the second sound reproduction parameter of the second phase and a third sound reproduction parameter have been applied, and the third sound reproduction parameter obtained by performing these steps is determined as the recommended parameter.<8> The information processing method according to <7>, wherein in the selection of the plurality of sound reproduction parameters, in the first phase, the selection of the plurality of sound reproduction parameters from the unpresented parameters is performed based on the user's past evaluation answers. <9> The information processing method according to <8>, wherein in the selection of the plurality of sound reproduction parameters, in the first phase, a sound reproduction parameter with a high predetermined similarity to the sound reproduction parameter determined as the first sound reproduction parameter based on the user's past evaluation answers is selected. <10> The information processing method according to any of <1> to <9>, further including generating a UI (user interface) for selecting an environment in which the user will listen to the sample sound. <11> The information processing method according to <10>, wherein the generation of the UI includes generating a UI for selecting the predetermined sound source. <12> The information processing method according to <10> or <11>, wherein the generation of the UI includes generating a UI for displaying an image related to the sample sound. <13> The information processing method according to <10>, wherein the generation of the UI includes generating a UI for presenting applied sound obtained by applying predetermined sound reproduction parameters to predetermined sound content. <14> The information processing method according to any one of <1> to <13>, wherein the sound reproduction parameter is at least one of a hearing aid parameter, a noise canceling parameter, an equalizer parameter, and an HRTF (head-related transfer function). <15> The information processing method according to <14>, wherein the sound reproduction parameter is the HRTF, and the HRTF applied to the predetermined sound source is an HRTF for sound source positions in one or more directions including the front. <16> The information processing method according to <15>, wherein the sample sound is a sound played from a sound source position in one direction only in the front, a sound played sequentially at sound source positions in two or more directions including the front, or a sound played simultaneously at sound source positions in two or more directions including the front.<17> The information processing method according to any one of <1> to <16>, further comprising: treating the recommended parameter as a target parameter to be adjusted, and adjusting the target parameter in accordance with the user's adjustment operation for the target parameter; presenting a first criterion for the user to use when providing an evaluation response for the sample sound; and presenting a second criterion different from the first criterion for the user to use when providing an adjustment operation for the target parameter. <18> The information processing method according to <17>, wherein the sound reproduction parameter is a head-related transfer function (HRTF), the first criterion is a sense of distance, and the second criterion is a direction of localization and / or timbre. <19> An information processing system comprising: a determination unit that determines a recommended parameter, which is a sound reproduction parameter to be recommended to the user, based on the user's evaluation response for a plurality of sample sounds obtained by applying a plurality of sound reproduction parameters related to sound reproduction to a predetermined sound source. <20> A program for causing a computer to function as a determination unit that determines recommended parameters, which are sound reproduction parameters to be recommended to a user, based on the user's evaluation responses to multiple sample sounds obtained by applying multiple sound reproduction parameters related to sound reproduction to a predetermined sound source.

[0494] <A1> An information processing method comprising: determining a recommended HRTF, which is an HRTF (head-related transfer function) recommended to a user; and adjusting the recommended HRTF as a target HRTF for adjustment in accordance with an adjustment operation by the user for the target HRTF. <A2> The information processing method described in <A1>, further comprising: determining the recommended HRTF based on a user's response to an evaluation of a plurality of sample sounds obtained by applying a plurality of HRTFs to a predetermined sound source; presenting a first judgment criterion when the user provides an evaluation of the sample sounds; and presenting a second judgment criterion different from the first judgment criterion when the user performs an adjustment operation for the target parameter. <A3> The information processing method according to <A2>, wherein the first determination criterion is a sense of distance, and the second determination criterion is a direction of localization and / or timbre, and wherein a listening comparison process is performed to determine the recommended HRTF based on a user's evaluation of a plurality of sample sounds obtained by applying each of the plurality of HRTFs to the predetermined sound source, and an adjustment process is performed to adjust the target HRTF in accordance with the user's adjustment operation for the target HRTF, using the recommended HRTF as the target HRTF. <A4> (Note: The predetermined plurality of directions refers to a plurality of directions determined in advance out of all directions.) The information processing method according to <A3>, wherein the listening comparison process generates the test sounds to which HRTFs for some of the HRTFs for the predetermined plurality of directions are applied, and the adjustment process adjusts the HRTFs for each of the predetermined plurality of directions. <A5> The information processing method according to <A3>, wherein the adjustment process adjusts an HRTF for adjusting the direction of localization, among the HRTFs for each of the plurality of directions as the target HRTF, using an HRTF for a direction near the direction of localization of the HRTF. <A6> The information processing method described in <A3>, wherein, in the adjustment process, the azimuth direction of the localization of an HRTF whose localization direction is to be adjusted, among HRTFs for each of a plurality of directions as the target HRTF, is adjusted by changing the interaural time difference and / or the interaural sound pressure difference of that HRTF.<A7> The information processing method described in <A3>, in which, in the adjustment process, the elevation angle direction of the localization of an HRTF for which the localization direction is to be adjusted, among HRTFs for each of a plurality of directions as the target HRTF, is adjusted by changing the notch frequency of that HRTF. <A8> The information processing method described in <A3>, in which a UI (user interface) is generated that allows the user to adjust the localization direction of the target HRTF. <A9> The information processing method described in <A3>, in which, in the adjustment process, the timbre of the target HRTF is adjusted by amplifying or attenuating a predetermined frequency band. <A10> The information processing method described in <A3>, in which, in the adjustment process, the timbre of the target HRTF is adjusted by adjusting the center frequency, bandwidth, and gain of the frequency band. <A11> The information processing method described in <A3>, in which, in the adjustment process, the timbre of the target HRTF is adjusted by changing the principal component score of the recommended HRTF of a predetermined principal component obtained by principal component analysis performed on a plurality of HRTFs stored in a database. <A12> The information processing method according to <A3>, wherein a UI (user interface) is generated for the user to adjust the timbre of the target HRTF. <A13> The information processing method according to <A3>, wherein, in the adjustment process, adjustment of HRTFs for one or more predetermined directions among the HRTFs for each of a plurality of directions serving as the target HRTF is performed direction by direction in turn. <A14> The information processing method according to <A13>, wherein, in the adjustment process, adjustment values ​​of HRTFs for some or all of the 1st to n-1th directions are used to set an initial value for adjustment of the HRTF for the nth direction. <A15> The information processing method according to <A13>, wherein, in the adjustment process, adjustment values ​​of HRTFs for some or all of the one or more predetermined directions are used to determine adjustment values ​​of HRTFs for directions other than the one or more predetermined directions among the plurality of directions.<A16> The information processing method described in <A3>, wherein in the listening comparison process, a localization direction and / or timbre of the HRTF applied to the predetermined sound source is adjusted in accordance with an adjustment operation by the user, and the adjusted HRTF is applied to the predetermined sound source to generate the sample sound. <A17> The information processing method described in <A16>, wherein an adjustment result of the localization direction and / or timbre of the HRTF is used to adjust the HRTF used to generate the subsequent sample sound in the listening comparison process, or adjust the target HRTF in the adjustment process.

[0495] <B1> An information processing device comprising: a determination unit that determines recommended parameters, which are sound reproduction parameters to be recommended to a user, based on the user's evaluation responses to a plurality of sample sounds obtained by applying a plurality of sound reproduction parameters related to sound reproduction to a predetermined sound source. <B2> The information processing device described in <B1>, further comprising a selection unit that selects the plurality of sound reproduction parameters. <B3> The information processing device described in <B2>, wherein the selection unit selects the plurality of sound reproduction parameters from unpresented parameters, which are sound reproduction parameters that have not been presented to the user to which the sample sounds have been applied. <B4> The information processing device described in <B2> or <B3>, wherein the selection unit selects the plurality of sound reproduction parameters from a database that stores sound reproduction parameters. <B5> The information processing device described in <B4>, wherein the selection unit selects a number of sound reproduction parameters that is smaller than the number of sound reproduction parameters stored in the database. <B6> The information processing device described in <B5>, wherein the selection unit selects two sound reproduction parameters. <B7> The information processing device according to <B3>, wherein the determination unit determines, as the recommended parameter, the third sound reproduction parameter obtained by performing: a first phase in which determining a first sound reproduction parameter from the plurality of sound reproduction parameters based on the user's evaluation response to the plurality of sample sounds to which each of the plurality of sound reproduction parameters selected from the unpresented parameters has been applied, the determination being repeated until a specified number of the first sound reproduction parameters has been reached; a second phase in which determining a second sound reproduction parameter from the first sound reproduction parameters of the first phase based on the user's evaluation response to the sample sounds to which the first sound reproduction parameter of the first phase has been applied; and a third phase in which determining a new third sound reproduction parameter based on the user's evaluation response to the sample sounds to which the second sound reproduction parameter and a third sound reproduction parameter of the second phase have been applied. <B8> The information processing device according to <B7>, wherein the selection unit selects the plurality of sound reproduction parameters from the unpresented parameters in the first phase based on the user's past evaluation response.<B9> The information processing device according to <B8>, wherein the selection unit selects, in the first phase, sound reproduction parameters that have a high predetermined degree of similarity to the sound reproduction parameters determined as the first sound reproduction parameters based on the user's past evaluation responses. <B10> The information processing device according to any one of <B1> to <B9>, further comprising a UI generation unit that generates a UI (user interface) for selecting an environment in which the user will listen to the preview sound. <B11> The information processing device according to <B10>, wherein the UI generation unit generates a UI for selecting the predetermined sound source. <B12> The information processing device according to <B10> or <B11>, wherein the UI generation unit generates a UI that displays an image related to the preview sound. <B13> The information processing device according to <B10>, wherein the UI generation unit generates a UI that presents an applied sound obtained by applying predetermined sound reproduction parameters to predetermined sound content. <B14> The information processing device according to any one of <B1> to <B13>, wherein the predetermined sound source is a monaural or multi-channel sound source having a substantially flat frequency characteristic in a frequency band of approximately 50 Hz to 16 kHz. <B15> The information processing device according to any one of <B1> to <B14>, wherein the sound reproduction parameter is at least one of a hearing aid parameter, a noise canceling parameter, an equalizer parameter, and an HRTF (head-related transfer function). <B16> The information processing device according to <B15>, wherein the sound reproduction parameter is an HRTF. <B17> The information processing device according to <B16>, wherein the HRTF applied to the predetermined sound source is an HRTF for sound source positions in one or more directions including the front. <B18> The information processing device according to <B17>, wherein the sample sound is a sound produced from a sound source position in one direction only in the front, a sound produced sequentially at sound source positions in two or more directions including the front, or a sound produced simultaneously at sound source positions in two or more directions including the front. <B19> An information processing method comprising: determining a recommended parameter, which is a sound reproduction parameter to be recommended to a user, based on a user's response to an evaluation of a plurality of sample sounds produced by applying a plurality of sound reproduction parameters related to sound reproduction to a predetermined sound source.<B20> A program for causing a computer to function as a determination unit that determines recommended parameters, which are sound reproduction parameters to be recommended to a user, based on the user's evaluation responses to multiple sample sounds obtained by applying multiple sound reproduction parameters related to sound reproduction to a predetermined sound source.

[0496] 10 Information processing device, 11 Database, 12 Selection unit, 13 Sound source storage unit, 14 Preview sound generation unit, 15 Decision unit, 16 Internal state storage unit, 17 Image storage unit, 18 UI generation unit, 21 Preview sound presentation unit, 22 Answer reception unit, 23 Image presentation unit, 50 Setting screen, 51 Image display area, 52 Environment selection box, 53 Audio setting button, 54 Optimization start button, 60 Optimization screen, 61 Image display area, 62 Preview sound playback button, 63 Answer selection button, 64 Next question button, 65 Return to previous question button, 66 Progress bar, 67 Preview sound selection box, 68 Keyboard shortcut list display button, 69 Expand button, 70 Keyboard shortcut assignment display unit, 71 Favorite button, 72 Ranking display unit, 73 Play button, 74 HRTF decision button, 75 Optimization stop button, 76 Optimization end button, 90 Experience screen, 91 Optimization result display area, 92 Experience button, 93 Selection box, 94 Save button, 95 Exit button, 111 Database, 112 Sound source storage unit, 113 Listening comparison unit, 114 Adjustment unit, 115 UI generation unit, 116 UI unit, 150 Adjustment screen, 151 Adjustment instruction message, 152 Localization direction adjustment unit, 153 Timbre adjustment unit, 154 Back button, 155 Next button, 161 Cursor, 162 Play button, 163 Reset button, 172 Play button, 173 Reset button, 180 Adjustment screen, 181 Adjustment instruction message, 182 Localization direction adjustment unit 183L, 183R Tone adjustment section, 184 Back button, 185 Next button, 191L, 191R Cursor, 192L, 192R Play button, 193L, 193R Reset button, 202L, 202R Play button, 203L, 203R Reset button, 210 Adjustment screen, 211 Adjustment instruction message, 212 Localization direction adjustment section, 213L, 213R Tone adjustment section, 214 Back button, 215 Next button, 221L, 221R Cursor, 222L, 222R Play button, 223L, 223R Reset button, 232L, 232R Play button233L, 233R Reset button, 250 Adjustment screen, 251 Instruction message, 252 Localization direction adjustment unit, 253 Timbre adjustment unit, 254 Complete button, 281 Cursor, 282 Play button, 283 Reset button, 292 Play button, 293 Reset button, 901 Bus, 902 CPU, 903 ROM, 904 RAM, 905 Hard disk, 906 Output unit, 907 Input unit, 908 Communication unit, 909 Drive, 910 Input / output interface, 911 Removable recording medium

Claims

1. An information processing method including determining recommended parameters, which are sound reproduction parameters to be recommended to a user, based on the user's evaluation responses to multiple sample sounds obtained by applying multiple sound reproduction parameters related to sound reproduction to a specified sound source.

2. The information processing method according to claim 1, further comprising selecting the plurality of sound reproduction parameters.

3. The information processing method according to claim 2, wherein in the selection of the plurality of sound reproduction parameters, the plurality of sound reproduction parameters are selected from unpresented parameters, which are sound reproduction parameters that have not been presented to the user when the applied preview sound is selected.

4. The information processing method according to claim 2, wherein the selection of the plurality of sound reproduction parameters includes selecting the plurality of sound reproduction parameters from a database that stores sound reproduction parameters.

5. The information processing method according to claim 4, wherein in selecting the plurality of sound reproduction parameters, a number of sound reproduction parameters that is less than the number of sound reproduction parameters stored in the database is selected.

6. The information processing method according to claim 5, wherein two sound reproduction parameters are selected in the selection of the plurality of sound reproduction parameters.

7. The information processing method according to claim 3, wherein the determination of the recommended parameters includes a first phase in which a first sound reproduction parameter is determined from the plurality of sound reproduction parameters based on the user's evaluation response to the plurality of sample sounds to which each of the plurality of sound reproduction parameters selected from the unpresented parameters has been applied, and this determination is repeated until a specified number of the first sound reproduction parameters is reached; a second phase in which a second sound reproduction parameter is determined from the first sound reproduction parameters of the first phase based on the user's evaluation response to the sample sounds to which the first sound reproduction parameter of the first phase has been applied; and a third phase in which a new third sound reproduction parameter is determined based on the user's evaluation response to the sample sounds to which the second sound reproduction parameter and a third sound reproduction parameter of the second phase have been applied, and the third sound reproduction parameter obtained by these phases is determined as the recommended parameter.

8. The information processing method according to claim 7, wherein in the selection of the plurality of sound reproduction parameters, in the first phase, the selection of the plurality of sound reproduction parameters from the unpresented parameters is made based on the user's past responses to evaluations.

9. The information processing method according to claim 8, wherein in the selection of the plurality of sound reproduction parameters, a sound reproduction parameter having a high predetermined similarity to the sound reproduction parameter determined as the first sound reproduction parameter based on the user's past evaluation responses in the first phase is selected.

10. The information processing method according to claim 1, further comprising generating a UI (user interface) for selecting an environment in which the user listens to the sample sound.

11. The information processing method according to claim 10, wherein the generation of the UI comprises generating a UI for selecting the predetermined sound source.

12. The information processing method according to claim 10, wherein the generation of the UI comprises generating a UI that displays an image related to the sample sound.

13. The information processing method according to claim 10, wherein the generation of the UI comprises generating a UI that presents an applied sound obtained by applying predetermined sound reproduction parameters to predetermined sound content.

14. The information processing method according to claim 1, wherein the sound reproduction parameter is at least one of a hearing aid parameter, a noise canceling parameter, an equalizer parameter, and a head-related transfer function (HRTF).

15. The information processing method according to claim 14, wherein the sound reproduction parameters are the HRTFs, and the HRTFs to be applied to the predetermined sound source are HRTFs for sound source positions in one or more directions including the front.

16. The information processing method according to claim 15, wherein the preview sound is a sound that is played from a sound source position in one direction only in front, a sound that is played in sequence from sound source positions in two or more directions including the front, or a sound that is played simultaneously from sound source positions in two or more directions including the front.

17. The information processing method of claim 1, further comprising: adjusting the target parameter according to the user's adjustment operation on the target parameter, with the recommended parameter being a target parameter to be adjusted; presenting a first judgment criterion when the user responds to evaluate the sample sound; and presenting a second judgment criterion different from the first judgment criterion when the user performs the adjustment operation on the target parameter.

18. The information processing method according to claim 17, wherein the sound reproduction parameter is a head-related transfer function (HRTF), the first judgment criterion is a sense of distance, and the second judgment criterion is a direction of localization and / or a tone color.

19. An information processing system comprising a determination unit that determines recommended parameters, which are sound reproduction parameters to be recommended to a user, based on the user's evaluation responses to multiple sample sounds obtained by applying multiple sound reproduction parameters related to sound reproduction to a specified sound source.

20. A program for causing a computer to function as a determination unit that determines recommended parameters, which are sound reproduction parameters to be recommended to a user, based on the user's evaluation responses to multiple sample sounds obtained by applying multiple sound reproduction parameters related to sound reproduction to a specified sound source.

Citation Information

Patent Citations

  • Method and apparatus for reproducing a virtual sound of two channels based on individual auditory characteristic

    KR1020080060640A

  • Audio mixing and equalization in gaming systems

    US20230060590A1

Cited By

  • Transmission space reproduction method and transmission space reproduction device

    US20240322924A1