Audio output device, control system and calibration method
The audio output device and calibration method improve XR experiences by allowing users to select optimal head-related transfer functions through interactive audio signal adjustments, ensuring accurate alignment of audio positions with user perceptions.
Patent Information
- Application Number
- JP2021159844
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-29
- Publication Date
- 2025-09-01
- Estimated Expiration
- 2041-09-29
AI Technical Summary
Conventional technologies struggle to accurately calibrate head-related transfer functions for individual users, leading to inconsistencies in audio environments in VR, AR, and MR experiences.
An audio output device and calibration method that utilize a storage unit for multiple audio processing patterns, a determination unit to select optimal patterns based on user interaction, and a control system to sequentially switch and adjust audio signals, allowing users to select the best head-related transfer function.
This approach enables precise calibration of head-related transfer functions, enhancing the realism and comfort of XR content by aligning audio positions with user perceptions.
Smart Images

Figure 0007731751000001 
Figure 0007731751000002 
Figure 0007731751000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an audio output device, a control system, and a calibration method. [Background technology]
[0002] Conventionally, there is known technology that provides users with digital content that includes virtual space experiences such as VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality), known as XR (Cross Reality) content, using devices such as HMDs (Head Mounted Displays). XR is a collective term that encompasses all virtual space technologies, including VR, AR, and MR, as well as SR (Substitutional Reality) and AV (Audio / Visual).
[0003] In addition, in this technology, for example, a technology has been proposed in which the audio environment is optimized for each user by adjusting the head-related transfer function according to the size and shape of the user's head and ears. (See, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2008-193382 Summary of the Invention [Problem to be solved by the invention]
[0005] However, in the conventional technology, head-related transfer functions vary greatly from person to person, and it is not easy to calibrate them to appropriate values.
[0006] The present invention has been made in view of the above, and has an object to provide an audio output device, a control system, and a calibration method that are capable of appropriately performing calibration related to head-related transfer functions. [Means for solving the problem]
[0007] In order to solve the above-mentioned problems and achieve the object, an audio output device according to the present invention includes an audio processing pattern storage unit, an audio output unit, and a determination unit. The audio processing pattern storage unit stores a plurality of audio processing patterns, each having a different adjustment parameter related to a head-related transfer function. The audio output unit sequentially switches between audio signals that have been acoustically processed using the audio processing patterns stored in the audio processing pattern storage unit and outputs the audio. The determination unit determines an audio processing pattern for audio processing settings based on a user's operation on the audio output output by the audio output unit. [Effects of the Invention]
[0008] According to the present invention, calibration relating to head-related transfer functions can be performed appropriately. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram showing an overview of the control system. [Figure 2] FIG. 2 is a diagram showing an outline of the control system. [Figure 3] FIG. 3 is a diagram showing an outline of the calibration method. [Figure 4] FIG. 4 is a block diagram of an information processing device. [Figure 5] FIG. 5 is a schematic diagram of an acoustic table. [Figure 6] FIG. 6 is a diagram showing an example of user settings. [Figure 7] FIG. 7 is an explanatory diagram of the answering operation. [Figure 8] FIG. 8 is an explanatory diagram of the answering operation. [Figure 9]FIG. 9 is an explanatory diagram of the answering operation. [Figure 10] FIG. 10 is a flowchart showing a processing procedure executed by the information processing device. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, a detailed description will be given of an embodiment (hereinafter referred to as "embodiment") of an audio output device, a control system, and a calibration method according to the present application, with reference to the drawings. Note that the audio output device, the control system, and the calibration method according to the present application are not limited to the embodiment.
[0011] First, an outline of an audio output device, a control system, and a calibration method according to an embodiment will be described with reference to Fig. 1 to Fig. 3. Fig. 1 and Fig. 2 are diagrams showing an outline of a control system. Fig. 3 is a diagram showing an outline of a calibration method. In the following embodiment, a case will be described in which the audio output device according to the embodiment is the information processing device 10 shown in Fig. 1.
[0012] As shown in FIG. 1, a control system 1 according to the embodiment includes a display unit 3, a speaker 4, a seat mechanism 6, and an information processing device 10.
[0013] As shown in Fig. 1, the display unit 3 is an HMD. The display unit 3 is an information processing terminal that presents XR content provided by the information processing device 10 to the user U and allows the user to enjoy an XR experience. The display unit 3 is a wearable computer that is worn on the head of the user U and is goggle-shaped in the example of Fig. 1. The HMD 31 may be glasses-shaped or hat-shaped.
[0014] The display unit 3 includes a display 31, a speaker 4, and a sensor unit 5. The display 31 is provided so as to be placed in front of the eyes of the user U, and displays images included in the XR content provided by the information processing device 10.
[0015] 1 shows an example in which one display 31 is provided in front of each of the left and right eyes of the user U, but there may be only one display. The display 31 may be a non-transmissive type that completely covers the field of view, or may be a video-transmissive type or an optical-transmissive type. In this embodiment, the display 31 is a non-transmissive type.
[0016] The sensor unit 5 is a device that detects changes in the inside and outside circumstances of the user U, and includes, for example, a camera, a motion sensor, and the like.
[0017] The speaker 4 is provided in the form of headphones, for example, as shown in FIG. 1, and is worn on the ears of the user U. The speaker 4 outputs audio included in the XR content provided by the information processing device 10. Note that the speaker 4 is not limited to a headphone type, and may be a box type speaker or the like that is placed around the user U. The speaker 4 may also be a multi-channel type speaker for stereo audio or multi-channel audio playback.
[0018] The seat mechanism 6 is, for example, a 6-axis motion base that moves the seat on which the user U sits according to the posture within the XR content (such as the posture of the user's avatar in the XR content or the posture of the moving body the user is riding on), and vibrates the seat or a vibration unit attached to the user according to the surrounding environment (such as the vibrations applied to the user's avatar in the XR content or the vibrations of the moving body the user is riding on), thereby reproducing the posture within the XR space.
[0019] The information processing device 10 is a computer that is connected to the HMD 31 via a wired or wireless connection and provides XR content to the HMD 31. The information processing device 10 also acquires changes in the situation detected by the sensor unit 5 as needed and reflects the changes in the situation in the XR content.
[0020] For example, the information processing device 10 can change the direction of the field of view in the virtual space of the XR content in accordance with changes in the head position or line of sight of the user U detected by the sensor unit 5.
[0021] Incidentally, when providing such XR content, for example, by adjusting the sound image position in conjunction with the display position of the object displayed as XR content, the user can be made to perceive the sound as if it were actually coming from the object.
[0022] On the other hand, for example, as shown in Figure 2, if there is a discrepancy between the position of the object Vo displayed as XR content and the actual sound image localization St, which is the sound image localization actually felt by the user U, there is a risk that the realism of the XR content will be reduced or the user will feel uncomfortable.
[0023] For example, these problems arise due to differences in head-related transfer functions for each user, so by adjusting the head-related transfer functions, it is possible to match the position of the object Vo with the actual sound image localization St.
[0024] For example, one technique for calibrating head-related transfer functions is to analyze the shapes of the head and ears from an image of a user U and then set a head-related transfer function suited to the user U. However, with such a technique, the precision of optimizing the head-related transfer function is not sufficient, and further improvement in the calibration of head-related transfer functions is still required.
[0025] Therefore, in the calibration method according to the embodiment, a plurality of patterns of acoustic signals generated by acoustic processing using different head-related transfer functions are sequentially switched and output as sound, and the user is allowed to listen to the acoustic signals processed using the plurality of head-related transfer functions, and the acoustic processing pattern for actual acoustic processing is determined based on the listening results.
[0026] For example, as shown in Fig. 3, in the calibration method according to the embodiment, calibration acoustic signals generated by performing acoustic processing using different head-related transfer functions are stored for each layer. Note that the example shown in Fig. 3 shows a case where calibration acoustic signal patterns for a total of N layers from the first layer to the Nth layer are stored. Also, the corresponding head-related transfer functions (or processing details using the head-related transfer functions) are stored in association with the calibration acoustic signals, and it becomes possible to read out the head-related transfer function (or processing details using the head-related transfer functions) corresponding to the calibration acoustic signal selected by the user and perform acoustic signal processing based on the read-out head-related transfer function on any acoustic signal.
[0027] For example, the acoustic signals for calibration in the first layer include pattern A, pattern B, and pattern C, and in the second layer related to pattern A, it is further divided into pattern A-1, pattern A-2, and pattern A-3.
[0028] For example, in the calibration method, audio is output while sequentially switching between first-layer acoustic signals (pattern A, pattern B, pattern C), and the actual audio is played back to the user U. Then, for example, the user U selects the optimal acoustic signal from among the first-layer acoustic signals.
[0029] Next, in the calibration method according to the embodiment, audio is output while switching between audio signals in a lower hierarchy than the audio signal selected by the user U, and the user U selects the optimal audio signal, and this process is repeated.
[0030] In the calibration method according to the embodiment, for example, these processes are repeated N times, and the finally selected acoustic signal is determined as the acoustic signal for acoustic setting. The example shown in the figure shows a case where pattern B is selected in the first layer, and acoustic signal BN under it is finally selected, that is, a case where the head-related transfer function corresponding to acoustic signal BN is determined as the head-related transfer function to be used in acoustic signal processing.
[0031] In this way, in the calibration method according to the embodiment, optimization of the head-related transfer functions (selection of the optimum head-related transfer functions) is performed by actually letting the user U hear, by voice, the acoustic signals that have been subjected to several acoustic signal processes. In this way, the calibration method according to the embodiment can appropriately perform calibration of the head-related transfer functions.
[0032] In this embodiment, acoustic signals generated by acoustic processing based on different head-related transfer functions are stored in a hierarchical structure and then played back for calibration. However, it is also possible to store different head-related transfer functions (or the acoustic processing content based on the head-related transfer functions) in a hierarchical structure and perform acoustic processing on the original sound for calibration based on these stored head-related transfer functions to generate and play back an acoustic signal for calibration.
[0033] Furthermore, in the control system 1 according to the embodiment, by performing calibration related to the head-related transfer function, the position of the object Vo can be matched with the real sound image localization St (see FIG. 2), thereby improving the sense of realism of the XR content.
[0034] Next, an example of the configuration of the information processing device 10 according to the embodiment will be described with reference to Fig. 4. Fig. 4 is a block diagram of the information processing device 10.
[0035] As shown in FIG. 4, for example, the information processing device 10 according to the embodiment is connected to a display unit 3, a speaker 4, a sensor unit 5, and a seat mechanism 6.
[0036] The information processing device 10 includes a storage unit 11 and a control unit 12. The storage unit 11 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. In the example of FIG. 4, the storage unit 11 stores an XR content database (DB) 11a, an audio table 11b, and user settings 11c.
[0037] The XR content DB 11a is a database that stores a group of XR contents to be displayed on the display unit 3. The acoustic table 11b is a table that stores adjustment parameters related to head-related transfer functions for a plurality of different acoustic processing contents. The adjustment parameters for the head-related transfer functions are read out based on a selection by the user U, and acoustic processing is performed on the acoustic signal using the head-related transfer signals for which the adjustment parameters have been set, resulting in audio output.
[0038] Fig. 5 is a diagram showing an example of the acoustic table 11b. As shown in Fig. 5, the acoustic table 11b stores information such as "acoustic pattern ID," hierarchical information ("first to fourth hierarchical levels"), "adjustment parameters," and "calibration acoustic signals" in association with one another.
[0039] The "acoustic pattern ID" is an identifier for identifying each acoustic processing pattern. The first to fourth hierarchical levels indicate the hierarchical level subordinate relationship patterns of the corresponding acoustic processing patterns. For example, an acoustic processing pattern identified by the acoustic pattern ID "A-1-1-2" indicates that the first hierarchical level is "A", the second hierarchical level is "1", the third hierarchical level is "1", and the fourth hierarchical level is "2".
[0040] The "adjustment parameters" are adjustment parameters for the head-related transfer functions used in the corresponding acoustic processing patterns. The adjustment parameters for calibration include various parameters that indicate characteristics such as delay time, volume (amplification level), and sound quality (frequency characteristics).
[0041] The "calibration audio signal" is an audio signal used when calibrating a head-related transfer function. For example, an audio signal obtained by previously processing an original sound for calibration using a head-related transfer function based on the corresponding adjustment parameter is stored as the calibration audio signal.
[0042] In the example of Fig. 6, each calibration audio signal in the first layer becomes a calibration audio signal for an audio pattern ID of "0" in the second layer, and each calibration audio signal in the second layer becomes a calibration audio signal for an audio pattern ID of 0 in the third layer. Also, each calibration audio signal in the third layer becomes a calibration audio signal for an audio pattern ID other than "0" in the third layer.
[0043] For example, first, the calibration acoustic signals corresponding to the first layer (the second layer is "0") are sequentially read out and audibly played back. After that, the calibration acoustic signals of the second layer, which are the same as the first layer selected by the user, are sequentially read out and audibly played back for the calibration acoustic signals of the first layer.
[0044] Next, the calibration acoustic signals for the third layer, which are the same as the first and second layers selected by the user for the first and second layers, are read out in sequence and audio is played back. Finally, the adjustment parameters for the calibration acoustic signals selected by the user are read out, set as head-related transfer functions, and used as head-related transfer functions during actual playback.
[0045] For example, when setting the adjustment parameters for a calibration acoustic signal identified by the acoustic pattern ID "A-1-2-2," the first layer is set to "A" by sequentially switching and outputting calibration acoustic signals whose second layer is "0."
[0046] Next, the second layer is set to "1" by sequentially switching and outputting calibration acoustic signals that satisfy the conditions that the first layer is "A", the second layer is other than "0", and the third layer is "0".
[0047] Next, the third layer is set to "2" by sequentially switching and outputting calibration acoustic signals that satisfy the conditions that the first layer is "A", the second layer is "1", the third layer is other than "0", and the fourth layer is "0".
[0048] Then, by sequentially switching and outputting calibration acoustic signals that satisfy the conditions that the first layer is "A," the second layer is "1," the third layer is "2," and the fourth layer is anything other than "0," the fourth layer will be set to "2," and the calibration acoustic signal will finally be identified as "A-1-2-2."
[0049] Returning to the explanation of Fig. 4, the user settings 11c will be explained. The user settings 11c are information about acoustic patterns for acoustic settings set for each user. Fig. 6 is a diagram showing an example of the user settings 11c.
[0050] 6, the user settings 11c are information in which information items such as "user ID," "identification information," and "acoustic pattern" are associated with each other. The "user ID" is an identifier for identifying the user U.
[0051] "Identification information" is information for identifying the corresponding user U, and may be, for example, biometric information, but may also be a user ID. For example, if the identification information is face information, which is one type of biometric information, an image of the user's face is taken with a camera, and the face image is compared with the identification information to identify the user (user ID).
[0052] The "acoustic processing pattern" is data representing acoustic processing details such as adjustment parameters for head-related transfer functions, and is an acoustic processing pattern for acoustic processing determined for the corresponding (linked) user U. For example, an acoustic processing pattern for the right ear and an acoustic processing pattern for the left ear of the user U may be stored in the user settings 11c. In other words, the information processing device 10 may determine an acoustic pattern for acoustic settings for each of the right ear and the left ear.
[0053] Using such table data, a user can be recognized, and an audio processing pattern suitable for the recognized user can be selected and audio processing can be performed. In addition to selecting an audio processing pattern suitable for the user and performing audio processing, an audio processing pattern suitable for the user can be used as an initial audio processing pattern, and then a calibration process, i.e., a process of selecting an audio processing pattern suitable for the user, can be performed. In this case, the audio processing pattern suitable for the user can be easily updated in response to changes in the user's physical and mental state and the surrounding environment.
[0054] Returning to the explanation of Fig. 4, the control unit 12 will be described. The control unit 12 is a controller, and is realized by, for example, a CPU (Central Processing Unit) or an MPU (Micro Processing Unit) executing various programs (not shown) stored in the storage unit 11 using RAM as a work area. The control unit 12 can also be realized by, for example, an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0055] The control unit 12 has a video output unit 12a, an acquisition unit 12b, an audio output unit 12c, a determination unit 12d, and a seat control unit 12e, and realizes or executes the functions and actions of information processing described below.
[0056] The video output unit 12a executes playback control to display and play back the XR content stored in the XR content DB 11a on the display unit 3. The video output unit 12a also reflects various situation changes detected by the sensor unit 5 and acquired by the acquisition unit 12b (described later) in the playback state of the XR content.
[0057] The acquisition unit 12b acquires sensing data from the sensor unit 5 at any time and outputs the data to the video output unit 12a and the determination unit 12d.
[0058] The audio output unit 12c executes playback control to output and play back audio included in the XR content stored in the XR content DB 11a to the speaker 4. At this time, the audio output unit 12c reads out an audio processing pattern (head related transfer function adjustment parameters, etc.) for audio processing set for the user U according to the identification result of the user U from the user settings 11c, sets the audio processing content, and outputs the audio that has been audio processed.
[0059] This simplifies the calibration operation for each user U regarding the head-related transfer function, and therefore makes it possible to provide an audio environment suited to the head-related transfer function of each user U efficiently.
[0060] Furthermore, during calibration of the sound processing pattern, the sound output unit 12c outputs sound by sequentially switching among the calibration sound signals stored in the sound table 11b.
[0061] The determination unit 12d determines an acoustic processing pattern for acoustic processing settings based on an operation of the user U on the audio output by the audio output unit 12c. For example, when calibrating the acoustic processing pattern, the determination unit 12d refers to the audio table 11b and determines an acoustic signal for calibration to be output by the audio output unit 12c.
[0062] For example, the decision unit 12d instructs the audio output unit 12c to switch the calibration acoustic signal in the first layer at any time, and then instructs the audio output unit 12c to switch the calibration acoustic signal in a layer lower than the calibration acoustic signal selected by the user U at any time.
[0063] Then, for example, the determination unit 12d determines the calibration acoustic signal of the lowest hierarchy as the acoustic processing pattern for the acoustic settings of the user U, and writes it as the user settings 11c to the storage unit 11. At this time, the acoustic processing pattern for the acoustic settings does not have to be the acoustic pattern of the lowest hierarchy, but may be an acoustic pattern of any hierarchy. In other words, if there is an acoustic pattern that matches the head-related transfer function of the user U, it may be determined as the acoustic pattern for the acoustic settings, without going all the way to the lowest hierarchy.
[0064] In this embodiment, when the calibration audio signal of the lowest hierarchical layer is played back, the calibration audio signals above it are also played back so that they can be selected in the same way as the calibration audio signal of the lowest hierarchical layer (as one of the calibration audio signals of the lowest hierarchical layer).
[0065] Furthermore, the determination unit 12d may narrow down the sound processing patterns based on other conditions before selecting a sound processing pattern using a calibration sound signal. For example, images that allow the sound processing patterns to be narrowed down may be presented, and sound output may begin from the layer of the sound processing pattern that corresponds to the image selected by the user U. In this case, for example, images of three different face shapes may be displayed to the user U, and the user U may select an image that is closest to their own face shape to narrow down the sound processing pattern (head related transfer function, etc.) (initial setting sound processing pattern).
[0066] Furthermore, for example, the determination unit 12d may narrow down the sound processing patterns based on a predetermined condition. For example, the determination unit 12d may analyze an image of the head of the user U, estimate the corresponding head-related transfer function of the user U, and then narrow down the sound processing patterns as described above.
[0067] In addition, the decision unit 12d may acquire the sound from a microphone placed inside the pinna of the user U, analyze the sound, estimate the head-related transfer function of the user U, and then narrow down the acoustic processing patterns.
[0068] In such a case, for example, the sound processing patterns can be narrowed down without any operation by the user U, which improves the convenience for the user U in setting the sound processing patterns for sound settings.
[0069] In addition, the determination unit 12d may narrow down the sound processing patterns in accordance with the user U's response operation regarding the relationship between the sound image localization set for each sound processing pattern and the actual sound image localization St actually felt by the user U.
[0070] Here, the answering operation will be described with reference to Fig. 7 to Fig. 9. Fig. 7 to Fig. 9 are explanatory diagrams of the answering operation. For example, the example shown in Fig. 7 shows an example in which the user U is asked to answer an acoustic pattern.
[0071] For example, as shown in the figure, audio signals processed with a plurality of audio processing patterns are output as audio, and the user U is asked to select the audio processing pattern that produces an audio listening (localization) image that is closest to the display object Vo1. In this case, the answering operation may be, for example, audio or gesture operation.
[0072] The object Vo1 is an image that indicates one or more parameters related to the sense of localization, and is preferably a figure whose center of gravity indicates the direction of localization and whose outline (size) indicates the sense of spaciousness. In other words, the object Vo1 is a display figure that visually indicates the set target sense of localization.
[0073] Then, when the determination unit 12d recognizes the result of the answering operation, it instructs the audio output unit 12c to output a sound pattern that is in a lower hierarchical level than the sound processing pattern specified by the answering operation, and then recognizes the result of the answering operation again. By repeatedly performing these processes, the sound processing pattern for sound processing setting is determined.
[0074] That is, in this case, the determination unit 12d narrows down the sound processing patterns for sound processing settings so that the actual sound image localization St matches the specified direction (that is, sound image localization) as a reference.
[0075] Next, the example of Fig. 8 will be described. In the example of Fig. 8, the user U is asked to answer the direction from which a sound is heard. In this case, the determination unit 12d recognizes the direction indicated by the answer operation of the user U by pointing a finger (the direction of the sense of position). Note that the answer operation here may be a gesture operation, and may be, for example, an operation on a touch panel display.
[0076] For example, in this case, the determination unit 12d narrows down the sound patterns according to the sound image localization indicated by the sound processing pattern of the output sound and the direction indicated by the answering operation (pointing) of the user U. Note that, similar to the method shown in Fig. 7, it is also possible to use a method in which the spread of the sense of localization by pointing, that is, the localization area, is recognized and used to narrow down the sound processing patterns.
[0077] That is, in this case, the sound processing patterns for sound processing settings are narrowed down so that the sound image localization matches the actual sound image localization St as a reference.
[0078] For example, as shown in FIG. 9, the determination unit 12d calculates the angle difference θ between the sound image localization Sp indicated by the sound processing pattern and the direction St indicated by the answer operation, and narrows down the sound processing patterns using the angle difference θ.
[0079] That is, in this case, the determination unit 12d determines the sound processing pattern for sound processing settings by successively narrowing down the sound processing patterns so as to find a sound processing pattern whose angular difference θ is closest to 0. Note that at this time, for example, the sound processing pattern to be output next does not have to be a sound processing pattern in a lower hierarchical layer than the sound processing pattern in question, or may be a sound processing pattern in a higher hierarchical layer than the sound processing pattern in question.
[0080] In this way, depending on the user U's response operation, the user can indicate and express the direction and area (spread) of the sense of localization that he or she feels, and based on that, an acoustic processing pattern for the acoustic processing settings can be determined, thereby appropriately matching the sound image localization with the actual sound image localization St that the user U actually feels.
[0081] Furthermore, in these processes, it is preferable that the determination unit 12d performs calibration using a sound processing pattern in which sound image localization that satisfies predetermined perception conditions is set. Here, the perception conditions refer to sound processing patterns in which sound image localization that makes it easy for the user U to perceive sound image localization is set, such as sound processing patterns corresponding to the front, back, left, and right directions of the user U. In other words, sound processing patterns that shift the sense of localization a predetermined distance in the front, back, left, and right directions based on a sound processing pattern in a higher hierarchy are stored as sound processing patterns in a lower hierarchy, and sound signals that have been subjected to this sound processing are sequentially played back, making it easier for the user U to perceive changes in the sense of localization and facilitating accurate selection of a sound processing pattern. Note that a sound processing pattern that shifts the sense of localization a predetermined distance can be determined, for example, by an appropriate distance for easily perceiving the sense of localization through experiments or the like, and that distance can be used.
[0082] In this way, by performing calibration using an acoustic pattern in which sound image localization that satisfies the perception conditions is set, it is possible to improve the accuracy of the calibration.
[0083] Returning to the explanation of Fig. 4, the seat control unit 12e will be described. The seat control unit 12e controls the posture of the seat mechanism 6 and the vibrations to be applied in accordance with posture information related to the posture in the XR content. For example, the seat control unit 12e controls the seat mechanism 6 in accordance with the posture (vibration, tilt, acceleration, etc.) in the XR content, thereby reproducing the sense of realism of the XR content for the user U.
[0084] Next, a processing procedure executed by the information processing device 10 according to the embodiment will be described with reference to Fig. 10. Fig. 10 is a flowchart showing the processing procedure executed by the information processing device 10, which is repeatedly executed at predetermined intervals or the like.
[0085] 10, first, the information processing device 10 determines whether the power of the control system 1 is turned on (step S101). Next, if the information processing device 10 determines that the power of the control system 1 is turned on (step S101; Yes), it identifies the user U (step S102).
[0086] Next, the information processing device 10 determines whether the sound processing pattern for sound settings of the identified user U has been set (step S103), and if it determines that it has not been set (step S103; No), it switches the sound processing pattern as needed and outputs audio (step S104).
[0087] Next, the information processing device 10 determines a sound processing pattern in response to the response operation to the voice (step S105), and determines whether or not to finalize the determined sound processing pattern as the sound processing pattern for sound processing settings (step S106).
[0088] If the information processing device 10 determines that the decision is not final (step S106: No), the process proceeds to step S104, and if the information processing device 10 determines that the decision is final (step S106; Yes), the process proceeds to step S107.
[0089] Then, the information processing device 10 proceeds to XR content playback start processing (step S107) and ends the processing.
[0090] As described above, the information processing device 10 (an example of an audio output device) according to the embodiment includes the acoustic table 11b (an example of an audio pattern storage unit), the audio output unit 12c, and the determination unit 12d. The acoustic table 11b stores adjustment parameters for different head-related transfer functions and calibration audio signals that have been subjected to audio signal processing using the head-related transfer functions based on the adjustment parameters.
[0091] The audio output unit 12c sequentially switches among the calibration audio signals stored in the audio table 11b to output audio. The determination unit 12d determines an audio processing pattern (adjustment parameters) for audio processing settings based on an operation by the user U on the calibration audio signals output by the audio output unit 12c. Therefore, the information processing device 10 according to the embodiment can appropriately perform calibration related to head-related transfer functions.
[0092] In this embodiment, the acoustic table 11b stores adjustment parameters relating to different head-related transfer functions and calibration acoustic signals that have been subjected to acoustic signal processing using the head-related transfer functions based on the adjustment parameters. However, it is also possible to use a method in which, without storing calibration acoustic signals, during calibration, a processed sound source signal (stored separately) is subjected to acoustic signal processing based on the adjustment parameters relating to the head-related transfer functions stored in the acoustic table 11b to generate calibration acoustic signals and play back the audio.
[0093] In the above-described embodiment, the information processing device 10 performs calibration related to acoustic settings, but the present invention is not limited to this. That is, the present invention may be applied to the acoustic settings of any audio system as long as it sets adjustment parameters related to head-related transfer functions.
[0094] Further advantages and modifications will readily occur to those skilled in the art. Therefore, the invention in its broader aspects is not limited to the specific details and representative embodiments shown and described above. Accordingly, various modifications may be made without departing from the spirit or scope of the general inventive concept as defined by the appended claims and their equivalents. [Explanation of symbols]
[0095] 1. Control System 3 Display section 4 speakers 5 Sensor section 6 Seat mechanism 10. Information processing equipment 11 Storage section 11b Acoustic Table 11c User Settings 12a Video output section 12b Acquisition part 12c Audio output section 12d Decision section 12e Seat control unit
Claims
1. A plurality of acoustic processing patterns each having different adjustment parameters related to head-related transfer functions are stored in a hierarchical structure, outputting audio by sequentially switching between audio signals that have been acoustically processed using the plurality of audio processing patterns at the same layer; determining an acoustic processing pattern for the acoustic processing setting of the layer based on a user's operation on the output acoustic signal; Audio output device.
2. storing a calibration acoustic signal that has been acoustically processed using the acoustic processing pattern in association with the acoustic processing pattern; outputting a calibration acoustic signal linked to the acoustic processing pattern; The audio output device according to claim 1 .
3. selecting the acoustic processing pattern based on a captured image of the user's head; The audio signal is output by sequentially switching between audio processing patterns at a lower level than the selected audio processing pattern. The audio output device according to claim 1 .
4. A target localization display unit is provided that displays a graphic that visually indicates the target localization.
4. The audio output device according to claim 1.
5. The user's response operation is a pointing action in a fixed direction.
5. The audio output device according to claim 1.
6. The user's response operation is a pointing action in the localized area.
6. The audio output device according to claim 1.
7. The audio signal is output by sequentially switching between audio processing patterns that satisfy predetermined perceptual conditions.
7. The audio output device according to claim 1.
8. storing the sound processing pattern determined as the sound processing pattern for sound settings for each user; The stored acoustic processing pattern is read out in response to a user and used in the process of determining the acoustic processing pattern for acoustic processing settings.
8. The audio output device according to claim 1.
9. a display device for displaying XR content; an audio output device that outputs audio of the XR content; a speaker that emits the sound output from the sound output device; Equipped with The audio output device is A plurality of acoustic processing patterns each having different adjustment parameters related to head-related transfer functions are stored in a hierarchical structure, outputting audio by sequentially switching between audio signals that have been acoustically processed using the plurality of audio processing patterns at the same layer; The acoustic processing pattern for the acoustic processing setting of the layer is determined based on a user's operation on the output acoustic signal. Control system.
10. Sequentially reading out acoustic processing patterns of the same layer from among a plurality of acoustic processing patterns each having different adjustment parameters related to head-related transfer functions stored in a hierarchical structure; outputting audio while sequentially switching between audio signals that have been acoustically processed using the read audio processing patterns; determining the acoustic processing pattern for the acoustic processing setting of the layer based on a user operation on the output acoustic signal; A calibration method including:
Citation Information
Patent Citations
Portable telephone set and sound adjustment method
JP2008193382A
Head transfer function selection device and acoustic reproduction apparatus
JP2014099797A
Information processing device, information processing method, program and information processing system
WO2020031486A1