Playback device, playback method, information processing device, information processing method, and program
By employing a hybrid HRTF strategy that combines personalized and common HRTFs, the playback device enhances audiobook realism and immersion while minimizing processing load, addressing the challenges of existing audiobook technologies.
Patent Information
- Application Number
- JP2022568171
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-09
- Filing Date
- 2021-11-25
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2041-11-25
AI Technical Summary
Existing audiobook technologies struggle to convey a realistic sense of scenes featuring multiple characters due to the heavy processing load imposed by using personalized Head-Related Transfer Functions (HRTFs) for all audio content.
A playback device and method that utilizes both personalized and common HRTFs to reduce processing load by selectively applying personalized HRTFs to critical audio elements and common HRTFs to less critical elements, enhancing realism while optimizing computational resources.
This approach provides a more realistic and immersive audiobook experience by accurately localizing sound images while reducing the processing burden on both content production and playback systems.
Smart Images

Figure 0007800443000001 
Figure 0007800443000002 
Figure 0007800443000003
Abstract
Description
[Technical Field]
[0001] In particular, the present technology relates to a playback device, a playback method, an information processing device, an information processing method, and a program that are capable of providing content with a sense of realism while reducing the processing load. [Background technology]
[0002] Audiobooks are content that contains recordings of narrators or voice actors reading books. By listening to audiobooks played on a smartphone with headphones, users can experience the sensation of reading a book without having to follow the text with their eyes. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2018-010654 [Patent Document 2] Japanese Patent Application Laid-Open No. 2009-260574 Summary of the Invention [Problem to be solved by the invention]
[0004] Since it is only audio of the book being read aloud, it is difficult to convey the realism of scenes featuring multiple characters.
[0005] It is thought that a sense of realism can be conveyed by reproducing the direction, distance, movement, etc. of sound using the head-related transfer function (HRTF). Calculations using the HRTF, which mathematically represents how sound travels from the sound source to the ear, make it possible to reproduce the sound heard through headphones in three dimensions.
[0006] Because HRTFs differ depending on factors such as the shape of the ear, a personalized HRTF must be used for each user in order to accurately localize the sound image. Using personalized HRTFs for all audio imposes a heavy processing load on both the content production and playback sides.
[0007] The present technology has been developed in light of these circumstances, and makes it possible to provide realistic content while reducing the processing load. [Means for solving the problem]
[0008] A playback device according to a first aspect of the present technology includes a playback processing unit that plays back object audio content including a first object that is played back using a personalized HRTF, which is an HRTF personalized for a listener, and a second object that is played back using a common HRTF, which is an HRTF commonly used by multiple listeners.
[0009] An information processing device according to a second aspect of the present technology includes: a metadata generation unit that generates metadata including a flag indicating whether the object is to be played back as a first object to be played back using a personalized HRTF, which is an HRTF personalized for a listener, or as a second object to be played back using a common HRTF, which is an HRTF commonly used by multiple listeners; and a content generation unit that generates object audio content including sound source data of multiple objects and the metadata.
[0010] An information processing device according to a third aspect of the present technology includes: a content acquisition unit that acquires object audio content including sound source data of a plurality of objects and metadata including a flag indicating whether the object is to be played as a first object to be played using a personalized HRTF, which is an HRTF personalized for a listener, or as a second object to be played using a common HRTF, which is an HRTF commonly used among a plurality of listeners; an audio processing unit that performs processing using the common HRTF on audio data of an object selected based on the flag to be played as the second object; and a transmission data generation unit that generates transmission data of the content including channel-based data generated by the processing using the common HRTF and the audio data of the first object.
[0011] In a first aspect of the present technology, object audio content is played back, which includes a first object played back using a personalized HRTF, which is an HRTF personalized for a listener, and a second object played back using a common HRTF, which is an HRTF commonly used among a plurality of listeners.
[0012] In a second aspect of the present technology, metadata is generated that includes a flag indicating whether the object is to be played as a first object to be played using a personalized HRTF, which is an HRTF personalized for a listener, or as a second object to be played using a common HRTF, which is an HRTF commonly used among multiple listeners, and object audio content is generated that includes sound source data and metadata of multiple objects.
[0013] In a third aspect of the present technology, object audio content is acquired, the object audio content including sound source data of a plurality of objects and metadata including a flag indicating whether the object is to be played back as a first object to be played back using a personalized HRTF, which is an HRTF personalized for a listener, or as a second object to be played back using a common HRTF, which is an HRTF commonly used among a plurality of listeners. Furthermore, processing using the common HRTF is performed on the audio data of the object selected to be played back as the second object based on the flag, and transmission data of the content is generated, the transmission data including channel-based data generated by processing using the common HRTF and the audio data of the first object. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a diagram illustrating an example of the configuration of a content distribution system according to an embodiment of the present technology. [Figure 2] FIG. 10 is a diagram showing a state when audiobook content is being listened to. [Figure 3] FIG. 2 is a diagram showing examples of objects constituting audiobook content. [Figure 4] FIG. 10 is a diagram illustrating an example of how to use HRTFs. [Figure 5] FIG. 10 is a diagram illustrating an example of sound image localization processing. [Figure 6] FIG. 1 is a diagram illustrating an example of HRTF measurement. [Figure 7] FIG. 1 is a diagram showing an example of the arrangement of speakers when measuring HRTFs. [Figure 8] FIG. 2 is a diagram showing the arrangement of speakers on each floor. [Figure 9] FIG. 10 is another diagram showing the arrangement of speakers on each floor. [Figure 10] FIG. 10 is a diagram illustrating an example of object placement. [Figure 11] FIG. 2 is a diagram schematically illustrating the position of an object. [Figure 12] FIG. 1 is a diagram showing the direction from which the sound of an object is heard. [Figure 13] FIG. 10 is a diagram illustrating an example of changing the listening position. [Figure 14] FIG. 2 is a block diagram showing an example of the hardware configuration of a content production device. [Figure 15] FIG. 2 is a block diagram illustrating an example of a functional configuration of the content production device. [Figure 16] FIG. 10 is a diagram illustrating an example of the configuration of audiobook content. [Figure 17] FIG. 2 is a block diagram illustrating an example of a functional configuration of a content management device. [Figure 18] FIG. 10 is a diagram illustrating an example of transmission data. [Figure 19] FIG. 2 is a block diagram illustrating an example of the hardware configuration of a playback device. [Figure 20] FIG. 2 is a block diagram illustrating an example of the functional configuration of a playback device. [Figure 21] A diagram showing an example of an inference model. [Figure 22] 10 is a flowchart illustrating processing of the content production device. [Figure 23] 10 is a flowchart illustrating processing of the content management device. [Figure 24] 10 is a flowchart illustrating processing performed by the playback device. [Figure 25] FIG. 10 is a block diagram showing another example configuration of the content production device. [Figure 26] 1A and 1B are diagrams showing the appearance of earphones having an external sound capture function. [Figure 27] FIG. 10 is a block diagram showing another example configuration of a playback device. [Figure 28] FIG. 10 is a block diagram showing another example configuration of a playback device. [Figure 29] FIG. 10 is a block diagram showing another example configuration of the content production device. DETAILED DESCRIPTION OF THE INVENTION
[0015] Hereinafter, embodiments of the present technology will be described in the following order. 1. Content distribution system 2. Principles of sound image localization using HRTF 3. Audiobook Content Details 4. Configuration of each device 5. Operation of each device 6. Variations 7. Other Examples
[0016] <<Content distribution system>> FIG. 1 is a diagram illustrating an example of the configuration of a content distribution system according to an embodiment of the present technology.
[0017] 1 is composed of a content production device 1, which is a device on the content production side, a content management device 2, which is a device on the content transmission side, and a playback device 11, which is a device on the content playback side. Headphones 12 are connected to the playback device 11.
[0018] The content management device 2 and the playback device 11 are connected via a network 21 such as the Internet. The content production device 1 and the content management device 2 may also be connected via the network 21.
[0019] The content distributed by the content distribution system of Figure 1 is so-called audiobook content, which is a recording of a book being read aloud. Audiobook content includes not only human voices such as narration and dialogue, but also various other sounds such as sound effects, environmental sounds, and background music. Hereinafter, when it is not necessary to distinguish between different types of sound, the sound of audiobook content will be referred to simply as "sound," but in reality, the sound of audiobook content also includes types of sound other than voice.
[0020] The content production device 1 generates audiobook content in response to operations by a producer. The audiobook content generated by the content production device 1 is supplied to the content management device 2.
[0021] The content management device 2 generates transmission data, which is data for transmitting content, based on the audiobook content generated by the content production device 1. In the content management device 2, multiple pieces of transmission data with different data amounts are generated, for example, for each title of the audiobook content. In response to a request from the playback device 11, the content management device 2 transmits a predetermined amount of transmission data to the playback device 11.
[0022] The playback device 11 is configured by a device such as a smartphone. However, the playback device 11 may also be configured by other devices such as a tablet terminal, a PC, or a TV. The playback device 11 receives and plays back audiobook content transmitted from the content management device 2 via the network 21. The playback device 11 transmits audio data obtained by playing back the audiobook content to the headphones 12, which then output the sound of the audiobook content.
[0023] The playback device 11 and the headphones 12 are connected wirelessly via communication of a predetermined standard such as wireless LAN or Bluetooth (registered trademark). The playback device 11 and the headphones 12 may also be connected by wire. In the example of Fig. 1, the device used by the user to listen to audiobook content is the headphones 12, but other devices such as speakers, earphones (in-ear headphones), etc. may also be used.
[0024] A playback device 11, which is a playback device, and headphones 12 are provided for each user who will listen to the audiobook content.
[0025] FIG. 2 is a diagram showing how audiobook content is listened to.
[0026] In the example on the left side of Fig. 2, users wearing headphones 12 are listening to audiobook content around a table T. The audio data constituting the audiobook content includes object audio data. By using object audio, it is possible to localize sound images such as the voices of characters in the audio at positions specified by position information.
[0027] By controlling the position of the sound image of each object according to the creator's intentions, the user will perceive sound images localized around them, as shown on the right side of Figure 2. In the example on the right side of Figure 2, the user perceives the character's line, "Play with me," as coming from the front left, and the birdsong as coming from above. The user also perceives the background music as coming from far away.
[0028] In this way, the audiobook content distributed in the content distribution system of FIG. 1 is content that can be reproduced as stereophonic sound.
[0029] The user can feel a greater sense of realism than when listening to audio that has not been subjected to stereophonic processing, and can also feel immersed in the audiobook content.
[0030] Stereophonic sound is realized using HRTF (Head-Related Transfer Function), which mathematically expresses how sound travels from a sound source to the ears. Calculations using HRTF are performed in the playback device 11, so that sound images are localized at positions set by the production team, and the sound of each object can be heard from its respective position.
[0031] FIG. 3 is a diagram showing an example of objects constituting audiobook content.
[0032] Audiobook content is composed of audio data for a plurality of objects (audio objects), some of which are played back using personalized HRTFs, and some of which are played back using common HRTFs.
[0033] A personalized HRTF is an HRTF that is optimized (personalized) according to the ear shape of the user who will be listening to the audiobook content. Normally, an HRTF differs depending on the shape of the listener's ears, head, etc. A personalized HRTF is a function that represents transfer characteristics that differ for each user.
[0034] On the other hand, a common HRTF is an HRTF that has not undergone such optimization. The same HRTF is commonly used by multiple listeners.
[0035] Sound played using a personalized HRTF has a higher accuracy of sound image localization than sound played using a common HRTF, resulting in a more realistic sound. On the other hand, sound played using a common HRTF has a lower accuracy of sound image localization than sound played using a personalized HRTF, resulting in a sound where the position of the sound image is perceived as more general.
[0036] In the example of Figure 3, objects #1 and #2 are personalized HRTF playback objects that are played back using personalized HRTFs, and objects #3 and #4 are common HRTF playback objects that are played back using a common HRTF.
[0037] FIG. 4 is a diagram showing an example of how HRTFs are used.
[0038] As shown in Figure 4A and B, personalized HRTFs are used to reproduce objects that require a high sense of localization, such as character dialogue or bird calls.
[0039] Furthermore, as shown in FIG. 4C, a common HRTF is used for reproducing objects that do not require a high sense of localization, such as background music.
[0040] Which object is to be reproduced as the personalized HRTF reproduction object or the common HRTF reproduction object is switched according to the information representing the importance of the sense of localization set by the producer. An object set to have an important sense of localization is reproduced as the personalized HRTF reproduction object.
[0041] Thus, the audio book content reproduced in the playback device 11 is the content of object audio in which the personalized HRTF reproduction object and the common HRTF reproduction object are mixed. Details of the data structure of the audio book content will be described later.
[0042] The playback device 11 is a device that reproduces the content of object audio in which the personalized HRTF reproduction object as the first object and the common HRTF reproduction object as the second object are mixed. Note that the reproduction of the content includes not only performing all the processes related to object audio for outputting the voices of the respective objects based on the data of the object audio, but also performing at least a part of all the processes.
[0043] <<Principle of Sound Image Localization Using HRTF>> HRTF is a transfer function that includes the physical effects received by reflection, diffraction, etc. until the sound wave propagates from the sound source to the eardrum. It is defined as the ratio between the sound pressure at a certain point in space in the state where the listener's head is absent and the sound pressure that propagates to the positions of the left and right eardrums of the listener when the center of the listener's head overlaps the same certain point. By superimposing HRTF on the audio signal and simulating the change in sound that occurs when the sound wave propagates through space from the sound source located in space and reaches the listener's eardrum, it becomes possible to localize the sound image for the listener.
[0044] FIG. 5 is a diagram showing an example of sound image localization processing.
[0045] For example, an audio signal after rendering using object position information is input to the convolution processing unit 31. Object audio data includes sound source data of each object as well as position information indicating the position of the object. Audio signals are signals of various sounds such as voice, environmental sound, and background music. Audio signals are composed of an audio signal L, which is a signal for the left ear, and an audio signal R, which is a signal for the right ear.
[0046] The convolution processing unit 31 processes the input audio signal so that the sound of the object is heard as if it were emitted from the positions of the left virtual speaker VSL and the right virtual speaker VSR indicated by the dashed lines on the right side of Fig. 5. In other words, the convolution processing unit 31 localizes the sound image of the sound output from the headphones 12 so that the user U perceives it as sound from the left virtual speaker VSL and the right virtual speaker VSR.
[0047] When there is no distinction between the left virtual speaker VSL and the right virtual speaker VSR, they are collectively referred to as the virtual speaker VS. In the example of Fig. 5, the virtual speakers VS are positioned in front of the user U, and there are two of them.
[0048] The convolution processing unit 31 performs sound image localization processing on the audio signal to output such sound, and outputs the audio signal L and audio signal R after sound image localization processing to the left unit 12L and right unit 12R, respectively, of the headphones 12. The left unit 12L constituting the headphones 12 is worn on the left ear of the user U, and the right unit 12R is worn on the right ear.
[0049] FIG. 6 is a diagram showing an example of HRTF measurement.
[0050] In a specified reference environment, the position of a dummy head DH is set as the position of a listener. Microphones are provided at the left and right ears of the dummy head DH. In addition, a left real speaker SPL and a right real speaker SPR are installed at the positions of the left and right virtual speakers where sound images are to be localized. The real speakers are speakers that are actually installed.
[0051] The sounds output from the left real speaker SPL and the right real speaker SPR are collected at both ears of the dummy head DH, and a transfer function indicating the change in characteristics when the sounds output from the left real speaker SPL and the right real speaker SPR reach both ears of the dummy head DH is measured in advance as an HRTF. Note that, instead of using a dummy head DH, the transfer function may be measured by actually having a person sit and placing microphones near their ears.
[0052] 6, it is assumed that the transfer function of sound from the left real speaker SPL to the left ear of the dummy head DH is M11, and the transfer function of sound from the left real speaker SPL to the right ear of the dummy head DH is M12. It is also assumed that the transfer function of sound from the right real speaker SPR to the left ear of the dummy head DH is M21, and the transfer function of sound from the right real speaker SPR to the right ear of the dummy head DH is M22.
[0053] The HRTF database 32 in FIG. 5 stores information on HRTFs (information on coefficients representing HRTFs), which are transfer functions measured in advance in this manner.
[0054] The convolution processing unit 31 reads and acquires HRTFs corresponding to the positions of the left virtual speaker VSL and the right virtual speaker VSR from the HRTF database 32, and sets them in the filters 41 to 44.
[0055] The filter 41 performs filtering by applying a transfer function M11 to the audio signal L, and outputs the filtered audio signal L to the adder 45. The filter 42 performs filtering by applying a transfer function M12 to the audio signal L, and outputs the filtered audio signal L to the adder 46.
[0056] Filter 43 performs a filtering process of applying a transfer function M21 to the audio signal R, and outputs the audio signal R after the filtering process to the adder 45. Filter 44 performs a filtering process of applying a transfer function M22 to the audio signal R, and outputs the audio signal R after the filtering process to the adder 46.
[0057] The adder 45, which is an adder for the left channel, adds the audio signal L after the filtering process by the filter 41 and the audio signal R after the filtering process by the filter 43, and outputs the added audio signal. The added audio signal is transmitted to the headphones 12, and a sound corresponding to the audio signal is output from the left unit 12L of the headphones 12.
[0058] The adder 46, which is an adder for the right channel, adds the audio signal L after the filtering process by the filter 42 and the audio signal R after the filtering process by the filter 44, and outputs the added audio signal. The added audio signal is transmitted to the headphones 12, and a sound corresponding to the audio signal is output from the right unit 12R of the headphones 12.
[0059] In this way, the convolution processing unit 31 performs convolution processing using the HRTF corresponding to the position (the position of the object) where the sound image is to be localized on the audio signal, and localizes the sound image from the headphones 12 so that the user U feels as if it is emitted from the virtual speaker VS.
[0060] <<Details of Audio Book Content>> <Measurement of HRTF> FIG. 7 is a diagram showing an example of the arrangement of speakers during the measurement of HRTF for audio book content.
[0061] As shown in Figure 7, the HRTF corresponding to each speaker position is measured with the speakers actually placed at various positions with different heights and directions, with the position of the person taking the measurements at the center. Each small circle represents the position of a speaker.
[0062] In the example of Figure 7, three layers are set: a HIGH layer, a MID layer, and a LOW layer, and speakers are placed on each layer at different heights. For example, the MID layer is set at the same height as the ears of the person being measured. The HIGH layer is set at a position higher than the ears of the person being measured. The LOW layer is set at a position lower than the ears of the person being measured.
[0063] The MID layer is composed of a MID_f layer and a MID_c layer set inside the MID_f layer, and the LOW layer is composed of a LOW_f layer and a LOW_c layer set inside the LOW_f layer.
[0064] Figures 8 and 9 show the placement of speakers on each floor. The circles surrounding the numbers indicate the position of the speakers on each floor. The top of Figures 8 and 9 is the front direction of the person taking the measurements.
[0065] Figure 8A shows the arrangement of the speakers in the HIGH layer from above. The HIGH layer speaker group consists of four speakers, Speaker TpFL, Speaker TpFR, Speaker TpLs, and Speaker TpRs, placed on a circle with a radius of 200 cm centered on the position of the person taking the test.
[0066] For example, the speaker TpFL is placed at a position 30 degrees to the left of the person directly in front of the person, and the speaker TpFR is placed at a position 30 degrees to the right of the person directly in front of the person.
[0067] FIG. 8B is a diagram showing the arrangement of the speakers on the MID layer from above.
[0068] The speaker group for the MID_f layer consists of five speakers placed on a circle with a radius of 200 cm centered on the position of the person taking the measurement. The speaker group for the MID_c layer consists of five speakers placed on a circle with a radius of 40 cm centered on the position of the person taking the measurement.
[0069] For example, speaker FL of the MID_f layer and speaker FL of the MID_c layer are placed 30 degrees to the left of the person directly in front of the person making the measurement. Speaker FR of the MID_f layer and speaker FR of the MID_c layer are placed 30 degrees to the right of the person directly in front of the person making the measurement. Each speaker group of the MID_f layer and MID_c layer is made up of five speakers placed in the same direction of the person directly in front of the person making the measurement.
[0070] As shown in A and B of FIG. 9, in the LOW_f layer and the LOW_c layer, speakers are also placed on the circumferences of two circles with different radii, each centered on the position of the person taking the measurement.
[0071] Based on the sound output from each speaker installed in this way, the HRTF corresponding to the position of each speaker is measured. The HRTF corresponding to a certain speaker position (sound source position) is found by measuring the impulse response from that position at the positions of both ears of the person being measured or both ears of a dummy head and expressing it on the frequency axis.
[0072] The HRTF measured using a dummy head may be used as the common HRTF, or a number of people may be actually seated in the position of the person measuring in Figure 7, and the average of the HRTFs measured at the positions of both ears may be used as the common HRTF.
[0073] By using the HRTFs measured in this way for playback, it becomes possible to localize the sound image of an object at various positions with different directions, heights, and distances, allowing users listening to audiobook content to feel as if they are immersed in the world of the story unfolding around them.
[0074] <Object example> FIG. 10 is a diagram showing an example of object placement.
[0075] As shown in Fig. 10, objects are placed at positions P1 and P2 on table T and at positions P3 to P6 around table T. Audio data used to play the audio of each object shown in Fig. 10 is included in the audiobook content.
[0076] In the example of Figure 10, the object placed at position P1 is the sound of building blocks, and the object placed at position P2 is the sound of tea being poured into a cup. The object placed at position P3 is the lines of a boy, and the object placed at position P4 is the lines of a girl. The object placed at position P5 is the lines of a chicken that speaks human language, and the object placed at position P6 is the chirping of a bird. Position P6 is above position P3.
[0077] Positions P11 to P14 are the positions of users as listeners. For example, a user at position P11 hears the lines of the boy at position P3 from the far left, and the lines of the girl at position P4 from the immediate left.
[0078] The objects shown in Fig. 10 placed near the user are each played back as, for example, personalized HRTF playback objects. Objects such as environmental sounds and background music, which are different from the objects shown in Fig. 10, are played back as common HRTF playback objects. Environmental sounds include the sound of a babbling brook and the sound of leaves rustling in the wind.
[0079] FIG. 11 is a diagram showing a schematic diagram of the position of the personalized HRTF playback object.
[0080] 11, it is assumed that users A to D are located at positions P11 to P14 around table T. Each circle represents an object described with reference to FIG.
[0081] For example, objects at positions P1 and P2 are reproduced using HRTFs corresponding to the positions of the speakers in the LOW_c layer. Objects at positions P3, P4, and P5 are reproduced using HRTFs corresponding to the positions of the speakers in the MID_c layer. An object at position P6 is reproduced using an HRTF corresponding to the position of the speaker in the HIGH layer. Position P7 is used as the position of the sound image when expressing the movement of the building block object at position P2.
[0082] Each object is placed so that positions P1 to P7 are absolute positions. That is, when the listening position is different, the sound of each object is heard from a different direction, as shown in FIG.
[0083] For example, in the example of Fig. 12, the sound of a building block placed at position P1 is heard by user A from the far left, and by user B from the near left, while user C hears the sound from the near right, and user D hears the sound from the far right.
[0084] By measuring the HRTF for each listening position using the speaker arrangement described with reference to FIG. 7, the absolute position of such an object can be fixed.
[0085] Since the absolute position of the object is fixed, the user can also change the listening position during playback of the audiobook content, as shown in FIG.
[0086] 13, user A, who was listening to audiobook content at position P11, moves to other positions P12 to P14. Every time user A changes his listening position, the HRTF used to play each object is switched.
[0087] By changing the listening position, users can get closer to the voice of their favorite voice actor or get closer to check the audio content. Also, by changing the listening position, users can enjoy the same audiobook content in different ways.
[0088] <<Configuration of each device>> <Configuration of content production device 1> FIG. 14 is a block diagram showing an example of the configuration of the content production device 1. As shown in FIG.
[0089] The content production device 1 is configured by a computer. The content production device 1 may be configured by one computer having the configuration shown in Fig. 14, or may be configured by multiple computers.
[0090] The CPU 101 , ROM 102 , and RAM 103 are connected to one another by a bus 104 .
[0091] An input / output interface 105 is also connected to the bus 104. An input unit 106 including a keyboard, a mouse, etc., and an output unit 107 including a display, a speaker, etc. are connected to the input / output interface 105. Using the input unit 106, the producer performs various operations related to the production of audiobook content.
[0092] Furthermore, the input / output interface 105 is connected to a storage unit 108 such as a hard disk or nonvolatile memory, a communication unit 109 such as a network interface, and a drive 110 that drives a removable medium 111 .
[0093] The content management device 2 has the same configuration as the content production device 1 shown in Fig. 14. In the following description, the configuration shown in Fig. 14 will be cited as the configuration of the content management device 2 where appropriate.
[0094] FIG. 15 is a block diagram showing an example of the functional configuration of the content production device 1. As shown in FIG.
[0095] The content production device 1 implements a content generation unit 121, an object sound source data generation unit 122, and a metadata generation unit 123. At least some of the functional units shown in Fig. 15 are implemented by the CPU 101 in Fig. 14 executing a predetermined program.
[0096] The content generating unit 121 generates audiobook content by linking the object sound source data generated by the object sound source data generating unit 122 with the object metadata generated by the metadata generating unit 123.
[0097] FIG. 16 is a diagram showing an example of the structure of audio book content.
[0098] As shown in FIG. 16, audiobook content is made up of object sound source data and object metadata.
[0099] 16, the object sound source data includes sound source data (waveform data) of five objects, Object #1 to #5. In the above example, the audio book content includes sound source data of sounds such as the sound of building blocks, the sound of pouring tea into a cup, and the boy's lines as object sound source data.
[0100] The object metadata includes location information, localization importance information, personalized layout information, and other information, such as volume information for each object.
[0101] The position information included in the object metadata is information that indicates the placement position of each object.
[0102] The localization importance information is a flag that indicates the importance of the localization of a sound image. The importance of the localization of a sound image is set for each object, for example, by the producer. The importance of the localization of a sound image is expressed, for example, as a numerical value from 1 to 10. As will be described later, whether each object is to be played back as an object for personalized HRTF playback or as an object for common HRTF playback is selected based on the localization importance information.
[0103] The personalized layout information is information that represents the speaker arrangement when the HRTF is measured. The speaker arrangement as described with reference to Fig. 7 etc. is represented by the personalized layout information. The personalized layout information is used by the playback device 11 when acquiring the personalized HRTF.
[0104] 15 generates sound source data for each object. The sound source data for each object recorded in a studio or the like is imported into the content production apparatus 1 and generated as object sound source data. The content generation unit 121 outputs the object sound source data to the content generation unit 121.
[0105] The metadata generation unit 123 generates object metadata and outputs it to the content generation unit 121. The metadata generation unit 123 is made up of a position information generation unit 131, a position sense importance information generation unit 132, and a personalized layout information generation unit 133.
[0106] The position information generating unit 131 generates position information that indicates the placement position of each object according to settings made by the creator.
[0107] The localization importance information generating unit 132 generates localization importance information in accordance with settings made by the creator.
[0108] The importance of the sense of localization is set according to, for example, the following: Distance from the listening position to the object's location Object height Object Orientation -Is it an object that you want people to focus on and listen to?
[0109] When the importance of localization is set according to the distance from the listening position to the object's position, for example, a high value representing the importance of localization is set for an object located closer to the listening position. In this case, using a threshold distance as a reference, an object located closer to the listening position is played back as an object for personalized HRTF playback, and an object located further away is played back as an object for common HRTF playback.
[0110] When the importance of localization is set according to the height of an object, for example, a high value representing the importance of localization is set for an object located at a high position. In this case, using a threshold height as a reference, objects located at a position higher than the reference are played back as objects for personalized HRTF playback, and objects located at a lower position are played back as objects for common HRTF playback.
[0111] When the importance of localization is set according to the direction of an object, for example, a high value representing the importance of localization is set for an object positioned closer to the front of the listener. In this case, objects positioned within a predetermined range based on the front of the listener are played back as objects for personalized HRTF playback, and objects positioned outside that range, such as behind the listener, are played back as objects for common HRTF playback.
[0112] When the importance of the sense of positioning is set depending on whether the object is one that you want people to listen to with concentration, for example, a high value representing the importance of the sense of positioning is set for an object that you want people to listen to with concentration.
[0113] The importance of the sense of localization may be set by combining at least one of the distance from the listening position to the object's position, the height of the object, the direction of the object, and whether the object is one that the listener wants to concentrate on listening to.
[0114] In this way, the content production device 1 functions as an information processing device that generates metadata for each object that includes a flag that is used as a criterion for selecting whether to play the object as content for personalized HRTF playback or content for common HRTF playback.
[0115] The personalized layout information generating unit 133 generates personalized layout information according to settings made by the creator.
[0116] <Configuration of Content Management Device 2> FIG. 17 is a block diagram showing an example of the functional configuration of the content management device 2. As shown in FIG.
[0117] The content management device 2 includes a content / common HRTF acquisition unit 151, an object audio processing unit 152, a 2ch mix processing unit 153, a transmission data generation unit 154, a transmission data storage unit 155, and a transmission unit 156. At least some of the functional units shown in Fig. 17 are realized by the CPU 101 of Fig. 14 constituting the content management device 2 executing a predetermined program.
[0118] The content / common HRTF acquisition unit 151 acquires the audiobook content generated by the content production device 1. For example, the content / common HRTF acquisition unit 151 controls the communication unit 109 to receive and acquire the audiobook content transmitted from the content production device 1.
[0119] Furthermore, the content / common HRTF acquisition unit 151 acquires a common HRTF for each object included in the audiobook content. Information about the common HRTF measured by, for example, the audiobook content production side is input to the content management device 2.
[0120] The common HRTF may be included in the audiobook content as data constituting the audiobook content, and the common HRTF may be acquired together with the audiobook content.
[0121] The content / common HRTF acquisition unit 151 selects personalized HRTF playback objects and common HRTF playback objects from among the objects constituting the audiobook content, based on the localization importance information included in the object metadata. If the importance of sound image localization is expressed as a numerical value from 1 to 10 as described above, for example, objects whose importance of sound image localization is equal to or greater than a threshold are selected as personalized HRTF playback objects, and objects whose importance is less than the threshold are selected as common HRTF playback objects.
[0122] The content / common HRTF acquisition unit 151 outputs the audio data of the object selected as the personalized HRTF playback object to the transmission data generation unit 154 .
[0123] Furthermore, the content / common HRTF acquisition unit 151 outputs the audio data of the object selected as the object for common HRTF reproduction to the rendering unit 161 of the object audio processing unit 152. The content / common HRTF acquisition unit 151 outputs the acquired common HRTF to the binaural processing unit 162.
[0124] When multiple transmission data sets are generated with different numbers of personalized HRTF playback objects, the selection of personalized HRTF playback objects and common HRTF playback objects based on the localization importance information is repeated by changing the threshold value.
[0125] The object audio processing unit 152 performs object audio processing on the audio data of the object for common HRTF playback supplied from the content / common HRTF acquisition unit 151. The object audio processing, which is a stereophonic process, includes object rendering processing and binaural processing.
[0126] That is, in this example, object audio processing of the common HRTF playback objects among the objects that make up the audiobook content is performed by the content management device 2, which is the device on the transmitting side. This makes it possible to reduce the processing load on the playback device 11 compared to when the audio data of all objects that make up the audiobook content is transmitted directly to the playback device 11.
[0127] The object audio processing unit 152 includes a rendering unit 161 and a binaural processing unit 162 .
[0128] The rendering unit 161 performs rendering processing of the common HRTF reproduction object based on the position information and the like, and outputs the audio data obtained by performing the rendering processing to the binaural processing unit 162. The rendering unit 161 performs rendering processing such as VBAP (Vector Based Amplitude Panning) based on the audio data of the common HRTF reproduction object.
[0129] The binaural processing unit 162 performs binaural processing using a common HRTF on the audio data of each object supplied from the rendering unit 161, and outputs the audio signal obtained by performing the binaural processing. The binaural processing performed by the binaural processing unit 162 includes the convolution processing described with reference to FIGS. 5 and 6.
[0130] When multiple objects are selected as common HRTF reproduction objects, the 2ch mix processing unit 153 performs 2ch mix processing on the audio signals generated based on the audio data of each object.
[0131] By performing the 2ch mix processing, channel-based audio data is generated that includes components of the audio signals L and R of each of the multiple common HRTF reproduction objects. The 2ch mixed audio data (channel-based audio data) obtained by performing the 2ch mix processing is output to the transmission data generation unit 154.
[0132] The transmission data generation unit 154 generates transmission data by linking the audio data of the personalized HRTF playback object supplied from the content / common HRTF acquisition unit 151 with the 2ch mixed audio data of the common HRTF playback object supplied from the 2ch mix processing unit 153. The audio data of the personalized HRTF playback object also includes object metadata such as position information and personalized layout information.
[0133] FIG. 18 is a diagram illustrating an example of transmission data.
[0134] As shown in FIG. 18, multiple types of transmission data with different numbers of personalized HRTF playback objects are generated for each title of audio book content.
[0135] Basically, the more personalized HRTF playback objects there are, the larger the amount of data transmitted. Also, because object audio processing is performed using personalized HRTFs, the more personalized HRTF playback objects there are, the more objects there are with high localization accuracy, and the more realistic the user will feel.
[0136] In the example of FIG. 18, transmission data D1 to D4 are generated as transmission data for audio book content #1 in descending order of data volume.
[0137] The transmission data D1 to D3 are transmission data with a large number of personalized HRTF reproduction objects, transmission data with a medium number of personalized HRTF reproduction objects, and transmission data with a small number of personalized HRTF reproduction objects, respectively. Each of the transmission data D1 to D3 includes audio data of a predetermined number of personalized HRTF reproduction objects as well as 2-channel mixed audio data of common HRTF reproduction objects.
[0138] The transmission data D4 does not include audio data of personalized HRTF reproduction objects, but is transmission data consisting only of 2-channel mixed audio data of common HRTF reproduction objects.
[0139] The transmission data storage unit 155 in FIG. 17 stores the transmission data generated by the transmission data generation unit 154.
[0140] The transmission unit 156 controls the communication unit 109 to communicate with the playback device 11 and transmits the transmission data stored in the transmission data storage unit 155 to the playback device 11. For example, of the transmission data of audiobook content to be played back on the playback device 11, transmission data according to the quality required by the playback device 11 is transmitted. If the playback device 11 requires high quality, transmission data with a large number of personalized HRTF playback objects is transmitted.
[0141] The transmission data may be transmitted in a download format or in a streaming format.
[0142] <Configuration of playback device 11> FIG. 19 is a block diagram showing an example of the hardware configuration of the playback device 11.
[0143] The playback device 11 is configured by connecting a control unit 201 to a communication unit 202, a memory 203, an operation unit 204, a camera 205, a display 206, and an audio output unit 207.
[0144] The control unit 201 is configured with a CPU, a ROM, a RAM, etc. The control unit 201 controls the overall operation of the playback device 11 by executing a predetermined program.
[0145] An application execution unit 201A is implemented in the control unit 201. Various applications (application programs) such as an application for playing audiobook content are executed by the application execution unit 201A.
[0146] The communication unit 202 is a communication module compatible with wireless communication of a mobile communication system such as 5G communication. The communication unit 202 receives radio waves output from a wireless base station and communicates with various devices such as the content management device 2 via the network 21. The communication unit 202 receives information transmitted from the content management device 2 and outputs the information to the control unit 201.
[0147] The memory 203 is configured by a flash memory, etc. The memory 203 stores various information such as applications executed by the control unit 201.
[0148] The operation unit 204 is configured with various buttons and a touch panel provided over the display 206. The operation unit 204 outputs to the control unit 201 information indicating the content of the user's operation.
[0149] The camera 205 takes a photograph in response to an operation by the user.
[0150] The display 206 is configured by an organic EL display, an LCD, etc. The display 206 displays various screens such as the screen of an application for playing audiobook content.
[0151] The audio output unit 207 transmits the audio data supplied from the control unit 201 after object audio processing and the like has been performed to the headphones 12, and outputs the sound of the audiobook content.
[0152] FIG. 20 is a block diagram showing an example of the functional configuration of the playback device 11.
[0153] The playback device 11 implements a content processing unit 221, a localization sense improvement processing unit 222, and a playback processing unit 223. At least some of the functional units shown in Fig. 20 are implemented by the application execution unit 201A in Fig. 19 executing an application for playing back audiobook content.
[0154] The content processing unit 221 is made up of a data acquisition unit 231 , a content storage unit 232 , and a reception setting unit 233 .
[0155] The data acquisition unit 231 controls the communication unit 202 to communicate with the content management device 2 and acquire transmission data of audiobook content.
[0156] For example, the data acquisition unit 231 transmits information about the quality set by the reception setting unit 233 to the content management device 2, and acquires transmission data transmitted from the content management device 2. Transmission data including audio data of the number of personalized HRTF playback objects corresponding to the quality set by the reception setting unit 233 is transmitted from the content management device 2. The transmission data acquired by the data acquisition unit 231 is output to and stored in the content storage unit 232. The data acquisition unit 231 functions as a transmission data acquisition unit that acquires transmission data to be played back from among multiple pieces of transmission data prepared in the content management device 2, which is the transmission source device.
[0157] The data acquisition unit 231 outputs audio data of the personalized HRTF playback object included in the transmission data stored in the content storage unit 232 to the object audio processing unit 251 of the playback processing unit 223. The data acquisition unit 231 also outputs 2ch mixed audio data of the common HRTF playback object to the 2ch mix processing unit 252. When audio data of the personalized HRTF playback object is included in the transmission data, the data acquisition unit 231 outputs personalized layout information included in the object metadata of the transmission data to the positioning improvement processing unit 222.
[0158] The reception setting unit 233 sets the quality of transmission data requested of the content management device 2, and outputs information indicating the set quality to the data acquisition unit 231. The quality of transmission data requested of the content management device 2 is set by, for example, the user.
[0159] The localization improvement processing unit 222 is made up of a localization improvement coefficient acquisition unit 241 , a localization improvement coefficient storage unit 242 , a personalized HRTF acquisition unit 243 , and a personalized HRTF storage unit 244 .
[0160] The localization improvement coefficient acquisition unit 241 acquires localization improvement coefficients to be used in the localization improvement process, based on the speaker layout represented by the personalized layout information supplied from the data acquisition unit 231. The localization improvement coefficients are acquired based on the listening position selected by the user as appropriate.
[0161] For example, information used for adding reverberation is acquired as the localization improvement coefficient. The localization improvement coefficient acquired by the localization improvement coefficient acquisition unit 241 is output to the localization improvement coefficient storage unit 242 and stored therein.
[0162] The personalized HRTF acquisition unit 243 acquires a personalized HRTF corresponding to the user's listening position based on the personalized layout information supplied from the data acquisition unit 231. The personalized HRTF is acquired, for example, from an external device that provides the personalized HRTF. The external device that provides the personalized HRTF may be the content management device 2, or may be a device different from the content management device 2.
[0163] For example, the personalized HRTF acquisition unit 243 controls the camera 205 to capture an image of the ear of the user who is the listener and acquires an ear image. The personalized HRTF acquisition unit 243 controls the communication unit 202 to transmit the personalized layout information and the ear image to an external device and to receive and acquire the personalized HRTF transmitted in response to the transmission of the personalized layout information and the ear image. The personalized HRTF acquired by the personalized HRTF acquisition unit 243 according to the listening position is output to and stored in the personalized HRTF storage unit 244.
[0164] FIG. 21 is a diagram showing an example of an inference model provided in an external device.
[0165] As shown in FIG. 21, an external device that provides personalized HRTFs is provided with a personalized HRTF inferor that receives speaker arrangement information represented by personalized layout information and an ear image as input and outputs personalized HRTFs.
[0166] The personalized HRTF inferrer is an inference model generated by machine learning using sets of HRTFs measured using speakers in various layouts, including the layout shown in Fig. 7, and an image of the subject's ear as training data. The external device infers personalized HRTFs based on the personalized layout information and ear image transmitted from the playback device 11, and transmits the inference results to the playback device 11.
[0167] The personalized HRTF for a listening position is obtained, for example, by correcting the HRTF obtained as the inference result of the personalized HRTF inferencing device according to the listening position. Information about the listening position may be input to the personalized HRTF inferencing device, and the HRTF output from the personalized HRTF inferencing device may be used as the personalized HRTF for the listening position.
[0168] A personalized HRTF inferrer may be provided in the playback device 11. In this case, the personalized HRTF is inferred by the personalized HRTF acquisition unit 243 using the ear image and personalized layout information.
[0169] The playback processing unit 223 in FIG. 20 is made up of an object audio processing unit 251 and a 2ch mix processing unit 252.
[0170] The object audio processing unit 251 performs object audio processing, including rendering processing and binaural processing, on the audio data of the personalized HRTF playback object supplied from the data acquisition unit 231 .
[0171] In this way, object audio processing of the common HRTF reproduction object is performed in the content management device 2 as described above, whereas object audio processing of the personalized HRTF reproduction object is performed in the reproduction device 11.
[0172] The object audio processing unit 251 is made up of a rendering unit 261 , a localization improvement processing unit 262 , and a binaural processing unit 263 .
[0173] The rendering unit 261 performs rendering processing of the personalized HRTF reproduction object based on the position information and the like, and outputs the audio data obtained by performing the rendering processing to the localization improvement processing unit 262. The rendering unit 261 performs rendering processing such as VBAP based on the audio data of the personalized HRTF reproduction object.
[0174] The localization improvement processing unit 262 reads out the localization improvement coefficients from the localization improvement coefficient storage unit 242, and performs localization improvement processing such as adding reverberation on the audio data supplied from the rendering unit 261. The localization improvement processing unit 262 outputs the audio data of each object obtained by performing the localization improvement processing to the binaural processing unit 263. It is also possible to prevent the localization improvement processing by the localization improvement processing unit 262 from being performed.
[0175] The binaural processing unit 263 reads and acquires the personalized HRTF from the personalized HRTF storage unit 244. The binaural processing unit 263 performs binaural processing using the personalized HRTF on the audio data of each object supplied from the localization improvement processing unit 262, and outputs the audio signal obtained by performing the binaural processing.
[0176] The 2ch mix processing unit 252 performs 2ch mix processing based on the 2ch mixed audio data of the common HRTF reproduction object supplied from the data acquisition unit 231 and the audio signal supplied from the binaural processing unit 263 .
[0177] By performing the 2ch mix processing, audio signals L and R containing audio components of the personalized HRTF reproduction object and the common HRTF reproduction object are generated. The 2ch mix processing unit 252 outputs the 2ch mixed audio data obtained by performing the 2ch mix processing to the audio output unit 207, and outputs the audio of each object of the audiobook content from the headphones 12.
[0178] <<Operation of each device>> The operation of each device having the above configuration will now be described.
[0179] <Operation of content production device 1> The processing of the content production device 1 will be described with reference to the flowchart of FIG.
[0180] In step S1, the object sound source data generating unit 122 generates object sound source data for each object.
[0181] In step S2, the position information generating unit 131 sets the placement position of each object in accordance with an operation by the creator, and generates position information.
[0182] In step S3, the localization importance information generating unit 132 sets the importance of the localization of each object in accordance with an operation by the creator, and generates localization importance information.
[0183] In step S4, the personalized layout information generating unit 133 generates personalized layout information for each listening position.
[0184] In step S5, the content generation unit 121 generates audiobook content by linking the object sound source data generated by the object sound source data generation unit 122 with object metadata including location information generated by the location information generation unit 131.
[0185] The above process is performed for each title of audio book content.
[0186] Because object audio data is included as audio data for audiobook content, producers can create audiobook content that is more realistic than if the audio of a book being read were recorded directly as stereo audio, for example.
[0187] Typically, the importance of spatialization differs for each object. Producers can set the importance of spatialization for each object and specify whether to play each object using a personalized HRTF or a common HRTF.
[0188] <Operation of Content Management Device 2> The processing of the content management device 2 will be described with reference to the flowchart of FIG.
[0189] In step S11, the content / common HRTF acquisition unit 151 acquires the audiobook content and the common HRTF.
[0190] In step S12, the content / common HRTF acquisition unit 151 focuses on one object from among the objects that make up the audiobook content.
[0191] In step S13, the content / common HRTF acquisition unit 151 checks the importance of the sense of localization of the object of interest based on the sense of localization importance information.
[0192] In step S14, the content / common HRTF acquisition unit 151 determines whether or not to reproduce the object of interest as a personalized HRTF reproduction object.
[0193] The determination of whether to play the object as a personalized HRTF playback object is made by comparing a value representing the importance of the localization of the object of interest with a threshold value. For example, if the value representing the importance of the localization of the object of interest is equal to or greater than the threshold value, the object is determined to be played as a personalized HRTF playback object.
[0194] If it is determined in step S14 that the object will not be reproduced as a personalized HRTF reproduction object, that is, that the object of interest will be reproduced as a personalized HRTF reproduction object, the process proceeds to step S15.
[0195] In step S15, the object audio processing unit 152 performs object audio processing on the audio data of the object for common HRTF reproduction supplied from the content / common HRTF acquisition unit 151. The object audio processing on the audio data of the object for common HRTF reproduction is performed using the common HRTF.
[0196] If it is determined in step S14 that the object of interest is to be reproduced as a personalized HRTF reproduction object, the process of step S15 is skipped.
[0197] In step S16, the content / common HRTF acquisition unit 151 determines whether or not attention has been paid to all of the objects that make up the audio book content.
[0198] If it is determined in step S16 that there is an object that is not being focused on, the process returns to step S12, and another object is focused on, and the above-described processing is repeated.
[0199] If it is determined in step S16 that all objects have been looked at, the 2ch mix processing unit 153 determines in step S17 whether or not there is an object that has already been rendered.
[0200] If it is determined in step S17 that there is a rendered object, in step S18, the 2ch mix processing unit 153 performs 2ch mix processing on the audio signal generated based on the audio data of the common HRTF reproduction object. The 2ch mixed audio data of the common HRTF reproduction object obtained by performing the 2ch mix processing is supplied to the transmission data generation unit 154.
[0201] If it is determined in step S17 that there is no object that has already been rendered, the process of step S18 is skipped.
[0202] In step S19, the transmission data generation unit 154 generates transmission data by linking the audio data of the personalized HRTF reproduction object with the 2ch mixed audio data of the common HRTF reproduction object. If there is no common HRTF reproduction object, transmission data consisting only of the audio data of the personalized HRTF reproduction object is generated. In this way, transmission data consisting only of the audio data of the personalized HRTF reproduction object may be generated.
[0203] The above process is repeated by changing the threshold value that serves as the basis for determining whether or not to play the data as a personalized HRTF playback object, thereby generating multiple transmission data sets with different numbers of personalized HRTF playback objects and preparing them in the content management device 2.
[0204] For objects for which the sense of positioning is important, by transmitting the object audio as is, the content management device 2 can cause the playback device 11 to perform object audio processing using the personalized HRTF.
[0205] If all objects are transmitted as object audio, it may not be possible to properly perform object audio processing in the playback device 11, depending on the performance of the playback side and the communication environment. By having the content management device 2 perform object audio processing for some objects for which the sense of positioning is not important, the content management device 2 can provide transmission data according to the performance of the playback device 11, allowing the playback device 11 to properly perform the processing.
[0206] <Operation of the playback device 11> The processing of the playback device 11 will be described with reference to the flowchart of FIG.
[0207] In step S31, the localization sense improvement processing unit 222 acquires the user's listening position. The user's listening position is acquired, for example, by the user selecting it on the screen of an application for playing audiobook content. The user's position may be analyzed based on sensor data, and the user's listening position may be acquired.
[0208] In step S32, the data acquisition unit 231 acquires, from among the transmission data prepared in the content management device 2, transmission data that corresponds to the quality setting.
[0209] In step S33, the data acquisition unit 231 determines whether or not the acquired transmission data includes audio data of a personalized HRTF reproduction object.
[0210] If it is determined in step S33 that audio data of a personalized HRTF playback object is included, then in step S34 the data acquisition unit 231 acquires personalized layout information included in the object metadata of the transmission data.
[0211] In step S35, the personalized HRTF acquisition unit 243 controls the camera 205 to capture an image of the user's ear. The ear image is captured in response to, for example, a user operation. Before listening to audiobook content, the user captures an image of their ear with the camera 205 of the playback device 11.
[0212] In step S36, the personalized HRTF acquisition unit 243 acquires personalized HRTFs from, for example, an external device, based on the personalized layout information and the ear image.
[0213] In step S37, the localization improvement processing unit 222 acquires a localization improvement coefficient to be used in the localization improvement process, based on the speaker arrangement and the like represented by the personalized layout information.
[0214] In step S38, the object audio processing unit 251 performs object audio processing on the audio data of the personalized HRTF reproduction object using the personalized HRTF.
[0215] In step S39, the 2ch mix processing unit 252 performs 2ch mix processing based on the 2ch mixed audio data of the common HRTF reproduction object and the audio signal of the personalized HRTF reproduction object obtained by the object audio processing.
[0216] In step S40, the 2ch mix processing unit 252 outputs the 2ch mixed audio data obtained by performing the 2ch mix processing, and outputs the audio of each object of the audiobook content from the headphones 12. The audio output from the headphones 12 includes the audio of the personalized HRTF reproduction object and the audio of the common HRTF reproduction object.
[0217] On the other hand, if it is determined in step S33 that the audio data of the personalized HRTF reproduction object is not included in the transmission data, the processes of steps S34 to S39 are skipped, and the transmission data will contain only the 2ch mixed audio data of the common HRTF reproduction object.
[0218] In this case, in step S40, the audio of the common HRTF reproduction object is output based on the 2ch mixed audio data of the common HRTF reproduction object included in the transmission data.
[0219] Through the above processing, the playback device 11 can localize the sound images of each object around the user.
[0220] For objects for which a sense of localization is important, playback is performed using an HRTF personalized for the user, allowing the playback device 11 to localize the sound image with high precision.
[0221] If object audio processing for all objects were performed in the content management device 2, the playback device 11 would not be able to change the sense of localization of the objects. By performing object audio processing for some objects for which the sense of localization is important on the playback side, the playback device 11 can perform playback using an HRTF personalized for the user, thereby enabling the playback device 11 to change the sense of localization of the objects.
[0222] The user can get a sense of realism compared to when listening to audio that has not been subjected to stereophonic processing.
[0223] Furthermore, by having the content management device 2 perform object audio processing for some objects for which the sense of positioning is not important, it is possible to reduce the processing load on the playback device 11. Generally, the greater the number of objects, the greater the load on computational resources, battery consumption, heat countermeasures, etc. Furthermore, the greater the number of objects, the greater the storage capacity that the playback device 11 needs to secure for storing personalized HRTF coefficients.
[0224] That is, the playback device 11 can provide audiobook content with a sense of realism while reducing the processing load.
[0225] <<Modifications>> <Variation 1: Example in which transmission data is prepared by the production side> The transmission data may be generated in the content production device 1. In this case, the configuration related to the generation of the transmission data, among the above-mentioned configuration of the content management device 2, is provided in the content production device 1.
[0226] FIG. 25 is a block diagram showing another example of the configuration of the content production device 1.
[0227] The content production device 1 shown in FIG. 25 is made up of a content generation processing unit 301 and a transmission data generation processing unit 302.
[0228] The content generation processing unit 301 includes the components described with reference to FIG.
[0229] 17, the transmission data generation processing unit 302 is provided with at least the components related to generation of transmission data, namely the content / common HRTF acquisition unit 151, the object audio processing unit 152, the 2ch mix processing unit 153, and the transmission data generation unit 154. The transmission data generated by the transmission data generation processing unit 302 is provided to the content management device 2.
[0230] The transmission unit 156 (FIG. 17) of the content management device 2 selects transmission data of a predetermined quality from the transmission data generated by the content production device 1, and transmits the selected transmission data to the playback device 11.
[0231] <Modification 2: Example of Transmission Data Selection> Although it has been assumed that information regarding quality is transmitted from the playback device 11 to the content management device 2, and that transmission data corresponding to the quality required by the playback device 11 is selected by the content management device 2, the transmission data to be transmitted may also be selected based on other criteria.
[0232] In this case, information regarding the network environment such as the throughput between the content management device 2 and playback device 11, the device configuration on the playback side, and the processing performance of playback device 11 is transmitted from playback device 11 to content management device 2. Based on the information transmitted from playback device 11, content management device 2 selects transmission data appropriate for the network environment and the like, and transmits it to playback device 11.
[0233] <Modification 3: Example of transmission data selection> Although the device worn by the user is assumed to be headphones 12, earphones capable of taking in external sounds as shown in Fig. 26 may also be used. In this case, audiobook content is listened to in a hybrid style that combines earphones with a speaker (actual speaker) provided on the playback side.
[0234] FIG. 26 is a diagram showing the appearance of earphones with an external sound capture function.
[0235] The earphones shown in FIG. 26 are so-called open-ear type earphones that do not seal the ear canal.
[0236] 26, the right unit of the earphone is configured by joining driver unit 311 and ring-shaped attachment part 313 via U-shaped sound conduit 312. The right unit is attached by pressing attachment part 313 against the area around the ear canal, with attachment part 313 and driver unit 311 sandwiching the right ear.
[0237] The left unit has the same configuration as the right unit, and the left and right units are connected by wire or wirelessly.
[0238] Driver unit 311 of the right unit receives an audio signal transmitted from playback device 11, and outputs sound corresponding to the audio signal from the tip of sound conduit 312, as shown by arrow #1. A hole is formed at the joint between sound conduit 312 and attachment part 313, which outputs sound toward the external ear canal.
[0239] Wearing part 313 has a ring shape. In addition to the sound of the content output from the tip of sound conduit 312, ambient sound also reaches the ear canal, as indicated by arrow #2.
[0240] FIG. 27 is a block diagram showing an example of the configuration of the playback device 11 when audiobook content is listened to in a hybrid style.
[0241] 27, the same components as those described with reference to Fig. 20 are denoted by the same reference numerals. Duplicate descriptions will be omitted as appropriate. The same applies to Fig. 28 described later.
[0242] The configuration of the playback device 11 shown in FIG. 27 differs from the configuration shown in FIG. 20 in that a real speaker configuration information acquisition unit 264 is additionally provided in the object audio processing unit 251, and a real speaker output control unit 271 is additionally provided in the playback processing unit 223.
[0243] The real speaker configuration information acquisition unit 264 acquires playback speaker configuration information, which is information about the configuration of real speakers provided on the playback side. The playback speaker configuration information includes, for example, information representing the arrangement of the real speakers. The arrangement of the real speakers represented by the playback speaker configuration information may be the same as or different from the arrangement of the speakers represented by the personalized layout information. The playback speaker configuration information acquired by the real speaker configuration information acquisition unit 264 is output to the rendering unit 261 and the localization improvement processing unit 262.
[0244] In the rendering unit 261, a rendering process is performed in which certain objects from the personalized HRTF playback objects are assigned to real speakers and other objects are assigned to earphones. The objects to be assigned to speakers / earphones may be selected based on the distance from the object to the listener and the importance of localization information.
[0245] The audio data obtained by the rendering process is subjected to localization improvement processing in a localization improvement processor 262. In the localization improvement process, for example, a process of adding reverberation according to the arrangement of the real speakers is performed. Of the audio data obtained by the localization improvement process, the audio data of the objects assigned to the real speakers is supplied to a real speaker output control unit 271, and the audio data of the objects assigned to the earphones is supplied to a binaural processing unit 263.
[0246] The real speaker output control unit 271 generates audio signals with a number of channels corresponding to the number of real speakers in accordance with the allocation by the rendering unit 261, and outputs the sounds of the objects allocated to the real speakers from the respective real speakers.
[0247] The audio of the object assigned to the earphone is output from the earphone after undergoing 2ch mix processing, etc., so that the user can hear the audio from the earphone along with the audio from the real speakers. In this way, it is possible to listen to audiobook content in a hybrid style that combines earphones and real speakers.
[0248] The audio book content may be listened to using only the actual speaker without using the headphones 12 or earphones.
[0249] <Modification 4: Example of head tracking> So-called head tracking, which corrects the HRTF in accordance with the user's head posture, may be performed when listening to audiobook content. The headphones 12 are equipped with sensors such as a gyro sensor and an acceleration sensor.
[0250] FIG. 28 is a block diagram showing another example of the configuration of the playback device 11.
[0251] The configuration of the playback device 11 shown in FIG. 28 differs from the configuration shown in FIG. 20 in that a head position / posture acquisition unit 281 is additionally provided in the playback processing unit 223.
[0252] The head position / posture acquisition unit 281 detects the position and posture of the user's head based on sensor data detected by a sensor mounted on the headphones 12. Information indicating the head position and posture detected by the head position / posture acquisition unit 281 is output to the rendering unit 261.
[0253] The rendering unit 261 corrects the position of each object represented by the position information so as to fix the absolute position of each object, based on the position and orientation of the head detected by the head position / orientation acquisition unit 281. For example, when the user's head is rotated 30 degrees to the right with respect to a certain direction, the rendering unit 261 corrects the position of each object by rotating the positions of all objects 30 degrees to the left.
[0254] In the localization improvement processing unit 262 and the binaural processing unit 263, for example, processing corresponding to the corrected position is performed on the audio data of each personalized HRTF reproduction object.
[0255] In this way, by performing head tracking when listening to audiobook content, the playback device 11 can localize the sound image of each object at a fixed position without being affected by the movement of the user's head.
[0256] <Variation 5: Example of determining the importance of localization> Although the importance of the sense of localization of each object is set manually by the creator, it may also be set automatically by the content production device 1.
[0257] FIG. 29 is a block diagram showing another example of the configuration of the content production device 1.
[0258] Of the components shown in Fig. 29, the same components as those explained with reference to Fig. 15 are denoted by the same reference numerals, and overlapping explanations will be omitted where appropriate.
[0259] The configuration of the content production device 1 shown in FIG. 29 differs from the configuration shown in FIG. 15 in that a sound source attribute information acquisition unit 291 is additionally provided in the metadata generation unit 123.
[0260] The sound source attribute information acquisition unit 291 acquires attribute information of the sound source data of each object. The attribute information includes, for example, the file name of the file storing the sound source data, the file contents such as the type of sound source data, information estimated from the file name and file contents, etc. The attribute information acquired by the sound source attribute information acquisition unit 291 is output to the localization importance information generation unit 132.
[0261] The localization importance information generating unit 132 sets the importance of the localization of each object based on the attribute information supplied from the sound source attribute information acquiring unit 291 .
[0262] For example, the type of sound source data of each object (such as sound effect or audio) is identified based on the attribute information. When automatic setting is performed with emphasis on sound effects, a high value is set for sound effect objects as a value representing the importance of the sense of localization. When automatic setting is performed with emphasis on audio, a high value is set for audio objects as a value representing the importance of the sense of localization.
[0263] In this way, the importance of the sense of positioning is automatically set, which makes it possible to reduce the burden involved in producing audiobook content.
[0264] <<Other examples>> <Example of process allocation> Object audio processing may be performed on not only the personalized HRTF reproduction object but also the common HRTF reproduction object in the playback device 11. In this case, the audio data of the common HRTF reproduction object is transmitted as object audio from the content management device 2 to the playback device 11 together with the audio data of the personalized HRTF reproduction object.
[0265] The playback device 11 acquires the common HRTF and performs object audio processing of the object for common HRTF playback, and also acquires the personalized HRTF and performs object audio processing of the object for personalized HRTF playback.
[0266] This also enables the playback device 11 to provide audiobook content with a sense of realism.
[0267] <Example of switching objects> The switching between the personalized HRTF playback object and the common HRTF playback object may be performed dynamically, in which case each object is dynamically switched between being played back as a personalized HRTF playback object or a common HRTF playback object.
[0268] For example, if the network environment between the content management device 2 and the playback device 11 is poor or if the processing load on the playback device 11 is heavy, many objects will be selected as objects for common HRTF playback. On the other hand, if the network environment between the content management device 2 and the playback device 11 is good or if the processing load on the playback device 11 is light, many objects will be selected as objects for personalized HRTF playback.
[0269] Depending on whether the audiobook content is transmitted by streaming or by download, the personalized HRTF playback object and the common HRTF playback object may be switched.
[0270] Instead of a flag indicating the importance of the sense of position, a flag indicating the priority of processing may be set for each object.
[0271] Furthermore, instead of a flag indicating the importance of the sense of localization, a flag may be set that directly indicates whether the object is to be reproduced as a personalized HRTF reproduction object or as a common HRTF reproduction object.
[0272] In this way, it is possible to switch between the personalized HRTF reproduction object and the common HRTF reproduction object based on various information specified by the production side.
[0273] <Example of interaction> HRTF acquisition Although the personalized HRTF is acquired based on an ear image or the like, the personalized HRTF may be acquired in response to an interaction by the user.
[0274] Example 1 of obtaining personalized HRTF For example, if a space with speakers arranged as shown in FIG. 7 is prepared on the playback side, the user may be seated at the center of the speakers as the measurer, and personalized HRTFs may be measured.
[0275] Example 2 of obtaining personalized HRTF The personalized HRTF may be measured by having the subject distinguish the direction of a sound, for example, in a game format, such as by having the subject react when hearing the sound of an insect directly in front of the subject.
[0276] Example 3 of obtaining personalized HRTF If multiple personalized HRTF samples are available, the user may be asked to listen to audio played using each sample and select which sample to use based on their perception of localization. The sample selected by the user is used as the personalized HRTF when listening to the audiobook content.
[0277] The personalized HRTF may be corrected in response to a game-style response, a response regarding sample reselection, a response made by the user while viewing content, or the like.
[0278] The personalized HRTF may be obtained by combining the acquisition of the personalized HRTF in a game format with the acquisition of the personalized HRTF using samples.
[0279] In this case, for example, the angular difference between the direction of the object's position represented by the position information and the direction of the position where the user feels the sound image is localized is calculated as the localization error. The localization error calculation is repeated by switching samples, and the sample with the smallest localization error is selected as the personalized HRTF to actually be used.
[0280] The personalized HRTFs may be acquired as described above while listening to audiobook content, and playback using the personalized HRTFs may be performed at the timing when the personalized HRTFs are acquired. In this case, for example, in the first half of the audiobook content, each object is played back as an object for playing back a common HRTF, and in the second half, a specific object is played back as an object for playing back a personalized HRTF.
[0281] Instead of ear images, personalized HRTFs calculated based on information about ear size and head circumference may be acquired.
[0282] When playing content The user may be allowed to perform operations related to the playback while the audiobook content is being played back.
[0283] Playback operations include the option to zoom in or out on the sound source (object) you want to immerse yourself in, whether or not to include sound effects, specifying the narration position, switching languages, whether or not to include the head tracking mentioned above, and selecting the story.
[0284] <Other> The above-described processing can be applied to various types of content including object audio data, such as music content, radio content, movie content, and television program content.
[0285] Computer configuration example The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the program constituting the software is installed from a program recording medium into a computer incorporated in dedicated hardware or a general-purpose personal computer.
[0286] The program to be installed is provided by being recorded on removable media 111 shown in FIG. 14, which may be an optical disk (CD-ROM (Compact Disc-Read Only Memory), DVD (Digital Versatile Disc), etc.) or a semiconductor memory. Alternatively, the program may be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting. The program can be installed in advance in ROM 102 or storage unit 108.
[0287] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0288] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device with multiple modules housed in a single housing, are both systems.
[0289] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0290] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the spirit of the present technology.
[0291] For example, this technology can be configured as cloud computing, in which a single function is shared and processed collaboratively by multiple devices via a network.
[0292] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by multiple devices.
[0293] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0294] <Configuration combination example> The present technology can also be configured as follows.
[0295] (1) The apparatus includes a playback processing unit that plays back object audio content including a first object that is played back using a personalized HRTF, which is an HRTF personalized for a listener, and a second object that is played back using a common HRTF, which is an HRTF commonly used among a plurality of listeners. playback device. (2) The audio system further includes an HRTF acquisition unit that acquires the personalized HRTF based on an image obtained by photographing the listener's ear and layout information that is information about the layout of speakers at the time of measuring the common HRTF and is included in the content. The playback device according to (1) above. (3) The transmission data of the content includes audio data of the first object and channel-based data generated in a device that transmits the content by processing the audio data of the second object using the common HRTF. The playback device according to (1) or (2) above. (4) The playback processing unit outputs sound according to channel-based data generated by performing processing using the personalized HRTF on the audio data of the first object and channel-based data included in the transmission data. The playback device according to (3) above. (5) The playback processing unit processes the audio data of the first object using the personalized HRTF, and processes the audio data of the second object using the common HRTF, and outputs sound corresponding to the channel-based data generated by performing each of the processes. The playback device according to (1) or (2) above. (6) a transmission data acquisition unit that acquires the transmission data to be played back from a plurality of transmission data sets prepared in the transmission source device and having different numbers of the first objects; The playback device according to (3) above. (7) The playback device Reproducing object audio content including a first object reproduced using a personalized HRTF, which is an HRTF personalized for a listener, and a second object reproduced using a common HRTF, which is an HRTF commonly used among a plurality of listeners. How to play. (8) On the computer, Reproducing object audio content including a first object reproduced using a personalized HRTF, which is an HRTF personalized for a listener, and a second object reproduced using a common HRTF, which is an HRTF commonly used among a plurality of listeners. A program for executing a process. (9) a metadata generating unit that generates metadata including a flag indicating whether the object is to be reproduced as a first object using a personalized HRTF, which is an HRTF personalized for a listener, or as a second object using a common HRTF, which is an HRTF commonly used among a plurality of listeners; a content generation unit that generates object audio content including sound source data of a plurality of objects and the metadata; An information processing device comprising: (10) The metadata generation unit generates the metadata further including layout information that is information about the layout of speakers when the common HRTF is measured. The information processing device according to (9) above. (11) The metadata generating unit generates the metadata including, as the flag, information indicating the importance of the sense of localization of the sound image of each object. The information processing device according to (9) or (10). (12) The importance is set based on at least one of the type of object, the distance from the listening position to the position of the object, the direction of the object at the listening position, and the height of the object. The information processing device according to (11) above. (13) The information processing device generating metadata including a flag indicating whether the object is to be played as a first object to be played using a personalized HRTF, which is an HRTF personalized for a listener, or as a second object to be played using a common HRTF, which is an HRTF commonly used among a plurality of listeners; Generate object audio content including sound source data of a plurality of objects and the metadata Information processing methods. (14) On the computer, generating metadata including a flag indicating whether the object is to be played as a first object to be played using a personalized HRTF, which is an HRTF personalized for a listener, or as a second object to be played using a common HRTF, which is an HRTF commonly used among a plurality of listeners; Generate object audio content including sound source data of a plurality of objects and the metadata A program for executing a process. (15) a content acquisition unit that acquires object audio content including sound source data of a plurality of objects and metadata including a flag that indicates whether the object is to be played as a first object to be played using a personalized HRTF, which is an HRTF personalized for a listener, or as a second object to be played using a common HRTF, which is an HRTF commonly used among a plurality of listeners; an audio processing unit that processes audio data of the object selected to be played back as the second object based on the flag, using the common HRTF; a transmission data generation unit that generates transmission data of the content, the transmission data including channel-based data generated by processing using the common HRTF and audio data of the first object; An information processing device comprising: (16) The transmission data generation unit generates a plurality of pieces of transmission data each having a different number of first objects for each piece of content. The information processing device according to (15) above. (17) The information processing device Acquire object audio content including sound source data of a plurality of objects and metadata including a flag indicating whether the object is to be played as a first object to be played using a personalized HRTF, which is an HRTF personalized for a listener, or as a second object to be played using a common HRTF, which is an HRTF commonly used among a plurality of listeners; performing processing using the common HRTF on audio data of the object selected to be played back as the second object based on the flag; generating transmission data for the content, the transmission data including channel-based data generated by processing using the common HRTF and audio data of the first object; Information processing methods. (18) On the computer, Acquire object audio content including sound source data of a plurality of objects and metadata including a flag indicating whether the object is to be played as a first object to be played using a personalized HRTF, which is an HRTF personalized for a listener, or as a second object to be played using a common HRTF, which is an HRTF commonly used among a plurality of listeners; performing processing using the common HRTF on audio data of the object selected to be played back as the second object based on the flag; generating transmission data for the content, the transmission data including channel-based data generated by processing using the common HRTF and audio data of the first object; A program for executing a process. [Explanation of symbols]
[0296] 1 Content production device, 2 Content management device, 11 Playback device, 12 Headphones, 121 Content generation unit, 122 Object sound source data generation unit, 123 Metadata generation unit, 131 Position information generation unit, 132 Localization importance information generation unit, 133 Personalized layout information generation unit, 151 Content / common HRTF acquisition unit, 152 Object audio processing unit, 153 2ch mix processing unit, 154 Transmission data generation unit, 155 Transmission data storage unit, 156 Transmission unit, 161 Rendering unit, 162 Binaural processing unit, 221 Content processing unit, 222 Localization improvement processing unit, 223 Object audio processing unit, 231 Data acquisition unit, 232 Content storage unit, 233 Reception setting unit, 241 Localization improvement coefficient acquisition unit, 242 Localization improvement coefficient storage unit 243 Personalized HRTF acquisition unit, 244 Personalized HRTF storage unit, 251 Object audio processing unit, 252 2ch mix processing unit, 261 Rendering unit, 262 Localization improvement processing unit, 263 Binaural processing unit
Claims
1. a playback processing unit that plays back object audio content including a first object that is played back using a personalized HRTF that is personalized for a listener, and a second object that is played back using a common HRTF that is used commonly by a plurality of listeners; The transmission data of the content includes audio data of the first object and channel-based data generated by processing the audio data of the second object using the common HRTF. playback device.
2. The audio system further includes an HRTF acquisition unit that acquires the personalized HRTF based on an image obtained by photographing the listener's ear and layout information, which is information about the layout of speakers at the time of measuring the common HRTF, included in the content. The playback device according to claim 1 .
3. The channel-based data is data generated by performing processing using the common HRTF in a device that transmits the content. The playback device according to claim 1 .
4. The playback processing unit outputs sound according to channel-based data generated by performing processing using the personalized HRTF on the audio data of the first object and channel-based data included in the transmission data. The playback device according to claim 1 .
5. The playback processing unit processes the audio data of the first object using the personalized HRTF, and processes the audio data of the second object using the common HRTF, and outputs sound corresponding to the channel-based data generated by performing each of the processes. The playback device according to claim 1 .
6. The apparatus further includes a transmission data acquisition unit that acquires the transmission data to be played back from a plurality of transmission data sets prepared in the transmission source device and having different numbers of the first objects. The playback device according to claim 3 .
7. The playback device Reproducing object audio content including a first object reproduced using a personalized HRTF, which is an HRTF personalized for a listener, and a second object reproduced using a common HRTF, which is an HRTF commonly used among a plurality of listeners; The transmission data of the content includes audio data of the first object and channel-based data generated by processing the audio data of the second object using the common HRTF. How to play.
8. On the computer, Reproducing object audio content including a first object reproduced using a personalized HRTF, which is an HRTF personalized for a listener, and a second object reproduced using a common HRTF, which is an HRTF commonly used among a plurality of listeners; The transmission data of the content includes audio data of the first object and channel-based data generated by processing the audio data of the second object using the common HRTF. A program for executing a process.
9. a metadata generating unit that generates metadata including a flag indicating whether the first object is to be reproduced using a personalized HRTF, which is an HRTF personalized for a listener, or the second object is to be reproduced using a common HRTF, which is an HRTF commonly used among a plurality of listeners; a content generation unit that generates object audio content including sound source data of a plurality of objects and the metadata; An information processing device comprising:
10. The metadata generation unit generates the metadata further including layout information that is information about the layout of speakers when the common HRTF is measured. The information processing device according to claim 9 .
11. The metadata generating unit generates the metadata including, as the flag, information indicating the importance of the sense of localization of the sound image of each object. The information processing device according to claim 9 .
12. The importance is set based on at least one of the type of object, the distance from the listening position to the position of the object, the direction of the object at the listening position, and the height of the object. The information processing device according to claim 11.
13. The information processing device Generate metadata including a flag indicating whether the object is to be played as a first object to be played using a personalized HRTF, which is an HRTF personalized for a listener, or as a second object to be played using a common HRTF, which is an HRTF commonly used among a plurality of listeners; and generate object audio content including the sound source data of the plurality of objects and the metadata. Information processing methods.
14. On the computer, Generate metadata including a flag indicating whether the object is to be played as a first object to be played using a personalized HRTF, which is an HRTF personalized for a listener, or as a second object to be played using a common HRTF, which is an HRTF commonly used among a plurality of listeners; and generate object audio content including the sound source data of the plurality of objects and the metadata. A program for executing a process.
15. a content acquisition unit that acquires object audio content including sound source data of a plurality of objects and metadata including a flag that indicates whether the object is to be played as a first object to be played using a personalized HRTF, which is an HRTF personalized for a listener, or as a second object to be played using a common HRTF, which is an HRTF commonly used among a plurality of listeners; an audio processing unit that processes audio data of the object selected to be reproduced as the second object based on the flag, using the common HRTF; a transmission data generation unit that generates transmission data of the content, the transmission data including channel-based data generated by processing using the common HRTF and audio data of the first object; An information processing device comprising:
16. The transmission data generation unit generates a plurality of pieces of transmission data each having a different number of first objects for each piece of content. The information processing device according to claim 15.
17. The information processing device Acquire object audio content including sound source data of a plurality of objects and metadata including a flag indicating whether the object is to be played as a first object to be played using a personalized HRTF, which is an HRTF personalized for a listener, or as a second object to be played using a common HRTF, which is an HRTF commonly used among a plurality of listeners; processing audio data of the object selected to be reproduced as the second object based on the flag using the common HRTF; generating transmission data for the content, the transmission data including channel-based data generated by processing using the common HRTF and audio data of the first object; Information processing methods.
18. On the computer, Acquire object audio content including sound source data of a plurality of objects and metadata including a flag indicating whether the object is to be played as a first object to be played using a personalized HRTF, which is an HRTF personalized for a listener, or as a second object to be played using a common HRTF, which is an HRTF commonly used among a plurality of listeners; processing audio data of the object selected to be reproduced as the second object based on the flag using the common HRTF; generating transmission data for the content, the transmission data including channel-based data generated by processing using the common HRTF and audio data of the first object; A program for executing a process.
Citation Information
Patent Citations
Sound signal processing device, sound signal processing method and mobile terminal equipped with the sound signal processing device
JP2009260574A
Sound processor, sound processing method and sound processing program
JP2015130550A
System and method for optimizing transfer of downloadable content
JP2018010654A
Head transfer function learning device and head transfer function inference device
JP2020170938A
System and method for processing audio across multiple audio spaces
JP2020174346A