Mobile reproduction device, fixed reproduction device, information processing device, and acoustic processing system
The integration of fixed and mobile playback devices with synchronized sound output and delay compensation allows for comprehensive sound image localization in various locations, overcoming the limitations of traditional surround sound systems.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SONY GROUP CORP
- Filing Date
- 2025-10-28
- Publication Date
- 2026-05-21
AI Technical Summary
Existing surround sound audio systems struggle to localize sound images in front of or behind speakers, especially when mobile devices are involved, making it difficult to achieve comprehensive sound image localization in various locations.
A system combining fixed and mobile playback devices that synchronize sound output based on position information, using a control device to render and select playback devices for optimal sound localization, compensating for communication and playback delays, and employing distance-based rendering algorithms.
Enables precise sound image localization in various locations within a content playback space, reducing the need for high-density speaker installation and maintaining cost-effectiveness while enhancing sound quality and synchronization.
Smart Images

Figure JP2025037730_21052026_PF_FP_ABST
Abstract
Description
Mobile playback device, stationary playback device, information processing device, and sound processing system
[0001] This technology relates to a mobile playback device, a fixed playback device, an information processing device, and an acoustic processing system, and more particularly to a mobile playback device, a fixed playback device, an information processing device, and an acoustic processing system that enable sound image localization in various locations by combining a speaker with a fixed installation position and a speaker with an unfixed installation position.
[0002] In a typical surround sound audio system, sound is output from multiple speakers placed around the listener, thereby achieving sound image localization to a predetermined location. For example, Patent Document 1 proposes a technology in which at least one of the multiple speakers constituting a surround sound audio system is replaced with a mobile device such as a smartphone.
[0003] Special table 2016-504824 publication
[0004] When speakers, including those for mobile devices, are placed around the listener, it is easy to localize the sound image between one speaker and another, but it is difficult to localize the sound image in front of or behind the speakers. Furthermore, when the listener is carrying a mobile device, that is, when at least one of the multiple speakers constituting the surround audio system is moving while emitting sound, achieving sound image localization has been difficult.
[0005] This technology was developed in consideration of these circumstances, and by combining speakers with fixed installation positions and speakers with flexible installation positions, it enables sound image localization in various locations.
[0006] The first aspect of this technology is a mobile playback device that can move while playing back the waveform signal of audio content in synchronization with a fixed playback device fixed in a content playback space, which is a space in which audio content is provided, and comprises a position acquisition unit that acquires position information of the device itself within the content playback space, and a playback unit that plays back the waveform signal of the audio content rendered based on information regarding the sound image localization of the audio content and the position information of the device itself.
[0007] The second aspect of this technology is a fixed playback device which is fixed within a content playback space, which is a space in which audio content is provided, and plays a second waveform signal of the audio content in synchronization with a movable playback device which is movable while playing a first waveform signal of the audio content, and comprises a playback unit which plays the second waveform signal rendered based on information regarding the sound image localization of the audio content and position information of the device itself within the content playback space.
[0008] The third aspect of this technology is an information processing device connected to a mobile playback device that can move within a content playback space, which is a space in which audio content is provided, while playing back a first waveform signal of the audio content, and a fixed playback device that is fixed within the content playback space and plays back a second waveform signal of the audio content in synchronization with the mobile playback device. The information processing device comprises a rendering unit that renders the first waveform signal and the second waveform signal based on information regarding the sound image localization of the audio content, the position information of the mobile playback device, and the position information of the fixed playback device, and a communication unit that transmits the first waveform signal to the mobile playback device and transmits the second waveform signal to the fixed playback device.
[0009] In the first aspect of this technology, a mobile playback device, which is movable while reproducing the waveform signal of the audio content in synchronization with a fixed playback device fixed within the content playback space, which is the space in which the audio content is provided, acquires the position information of the device itself within the content playback space, and reproduces the waveform signal of the audio content rendered based on the sound image localization information of the audio content and the position information of the device itself.
[0010] In a second aspect of this technology, a fixed playback device plays a second waveform signal of the audio content in synchronization with a mobile playback device, which is fixed within a content playback space where audio content is provided and can move while playing a first waveform signal of the audio content. The second waveform signal is rendered based on information regarding the sound image localization of the audio content and the position information of the device itself within the content playback space.
[0011] In a third aspect of this technology, a mobile playback device that can move while playing a first waveform signal of the audio content within a content playback space, which is a space in which audio content is provided, and an information processing device connected to a fixed playback device fixed within the content playback space and synchronizing with the mobile playback device to play a second waveform signal of the audio content, renders the first waveform signal and the second waveform signal based on information regarding the sound image localization of the audio content, the position information of the mobile playback device, and the position information of the fixed playback device. The first waveform signal is then transmitted to the mobile playback device, and the second waveform signal is transmitted to the fixed playback device.
[0012] This is a diagram illustrating the outline of a content playback system to which this technology is applied. This is a diagram illustrating a general surround audio system. This is a diagram illustrating an example of sound effects provided by a content playback system. This is a block diagram illustrating example configurations of a control device, a fixed playback device, and a portable playback terminal. This explains the processing performed by a content playback system having the configuration shown in Figure 4. This is a diagram illustrating the selection algorithm for playback devices that output sound. This is a diagram illustrating an example of sound pressure gain according to the distance between the sound source and the playback terminal. This is a diagram illustrating a distance-based rendering algorithm. This is a diagram illustrating the flow of communication delay compensation. This is the first diagram illustrating the flow of playback delay compensation. This is the second diagram illustrating the flow of playback delay compensation. This is a block diagram illustrating the first modified configuration of a control device, a fixed playback device, and a portable playback terminal. This is a diagram illustrating an example of the sound image range of a certain sound source. This is a diagram illustrating another example of the sound image range of a certain sound source. This is a diagram illustrating the appearance of a content playback space such as a theme park or event venue. This is a diagram illustrating an algorithm for selecting a predetermined number of playback devices. This is a diagram illustrating the designation of the localization position of the sound image of each sound source by a person with special authority. This is a diagram illustrating the addition of a sense of sound localization using grouped portable playback terminals 2. This is a block diagram illustrating a second modified configuration of a control device, a fixed playback device, and a portable playback terminal. This is a sequence diagram illustrating the processing performed by a content playback system having the configuration shown in Figure 19. This is a diagram illustrating distance-based rendering algorithms and vector-based rendering algorithms. This is the first diagram illustrating a use case for the content playback system. This is the second diagram illustrating a use case for the content playback system. This is a block diagram illustrating an example of computer hardware configuration.
[0013] The following describes the configurations for implementing this technology. The explanation will proceed in the following order: 1. Overview of the content playback system 2. Configuration and operation of each device 3. Variations
[0014] <1. Overview of the Content Playback System> Figure 1 is a diagram illustrating the overview of a content playback system to which this technology is applied.
[0015] The content playback system shown in Figure 1 is an audio processing system that provides content, such as the voices of performers and the music they play, to the audience (listeners) in a live music venue where a live performance is taking place. The content provided by this technology's content playback system is, for example, audio content consisting of sounds from one or more sound sources, and includes any sound such as the voices of people or characters, the sounds of musical instruments, music, sound effects, and background sounds. Hereafter, the space where the content is provided to the listener, such as a live music venue, will also be referred to as the content playback space.
[0016] The content playback system is configured such that, for example, a fixed playback device 1 and a portable playback terminal (not shown) carried by the audience are connected to a control device (not shown) that controls the fixed playback device 1 and the portable playback terminal.
[0017] The fixed playback device 1 is configured to include a speaker installed at the live venue. The portable playback terminal is configured as a smartphone, tablet, PC, etc., and is configured to include a speaker. Note that the portable playback terminal is an example of a mobile playback device in this disclosure. The mobile playback device is a device that can move while outputting sound, and may be configured as a speaker that can change position along a rail, for example. Hereinafter, when there is no need to particularly distinguish between the fixed playback device 1 and the portable playback terminal, they will simply be referred to as playback devices.
[0018] A playback device outputs multi-channel sound from its speakers, which is a mix of multiple sound sources, such as the voices of performers at a live venue (e.g., on stage) and the music they are playing. In other words, a content playback system can be described as a speaker system that supports multi-channel audio.
[0019] Figure 2 is a diagram illustrating a typical surround sound audio system.
[0020] Figure 2 shows a speaker system corresponding to 5.1ch as an example of a typical surround audio system. As shown in Figure 2, the surround audio system consists of a center speaker 1C, front speakers 1FL and 1FR, rear speakers 1RL and 1RR, and a subwoofer 1W, with each speaker positioned to surround the listener U1.
[0021] In such surround audio systems, sound from each corresponding channel is output from each speaker, creating an acoustic effect where the sound image of each sound source is localized in various locations. When creating such an acoustic effect in a surround audio system, it is easy to localize the sound image on the region connecting the positions of each speaker (for example, the circumference C1), but it is difficult to localize the sound image in the region inside the circumference C1 (A2) or in the region outside of the circumference C1.
[0022] By densely installing speakers within area A2, it is conceivable to achieve an acoustic effect where the sound image of each sound source is localized in various locations, but this would be costly. The same applies to live venues; with only the speakers installed in a live venue, it is difficult to localize sound images in areas closer to the listener (foreground) or areas further away from the listener (background) than the area connecting the speakers.
[0023] Therefore, the content playback system of this technology renders a waveform signal for the channel corresponding to the mobile playback device (first waveform signal) and a waveform signal for the channel corresponding to the fixed playback device (second waveform signal) based on the position information of a mobile playback device (portable playback terminal) that can move while outputting sound within the content playback space, the position information of a fixed playback device 1 that is fixed within the content playback space and outputs sound in synchronization with the mobile playback device, and information regarding the localization position of the sound image of the sound source of the content. The mobile playback device plays the first waveform signal and outputs the sound indicated by the first waveform signal, and the fixed playback device 1 plays the second waveform signal and outputs the sound indicated by the second waveform signal.
[0024] During a music live performance, it is assumed that there are a large number of audiences in areas within the live venue where speakers are not installed. By outputting sound from the speakers of the portable playback terminals that exist densely within the live venue, the content playback system can achieve sound image localization not only in the area connecting the speakers of the fixed playback device 1 but also in areas closer to the listeners than that area.
[0025] For example, as shown in A of FIG. 3, the content playback system can provide the audience with an acoustic effect such that the voice of the performer and the music can be heard from various positions such as the audience seats and other locations. Also, for example, as shown in B of FIG. 3, the content playback system can provide the audience with an acoustic effect (an acoustic effect like a wave) where the sound image of the sound source that was localized at the position of the speaker of the fixed playback device 1 in the front right as seen from the audience moves around the audience seats.
[0026] <2. Configuration and Operation of Each Device> FIG. 4 is a block diagram showing configuration examples of the control device 12, the fixed playback device 1, and the portable playback terminal 2 respectively.
[0027] The control device 12 is an information processing device such as a PC that functions as an audio mixer, for example. As shown in FIG. 4, the control device 12 is composed of a control unit 51, a rendering unit 52, an input unit 53, and a communication unit 54.
[0028] By executing a program stored in a memory (not shown), etc., the control unit 51 realizes a functional block including a synchronization control unit 61, a playback control unit 62, and a speaker selection unit 63.
[0029] The synchronization control unit 61 executes processing for synchronizing the sounds output from the speakers of each playback device. For example, the synchronization control unit 61 calculates the standby time for communication delay compensation and the standby time for playback delay compensation for each playback terminal, and transmits a playback trigger to each playback device. Details of the processing for synchronizing the sounds output from the speakers of each playback device will be described later with reference to FIGS. 9 to 11.
[0030] The playback control unit 62 determines the content to be provided to the audience in the content playback space. The playback control unit 62 receives content specifications, for example, from a mixing engineer, via the input unit 53.
[0031] The playback control unit 62 determines the localization position of the sound image of each sound source in the content provided to the audience. The playback control unit 62 receives, for example, a specification of the localization position of the sound image of each sound source from the mixing engineer via the input unit 53. The specification of the localization position of each sound image may be done during the live music performance or in advance.
[0032] The playback control unit 62 supplies metadata indicating the localization position of the sound image of each sound source to the speaker selection unit 63 and the rendering unit 52. The metadata can be described as information regarding the localization position of the sound image of each sound source in the content.
[0033] The speaker selection unit 63 receives position and orientation information, which indicates the position and orientation of each portable playback terminal 2, transmitted from each portable playback terminal 2 via the communication unit 54. Based on the localization position of the sound image of each sound source, as well as the position and orientation of each playback device, determined by the playback control unit 62, the speaker selection unit 63 selects a playback device (speaker) that outputs sound (reproduces waveform signals) from among a plurality of playback devices located in the content playback space.
[0034] Here, it is assumed that the position and orientation of each fixed playback device 1 are known. For example, after the organizer of a music live event actually installs the fixed playback device 1 (speaker) in the live venue, they input the coordinates indicating the position of the fixed playback device 1 on the UI screen of the control device 12. The coordinates entered by the music live event organizer are stored in the control device 12 as position and orientation information of the fixed playback device 1. The coordinates indicating the position of the fixed playback device 1 can be obtained, for example, by scanning the live venue using a camera, or by obtaining them from a diagram of the live venue.
[0035] The rendering unit 52 renders waveform signals (content waveform signals) for channels corresponding to each of the multiple playback devices that output sound. The rendering of the waveform signals is performed based on the localization position of the sound image of each sound source determined by the playback control unit 62, and the position and orientation of the playback devices selected by the speaker selection unit 63 as playback devices that output sound. The rendering unit 52 transmits the waveform signals for channels corresponding to the destination playback device to each playback device that outputs sound via the communication unit 54.
[0036] The input unit 53 is configured as various input devices for operating the control device 12 and inputting various data into the control device 12. The input unit 53 accepts input from operators such as mixing staff and music live event organizers.
[0037] The communication unit 54 communicates with external devices via the network.
[0038] The fixed playback device 1 is composed of a control unit 71, a signal processing unit 72, a playback unit 73, and a communication unit 74.
[0039] The control unit 71 implements a functional block including the synchronization control unit 81 by executing a program stored in a memory (not shown) or the like.
[0040] The synchronization control unit 81 performs processing to synchronize the sound output from the speakers of each playback device. For example, the synchronization control unit 81 measures the RTT of its own device, controls the output and recording of test sounds to calculate the waiting time for playback delay compensation, and receives playback triggers transmitted from the control device 12 to control the timing of sound output. Details of the processing to synchronize the sound output from the speakers of each playback device will be described later with reference to Figures 9 to 11.
[0041] The signal processing unit 72 receives the waveform signal transmitted from the control device 12 via the communication unit 74 and performs playback processing such as D / A conversion and amplification on the waveform signal. The waveform signal transmitted from the control device 12 can be said to be a signal that works in conjunction with other playback devices (especially the portable playback terminal 2) to localize the sound image of the content's sound source to a predetermined localization position.
[0042] The playback unit 73 is configured, for example, as a speaker installed in a live venue, and reproduces the waveform signal supplied from the signal processing unit 72 to output the sound indicated by the waveform signal. Alternatively, the playback unit 73 may be configured as multiple speakers installed in various locations within the live venue. In this case, the control device 12 transmits waveform signals for channels corresponding to each of the multiple speakers to the fixed playback device 1.
[0043] The communication unit 74 communicates with external devices via the network.
[0044] The portable playback terminal 2 is composed of a control unit 91, a signal processing unit 92, a playback unit 93, an input unit 94, and a communication unit 95.
[0045] The control unit 91 implements a functional block including the synchronization control unit 101 and the position acquisition unit 102 by executing a program stored in a memory (not shown) or the like.
[0046] The synchronization control unit 101 performs processing to synchronize the sound output from the speakers of each playback device. For example, the synchronization control unit 101 measures the RTT of its own device, controls the output and recording of test sounds to calculate the waiting time for playback delay compensation, and receives playback triggers transmitted from the control device 12 to control the timing of sound output. Details of the processing to synchronize the sound output from the speakers of each playback device will be described later with reference to Figures 9 to 11.
[0047] The position acquisition unit 102 acquires position and orientation information indicating the position and orientation of its own device using technologies such as VPS (Visual Positioning System) and SLAM (Simultaneous Localization and Mapping). VPS and SLAM are technologies that estimate the position and orientation of the portable playback terminal 2 based on camera images taken by a camera mounted on the portable playback terminal 2. The position acquisition unit 102 transmits the acquired position and orientation information to the control device 12 via the communication unit 95.
[0048] The signal processing unit 92 receives the waveform signal transmitted from the control device 12 via the communication unit 95 and performs playback processing such as D / A conversion and amplification on the waveform signal. The waveform signal transmitted from the control device 12 can be said to be a signal that works in conjunction with other playback devices (especially the fixed playback device 1) to localize the sound image of the content's sound source to a predetermined localization position.
[0049] The playback unit 93 is configured, for example, as a speaker mounted on a portable playback terminal 2, and plays back a waveform signal supplied from the signal processing unit 92 and outputs the sound represented by the waveform signal.
[0050] The input unit 94 is configured as various input devices for operating the portable playback terminal 2 and inputting various data into the portable playback terminal 2. The input unit 94 accepts input from the listener.
[0051] The communication unit 95 communicates with external devices via the network.
[0052] Next, referring to the sequence diagram in Figure 5, we will explain the processes performed by the content playback system having the above configuration.
[0053] In step S1, the playback control unit 62 of the control device 12 determines the content to be provided to the audience in the content playback space.
[0054] In step S21, the position acquisition unit 102 of the portable playback terminal 2 acquires position and orientation information of its own device and transmits it to the control device 12. The speaker selection unit 63 of the control device 12 receives the position and orientation information transmitted from the portable playback terminal 2.
[0055] In step S2, the playback control unit 62 determines the localization position of the sound image of each sound source in the content provided to the audience. The playback control unit 62 also sets a selection algorithm for the playback terminal that outputs sound and a waveform signal rendering algorithm based on the localization position of the sound image of each sound source, the position and orientation of each fixed playback device 1, and the position and orientation of each portable playback terminal 2.
[0056] In step S3, the speaker selection unit 63 of the control device 12 selects a playback device (speaker) that outputs sound from among a plurality of playback devices located in the content playback space based on the sound image localization positions of each sound source and the positions and orientations of each playback device.
[0057] FIG. 6 is a diagram for explaining an algorithm for selecting a playback device that outputs sound. In FIG. 6, the solid black speaker icons indicate the positions of the fixed playback devices 1FL, 1FR, 1RL, 1RR (speakers), and the hollow speaker icons indicate the positions of the portable playback terminals 2 (speakers).
[0058] Assuming that the variable for identifying the sound source is j, as shown in FIG. 6, the speaker selection unit 63 selects, as the playback device that outputs the sound of the sound source j, the playback devices located within a circular sound image range A11 with a radius r centered on the sound image localization position (coordinate position) P of the sound source j. The localization position P, and the radius r corresponding to the size of the sound image of the sound source j are set for each sound source. Sj centered on the sound image localization position (coordinate position) P of the sound source j j of the sound source j is selected as the playback device that outputs the sound of the sound source j. The localization position P Sj and the radius r corresponding to the size of the sound image of the sound source j j are set for each sound source.
[0059] Specifically, assuming that the variable for identifying the playback device is i and the position of the playback device i is P Di the speaker selection unit 63 selects, as the playback device that outputs the sound of the sound source j, the playback device i that satisfies the following formula (1).
[0060] Dist(P Sj , P Di ) < r j ... (1)
[0061] In formula (1), Dist(P Si , P Di ) indicates the distance between the sound image localization position P Sj of the sound source j and the position P Di of the playback device i.
[0062] Returning to FIG. 5, in step S4, the rendering unit 52 of the control device 12 renders the waveform signals of the channels corresponding to the playback devices that output sound using the waveform signals indicating the sounds of each sound source. Specifically, the rendering unit 52 uses the sound source j 1 and the sound source j2 When the playback device i included in the sound source range is selected as the processing target, the sound source j 1 The sound source j is multiplied by a sound pressure gain corresponding to the distance between it and the playback device i. 1 The waveform signal, and the sound source j 2 The sound source j is multiplied by a sound pressure gain corresponding to the distance between it and the playback device i. 2 The waveform signals are combined to generate the waveform signal for the channel corresponding to the playback device i.
[0063] Figure 7 shows an example of sound pressure gain depending on the distance between the sound source and the playback device.
[0064] As shown in Figure 7, for example, the sound pressure gain is multiplied by the waveform signal of the sound source, with the gain set to increase as the distance between the sound source and the playback device decreases. In other words, among the playback devices that output sound from the sound source, the closer the playback device is to the localization position of the sound image of the sound source, the higher the sound pressure output of the sound source.
[0065] Returning to Figure 5, the rendering unit 52 transmits waveform signals for the channels corresponding to the destination playback devices to the portable playback terminal 2 and the fixed playback device 1 that output sound. The signal processing unit 92 of the portable playback terminal 2 and the signal processing unit 72 of the fixed playback device 1 receive the waveform signals transmitted from the control device 12.
[0066] In step S22, the signal processing unit 92 performs waveform signal playback processing, and the playback unit 93 of the portable playback terminal 2 outputs the sound represented by the waveform signal.
[0067] In step S41, the signal processing unit 72 performs waveform signal playback processing, and the playback unit 73 of the fixed playback device 1 outputs the sound indicated by the waveform signal.
[0068] Through the above process, it becomes possible to achieve sound localization to various locations using the portable playback device 2 carried by the audience. Since it is not necessary to install speakers at high density in the live venue, it is possible to create sound effects that localize the sound images of performers' voices and music to various locations while keeping costs down.
[0069] In the example described above, the selection of the playback device that outputs sound and the rendering of the waveform signal were performed without any particular distinction between the fixed playback device 1 and the portable playback terminal 2. However, different processing may be performed for the fixed playback device 1 and the portable playback terminal 2.
[0070] Generally, the sound quality of the speakers in the fixed playback device 1 is higher than that of the speakers in the portable playback terminals 2. For example, some of the portable playback terminals 2 located within the live venue may be selected as playback terminals that output sound, while all of the fixed playback devices 1 installed within the live venue may be selected as playback terminals that output sound.
[0071] In this case, the rendering unit 52 renders the waveform signals of the channels corresponding to each fixed playback device 1 based on a distance-based rendering algorithm, for example, DBAP (Distance-based amplitude panning).
[0072] Figure 8 illustrates a distance-based rendering algorithm.
[0073] For example, when the sound from one sound source j is output from the speakers of the fixed playback devices 1FL, 1FR, 1RL, and 1RR, the rendering unit 52, as shown in Figure 8A, will render the sound source j (localization position P Sj ) and the distance d between the fixed regeneration device 1FR 1 The sound pressure gain corresponding to the distance d between the sound source j and the fixed playback device 1FL is multiplied by the waveform signal representing the sound from the sound source j to render the waveform signal of the channel corresponding to the fixed playback device 1FL. The rendering unit 52 calculates the distance d between the sound source j and the fixed playback device 1FL. 2 The sound pressure gain corresponding to the distance d between the sound source j and the fixed playback device 1RL is multiplied by the waveform signal representing the sound from the sound source j and the waveform signal of the channel corresponding to the fixed playback device 1FL is rendered. The rendering unit 52 calculates the distance d between the sound source j and the fixed playback device 1RL. 3 The sound pressure gain corresponding to the distance d between the sound source j is multiplied by the waveform signal representing the sound from the sound source j and the waveform signal of the channel corresponding to the fixed playback device 1RL is rendered. The rendering unit 52 calculates the distance d between the sound source j and the fixed playback device 1RL. 4The corresponding sound pressure gain is multiplied by the waveform signal representing the sound from sound source j to render the waveform signal for the channel corresponding to the fixed playback device 1RR.
[0074] By combining the addition of a sense of sound localization by the fixed playback device 1 and the addition of a sense of sound localization by the portable playback terminal 2, it is expected that effective sound effects can be achieved.
[0075] As shown in Figure 8B, the sound pressure gain corresponding to the distance between the outer edge of the sound image range A11 of the sound source j and each fixed playback device 1 may be multiplied by the waveform signal representing the sound of the sound source j, so that the waveform signal of the channel corresponding to each fixed playback device 1 is rendered.
[0076] Next, we will describe the process performed by the fixed playback device 1, the portable playback terminal 2, and the control device 12 (each with its own synchronization control unit) to synchronize the sound output from the speakers of each playback device. Content provision is achieved, for example, by the fixed playback device 1 and the portable playback terminal 2 receiving a playback trigger transmitted from the control device 12 via a predetermined communication path, and by synchronizing the output of sound, indicated by a waveform signal, according to the playback trigger.
[0077] Here, a communication delay occurs between the time the playback trigger is sent and received, which varies depending on the communication path and other factors. Additionally, a playback delay occurs between the time the playback device starts processing the waveform signal and the time the sound represented by the waveform signal is output from the speaker, which varies depending on the performance of the playback device and how the speaker is connected to the playback device.
[0078] Because communication paths, playback device performance, and speaker connection methods differ for each playback device, the communication delay and playback delay that occur when providing content also differ between the fixed playback device 1 and the portable playback terminal 2, and also differ for each portable playback terminal 2. Therefore, the content playback system of this technology synchronizes each playback device so that the audio from the same sound source is output from the speakers of each playback device approximately simultaneously by compensating for both communication delay and playback delay for each playback device.
[0079] Figure 9 is a diagram illustrating the flow of compensation for communication delays.
[0080] As described above, the communication delay is the time required to send and receive playback triggers between the control device 12 and each playback device. In the content playback system, the round-trip time (RTT), which is the communication round-trip time between the playback device and the control device 12, is measured for each playback device, and compensation for the communication delay is performed based on the RTT of each playback device. Here, we will explain how to compensate for the sound delay between the portable playback terminals 2 that occurs due to the difference in communication delay.
[0081] To compensate for communication delays, first, as shown in the upper part of Figure 9, the portable playback terminals 2A to 2C synchronize their system time with the NTP time provided by the NTP server 11. Since all portable playback terminals 2 operate in synchronization with the NTP time, the time difference between the portable playback terminals 2 is minimized.
[0082] Next, as shown at the tip of the white arrow #11 in Figure 9, for example, portable playback terminals 2A to 2C measure the RTT of their own devices. In the example in Figure 9, the RTT of portable playback terminal 2A is 120 ms, the RTT of portable playback terminal 2B is 10 ms, and the RTT of portable playback terminal 2C is 70 ms. Half of the RTT corresponds to the communication delay for each portable playback terminal 2. Portable playback terminals 2A to 2C notify the control device 12 of the measured RTT.
[0083] Next, the control device 12 calculates the waiting time for each portable playback terminal 2, based on the round-time time (RTT) measured for each terminal 2, from the time it receives a playback trigger until it starts outputting sound. This waiting time can be called the waiting time for communication delay compensation. The control device 12 notifies each portable playback terminal 2 of the waiting time for communication delay compensation. In the example in Figure 9, the waiting time for portable playback terminal 2A is set to 0 ms, the waiting time for portable playback terminal 2B is set to 55 ms, and the waiting time for portable playback terminal 2C is set to 25 ms.
[0084] Next, as shown at the tip of the white arrow #12 in Figure 9, the control device 12 sets the time to start outputting sound (output start time) to a time later (future) than the current time by the largest communication delay among the communication delays of each portable playback terminal 2, and sends a playback trigger indicating the output start time to each portable playback terminal 2. In the example in Figure 9, since the communication delay of portable playback terminal 2A is 60 ms, the output start time is set to a time 60 ms later than the current time.
[0085] Each portable playback terminal 2 waits for a specified waiting time after receiving a playback trigger transmitted from the control device 12 before starting to output sound. In the example in Figure 9, portable playback terminal 2A starts outputting sound immediately after receiving the playback trigger. Portable playback terminal 2B starts outputting sound 55 ms after receiving the playback trigger, and portable playback terminal 2C starts outputting sound 25 ms after receiving the playback trigger.
[0086] As described above, the audio delay between the portable playback devices 2 caused by differences in communication delay is compensated. Furthermore, by replacing at least one of the portable playback devices 2 described above with a fixed playback device 1, the communication delay between the fixed playback device 1 and the playback device including the portable playback devices 2 can be compensated.
[0087] Figures 10 and 11 illustrate the flow of compensation for playback delay.
[0088] As mentioned above, playback delay is the time from when the waveform signal playback processing begins until the sound is actually output, and it varies depending on the performance of the playback device and the way the speakers are connected.
[0089] In this technology's content playback system, a reader terminal and follower terminals are set up, and a relative playback delay (relative delay) is measured for each follower terminal based on the playback delay of the reader terminal. For example, mobile playback terminal 2A is set as the reader terminal, and mobile playback terminal 2B is set as the follower terminal. Hereafter, mobile playback terminal 2A will also be referred to as reader terminal 2A, and mobile playback terminal 2B will also be referred to as follower terminal 2B. Based on the relative playback delay measured for each follower terminal, compensation for the playback delay of each mobile playback terminal 2 is performed. Here, we will explain the compensation for the sound delay between mobile playback terminals 2 that occurs due to the difference in playback delay.
[0090] To compensate for playback delay, first, as shown in the upper part of Figure 10, the leader terminal 2A and the follower terminal 2B synchronize their respective system times with the NTP time provided by the NTP server 11. Here, the follower terminal 2B is located near the leader terminal 2A.
[0091] Next, as shown at the tip of the white arrow #21 in Figure 10, the playback delay of the reader terminal 2A is measured.
[0092] Specifically, the control device 12 notifies the leader terminal 2A and the follower terminal 2B of the playback start time T1, which is the time when the leader terminal 2A starts processing the playback of the test sound (a waveform signal representing the test sound). In order to allow for communication delay, the time required for the follower terminal 2B to start recording, and the time required for the leader terminal 2A to output the test sound, the playback start time T1 is set to a time several hundred milliseconds to several seconds after the current time.
[0093] The leader terminal 2A starts the playback process of the test sound at the playback start time T1 and outputs the test sound from the speaker mounted on the leader terminal 2A. The follower terminal 2B picks up the test sound with the microphone mounted on the follower terminal 2B and records it. During recording, the follower terminal 2B embeds a pattern (flag information) indicating the playback start time T1 into the recording data.
[0094] As shown in the speech bubble B1 of Figure 10, the follower terminal 2B calculates the difference time D1 from the playback start time T1 to the time when the test sound was picked up by the microphone by analyzing the waveform of the recorded data. The difference time D1 corresponds to the playback delay of the leader terminal 2A.
[0095] Next, as shown in the upper part of Figure 11, the playback delay of the follower terminal 2B is measured.
[0096] Specifically, at an arbitrary playback start time T2, the follower terminal 2B starts processing the playback of a test sound (a waveform signal representing it) and outputs the test sound from the speaker mounted on the follower terminal 2B. The follower terminal 2B captures and records the test sound using the microphone mounted on the follower terminal 2B. During recording, the follower terminal 2B embeds a pattern (flag information) indicating the playback start time T2 into the recorded data.
[0097] As shown in the speech bubble B2 in Figure 11, the follower terminal 2B analyzes the waveform of the recorded data to calculate the difference time D2 from the playback start time T2 to the time when the test sound was picked up by the microphone. The difference time D2 corresponds to the playback delay of the follower terminal 2B. The follower terminal 2B notifies the control device 12 of the difference times D1 and D2.
[0098] Next, the control device 12 calculates the difference between the differential time D1 (playback delay of the leader terminal 2A) and the differential time D2 (playback delay of the follower terminal 2B) as the relative delay of the follower terminal 2B based on the playback delay of the leader terminal 2A. Based on the relative delay of the follower terminal 2B, the control device 12 calculates the waiting time for each mobile playback terminal 2 from the reference time until the waveform signal playback processing actually starts. This waiting time can be called the waiting time for playback delay compensation. The control device 12 notifies each mobile playback terminal 2 of the waiting time for playback delay compensation. In the example in Figure 11, the waiting time for the leader terminal 2A is set to Ws, and the waiting time for the follower terminal 2B is set to 0s.
[0099] Next, as shown at the tip of the white arrow #22 in Figure 11, the control device 12 transmits a playback trigger indicating a reference time to each mobile playback terminal 2. Each mobile playback terminal 2 receives the playback trigger transmitted from the control device 12, waits for a waiting time from the reference time indicated by the playback trigger, and then starts the content playback process. In the example in Figure 11, the reader terminal 2A starts the waveform signal playback process at a time Ws after the reference time and outputs sound from its speaker. The follower terminal 2B starts the waveform signal playback process at the reference time and outputs sound from its speaker.
[0100] As described above, the start time for the portable playback terminal 2A (first mobile playback device) to begin processing the waveform signal is controlled based on the measurement result of the delay from when the portable playback terminal 2A (first mobile playback device) starts processing the waveform signal until the sound is output from its speaker. Similarly, the start time for the portable playback terminal 2B (second mobile playback device) to begin processing the waveform signal is controlled based on the measurement result of the delay from when the portable playback terminal 2B (second mobile playback device) starts processing the waveform signal until the sound is output from its speaker. This compensates for the sound delay between the two portable playback terminals caused by the difference in playback delay. Furthermore, by considering the fixed playback device 1 as the reader terminal 2A described above, playback delays between the fixed playback device 1 and the playback devices including the portable playback terminals 2 can be compensated.
[0101] Specifically, the portable playback terminal 2 records a test sound output from the speaker of the fixed playback device 1 before, for example, a live music performance begins, and measures the delay that occurs when the speaker of the fixed playback device 1 outputs sound as the playback delay of the reader terminal. Next, the portable playback terminal 2 measures the playback delay of its own device.
[0102] After the concert begins, the portable playback device 2 plays the performers' voices, music, etc., while compensating for the playback delay based on the delay that occurs when the speakers of the fixed playback device 1 output sound and the playback delay of its own device.
[0103] As described above, by compensating for communication delay and playback delay, this technology's content playback system makes it possible to synchronize the sound output from the speaker of the fixed playback device 1 with the sound output from the speaker of the portable playback terminal 2 without any delay, and also to synchronize the sound output from the speakers of each portable playback terminal 2 without any delay. By having each playback device work together without any delay, sound image localization can be achieved with higher precision.
[0104] <3. Modified Examples> Figure 12, an example of rendering waveform signals in each playback device, is a block diagram showing the first modified configurations of the control device 12, the fixed playback device 1, and the portable playback terminal 2. In Figure 12, components identical to those in Figure 4 are denoted by the same reference numerals. Repetitive explanations are omitted as appropriate.
[0105] The control device 12 in Figure 12 differs from the control device 12 in Figure 4 in that it does not have a speaker selection unit 63 and a rendering unit 52. Also, the fixed playback device 1 in Figure 12 differs from the fixed playback device 1 in Figure 4 in that it has a speaker selection unit 151 and a rendering unit 152. Furthermore, the portable playback terminal 2 in Figure 12 differs from the portable playback terminal 2 in Figure 4 in that it has a speaker selection unit 171 and a rendering unit 172.
[0106] The playback control unit 62 of the control device 12 transmits content specification information and metadata, which specify the content to be provided to the audience in the content playback space, to the fixed playback device 1 and the portable playback terminal 2 via the communication unit 54. The metadata includes information indicating the localization position of the sound image of each sound source of the content specified in the content specification information, and information indicating the sound image range of each sound source.
[0107] The speaker selection unit 151 of the fixed playback device 1 is implemented, for example, by the control unit 71 executing a predetermined program. The speaker selection unit 151 receives content specification information and metadata transmitted from the control device 12 via the communication unit 74, and selects whether or not to output sound with the playback unit 73 (whether or not to reproduce the waveform signal) based on the localization position of the sound image of each sound source of the content provided to the audience in the content playback space, as well as the position and orientation of the device itself. The position and orientation of the fixed playback devices 1 are notified to each fixed playback device 1 from the control device 12, for example, before the start of a live music performance.
[0108] The rendering unit 152 renders waveform signals corresponding to the channels of its own device based on the localization position of the sound images of each sound source of the content provided to the audience in the content playback space, as well as the position and orientation of its own device.
[0109] If the speaker selection unit 151 selects that the playback unit 73 output sound, the signal processing unit 72 performs playback processing of the waveform signal rendered by the rendering unit 152, and outputs the sound indicated by the waveform signal from the playback unit 73.
[0110] The speaker selection unit 171 of the portable playback terminal 2 is implemented, for example, by the control unit 91 executing a predetermined program. The speaker selection unit 171 receives content specification information and metadata transmitted from the control device 12 via the communication unit 95, and selects whether or not to output sound with the playback unit 93 (whether or not to reproduce the waveform signal) based on the localization position of the sound image of each sound source of the content provided to the audience in the content playback space, as well as the position and orientation of the device itself.
[0111] The rendering unit 172 renders waveform signals corresponding to the channels of its own device based on the localization position of the sound image of each sound source of the content provided to the audience in the content playback space, as well as the position and orientation of its own device estimated by the position acquisition unit 102.
[0112] If the speaker selection unit 171 selects that the playback unit 93 output sound, the signal processing unit 92 performs playback processing of the waveform signal rendered by the rendering unit 172, and outputs the sound indicated by the waveform signal from the playback unit 93.
[0113] As described above, by rendering the waveform signal in each playback device, it becomes unnecessary for the portable playback terminal 2 to transmit position and orientation information to the control device 12, and for the control device 12 to stream the waveform signal to the playback device. Therefore, it is expected that the latency of the sound output will be reduced, and the communication load between the playback device and the control device 12 will be reduced.
[0114] • Variations of the sound image range of a sound source The above describes an example where the shape of the sound image range, including the localization position of the sound source, is circular, but the shape and size of the sound image range can be any shape and size.
[0115] Figure 13 shows an example of the sound image range of a certain sound source.
[0116] As shown in Figure 13, the sound image range A12 of a sound source may be set to a vertically elongated rectangular area. As indicated by the white arrow in Figure 13, by moving the sound image range A12 across the entire audience seating area of a live venue, it is possible to create an acoustic effect in which a linear sound image traverses the entire audience seating area.
[0117] When moving the sound image range, the speaker selection unit 63 can select a playback device that outputs sound from among multiple playback devices located within the live venue, using a collision detection algorithm commonly used in game engines and the like. For example, the speaker selection unit 63 selects a playback device that outputs sound based on the result of collision detection between the polygon corresponding to the sound image range and the polygon corresponding to each playback device.
[0118] Furthermore, it is desirable that an appropriate algorithm be set as the waveform signal rendering algorithm depending on the shape of the sound image range.
[0119] Figure 14 shows another example of the sound image range of a given sound source.
[0120] For example, the playback control unit 62 can change the square-shaped sound image range A13, which is set approximately in the center of the live venue as shown on the left side of Figure 14, into a roughly hexagonal shape during a live music performance, as shown on the right side of Figure 14.
[0121] Thus, during a live music performance, the playback control unit 62 may change the sound image range of each sound source in the content to any shape or change the size of the sound image range.
[0122] - In cases where portable playback devices 2 are not present at high density: In the content playback space, if portable playback devices 2 are not present at high density, in other words, if listeners of the content are scattered, the number of playback devices included in the sound image range of each sound source may be small, or there may be no playback devices within the sound image range.
[0123] For example, in content playback spaces such as theme parks and event venues, as shown in Figure 15, the speakers of the fixed playback device 1 are easily installed in a scattered manner, and the listeners of the content are also easily scattered.
[0124] If playback devices are not densely located in the content playback space, the speaker selection unit 63 selects playback devices that output sound using an algorithm that selects a predetermined number (for example, n) of playback devices, rather than an algorithm that selects playback devices located within the sound image range of each sound source.
[0125] Figure 16 illustrates an algorithm for selecting a predetermined number of playback devices.
[0126] For example, in an algorithm for selecting three playback devices, as shown in Figure 16A, the sound image localization position P of the sound source j is selected from among nine playback devices 201-1 to 201-9. Sj The top three playback devices 201-1 to 201-3 with the smallest distance from the sound source j are selected as playback devices that output the sound from the sound source j. In this case, the distance from the sound source j (d 1 from d 3The sound pressure gain corresponding to the sound source j is multiplied by the waveform signal representing the sound, and the waveform signals for the channels corresponding to playback devices 201-1 to 201-3 are rendered. In other words, a distance-based panning algorithm is used to render the waveform signals.
[0127] For example, in an algorithm for selecting five playback devices, as shown in Figure 16B, the sound image localization position P of the sound source j is selected from among nine playback devices 201-1 to 201-9. Sj The top five playback devices 201-1 to 201-5 with the smallest distance from the sound source j are selected as playback devices that output the sound from the sound source j. In this case, the distance (d) from the sound image of the sound source j is selected. 1 from d 5 The sound pressure gain corresponding to the sound source j is multiplied by the waveform signal representing the sound, and the waveform signals for the channels corresponding to playback devices 201-1 to 201-5 are rendered.
[0128] By utilizing such a selection algorithm, even in content playback spaces where playback devices are not densely distributed, it is possible to provide a listener with a sense of localization of the sound image of the content's sound source by using a portable playback device 2 carried by a specific listener, a portable playback device 2 carried by a companion of the listener, or a portable playback device 2 carried by another person around the listener.
[0129] Furthermore, a clustering algorithm such as the k-nearest neighbor method may be applied to the algorithm for selecting the playback device that outputs sound. A threshold distance from the sound source may be set to prevent playback devices that are extremely far from the sound source from being selected as playback devices that output the sound from the sound source. In addition, among the multiple areas into which the content playback space is divided, a playback device that outputs the sound from the sound source may be selected from among the playback devices in the area that includes the localization position of the sound image of the sound source.
[0130] Regarding the rendering algorithm, in the content playback system of this technology, any algorithm can be used as the waveform signal rendering algorithm. For example, a vector-based algorithm such as VBAP (Vector Base Amplitude Panning) may be applied as the waveform signal rendering algorithm. In VBAP, for example, the content listening position is assumed to be the center of the content playback space, and the waveform signal is rendered accordingly.
[0131] The rendering unit 52 may apply various effects processing to the waveform signal to make the sense of distance and localization of the sound image easier to understand.
[0132] - Example of equalizing the speaker volume of portable playback device 2: In actual use cases, even if the same waveform signal is processed for playback on multiple portable playback devices 2, it is expected that differences in the output volume will occur between the portable playback devices 2 due to differences in speaker performance, volume settings, etc.
[0133] Therefore, each portable playback terminal 2 may issue a notification to the listener prompting them to set the speaker volume to a predetermined value. When a portable playback terminal 2 reads a two-dimensional code indicating the speaker volume setting value displayed in the content playback space, the speaker volume may be automatically set on each portable playback terminal 2. The sound pressure gain set based on a database in which the speaker performance and volume settings of each portable playback terminal 2 are registered may be multiplied by the waveform signal of the channel corresponding to each portable playback terminal 2.
[0134] Regarding the method for acquiring the position of the portable playback terminal 2, the position and orientation of the portable playback terminal 2 may be acquired by a method other than one based on camera images captured by the camera mounted on the portable playback terminal 2 (such as VPS or SLAM). For example, the position and orientation of the portable playback terminal 2 may be acquired by a method using GNSS (Global Navigation Satellite System) positioning signals, such as GPS (Global Positioning System) positioning, local area RTK (Real Time Kinematic) positioning, or PPP (Precise Point Positioning)-RTK positioning.
[0135] In addition, the position and orientation of the portable playback terminal 2 can be acquired by combining methods such as using beacons, transmitting an inaudible sound signal from the speaker of the fixed playback device 1 and receiving the signal with a microphone mounted on the portable playback terminal 2, and a method called PDR (Pedestrian Dead Reckoning).
[0136] In a music concert use case, tickets may assign seats within the concert venue to each listener. In this case, when the ticket is read by the portable playback device 2, the content playback system may register the location of the seat assigned to the listener carrying the portable playback device 2 as the location of the portable playback device 2.
[0137] - An example of providing an experience based on the listener's registration information: In the music live concert use case described above, the same experience was provided to all listeners in the concert venue. However, it is also possible to provide an experience tailored to each listener based on the information they have registered.
[0138] For example, listeners who purchase paid tickets or make donations can experience special effects. For instance, at an idol music concert, the sound image of the idol's voice will be localized near the listener. For instance, at a sports venue where a baseball game is being played, the sound image of the team's theme song or cheering sound effects will be localized near the listener.
[0139] In the content playback system, information such as whether or not a charge is made, the name of the idol the listener supports, and the sports team they support is registered as registration information for each listener. The content playback system renders the waveform signals of the channels corresponding to each playback device so that the sound image of the sound source corresponding to the registration information is localized to the localized position set based on the registration information. Here, the playback control unit 62 of the control device 12 functions as a setting unit that sets the localized position of the sound image of the content's sound source based on the registration information.
[0140] When a listener performs a predetermined action, such as holding a ticket over a dedicated device installed in the content playback space, an acoustic effect may be performed in which a sound image of a predetermined sound source is localized near the device. The dedicated device may include a 2D code reader or an RFID (Radio Frequency Identification) reader.
[0141] A person with special privileges, such as a paying listener, may operate their portable playback device 2 to specify the localization position of each sound image of the content provided in the content playback space. In other words, a portable playback device 2 carried by a person with special privileges can accept the person's specification of the localization position of each sound image.
[0142] For example, as shown on the left side of Figure 17, the display of the listener's portable playback device 2 shows a UI screen that shows at least a portion of the seating area of the sports venue where the listener is located. When the listener slides their finger on the display of the portable playback device 2 to input the trajectory of the sound image's localization position, sound effects are created such that the sound image of, for example, a cheering song moves within the actual seating area along the trajectory specified by the listener, as shown on the right side of Figure 17. The listener's specification of the sound image's localization position is reflected, for example, in real time at the actual sports venue.
[0143] The localization position of each sound source of content provided at limited times, such as during changes of possession, may be specified by listeners at the sports venue. It is also possible for a producer, such as a disc jockey (DJ), to operate the control device 12 or the portable playback terminal 2 in accordance with the situation of the live music performance to move the localization position of each sound source of content provided at the live venue in real time.
[0144] The localization position of each sound source in the game content provided in the content playback space may be specified by a person in a remote location. In this case, the person in the remote location can specify the localization position of each sound source in the game content while viewing the content playback space on a screen. For example, the person in the remote location can bring the footsteps of a zombie character closer to a specific player in the content playback space, thereby triggering an event where that player runs away from or fights the zombie character.
[0145] • Example of a portable playback terminal 2: While an example of a portable playback terminal 2 being composed of a device that the listener normally uses, such as a smartphone, has been described, the portable playback terminal 2 may also be composed of a device dedicated to the content playback space.
[0146] For example, by equipping a device such as a penlight used at a music concert with a speaker, the penlight can function as a portable playback terminal 2 for the content playback system of this technology. In this case, the light emitted by the light-emitting part such as an LED (Light Emitting Diode) mounted on the penlight may be synchronized with the sound output from the speaker, or the vibration by the vibration part such as an actuator mounted on the penlight may be synchronized with the sound output from the speaker. The penlight may also emit light in the color of the team that the listener carrying it is supporting.
[0147] Furthermore, by equipping props such as lanterns used in night walks with speakers, the props can function as portable playback terminals 2 for the content playback system of this technology.
[0148] Furthermore, by equipping props such as weapons used in immersive games with speakers, these props can function as portable playback terminals 2 for the content playback system of this technology. In this case, for example, by positioning the sound image of sound effects around the player carrying the prop, it becomes possible to provide the player with an immersive experience.
[0149] When using a dedicated device for content playback as the portable playback terminal 2, it becomes unnecessary to consider delays in sound output due to differences in performance between the portable playback terminals 2. Since the listener does not use their usual smartphone or other device, it is possible to prevent situations that would disrupt the experience, such as notification sounds playing during the content experience.
[0150] - Example of achieving sound localization using grouped portable playback devices 2: Instead of adding a sense of sound localization using all portable playback devices 2 present in the content playback space, the sense of sound localization may be added using grouped portable playback devices 2.
[0151] Figure 18 illustrates the addition of a sense of sound localization using grouped portable playback devices 2.
[0152] As shown in Figure 18, for example, suppose three listeners U51 to U53, who are acquaintances, enter a content playback space such as a theme park or event venue. Listeners U51 to U53 have previously grouped their portable playback devices 2. In this case, the content playback system groups each portable playback device 2 carried by listeners U51 to U53 into the same group.
[0153] In the content playback space shown in Figure 18, eight speakers of fixed playback devices 1 are installed in a scattered manner. When content is provided from listener U51 to U53, for example, the sound of a certain sound source is output from the speakers of each playback device so that the sound image Ph11 of that sound source is localized to a predetermined localization position.
[0154] In the example shown in Figure 18, from among the eight fixed playback devices 1-1 to 1-8, the top three fixed playback devices 1-1 to 1-3 with the smallest distance from the sound image Ph1 are selected as playback devices that output sound. Additionally, from among the portable playback terminals 2 present in the content playback space, portable playback terminals 2 grouped into the same group (portable playback terminals 2 carried by listeners U51 to U53) are selected as playback devices that output sound.
[0155] In this case, the distance (d) from the sound image Ph1 1 from d 3 The sound, multiplied by a sound pressure gain corresponding to the distance (d) from the fixed playback devices 1-1 to 1-3, is output from the speakers, and the distance (d) from the sound image Ph1 is output. 4 from d 6 The sound, multiplied by a sound pressure gain corresponding to the sound, is output from the speakers of each of the portable playback terminals 2 of listeners U51 to U53, thereby giving listeners U51 to U53 a sense of localization of the sound image Ph1.
[0156] In this way, by using only the portable playback terminal 2 carried by listeners U51 to U53 and the fixed playback device 1, it is possible to provide listeners U51 to U53 with a sense of localization of the sound image of the content's sound source.
[0157] Figure 19 is a block diagram showing a second modified configuration of the control device 12, the fixed playback device 1, and the portable playback terminal 2. In Figure 19, components identical to those in Figure 12 are denoted by the same reference numerals. Repetitive explanations are omitted as appropriate.
[0158] The control device 12 in Figure 19 differs from the control device 12 in Figure 12 in that it is provided with a group selection unit 301. The configuration of the fixed playback device 1 and portable playback terminal 2 in Figure 19 is the same as the configuration of the fixed playback device 1 and portable playback terminal 2 in Figure 12.
[0159] The group selection unit 301 of the control device 12 is implemented, for example, by the control unit 51 executing a predetermined program. The group selection unit 301 selects the group that outputs the sound from each sound source of the content.
[0160] The playback control unit 62 of the control device 12 transmits content specification information and metadata via the communication unit 54 to the fixed playback device 1 and the portable playback terminals 2, which are grouped into groups that output the sound of each sound source of the content.
[0161] Since the portable playback terminal 2, which has received the content specification information and metadata, will definitely play the sound of each audio source in the content, the speaker selection unit 171, which has received the content specification information and metadata, selects to output sound using the playback unit 93.
[0162] Next, referring to the sequence diagram in Figure 20, we will explain the processes performed by a content playback system having the configuration shown in Figure 19.
[0163] In step S101, the playback control unit 62 of the control device 12 determines the content to be provided to listeners carrying portable playback terminals 2 that are grouped into the same group in the content playback space.
[0164] In step S102, the playback control unit 62 determines the localization position of the sound image of each sound source of the content provided to the listener, and transmits content designation information and metadata to the portable playback terminal 2 and the fixed playback device 1. The portable playback terminal 2 and the fixed playback device 1 receive the content designation information and metadata transmitted from the control device 12.
[0165] In step S121, the position acquisition unit 102 of the portable playback terminal 2 acquires position and orientation information of its own device.
[0166] In step S122, the speaker selection unit 171 of the portable playback terminal 2 selects whether or not to output sound with the playback unit 93 based on the localization position of the sound image of each sound source of the content, as well as the position and orientation of the device itself.
[0167] In step S141, the speaker selection unit 151 of the fixed playback device 1 selects whether or not to output sound with the playback unit 73 based on the localization position of the sound image of each sound source of the content, as well as the position and orientation of the device itself.
[0168] In step S123, the rendering unit 172 of the portable playback terminal 2 switches (determines) the waveform signal rendering algorithm based on the localization position of each sound source in the content, as well as the position and orientation of the device itself. For example, if the distance between the center of gravity of the portable playback terminal 2 and the localization position of the sound source is less than or equal to a threshold, the rendering unit 172 uses a distance-based rendering algorithm; if the distance exceeds the threshold, it uses a vector-based rendering algorithm.
[0169] Figure 21 illustrates distance-based rendering algorithms and vector-based rendering algorithms.
[0170] If we attempt to localize the sound image Ph2 of a sound source at a localization position away from listeners U61 to U63 by outputting the sound of a certain sound source from the speakers of each portable playback terminal 2 carried by listeners U61 to U63, as shown in Figure 21A, the distance between the sound image Ph2 and each portable playback terminal 2 (d 1 from d 3 There is almost no difference in the output sound pressure between the two portable playback devices. Therefore, when waveform signals are rendered using a distance-based rendering algorithm, there is almost no difference in the sound pressure of the output sound between the two portable playback devices. Because there is almost no difference in the sound pressure of the output sound between the two portable playback devices, listeners U61 to U63 cannot obtain a sense of localization of the sound image Ph2 using a distance-based rendering algorithm.
[0171] Therefore, when the distance between listener U61 to U63 and the sound image Ph2 is large, the waveform signal is rendered using a vector-based rendering algorithm that does not depend on the distance between listener U61 to U63 and the sound image Ph2, as shown in Figure 21B. In the vector-based rendering algorithm, first, the circumcenter P1 of the triangle connecting the respective portable playback terminals 2 of listeners U61 to U63 is set as the position of the virtual listener. When the direction from the circumcenter P1 toward the sound image Ph2 is considered the front direction, the sound is output from the speaker of each portable playback terminal 2 multiplied by the sound pressure gain corresponding to the angular direction (θ1 to θ3) of each portable playback terminal 2 as seen from the virtual listener.
[0172] In the example shown in Figure 21, the angular direction θ of the listener U61's portable playback terminal 2. 1 and the angular direction θ of listener U63's portable playback terminal 2 3 Since the values are similar, sounds with similar sound pressure levels are output from the speakers of these portable playback devices 2. The angular direction θ of the listener U62's portable playback device 2 2 The angular direction θ of the other portable playback device 2 1 and θ 3 Because it is larger, the speaker of listener U62's portable playback device 2 outputs sound at a lower sound pressure than the speaker of other portable playback devices 2.
[0173] A vector-based rendering algorithm cannot convey a sense of distance to the sound image Ph2 from listener U61 to U63, but it can convey a sense of direction. To convey a sense of distance to the sound image Ph2 from listener U61 to U62, the waveform signal may be processed with effects. For example, the rendering unit 172 may increase the reverb applied to the sound source as the distance between the listener and the sound image increases, change the sound source itself according to the distance, or increase the number of sounds in the sound source according to the distance.
[0174] Returning to Figure 20, in step S124, the rendering unit 172 renders the waveform signal of the channel corresponding to its own device.
[0175] In step S142, the rendering unit 152 of the fixed playback device 1 renders the waveform signal of the channel corresponding to its own device.
[0176] In step S125, the signal processing unit 92 of the portable playback terminal 2 performs waveform signal playback processing, and the playback unit 93 of the portable playback terminal 2 outputs sound represented by the waveform signal.
[0177] In step S143, the signal processing unit 72 of the fixed playback device 1 performs waveform signal playback processing, and the playback unit 73 of the fixed playback device 1 outputs the sound indicated by the waveform signal.
[0178] Figures 22 and 23 illustrate use cases for a content playback system that achieves sound localization using grouped portable playback terminals 2.
[0179] As shown in Figure 22A, the content playback system of this technology can be used to guide listeners U51 to U53 in a specific direction by orienting the sound image of a guide voice that directs users to a theme park or event venue in the direction the listener is desired to move. By following the leading guide voice, listeners U51 to U53 can fully experience the theme park or event venue.
[0180] This technology's content playback system can be used to encourage listeners to notice recommended spots located in a different direction from their direction of travel, for example, by orienting the sound image of guide voices that direct listeners to recommended spots in their vicinity towards those spots.
[0181] As shown in Figure 22B, the content playback system of this technology can be used, for example, to guide listeners U51 to U53 in a direction other than the direction in which they are not wanted to move by orienting the sound image of the voice of an enemy character appearing in game content in a direction that the listener is not wanted to move.
[0182] As shown in Figure 23A, even if the speakers of the fixed playback device 1 are not installed around listeners U51 to U53, the content playback system of this technology can, for example, output guide voice from the speaker of the portable playback terminal 2 carried by listeners U51 to U53, thereby localizing the sound image of the guide voice in the direction to which the listener is to move.
[0183] As shown in Figure 23B, even if the speakers of the fixed playback device 1 are not installed around listeners U51 to U53, the content playback system of this technology can, for example, output the voice of an enemy character from the speaker of a portable playback terminal 2 carried by listeners U51 to U53, thereby localizing the sound image of the enemy character's voice in a direction that listeners do not want to move.
[0184] As shown in Figure 23C, the content playback system of this technology can localize the sound images of guide voices and enemy character voices within the area surrounded by listeners U51 to U53. Listeners U51 to U53 can have experiences such as walking with a virtual navigator or capturing enemy characters.
[0185] As shown in Figure 23, the content playback system of this technology may provide content using only the portable playback terminal 2. Even in spaces where it is difficult to install speakers, such as forests where night walks are held, World Heritage sites, and public squares, it is possible to create sound effects that localize the sound image of the content's sound source in various locations.
[0186] As described above, this technology's content playback system can, for example, represent the approach of an enemy character through sound localization, providing listeners with experiences such as searching for an enemy character, fighting an enemy character, running away from an enemy character, or chasing an enemy character. For example, it is assumed that listeners will pay attention to the direction of a rustling sound, so by localizing the sound image of such a sound to a certain location, an event of searching for something in that location can be triggered. For example, by localizing the sound image of a zombie enemy character's footsteps to a predetermined location, events such as running away from a zombie, fighting a zombie, or chasing a zombie can be triggered.
[0187] As described above, by equipping props such as lanterns or weapons with speakers, these props can function as portable playback terminals 2 of this technology. For example, if a player carrying a lantern points in a direction that you want to draw attention to, the lantern's light may become brighter or its color may change. For example, if a player carrying a prop points in a direction that you want to draw attention to, the prop may vibrate. These effects can direct the player's attention in the desired direction.
[0188] By combining this technology with AR (Augmented Reality) technology, it is possible to provide content that is easy to understand both visually and aurally. For example, a content playback system can display camera images and store information superimposed on the smartphone's display, while also orienting the sound image of a guide voice that directs users through the store in the direction the smartphone is pointed. For example, a content playback system can display camera images and a character superimposed on the smartphone's display, while also orienting the sound image of the character's voice in the direction the smartphone is pointed.
[0189] • Example of computer configuration: The series of processes described above can be executed by hardware or by software. When the series of processes are executed by software, the programs that make up the software are installed from a program storage medium onto a computer that is built into dedicated hardware, or a general-purpose personal computer.
[0190] Figure 24 is a block diagram showing an example of the hardware configuration of a computer that executes the series of processes described above using a program.
[0191] The processing circuit 501, ROM (Read Only Memory) 502, and RAM (Random Access Memory) 503 are interconnected by a bus 504.
[0192] An input / output interface 505 is further connected to the bus 504. An input unit 506 consisting of a keyboard, mouse, etc., and an output unit 507 consisting of a display, speakers, etc. are connected to the input / output interface 505. In addition, a storage unit 508 consisting of a hard disk, non-volatile memory, etc., a communication unit 509 consisting of a network interface, etc., and a drive 510 that drives removable media 511 are connected to the input / output interface 505.
[0193] In a computer configured as described above, the processing circuit 501 loads, for example, a program stored in the memory unit 508 into the RAM 503 via the input / output interface 505 and the bus 504, and executes it, thereby performing the series of processes described above.
[0194] The program executed by the processing circuit 501 is recorded on, for example, removable media 511, or provided via a wired or wireless transmission medium such as a local area network, the internet, or digital broadcasting, and installed in the storage unit 508.
[0195] The programs executed by the computer may be programs that are processed chronologically in the order described herein, or they may be programs that are processed in parallel or at necessary times, such as when a call is made.
[0196] Note that the processing circuit 501 is an example of an integrated circuit, and CPU (Central Processing Unit), MPU (Micro Processing Unit), GPU (Graphical Processing Unit), APU (Accelerated Processing Unit), ASIC (Application Specific Integrated Circuit), and FPGA (Field Programmable Gate Array) can all be considered integrated circuits.
[0197] In this specification, a system means a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device containing multiple modules in one enclosure, are both considered systems.
[0198] The effects described herein are illustrative and not limited to those described herein, and other effects may also occur.
[0199] The embodiments of this technology are not limited to those described above, and various modifications are possible without departing from the spirit of this technology.
[0200] For example, this technology can be configured as cloud computing, where a single function is shared and processed collaboratively by multiple devices via a network.
[0201] Furthermore, each step described in the flowchart above can be performed by a single device, or it can be divided and performed by multiple devices.
[0202] Furthermore, if a single step includes multiple processes, those processes can be executed by a single device or shared among multiple devices.
[0203] <Examples of configuration combinations> This technology can also be configured as follows:
[0204] (1) A mobile playback device that can move while playing back the waveform signal of audio content in synchronization with a fixed playback device fixed in a content playback space which is a space in which audio content is provided, the mobile playback device comprising: a position acquisition unit that acquires position information of the device itself in the content playback space; and a playback unit that plays back the waveform signal of the audio content rendered based on information regarding the sound image localization of the audio content and the position information of the device itself. (2) The mobile playback device according to (1), further comprising a communication unit that transmits the position information of the device itself to an external information processing device that renders the waveform signal of the audio content based on information regarding the sound image localization of the audio content and the position information of the device itself, and receives the waveform signal of the audio content transmitted from the information processing device, wherein the playback unit plays back the waveform signal of the audio content received by the communication unit. (3) The mobile playback device according to (1), further comprising a rendering unit that renders the waveform signal of the audio content based on information regarding the sound image localization of the audio content and the position information of the device itself. (4) The mobile playback device according to (3), wherein the rendering unit renders the waveform signal of the audio content based on a rendering algorithm that corresponds to the distance between the localization position of the sound image of the audio content and the position of the device itself. (5) The mobile playback device according to (3) or (4), further comprising a selection unit that selects whether or not to play the waveform signal of the audio content in the playback unit based on information regarding the localization of the sound image of the audio content and the position information of the device itself. (6) The mobile playback device according to any one of (1) to (5), wherein the position acquisition unit acquires the position information of the device itself based on a camera image taken by a camera mounted on the device itself. (7) The mobile playback device according to any one of (1) to (5), wherein the position acquisition unit acquires the position information of the device itself based on a seat assigned to a listener of the audio content who is carrying the device itself.(8) A mobile playback device according to any one of (1) to (7), further comprising an input unit for receiving a specification of the localization position of the sound image of the audio content by a listener of the audio content. (9) A mobile playback device according to any one of (1) to (8), further comprising a light-emitting unit or a vibrating unit. (10) A fixed playback device that is fixed in a content playback space, which is a space in which audio content is provided, and plays a second waveform signal of the audio content in synchronization with a mobile playback device that is movable while playing a first waveform signal of the audio content, the fixed playback device comprising a playback unit that plays the second waveform signal rendered based on information relating to the sound image localization of the audio content and position information of the device itself in the content playback space. (11) A fixed playback device according to (10), further comprising a communication unit that receives the second waveform signal transmitted from an external information processing device that renders the waveform signal of the audio content based on information relating to the sound image localization of the audio content and position information of the device itself, wherein the playback unit plays the second waveform signal received by the communication unit. (12) The fixed playback device according to (10), further comprising a rendering unit that renders the second waveform signal based on information relating to the sound image localization of the audio content and the position information of the device itself. (13) The fixed playback device according to (12), further comprising a selection unit that selects whether or not to play the waveform signal of the audio content in the playback unit based on information relating to the sound image localization of the audio content and the position information of the device itself.(14) An information processing device connected to a mobile playback device that can move within a content playback space, which is a space in which audio content is provided, while playing a first waveform signal of the audio content, and a fixed playback device fixed within the content playback space and playing a second waveform signal of the audio content in synchronization with the mobile playback device, the information processing device comprising: a rendering unit that renders the first waveform signal and the second waveform signal based on information relating to the sound image localization of the audio content, the position information of the mobile playback device, and the position information of the fixed playback device; and a communication unit that transmits the first waveform signal to the mobile playback device and transmits the second waveform signal to the fixed playback device. (15) The information processing device according to (14), further comprising a selection unit that selects a playback device to play the waveform signal of the audio content from among a plurality of playback devices including at least one mobile playback device and at least one fixed playback device, based on information relating to the sound image localization of the audio content, the position information of the mobile playback device, and the position information of the fixed playback device. (16) The information processing device according to (15), wherein the selection unit selects from among a plurality of playback devices a playback device located within a predetermined sound image range including the localization position of the sound image of the audio content as a playback device for playing the waveform signal of the audio content. (17) The information processing device according to (15), wherein the selection unit selects from among a plurality of playback devices a predetermined number of playback devices as playback devices for playing the waveform signal of the audio content, based on the distance between the localization position of the sound image of the audio content and the position of each playback device. (18) The information processing device according to any one of (15) to (17), wherein the selection unit selects from among a plurality of playback devices at least a plurality of mobile playback devices grouped into the same group as playback devices for playing the waveform signal of the audio content. (19) The information processing device according to any one of (14) to (18), further comprising a setting unit for setting information relating to the sound image localization of the audio content based on registration information which is information registered about the listener of the audio content.(20) A mobile playback device that can move within a content playback space, which is a space in which audio content is provided, while playing a first waveform signal of the audio content; a fixed playback device fixed within the content playback space and playing a second waveform signal of the audio content in synchronization with the mobile playback device; and an information processing device, wherein the information processing device has a rendering unit that renders the first waveform signal and the second waveform signal based on information relating to the sound image localization of the audio content, the position information of the mobile playback device, and the position information of the fixed playback device; a first communication unit that receives the position information of the mobile playback device transmitted from the mobile playback device, transmits the first waveform signal to the mobile playback device, and transmits the second waveform signal to the fixed playback device; the mobile playback device has a position acquisition unit that acquires its own position information within the content playback space; a second communication unit that transmits its own position information to the information processing device and receives the first waveform signal transmitted from the information processing device; and a first playback unit that plays the first waveform signal received by the second communication unit. The fixed regeneration device is an acoustic processing system comprising a third communication unit that receives the second waveform signal transmitted from the information processing device, and a second regeneration unit that regenerates the second waveform signal received by the third communication unit.
[0205] 1 Fixed playback device, 2 Portable playback terminal, 11 NTP server, 12 Control device, 51 Control unit, 52 Rendering unit, 53 Input unit, 54 Communication unit, 61 Synchronization control unit, 62 Playback control unit, 63 Speaker selection unit, 71 Control unit, 72 Signal processing unit, 73 Playback unit, 74 Communication unit, 81 Synchronization control unit, 91 Control unit, 92 Signal processing unit, 93 Playback unit, 94 Input unit, 95 Communication unit, 101 Synchronization control unit, 102 Position acquisition unit, 151 Speaker selection unit, 152 Rendering unit, 171 Speaker selection unit, 172 Rendering unit, 201-1 to 201-9 Playback device, 301 Group selection unit
Claims
1. A mobile playback device that can move while reproducing the waveform signal of audio content in synchronization with a fixed playback device fixed within a content playback space, which is a space in which audio content is provided, the mobile playback device comprising: a position acquisition unit that acquires position information of the device itself within the content playback space; and a playback unit that reproduces the waveform signal of the audio content rendered based on information regarding the sound image localization of the audio content and the position information of the device itself.
2. The mobile playback device according to claim 1, further comprising a communication unit that transmits the location information of the device to an external information processing device that renders a waveform signal of the audio content based on information regarding the sound image localization of the audio content and the location information of the device itself, and receives the waveform signal of the audio content transmitted from the information processing device, wherein the playback unit plays back the waveform signal of the audio content received by the communication unit.
3. The mobile playback device according to claim 1, further comprising a rendering unit that renders a waveform signal of the audio content based on information relating to the sound image localization of the audio content and the position information of the device itself.
4. The mobile playback device according to claim 3, wherein the rendering unit renders the waveform signal of the audio content based on a rendering algorithm that corresponds to the distance between the localization position of the sound image of the audio content and the position of the device itself.
5. The mobile playback device according to claim 3, further comprising a selection unit that selects whether or not to play back the waveform signal of the audio content in the playback unit based on information relating to the sound image localization of the audio content and the position information of the device itself.
6. The mobile playback device according to claim 1, wherein the position acquisition unit acquires position information of the device based on a camera image taken by a camera mounted on the device.
7. The mobile playback device according to claim 1, wherein the position acquisition unit acquires position information of the device based on the seat assigned to the listener of the audio content who is carrying the device.
8. The mobile playback device according to claim 1, further comprising an input unit for receiving a specification of the localization position of the sound image of the audio content by a listener of the audio content.
9. The mobile regeneration device according to claim 1, further comprising a light-emitting unit or a vibrating unit.
10. A fixed playback device that is fixed within a content playback space, which is a space in which audio content is provided, and plays a second waveform signal of the audio content in synchronization with a mobile playback device that is movable while playing a first waveform signal of the audio content, the fixed playback device comprising a playback unit that plays the second waveform signal rendered based on information relating to the sound image localization of the audio content and position information of the device itself within the content playback space.
11. The fixed playback device according to claim 10, further comprising a communication unit that receives the second waveform signal transmitted from an external information processing device that renders the waveform signal of the audio content based on information relating to the sound image localization of the audio content and position information of the device itself, wherein the playback unit plays back the second waveform signal received by the communication unit.
12. The fixed playback device according to claim 10, further comprising a rendering unit that renders the second waveform signal based on information relating to the sound image localization of the audio content and the position information of the device itself.
13. The fixed playback device according to claim 12, further comprising a selection unit that selects whether or not to play back the waveform signal of the audio content in the playback unit based on information relating to the sound image localization of the audio content and the position information of the device itself.
14. An information processing device connected to a mobile playback device that can move within a content playback space, which is a space in which audio content is provided, while playing a first waveform signal of the audio content, and a fixed playback device fixed within the content playback space and playing a second waveform signal of the audio content in synchronization with the mobile playback device, the information processing device comprising: a rendering unit that renders the first waveform signal and the second waveform signal based on information relating to the sound image localization of the audio content, position information of the mobile playback device, and position information of the fixed playback device; and a communication unit that transmits the first waveform signal to the mobile playback device and transmits the second waveform signal to the fixed playback device.
15. The information processing device according to claim 14, further comprising a selection unit that selects a playback device for playing back the waveform signal of the audio content from among a plurality of playback devices, including at least one mobile playback device and at least one fixed playback device, based on information regarding the sound image localization of the audio content, position information of the mobile playback device, and position information of the fixed playback device.
16. The information processing apparatus according to claim 15, wherein the selection unit selects from among a plurality of playback devices a playback device located within a predetermined sound image range including the localization position of the sound image of the audio content as a playback device for playing the waveform signal of the audio content.
17. The information processing apparatus according to claim 15, wherein the selection unit selects a predetermined number of playback devices from among a plurality of playback devices as playback devices for playing the waveform signal of the audio content, based on the distance between the localization position of the sound image of the audio content and the position of each playback device.
18. The information processing apparatus according to claim 15, wherein the selection unit selects at least a plurality of the mobile playback devices, grouped into the same group, from among a plurality of playback devices, as playback devices for playing the waveform signal of the audio content.
19. The information processing apparatus according to claim 14, further comprising a setting unit that sets information relating to the sound image localization of the audio content based on registration information, which is information registered about the listener of the audio content.
20. A mobile playback device that can move within a content playback space, which is a space in which audio content is provided, while playing a first waveform signal of the audio content; a fixed playback device fixed within the content playback space and playing a second waveform signal of the audio content in synchronization with the mobile playback device; and an information processing device, wherein the information processing device includes a rendering unit that renders the first waveform signal and the second waveform signal based on information relating to the sound image localization of the audio content, the position information of the mobile playback device, and the position information of the fixed playback device; a first communication unit that receives the position information of the mobile playback device transmitted from the mobile playback device, transmits the first waveform signal to the mobile playback device, and transmits the second waveform signal to the fixed playback device; the mobile playback device includes a position acquisition unit that acquires its own position information within the content playback space; a second communication unit that transmits its own position information to the information processing device and receives the first waveform signal transmitted from the information processing device; and a first playback unit that plays the first waveform signal received by the second communication unit; and the fixed playback device is An acoustic processing system comprising: a third communication unit that receives the second waveform signal transmitted from the information processing device; and a second playback unit that reproduces the second waveform signal received by the third communication unit.