Information processing device, information processing method, and program

The information processing device and method enhance the unity of video and audio by mapping audio data to display and speaker units, addressing the limited viewing positions of phantom sound imaging, achieving accurate sound image localization across a broader area.

JP7718422B2Active Publication Date: 2025-08-05SONY GROUP CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022551197
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-25
Filing Date
2021-08-19
Publication Date
2025-08-05
Estimated Expiration
2041-08-19

AI Technical Summary

Technical Problem

Existing phantom sound imaging methods limit the range of viewing positions where sound images can be accurately reproduced, making it difficult to achieve a unified sense of video and audio.

Method used

An information processing device and method that extracts audio data corresponding to different sound sources and maps them to combinable display and speaker units, utilizing a control system to render and synchronize audio with video content, including a sound source extraction unit, mapping processing unit, and a DNN engine for sound source position estimation.

Benefits of technology

Enhances the sense of unity between video and audio by accurately localizing sound images across a wider range of viewing positions, reducing discrepancies and improving sound image localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007718422000001
    Figure 0007718422000001
  • Figure 0007718422000002
    Figure 0007718422000002
  • Figure 0007718422000003
    Figure 0007718422000003
Patent Text Reader

Abstract

An information processing device (30) comprises a sound source extraction unit (341) and a mapping processing unit (343). The sound source extraction unit (341) extracts one or more pieces of audio data (AD), corresponding to different sound sources, from audio content (AC). The mapping processing unit (343) selects, for each piece of audio data (AD) and from among one or more combinable display units (12) having a sound-producing mechanism, one or more display units (12) that serve as a mapping destination of the audio data (AD)..
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] There are known techniques for linking a sound field with an image by using multiple speakers. For example, Patent Document 1 discloses a system that controls the position of a phantom sound image in conjunction with the position of a sound source displayed on a display. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2011-259298 Summary of the Invention [Problem to be solved by the invention]

[0004] With the phantom sound imaging method, the range of viewing positions in which the sound image can be accurately reproduced is narrow, making it difficult to achieve a sense of unity between the video and audio.

[0005] Therefore, the present disclosure proposes an information processing device, an information processing method, and a program that make it easier to achieve a sense of unity between video and audio. [Means for solving the problem]

[0006] According to the present disclosure, there is provided an information processing device having a sound source extraction unit that extracts one or more audio data corresponding to different sound sources from audio content, and a mapping processing unit that selects, for each piece of audio data, one or more display units to which the audio data is to be mapped from one or more combinable display units having sound generation mechanisms.The present disclosure also provides an information processing method in which information processing of the information processing device is executed by a computer, and a program that causes a computer to realize the information processing of the information processing device. [Brief explanation of the drawings]

[0007] [Figure 1] 1 is a diagram showing a schematic configuration of an audio / video content output system. [Figure 2] FIG. 1 is a diagram illustrating a configuration of a control system. [Figure 3] FIG. 2 is a diagram illustrating the configuration of an audio decoder. [Figure 4] FIG. 1 is a diagram showing a schematic configuration of a tiling display. [Figure 5] 3A and 3B are diagrams illustrating an example of the configuration and arrangement of a display unit. [Figure 6] FIG. 10 is an explanatory diagram of a tiling display and a speaker unit reproduction frequency. [Figure 7] FIG. 10 is a diagram showing the relationship between the reproduction frequency of the display unit and the magnitude of vibration during reproduction. [Figure 8] FIG. 10 is a diagram illustrating the logical numbers of the display units. [Figure 9] FIG. 10 is a diagram illustrating the logical numbers of the display units. [Figure 10] FIG. 10 is a diagram illustrating the logical numbers of the display units. [Figure 11] FIG. 2 is a diagram illustrating an example of a connection configuration between a cabinet and a control system. [Figure 12] FIG. 2 is a diagram illustrating an example of a connection configuration between a cabinet and a control system. [Figure 13] FIG. 2 is a diagram illustrating an example of a connection configuration between a cabinet and a display unit. [Figure 14] FIG. 1 is a diagram showing an example in which an audio / video content output system is applied to a theater. [Figure 15] FIG. 10 is a diagram illustrating an example of a mapping process for audio data of channel-based audio. [Figure 16] FIG. 10 is a diagram illustrating an example of a mapping process for audio data in object-based audio. [Figure 17]FIG. 10 is a diagram illustrating an example of a mapping process for audio data in object-based audio. [Figure 18] FIG. 10 is a diagram illustrating another example of mapping processing of audio data for channel-based audio. [Figure 19] 10A and 10B are diagrams illustrating a method for controlling a sound image in the depth direction. [Figure 20] 10A and 10B are diagrams illustrating a method for controlling a sound image in the depth direction. [Figure 21] 10A and 10B are diagrams illustrating a method for controlling a sound image in the depth direction. [Figure 22] 10A and 10B are diagrams illustrating a method for controlling a sound image in the depth direction. [Figure 23] 10A and 10B are diagrams illustrating another example of a sound image localization enhancement control technique. [Figure 24] FIG. 2 is a diagram illustrating an example of the arrangement of speaker units. [Figure 25] FIG. 10 is a diagram illustrating an example of a method for detecting the position of a display unit. [Figure 26] FIG. 10 is a diagram showing the arrangement of microphones used to detect the position of the display unit. [Figure 27] 10A and 10B are diagrams illustrating another example of a method for detecting the physical position of a display unit. [Figure 28] FIG. 10 is a diagram illustrating the control of the directionality of reproduced sound. [Figure 29] FIG. 10 is a diagram showing an example of allocating different playback sounds to different viewers. [Figure 30] 10 is a flowchart illustrating an example of an information processing method performed by the control system. [Figure 31] FIG. 1 is a diagram showing an example in which an audio / video content output system is applied to a theater. [Figure 32] FIG. 1 is a diagram showing an example in which an audio / video content output system is applied to a theater. [Figure 33] FIG. 2 is a diagram illustrating an example of the arrangement of speaker units. [Figure 34] FIG. 10 is a diagram showing another example of the arrangement of speaker units. [Figure 35] FIG. 1 is a diagram showing the arrangement of microphones used to measure spatial characteristics. [Figure 36] This is a diagram showing an example in which an audio / video content output system is applied to a telepresence system. [Figure 37] 10A and 10B are diagrams illustrating an example of a sound collection process and a reproduction process of an object sound. [Figure 38] FIG. 10 is a diagram showing an example in which an audio / video content output system is applied to a digital signage system. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.

[0009] The explanation will be given in the following order. [1. Overview of the Audio / Video Content Output System] [1-1. System configuration example] [1-2. Control system configuration] [1-3. Display unit configuration and layout] [1-4. Logical number of the display unit] [1-5. Connection between cabinet and control system] [1-6. Connection between the cabinet and the display unit] [2. First embodiment] [2-1. System Image] [2-2. Audio data mapping process for channel-based audio] [2-3. Audio data mapping process for object-based audio] [2-4. Sound source placement using DNN engine] [2-5. Controlling the sound image in the depth direction] [2-6. Sound image localization emphasis control] [2-6-1. Strengthening sound image localization by expanding the frequency band] [2-6-2. Enhancement of sound localization ability by precedence effect] [2-7. Speaker unit placement] [2-8. How to detect the position of the display unit] [2-9. Controlling the directionality of reproduced sound] [2-10. Information processing method] [2-11.Effects] 3. Second Embodiment [3-1. System Image] [3-2. Speaker unit placement] [3-3. Measurement of spatial characteristics and reverberation cancellation using built-in microphone] 4. Third Embodiment [4-1. System Image] [4-2. Object sound collection and playback] 5. Fourth Embodiment

[0010] [1. Overview of the Audio / Video Content Output System] [1-1. System configuration example] FIG. 1 is a diagram showing a schematic configuration of an audio / video content output system 1. As shown in FIG.

[0011] The audio / video content output system 1 is a system that plays back audio / video content from content data CDs and presents it to a viewer U. The audio / video content output system 1 includes a tiling display 10, a plurality of speaker units 20, and a control system 30.

[0012] The tiling display 10 has multiple display units 12 arranged in a tiled pattern. The tiling display 10 has a single large screen SCR formed by one or more display units 12 that can be combined in a matrix pattern. The display units 12 reproduce both video and audio. The tiling display 10 outputs audio related to the video from the display units 12 that display the video. In the following description, the vertical direction will be referred to as the height direction of the tiling display 10. The arrangement direction of the display units 12 that is perpendicular to the height direction will be referred to as the width direction of the tiling display 10. The direction perpendicular to the height and width directions will be referred to as the depth direction of the tiling display 10.

[0013] A plurality of speaker units 20 are arranged around the tiling display 10. In the example of FIG. 1, the plurality of speaker units 20 include a first array speaker 21, a second array speaker 22, and a subwoofer 23. The first array speaker 21 and the second array speaker 22 are line array speakers in which a plurality of speakers ASP (see FIG. 15) are arranged in a line. The first array speaker 21 is arranged along the top side of the tiling display 10. The second array speaker 22 is arranged along the bottom side of the tiling display 10. The plurality of speaker units 20, together with the tiling display 10, output sounds related to the displayed video.

[0014] The control system 30 is an information processing device that processes various information extracted from the content data CD. The control system 30 extracts one or more pieces of audio data AD (see FIG. 3) corresponding to different sound sources from the content data CD. The control system 30 acquires playback environment information 352 (see FIG. 3) related to the configuration of the multiple display units 12 and multiple speaker units 20 that form the playback environment. The control system 30 performs rendering based on the playback environment information 352 and maps each piece of audio data AD to the playback environment.

[0015] [1-2. Control system configuration] FIG. 2 is a diagram showing the configuration of the control system 30.

[0016] The control system 30 includes a demultiplexer 31, a video decoder 32, and an audio decoder 33. The demultiplexer 31 acquires content data CD from an external device. The content data CD includes information about video content VC and information about audio content AC. The demultiplexer 31 separates and generates the video content VC and the audio content AC from the content data CD.

[0017] The video decoder 32 generates a video output signal from the video content VC and outputs it to the plurality of display units 12 via a video output signal line VL. The audio decoder 33 extracts one or more pieces of audio data AD from the audio content AC. The audio decoder 33 maps each piece of audio data AD to the plurality of display units 12 and the plurality of speaker units 20. The audio decoder 33 outputs an audio output signal generated based on the mapping to the plurality of display units 12 and the plurality of speaker units 20 via an audio output signal line AL.

[0018] The control system 30 can handle audio content AC in various formats, such as channel-based audio, object-based audio, and scene-based audio. The control system 30 performs rendering processing on the audio content AC based on the playback environment information 352. As a result, the audio data AD is mapped to the multiple display units 12 and multiple speaker units 20 that form the playback environment.

[0019] For example, audio content AC for channel-based audio includes one or more pieces of audio data AD generated for each channel. The control system 30 selects a mapping destination for the audio data AD for the channels other than the subwoofer 23 from among the multiple display units 12 and multiple speakers ASP based on the channel arrangement.

[0020] The audio content AC of object-based audio includes one or more pieces of audio data generated for each object (material sound) and meta information. The meta information includes information such as the position OB, sound spread, and various effects for each object. The control system 30 selects a mapping destination for the audio data AD from among multiple display units 12 and multiple speakers ASP based on the object's position OB defined in the meta information. The control system 30 changes the display unit 12 to which the object's audio data AD is mapped in accordance with the movement of the object's position OB.

[0021] Scene-based audio is a method of recording and playing back physical information of the entire space surrounding a viewer U in a 360-degree celestial sphere. Audio content AC of scene-based audio includes four audio data AD corresponding to channels W (omnidirectional component), X (front-to-back spread component), Y (left-to-right spread component), and Z (up-to-down spread component). Based on the recorded physical information, the control system 30 selects a mapping destination for the audio data AD from among multiple display units 12 and multiple speakers ASP.

[0022] FIG. 3 is a diagram showing the configuration of the audio decoder 33.

[0023] The audio decoder 33 includes a calculation unit 34 and a storage unit 35. The calculation unit 34 includes a sound source extraction unit 341, a band division unit 342, a mapping processing unit 343, a position detection unit 344, and a sound source position estimation unit 345.

[0024] The sound source extraction unit 341 extracts one or more pieces of audio data AD from the audio content AC. For example, the audio data AD is generated for each sound source. For example, from the channel-based audio content AC, one or more pieces of audio data AD generated for each channel that serves as a sound source are extracted. From the object-based audio content AC, one or more pieces of audio data AD generated for each object that serves as a sound source are extracted.

[0025] The band dividing unit 342 divides the audio data AD into frequency bands. The band dividing process is performed, for example, after cutting out the deep bass components of the audio data AD. The band dividing unit 342 outputs one or more pieces of waveform data PAD obtained by dividing the audio data AD to the mapping processing unit 343. The band dividing process is performed on audio data AD that has frequency components other than deep bass. The audio data AD containing only deep bass is mapped from the sound source extraction unit 341 to the subwoofer 23 via the mapping processing unit 343.

[0026] The mapping processing unit 343 maps one or more pieces of waveform data PAD output from the band dividing unit 342 to the tiling display 10 (display unit 12) and the plurality of speaker units 20 according to the frequency band.

[0027] The mapping processing unit 343 selects, for each piece of audio data AD, one or more display units 12 or one or more speaker ASPs, or one or more display units 12 and one or more speaker ASPs, from the plurality of display units 12 and the plurality of speaker ASPs, to which the audio data AD is to be mapped.

[0028] For example, if the audio data AD is audio data for multi-channel speakers extracted from audio content AC of channel-based audio, the mapping processing unit 343 selects one or more display units 12 or one or more speaker ASPs, or one or more display units 12 and one or more speaker ASPs, as the mapping destination, depending on the arrangement of the multi-channel speakers.

[0029] If the audio data AD is audio data of an object extracted from audio content AC of object-based audio, the mapping processing unit 343 selects one or more display units 12 or one or more speaker ASPs, or one or more display units 12 and one or more speaker ASPs, corresponding to the position OB of the object extracted from the audio content AC as the mapping destination.

[0030] The position detection unit 344 detects the spatial arrangement of the multiple display units 12. The detection of the spatial arrangement is performed based on measurement data MD such as sound or video output from the display units 12. The position detection unit 344 assigns a logical number LN to each display unit 12 based on the detected spatial arrangement. The mapping processing unit 343 identifies the mapping destination based on the logical number LN.

[0031] The sound source position estimation unit 345 estimates the display position of the sound source of each piece of audio data AD. When audio data AD that does not have position information of the sound source is input, the sound source position estimation unit 345 is used to identify the position of the sound source within the video. The mapping processing unit 343 selects one or more display units 12 corresponding to the display position of the sound source as the mapping destination.

[0032] For example, the sound source position estimation unit 345 applies one or more pieces of audio data AD and video content AC extracted by the sound source extraction unit 341 to an analysis model 351. The analysis model 351 is a DNN (Deep Neural Network) engine that has learned the relationship between the audio data AD and the position of a sound source in a video by machine learning. Based on the analysis result by the analysis model 351, the sound source position estimation unit 345 estimates the position of the sound source within the screen SCR where the sound source is displayed.

[0033] The storage unit 35 stores, for example, a program 353 executed by the calculation unit 34, an analytical model 351, and playback environment information 352. The program 353 is a program that causes a computer to execute information processing performed by the control system 30. The calculation unit 34 performs various processes in accordance with the program 353 stored in the storage unit 35. The storage unit 35 may be used as a work area for temporarily storing processing results of the calculation unit 34. The storage unit 35 includes, for example, any non-transitory storage medium such as a semiconductor storage medium or a magnetic storage medium. The storage unit 35 includes, for example, an optical disk, a magneto-optical disk, or a flash memory. The program 353 is stored, for example, in a non-transitory storage medium readable by a computer.

[0034] The calculation unit 34 is, for example, a computer configured with a processor and a memory. The memory of the calculation unit 34 includes a RAM (Random Access Memory) and a ROM (Read Only Memory). The calculation unit 34 executes a program 353 to function as a sound source extraction unit 341, a band division unit 342, a mapping processing unit 343, a position detection unit 344, and a sound source position estimation unit 345.

[0035] [1-3. Display unit configuration and layout] FIG. 4 is a diagram showing a schematic configuration of the tiling display 10. As shown in FIG.

[0036] The tiling display 10 has a plurality of cabinets 11 combined in a tiled pattern. A plurality of display units 12 are attached to the cabinets 11 and arranged in a tiled pattern. There is no frame area around the periphery of the display units 12. The pixels of the plurality of display units 12 are arranged continuously across the boundaries of the display units 12 while maintaining the pixel pitch. This forms a tiling display 10 having a single screen SCR spanning the plurality of display units 12.

[0037] The number and arrangement of display units 12 attached to one cabinet 11 are arbitrary. The number and arrangement of cabinets 11 constituting the tiling display 10 are also arbitrary. For example, in the example of FIG. 4, a total of 32 cabinets are arranged two-dimensionally, with four columns in the height direction and eight columns in the width direction. A total of six display units 12 are attached to one cabinet 11, with two columns in the height direction and three columns in the width direction. Therefore, the tiling display 10 is composed of a total of 192 display units 12 arranged in eight columns in the height direction and 24 columns in the width direction.

[0038] FIG. 5 is a diagram showing an example of the configuration and arrangement of the display unit 12. As shown in FIG.

[0039] The display unit 12 has a display panel 121, an actuator 122, and a control circuit 123. The display panel 121 is a self-luminous thin display panel that does not have a backlight. In this embodiment, an LED panel is used as the display panel 121, in which three types of micro LEDs (Light Emitting Diodes), red, green, and blue, are arranged for each pixel. The actuator 122 vibrates the display panel 121 to output sound from the surface of the display panel 121. The control circuit 123 has a pixel drive circuit that drives the pixels and an actuator drive circuit that drives the actuator 122. The actuator 122 and the actuator drive circuit function as a sound generation mechanism for generating sound from the display unit 12.

[0040] The cabinet 11 has a housing 111, a connection board 112, and a cabinet board 113. The connection board 112 is a board that connects the control circuit 123 and the cabinet board 113. The connection board 112 is fixed to the housing 111. The display panel 121 is fixed to the connection board 112 by screws or the like. In this way, the display unit 12 is supported by the cabinet 11. The cabinet board 113 is connected to the control system 30. The control system 30 outputs a video output signal and an audio output signal to the control circuit 123 via the cabinet board 113.

[0041] Fig. 6 is an explanatory diagram of the reproduction frequency of the tiling display 10 and the speaker unit 20. Fig. 7 is a diagram showing the relationship between the reproduction frequency of the display unit 12 and the magnitude of vibration during reproduction.

[0042] Sounds related to the displayed images of the tiling display 10 are reproduced by the tiling display 10 (display unit 12) and multiple speaker units 20. As shown in FIG. 6, the reproduction frequency bands are divided into four: high frequency HF, mid frequency MF, low frequency LF, and very low frequency VLF (deep bass). The high frequency HF is a frequency band equal to or higher than a first frequency FH. The mid frequency MF is a frequency band equal to or higher than a second frequency FM and lower than the first frequency FH. The low frequency LF is a frequency band equal to or higher than a third frequency FL and lower than the second frequency FM. The very low frequency VLF is a frequency band lower than the third frequency FL. For example, the first frequency FH is 1 kHz. The second frequency FM is 500 Hz. The third frequency FL is 100 Hz.

[0043] The band dividing unit 342 divides the audio data AD into three waveform data PAD: high frequency HF, mid frequency MF, and low frequency LF. The waveform data for the very low frequency VLF is also divided by the band dividing unit 342. The mapping processing unit 343 maps the waveform data PAD for the high frequency HF, mid frequency MF, and low frequency LF to the display unit 12 or the speaker ASP.

[0044] The sound image localization ability that allows the position of a sound image to be sensed varies depending on the frequency of the sound. The higher the frequency of the sound, the higher the sound image localization ability. Therefore, the mapping processing unit 343 maps the waveform data PAD of the highest frequency range HF among the audio data AD to one or more display units 12 as the mapping destination. By outputting high-frequency range HF sound with high sound image localization ability from the display unit 12, it becomes less likely that a discrepancy will occur between the position of the sound source and the position of the sound image.

[0045] 7, as the reproduction frequency of the display unit 12 decreases, the amount of vibration of the display unit 12 increases. Therefore, when low-frequency sound is reproduced on the display unit 12, the viewer U may perceive shaking of the image due to the vibration. For this reason, the mapping processing unit 343 maps the waveform data PAD of the mid-range MF and low-range LF (mid-low range) to the first array speaker 21 and the second array speaker 22.

[0046] For example, the mapping processing unit 343 selects one or more speakers ASP that correspond to the positions of the sound sources of the audio data AD from among the multiple speakers ASP arranged around the tiling display 10. The mapping processing unit 343 maps the waveform data PAD of the low frequency LF, which has the lowest frequency, of the audio data AD, and the waveform data PAD of the mid frequency MF, which is between the high frequency HF and the low frequency LF, to the selected one or more speakers ASP.

[0047] The magnitude of vibration of the display unit 12 varies depending on the sound pressure (volume) of the sound being reproduced. The greater the sound pressure, the greater the vibration, and the lower the sound pressure, the smaller the vibration. Therefore, even if the sound pressure is low, the mapping processing unit 343 can map the mid-low frequency waveform data PAD to the display unit 12. For example, when the amplitude of the mid-frequency MF waveform data PAD, which has high sound image localization ability, among the mid-low frequency waveform data PAD, is equal to or less than a preset threshold, the mapping processing unit 343 maps the mid-frequency MF waveform data PAD to the display unit 12. This makes it possible to improve sound image localization ability while suppressing image shaking caused by vibration of the display unit 12.

[0048] Even when the sound pressure is high, increasing the number of display units 12 to be vibrated can reduce the magnitude of vibration of each display unit 12. Therefore, the mapping processing unit 343 makes the number of display units 12 onto which the mid-frequency MF waveform data PAD is mapped greater than the number of display units 12 onto which the high-frequency HF waveform data PAD is mapped. This configuration also makes it possible to improve sound image localization ability while suppressing image shaking caused by vibration of the display units 12.

[0049] [1-4. Logical number of the display unit] 8 to 10 are diagrams for explaining the logical numbers of the display unit 12. FIG.

[0050] As shown in FIG. 8, a logical number L1 based on the position of each cabinet 11 is assigned to each of the multiple cabinets 11. In the example of FIG. 8, an XY coordinate system is set, with the width direction being the X direction and the height direction being the Y direction. A logical number L1 is assigned to each cabinet 11 based on its position on the XY coordinate system. For example, the cabinet 11 located in the first row and first column is assigned the logical number L1 "CLX1CLY1." The cabinet 11 located in the second column and fifth row is assigned the logical number L1 "CLX5CLY2."

[0051] As shown in Fig. 9, a plurality of display units 12 are attached to one cabinet 11. The plurality of display units 12 attached to the same cabinet 11 are assigned logical numbers L2 based on their respective positions within the cabinet 11. For example, the display unit 12 located in the first row and first column of the cabinet 11 is assigned the logical number L2 "ULX1ULY1." The display unit 12 located in the second column and third row of the cabinet 11 is assigned the logical number L2 "ULX3ULY2."

[0052] 10, a logical number LN is assigned to each display unit 12 based on the position of the cabinet 11 to which the display unit 12 belongs and the position of the display unit 12 within the cabinet 11. For example, the logical number LN "CLX1CLY1-ULX1ULY1" is assigned to the display unit 12 in the first row, first column of the cabinet 11. The logical number LN "CLX1CLY1-ULX2ULY1" is assigned to the display unit 12 in the first row, second column of the cabinet 11 in the first row, first column.

[0053] [1-5. Connection between cabinet and control system] 11 and 12 are diagrams showing an example of a connection between the cabinet 11 and the control system 30. In FIG.

[0054] The multiple cabinets 11 are connected to the control system 30 by serial connection, parallel connection, or a combination of these. For example, in the example of Fig. 11, the multiple cabinets 11 are serially connected to the control system 30. Two adjacent cabinets 11 are connected by connecting their cabinet boards 113 together. The multiple cabinets 11 are assigned cabinet connection numbers CEk (k is an integer from 1 to 32). Video output signals and audio output signals are output from the control system 30 to the multiple cabinets 11 in accordance with the cabinet connection numbers.

[0055] 12, the multiple cabinets 11 are connected to the control system 30 using a method that combines serial connection and parallel connection. The multiple cabinets 11 are assigned cabinet connection numbers CE1,m (where 1 is an integer from 1 to 8, and m is an integer from 1 to 4). The control system 30 outputs video output signals and audio output signals to the multiple cabinets 11 in accordance with the cabinet connection numbers.

[0056] [1-6. Connection between the cabinet and the display unit] FIG. 13 is a diagram showing an example of a connection configuration between the cabinet 11 and the display unit 12. As shown in FIG.

[0057] The plurality of display units 12 supported by the same cabinet 11 are connected in parallel to a cabinet substrate 113. The plurality of display units 12 are electrically connected to the control system 30 via the cabinet substrate 113. Unit connection numbers UE1 to UE6 are assigned to the plurality of display units 12. Video output signals and audio output signals are output from the cabinet substrate 113 to the plurality of display units 12 in accordance with the unit connection numbers.

[0058] [2. First embodiment] [2-1. System Image] FIG. 14 is a diagram showing an example in which the audio / video content output system 1 is applied to a theater CT1.

[0059] The theater CT1 uses channel-based audio content AC. Figure 14 shows hypothetical positions of multi-channel speakers for the left channel LCH, center channel CCH, and right channel RCH.

[0060] In theaters that use a sound screen, multi-channel speakers are placed behind the sound screen, which has many tiny sound holes. The sound output from the multi-channel speakers is output to the audience (the front side of the sound screen) through the sound holes.

[0061] However, in the tiling display 10, multiple display units 12 are closely spaced, so it is not possible to provide holes such as sound holes in the tiling display 10. One possible method would be to arrange multi-channel speakers around the tiling display 10 to generate a phantom sound image, but this method limits the range of viewing positions in which the sound image can be correctly localized to a small area.

[0062] Therefore, in theater CT1, audio data AD for the left channel LCH, center channel CCH, and right channel RCH are mapped to the tiling display 10 (display unit 12). By directly playing back the audio data AD for the multi-channel speakers on the screen SCR, a sense of unity between video and audio, like that of a sound screen, is achieved.

[0063] [2-2. Audio data mapping process for channel-based audio] FIG. 15 is a diagram showing an example of mapping processing of audio data AD of channel-based audio.

[0064] The calculation unit 34 receives audio content AC of channel-based audio. The audio content AC includes one or more pieces of audio data AD generated for each channel. The sound source extraction unit 341 extracts audio data AD from the audio content AC for each channel that serves as a sound source. In the example of FIG. 15, four pieces of audio data AD corresponding to the left channel LCH, the center channel CCH, the right channel RCH, and the bass enhancement channel LFE are extracted.

[0065] The audio data AD of the left channel LCH, center channel CCH, and right channel RCH is assigned sounds in the frequency band from the high frequency range HF to the low frequency range LF. The audio data AD of the bass enhancement channel LFE is assigned sounds in the very low frequency range VLF. The sound source extraction unit 341 outputs the audio data AD of the left channel LCH, center channel CCH, and right channel RCH to the band division unit 342. The sound source extraction unit 341 outputs the audio data AD of the bass enhancement channel LFE to the subwoofer 23.

[0066] The band dividing unit 342 divides the audio data AD of the channels other than the bass enhancement channel LFE (left channel LCH, center channel CCH, and right channel RCH) into frequency bands. For example, the band dividing unit 342 divides each of the audio data AD of the left channel LCH, center channel CCH, and right channel RCH into high-frequency waveform data PAD and mid-low-frequency waveform data, and outputs the data to the mapping processing unit 343.

[0067] The mapping processing unit 343 maps the waveform data PAD of the high frequency range HF and the low-mid frequency range extracted from the audio data AD of each channel to one or more display units 12 and one or more speakers ASP determined by the position of the multi-channel speaker. The position of the multi-channel speaker is extracted from the reproduction environment information 352. The reproduction environment information 352 specifies, for example, the coordinates on the screen SCR where the center of the multi-channel speaker is located as the position of the multi-channel speaker. The mapping processing unit 343 extracts a predetermined area on the screen SCR centered on these coordinates as the sound source area SR.

[0068] For example, the mapping processing unit 343 extracts a sound source region LSR of the left channel LCH, a sound source region CSR of the center channel CCH, and a sound source region RSR of the right channel RCH as the sound source region SR of each channel from the playback environment information 352. In the example of Fig. 15, the region indicated by thick hatching (in the example of Fig. 15, the region spanning eight display units 12) is extracted as the sound source region SR.

[0069] The mapping processing unit 343 maps the waveform data PAD of the high frequency HF of the left channel LCH to one or more display units 12 arranged in the sound source region LSR of the left channel LCH. The mapping processing unit 343 maps the waveform data PAD of the low-mid frequency of the left channel LCH to one or more speakers ASP arranged at the same position on the X axis as the sound source region LSR of the left channel LCH.

[0070] When the sound pressure of the high frequency HF of the left channel LCH is high, if the set sound pressure is to be achieved using only the display units 12 arranged in the sound source region LSR, the vibration of each display unit 12 will increase. If the vibration of the display units 12 increases, the viewer U may perceive shaking of the image.

[0071] Therefore, the mapping processing unit 343 expands the mapping destination to the periphery of the sound source region LSR. The mapping processing unit 343 also maps the waveform data PAD to one or more display units 12 (five display units 12 indicated by light hatching in the example of FIG. 15) arranged around the sound source region LSR. The mapping processing unit 343 also expands the mapping destination of the waveform data PAD for the low-mid range in accordance with the expansion of the mapping destination of the waveform data PAD for the high HF range. This makes it less likely that a discrepancy will occur between the sound image for the high HF range and the sound image for the low-mid range.

[0072] The waveform data PAD of the center channel CCH and the right channel RCH are also mapped in a similar manner.

[0073] That is, the mapping processing unit 343 maps the high HF waveform data PAD of the center channel CCH to one or more display units 12 arranged in the sound source region CSR of the center channel CCH. The mapping processing unit 343 maps the low-mid frequency waveform data PAD of the center channel CCH to one or more speakers ASP arranged at the same position on the X axis as the sound source region CSR. If the sound pressure of the high HF frequency of the left channel LCH is high, the mapping processing unit 343 expands the mapping destination to the periphery of the sound source region CSR. The mapping processing unit 343 also expands the mapping destination of the low-mid frequency waveform data PAD in accordance with the expansion of the mapping destination of the high HF waveform data PAD.

[0074] The mapping processing unit 343 maps the waveform data PAD of the high HF range of the right channel RCH to one or more display units 12 arranged in the sound source region RSR of the right channel RCH. The mapping processing unit 343 maps the waveform data PAD of the mid-low range of the right channel RCH to one or more speakers ASP arranged at the same position on the X axis as the sound source region RSR. If the sound pressure of the high HF range of the right channel RCH is large, the mapping processing unit 343 expands the mapping destination to the periphery of the sound source region RSR. The mapping processing unit 343 also expands the mapping destination of the mid-low range waveform data PAD in accordance with the expansion of the mapping destination of the high HF range waveform data PAD.

[0075] The mapping processing unit 343 serializes the waveform data PAD mapped to each display unit 12. The mapping processing unit 343 outputs the acoustic output signal for the display unit 12 generated by the serialization processing to the tiling display 10. The mapping processing unit 343 generates an acoustic output signal for the speaker ASP based on the waveform data PAD mapped to each speaker ASP, and outputs it to the first array speaker 21 and the second array speaker 22.

[0076] [2-3. Audio data mapping process for object-based audio] 16 and 17 are diagrams showing an example of mapping processing of audio data AD of object-based audio.

[0077] 16, object-based audio content AC is input to the calculation unit 34. The audio content AC includes one or more pieces of audio data AD generated for each object. The sound source extraction unit 341 extracts audio data AD from the audio content AC for each object that serves as a sound source.

[0078] In the example of Fig. 16, a video of a character snapping his fingers is displayed on the screen SCR. The audio content AC includes audio data AD of the sound of the snapping (object) and meta information indicating the position of the snapping (object position OB). In the example of Fig. 16, there is one object, but the number of objects is not limited to one. As shown in Fig. 17, different objects may be placed at multiple positions OB. In this case, the sound source extraction unit 341 extracts multiple pieces of audio data AD corresponding to different objects from the audio content AC.

[0079] The band dividing unit 342 divides the waveform data of the audio data AD above the low frequency band LF into frequency bands. For example, the band dividing unit 342 divides the audio data AD of the object into waveform data PAD of the high frequency band HF and waveform data of the low-mid frequency band, and outputs them to the mapping processing unit 343.

[0080] The mapping processor 343 maps the high-frequency (HF) and low-mid-frequency (MHF) waveform data PAD extracted from the object's audio data AD to one or more display units 12 and one or more speakers ASP corresponding to the object's position OB. The object's position OB is defined in the meta information as, for example, information on the horizontal angle, elevation angle, and distance from a preset viewing position. The mapping processor 343 extracts a predetermined area on the screen SCR centered on the position OB as a sound source area OSR. In the example of FIG. 16, the sound source area OSR is extracted as an area the size of one display unit 12, indicated by thick hatching.

[0081] FIG. 16 shows a state in which the sound source regions LSR, CSR, and RSR of each channel and the sound source region OSR of the object coexist as the sound source region SR.

[0082] The mapping processing unit 343 maps the waveform data PAD of the high frequency range of the object to one or more display units 12 arranged in the sound source region SR of the object. The mapping processing unit 343 maps the waveform data PAD of the mid-low frequency range of the object to one or more speakers ASP arranged at the same position on the X axis as the sound source region OSR of the object.

[0083] When the sound pressure of the high HF frequency of the object is large, the mapping processing unit 343 expands the mapping destination to the periphery of the sound source region SR (the three display units 12 indicated by light hatching in the example of FIG. 16). The mapping processing unit 343 also expands the mapping destination of the mid-low frequency waveform data PAD in accordance with the expansion of the mapping destination of the high HF frequency waveform data PAD.

[0084] The mapping processing unit 343 serializes the waveform data PAD mapped to each display unit 12. The mapping processing unit 343 outputs the acoustic output signal for the display unit 12 generated by the serialization processing to the tiling display 10. The mapping processing unit 343 serializes the waveform data PAD mapped to each speaker ASP. The mapping processing unit 343 outputs the acoustic output signal for the speaker ASP generated by the serialization processing to the first array speaker 21 and the second array speaker 22.

[0085] [2-4. Sound source placement using DNN engine] FIG. 18 is a diagram showing another example of the mapping process of the audio data AD of the channel-based audio.

[0086] The calculation unit 34 receives input of audio content AC of channel-based audio. The sound source extraction unit 341 uses a sound source separation technique to extract audio data AD for each sound source SS from the audio content AC. As the sound source separation technique, a known sound source separation technique such as blind source separation is used. In the example of FIG. 18, each character displayed on the screen SCR is the sound source SS. For each sound source SS, the sound source extraction unit 341 extracts the speaking voice of the character that is the sound source SS as audio data AD. Note that in the example of FIG. 18, sound sources SS1, SS2, and SS3 are extracted as the sound sources SS. However, the number N of sound sources SS is not limited to this. The number N of sound sources SS can be any number equal to or greater than 1.

[0087] The position of the sound source SS is estimated by a sound source position estimation unit 345. The sound source position estimation unit 345 applies one or more pieces of audio data AD and video content AC extracted by the sound source extraction unit 341 to an analysis model 351 using, for example, a DNN engine. Based on the analysis result by the analysis model 351, the sound source extraction unit 341 estimates, for each sound source SS, the position on the screen SCR where the sound source SS is displayed as a sound source region SR.

[0088] For each sound source SS, the mapping processing unit 343 maps the audio data AD of the sound source SS to one or more display units 12 arranged at the position of the sound source SS. The mapping processing unit 343 serializes the audio data AD of each sound source SS based on the mapping result. The mapping processing unit 343 outputs the acoustic output signal obtained by the serialization processing to the tiling display 10.

[0089] 18, the sound source region SR1 of the sound source SS1 is estimated as a region spanning four display units 12. When the speaking volume of the sound source SS1 is low, the mapping processing unit 343 selects the four display units 12 in which the sound source region SR1 is located as the mapping destination of the audio data AD of the sound source SS1.

[0090] The sound source region SR2 of the sound source SS2 is estimated as a region spanning two display units 12. When the speech of the sound source SS2 is loud, the mapping processing unit 343 selects the two display units 12 in which the sound source region SR2 is located (display units 12 with dark hatching) and the five display units 12 located around them (display units 12 with light hatching) as mapping destinations for the audio data AD of the sound source SS2.

[0091] The sound source region SR3 of the sound source SS3 is estimated as a region spanning two display units 12. If the speech volume of the sound source SS3 is low, the mapping processing unit 343 selects the two display units 12 in which the sound source region SR3 is located as the mapping destination for the audio data AD of the sound source SS3.

[0092] [2-5. Controlling the sound image in the depth direction] 19 to 22 are diagrams for explaining a method for controlling a sound image in the depth direction.

[0093] The position of the sound image in the depth direction is controlled by known signal processing such as monopole synthesis, wave field synthesis (WFS), spectral division method, and mode matching.

[0094] For example, assume that multiple point sound sources PS are arranged on a reference plane RF, as shown in Figures 20 and 21. When the sound pressure and phase of the multiple point sound sources PS are appropriately controlled, a sound field is generated with a focal point FS located away from the reference plane RF. The sound image is localized at the focal point FS. As shown in Figure 20, when the focal point FS moves further back than the reference plane RF, a sound image is generated that seems to move away from the viewer U. As shown in Figure 21, when the focal point FS moves in front of the reference plane RF, a sound image is generated that seems to move closer to the viewer U.

[0095] The point sound source PS corresponds to each display unit 12 or speaker ASP. The reference plane RF corresponds to the screen SCR of the tiling display 10 or the sound output surface of the array speaker (first array speaker 21, second array speaker 22).

[0096] As shown in FIG. 19, the mapping processing unit 343 uses an FIR (Finite Impulse Response) filter to control the sound pressure and phase of the sound output from the display unit 12 and speaker ASP, which are the mapping destinations.

[0097] 16 except that digital filtering using an FIR filter is performed on the waveform data PAD. That is, the audio data AD extracted by the sound source extraction unit 341 is divided by the band division unit 342 into high-frequency HF waveform data PAD and mid-low frequency waveform data PAD. The high-frequency HF waveform data PAD is mapped to n (n is an integer of 2 or greater) display units 12 corresponding to the object positions OB. The mid-low frequency waveform data PAD is mapped to m (m is an integer of 2 or greater) speakers ASP corresponding to the object positions OB.

[0098] The mapping processing unit 343 performs digital filtering using an FIR filter on the high-frequency HF waveform data PAD. The mapping processing unit 343 uses digital filtering to adjust the sound pressure and phase of the sound output from the n display units 12 to which the high-frequency HF waveform data PAD is mapped, for each display unit 12. The mapping processing unit 343 controls the position of the sound image in the depth direction by adjusting the sound pressure and phase of the sound output from each display unit 12.

[0099] The mapping processing unit 343 performs digital filtering using an FIR filter on the mid-low frequency waveform data PAD. The mapping processing unit 343 uses digital filtering to adjust the sound pressure and phase of the sound output from the m speakers ASP to which the mid-low frequency waveform data PAD is mapped, for each speaker ASP. The mapping processing unit 343 controls the position of the sound image in the depth direction by adjusting the sound pressure and phase of the sound output from each speaker ASP.

[0100] [2-6. Sound image localization emphasis control] [2-6-1. Strengthening sound image localization by expanding the frequency band] FIG. 22 is a diagram showing an example of a sound image localization enhancement control technique.

[0101] FIG. 22 shows audio data AD with a low sound pressure level in the high HF range. When the audio data AD is divided into bands, waveform data PAD with a low sound pressure in the high HF range is generated. The sound image localization ability varies depending on the sound pressure of the waveform data PAD in the high HF range. Therefore, the mapping processing unit 343 uses high-frequency interpolation technology to generate corrected audio data CAD with a sound pressure level in the high HF range equal to or greater than the threshold TH from the audio data AD with a sound pressure level in the high HF range lower than the threshold TH. The mapping processing unit 343 maps the waveform data PAD in the high HF range of the corrected audio data CAD to one or more display units 12 as mapping destinations.

[0102] [2-6-2. Enhancement of sound localization ability by precedence effect] FIG. 23 is a diagram showing another example of a sound image localization enhancement control technique.

[0103] Figure 23 shows the relationship between the frequency band and phase of the audio data AD. The phase is related to the timing of sound output. With the original audio data AD, mid-low and very low frequency (VLF) sounds, which have low sound image localization capability, and high frequency (HF) sounds, which have high sound image localization capability, are output simultaneously.

[0104] Therefore, the mapping processing unit 343 outputs the high HF waveform data PAD simultaneously with the output of the low-mid and ultra-low VLF waveform data PAD, or earlier than the output of the low-mid and ultra-low VLF waveform data PAD. By outputting the high HF sound first, the listener U can quickly recognize the position of the sound image. During the period when the low-mid and ultra-low VLF sounds are being output, the listener U can recognize the sound image at the position localized by the high HF sound, which is the preceding sound.

[0105] [2-7. Speaker unit placement] FIG. 24 is a diagram showing an example of the arrangement of the speaker units 20. As shown in FIG.

[0106] An enclosure that houses a first array speaker 21 is attached to the top cabinet 11 of the tiling display 10. An enclosure that houses a second array speaker 22 is attached to the bottom cabinet 11 of the tiling display 10. A slit that serves as a sound guide section SSG is provided in the enclosure. The width of the slit is narrower than the diameter of the speaker ASP. Sound output from the speaker ASP is emitted to the outside of the enclosure via the sound guide section SSG. The sound guide section SSG is positioned close to the edge of the tiling display 10. Since sound is output from right at the edge of the tiling display 10, high sound image localization ability is achieved.

[0107] As shown in the enlarged view, the speaker ASP may be housed in a cabinet 11. In this case, dedicated end cabinets with built-in speakers and sound guide sections SSG are arranged at the top and bottom levels of the tiling display 10.

[0108] [2-8. How to detect the position of the display unit] Fig. 25 is a diagram showing an example of a method for detecting the position of the display unit 12. Fig. 26 is a diagram showing the arrangement of microphones MC used to detect the position of the display unit 12.

[0109] As shown in Fig. 25, display units with microphones 12M are arranged at the four corners of the tiling display 10. As shown in Fig. 26, a microphone MC is attached to the rear surface of the display unit with microphone 12M. A cutout that serves as a sound guide section CSG is formed in one corner of the display unit with microphone 12M. The microphone MC is arranged near the corner of the display unit with microphone 12M where the cutout is formed.

[0110] The position detection unit 344 detects the spatial position of the display unit 12 based on the time it takes for a sound (impulse) output from the display unit 12 to reach each of the microphones MC provided at multiple locations. The position detection unit 344 assigns a logical number LN to each display unit 12 based on the spatial arrangement of each display unit 12.

[0111] For example, the position detection unit 344 selects one display unit 12 for each cabinet 11 and outputs a sound (impulse) from the selected display unit 12. The position detection unit 344 acquires measurement data MD relating to the propagation time of the sound from each microphone MC. The position detection unit 344 detects the spatial position of the cabinet 11 based on the measurement data MD acquired from each microphone MC.

[0112] The arrangement of the display units 12 in the cabinet 11 is defined in the playback environment information 352. The position detection unit 344 detects the relative position of the cabinet 11 and each display unit 12 held in the cabinet 11 based on the information on the arrangement defined in the playback environment information 352. The position detection unit 344 detects the position of each display unit 12 based on the position of the cabinet 11 and the relative position of each display unit 12 with respect to the cabinet 11.

[0113] If there is an obstacle in front of the tiling display 10 that reflects sound, accurate measurements may not be possible. In such cases, measurement accuracy can be improved by installing microphones MC on all display units 12 or on multiple display units 12 arranged at a certain density. The microphones MC can also be used for acoustic correction of the sound output from the display units 12.

[0114] FIG. 27 is a diagram showing another example of a method for detecting the position of the display unit 12. In FIG.

[0115] 27, multiple microphones MC are placed outside the tiling display 10. Although the positions of the microphones MC are different, the position detection unit 344 can detect the position of each display unit 12 in the same manner as described in FIG. 25. In the example of FIG. 27, there is no need to provide the tiling display 10 with a sound guide CSG for transmitting sound to the microphones MC. Therefore, degradation of image quality due to the sound guide CSG is unlikely to occur.

[0116] [2-9. Controlling the directionality of reproduced sound] FIG. 28 is a diagram for explaining the directivity control of the reproduced sound DS.

[0117] The directivity of the reproduced sound DS is controlled by utilizing the interference of the wavefronts of multiple arranged point sound sources. For example, the directivity of the reproduced sound DS in the height direction is controlled by the interference of the wavefronts of multiple point sound sources lined up in the height direction. The directivity of the reproduced sound DS in the width direction is controlled by the interference of the wavefronts of multiple point sound sources lined up in the width direction. The point sound sources correspond to individual display units 12 or speakers ASP. For example, the mapping processing unit 343 uses an FIR filter to individually control the sound pressure and phase of the sound output from each display unit 12 and speaker ASP that is the mapping destination.

[0118] 15 except that digital filtering using an FIR filter is performed on the waveform data PAD. That is, the audio data AD extracted by the sound source extraction unit 341 is divided by the band division unit 342 into high-frequency HF waveform data PAD and mid-low frequency waveform data PAD. The high-frequency HF waveform data PAD is mapped to n (n is an integer of 2 or greater) display units 12 corresponding to the positions of the multi-channel speakers. The mid-low frequency waveform data PAD is mapped to m (m is an integer of 2 or greater) speakers ASP corresponding to the positions of the multi-channel speakers.

[0119] The mapping processing unit 343 performs digital filtering using an FIR filter on the high-frequency HF waveform data PAD. The mapping processing unit 343 uses digital filtering to adjust the sound pressure and phase of the sound output from the n display units 12 to which the high-frequency HF waveform data PAD is mapped, for each display unit 12. By adjusting the sound pressure and phase of the sound output from each display unit 12, the mapping processing unit 343 controls acoustic characteristics such as the directivity and sound pressure uniformity of the reproduced sound DS within the viewing area VA.

[0120] The mapping processing unit 343 performs digital filtering using an FIR filter on the mid-low frequency waveform data PAD. The mapping processing unit 343 uses digital filtering to adjust the sound pressure and phase of the sound output from the m speakers ASP to which the mid-low frequency waveform data PAD is mapped, for each speaker ASP. By adjusting the sound pressure and phase of the sound output from each speaker ASP, the mapping processing unit 343 controls acoustic characteristics such as the directivity and sound pressure uniformity of the reproduced sound DS within the listening area VA.

[0121] FIG. 29 is a diagram showing an example in which different playback sounds DS are allocated to each viewer U.

[0122] One or more cameras CA are installed near the tiling display 10. The cameras CA are wide-angle cameras that can capture an image in front of the tiling display 10. In the example of Fig. 29, one camera CA is installed on each side in the width direction of the tiling display 10 in order to cover the entire viewing area VA of the tiling display 10.

[0123] The control system 30 detects the number of viewers U present in the viewing area VA and the positions of each viewer U based on the shooting data acquired from each camera CA. The tiling display 10 displays images of multiple sound sources SS set for each viewer U at different positions on the screen SCR. For each sound source SS, the mapping processing unit 343 selects multiple display units 12 corresponding to the display positions of the sound source SS as mapping destinations for the audio data AD of the sound source SS. Based on the position information of each viewer U, the mapping processing unit 343 generates and outputs, for each viewer U, a playback sound DS with high directionality from the sound source SS toward the viewer U.

[0124] [2-10. Information processing method] FIG. 30 is a flowchart showing an example of an information processing method performed by the control system 30.

[0125] In step S1, the sound source extraction unit 341 extracts one or more pieces of audio data AD from audio content AC. Various types of audio content AC can be used as the audio content AC, such as channel-based audio, object-based audio, and scene-based audio. For example, the sound source extraction unit 341 extracts one or more pieces of audio data AD from the audio content AC, which are generated for each channel or object that serves as a sound source.

[0126] In step S2, the mapping processing unit 343 selects, for each piece of audio data AD, one or more display units 12 and one or more speakers ASP to which the audio data AD is to be mapped. For example, the mapping processing unit 343 detects a sound source region SR on the screen SCR that corresponds to the position of a multi-channel speaker or the position OB of an object. The mapping processing unit 343 selects one or more display units 12 and one or more speakers ASP that correspond to the sound source region SR as the mapping destination. The mapping processing unit 343 expands the mapping destination outside the sound source region SR based on the sound pressure of the audio data AD, the depth position of the sound image, the directivity of the reproduced sound DS, and the like.

[0127] In step S3, the mapping processing unit 343 outputs the audio data AD to one or more display units 12 and one or more speakers ASP as mapping destinations, and localizes the sound image at a position associated with the sound source (the sound source region SR or a position shifted in the depth direction from the sound source region SR).

[0128] [2-11.Effects] The control system 30 has a sound source extraction unit 341 and a mapping processing unit 343. The sound source extraction unit 341 extracts one or more pieces of audio data AD corresponding to different sound sources from audio content AC. The mapping processing unit 343 selects, for each piece of audio data AD, one or more display units 12 to which the audio data AD is to be mapped from one or more combinable display units 12 each having a sound generation mechanism. In the information processing method of this embodiment, the processing of the control system 30 described above is executed by a computer. The program of this embodiment causes a computer to realize the processing of the control system 30 described above.

[0129] According to this configuration, the audio data AD is directly reproduced on the display unit 12. This makes it easier to obtain a sense of unity between the video and audio.

[0130] The audio data AD is audio data for multi-channel speakers extracted from the channel-based audio content AC. The mapping processing unit 343 selects one or more display units 12 as mapping destinations, which are determined depending on the arrangement of the multi-channel speakers.

[0131] This configuration produces a powerful sound, as if multi-channel speakers were placed in front of the screen SCR.

[0132] The audio data AD is audio data of an object extracted from the object-based audio content AC. The mapping processing unit 343 selects one or more display units 12 corresponding to the position OB of the object extracted from the audio content AC as a mapping destination.

[0133] According to this configuration, the sound image of the object can be localized at the position OB of the object.

[0134] The control system 30 includes a sound source position estimation unit 345. The sound source position estimation unit 345 estimates, for each piece of audio data AD, the position at which the sound source SS of the audio data AD is displayed. The mapping processing unit 343 selects, as mapping destinations, one or more display units 12 corresponding to the positions at which the sound source SS is displayed.

[0135] According to this configuration, the sound image of the sound source SS can be localized at the position where the sound source SS is displayed.

[0136] The mapping processing unit adjusts the sound pressure and phase of the sound output from the plurality of display units 12 that are the mapping destinations for each display unit 12, thereby controlling the position of the sound image in the depth direction.

[0137] According to this configuration, the position of the sound image in the depth direction can be easily controlled.

[0138] The control system 30 includes a band dividing unit 342. The band dividing unit 342 divides the audio data AD into frequency bands. The mapping processing unit 343 maps waveform data PAD of the highest frequency band HF of the audio data AD to one or more display units 12 as mapping destinations.

[0139] According to this configuration, high-frequency HF sounds with high sound image localization capability are output from the display unit 12. Therefore, a discrepancy between the position of the sound source and the position of the sound image is unlikely to occur.

[0140] The mapping processing unit 343 selects one or more speakers ASP corresponding to the positions of the sound sources of the audio data AD from the multiple speakers ASP arranged around the multiple display units 12. The mapping processing unit 343 maps the waveform data PAD of the low frequency LF, which has the lowest frequency, of the audio data AD, and the waveform data PAD of the mid frequency MF, which is between the high frequency HF and the low frequency LF, to the selected one or more speakers ASP.

[0141] With this configuration, mid-range MF and low-range LF sounds, which have lower sound image localization ability than the high-range HF sounds, are output from the speaker ASP. Because the only sounds output from the display unit 12 are high-range HF sounds, vibrations of the display unit 12 when sounds are output are kept to a minimum.

[0142] The mapping processing unit 343 generates corrected audio data CAD whose high HF sound pressure level is equal to or greater than a threshold from audio data AD whose high HF sound pressure level is less than a threshold. The mapping processing unit 343 maps the high HF waveform data PAD of the corrected audio data CAD to one or more display units 12 as mapping destinations.

[0143] According to this configuration, high sound image localization performance can be obtained even for audio data AD with a low sound pressure level in the high frequency range HF.

[0144] The mapping processing unit 343 outputs the waveform data PAD for the high frequency band HF at the same time as the waveform data PAD for the mid frequency band MF and low frequency band LF is output, or earlier than the timing at which the waveform data PAD for the mid frequency band MF and low frequency band LF is output.

[0145] According to this configuration, the output timing of the waveform data PAD of the high frequency band HF, which has a high sound image localization capability, is advanced, and therefore the sound image localization capability of the audio data AD is improved due to the precedence effect.

[0146] The control system 30 includes a position detection unit 344. The position detection unit 344 detects the spatial arrangement of the multiple display units 12. The position detection unit 344 assigns a logical number LN to each display unit 12 based on the detected spatial arrangement. The mapping processing unit 343 identifies a mapping destination based on the logical number LN.

[0147] According to this configuration, the addressing of the display unit 12 can be performed automatically.

[0148] The position detection unit 344 detects the spatial arrangement of the display unit 12 based on the time it takes for the sound output from the display unit 12 to reach each of the microphones MC provided at a plurality of locations.

[0149] According to this configuration, the spatial arrangement of the display unit 12 can be easily detected.

[0150] The mapping processing unit controls the directivity of the reproduced sound DS by adjusting the sound pressure and phase of the sound output from the plurality of display units 12 that are the mapping destinations for each display unit 12.

[0151] According to this configuration, the directivity of the reproduced sound DS is controlled by the interference of the wavefronts output from the display units 12.

[0152] 3. Second Embodiment [3-1. System Image] 31 and 32 are diagrams showing an example in which the audio / video content output system 1 is applied to a theater CT2.

[0153] As shown in Fig. 31, theater CT2 is a theater capable of displaying spherical images. As shown in Fig. 32, tiling displays 10 are arranged to cover the entire front, left and right sides, ceiling, and floor of the audience seats ST. Sound is reproduced from all directions by the numerous display units 12 installed in all directions.

[0154] [3-2. Speaker unit placement] FIG. 33 is a diagram showing an example of the arrangement of the speaker units 20. As shown in FIG.

[0155] In the theater CT2, a large number of display units 12 are arranged closely together in all directions. This limits the space available for installing the speaker units 20. For example, in the first embodiment, the mid-low range speaker units 20 (first array speaker 21, second array speaker 22) are installed along the top and bottom sides of the tiling display 10. However, in the theater CT2, the tiling display 10 is installed in all directions, so there is no space available for installing the first array speaker 21 and the second array speaker 22.

[0156] For this reason, in theater CT2, woofers 24 are installed in the shoulder areas of the seats of the audience ST as speaker units 20 for mid-low frequencies. Subwoofers 23, which are speaker units 20 for VLF frequencies, are installed under the seats. High-frequency HF sounds, which have high sound image localization capabilities, are output from the display unit 12. By installing the speaker units 20 in the seats, the distance from the speaker units 20 to the audience U is shortened. Therefore, there is no need to reproduce excessive sound pressure. This reduces unnecessary reverberation in theater CT2.

[0157] FIG. 34 is a diagram showing another example of the arrangement of the speaker units 20. In FIG.

[0158] In the example of FIG. 34, open-type earphones EP are worn on the ears UE of the viewer U as the mid-low frequency speaker units 20. The earphones EP have openings OP in the ear canals. The viewer U can hear and listen to the sound output from the display unit 12 through the openings OP. The speaker units 20 do not necessarily have to be earphones EP, but can be any wearable acoustic device (open headphones, shoulder speakers, etc.) that can be worn by the viewer U. In the example of FIG. 34 as well, the distance from the speaker units 20 to the viewer U is shortened. This eliminates the need to reproduce excessive sound pressure, and unnecessary reverberation is suppressed.

[0159] [3-3. Measurement of spatial characteristics and reverberation cancellation using built-in microphone] FIG. 35 is a diagram showing the arrangement of microphones MC used to measure spatial characteristics.

[0160] Because the tiling display 10 covers all directions, sound reflections occur between screen portions facing each other, which can reduce the sense of sound positioning. Therefore, the control system 30 controls the sound pressure and phase of each display unit 12 based on the spatial characteristics of the theater CT2 measured in advance, thereby reducing reverberation. The placement of the microphones MC is similar to that described in FIG. 26. In the example of FIG. 26, microphones MC are installed only on specific display units 12, but in this embodiment, microphones MC are installed on all display units 12.

[0161] The spatial characteristics of the theater CT2 are measured using a microphone MC built into each display unit 12. For example, in the theater CT2, the output characteristics of the output sound of each display unit 12 relative to all other display units (microphones MC) are measured for each display unit 12. This measurement measures the transfer characteristics of the wavefront (transfer characteristics with frequency and sound pressure as variables, and transfer characteristics with frequency and phase (including transfer time) as variables). The spatial characteristics of the theater CT2 are detected based on the transfer characteristics. The spatial characteristics of the theater CT2 are stored in the memory unit 35 as playback environment information 352.

[0162] The mapping processing unit 343 adjusts the sound pressure and phase of the sound output from the multiple display units 12 that are to be mapped, for each display unit 12, based on the spatial characteristics of the theater CT2, to reduce reverberation. For example, the display units 12 selected as the mapping destination are designated as mapping destination units, and the display units 12 that are not selected as mapping destinations are designated as non-mapping destination units. When the sound output from the mapping destination units reaches and reflects off the non-mapping destination units, the mapping processing unit 343 reproduces sound that is in opposite phase to the primary reflected wavefront at the non-mapping destination units. This reduces reverberation caused by reflection at the non-mapping destination units.

[0163] 4. Third Embodiment [4-1. System Image] FIG. 36 is a diagram showing an example in which the audio / video content output system 1 is applied to a telepresence system TP.

[0164] The telepresence system TP is a system that connects remote locations and conducts two-way video and audio conferences. An entire wall is a tiling display 10 that displays video from the remote locations. The video and audio of viewer U1 at a first remote location are output to viewer U2 from tiling display 10B at a second remote location. The video and audio of viewer U2 at the second remote location are output to viewer U1 from tiling display 10A at the first remote location.

[0165] [4-2. Object sound collection and playback] FIG. 37 is a diagram showing an example of the sound collection process and playback process of an object sound.

[0166] One or more cameras CA are installed near the tiling display 10. The cameras CA are wide-angle cameras that can capture an image in front of the tiling display 10. In the example of Fig. 37, one camera CA is installed on each side of the width direction of the tiling display 10 in order to cover the entire viewing area VA of the tiling display 10.

[0167] At the first remote location, the number of viewers U1 present in the viewing area VA, the position of each viewer U1, and the mouth movements of each viewer U1 are detected based on the image data from each camera CA. The voices of the viewers U1 are collected as input sound IS by a highly directional microphone built into each display unit 12. The control system 30A inputs the collected sound data and the image data from camera CA to a DNN to perform sound source separation and generate audio content AC in which the voice of viewer U1, which is the sound source, is used as an object. The control system 30A generates content data CD using video content generated using the image data from camera CA and audio content AC generated using the input sound IS.

[0168] The control system 30B at the second remote location acquires the content data CD generated by the control system 30A at the first remote location via the network NW. The control system 30B separates the content data CD into audio content AC and video content VC. The control system 30B uses the video content VC to play back the video of the viewer U1 at the first remote location on the tiling display 10B. The control system 30B uses the audio content AC to play back the audio of the viewer U1 at the first remote location on the tiling display 10B and multiple speaker units 20B. The playback process of the audio content AC is similar to that shown in FIG. 16.

[0169] When playing back audio content AC, the control system 30B detects the number of viewers U2 present in the viewing area VA and the positions of each viewer U2 based on the image data acquired from each camera CA. The tiling display 10B displays on the screen SCR an image of a viewer U1 at a first remote location who is the sound source of the object. The mapping processor 343 selects multiple display units 12 corresponding to the position of the object (audio of viewer U1) as a mapping destination for the audio data AD of the object. Based on the position information of each viewer U2, the mapping processor 343 generates and outputs, for each viewer U2, a playback sound DS with high directionality toward viewer U2 from the multiple display units 12 that are the mapping destinations. The method for controlling the directionality of the playback sound DS is the same as that shown in FIG. 29.

[0170] 5. Fourth Embodiment FIG. 38 is a diagram showing an example in which the audio / video content output system 1 is applied to a digital signage system DSS.

[0171] The digital signage system DSS is a system that transmits information using digital video equipment instead of conventional billboards and paper posters. The walls of buildings and corridors serve as tiling displays 10 that project images. In the digital signage system DSS, a digital advertisement DC is generated for each viewer U. The tiling display 10 displays multiple digital advertisements DC generated for each viewer U at different positions on the screen SCR. For each digital advertisement DC that serves as a sound source, the mapping processor 343 selects multiple display units 12 corresponding to the display position of the digital advertisement DC as the mapping destination for the audio data AD of the digital advertisement DC. Based on the position information of each viewer U, the mapping processor 343 generates and outputs, for each viewer U, a highly directional playback sound from the display position of the digital advertisement DC toward the viewer U.

[0172] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0173] [Note] The present technology can also be configured as follows. (1) a sound source extraction unit for extracting one or more pieces of audio data corresponding to different sound sources from the audio content; a mapping processing unit that selects, for each piece of audio data, one or more display units to which the audio data is to be mapped from among one or more combinable display units each having a sound generating mechanism; An information processing device having the above. (2) the audio data is audio data for multi-channel speakers extracted from the audio content of channel-based audio; The mapping processing unit selects one or more display units determined by the arrangement of the multi-channel speakers as the mapping destination. The information processing device according to (1) above. (3) the audio data is audio data of an object extracted from the audio content of object-based audio; The mapping processing unit selects one or more display units corresponding to the position of the object extracted from the audio content as the mapping destination. The information processing device according to (1) above. (4) a sound source position estimation unit that estimates a position where a sound source of the audio data is displayed for each of the audio data; The mapping processing unit selects one or more display units corresponding to the position where the sound source is displayed as the mapping destination. The information processing device according to (1) above. (5) The mapping processing unit adjusts the sound pressure and phase of the sound output from the plurality of display units to be mapped for each display unit, thereby controlling the position of the sound image in the depth direction. The information processing device according to (3) or (4) above. (6) a band dividing unit that divides the audio data into frequency bands, The mapping processing unit maps waveform data of the highest frequency range among the audio data to the one or more display units that are the mapping destinations. The information processing device according to any one of (1) to (5) above. (7) The mapping processing unit selects one or more speakers corresponding to the position of a sound source of the audio data from a plurality of speakers arranged around the plurality of display units, and maps waveform data of a low frequency range having the lowest frequency of the audio data and waveform data of a mid-frequency range between the high frequency range and the low frequency range to the one or more selected speakers. The information processing device according to (6) above. (8) The mapping processing unit generates corrected audio data whose high-frequency sound pressure level is equal to or greater than a threshold from audio data whose high-frequency sound pressure level is less than a threshold, and maps waveform data of the high-frequency range of the corrected audio data onto the one or more display units that are the mapping destinations. The information processing device according to (6) or (7) above. (9) The mapping processing unit outputs the high-frequency waveform data at the same time as the mid-frequency and low-frequency waveform data are output, or earlier than the mid-frequency and low-frequency waveform data are output. The information processing device according to (7) above. (10) a position detection unit that detects a spatial arrangement of the plurality of display units and assigns a logical number to each display unit based on the spatial arrangement; The mapping processing unit identifies the mapping destination based on the logical number. The information processing device according to any one of (1) to (9) above. (11) The position detection unit detects the spatial arrangement of the display unit based on the time it takes for the sound output from the display unit to reach each of the microphones provided at a plurality of locations. The information processing device according to (10) above. (12) The mapping processing unit controls the directivity of the reproduced sound by adjusting the sound pressure and phase of the sound output from the plurality of display units to which the mapping is applied for each display unit. The information processing device according to any one of (1) to (11) above. (13) The mapping processing unit reduces reverberation by adjusting the sound pressure and phase of the sound output from the plurality of display units to which the mapping is applied for each display unit. The information processing device according to any one of (1) to (12) above. (14) extracting from the audio content one or more audio data corresponding to different audio sources; selecting, for each piece of audio data, one or more display units to which the audio data is to be mapped from among one or more combinable display units each having a sound generating mechanism; 10. A computer-implemented information processing method comprising: (15) extracting from the audio content one or more audio data corresponding to different audio sources; selecting, for each piece of audio data, one or more display units to which the audio data is to be mapped from among one or more combinable display units each having a sound generating mechanism; A program that makes a computer do something. [Explanation of symbols]

[0174] 12 Display unit 30 Control system (information processing device) 341 Sound source extraction section 342 Band division section 343 Mapping processing section 344 Position detection unit 345 Sound source position estimation section AC Audio Content AD Audio Data

Claims

1. a sound source extraction unit for extracting one or more pieces of audio data corresponding to different sound sources from the audio content; a mapping processing unit that selects, for each piece of audio data, one or more display units to which the audio data is to be mapped from a plurality of combinable display units each having a sound generating mechanism; a band dividing unit that divides the audio data into frequency bands; and the mapping processing unit maps waveform data of the highest frequency band among the audio data to the one or more display units that are the mapping destinations; Information processing device.

2. the audio data is audio data for multi-channel speakers extracted from the audio content of channel-based audio; The mapping processing unit selects one or more display units determined depending on the arrangement of the multi-channel speakers as the mapping destination. The information processing device according to claim 1 .

3. the audio data is audio data of an object extracted from the audio content of object-based audio; The mapping processing unit selects one or more display units corresponding to the position of the object extracted from the audio content as the mapping destination. The information processing device according to claim 1 .

4. a sound source position estimation unit that estimates a position where a sound source of the audio data is displayed for each of the audio data; The mapping processing unit selects one or more display units corresponding to the position where the sound source is displayed as the mapping destination. The information processing device according to claim 1 .

5. The mapping processing unit adjusts the sound pressure and phase of the sound output from the plurality of display units to be mapped for each display unit, thereby controlling the position of the sound image in the depth direction. The information processing device according to claim 3 .

6. The mapping processing unit selects one or more speakers corresponding to the position of a sound source of the audio data from a plurality of speakers arranged around the plurality of display units, and maps waveform data of a low frequency range having the lowest frequency of the audio data and waveform data of a mid-frequency range between the high frequency range and the low frequency range to the one or more selected speakers. The information processing device according to claim 1 .

7. The mapping processing unit generates corrected audio data whose high-frequency sound pressure level is equal to or greater than a threshold from audio data whose high-frequency sound pressure level is smaller than a threshold, and maps waveform data of the high-frequency sound of the corrected audio data onto the one or more display units that are the mapping destinations. The information processing device according to claim 1 .

8. The mapping processing unit outputs the high-frequency waveform data at the same time as the mid-frequency and low-frequency waveform data are output, or earlier than the mid-frequency and low-frequency waveform data are output. The information processing device according to claim 6 .

9. a position detection unit that detects a spatial arrangement of the plurality of display units and assigns a logical number to each display unit based on the spatial arrangement; The mapping processing unit identifies the mapping destination based on the logical number. The information processing device according to claim 1 .

10. The position detection unit detects the spatial arrangement of the display unit based on the time it takes for sound output from the display unit to be transmitted to each of microphones provided at a plurality of locations. The information processing device according to claim 9 .

11. The mapping processing unit controls the directivity of the reproduced sound by adjusting the sound pressure and phase of the sound output from the plurality of display units to which the mapping is applied for each display unit. The information processing device according to claim 1 .

12. The mapping processing unit reduces reverberation by adjusting the sound pressure and phase of the sound output from the plurality of display units to which the mapping is applied for each display unit. The information processing device according to claim 1 .

13. extracting one or more audio data corresponding to different sound sources from the audio content; selecting, for each piece of audio data, one or more display units to which the audio data is to be mapped from a plurality of combinable display units each having a sound generating mechanism; Dividing the audio data into frequency bands; mapping waveform data of the highest frequency band among the audio data onto the one or more display units that are the mapping destinations; 10. A computer-implemented information processing method comprising:

14. extracting one or more audio data corresponding to different sound sources from the audio content; selecting, for each piece of audio data, one or more display units to which the audio data is to be mapped from a plurality of combinable display units each having a sound generating mechanism; Dividing the audio data into frequency bands; mapping waveform data of the highest frequency band among the audio data onto the one or more display units that are the mapping destinations; A program that makes a computer do something.

Citation Information

Patent Citations

  • Information transmission system

    JP2001078282A

  • Video / audio reproducing method for outputting audio from display area of sound source video

    JP2004187288A

  • Display device and sound output method thereof

    JP2008011475A

  • Video sound output apparatus, and video sound output method

    JP2008167032A

  • Three-dimensional sound output device

    JP2011259298A