Information processing device and method, and program

The information processing system addresses inconsistent content presentation in venues by generating and distributing venue-specific corrected video data, ensuring optimal content alignment with facility-specific adjustments for enhanced viewer experience.

WO2025158716A1PCT designated stage expired Publication Date: 2025-07-31SONY MUSIC ENTERTAINMENT (JAPAN) INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/035989
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-22
Filing Date
2024-10-08
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing content distribution systems struggle to appropriately present content in venues with varying facilities, such as planetariums, due to differences in orientation, size, and layout, leading to inconsistent viewer experiences.

Method used

An information processing system that includes a generation unit to process video data based on venue-specific identification information, generating corrected video data for each venue, and an output unit to distribute this data, utilizing planetarium video and audio conversion units to adjust content format and layout for optimal presentation.

Benefits of technology

Ensures consistent and appropriate content presentation across venues with different facilities, enhancing viewer experience by aligning content with venue-specific characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024035989_31072025_PF_FP_ABST
    Figure JP2024035989_31072025_PF_FP_ABST
Patent Text Reader

Abstract

The present technology relates to an information processing device and method, and a program that make it possible to appropriately present content regardless of venue. This information processing device comprises: a correction unit that performs a correction process on content data of content on the basis of venue facility information pertaining to a facility of a venue where the content is played; and an output unit that outputs corrected content data obtained through the correction process. The present technology can be applied to information processing systems.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, method, and program

[0001] The present technology relates to an information processing device, method, and program, and in particular to an information processing device, method, and program that enable content to be presented appropriately regardless of the venue.

[0002] 2. Description of the Related Art Conventionally, a content distribution system is known that distributes video of a live performance at a live venue or in a metaverse space to terminal devices of a plurality of audience members connected via a network.

[0003] As a technology related to such content distribution systems, a technology has been proposed that reduces communication delays throughout the system by transmitting media signals that take into account the context of each remote audience member (see, for example, Patent Document 1).

[0004] International Publication No. 2023 / 120244

[0005] However, when the destination of content data is a venue that can accommodate multiple spectators, such as a planetarium, depending on the venue's facilities, even if the same data is distributed to multiple venues as is, the content may not be presented properly in each venue.

[0006] This technology was developed in light of these circumstances, and makes it possible to present content appropriately regardless of the venue.

[0007] An information processing device according to a first aspect of the present technology includes a generation unit that performs first information processing on first video data constituting content data of the content based on identification information related to the venue where the content is to be played, and generates second video data, and an output unit that outputs the second video data.

[0008] An information processing method or program according to a first aspect of the present technology includes the steps of performing information processing on first video data constituting content data of the content based on identification information related to a venue where the content is to be played, generating second video data, and outputting the second video data.

[0009] In a first aspect of the present technology, information processing is performed on first video data that constitutes content data of the content based on identification information related to the venue where the content is played, second video data is generated, and the second video data is output.

[0010] 1 is a diagram illustrating an example of the configuration of an information processing system. FIG. 1 is a diagram illustrating an example of the configuration of a venue. FIG. 1 is a diagram illustrating types of planetarium venues. FIG. 1 is a diagram illustrating types of planetarium venues. FIG. 1 is a diagram illustrating types of planetarium venues. FIG. 1 is a diagram illustrating content correction. FIG. 1 is a diagram illustrating content correction. FIG. 1 is a diagram illustrating content correction. FIG. 1 is a diagram illustrating content correction. FIG. 1 is a diagram illustrating content video capture. FIG. 1 is a diagram illustrating content video capture. FIG. 1 is a diagram illustrating content video capture. FIG. 1 is a diagram illustrating content video capture. FIG. 1 is a diagram illustrating content audio capture. A flowchart illustrating distribution processing. A flowchart illustrating data correction processing. A flowchart illustrating playback processing. A diagram illustrating an example of displaying audience video in the metaverse space. A diagram illustrating an example of displaying audience video at other venues in the same venue. A diagram illustrating an example of performance in the metaverse space. A diagram illustrating an example of audience seat allocation for each venue in the metaverse space. A diagram illustrating an example of audience display in the metaverse space. A diagram illustrating an example of performance in the metaverse space. A diagram illustrating another example of the configuration of an information processing system. A diagram illustrating another example of the configuration of an information processing system. A diagram illustrating another example of the configuration of an information processing system. A diagram illustrating an example of the configuration of a studio. A diagram illustrating an example of the configuration of a communication / rendering PC. FIG. 1 is a diagram illustrating an example of the configuration of a venue. FIG. 2 is a diagram illustrating an example of the configuration of a communication / rendering PC. FIG. 3 is a diagram illustrating transmission of content data. FIG. 4 is a diagram illustrating an example of a UI display. FIG. 5 is a diagram illustrating an example of a UI display. FIG. 6 is a diagram illustrating an example of a UI display. FIG. 7 is a diagram illustrating an example of a UI display. FIG. 8 is a flowchart illustrating distribution processing. FIG. 9 is a flowchart illustrating venue video display processing. FIG. 10 is a flowchart illustrating playback processing. FIG. 11 is a flowchart illustrating venue video transmission processing. FIG. 12 is a diagram illustrating video clipping of content. FIG. 13 is a diagram illustrating video clipping for each ID. FIG. 14 is a diagram illustrating video clipping for each ID. FIG. 15 is a diagram illustrating video clipping for each ID.It is a figure explaining the cutting out of the video for each item of the playlist.It is a figure explaining the presentation of the sound image of the content.It is a figure explaining the presentation of the sound image of the content.It is a figure explaining the configuration example of the computer.It is a figure explaining the cutting out of the video for each item of the playlist.It is a figure explaining the presentation of the sound image of the content.It is a figure explaining the configuration example of the computer.

[0011] Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.

[0012] First Embodiment Configuration Example of Information Processing System FIG. 1 is a diagram showing a configuration example of an embodiment of an information processing system to which the present technology is applied.

[0013] The information processing system 11 shown in FIG. 1 includes a metaverse login PC (Personal Computer) 21, a cloud 22, and venues 23-1 to 23-N (however, venues 23-3 to 23-(N-1) are not shown).

[0014] In the following description, when there is no need to particularly distinguish between the venues 23-1 to 23-N, they will also be simply referred to as venues 23.

[0015] The information processing system 11 is a content distribution system that distributes content consisting of video and audio of a live performance or the like that is held within the metaverse space from a metaverse login PC 21 to each venue 23 .

[0016] The content to be distributed may be content consisting of only video or audio, or content consisting of video and accompanying audio, but the following description will be given assuming that content consisting of video and audio is being distributed. In particular, the content here is assumed to be live content obtained by capturing a live performance taking place in the metaverse space (a virtual three-dimensional space).

[0017] The metaverse login PC 21 distributes content data, which is data for reproducing (presenting) content, to each venue 23 via the cloud 22 .

[0018] The content data includes at least video data and sound image data (audio data) for playing back the content. For example, the video data constituting the content data may be data of a celestial sphere video (a celestial sphere image) that is a video (image) in all directions (up, down, left, and right). Note that the content data may also include haptic data for providing haptic sensations.

[0019] The cloud 22 is made up of, for example, one or more information processing devices on a network. The cloud 22 corrects the content data transmitted (distributed) from the metaverse login PC 21 according to the facilities of each venue 23, and transmits the corrected content data to each venue 23.

[0020] The venue 23 is a planetarium (planetarium venue) that can accommodate multiple spectators and has a planetarium apparatus 24 consisting of one or more devices for presenting content. As will be described later, each venue 23 has different facilities, such as the orientation and angle of the seats, and the orientation and size of the dome portion that serves as the screen (display unit).

[0021] Here, we will explain an example in which venue 23 is a planetarium, but venue 23 may be any facility that can present content to multiple spectators, such as a facility with a dome-shaped display unit.

[0022] Each venue 23 (planetarium device 24) receives the content data transmitted from the cloud 22 and presents the content to one or more audience members present within the venue 23 based on the content data.

[0023] The metaverse login PC 21 includes a metaverse space capture unit 31 , a spatial audio capture unit 32 , a video and audio image distribution unit 33 , an interaction information communication unit 34 , and an interaction feedback unit 35 .

[0024] The metaverse space capture unit 31 acquires (captures) video data for playing back video of the metaverse space to which the metaverse login PC 21 is logged in.

[0025] For example, the metaverse space capture unit 31 can place a virtual camera (hereinafter referred to as a virtual camera) at any position in the metaverse space.

[0026] The metaverse space capture unit 31 can appropriately connect to a server (not shown) that provides services related to the metaverse space (hereinafter also referred to as the metaverse server) and, by specifying the placement position of the virtual camera, etc., obtain video data from the metaverse server for reproducing the appearance of the metaverse space when the position and orientation of the virtual camera are the viewpoint position and line of sight direction.

[0027] For example, the metaverse server has CG (Computer Graphics) data and the like prepared in advance for generating images of the metaverse space, and the metaverse server generates image data based on the position and orientation of a specified virtual camera. However, if the virtual camera is a camera capable of capturing images in all directions (spherical images) up, down, left, and right, it is not necessarily necessary to specify the orientation of the virtual camera.

[0028] In addition, CG data, etc. may be supplied from the metaverse server to the metaverse space capture unit 31, and the metaverse space capture unit 31 may generate video data based on the CG data, etc. and the position and orientation of the virtual camera.

[0029] The metaverse space capture unit 31 acquires (captures) video data for the positions of one or more virtual cameras, and supplies the acquired video data to the video and audio image distribution unit 33 .

[0030] At this time, the metaverse space capture unit 31 can also add, as appropriate, information such as the viewpoint position and line of sight of the video data, i.e., coordinates indicating the placement position and orientation of the virtual camera in the metaverse space, as camera position and direction information to the video data. In other words, the video data to which the camera position and direction information has been added may be supplied to the video and audio image distribution unit 33.

[0031] The spatial sound capture unit 32 acquires (captures) sound image data for reproducing voice (sound) in the metaverse space to which the metaverse login PC 21 is logged in.

[0032] For example, the spatial sound capture unit 32 can place a virtual microphone (hereinafter referred to as a virtual microphone) at any position in the metaverse space in any orientation. Basically, the placement position and orientation of the virtual microphone are the same as the placement position and orientation of the virtual camera. Note that for the virtual microphone, only the placement position may be specified, and the orientation of the virtual microphone may not be specified.

[0033] The spatial sound capture unit 32 can connect to the metaverse server as appropriate and, by specifying the placement position of the virtual microphone, etc., obtain sound image data from the metaverse server for reproducing sound in the metaverse space when the position of the virtual microphone is the listening position.

[0034] In this case, for example, the metaverse server stores sound image data (hereinafter also referred to as sound image source data) for generating sounds to be heard at each position and orientation in the metaverse space, and the metaverse server generates sound image data for each virtual microphone position based on the position and orientation of the specified virtual microphone. Note that sound image data may be prepared in advance for each position and orientation in the metaverse space, or sound image data may be generated on the spatial sound capture unit 32 side as appropriate.

[0035] The spatial audio capture unit 32 acquires (captures) sound image data for the positions of one or more virtual microphones, and supplies the acquired sound image data to the video and audio image distribution unit 33 .

[0036] At this time, the spatial audio capture unit 32 can also appropriately add information such as coordinates indicating the placement position and orientation of the virtual microphone of the sound image data as microphone position / orientation information to the sound image data. In other words, the sound image data to which the microphone position / orientation information has been added may be supplied to the video and audio image distribution unit 33. As described above, the placement positions and orientations of the virtual camera and virtual microphone are basically the same, so the camera position / orientation information and the microphone position / orientation information are the same information.

[0037] The sound image data is composed of data for each speaker, i.e., data for each channel corresponding to a speaker, for reproducing sound in a speaker system consisting of multiple speakers. Specifically, the sound image data is multi-channel data (channel-based audio data) such as 5.1 channels.

[0038] The sound image data may be data for each object (audio object) such as an avatar in the Metaverse space, i.e., object-based audio data for playing the sound of the object. In such a case, for example, object position information indicating the position of the object in the Metaverse space may be added to the audio data for each object as sound image data.

[0039] The video and audio image distribution unit 33 uses the data consisting of the video data supplied from the metaverse space capture unit 31 and the audio image data supplied from the spatial audio capture unit 32 as content data for reproducing the live performance taking place in the metaverse space. In this case, content data is obtained for one or more positions in the metaverse space, i.e., for each position of the virtual camera.

[0040] The video and audio image distribution unit 33 transmits the content data to the cloud 22 via a network or the like. In other words, the content is distributed.

[0041] It should be noted that the video and sound image distribution unit 33 may also prevent sound image data from being included in the content data. For example, if the spatial audio capture unit 32 can only obtain stereo sound image data and the cloud 22 has previously stored sound image data and sound image source data for each position and orientation in the metaverse space, the content data may not include sound image data.

[0042] In response to the distribution of content, each venue 23 outputs information about the state of the venue 23, such as the audience's reaction to the content presentation, as interaction information. That is, the interaction information transmitted from the venue 23 is supplied to the metaverse login PC 21 via the cloud 22. As an example, the interaction information may include video data for playing back a video of the audience seats in the venue 23 and venue identification information that uniquely identifies the venue 23.

[0043] The interaction information communication unit 34 receives interaction information transmitted from the cloud 22 and supplies it to the interaction feedback unit 35. The interaction information communication unit 34 also transmits to the cloud 22 reflection information supplied from the interaction feedback unit 35 in accordance with the interaction information.

[0044] The reflection information is information for reflecting in the content the effects and the like according to the interaction information, and may be information including, for example, venue identification information for one or more venues 23 and video data for playing back video of the audience seats in the venues 23. The reflection information may also include information indicating the display position of the video of the audience seats for each venue 23.

[0045] The interaction feedback unit 35 performs feedback processing in accordance with the interaction information supplied from the interaction information communication unit 34 to reflect the reactions to the content at each venue 23 in the content through production or the like.

[0046] For example, the interaction feedback unit 35 performs feedback processing such as generating reflection information according to the interaction information and supplying it to the interaction information communication unit 34, or requesting the metaverse server to execute processing according to the interaction information.

[0047] The cloud 22 has an image / sound image receiving unit 41, a spatial audio output unit 42, an image / sound image distribution unit 43, planetarium image / sound image conversion units 44-1 to 44-M, image / sound image distribution units 45-1 to 45-M, a feedback communication unit 46, and a feedback processing unit 47.

[0048] In FIG. 1, the planetarium image / sound conversion unit 44-2 to the planetarium image / sound conversion unit 44-(M-1) and the image / sound distribution unit 45-2 to the image / sound distribution unit 45-(M-1) are not shown.

[0049] Hereinafter, when there is no need to particularly distinguish between the planetarium image-sound-image conversion units 44-1 to 44-M, they will also be simply referred to as planetarium image-sound-image conversion units 44, and when there is no need to particularly distinguish between the image-sound-image distribution units 45-1 to 45-M, they will also be simply referred to as image-sound-image distribution units 45.

[0050] The video / audio image receiving unit 41 functions as a receiving unit that receives content data transmitted from the video / audio image distribution unit 33 of the metaverse login PC 21 .

[0051] That is, the video and audio image receiving unit 41 receives the content data transmitted from the metaverse login PC 21 and supplies it to the video and audio image distributing unit 43 .

[0052] Furthermore, if the received content data does not contain sound image data or if stereo sound image data is included, the video / sound image receiving unit 41 supplies the camera position / direction information added to the video data constituting the content data to the spatial audio output unit 42.

[0053] The spatial audio output unit 42 stores pre-prepared sound image source data for generating sounds to be heard at various positions and directions in the metaverse space, or multi-channel sound image data prepared for each position and direction in the metaverse space, of sounds to be heard at those positions and directions. This sound image source data and multi-channel sound image data are pre-recorded.

[0054] The spatial audio output unit 42 uses the camera position and direction information supplied from the video and audio image receiving unit 41 as microphone position and direction information, and supplies multi-channel sound image data with a position and direction corresponding to the microphone position and direction information to the video and audio image distribution unit 43.

[0055] Furthermore, for example, if the content data includes stereo sound image data, the spatial audio output unit 42 may acquire the stereo sound image data from the video / audio image receiving unit 41 and generate multi-channel sound image data based on the acquired sound image data.

[0056] As an example, the spatial audio output unit 42 extracts sound image data for each sound source (audio object) by performing processing such as sound source separation based on stereo sound image data. The spatial audio output unit 42 then generates multi-channel sound image data for stereophonic playback by performing any rendering processing such as VBAP (Vector Based Amplitude Panning) based on the extracted sound image data for each sound source and microphone position and direction information, and supplies the generated data to the video and audio image distribution unit 43. Note that the position of each sound source within the space (metaverse space) may be used during the rendering processing. In such cases, the position (direction) of each sound source may be determined based on, for example, the results of processing such as sound source separation or the results of image recognition based on video data included in the content data. Alternatively, information indicating the position of each sound source may be acquired by some other method, such as by adding information indicating the position of each sound source to the stereo sound image data.

[0057] The video / audio image distribution unit 43 distributes (supplies) the content data supplied from the video / audio image receiving unit 41 to each planetarium video / audio image conversion unit 44 .

[0058] In this case, when sound image data is supplied from the spatial audio output unit 42, the video / sound image distribution unit 43 supplies data consisting of the sound image data and the video data that constitutes the content data supplied from the video / sound image receiving unit 41 to the planetarium video / sound image conversion unit 44 as final content data.

[0059] The planetarium video / audio conversion unit 44 stores venue facility information relating to the facilities of the venue 23 where the content is played back, and functions as a correction unit that performs content correction processing on the content data based on the venue facility information. Note that the cloud 22 may receive the venue facility information from each venue 23.

[0060] For example, the venue facility information includes at least one of information about the video projection equipment in venue 23, information about audio equipment such as a speaker system installed in venue 23, information about the seats (audience seats) in venue 23, and information about the dome portion that serves as a display unit (presentation unit) on which the video of the content is displayed in venue 23. The venue facility information may also include venue identification information that indicates which venue 23 the venue facility information belongs to.

[0061] For example, the information about the projection equipment may be information such as the installation location of the video projection equipment in the venue 23, the video formats that can be handled by the video projection equipment, etc. Furthermore, for example, the information about the sound equipment may be information indicating the channel configuration of the speaker system installed in the venue 23, the placement positions of each speaker that makes up the speaker system, etc.

[0062] For example, information about the seats in the venue 23 may be information such as the layout type of the seats in the venue 23, the angle of the seats relative to the horizontal plane or the dome part, etc. Furthermore, information about the dome part that serves as the display unit may be information such as the radius (size) of the dome part, the tilt angle which is the angle (inclination) of the dome part relative to the horizontal plane (ground plane) or the seats, and the position of the image center in the dome part.

[0063] In the information processing system 11, since the equipment installed in each venue 23 differs, a planetarium image and sound image converter 44 is provided for each of these different pieces of equipment. In particular, in this example, since there are M types of equipment in the venue 23, M planetarium image and sound image converters 44 are provided, and each planetarium image and sound image converter 44 holds different venue equipment information.

[0064] When distributing content data by the video / audio-image distribution unit 43, the video / audio-image distribution unit 43 may refer to the venue facility information stored in each planetarium video / audio-image conversion unit 44 as appropriate. This allows content data such as the position of a virtual camera corresponding to a specific dome radius to be supplied (distributed) to the planetarium video / audio-image conversion unit 44 that stores the venue facility information for the venue 23 having that specific dome radius.

[0065] The planetarium video / audio image conversion unit 44 performs content correction processing on the content data supplied from the video / audio image distribution unit 43 in accordance with the venue equipment information stored in advance, and supplies the resulting content data (hereinafter also referred to as corrected content data) to the video / audio image distribution unit 45.

[0066] The content correction process is a process of applying corrections to content data so that the content can be appropriately played (presented) on the equipment in the venue 23. In other words, the content correction process can be said to be a process of generating corrected content data from content data by converting the video format or the like, that is, a process of converting content data into corrected content data.

[0067] Specifically, for example, content correction processing involves converting the video format of video data and converting the channel configuration of sound image data, such as downmixing or upmixing.

[0068] In addition, when the sound image data constituting the content data is object-based audio data, rendering processing such as VBAP based on venue facility information, information indicating the position of the object, and sound image data may be performed as content correction processing.

[0069] In addition, the planetarium image / sound image conversion unit 44 performs processing such as embedding (synthesis processing) images based on reflection information and interaction information into images based on the video data of the content, as necessary, under the control of the feedback processing unit 47.

[0070] The video and sound image distribution unit 45 transmits (distributes) the corrected content data supplied from the planetarium video and sound image conversion unit 44 to one or more venues 23 that have equipment corresponding to the venue equipment information held by the supplying planetarium video and sound image conversion unit 44. In other words, the video and sound image distribution unit 45 functions as an output unit that outputs (transmits) the corrected content data to the venues 23. Note that the venues 23 to which the corrected content data is sent differ for each video and sound image distribution unit 45.

[0071] 1, the video and audio image distribution unit 45-1 distributes corrected content data to the venues 23-1 and 23-2. By each video and audio image distribution unit 45 transmitting the corrected content data, content is distributed to multiple venues 23 simultaneously.

[0072] The feedback communication unit 46 functions as a communication unit that transmits and receives interaction information and the like.

[0073] That is, the feedback communication unit 46 receives interaction information transmitted from each venue 23 and transmits it to the metaverse login PC 21, and receives reflection information transmitted from the metaverse login PC 21. The feedback communication unit 46 also supplies the received reflection information and interaction information to the feedback processing unit 47 as necessary.

[0074] The feedback processing unit 47 controls the planetarium image and sound image conversion unit 44 based on the reflection information and interaction information supplied from the feedback communication unit 46, and reflects the reactions to the content in each venue 23 in the content through production etc. In other words, the feedback processing unit 47 causes the planetarium image and sound image conversion unit 44 to execute processing described below.

[0075] When the feedback processing unit 47 controls the planetarium image and sound conversion units 44, the feedback processing unit 47 may refer to the venue facility information stored in each planetarium image and sound conversion unit 44 as appropriate in order to identify the planetarium image and sound conversion unit 44 to be controlled.

[0076] The venue 23, which is a planetarium venue, has a planetarium device 24 as a facility for presenting the distributed content to the audience.

[0077] The planetarium device 24 includes an image and sound receiving unit 51 , an image projection unit 52 , a spatial audio output unit 53 , a scene analysis unit 54 , and a feedback communication unit 55 .

[0078] The video and sound image receiving unit 51 receives the corrected content data transmitted from the video and sound image distribution unit 45 of the cloud 22. The video and sound image receiving unit 51 also extracts video data and sound image data from the received corrected content data, and supplies the video data to the video projection unit 52 and the sound image data to the spatial audio output unit 53.

[0079] The video projection unit 52 is made up of video projection equipment such as a projector for displaying (presenting) the video of the content, and displays the video of the content based on the video data supplied from the video and sound image receiving unit 51. The spatial audio output unit 53 is made up of, for example, a multi-channel speaker system, and plays (outputs) the sound of the content based on the sound image data supplied from the video and sound image receiving unit 51.

[0080] The video projection unit 52 and the spatial audio output unit 53 function as a playback unit that plays back content based on the corrected content data.

[0081] The scene analysis unit 54 is made up of, for example, a camera that captures images of the audience seats in the venue 23, a microphone that picks up the sounds of the audience in the venue 23, a processor, and the like.

[0082] The scene analysis unit 54 generates interaction information by taking photographs and collecting audio of the audience seats and other areas of the venue 23, and performing information processing (analysis processing) such as scene analysis of the video and audio of the audience seats and other areas of the venue 23, and supplies the interaction information to the feedback communication unit 55.

[0083] The feedback communication unit 55 transmits the interaction information supplied from the scene analysis unit 54 to the cloud 22 .

[0084] <Example of venue configuration> In FIG. 1, the information processing system 11 for presenting content is configured by the metaverse login PC 21, the cloud 22, and the venue 23, but the content may also be presented by an on-premise planetarium projection system.

[0085] In such a case, the venue 23 may be configured as shown in Figure 2. In Figure 2, the same reference numerals are used to designate parts that correspond to those in Figure 1, and the description thereof will be omitted where appropriate.

[0086] The venue 23 shown in FIG. 2, that is, the planetarium device 24, has a metaverse login PC 21, a video projection unit 52, a spatial audio output unit 53, and a scene analysis unit .

[0087] The metaverse login PC 21 also has a metaverse space capture unit 31 , a spatial sound capture unit 32 , a spatial sound output unit 42 , a planetarium video / sound image conversion unit 44 , a feedback communication unit 46 , and an interaction feedback unit 35 .

[0088] In this example, the metaverse login PC 21 has the functions of the metaverse login PC 21 in FIG. 1 as well as the functions of the cloud 22 shown in FIG.

[0089] 2, the content data obtained by capture in the metaverse space capture unit 31 and the spatial audio capture unit 32 is subjected to content correction processing in the planetarium image and audio conversion unit 44. The video data and audio data constituting the corrected content data obtained by the content correction processing are supplied to the video projection unit 52 and the spatial audio output unit 53, and the content is played (presented).

[0090] Furthermore, the interaction information obtained by the scene analysis unit 54 is supplied to the interaction feedback unit 35 via the feedback communication unit 46, and the interaction feedback unit 35 performs feedback processing based on the interaction information. In this example, as the feedback processing, in addition to the processing performed by the interaction feedback unit 35 of the information processing system 11 shown in Figure 1, processing performed by the feedback processing unit 47 in Figure 1 is also executed.

[0091] 1 may be performed at the venue 23, and a metaverse login PC 21 may be provided separately from the venue 23. In such a case, for example, the venue 23 may communicate with the metaverse login PC 21 and perform processes such as receiving content data as appropriate.

[0092] <Differences in venue facilities and processing by each unit of the information processing system> Next, differences in facilities for each venue 23 shown in FIG. 1 and processing performed by each unit of the information processing system 11 will be described.

[0093] First, differences in facilities for each venue 23 will be described.

[0094] In the planetarium venue 23, the inside (internal side) of the dome portion functions as a display section (screen) that displays images of the content, but the size of this dome portion may differ from venue to venue 23.

[0095] For example, there is a venue 23 where the radius of the dome part is 5 m as indicated by arrow Q11 in FIG. 3, and there is also a venue 23 where the radius of the dome part is 35 m as indicated by arrow Q12.

[0096] Therefore, the venue facility information held by the planetarium video / audio image conversion unit 44 includes information indicating the radius of the dome portion (hereinafter also referred to as the dome radius), and the content correction process adjusts (corrects) the difference in dome radius for each venue 23.

[0097] Furthermore, as shown in FIG. 4, for example, the inclination of the dome portion that functions as a display section, that is, the inclination angle of the dome portion with respect to the horizontal plane (ground plane), may differ for each venue 23.

[0098] Specifically, in the example shown by arrow Q21, the hemispherical dome portion is positioned so that it is parallel to the horizontal plane. In other words, the dome portion is positioned so that the plane of the hemisphere in the dome portion is parallel to the horizontal plane (parallel to the ground).

[0099] Hereinafter, the dome arrangement indicated by arrow Q21 will be referred to as a horizontal arrangement, i.e., a horizontal dome. In a horizontal dome arrangement, the hemispherical plane of the dome may be parallel to the plane on which the spectator seats are installed, or the hemispherical plane of the dome may be inclined relative to the plane on which the seats are installed.

[0100] In the example shown by arrow Q22, the hemispherical dome portion is positioned so that it is tilted at a predetermined angle relative to the horizontal plane, i.e., the dome portion is positioned so that the plane of the hemisphere in the dome portion forms a predetermined angle with the horizontal plane (ground).

[0101] Hereinafter, the dome portion arrangement indicated by arrow Q22 will be referred to as a tilted arrangement, i.e., a tilted dome. With a tilted arrangement, it is possible to display images not only in the portion above the center of the dome, but also in the portion below the center of the dome, i.e., the portion indicated by arrow A11.

[0102] The venue equipment information includes angle information indicating the inclination (tilt angle) of the dome portion, and the content correction process performs correction according to the inclination of the dome portion, i.e., differences in the inclination of the dome portion, such as whether it is horizontal or inclined.

[0103] Furthermore, as shown in Fig. 5, for example, the seating arrangement (seat layout type) may differ for each venue 23. Note that Fig. 5 shows the area in which the seats are arranged in the venue 23 as viewed from above.

[0104] For example, in the example shown by arrow Q31, the seats are arranged concentrically around the center of the venue 23 (hereinafter also referred to as the venue center) so that each seat faces the center. Hereinafter, the seat arrangement shown by arrow Q31 will also be referred to as the concentric arrangement.

[0105] In the example shown by arrow Q32, the seats are arranged in a fan shape so that they face a specific direction (forward) in the venue 23. In particular, in this example, the upper side in the drawing is the front, and the seats are arranged facing forward. Hereinafter, the seat arrangement shown by arrow Q32 will also be referred to as a fan-shaped arrangement.

[0106] The venue facility information includes information indicating the seating layout type, such as a concentric circle arrangement or a fan-shaped arrangement, and the content correction process performs correction according to the difference in layout type.

[0107] Next, an example of content correction processing based on venue facility information, which is performed in the planetarium image / sound image conversion unit 44, will be described.

[0108] For example, the video formats that can be handled (processed) may differ depending on the video projection unit 52 installed in the venue 23 .

[0109] For example, there are various video formats, such as the dome master format, the equirectangular format, and other formats that allow unique transmission.

[0110] The venue facility information includes information indicating the video format, such as a dome master format, used by the video projection unit 52, and the content correction process corrects the content data according to the video format. That is, the content correction process converts the video format of the video data that constitutes the content data into the video format indicated by the venue facility information.

[0111] If the seat layout type differs depending on the venue 23, the position of the image center in the dome portion will also differ, as shown in FIG. 6, for example.

[0112] In Figure 6, the part indicated by arrow Q41, i.e., the upper part of the figure, shows the fan-shaped seating area of ​​venue 23 as seen from above on the left, and the fan-shaped seating area of ​​venue 23 as seen from directly to the side on the right.

[0113] Also, in the part indicated by arrow Q42, i.e., the lower part of the figure, the left side shows the seating area of ​​venue 23 arranged in a concentric circle as seen from above, and the right side shows the seating area of ​​venue 23 arranged in a concentric circle as seen from the side.

[0114] For example, as indicated by arrow Q41, when the layout type is a fan-shaped arrangement, the seats of all spectators in the venue 23 face a predetermined position in the dome portion serving as the display unit, that is, the front portion of the venue 23.

[0115] Therefore, it is possible to position the center of the image, that is, the center position of the content image, at the front of the venue 23 so that all spectators can see the center of the image. In other words, even if one image is displayed in the dome part, all spectators can view the image from the appropriate direction.

[0116] In contrast, as shown by arrow Q42, when the layout type is concentric, the direction of the audience in each seat is different. In other words, the audience will be looking at a position diagonally above the seat in the dome. Therefore, the audience's viewing position (image) will be different for each seat.

[0117] Therefore, when distributing content to the concentrically arranged venue 23, it is necessary to process (correct) the image of the content so that all spectators can see the center of the image.

[0118] For example, assume that the video data constituting the content data supplied (transmitted) from the metaverse login PC 21 to the cloud 22 is video data of a celestial sphere, which is a video of all directions, up, down, left, and right. The celestial sphere video can be said to be a video of the entire inner surface of a sphere.

[0119] In addition, in each venue 23, a full-dome image, which is a hemispherical image of the upper half of the full-dome image, is displayed on a hemispherical dome portion.

[0120] In such a case, the content correction process for generating corrected content data for the concentrically arranged venues 23 may be performed by carrying out the processes shown in FIG. 7 or FIG.

[0121] The content correction process shown in FIGS. 7 and 8 is a process of arranging a plurality of images (cut-out images) cut out from a celestial sphere image based on video data constituting the content data adjacent to each other (arranging them in a plurality of locations) and synthesizing them in accordance with the seat layout type of venue 23 indicated by the venue facility information.

[0122] In the example shown in FIG. 7, a quarter area of ​​the full-dome image is cut out, and the cut-out image is projected in front of the seats in each of the four divided areas of the dome.

[0123] That is, as shown on the left side of Figure 7, the image obtained by cutting out a quarter area W11 in the front part of the full-dome image of the content to be displayed in the fan-shaped venue 23 will be called the cut-out image.

[0124] Also, as shown on the right side of the figure, let us assume that the hemispherical dome portion of venue 23, which is arranged in concentric circles, is divided into four areas, areas W21 to W24. In this case, content correction processing is performed so that a cut-out video is displayed in each of areas W21 to W24. In other words, content correction processing is performed by arranging (side-by-side) one cut-out video in each of the four areas and combining them to generate a single full-dome video.

[0125] This allows each spectator in the concentrically arranged venue 23 to see the same image, particularly the same center of the image.

[0126] In the example shown in FIG. 8, half of the full-dome image is cut out, and the cut-out image is projected in front of the seats in each of the two divided areas of the dome.

[0127] That is, as shown on the left side of FIG. 8, of the full-dome video of the content to be displayed in the fan-shaped venue 23, a half (1 / 2) area W41 in the front part is cut out and used as the cut-out video.

[0128] Furthermore, as shown on the right side of the figure, the hemispherical dome portion of the concentrically arranged venue 23 is divided into two areas, area W51 and area W52, and content correction processing is performed so that a cut-out image is displayed in each of areas W51 and W52. In other words, the content correction processing involves placing one cut-out image in each of the two areas and combining them to generate a single full-dome image. This allows each spectator in the concentrically arranged venue 23 to see the same image, particularly the same center of the image, as in the example of Figure 7.

[0129] Furthermore, as described above, the inclination of the dome portion differs for each venue 23. That is, some venues 23 are horizontal, while others are inclined, and the image center needs to be changed depending on the inclination angle of the dome portion of the venue 23.

[0130] Therefore, as shown in FIG. 9 , for example, content correction processing is performed in which a part of the omnidirectional image based on the video data constituting the content data is cut out as a omnidirectional image based on the inclination (tilt angle) of the dome portion indicated by the venue facility information.

[0131] In FIG. 9, a spherical image D11 can be displayed based on the video data that constitutes the content data supplied to the cloud 22.

[0132] For example, as shown in the upper right of the figure, it is assumed that the venue 23 is identified as a horizontal venue 23 from the venue facility information. In such a case, in the content correction process, the upper half of the omnidirectional image D11 in the figure is cut out to create a cut-out image that is a omnidirectional image.

[0133] Then, the video data for displaying the cut-out video is directly used as the video data that constitutes the corrected content data, or the processing described with reference to Figures 7 and 8 is performed based on the cut-out video to generate the video data that constitutes the corrected content data.

[0134] In contrast to this, for example, as shown in the lower right of the figure, it is assumed that the venue 23 is identified as an inclined venue 23 from the venue facility information. In such a case, in the content correction process, an area corresponding to the inclined hemisphere is cut out from the omnidirectional image D11 to form a cutout image that is a omnidirectional image. In this example, the upper left half of the omnidirectional image D11 is cut out as the cutout image in accordance with the inclination of the dome portion of the venue 23.

[0135] Then, the video data for displaying the cut-out video is directly used as the video data that constitutes the corrected content data, or the processing described with reference to Figures 7 and 8 is performed based on the cut-out video to generate the video data that constitutes the corrected content data.

[0136] It is possible that the metaverse login PC 21 captures a spherical image as an image of the metaverse space at a higher resolution than the resolution of an image that can be projected by the image projection unit 52 of the venue 23. For example, if the image projection unit 52 is capable of projecting (displaying) an HD (High Definition) image, the metaverse login PC 21 may capture a 4K spherical image.

[0137] As described with reference to FIG. 9 , by extracting the panoramic image obtained in this manner at an optimal angle for each venue 23 according to the inclination angle of the dome portion of venue 23, it is possible to obtain panoramic images for each of multiple venues 23 from a single panoramic image.

[0138] Next, the capturing of images and sounds in the metaverse space performed by the metaverse login PC 21 will be described.

[0139] The appearance of the content video differs depending on the size (dome radius) of the dome part of the venue 23. In particular, the size of people in the video changes significantly depending on the dome radius.

[0140] Specifically, for example, as shown in Figure 10, in a venue 23 with a dome radius of 5 m, person H11 in the content image is displayed at the size indicated by arrow Z11, and in a venue 23 with a dome radius of 35 m, person H11 is displayed at the size indicated by arrow Z12.

[0141] In this case, for example, person H11 will be displayed at an appropriate size in venue 23 where the dome radius is 5 m, but person H11 will be too large in venue 23 where the dome radius is 35 m, and the appearance of the subject (object) will change depending on the dome radius.

[0142] Therefore, taking advantage of the fact that the content is an image of the Metaverse space, it is conceivable to place a virtual camera at a location that provides an angle that is appropriate for the size of the dome part of each venue 23 (dome radius).

[0143] In such a case, the dome radius of each venue 23, i.e., the dome radius required by the information processing system 11, is determined in advance by some method, and information indicating the required dome radius is stored in the metaverse space capture unit 31 of the metaverse login PC 21.

[0144] The metaverse space capture unit 31 then controls the position of the virtual camera based on one or more dome radii, and captures (photographs) images from the required angles for each venue 23 .

[0145] Specifically, for example, suppose that a virtual person H11, such as an avatar, is present in the metaverse space as the subject, as shown in Fig. 11. Also, suppose that the information processing system 11 includes a venue 23 with a dome radius of 5 m and a venue 23 with a dome radius of 35 m.

[0146] In such a case, for example, the metaverse space capture unit 31 places a virtual camera at position P11 in the metaverse space to capture, and uses the resulting video data as video data for the venue 23, which has a dome radius of 5 m.

[0147] Similarly, for example, the metaverse space capture unit 31 places a virtual camera at position P12 in the metaverse space, which is farther from person H11 than position P11, and performs capture, and the resulting video data is used as video data for venue 23, which has a dome radius of 35 m.

[0148] The placement position of the virtual camera may be determined based on a specific position, such as the position of a live performer in the metaverse space or the position of the stage where the live performance is taking place.

[0149] However, the reference position for determining the position of the virtual camera may change, for example, when a stage in the metaverse space is destroyed and a new stage is set up (created), when live performers move within the metaverse space, or when the space that is the subject of the content changes over time.

[0150] In such a case, it is possible to dynamically change the reference position for determining the placement position of the virtual camera, as shown in Figures 12 and 13. This allows the position of each of the multiple virtual cameras to be dynamically controlled, making it possible to obtain images from the required angles.

[0151] In FIG. 12, the shooting angle is controlled in association with the timeline of the live performance.

[0152] That is, in the example of FIG. 12, there are a plurality of areas, including one area E11 and another area E12, as areas such as live music venues in the metaverse space.

[0153] For example, suppose that from a predetermined time t=a, an image in area E11 is shot as a content image, and from a subsequent time t=b, an image in area E12 is shot as a content image.

[0154] In this case, the metaverse space capture unit 31 pre-stores, in addition to information indicating one or more dome radii, timeline information indicating each time (content playback time) and the reference position (reference position) for shooting at each of those times.

[0155] Then, the metaverse space capture unit 31 determines the placement position of the virtual camera based on the timeline information and information indicating the dome radius, and captures the video.

[0156] In the example shown in FIG. 13, when a person H11, such as a live performer, moves within the metaverse space, the virtual camera tracks the person H11.

[0157] In this example, the placement positions of one or more virtual cameras are determined according to the dome radius and the position of person H11. Here, virtual camera C11 is placed at a position a predetermined distance away from person H11, and another virtual camera C12 is placed at a position farther away than virtual camera C11 as viewed from person H11.

[0158] For example, the distance from the person H11 to the virtual camera C11 is set to distance CL11, and the distance from the person H11 to the virtual camera C12 is set to distance CL12.

[0159] In this state, when the person H11 moves, the virtual camera C11 basically moves (tracks) along with the person H11 while maintaining the distance CL11, and the virtual camera C12 also moves along with the person H11 while maintaining the distance CL12.

[0160] However, a certain dead zone is provided in the tracking of virtual camera C11 and virtual camera C12. That is, virtual camera C11 does not move until the distance to person H11 becomes a predetermined distance greater than distance CL11, and starts moving when the distance to person H11 exceeds the predetermined distance. Similarly, virtual camera C12 also starts moving when the distance to person H11 becomes a predetermined distance greater than distance CL12.

[0161] As described above, the spatial sound capture unit 32 can capture sounds in the metaverse space by placing virtual microphones in the metaverse space and reproduce the sounds (sound images) of each sound source in the metaverse space. In this case, the virtual microphones are basically placed in the same positions as the virtual cameras.

[0162] For example, as shown in Figure 14, suppose that a spectator U11, i.e., a virtual camera, is placed at a predetermined position in the metaverse space, and a person H31, such as a live performer, is located directly in front of the spectator U11, and a person H32 is located diagonally forward and to the left of the spectator U11.

[0163] In such a case, when the sound of the content is played back based on the multi-channel sound image data obtained by capture, the sound of person H31 will be heard from directly in front of the audience, and the sound of person H32 will be heard from diagonally forward and to the left of the audience.

[0164] In each venue 23, a plurality of speakers constituting a speaker system as a spatial sound output unit 53 are arranged to surround the audience.

[0165] For example, multiple speakers are arranged in a ring along the dome at multiple different heights, such as at approximately the same height as the audience's ears, at a higher height than the audience's ears, or near the top of the dome.

[0166] Depending on the metaverse platform, it may not be possible to output the sound of the metaverse space in multi-channel. In such a case, if the equipment of venue 23 (spatial audio output unit 53), i.e., the playback environment of venue 23, is compatible with multi-channel, a stereophonic space is created separately from the metaverse space, and sound (sound image data) is acquired from the same coordinate points in that space.

[0167] That is, as described above, sound image source data for generating sounds to be heard at each position and orientation in the metaverse space, which has been obtained by pre-recording, or multi-channel sound image data for each position and orientation in the metaverse space, is held in the spatial audio output unit 42. Then, multi-channel sound image data according to the position of the virtual microphone (virtual camera) is supplied from the spatial audio output unit 42 to the video and audio image distribution unit 43.

[0168] As explained with reference to Figures 10 to 14, by determining the positions of the virtual camera and virtual microphone according to the dome radius of venue 23, etc., the corrected content data transmitted from cloud 22 to venue 23 will have different viewpoint positions, etc. depending on the size (radius) of the dome part of venue 23.

[0169] <Description of Distribution Processing> Next, the operation of the information processing system 11 will be described.

[0170] First, the distribution process performed by the metaverse login PC 21 will be described with reference to the flowchart of FIG.

[0171] In step S11, the metaverse login PC 21 captures video and audio images of the metaverse space.

[0172] For example, the metaverse space capture unit 31 determines the placement position and orientation of the virtual camera for each dome radius based on information indicating the dome radius stored in advance. At this time, timeline information or information indicating the position of an object (subject) of interest, such as a performer, in the metaverse space may be used, as described with reference to FIG.

[0173] The metaverse space capture unit 31 transmits the positions and orientations of one or more virtual cameras to the metaverse server, thereby acquiring video data for each position and orientation of the virtual cameras from the metaverse server. The metaverse space capture unit 31 adds, as appropriate, information indicating camera position and direction information and dome radius to the acquired video data for each virtual camera, and supplies the video and audio image distribution unit 33 with the video data.

[0174] Furthermore, for example, the spatial audio capture unit 32 sets the position and orientation of a virtual microphone to the same position and orientation as the virtual camera, and transmits the positions and orientations of one or more virtual microphones to the metaverse server, thereby acquiring sound image data for each position and orientation of the virtual microphone from the metaverse server. The spatial audio capture unit 32 adds microphone position and orientation information to the acquired sound image data for each virtual microphone as appropriate, and supplies the sound image data to the video and audio image distribution unit 33.

[0175] In step S12, the video and audio image distribution unit 33 distributes the content data.

[0176] That is, the video and audio image distribution unit 33 converts data consisting of the video data supplied from the metaverse space capture unit 31 and the audio image data supplied from the spatial audio capture unit 32 into content data, and transmits the content data to the cloud 22.

[0177] When the content data is distributed (transmitted), interaction information is transmitted from each venue 23 via the cloud 22 .

[0178] In step S13 , the interaction information communication unit 34 receives the interaction information transmitted from the cloud 22 and supplies it to the interaction feedback unit 35 .

[0179] In step S14, the interaction feedback section 35 performs processing according to the interaction information supplied from the interaction information communication section 34.

[0180] As an example, the interaction feedback unit 35 identifies the level of excitement at each venue 23 from the interaction information for that venue 23, and requests the metaverse server to add a dramatic effect according to the level of excitement. As a result, in the next step S11, a video in which a dramatic effect according to the interaction information is being applied is captured in the metaverse space.

[0181] As another example, the interaction feedback unit 35 generates reflection information in accordance with the interaction information of each venue 23 and supplies the reflection information to the interaction information communication unit 34. The interaction information communication unit 34 transmits the reflection information supplied from the interaction feedback unit 35 to the cloud 22.

[0182] A specific example of processing according to interaction information will be described later.

[0183] In step S15, the metaverse login PC 21 determines whether or not to end the process of distributing the content data. For example, if the distribution of the content data has ended, it is determined that the process is to end.

[0184] If it is determined in step S15 that the process is not yet finished, the process then returns to step S11, and the above-described process is repeated.

[0185] On the other hand, if it is determined that the processing should be ended, the metaverse login PC 21 stops the processing of each unit, and the distribution processing ends.

[0186] In this way, the metaverse login PC 21 distributes content data while performing processing according to the interaction information received from each venue 23. In this way, the reactions at each venue 23 can be reflected in the content through live performances, etc., and more realistic content can be presented to the audience, allowing them to enjoy the content even more and improving their satisfaction.

[0187] <Description of Data Correction Process> When the distribution process of Fig. 15 is started by the metaverse login PC 21, the data correction process shown in Fig. 16 is performed in the cloud 22. The data correction process by the cloud 22 will be described below with reference to the flowchart of Fig. 16.

[0188] In step S51, the video / audio image receiving unit 41 receives the content data transmitted by the metaverse login PC 21 in step S12 of FIG.

[0189] In addition, if the content data does not contain sound image data or if stereo sound image data is included, the video / sound image receiving unit 41 supplies the camera position / direction information added to the video data constituting the content data to the spatial audio output unit 42.

[0190] When the spatial audio output unit 42 receives camera position and direction information from the video and audio image receiving unit 41, the spatial audio output unit 42 uses the camera position and direction information as microphone position and direction information. The spatial audio output unit 42 then reads out the multi-channel sound image data for the position and direction indicated by the microphone position and direction information from the multi-channel sound image data for each position and direction stored in advance, and supplies this to the video and audio image distribution unit 43. Note that the sound image data for each position and direction may be generated from sound image source data.

[0191] In step S52 , the video / audio image distribution unit 43 distributes (supplies) the content data supplied from the video / audio image receiving unit 41 to each of the M planetarium video / audio image conversion units 44 .

[0192] For example, if there are multiple content data with different positions and orientations of virtual cameras (virtual microphones), the video / audio image distribution unit 43 supplies appropriate content data to each planetarium video / audio image conversion unit 44 based on the position and orientation of the virtual camera, information indicating the dome radius added to the video data, and the like.

[0193] In this case, venue facility information may be appropriately acquired from the planetarium image and sound conversion unit 44 and used to determine the distribution destination of the content data. Specifically, for example, based on the dome radius specified by the venue facility information and information indicating the dome radius added to the video data of the content data, content data for a predetermined dome radius is supplied to the planetarium image and sound conversion unit 44 for the venue 23 with the predetermined dome radius.

[0194] When sound image data is supplied from the spatial audio output unit 42, the video / sound image distribution unit 43 supplies data consisting of the sound image data and the video data constituting the content data supplied from the video / sound image receiving unit 41 as final content data to the planetarium video / sound image conversion unit 44. The content data may also include haptic data.

[0195] In step S53, the planetarium image / sound image conversion unit 44 performs content correction processing on the content data supplied from the image / sound image distribution unit 43 based on the venue equipment information stored in advance, and supplies the resulting corrected content data to the image / sound image distribution unit 45.

[0196] Specifically, for example, based on the venue facility information, planetarium image-sound-image converter 44 identifies the tilt angle of the dome portion of venue 23 and the position of the image center. Then, based on the identified tilt angle and image center, planetarium image-sound-image converter 44 cuts out, as a panoramic image, a hemispherical portion with a predetermined tilt angle from the panoramic image based on the video data of the content, as described with reference to Fig. 9 for example.

[0197] Furthermore, when the seat layout type of venue 23 specified by the venue facility information is a concentric arrangement, planetarium image-sound-image conversion unit 44 cuts out a portion of the full-dome image to create a cut-out image, as described with reference to Figures 7 and 8, and then arranges and combines the cut-out images to generate a single full-dome image.

[0198] The planetarium image and sound conversion unit 44 converts the image format of the obtained full-dome image data, for example, from the dome master format to the equirectangular format, as necessary, and uses the resulting image data as the image data of the corrected content data. The image format used by the image projection unit 52 in the venue 23 can be identified from the venue facility information.

[0199] Furthermore, based on the venue equipment information, the planetarium image sound-image conversion unit 44 identifies the channel configuration of the speaker system serving as the spatial audio output unit 53 of the venue 23. Then, depending on the result of identifying the channel configuration, the planetarium image sound-image conversion unit 44 downmixes or upmixes the sound image data of the content data as appropriate, and uses the resulting multi-channel sound image data as sound image data of the corrected content data.

[0200] That is, the content correction process is a process of converting the sound image data constituting the content data into multi-channel sound image data corresponding to the channel configuration of the speaker system of the venue 23 indicated by the venue equipment information.

[0201] In step S54, the video / audio image distribution unit 45 transmits (distributes) the corrected content data supplied from the planetarium video / audio image conversion unit 44 to one or more venues 23.

[0202] Furthermore, when the corrected content data is distributed, interaction information is transmitted from each venue 23 .

[0203] In step S55, the feedback communication unit 46 receives the interaction information transmitted from each venue 23.

[0204] In step S56, the feedback communication unit 46 transmits the interaction information received from each venue 23 to the metaverse login PC 21. Note that the feedback communication unit 46 may also supply the interaction information to the feedback processing unit 47 as appropriate.

[0205] When the interaction information is sent, the metaverse login PC 21 sends reflection information generated according to the interaction information as appropriate, and the feedback communication unit 46 receives the sent reflection information and supplies it to the feedback processing unit 47.

[0206] In step S57, the feedback processing unit 47 performs processing according to the interaction information generated in each venue 23.

[0207] That is, based on the interaction information or reflection information supplied from the feedback communication section 46, the feedback processing section 47 instructs the planetarium video / audio image conversion section 44 to process the corrected content data as appropriate.

[0208] Then, in the next step S53, the planetarium video / audio image conversion unit 44 processes the corrected content data in accordance with instructions from the feedback processing unit 47, and supplies the resulting corrected content data to the video / audio image distribution unit 45.

[0209] Specific examples of the processing will be described later, but for example, the processing involves combining a portion of a panoramic image based on the video data that constitutes the corrected content data with video obtained by filming the audience seats in venue 23, which is included in the interaction information or reflection information.

[0210] In step S58, the cloud 22 determines whether or not to end the process of distributing the corrected content data (content data).

[0211] If it is determined in step S58 that the process is not yet finished, the process then returns to step S51, and the above-described process is repeated.

[0212] On the other hand, if it is determined in step S58 that the processing is to be ended, the cloud 22 stops the processing of each unit, and the data correction processing ends.

[0213] In this way, the cloud 22 performs content correction processing on the content data based on the venue facility information, and distributes the resulting corrected content data to each venue 23. In this way, it is possible to distribute corrected content data that is suited to the facilities of each venue 23. This makes it possible to present content appropriately at each venue 23, regardless of the facilities of the venue 23.

[0214] <Explanation of Playback Processing> After the cloud 22 performs the data correction processing of Fig. 16, the playback processing shown in Fig. 17 is performed in each venue 23. Hereinafter, the playback processing by the venue 23 (planetarium device 24) will be described with reference to the flowchart of Fig. 17.

[0215] In step S91, the video / audio image receiving unit 51 receives the corrected content data transmitted from the cloud 22 in step S54 of FIG.

[0216] Furthermore, the video and sound image receiving unit 51 extracts video data and sound image data from the received corrected content data, and supplies the video data to the video projection unit 52 and the sound image data to the spatial audio output unit 53 .

[0217] In step S92, the video projection unit 52 and the spatial audio output unit 53 play back (present) the content.

[0218] For example, a projector serving as the video projection unit 52 projects light based on the video data supplied from the video and sound image receiving unit 51 onto a dome portion (display unit) of the venue 23 to display the video of the content. That is, the video of the content is projected onto the dome portion. Furthermore, a speaker system serving as the spatial audio output unit 53 reproduces (outputs) the sound of the content based on the sound image data supplied from the video and sound image receiving unit 51.

[0219] As a result, content consisting of video and audio is presented to the audience in the venue 23.

[0220] In step S93, the scene analysis unit 54 generates interaction information, which is information relating to the reactions of the audience in the venue 23, and supplies it to the feedback communication unit 55.

[0221] For example, the scene analysis unit 54 generates interaction information by taking photographs and collecting sound of the audience seats and the like in the venue 23, and by performing scene analysis of the images and sounds of the audience seats and the like in the venue 23.

[0222] The interaction information may include, for example, video data of a video including at least the audience seats in venue 23 as a subject (hereinafter also referred to as venue video data), sound image data obtained by capturing sounds in venue 23 (hereinafter also referred to as venue sound image data), venue identification information that identifies venue 23, and the like.

[0223] In addition, the interaction information may include, for example, information indicating the level of excitement in the venue 23 obtained by analysis processing based on at least one of the venue video data and the venue sound image data, the number of spectators and their movements detected from video based on the venue video data, the results of skeletal structure estimation processing of spectators in the venue 23, the color and movement of penlights, etc. It can be said that such information indicating the level of excitement (estimated results of the level of excitement), the number of spectators, the movements of spectators, the results of skeletal structure estimation processing, etc. are the results of analysis processing based on the venue video data and the venue sound image data.

[0224] In step S94, the feedback communication unit 55 transmits the interaction information supplied from the scene analysis unit 54 to the cloud 22. The interaction information transmitted in step S94 is received by the cloud 22 in step S55 of FIG.

[0225] In step S95, the venue 23 (planetarium device 24) determines whether or not to end the content reproduction process.

[0226] If it is determined in step S95 that the process is not yet finished, the process then returns to step S91, and the above-described process is repeated.

[0227] On the other hand, if it is determined in step S95 that the processing is to be ended, the venue 23 (planetarium device 24) stops the processing of each unit, and the playback processing ends.

[0228] In this way, the venue 23 (planetarium device 24 ) plays back content based on the corrected content data, generates interaction information based on the state of the venue 23 , etc., and transmits it to the cloud 22 .

[0229] In this way, by receiving corrected content data obtained by the content correction process from the cloud 22 and playing back the content, it is possible to present the content appropriately regardless of the facilities at the venue 23 .

[0230] Furthermore, by generating and transmitting interaction information, it becomes possible to reflect the reactions and the like at each venue 23 in the content. This makes it possible to present more realistic content to the audience, and to enable the audience to enjoy the content more, thereby improving their satisfaction.

[0231] <Processing in Response to Audience Reactions, etc.> Here, an example of the processing performed in step S14 of FIG. 15 and step S57 of FIG. 16, that is, the processing performed by the interaction feedback unit 35 and the feedback processing unit 47, will be described.

[0232] For example, in step S14 of Fig. 15, the metaverse login PC 21, more specifically the metaverse login PC 21 or the metaverse server, processes the content data according to the interaction information (processing). In this case, in step S12, which is performed next in the distribution process of Fig. 15, the content data that has been processed based on the interaction information, such as processing, is distributed (transmitted).

[0233] Processing based on interaction information here refers to, for example, at least one of the following: synthesis of images based on venue video data, synthesis of sounds based on venue sound image data, addition or modification of effects according to the results of analysis processing, synthesis of images of virtual audience members such as avatars according to the results of analysis processing, change of the viewpoint position of images based on the video data that constitutes the content data, and sending stamps to a comment section displayed on a device such as a smartphone within the metaverse space or as an area outside the metaverse space.

[0234] In contrast to this, for example, in step S57 of Figure 16, on the cloud 22 side, based on the control of the feedback processing unit 47, the planetarium image and sound image conversion unit 44 processes the corrected content data according to the interaction information.

[0235] The processing may be performed according to the interaction information itself, or according to the reflection information generated from the interaction information, but in either case, the processing can be said to be processing performed according to the interaction information.

[0236] Therefore, it can be said that the processing is carried out based on at least one of the venue video data, venue sound image data, and the results of analysis processing based on at least one of the venue video data and venue sound image data obtained for one or more venues 23.

[0237] The processing in step S57 etc. can be, for example, at least one of the following: synthesis of images based on venue image data, synthesis of sounds based on venue sound image data, synthesis of images of virtual audience members such as avatars according to the results of the analysis processing, and addition or modification of performance according to the results of the analysis processing.

[0238] A more specific example of processing (processing) of content data in accordance with interaction information will be described below.

[0239] For example, it is conceivable to display images of the audience at each venue 23 in the metaverse space as feedback regarding the reaction of the venue 23 to a live performer (artist) in the metaverse space.

[0240] In this case, for example, the interaction information transmitted from each venue 23 includes venue image data and venue identification information.

[0241] In the process of step S14 of FIG. 15, the interaction feedback unit 35 transmits the venue image data and venue identification information of each venue 23 to the metaverse server, and requests the metaverse server to composite the images of each venue 23.

[0242] The metaverse server then synthesizes an image based on the venue image data of the venue 23 in a position (area) in the metaverse space that is determined for the venue 23 indicated by the venue identification information.

[0243] As a result, the image of the metaverse space captured next time in step S11 of FIG. 15 will display an image of the spectator seats in the venue 23, for example, as shown in FIG.

[0244] The part indicated by arrow Q81 in Figure 18 shows a live performance in the metaverse space, captured by the metaverse space capture unit 31. In this example, a stage is placed in the center of the figure, and performers are performing live on the stage.

[0245] Additionally, area W71 at the back of the stage displays an image based on the venue video data of the audience seats in venue 23-1, as indicated by arrow Q82. Similarly, area W72 displays an image based on the venue video data of the audience seats in venue 23-2, as indicated by arrow Q82.

[0246] By displaying the state of the audience at each venue 23 in the metaverse space in this way, performers at a live performance can grasp the state of each venue 23, allowing for a more exciting live performance.

[0247] When combining images of each venue 23 into the metaverse space, sounds based on venue sound image data of each venue 23 may also be combined as sounds in the metaverse space.

[0248] In such a case, the interaction information is made to include venue sound image data, and the venue sound image data is also supplied from the interaction feedback unit 35 to the metaverse server, and synthesis of sound based on the venue sound image data for the sound in the metaverse space is also requested.

[0249] As a result, the sound based on the sound image data of the content captured in step S11 of Figure 15 the next time will include not only the sounds of live performers, etc., but also the actual voices of the audience at each venue 23 as cheers in the metaverse space.

[0250] The synthesis of the video of the audience seats in the venue 23 based on the venue video data may be performed in the cloud 22 instead of the metaverse server.

[0251] In such a case, for example, in step S14 of Figure 15, reflection information is generated for each venue 23, which includes venue image data and venue identification information for each venue 23 other than that venue 23, and requests the synthesis of an image of the venue 23 based on the venue image data.

[0252] 16, or more specifically, in step S53, the planetarium image sound conversion unit 44 synthesizes an image based on the venue image data into a predetermined area of ​​an image based on the image data constituting the corrected content data, based on the reflection information supplied from the feedback processing unit 47, to produce final image data of the corrected content data. In other words, the image synthesis process is performed as the above-mentioned processing process.

[0253] Specifically, for example, it is assumed that processing is performed on the video data of the corrected content data for a venue 23 as shown in FIG.

[0254] In this example, an image of another venue 23 is composited with the rear area of ​​the full-dome image based on the image data of the corrected content data, that is, the area displayed at the rear of the venue 23 .

[0255] Here, the video data is processed (composite) so that an image of the spectator seats of another venue 23 is displayed in area W81 behind venue 23, and similarly an image of the spectator seats of another venue 23 is displayed in area W82.

[0256] Therefore, at the venue 23, images of the content are displayed together with images of the other venues 23. In other words, the audience can also see what is happening at the other venues 23, and can enjoy the live performance even more.

[0257] The reflection information may also include venue sound image data for each venue 23, and the planetarium video / sound-image conversion unit 44 may perform processing such that audio synthesis is performed based on the audio image data constituting the corrected content data and the venue audio image data for each venue 23. In such a case, in the example of Fig. 19, when content is played back in venue 23, audio synthesis is performed so that the audio of the audience in another venue 23 is heard from the direction of area W81 and the audio of the audience in another venue 23 is heard from the direction of area W82.

[0258] The interaction information may include information indicating the level of excitement of the audience at the venue 23 (estimated result of the level of excitement).

[0259] In such a case, for example, in step S93 of Figure 17, the scene analysis unit 54 performs processing using AI (Artificial Intelligence) based on the venue video data and venue sound image data, and estimates the level of excitement in the venue 23 from the processing results.

[0260] The processing using AI here refers to, for example, skeleton estimation processing using a skeleton estimator obtained through machine learning, emotion estimation processing using an emotion estimator obtained through machine learning, and excitement estimation processing using audio processing such as voice recognition.

[0261] For example, in the skeleton estimation process, the skeleton of each audience member, in other words, the movements and postures of the audience members, are obtained as estimation results, and the degree of excitement in the venue 23 can be estimated from these estimation results. In addition, in the emotion estimation process, the emotion of each audience member is obtained as estimation results, and the degree of excitement in the venue 23 can be estimated from these estimation results. In addition, in voice-related processing such as voice recognition, it is possible to estimate the degree of excitement from the frequency and sound pressure of the audience members' voices.

[0262] The scene analysis unit 54 generates interaction information including information indicating the degree of excitement in the venue 23 and venue identification information.

[0263] Also, in step S14 of FIG. 15, for example, the interaction feedback unit 35 changes the presentation in the metaverse space based on information indicating the degree of excitement included in the interaction information.

[0264] That is, the interaction feedback unit 35 requests the metaverse server to add a performance determined by the degree of excitement, and the metaverse server adds the performance in accordance with the request from the interaction feedback unit 35 within the metaverse space.

[0265] As a result, the image of the Metaverse space obtained by the next capture in step S11 of FIG. 15 will be an image in which effects are performed according to the degree of excitement, as shown in FIG. 20, for example.

[0266] In the example of Fig. 20, a stage is placed in the center of the figure, and performers are performing live on the stage. Objects such as stars are displayed near the stage as a production effect according to the level of excitement.

[0267] For example, the interaction feedback unit 35 changes the number of stars displayed as a performance, the brightness of the stars, the movement of the stars, etc., depending on the level of excitement in the venue 23.

[0268] Note that the effects according to the level of excitement are not limited to stars, and may include video objects such as flames and fireworks, as well as sound effects, etc. In this case, for example, it is possible to change the height and number of video objects, the rate of change of movement speed, etc., the movement of the video objects such as shaking, the volume and number of sound effects, etc., according to the level of excitement.

[0269] Furthermore, as a production effect according to the level of excitement, the color of any video object such as a penlight in the Metaverse space may be displayed in a color, for example, that is determined for the venue 23 with the highest level of excitement.

[0270] Another example of a performance based on the level of excitement is a performance in which the venue for the live performance changes depending on the level of excitement. For example, a user may be moved (transferred) to a venue determined based on the level of excitement in the metaverse space, with the destination space being the live venue. In this case, the destination space may be a space that simulates the venue 23 with the highest level of excitement. Furthermore, in response to a request from the interaction feedback unit 35 to the metaverse server, stamps (marks) such as hearts may be displayed in a predetermined area, such as a comment section, displayed in the metaverse space depending on the level of excitement. Note that there may be users who log in to the metaverse space using devices such as smartphones and view content. In such cases, the metaverse server may transmit data for displaying stamps to the device, and the device may display stamps such as hearts in a display area, such as a comment section, that is provided separately from the area where the video (content) of the metaverse space is displayed.

[0271] In addition, it is also possible to transmit information according to the situation in the metaverse space to the venue 23.

[0272] For example, it is conceivable to prepare virtual spectator seats (seats) in the metaverse space and virtually allocate spectators to the seats in the metaverse space for each venue 23. In this case, the viewpoint in the metaverse space, that is, the position and direction of the virtual camera, can be changed for each venue 23.

[0273] As a specific example, let us assume that a live concert venue in the metaverse space is a space as shown in FIG.

[0274] 21, audience seats are provided in front of a stage ST11 where the live performance will be held. Also, for example, an area R51 of the audience seats is allocated to one venue 23 (hereinafter also referred to as venue A), and an area R52 is allocated to another venue 23 (hereinafter also referred to as venue B).

[0275] In such a case, the content image will be played in venue A with the position of area R51 as the viewpoint, i.e., the position of the virtual camera, and the content image will be played in venue B with the position of area R52 as the viewpoint.

[0276] Which area of ​​the audience seats is allocated (assigned) to which venue 23 may be determined in advance, or may be dynamically determined by the interaction feedback unit 35 based on information contained in the interaction information, such as the level of excitement.

[0277] There are several possible methods for realizing a performance in which the viewpoint changes for each venue 23 in this way.

[0278] For example, in a first method, the interaction feedback unit 35 appropriately uses interaction information or the like to determine the area of ​​the audience seats to be allocated to each venue 23, and notifies (supplies) the determination result to the metaverse space capture unit 31 and the spatial sound capture unit 32. In this case, the metaverse space capture unit 31 and the spatial sound capture unit 32 determine the positions of the virtual cameras and virtual microphones based on the positions of the audience seating areas indicated by the determination result supplied, and capture video and audio.

[0279] Furthermore, the video data and sound image data obtained by capture may be appropriately added with venue identification information indicating the corresponding venue 23. This enables the video and sound image distribution unit 43 of the cloud 22 to identify the planetarium video and sound image conversion unit 44 to which each content data is to be distributed.

[0280] The second method is a method in which the interaction feedback unit 35 appropriately uses interaction information or the like to determine the area of ​​spectator seats to be allocated to each venue 23, and generates reflection information indicating the result of the determination for each venue 23. For example, the reflection information may store venue identification information indicating the venue 23 and the position of the spectator seats allocated to that venue 23, that is, the position of the virtual camera, in association with each other.

[0281] In this case, for example, the metaverse space capture unit 31 and the spatial audio capture unit 32 capture video and audio in advance at positions and orientations of multiple virtual cameras and virtual microphones corresponding to each of multiple areas of the audience seats.

[0282] Furthermore, in the cloud 22, the feedback processing unit 47 supplies the reflection information to the video / audio image distribution unit 43 or the planetarium video / audio image conversion unit 44, causing the appropriate content data to be selected. In this example, the venue 23 can be appropriately associated with the content data for that venue 23 based on the camera position / direction information added to the video data of the content data and the reflection information. In other words, appropriate content data can be selected for each venue 23.

[0283] In addition, for example, when a live performer (artist) calls to venue A, it is possible to have the special effects only be applied to venue A.

[0284] For example, it is conceivable that which venue 23 the performer has called out to can be identified by image recognition or voice recognition based on content data, input operations by the performer or live event management staff, pre-prepared event information indicating the timing of the call, etc. The event information can be, for example, information that associates the playback time of the content with venue identification information of the venue 23 to be called out to at that playback time.

[0285] When the interaction feedback section 35 determines by some method that the performer has made a call to a specific venue 23, it performs processing to add a performance based on venue identification information indicating the venue 23 to which the call is addressed and information indicating the content of the performance to be given in response to the call.

[0286] Specifically, for example, the interaction feedback unit 35 generates reflection information including video data for adding effects (hereinafter also referred to as effect video data) and venue identification information indicating the venue 23 to which the call is addressed, and transmits the reflection information to the cloud 22 via the interaction information communication unit 34.

[0287] The feedback processing unit 47 then supplies the reflection information to the appropriate planetarium image / sound image conversion unit 44, which synthesizes an effect image based on the effect image data with an image based on the image data constituting the corrected content data. When adding an effect, in addition to synthesizing an effect image using the effect image data, effect sound image data for reproducing effect sound (sound effects) may be used to synthesize effect sound with the sound of the content. Also, the effect may be added only as effect sound.

[0288] In the example described with reference to Figure 18, it was explained that the sounds (cheers) of spectators and the like in venue 23 may be synthesized as sounds in the metaverse space, but in the example shown in Figure 21, similar to the example of Figure 18, the sounds of spectators and the like in venue 23 may be synthesized as sounds in the metaverse space.

[0289] In such a case, the interaction feedback unit 35 supplies the metaverse server with not only the venue sound image data of the venue 23, but also information indicating the location of the area of ​​the audience seats allocated to that venue 23, so that the sound based on the venue sound image data is localized at the location of the audience seats.

[0290] In this way, for example, in the example of Fig. 21, the sound image data of the content will be captured in step S11 of Fig. 15 so that the cheers from venue A are emitted from area R51 and the cheers from venue B are emitted from area R52. Note that sound synthesis based on the venue sound image data may be performed by the planetarium image and sound image conversion unit 44 of the cloud 22.

[0291] As described above, the scene analysis unit 54 appropriately performs skeleton estimation processing on the venue video data to estimate the movements and postures of the audience.

[0292] Therefore, using the results of the skeleton estimation process, feedback may be provided in the metaverse space to display virtual spectators, such as avatars, that resemble spectators who are actually present at the venue 23.

[0293] In such a case, the interaction information may include, for example, the results of the skeleton estimation process and information indicating the number of spectators.

[0294] The interaction feedback unit 35 supplies the metaverse server with information indicating the results of the skeleton estimation process and the number of spectators at each venue 23, and requests the display of virtual spectators such as avatars with movements and postures determined by the results of the skeleton estimation process at any position within the metaverse space.

[0295] Then, in response to a request from the interaction feedback unit 35, the metaverse server displays (projects) a predetermined number of virtual audience members such as avatars at any position in the metaverse space, such as a position near the stage, based on the results of the skeleton estimation process, etc. In other words, a predetermined number of images of audience members making predetermined movements and poses are synthesized into the image of the metaverse space.

[0296] As a result, in the image of the metaverse space captured next time in step S11 in Fig. 15, one or more images of spectators, including spectator M11, will be newly displayed near the stage as shown in Fig. 22. This provides feedback that makes it seem as if the spectators in the actual venue 23 are in the metaverse space.

[0297] In addition, the metaverse space may display not only virtual spectators such as avatars that resemble actual spectators, but also the actual number of spectators at each venue 23 and the total number of spectators at all venues 23.

[0298] In addition, the example described with reference to FIG. 21 and the example described with reference to FIG. 22 may be combined to realize a performance that will liven up the live performance at each venue 23 together.

[0299] Specifically, for example, the performance in the metaverse space can be made to move in conjunction with the movements of the audience at the actual venue 23.

[0300] In this case, for example, if the interaction information includes the results of the skeleton estimation process, the metaverse server can change the movement of visual objects such as flames in the metaverse space and sound effects in accordance with the movements of the audience identified by the results of the skeleton estimation process, in response to a request from the interaction feedback unit 35.

[0301] 23, for example, when an audience member waves a penlight at an actual venue 23, the penlight held by a virtual audience member can also move in the Metaverse space. In FIG. 23, parts corresponding to those in FIG. 21 are assigned the same reference numerals, and their explanation will be omitted as appropriate.

[0302] In the example of Fig. 23, an audience member at Venue A is waving a penlight, as shown in the upper right corner of the figure. This movement can be identified by skeletal structure estimation processing, image recognition, etc., and the interaction information includes the results of the skeletal structure estimation processing, image recognition, etc.

[0303] The metaverse server controls the display of the video in area R51 so that the penlights held by spectators displayed in area R51 of the spectator seats allocated to venue A in the metaverse space move in accordance with the movement of the penlights identified by the results of skeletal structure estimation processing, image recognition, etc., in response to a request (control) from the interaction feedback unit 35. As a result, the penlights of the virtual spectators in area R51 move in the same way as the penlights of the actual spectators in venue A.

[0304] Furthermore, the colors of virtual audience members such as avatars that represent audience members in the metaverse space, the colors of penlights, etc. may be changed depending on the results of the skeleton estimation process for the audience members in the venue 23. In addition, when a live song is being played, the display format of the colors of video objects in the metaverse space, such as the virtual audience members and penlights, may be changed depending on the parts such as vocals and guitar, or the level of excitement in the venue 23.

[0305] Additionally, it may be possible to allow the audience at the venue 23 to react in the same way as the audience participating in the live performance in the metaverse space.

[0306] For example, when the scene analysis unit 54 performs analytical processes such as skeletal estimation and image recognition on the venue video data to detect pre-specified movements of the audience, it is possible to have a reaction, i.e., a performance, be carried out in the metaverse space in accordance with the detected movements.

[0307] In such a case, for example, the interaction information may include the detection results of a pre-specified movement, i.e., information indicating whether or not movement is present, and the interaction feedback unit 35 requests the metaverse server to add a performance that corresponds to the movement based on the detection results of the movement.

[0308] As an example of a performance that responds to the movements of the audience, for example, when an audience member at venue 23 makes a gesture (movement) of making a heart symbol with their hands, a performance in which the heart symbol flies within the metaverse space could be performed.

[0309] As another example, a performance in which a picture drawn by an audience member at venue 23 using a penlight flies through the Metaverse space is conceivable. In this case, it is possible to identify the trajectory of the movement of the picture drawn by the audience member, i.e., the penlight, by analytical processing such as skeletal structure estimation processing and image recognition processing.

[0310] Furthermore, for example, when a spectator operates a product purchased at the venue 23, a reaction (performance) associated with the product can be performed in the Metaverse space. In this case, the interaction information can include information indicating whether or not the product has been operated as a result of the analysis process. The performance corresponding to the operation of the product can be, for example, a performance in which an icon of a specific performer (such as a performer the audience is cheering for) at a live concert is displayed.

[0311] Spectators in the venue 23 may be able to confirm that their own reactions (movements) are reflected in the Metaverse space. Specifically, for example, if the dome portion of the venue 23 is inclined, it may be possible to display the heads of virtual spectators, such as avatars in the Metaverse space, corresponding to the actual spectators, at the bottom of the dome portion.

[0312] Although the above description has been given of an example in which content is distributed in real time in the information processing system 11, pre-recorded content may also be distributed.

[0313] In such a case, the metaverse space capture unit 31 and the spatial audio capture unit 32 of the metaverse login PC 21 function as a recording unit (storage). That is, the metaverse space capture unit 31 reads and outputs video data with an appropriate virtual camera position and orientation from pre-recorded video data, and the spatial audio capture unit 32 reads and outputs sound image data with an appropriate virtual microphone position and orientation from pre-recorded sound image data.

[0314] Furthermore, by replacing the metaverse space capture unit 31 and spatial audio capture unit 32 of the metaverse login PC 21 with a real celestial sphere camera and microphone, it is also possible to distribute real live performances that take place in real space as content.

[0315] Furthermore, it is also conceivable to use haptics (tactile presentation) as a method for feeding back the excitement of the live performance to the venue 23 .

[0316] In such a case, for example, the interaction feedback unit 35 instructs the video / audio image distribution unit 33 to distribute haptic data of vibration intensity and vibration pattern according to the level of excitement, etc., included in the interaction information. In response, the video / audio image distribution unit 33 distributes content data consisting of video data, audio image data, and haptic data.

[0317] At the venue 23, the video / audio image receiving unit 51 supplies the haptic data included in the corrected content data to vibration devices attached to seats, etc. in the venue 23, and vibrates the vibration devices to provide a haptic presentation to the audience. The haptic presentation may also be achieved by vibrating devices such as penlights held by the audience. In this way, the seats, etc. in the venue 23 will vibrate in conjunction with the excitement of the live performance. Alternatively, the haptic data may be stored in reflection information and provided to the cloud 22.

[0318] <Other Modifications> <Configuration Examples of Information Processing System> Figures 24 to 26 show other configuration examples of the information processing system 11. Note that in Figures 24 to 26, parts corresponding to those in Figure 1 are given the same reference numerals, and descriptions thereof will be omitted as appropriate.

[0319] Figure 24 shows an example in which multiple metaverse login PCs are prepared so that if any abnormality occurs during content distribution, such as a communication failure or video abnormality, the system can automatically switch to a backup metaverse login PC.

[0320] The information processing system 11 shown in FIG. 24 includes a metaverse login PC 21, a metaverse login PC 101, a cloud 22, and venues 23-1 to 23-N.

[0321] The metaverse login PC 21 has the same configuration as that in FIG. 1, and the metaverse login PC 101 has the same configuration as the metaverse login PC 21.

[0322] The cloud 22 further includes a selector 111 in addition to the components from the video / audio image receiving unit 41 to the feedback processing unit 47 shown in FIG.

[0323] The selector 111 supplies the content data received from either the metaverse login PC 21 or the metaverse login PC 101 to the video and audio image receiving unit 41 .

[0324] Similarly, the selector 111 transmits the interaction information supplied from the feedback communication unit 46 to either the metaverse login PC 21 or the metaverse login PC 101. The selector 111 also supplies the feedback communication unit 46 with either the reflection information received from the metaverse login PC 21 or the reflection information received from the metaverse login PC 101.

[0325] In this way, the metaverse login PC that is the content distribution source can be appropriately switched to prevent interruption of content distribution. Note that three or more metaverse login PCs that can be content distribution sources may be prepared.

[0326] FIG. 25 shows an example in which a live performance or the like in the real world, rather than in the metaverse space, is distributed as content.

[0327] The information processing system 11 shown in FIG. 25 includes a PC 201, a cloud 22, and venues 23-1 to 23-N.

[0328] In the example shown in FIG. 25, the configuration of the cloud 22 and the configuration of each venue 23 are the same as those in FIG.

[0329] The PC 201 includes a real space capture unit 211 , a spatial audio capture unit 212 , a video and audio image distribution unit 33 , an interaction information communication unit 34 , and an interaction feedback unit 35 .

[0330] The PC 201 is configured such that a real space capture unit 211 and a spatial sound capture unit 212 are provided instead of the metaverse space capture unit 31 and the spatial sound capture unit 32 in the metaverse login PC 21 of FIG.

[0331] The real space capture unit 211 is composed of one or more omnidirectional cameras or the like, captures an image of a subject such as a live performance venue, which is a real space, and supplies the resulting video data of the omnidirectional video to the video and audio image distribution unit 33 as video data of the content.

[0332] The spatial audio capture unit 212 consists of one or more microphones, and picks up sounds in a real space such as a live venue, and supplies the resulting sound image data to the video and audio image distribution unit 33 as sound image data of the content.

[0333] 26 shows an example in which a metaverse space is displayed in a planetarium, which is a facility (venue) in the real world, and performers give a live performance or the like in the real world where the metaverse space is displayed. In this example, a live performance or the like in the real world that is held in the venue where the metaverse space is displayed is distributed as content.

[0334] The information processing system 11 shown in Fig. 26 includes a planetarium venue 301, a cloud 22, and venues 23-1 to 23-N. Note that in Fig. 26, parts corresponding to those in Fig. 25 are given the same reference numerals, and descriptions thereof will be omitted as appropriate.

[0335] The planetarium venue 301 is a venue where live performances and the like are held, that is, a venue where content is filmed and the like, and the planetarium venue 301 has a planetarium device for distributing content.

[0336] The planetarium venue 301 (planetarium device) has a real space capture unit 211 , a spatial sound capture unit 212 , a video and audio image distribution unit 33 , an interaction information communication unit 34 , and an interaction feedback unit 35 .

[0337] The planetarium venue 301 has the same configuration as the PC 201 in FIG. 25, and the content data transmitted from the planetarium venue 301 is received by the cloud 22 .

[0338] The above description mainly deals with an example in which the audience (users) at each venue 23 are the viewers of the content. However, this is not limiting, and it is also possible for users at home to be able to view the content, that is, for individual users outside the venue 23 to be able to participate in a live performance or the like in the metaverse space.

[0339] In such a case, for example, each individual user will use their own metaverse login PC to connect to a metaverse server that provides services related to the metaverse space. For example, each individual user's metaverse login PC has functions similar to the metaverse space capture unit 31 and spatial audio capture unit 32 shown in Figure 1, and can display video of live performances or other content on a display and play sounds from a speaker.

[0340] In this way, by allowing not only users (audience) at the venue 23 but also individual users to participate in the live performance in the metaverse space, more users can enjoy the content. Furthermore, since processing is performed according to the interaction information as described above, not only users (audience) at the venue 23 but also users outside the venue 23, such as at home, can feel a sense of unity with users at the venue 23 and get excited. In other words, the content can be enjoyed even more.

[0341] Second Embodiment Example of Studio Configuration The above describes an example in which content data is made up of video data and sound image data (audio data) of the content. However, the present invention is not limited to this, and the content data may be data that includes at least motion data for displaying the video of the content and sound image data.

[0342] By transmitting motion data for generating video data at the venue, rather than transmitting the video data itself as data for displaying the content video, the amount of content data transmitted can be significantly reduced.

[0343] When content data includes motion data and sound image data, for example, an information processing system to which the present technology is applied is a content distribution system having a studio that distributes content and one or more venues that receive the content distribution.

[0344] In this case, the studio that constitutes the information processing system (content distribution system) is configured as shown in FIG. 27, for example.

[0345] In the example shown in FIG. 27, the information processing system has a studio 401, a venue 402-1, and a venue 402-2, and the studio 401, the venue 402-1, and the venue 402-2 are interconnected via a cloud 403 on a network.

[0346] Note that although venues 402-1 and 402-2 are shown here as destinations for content distribution, there may be any number of venues to which content is distributed. Hereinafter, when there is no need to particularly distinguish between venues 402-1 and 402-2, they will also be simply referred to as venue 402. Hereinafter, venue 402-1 will also be particularly referred to as venue A, and venue 402-2 will also be particularly referred to as venue B.

[0347] In studio 401, a performance is performed by one or more performers, motion data, sound image data, and facial expression data showing the performers' facial expressions related to the performance are generated, and content data including these data is transmitted (distributed) to each venue 402.

[0348] In the following description, it is assumed that a performance such as a live music concert is being performed by multiple performers in studio 401. In this case, the performers are subjects (objects) to be displayed in the content.

[0349] Venue 402 is a facility capable of presenting content (a facility capable of playing content) to a plurality of audience members (users) who will be viewers of the content, such as a planetarium venue similar to venue 23. In the following, the description will be continued assuming that venue 402 is a planetarium venue similar to venue 23.

[0350] At each venue 402, information processing including rendering processing is performed based on the content data acquired (received) from studio 401, and content is presented to the audience based on the resulting data (hereinafter also referred to as content data for presentation). For example, the content presented at venue 402 may be a video with audio in which an avatar corresponding to a performer performs in a virtual space such as a metaverse space. Note that venues 402-1 and 402-2 may have different equipment, such as projection equipment, for presenting content.

[0351] Furthermore, the above-mentioned interaction information is transmitted from each venue 402 to the studio 401 as appropriate.

[0352] In this embodiment, the interaction information includes venue video data showing the audience at venue 402, i.e., venue seats as subjects, venue sound image data obtained by capturing sounds from venue 402, and venue identification information that identifies venue 402.

[0353] As in the first embodiment, the interaction information may include information obtained as a result of analysis processing based on at least one of the venue video data and the venue sound image data, such as information indicating the degree of excitement.

[0354] Additionally, when any operation is performed on the venue 402 side in relation to the presentation of content, such as in response to an emergency, a control signal relating to that operation is transmitted to the studio 401 as appropriate.

[0355] As an example, when the content video cannot be temporarily presented due to a malfunction of equipment such as a PC that performs rendering processing in venue 402, a pre-prepared image (still image or moving image (video)) called a cover image is displayed instead of the content video. In such a case, a control signal is transmitted to indicate that the content video cannot be temporarily presented and that a predetermined cover image is being displayed. The control signal can be said to be information indicating the status (hereinafter also referred to as the progress status) regarding the presentation of the content in venue 402.

[0356] Although an example will be described here in which the distribution destination of the content data is only the venue 402 that can accommodate multiple spectators, the content data may also be distributed (transmitted) from the studio 401 to a user's personal PC or the like via the cloud 403. In such a case, the user can also view the content at home.

[0357] The studio 401 includes a headset 411, a motion capture system 412, a camera 413, a communication / rendering PC 414, a mixer 415, wireless equipment 416, a monitor 417-1, a monitor 417-2, a monitor 418, a facial expression controller 419, and a host PC 420 as equipment.

[0358] The headset 411 is worn on the head of a performer performing on a stage or the like in the studio 401 .

[0359] The headset 411 has a wireless communication function, a microphone, and earphones, and picks up the sounds (voice) emitted by the performer and transmits the resulting sound image data (hereinafter also referred to as performer sound image data) to the wireless device 416 via wireless communication.

[0360] The headset 411 also receives sound image data transmitted by wireless communication from the wireless device 416, and outputs sound based on the received sound image data, thereby presenting the sound to the performers. For example, the sounds presented to the performers by the headset 411 may include voice instructions from the production director / manager, voices of other performers, voices from each venue 402, and sound for musical accompaniment or performance (hereinafter also referred to as performance sound). The headset 411 may be connected to a specific device by wire, and the device may be connected to the mixer 415, or a wired or wireless microphone and headphones may be worn on the performer's head, etc., instead of the headset 411.

[0361] The motion acquisition system 412 comprises, for example, a motion sensor worn by the performer, a unit for wireless communication with the motion sensor, a processor, and the like.

[0362] The motion acquisition system 412 acquires sensing data indicating the movement of each part of the performer using motion sensors attached to the performer, and generates motion data indicating the movement of the performer by performing signal processing on the sensing data as needed. For example, the motion data consists of vector information indicating the movement of each part of the performer.

[0363] The motion capture system 412 provides the captured actor motion data to a communication / rendering PC 414 .

[0364] The motion acquisition system 412 is not limited to a type in which a motion sensor is attached to the performer, but may be any type, such as a type that captures video of the performer and obtains motion data by analyzing the video.

[0365] The camera 413 is a camera for acquiring facial expression data showing the facial expressions of the performer. It photographs the performer's face as the subject and supplies the resulting video data (hereinafter also referred to as facial expression video) to the communication / rendering PC 414.

[0366] For example, the facial expression data may be information indicating one of a plurality of predetermined facial expressions, such as a smile. More specifically, the facial expression data may be an facial expression identification ID indicating the facial expression of the performer, such as a smile.

[0367] It should be noted that if facial expression data is input by a facial expression controller 419 (described later), the camera 413 does not necessarily have to be provided.

[0368] The communication / rendering PC 414 is made up of one or more PCs and controls the overall operation of the studio 401 .

[0369] For example, the communication / rendering PC 414 generates facial expression data for each performer by performing an analysis process based on the video data of the facial expression video supplied from the camera 413. The communication / rendering PC 414 also generates content data including the facial expression data, motion data from the motion acquisition system 412, and performer sound image data supplied from the wireless device 416, and transmits the content data to each venue 402. Note that the content data may also include musical accompaniment and background music, sound image data of performance sounds for performances, and video data of performance effects in the virtual space.

[0370] Furthermore, for example, the communication / rendering PC 414 performs rendering processing based on the content data, and generates confirmation content data made up of video data and sound image data for presenting the content.

[0371] The confirmation content data generated by the communication / rendering PC 414 is video data and audio image data for reproducing video and audio of a performer, more specifically, an avatar corresponding to the performer, viewed from the front in the virtual space. In other words, the confirmation content data is generated as video data and audio image data that presents video and audio when a virtual camera and a virtual microphone are placed in front of an avatar corresponding to the performer in the virtual space. Note that in the following explanation, the virtual space in which the avatar is placed is assumed to be the Metaverse space.

[0372] The mixer 415 outputs various types of sound image data in response to the operation of the mixing engineer. For example, the mixer 415 supplies one or more of various types of sound image data, such as sound image data of the voice of instructions from the production director manager supplied from a predetermined microphone, the voice of other performers supplied from the communication / rendering PC 414, the voice of each venue 402, and sound image data of the performance sound, to the headset 411 via the wireless device 416. In more detail, when there are multiple pieces of sound image data to be output, the mixer 415 mixes these multiple pieces of sound image data to create one piece of sound image data and outputs it to the wireless device 416, etc.

[0373] The wireless device 416 transmits the sound image data supplied from the mixer 415 to the headset 411 via wireless communication, and receives the performer sound image data transmitted from the headset 411 via wireless communication and supplies it to the communication / rendering PC 414. Note that the performer sound image data obtained by the headset 411 may be supplied to the communication / rendering PC 414 via the wireless device 416 and the mixer 415.

[0374] Monitor 417-1 and monitor 417-2 display video of venue 402 based on venue video data supplied from communication / rendering PC 414. As an example, monitor 417-1 displays video of venue 402-1 (venue A), and monitor 417-2 displays video of venue 402-2 (venue B). Furthermore, in studio 401, sound from venue A or venue B may be reproduced based on venue sound image data as needed.

[0375] Hereinafter, when there is no need to particularly distinguish between the monitor 417-1 and the monitor 417-2, they will also be simply referred to as monitor 417.

[0376] For example, when a live performance is taking place, the production manager, the progress operator, the performers, etc. can check (understand) the situation at each venue 402 in real time by watching the video displayed on the monitor 417.

[0377] The monitor 418 displays an image of the metaverse space as seen from the front of the avatar (hereinafter also referred to as the confirmation image) based on the confirmation content data supplied from the communication / rendering PC 414. Like the monitor 417, the monitor 418 is also positioned so that it can be visually confirmed by the performers, production supervisor manager, and progress operator in the studio 401. By viewing the confirmation image displayed on the monitor 418, the performers can confirm how content such as a live performance is being presented to the audience in each venue 402.

[0378] In response to operations by the facial expression operator, the facial expression controller 419 supplies facial expression data indicating the facial expressions of each performer performing a live performance or the like to the communication / rendering PC 414. Note that it is sufficient for the studio 401 to obtain facial expression data from either the camera 413 or the facial expression controller 419. Therefore, it is sufficient for only one of the camera 413 and the facial expression controller 419 to be provided in the studio 401.

[0379] The host PC 420 is a PC operated by a host operator who controls the progress of the live performance, or more specifically the distribution of content data, and has a display unit 431 , an input unit 432 , and a control unit 433 .

[0380] The input unit 432 is made up of, for example, a mouse, keyboard, switches, etc., and supplies signals to the control unit 433 in response to operations by the progress operator.

[0381] In the promoter PC 420, the operation of the entire promoter PC 420 is controlled by a control unit 433. For example, the control unit 433 performs predetermined processing in response to signals supplied from the input unit 432, and controls the display of various images on the display unit 431.

[0382] Furthermore, for example, the control unit 433 displays the progress of content presentation at each venue 402 on the display unit 431 based on control signals for each venue 402 supplied from the communication / rendering PC 414, and supplies progress information corresponding to signals from the input unit 432 to the communication / rendering PC 414.

[0383] The progress information is, for example, information indicating the progress of content distribution from the studio 401 to each venue 402, i.e., the progress of a live performance or the like related to the content. In the communication / rendering PC 414, the progress information supplied from the progress PC 420 is treated as one piece of data constituting the content data to be distributed (transmitted) to each venue 402. That is, in addition to the motion data and performer sound image data as content data, the progress information is also distributed to each venue 402.

[0384] Details will be described later, but for example, in the proceeding PC 420, the control unit 433 displays on the display unit 431 a playlist of live performances and the like provided as content.

[0385] A playlist is information that shows the overall progress schedule of a live performance or the like. In other words, a playlist lists multiple parts that make up the content of a live performance or the like in the order in which they will be performed. Note that a playlist may list multiple parts, including parts that will be performed in the live performance or the like as content, and parts that are not originally scheduled to be performed in the live performance or the like but are prepared as replacements in case of an emergency.

[0386] Specifically, the playlist displays the parts of a performance, such as the MC (master of ceremony) part where the performer gives a talk, or the singing part where the performer plays or sings a song, in order as list items. Each list item also displays the time allocated to the corresponding part, if necessary. Therefore, the playlist can be considered information indicating what parts will be performed in what order and for how long each part will last during a live performance.

[0387] Hereinafter, the content playlist data will also be referred to as playlist data. The playlist data is generated in advance before the content distribution starts and recorded on the host PC 420. The playlist data is also shared (recorded) in advance at each venue 402.

[0388] Each time a part in the performance given by the performers in studio 401 is switched, the progress operator operates input unit 432 to specify a list item indicating the new part from among the list items in the playlist displayed on display unit 431. More specifically, an operation to instruct the switching of list items (parts) is performed.

[0389] Then, based on the signal supplied from the input unit 432, the control unit 433 generates progress information indicating the part corresponding to the list item specified by the progress operator, and outputs (supplies) it to the communication / rendering PC 414. For example, the progress information may be a part identification ID indicating the currently performed part, such as a vocal part.

[0390] <Example of Functional Configuration of Communication / Rendering PC> As described above, the communication / rendering PC 414 is configured by one or more PCs, etc. Furthermore, the communication / rendering PC 414 has, for example, the functional configuration shown in Fig. 28.

[0391] In the example shown in FIG. 28, the communication / rendering PC 414 includes a communication unit 461 , a control unit 462 , and a rendering processing unit 463 .

[0392] The communication unit 461 communicates with each venue 402 via the cloud 403. That is, the communication unit 461 transmits content data to the venue 402 (cloud 403) and receives interaction information transmitted from the venue 402.

[0393] Furthermore, for example, the communication unit 461 may communicate with each device in the studio 401 via a device such as a router. Specifically, for example, the communication unit 461 may receive progress information transmitted from the progress PC 420 and performer sound image data transmitted from the wireless device 416.

[0394] The control unit 462 controls the overall operation of the communication / rendering PC 414. For example, the control unit 462 generates facial expression data based on the video data of the facial expression video supplied from the camera 413, and supplies venue video data of each venue 23 and video data of the confirmation video to the monitor 417 and the monitor 418 to display the video.

[0395] The rendering processing unit 463 performs rendering processing based on the motion data, facial expression data, and performer sound image data of each performer, as well as image data (CG data) of the metaverse space in which the avatars corresponding to the performers are placed, to generate confirmation content data. Note that the rendering processing unit 463 may also use venue video data and venue sound image data for the rendering processing.

[0396] <Example of Venue Configuration> A venue 402 that constitutes an information processing system (content distribution system) is configured, for example, as shown in Fig. 29. In the example shown in Fig. 29, the venue 402 is a planetarium venue.

[0397] In the example shown in FIG. 29, the venue 402 has a camera 501 , a microphone 502 , a communication / rendering PC 503 , a monitor 504 , a planetarium projector 505 , planetarium sound equipment 506 , and a scripting device 507 .

[0398] The camera 501 consists of one or more cameras placed within the venue 402, and captures at least the audience seating area within the venue 402, i.e., the audience in the audience seating area, as its subject, and supplies the resulting venue video data to the communication / rendering PC 503.

[0399] The microphone 502 consists of one or more microphones placed within the venue 402, and picks up sounds within the venue 402, particularly the voices of the audience in the auditorium area within the venue 402, and supplies the resulting venue sound image data to the communication / rendering PC 503.

[0400] The communication / rendering PC 503 is made up of one or more PCs and controls the operation of the entire venue 402, that is, the entire planetarium apparatus made up of equipment installed in the planetarium venue which is the venue 402.

[0401] For example, the communication / rendering PC 503 receives (acquires) content data transmitted from the studio 401, and transmits interaction information including venue video data and venue audio image data supplied from the camera 501 and microphone 502 to the studio 401. In addition, for example, the communication / rendering PC 503 transmits to the studio 401 a control signal indicating the progress of content presentation in the venue 402.

[0402] Furthermore, for example, the communication / rendering PC 503 performs information processing based on the content data received from the studio 401 and pre-prepared venue facility information for the venue 402 to generate content data (presentation content data) in a format suitable for the facilities of the venue 402. The presentation content data is composed of at least one of video data for displaying the video of the content and sound image data for reproducing the sound of the content. Note that the presentation content data may also include haptic data.

[0403] For example, when generating content data for presentation, information processing such as rendering processing based on the content data and capture processing based on venue facility information is performed.

[0404] For example, in the rendering process, an avatar with movements and expressions corresponding to the motion data and facial expression data is placed in the metaverse space, and in the capture process, a virtual camera and a virtual microphone are placed in the metaverse space to capture video and sound, and content data for presentation is generated. In other words, content data is converted into content data for presentation.

[0405] More specifically, when generating the content data for presentation, video clipping and synthesis, downmixing and upmixing of sound image data, and audio rendering processing such as VBAP are performed as appropriate so that video and audio can be obtained in a video format and channel configuration that corresponds to venue facility information such as the dome master format. Therefore, although the content presented at each venue 402 is the same, the video data and sound image data that make up the content data for presentation may differ from venue to venue 402.

[0406] In addition, the communication / rendering PC 503 also performs processing necessary for presenting the content as appropriate, such as applying delay processing to the sound image data that constitutes the presentation content data so that the video and audio of the content based on the presentation content data are synchronized.

[0407] The monitor 504 displays various images (video) based on image data (video data) etc. supplied from the communication / rendering PC 503. For example, the monitor 504 displays video based on presentation content data, a playlist based on playlist data, etc.

[0408] An on-site operator at the venue 402 can respond to emergencies while checking the video and playlist based on the presentation content data displayed on the monitor 504. For example, if communication is interrupted, the on-site operator can have the audience see pre-prepared emergency recorded video or still images instead of the video of the metaverse space that had been displayed up until then. In other words, video switching can be performed as an emergency response.

[0409] The planetarium projector 505 corresponds to the image projection unit 52 shown in FIG. 1, and is made up of image projection equipment such as a projector for displaying (presenting) images of content.

[0410] The planetarium projector 505 displays the content image on a dome-shaped display unit (screen) installed in the venue 402 based on the video data constituting the presentation content data supplied from the communication / rendering PC 503 via a projection control unit (not shown).

[0411] The planetarium sound equipment 506 corresponds to the spatial sound output unit 53 shown in FIG. 1 and is made up of, for example, a speaker system with a multi-channel configuration.

[0412] The planetarium sound equipment 506 reproduces (outputs) the sound of the content based on the sound image data constituting the presentation content data supplied from the communication / rendering PC 503 via a reproduction control unit (not shown). The planetarium projector 505 and the planetarium sound equipment 506 function as a reproduction unit that reproduces the content.

[0413] The script device 507 controls various devices such as lighting and video projection in response to operations by the facility operator.

[0414] <Example of Functional Configuration of Communication / Rendering PC> The communication / rendering PC 503 is configured from one or more PCs and has, for example, the functional configuration shown in FIG.

[0415] In the example shown in FIG. 30, the communication / rendering PC 503 includes a communication unit 531 , an input unit 532 , a control unit 533 , a rendering processing unit 534 , and a video / audio image conversion unit 535 .

[0416] The communication unit 531 communicates with the studio 401 via the cloud 403. That is, the communication unit 531 receives content data from the studio 401 and transmits interaction information to the studio 401.

[0417] The input unit 532 is composed of, for example, a mouse, a keyboard, a switch, etc., and supplies a signal to the control unit 533 in response to an operation by an operator on-site, etc.

[0418] The control unit 533 controls the overall operation of the communication / rendering PC 503. For example, the control unit 533 displays a UI (User Interface) including a playlist based on pre-prepared playlist data on the monitor 504. The control unit 533 also functions as an output unit that supplies the video data and sound image data that make up the presentation content data to the planetarium projector 505 and the planetarium sound equipment 506.

[0419] The rendering processing unit 534 performs rendering processing based on the motion data, facial expression data, and sound image data of each performer, as well as image data (CG data) of the metaverse space in which the avatars corresponding to the performers are placed.

[0420] 1, and generates content data for presentation by performing information processing based on the rendering result by the rendering processing unit 534. The video-sound-image conversion unit 535 performs, as information processing, a conversion process for converting content data into content data for presentation in a format compatible with the venue equipment of the venue 402. More specifically, the video-sound-image conversion unit 535 performs, as information processing, the above-mentioned capture process, etc.

[0421] <Regarding Transmission and Reception of Content Data and Interaction Information> Transmission and reception of content data and interaction information will be further described with reference to Fig. 31. Note that in Fig. 31, parts corresponding to those in Fig. 27 or 29 are given the same reference numerals, and their description will be omitted as appropriate.

[0422] 31, the studio 401 is provided with new functional blocks, such as an expression application 571 (expression application program 571), a transmission application 572 (transmission application program 572), a transmission unit 573, a rendering PC 574, and a receiving PC 575. These expression application 571 to receiving PC 575 are functional blocks that make up the communication / rendering PC 414 shown in FIG.

[0423] For example, the facial expression application 571 and the transmission application 572 are functional blocks that are realized by the control unit 462 executing predetermined application programs. In particular, the transmission application 572 functions as the communication unit 461 shown in FIG.

[0424] The transmitting unit 573 and the receiving PC 575 function as the communication unit 461 shown in FIG. 28, and the rendering PC 574 functions as the rendering processing unit 463 shown in FIG.

[0425] The cloud 403 also has a server 581 that manages the connection of each venue 402 to the metaverse space.

[0426] Furthermore, the venue 402 is provided with new functional blocks: a sending PC 591, an on-site operator PC 592, a rendering PC 593, a receiving unit 594, a projection control unit 595, and a playback control unit 596. In particular, the sending PC 591 to the receiving unit 594 are functional blocks that make up the communication / rendering PC 503 shown in FIG.

[0427] The sending PC 591 and the receiving unit 594 function as the communication unit 531 shown in Fig. 30, and the on-site operator PC 592 functions as the input unit 532 and the control unit 533 shown in Fig. 30. Furthermore, the rendering PC 593 functions as the communication unit 531, the control unit 533, the rendering processing unit 534, and the video / audio image conversion unit 535 shown in Fig. 30.

[0428] The projection control unit 595 and the playback control unit 596 are provided between the communication / rendering PC 503 in FIG. 29 and the planetarium projector 505 and the planetarium sound equipment 506 .

[0429] In the example shown in FIG. 31, when the motion acquisition system 412 acquires the motion data of each performer, the motion data is supplied to the facial expression application 571 .

[0430] The facial expression application 571 supplies the motion data supplied from the motion acquisition system 412 and the facial expression data obtained by the control unit 462 of the communication / rendering PC 414 to the transmission application 572 .

[0431] In this case, facial expression data may be added to the motion data, i.e., the motion data and facial expression data may be associated with each other for each performer and supplied to the transmission application 572. The motion data may also include data indicating the performer's facial expression. In such a case, the facial expression application 571 may replace (overwrite) the data portion of the motion data indicating the performer's facial expression with the facial expression data, thereby generating motion data including the facial expression data, and supplying the motion data to the transmission application 572.

[0432] The transmission application 572 adds progress information indicating the part currently being performed, which is supplied from the proceeding PC 420, to the motion data and facial expression data of each performer supplied from the facial expression application 571. The transmission application 572 also transmits the motion data and facial expression data with the added progress information to the server 581 that constitutes the cloud 403 as data that constitutes part of the content data.

[0433] Additionally, in the studio 401, performer sound image data obtained by the headset 411 is supplied to the control unit 462 of the communication / rendering PC 414 via the wireless device 416 and the mixer 415. The control unit 462 adds a time code supplied from the motion acquisition system 412 to the performer sound image data of each performer, and then transmits the performer sound image data to the cloud 403 via the transmission unit 573. By adding a time code to the performer sound image data, it is possible to synchronize the performer sound image data and motion data of each performer.

[0434] In addition, along with the motion data or performer sound image data transmitted from studio 401, the above-mentioned sound image data of the performance sound and video data of the performance effects may also be transmitted to venue 402 via cloud 403 by communication unit 461.

[0435] The server 581 of the cloud 403 receives the motion data, facial expression data, and progress information transmitted from the transmission application 572 of the studio 401, and transmits them to the venue 402. The cloud 403 also receives the performer sound image data transmitted from the transmission unit 573 of the studio 401, and transmits them to the venue 402. For example, the performer sound image data is transmitted to the venue 402 without going through the server 581.

[0436] The rendering PC 593 in the venue 402 receives the motion data, facial expression data, and progress information transmitted from the server 581 in the cloud 403. The receiving unit 594 in the venue 402 also receives the performer sound image data transmitted from the cloud 403 and supplies it to the rendering PC 593.

[0437] The rendering PC 593 performs rendering and capture processes based on the received motion data, facial expression data, and progress information, and the performer sound image data supplied from the receiving unit 594, and generates content data for presentation.

[0438] At this time, the rendering PC 593 communicates with the server 581 as appropriate to identify the metaverse space, or more specifically, the world within the metaverse space, that corresponds to the part indicated by the progress information. For example, the rendering PC 593 stores in advance, for each world, background image data (CG data) and the like for displaying the backgrounds, objects, and the like of that world.

[0439] The rendering PC 593 transmits progress information to the server 581, and receives, as a response from the server 581, information indicating the world in which an avatar or the like is present when the part indicated by the progress information is being performed, thereby identifying the world corresponding to the current part.

[0440] The rendering PC 593 uses the background video data of the identified world for rendering, i.e., for generating content data for presentation, thereby enabling the presentation of content in which an avatar performs in an appropriate world defined for each part.

[0441] The rendering PC 593 supplies the video data constituting the presentation content data to the planetarium projector 505 via the projection control unit 595, causing the image of the content to be displayed. At the same time, the rendering PC 593 supplies the sound image data constituting the presentation content data to the planetarium sound equipment 506 via the playback control unit 596, causing the sound of the content to be output (played back).

[0442] Furthermore, for example, the rendering PC 574 in the studio 401 acquires motion data and performer sound image data from the server 581 or the like, performs rendering processing, etc., and presents content based on the resulting confirmation content data. That is, the rendering PC 574 supplies video data constituting the confirmation content data to the monitor 418 to display the video. Also, for example, the rendering PC 574 appropriately supplies sound image data constituting the confirmation content data to the headset 411 via the mixer 415 and wireless device 416 to present sound to the performer, or supplies the sound image data to speakers in the studio 401 to play the sound.

[0443] In addition, in a part where the performers interact with the audience in the venues 402, such as a call and response, interaction information is transmitted from all venues 402 or from a specific venue 402 to the studio 401.

[0444] In such a case, the transmitting PC 591 generates interaction information including venue video data from the camera 501, venue sound image data from the microphone 502, and pre-recorded venue identification information, and transmits it to the studio 401 via the cloud 403.

[0445] In the studio 401, a receiving PC 575 receives the interaction information transmitted from the venue 402 via the cloud 403. The receiving PC 575 supplies venue image data included in the interaction information to a monitor 417 to display an image of the venue 402, and supplies venue sound image data included in the interaction information to a headset 411 via a mixer 415 and a wireless device 416 to present sounds from the venue 402 to the performers. In addition, sounds based on the venue sound image data may be reproduced on speakers in the studio 401, and sounds from the venue 402 may be presented to a production director / manager or the like.

[0446] In addition, data such as information indicating the level of excitement in the venue 402 and video and sound effects corresponding to the level of excitement, i.e., the above-mentioned reflection information, may be transmitted from the receiving PC 575 to the transmitting PC 591 or rendering PC 593 of each venue 402 via the server 581. In this case, the reflection information may be data (information) obtained by processing the receiving PC 575 (control unit 462) based on interaction information from one or more venues 402, such as video data of video based on venue video data or audio image data of audio based on venue audio image data, data for adding effects (video data, audio image data, etc.) that is added or changed according to the results of the above-mentioned analysis processing, video data of virtual audience images according to the results of the analysis processing, or data indicating a change in the viewpoint position of the content video. By receiving the reflection information at the venue 402, each venue 402 can reflect the effects, etc., corresponding to the excitement level of each venue 402 in the content data to be presented.

[0447] Also, suppose an emergency response becomes necessary during playback of content at the venue 402, and the on-site operator operates the on-site operator PC 592. In this case, a control signal indicating the progress of content presentation is supplied from the on-site operator PC 592 to the rendering PC 593. The rendering PC 593 then switches the video and audio presented at the venue 402 to pre-recorded data as appropriate. In other words, the above-mentioned cover image or the like is displayed.

[0448] In this case, a control signal for emergency response is transmitted from the communication unit 531 of the communication / rendering PC 503 to the studio 401 via the cloud 403 .

[0449] <UI Display Example> A display example of a UI including a playlist displayed in the studio 401 and the venue 402 will be described.

[0450] In the host PC 420 in the studio 401, for example, a UI screen (display screen) shown in FIG.

[0451] In the example of FIG. 32, the current time is displayed together with a clock mark at the top center of the display screen, and a playlist is displayed on the left side of the drawing.

[0452] In the playlist, the parts are arranged in the order in which they are to be performed, i.e., in chronological order. Each list item, represented by a rectangle on the playlist, is marked with a letter indicating the part that corresponds to that list item.

[0453] For example, list item LI11 has the characters "OP~1st song" written on it, indicating that it is the first song part, which will be performed from the opening (OP) to the first song of the live performance.

[0454] For example, possible parts to be displayed in the list items include a singing part where songs are sung, an MC part where MCs are present, an audience entry cover image part where an image (video) for admitting the audience is displayed as the audience enters the venue 402, an end credits part where an end credits is displayed at the end of the live performance, an audience exit cover image part where an image (video) for escaping the audience is displayed as the audience leaves after the live performance has ended, and an emergency cover image part where replacement video is displayed in the event of an emergency (trouble).

[0455] The on-site operator can control the playback and stopping of the playlist by operating the input unit 432 of the host PC 420. Furthermore, by operating the playlist or other buttons on the display screen, it is possible to execute a fail-safe function in the event of an emergency such as a network failure. The emergency fail-safe function can also be applied to each venue 402, for example, to replace the MC part with a recorded video. It is also possible to perform on-site playlist operations, such as displaying an emergency cover image at a specific venue 402.

[0456] When playback of a playlist begins, the part currently being played is highlighted. For example, in the example of Fig. 32, list item IL11 is highlighted, indicating that the opening and first song are currently being sung in studio 401.

[0457] At this time, the control unit 433 of the proceeding PC 420 outputs the part identification ID indicating the highlighted part and other information as progress information to the transmission application 572. The progress information is transmitted to the venue 402 via the cloud 403 as data constituting the content data.

[0458] The on-site operator operates the input unit 432 of the proceeding PC 420 and operates button BT11 on the display screen to switch between list items that are in a selected state, that is, that are highlighted. For example, if button BT11 is operated while list item IL11 is highlighted, the highlighting state of list item IL11 is canceled, and the next list item, that is, the list item located below list item IL11 in the figure, is highlighted.

[0459] The parts corresponding to each list item may present an image (still image) and background music, an image only, a video (video output), an MC (call and response), etc. In other words, the type of processing to be performed in each part is predetermined, and when creating a playlist, each part is specified according to the content of the performance to be performed at a live concert, etc.

[0460] For example, in a list item that specifies a part for which video output is to be performed, such as a singing part, the total time (duration) of the part is displayed below the diagram in the list item. Also, when the list item is being played (the part is being performed), in other words, when it is highlighted, the current timeline, i.e., the time elapsed since the start of the part, is displayed in addition to the total time.

[0461] In the example of FIG. 32, for example, the time required until the end of the part (total duration) and the current elapsed time are displayed below list item IL11.

[0462] Furthermore, when the elapsed time of the highlighted list item becomes the same as the time of the entire part and the end time of the part is reached, playback of the next list item begins, i.e., the next list item is highlighted.

[0463] For example, in a part where video output is performed, such as a singing part, motion data, facial expression data, performer sound image data, and the like are transmitted to each venue 402 as content data.

[0464] For list items that specify parts in which images and background music are presented, content data such as video data, e.g., still images, and sound image data, e.g., background music, is transmitted to each venue 402. In each venue 402, loop playback is performed in which the background music is repeatedly played until the end of the part, even if playback of the background music has finished.

[0465] During the MC part, a call and response is performed connecting the studio 401 and each venue 402. In this case, the MC part list item displays the elapsed time since the MC part began in the lower part of the list item.

[0466] During the MC part, content data such as motion data, facial expression data, and performer sound image data is transmitted to each venue 402, and interaction information is transmitted from each venue 402 to the studio 401.

[0467] Furthermore, if an emergency occurs in the studio 401, the on-site operator can also switch the playlist item to be played back.

[0468] For example, in the state shown in Fig. 32, if a list item other than the currently highlighted list item IL11, i.e., the currently being played list item IL11, is selected by clicking or the like, the selected list item is highlighted. At this time, the currently being played list item IL11 remains highlighted.

[0469] In this way, when a list item that is not currently being played is selected (highlighted), the apply button BT12 on the display screen is activated and becomes operable (active).

[0470] When an on-site operator or the like presses (operates) the apply button BT12, a selected list item that is not currently being played becomes the new playback target, and playback of that list item begins. In this case, data (video data and sound image data) corresponding to the list item whose playback has begun is transmitted to the venue 402 as content data.

[0471] Therefore, if a problem occurs during playback of list item IL11, for example, playback can be switched to the next list item, "MC1," and the MC part can be played. In this case, it is also possible to return to the song part indicated by list item IL11 during or after the MC part has finished.

[0472] The on-site operator or the like can close the display screen (UI) including the playlist by operating the close button BT13 on the display screen, and can terminate the playlist control application.

[0473] For example, when the close button BT13 is operated, a dialogue screen shown in FIG. 33 is displayed on the display unit 431 of the proceeding PC 420.

[0474] When the button marked with the letters "OK" on the dialog screen shown in FIG. 33 is operated, the screen shown in FIG. 32 is closed and the playlist control application is terminated.

[0475] On the other hand, when the button marked with the word "Cancel" on the dialog screen shown in FIG. 33 is operated, the dialog screen is closed and the screen shown in FIG. 32 continues to be displayed.

[0476] Furthermore, on the display screen shown in FIG. 32, it is possible to set up to a predetermined number of list items in a playlist, such as 100, and when not all list items can be displayed in the playlist, a scroll bar is displayed at the end of the playlist.

[0477] In this case, for example, the list is automatically scrolled each time the button BT11 is operated so that the highlighted, that is, activated, list item can be seen.

[0478] In addition, the content data may be faded in and out when the list item to be played is switched.

[0479] In this case, processing is performed based on the video data and sound image data so that the video and audio of the part corresponding to the list item before the switch fades out just before the switch, and the video and audio of the part corresponding to the list item after the switch fades in just after the switch.

[0480] At this time, the fade-in and fade-out processing may be performed by the control unit 462 or the like on the studio 401 side, or may be performed by the control unit 533 on the venue 402 side. In particular, when motion data is transmitted as content data, the fade-in and fade-out processing is performed on the venue 402 side. Furthermore, the length of the period during which the fade-in and fade-out are performed, that is, the length of the period during which the video and audio transition, may be set for each list item on the playlist.

[0481] Alternatively, it may be possible to set whether or not to perform fade-in and fade-out, or fade-in and fade-out may be performed according to some conditions.

[0482] The same playlist as that displayed in the studio 401 is also displayed on the venue 402 side. Specifically, for example, the on-site operator PC 592 in the venue 402 displays a UI screen (display screen) shown in FIG.

[0483] In FIG. 34, the current time and the current communication status are displayed together with a clock mark at the top center of the display screen, and a playlist is displayed on the left side of the drawing.

[0484] In particular, in this example, the communication status is "good," indicating that the communication status between the venue 402 and the studio 401 is good. When the communication status is good, for example, the studio 401 can control the content playback in the venue 402.

[0485] On the other hand, if the communication state between the venue 402 and the studio 401 is not good, that is, if the venue 402 and the studio 401 cannot communicate with each other, the studio 401 may be unable to control the content playback in the venue 402. In such a case, the communication state is displayed as "not connected" on the display screen.

[0486] In the drawing, the playlist shown on the left side is the same as the playlist displayed on the studio 401 side shown in Fig. 32, and basically the same list items are highlighted as on the studio 401 side. In this example, list item IL31 corresponding to list item IL11 in Fig. 32 is highlighted.

[0487] In addition, in the example of Fig. 34, there are provided buttons BT31, BT32, and BT33, which correspond to the buttons BT11, BT12, and BT13 in Fig. 32. Furthermore, in the example of Fig. 34, there are also provided buttons BT34 and BT35 for emergency response.

[0488] For example, suppose that an MC part is being performed in studio 401 and content data related to that MC part is being sent (transmitted) to each venue 402, but then communication becomes impossible for some reason, i.e., the network connection becomes impossible.

[0489] In such a case, when an on-site operator at the venue 402 presses (operates) the emergency response button BT34, the video and audio of the MC part that had been presented to the audience at the venue 402 is switched to emergency cover video and audio prepared in advance for emergencies.

[0490] For example, the video and audio of the MC part fade out, then the emergency cover image video and audio fade in, and the emergency cover image video and audio are repeatedly played (looped).

[0491] In particular, while the emergency cover image video is being played, button BT34 is in an inactive state where it cannot be operated, and button BT35 is in an active state where it can be operated. Note that button BT35 is in an active state only when button BT34 is operated, that is, when the emergency cover image video is being played, and is in an inactive state at other times.

[0492] In addition, even while the emergency cover image video is being played, if the venue 402 and the studio 401 are able to communicate, i.e., if the communication conditions are good, the venue 402 can receive the content data transmitted from the studio 401.

[0493] Therefore, if the communication state is good even while the emergency cover image video is being played, the highlight display in the playlist moves in sync with the highlight display on the playlist displayed in studio 401. That is, the emergency cover image video is being presented in venue 402, but the highlighted list items are switched in the same way as the playlist in studio 401. In this way, the on-site operator in venue 402 can know which parts are being performed in venues other than the one in which he or she is present.

[0494] When the emergency situation such as a network disconnection is resolved and content presentation becomes possible in the venue 402, the on-site operator presses (operates) the button BT35, which resumes content presentation based on the content data received from the studio 401. That is, the rendering PC 593 performs rendering processing and the like based on the content data received from the studio 401 to generate content data for presentation, and presents the content based on the content data for presentation.

[0495] When the close button BT33 is operated, a dialogue screen shown in FIG.

[0496] When the button marked with the letters "OK" on the dialog screen shown in FIG. 35 is operated, the screen shown in FIG. 34 is closed and the playlist control application is terminated.

[0497] On the other hand, when the button marked with the word "Cancel" on the dialog screen shown in FIG. 35 is operated, the dialog screen is closed and the screen shown in FIG. 34 continues to be displayed.

[0498] As described above, the same display screen including the playlist is displayed on the studio 401 side and the venue 402 side. Note that the display screen shown in Fig. 34 may also be displayed on the studio 401 side. In such a case, however, the emergency response buttons BT34 and BT35 are always inactive.

[0499] <Description of Distribution Processing> The operation of the information processing system (content distribution system) described with reference to FIGS. 27 to 31 will now be described.

[0500] First, the distribution process by the studio 401 will be described with reference to the flowchart of Fig. 36. This distribution process starts when content distribution starts.

[0501] In step S 201 , the motion acquisition system 412 acquires motion data indicating the movements of the performer and supplies it to the facial expression application 571 , and also supplies the time code of the motion data to the control unit 462 of the communication / rendering PC 414 .

[0502] In step S202, the headset 411 picks up the sound of the performer and acquires sound image data of the performer (performer sound image data).

[0503] The performer sound image data acquired by the headset 411 is supplied to the control unit 462 of the communication / rendering PC 414 via the wireless device 416 and the mixer 415. The control unit 462 adds a time code supplied from the motion acquisition system 412 to the performer sound image data of each performer, and supplies the data to the transmission unit 573 (communication unit 461).

[0504] The control unit 462 may identify the position of each performer, more specifically, the avatar corresponding to each performer, in the metaverse space based on video data obtained by the camera 413, and use the identification result as object position information indicating the sound image localization position of the sound based on the performer sound image data. In such a case, the control unit 462 adds a time code and object position information to the performer sound image data and supplies the data to the transmission unit 573. Alternatively, the object position information may be input by, for example, the facial expression operator operating the facial expression controller 419, or may be predetermined for each performer.

[0505] In step S203, the facial expression application 571 acquires facial expression data of each performer.

[0506] For example, the facial expression application 571 generates facial expression data for each performer by performing analytical processing based on the video data of the facial expression video supplied from the camera 413, or acquires facial expression data for each performer output from the facial expression controller 419 in response to the operation of the facial expression operator.

[0507] The facial expression application 571 associates the motion data supplied from the motion acquisition system 412 with the facial expression data obtained in step S203 and supplies the data to the transmission application 572.

[0508] In step S204, the communication unit 461 of the communication / rendering PC 414 distributes the content data.

[0509] That is, the transmission application 572 functioning as the communication unit 461 associates the motion data and facial expression data supplied from the facial expression application 571 with the progress information supplied from the control unit 433 of the host PC 420. The transmission application 572 then transmits the motion data, facial expression data, and progress information to the venue 402 via the cloud 403 (server 581). The transmission unit 573 functioning as the communication unit 461 also transmits the performer sound image data supplied from the control unit 462 to the venue 402 via the cloud 403.

[0510] In this case, for example, when the display screen (playlist) shown in FIG. 32 is displayed and the progress operator operates button BT11 to switch the list item being played, or when playback of a list item ends and the list item to be played switches, the control unit 433 of the progress PC 420 supplies progress information indicating the part corresponding to the list item after the switch to the transmission application 572.

[0511] Note that sound image data of the sound effects and video data of the performance effects may also be transmitted as content data by the communication unit 461. In this case, the sound image data of the sound effects and video data of the performance effects may be used to reflect, for example, the above-mentioned reflection information, such as the state of excitement in each venue 402, in the content.

[0512] Furthermore, in order to respond to an emergency or the like that occurs in studio 401, when the progress operator selects any list item and operates apply button BT12, video data of an emergency cover image or the like is transmitted. Specifically, for example, communication unit 461, i.e., at least one of transmission application 572 and transmission unit 573, transmits video data for the cover image of the part corresponding to the selected list item, sound image data such as background music, and the like to venue 402 via cloud 403. In this case, progress information is also transmitted together with the video data and the like.

[0513] In addition, for example, when the list item being played is a part for which motion data and performer sound image data are not transmitted, such as a cover image part for admitting the audience, the communication unit 461 transmits video data etc. corresponding to that part and progress information to the venue 402 via the cloud 403.

[0514] In step S205, the rendering PC 574, that is, the rendering processing unit 463 of the communication / rendering PC 414, performs rendering processing to generate confirmation content data.

[0515] For example, the rendering process uses the motion data, facial expression data, and performer sound image data of each performer, as well as image data (CG data) of the metaverse space in which the avatars corresponding to the performers are placed, which were sent in step S204 and transmitted from the cloud 403. Note that the motion data, facial expression data, and performer sound image data used in the rendering process are not limited to those received by the rendering PC 574 or the like from the cloud 403, but may be those obtained in steps S201 to S203.

[0516] In step S206, the rendering PC 574 (control unit 462) supplies the video data constituting the confirmation content data obtained in step S205 to the monitor 418 to display the video (rendered video). In this case, sound based on the sound image data constituting the confirmation content data may also be presented to the performers or the production director / manager.

[0517] In step S207, the control unit 462 of the communication / rendering PC 414 determines whether or not to end the process of distributing the content data.

[0518] If it is determined in step S207 that the distribution is not yet to be ended, the process returns to step S201, and the above-described processes are repeated.

[0519] On the other hand, if it is determined in step S207 that the distribution is to be ended, the control unit 462 stops the processing of each unit in the studio 401, and the distribution processing ends.

[0520] In this way, the studio 401 transmits the motion data, facial expression data, performer sound image data, and progress information required for presenting the content. In this way, each venue 402 can perform rendering processing based on the motion data and performer sound image data, and obtain content data for presentation that is suited to the equipment of each venue 402. This allows each venue 402 to present content appropriately, regardless of the equipment of the venue 402.

[0521] <Explanation of Venue Video Display Processing> Furthermore, for example, when an MC part, i.e., a call and response, is being performed, an exchange takes place between the performer in studio 401 and the audience in venue 402. In such a case, in studio 401, the venue video display processing shown in Fig. 37 is performed.

[0522] Hereinafter, the venue image display process by the studio 401 will be described with reference to the flowchart of FIG.

[0523] In step S 231 , the communication unit 461 of the communication / rendering PC 414 , that is, the receiving PC 575 , receives the video data and sound image data of the venues 402 transmitted from each venue 402 .

[0524] Specifically, the receiving PC 575 receives interaction information transmitted from the venue 402, including venue image data, venue sound image data, and venue identification information.

[0525] In step S232, the control unit 462 of the communication / rendering PC 414, that is, the receiving PC 575, presents images and sounds of each venue 402 based on the received interaction information.

[0526] Specifically, the receiving PC 575 supplies the venue video data included in the interaction information to the monitor 417 to display the video of the venue 402. The receiving PC 575 also supplies the venue sound image data included in the interaction information to the headset 411 via the mixer 415 and the wireless device 416, or to a speaker (not shown), thereby reproducing the sound of the venue 402. Note that it is possible to identify which venue 402 each piece of venue video data etc. belongs to by using the venue identification information.

[0527] Once the images and sounds from each venue 402 have been presented, the venue image display process ends.

[0528] In this way, the studio 401 receives interaction information from each venue 402 and displays the images and sounds of the venue 402. This allows the performers to know the reactions of each venue 402 in real time. This allows for exchanges such as call and response, making the live performance even more exciting.

[0529] <Description of Distribution Processing> When the distribution processing of Fig. 36 is started, the playback processing shown in Fig. 38 is started at each venue 402. The playback processing at the venue 402 will be described below with reference to the flowchart of Fig. 38.

[0530] In step S271, the communication unit 531 of the communication / rendering PC 503 in the hall 402 receives the content data transmitted from the studio 401 in step S204 of FIG.

[0531] That is, the rendering PC 593 serving as the communication unit 531 receives motion data, facial expression data, and progress information as content data. At this time, if video data such as reflection information and dramatic effects is also transmitted as content data, the rendering PC 593 also receives the video data.

[0532] Furthermore, the rendering PC 593 transmits the received progress information to the server 581 of the cloud 403, and by receiving a response to the transmission of the progress information, identifies the world in the metaverse space corresponding to the part indicated by the progress information, and reads background video data (CG data) of that world, etc. Note that the background video data, etc. may be pre-recorded in the rendering PC 593, or may be obtained from an external device such as the server 581.

[0533] Furthermore, the receiving unit 594 serving as the communication unit 531 receives performer sound image data as content data and supplies it to the rendering PC 593. At this time, when reflection information or sound image data of performance sound is also transmitted as content data, the receiving unit 594 also receives the sound image data and supplies it to the rendering PC 593.

[0534] In step S272, the rendering PC 593, that is, the rendering processing unit 534, performs rendering processing based on the content data received in step S271.

[0535] The rendering process here mainly refers to the rendering process of video. For example, the rendering processing unit 534 generates a video of the entire world (metaverse space) in which avatars with predetermined movements and expressions are placed, based on content data such as motion data, facial expression data, and background video data of the world. At this time, venue video and performance video based on reflection information and venue video data may be placed (displayed) at predetermined positions in the world, as appropriate. The data (video data) of the entire world obtained by the rendering process of video based on motion data, background video data, etc. may be 3D data or video data of a celestial sphere obtained from the 3D data, as long as it can display the entire world. In addition, sound (sound image) based on performer sound image data is also placed at the position of an avatar, etc. in the world. The placement position of the performer's sound is, for example, an absolute position in the world indicated by object position information added to the performer sound image data. Note that the video data referred to in this specification also includes 3D data of the entire world obtained by the rendering process of video based on motion data, background video data, etc.

[0536] In step S273, the rendering PC 593, i.e., the video / audio conversion unit 535, generates content data for presentation by performing conversion processing on the content data obtained in step S272 based on the venue facility information of the venue 402. Here, information processing including video and audio capture processing is performed as the conversion processing.

[0537] Specifically, for example, based on the venue facility information, the video / audio conversion unit 535 identifies the placement positions of the virtual camera and virtual microphone for each venue 402. In this case, the placement positions of the virtual camera and the virtual microphone are basically set to the same positions.

[0538] The image-sound-image conversion unit 535 places a virtual camera at a specified position on the image of the world based on the 3D data or the image data of the omnidirectional image obtained in step S272, and performs a capture process with the position of the virtual camera as the viewpoint position, thereby obtaining image data of the world as seen from the virtual camera.

[0539] At this time, the image-sound-image conversion unit 535 generates video data of a hemispherical, full-dome image as seen from the position of the virtual camera by cutting out a portion of the world image and creating a cut-out image based on the venue equipment information, or by arranging and combining multiple cut-out images, and uses this video data as video data that constitutes the content data to be presented.

[0540] In addition, the image-sound-image conversion unit 535 places a virtual microphone at the specified position on the image of the world obtained in step S272, and performs a capture process with the position of the virtual microphone as the listening position, thereby obtaining sound image data of the sound of the world that is heard at the position of the virtual microphone.

[0541] Specifically, for example, the video / audio-image conversion unit 535 calculates avatar position information indicating the relative position of the avatar as seen from the position of the virtual microphone for each performer, based on the position of the virtual microphone and the position of each performer's avatar in the world indicated by the object position information.The video / audio-image conversion unit 535 then generates sound image data of the presentation content data by performing audio rendering processing for VBAP or the like, based on the venue facility information, avatar position information, and performer sound image data.

[0542] The sound image data of the presentation content data is data with the same channel configuration as that of the speaker system serving as the planetarium sound equipment 506 installed in the venue 402. Note that the sound image data constituting the presentation content data may be generated by capturing a sound image according to the position of a virtual microphone and performing processing such as downmixing or upmixing on the sound image data obtained by the capture.

[0543] Furthermore, if reflection information or sound image data of the performance sound is also received as content data, that sound image data is also used to generate the sound image data of the content data for presentation. In this case, the sound based on the sound image data of the content data for presentation includes not only the sounds of the performers and the accompaniment of the music, but also the sounds (voices) of the audience in other venues 402 and the performance sound. When processing based on the reflection information is performed during generation of the content data for presentation, such as in step S272 or step S273, the rendering PC 593 performs processing similar to the processing performed in the planetarium image sound-image conversion unit 44 of the first embodiment.

[0544] In step S274, the rendering PC 593, that is, the control unit 533, reproduces the content based on the presentation content data.

[0545] That is, the control unit 533 supplies the video data and sound image data that constitute the presentation content data to the projection control unit 595 and the playback control unit 596 .

[0546] The projection control unit 595 supplies the video data supplied from the control unit 533 to the planetarium projector 505, which then outputs light corresponding to the video data (projects an image), thereby displaying the video of the content on a dome-shaped screen in the venue 402. The playback control unit 596 supplies the sound image data supplied from the control unit 533 to the planetarium sound equipment 506, which then outputs sound based on the sound image data from the speakers.

[0547] While the content is being presented in the venue 402, the control unit 533 of the on-site operator PC 592, i.e., the communication / rendering PC 503, displays, for example, a display screen (playlist) shown in Figure 34 on the monitor 504 based on the playlist data.

[0548] For example, if an on-site operator operates button BT34 to respond to an emergency that occurs in venue 402 during content presentation, an emergency cover image or the like is presented. That is, when button BT34 on a display screen including a playlist is operated, the control unit 533 (rendering PC 593) as an output unit temporarily stops outputting the presentation content data to the planetarium projector 505 and planetarium sound equipment 506. Then, the control unit 533 outputs video data (image data) and sound image data of an emergency cover image or the like that have been prepared in advance to the planetarium projector 505 and planetarium sound equipment 506.

[0549] Specifically, the control unit 533 supplies video data such as a cover image for emergency use prepared in advance to the planetarium projector 505 via the projection control unit 595, and displays the cover image based on the video data. The control unit 533 also supplies sound image data such as background music for emergency use prepared in advance to the planetarium sound equipment 506 via the playback control unit 596, and outputs sound based on the sound image data. The images presented in an emergency response are not limited to cover images, and may be pre-recorded parts of video based on pre-recorded data prepared in advance. Furthermore, switching of images during an emergency response may be performed not only by operating the button BT34, but also by operating the apply button BT32.

[0550] When the emergency response is completed and the button BT35 is operated, the control unit 533 (rendering PC 593) causes the rendering processing unit 534 and the video / audio image conversion unit 535 to execute processing to generate content data for presentation based on the received content data. As a result, video and sound based on the content data for presentation are once again presented.

[0551] In step S275, the control unit 533 determines whether or not to end the process of receiving the content data and presenting the content.

[0552] If it is determined in step S275 that the process is not yet finished, the process returns to step S271, and the above-described process is repeated.

[0553] On the other hand, if it is determined in step S275 that the processing is to be ended, the control unit 533 stops the processing of each unit in the venue 402, and the playback processing ends.

[0554] In this way, the venue 402 receives the content data, generates content data for presentation from the content data according to the venue facility information, and presents the content. In this way, the content can be presented appropriately at each venue 402 regardless of the venue's 402 facilities.

[0555] <Description of Venue Video Transmission Process> When communication takes place between the performer in the studio 401 and the audience in the venue 402, the venue video transmission process shown in FIG. 39 is carried out in each venue 402.

[0556] Hereinafter, the venue video transmission process by the venue 402 will be described with reference to the flowchart of FIG.

[0557] In step S 311 , the control unit 533 of the sending PC 591 , that is, the communication / rendering PC 503 , acquires the video data and sound image data of the venue 402 .

[0558] Specifically, the control unit 533 acquires venue video data from the camera 501 and venue sound image data from the microphone 502. The control unit 533 also generates interaction information including the acquired venue video data and venue sound image data, as well as pre-recorded venue identification information. Note that the control unit 533 may perform the same processing as the scene analysis unit 54 in the first embodiment, i.e., an analysis processing based on at least one of the venue video data and the venue sound image data, and generate interaction information including information obtained as a result of the processing.

[0559] In step S312, the sending PC 591, i.e., the communication unit 531 of the communication / rendering PC 503, transmits the video data and sound image data of the venue 402 to the studio 401 via the cloud 403. Specifically, the communication unit 531 transmits the interaction information obtained in step S311. Once the interaction information is transmitted, the venue video transmission process ends.

[0560] In this way, the venue 402 acquires video data and audio image data of the audience within the venue 402 and transmits them as interaction information to the studio 401. In this way, the studio 401 can know the reactions of each venue 402 in real time. This allows for exchanges such as call and response, which can further liven up a live performance.

[0561] <Processing According to Venue Facility Information> The content correction process performed by the planetarium image / sound converter 44 in the first embodiment and the process of generating presentation content data performed by the communication / rendering PC 503 in the second embodiment can be considered to be information processing for generating content data from original content data in a format according to venue facility information. In other words, both the planetarium image / sound converter 44 and the communication / rendering PC 503 (the rendering processor 534 and the image / sound converter 535) function as generators that generate different content data with the same content content by processing the original content data based on venue facility information. Note that, even in the first embodiment, motion data and facial expression data may be transmitted to the cloud 22 as data related to the video that constitutes the content data. In other words, the video data that constitutes the content data may be motion data and facial expression data. In such a case, rendering processing is performed in the planetarium image / sound converter 44, and the video data of the content is generated.

[0562] For example, in the first embodiment, it is possible to hold a real live performance on a real stage, etc., and at the same time, distribute to each venue 23 a virtual live performance corresponding to that live performance, that is, a live performance in the metaverse space, as content.

[0563] In the first embodiment, for example, as shown by arrow Q101 in FIG. 40, the metaverse login PC 21 captures the video of the content.

[0564] That is, a virtual camera facing a predetermined direction is placed at a predetermined position in the Metaverse space, and the video (video data) of the content is captured from the viewpoint of the virtual camera. For example, the position and orientation of the virtual camera are specified by coordinates in a Cartesian coordinate system and rotation angles such as the roll angle, pitch angle, and yaw angle of the virtual camera.

[0565] By capturing with the virtual camera, video data of a celestial sphere video in all directions, up, down, left, and right, can be obtained, for example, as indicated by arrow Q102.

[0566] Then, in the cloud 22 , the planetarium image / sound image converter 44 performs content correction processing in accordance with the venue facility information of each venue 23 , and corrected content data is generated for each venue 23 .

[0567] As a result, as shown by arrow Q103, hemispherical full-dome video data that matches the facilities of venue 23, such as the inclination angle of the dome portion, is obtained as video data that constitutes the corrected content data.

[0568] That is, a partial area of ​​the omnidirectional image is cut out, and a omnidirectional image is generated from the cut-out image. Here, the omnidirectional image for a horizontal dome is shown on the upper side of the figure, and the omnidirectional image for a tilted dome is shown on the lower side of the figure.

[0569] As described above, in the first embodiment, a celestial sphere image is first generated by capturing with the virtual camera, and then, in the cloud 22, a part of the celestial sphere image is cut out in accordance with the venue facility information, and a celestial sphere image for each venue 23 is generated.

[0570] In contrast to this, in the second embodiment, the video of the content is captured in each venue 402, as shown by arrow Q111 in FIG.

[0571] In this case, a virtual camera facing a predetermined direction is placed at a predetermined position in the metaverse space according to venue facility information of venue 402, and video (video data) of the content is captured from the viewpoint of the virtual camera. In this case as well, the position and orientation of the virtual camera are specified by, for example, coordinates in a Cartesian coordinate system and the rotation angle of the virtual camera.

[0572] In particular, in the second embodiment, capturing is performed by a virtual camera in each venue 402, and therefore, unlike the first embodiment, it is not necessary to first capture an all-directional spherical image and then extract an area from the spherical image that corresponds to the venue facilities as a panoramic image.

[0573] Therefore, in the second embodiment, a virtual camera captures an area of ​​the entire metaverse space corresponding to the facilities of venue 402, as shown by arrow Q112, and video data of the full-dome image is obtained from the image of that area (cut-out image).

[0574] The venue facility information used to generate video data according to the facilities of venue 23 or venue 402 may be any information about the venue, such as venue 23 or venue 402, that is, information about the venue facilities.

[0575] For example, the venue facility information may consist of one or more pieces of information relating to the venue facilities, such as the angle information indicating the inclination (tilt angle) of the dome portion described above.

[0576] Furthermore, for example, the venue facility information may be identification information relating to the venue, such as an ID indicating one of a plurality of groups according to the venue's facilities, or an ID indicating the venue that is assigned to the venue itself.

[0577] In the following, a case where the venue facility information is identification information relating to venue facilities, that is, a case where the venue facility information is an ID indicating a group corresponding to the venue facilities, will be described.

[0578] When the venue facility information is identification information, the characteristics of the facilities at each venue 23 (venue 402) are identified by measurement or the like, for example, by inspecting the venue 23 (or venue 402) in advance, and classification is performed according to the identification results. In other words, multiple venues 23 (venues 402) are grouped based on the characteristics of the venue facilities.

[0579] The characteristics of the venue equipment here refer to, for example, the diameter and tilt angle of the dome part, the video format that can be displayed, the seating position, the seating layout type, the channel configuration of the speaker system, and the placement position of each speaker that makes up the speaker system.

[0580] By dividing the venues into groups in this way, each venue 23 (venue 402) will belong to one of the multiple groups. Each group (hereinafter also referred to as a facility group) is assigned identification information (ID) that identifies the group as venue facility information.

[0581] Next, for each equipment group, the details of the processing (processing details) to be performed when generating corrected content data (or presentation content data) from the content data, more specifically, parameters for performing each process such as capture processing, are determined. The parameters here include, for example, the angle at which the image is cut out during the capture processing, in other words, the camera angle of the virtual camera.

[0582] For example, when the process content (processing parameters) is determined for each equipment group, a table showing the determination results (hereinafter also referred to as a process table) is created.

[0583] An example of a processing table in which identification information (ID) of a facility group as venue facility information is associated with processing content defined for that identification information is shown in Fig. 42. Note that, for simplicity of explanation, only the video clipping angle and the method of combining the clipped video are shown as processing content.

[0584] In the example shown in FIG. 42, ID=0, ID=1, and ID=2 are shown as identification information (ID) of facility groups.

[0585] In the equipment group with ID=0, the cutout angle, that is, the video cutout is performed at a tilt angle of 24 degrees, and the cutout video is not synthesized.

[0586] A cutout angle of 24 degrees here means that the plane portion including the center of the hemisphere that forms the shape of the cutout image is tilted at 24 degrees with respect to the horizontal plane of the Metaverse space (virtual space). Note that as a process defined for the identification information (ID) of a facility group, only cutout of the image at a predetermined tilt angle may be specified, and the image composition method may be processed as needed for each venue 23 or venue 402.

[0587] In the equipment group with ID=1, video is cut out at a cutout angle (tilt angle) of 0 degrees, and the cut out video is not synthesized.

[0588] In the equipment group with ID=2, video is cut out at a cutout angle (tilt angle) of 0 degrees, and half of the cutout video is composited facing each other.

[0589] For example, in the first embodiment, when the planetarium image / sound image conversion unit 44 generates corrected content data to be transmitted to the venue 23 that belongs to (is classified into) the facility group with ID=0 shown in FIG. 42, the image is cropped at a cropping angle of 24 degrees as shown in FIG. 43.

[0590] That is, in this example, as shown on the left side of the drawing, the omnidirectional image of the content data before correction is tilted by 24 degrees to cut out a hemispherical area, which is used as the omnidirectional image of the corrected content data. In other words, a hemispherical area tilted by 24 degrees is cut out from the omnidirectional image and used as the image of the corrected content data. Therefore, the image of the corrected content data is a hemispherical omnidirectional image tilted by 24 degrees.

[0591] Also, for example, in the first embodiment, when the planetarium image / sound image conversion unit 44 generates corrected content data to be transmitted to the venue 23 belonging to the facility group with ID=2 shown in FIG. 42, image clipping and synthesis of the clipped image are performed as shown in FIG. 44.

[0592] In this example, an image is cut out from the spherical image of the content data before correction shown on the left side of the figure at an inclination angle of "0 degrees" as shown in the lower right of the figure. That is, the upper half of the spherical image is cut out. Furthermore, as shown in the upper right of the figure, half of the entire area in the front part of the cut out hemispherical image (panoramic image) is cut out and used as a cutout image, and this one cutout image is arranged in two places and combined to generate one hemispherical panoramic image. The panoramic image generated in this way is used as the image of the corrected content data.

[0593] The method for cutting out and combining such images is the same as that described with reference to Fig. 8. Since the seat layout type of venue 23, which belongs to the facility group with ID=2, is a concentric circle arrangement, the same cut-out images are arranged and combined so that each spectator in venue 23 can see the same image.

[0594] Furthermore, for example, in the second embodiment, when the video / audio conversion unit 535 generates content data to be presented for the venue 402 belonging to the facility group of each ID shown in FIG. 42, video clipping is performed as shown in FIG. 45.

[0595] That is, in the example shown in Figure 45, a virtual camera is placed in the metaverse space at a position and orientation determined for the ID value, which is the venue equipment information for venue 402, and content video (video data) is captured from the viewpoint of the virtual camera.

[0596] In this case, for example, the position and orientation of the virtual camera are determined by the coordinates and rotation angle of the Cartesian coordinate system, and the image is captured by the virtual camera as described with reference to Fig. 41. That is, the image is captured at an angle determined for the ID value, which is the venue facility information of venue 402. The image captured at this time may be an image of any shape, such as a hemispherical image (full-dome image) or an image of half the entire area in the front part of the full-dome image, as in the example described with reference to Fig. 44.

[0597] When the virtual camera performs a capture process and cuts out the video, the cut-out video is synthesized as necessary to form the final video of the content data for presentation.

[0598] The processing defined for the identification information (ID) of an equipment group as venue equipment information may be any processing, including but not limited to a processing of cutting out a portion of an image based on content data, such as an image obtained by a rendering processing, at an inclination angle corresponding to the above-mentioned identification information (ID) (an inclination angle associated with the identification information), or a processing of arranging and synthesizing multiple images cut out from an image based on content data according to the seat layout type corresponding to the identification information (ID).

[0599] For example, a process may be performed to convert the video format of the video data constituting the content or the video format of the video data obtained by capture by a virtual camera into a video format corresponding to the identification information (ID) of the facility group. Furthermore, the viewpoint position of the virtual camera may be varied depending on the size (radius or diameter) of the dome portion identified by the identification information (ID) of the facility group. That is, a process for capturing video with different viewpoint positions and orientations of the virtual camera may be defined for each facility group ID. In this case, for example, in the second embodiment, the video based on the video data constituting the presentation content data will be video with a different viewpoint position depending on the size of the dome portion of the venue corresponding to the identification information (ID) of the facility group.

[0600] Furthermore, although we have explained the case where an ID is assigned to each group of venue equipment, in the case where an ID is assigned to each venue 23 (venue 402), correction content data and presentation content data can be obtained in the same manner by performing the processing specified for each ID.

[0601] Furthermore, for example, in the second embodiment, the angle of the virtual camera may be different for each combination of the identification information (ID) of the facility group and the part indicated by the progress information. In other words, the video / audio-image conversion unit 535 may perform capture processing at an angle determined for the ID of the facility group as venue facility information and the part indicated by the progress information.

[0602] As an example, suppose that an MC part "MC1" corresponding to list item LI41 and an MC part "MC2" corresponding to list item LI42 in the playlist as shown in FIG. 46 are performed as content in a live performance.

[0603] In this case, it is conceivable to make the angle of the virtual camera, that is, the placement position and orientation of the virtual camera in the metaverse space (world), different between the MC part "MC1" and the MC part "MC2."

[0604] Furthermore, in the second embodiment, for example, the video / audio conversion unit 535 performs rendering processing of audio such as VBAP, and multi-channel sound image data is generated as sound image data constituting the content data for presentation. Similarly, in the first embodiment, for example, the spatial audio output unit 42 also performs rendering processing of audio such as VBAP, and multi-channel sound image data is generated as appropriate.

[0605] By performing audio rendering processing such as VBAP, multi-channel sound image data for stereophonic playback can be obtained.

[0606] In stereophonic sound reproduction, sound source objects can be placed at any position in all directions of 360 degrees as seen by the listener, as shown in Fig. 47. In other words, sounds based on sound image data can be presented so that sounds from objects in the Metaverse space, such as sounds from avatars corresponding to performers, come from any direction as seen by the listener.

[0607] The position of the listener here is the position of the virtual microphone in the metaverse space, and for example, the position of the virtual microphone corresponds in real space to a predetermined fixed position of the audience seats in venue 402. Therefore, by moving the position of the virtual microphone in the metaverse space, listeners in venue 402, which is real space, can receive audio presentation as if they were at the position of the virtual microphone in the metaverse space.

[0608] As described above, if the venue equipment information is an ID (identification information) of an equipment group or the like, and processing according to that ID is performed when generating sound image data that constitutes corrected content data or presentation content data, it is possible to make the sound output angle different for each ID.

[0609] That is, for example, as shown in FIG. 48, if the image cutout angle, i.e., the range from which the image is cut out, differs for each equipment group ID, the sound output angle based on the sound image data may also differ (be determined) depending on the cutout range.

[0610] In this case, a process for adjusting the output angle of sound based on the sound image data constituting the corrected content data or the presentation content data is performed according to the ID of the equipment group. That is, the spatial audio output unit 42 and the video / sound-image conversion unit 535 perform an audio rendering process for presenting sound at a predetermined output angle according to the ID of the equipment group. This rendering process can be said to be a process for generating the sound image data constituting the corrected content data or the presentation content data by rotating the position (localization position) of the sound image based on the sound image data constituting the content data, for example, by an inclination angle associated with the ID of the equipment group, i.e., by the video cropping angle (inclination angle) determined for the ID of the equipment group. In this case, the sound image of the sound based on the sound image data constituting the corrected content data or the presentation content data is localized at the position after rotation.

[0611] For example, in the second embodiment, when VBAP is performed as the audio rendering process, the output angle can be adjusted by adjusting avatar position information that indicates the relative position of the avatar as viewed from the position of the virtual microphone. In this case, the video / audio conversion unit 535 may adjust the sound output angle by adjusting the avatar position information, that is, by rotating the position indicated by the avatar position information around the listening position, or by adjusting object position information that indicates the absolute position of the avatar corresponding to the performer in the Metaverse space.

[0612] Furthermore, the placement positions and speaker channel configurations of the speakers constituting the speaker system as the spatial audio output unit 53 and the planetarium sound equipment 506 differ depending on the venue 23 or 402. Therefore, as a process defined for the identification information (ID) of an equipment group as venue equipment information, an audio rendering process corresponding to the speaker placement and channel configuration according to the equipment group ID is performed. That is, a rendering process is performed to generate sound image data of the channel configuration and speaker placement according to the identification information (ID) of the equipment group based on the sound image data constituting the content data as sound image data constituting the corrected content data or the content data for presentation.

[0613] Although the above description has been given of a case where the content image is two-dimensional (2D), i.e., where one virtual camera is placed in the metaverse space, the content image may also be three-dimensional (3D). In such a case, two virtual cameras are placed in the metaverse space (virtual space), and a pair of images having a parallax between them are captured by the virtual cameras. In this way, a three-dimensional image can be presented based on the image data.

[0614] <Example of Computer Configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the computer includes a computer built into dedicated hardware, and a general-purpose personal computer, for example, that can execute various functions by installing various programs.

[0615] FIG. 49 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes using a program.

[0616] In the computer, a CPU (Central Processing Unit) 801, a ROM (Read Only Memory) 802, and a RAM (Random Access Memory) 803 are interconnected by a bus 804.

[0617] An input / output interface 805 is further connected to the bus 804. An input unit 806, an output unit 807, a recording unit 808, a communication unit 809, and a drive 810 are connected to the input / output interface 805.

[0618] The input unit 806 includes a keyboard, a mouse, a microphone, an image sensor, etc. The output unit 807 includes a display, a speaker, etc. The recording unit 808 includes a hard disk, a non-volatile memory, etc. The communication unit 809 includes a network interface, etc. The drive 810 drives a removable recording medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0619] In a computer configured as described above, the CPU 801 performs the above-described series of processes by, for example, loading a program recorded in the recording unit 808 into the RAM 803 via the input / output interface 805 and the bus 804 and executing it.

[0620] The program executed by the computer (CPU 801) can be provided by being recorded on a removable recording medium 811 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0621] In a computer, the program can be installed in the recording unit 808 via the input / output interface 805 by inserting a removable recording medium 811 into the drive 810. The program can also be received by the communication unit 809 via a wired or wireless transmission medium and installed in the recording unit 808. Alternatively, the program can be installed in the ROM 802 or the recording unit 808 in advance.

[0622] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0623] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.

[0624] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.

[0625] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.

[0626] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0627] Furthermore, the present technology can also be configured as follows.

[0628] (1) An information processing device comprising: a correction unit that performs a correction process on content data of the content based on venue facility information regarding facilities of a venue where the content is played back; and an output unit that outputs the corrected content data obtained by the correction process. (2) The information processing device described in (1), in which the venue is a planetarium venue. (3) The information processing device described in (2), in which the venue facility information includes at least one of information regarding video projection equipment of the venue, information regarding audio equipment of the venue, information regarding seats in the venue, and information regarding a dome portion of the venue where the video of the content is displayed. (4) The information processing device described in (3), in which the correction process is a process of converting the video format of video data constituting the content data into the video format indicated by the venue facility information. (5) The information processing device described in (3) or (4), in which the correction process is a process of cutting out a portion of a video based on the video data constituting the content data, based on the tilt angle of the dome portion indicated by the venue facility information. (6) The information processing device described in any one of (3) to (5), wherein the correction process is a process of arranging and synthesizing a plurality of cut-out images cut out from images based on video data constituting the content data in accordance with a seating layout type of the venue indicated by the venue facility information. (7) The information processing device described in any one of (3) to (6), wherein the images based on video data constituting the corrected content data are images whose viewpoint position varies depending on the size of the dome part of the venue. (8) The information processing device described in any one of (3) to (7), wherein the correction process is a process of converting sound image data constituting the content data into multi-channel sound image data corresponding to the channel configuration of a speaker system of the venue indicated by the venue facility information.(9) The information processing device according to any one of (3) to (8), wherein the output unit outputs the corrected content data obtained by the correction process based on the venue facility information to the venue corresponding to the venue facility information, and outputs the corrected content data obtained by the correction process based on other venue facility information to other venues corresponding to the other venue facility information. (10) The information processing device according to any one of (3) to (9), further comprising a communication unit that receives interaction information from the venue, the interaction information including at least one of venue video data including audience seats at the venue as a subject, venue sound image data obtained by collecting sounds from the venue, and results of analysis processing based on at least one of the venue video data and the venue sound image data, and transmits the interaction information to a device that distributes the content data. (11) The information processing device according to (10), further comprising a receiving unit that receives the content data that has been processed based on the interaction information of one or more venues transmitted by the device. (12) The information processing device according to (11), wherein the processing based on the interaction information is at least one of synthesizing video based on the venue video data, synthesizing sound based on the venue sound image data, adding or changing a performance in accordance with the result of the analysis processing, synthesizing a video of a virtual audience in accordance with the result of the analysis processing, and changing a viewpoint position of video based on video data constituting the content data. (13) The information processing device according to any one of (10) to (12), wherein the result of the analysis processing is a result of a bone structure estimation process for audience members at the venue, or an estimation result of the level of excitement at the venue. (14) The information processing device according to any one of (3) to (9), wherein the correction unit performs processing on the corrected content data based on at least one of venue video data obtained for one or more of the venues, the venue video data including audience seats at the venue as a subject, venue sound image data obtained by collecting sounds from the venue, and a result of an analysis process based on at least one of the venue video data and the venue sound image data.(15) The information processing device according to (14), wherein the processing is at least one of synthesizing an image based on the venue image data, synthesizing a sound based on the venue sound image data, synthesizing an image of a virtual audience in accordance with the result of the analysis processing, and adding or changing a performance in accordance with the result of the analysis processing. (16) The information processing device according to (14) or (15), wherein the result of the analysis processing is the result of a skeletal structure estimation process for the audience at the venue, or the result of an estimation of the level of excitement at the venue. (17) The information processing device according to any one of (3) to (16), wherein the corrected content data includes at least one of video data, sound image data, and haptic data. (18) An information processing method, wherein an information processing device performs a correction process on content data of the content based on venue facility information related to the facilities of the venue where the content will be played back, and outputs the corrected content data obtained by the correction process. (19) A program causing a computer to execute a process including the steps of: correcting content data of content based on venue facility information related to facilities at a venue where the content is to be played back; and outputting the corrected content data obtained by the correction process. (20) An information processing system having one or more venues where content is to be played back, comprising: a correction unit that performs correction processing on content data of content based on venue facility information related to facilities at the venues; and a playback unit that plays back the content based on the corrected content data obtained by the correction process.

[0629] 11 Information processing system, 21 Metaverse login PC, 22 Cloud, 23-1 to 23-N, 23 Venue, 44-1 to 44-M, 44 Planetarium image and sound image conversion unit, 45-1 to 45-M, 45 Image and sound image distribution unit, 46 Feedback communication unit, 47 Feedback processing unit, 401 Studio, 402-1, 402-2, 402 Venue, 503 Communication / rendering PC, 533 Control unit

Claims

1. A program for causing a computer to execute a process including the steps of performing first information processing on first video data constituting content data of the content, which is generated using motion data, based on identification information regarding a venue where the content is played, generating second video data, and outputting the second video data.

2. The program according to claim 1, wherein the venue is a planetarium venue.

3. The program according to claim 1, wherein the first video data is 3D data obtained by rendering processing based on the motion data and background video data, or video data of an omnidirectional video obtained from the 3D data.

4. The program according to claim 2, wherein the identification information is identification information regarding the facilities of the venue or identification information indicating the venue.

5. The program according to claim 4, wherein the first information processing is a process of converting the video format of the first video data into a video format corresponding to the identification information.

6. The program according to claim 4, wherein the first information processing is a process of cutting out a part of the video based on the first video data at an inclination angle associated with the identification information.

7. The program according to claim 4, wherein the first information processing is a process of arranging and synthesizing a plurality of cut-out videos cut out from the video based on the first video data according to the layout type of the seats in the venue corresponding to the identification information.

8. The program according to claim 4, wherein the video based on the second video data is a video with different viewpoint positions depending on the size of the dome portion of the venue corresponding to the identification information.

9. The program according to claim 3, wherein second information processing is performed on first audio-visual data constituting the content data based on the identification information, and second audio-visual data is generated.

10. The program according to claim 9, wherein the second information processing is a process of generating second audio-visual data having a channel configuration corresponding to the identification information based on the first audio-visual data.

11. The program according to claim 9, wherein the second information processing is a process of adjusting the output angle of the sound based on the second audio-visual data according to the identification information.

12. The second information processing is a process of generating the second audio-visual data by rotating the audio-visual based on the first audio-visual data by the inclination angle associated with the identification information. The program according to claim 9.

13. Causing interaction information including at least any one of venue video data including the audience seats in the venue as a subject, venue audio-visual data obtained by collecting the sound of the venue, and the result of analysis processing based on at least any one of the venue video data and the venue audio-visual data to be transmitted to a device that distributes the content data or data for obtaining the content data. The program according to claim 4.

14. Receiving data obtained by processing based on the interaction information of one or more of the venues transmitted by the device. The program according to claim 13.

15. The data obtained by processing based on the interaction information is at least any one of video data based on the venue video data, audio-visual data based on the venue audio-visual data, data for adding an effect according to the result of the analysis processing, virtual audience video data according to the result of the analysis processing, and data indicating a change in the viewpoint position of the video of the content. The program according to claim 14.

16. Acquiring the motion data from a studio, and generating the first video data using the acquired motion data. The program according to claim 2.

17. Displaying a playlist in which a plurality of the parts including the part implemented in the content are listed. The program according to claim 1.

18. When a predetermined operation is performed on the screen including the playlist, temporarily stopping the output of the second video data and outputting prepared video data or audio-visual data. The program according to claim 17.

19. An information processing method in which an information processing device performs information processing on first video data constituting content data of the content generated using motion data based on identification information regarding a venue where the content is played, generates second video data, and outputs the second video data.

20. An information processing apparatus comprising: a generation unit that performs information processing on first video data constituting content data of the content, which is generated using motion data, based on identification information regarding a venue where the content is played, to generate second video data; and an output unit that outputs the second video data.

Citation Information

Patent Citations

  • Video conversion method and image converter, and multi-projection system

    JP2005347813A

  • Dome screen projection facility

    JP2016170252A

  • Sign language translation device and program

    JP2023156082A

  • System for high-resolution content playback

    US20190082152A1

  • Information-processing device and method

    WO2016147888A1