Content viewing system, content viewing method, and program
The content viewing system addresses user discomfort in VR by displaying 2D content on a virtual aerial display within a VR space, using generative AI to match audio effects with the VR space's characteristics, thereby enhancing immersion.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- DENSO TEN LTD
- Filing Date
- 2024-10-07
- Publication Date
- 2026-04-17
AI Technical Summary
In virtual reality (VR) spaces, displaying 2D content on a virtual aerial display without matching the acoustic characteristics of the VR space leads to user discomfort and reduced immersion due to inconsistent sound perception.
A content viewing system that includes an information processing device to display 2D content on a virtual aerial display within a VR space and applies audio effects processing based on the VR space's characteristics, using a generative AI server to determine reverb parameters.
Enhances user immersion by aligning the VR space's type with sound perception, reducing discomfort and improving the overall experience.
Smart Images

Figure 2026066501000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to content viewing technology.
Background Art
[0002] Conventionally, a technology has been known that uses computer graphics (CG) or the like to generate a virtual reality (VR) space, and allows a user to perceive and experience the VR space through a VR goggle or the like.
[0003] For example, Patent Document 1 proposes a landscape simulator that allows a subject to visually recognize a landscape that changes as they move within a VR space and provides a pseudo-experience as if they were actually moving.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, 2D (2-dimensional) content is still the mainstream of the content available in the world. Therefore, there is a need to provide a virtual aerial display in a VR space and display the video of 2D content on such an aerial display.
[0006] In the real world, the reverberation characteristics change depending on the type of space, and the way sound is heard changes. Therefore, when providing a virtual aerial display in a VR space and displaying the video of 2D content on such an aerial display, if the type of VR space and the way sound is heard are not consistent, the user will feel a sense of discomfort in the way sound is heard, and the user's immersion will decrease.
[0007] In view of the above, the present invention aims to provide content viewing technology that can enhance user immersion. [Means for solving the problem]
[0008] An exemplary content viewing system of the present invention comprises an information processing device, a display device, and an audio output device. The information processing device displays at least a portion of the VR space, generates a display signal for displaying the image of 2D content on a virtual aerial display provided in the VR space and outputs it to the display device, and generates an audio signal that has been processed with effects according to the characteristics of the VR space for the audio of the 2D content and outputs it to the audio output device. [Effects of the Invention]
[0009] According to an exemplary example of the present invention, the video of 2D content is displayed on a virtual aerial display set up in a VR space, and effect processing is applied to the audio of the 2D content according to the characteristics of the VR space, thereby matching the type of VR space with how the sound is heard. This can enhance the user's sense of immersion. [Brief explanation of the drawing]
[0010] [Figure 1] Diagram showing the configuration of the content viewing system according to the embodiment. [Figure 2] Diagram showing the configuration of the content delivery server. [Figure 3] Diagram showing the configuration of a server equipped with a generative AI. [Figure 4] Diagram showing the configuration of the information processing device. [Figure 5] Diagram showing the prompt template table [Figure 6] This diagram shows the relationship between each reverb parameter and the reverb signal generated by the reverb process. [Figure 7] Flowchart of the content playback process executed by the controller of the information processing device. [Figure 8] A diagram showing an example of a VR image. [Figure 9] Figure showing the first example of the input image [Figure 10] Figure showing the second example of the input image [Figure 11] Figure showing the third example of the input image [Figure 12] Figure showing the fourth example of the input image [Figure 13] Table showing the response from the server equipped with generative AI
Mode for Carrying Out the Invention
[0011] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the drawings.
[0012] <Configuration of the Content Viewing System> FIG. 1 is a diagram showing the configuration of a content viewing system SYS1 according to an embodiment of the present invention. As shown in FIG. 1, the content viewing system SYS1 includes an information processing device 1, a display device 2, and an audio output device 3.
[0013] The information processing device 1 is, for example, a PC (personal computer). The display device 2 is, for example, VR goggles worn on the user's head. The VR goggles include a left-eye display, a right-eye display, and a sensor for sensing the orientation of the user's head. The audio output device 3 is, for example, stereo earphones worn on both ears of the user. Note that the display device 2 and the audio output device 3 may be separate or integrated like a VR headset.
[0014] The information processing device 1 is wirelessly connected to each of the display device 2 and the audio output device 3 by, for example, Bluetooth (registered trademark). Note that the information processing device 1 and the display device 2 and the audio output device 3 may be connected by wire instead of wireless connection.
[0015] The information processing device 1 and the content providing server SV1 communicate with each other via the network NT1. The content providing server SV1 may be a physical server or a virtual server. The content providing server SV1 may be composed of one server or a plurality of servers.
[0016] The content providing server SV1 is a server capable of providing a plurality of 2D contents and a plurality of VR data. The 2D content includes a two-dimensional video and stereo audio. The VR data includes a 360-degree image for generating a VR space. The 360-degree image may be a still image or a video. Also, the VR data may include audio. The information processing device 1 downloads arbitrary 2D content and arbitrary VR data from the content providing server SV1. In this embodiment, although arbitrary 2D content and arbitrary VR data are provided to the information processing device 1 via the network NT1, the arbitrary 2D content and arbitrary VR data may be stored in a portable storage medium and provided to the information processing device 1 via the portable storage medium.
[0017] The information processing device 1 and the generative AI (Artificial Intelligence) - equipped server SV2 communicate with each other via the network NT1. The generative AI - equipped server SV2 may be a physical server or a virtual server. The generative AI - equipped server SV2 may be composed of one server or a plurality of servers.
[0018] The information processing device 1 transmits a prompt to be given to the generative AI to the generative AI - equipped server SV2. The generative AI - equipped server SV2 transmits an answer generated by the generative AI composed of, for example, LLM (Large Language Models) to the information processing device 1.
[0019] <Configuration of the Content Providing Server> FIG. 2 is a diagram showing the configuration of the content providing server SV1. The content providing server SV1 includes a communication unit 11, a storage unit 12, and a controller 13.
[0020] The communication unit 11 transmits and receives arbitrary signals with the information processing device 1. The controller 13 can also transmit and receive arbitrary information with the information processing device 1 using the communication unit 11; however, the description of the communication unit 11 may be omitted below.
[0021] The storage unit 12 is configured to have non-volatile memory such as ROM (Read-only memory) or flash memory, and volatile memory such as RAM (Random Access Memory). The storage unit 12 stores a program 121, a plurality of 2D contents 122, and a plurality of VR data 123.
[0022] Program 121 is a program that should be executed by controller 13.
[0023] The controller 13 comprehensively controls the operation of each component in the content provision server SV1. The controller 13 is equipped with a processing unit including a CPU (Central Processing Unit) as a hardware resource. The controller 13 has a provisioning unit 131.
[0024] When the provisioning unit 131 receives a download request from the information processing device 1, it transmits any 2D content 122 or VR data 123 corresponding to the download request to the information processing device 1.
[0025] <Configuration of a server equipped with a generation AI system> Figure 3 shows the configuration of the generative AI-equipped server SV2. The generative AI-equipped server SV2 comprises a communication unit 21, a storage unit 22, and a controller 23.
[0026] The communication unit 21 transmits and receives arbitrary signals with the information processing device 1. The controller 23 can also transmit and receive arbitrary information with the information processing device 1 using the communication unit 21; however, the description of the communication unit 21 may be omitted below.
[0027] The memory unit 22 is configured to have non-volatile memory such as ROM or flash memory, and volatile memory such as RAM. The memory unit 22 stores the program 221 and the generation system AI 222.
[0028] Program 221 is a program to be executed by controller 23. Generative AI (AI model) 222 is a trained machine learning model for generating responses to prompts sent from information processing device 1. Generative AI 222 is composed of, for example, LLM.
[0029] The controller 23 comprehensively controls the operation of each component in the generative AI-equipped server SV2. The controller 23 is equipped with a processing unit including a CPU as a hardware resource. The controller 23 has a generative AI execution unit 231.
[0030] When the generation AI execution unit 231 receives a prompt from the information processing device 1, it uses the prompt as input data to execute the generation AI 222 and causes the generation AI 222 to generate a response. The controller 23 transmits the response generated by the generation AI 222 to the information processing device 1.
[0031] <Configuration of the information processing device> Figure 4 shows the configuration of the information processing device 1. It comprises a communication unit 31, a storage unit 32, a controller 33, and an operation unit 34.
[0032] The communication unit 31 transmits and receives arbitrary signals to and from the content provision server SV1 and the generation AI-equipped server SV2, respectively. Although the controller 33 can transmit and receive arbitrary information to and from the content provision server SV1 and the generation AI-equipped server SV2 using the communication unit 31, the description of the communication unit 31 may be omitted below.
[0033] The communication unit 31 transmits and receives arbitrary signals to and from the display device 2 and the audio output device 3, respectively. Although the controller 33 can transmit and receive arbitrary information to and from the display device 2 and the audio output device 3 using the communication unit 31, the description of the communication unit 31 may be omitted below.
[0034] The storage unit 32 is configured to have non-volatile memory such as ROM or flash memory, and volatile memory such as RAM. The storage unit 32 stores the program 321, 2D content 322, VR data 323, prompt template table 324, and reverb parameters 325.
[0035] Program 321 is a program that should be executed by the controller 33. Program 321 is a program that causes the controller 33 to perform the following actions: display at least a portion of the VR space, generate a display signal for displaying the video of the 2D content to be viewed on a virtual aerial display set up in the VR space, and generate an audio signal that applies effect processing to the audio of the 2D content to be viewed according to the characteristics of the VR space.
[0036] The 2D content 322 is content downloaded from the content provision server SV1. Although only one 2D content 322 is shown in Figure 4, the storage unit 32 may store multiple 2D content 322. When the storage unit 32 stores multiple 2D content 322, the 2D content to be viewed is selected from among the multiple 2D content 322 in response to input from the operation unit 34.
[0037] VR data 323 is data downloaded from the content provision server SV1. Although only one VR data 323 is shown in Figure 4, the storage unit 32 may store multiple VR data 323. When the storage unit 32 stores multiple VR data 323, data is selected from among the multiple VR data 323 in response to input from the operation unit 34 to generate a VR space where a virtual aerial display showing the video of the 2D content to be viewed is provided.
[0038] The prompt template table 324 is data that shows an example of a prompt structure, and as shown in Figure 5, it is a template arranged in the order of "instruction," "constraints," one or more "questions," and "output format."
[0039] An example of an "instruction sheet" is: "You are a sound engineer with extensive field measurement experience. Please compile the answers to the questions from the input image into a table and output it."
[0040] An example of a "constraint" is the following sentence: "Unless the answer is unknown, the output must always be a positive number. Please only provide answers to the questions, following the output format."
[0041] An example of the "first question" is: "Please estimate the volume of space in the input image. The unit is cubic meters. If it is difficult to estimate, an approximate value is acceptable. If you cannot specify a value, please answer 'unknown'."
[0042] An example of the "second question" is: "Capture the characteristics of the floor and walls in the input image and predict the equivalent sound absorption area. The unit is square meters. If prediction is difficult, an approximate value is acceptable. If you cannot specify a value, please answer 'unknown'."
[0043] An example of the "third question" is: "Capture the characteristics of the floor in the input image's space and predict the sound absorption coefficient of the floor surface (average value across all frequency bands). If prediction is difficult, an approximate value is acceptable. If you cannot specify a value, please answer 'unknown'."
[0044] An example of the "fourth question" is: "Identify the characteristics of the walls in the input image's spatial context and predict the sound absorption coefficient of the wall surface (average value across all frequency bands). If prediction is difficult, an approximate value is acceptable. If you cannot determine a specific value, please answer 'unknown'."
[0045] An example of the "fifth question" is: "Using the volume and equivalent sound absorption area estimated in the first and second questions, calculate the reverberation time using Sabin's formula. The unit is seconds. If you cannot calculate it, please answer 'unknown'."
[0046] An example of the "output format" is "|Question|Answer| |First Question|m^{3}| |Second Question|m^{2}| |Third Question|%| |Fourth Question|%| |Fifth Question|sec|".
[0047] The reverb parameter 325 is a parameter used when applying reverb processing, which is one of the effect processing methods, to an audio signal. The reverb parameter 325 is determined by the reverb parameter determination unit 333, which will be described later. In this embodiment, the reverb parameter 325 includes "Pre-Delay", "Early Reflections", "Early Reflections Level", and "Decay". The relationship between each parameter of the reverb parameter 325 and the reverb signal RVB generated by the reverb processing is as shown in Figure 6.
[0048] The controller 33 comprehensively controls the operation of each part of the information processing device 1. The controller 33 is equipped with a processing unit including a CPU as a hardware resource. The controller 33 has functional blocks 331 to 334.
[0049] The controller 33 is a program execution device (computer) capable of executing any program. The controller 33 executes program 321, thereby realizing each function of the controller 33 (including the functions of function blocks 331 to 334). All operations of the controller 33 described in this embodiment may be operations realized by the controller 33 executing program 321. Program 321 may consist of multiple programs.
[0050] The operation unit 34 is an input device that accepts user input. The operation unit 34 is, for example, a keyboard, touch panel, mouse, etc.
[0051] The functional blocks of the controller 33 will now be described. Functional blocks 331, 332, 333, and 334 are the acquisition unit, image processing unit, reverb parameter determination unit, and audio processing unit, respectively.
[0052] The acquisition unit 331 acquires 2D content and VR data from the content provision server SV1. The acquisition unit 331 also acquires information indicating the user's head orientation from the display device 2.
[0053] The image processing unit 332 generates a VR space from the 360-degree image contained in the selected VR data 323, extracts a portion of the VR space according to the user's head orientation, and generates left-eye and right-eye image signals with parallax corresponding to the extracted region. The image processing unit 332 superimposes the image signal corresponding to the video of the 2D content 322 to be viewed onto the left-eye and right-eye image signals, respectively, so that the video of the 2D content 322 to be viewed is displayed on a virtual aerial display set up in the VR space. The image processing unit 332 outputs the left-eye and right-eye image signals, respectively, after the image signal corresponding to the video of the 2D content 322 has been superimposed, to the display device 2.
[0054] The reverb parameter determination unit 333 determines reverb parameters 325 according to the characteristics of the VR space generated by the image processing unit 332.
[0055] The audio processing unit 334 generates an audio signal to which the audio of the 2D content 332 to be viewed has been processed with effects according to the characteristics of the VR space generated by the image processing unit 332, and outputs it to the audio output device 3. If the selected VR data 323 includes audio, the audio processing unit 334 may apply effects to the combined audio of the audio of the 2D content 332 to be viewed and the audio of the selected VR data 323, or it may mute the audio of the selected VR data 323 and apply effects only to the audio of the 2D content 332 to be viewed.
[0056] In this embodiment, reverb processing is used as the effect processing, and the audio processing unit 334 applies reverb processing to the audio of the 2D content 332 to be viewed using the reverb parameter 325 determined by the reverb parameter determination unit 333.
[0057] <Operation of the Information Processing Device> Figure 7 is a flowchart of the content playback process executed by the controller 33 of the information processing device 1. This content playback process is realized when the controller 33 executes the program 321 described above. The flow shown in Figure 7 starts when the 2D content 322 (the 2D content to be viewed 322) and VR data 323 used in the content playback process are selected by user operation on the operation unit 34 and content playback is instructed.
[0058] First, in step S10, the image processing unit 332 of the controller 33 generates a VR space from the 360-degree images included in the selected VR data 323, and then proceeds to step S20.
[0059] In step S20, the acquisition unit 331 of the controller 33 acquires information indicating the direction of the user's head from the display device 2, and proceeds to step S30.
[0060] In step S30, the image processing unit 332 of the controller 33 extracts a portion of the VR space according to the orientation of the user's head, generates a left-eye image signal and a right-eye image signal with parallax corresponding to the extracted region, and proceeds to step S40.
[0061] In step S40, the image processing unit 332 of the controller 33 superimposes the image signals corresponding to the video of the 2D content 322 to be viewed onto the left-eye image signal and the right-eye image signal, respectively, so that the video of the 2D content 322 to be viewed is displayed on a virtual aerial display set up in the VR space. The superimposed left-eye image signal and the right-eye image signal are then output to the display device 2, and the process proceeds to step S50. After the processing in step S40 is completed, the user can view a VR image as shown in Figure 8 by looking at the screen of the display device 2 with both eyes. The VR image shown in Figure 8 is an image in which the video of the 2D content 322 to be viewed, displayed on a virtual aerial display 42 set up in the VR space, is superimposed onto a three-dimensional VR image 41 as seen from the virtual user's position in the VR space.
[0062] In step S50, the reverb parameter determination unit 333 of the controller 33 refers to the prompt template table 324 and sends the prompt to the generation AI-equipped server SV2, then proceeds to step S60. When the reverb parameter determination unit 333 of the controller 33 sends the prompt to the generation AI-equipped server SV2, it also sends an input image corresponding to the VR space generated by the image processing unit 332 along with the prompt to the generation AI-equipped server SV2. In this embodiment, the input image corresponding to the VR space generated by the image processing unit 332 is a two-dimensional image of the VR space viewed from one direction.
[0063] If the VR space generated by the image processing unit 332 is a theater, the input image will be, for example, the two-dimensional image 51 shown in Figure 9. If the VR space generated by the image processing unit 332 is a Japanese-style room, the input image will be, for example, the two-dimensional image 52 shown in Figure 10. If the VR space generated by the image processing unit 332 is a cafe space, the input image will be, for example, the two-dimensional image 53 shown in Figure 11. If the VR space generated by the image processing unit 332 is an anechoic chamber, the input image will be, for example, the two-dimensional image 54 shown in Figure 12.
[0064] In step S60, the reverb parameter determination unit 333 of the controller 33 receives a response from the generative AI-equipped server SV2 and proceeds to step S70. Figure 13 is a table showing the responses from the generative AI-equipped server SV2. As shown in Figure 13, the generative AI-equipped server SV2 provides responses that are reasonably appropriate for each VR space. Here, in this invention, the aim is to ensure that the type of VR space and the way the sound is heard are matched to such an extent that the user does not feel any discomfort with how the sound is heard, so it is not necessary to provide responses that are strictly (highly accurate) for the characteristics of each VR space. Furthermore, since the characteristics of the VR space generated by the image processing unit 332 are obtained from the input image (an image based on the VR space generated by the image processing unit 332), the characteristics of the VR space generated by the image processing unit 332 can be obtained even if the VR space is an unrealistic space where it does not actually exist.
[0065] In step S70, the reverb parameter determination unit 333 of the controller 33 determines the reverb parameters and proceeds to step S80.
[0066] The reverb parameter determination unit 333 of the controller 33 determines the value of "Pre-Delay" according to the distance between the user's virtual position in the VR space generated by the image processing unit 332 and the virtual aerial display provided in the VR space generated by the image processing unit 332 (hereinafter abbreviated as aerial display distance). The aerial display distance is a parameter used by the image processing unit 332 when providing the aerial display in the VR space. The relationship between the aerial display distance and the value of "Pre-Delay" may be stored in the storage unit 32, for example, in the form of a data table or calculation formula. By determining the value of "Pre-Delay" according to the aerial display distance, the controller 33 can apply effect processing (more specifically, reverb processing) according to the aerial display distance, thereby reducing the risk that the user may feel uncomfortable with the way sounds related to "Pre-Delay" are heard.
[0067] The reverb parameter determination unit 333 of the controller 33 determines the value of "Early Reflections" according to the response from the generative AI-equipped server SV2 to the first question. The response from the generative AI-equipped server SV2 to the first question is the size (volume) of the VR space generated by the image processing unit 332. Since the value of "Early Reflections" is determined using the response from the AI-equipped server SV2, the processing load on the information processing device 1 required to determine the value of "Early Reflections" can be reduced. The relationship between the size (volume) of the VR space generated by the image processing unit 332 and the value of "Early Reflections" may be stored in the storage unit 32, for example, in the form of a data table or calculation formula. By determining the value of "Early Reflections" according to the size (volume) of the VR space generated by the image processing unit 332, the controller 33 can apply effect processing (more specifically, reverb processing) according to the size (volume) of the VR space, thereby reducing the risk that the user may feel uncomfortable with the way sounds related to "Early Reflections" are heard.
[0068] The reverb parameter determination unit 333 of the controller 33 determines the value of "Early Reflections Level" based on the answers from the generative AI-equipped server SV2 to the first question, the third question, and the fourth question. The answer from the generative AI-equipped server SV2 to the first question is the size (volume) of the VR space generated by the image processing unit 332. The answer from the generative AI-equipped server SV2 to the third question is the sound absorption coefficient of the floor surface of the VR space generated by the image processing unit 332. The answer from the generative AI-equipped server SV2 to the fourth question is the sound absorption coefficient of the walls of the VR space generated by the image processing unit 332. Since the value of "Early Reflections Level" is determined using the answers from the AI-equipped server SV2, the processing load on the information processing device 1 required to determine the value of "Early Reflections Level" can be reduced. The relationship between the size (volume) of the VR space, the sound absorption coefficient of the floor, and the sound absorption coefficient of the walls generated by the image processing unit 332, and the value of "Early Reflections Level" may be stored in the storage unit 32, for example, in the form of a data table or calculation formula. By determining the value of "Early Reflections Level" according to the size (volume) of the VR space, the sound absorption coefficient of the floor, and the sound absorption coefficient of the walls generated by the image processing unit 332, the controller 33 can apply effect processing (more specifically, reverb processing) according to the size (volume) of the VR space, the sound absorption coefficient of the floor, and the sound absorption coefficient of the walls, thereby reducing the risk that the user may feel uncomfortable with the way sounds related to "Early Reflections Level" are heard.
[0069] The reverb parameter determination unit 333 of the controller 33 determines the value of "Decay" according to the answer from the generation AI-equipped server SV2 to the fifth question. The answer from the generation AI-equipped server SV2 to the fifth question is the reverberation time of the VR space generated by the image processing unit 332. Since the value of "Decay" is determined using the answer from the AI-equipped server SV2, the processing load on the information processing device 1 required to determine the value of "Decay" can be reduced. The relationship between the reverberation time of the VR space generated by the image processing unit 332 and the value of "Decay" may be stored in the storage unit 32 in the form of a data table or calculation formula, for example. By determining the value of "Decay" according to the reverberation time of the VR space generated by the image processing unit 332, the controller 33 can apply effect processing (more specifically, reverb processing) according to the reverberation time of the VR space, thereby reducing the risk that the user may feel uncomfortable with the way sounds related to "Decay" are heard. In contrast to this embodiment, the information processing device 1 may not include a fifth question in the prompt, and the controller 33 may use the answers from the generative AI-equipped server SV2 to the first question and the answers from the generative AI-equipped server SV2 to the second question to calculate the reverberation time of the VR space generated by the image processing unit 332 based on Sabine's reverberation formula.
[0070] In step S80, the audio processing unit 334 of the controller 33 adjusts the left-right balance of the audio of the 2D content 322 being viewed according to the user's head orientation, and further applies effect processing (more specifically, reverb processing using the reverb parameters determined by the reverb parameter determination unit 333) according to the characteristics of the VR space generated by the image processing unit 332. This audio signal is then output to the audio output device 3, and the process proceeds to step S90. This makes it possible to match the type of VR space generated by the image processing unit 332 with how the sound is heard, thereby enhancing the user's sense of immersion.
[0071] In step S80, the controller 33 determines whether or not to terminate playback of the 2D content 322 to be viewed. The controller 33 terminates processing if the 2D content 322 to be viewed is played to completion, or if the user instructs the end of playback of the 2D content 322 to be viewed via the operation unit 34. On the other hand, if playback of the 2D content 322 to be viewed is not terminated, the controller 33 repeats the processing in steps S20 to S80.
[0072] <Notes, etc.> The various technical features disclosed in the embodiments for carrying out the invention as specified herein can be modified in various ways without departing from the spirit of the technical creation. Furthermore, the multiple embodiments and modifications disclosed in the embodiments for carrying out the invention as specified herein may be combined to the extent possible.
[0073] For example, the VR space-based image accompanying the prompt may be multiple two-dimensional images viewed from multiple different directions (e.g., forward, backward, right, left, etc., relative to the virtual user's position in the VR space). This is expected to improve the accuracy of the response from the generative AI-equipped server SV2. The prompt may also request the average value obtained from the multiple two-dimensional images, or the responses obtained from each of the multiple two-dimensional images may be averaged in the information processing device 1.
[0074] For example, the VR-based image accompanying the prompt may be a 360-degree image of the VR space viewed from a virtual user's position within that VR space. This allows for improved accuracy of responses from the SV2 server equipped with the generative AI, provided that the generative AI 222 has already learned 360-degree images. [Explanation of symbols]
[0075] 1. Information Processing Device 2...Display device 3. Audio output device SV1...Content delivery server SV2... AI-powered server for data generation SYS1...Content viewing system
Claims
1. Information processing equipment and Display device and It includes an audio output device, The aforementioned information processing device is The system displays at least a portion of the VR space, generates a display signal for displaying the image of 2D content on a virtual aerial display provided in the VR space, and outputs it to the display device. The system generates an audio signal by applying effect processing to the audio of the 2D content according to the characteristics of the VR space, and outputs it to an audio output device. Content viewing system.
2. The content viewing system according to claim 1, wherein the information processing device performs the effect processing according to the distance between the user's virtual position in the VR space and the aerial display.
3. The content viewing system according to claim 1 or 2, wherein the information processing device performs the effect processing according to the characteristics of the VR space using the response from a generative AI that has been given a prompt accompanied by an image based on the VR space.
4. The content viewing system according to claim 3, wherein the response from the generation AI includes at least one of the size of the VR space, the equivalent sound absorption area of the VR space, the sound absorption coefficient of the floor surface of the VR space, the sound absorption coefficient of the walls of the VR space, and the reverberation time in the VR space.
5. The content viewing system according to claim 3, wherein the image based on the VR space associated with the prompt is a two-dimensional image of the VR space viewed from one direction.
6. The content viewing system according to claim 3, wherein the image based on the VR space associated with the prompt is a plurality of two-dimensional images of the VR space viewed from a plurality of different directions.
7. The content viewing system according to claim 3, wherein the image based on the VR space associated with the prompt is a 360-degree image of the VR space viewed from all angles.
8. At least a portion of the VR space is displayed on a display device, and the image of the 2D content is displayed on a virtual aerial display provided in the VR space. The audio output device outputs audio based on an audio signal that has been processed with effects according to the characteristics of the VR space for the audio of the 2D content. How to view the content.
9. To display at least a portion of the VR space and generate a display signal for displaying the image of 2D content on a virtual aerial display provided in the VR space. To generate an audio signal by applying effect processing to the audio of the 2D content according to the characteristics of the VR space, A program that causes a computer to execute something.
Citation Information
Patent Citations
Scene simulator
JP1999327422A