Information processing device, information processing method, and information processing program

The information processing device automates the setting of realism parameters in XR content by detecting scenes and calculating data, enhancing the efficiency and optimization of realism enhancement in XR content.

JP7843123B2Active Publication Date: 2026-04-09DENSO TEN LTD
View PDF 12 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing technologies require significant manual effort to set presence parameters for enhancing realism in XR content, leading to inefficiency.

Method used

An information processing device with a storage unit, calculation unit, and scene detection unit automates the setting of realism parameters by detecting scenes from input content using conditional expressions and calculating corresponding data.

Benefits of technology

Improves the efficiency of setting realism parameters, allowing for automated and optimized enhancement of XR content realism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007843123000001
    Figure 0007843123000001
  • Figure 0007843123000002
    Figure 0007843123000002
  • Figure 0007843123000003
    Figure 0007843123000003
Patent Text Reader

Abstract

To provide an information processing device capable of improving the efficiency in realism parameter setting for improving the realism of a content.SOLUTION: An information processing device in an embodiment includes a storage unit, a calculation unit, and a scene detection unit. The storage unit stores a conditional expression for detecting a scene from a content. The calculation unit calculates condition data for the item of the conditional expression from the content. The scene detection unit detects a scene from a content using the conditional expression stored in the storage unit and the condition data calculated by the calculation unit.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0006] ,

[0001] The present invention relates to an information processing apparatus, information processing method and information processing program related thereto.

Background Art

[0002] Conventionally, technologies for providing digital content including virtual space experiences such as VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality) to users using an HMD (Head Mounted Display) or the like, so-called XR (Cross Reality) content, are known. XR is an expression that encompasses all virtual space technologies including SR (Substitutional Reality), AV (Audio / Visual), etc. in addition to VR, AR, and MR.

[0003] Also, for example, a technology has been proposed to enhance the sense of presence with respect to video by giving vibrations to the user according to the video the user is viewing (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the prior art, the presence parameters related to the sense of presence had to be set manually in advance, and a huge amount of manual work was required for setting the presence parameters.

[0006] The present invention has been made in view of the above, and aims to provide an information processing device, an information processing system, and an information processing method that can improve the efficiency of setting realism parameters related to enhancing the realism of content. [Means for solving the problem]

[0007] To solve the above-mentioned problems and achieve the objective, the information processing device according to the present invention comprises a storage unit, a calculation unit, and a scene detection unit. The storage unit stores a conditional expression for detecting a scene from input content. The calculation unit calculates conditional data for the items of the conditional expression from the content. The scene detection unit detects a scene in the content using the conditional expression stored in the storage unit and the conditional data calculated by the calculation unit. [Effects of the Invention]

[0008] According to the present invention, it is possible to improve the efficiency of setting realism parameters related to enhancing the realism of content. [Brief explanation of the drawing]

[0009] [Figure 1] Figure 1 is a diagram illustrating the overview of the information processing system. [Figure 2] Figure 2 shows an overview of the information processing method. [Figure 3] Figure 3 is a block diagram of the information processing device. [Figure 4] Figure 4 shows an example of a scene information database. [Figure 5] Figure 5 shows an example of a scene information database. [Figure 6] Figure 6 shows an example of a scene information database. [Figure 7] Figure 7 shows an example of a priority information database. [Figure 8] Figure 8 shows an example of a parameter information database. [Figure 9] Figure 9 is a block diagram of the scene detection unit. [Figure 10] FIG. 10 is a block diagram of the priority setting unit. [Figure 11] FIG. 11 is a block diagram of the parameter extraction unit. [Figure 12] FIG. 12 is a block diagram of the output unit. [Figure 13] FIG. 13 is a flowchart showing the processing procedure executed by the information processing apparatus.

BEST MODE FOR CARRYING OUT THE INVENTION

[0010] Hereinafter, embodiments of the information processing apparatus, information processing system, and information processing method disclosed in the present application will be described in detail with reference to the accompanying drawings. Note that the present invention is not limited by the embodiments described below.

[0011] First, the outline of the information processing system and information processing method according to the embodiment will be described using FIGS. 1 and 2. FIG. 1 is a diagram showing the outline of the information processing system. FIG. 2 is a diagram showing the outline of the information processing method. Hereinafter, the case where the XR space (virtual space) is the VR space will be described.

[0012] As shown in FIG. 1, the information processing system 1 includes a display device 3, a speaker 4, and a vibration device 5.

[0013] The display device 3 is, for example, a head-mounted display, and is an information processing terminal that presents video data related to XR content provided from the information processing apparatus 10 to the user and allows the user to enjoy a VR experience.

[0014] Note that the display device 3 may be a non-transmissive type that completely covers the field of view, or may be a video transmissive type or an optical transmissive type. Further, the display device 3 has a device that detects changes in the internal and external situations of the user by a sensor unit, such as a camera or a motion sensor.

[0015] The speaker 4 is an audio output device that outputs sound. For example, it is provided in a headphone type and is worn on the user's ear. The speaker 4 generates the audio data provided from the information processing device 10 as sound. Note that the speaker 4 is not limited to the headphone type and may be of a box type (installed on the floor or the like). Also, the speaker 4 may be of a stereo audio or multi-channel audio type.

[0016] The vibration device 5 is composed of an electric vibration converter composed of an electromagnetic circuit or a piezoelectric element. For example, it is provided on the seat on which the user sits and vibrates in accordance with the vibration data provided from the information processing device 10. Note that, for example, a plurality of vibration devices 5 are provided for the seat, and the information processing device 10 controls each vibration device 5 individually.

[0017] By adapting the sound from these speakers 4, the vibration of the vibration device 5, that is, the wave by the wave device, to the reproduced video and applying it to the content user, it is possible to increase the sense of presence regarding video playback.

[0018] The information processing device 10 is composed of a computer and is connected to the display device 3 by wire or wirelessly, and provides the video of the XR content to the display device 3. Also, the information processing device 10 acquires, at any time, the change in the situation detected by, for example, the sensor unit provided in the display device 3, and reflects such a change in the situation in the XR content.

[0019] For example, the information processing device 10 can change the direction of the field of view in the virtual space of the XR content according to the change in the user's head or line of sight detected by the sensor unit.

[0020] By emphasizing the sound generated from the speaker 4 according to the scene or vibrating the vibration device 5 according to the scene when providing the XR content, it is possible to improve the sense of presence of the XR content.

[0021] However, the parameters used to control the sense of presence in order to enhance this sense of presence (hereinafter referred to as "sense of presence parameters") had to be set manually after the XR content was created, and setting these sense of presence parameters required an enormous amount of work.

[0022] Therefore, the information processing method aims to automate the setting of these realism parameters. For example, as shown in Figure 2, first, in the information processing method according to the embodiment, a scene that satisfies predetermined conditions is detected from the video data and audio data related to the XR content (step S1).

[0023] The predetermined conditions here refer to, for example, conditions relating to whether the corresponding video or audio data is a scene requiring the setting of presence parameters, and are defined, for example, by a conditional expression relating to the situation within the XR content.

[0024] In other words, the information processing method detects a scene as satisfying a predetermined condition when the situation within the XR content satisfies the conditions defined by the conditional expression. This eliminates the need for detailed analysis of video data in the information processing method, thereby reducing the processing load of scene detection.

[0025] Next, in the information processing method, a priority is set for the scenes detected by scene detection (step S2). Here, priority indicates the order in which the presence parameters of each scene should be given priority. In other words, in the information processing method, when multiple scenes overlap in time, the priority of which scene's presence parameters should be given is defined in advance for each scene.

[0026] This allows for the provision of an appropriate sense of realism to the user, even when multiple scenes overlap. As will be described later, the information processing method sets separate priorities for audio and vibration.

[0027] Next, the information processing method extracts presence parameters for each scene (step S3). For example, the information processing method extracts presence parameters for each scene using parameter information in which the relationship between the scene and the presence parameters is defined in advance.

[0028] In this process, the information processing method extracts the corresponding realism parameters according to their priority. Specifically, for example, if a low-priority scene and a high-priority scene overlap, the information processing method will extract the realism parameters from the high-priority scene.

[0029] In the information processing method, the audio data is enhanced using the audio enhancement parameter from the extracted presence parameters (step S4), and then output to speaker 4. In addition, the information processing method performs a vibration conversion process to convert the audio data into vibration data, and the vibration data is enhanced using the vibration parameter from the extracted presence parameters (step S5), and then output to vibration device 5.

[0030] This allows the information processing method to provide the user with audio that is emphasized according to the scene they are viewing, as well as vibrations that correspond to the scene.

[0031] Thus, in the information processing method according to this embodiment, scenes are detected from XR content, priorities are set, and then presence parameters related to wave control, including audio processing and vibration processing, are extracted for each scene. Therefore, according to the information processing method according to this embodiment, the setting of presence parameters related to improving the presence of content can be automated.

[0032] Next, an example of the configuration of the information processing device 10 according to the embodiment will be described using Figure 3. Figure 3 is a block diagram of the information processing device 10. As shown in Figure 3, the information processing device 10 comprises a control unit 120 and a storage unit 130.

[0033] The storage unit 130 is implemented by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or by storage devices such as hard disks and optical discs. In the example shown in Figure 3, the storage unit 130 includes an XR content DB (Database) 131, a scene information DB 132, a priority information DB 133, and a parameter information DB 134.

[0034] The XR content DB131 is a database that stores the XR content set to be displayed on the display device 3. The scene information DB132 is a database that stores various information about the scene to be detected.

[0035] Figures 4 to 6 show an example of the scene information DB132. As shown in Figure 4, for example, the scene information DB132 stores information on items such as "detected scene," "condition category," "object," "condition parameter," "threshold," and "condition expression" in a manner that is associated with each other.

[0036] "Detected Scene" indicates the name of the scene to be detected. While "Detected Scene" typically uses numerical codes as identification symbols, this example uses a name (no duplicates allowed) for clarity. "Condition Category" indicates the category of information used to detect the scene. In the example shown, these categories are broadly divided into the positional relationship between the user and the object, the user's actions, the spatial information of the user's presence, or the temporal information of the user's presence. Here, "user" refers to the operator within the XR space.

[0037] "Object" refers to the object used for scene detection. In the example shown in the figure, information such as Object 1, Object 2, User, Space 1, Space 1 + Object 3, and Content 1 correspond to the object. Here, Object 1, Object 2, and Object 3 each represent different objects in the XR space. Space 1, for example, represents the space in the XR space where the user is located, and Content 1, for example, represents a predetermined event in the XR space.

[0038] "Conditional parameters" indicate the conditions for the parameters used in scene detection. As shown in the figure, information such as distance, angle, velocity, acceleration, rotational speed, presence in space, presence of objects, quantity, and start time to end time can be associated with these parameters.

[0039] The "threshold" indicates the threshold corresponding to the condition parameter. The "condition expression" indicates the condition expression for detecting the detection scene; for example, the relationship between the condition parameter and the threshold is defined as the condition expression.

[0040] Furthermore, the information processing device 10 may detect scenes by combining, for example, the condition categories or condition parameters shown in Figure 4. For example, as shown in Figure 5, a detection scene may be set by combining condition categories for multiple scenes, or as shown in Figure 6, a detection scene may be set by combining condition parameters for multiple scenes.

[0041] For example, by combining condition categories and condition parameters in this way, the setup of new detection scenes can be simplified.

[0042] Returning to the explanation of Figure 3, let's describe the priority information DB133. For example, in the information processing device 10 according to this embodiment, priority is set for each scene in a rule-based manner. The priority information DB133 stores various information regarding the priority of the realism parameters. Figure 7 shows an example of the priority information DB133.

[0043] As shown in Figure 7, for example, the priority information DB133 stores information such as "rule number" and "priority rule" in a related manner. The "rule number" indicates a number used to identify a priority rule, and the "priority rule" indicates a rule related to priority.

[0044] The options shown in the diagram, "Prioritize the scene detected first" and "Prioritize the scene detected later (switch when it becomes a later scene)," indicate that the presence parameters of the scene that appears earlier or later in time will be prioritized, respectively. This makes it easier to set rules, for example, when setting scene priorities.

[0045] Furthermore, "Prioritize the parameter with the larger weight" indicates that the presence parameter of the scene with the larger of either the sound enhancement parameter or the vibration parameter will be prioritized.

[0046] In other words, in this case, the sense of presence parameter extracted from the scene with the larger audio emphasis parameter or vibration parameter is set, so that a sense of presence parameter linked to the audio data or vibration data to be emphasized can be provided.

[0047] Furthermore, "prioritize the parameter with the larger weight" means that, among the presence parameters, the presence parameter of the scene with the larger weight among the audio enhancement parameters or among the vibration parameters will be prioritized. In this rule, it is possible that parameters from different scenes may be used for the audio enhancement parameter and the vibration parameter.

[0048] In other words, in this case, both vibration data and audio data can be emphasized with a large presence parameter, thereby improving the sense of presence of both the vibration data and the audio data. Here, the magnitude of the weights refers, for example, to the magnitude of the parameter values.

[0049] Furthermore, "Prioritize parameters of the shorter scene" means that the immersion parameters of the shorter scene will be prioritized. When a shorter scene interrupts a longer scene during playback, the immersion parameters of the shorter scene will be prioritized during the playback of that scene.

[0050] This allows, for example, to appropriately highlight shorter scenes. Alternatively, you could set a rule that prioritizes the parameters of longer scenes.

[0051] Returning to the explanation of Figure 3, let's describe the parameter information DB134. The parameter information DB134 is a database that stores information about the realism parameters for each scene. Figure 8 shows an example of the parameter information DB134.

[0052] As shown in Figure 8, the parameter information DB134 stores information on items such as "scene name," "sound enhancement parameter," and "vibration parameter" in a manner that associates them with each other.

[0053] The "scene name" indicates the name of the detected scene as described above, and corresponds to the "detected scene" shown in Figure 4, for example. For the sake of clarity, the "scene name" is shown here as "explosion scene" or "concert hall scene."

[0054] The "Audio Enhancement Parameters" indicate the audio enhancement parameters to be set in the corresponding scene. For example, as shown in Figure 8, the audio enhancement parameters store individual parameters for each speaker 4, such as "For Speaker 1" and "For Speaker 2," depending on the number of speakers 4.

[0055] Furthermore, for each speaker 4, the system stores parameter values ​​for audio processing items such as "delay" and "bandwidth enhancement / attenuation." For example, "delay" indicates a parameter related to the delay time, and "bandwidth enhancement / attenuation" indicates parameters such as which frequency bands to enhance or attenuate and to what extent.

[0056] The "Vibration Parameters" indicate the audio enhancement parameters to be set in the corresponding scene. Similar to the "Audio Enhancement Parameters," individual parameters are stored for each vibration device 5, depending on the number of vibration devices 5. For example, parameters for items such as "LPF (Low Pass Filter)," "Delay," and "Amplification" are stored as "Vibration Parameters."

[0057] "LPF" indicates parameters related to the low-pass filter, and "Delay" indicates parameters related to the delay time. "Amplification" indicates parameters related to vibration processing, such as the degree of amplification or attenuation.

[0058] Returning to the explanation of Figure 3, let's describe the control unit 120. The control unit 120 is a controller, and is realized, for example, by a CPU (Central Processing Unit) or MPU (Micro Processing Unit) executing various programs (not shown) stored in the memory unit 11 using RAM as the working area. Alternatively, the control unit 120 can also be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array).

[0059] The control unit 120 includes a content generation unit 121, a rendering processing unit 122, a scene detection unit 123, a priority setting unit 124, a parameter extraction unit 125, and an output unit 126, and realizes or executes the information processing functions and operations described below.

[0060] The content generation unit 121 generates a 3D model of the space within the XR content. For example, the content generation unit 121 refers to the XR content DB 131 and generates a 3D model of the space within the XR content to match the user's current field of view within the XR content. The content generation unit 121 passes the generated 3D model to the rendering processing unit 122.

[0061] The rendering processing unit 122 performs rendering processing to convert the 3D model received from the content generation unit 121 into video data and audio data. For example, the rendering processing unit 122 outputs the converted video data to the display device 3 (see Figure 1) and also passes it to the scene detection unit 123. The rendering processing unit 122 also passes the converted audio data to the output unit 126 and the scene detection unit 123. The content generation unit 121 and the rendering processing unit 122 also function as calculation units that calculate condition data for items in conditional expressions from the content.

[0062] The scene detection unit 123 detects scenes that meet predetermined conditions from the input content. For example, the scene detection unit 123 uses video data input from the rendering processing unit 122 and conditional expressions stored in the scene information DB 132 to detect scenes for which the realism parameter should be set.

[0063] In this process, for example, the scene detection unit 123 receives coordinate information and object type information of objects in the XR space from the rendering processing unit 122, and uses a conditional expression to detect a scene for which the sense of presence parameters should be set.

[0064] Furthermore, if the XR content is MR content, the scene detection unit 123 may, for example, perform image analysis on images taken within the MR space to recognize objects in the MR space or calculate the coordinates of objects.

[0065] Figure 9 is a block diagram of the scene detection unit 123. As shown in Figure 9, for example, the scene detection unit 123 includes a scene determination unit 123a and a condition setting unit 123b. The scene determination unit 123a uses the condition data (conditional expression) for scene determination stored in the scene information DB 132 to determine whether the situation in the video data satisfies the detection conditions for each scene.

[0066] More specifically, as shown in Figure 4, the scene determination unit 123a determines whether the current state of the XR space corresponds to one of the predefined detection scenes, based on data (calculated by the content generation unit 121 or the rendering processing unit 122 from the content) for items in the conditional expression such as the positional relationship between the user and the object (object in the XR space), the user's actions, and spatial information in which the user exists.

[0067] Here, the scene determination unit 123a performs scene detection processing using textual data already calculated by the content generation unit 121 or the rendering processing unit 122, such as user movement in the XR space, object coordinate information and object type information, and spatial information.

[0068] This makes it possible to perform processing such as scene detection and extraction of realism parameters in parallel with relatively computationally intensive processing such as rendering by the rendering processing unit 122, even when the CPU performance is relatively low.

[0069] Furthermore, in this case, for example, the scene determination unit 123a may determine whether the current state of the XR space corresponds to each detected scene based on scene determination information that includes, for example, a combination of condition categories as shown in Figure 5, or a combination of condition parameters as shown in Figure 6.

[0070] Then, if the scene determination unit 123a determines that the video data corresponds to a detected scene, it passes the detected scene information for that video data to the priority setting unit 124 (see Figure 3). If the scene determination unit 123a determines that the video data does not correspond to any of the detected scenes, the presence parameter is reset to its initial state (the presence parameter for when the video data does not correspond to a detected scene). Furthermore, if the scene determination unit 123a determines that the current state of the XR space corresponds to multiple detected scenes, it passes the determined multiple detected scenes to the priority setting unit 124.

[0071] Furthermore, although this explanation describes a case where the scene determination unit 123a determines whether or not a scene is detected based on video data, the scene determination unit 123a may also determine whether or not a scene is detected based on audio data.

[0072] The condition setting unit 123b sets various conditional expressions for scene detection. The condition setting unit 123b sets conditional expressions based, for example, on information input from the creator of the XR content or from the user.

[0073] For example, the condition setting unit 123b receives input from the creator or user regarding what kind of realism parameters they want to set for what kind of scene, and translates the situation of that scene into a conditional expression. Then, for each conditional expression set, the condition setting unit 123b writes information related to the conditional expression to the scene information DB 132 and writes the corresponding realism parameters to the parameter information DB 134.

[0074] As a result, the information processing device 10 can detect scenes requested by the creator or user, and set the sense of presence parameters requested by the creator or user for the detected scenes.

[0075] Returning to the explanation of Figure 3, let's describe the priority setting unit 124. The priority setting unit 124 sets priorities for the scenes detected by the scene detection unit 123.

[0076] For example, the priority setting unit 124 refers to the priority information DB 133 and selects which scene to prioritize processing if the scene detection unit 123 detects multiple types of scenes simultaneously. If the scene detection unit 123 detects only one scene, that scene will have the highest priority.

[0077] Figure 10 is a block diagram of the priority setting unit 124. For example, as shown in Figure 10, the priority setting unit 124 includes a timing detection unit 124a and a rule setting unit 124b.

[0078] The timing detection unit 124a detects the timing at which scenes detected by the scene detection unit 123 occur and the timing at which they end. For example, based on the scene information from the scene detection unit 123 at each point in time, the timing detection unit 124a detects each scene that exists at each point in time (including duplicates), the timing at which existing scenes occur, and the timing at which existing scenes are deleted. In other words, the timing detection unit 124a understands the state of all scenes that exist at each point in time, including their order of occurrence.

[0079] The rule setting unit 124b sets the priority of scenes to be used to determine the sense of presence parameters for the scenes detected by the scene detection unit 123. In other words, based on the state of all existing scenes grasped by the timing detection unit 124a, it sets a priority for the detected scenes in order to determine which scene's associated parameters should be used preferentially for the sense of presence parameters to be used at that time. This allows the information processing device 10 to set the sense of presence parameters according to the said priority.

[0080] In other words, the information processing device 10 can set priority conditions for each scene in advance, so that when scene A and scene B overlap in time, it can appropriately determine which scene's sense of presence parameters should be prioritized.

[0081] For example, the rule setting unit 124b refers to the priority information DB 133 and sets the priority of the scenes that determine which parameters to use for each of the audio enhancement parameters and vibration parameters. In this case, the rule setting unit 124b may set the scenes to be used for parameter selection based on independent priority rules for each speaker 4 and each vibration device 5.

[0082] As a result, each speaker 4 and each vibration device 5 has its own rules for setting the sense of presence parameters, which allows for a further improvement in the sense of presence compared to setting the sense of presence parameters uniformly.

[0083] Furthermore, the rule setting unit 124b associates the set rule information with the video data and audio data and passes it to the parameter extraction unit 125 (see Figure 3).

[0084] Returning to the explanation of Figure 3, let's describe the parameter extraction unit 125. The parameter extraction unit 125 extracts presence parameters from the scene detected by the scene detection unit 123.

[0085] Figure 11 is a block diagram of the parameter extraction unit 125. As shown in Figure 11, the parameter extraction unit 125 includes a vibration parameter extraction unit 125a, a speech enhancement parameter extraction unit 125b, and a learning unit 125c.

[0086] The vibration parameter extraction unit 125a refers to the parameter information DB 134 and extracts vibration parameters corresponding to the scene that has been given the highest priority by the priority setting unit 124. For example, the vibration parameter extraction unit 125a extracts vibration parameters corresponding to the scene by extracting from the parameter information DB 134 the vibration parameters corresponding to the highest priority "detected scene" received from the priority setting unit 124.

[0087] In this process, the vibration parameter extraction unit 125a extracts the corresponding vibration parameters for each vibration device 5. This allows for a further improvement in the sense of realism compared to extracting vibration parameters uniformly.

[0088] The audio enhancement parameter extraction unit 125b refers to the parameter information DB 134 and extracts the audio enhancement parameters corresponding to the scene that has been given the highest priority by the priority setting unit 124. The audio enhancement parameter extraction unit 125b extracts audio enhancement parameters individually for each speaker 4 and, similar to the vibration parameter extraction unit 125a, determines the audio enhancement parameters to extract based on the priority set by the priority setting unit 124 (based on the scene with the highest priority).

[0089] The learning unit 125c learns the relationship between scenes stored in the parameter information DB 134 and the sense of presence parameters. For example, the learning unit 125c learns the relationship between each scene stored in the parameter information DB 134 and each corresponding sense of presence parameter by performing machine learning using user responses to the sense of presence control by those parameters as training data.

[0090] In this case, for example, the learning unit 125c may use user evaluations of the sense of presence parameters (user adjustment operations after sense of presence control, or user input such as questionnaires) as learning data. That is, the learning unit 125c may learn the relationship between scenes and sense of presence parameters from the perspective of what kind of scenes and what kind of sense of presence parameters to set to obtain high user evaluations (i.e., whether a high sense of presence was obtained).

[0091] Furthermore, the learning unit 125c can also determine from the learning results what realism parameters should be set when a new scene is input. For example, the realism parameters for a fireworks scene can be determined using the learning results of realism control for similar situations such as explosion scenes. It is also possible to learn rules regarding priority based on the presence and degree of elements that change priority in user adjustment operations after realism control or user input such as questionnaires (for example, when user adjustment operations bring parameters closer to those corresponding to other scenes that exist simultaneously, or when there are answers in the questionnaire that should prioritize other scenes).

[0092] This allows the information processing device 10 to automatically perform tasks such as setting priority rules and optimizing realism parameters.

[0093] Returning to the explanation of Figure 3, let's describe the output unit 126. The output unit 126 outputs the presence parameters extracted by the parameter extraction unit 125 to the speaker 4 and the vibration device 5.

[0094] Figure 12 is a block diagram of the output unit 126. As shown in Figure 12, the output unit 126 includes a voice enhancement processing unit 126a and a voice vibration conversion processing unit 126b.

[0095] The audio enhancement processing unit 126a performs enhancement processing on the audio data received from the rendering processing unit 122 using audio enhancement parameters extracted by the parameter extraction unit 125. For example, the audio enhancement processing unit 126a performs enhancement processing on the audio data by applying delay or bandwidth enhancement / attenuation processing based on the audio enhancement parameters.

[0096] In this process, the audio enhancement processing unit 126a performs audio enhancement processing for each speaker 4 and outputs the enhanced audio data to the corresponding speaker 4.

[0097] The audio-vibration conversion processing unit 126b converts the audio data received from the rendering processing unit 122 into vibration data by performing bandwidth limiting processing suitable for vibration, such as LPF. The audio-vibration conversion processing unit 126b also performs enhancement processing on the converted vibration parameters according to the vibration parameters extracted by the parameter extraction unit 125.

[0098] For example, the audio-vibration conversion processing unit 126b performs enhancement processing on vibration data by applying frequency characteristic addition processing such as low-frequency emphasis, delay, and amplification according to the vibration parameters.

[0099] In this process, the audio-vibration conversion processing unit 126b performs vibration coordination processing for each vibration device 5 and outputs the vibration data that has undergone vibration coordination processing to the corresponding vibration device 5.

[0100] Next, the processing procedure executed by the information processing device 10 according to this embodiment will be explained using Figure 13. Figure 13 is a flowchart of the processing procedure executed by the information processing device 10. Note that the processing procedure shown below is repeatedly executed by the control unit 120.

[0101] The flowchart shown in Figure 13 is executed when the power to the information processing system 1 is turned on (step S101). First, the XR content setting process is executed (step S102). The XR content setting process here includes various processes such as the initial settings of the device for XR content playback and the user's selection of XR content.

[0102] Next, the information processing device 10 starts playing the XR content (step S103) and performs scene detection processing on the XR content being played (step S104). Subsequently, the information processing device 10 performs priority setting processing on the results of the scene detection processing (step S105) and executes the realism parameter extraction processing (step S106).

[0103] Then, the information processing device 10 performs output processing of various vibration data or sound data that reflects the processing results of the presence parameter extraction process (step S107). Then, the information processing device 10 determines whether or not the XR content has ended (step S108), and if it determines that the XR content has ended (step S108; Yes), it terminates the process.

[0104] Furthermore, if the information processing device 10 determines in step S108 that the XR content has not finished (step S108; No), it proceeds back to the process in step S104.

[0105] As described above, the information processing device 10 according to the embodiment includes a scene information DB 132 (an example of a storage unit), a content generation unit 121 and a rendering processing unit 122 (an example of a calculation unit), and a scene detection unit 123. The scene information DB 132 stores conditional expressions for detecting scenes from input content. The content generation unit 121 and the rendering processing unit 122 calculate conditional data for items of the conditional expressions from the content.

[0106] The scene detection unit 123 detects scenes in the content using conditional expressions stored in the scene information DB 132 and conditional data calculated by the content generation unit 121 and the rendering processing unit 122. Therefore, according to the information processing device 10 of this embodiment, it is possible to improve the efficiency of setting realism parameters related to improving the realism of the content.

[0107] By the way, the above-described embodiment explained the case where the content is XR content, but it is not limited to this. In other words, the content may be 2D video and audio, or video only, or audio only.

[0108] Further effects and modifications can be readily derived by those skilled in the art. Therefore, broader aspects of the present invention are not limited to the specific details and representative embodiments expressed and described above. Accordingly, various modifications are possible without departing from the spirit or scope of the overall concept of the invention as defined by the appended claims and equivalents. [Explanation of Symbols]

[0109] 1. Information Processing System 3 Display device 4 speakers 5. Vibration Devices 10 Information Processing Devices 121 Content Generation Department 122 Rendering Processing Unit 123 Scene detection unit 123a Scene determination unit 123b Condition setting section 124 Priority Setting Section 124a Timing detection unit 124b Rule setting section 125 Parameter Extraction Unit 125a Vibration parameter extraction unit 125b Audio enhancement parameter extraction unit 125c Learning Department 126 Output section 126a Speech enhancement processing unit 126b Audio-Vibration Conversion Processing Unit 131 XR Content Database 132 Scene Information Database 133 Priority Information Database 134 Parameter Information DB

Claims

1. An information processing device that generates vibration data corresponding to a scene in content and outputs it to a vibration device, comprising a controller, The aforementioned controller, The spatial information of the user's location within the aforementioned content is detected, By comparing the aforementioned spatial information as one of the parameters with a conditional expression associated with the scene, the corresponding scene is detected. The aforementioned conditional expression includes a combination of multiple conditional expressions that are associated with multiple other scenes. Information processing device.

2. An information processing device that generates vibration data corresponding to a scene in content and outputs it to a vibration device, comprising a controller, The aforementioned controller, The user's presence time in the space within the aforementioned content is detected, By using the user's presence time as one of the parameters and comparing it with a conditional expression associated with the scene, the corresponding scene is detected. The aforementioned conditional expression includes a combination of multiple conditional expressions that are associated with multiple other scenes. Information processing device.

3. An information processing device that generates vibration data corresponding to a scene in content and outputs it to a vibration device, comprising a controller, The aforementioned controller, The user's movement speed in the space within the aforementioned content is detected, By using the user's movement speed as one of the parameters and comparing it with a conditional expression associated with the scene, the corresponding scene is detected. The aforementioned conditional expression includes a combination of multiple conditional expressions that are associated with multiple other scenes. Information processing device.

4. An information processing method that generates vibration data corresponding to a scene in content and outputs it to a vibration device, The spatial information of the user's location within the aforementioned content is detected, By comparing the aforementioned spatial information as one of the parameters with a conditional expression associated with the scene, the corresponding scene is detected. The aforementioned conditional expression includes a combination of multiple conditional expressions that are associated with multiple other scenes. Information processing methods.

5. An information processing method that generates vibration data corresponding to a scene in content and outputs it to a vibration device, The user's presence time in the space within the aforementioned content is detected, By using the user's presence time as one of the parameters and comparing it with a conditional expression associated with the scene, the corresponding scene is detected. The aforementioned conditional expression includes a combination of multiple conditional expressions that are associated with multiple other scenes. Information processing methods.

6. An information processing method that generates vibration data corresponding to a scene in content and outputs it to a vibration device, The user's movement speed in the space within the aforementioned content is detected, By using the user's movement speed as one of the parameters and comparing it with a conditional expression associated with the scene, the corresponding scene is detected. The aforementioned conditional expression includes a combination of multiple conditional expressions that are associated with multiple other scenes. Information processing methods.

7. An information processing program that generates vibration data corresponding to a scene in the content and outputs it to a vibration device, The steps include detecting spatial information where the user is located within the aforementioned content, The steps include: detecting a scene by comparing the aforementioned spatial information as one of the parameters with a conditional expression associated with the scene; and Have the computer run it, The aforementioned conditional expression includes a combination of multiple conditional expressions that are associated with multiple other scenes. Information processing program.

8. An information processing program that generates vibration data corresponding to a scene in the content and outputs it to a vibration device, The steps include detecting the user's presence time in the space within the content, The steps include: detecting a scene by comparing the user's presence time, which is used as one of the parameters, with a conditional expression associated with the scene; and Have the computer run it, The aforementioned conditional expression includes a combination of multiple conditional expressions that are associated with multiple other scenes. Information processing program.

9. An information processing program that generates vibration data corresponding to a scene in the content and outputs it to a vibration device, The steps include detecting the user's movement speed within the space of the content, The process involves detecting a scene by comparing the user's movement speed, used as one of the parameters, with a conditional expression associated with the scene. Have the computer run it, The aforementioned conditional expression includes a combination of multiple conditional expressions that are associated with multiple other scenes. Information processing program.

Citation Information

Patent Citations

  • Bodily feeling experiencing video / sound system

    JP2004081357A

  • Game device, game device control method and program

    JP2008295786A

  • Game image display control program, game device, and game image display control method

    JP2010115244A

  • Image capturing apparatus, image processing method, and program

    JP2011015256A

  • Virtual sensor in virtual environment

    JP2016126766A