Information processing device, information processing system, and information processing method

The information processing device automates the setting of realism parameters in XR content by detecting scenes and prioritizing audio and vibration parameters, enhancing the realism of XR content efficiently.

JP7778516B2Active Publication Date: 2025-12-02DENSO TEN LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021161742
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2025-12-02
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

Conventional methods for setting realistic sensation parameters in XR content require significant manual effort, which is inefficient.

Method used

An information processing device and method that automates the setting of realism parameters by detecting scenes from XR content, setting priorities, and extracting audio and vibration parameters accordingly.

Benefits of technology

Improves the efficiency of setting realistic sensation parameters, enhancing the realism of XR content by automating the process and ensuring appropriate audio and vibration outputs based on scene analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007778516000001
    Figure 0007778516000001
  • Figure 0007778516000002
    Figure 0007778516000002
  • Figure 0007778516000003
    Figure 0007778516000003
Patent Text Reader

Abstract

To increase efficiency of setting of a realistic sensation parameter relating to improvement of a realistic sensation of content.SOLUTION: An information processor includes a scene detection part, a parameter extraction part, and an output part. The scene detection part detects a scene from input content. The parameter extraction part extracts a realistic sensation parameter relating to wave control corresponding to a scene detected by the scene detection part. The output part outputs a wave signal relating to the content subjected to emphasis processing by the realistic sensation parameter extracted by the parameter extraction part.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing system, and an information processing method. [Background technology]

[0002] Conventionally, there is known technology that provides users with digital content that includes virtual space experiences such as VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality), known as XR (Cross Reality) content, using devices such as HMDs (Head Mounted Displays). XR is a collective term that encompasses all virtual space technologies, including VR, AR, and MR, as well as SR (Substitutional Reality) and AV (Audio / Visual).

[0003] Furthermore, a technique has been proposed in which vibrations corresponding to the video being viewed by the user are applied to the user, thereby improving the sense of realism of the video (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2004-081357 Summary of the Invention [Problem to be solved by the invention]

[0005] However, in the conventional technology, the realistic sensation parameters relating to the realistic sensation must be set manually in advance, and setting the realistic sensation parameters requires a huge amount of manual work.

[0006] The present invention has been made in consideration of the above, and aims to provide an information processing device, an information processing system, and an information processing method that can improve the efficiency of setting realism parameters related to improving the realism of content. [Means for solving the problem]

[0007] In order to solve the above-mentioned problems and achieve the object, an information processing device according to the present invention includes a scene detection unit, a parameter extraction unit, and an output unit. The scene detection unit detects scenes from input content. The parameter extraction unit extracts realism parameters related to wave control corresponding to the scenes detected by the scene detection unit. The output unit outputs a wave signal related to the content that has been emphasized using the realism parameters extracted by the parameter extraction unit. [Effects of the Invention]

[0008] According to the present invention, it is possible to improve the efficiency of setting realistic sensation parameters related to improving the realistic sensation of content. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating an overview of an information processing system. [Figure 2] FIG. 2 is a diagram showing an outline of the information processing method. [Figure 3] FIG. 3 is a block diagram of an information processing device. [Figure 4] FIG. 4 is a diagram showing an example of the scene information DB. [Figure 5] FIG. 5 is a diagram showing an example of the scene information DB. [Figure 6] FIG. 6 is a diagram showing an example of the scene information DB. [Figure 7] FIG. 7 is a diagram illustrating an example of the priority information DB. [Figure 8] FIG. 8 is a diagram illustrating an example of the parameter information DB. [Figure 9]FIG. 9 is a block diagram of the scene detection unit. [Figure 10] FIG. 10 is a block diagram of the priority setting unit. [Figure 11] FIG. 11 is a block diagram of the parameter extraction unit. [Figure 12] FIG. 12 is a block diagram of the output section. [Figure 13] FIG. 13 is a flowchart showing a processing procedure executed by the information processing device. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of an information processing device, an information processing system, and an information processing method disclosed in the present application will be described in detail with reference to the accompanying drawings. Note that the present invention is not limited to the embodiments described below.

[0011] First, an overview of an information processing system and an information processing method according to an embodiment will be described with reference to Figures 1 and 2. Figure 1 is a diagram showing an overview of the information processing system. Figure 2 is a diagram showing an overview of the information processing method. Note that the following describes a case where the XR space (virtual space) is a VR space.

[0012] As shown in FIG. 1, the information processing system 1 includes a display device 3, a speaker 4, and a vibration device 5.

[0013] The display device 3 is, for example, a head-mounted display, and is an information processing terminal that presents video data relating to XR content provided by the information processing device 10 to the user, allowing the user to enjoy a VR experience.

[0014] The display device 3 may be a non-transparent type that completely covers the field of view, or a video-transparent type or an optical-transparent type. The display device 3 also has a device, such as a camera or a motion sensor, that detects changes in the user's internal and external circumstances using a sensor unit.

[0015] The speaker 4 is an audio output device that outputs audio, and is provided, for example, as a headphone type, and is worn on the user's ear. The speaker 4 generates audio data provided from the information processing device 10 as audio. Note that the speaker 4 is not limited to a headphone type, and may also be a box type (installed on the floor, etc.). The speaker 4 may also be a stereo audio or multi-channel audio type.

[0016] The vibration device 5 is composed of an electric vibration converter made up of an electric / magnetic circuit and a piezoelectric element, and is provided, for example, on a seat on which a user sits, and vibrates in accordance with vibration data provided from the information processing device 10. Note that, for example, a plurality of vibration devices 5 are provided on the seat, and the information processing device 10 controls each vibration device 5 individually.

[0017] By applying the sound from the speaker 4 and the vibration from the vibration device 5, that is, the waves from the wave device, to the content user in a manner that is suited to the reproduced video, it is possible to enhance the sense of realism of the video reproduction.

[0018] The information processing device 10 is configured by a computer, is connected to the display device 3 via a wired or wireless connection, and provides images of XR content to the display device 3. The information processing device 10 also acquires changes in the situation detected by a sensor unit provided in the display device 3 at any time, and reflects such changes in the situation in the XR content.

[0019] For example, the information processing device 10 can change the direction of the field of view in the virtual space of the XR content in accordance with changes in the user's head or line of sight detected by the sensor unit.

[0020] Incidentally, when providing XR content, the sense of realism of the XR content can be improved by emphasizing the sound emitted from the speaker 4 in accordance with the scene, or by vibrating the vibration device 5 in accordance with the scene.

[0021] However, the parameters used to control the sense of realism to improve the sense of realism (hereinafter referred to as "sense of realism parameters") had to be set manually after the XR content was produced, which required a huge amount of work to set the sense of realism parameters.

[0022] Therefore, the information processing method aims to automate the setting of these realism parameters. For example, as shown in Figure 2, the information processing method according to the embodiment first detects scenes that satisfy predetermined conditions from video data and audio data related to XR content (step S1).

[0023] The predetermined condition here is, for example, a condition regarding whether the corresponding video data or audio data is a scene that requires the setting of a realism parameter, and is defined, for example, by a conditional expression regarding the situation within the XR content.

[0024] In other words, in the information processing method, when the situation within the XR content satisfies the conditions defined by the conditional expressions, the scene is detected as satisfying the predetermined conditions. This eliminates the need for detailed analysis of the video data, thereby reducing the processing load of scene detection.

[0025] Next, in the information processing method, a priority order is set for the scenes detected by the scene detection (step S2). Here, the priority order indicates the order in which the realism parameter of a scene should be prioritized. That is, in the information processing method, when multiple scenes overlap in time, the realism parameter of a scene that should be prioritized is defined in advance for each scene.

[0026] As a result, even when multiple scenes overlap, it is possible to provide the user with an appropriate sense of realism. As will be described later, in the information processing method, priorities for audio and vibration are set separately.

[0027] Next, in the information processing method, a realism parameter is extracted for each scene (step S3). For example, in the information processing method, a realism parameter is extracted for each scene using parameter information in which the relationship between the scene and the realism parameter is defined in advance.

[0028] In this case, the information processing method extracts a corresponding realism parameter according to the priority. Specifically, for example, in the case where a scene with a low priority overlaps with a scene with a high priority, the information processing method extracts the realism parameter of the scene with the high priority.

[0029] In the information processing method, a voice enhancement process is performed to enhance the voice data using a voice enhancement parameter from among the extracted realism parameters (step S4), and the voice data is output to the speaker 4. In addition, in the information processing method, a vibration conversion process is performed to convert the voice data into vibration data, and the vibration data is enhanced using a vibration parameter from among the extracted realism parameters (step S5), and the vibration data is output to the vibration device 5.

[0030] As a result, the information processing method can provide the user with sound that is emphasized in accordance with the scene that the user is viewing, and vibration that is appropriate for the scene.

[0031] In this way, the information processing method according to the embodiment detects scenes from XR content, sets priorities, and then extracts realistic sensation parameters related to wave control, including audio processing and vibration processing, for the scenes. Therefore, the information processing method according to the embodiment can automate the setting of realistic sensation parameters related to improving the realistic sensation of content.

[0032] Next, an example of the configuration of the information processing device 10 according to the embodiment will be described with reference to Fig. 3. Fig. 3 is a block diagram of the information processing device 10. As shown in Fig. 3, the information processing device 10 includes a control unit 120 and a storage unit 130.

[0033] The storage unit 130 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. In the example of FIG. 3, the storage unit 130 has an XR content DB (Database) 131, a scene information DB 132, a priority order information DB 133, and a parameter information DB 134.

[0034] The XR content DB 131 is a database that stores a group of XR contents to be displayed on the display device 3. The scene information DB 132 is a database that stores various information related to the scene to be detected.

[0035] 4 to 6 are diagrams showing an example of the scene information DB 132. As shown in Fig. 4, for example, the scene information DB 132 stores information on items such as "detected scene," "condition category," "object," "condition parameter," "threshold," and "condition formula" in association with one another.

[0036] "Detected Scene" indicates the name of the scene to be detected. Note that "Detected Scene" acts as an identification symbol, and although a code such as a number is normally used, in this example, a name (duplication prohibited) is used to make the explanation easier to understand. "Condition Category" indicates the category based on what information is used to detect the scene. In the example shown in the figure, the categories are broadly divided into categories based on the positional relationship between the user and the object, the user's actions, and information about the space or time in which the user is present. Note that the user here refers to the operator in the XR space.

[0037] "Object" refers to an object for scene detection. In the example shown in the figure, information such as Object 1, Object 2, User, Space 1, Space 1 + Object 3, and Content 1 corresponds to the object. Here, Object 1, Object 2, and Object 3 each represent different objects in the XR space. Furthermore, Space 1 represents, for example, the space in the XR space where the user exists, and Content 1 represents, for example, a predetermined event in the XR space.

[0038] "Condition parameters" indicate conditions related to parameters, such as which parameters are used for scene detection. As shown in the figure, information such as distance, angle, speed, acceleration, rotation speed, space, object presence, quantity, start time, end time, etc. is associated with each parameter.

[0039] The "threshold" indicates a threshold value corresponding to the condition parameter. The "condition expression" indicates a condition expression for detecting a detected scene, and for example, the relationship between the condition parameter and the threshold value is defined as a condition expression.

[0040] Furthermore, the information processing device 10 may detect a scene by combining the condition categories or condition parameters shown in Fig. 4. For example, as shown in Fig. 5, a detected scene may be set by combining the condition categories of a plurality of scenes, or as shown in Fig. 6, a detected scene may be set by combining the condition parameters of a plurality of scenes.

[0041] For example, by combining condition categories and condition parameters in this way, it is possible to simplify the setting of a new detection scene.

[0042] Returning to the explanation of Fig. 3, the priority information DB 133 will be described. For example, in the information processing device 10 according to the embodiment, a priority is set for each scene on a rule basis. The priority information DB 133 stores various information related to the priority of the realism parameters. Fig. 7 is a diagram showing an example of the priority information DB 133.

[0043] 7, for example, the priority information DB 133 stores information such as "rule number" and "priority rule" in association with each other. The "rule number" indicates a number for identifying a priority rule, and the "priority rule" indicates a rule regarding priority.

[0044] The "give priority to the first detected scene" and "give priority to the last detected scene (switch to the next scene)" shown in the figure indicate that the realism parameters of the first or last scene are given priority, respectively. This makes it possible to simplify the rules for setting scene priorities, for example.

[0045] Furthermore, "give priority to a specific parameter with a larger weight" indicates that priority is given to the realism parameter of a scene in which either the voice emphasis parameter or the vibration parameter is larger among the realism parameters.

[0046] In other words, in this case, the extracted realism parameter is set for the scene in which the voice emphasis parameter or the vibration parameter is larger, so that it is possible to provide a realism parameter that is linked to the voice data or vibration data to be emphasized.

[0047] Furthermore, "prioritize the parameter with the larger weight" indicates that among the realism parameters, the realism parameter of the scene with the larger voice emphasis parameters or the larger vibration parameters is given priority. In this rule, different scene parameters may be used for the voice emphasis parameters and the vibration parameters.

[0048] In other words, in this case, the vibration data and the audio data can be emphasized by the realism parameter having a large value, so that the realism of each of the vibration data and the audio data can be improved. Note that the magnitude of the weight here indicates, for example, the magnitude of the parameter value.

[0049] Furthermore, "prioritize parameters of shorter scenes" indicates that priority is given to the realism parameters of scenes with shorter durations. When a shorter scene interrupts playback of a longer scene, the realism parameters of the shorter scene are given priority during playback.

[0050] This allows, for example, scenes with a short duration to be appropriately emphasized. Note that a rule may be set to give priority to parameters for scenes with a long duration.

[0051] Returning to the explanation of Fig. 3, the parameter information DB 134 will be explained. The parameter information DB 134 is a database that stores information relating to the realism parameters for each scene. Fig. 8 is a diagram showing an example of the parameter information DB 134.

[0052] As shown in FIG. 8, the parameter information DB 134 stores information items such as "scene name," "voice emphasis parameter," and "vibration parameter" in association with one another.

[0053] The "scene name" indicates the name of the detected scene described above, and corresponds to, for example, the "detected scene" shown in Figure 4. Note that, for ease of understanding, the "scene name" is shown here as an explosion scene or a concert hall scene.

[0054] "Speech enhancement parameters" indicate speech enhancement parameters to be set for the corresponding scene. For example, as shown in Fig. 8, speech enhancement parameters are stored for each speaker 4 according to the number of speakers 4, such as "for speaker 1," "for speaker 2," etc.

[0055] Also, for each speaker 4, parameter values ​​related to audio processing, such as "delay" and "band emphasis / attenuation," are stored. For example, "delay" indicates a parameter related to the delay time, and "band emphasis / attenuation" indicates a parameter related to the extent to which sound in which band is emphasized or attenuated.

[0056] The "vibration parameters" indicate voice enhancement parameters to be set in the corresponding scene, and similarly to the "voice enhancement parameters," individual parameters are stored for each vibration device 5 according to the number of vibration devices 5. As the "vibration parameters," for example, parameters for items such as "LPF (Low Pass Filter)," "delay," and "amplification" are stored.

[0057] "LPF" indicates a parameter related to a low-pass filter, "delay" indicates a parameter related to a delay time, and "amplification" indicates a parameter related to vibration processing, such as the degree of amplification or attenuation.

[0058] Returning to the explanation of Fig. 3, the control unit 120 will be described. The control unit 120 is a controller, and is realized by, for example, a CPU (Central Processing Unit) or an MPU (Micro Processing Unit) executing various programs (not shown) stored in the storage unit 11 using a RAM as a work area. The control unit 120 can also be realized by, for example, an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0059] The control unit 120 has a content generation unit 121, a rendering processing unit 122, a scene detection unit 123, a priority setting unit 124, a parameter extraction unit 125, and an output unit 126, and realizes or executes the information processing functions and actions described below.

[0060] The content generation unit 121 generates a 3D model of the space within the XR content. For example, the content generation unit 121 references the XR content DB 131 and generates a 3D model of the space within the XR content in accordance with the user's current field of view within the XR content. The content generation unit 121 passes the generated 3D model to the rendering processing unit 122.

[0061] The rendering processing unit 122 performs rendering processing to convert the 3D model received from the content generation unit 121 into video data and audio data. For example, the rendering processing unit 122 outputs the converted video data to the display device 3 (see FIG. 1 ) and passes it to the scene detection unit 123. The rendering processing unit 122 also passes the converted audio data to the output unit 126 and the scene detection unit 123. The content generation unit 121 and the rendering processing unit 122 function as calculation units that calculate condition data for items of a conditional expression from the content.

[0062] The scene detection unit 123 detects scenes that satisfy predetermined conditions from the input content. For example, the scene detection unit 123 detects scenes for which a sense of presence parameter should be set, using video data input from the rendering processing unit 122 and a conditional expression stored in the scene information DB 132.

[0063] At this time, for example, the scene detection unit 123 receives coordinate information of objects in the XR space and information about the object type from the rendering processing unit 122, and detects a scene for which a realism parameter should be set using a conditional expression.

[0064] In addition, for example, when the XR content is MR content, the scene detection unit 123 may recognize objects in the MR space or calculate the coordinates of the objects by performing image analysis on images captured in the MR space.

[0065] Fig. 9 is a block diagram of scene detection unit 123. As shown in Fig. 9, scene detection unit 123 includes, for example, scene determination unit 123a and condition setting unit 123b. Scene determination unit 123a uses each condition data (conditional formula) for scene determination stored in scene information DB 132 to determine whether or not the situation in the video data satisfies the detection condition for each scene.

[0066] More specifically, for example, as shown in FIG. 4, the scene determination unit 123a determines whether the current situation in the XR space corresponds to each predefined detected scene based on data (calculated from the content by the content generation unit 121 or the rendering processing unit 122) for items in a conditional expression such as the positional relationship between the user and the target object (an object in the XR space), the user's actions, and spatial information in which the user exists.

[0067] Here, the scene determination unit 123a performs scene detection processing using text information-like data already calculated by the content generation unit 121 or the rendering processing unit 122, such as user movement in the XR space, object coordinate information, information about the object type, and spatial information.

[0068] This makes it possible to perform processes such as scene detection and extraction of realistic parameters in parallel with processes that have a relatively heavy processing load, such as rendering by the rendering processing unit 122, even if the CPU performance is relatively low.

[0069] In addition, at this time, for example, the scene determination unit 123a may determine whether or not the current situation in the XR space corresponds to each detected scene based on scene determination information that also includes a combination of condition categories as shown in FIG. 5, or a combination of condition parameters as shown in FIG. 6.

[0070] If the scene determination unit 123a determines that the video data corresponds to a detected scene, it passes the detected scene information for the video data to the priority order setting unit 124 (see FIG. 3). If the scene determination unit 123a determines that the video data does not correspond to any detected scene, it considers the detected scene to be not the relevant detected scene and returns the realism parameters to their initial states (realism parameters for the case where the detected scene is not the relevant detected scene). If the scene determination unit 123a determines that the current situation in the XR space corresponds to multiple detected scenes, it passes the determined multiple detected scenes to the priority order setting unit 124.

[0071] Furthermore, although the case where the scene determination unit 123a determines whether or not a scene is detected based on video data has been described here, the scene determination unit 123a may also determine whether or not a scene is detected based on audio data.

[0072] The condition setting unit 123b sets various conditional expressions for scene detection. The condition setting unit 123b sets the conditional expressions based on, for example, information input by the creator of the XR content or the user.

[0073] For example, the condition setting unit 123b receives input from a producer or a user, such as information about what kind of realism parameters are desired to be set for what kind of scenes, and converts the situation of the scenes into conditional expressions. Then, for each setting of a conditional expression, the condition setting unit 123b writes information about the conditional expression into the scene information DB 132 and writes the corresponding realism parameters into the parameter information DB 134.

[0074] This makes it possible for the information processing device 10 to detect scenes desired by the producer or user, and to set the sense of presence parameters desired by the producer or user for the detected scenes.

[0075] 3, the priority setting unit 124 will be described. The priority setting unit 124 sets priorities for the scenes detected by the scene detection unit 123.

[0076] For example, the priority setting unit 124 refers to the priority information DB 133 and selects which scene should be given priority in processing when multiple types of scenes are detected simultaneously by the scene detection unit 123. Note that when the scene detection unit 123 detects only one scene, that scene is given the highest priority.

[0077] Fig. 10 is a block diagram of the priority setting unit 124. For example, as shown in Fig. 10, the priority setting unit 124 includes a timing detection unit 124a and a rule setting unit 124b.

[0078] The timing detection unit 124a detects the timing at which a scene detected by the scene detection unit 123 occurs and the timing at which it ends. For example, the timing detection unit 124a detects each scene that exists at each time point (including whether it is overlapping), the timing at which an existing scene occurs, the timing at which an existing scene is deleted, etc., based on the scene information at each time point from the scene detection unit 123. In other words, the timing detection unit 124a grasps the state of all scenes that exist at each time point, including the order in which they occur.

[0079] The rule setting unit 124b sets a priority order for the scenes detected by the scene detection unit 123 to be used to determine the realism parameters. That is, based on the states of all existing scenes grasped by the timing detection unit 124a, the rule setting unit 124b sets a priority order for the detected scenes to determine which parameters associated with which scenes should be used preferentially as the realism parameters to be used at that time. This allows the information processing device 10 to set realism parameters according to the priority order.

[0080] That is, by setting a priority condition for each scene in advance, the information processing device 10 can appropriately determine which scene's realism parameter should be used preferentially when scene A and scene B overlap in time.

[0081] For example, the rule setting unit 124b sets the priority of the scene for determining the parameter to be used for each of the voice enhancement parameter and the vibration parameter with reference to the priority information DB 133. In this case, the rule setting unit 124b may set the scene to be used for parameter selection based on an independent priority rule for each speaker 4 and each vibration device 5, for example.

[0082] As a result, the realism parameters are set in accordance with unique rules for each speaker 4 and each vibration device 5, so that the realism can be further improved compared to when the realism parameters are set uniformly.

[0083] Furthermore, rule setting unit 124b associates information about the set rules with the video data and audio data and passes the information to parameter extraction unit 125 (see FIG. 3).

[0084] 3, the parameter extraction unit 125 will be described. The parameter extraction unit 125 extracts a sense of presence parameter for the scene detected by the scene detection unit 123.

[0085] Fig. 11 is a block diagram of the parameter extraction unit 125. As shown in Fig. 11, the parameter extraction unit 125 includes a vibration parameter extraction unit 125a, a voice emphasis parameter extraction unit 125b, and a learning unit 125c.

[0086] The vibration parameter extraction unit 125a refers to the parameter information DB 134 and extracts vibration parameters corresponding to the scene that has been set to the highest priority by the priority setting unit 124. For example, the vibration parameter extraction unit 125a extracts vibration parameters corresponding to the "detected scene" that has been received from the priority setting unit 124 and has the highest priority, from the parameter information DB 134, thereby extracting vibration parameters corresponding to the scene.

[0087] At this time, the vibration parameter extraction unit 125a extracts vibration parameters corresponding to each of the vibration devices 5. This makes it possible to further improve the sense of realism compared to the case where vibration parameters are uniformly extracted.

[0088] The voice enhancement parameter extraction unit 125b refers to the parameter information DB 134 and extracts a voice enhancement parameter corresponding to the scene that has been assigned the highest priority by the priority setting unit 124. The voice enhancement parameter extraction unit 125b extracts a voice enhancement parameter individually for each speaker 4, and determines the voice enhancement parameter to be extracted based on the priority set by the priority setting unit 124 (based on the scene with the highest priority), similar to the vibration parameter extraction unit 125a.

[0089] The learning unit 125c learns the relationship between the scenes and the realism parameters stored in the parameter information DB 134. For example, the learning unit 125c learns the relationship between the scenes and the realism parameters by performing machine learning on each scene stored in the parameter information DB 134 and each corresponding realism parameter using, as learning data, the user's reaction to the realism control by the parameter, etc.

[0090] In this case, for example, the learning unit 125c may use user evaluations of the realism parameters (user adjustment operations after realism control, user input such as questionnaires) as learning data. That is, the learning unit 125c may learn the relationship between scenes and realism parameters from the viewpoint of what kind of realism parameters should be set for what kind of scenes to obtain high user evaluations (i.e., whether a high sense of realism is obtained).

[0091] Furthermore, the learning unit 125c can determine what realism parameters should be set when a new scene is input, based on the learning results. As a specific example, the realism parameters for a fireworks scene can be determined using the learning results of realism control for a similar situation, such as an explosion scene. It is also possible to learn rules regarding priority based on the presence or absence and degree of factors that change the priority in user adjustment operations after realism control or user input such as a questionnaire (for example, when the user adjustment operation brings the parameters closer to those corresponding to other scenes that exist simultaneously, or when there is a response in the questionnaire indicating that other scenes should be given priority).

[0092] This allows the information processing device 10 to automatically optimize, for example, rules regarding priority and realistic parameters.

[0093] 3, the following describes the output unit 126. The output unit 126 outputs the realism parameters extracted by the parameter extraction unit 125 to the speaker 4 and the vibration device 5.

[0094] Fig. 12 is a block diagram of the output unit 126. As shown in Fig. 12, the output unit 126 has a voice emphasis processing unit 126a and a voice vibration conversion processing unit 126b.

[0095] The voice enhancement processing unit 126a performs enhancement processing on the voice data received from the rendering processing unit 122 using the voice enhancement parameters extracted by the parameter extraction unit 125. For example, the voice enhancement processing unit 126a performs enhancement processing on the voice data by performing delay or band enhancement / attenuation processing based on the voice enhancement parameters.

[0096] At this time, the voice enhancement processing unit 126a performs voice enhancement processing for each speaker 4, and outputs the voice data that has been subjected to the voice enhancement processing to each corresponding speaker 4.

[0097] The sound vibration conversion processing unit 126b converts the sound data received from the rendering processing unit 122 into vibration data by performing band limiting processing suitable for vibration, such as LPF, etc. Furthermore, the sound vibration conversion processing unit 126b performs emphasis processing on the converted vibration parameters in accordance with the vibration parameters extracted by the parameter extraction unit 125.

[0098] For example, the sound vibration conversion processing unit 126b performs emphasis processing on the vibration data by performing frequency characteristic adding processing such as low frequency emphasis, and emphasis processing such as delay and amplification according to the vibration parameters.

[0099] At this time, the sound vibration conversion processing unit 126b performs vibration coordination processing for each vibration device 5, and outputs the vibration data that has been subjected to the vibration coordination processing to each corresponding vibration device 5.

[0100] Next, a processing procedure executed by the information processing device 10 according to the embodiment will be described with reference to Fig. 13. Fig. 13 is a flowchart showing the processing procedure executed by the information processing device 10. Note that the processing procedure shown below is repeatedly executed by the control unit 120.

[0101] The process of the flowchart shown in Fig. 13 is executed when the information processing system 1 is powered on (step S101). Then, first, an XR content setting process is executed (step S102). Note that the XR content setting process here includes, for example, various processes related to the initial settings of the device for playing XR content, the selection of XR content by the user, etc.

[0102] Next, the information processing device 10 starts playing the XR content (step S103) and performs scene detection processing on the XR content being played (step S104). Next, the information processing device 10 performs priority setting processing on the results of the scene detection processing (step S105) and executes realism parameter extraction processing (step S106).

[0103] The information processing device 10 then executes output processing of various vibration data or audio data that reflects the processing results of the realism parameter extraction processing (step S107). The information processing device 10 then determines whether the XR content has ended (step S108), and if it determines that the XR content has ended (step S108; Yes), ends the processing.

[0104] Furthermore, when the information processing device 10 determines in step S108 that the XR content has not ended (step S108; No), the information processing device 10 again proceeds to the processing of step S104.

[0105] As described above, the information processing device 10 according to the embodiment includes the scene detection unit 123, the parameter extraction unit 125, and the output unit 126. The scene detection unit 123 detects a scene from input content. The parameter extraction unit 125 extracts a sense of presence parameter related to wave control corresponding to the scene detected by the scene detection unit 123.

[0106] The output unit 126 outputs a wave signal related to the content that has been emphasized using the realism parameters corresponding to the scenes extracted by the parameter extraction unit 125. Therefore, according to the information processing device 10 according to the embodiment, it is possible to improve the efficiency of setting the realism parameters related to improving the realism of the content.

[0107] In the above-described embodiment, the content is XR content, but the present invention is not limited to this. That is, the content may be 2D video and audio, or video only, or audio only.

[0108] Further advantages and modifications will readily occur to those skilled in the art. Therefore, the invention in its broader aspects is not limited to the specific details and representative embodiments shown and described above. Accordingly, various modifications may be made without departing from the spirit or scope of the general inventive concept as defined by the appended claims and their equivalents. [Explanation of symbols]

[0109] 1. Information Processing Systems 3 Display device 4 speakers 5. Vibration Devices 10. Information processing equipment 121 Content Generation Department 122 Rendering Processor 123 Scene detection section 123a Scene determination section 123b Condition setting section 124 Priority Setting Section 124a Timing detection unit 124b Rule setting section 125 Parameter Extraction Unit 125a Vibration parameter extraction unit 125b Speech enhancement parameter extraction unit 125c Learning Department 126 Output section 126a Speech enhancement processing unit 126b Voice vibration conversion processing unit 131 XR Content DB 132 Scene Information DB 133 Priority Information DB 134 Parameter Information DB

Claims

1. An information processing device that outputs vibration data corresponding to content to a vibration device and vibrates the vibration device, the information processing device comprising: a control unit; The control unit Detecting scenes included in the input content from the input content; determining whether the scene detected at the same time during playback of the content is a single scene or a plurality of scenes; If the determination is single, the detected scene is determined as a scene to be used for extracting vibration parameters; If the determination is made in a plurality of cases, a scene to be used for extracting vibration parameters is determined from the plurality of detected scenes based on a predetermined priority rule; extracting vibration parameters corresponding to the determined scene from a vibration parameter database in which scenes and vibration parameters are stored in association with each other; generating the vibration data by processing the audio data in the content based on the extracted vibration parameters; Information processing device.

2. The control unit A scene information database in which scenes and conditions related to the contents are stored in association with each other is collated with the input contents to detect the scene, If the number of detected scenes is multiple, extracting the priority rule by comparing a priority information database in which priority rules for scenes are stored with the occurrence status of the detected multiple scenes. The information processing device according to claim 1 .

3. The control unit learning a relationship between the scene and a vibration parameter based on a user's reaction to the vibration of the vibration device, and updating the vibration parameter database; The information processing device according to claim 2 .

4. The vibration device is a plurality of vibration devices that impart vibration to the seat, The control unit extracting the vibration parameters for each of the plurality of vibration devices; generating the vibration data using vibration parameters corresponding to each of the plurality of vibration devices; 4. The information processing device according to claim 1, 2 or 3.

5. The control unit extracting speech enhancement parameters by comparing the determined scene with a speech enhancement parameter database in which scenes and speech enhancement parameters are stored in association with each other; generating audio data by processing the audio data in the content based on the extracted audio enhancement parameters; outputting the generated audio data to a speaker; 5. The information processing device according to claim 1.

6. an information processing device that plays XR content; a display device that displays an image in accordance with the image signal output from the information processing device; an audio output device that generates audio in response to an audio signal output from the information processing device; a vibration device that vibrates in response to a vibration signal output from the information processing device; Equipped with The information processing device includes: Detecting scenes included in the input content from the input content; determining whether the scene detected at the same time during playback of the content is a single scene or a plurality of scenes; If the determination is single, the detected scene is determined as a scene to be used for extracting vibration parameters; If the determination is made in a plurality of cases, a scene to be used for extracting vibration parameters is determined from the plurality of detected scenes based on a predetermined priority rule; extracting vibration parameters corresponding to the determined scene from a vibration parameter database in which scenes and vibration parameters are stored in association with each other; generating vibration data by processing audio data in the content based on the extracted vibration parameters; outputting a vibration signal based on the vibration data to the vibration device; Information processing system.

7. 1. An information processing method for outputting vibration data corresponding to content to a vibration device and vibrating the vibration device, comprising: detecting scenes included in the input content from the input content; determining whether the number of scenes detected at the same time during playback of the content is one or more; If the determination is that there is only one scene, determining the detected scene as a scene to be used for extracting vibration parameters; If the determination is made in a plurality of cases, a step of determining a scene to be used for extracting vibration parameters from the plurality of detected scenes based on a predetermined priority rule; extracting vibration parameters corresponding to the determined scene from a vibration parameter database in which scenes and vibration parameters are stored in association with each other; generating the vibration data by processing the audio data in the content based on the extracted vibration parameters; The information processing method performed by the controller.

Citation Information

Patent Citations

  • Bodily feeling experiencing video / sound system

    JP2004081357A

  • Airbag deployment mode notification device

    JP2010228581A

  • Acoustic processor and acoustic processing method

    JP2013243619A

  • Vibration device and vibration method

    JP2015231098A

  • Information processing apparatus, information processing method, and program

    WO2015198716A1