Information processing apparatus, information processing system, and information processing method
Patent Information
- Application Number
- CN202210266841.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-02-09
- Filing Date
- 2022-03-17
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2042-03-17
AI Technical Summary
[0004]然而,在传统技术中,用于真实感的真实感参数必须由人工初步设置,因此需要大量人工来设置这样的真实感参数
Smart Images

Figure CN115877945B_ABST
Abstract
Description
Technical Field
[0001] The disclosed embodiments relate to information processing apparatus, information processing system, and information processing method. Background Technology
[0002] Technologies that provide users with digital content, including virtual space experiences, through head-mounted displays (HMDs) are traditionally known. These virtual space experiences include, for example, virtual reality (VR), augmented reality (AR), and / or mixed reality (MR), or so-called cross-reality (XR) content. XR is an expression that incorporates all virtual space technologies, including not only VR, AR, and MR, but also alternative reality (SR), audio / visual (AV), and so on.
[0003] Furthermore, for example, techniques have been proposed to provide users with vibrations that depend on the video being viewed by the user, in order to enhance the realism of such videos (see, for example, Japanese Patent Application Publication No. 2004-081357).
[0004] However, in traditional techniques, the realism parameters used to achieve a sense of realism must be initially set manually, thus requiring a large amount of manual labor to set these parameters. Summary of the Invention
[0005] An information processing apparatus according to one aspect of an embodiment includes a control unit configured to perform: scene detection processing to detect a scene from input content; parameter extraction processing to extract realistic parameters for wave control corresponding to the scene detected by the scene detection processing; and output processing to output a wave signal of the content, which is generated by processing the sound data of the input content using the realistic parameters extracted by the parameter extraction processing. Attached Figure Description
[0006] Figure 1 This is a diagram showing an overview of an information processing system.
[0007] Figure 2 This is a diagram illustrating an overview of an information processing method.
[0008] Figure 3 This is a block diagram of an information processing device.
[0009] Figure 4 This is a diagram showing an example of a scene information database.
[0010] Figure 5 This is a diagram showing an example of a scene information database.
[0011] Figure 6 This is a diagram showing an example of a scene information database.
[0012] Figure 7 This is a diagram showing an example of a priority information database.
[0013] Figure 8 This is a diagram showing an example of parameter information DB.
[0014] Figure 9 This is a block diagram of the scene detection unit.
[0015] Figure 10 This is a block diagram of the priority setting unit.
[0016] Figure 11 This is a block diagram of the parameter extraction unit.
[0017] Figure 12 This is a block diagram of the output unit.
[0018] Figure 13 It is a flowchart illustrating the processing procedure performed by the information processing device.
[0019] Figure 14 This is a diagram illustrating an example of a method for determining priority objects. Detailed Implementation
[0020] In the following, embodiments of the information processing apparatus, information processing system, and information processing method disclosed in this application will be described in detail with reference to the accompanying drawings. However, the present invention is not limited to the embodiments shown below.
[0021] First, it will be through the use of Figure 1 and Figure 2 This section outlines the information processing system and information processing method according to embodiments. Figure 1 This is a diagram showing an overview of an information processing system. Figure 2 This is a diagram illustrating an overview of the information processing method. Additionally, the following will explain the case where XR space (virtual space) is VR space.
[0022] like Figure 1 As shown, the information processing system 1 includes a display device 3, a speaker 4, and a vibration device 5.
[0023] The display device 3 is, for example, a head-mounted display, and is an information processing terminal for presenting video data of XR content provided by the information processing device 10 to the user so that the user can enjoy a VR experience.
[0024] Furthermore, the display device 3 can be a non-transmission type that completely covers the field of view, or it can be a video transmission type and / or an optical transmission type. Additionally, the display device 3 has a device for detecting changes in the user's internal and / or external environment via a sensor unit (e.g., a camera, motion sensor, etc.).
[0025] Speaker 4 is a sound output device that outputs sound, and is configured as, for example, an earphone type and worn on the user's ears. Speaker 4 generates sound data provided from information processing device 10 as sound. In addition, speaker 4 is not limited to earphone type, but can be a box type (which is mounted on the floor, etc.). Furthermore, speaker 4 can be stereo audio or multi-channel audio type.
[0026] The vibration device 5 consists of an electric vibration transducer, which is installed, for example, on a seat where a user sits, and vibrates according to vibration data provided from the information processing device 10. The electric vibration transducer consists of circuits, magnetic circuits, and / or piezoelectric elements. Alternatively, for example, multiple vibration devices 5 are installed on the seat, and the information processing device 10 controls each vibration device 5 individually.
[0027] The sound provided by the speaker 4 and / or the vibration of the vibration device 5 (i.e., the wave provided by the wave device) are suitable for reproducing the video and applied to the content user, thereby further increasing the realism of the video reproduction.
[0028] The information processing device 10, consisting of a computer, is connected to the display device 3 via wired or wireless means and provides XR content video to the display device 3. Furthermore, for example, the information processing device 10 acquires changes in conditions detected by sensors installed on the display device 3 as needed and reflects these changes in the XR content.
[0029] For example, the information processing device 10 can change the direction of the field of view in the virtual space of the XR content based on changes in the user's head and / or gaze detected by the sensor unit.
[0030] At the same time, when providing XR content, the sound generated from the speaker 4 can be enhanced according to the scene, or the vibration device 5 can be vibrated according to the scene, thereby improving the realism of the XR content.
[0031] However, the parameters used to achieve this enhanced realism (hereinafter referred to as realism parameters) must be manually set after the XR content is generated, which requires a lot of work to set the realism parameters.
[0032] Therefore, the setting of such realism parameters has been automated in information processing methods. For example, such as Figure 2 As shown, in the information processing method according to the embodiment, firstly, a scene that meets predetermined conditions is detected from the video data and audio data of the XR content (step S1).
[0033] The preconditions in this article are, for example, conditions regarding whether the corresponding video or audio data is a scene for which realism parameters need to be set, and are defined, for example, by conditional expressions for the internal conditions of XR content.
[0034] In other words, in the information processing method, if the internal conditions of the XR content satisfy the conditions defined by the conditional expression, it is detected as a scene that meets predetermined conditions. Therefore, the information processing method does not require detailed analysis of video data, thereby reducing the processing load of scene detection.
[0035] Then, in the information processing method, priorities are set for the scenes detected through scene detection (step S2). In this paper, priority represents the level of scenes with realistic parameters that should be given priority. That is, in the information processing method, when multiple scenes overlap in time, for each scene, the scenes with realistic parameters that should be given priority are initially defined.
[0036] Therefore, even when multiple scenes overlap, a suitable sense of realism can be provided to the user. Furthermore, as described later, in the information processing method, the priorities of sound and vibration are set separately.
[0037] Then, in the information processing method, realism parameters are extracted for each scene (step S3). For example, in the information processing method, realism parameters are extracted for each scene by using parameter information that initially defines the relationship between the scene and the realism parameters.
[0038] In this paper, the information processing method extracts corresponding realism parameters based on priority. Specifically, in the information processing method, for example, when low-priority scenes overlap with high-priority scenes, the realism parameters of the high-priority scenes are extracted.
[0039] In the information processing method, sound enhancement processing is performed to enhance the sound data by using sound enhancement parameters from the extracted realism parameters (step S4), and output is performed to the speaker 4. Furthermore, in the information processing method, after performing vibration conversion processing to convert the sound data into vibration data, and enhancing this vibration data by using vibration parameters from the extracted realism parameters (step S5), output is performed to the vibration device 5.
[0040] Therefore, in information processing methods, users can be provided with enhanced sound based on the scene the user is viewing and / or vibrations dependent on the scene.
[0041] Therefore, in the information processing method according to the embodiment, after detecting the scene from the XR content and setting a priority, realism parameters for wave control are extracted, which includes sound processing and vibration processing for the scene. Thus, in the information processing method according to the embodiment, realism parameters for improving the realism of the content can be automatically set.
[0042] Next, we will use Figure 3 Here is an example of the structure of the information processing apparatus 10 according to an embodiment. Figure 3 This is a block diagram of the information processing device 10. (For example...) Figure 3 As shown, the information processing device 10 includes a control unit 120 and a storage unit 130.
[0043] Storage unit 130 is implemented, for example, by semiconductor storage elements such as random access memory (RAM) and / or flash memory, or storage devices such as hard disks and / or optical disks. Figure 3 In the example, storage unit 130 has an XR content database (DB) 131, scene information DB 132, priority information DB 133, and parameter information DB 134.
[0044] XR Content DB 131 is a database that stores the XR content group displayed on the display device 3. Scene Information DB 132 is a database that stores various information about the detected scene.
[0045] Figures 4 to 6 This is a diagram illustrating an example of scene information DB 132. (As shown...) Figure 4 As shown, for example, the scene information DB132 stores information about the items "detection scene", "condition category", "object", "condition parameter", "threshold" and "condition expression" in a corresponding manner.
[0046] "Detected Scene" indicates the name of the scene being detected. Additionally, "Detected Scene" serves as an identification symbol; although codes such as numerical values are typically used, the name is used in this example for clarity (and duplication is prohibited). "Condition Category" indicates the category of information upon which scene detection is based. In the example shown in the same figure, a broad classification is performed, categorized into categories such as the positional relationship between the user and objects, user movement, information about the space where the user appears, and information about the time of the user's appearance. Furthermore, in this text, "user" refers to the operator themselves in the XR space.
[0047] "Object" refers to the object detected in the scene. In the example shown in the same figure, information such as object 1, object 2, user, space 1, space 1 + object 3, and / or content 1 correspond to objects. In this paper, object 1, object 2, and object 3 represent different objects in the XR space. Furthermore, space 1 represents, for example, the space where a user appears in the XR space, and content 1 represents, for example, a predetermined event in the XR space.
[0048] "Conditional parameters" refer to the conditions of a parameter, such as those used when performing scene detection. As shown in the same figure, such information includes, for example, distance, angle, velocity, acceleration, rotational speed, interior of space, presence and / or number of objects, and / or time from start to end.
[0049] "Threshold" refers to the threshold corresponding to the conditional parameter. Furthermore, "conditional expression" refers to the conditional expression used to detect the scene; for example, the relationship between the conditional parameter and the threshold is defined as a conditional expression.
[0050] Furthermore, in the information processing device 10, a scene can be detected, for example, by combining condition categories or condition parameters. Figure 4 As shown. For example, as Figure 5 As shown, detection scenarios can be set by combining condition categories from multiple scenarios. Furthermore, as... Figure 6 As shown, the detection scenario can be set by combining the condition parameters of multiple scenarios.
[0051] For example, condition categories and / or condition parameters are thus combined, which simplifies the setup of new detection scenarios.
[0052] Return to Figure 3 The following description will explain the priority information DB 133. For example, in the information processing apparatus 10 according to an embodiment, the priority of each scene is set based on rules. The priority information DB 133 stores various information about the priority of the realism parameters. Figure 7 This is a diagram showing an example of priority information DB 133.
[0053] For example, such as Figure 7 As shown, priority information DB 133 stores the item "rule number" and "priority rule" information in a corresponding manner. "Rule number" represents the number used to identify the priority rule, while "priority rule" represents the rule with priority.
[0054] The terms "prioritize scenes detected earlier" and "prioritize scenes detected later (switch when subsequent scenes are provided)" shown in the same diagram represent prioritizing the realism parameters of scenes provided earlier or later in time, respectively. This, for example, simplifies the rules for setting scene priorities.
[0055] In addition, "prioritize scenes with larger weights for specific parameters" means that among the realism parameters, scenes with larger sound enhancement parameters or vibration parameters are given priority in terms of realism.
[0056] That is, in this case, setting realistic parameters for scenes with large sound enhancement or vibration parameters can provide realistic parameters associated with the sound or vibration data that should be enhanced.
[0057] Furthermore, "prioritizing scenes with larger weights for each parameter" means that, among the realism parameters, each realism parameter of a scene with larger sound enhancement and vibration parameters is given priority. Under this rule, parameters from mutually different scenes can be used for sound enhancement and vibration parameters.
[0058] That is, in this case, each of the vibration and sound data can be enhanced by using a realism parameter with a larger value, thereby improving the realism of each of these vibration and sound data. Furthermore, in this paper, larger or smaller weights represent, for example, larger or smaller values of a parameter.
[0059] Furthermore, "prioritize parameters for shorter scenes" means prioritizing the realism parameters of scenes with shorter durations. In cases where a scene with a shorter duration interrupts a scene with a longer duration during its reproduction, the realism parameters for the shorter duration scene are prioritized during that scene.
[0060] Therefore, for example, scenarios with shorter time periods can be appropriately enhanced. Additionally, rules can be set to prioritize parameters for longer scenarios.
[0061] Return to Figure 3 The following explanation will cover parameter information DB 134. Parameter information DB 134 is a database that stores information on the realism parameters of each scene. Figure 8 This is a diagram showing an example of parameter information DB 134.
[0062] like Figure 8 As shown, for example, parameter information DB 134 stores information about the project "scene name", "sound enhancement parameters" and "vibration parameters" in a corresponding manner.
[0063] "Scene Name" refers to the name of the detection scene as described above, and corresponds to, for example, Figure 4 The "detection scenarios" are shown below. Furthermore, for clarity of explanation, the "scenario names" are shown herein as explosion scenarios and / or concert hall scenarios.
[0064] "Sound enhancement parameters" refers to the sound enhancement parameters set in the corresponding scene. For example, such as Figure 8 As shown, the sound enhancement parameters store individual parameters for each speaker 4 according to the number of speakers 4, such as "for speaker 1", "for speaker 2", etc.
[0065] Furthermore, for example, for each speaker 4, parameter values from sound processing items such as "delay" or "bandwidth enhancement / attenuation" are stored. For example, "delay" refers to the parameter of the delay time, while "bandwidth enhancement / attenuation" refers to parameters such as the frequency bands by which the sound is enhanced or attenuated and the degree of enhancement.
[0066] "Vibration parameters" refers to the sound enhancement parameters set in the corresponding scenario, and the individual parameters of each vibration device 5 are stored according to the number of vibration devices 5, similar to "sound enhancement parameters". For example, the parameters in each of the items "LPF (low-pass filter)", "delay" and "amplification" are stored as "vibration parameters".
[0067] "LPF" represents the parameters of the low-pass filter, while "delay" represents the parameter of the delay time. In addition, "amplification" represents the parameters used for vibration processing, such as the degree of amplification or attenuation performed.
[0068] Return to Figure 3 The following description will explain the control unit 120. The control unit 120 is a controller, implemented by, for example, a central processing unit (CPU), a microprocessor unit (MPU), etc., in which various programs (not shown) stored in storage unit 130 are executed in RAM, which serves as the workspace. Furthermore, the control unit 120 may also be implemented using integrated circuits such as application-specific integrated circuits (ASICs) and / or field-programmable gate arrays (FPGAs).
[0069] The control unit 120 has a content generation unit 121, a presentation processing unit 122, a scene detection unit 123, a priority setting unit 124, a parameter extraction unit 125, and an output unit 126, and implements or performs information processing functions and / or actions as described below.
[0070] Content generation unit 121 generates a 3D model of the space within the XR content. For example, content generation unit 121 references XR content DB 131 and generates a 3D model of the space within the XR content based on the user's current field of view. Content generation unit 121 then transmits the generated 3D model to presentation processing unit 122.
[0071] The rendering processing unit 122 performs rendering processing, converting the 3D model received from the content generation unit 121 into video data and / or audio data. For example, the rendering processing unit 122 outputs the converted video data to a display device (see [link]). Figure 1 The content generation unit 121 and the presentation processing unit 122 transmit the converted audio data to the scene detection unit 123. Furthermore, the presentation processing unit 122 transmits the converted audio data to the output unit 126 and the scene detection unit 123. Additionally, the content generation unit 121 and the presentation processing unit 122 function as calculation units, which calculate the conditional data of the items in the conditional expression from the content.
[0072] The scene detection unit 123 detects scenes that meet predetermined conditions from the input content. For example, by using video data input from the rendering processing unit 122 and conditional expressions stored in the scene information DB132, the scene detection unit 123 detects scenes for which realism parameters should be set.
[0073] In this paper, for example, the scene detection unit 123 receives, for example, the coordinate information of the object in XR space and the object type information from the rendering processing unit 122, and detects the scene for which the realism parameters should be set by using conditional expressions.
[0074] Additionally, for example, when the XR content is MR content, the scene detection unit 123 can perform image analysis on images provided by capturing the interior of the MR space, in order to identify objects in such MR space or calculate the coordinates of such objects.
[0075] Figure 9 This is a block diagram of scene detection unit 123. For example, as shown... Figure 9 As shown, the scene detection unit 123 includes a scene determination unit 123a and a condition setting unit 123b. By using each of the condition data (condition expressions) for scene determination stored in the scene information DB 132, the scene determination unit 123a determines whether the conditions in the video data meet the detection conditions for each scene.
[0076] More specifically, for example, such as Figure 4As shown, based on the positional relationship between the user and the object (the object in the XR space) calculated from the content by the content generation unit 121 or the presentation processing unit 122, the user's movement, and / or the data of the items in the conditional expression (e.g., information about the space where the user appears), the scene determination unit 123a determines whether the current situation in the XR space corresponds to each initially defined detection scene.
[0077] In this paper, the scene determination unit 123a performs scene detection processing by using text information data (e.g., user movement in XR space, object coordinate information, object type information, spatial information, etc.) that has been calculated by the content generation unit 121 or the presentation processing unit 122.
[0078] Therefore, for example, even when the CPU performance is relatively low, processes such as scene detection and realism parameter extraction can be performed in parallel with processes that have a relatively heavy processing load (e.g., rendering processes performed by the rendering processing unit 122).
[0079] Furthermore, in this article, for example, based on, for example Figure 5 The combination of condition categories shown may also include, for example, Figure 6 The scene determination unit 123a can determine whether the current situation in the XR space corresponds to each detection scene, including the combination of the condition parameters shown.
[0080] Then, when the scene determination unit 123a determines that the current situation in the XR space corresponds to the detection scene, the detection scene information of this video data is transmitted to the priority setting unit 124 (see [link]). Figure 3 Additionally, if the scene determination unit 123a determines that the current state in the XR space does not correspond to any detection scene, it is not a corresponding detection scene, and the realism parameter is returned to its initial state (the realism parameter in the case of not corresponding detection scene). Furthermore, if the scene determination unit 123a determines that the current state in the XR space corresponds to multiple detection scenes, the determined multiple detection scenes are passed to the priority setting unit 124.
[0081] Furthermore, although this paper has explained that the scene determination unit 123a determines whether the current state of the XR space is a detection scene based on video data, the scene determination unit 123a can also determine whether it is a detection scene based on audio data.
[0082] The condition setting unit 123b sets various condition expressions for scene detection. The condition setting unit 123b sets the condition expressions based on information, for example, from the creator of the XR content and / or user input.
[0083] For example, the condition setting unit 123b receives input from the creator or user, such as realism parameters and the scene in which those parameters are set, and places the details of that scene into a conditional expression. Then, for each setting in the conditional expression, the condition setting unit 123b writes the information of the conditional expression into the scene information DB 132 and writes the corresponding realism parameter into the parameter information DB 134.
[0084] Therefore, in the information processing device 10, a scene requested by the creator or user can be detected, and a realism parameter requested by such a creator or user can be set for the detected scene.
[0085] Return to Figure 3 The following description will explain the priority setting unit 124. The priority setting unit 124 sets a priority for the scene detected by the scene detection unit 123.
[0086] For example, if the scene detection unit 123 determines that multiple types of scenes are detected simultaneously, the priority setting unit 124 refers to the priority information DB 133 and selects the scene to be given priority. Furthermore, if the scene detection unit 123 determines that only one scene is detected, that scene has the highest priority.
[0087] Figure 10 This is a block diagram of priority setting unit 124. For example, as... Figure 10 As shown, the priority setting unit 124 has a timing detection unit 124a and a rule setting unit 124b.
[0088] The timing detection unit 124a detects the timing at which a scene is generated and the timing at which it ends, as detected by the scene detection unit 123. For example, based on scene information from the scene detection unit 123 at each time point, the timing detection unit 124a detects each scene that appears at each time point (and also detects overlapping states), the timing at which the scene is generated, the timing at which the scene is deleted, etc. That is, the timing detection unit 124a detects the state of all scenes that appear at each time point, including the order in which they are generated.
[0089] For each scene detected by the scene detection unit 123, the rule setting unit 124b sets a priority for scenes used to determine realism parameters. That is, based on the state of all scenes that appear and are detected by the timing detection unit 124a, scenes with parameters linked to and preferentially used by the scene are identified as having realism parameters used at their respective points in time, thereby setting a priority for the detected scenes. Therefore, in the information processing device 10, realism parameters dependent on such priorities can be set.
[0090] That is, in the information processing device 10, priority conditions are initially set for each scene so that when scene A and scene B overlap in time, the scene with the realism parameters that should be used first can be appropriately determined.
[0091] For example, the rule setting unit 124b references the priority information DB 133 and sets the priority for each setting scenario among the sound enhancement parameters and vibration parameters, which is the scenario in which the parameters to be used are determined to be set. In this document, the rule setting unit 124b can set scenarios for parameter selection based on, for example, independent priority rules for each speaker 4 and / or each vibration device 5.
[0092] Therefore, in each speaker 4 and each vibration device 5, the realism parameters are set according to their respective rules, so that the realism can be further improved compared with the case of uniformly setting the realism parameters.
[0093] In addition, the rule setting unit 124b transmits the information of the set rules to the parameter extraction unit 125 (see [reference]). Figure 3 ( ), so as to correspond to video data and audio data.
[0094] return Figure 3 The following description will explain the parameter extraction unit 125. The parameter extraction unit 125 extracts the realism parameters of the scene detected by the scene detection unit 123.
[0095] Figure 11 This is a block diagram of parameter extraction unit 125. For example... Figure 11 As shown, the parameter extraction unit 125 includes a vibration parameter extraction unit 125a, a sound enhancement parameter extraction unit 125b, and a learning unit 125c.
[0096] The vibration parameter extraction unit 125a references the parameter information DB 134 and extracts the vibration parameters corresponding to the scene with the highest priority provided by the priority setting unit 124. For example, the vibration parameter extraction unit 125a extracts the vibration parameters corresponding to the "detection scene" with the highest priority received from the priority setting unit 124 from the parameter information DB 134, so as to extract the vibration parameters corresponding to the scene.
[0097] That is, when the scene detection unit 123 detects multiple overlapping scenes in which the objects generating sound are different from each other, the parameter extraction unit 125 can select the scene with high priority (i.e., the scene with a priority that is estimated to provide a more realistic feeling for the user through vibration) and extract the parameters for vibration generation corresponding to this scene. As a result, even during the time period used to reproduce the content of multiple overlapping scenes, vibrations with rich realism can be generated with appropriate parameters.
[0098] Specifically, through, as Figure 7 The priority information DB shows the content of the priority rules and the priority conditions for each scenario (which are set and stored as follows). Figure 4 (As shown in the scene information DB), the scene detection unit 123 can perform this scene selection processing.
[0099] For example, if the scene detection unit 123 detects a scene where an elephant makes a walking sound (elephant walking scene) and a scene where a horse makes a walking sound (horse walking scene), the parameter extraction unit 125 prioritizes the elephant walking scene according to the rule of "prioritizing larger amplitudes in lower frequency bands". Therefore, vibrations that reproduce the vibrations caused by the elephant's walking and are also primarily felt in the real world are applied to the user in the content reproduction (e.g., virtual space), so that the user can obtain a vibrational sensation with rich realism (i.e., a near-realistic sense of reality).
[0100] Furthermore, when the scene detection unit 123 detects multiple scenes in which the objects generating the sound are different from each other and overlap in time, the parameter extraction unit 125 can also apply a method that extracts parameters corresponding to the scene selected from the multiple scenes based on the type and position of the object corresponding to each of the multiple scenes in the image included in the content.
[0101] Specifically, for such Figure 7 The priority information shown in the DB contains the settings of priority rules and the priority conditions for each scenario (which are set and stored in, for example...). Figure 4 In the scene information DB shown, settings are made (in this example, the type of the object (m) and the function value F(M,d) of the distance (d) to such an object are added to this priority condition, and the condition provided by the function value F(M,d) (e.g., giving preference to larger function values "F(M,d)") is added to the priority rule) so that the scene detection unit 123 can perform this scene selection processing.
[0102] Will be used as Figure 14 The specific examples shown illustrate a method for prioritizing scenarios based on the location of objects. Figure 14 This is a diagram illustrating an example of a method for determining priority objects.
[0103] like Figure 14As shown, the display device 3 displays an image 31 of the content during content reproduction. Objects 311 (horse) and 312 (elephant) are seen in image 31. Here, the scene detection unit 123 detects both the horse walking scene and the elephant walking scene as object scenes for vibration control, both meeting certain conditions.
[0104] Furthermore, a distance L1 from a reference position (the user's position in the content image, for example, the position of the avatar corresponding to the user in the XR content) to object 311 is provided. On the other hand, a distance L2 from this reference position to object 312 is provided. In addition, reference vibration intensities V1 and V2 (the intensities of the low-frequency components of the sound signal of the object in the content) are provided for objects 311 and 312, respectively. Furthermore, an example is provided where the priority condition of "preferring the larger value of the function F(Ln, Vn) = Vn / (Ln·Ln)" is set.
[0105] Additionally, the distance from the reference position to the object is calculated based on information added to the content, etc. (e.g., calculated based on the position information of each object used for video reproduction in the XR content). Furthermore, the reference vibration intensity of the object can be determined by: depending on the type of object, by reading the reference vibration intensity from a data table that stores preliminary settings for reference vibration intensity for each object type; by adding the reference vibration intensity as content information to the content; and so on. Furthermore, sound data used for sound reproduction is often added to the content, allowing the reference vibration intensity to be calculated based on the lower frequency characteristics of this sound data (sound intensity level, lower frequency signal level, etc.) (vibration patterns are highly correlated with the lower frequency components of sound, and vibrations are often generated based on these lower frequency components of sound).
[0106] Therefore, the information processing device 10 can estimate the lower frequency band characteristics of the sound generated by the vibration generating object in the content. In this case, the information processing device 10 selects the vibration generating object based on the estimated lower frequency band characteristics. Therefore, a more suitable vibration generating object can be selected.
[0107] For example, the lower frequency band characteristic of sound is the lower frequency band signal level. In this case, the information processing device 10 selects vibration generating objects whose estimated lower frequency band signal level is greater than a threshold. The information processing device 10 can extract the lower frequency band signal level from the sound data. Therefore, vibration generating objects can be easily selected by using the lower frequency band signal level included in the sound data.
[0108] Furthermore, the threshold for lower frequency band signal levels should be set depending on the content type. As mentioned earlier, vibration is often preferred in music videos, even for the same object, compared to animal documentaries. Therefore, a vibration object suitable for the content type (music video, animal documentary, etc.) can be selected.
[0109] In this scenario, if the relationship between the function values of object 311 (horse) and object 312 (elephant) is function F(L1, V1) > function F(L2, V2), then the scene where object 311 produces sound (vibration) (i.e., the horse walking scene) is preferentially selected, and the parameter extraction unit 125 extracts the vibration parameters corresponding to this horse walking scene. The vibration corresponding to the horse walking scene is then applied to the user. Subsequently, for example, if object 312 (elephant) approaches the reference position and the relationship becomes function F(L1, V1) < function F(L2, V2), then the scene where object 312 produces sound (vibration) (i.e., the elephant walking scene) is preferentially selected, and the parameter extraction unit 125 extracts the vibration parameters corresponding to this elephant walking scene. The vibration corresponding to the elephant walking scene is then applied to the user.
[0110] Furthermore, if the function F(Ln, Vn) is less than a pre-determined or predetermined threshold, that is, if the vibration caused by an object at the user's location in the content (such as the virtual space of a game) is weak (providing less sensation to the user, i.e., requiring less vibration to be applied), then it is also effective not to select that object as the object generating vibration. In other words, it is also effective to select only objects that have a vibration caused by an object at the user's location in the content (such as the virtual space of a game) with a certain intensity (so strong that if the vibration is reproduced, the sense of realism is enhanced). That is, objects that significantly affect the vibration signal generated from the candidate objects provided as vibration generating objects (vibration objects that the user strongly feels the vibration of).
[0111] Therefore, the information processing device 10 can estimate the object candidates that significantly affect the vibration signals generated from such object candidates provided as vibration generating objects, and select them as such vibration generating objects. As a result, vibrations that match the user's sensations in real space are applied to the user, making it possible to reproduce content with a rich sense of realism.
[0112] In this case, when the object selected as the source of vibration is chosen, the threshold is preferably adjusted based on the content type. That is, since the determined content (determined level) of the object that generates vibration is preferably adjusted, the reproduction of vibration caused by the object appearing in the content can preferably be paused or emphasized depending on the entity of such content.
[0113] In other words, the principle of vibration generation is as follows: Based on the entity of the content, the object that generates vibration in the content (in each case) is determined. Then, based on the audio signal corresponding to the determined object (sound data of the object included in the content or sound data of the object generated from sound data in such a scene (obtained by, for example, filtering the lower frequency domain)), a vibration signal (vibration data) is generated (vibration signal is generated by acquiring and appropriately amplifying the lower frequency component of the object's sound signal, etc.).
[0114] Furthermore, in the method for determining the object that generates vibration, the lower frequency band characteristics (e.g., volume level) of the sound-generating object in the content are estimated (in the example described above, this is estimated based on a reference vibration intensity, which is based on the type of object and the distance between a reference location (the location where the user exists in the virtual space of the content, etc.) and such an object), and the object is determined (the sound-generating object with a large lower frequency band volume level of the sound is determined as the object that generates vibration).
[0115] Therefore, by determining the priority of a scene based on the location of an object, vibrations that also adapt to the user's visual intuition (i.e., vibrations that match the user's perception in real space) are applied to the user, and content with a rich sense of realism can be reproduced.
[0116] In this paper, the vibration parameter extraction unit 125a extracts the vibration parameters corresponding to each vibration device 5. Therefore, compared with the case of uniformly extracting vibration parameters, a further improvement in realism can be achieved.
[0117] The sound enhancement parameter extraction unit 125b references parameter information DB 134 and extracts sound enhancement parameters corresponding to the scene with the highest priority provided by the priority setting unit 124. The sound enhancement parameter extraction unit 125b extracts sound enhancement parameters for each speaker 4 separately and determines the sound enhancement parameters extracted based on the priority set by the priority setting unit 124 (or based on the scene with the highest priority), similar to the vibration parameter extraction unit 125a.
[0118] The learning unit 125c learns the relationship between scenes and realism parameters stored in the parameter information DB 134. For example, the learning unit 125c performs machine learning on each scene and each corresponding realism parameter stored in the parameter information DB 134, while providing user responses to realism controls performed by such parameters as learning data, so as to learn the relationship between scenes and realism parameters.
[0119] In this paper, for example, the learning unit 125c can use user evaluations of the realism parameters (user adjustments after realism control and / or user input such as questionnaires) as learning data. That is, the learning unit 125c can learn the relationship between the scene and the realism parameters set for it from the perspective of the scene and the realism parameters set for it, in order to obtain higher user evaluations (i.e., in order to obtain higher realism).
[0120] Furthermore, the learning unit 125c can determine the realism parameters that should be set based on the learning results when a new scene is input to it. For example, the realism parameters for a fireworks scene can be determined using the learning results from realism control for similar situations such as explosion scenes. Additionally, priority rules can be learned based on the presence or absence and / or degree of prioritization of elements in user input such as questionnaires, following realism control (e.g., when the user's adjustment is close to the parameters corresponding to another scene occurring simultaneously, or when a questionnaire provides an answer suggesting that another scene should be prioritized).
[0121] Therefore, in the information processing device 10, for example, optimizations for priority rules and / or realism parameters can be performed automatically.
[0122] return Figure 3 The following description will explain the output unit 126. The output unit 126 outputs the realistic parameters extracted by the parameter extraction unit 125 to the speaker 4 and the vibration device 5.
[0123] Figure 12 This is a block diagram of output unit 126. (Example) Figure 12 As shown, the output unit 126 has a sound enhancement processing unit 126a and a sound-vibration conversion processing unit 126b.
[0124] The sound enhancement processing unit 126a performs enhancement processing on the sound data received from the presentation processing unit 122 using sound enhancement parameters extracted by the parameter extraction unit 125. For example, the sound enhancement processing unit 126a performs delay or band enhancement / attenuation processing based on the sound enhancement parameters in order to perform enhancement processing on the sound data.
[0125] In this paper, the sound enhancement processing unit 126a performs sound enhancement processing on each speaker 4 and outputs the sound data with applied sound enhancement processing to each corresponding speaker 4.
[0126] The sound-vibration conversion processing unit 126b performs vibration-appropriate frequency band limiting processing, etc., on the sound data received from the presentation processing unit 122, so as to convert it into vibration data. Furthermore, depending on the vibration parameters extracted by the parameter extraction unit 125, the sound-vibration conversion processing unit 126b performs enhancement processing on the converted vibration parameters.
[0127] For example, the sound-vibration conversion processing unit 126b performs enhancement processing on the vibration data based on the vibration parameters, such as additional frequency characteristic processing like low-frequency band enhancement, delay, and amplification, so as to perform such enhancement processing on such vibration data.
[0128] In this paper, the sound-vibration conversion processing unit 126b performs vibration enhancement processing on each vibration device 5 and outputs the vibration data with such vibration enhancement processing to each corresponding vibration device 5.
[0129] Next, we will use Figure 13 To illustrate the processing procedure performed by the information processing device 10 according to the embodiment. Figure 13 This is a flowchart illustrating the processing procedure performed by the information processing device 10. Furthermore, the processing procedure shown below is repeatedly performed by the control unit 120.
[0130] When the power is turned on, the information processing system 1 executes the following: Figure 13 The flowchart shown illustrates the process (step S101). Then, XR content setting process is first executed (step S102). In addition, the XR content setting process of this document includes, for example, each initial setting of the device for XR content reproduction, and various processes such as selection of XR content performed by the user.
[0131] Subsequently, the information processing device 10 begins the reproduction of the XR content (step S103) and performs scene detection processing on the XR content during the reproduction (step S104). Then, the information processing device 10 performs priority setting processing on the result of the scene detection processing (step S105) and performs realism parameter extraction processing (step S106).
[0132] Then, the information processing device 10 performs output processing on various vibration data or sound data that reflect the processing results of the realism parameter extraction processing (step S107). Then, the information processing device 10 determines whether the XR content has ended (step S108), and ends the processing if it is determined that such XR content has ended (step S108; Yes).
[0133] Furthermore, if the information processing device 10 determines in step S108 that the XR content has not ended (step S108; no), it will transfer to step S104 for processing again.
[0134] As described above, the information processing apparatus 10 according to the embodiment includes a scene detection unit 123, a parameter extraction unit 125, and an output unit 126. The scene detection unit 123 detects a scene from the input content. The parameter extraction unit 125 extracts realistic parameters for wave control corresponding to the scene detected by the scene detection unit 123.
[0135] The output unit 126 outputs a waveform signal of the content, which has been enhanced by realism parameters corresponding to the scene and extracted by the parameter extraction unit 125. Therefore, in the information processing apparatus 10 according to the embodiment, the efficiency of setting realism parameters to improve the realism of the content can be improved.
[0136] As described above, the information processing apparatus 10 according to the embodiment includes a scene information DB 132 (an example of a storage unit), a content generation unit 121, a presentation processing unit 122 (an example of a calculation unit), and a scene detection unit 123. The scene information DB 132 stores conditional expressions for detecting scenes from input content. The content generation unit 121 and the presentation processing unit 122 calculate conditional data for items of the conditional expressions from the content.
[0137] The scene detection unit 123 detects the scene of the content by using conditional expressions stored in the scene information DB132 and conditional data calculated by the content generation unit 121 and the presentation processing unit 122. Therefore, in the information processing apparatus 10 according to the embodiment, the efficiency of setting realism parameters to improve the realism of the content can be improved.
[0138] As described above, the information processing apparatus 10 according to the embodiment includes a scene detection unit 123, a priority setting unit 124, and a parameter extraction unit 125. The scene detection unit 123 detects scenes from the content. The priority setting unit 124 sets a priority for the scenes detected by the scene detection unit 123.
[0139] The parameter extraction unit 125 extracts realism parameters corresponding to the scene determined by the priority set by the priority setting unit 124, and uses these as realism parameters for realism control. Therefore, in the information processing apparatus 10 according to the embodiment, the efficiency of setting realism parameters to improve the realism of content can be improved.
[0140] Furthermore, although the above embodiments have described the case where the content is XR content, this is not limiting. That is, the content can be 2D video and sound, or it can be video only, or sound only.
[0141] One aspect of the embodiments aims to provide an information processing apparatus, an information processing system, and an information processing method, wherein efficiency improvements can be achieved in setting realism parameters to enhance the realism of content.
[0142] An information processing apparatus according to one aspect of an embodiment includes a scene detection unit, a parameter extraction unit, and an output unit. The scene detection unit detects a scene from input content. The parameter extraction unit extracts realistic parameters for wave control corresponding to the scene detected by the scene detection unit. The output unit outputs a wave signal of the content, which has been enhanced by the realistic parameters extracted by the parameter extraction unit.
[0143] An information processing apparatus according to one aspect of an embodiment includes a storage unit, a computing unit, and a scene detection unit. The storage unit stores conditional expressions for detecting scenes from input content. The computing unit calculates conditional data for items corresponding to the conditional expressions from the content. The scene detection unit detects scenes in the content by using the conditional expressions stored in the storage unit and the conditional data calculated by the computing unit.
[0144] An information processing apparatus according to one aspect of an embodiment includes a scene detection unit, a priority setting unit, and a parameter extraction unit. The scene detection unit detects scenes from content. The priority setting unit sets a priority for the scenes detected by the scene detection unit. The parameter extraction unit extracts realism parameters corresponding to the scenes determined by the priority set by the priority setting unit, as realism parameters for realism control.
[0145] According to one aspect of the embodiments, the efficiency of setting realism parameters for improving the realism of content can be improved.
[0146] Example (1-1):
[0147] An information processing apparatus, comprising:
[0148] The scene detection unit detects scenes from the input content;
[0149] The parameter extraction unit extracts realistic parameters for wave control corresponding to the scene detected by the scene detection unit; and
[0150] The output unit outputs a wave signal of the content, which is enhanced by realism parameters extracted by the parameter extraction unit.
[0151] Examples (1-2):
[0152] According to the information processing apparatus of embodiment (1-1), wherein:
[0153] The parameter extraction unit includes a vibration parameter extraction unit, which extracts vibration parameters as realism parameters based on the content. These vibration parameters control the vibration device that applies vibration to the user.
[0154] The output unit outputs a vibration signal that has been enhanced using vibration parameters to the vibration device.
[0155] Examples (1-3):
[0156] According to the information processing apparatus of embodiments (1-2), wherein:
[0157] Vibration equipment is a device that provides vibration to a seat;
[0158] The parameter extraction unit extracts vibration parameters for each of the multiple vibration devices installed on the seat; and
[0159] The output unit outputs a vibration signal that has been enhanced by using the vibration parameters corresponding to each of the vibration devices.
[0160] Examples (1-4):
[0161] The information processing apparatus according to Embodiment (1-1), Embodiment (1-2), or Embodiment (1-3) includes:
[0162] The parameter extraction unit includes a sound parameter extraction unit, which extracts sound parameters to enhance the sound data of the content; and
[0163] The output unit outputs a vibration signal that has been enhanced using sound parameters to the sound output device.
[0164] Examples (1-5):
[0165] The information processing apparatus (10) according to any one of embodiments (1-1) to (1-4) includes:
[0166] Learning units that study the relationship between learning scenarios and realism parameters.
[0167] Examples (1-6):
[0168] An information processing system, comprising:
[0169] Information processing device to reproduce XR content;
[0170] The display device displays video based on the video signal output from the information processing device;
[0171] Sound output devices generate sound based on sound signals output from information processing devices; and
[0172] Vibrating equipment vibrates based on vibration signals output from an information processing device, wherein...
[0173] Information processing devices include:
[0174] The scene detection unit detects scenes from the input XR content;
[0175] The parameter extraction unit extracts realistic parameters for sound and vibration processing, which correspond to the scene detected by the scene detection unit; and
[0176] The output unit outputs sound data and vibration data, which have been enhanced using realistic parameters extracted by the parameter extraction unit, to the sound output device and vibration device, respectively.
[0177] Examples (1-7):
[0178] An information processing method, wherein, based on a content-driven scenario, the wave signal of a wave device is enhanced.
[0179] Example (2-1):
[0180] An information processing apparatus, comprising:
[0181] Storage unit, storing conditional expressions used to detect scenarios from content;
[0182] The calculation unit calculates conditional data for items against a conditional expression from the content; and
[0183] The scene detection unit detects the scene of the content by using conditional expressions stored in the storage unit and conditional data calculated by the calculation unit.
[0184] Example (2-2):
[0185] According to the information processing apparatus of embodiment (2-1), wherein:
[0186] The scene detection unit includes a condition setting unit for setting conditional expressions; and
[0187] The storage unit stores the conditional expressions set by the condition setting unit.
[0188] Examples (2-3):
[0189] The information processing apparatus according to embodiment (2-1) or embodiment (2-2), wherein
[0190] The items in a conditional expression are the positional relationships between the user and objects in the content.
[0191] Examples (2-4):
[0192] The information processing apparatus according to Embodiment (2-1), Embodiment (2-2) or Embodiment (2-3), wherein
[0193] The item in the conditional expression is the user's movement within the content.
[0194] Examples (2-5):
[0195] The information processing apparatus according to any one of embodiments (2-1) to (2-4), wherein
[0196] The items in a conditional expression are the spaces in the content where the user appears.
[0197] Examples (2-6):
[0198] The information processing apparatus according to any one of embodiments (2-1) to (2-5), wherein
[0199] The items in the conditional expression are information about the time when the user appears in the content.
[0200] Examples (2-7):
[0201] An information processing system, comprising:
[0202] Information processing device to reproduce XR content;
[0203] The display device displays video based on the video data output from the information processing device;
[0204] Sound output devices generate sound based on sound data output from information processing devices; and
[0205] Vibration equipment vibrates based on vibration data output from an information processing device, wherein...
[0206] Information processing devices include:
[0207] Storage unit, storing conditional expressions used to detect scenarios from content;
[0208] The calculation unit calculates conditional data for items against a conditional expression from the content; and
[0209] The scene detection unit detects the scene of the content by using conditional expressions stored in the storage unit and conditional data calculated by the calculation unit.
[0210] Examples (2-8):
[0211] An information processing method, wherein:
[0212] Calculate conditional data for items used to perform conditional expressions for detecting scenes from content; and
[0213] A scenario where content is detected using conditional expressions and calculated conditional data.
[0214] Example (3-1):
[0215] An information processing apparatus, comprising:
[0216] The scene detection unit detects scenes from the content.
[0217] The priority setting unit sets the priority for scenes detected by the scene detection unit; and
[0218] The parameter extraction unit extracts realism parameters corresponding to the scene, which depend on the priority set by the priority setting unit, and uses them as realism parameters for realism control.
[0219] Example (3-2):
[0220] According to the information processing apparatus of embodiment (3-1), wherein:
[0221] The parameter extraction unit includes a sound parameter extraction unit and a vibration parameter extraction unit. The sound parameter extraction unit extracts sound enhancement parameters for sound processing, and the vibration parameter extraction unit extracts vibration parameters for vibration processing.
[0222] The priority setting unit sets the priority of the scene for each of the sound enhancement parameters and vibration parameters.
[0223] Example (3-3):
[0224] The information processing apparatus according to embodiment (3-1) or embodiment (3-2), wherein
[0225] The priority setting unit sets the priority of scene-based detection timing.
[0226] Examples (3-4):
[0227] The information processing apparatus according to Embodiment (3-1), Embodiment (3-2) or Embodiment (3-3), wherein
[0228] The priority setting unit sets the priority of weights based on the realism parameters.
[0229] Examples (3-5):
[0230] The information processing apparatus according to any one of Embodiments (3-1) to (3-4), wherein
[0231] The priority setting unit sets the priority based on the time length of the scenario.
[0232] Examples (3-6):
[0233] An information processing system, comprising:
[0234] Information processing device to reproduce XR content;
[0235] The display device displays video based on the video data output from the information processing device;
[0236] Sound output devices generate sound based on sound data output from information processing devices; and
[0237] Vibration equipment vibrates based on vibration data output from an information processing device, wherein...
[0238] Information processing devices include:
[0239] The scene detection unit detects scenes from XR content;
[0240] The priority setting unit sets the priority for scenes detected by the scene detection unit;
[0241] The parameter extraction unit extracts realism parameters corresponding to the scene, which depend on the priority set by the priority setting unit, as realism parameters for realism control; and
[0242] The output unit outputs sound data and vibration data, which have been enhanced using realistic parameters extracted by the parameter extraction unit, to the sound output device and vibration device.
[0243] Examples (3-7):
[0244] An information processing method, wherein:
[0245] Prioritize scenarios detected from the content; and
[0246] Extract the realism parameters corresponding to the scene, which are determined based on the set priority, and use them as realism parameters for realism control.
Claims
1. An information processing apparatus, comprising: The control unit is configured to perform: Scene detection processing detects scenes from the input content; Parameter extraction processing: extracting realistic parameters for wave control corresponding to the scene detected by the scene detection processing; as well as Output processing outputs a vibration signal of the content, which is generated by processing the sound data of the input content using realism parameters extracted by the parameter extraction process. The control unit is configured to: Detect the scene included in the input content; Determine whether the scene detected during content reproduction is a single scene or multiple scenes; In the case of a single scene, the detected scene is identified as the scene for extracting vibration parameters; When multiple scenarios are identified, the scenario used to extract vibration parameters is determined from the detected scenarios based on a preset priority rule. From a vibration parameter database that stores scenes and vibration parameters in a corresponding manner, extract the vibration parameters corresponding to the determined scene; and The vibration signal is generated by processing the sound data of the content based on the extracted vibration parameters.
2. The information processing apparatus according to claim 1, wherein: The parameter extraction process includes vibration parameter extraction, which extracts vibration parameters based on the content as the realism parameters. These vibration parameters control at least one vibration device to apply vibration to the user. The output processing includes outputting a vibration signal to at least one vibration device, the vibration signal being generated by processing the sound data of the input content using the vibration parameters.
3. The information processing apparatus according to claim 2, wherein: The at least one vibration device is a plurality of vibration devices that provide vibration to the seat; The parameter extraction process includes extracting the vibration parameters for each of a plurality of vibration devices mounted on the seat; and The output processing includes outputting the vibration signal, which is generated by processing the sound data of the input content using vibration parameters corresponding to each of the plurality of vibration devices.
4. The information processing apparatus according to claim 1, wherein: The parameter extraction process includes audio parameter extraction process, which extracts audio parameters for processing the audio data of the content; as well as The output processing includes outputting a vibration signal to a sound output device, the vibration signal being generated by processing the sound data of the content using the sound parameters.
5. The information processing apparatus according to claim 1, wherein: The control unit is configured to perform a learning process that learns the relationship between the scene and the realism parameters.
6. The information processing apparatus according to any one of claims 1 to 5, comprising: A storage unit stores conditional expressions used to detect scenes from content, where: The control unit is configured to perform computational processing, which calculates conditional data for items against a conditional expression from the content; and The scene detection process includes detecting scenes of content by using conditional expressions stored in the storage unit and conditional data calculated by the computation process.
7. The information processing apparatus according to claim 6, wherein: The scene detection process includes setting the condition expression; as well as The storage unit stores the conditional expressions set by the condition setting process.
8. The information processing apparatus according to claim 6, wherein The items in the conditional expression are the positional relationships between the user and objects in the content.
9. The information processing apparatus according to claim 6, wherein The item in the conditional expression is the user's movement within the content.
10. The information processing apparatus according to claim 6, wherein The item in the conditional expression is the space in the content where the user appears.
11. The information processing apparatus according to any one of claims 1 to 5, wherein The control unit is also configured to perform a selection process that selects an object for wave control in a scene detected by the scene detection process. The parameter extraction process includes: Extract the realistic parameters for wave control corresponding to the object selected for wave control by the selection process; as well as The output processing includes: outputting a vibration signal of the content, which is generated by processing sound data of an object selected by the selection processing for wave control in the input content using realism parameters extracted by the parameter extraction processing.
12. The information processing apparatus according to any one of claims 1 to 5, wherein: The control unit is configured to perform a priority setting process, which sets a priority for the scene detected by the scene detection process.
13. The information processing apparatus according to claim 12, wherein: The parameter extraction process includes sound parameter extraction process and vibration parameter extraction process. The sound parameter extraction process extracts sound parameters for sound processing, and the vibration parameter extraction process extracts vibration parameters for vibration processing. as well as The priority setting process includes setting the priority of the scene for each of the sound parameters and the vibration parameters.
14. The information processing apparatus according to claim 12, wherein The priority setting process includes setting the priority of scene-based detection timing.
15. The information processing apparatus according to claim 12, wherein The priority setting process includes setting priorities based on weights of realism parameters.
16. The information processing apparatus according to claim 12, wherein The priority setting process includes setting priorities based on the time length of the scenario.
17. An information processing system, comprising: Information processing device to reproduce XR content; The display device displays video based on the video signal output from the information processing device; A sound output device generates sound based on a sound signal output from the information processing device; and The vibrating device vibrates based on the vibration signal output from the information processing device, wherein... The information processing device includes a control unit, which is configured to perform: Scene detection processing: Detecting scenes from the input XR content; The parameter extraction process extracts realistic parameters for sound processing and vibration processing, the realistic parameters corresponding to the scene detected by the scene detection process; as well as Output processing involves outputting sound data and vibration data to the sound output device and the vibration device, respectively. The sound data and vibration data are generated by processing the sound data of the input XR content using realism parameters extracted by the parameter extraction process. The control unit is configured to: Detect the scene included in the input content; Determine whether the scene detected during content reproduction is a single scene or multiple scenes; In the case of a single scene, the detected scene is identified as the scene for extracting vibration parameters; When multiple scenarios are identified, the scenario for extracting vibration parameters is determined from the detected scenarios based on a preset priority rule. From a vibration parameter database that stores scenes and vibration parameters in a corresponding manner, extract the vibration parameters corresponding to the determined scene; and The vibration data is generated by processing the sound data of the content based on the extracted vibration parameters.
Citation Information
Patent Citations
Bodily feeling experiencing video / sound system
JP2004081357A
Haptic feedback sensations based on audio output from computer devices
CN1599925A
Intelligent augmented reality (IAR) platform-based communication system
US20180047196A1