Information processing device, information processing method, program, and information processing system
The information processing device automates the setting of realistic sensation parameters in XR content by detecting scenes, setting priorities, and extracting audio and vibration parameters, enhancing user experience through efficient parameter setting.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-09
- Publication Date
- 2026-03-05
AI Technical Summary
Conventional methods for setting realistic sensation parameters in XR content require significant manual effort, leading to inefficiencies.
An information processing device that automates the setting of realistic sensation parameters by detecting scenes from XR content, setting priorities, and extracting audio and vibration parameters based on predefined conditions and rules.
Improves the efficiency of setting realistic sensation parameters, ensuring appropriate audio and vibration enhancements for enhanced user experience in XR environments.
Smart Images

Figure 0007824781000001 
Figure 0007824781000002 
Figure 0007824781000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, Information processing method, program, and Information Processing System Mu Regarding. [Background technology]
[0002] Conventionally, there is known technology that provides users with digital content that includes virtual space experiences such as VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality), known as XR (Cross Reality) content, using devices such as HMDs (Head Mounted Displays). XR is a collective term that encompasses all virtual space technologies, including VR, AR, and MR, as well as SR (Substitutional Reality) and AV (Audio / Visual).
[0003] Furthermore, a technique has been proposed in which vibrations corresponding to the video being viewed by the user are applied to the user to improve the sense of realism of the video (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2004-081357 Summary of the Invention [Problem to be solved by the invention]
[0005] However, in the conventional technology, the realistic sensation parameters relating to the realistic sensation must be set manually in advance, and setting the realistic sensation parameters requires a huge amount of manual work.
[0006] The present invention has been made in view of the above, and provides an information processing device that can improve the efficiency of setting realistic sensation parameters related to improving the realistic sensation of content; Information processing method, program, and Information Processing System M The purpose is to provide. [Means for solving the problem]
[0007] In order to solve the above-mentioned problems and achieve the object, an information processing device according to the present invention comprises: An information processing device that outputs a vibration signal corresponding to input content includes a controller that selects a vibration generating object from the content as a vibration source, extracts an acoustic signal corresponding to the selected vibration generating object from the content, and processes the extracted acoustic signal to generate a vibration signal. [Effects of the Invention]
[0008] According to the present invention, it is possible to improve the efficiency of setting realistic sensation parameters related to improving the realistic sensation of content. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram illustrating an overview of an information processing system. [Figure 2] FIG. 2 is a diagram showing the flow of data in the information processing system. [Figure 3] FIG. 3 is a diagram showing an outline of the information processing method. [Figure 4] FIG. 4 is a block diagram of an information processing device. [Figure 5] FIG. 5 is a diagram showing an example of the scene information DB. [Figure 6] FIG. 6 is a diagram showing an example of the scene information DB. [Figure 7] FIG. 7 is a diagram showing an example of the scene information DB. [Figure 8] FIG. 8 is a diagram illustrating an example of the priority information DB. [Figure 9] FIG. 9 is a diagram illustrating an example of the parameter information DB. [Figure 10] FIG. 10 is a block diagram of the scene detection unit. [Figure 11] FIG. 11 is a block diagram of the priority setting unit. [Figure 12]FIG. 12 is a diagram showing an example of a method for determining a priority object. [Figure 13] FIG. 13 is a block diagram of the parameter extraction unit. [Figure 14] FIG. 14 is a block diagram of the output section. [Figure 15] FIG. 15 is a flowchart showing a processing procedure executed by the information processing device. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of an information processing device, an information processing system, and an information processing method disclosed in the present application will be described in detail with reference to the accompanying drawings. Note that the present invention is not limited to the embodiments described below.
[0011] First, an overview of an information processing system and an information processing method according to an embodiment will be described with reference to Figures 1, 2, and 3. Figure 1 is a diagram showing an overview of the information processing system. Figure 2 is a diagram showing a data flow in the information processing system. Figure 3 is a diagram showing an overview of the information processing method. Note that the following description will be given assuming that the XR space (virtual space) is a VR space.
[0012] As shown in FIG. 1, the information processing system 1 includes a display device 3, a speaker 4, and a vibration device 5.
[0013] 2, the information processing device 10 provides video data to the display device 3. The information processing device 10 also provides audio data to the speaker 4. The information processing device 10 also provides vibration data to the vibration device 5.
[0014] 1, the display device 3 is, for example, a head-mounted display. The display device 3 is an information processing terminal that presents video data related to XR content provided by the information processing device 10 to the user, allowing the user to enjoy a VR experience.
[0015] The display device 3 may be a non-transparent type that completely covers the field of view, or a video-transparent type or an optical-transparent type. The display device 3 also has a device, such as a camera or a motion sensor, that detects changes in the user's internal and external circumstances using a sensor unit.
[0016] The speaker 4 is an audio output device that outputs audio, and is provided, for example, as a headphone type, and is worn on the user's ear. The speaker 4 generates audio data provided from the information processing device 10 as audio. Note that the speaker 4 is not limited to a headphone type, and may also be a box type (installed on the floor, etc.). The speaker 4 may also be a stereo audio or multi-channel audio type.
[0017] The vibration device 5 is composed of an electric vibration converter made up of an electric / magnetic circuit and a piezoelectric element, and is provided, for example, on a seat on which a user sits, and vibrates in accordance with vibration data provided from the information processing device 10. Note that, for example, a plurality of vibration devices 5 are provided on the seat, and the information processing device 10 controls each vibration device 5 individually.
[0018] By applying the sound from the speaker 4 and the vibration from the vibration device 5, that is, the waves from the wave device, to the content user in a manner that is suited to the reproduced video, it is possible to increase the sense of realism regarding the video reproduction.
[0019] The information processing device 10 is configured by a computer, is connected to the display device 3 via a wired or wireless connection, and provides images of XR content to the display device 3. The information processing device 10 also acquires changes in the situation detected by a sensor unit provided in the display device 3 at any time, and reflects such changes in the situation in the XR content.
[0020] For example, the information processing device 10 can change the direction of the field of view in the virtual space of the XR content in accordance with changes in the user's head or line of sight detected by the sensor unit.
[0021] Incidentally, when providing XR content, the sense of realism of the XR content can be improved by emphasizing the sound emitted from the speaker 4 in accordance with the scene, or by vibrating the vibration device 5 in accordance with the scene.
[0022] However, the parameters used to control the sense of realism to improve the sense of realism (hereinafter referred to as "sense of realism parameters") had to be set manually after the XR content was produced, which required a huge amount of work to set the sense of realism parameters.
[0023] Therefore, the information processing method aims to automate the setting of these realism parameters. For example, as shown in Fig. 3, the information processing method according to the embodiment first detects scenes that satisfy predetermined conditions from video data and audio data related to XR content (step S1).
[0024] The predetermined condition here is, for example, a condition regarding whether the corresponding video data or audio data is a scene that requires the setting of a realism parameter, and is defined, for example, by a conditional expression regarding the situation within the XR content.
[0025] In other words, in the information processing method, when the situation within the XR content satisfies the conditions defined by the conditional expressions, the scene is detected as satisfying the predetermined conditions. This eliminates the need for detailed analysis of the video data, thereby reducing the processing load of scene detection.
[0026] Next, in the information processing method, a priority order is set for the scenes detected by the scene detection (step S2). Here, the priority order indicates the order in which the realism parameter of a scene should be prioritized. That is, in the information processing method, when multiple scenes overlap in time, the realism parameter of a scene that should be prioritized is defined in advance for each scene.
[0027] As a result, even when multiple scenes overlap, it is possible to provide the user with an appropriate sense of realism. As will be described later, in the information processing method, the priority order for audio and the priority order for vibration are set separately.
[0028] Next, in the information processing method, a realism parameter is extracted for each scene (step S3). For example, in the information processing method, a realism parameter is extracted for each scene using parameter information in which the relationship between the scene and the realism parameter is defined in advance.
[0029] In this case, the information processing method extracts a corresponding realism parameter according to the priority. Specifically, for example, in the case where a scene with a low priority overlaps with a scene with a high priority, the information processing method extracts the realism parameter of the scene with the high priority.
[0030] In the information processing method, a voice enhancement process is performed to enhance the voice data using a voice enhancement parameter from among the extracted realism parameters (step S4), and the voice data is output to the speaker 4. In addition, in the information processing method, a vibration conversion process is performed to convert the voice data into vibration data, and the vibration data is enhanced using a vibration parameter from among the extracted realism parameters (step S5), and the vibration data is output to the vibration device 5.
[0031] As a result, the information processing method can provide the user with sound that is emphasized in accordance with the scene that the user is viewing, and vibration that is appropriate for the scene.
[0032] In this way, the information processing method according to the embodiment detects scenes from XR content, sets priorities, and then extracts realistic sensation parameters related to wave control, including audio processing and vibration processing, for the scenes. Therefore, the information processing method according to the embodiment can automate the setting of realistic sensation parameters related to improving the realistic sensation of content.
[0033] Next, an example of the configuration of the information processing device 10 according to the embodiment will be described with reference to Fig. 4. Fig. 4 is a block diagram of the information processing device 10. As shown in Fig. 4, the information processing device 10 includes a control unit 120 and a storage unit 130.
[0034] The storage unit 130 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. In the example of FIG. 4, the storage unit 130 has an XR content DB (Database) 131, a scene information DB 132, a priority order information DB 133, and a parameter information DB 134.
[0035] The XR content DB 131 is a database that stores a group of XR contents to be displayed on the display device 3. The scene information DB 132 is a database that stores various information related to the scene to be detected.
[0036] 5 to 7 are diagrams showing an example of the scene information DB 132. As shown in Fig. 5, for example, the scene information DB 132 stores information on items such as "detected scene," "condition category," "object," "condition parameter," "threshold," and "condition formula" in association with one another.
[0037] "Detected Scene" indicates the name of the scene to be detected. Note that "Detected Scene" acts as an identification symbol, and while a code such as a number is usually used, in this example, a name (duplication prohibited) is used to make the explanation easier to understand. "Condition Category" indicates the category based on what information is used to detect the scene. In the example shown in the figure, the categories are broadly divided into the positional relationship between the user and the object, the user's actions, spatial information about the user's location, time information about the user's location, and sound emitted from the object. Note that the user here refers to the operator in the XR space.
[0038] "Object" refers to an object for scene detection. In the example shown in the figure, information such as Object 1, Object 2, User, Space 1, Space 1 + Object 3, Content 1, Object 4, Object 5, and Object 6 corresponds to the object. Here, Object 1, Object 2, Object 3, Object 4, Object 5, and Object 6 each represent a different object in the XR space. Furthermore, Space 1 represents, for example, the space in the XR space where the user exists, and Content 1 represents, for example, a predetermined event in the XR space.
[0039] "Condition parameters" indicate conditions related to parameters, such as which parameters are used for scene detection. As shown in the figure, information such as distance, angle, speed, acceleration, rotation speed, space, object presence, quantity, start time, end time, and sound pattern is associated with each parameter.
[0040] The "threshold" indicates a threshold value corresponding to the condition parameter. The "condition expression" indicates a condition expression for detecting a detected scene, and for example, the relationship between the condition parameter and the threshold value is defined as a condition expression.
[0041] In Figure 5, for the purpose of explanation, each item value is represented using symbols such as "W," "4," and "w," such as "Scene W," "Object 4," and "Pattern w," but in reality, each item value will be stored as data in a manner that allows the specific meaning to be understood.
[0042] For example, "Scene W," "Scene X," "Scene Y," and "Scene Z" would actually represent data such as an "elephant walking scene," a "horse walking scene," a "car driving scene," and a "car making a sharp turn scene," respectively.
[0043] In this case, "Object 4," "Object 5," and "Object 6" would actually be data such as "Horse," "Elephant," and "Car," respectively.
[0044] Furthermore, "pattern w," "pattern x," "pattern y," and "pattern z" are actually data such as "pattern of horse walking sounds," "pattern of elephant walking sounds," "pattern of car driving sounds," and "pattern of tire squealing sounds," respectively.
[0045] The speech pattern is represented by, for example, a feature vector having speech features as elements. For example, the features may be obtained by performing spectral decomposition on the speech signal (e.g., Mel filter bank or cepstrum).
[0046] If the similarity (for example, cosine similarity or Euclidean distance) between the feature vectors corresponding to two voice patterns is equal to or greater than a threshold, the two voice patterns can be said to be similar.
[0047] For example, "the audio pattern is similar to pattern w" means that the similarity between the feature vector calculated from the audio occurring in the scene and the feature vector of the audio corresponding to pattern w is equal to or greater than a threshold value.
[0048] The threshold value for the similarity of the sound pattern may also be included in the “threshold value” of the scene information DB 132 .
[0049] Furthermore, the information processing device 10 may detect a scene by combining, for example, the condition categories or condition parameters shown in Fig. 5. For example, as shown in Fig. 6, a detected scene may be set by combining condition categories of multiple scenes, or as shown in Fig. 7, a detected scene may be set by combining condition parameters of multiple scenes.
[0050] For example, by combining condition categories and condition parameters in this way, it is possible to simplify the setting of a new detection scene.
[0051] Returning to the description of Fig. 4, the priority information DB 133 will be described. For example, in the information processing device 10 according to the embodiment, a priority is set for each scene on a rule basis. The priority information DB 133 stores various information related to the priority of the realism parameters. Fig. 8 is a diagram showing an example of the priority information DB 133.
[0052] 8, for example, the priority information DB 133 stores information items such as "rule number" and "priority rule" in association with each other. The "rule number" indicates a number for identifying a priority rule, and the "priority rule" indicates a rule related to priority.
[0053] The "Give priority to the first detected scene" and "Give priority to the last detected scene (switch to the next scene)" shown in the figure indicate that the realism parameters of the first or last scene are given priority, respectively. This makes it possible to simplify the rules for setting scene priorities, for example.
[0054] Furthermore, "give priority to a specific parameter with a larger weight" indicates that priority is given to the realism parameter of a scene in which either the voice emphasis parameter or the vibration parameter is larger among the realism parameters.
[0055] In other words, in this case, the extracted realism parameter is set for the scene in which the voice emphasis parameter or the vibration parameter is larger, so that it is possible to provide a realism parameter that is linked to the voice data or vibration data to be emphasized.
[0056] Furthermore, "prioritize the parameter with the larger weight" indicates that among the realism parameters, the realism parameter of the scene with the larger voice emphasis parameters or the larger vibration parameters is given priority. In this rule, different scene parameters may be used for the voice emphasis parameters and the vibration parameters.
[0057] In other words, in this case, the vibration data and the audio data can be emphasized by using a realism parameter with a large value, so that the realism of each of the vibration data and the audio data can be improved. Note that the magnitude of the weight here indicates, for example, the magnitude of the parameter value.
[0058] Furthermore, "prioritize parameters of shorter scenes" indicates that priority is given to the realism parameters of scenes with shorter durations. When a shorter scene interrupts playback of a longer scene, the realism parameters of the shorter scene are given priority during playback.
[0059] This allows, for example, scenes with a short duration to be appropriately emphasized. Note that a rule may be set such that priority is given to parameters for scenes with a long duration.
[0060] Additionally, "give priority to the one with the larger amplitude in the low frequency range" indicates that when multiple scenes in which objects are emitting sounds occur simultaneously, priority is given to the scene corresponding to the object emitting the sound with the larger amplitude in the low frequency range (for example, below 500 Hz).
[0061] Generally, it is considered that the larger the creature, the greater the amplitude of the low-frequency components of its walking sounds. For this reason, for example, if a walking scene of an elephant and a walking scene of a horse are detected, the walking scene of the elephant will be prioritized according to the rule that "the one with the larger low-frequency components has priority."
[0062] Furthermore, "give priority to scenes with large temporal fluctuations in sound and video" indicates that priority is given to scenes with large fluctuations per unit time in the volume of the sound produced by an object or the position of an object in the video.
[0063] Additionally, "Give priority to scenes of objects closer to the center of the field of view" indicates that scenes of objects located closer to the center of the screen are given priority in the content video. This rule will be explained later with reference to Figure 12.
[0064] Also, "Give priority to scene X over scene W" indicates that if scene W and scene X are detected, scene X will be given priority. In this way, a person (designer, developer) may manually define a priority rule in advance for two or more specific scenes.
[0065] Returning to the explanation of Fig. 4, the following will explain the parameter information DB 134. The parameter information DB 134 is a database that stores information related to realistic parameters for each scene. Fig. 9 is a diagram showing an example of the parameter information DB 134.
[0066] As shown in FIG. 9, the parameter information DB 134 stores information items such as "scene name," "voice emphasis parameter," and "vibration parameter" in association with one another.
[0067] The "scene name" indicates the name of the detected scene described above, and corresponds to, for example, the "detected scene" shown in Fig. 5. Note that, for ease of understanding, the "scene names" are shown here as an explosion scene, a concert hall scene, an elephant walking scene, a horse walking scene, a car driving scene, and a car making a sharp turn scene.
[0068] "Speech enhancement parameters" indicate speech enhancement parameters to be set for the corresponding scene. For example, as shown in Fig. 9, speech enhancement parameters are stored for each speaker 4 according to the number of speakers 4, such as "for speaker 1," "for speaker 2," etc.
[0069] Also, for each speaker 4, parameter values related to audio processing, such as "delay" and "band emphasis / attenuation," are stored. For example, "delay" indicates a parameter related to the delay time, and "band emphasis / attenuation" indicates a parameter related to the extent to which sound in which band is emphasized or attenuated.
[0070] The "vibration parameters" indicate the parameters of vibration to be set in the corresponding scene, and similarly to the "voice enhancement parameters," individual parameters are stored for each vibration device 5 according to the number of vibration devices 5. As the "vibration parameters," for example, parameters for items such as "LPF (Low Pass Filter)," "delay," and "amplification" are stored.
[0071] "LPF" indicates a parameter related to the low-pass filter (cutoff frequency of the low-pass filter), "delay" indicates a parameter related to the delay time, and "amplification" indicates a parameter related to vibration processing, such as the degree of amplification or attenuation.
[0072] Returning to the explanation of Fig. 4, the control unit 120 will be described. The control unit 120 is a controller, and is realized by, for example, a CPU (Central Processing Unit) or an MPU (Micro Processing Unit) executing various programs (not shown) stored in the storage unit 11 using a RAM as a work area. The control unit 120 can also be realized by, for example, an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0073] The control unit 120 has a content generation unit 121, a rendering processing unit 122, a scene detection unit 123, a priority setting unit 124, a parameter extraction unit 125, and an output unit 126, and realizes or executes the information processing functions and actions described below.
[0074] The content generation unit 121 generates a 3D model of the space within the XR content. For example, the content generation unit 121 references the XR content DB 131 and generates a 3D model of the space within the XR content in accordance with the user's current field of view within the XR content. The content generation unit 121 passes the generated 3D model to the rendering processing unit 122.
[0075] The rendering processing unit 122 performs rendering processing to convert the 3D model received from the content generation unit 121 into video data and audio data. The rendering processing unit 122 outputs the converted video data to the display device 3 (see FIG. 2) and passes it to the scene detection unit 123. The rendering processing unit 122 also passes the converted audio data to the output unit 126 and the scene detection unit 123. The content generation unit 121 and the rendering processing unit 122 function as calculation units that calculate condition data for items of a conditional expression from the content.
[0076] The scene detection unit 123 detects scenes that satisfy predetermined conditions from the input content. For example, the scene detection unit 123 detects scenes for which a sense of presence parameter should be set, using video data input from the rendering processing unit 122 and conditional expressions stored in the scene information DB 132.
[0077] At this time, for example, the scene detection unit 123 receives coordinate information of objects in the XR space and information about the object type from the rendering processing unit 122, and detects a scene for which a realism parameter should be set using a conditional expression.
[0078] In addition, for example, when the XR content is MR content, the scene detection unit 123 may recognize objects in the MR space or calculate the coordinates of the objects by performing image analysis on images captured in the MR space.
[0079] Fig. 10 is a block diagram of the scene detection unit 123. As shown in Fig. 10, for example, the scene detection unit 123 includes a scene determination unit 123a and a condition setting unit 123b. The scene determination unit 123a uses each condition data (conditional formula) for scene determination stored in the scene information DB 132 to determine whether or not the situation in the video data satisfies the detection condition for each scene.
[0080] More specifically, for example, as shown in FIG. 5, the scene determination unit 123a determines whether the current situation in the XR space corresponds to each predefined detected scene based on data (calculated from the content by the content generation unit 121 or the rendering processing unit 122) for items in a conditional expression such as the positional relationship between the user and the target object (an object in the XR space), the user's actions, and spatial information in which the user exists.
[0081] Here, the scene determination unit 123a performs scene detection processing using text information-like data that has already been calculated by the content generation unit 121 or the rendering processing unit 122, such as user movement in the XR space, object coordinate information, information about the object type, and spatial information.
[0082] This makes it possible to perform processes such as scene detection and extraction of realistic parameters in parallel with processes that have a relatively heavy processing load, such as rendering by the rendering processing unit 122, even if the CPU performance is relatively low.
[0083] In addition, at this time, for example, the scene determination unit 123a may determine whether or not the current situation in the XR space corresponds to each detected scene based on scene determination information that also includes a combination of condition categories as shown in FIG. 6, or a combination of condition parameters as shown in FIG. 7.
[0084] If the scene determination unit 123a determines that the video data corresponds to a detected scene, it passes the detected scene information for the video data to the priority order setting unit 124 (see FIG. 4). If the scene determination unit 123a determines that the video data does not correspond to any of the detected scenes, it considers the detected scene to be not the relevant detected scene and returns the realism parameters to their initial states (realism parameters for the case where the detected scene is not the relevant detected scene). If the scene determination unit 123a determines that the current situation in the XR space corresponds to multiple detected scenes, it passes the determined multiple detected scenes to the priority order setting unit 124.
[0085] Furthermore, although the case where the scene determination unit 123a determines whether or not a scene is detected based on video data has been described here, the scene determination unit 123a may also determine whether or not a scene is detected based on audio data.
[0086] For example, the scene determination unit 123a detects a scene in which sound is emitted from an object in the input content. In this case, the detected scene has the condition category of "sound is emitted from an object," so the candidate scenes are scene W, scene X, scene Y, and scene Z in Fig. 5 (a scene of an elephant walking, a scene of a horse walking, a scene of a car running, and a scene of a car making a sharp turn).
[0087] Furthermore, the scene determination unit 123a calculates the similarity between the feature vector obtained from the audio signal of the content and a feature vector of a predetermined audio in the candidate scene (e.g., pattern w, etc.), determines whether the similarity is equal to or greater than a threshold, and judges whether the audio pattern satisfies the audio pattern condition of the candidate scene. Furthermore, the scene determination unit 123a determines whether the distance to the object in the input content is equal to or less than a threshold in the candidate scene (e.g., 20 m or less), and judges whether the distance to the object satisfies the threshold condition of the candidate scene. If these conditions are met, the scene determination unit 123a determines the candidate scene as the detected scene (e.g., scene W).
[0088] The condition setting unit 123b sets various conditional expressions for scene detection. The condition setting unit 123b sets the conditional expressions based on, for example, information input by the creator of the XR content or the user.
[0089] For example, the condition setting unit 123b receives input from a producer or a user about what kind of realism is desired to be set for what kind of scene, and converts the situation of the scene into a conditional formula. Then, for each setting of a conditional formula, the condition setting unit 123b writes information about the conditional formula into the scene information DB 132 and writes the corresponding realism parameters into the parameter information DB 134.
[0090] In addition, the scene information DB 132 and the parameter information DB 134 may be registered and stored for each piece of content in a cloud server or the like, and the condition setting unit 123b may search for and retrieve the scene information and the parameter information 34 from the cloud server or the like based on information about the content to be viewed by the user before the content is viewed, and set them in the scene information DB 132 and the parameter information DB 134.
[0091] The condition setting unit 123b can set conditions for detecting a scene in which an object emits sound in a specified low-frequency range, and specifically sets the conditions as follows: For example, as a condition for detecting a scene including the walking sound (including sound in a low-frequency range) of an elephant located within 20 meters, the condition setting unit 123b sets the following: "Condition category": "Sound emitted from object," "Object": "Elephant," "Condition parameters": "Distance" and "Sound pattern," "Threshold": "20 m" (not shown), and "Acceptable threshold for difference between reference sound pattern (elephant in this case) and sound pattern (maximum difference that can be determined as similar)," and "Condition formula": "Distance is smaller than threshold" and "Sound pattern difference is equal to or smaller than acceptable threshold (sound patterns are similar)," and adds a record of the set data to the scene information DB 132 (corresponding to the record of scene W in FIG. 5).
[0092] In the above example, the condition setting unit 123b sets the conditions for determining, using an audio pattern, whether an object (e.g., an elephant) is present in a scene in the content. However, by setting conditions for determining using video analysis (using a video pattern) (for example, the "condition parameter" is "video pattern," the "threshold" is "allowable threshold for the difference between the reference video pattern (here, an elephant) and the video pattern (maximum difference that can be determined as similar)," and the "condition formula" is "video pattern difference is equal to or less than the allowable threshold (video patterns are similar)"), it is possible to similarly determine from the video whether an object (e.g., an elephant) is present in a scene.
[0093] Furthermore, the condition setting unit 123b sets (initializes or changes) the value of the "vibration parameter" in the parameter information DB 134. The main setting methods are the parameter setting method based on the input information by the creator or user as described above, and an automatic setting method based on the content type, etc.
[0094] Specifically, the method for setting parameters based on information input by the user involves the user selecting a scene for which the parameters are to be set (changed) and a type of parameter to be set (adjusted), and then changing the parameters to be set in that scene by operating the up / down operation button, etc. When setting, it is preferable to display a test image of the scene for which the parameters are to be set, and generate vibrations based on the parameters being set, so that the user can feel the vibrations while setting.
[0095] Furthermore, an automatic setting method based on content type, etc., detects the type of content to be played (determined from the content name and type information, etc., added to the content information, or estimated by analyzing part of the content video and audio), and corrects each parameter according to the detected content type. The correction values are obtained from correction value information (stored in a memory, etc., set by a device designer, etc., in the device, or obtained from a server (which collects correction value information from each device and stores appropriate correction values according to the content type by performing statistical processing, etc.)) previously set according to the content type.
[0096] This allows the scene information DB 132 and the parameter information DB 134 to be set more appropriately.
[0097] Furthermore, it is efficient for the condition setting unit 123b to set a condition for a scene in the content in which the amplitude of a sound in the low frequency range generated from an object exceeds a threshold value.
[0098] In other words, vibrations in the low frequency range have a large impact on the sense of realism felt by the user (person), so scenes in which such low frequency vibrations are relatively large (for example, vibrations that exceed an intensity threshold (it is preferable to add a moderate offset) that is perceived as noise) are selected as scenes to be subject to vibration control, and parameters for those scenes are set.
[0099] Such scenes can be set by the user or content creator, or obtained from a server (which collects scene information, parameter information, etc. for various content from each device and stores appropriate scene information and parameter information after performing statistical processing, etc.).
[0100] The intensity threshold may be determined based on the type (details) of content. Specifically, a data table of content types (details) and intensity thresholds is created in advance, and when selecting a scene for which conditions are to be set, the intensity threshold corresponding to the target content is searched for in the data table, and the scene for which conditions are to be set is selected using the searched intensity threshold.
[0101] For example, types of content include music videos that mainly allow users to listen to music, animal documentaries that explain the living creatures of animals, and so on.
[0102] In music videos, it is often best not to generate excessive vibrations when an elephant is walking, so as not to interfere with the music. On the other hand, in animal documentaries, it is often best to generate vibrations when an elephant is walking, to create a sense of realism.
[0103] For this reason, the threshold value for music videos is set lower than the threshold value for animal documentaries. As a result, the condition setting unit 123b is less likely to set a scene in which an elephant is walking in a music video as a scene in which vibration should be generated than a scene in which an elephant is walking in an animal documentary, and application of unnecessary vibrations to a scene in which an elephant is walking in a music video is suppressed.
[0104] This makes it possible to generate vibrations for each scene that are appropriate for the content that includes that scene.
[0105] The setting process of the above-mentioned scene information DB132 and parameter information DB134 may be performed by setting a new parameter value (for example, the adjustment value itself or a value to which an offset or the like has been added) based on various vibration adjustments (delay values, etc.) that the user actually made while watching the content.
[0106] This makes it possible for the information processing device 10 to detect scenes desired by the producer or user, and to set the sense of presence parameters desired by the producer or user for the detected scenes.
[0107] 4, the priority setting unit 124 will be described. The priority setting unit 124 sets priorities for the scenes detected by the scene detection unit 123.
[0108] For example, the priority setting unit 124 refers to the priority information DB 133 and selects which scene should be given priority in processing when multiple types of scenes are detected simultaneously by the scene detection unit 123. Note that when the scene detection unit 123 detects only one scene, that scene is given the highest priority.
[0109] Fig. 11 is a block diagram of the priority setting unit 124. For example, as shown in Fig. 11, the priority setting unit 124 includes a timing detection unit 124a and a rule setting unit 124b.
[0110] The timing detection unit 124a detects the timing at which a scene detected by the scene detection unit 123 occurs and the timing at which it ends. For example, the timing detection unit 124a detects each scene that exists at each time point (including whether it is overlapping), the timing at which an existing scene occurs, the timing at which an existing scene is deleted, etc., based on the scene information at each time point from the scene detection unit 123. In other words, the timing detection unit 124a grasps the state of all scenes that exist at each time point, including the order in which they occur.
[0111] The rule setting unit 124b sets a priority order for the scenes detected by the scene detection unit 123 to be used to determine the realism parameters. That is, based on the states of all existing scenes grasped by the timing detection unit 124a, the rule setting unit 124b sets a priority order for the detected scenes to determine which parameters associated with which scenes should be used preferentially as the realism parameters to be used at that time. This allows the information processing device 10 to set realism parameters according to the priority order.
[0112] That is, by setting a priority condition for each scene in advance, the information processing device 10 can appropriately determine which scene's realism parameter should be used preferentially when scene A and scene B overlap in time.
[0113] For example, the rule setting unit 124b sets the priority of the scene for determining the parameter to be used for each of the voice enhancement parameter and the vibration parameter with reference to the priority information DB 133. In this case, the rule setting unit 124b may set the scene to be used for parameter selection based on an independent priority rule for each speaker 4 and each vibration device 5, for example.
[0114] As a result, the realism parameters are set in accordance with unique rules for each speaker 4 and each vibration device 5, so that the realism can be further improved compared to when the realism parameters are set uniformly.
[0115] Furthermore, rule setting unit 124b associates information about the set rules with the video data and audio data and passes them to parameter extraction unit 125 (see FIG. 4).
[0116] 4, the parameter extraction unit 125 will be described. The parameter extraction unit 125 extracts a sense of presence parameter for the scene detected by the scene detection unit 123.
[0117] Fig. 13 is a block diagram of the parameter extraction unit 125. As shown in Fig. 13, the parameter extraction unit 125 has a vibration parameter extraction unit 125a, a voice emphasis parameter extraction unit 125b, and a learning unit 125c.
[0118] The vibration parameter extraction unit 125a refers to the parameter information DB 134 and extracts vibration parameters corresponding to the scene that has been set to the highest priority by the priority setting unit 124. For example, the vibration parameter extraction unit 125a extracts vibration parameters corresponding to the "detected scene" that has been received from the priority setting unit 124 and has the highest priority from the parameter information DB 134, thereby extracting vibration parameters corresponding to the scene.
[0119] In other words, when the scene detection unit 123 detects multiple scenes in which different objects emitting sounds overlap in time, the parameter extraction unit 125 can select a scene with a high priority, i.e., a scene that is estimated to give the user a more realistic feeling from vibration, and extract parameters for vibration generation corresponding to that scene. As a result, realistic vibrations can be generated using appropriate parameters even during content playback periods in which multiple scenes overlap.
[0120] Specifically, the scene detection unit 123 can perform such scene selection processing by using the priority rules of the priority information DB shown in Figure 8 and the priority conditions for each scene (which are set and stored in the scene information DB shown in Figure 5).
[0121] For example, if the scene detection unit 123 detects a scene in which an elephant makes a walking sound (an elephant walking scene) and a scene in which a horse makes a walking sound (a horse walking scene), the parameter extraction unit 125 prioritizes the elephant walking scene according to the rule that "prioritize the scene with a larger low-frequency amplitude." As a result, vibrations that reproduce the vibrations caused by an elephant's walking, which are primarily felt in the real world, are applied to the user in the content playback (for example, in a virtual space), and the user can experience a highly realistic, i.e., realistic, vibration sensation.
[0122] In addition, when the scene detection unit 123 detects multiple scenes that overlap in time in which different objects emitting sounds are present, the parameter extraction unit 125 can also apply a method of extracting parameters corresponding to a scene selected from the multiple scenes based on the type and position of the object corresponding to each of the multiple scenes in the images included in the content.
[0123] Specifically, by setting the priority rules of the priority information DB shown in Figure 8 and the priority conditions for each scene (which are set and stored in the scene information DB shown in Figure 5) (in this example, the function value F(M, d) of the type of object (m) and the distance to the object (d) is added to the priority conditions, and a condition based on the function value F(M, d) is set to the priority rules (for example, the larger the function value "F(M, d)" is, the higher the priority)), the scene detection unit 123 can perform such scene selection processing.
[0124] A method for determining a prioritized scene based on the position of an object will be described using a specific example shown in Fig. 12. Fig. 12 is a diagram showing an example of a method for determining a prioritized object.
[0125] 12, it is assumed that an image 31 of content being played is displayed on the display device 3. An object 311 (a horse) and an object 312 (an elephant) are shown in the image 31. At this time, it is assumed that the scene detection unit 123 has detected both a horse walking scene and an elephant walking scene that satisfy the conditions as target scenes for vibration control.
[0126] Also, assume that the distance from the reference position (the user's position relative to the content image, for example, the position of an avatar corresponding to the user in XR content) to object 311 is L1. Meanwhile, assume that the distance from the reference position to object 312 is L2. Also, assume that the reference vibration intensities (the intensities of the low-frequency components of the audio signals of the objects in the content) of object 311 and object 312 are V1 and V2, respectively. Furthermore, assume that the priority condition is set as "prioritize the one with the larger value of the function F(Ln, Vn) = Vn / (Ln Ln)."
[0127] The distance from the reference position to the object is calculated based on information added to the content (for example, calculated based on the position information of each object used to generate an image in XR content). The reference vibration intensity of the object can be determined by reading it from a data table that stores preset reference vibration intensities for each object type according to the type of target object, or by adding it to the content as content information. In addition, since content often includes audio data for audio playback, it is possible to calculate the reference vibration intensity based on the low-frequency characteristics of the audio data (audio intensity level, low-frequency signal level, etc.) (the vibration pattern is highly correlated with the low-frequency components of the audio, and vibrations are often generated based on the low-frequency components of the audio).
[0128] In this way, the information processing device 10 can estimate the low-frequency characteristics of the sound generated by the vibration generating object in the content. In this case, the information processing device 10 selects the vibration generating object based on the estimated low-frequency characteristics. This makes it possible to select a more appropriate vibration generating object.
[0129] For example, the low-frequency characteristic of the sound is the low-frequency signal level. In this case, the information processing device 10 selects a vibration generating object whose estimated low-frequency signal level exceeds a threshold. The information processing device 10 can extract the low-frequency signal level from the sound data. This makes it possible to easily select a vibration generating object using the low-frequency signal level included in the sound data.
[0130] In addition, the threshold for the low-frequency signal level is set according to the type of content. As mentioned above, it is often better to generate vibrations in a music video than in an animal documentary, even though the subject is the same. In this way, it is possible to select a vibration target that is appropriate for the content type (music video, animal documentary, etc.).
[0131] In this case, if the relationship between the function values of object 311 (horse) and object 312 (elephant) is function F(L1, V1)>function F(L2, V2), a scene in which object 311 generates sound (vibration), i.e., a scene in which the horse is walking, is preferentially selected, and the parameter extraction unit 125 extracts vibration parameters corresponding to the scene in which the horse is walking. Then, vibrations corresponding to the scene in which the horse is walking are applied to the user. Thereafter, for example, if object 312 (elephant) approaches the reference position and the relationship changes to function F(L1, V1)<function F(L2, V2), a scene in which object 311 generates sound (vibration), i.e., a scene in which the elephant is walking, is preferentially selected, and the parameter extraction unit 125 extracts vibration parameters corresponding to the scene in which the elephant is walking. Then, vibrations corresponding to the scene in which the elephant is walking are applied to the user.
[0132] Note that if the function F(Ln, Vn) is smaller than a predetermined threshold, that is, if the vibration caused by the object at the user's position in the content (such as the virtual space of a game) is small (the user does not feel it much, i.e., there is little need to apply vibration), it is also effective to not select it as an object that generates vibration. In other words, it is also effective to select only objects in content that cause a certain level of vibration caused by the object at the user's position in the content (such as the virtual space of a game) (to the extent that reproducing the vibration would improve the sense of realism) as objects that generate vibration. In other words, it is effective to select objects that have a large influence on the vibration signal generated from the candidate object that is the vibration-generating object (vibration objects whose vibrations the user feels strongly).
[0133] This allows the information processing device 10 to estimate candidate objects that have a large influence on the vibration signal generated from the candidate objects that are candidates for vibration generating objects, and select them as vibration generating objects. As a result, vibrations that match the sensations of the user in real space can be applied to the user, making it possible to play content with a rich sense of presence.
[0134] In this case, it is preferable to change the threshold value for selecting an object that generates vibration based on the type of content. That is, depending on the content, it may be preferable to suppress or emphasize the reproduction of vibrations caused by objects that appear in the content, and it is therefore preferable to adjust the determination content (judgment level) of the object that generates vibration.
[0135] In other words, the principle of vibration generation is as follows: The object that will generate vibration in the content (for each scene) is determined based on the content of the content. Then, a vibration signal (vibration data) is generated based on the acoustic signal corresponding to the determined object (audio data of the object included in the content, or audio data of the object generated from audio data in that scene (for example, extracted by filtering the low-frequency range)). (The vibration signal (vibration data) is generated by extracting the low-frequency components of the audio signal of the object and amplifying it appropriately.)
[0136] In addition, a method for determining the object that generates vibrations is to estimate the low-frequency characteristics (e.g., volume level) of the vocalizations of the sound-generating object in the content (in the above example, estimation is based on the reference vibration strength based on the type of object and the distance between the reference position (such as the user's location in the virtual space of the content) and the object), and then determine the object (the sound-generating object with the higher low-frequency volume level of the vocalizations is determined as the object that generates vibrations).
[0137] In this way, by determining the priority scene based on the position of the object, vibrations that are more suited to the user's visual intuition, that is, vibrations that match the user's sensations in real space, can be applied to the user, making it possible to play content that is more immersive.
[0138] At this time, the vibration parameter extraction unit 125a extracts vibration parameters corresponding to each of the vibration devices 5. This makes it possible to further improve the sense of realism compared to the case where vibration parameters are uniformly extracted.
[0139] The voice enhancement parameter extraction unit 125b refers to the parameter information DB 134 and extracts a voice enhancement parameter corresponding to the scene that has been assigned the highest priority by the priority setting unit 124. The voice enhancement parameter extraction unit 125b extracts a voice enhancement parameter individually for each speaker 4, and determines the voice enhancement parameter to be extracted based on the priority set by the priority setting unit 124 (based on the scene with the highest priority), similar to the vibration parameter extraction unit 125a.
[0140] The learning unit 125c learns the relationship between the scenes and the realism parameters stored in the parameter information DB 134. For example, the learning unit 125c learns the relationship between the scenes and the realism parameters by performing machine learning on each scene stored in the parameter information DB 134 and each corresponding realism parameter using, as learning data, the user's reaction to the realism control by the parameter, etc.
[0141] In this case, for example, the learning unit 125c may use user evaluations of the realism parameters (user adjustment operations after realism control, user input such as questionnaires) as learning data. That is, the learning unit 125c may learn the relationship between scenes and realism parameters from the viewpoint of what kind of realism parameters should be set for what kind of scenes to obtain high user evaluations (i.e., whether a high sense of realism is obtained).
[0142] Furthermore, the learning unit 125c can determine what realism parameters should be set when a new scene is input, based on the learning results. As a specific example, the realism parameters for a fireworks scene can be determined using the learning results of realism control for a similar situation, such as an explosion scene. It is also possible to learn rules regarding priority based on the presence or absence and degree of factors that change the priority in user adjustment operations after realism control or user input such as a questionnaire (for example, when the user adjustment operation brings the parameters closer to those corresponding to other scenes that exist simultaneously, or when there is a response in the questionnaire indicating that other scenes should be given priority).
[0143] This allows the information processing device 10 to automatically optimize, for example, rules regarding priority and realistic parameters.
[0144] 4, the following describes the output unit 126. The output unit 126 outputs the realism parameters extracted by the parameter extraction unit 125 to the speaker 4 and the vibration device 5.
[0145] Fig. 14 is a block diagram of the output unit 126. As shown in Fig. 14, the output unit 126 has a voice emphasis processing unit 126a and a voice vibration conversion processing unit 126b.
[0146] The voice enhancement processing unit 126a performs enhancement processing on the voice data received from the rendering processing unit 122 using the voice enhancement parameters extracted by the parameter extraction unit 125. For example, the voice enhancement processing unit 126a performs enhancement processing on the voice data by performing delay or band enhancement / attenuation processing based on the voice enhancement parameters.
[0147] At this time, the voice enhancement processing unit 126a performs voice enhancement processing for each speaker 4, and outputs the voice data that has been subjected to the voice enhancement processing to each corresponding speaker 4.
[0148] The sound vibration conversion processing unit 126b converts the sound data received from the rendering processing unit 122 into vibration data by performing band limiting processing suitable for vibration, such as LPF, etc. Furthermore, the sound vibration conversion processing unit 126b performs emphasis processing on the converted vibration parameters in accordance with the vibration parameters extracted by the parameter extraction unit 125.
[0149] For example, the sound vibration conversion processing unit 126b performs emphasis processing on the vibration data by performing emphasis processing on the vibration data, such as frequency characteristic addition processing such as low frequency emphasis, and delay and amplification, according to the vibration parameters. In this way, the sound vibration conversion processing unit 126b processes the signal of the sound generated from the object to obtain a signal suitable for the vibration, and outputs the signal (vibration data) that has been emphasis-processed using the vibration parameters to the vibration device.
[0150] At this time, the sound vibration conversion processing unit 126b performs vibration emphasis processing for each vibration device 5, and outputs the vibration data that has been subjected to the vibration emphasis processing to each corresponding vibration device 5.
[0151] Note that the vibration parameters are set according to the scene (for example, "elephant walking scene") as shown in Figure 9, but it is also effective to make corrections according to the detailed situation in the scene (which can also be called the detailed scene type). For example, depending on the distance between the user and the elephant in the content (virtual space) (according to the detailed scene by distance), the values of the vibration generation parameters "IPF (cutoff frequency)," "delay time," and "amplification" can be increased or decreased to adjust the vibration characteristics.
[0152] This allows the information processing device 10 to determine the scene type of a scene in which a vibration generating object appears, and to process the acoustic signal corresponding to the vibration generating object according to the determined scene type to generate a vibration signal, thereby enabling the generation of a vibration signal that is precisely adapted to each scene.
[0153] Next, a processing procedure executed by the information processing device 10 according to the embodiment will be described with reference to Fig. 15. Fig. 15 is a flowchart showing the processing procedure executed by the information processing device 10. Note that the processing procedure shown below is repeatedly executed by the control unit 120.
[0154] The processing of the flowchart shown in Fig. 15 is executed when the information processing system 1 is powered on. First, an XR content setting process is executed (step S101). Note that the XR content setting process here includes, for example, various processes related to the initial settings of the device for playing XR content, the selection of XR content by the user, etc.
[0155] Next, the information processing device 10 starts playing the XR content (step S102) and performs a scene detection process on the XR content being played (step S103). Next, the information processing device 10 performs a priority setting process on the results of the scene detection process (step S104), and executes a realism parameter extraction process for the scene to be subjected to vibration control based on the priority setting content (step S105).
[0156] Then, the information processing device 10 executes an output process of various vibration data or audio data generated based on the extracted realism parameters (step S106). As a result, vibrations that provide a realism are output from the vibration device 5, and audio is output from the speaker 4.
[0157] Then, the information processing device 10 determines whether or not the XR content has ended (step S107), and if it determines that the XR content has ended (step S107; Yes), ends the processing.
[0158] Furthermore, when the information processing device 10 determines in step S107 that the XR content has not ended (step S107; No), the information processing device 10 again proceeds to the processing of step S103.
[0159] As described above, the information processing device 10 according to the embodiment selects a vibration-generating object from the content to be used as a vibration source, processes an acoustic signal corresponding to the selected vibration-generating object to generate a vibration signal, and outputs the generated vibration signal as a vibration signal corresponding to the content.
[0160] Therefore, vibration control is performed based on a vibration source that is suitable for providing vibration during content playback, thereby efficiently improving the sense of realism.
[0161] In the above-described embodiment, the content is XR content, but the present invention is not limited to this. That is, the content may be 2D video and audio, or video only, or audio only.
[0162] Further advantages and modifications will readily occur to those skilled in the art. Therefore, the invention in its broader aspects is not limited to the specific details and representative embodiments shown and described above. Accordingly, various modifications may be made without departing from the spirit or scope of the general inventive concept as defined by the appended claims and their equivalents. [Explanation of symbols]
[0163] 1. Information Processing Systems 3 Display device 4 speakers 5. Vibration Devices 10. Information processing equipment 31 images 311, 312 Objects 121 Content Generation Department 122 Rendering Processor 123 Scene detection section 123a Scene determination section 123b Condition setting section 124 Priority Setting Section 124a Timing detection unit 124b Rule setting section 125 Parameter Extraction Unit 125a Vibration parameter extraction unit 125b Speech enhancement parameter extraction unit 125c Learning Department 126 Output section 126a Speech enhancement processing unit 126b Voice vibration conversion processing unit 131 XR Content DB 132 Scene Information DB 133 Priority Information DB 134 Parameter Information DB
Claims
1. An information processing device that outputs a vibration signal according to input content, the information processing device having a controller, The controller selecting an object in the content whose low-frequency signal level of a sound generated by the object exceeds a threshold as a vibration-generating object that is a source of vibration; extracting an acoustic signal corresponding to the selected vibration generating object from the content; processing the extracted acoustic signal to generate a vibration signal; The threshold value is set according to the content type of the content. Information processing device.
2. The controller determining a scene type in a scene in which a vibration-generating object appears; The acoustic signal corresponding to the vibration-generating object is processed according to the determined scene type to generate a vibration signal. The information processing device according to claim 1 .
3. The controller selecting, as the vibration generating object, a candidate object that has a large influence on the vibration signal generated from the candidate object that is the vibration generating object; 3. The information processing device according to claim 1 or 2.
4. The processing of the acoustic signal includes at least one of band limiting, delaying, and amplifying the acoustic signal. The information processing device according to claim 1 .
5. a controller of an information processing device that outputs a vibration signal according to input content, selecting an object in the content whose low-frequency signal level of a sound generated by the object exceeds a threshold as a vibration-generating object that is a source of vibration; extracting an acoustic signal corresponding to the selected vibration generating object from the content; processing the extracted acoustic signal to generate a vibration signal; The threshold value is set according to the content type of the content. Information processing methods.
6. A controller of an information processing device that outputs a vibration signal according to input content, selecting an object in the content whose low-frequency signal level of a sound generated by the object exceeds a threshold as a vibration-generating object that is a source of vibration; extracting an acoustic signal corresponding to the selected vibration generating object from the content; processing the extracted acoustic signal to generate a vibration signal; The threshold value is set according to the content type of the content. A program that executes a process.
7. an information processing device that plays XR content, outputs a vibration signal according to the input content, and has a controller; a display device that displays an image in accordance with the image signal output from the information processing device; an audio output device that generates audio in response to an audio signal output from the information processing device; a vibration device that applies vibration to a user in response to a vibration signal output from the information processing device; Equipped with The controller selecting an object in the content whose low-frequency signal level of a sound generated by the object exceeds a threshold as a vibration-generating object that is a source of vibration; extracting an acoustic signal corresponding to the selected vibration generating object from the content; processing the extracted acoustic signal to generate a vibration signal; The threshold value is set according to the content type of the content. Information processing system.
Citation Information
Patent Citations
Bodily feeling experiencing video / sound system
JP2004081357A
Vibration feeling device
JP2019121813A
Method and device for audio signal processing, and storage medium
US20200211577A1
Electronic device
WO2013168732A1