Information processing device, information processing method, and information processing program

The information processing device improves XR content realism by analyzing audio signals to amplify low-frequency bands and adjust vibrations based on user characteristics, addressing inconsistencies in conventional technologies.

JP7808949B2Active Publication Date: 2026-01-30DENSO TEN LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2021175611
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-10-27
Publication Date
2026-01-30
Estimated Expiration
2041-10-27

AI Technical Summary

Technical Problem

Conventional technologies for providing vibration stimulation during XR content playback lack realism due to the removal of low-frequency bands in audio recordings and variations in chair materials and user physiques, leading to inconsistent vibration transmission.

Method used

An information processing device that generates a vibration stimulation signal by analyzing audio signals, amplifying low-frequency bands using low-pass filters when necessary, and adjusting vibrations based on user characteristics and application environments.

Benefits of technology

Enhances the sense of realism during XR content playback by ensuring appropriate vibration transmission, regardless of chair materials and user physiques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007808949000001
    Figure 0007808949000001
  • Figure 0007808949000002
    Figure 0007808949000002
  • Figure 0007808949000003
    Figure 0007808949000003
Patent Text Reader

Abstract

To more increase realistic feeling due to vibration stimulation at the time of content reproduction.SOLUTION: An information processing device includes a control unit that generates a vibration stimulation signal to be given to a user, based on a sound signal in a content. The control unit obtains data of the content including the sound signal; performs analysis processing of the sound signal; and generates a vibration stimulation signal to be given to the user by conversion processing for the sound signal according to a result of the analysis processing.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The disclosed embodiments relate to an information processing device, an information processing system, and an information processing method. [Background technology]

[0002] Conventionally, a technology has been known that uses a head-mounted display (HMD) or the like to provide a user in a remote location with live experiential XR (Cross Reality) content, including video and audio recorded live at an event venue or the like.

[0003] XR is a collective term for all virtual space technologies, including VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), SR (Substitutional Reality), and AV (Audio / Visual).

[0004] There is also known a technology that, when playing such XR content, activates a vibrator such as an exciter built into a chair or the like, allowing the user to simulate a sense of vibration or impact corresponding to the video and audio being played (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2007-324829 Summary of the Invention [Problem to be solved by the invention]

[0006] However, the conventional technology has room for further improvement in terms of further enhancing the sense of realism provided by vibration stimulation during content playback.

[0007] For example, when recording audio in an external environment, it is common to cut low-frequency bands using a high-pass filter (HPF) to remove noise such as footsteps, wind noise, etc. As a result, audio recorded in an external environment lacks low-frequency bands, and even if vibrations are generated based on such audio, it is difficult for the user to feel a sense of realism.

[0008] Furthermore, since the objects to which vibrations are applied, such as chairs and users, vary in material and type for chairs, and the physiques of users vary, the characteristics of the same vibration stimulus usually differ. As a result, the vibrations may not be transmitted as intended, and the user may not feel a sense of realism.

[0009] One aspect of the embodiment has been made in consideration of the above, and aims to provide an information processing device, an information processing system, and an information processing method that can further improve the sense of realism provided by vibration stimulation when content is played. [Means for solving the problem]

[0010] According to one aspect of the present invention, there is provided an information processing device having a control unit that generates a vibration stimulation signal to be provided to a user based on an audio signal in content, The audio signal in the content has a low frequency band cut by a high-pass filter, The control unit The aforementioned Acquiring an audio signal, and cutting a high frequency band in the acquired audio signal with a low pass filter; The lowest frequency in the frequency band that is not cut by the high-pass filter When the level of the audio signal at the cutoff frequency exceeds a predetermined threshold, a signal having a frequency equal to or lower than the cutoff frequency and having a high frequency band in the audio signal cut off by a low-pass filter is amplified to generate a vibration signal, High Pass If the level of the audio signal at the cutoff frequency of the filter does not exceed a predetermined threshold, the audio signal whose high frequency band is cut by a low-pass filter is used as the vibration signal. [Effects of the Invention]

[0011] According to one aspect of the embodiment, it is possible to further improve the sense of realism provided by vibration stimulation during content playback. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a diagram illustrating an outline of an information processing method according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of the information processing system according to the first embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of the configuration of the on-site device according to the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of the configuration of the remote device according to the first embodiment. [Figure 5] FIG. 5 is a diagram illustrating an example of the configuration of the vibration output unit. [Figure 6] FIG. 6 is a block diagram of the remote device according to the first embodiment. [Figure 7] FIG. 7 is a block diagram of the sound vibration conversion processing unit. [Figure 8] FIG. 8 is a supplementary explanatory diagram (part 1) of the audio signal conversion process according to the first embodiment. [Figure 9] FIG. 9 is a supplementary explanatory diagram (part 2) of the audio signal conversion process according to the first embodiment. [Figure 10] FIG. 10 is a supplementary explanatory diagram (part 3) of the audio signal conversion process according to the first embodiment. [Figure 11] FIG. 11 is a flowchart (part 1) illustrating a processing procedure executed by the remote device according to the first embodiment. [Figure 12] FIG. 12 is a flowchart (part 2) illustrating the processing procedure executed by the remote device according to the second embodiment. [Figure 13] FIG. 13 is a block diagram of a remote device according to the second embodiment. [Figure 14] FIG. 14 is a supplementary explanatory diagram (part 1) of the audio signal conversion process according to the second embodiment. [Figure 15]FIG. 15 is a supplementary explanatory diagram (part 2) of the audio signal conversion process according to the second embodiment. [Figure 16] FIG. 16 is a supplementary explanatory diagram (part 3) of the audio signal conversion process according to the second embodiment. [Figure 17] FIG. 17 is a block diagram of a remote device according to the third embodiment. [Figure 18] FIG. 18 is a flowchart illustrating a processing procedure executed by a remote device according to the third embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, embodiments of an information processing device, an information processing system, and an information processing method disclosed in the present application will be described in detail with reference to the accompanying drawings. Note that the present invention is not limited to the following embodiments.

[0014] In the following description, multiple components having substantially the same functional configuration may be distinguished by adding a different number with a hyphen after the same reference symbol. For example, multiple components having substantially the same functional configuration may be distinguished as necessary, such as remote device 100-1 and remote device 100-2. However, when there is no need to particularly distinguish between multiple components having substantially the same functional configuration, only the same reference symbol is used. For example, when there is no need to particularly distinguish between remote device 100-1 and remote device 100-2, they will simply be referred to as remote device 100.

[0015] First, an overview of an information processing method according to an embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram illustrating an overview of an information processing method according to an embodiment.

[0016] The information processing system 1 according to the embodiment is a system that provides live experience-type XR content including on-site video and audio from an event venue such as an exhibition venue, concert venue, fireworks venue, or e-sports tournament venue to a remote location outside the venue. The XR content is an example of "content."

[0017] 1, the information processing system 1 includes a local device 10 and one or more remote devices 100. The local device 10 and the remote devices 100 are provided so as to be able to communicate with each other via a network N such as the Internet.

[0018] The example in Figure 1 shows a situation in which a local device 10 live-streams XR content including video and audio of an event being held at an event venue 1000 in the Kansai region to remote devices 100 in various locations.

[0019] The example in FIG. 1 also shows a situation in which the remote device 100-1 presents XR content delivered from the local device 10 to a user U1 located in the Kanto region via an HMD.

[0020] The HMD is an information processing terminal that presents XR content to the user U1 and allows the user U1 to enjoy an XR experience. The HMD is a wearable computer that is worn on the head of the user U1, and is goggle-shaped in the example of FIG. 1. The HMD may also be in the form of glasses or a hat.

[0021] The HMD includes a video output unit 110 and an audio output unit 120. The video output unit 110 displays video included in the XR content provided by the on-site device 10. In the example of FIG. 1, the video output unit 110 of the HMD is provided so as to be placed in front of the eyes of the user U1.

[0022] The audio output unit 120 outputs audio included in the XR content provided by the on-site device 10. In the example of Fig. 1, the audio output unit 120 of the HMD is provided in the form of, for example, an earphone, and is worn on the ear of the user U1.

[0023] 1 shows a situation in which the remote device 100-2 presents XR content delivered from the local device 10 via a satellite dome D to a user U2 located in the Kyushu region.

[0024] Satellite dome D is a viewing facility for XR content, is dome-shaped, and includes a video output unit 110 and an audio output unit 120. In the example of FIG. 1, video output unit 110 of satellite dome D is provided on a wall surface. Video output unit 110 is realized, for example, by installing a thin liquid crystal display or an organic EL (Electro Luminescence) display on the wall surface, or by projecting an image onto the wall surface using a projector. Furthermore, audio output unit 120 of satellite dome D is provided near the head position of seated user U2.

[0025] Although not shown in FIG. 1, a vibration output unit 130 (see FIG. 4 and subsequent figures) is provided near the users U1 and U2. The vibration output unit 130 outputs vibrations corresponding to the audio included in the XR content, providing vibration stimulation to the users U1 and U2. The vibration output unit 130 is realized by a vibrator such as an exciter, and is provided inside the chairs on which the users U1 and U2 sit, or worn by the users U1 and U2.

[0026] Incidentally, when recording audio in an external environment, it is common to cut the low frequency band using an HPF to remove noise such as footsteps, wind noise, etc. However, even if an audio signal passed through such an HPF is input to the vibration output unit 130 to provide vibration stimulation to users U1 and U2, it is difficult for users U1 and U2 to get a sense of realism due to the lack of the low frequency band.

[0027] Therefore, for example, it is conceivable to use existing technology to input the audio signal to the vibration output unit 130 via an equalizer that boosts the low frequency band. However, when boosting the low frequency band with an equalizer, there are problems such as boosting low frequency noise that was not cut off, and the equalizer alone cannot boost the cut frequency band sufficiently to contribute to improving the sense of realism.

[0028] Furthermore, since the objects to which vibrations are applied, such as chairs and users, vary in material and type for chairs, and the physiques of users vary, the characteristics of the same vibration stimulus usually differ. As a result, the vibrations may not be transmitted as intended, and the user may not feel a sense of realism.

[0029] In this regard, when using existing technology, it is conceivable to adjust the vibration based on the user's sense, or to adjust the vibration based on actual measured values ​​from an acceleration sensor, etc. However, when adjusting based on the user's sense, it is difficult to reproduce the vibration that the stimulus presenter wants to give, and when the stimulus presenter adjusts the vibration based on actual measured values, a stimulus presenter with the necessary know-how is always required.

[0030] Therefore, in the information processing method according to the embodiment, XR content including an audio signal is acquired, an analysis process is performed to convert the audio signal into vibration, and a vibration stimulus is generated to be given to the user according to the results of the analysis process.

[0031] 1, in the information processing method according to the embodiment, the local device 10 first provides XR content from the local location (step S1). The remote device 100 then performs an analysis process for converting the audio signal in the provided XR content into vibrations, and generates a vibration pattern to be provided to the user according to the analysis result (step S2).

[0032] For example, in the information processing method according to the embodiment, (1) a frequency analysis of an audio signal is performed using a technique such as FFT (Fast Fourier Transform). If the result shows that the level of a predetermined low frequency band is below a preset threshold, the frequency is divided by N (1 / N) by pitch shifting; if the level is not below the threshold, the frequency is output as is.

[0033] Furthermore, for example, in the information processing method according to the embodiment, (2) a sound source of an audio signal is estimated using an AI (Artificial Intelligence) inference model that estimates a sound source from an audio signal. As a result, if the sound source is subject to frequency division in the settings, the frequency is divided by N by pitch shifting, and if the sound source is not subject to frequency division, the frequency is output as is.

[0034] Furthermore, for example, in the information processing method according to the embodiment, (3) as a method other than pitch shifting, a threshold is set at frequency A, the lowest frequency band that is not cut, and when a sound whose volume exceeds the threshold is input, a signal composed of frequencies equal to or lower than frequency A is input, and the low frequency band is enhanced.

[0035] The above (1) to (3) will be described later as a first embodiment with reference to FIGS. 2 to 12. FIG.

[0036] Furthermore, for example, in the information processing method according to the embodiment, (4) calibration of vibration characteristics is performed according to the difference between the objects to which vibration is applied and the state of the objects. This (4) will be described later as a second embodiment with reference to FIGS. 13 to 16.

[0037] Furthermore, for example, in the information processing method according to the embodiment, (5) a specific scene is detected from the input video signal and audio signal. Then, according to the detected scene, the frequency is divided by N by pitch shifting based on a preset vibration parameter. This (5) will be described later as a third embodiment using FIGS. 17 and 18.

[0038] That is, in the information processing method according to the embodiment, a vibration pattern to be given to the user is generated according to the analysis processing results as described above in (1) to (5). Then, the remote device 100 drives the vibration output unit 130 based on the generated vibration pattern, and gives a vibration stimulus with an enhanced low frequency band, or a vibration stimulus according to the target, for example.

[0039] This can further improve the sense of realism provided by vibration stimulation during content playback.

[0040] In this way, in the information processing method according to the embodiment, XR content including an audio signal is acquired, an analysis process is performed to convert the audio signal into vibration, and a vibration stimulus is generated to be given to the user according to the results of the analysis process.

[0041] Therefore, according to the information processing method of the embodiment, it is possible to further improve the sense of realism caused by vibration stimulation during content playback. Hereinafter, each embodiment of the information processing system 1 to which the information processing method of the embodiment is applied will be described in more detail.

[0042] First Embodiment

[0043] Fig. 2 is a diagram showing an example of the configuration of an information processing system 1 according to the first embodiment. Fig. 3 is a diagram showing an example of the configuration of a local device 10 according to the first embodiment. Fig. 4 is a diagram showing an example of the configuration of a remote device 100 according to the first embodiment. Fig. 5 is a diagram showing an example of the configuration of a vibration output unit 130.

[0044] 2, the information processing system 1 includes a local device 10 and one or more remote devices 100. The local device 10 and the remote device 100 are examples of "information processing devices" and are each realized by a computer. The local device 10 and the remote device 100 are provided so as to be able to communicate with each other via a network N, which may be the Internet, a dedicated line network, a mobile phone network, or the like.

[0045] As shown in FIG. 3 , the local device 10 has one or more cameras 11 and one or more microphones 12. The cameras 11 record images of the external environment. The microphones 12 record audio of the external environment. The local device 10 generates XR content including the images recorded by the cameras 11 and the audio recorded by the microphones 12, and provides the content to the remote device 100.

[0046] 4, the remote device 100 has a video output unit 110, an audio output unit 120, and a vibration output unit 130. The video output unit 110 displays video included in the XR content provided by the local device 10. The audio output unit 120 outputs audio included in the XR content.

[0047] The vibration output unit 130 outputs vibrations corresponding to audio included in the XR content. As already mentioned, as shown in FIG. 5, the vibration output unit 130 is provided inside the chair S on which the user U sits. The vibration output unit 130 may be attached to the user U, for example, by being embedded in clothing or a seat belt. The vibration output unit 130 is a known electric vibration conversion device, for example, an electric vibration converter formed from a magnet (magnetic circuit) and a coil through which a drive current flows, or a piezoelectric element, and has a built-in power amplifier that amplifies the signal to a level required for driving.

[0048] Next, Fig. 6 is a block diagram of the remote device 100 according to the first embodiment. Fig. 7 is a block diagram of the sound vibration conversion processing unit 103b. Note that Figs. 6 and 7, as well as Figs. 13 and 17 shown later, show only components necessary for explaining the features of the embodiment, and general components are omitted.

[0049] In other words, the components shown in Figures 6, 7, 13, and 17 are conceptual functional components and do not necessarily have to be physically configured as shown. For example, the specific form of distribution and integration of each block is not limited to that shown, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0050] In addition, in the description using FIGS. 6, 7, 13, and 17, the description of components that have already been described may be simplified or omitted.

[0051] 6, the remote device 100 according to the embodiment includes a communication unit 101, a storage unit 102, and a control unit 103. The remote device 100 is also connected to a vibration output unit 130. Note that the video output unit 110 and the audio output unit 120 are intentionally omitted to more clearly explain the features of the embodiment.

[0052] The communication unit 101 is realized by, for example, a network interface card (NIC) etc. The communication unit 101 is connected to a network N by wire or wirelessly, and transmits and receives information to and from the on-site device 10 via the network N.

[0053] The storage unit 102 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. In the example of Fig. 6, the storage unit 102 stores vibration parameter information 102a and a sound source estimation model 102b.

[0054] The vibration parameter information 102a is information including various parameters related to the vibration to be output to the vibration output unit 130, and includes, for example, various thresholds used for determination described below. The sound source estimation model 102b is an AI inference model that estimates a sound source from the above-mentioned audio signal.

[0055] The sound source estimation model 102b receives an audio signal as input, passes it through a trained neural network, and outputs the sound source class with the highest probability from the probability distribution of the final layer as a result. Learning is performed using the audio signal and information on the sound source class assigned to the audio signal as correct answer data, so as to minimize the difference (cost) between the output result of the classifier and the correct answer data. The correct answer data is collected, for example, through manual annotation.

[0056] The control unit 103 is a controller, and is realized, for example, by a CPU (Central Processing Unit) or an MPU (Micro Processing Unit) executing various programs (not shown) stored in the storage unit 102 using a RAM as a work area. The control unit 103 can also be realized, for example, by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0057] The control unit 103 has an acquisition unit 103a and a sound vibration conversion processing unit 103b, and realizes or executes the functions and actions of information processing described below.

[0058] The acquisition unit 103a acquires the XR content provided from the on-site device 10 via the communication unit 101.

[0059] The sound vibration conversion processing unit 103b inputs the sound signal included in the XR content acquired by the acquisition unit 103a and performs analysis processing for vibration conversion. The sound vibration conversion processing unit 103b also generates a vibration pattern to be given to the user according to the analysis processing result.

[0060] As shown in FIG. 7, the sound vibration conversion processing unit 103b includes a high-frequency cut unit 103ba, a determination unit 103bb, a pitch shift unit 103bc, and an amplification unit 103bd.

[0061] The high-frequency cut unit 103ba uses a low-pass filter (LPF) to cut out high-frequency bands that are unnecessary for vibration conversion as preprocessing for the audio signal, which has already had low frequencies cut by an HPF during recording. This is because humans mainly sense vibrations in the low frequency range. The judgment unit 103bb inputs and analyzes the audio signal from which the high frequencies have been cut, and determines whether pitch shifting is necessary or not.

[0062] The determination unit 103bb determines whether pitch shifting is necessary or not by, for example, frequency analysis using FFT, etc. Furthermore, the determination unit 103bb inputs, for example, a speech signal to the sound source estimation model 102b, and determines whether pitch shifting is necessary or not based on the output result of the sound source estimation model 102b in response to the input.

[0063] The pitch shifter 103bc performs pitch shifting on the audio signal when the decision unit 103bb decides that pitch shifting is necessary. The amplifier 103bd amplifies the audio signal and outputs it to the vibration output unit 130 as a vibration signal.

[0064] Here, the audio signal conversion process will be further explained with reference to Figures 8 to 10. Figures 8 to 10 are supplementary explanatory diagrams (part 1) to (part 3) of the audio signal conversion process according to the first embodiment.

[0065] As shown in the upper and middle sections of Figure 8, actual environmental sounds recorded on-site have their low frequencies cut by an HPF. As a result, if the signal level of a predetermined frequency band is below a set threshold, for example, if the average level at 20 Hz is -20 dB or less, the frequency is divided by N (here, N = 2) by pitch shifting in the audio signal conversion process, as shown in the lower section of Figure 8. The example in Figure 8 corresponds to (1) above. The predetermined frequency band and threshold are set to appropriate values ​​based on, for example, the results of a sensitivity test.

[0066] Also, as shown in Figure 9, a threshold is set at the cutoff frequency, which is the lowest frequency in the frequency band that is not cut, and if the sound level of the cutoff frequency does not exceed that threshold, the audio signal conversion process does not boost the low frequency band.

[0067] On the other hand, as shown in Figure 10, when a sound at the cutoff frequency exceeds the threshold, the audio signal conversion process simultaneously inputs a signal composed of frequencies below the cutoff frequency (for example, a signal that provides an appropriate sensation created based on the results of a sensitivity test, etc.) to boost the low-frequency band. The examples in Figures 9 and 10 correspond to (3) above. This method is expected to be used when frequency division is not effective but boosting the low-frequency band contributes to improving the sense of presence.

[0068] Next, the processing procedure executed by the remote device 100 will be described with reference to Fig. 11 and Fig. 12. Fig. 11 is a flowchart (part 1) showing the processing procedure executed by the remote device 100 according to the first embodiment. Fig. 12 is a flowchart (part 2) showing the processing procedure executed by the remote device 100 according to the first embodiment.

[0069] 11 and 12 mainly show the processing procedure of the audio signal conversion process. Fig. 11 corresponds to (1) above. Fig. 12 corresponds to (2) above.

[0070] In the case of the above (1), as shown in FIG. 11, the sound vibration conversion processing unit 103b first receives an audio signal (step S101) and performs frequency analysis of the audio signal (step S102).

[0071] Then, it is determined whether the signal level of the predetermined low frequency band is below the set threshold value (step S103). If it is below the threshold value (step S103, Yes), frequency division is performed (step S104), and the divided audio signal is output to the vibration output unit 130 as a vibration signal (step S105). Then, the process ends.

[0072] On the other hand, if the threshold value is exceeded (step S103, No), the audio signal is output as it is to the vibration output unit 130 as a vibration signal (step S105). Then, the process ends. In the case of (3) above, the process of step S104 is a process of adding a signal composed of frequencies equal to or lower than the cutoff frequency.

[0073] In the case of (2) above, as shown in FIG. 12, the sound vibration conversion processing unit 103b first receives an audio signal (step S201), and performs inference on the audio signal using the sound source estimation model 102b (step S202).

[0074] Then, as a result of the inference, it is determined whether the sound source is a target for frequency division (step S203). If the sound source is a target for frequency division (step S203, Yes), frequency division is performed (step S204), and the audio signal after frequency division is output to the vibration output unit 130 as a vibration signal (step S205). Then, the process ends.

[0075] On the other hand, if the sound source is not a target for frequency division (No at step S203), the audio signal is output as it is to the vibration output unit 130 as a vibration signal (step S205), and the process then ends.

[0076] <Second embodiment> Next, a second embodiment corresponding to (4) above will be described. Fig. 13 is a block diagram of a remote device 100A according to the second embodiment. Since Fig. 13 corresponds to Fig. 6, differences from Fig. 6 will be mainly described here. Figs. 14 to 16 are supplementary explanatory diagrams (parts 1 to 3) of the audio signal conversion process according to the second embodiment.

[0077] 13, the remote device 100A differs from the first embodiment in that it further includes an acceleration sensor 140 and a calibration unit 103c. The calibration unit 103c calibrates vibration characteristics according to differences between objects to which vibration is applied and the state of the objects.

[0078] First, we will explain the case where the difference between objects to which vibration is applied is calibrated. In this case, the calibration unit 103c acquires the vibration characteristics when a predetermined reference signal is applied to the reference object before presenting the actual vibration. For example, an acceleration sensor 140 is installed on the seat of the reference chair α, and the actual vibration characteristics of the chair α when the reference signal is applied are acquired. The vibration signal is generated based on the characteristics of this reference chair α to achieve the desired vibration. The example in FIG. 14 shows the vibration characteristics of the reference chair α, and if the vibration characteristics of the chair or the like being used are close to these characteristics, the desired vibration can be applied to the user.

[0079] On the other hand, the calibration unit 103c acquires vibration characteristics when the same reference signal is applied to a vibration device used by a user that receives actual vibration. In this case, for example, an acceleration sensor 140 is installed on the seat of the wheelchair β, and actual vibration characteristics are acquired when the reference signal is input to the wheelchair β. Figure 15 shows the vibration characteristics of such wheelchair β.

[0080] Then, the calibration unit 103c adjusts the output level at each frequency of the vibration signal to be output to the wheelchair β so as to reduce the difference between the vibration characteristics of the chair α and the vibration characteristics of the wheelchair β.

[0081] For example, as shown in Figures 14 and 15, assume that wheelchair β has vibration characteristics in which vibrations at 40 Hz are extremely attenuated compared to chair α. In this case, the calibration unit 103c adjusts the vibration signal to wheelchair β using an equalizer to increase the level of 40 Hz by +2 dB or more, as shown in Figure 16. Then, the calibration unit 103c stores such adjustment characteristics in the vibration parameter information 102a, and the voice vibration conversion processing unit 103b uses this to make adjustments when actually applying vibrations to wheelchair β.

[0082] Next, we will explain the case of calibration based on the state of the subject. Although it is difficult to measure how a person feels vibrations on the skin, it is generally known that the strength of vibration stimuli felt by a person is related to the amount of body fat.

[0083] Therefore, the calibration unit 103c stores in advance each parameter for vibration adjustment for each weight level in 10 kg increments, for example. The calibration unit 103c measures the weight of the subject who will actually receive the vibration. For example, subject C weighs 80 kg.

[0084] Then, the calibration unit 103c adjusts the vibration characteristics for subject C so that subject C, that is, a person weighing 80 kg, feels the same vibration as subject B, who has an appropriate weight. For example, it can be estimated that a person weighing 80 kg feels vibration less than a person weighing 60 kg, so in this case, the calibration unit 103c adjusts the vibration output level for subject C to be, for example, +2 dB compared to a person weighing 60 kg. Note that in this example, the vibration level (amplitude) is adjusted according to weight, but various parameters for vibration adjustment, such as the vibration frequency level characteristics, may also be adjusted according to weight.

[0085] In this way, by calibrating the vibration characteristics according to the differences between the objects to which vibration is applied and the state of the objects, it is possible to improve the sense of realism provided by vibration stimulation, regardless of the object. Note that calibration in which vibration (signal) is actually applied to the object to be vibrated and the reaction is measured and based on the results can be called actual measurement type, while calibration in which the state of the object (weight, etc.) is detected and based on the detection results can be called estimation type.

[0086] Furthermore, in the example of the estimation type, weight is used as an example of the condition of the target to which vibration is to be applied, but this is not limited to this, and other conditions such as bone density, age, and sex may also be used.

[0087] <Third embodiment> Next, a third embodiment corresponding to the above (5) will be described. Fig. 17 is a block diagram of a remote device 100B according to the third embodiment. Note that Fig. 17 corresponds to Fig. 6, just like Fig. 13, and therefore differences from Fig. 6 will be mainly described here.

[0088] As shown in FIG. 17, the remote device 100B differs from the first embodiment in that it further includes a scene detection unit 103d and an extraction unit 103e.

[0089] The scene detection unit 103d detects a specific scene from the video signal and audio signal of the XR content acquired by the acquisition unit 103a. The scene detection unit 103d detects a scene, for example, when a preset time arrives. In this case, the occurrence time of the specific scene (XR content playback position time) must be specified manually in advance. This occurrence time can be specified by directly specifying the time, or by specifying the type of scene to be detected and matching the scene estimated from the scene data or image and audio data included in the XR content data with the playback position time data.

[0090] The scene detection unit 103d also detects scenes from their positional relationships with objects in the XR content. For example, this occurs when the user approaches fireworks at a certain distance. Approaching a certain distance is determined, for example, by the object (type) and its position data included in the XR content data. The scene detection unit 103d also detects scenes from changes in the situation in the XR content. For example, this occurs when the user enters a concert hall in the virtual space of the XR content. The scene detection unit 103d also detects scenes from their contact relationships with objects in the XR content. For example, this occurs when the user collides with something in the virtual space of the XR content. This collision detection is also determined, for example, by the object (type) and its position data included in the XR content data.

[0091] The vibration parameters for each scene are set in advance in the vibration parameter information 102a, and the extraction unit 103e extracts the vibration parameters according to the scene detected by the scene detection unit 103d.

[0092] Then, the sound vibration conversion processing unit 103b executes sound signal conversion processing based on the vibration parameters extracted by the extraction unit 103e.

[0093] Next, a processing procedure executed by the remote device 100B will be described with reference to Fig. 18. Fig. 18 is a flowchart showing a processing procedure executed by the remote device 100B according to the third embodiment.

[0094] 18, in the third embodiment, the scene detection unit 103d detects a scene based on a video signal, an audio signal, etc. of the XR content (step S301). Also, the sound vibration conversion processing unit 103b inputs an audio signal (step S302).

[0095] Then, it is determined whether the scene detected by the scene detection unit 103d is a scene to be divided (whether it is a scene to be subjected to vibration emphasis processing) (step S303). If it is a scene to be divided (step S303, Yes), frequency division is performed (step S304), and the divided audio signal is output to the vibration output unit 130 as a vibration signal (step S305). Then, the processing ends.

[0096] On the other hand, if the scene is not a target for frequency division (No at step S303), the audio signal is output as it is to the vibration output unit 130 as a vibration signal (step S305), and the process then ends.

[0097] As described above, the remote devices 100, 100A, and 100B are information processing devices having a control unit 103 that generates a vibration stimulation signal to be given to the user based on an audio signal in the content. The control unit 103 acquires data of the XR content (corresponding to an example of "content") that includes the audio signal, analyzes the audio signal, and generates a vibration stimulation signal to be given to the user by converting the audio signal according to the results of the analysis process.

[0098] Therefore, the remote devices 100, 100A, and 100B can further improve the sense of realism by providing appropriate vibration stimulation based on the analysis results during playback of XR content.

[0099] The conversion process is a process of emphasizing the low frequency band in the vibration stimulation signal according to the result of the analysis process.

[0100] Therefore, according to the remote devices 100, 100A, and 100B, the low frequency range of the vibration stimulus can be emphasized to create an appropriate state when playing back XR content, thereby further improving the sense of realism.

[0101] The emphasis process is a frequency division process of the audio signal used in the conversion process.

[0102] Therefore, according to the remote devices 100, 100A, and 100B, by dividing the frequency of the audio signal, the low frequency range of the vibration stimulus when playing XR content can be emphasized to make it appropriate, thereby further improving the sense of realism.

[0103] Furthermore, the frequency division process divides the frequency of the audio signal by pitch shifting in accordance with the result of the analysis process.

[0104] Therefore, according to the remote devices 100, 100A, and 100B, by dividing the audio signal by pitch shifting, the low frequency range of the vibration stimulation when playing XR content can be emphasized appropriately according to the audio state of the XR content, thereby further improving the sense of realism.

[0105] Furthermore, the conversion process synthesizes a vibration signal that is composed of signals in a predetermined low frequency band.

[0106] Therefore, the remote devices 100, 100A, and 100B can generate vibrations that emphasize the low frequency range by a method other than pitch shifting, thereby further improving the sense of realism provided by vibration stimulation when playing XR content.

[0107] Furthermore, the control unit 103 performs the above-mentioned emphasis process when the level of a predetermined low frequency band in the audio signal is below a preset threshold value.

[0108] Therefore, according to the remote devices 100, 100A, and 100B, the level of a predetermined low-frequency band is used to determine whether low-frequency band emphasis processing is necessary in the applied vibration in the voice vibration conversion processing unit 103b, and a vibration signal is generated accordingly, making it possible to provide a moderate vibration that does not undergo excessive vibration enhancement.

[0109] Furthermore, the control unit 103 estimates the sound source using an AI inference model that estimates the sound source of the audio signal, and performs the above conversion process corresponding to the estimated sound source.

[0110] Therefore, the remote devices 100, 100A, and 100B can generate vibrations according to the inferred sound source, thereby further improving the sense of realism provided by vibration stimulation when playing XR content.

[0111] Furthermore, the control unit 103 of the remote device 100B detects a specific scene from the XR content and performs the above conversion process corresponding to the detected scene.

[0112] Therefore, the remote device 100B can generate vibrations that emphasize the low frequency range in accordance with the detected scene, thereby further improving the sense of realism provided by vibration stimulation when playing XR content.

[0113] Furthermore, the control unit 103 of the remote device 100A calibrates the conversion process in accordance with the vibration application environment.

[0114] Therefore, the remote device 100A can generate vibrations adjusted according to the situation of the target, and can improve the sense of realism caused by vibration stimulation regardless of the target.

[0115] In the above-described embodiment, the voice-to-vibration conversion process is performed on the remote device side, but it may also be performed on the local device side. In such a case, the provided XR content will include a vibration signal for providing vibration stimulation. In addition, data required for calibration, etc. will be communicated between the remote device and the local device.

[0116] Further advantages and modifications will readily occur to those skilled in the art. Therefore, the invention in its broader aspects is not limited to the specific details and representative embodiments shown and described above. Accordingly, various modifications may be made without departing from the spirit or scope of the general inventive concept as defined by the appended claims and their equivalents. [Explanation of symbols]

[0117] 1. Information Processing Systems 10 On-site equipment 11 Camera 12. Microphone 100, 100A, 100B Remote equipment 101 Communications Department 102 Storage section 102a Vibration parameter information 102b Sound source estimation model 103 Control Unit 103a Acquisition Department 103b Voice vibration conversion processing unit 103c Calibration section 103d Scene detection section 103e Extraction part 110 Video output section 120 Audio output section 130 vibration output unit 140 Accelerometer

Claims

1. An information processing device having a control unit that generates a vibration stimulus signal to be given to a user based on an audio signal in content, The audio signal in the content has a low frequency band cut by a high-pass filter, The control unit Acquire the audio signal in the content; A high frequency band in the acquired audio signal is cut by a low pass filter; When the level of the audio signal at a cutoff frequency, which is the lowest frequency in a frequency band that is not cut off by the high-pass filter, exceeds a predetermined threshold, a signal having a frequency equal to or lower than the cutoff frequency that is created in advance and in which a high-frequency band in the audio signal is cut off by a low-pass filter is amplified to generate a vibration signal; When the level of the audio signal at the cutoff frequency of the high-pass filter does not exceed a predetermined threshold, a signal obtained by cutting a high frequency band of the audio signal using a low-pass filter is used as a vibration signal. Information processing device.

2. An information processing method executed by a control unit for generating a vibration stimulus signal to be given to a user based on an audio signal in content, The audio signal in the content has a low frequency band cut by a high-pass filter, Acquire the audio signal in the content; A high frequency band in the acquired audio signal is cut by a low pass filter; When the level of the audio signal at a cutoff frequency, which is the lowest frequency in a frequency band that is not cut off by the high-pass filter, exceeds a predetermined threshold, a signal having a frequency equal to or lower than the cutoff frequency that is created in advance and in which a high-frequency band in the audio signal is cut off by a low-pass filter is amplified to generate a vibration signal; When the level of the audio signal at the cutoff frequency of the high-pass filter does not exceed a predetermined threshold, a signal obtained by cutting a high frequency band of the audio signal using a low-pass filter is used as a vibration signal. Information processing methods.

3. A program for generating a vibration stimulus signal to be given to a user based on an audio signal in content, The audio signal in the content has a low frequency band cut by a high-pass filter, acquiring the audio signal in the content; a step of cutting high frequency bands in the acquired audio signal using a low pass filter; When the level of the audio signal at a cutoff frequency, which is the lowest frequency in a frequency band not cut off by the high-pass filter, exceeds a predetermined threshold, a signal having frequencies equal to or lower than the cutoff frequency, in which a high-frequency band in the audio signal is cut off by a low-pass filter, is amplified to generate a vibration signal; a step of using a low-pass filter to cut off a high frequency band of the audio signal, when the level of the audio signal at the cutoff frequency of the high-pass filter does not exceed a predetermined threshold, as a vibration signal; An information processing program executed by a control unit including the

Citation Information

Patent Citations

  • Vibration waveform signal output device

    JP2002078066A

  • Device and method for reproducing physical feeling of vibration

    JP2007324829A

  • Vibration output device and program for outputting vibration

    JP2020198591A

  • Production control device, direction system, and program

    WO2018199115A1

  • Information processing device, information processing method, and recording medium

    WO2020054415A1