Tactile presentation device and program

JP7900968B2Active Publication Date: 2026-08-05NIPPON HOSO KYOKAI
View PDF 15 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NIPPON HOSO KYOKAI
Filing Date
2022-07-28
Publication Date
2026-08-05

AI Technical Summary

Benefits of technology

【0029】 以上のように、本発明によれば、視聴者が一人称視点映像を視聴する際に、没入感向上に寄与する触覚刺激を提示するための情報を生成することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007900968000001
    Figure 0007900968000001
  • Figure 0007900968000002
    Figure 0007900968000002
  • Figure 0007900968000003
    Figure 0007900968000003
Patent Text Reader

Abstract

To generate information for presenting haptic stimulus that contributes to improvement of immersiveness when a viewer views first-person view video.SOLUTION: A self-speed estimation unit 11 of a haptic presentation device 1 acquires a plurality of frames by sampling first-person view video E, generates a mask image by masking a predetermined object, estimates, using NN31, a translation vector on the basis of a predetermined number of frames and the same number of mask images corresponding thereto, and estimates a self speed v. A sound volume control unit 12 extracts a maximum speed vmax from the self speed v of all the frames, extracts a low-frequency sound signal S from the first-person view video E, and reduces, when the self-speed v is equal to or lower than a predetermined threshold, on the basis of the maximum speed vmax and the self-speed v, volume A of the low-frequency sound signal S corresponding to a frame of the self-speed v, to generate a new low-frequency sound signal S'. A haptic presentation unit 13 outputs the low-frequency sound signal S' to a haptic device 7.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to a technology for generating information to present tactile stimuli to viewers of first-person perspective video via a haptic device. [Background technology]

[0002] Traditionally, video has been a medium primarily focused on sight and hearing, but technologies that present tactile stimuli as a third sense are gaining attention. For example, there are tactile sound systems that incorporate mechanisms in chairs to present tactile stimuli synchronized with music, and tactile devices that convert sound into vibrations are also known.

[0003] Specifically, this tactile sound system incorporates vibrators into a chair, extracts low-frequency components from music, and converts these low-frequency components into tactile vibration information using the vibrators, thereby presenting tactile vibration stimuli to the music listener (see, for example, Patent Document 1).

[0004] In addition to tactile sound systems that incorporate vibrators into chairs, technologies are known that present tactile stimuli such as vibrations and somatosensory stimuli such as the sense of movement, in addition to conventional video and audio, in theme parks, movie theaters, etc. Furthermore, technologies are known that use broadcast-communication linked services to transmit recorded tactile information of vibrations via communication in addition to the video and audio of television broadcasts.

[0005] Furthermore, in game devices that use visual displays, there are known technologies that provide players with a tactile sensation in addition to the visual display (see, for example, Patent Document 2). Specifically, this game device outputs a signal amplified by a high-power amplifier to a low-frequency speaker at a specific timing of the visual display, thereby presenting the player with a tactile sensation of a low-frequency sound source via the low-frequency speaker.

[0006] Furthermore, in tactile sound systems, there is a known technology that does not cause discomfort or a feeling of pressure to the viewer even when used for a long period of time (see, for example, Patent Document 3). Specifically, this tactile sound system includes a seat having a backrest and a seat, a band division circuit that band-divides an input audio signal and outputs a first audio signal and a second audio signal, a first vibrating element arranged in the backrest that vibrates in accordance with the first audio signal and whose vibration direction is parallel to the user-side surface of the backrest, and a second vibrating element arranged in the seat that vibrates in accordance with the second audio signal and whose vibration direction is parallel to the user-side surface of the seat.

[0007] In this way, by stimulating the third sense, touch, in addition to sight and hearing, while watching a video, it is possible to achieve a more immersive and realistic viewing experience. In other words, by inputting audio signals, converting them into tactile information, and continuously presenting tactile stimuli, it is possible to enhance the sense of immersion and realism in video content.

[0008] Attempts to convert such audio signals into tactile information and present tactile stimuli to viewers have been made for some time. Hereafter, the method of inputting audio signals, converting them into tactile information, and presenting tactile stimuli will be referred to as the "audio input method."

[0009] An example of this "voice input method" is a chair-type haptic feedback system. This chair-type haptic feedback system works by linking a chair-type haptic device to a first-person perspective video of a vehicle, such as a tram, displayed with a 180-degree field of view using a flexible display. The device converts audio signals into haptic information and presents tactile stimuli. As a result, a high level of immersion can be achieved through visual and auditory stimuli received from the first-person perspective video, as well as tactile stimuli received from the seat and footwell.

[0010] On the other hand, an apparatus for estimating the speed of a vehicle based on an image taken from the outside of a vehicle such as a tram is known (see, for example, Patent Document 4). This speed estimation apparatus extracts feature points on a subject around the vehicle from past images and current images, respectively, estimates the translational azimuth angle of the vehicle based on the feature points, and estimates the speed of the vehicle based on the translational azimuth angle.

Prior Art Documents

Patent Documents

[0011]

Patent Document 1

Patent Document 2

Patent Document 3

Patent Document 4

Summary of the Invention

Problems to be Solved by the Invention

[0012] The voice input method using the above-described somatosensory sound system converts a voice signal into tactile information and presents a tactile stimulus. In this voice input method, since it is based on actual voice, a tactile stimulus without a sense of incongruity can be presented for video voice.

[0013] However, actual voice often includes background sounds such as environmental sounds and BGM, and using such voice causes extra vibrations.

[0014] For example, in a first-person perspective video such as an in-vehicle camera video taken from the outside of a vehicle such as a tram, even when the vehicle is moving slowly or stopped, tactile stimuli may continue to be presented due to background sounds such as environmental sounds and BGM, resulting in a problem of a mismatch between the vision by the video and the touch by the vibration.

[0015] Here, for example, in order to determine whether a vehicle is moving at a low speed or stopped, it is conceivable to use the aforementioned speed estimation device. This speed estimation device estimates the speed of the vehicle (the self-speed of the person riding in the vehicle) by detecting feature points on the subject in the video. For this reason, there is a problem that when the subject on which the feature points are detected moves independently of the vehicle, the speed of the vehicle may not be accurately estimated.

[0016] Thus, in the voice input method using the immersive audio system, in order to further improve the sense of immersion, it has been desired to perform vibration control linked to the content of the video even when the vehicle is moving at a low speed or stopped. Also, when estimating the vehicle speed (the self-speed of the person riding in the vehicle), it has been desired to achieve highly accurate speed estimation.

[0017] Therefore, the present invention has been made to solve the above problems, and an object thereof is to provide a tactile presentation device and a program that generate information for presenting a tactile stimulus that contributes to improving the sense of immersion when a viewer watches a first-person perspective video.

Means for Solving the Problems

[0019] In order to solve the above problems, the tactile presentation device according to claim 1 inputs a first-person perspective video captured by moving the camera with the position of the camera as the position of the viewer, extracts a low-frequency sound signal from the first-person perspective video, and based on the low-frequency sound signal, in a tactile presentation device that generates information for presenting a tactile stimulus to the viewer via a tactile device, The speed corresponding to each of the multiple time-series frames in the aforementioned first-person perspective video, wherein the viewer's movement speed accompanying the camera's movement is considered as its own speed. Assuming that it is preset, a volume control unit that extracts the low-frequency sound signal from the first-person perspective video and reduces the volume of the low-frequency sound signal corresponding to the frame when the self-speed in the frame is below a predetermined threshold, and a tactile presentation unit that outputs the low-frequency sound signal whose volume has been reduced by the volume control unit to the tactile device, characterized by comprising.

[0020] Furthermore, the haptic presentation device of claim 2 is a haptic presentation device that inputs a first-person viewpoint video captured by the camera moving with the camera's position as the viewer's position, extracts a low-frequency audio signal from the first-person viewpoint video, and generates information for presenting haptic stimuli to the viewer via a haptic device based on the low-frequency audio signal, wherein for each of the multiple time-series frames included in the first-person viewpoint video, the viewer's movement speed in that frame is taken as the self-speed, and for each of the multiple frames, The aforementioned movement is independent of its own speed. The device is characterized by comprising: a self-velocity estimation unit that generates a mask image with an object masked, estimates the viewer's translation vector in the first-person view video based on a predetermined number of consecutive frames and the same number of mask images corresponding to the predetermined number of frames using a predetermined neural network, and calculates the viewer's own velocity in the frame based on the translation vector; a volume control unit that extracts the low-frequency audio signal from the first-person view video, and if the viewer's own velocity in the frame calculated by the self-velocity estimation unit is below a predetermined threshold, lowers the volume of the low-frequency audio signal corresponding to the frame; and a haptic presentation unit that outputs the low-frequency audio signal, whose volume has been lowered by the volume control unit, to the haptic device.

[0021] Furthermore, the haptic presentation device of claim 3 is the haptic presentation device of claim 1 or 2, wherein the volume control unit includes a low-frequency sound extraction unit that extracts the low-frequency sound signal, the video signal, and sound signals other than the low-frequency sound signal from the first-person viewpoint video, and when the self-velocity in the frame is less than or equal to the predetermined threshold, the volume of the low-frequency sound signal extracted by the low-frequency sound extraction unit corresponding to the frame is reduced to generate a new low-frequency sound signal, and when the self-velocity in the frame is greater than the predetermined threshold, The haptic presentation unit is characterized by comprising: a volume reduction control unit that sets the low-frequency audio signal corresponding to the frame as the new low-frequency audio signal; a synthesis unit that synthesizes the new low-frequency audio signal generated or set by the volume reduction control unit, the video signal extracted by the low-frequency audio extraction unit, and audio signals other than the low-frequency audio signal to obtain a volume-controlled video; and the haptic presentation unit extracts the new low-frequency audio signal from the volume-controlled video obtained by the synthesis unit and outputs the new low-frequency audio signal to the haptic device.

[0022] Furthermore, the tactile presentation device of claim 4 is the tactile presentation device of claim 3, wherein the volume control unit further comprises the plurality of frames So The system includes a maximum speed extraction unit that extracts the maximum speed from each of the aforementioned self-speeds, and the volume reduction control unit controls the maximum speed extracted by the maximum speed extraction unit v max Let v be the self-velocity in the frame, let A be the volume of the low-frequency audio signal, and let A be the volume of the new low-frequency audio signal. new As such, if the self-velocity in the frame is less than or equal to a predetermined threshold, then the following formula: A new =(v / v max )A reduces the volume of the low-frequency audio signal and generates the new low-frequency audio signal.

[0023] Furthermore, the tactile presentation device of claim 5 is the tactile presentation device of claim 3, wherein the volume reduction control unit sets a preset maximum speed v maxLet v be the self-velocity in the frame, let A be the volume of the low-frequency audio signal, and let A be the volume of the new low-frequency audio signal. new As such, if the self-velocity in the frame is less than or equal to a predetermined threshold, then the following formula: A new =(v / v max )A is used to reduce the volume of the low-frequency audio signal and generate the new low-frequency audio signal.

[0024] Furthermore, the haptic presentation device of claim 6 is characterized in that, in the haptic presentation device of claim 2, the self-velocity estimation unit trims each of the plurality of frames within a predetermined fixed coordinate frame, generates a mask image for each of the plurality of frames after trimming, estimates the translation vector based on a predetermined number of frames and the same number of mask images corresponding to the predetermined number of frames using the predetermined NN, and calculates the self-velocity in the frame based on the translation vector.

[0025] Furthermore, the program of claim 7 includes a computer that constitutes a haptic presentation device, which inputs first-person perspective video captured by moving the camera with the camera's position as the viewer's position, extracts a low-frequency audio signal from the first-person perspective video, and generates information for presenting haptic stimuli to the viewer via a haptic device based on the low-frequency audio signal. The speed corresponding to each of the multiple time-series frames in the aforementioned first-person perspective video, wherein the viewer's movement speed accompanying the camera's movement is considered as its own speed. The system is characterized by the following functions: a volume control unit that extracts the low-frequency audio signal from the first-person perspective video, assuming it is pre-configured, and, if the self-velocity in the frame is below a predetermined threshold, reduces the volume of the low-frequency audio signal corresponding to the frame; and a haptic presentation unit that outputs the low-frequency audio signal, whose volume has been reduced by the volume control unit, to the haptic device. [Effects of the Invention]

[0029] As described above, according to the present invention, when a viewer watches first-person perspective video, it is possible to generate information for presenting tactile stimuli that contribute to improving immersion. [Brief explanation of the drawing]

[0030] [Figure 1] This is a block diagram showing an example configuration of the first tactile presentation device. [Figure 2] Figure 1 is a flowchart illustrating an example of processing by the tactile presentation device. [Figure 3] This is a block diagram showing an example configuration of the self-speed estimation unit. [Figure 4] Figure 3 is a flowchart showing an example of processing by the self-velocity estimation unit. [Figure 5] This is a block diagram showing an example configuration of the self-speed estimation processing unit. [Figure 6] This is a block diagram showing another example configuration of the self-speed estimation unit. [Figure 7] This is a block diagram showing an example configuration of a volume control unit. [Figure 8] Figure 7 is a flowchart showing an example of processing by the volume control unit. [Figure 9] This figure shows an example configuration of the haptic feedback unit when playing a first-person perspective video E in 5.1ch format. [Figure 10] This figure shows an example of a frame and a mask image from first-person perspective video E. [Figure 11] This figure shows the estimated results of the self-velocity v. [Figure 12] This figure shows an example of a frame from first-person perspective video E1 and an example of a frame after cropping. [Figure 13] This is a block diagram showing an example configuration of a second haptic presentation device. [Figure 14] This is a block diagram showing an example configuration of the first self-velocity estimation device. [Figure 15] This is a block diagram showing an example configuration of the second self-velocity estimation device. [Modes for carrying out the invention]

[0031] The embodiments for carrying out the present invention will be described in detail below with reference to the drawings. [Haptic presentation device / First example] First, let's describe the first haptic presentation device. Figure 1 is a block diagram showing an example of the configuration of the first haptic presentation device, and Figure 2 is a flowchart showing an example of the processing of the haptic presentation device shown in Figure 1.

[0032] This haptic presentation device 1 comprises a self-velocity estimation unit 11, a volume control unit 12, and a haptic presentation unit 13. The haptic presentation device 1 estimates its own velocity v from a first-person perspective video E, controls the volume A of the low-frequency audio signal S that is the source of vibration based on the self-velocity v to generate a volume-controlled video E', and extracts the controlled low-frequency audio signal S' from the volume-controlled video E' and outputs it to the haptic device 7. This improves the viewer's sense of immersion when viewing the first-person perspective video E.

[0033] The self-velocity estimation unit 11 receives a first-person view video E as input (step S201), samples the first-person view video E, detects a predetermined object for each of the multiple time-series frames contained in the first-person view video E, and generates a mask image with the predetermined object masked (step S202).

[0034] The multiple time-series frames included in the first-person perspective video E may be all the frames that make up the first-person perspective video E, or they may be a group of frames sampled at predetermined intervals. The first-person perspective video E is video footage captured by a camera, with the camera's position being the position of the viewer watching the first-person perspective video E, and is video footage captured as the camera moves.

[0035] The self-velocity estimation unit 11 estimates the viewer's translation vector in the first-person view video E for each of the multiple frames, using a neural network (NN) 31 (described later), based on a predetermined number of consecutive frames and the same number of mask images corresponding to those predetermined frames, and estimates the self-velocity v based on the translation vector (step S203). The self-velocity estimation unit 11 then outputs the self-velocity v for each sampled frame to the volume control unit 12.

[0036] Here, the self-velocity v is the relative speed (viewer's movement speed in the first-person perspective video E) due to the movement of the camera (viewer's viewpoint) that filmed the first-person perspective video E. For example, if the first-person perspective video E is footage from a tram's onboard camera, the self-velocity v = 0 when the tram is completely stopped, and the self-velocity v > 0 when the tram is moving.

[0037] This allows the self-velocity v to be calculated for each of the multiple time-series frames contained in the first-person perspective video E, from the first frame to the last frame. Details of the self-velocity estimation unit 11 will be described later.

[0038] The volume control unit 12 receives the self-velocity v for each frame sampled from the self-velocity estimation unit 11 and stores the self-velocity v in the memory 41, which will be described later (step S204). As a result, the self-velocity v for each of the multiple frames included in the first-person view video E is stored in the memory 41.

[0039] The volume control unit 12 determines, in accordance with the viewer's operation, whether or not an operation to start viewing the first-person view video E (the same video as the first-person view video E input by the self-velocity estimation unit 11) has been performed (step S205). If the volume control unit 12 determines in step S205 that there has been no operation to start viewing (step S205:N), it waits until such operation is performed.

[0040] If the volume control unit 12 determines in step S205 that a viewing start operation has been performed (step S205:Y), it inputs the first-person view video E (step S206). The volume control unit 12 then extracts the low-frequency audio signal S, the video signal, and audio signals other than the low-frequency audio signal S from the first-person view video E (step S207).

[0041] For example, if the first-person view video E contains an audio signal of a low-frequency channel, the volume control unit 12 extracts the audio signal of that channel from the first-person view video E, thereby extracting the audio signal of that channel as a low-frequency audio signal S. Also, if the first-person view video E consists of a video signal and an audio signal, and the audio signal contains both high-frequency and low-frequency components, the volume control unit 12 extracts the low-frequency component from the audio signal contained in the first-person view video E, thereby extracting the low-frequency component as a low-frequency audio signal S.

[0042] The volume control unit 12 reads the self-velocity v of the video signal frame corresponding to the low-frequency audio signal S of the first-person view video E extracted in step S207 from the memory 41 (step S208). As a result, the self-velocity v stored in the memory 41 are read out in order, corresponding to the low-frequency audio signal S.

[0043] The volume control unit 12, when reading a self-velocity v from the memory 41, generates a new low-frequency audio signal S' by lowering the volume A of the low-frequency audio signal S corresponding to the frame of the self-velocity v if the self-velocity v is below a predetermined threshold (step S209). Here, the low-frequency audio signal S corresponding to the self-velocity v is the audio signal from the frame of the self-velocity v up to just before the next frame (the next frame among the multiple frames stored in the memory 41).

[0044] The volume control unit 12 obtains a volume-controlled video E' by combining the low-frequency audio signal S' generated in step S209, the video signal extracted in step S207, and audio signals other than the low-frequency audio signal S (step S210). The volume control unit 12 then outputs the volume-controlled video E' to the tactile presentation unit 13. Details of the volume control unit 12 will be described later.

[0045] The haptic presentation unit 13 receives a volume-controlled video E' from the volume control unit 12, extracts a low-frequency audio signal S' from the volume-controlled video E', and outputs the low-frequency audio signal S' to the haptic device 7 (step S211). Details of the haptic presentation unit 13 will be described later.

[0046] As a result, when the self-velocity v is below a predetermined threshold, a low-frequency audio signal S' with reduced volume A is input to the tactile device 7, thereby reducing vibration. In other words, when the self-velocity v is 0 or small, tactile stimulation caused by vibrations reflecting ambient sounds and background sounds such as BGM can be suppressed.

[0047] Therefore, when viewing the first-person perspective video E, the haptic presentation device 1 can generate information to present haptic stimuli that contribute to improving immersion, and the viewer can receive vibration stimuli with reduced influence from background sound when their own speed v is slow or stopped, thereby improving immersion.

[0048] (Self-speed estimation unit 11) Next, the self-velocity estimation unit 11 shown in Figure 1 will be described in detail. Figure 3 is a block diagram showing an example of the configuration of the self-velocity estimation unit 11, and Figure 4 is a flowchart showing an example of the processing of the self-velocity estimation unit 11 shown in Figure 3.

[0049] This self-velocity estimation unit 11-1 includes a frame sampling processing unit 21, a mask image generation processing unit 22, and a self-velocity estimation processing unit 23.

[0050] The frame sampling processing unit 21 receives first-person view video E as input (step S401) and samples the first-person view video E into multiple time-series frames at predetermined intervals (step S402). By sampling at predetermined intervals, the computational load of subsequent processing can be reduced. Here, the frame sampling processing unit 21 may sample all frames that make up the first-person view video E.

[0051] The frame sampling processing unit 21 outputs each of the multiple frames after sampling (frames 0, ..., n, ..., N) to the mask image generation processing unit 22 and the self-velocity estimation processing unit 23. N is an integer greater than or equal to 1, and n is 0 ≤ n ≤ N. Frame n indicates the frame with frame number n.

[0052] The mask image generation processing unit 22 receives each of the multiple frames from the frame sampling processing unit 21 and generates a mask image for each of the multiple frames that masks a predetermined object (step S403). The mask image generation processing unit 22 then outputs the mask image to the self-velocity estimation processing unit 23.

[0053] This generates N mask images from N frames (mask image 0 is generated from frame 0, ..., mask image n is generated from frame n, ..., mask image N is generated from frame N).

[0054] Specifically, the mask image generation processing unit 22 generates a mask image by applying a common instance segmentation method such as Mask R-CNN to each of multiple frames and specifying the type of object to be masked. The mask image is an image in which pre-defined objects (objects that move independently of the self-velocity v (moving objects), i.e., oncoming cars, pedestrians, etc., which may affect (become noise) when calculating the self-velocity v) are masked. For example, cars and people are specified as the types of objects to be masked. Objects that do not move, such as signs and trees, do not affect the calculation of the self-velocity v and therefore do not need to be specified.

[0055] Figure 10 shows an example of a frame from first-person perspective video E and an example of a mask image. The left side of Figure 10 shows an example of a frame from first-person perspective video E, and the right side shows an example of a mask image generated from the frame shown on the left.

[0056] The mask image generation processing unit 22 detects multiple oncoming vehicles from the frame shown on the left, and generates a mask image in which each of these multiple oncoming vehicles is represented by a distinct shape.

[0057] Returning to Figures 3 and 4, the self-velocity estimation processing unit 23 receives a frame from the frame sampling processing unit 21 and a mask image corresponding to that frame from the mask image generation processing unit 22.

[0058] The self-velocity estimation processing unit 23 uses the NN31 to estimate translation vectors translation(x,y,z), etc., from a predetermined number (e.g., 3) frames and the same number of corresponding mask images, in order to exclude objects masked by the mask images (step S404). The translation vector translation(x,y,z) is data indicating the direction of the viewer's movement in a predetermined number of frames.

[0059] As a result, by using a mask image, oncoming vehicles, pedestrians, etc., which may affect the estimation of the translation vector translation(x,y,z) used to calculate the self-velocity v, are excluded, and a highly accurate translation vector translation(x,y,z) can be obtained. Therefore, in step S405 described later, a highly accurate self-velocity v can be calculated.

[0060] The self-velocity estimation processing unit 23 calculates the self-velocity v based on the translation vector translation(x,y,z) (step S405). As a result, the self-velocity v is calculated for each of the multiple frames input from the frame sampling processing unit 21. The self-velocity estimation processing unit 23 then outputs the self-velocity v to the volume control unit 12 (step S406).

[0061] Figure 5 is a block diagram showing an example configuration of the self-velocity estimation processing unit 23. This self-velocity estimation processing unit 23 includes NN31 and a self-velocity calculation unit 32.

[0062] NN31 receives, for example, three consecutive frames from the frame sampling processing unit 21 and three mask images corresponding to the three frames from the mask image generation processing unit 22. Then, NN31 performs the operations of NN31, estimates the translation vector translation(x, y, z) and the depth image, and outputs the translation vector translation(x, y, z) to the self-speed calculation unit 32.

[0063] For example, from three frames 0, 1, 2 out of the frames 0, ···, n, ···, N and three mask images 0, ···, n, ···, N corresponding thereto, the translation vector translation(x, y, z) for frame 1 is estimated. Then, the translation vector translation(x, y, z) for frame 2 is estimated from three frames 1, 2, 3 and three mask images 1, 2, 3, ···, and the translation vector translation(x, y, z) for frame N - 1 is estimated from three frames N - 2, N - 1, N and three mask images N - 2, N - 1, N.

[0064] The self-speed calculation unit 32 receives the translation vector translation(x, y, z) from NN31 and calculates the self-speed v by the following formula. [Equation 1] v = √(x 2 + y 2 + z 2 ) ···(1)

[0065] Note that NN31 used for estimating the self-speed v is not limited to a specific network configuration. For example, a neural network having the same configuration as the monocular camera depth estimation model shown in the following literature for estimating the translation vector translation(x, y, z), or an improved one based on these architectures is used. [Non-patent literature] Vincent Casser, Soeren Pirk Reza, Mahjourian, Anelia Angelova, “Depth Prediction Without the Sensors: Leveraging Structure for Unsupervised Learning from Monocular Videos.”, the AAAI Conference on Artificial Intelligence, Vol. 33, pp. 8001-8008 (2019)

[0066] As a result, the self-velocity estimation unit 11-1 estimates the self-velocity v, which represents the viewer's movement speed in the first-person view video E, for each of the multiple time-series frames sampled at predetermined intervals from the first-person view video E.

[0067] Figure 11 shows the estimated self-velocity v, and represents the self-velocity v estimated by the self-velocity estimation unit 11-1. The vertical axis represents the self-velocity v, and the horizontal axis represents time (frame number).

[0068] The self-velocity estimation unit 11-1 estimates the self-velocity v shown in Figure 11, and the self-velocity v is output to the subsequent volume control unit 12.

[0069] (Another example of the self-speed estimation unit 11) Figure 6 is a block diagram showing another example configuration of the self-velocity estimation unit 11 shown in Figure 1. This self-velocity estimation unit 11-2 includes a frame sampling processing unit 21, a mask image generation processing unit 22, a self-velocity estimation processing unit 23, and a trimming processing unit 24.

[0070] Comparing the self-velocity estimation unit 11-1 shown in Figure 3 with this self-velocity estimation unit 11-2, both self-velocity estimation units 11-1 and 11-2 are common in that they are equipped with a frame sampling processing unit 21, a mask image generation processing unit 22, and a self-velocity estimation processing unit 23. However, self-velocity estimation unit 11-2 differs from self-velocity estimation unit 11-1 in that it is further equipped with a trimming processing unit 24. In Figure 6, parts common with Figure 3 are denoted by the same reference numerals as in Figure 3, and their detailed explanations are omitted.

[0071] The self-velocity estimation unit 11-2 receives a first-person view video E1 (video containing images that may adversely affect the self-velocity v estimation process) that is different from the first-person view video E (first-person view video E shown in Figure 10) that is input to the self-velocity estimation unit 11-1.

[0072] The frame sampling processing unit 21 receives the first-person perspective video E1 as input, performs the same processing as the frame sampling processing unit 21 shown in Figure 3, and outputs multiple time-series frames sampled at predetermined intervals to the trimming processing unit 24.

[0073] The trimming processing unit 24 receives multiple frames from the frame sampling processing unit 21 and generates a trimmed frame by trimming each of the multiple frames within a predetermined fixed coordinate frame. The trimming processing unit 24 then outputs the trimmed frame to the mask image generation processing unit 22 and the self-velocity estimation processing unit 23.

[0074] The pre-defined fixed coordinate frame is specified within the first-person perspective video E1 by the coordinates of the top-left corner of the image to be trimmed, the width in the x-axis direction, and the height in the y-axis direction. The trimming process can be performed using existing video editing software.

[0075] For example, if the first-person view video E1 is footage from a tram's onboard camera, too much of the tram's interior is likely to negatively affect the estimation of its own speed v. Therefore, the trimming processing unit 24 trims the frames of the first-person view video E1 using a pre-set fixed coordinate frame so that the image mainly shows what is outside the window.

[0076] Furthermore, the pre-set fixed coordinate frame will not be changed for all frames input by the trimming processing unit 24.

[0077] Figure 12 shows an example of a frame from first-person perspective video E1 and an example of a frame after cropping.

[0078] The trimming processing unit 24 extracts the trimmed frame from the first-person view video E1, as shown in Figure 12. This frame of the first-person view video E1 contains too much of the inside of the tram, and therefore includes objects (such as the vehicle frame) that should be excluded when considering the accuracy of the self-velocity v estimation. For this reason, trimming is performed using a pre-set fixed coordinate frame as shown in Figure 12, and the trimmed frame that takes into account the accuracy of the self-velocity v estimation is extracted.

[0079] Returning to Figure 6, the mask image generation processing unit 22 receives the trimmed frame from the trimming processing unit 24 and performs the same processing as the mask image generation processing unit 22 shown in Figure 3.

[0080] The self-velocity estimation processing unit 23 receives a predetermined number (for example, 3) of trimmed frames from the trimming processing unit 24, and the same number of mask images corresponding to the trimmed frames from the mask image generation processing unit 22. Then, the self-velocity estimation processing unit 23 performs the same processing as the self-velocity estimation processing unit 23 shown in Figure 3, and outputs the self-velocity v to the volume control unit 12.

[0081] As a result, the self-velocity estimation unit 11-2 trims each of the multiple time-series frames sampled at predetermined intervals from the first-person view video E1, and estimates the self-velocity v, which represents the viewer's movement speed in the first-person view video E1, for the trimmed frames. In other words, by trimming, regions that are likely to negatively affect the estimation of self-velocity v can be excluded from the frame, thus enabling the acquisition of a more accurate self-velocity v.

[0082] (Volume control unit 12) Next, the volume control unit 12 shown in Figure 1 will be described in detail. Figure 7 is a block diagram showing an example configuration of the volume control unit 12, and Figure 8 is a flowchart showing an example of processing by the volume control unit 12 shown in Figure 7.

[0083] This volume control unit 12 includes a memory 41, a maximum speed extraction unit 42, a low-frequency sound extraction unit 43, a volume reduction control unit 44, and a synthesis unit 45.

[0084] The volume control unit 12 receives the self-velocity v for each frame sampled from the self-velocity estimation unit 11 and stores the self-velocity v in the memory 41 (step S801). As a result, the self-velocity v of multiple time-series frames (all frames) sampled from the first-person perspective video E is stored in the memory 41.

[0085] The maximum velocity extraction unit 42 reads the self-velocities v of all sampled frames from memory 41 once the self-velocities v of all sampled frames have been stored in memory 41, and extracts the maximum velocity v from these self-velocities v. max Extracts (step S802). Then, the maximum speed extraction unit 42 extracts the maximum speed v max This is output to the volume reduction control unit 44.

[0086] The volume control unit 12 determines, in accordance with the viewer's operation, whether or not an operation to start viewing the first-person view video E (the same video as the first-person view video E input by the self-velocity estimation unit 11 shown in Figure 1) has been performed (step S803). If the volume control unit 12 determines in step S803 that no operation to start viewing has been performed (step S803:N), it waits until such operation is performed.

[0087] If the volume control unit 12 determines in step S803 that a viewing start operation has been performed (step S803:Y), the low-frequency sound extraction unit 43 receives the first-person view video E as input and extracts the low-frequency sound signal S, the video signal, and sound signals other than the low-frequency sound signal S from the first-person view video E (step S804).

[0088] In this case, the low-frequency audio extraction unit 43 may extract the audio signal of a sound source that has been pre-recorded with only low frequencies as the low-frequency audio signal S.

[0089] For example, in a first-person perspective video E consisting of an 8K video signal and a 22.2ch audio signal, a low-frequency sound below 120Hz called an LFE (Low Frequency Effect) channel is used in channel 0.2. In this case, the low-frequency sound extraction unit 43 extracts this LFE audio signal as a low-frequency audio signal S, and by utilizing the LFE audio signal in the subsequent haptic device 7, the viewer can obtain haptic stimuli suitable for vibration.

[0090] Furthermore, if a low-frequency audio signal such as an LFE cannot be prepared, a mixed audio signal may be used. In other words, the low-frequency audio extraction unit 43 can artificially generate a low-frequency audio signal S by equalizing the mixed audio signal to emphasize the low-frequency components and suppress the high-frequency components.

[0091] The low-frequency audio extraction unit 43 outputs the low-frequency audio signal S to the volume reduction control unit 44, and outputs the video signal and audio signals other than the low-frequency audio signal S (other audio signals) to the synthesis unit 45.

[0092] The volume reduction control unit 44 receives the maximum speed v from the maximum speed extraction unit 42. max Along with inputting the signal, the low-frequency audio signal S is input from the low-frequency audio extraction unit 43. The volume reduction control unit 44 reads out the self-velocity v from the memory 41 for all sampled frames in order from the beginning (step S805). The volume reduction control unit 44 then determines whether the self-velocity v read out from the memory 41 is below a predetermined (pre-set) threshold (step S806).

[0093] In step S806, if the volume reduction control unit 44 determines that its own speed v is below a threshold (step S806:Y), it controls its own speed v and maximum speed v max Based on this, the volume A of the low-frequency audio signal S of the frame corresponding to the self-velocity v is reduced to generate a new low-frequency audio signal S' (step S807), and the process proceeds to step S809. The volume reduction control unit 44 then outputs the new low-frequency audio signal S' to the synthesis unit 45.

[0094] For example, the volume reduction control unit 44 reduces the volume A of the low-frequency audio signal S of the frame corresponding to its own velocity v in proportion to the magnitude of its own velocity v, according to the following equation, and then reduces the volume A new A new low-frequency audio signal S' having the following characteristics is generated. [Math 2] A new =(v / v max )A ···(2)

[0095] Here, the low-frequency audio signal S of the frame corresponding to the self-velocity v is the audio signal from the time of the frame corresponding to the self-velocity v to the time immediately preceding the next frame.

[0096] On the other hand, if the volume reduction control unit 44 determines in step S806 that its own velocity v is not below the threshold, that is, that its own velocity v is greater than the threshold (step S806:N), it sets the low-frequency audio signal S as a new low-frequency audio signal S' without changing the volume A of the low-frequency audio signal S (step S808), and proceeds to step S809. In other words, the volume reduction control unit 44 sets the low-frequency audio signal S as a new low-frequency audio signal S' without changing the volume A of the low-frequency audio signal S. new The new low-frequency audio signal S' having =A is set. The volume reduction control unit 44 then outputs the new low-frequency audio signal S' to the synthesis unit 45.

[0097] As a result, when the self-velocity v is below the threshold, the listener perceives that the device is decelerating, stopping, accelerating, etc., at this speed, and the volume A of the low-frequency audio signal S decreases. In other words, the volume reduction control unit 44 generates a low-frequency audio signal S' with reduced volume A, corresponding to deceleration, stopping, accelerating, etc., in conjunction with the self-velocity v. On the other hand, when the self-velocity v is greater than the threshold, the volume A of the low-frequency audio signal S is changed, resulting in the same volume A. new Set to =A.

[0098] The synthesis unit 45, moving from steps S807 and S808, receives the low-frequency audio signal S' from the volume reduction control unit 44, and also receives the video signal and audio signals other than the low-frequency audio signal S from the low-frequency audio extraction unit 43. The synthesis unit 45 then synthesizes the low-frequency audio signal S', the video signal, and the audio signals other than the low-frequency audio signal S to obtain the volume-controlled video E' (step S809). The volume control unit 12 outputs the volume-controlled video E' to the tactile presentation unit 13 (step S810).

[0099] As a result, the volume control unit 12 reduces the volume A of the low-frequency audio signal S in the range where the self-velocity v in the frame of the first-person view video E is slow or stopped, thereby generating a volume-controlled video E' that includes a low-frequency audio signal S' with reduced volume A.

[0100] Furthermore, the volume control unit 12 shown in Figure 7 may also include a smoothing unit prior to the memory 41. The smoothing unit receives the self-velocity v for each frame sampled from the self-velocity estimation unit 11.

[0101] If the self-velocity v is unstable (for example, if the rate of change of the self-velocity v is greater than or equal to a predetermined value), the smoothing unit uses a predetermined number of frames in the vicinity (before and after) to smooth the self-velocity v of the frame in question. The smoothing unit then stores the smoothed self-velocity v in the memory 41.

[0102] (Tactile presentation unit 13) Next, the haptic feedback unit 13 shown in Figure 1 will be described in detail. Figure 9 shows an example configuration of the haptic feedback unit 13 when playing a first-person perspective video E in 5.1ch format. This example shows the use of three audio signals, L, R, and LFE, out of the 5.1ch format (L, R, C, SL, SR, LFE) audio signal.

[0103] This tactile presentation unit 13 includes an extraction unit 51 and an amplification unit 52. Note that in Figure 9, the configuration for amplifying the video signal and the L and R audio signals is omitted.

[0104] The extraction unit 51 receives the volume-controlled video E' from the volume control unit 12, extracts the LFE audio signal as a low-frequency audio signal S' from the volume-controlled video E', and also extracts the video signal and the L (left) and R (right) audio signals. The extraction unit 51 outputs the LFE audio signal to the amplification unit 52, which amplifies the LFE audio signal and outputs it to the haptic device 7 and the speaker 9. The extraction unit 51 also outputs the video signal to the display 8 and the L and R audio signals to the speaker 9.

[0105] The haptic device 7 receives the audio signal from the LFE from the amplification unit 52, and the volume A of the audio signal from the LFE new The smaller the value, the smaller the vibration presented to the viewer, and the volume A of the LFE audio signal. new The larger the value, the greater the vibration presented to the viewer.

[0106] As a result, viewers can view the first-person perspective video E via a haptic device 7 that receives an LFE audio signal from the haptic presentation unit 13, a display 8 that receives a video signal, and a speaker 9 that receives L, R, and LFE audio signals, and at the same time receive haptic stimuli linked to the first-person perspective video E.

[0107] In particular, in the range where the self-velocity v in the first-person perspective video E is high, the volume A of the low-frequency audio signal S remains unchanged, allowing the viewer to receive normal tactile stimulation. On the other hand, in the range where the self-velocity v in the first-person perspective video E is low or stationary, the volume A of the low-frequency audio signal S decreases, allowing the viewer to receive weaker tactile stimulation than usual.

[0108] Here, the low-frequency audio signal S', which is the audio signal of the LFE, is output to the tactile device 7, and the low-frequency audio signal S' is converted into a tactile stimulus because, generally speaking, humans can only receive tactile stimuli at low frequencies of around 200 Hz or less, and they cannot receive suitable tactile stimuli when the frequency of the audio signal is high.

[0109] Furthermore, the extraction unit 51 may extract the LFE audio signal as a low-frequency audio signal S' from the volume-controlled video E', as well as the video signal and the L (left) and R (right) audio signals. The LFE audio signal may be output to the haptic device 7 via the amplification unit 52, the video signal to the display 8, and the L (left) and R (right) audio signals to the speaker 9.

[0110] As described above, according to the first haptic presentation device 1, the self-velocity estimation unit 11 samples the first-person viewpoint video E to obtain multiple time-series frames, and generates a mask image for each of the multiple frames that masks a predetermined object. Then, the self-velocity estimation unit 11 uses NN31 to estimate a translation vector translation(x,y,z) based on a predetermined number of consecutive frames and the same number of corresponding mask images, and estimates the self-velocity v based on the translation vector translation(x,y,z).

[0111] The volume control unit 12 calculates the maximum velocity v from its own velocity v of all sampled frames. max The system extracts the audio signal S, etc., from the first-person perspective video E when the operation to start viewing is performed. Then, the volume control unit 12 sets the maximum speed v if its own speed v is below a predetermined threshold. max Based on the self-velocity v, a new low-frequency audio signal S' is generated by lowering the volume A of the low-frequency audio signal S corresponding to the frame at the self-velocity v, and a volume-controlled video E' containing the low-frequency audio signal S' is synthesized.

[0112] The haptic presentation unit 13 extracts a low-frequency audio signal S' from the volume-controlled video E' and outputs the low-frequency audio signal S' to the haptic device 7.

[0113] As a result, when the self-velocity v is below a predetermined threshold (in the range of low speed or stopped), a low-frequency audio signal S' with reduced volume A is input to the haptic device 7, thereby reducing vibration. In other words, when the self-velocity v is 0 or small, tactile stimulation caused by vibrations reflecting ambient sounds and background sounds such as BGM can be suppressed, and a matching between visual information from images and tactile sensations from vibrations can be achieved.

[0114] Therefore, when a viewer watches a first-person perspective video E, the haptic presentation device 1 can generate information to present haptic stimuli that contribute to improving immersion. The viewer can receive vibration stimuli that are less affected by background noise when their own speed v is low or stopped, thereby improving immersion.

[0115] [Haptic presentation device / Second example] Next, a second haptic presentation device will be described. Figure 13 is a block diagram showing an example configuration of the second haptic presentation device. This haptic presentation device 2 includes a volume control unit 12 and a haptic presentation unit 13.

[0116] The haptic presentation device 2 receives a preset self-velocity v of multiple time-series frames contained in the first-person perspective video E, and, based on the self-velocity v, controls the volume A of the low-frequency audio signal S that is the source of vibration to generate a volume-controlled video E'. It then extracts the controlled low-frequency audio signal S' from the volume-controlled video E' and outputs it to the haptic device 7. This improves the viewer's sense of immersion when viewing the first-person perspective video E.

[0117] Comparing the tactile presentation device 1 shown in Figure 1 with this tactile presentation device 2, both tactile presentation devices 1 and 2 are similar in that they are equipped with a volume control unit 12 and a tactile presentation unit 13, but tactile presentation device 2 differs from tactile presentation device 1 in that it does not have a self-speed estimation unit 11.

[0118] In other words, the haptic presentation device 1 shown in Figure 1 receives a first-person view video E as input in its self-velocity estimation unit 11 and estimates its own velocity v for each frame sampled from the first-person view video E. In contrast, the haptic presentation device 2 does not perform the process of estimating its own velocity v, but instead receives a pre-set self-velocity v for each frame sampled from the first-person view video E.

[0119] The volume control unit 12 and the tactile presentation unit 13 perform the same processing as the volume control unit 12 and the tactile presentation unit 13 shown in Figure 1, so a detailed explanation is omitted.

[0120] As described above, the second haptic presentation device 2, like the first haptic presentation device 1, can generate information to present haptic stimuli that contribute to improved immersion when a viewer is watching a first-person perspective video E. Viewers can receive vibration stimuli with reduced influence from background noise when their own speed v is low or stopped, thereby improving immersion.

[0121] [Self-Speed ​​Estimation Device / First Example] Next, we will describe a self-velocity estimation device. A self-velocity estimation device is a device belonging to the technical field of estimating the viewer's movement speed in a first-person perspective video E as their own velocity v. First, we will describe the first self-velocity estimation device.

[0122] Figure 14 is a block diagram showing an example configuration of the first self-velocity estimation device. This self-velocity estimation device 3 includes a frame sampling processing unit 21, a mask image generation processing unit 22, and a self-velocity estimation processing unit 23.

[0123] The self-velocity estimation device 3 generates a mask image that masks a predetermined object for each of multiple time-series frames sampled at predetermined intervals from first-person view video E. Using NN31, it estimates a translation vector translation(x,y,z) based on a predetermined number of consecutive frames and the same number of corresponding mask images, and estimates the self-velocity v based on the translation vector translation(x,y,z), thereby obtaining a highly accurate self-velocity v.

[0124] This self-velocity estimation device 3 performs the same processing under the same configuration as the self-velocity estimation unit 11-1 shown in Figure 3, and outputs the self-velocity v for each frame sampled from the first-person view video E.

[0125] The frame sampling processing unit 21, mask image generation processing unit 22, and self-velocity estimation processing unit 23 are the same as those shown in Figure 3, so a detailed explanation is omitted.

[0126] As described above, according to the first self-velocity estimation device 3, the frame sampling processing unit 21 samples multiple time-series frames from the first-person viewpoint video E at predetermined intervals, and the mask image generation processing unit 22 generates a mask image for each of the multiple frames, masking a predetermined object.

[0127] The self-velocity estimation processing unit 23 uses NN31 to estimate translation vectors translation(x,y,z) etc. from a predetermined number (e.g., 3) frames and the same number of corresponding mask images, in order to exclude objects masked by the mask images, and calculates the self-velocity v based on the translation vectors translation(x,y,z).

[0128] By using a mask image, oncoming vehicles, pedestrians, etc., which may affect the estimation of the translation vector translation(x,y,z) used to calculate the self-velocity v, are excluded. As a result, a highly accurate translation vector translation(x,y,z) can be obtained, and consequently, a highly accurate self-velocity v can be obtained.

[0129] [Self-Speed ​​Estimation Device / Second Example] Next, a second self-velocity estimation device will be described. Figure 15 is a block diagram showing an example configuration of the second self-velocity estimation device. This self-velocity estimation device 4 includes a frame sampling processing unit 21, a mask image generation processing unit 22, a self-velocity estimation processing unit 23, and a trimming processing unit 24.

[0130] The self-velocity estimation device 4 receives a first-person view video E1 as input, trims each of the multiple time-series frames sampled from the first-person view video E1 at predetermined intervals, generates a mask image by masking a predetermined object from the trimmed frames, estimates a translation vector translation(x,y,z) based on a predetermined number of consecutive trimmed frames and the same number of corresponding mask images using NN31, and estimates the self-velocity v based on the translation vector translation(x,y,z) to obtain a more accurate self-velocity v.

[0131] This self-velocity estimation device 4 performs the same processing under the same configuration as the self-velocity estimation unit 11-2 shown in Figure 6, and outputs the self-velocity v for each frame sampled and trimmed from the first-person view video E1.

[0132] The frame sampling processing unit 21, mask image generation processing unit 22, self-velocity estimation processing unit 23, and trimming processing unit 24 are the same as those shown in Figure 6, so a detailed explanation is omitted.

[0133] As described above, according to the second self-velocity estimation device 4, the frame sampling processing unit 21 samples multiple time-series frames from the first-person viewpoint video E1 at predetermined intervals, the trimming processing unit 24 performs trimming on each of the multiple frames, and the mask image generation processing unit 22 generates a mask image that masks a predetermined object for each of the multiple frames after trimming.

[0134] The self-velocity estimation processing unit 23 uses NN31 to estimate translation vectors translation(x,y,z) etc. from a predetermined number (e.g., 3) frames and the same number of corresponding mask images, in order to exclude objects masked by the mask images, and calculates the self-velocity v based on the translation vectors translation(x,y,z).

[0135] As a result, similar to the self-velocity estimation device 3, by using a mask image, oncoming vehicles, pedestrians, etc., which may affect the estimation of the translation vector translation(x,y,z) used to calculate the self-velocity v, can be excluded. Therefore, a highly accurate translation vector translation(x,y,z) can be obtained, and as a result, a highly accurate self-velocity v can be obtained.

[0136] Furthermore, the trimming processing unit 24 can remove regions from the frame that are likely to negatively affect the estimation of the self-velocity v, thereby enabling the acquisition of a more accurate self-velocity v.

[0137] Although the present invention has been described above with reference to embodiments, the present invention is not limited to the above embodiments and can be modified in various ways without departing from the technical concept.

[0138] For example, in the examples shown in Figures 2 and 8, the volume control unit 12 of the haptic presentation device 1 stores the self-velocity v of all sampled frames in the memory 41. Then, the maximum velocity extraction unit 42 of the volume control unit 12 reads the self-velocity v of all sampled frames from the memory 41 and extracts the maximum velocity v max When the viewer initiates playback, the volume reduction control unit 44 extracts the data, and if the viewer initiates playback, it sets the self-speed v and the maximum speed v to a threshold value. max Based on this, the volume A of the low-frequency audio signal S was reduced.

[0139] In response, the volume reduction control unit 44 uses the maximum speed v extracted by the maximum speed extraction unit 42. max Instead of using a predetermined maximum speed v max The volume A of the low-frequency audio signal S may be reduced using this method. In this case, the volume control unit 12 does not need to include the memory 41 and the maximum speed extraction unit 42 in the configuration example shown in Figure 7.

[0140] In other words, without the volume control unit 12 waiting for the viewer to initiate playback, the volume reduction control unit 44, when the self-speed v is below a threshold, uses the self-speed v estimated by the self-speed estimation unit 11 and a preset maximum speed v max Based on this, the volume A of the low-frequency audio signal S is reduced.

[0141] This enables a series of processes to be carried out in real time, from the self-velocity estimation unit 11 estimating the self-velocity v by inputting a first-person perspective video E, to the volume control unit 12 lowering the volume A of the low-frequency audio signal S to generate a volume-controlled video E', and finally to the haptic presentation unit 13 outputting the low-frequency audio signal S' to the haptic device 7.

[0142] Furthermore, in the examples shown in Figures 2 and 8, the volume control unit 12 of the haptic presentation device 1, when it determines that the viewer has initiated viewing, extracts a low-frequency audio signal S from the first-person view video E, lowers the volume A of the low-frequency audio signal S to generate a volume-controlled video E', and the haptic presentation unit 13 extracts the low-frequency audio signal S' from the volume-controlled video E' and outputs it to the haptic device 7.

[0143] Alternatively, the volume control unit 12 may store the generated volume-controlled video E' in a memory (not shown in Figure 7), and the haptic presentation unit 13 may repeatedly use the volume-controlled video E' stored in the memory each time the viewer initiates viewing.

[0144] Furthermore, a standard computer can be used as the hardware configuration for the haptic presentation devices 1 and 2. The haptic presentation devices 1 and 2 consist of a computer equipped with a CPU, volatile storage media such as RAM, non-volatile storage media such as ROM, and interfaces. The same applies to the self-velocity estimation devices 3 and 4.

[0145] The functions of the self-velocity estimation unit 11, volume control unit 12, and haptic presentation unit 13 in the haptic presentation device 1 are realized by having the CPU execute a program that describes these functions. Similarly, the functions of the volume control unit 12 and haptic presentation unit 13 in the haptic presentation device 2 are also realized by having the CPU execute a program that describes these functions.

[0146] Furthermore, the functions of the frame sampling processing unit 21, mask image generation processing unit 22, and self-velocity estimation processing unit 23 provided in the self-velocity estimation device 3, as well as the functions of the frame sampling processing unit 21, mask image generation processing unit 22, self-velocity estimation processing unit 23, and trimming processing unit 24 provided in the self-velocity estimation device 4, are each realized by having the CPU execute a program that describes these functions.

[0147] These programs are stored in the aforementioned storage medium and are read and executed by the CPU. These programs can also be stored and distributed on storage media such as magnetic disks (floppy disks, hard disks, etc.), optical disks (CD-ROMs, DVDs, etc.), and semiconductor memory, and can be transmitted and received via a network. [Explanation of Symbols]

[0148] 1,2 Tactile presentation device 3,4 Self-speed estimation device 7. Haptic devices 8 Display 9 speakers 11 Self-speed estimation part 12 Volume control unit 13. Tactile presentation section 21 Frame sampling processing unit 22 Mask Image Generation Processing Unit 23 Self-speed estimation processing unit 24 Trimming Processing 31. Neural Network (NN) 32 Self-speed calculation section 41 memory 42 Maximum speed extraction part 43 Low-frequency sound extraction unit 44 Volume Reduction Control Unit 45 Synthesis part 51 Extraction part 52 Amplification section E,E1 First-person perspective video E' Volume-controlled video v Self-speed v max maximum speed S,S' Low-frequency audio signal A,A new volume translation(x,y,z) Translation vector

Claims

1. In a haptic presentation device, the camera's position is set to the viewer's position, and the device receives first-person perspective video footage captured by the camera's movement. The device extracts a low-frequency audio signal from the first-person perspective video footage and generates information for presenting haptic stimuli to the viewer via a haptic device based on the low-frequency audio signal. The speed corresponding to each of the multiple frames in the time series in the aforementioned first-person perspective video, where the viewer's movement speed accompanying the camera's movement is set in advance as the viewer's own speed, A volume control unit that extracts the low-frequency audio signal from the first-person perspective video and, if the self-velocity in the frame is below a predetermined threshold, reduces the volume of the low-frequency audio signal corresponding to the frame. A tactile presentation unit outputs the low-frequency audio signal, whose volume has been reduced by the volume control unit, to the tactile device. A tactile presentation device characterized by having the following features.

2. In a haptic presentation device, the camera's position is set to the viewer's position, and the device receives first-person perspective video footage captured by the camera's movement. The device extracts a low-frequency audio signal from the first-person perspective video footage and generates information for presenting haptic stimuli to the viewer via a haptic device based on the low-frequency audio signal. For each of the multiple time-series frames included in the aforementioned first-person perspective video, the viewer's movement speed in that frame is taken as the viewer's own speed. A self-velocity estimation unit generates a mask image for each of the plurality of frames that masks objects moving independently of the self-velocity, estimates the viewer's translation vector in the first-person view video based on a predetermined number of consecutive frames and the same number of mask images corresponding to the predetermined number of frames using a predetermined neural network, and calculates the self-velocity in the frame based on the translation vector. A volume control unit extracts the low-frequency audio signal from the first-person perspective video and, if the self-velocity in the frame calculated by the self-velocity estimation unit is below a predetermined threshold, reduces the volume of the low-frequency audio signal corresponding to the frame. A tactile presentation unit outputs the low-frequency audio signal, whose volume has been reduced by the volume control unit, to the tactile device. A tactile presentation device characterized by having the following features.

3. In the tactile presentation device according to claim 1 or 2, The volume control unit, A low-frequency audio extraction unit extracts the low-frequency audio signal, video signal, and audio signals other than the low-frequency audio signal from the first-person perspective video. A volume reduction control unit that, when the self-velocity in the frame is less than or equal to the predetermined threshold, reduces the volume of the low-frequency audio signal extracted by the low-frequency audio extraction unit corresponding to the frame to generate a new low-frequency audio signal, and when the self-velocity in the frame is greater than the predetermined threshold, sets the low-frequency audio signal corresponding to the frame as is and uses it as the new low-frequency audio signal. The system comprises a combining unit that combines the new low-frequency audio signal generated or set by the volume reduction control unit, the video signal extracted by the low-frequency audio extraction unit, and audio signals other than the low-frequency audio signal to obtain a volume-controlled video, The tactile presentation unit is, A tactile presentation device characterized by extracting a new low-frequency audio signal from the volume-controlled video obtained by the synthesis unit, and outputting the new low-frequency audio signal to the tactile device.

4. In the tactile presentation device described in claim 3, The volume control unit further, The system includes a maximum speed extraction unit that extracts the maximum speed from the self-speed of each of the multiple frames, The volume reduction control unit, The maximum speed extracted by the maximum speed extraction unit is v max Let v be the self-velocity in the frame, let A be the volume of the low-frequency audio signal, and let A be the volume of the new low-frequency audio signal. new Assuming that the self-velocity in the frame is below a predetermined threshold, the following formula applies: A new =(v / v max )A A tactile presentation device characterized by reducing the volume of the low-frequency audio signal and generating a new low-frequency audio signal.

5. In the tactile presentation device described in claim 3, The volume reduction control unit, The pre-set maximum speed v max Let v be the self-velocity in the frame, let A be the volume of the low-frequency audio signal, and let A be the volume of the new low-frequency audio signal. new Assuming that the self-velocity in the frame is below a predetermined threshold, the following formula applies: A new =(v / v max )A A tactile presentation device characterized by reducing the volume of the low-frequency audio signal and generating a new low-frequency audio signal.

6. In the tactile presentation device according to claim 2, The self-velocity estimation unit, A tactile presentation device characterized by: trimming each of the plurality of frames within a predetermined fixed coordinate frame; generating a mask image for each of the plurality of frames after trimming; estimating the translation vector using a predetermined NN based on a predetermined number of frames and the same number of mask images corresponding to the predetermined number of frames; and calculating the self-velocity in the frame based on the translation vector.

7. A computer comprising a haptic presentation device that inputs first-person perspective video captured by moving the camera with the camera's position as the viewer's position, extracts a low-frequency audio signal from the first-person perspective video, and generates information for presenting haptic stimuli to the viewer via a haptic device based on the low-frequency audio signal, The speed corresponding to each of the multiple frames in the time series in the aforementioned first-person perspective video, where the viewer's movement speed accompanying the camera's movement is set in advance as the viewer's own speed, A volume control unit that extracts the low-frequency audio signal from the first-person perspective video and, if the self-velocity in the frame is below a predetermined threshold, reduces the volume of the low-frequency audio signal corresponding to the frame, and A program for causing the low-frequency audio signal, whose volume has been reduced by the volume control unit, to function as a tactile presentation unit that outputs it to the tactile device.