Video distribution method, video playback method, video distribution device, and distribution data structure

By generating and distributing video streams for multiple viewpoints on a celestial sphere, the system addresses server load issues caused by user line-of-sight changes, ensuring efficient and high-quality video distribution.

JP7811787B2Active Publication Date: 2026-02-06OHMI DIGITAL FAB CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022503743
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-29
Filing Date
2021-02-26
Publication Date
2026-02-06
Estimated Expiration
2041-02-26

AI Technical Summary

Technical Problem

Existing video distribution systems experience increased server load due to changes in user line of sight, particularly when handling multiple user requests, as they require real-time processing of high-resolution images.

Method used

The system generates and distributes video streams for multiple viewpoints on a celestial sphere, allowing distribution of a video stream corresponding to a viewpoint other than the nearest to the user's line of sight, reducing the need for real-time high-resolution image processing on the server.

Benefits of technology

This approach reduces server load by pre-generating video streams for various viewpoints, enabling efficient distribution and maintaining image quality even with changes in user gaze without significant increases in server processing demands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007811787000008
    Figure 0007811787000008
  • Figure 0007811787000009
    Figure 0007811787000009
  • Figure 0007811787000010
    Figure 0007811787000010
Patent Text Reader

Abstract

[Problem] To provide a video delivery method and a video delivery device that alleviate an increase in server load due to a change in a user's line of sight. [Solution] A video delivery method comprising a step for storing, with respect to each of a plurality of points of view defined in a celestial sphere having a camera 10 as an observation point, a video stream 44 including the celestial sphere, and a delivery step for delivering the video stream 44 to a user terminal 14, the video delivery method characterized in that the delivery step delivers the video stream 44 of the points of view other than a closest point of view in the celestial sphere corresponding to a line of sight determined by the user terminal 14.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a video distribution method for distributing videos, a video playback method, a video distribution device, and a distribution data structure. [Background technology]

[0002] Distribution systems that distribute still images and videos are known. For example, the distribution system disclosed in Patent Document 1 includes a server and a client, and the server's memory stores key frame images and difference frame images that constitute the video to be distributed. When the server receives a request from the client, it distributes the key frame images and difference frame images stored in the memory to the client.

[0003] The panoramic video distribution system described in Non-Patent Document 1 includes a server that distributes the entire background as a low-resolution image and cuts out and distributes the portion corresponding to the user's line of sight as a high-resolution image. The client that receives the low-resolution image and the high-resolution image synthesizes these images and displays them on the screen, enabling the portion the user is looking at to be displayed in high quality. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent No. 6149967 [Non-patent literature]

[0005] [Non-Patent Document 1] NTT Technocross "Panorama Chou Player / Panorama Chou Engine," [Retrieved February 23, 2020], Internet<https: / / www.ntt-tx.co.jp / products / panocho / > Summary of the Invention [Problem to be solved by the invention]

[0006] However, since the server must perform processing to crop high-resolution images in accordance with the user's viewpoint, which changes from moment to moment, the load on the server increases accordingly when a large number of users access the server. Furthermore, as in Patent Document 1, if key frame images and difference frame images are generated for the video to be transmitted (high-quality video), the load on the server will increase even further.

[0007] An object of the present invention is to provide a video distribution method, a video playback method, a video distribution device, and a distribution data structure that reduce an increase in server load caused by changes in a user's line of sight. [Means for solving the problem]

[0008] In order to achieve the above object, the video distribution method of the present invention includes a step of storing a video stream including a celestial sphere for each of a plurality of viewpoints defined on the celestial sphere with a camera as an observation point, and a distribution step of distributing the video stream to a user's terminal, wherein the distribution step distributes a video stream of a viewpoint other than the nearest viewpoint on the celestial sphere that corresponds to the line of sight determined on the user's terminal.

[0009] In order to achieve the above object, the video playback method of the present invention includes a step of storing a video stream including a celestial sphere for each of a plurality of viewpoints defined on the celestial sphere with a camera as an observation point, and a playback step of playing back the video stream on a user's terminal, wherein the playback step plays back a video stream of a viewpoint other than the nearest viewpoint on the celestial sphere that corresponds to the line of sight determined on the user's terminal.

[0010] In addition, in order to achieve the above-mentioned object, the video distribution device of the present invention includes a memory unit that stores a video stream including a celestial sphere for each of a plurality of viewpoints defined on the celestial sphere with a camera as an observation point, and a distribution unit that distributes the video stream to a user's terminal, wherein the distribution unit distributes the video stream of a viewpoint other than the nearest viewpoint on the celestial sphere that corresponds to the line of sight determined on the user's terminal.

[0011] Furthermore, the distribution data structure of the present invention comprises a video stream that includes, in its center, an image on a line of sight directed from a specific observation point, and, outside the center, an image of the celestial sphere photographed from the observation point, the video stream including a first video stream that includes, in its center, an image of a viewpoint on a first line of sight directed from the specific observation point, and a second video stream that includes, in its center, an image of a viewpoint on a second line of sight directed from the observation point. [Effects of the Invention]

[0012] According to the video distribution method, video playback method, video distribution device, and distribution data structure of the present invention, it is possible to reduce an increase in server load caused by changes in the user's line of sight. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a schematic diagram of a video distribution system according to an embodiment of the present invention; [Figure 2] (a) A hardware schematic diagram of a user terminal of the video distribution system, (b) a hardware schematic diagram of a camera of the video distribution system, and (c) a hardware schematic diagram of a server of the video distribution system. [Figure 3] (a) A flow diagram of a video generation and distribution program executed on the server. (b) A flow diagram of a generation process in the video generation program. [Figure 4] FIG. 10 is a diagram showing an image generated in the generation process. [Figure 5] A diagram showing the position of the viewpoint. [Figure 6] A diagram showing the pixel extraction process in the video stream generation process. [Figure 7] A diagram showing the correspondence between keyframe images for each viewpoint and a virtual sphere [Figure 8] A diagram showing an example of a function of angle of view information [Figure 9] FIG. 10 is a diagram showing the correspondence relationship when generating low-quality parts in viewpoint-specific keyframe images. [Figure 10] A diagram showing an example of a group of video streams for different viewpoints. [Figure 11] User terminal flow diagram DETAILED DESCRIPTION OF THE INVENTION

[0014] [First embodiment]

[0015] Hereinafter, a video distribution system and a video distribution method according to an embodiment of the present invention will be described with reference to the drawings.

[0016] 1, the video distribution system 1 of the first embodiment is a system that distributes video (video stream 44) to a user terminal 14 (hereinafter referred to as user terminal 14), and includes a camera 10 that generates images, and a server 12 that functions as a distribution device that generates video for distribution based on images acquired from the camera 10. The camera 10, server 12, and user terminal 14 are connected to a network typified by an Internet communication line, and the server 12 is capable of communicating with the camera 10 and the user terminal 14.

[0017] The user terminal 14 is a mobile information terminal such as a publicly known smartphone or tablet terminal, and as shown in Figure 2(a), it is equipped with a communication module 16 (communication unit) which is an interface for connecting to an Internet communication line, an LCD display 18 (display unit) which displays video received from the server 12, a touch panel 20 (input unit) which is superimposed on the LCD display 18 and accepts input from the user, an angular velocity sensor 22 (detection unit) which detects the attitude of the terminal, and a CPU 26 (control unit) which controls the LCD display 18, touch panel 20, and angular velocity sensor 22 by executing a program stored in memory 24.

[0018] The camera 10 is a device that generates at least a hemispherical image. As shown in FIG. 2(b), the camera 10 includes an image sensor 28, a fisheye lens, which is an optical component that forms an image circle of a virtual hemisphere of infinite radius, with the image sensor 28 as the observation point, on the light-receiving surface of the image sensor 28, a CPU 30 that controls the image sensor 28 and generates the hemispherical image based on the electrical signal output from the image sensor 28, and a communication module 32 for connecting to an Internet communication line. The camera 10 generates the hemispherical image at a frame rate of 60 fps (frames per second). The multiple consecutive hemispherical images thus generated are stored in a memory 34 in chronological order. After multiple hemispherical images (a group of hemispherical images) captured and generated over a certain period of time are accumulated in the memory 34, the camera 10 transmits the group of hemispherical images stored in the memory 34 to the server 12 via the Internet communication line.

[0019] The server 12 is a terminal that distributes video (video stream 44), which is distribution data generated based on the group of hemispherical images, to the user terminal 14, and as shown in FIG. 2(c), includes a communication module 36 connected to an Internet communication line, a memory 38 in which a video generation and distribution program is stored, and a CPU 40 that executes the video generation and distribution program.

[0020] 3(a), the video generation and distribution program is a program that causes the server 12 to execute an acquisition process (s10) of acquiring a group of hemispherical images from the camera 10, a generation process (s20) of generating a video stream 44 to be distributed from the acquired group of hemispherical images, and a distribution process (s30) of distributing the video stream 44 in response to a request from the user terminal 14 to the user terminal 14. In the present embodiment, the video stream 44 generated in the generation process (s20) is generated for each predetermined viewpoint on a virtual celestial sphere with an infinite radius, with the image sensor 28 of the camera 10 as the base point, as shown in FIG. 5. In other words, when the hemispherical images generated by the camera 10 are mapped onto the virtual celestial sphere, a plurality of viewpoints on the celestial sphere are set as the user's viewpoint when observing the celestial sphere from the viewpoint, and the video stream 44 is generated for each of the plurality of viewpoints. Then, in the distribution process (s30), a video stream 44 of one viewpoint that corresponds to or is similar to the user's line of sight information included in the request from the user terminal 14 is transmitted to the user terminal 14. A specific description will be given below.

[0021] The acquisition process (s10) is a process of acquiring a group of hemispherical images 42 (FIGS. 1 and 4) from the camera 10, and the received group of hemispherical images 42 are stored in chronological order in the memory 38. In this way, the server 12 functions as an acquisition unit that acquires the group of hemispherical images 42 from the camera 10. The memory 38 of the server 12 also functions as a storage unit that stores the group of hemispherical images 42.

[0022] After the acquisition process (s10) is executed, a generation process (s20) is executed. As shown in Fig. 4, the generation process (s20) is a process for generating a video stream 44 (continuous images that are successive in time series) for each predetermined viewpoint based on a group of hemispherical images 42 stored in memory 38, and includes an intermediate image generation process (s21) and a video stream generation process (s22).

[0023] The intermediate image generation process (s21) extracts a group of hemispherical images 42 stored in the memory 38 and generates a group of intermediate images 46 from the extracted group of hemispherical images 42 ( FIG. 4 ). The group of intermediate images 46 includes key frame images 46a and difference frame images 46b generated using known inter-frame prediction. In this embodiment, the hemispherical images 42a extracted from the group of hemispherical images 42 for each predetermined number of frames (every 60 frames in this embodiment) are used as key frame images. Furthermore, for a plurality of other hemispherical images 42b in the frame following the hemispherical image 42a (key frame image), differences between the hemispherical image of the previous frame are calculated to generate difference frame images 46b. The group of intermediate images 46 generated by the intermediate image generation process (s21) is stored in the memory 38 of the server 12. The last frame of the group of hemispherical images 42 is extracted as a key frame image 46a to be placed in the last frame of the group of intermediate images 46. In this way, the CPU 38 of the server 12 functions as an intermediate image generation unit that generates the group of intermediate images 46, and the memory 38 of the server 12 functions as a storage unit that stores the group of intermediate images 46. Hereinafter, the key frame image 46a and the difference frame image 46b generated in the intermediate image generation process (s21) will be referred to as intermediate key frame image 46a and intermediate difference frame image 46b, respectively.

[0024] The video stream generation process (s22) is a process of generating a video stream 44 for each viewpoint based on a group of intermediate images 46. The video stream 44 is a series of images delivered to the user terminal 14, and includes viewpoint-specific key frame images 44a and viewpoint-specific difference frame images 44b. As described above, the video stream 44 is generated corresponding to each of a plurality of predetermined viewpoints. The plurality of predetermined viewpoints are, as shown in FIG. 5 as described above, a plurality of points determined on a virtual celestial sphere including the celestial sphere viewed from the image sensor 28 of the camera 10 as the observation point (base point), and each viewpoint is defined by viewpoint information consisting of a roll angle (α), a pitch angle (β), and a yaw angle (γ) with respect to the observation point as the base point. For example, for viewpoint a, (α a ,βa ,γ a ), and this viewpoint information is stored in memory 38 in association with viewpoint identification information assigned to viewpoint a. Also, a video stream 44 generated for viewpoint a is stored in memory 38 in association with the viewpoint identification information. That is, the viewpoint information for each viewpoint and the video stream 44 generated for each viewpoint are stored in memory 38 in association with the viewpoint identification information.

[0025] As shown in Fig. 6(a), the viewpoint-specific key frame images 44a and viewpoint-specific difference frame images 44b constituting the video stream 44 are compressed so that the image quality gradually decreases from the center of the image toward the outside when they are expanded on the user terminal 14. Taking the viewpoint-specific key frame image 44a as an example, as shown in Fig. 6(b), the viewpoint-specific key frame image 44a is compressed so that, with the center of the viewpoint-specific key frame image 44a as the base point, the image quality is high inside the inscribed circle inscribed on the four sides (edges) of the viewpoint-specific key frame image 44a, and low outside the inscribed circle (the four corners of the image). The process of generating such viewpoint-specific key frame images 44a will be described below using viewpoint a as an example.

[0026] As shown in Figures 6(c) and 6(d), a viewpoint-specific keyframe image 44a at viewpoint a (hereinafter referred to as viewpoint a keyframe image 44a) is generated by extracting pixels from a virtual sphere 56 onto which an intermediate keyframe image 46a is virtually mapped. Specifically, as shown in Figure 3(b), for each pixel that constitutes the viewpoint a keyframe image 44a, a corresponding first coordinate on the surface of the virtual sphere 56 is calculated using a correspondence equation (first calculation process (s221)), a rotation equation including viewpoint information of viewpoint a is applied to the first coordinate to calculate a second coordinate (second calculation process (s222)), and the pixel on the surface of the virtual sphere 56 located at the second coordinate is extracted. Note that the coordinates of the viewpoint a keyframe image 44a are represented by XY Cartesian coordinates with the center as the origin, as shown in Figure 7(b), and the horizontal direction (X coordinate) of the viewpoint a keyframe image 44a takes a value of -1 ≤ X ≤ 1, and the vertical direction (Y coordinate) takes a value of -1 ≤ Y ≤ 1. As shown in FIG. 7(a), the coordinates of the virtual sphere 56 are expressed by an XYZ Cartesian coordinate system with the center of the virtual sphere 56 as the origin, and the radius r of the virtual sphere 56 is set to 1.

[0027] The first calculation process (s221) includes a spherical coordinate calculation process for calculating spherical coordinates (r, θ, φ) on the virtual sphere 56 based on the coordinates of the viewpoint a keyframe image 44a and field of view information, and an orthogonal coordinate calculation process for calculating orthogonal coordinates (x, y, z) corresponding to the spherical coordinates. The field of view information is information indicating the range to be displayed on the liquid crystal display 18 of the user terminal 14, and is set to 30° in this embodiment.

[0028] As shown in FIG. 7, the spherical coordinate calculation process will be described using pixel P included in viewpoint a keyframe image 44a as an example. The angle θp' with respect to the Z axis and the angle φp' with respect to the X axis of the virtual sphere are calculated as follows. As described above, the radius r of virtual sphere 56 is 1. The angle θp' is determined based on the distance Pr from the origin to pixel P in the XY Cartesian coordinate system of viewpoint a keyframe image 44a and predetermined angle of view information. The distance Pr is determined by the following correspondence equation based on the coordinate values ​​(Px, Py) of pixel P.

[0029]

number

[0030] The calculated value of the distance Pr is then input into a function f(Pr) that is predetermined according to the angle of view information to determine the angle θp'. As shown in FIG. 8(a), this function defines the relationship between the distance Pr and the angle θp'. For example, when the angle of view information is set to 30°, the function is defined so that θ becomes 30° when Pr=1. The distance Pr determined in Equation 1 above is substituted into this function to determine the angle θ at point P. That is, the function is defined so that the boundary between the high pixel area and the low pixel area in the viewpoint a keyframe image 44a corresponds to the angle of view information. The angle of view information and the function may be defined so that θ becomes 90° when Pr=1 when the angle of view information is 90°, as shown in FIG. 8(b). Alternatively, the function may be a linear function as shown in FIG. 8(c).

[0031] The angle φp' is the same as φp in the XY Cartesian coordinate system of the viewpoint a keyframe image 44a, and φp is calculated based on the coordinates (Px, Py) of the point P by the following equation.

[0032]

number

[0033] Here, as shown in FIG. 9(b), if the angle φ is calculated for pixels constituting a low-image-quality portion, for example, pixels on the circumference C, in the same manner as in the correspondence equation (Equation 2), the pixels located on the dashed arc (dashed arc) are not taken into account, and only pixels corresponding to the arc shown by the dashed line are taken into account, resulting in biased extraction of pixel information. Therefore, in this embodiment, based on the ratio of the dashed arc to the circumference C, points on the circumference C, including the dashed line portion, are evenly arranged on the dashed line, thereby thinning out and extracting pixel information without bias, thereby reducing the amount of information in the viewpoint a keyframe image 44a (video stream). Therefore, for example, pixel information corresponding to pixel Q' is extracted for pixel Q on the circumference C. The correspondence equation for achieving such an even arrangement is as follows:

[0034]

number

[0035] where φ i is the angle used to calculate the ratio (proportion) of the dashed arc to the circumference C.

[0036] Once the spherical coordinates (1, θ, φ) for each pixel in the viewpoint a keyframe image 44a are calculated as described above, the first coordinates (x1, y1, z1) for each pixel are calculated using the following conversion formula in the orthogonal coordinate calculation process.

[0037]

number

number

number

[0038] After the orthogonal coordinate calculation process is performed, the second calculation process is performed. In the second calculation process, viewpoint information (α a ,β a ,γ a), the second coordinate (x2, y2, z2) is obtained by applying a rotation formula including

[0039]

number

[0040] The second calculation process identifies pixels to be extracted on the virtual sphere. Information about the identified pixels is then extracted, and the extracted pixel information is assigned to each corresponding pixel in the viewpoint a keyframe image 44a. In this way, within the inscribed circle, which is the high-image-quality portion, pixels on the virtual sphere are extracted in a fisheye image-like manner according to the angle of view, and outside the inscribed circle, which is the low-image-quality portion, pixels on the virtual sphere outside the angle of view are thinned out and extracted, resulting in the generated viewpoint a keyframe image 44a.

[0041] As described above, the process for generating the viewpoint-specific key frame image 44a for viewpoint a has been described, but the viewpoint-specific difference key frame image 44b for viewpoint a is also generated by a similar process. In this manner, the video stream 44 for viewpoint a is generated. For the other viewpoints, video streams 44 (viewpoint-specific key frame images 44a and viewpoint-specific difference frame images 44b) are generated by a process similar to that for viewpoint a, and the generated video streams 44 are associated with the viewpoint information (associated with the viewpoint information by being associated with the viewpoint identification information) and stored in the memory 38 of the server 12. In this manner, the memory 38 of the server 12 functions as a storage unit that stores the video streams 44 for each viewpoint in association with the viewpoint information.

[0042] As described above, a video stream 44 for each viewpoint is generated, but in this embodiment, the viewpoint-specific key frame images 44a constituting the video stream 44 are not synchronized between viewpoints, and the viewpoint-specific key frame images 44a for one viewpoint and the viewpoint-specific key frame images 44a for another viewpoint are arranged at different timings in time series and stored in the memory 38. That is, in each of the video streams 44, the viewpoint-specific key frame images 44a and viewpoint-specific difference frame images 44b are arranged such that they are asynchronous with each other in time series. 10, in a video stream 44 for viewpoints a to d, viewpoint-specific keyframe images KF002a, KF002b, KF002c, and KF002d for each of viewpoints a to d are images generated from the intermediate keyframe image KF002, but with respect to the viewpoint a keyframe image KF002a, the viewpoint b keyframe image KF002b is arranged so as to be delayed by four frames, the viewpoint c keyframe image KF002c is arranged so as to be delayed by nine frames, and the viewpoint d keyframe image KF002d is arranged so as to be delayed by 14 frames. As described above, in the video stream 44 for viewpoint b, for example, the viewpoint b keyframe image KF001b (the first viewpoint-specific keyframe image 44a) is arranged continuously over frames 1 to 4 so that the video streams 44 for each viewpoint are asynchronous with each other.

[0043] Next, the distribution process (s30) to the user terminal 14 will be described.

[0044] Prior to the distribution process (s30), a peer-to-peer connection is established between the server 12 and the user terminal 14 by a signaling server 12 (not shown), enabling mutual communication and receiving a request (s40) from the user terminal 14 (FIG. 11). The request is information requesting the server 12 to distribute a video, and includes line-of-sight information of the user terminal 14. The line-of-sight information is information indicating the user's line of sight (the center of the image to be displayed on the user terminal 14), and includes a roll angle (α), pitch angle (β), and yaw angle (γ) determined by the CPU 26 of the user terminal 14 based on the output signal of the angular velocity sensor 22.

[0045] When the server 12 receives a request from the user terminal 14, it compares the gaze information included in the request with multiple viewpoint information stored in memory, and delivers to the user terminal 14 a video stream 44 corresponding to the viewpoint information that matches or is close to the gaze information.

[0046] As shown in FIG. 11, the user terminal 14 performs unfolding processing (s60) when it receives the video stream 44. In the unfolding processing (s60), first, key frame images for unfolding and difference frame images are generated based on the received video stream 44. In the center of the key frame image for unfolding, pixels of the high-quality image part in the viewpoint-specific key frame image 44a are placed as they are. Images of the low-quality image part in the viewpoint-specific key frame image 44a are placed around the high-quality image part. Here, the corner pixels of the low-quality image part are not placed as they are, but are calculated by using the above formula 4 to obtain φ Q The position of ' is identified, and the pixel is placed at the identified position. Q Since pixels are not arranged consecutively on the circumference C including ', an interpolation process is performed to interpolate between each pixel. There are no particular limitations on the interpolation process, but for example, pixels that are similar to each other are arranged between pixels on the same circumference. A difference frame image for expansion is generated by the same process as for the key frame image for expansion.

[0047] After the interpolation process has been completed and key frame images for expansion and differential frame images for expansion are generated, key frame images for display and differential frame images for display are generated using a known panoramic expansion process, and a video is generated based on these images, and the video is displayed on the user terminal 14.

[0048] While a video is being displayed (played) on the user terminal 14, the CPU 26 of the user terminal 14 monitors the user's line of sight by checking the output of the angular velocity sensor 22, and shifts the display coordinates of the video in accordance with the amount of change in the line of sight. The user terminal 14 also updates the line of sight information and transmits the updated line of sight information to the server 12.

[0049] Each time the server 12 receives gaze information, it extracts viewpoint information from the video stream 44 in which key frames are arranged close to each other in time series, compares the received gaze information with the extracted viewpoint information to search for the most similar viewpoint, and transmits the video stream 44 corresponding to the most similar viewpoint to the user terminal 14.

[0050] Here, while the user's line of sight changes from line of sight a to line of sight f, specifically, while the posture of the user terminal 14 changes due to the user's terminal operation and the user's line of sight detected based on the change in posture changes from viewpoint a to viewpoint f, the video stream 44 is delivered as follows.

[0051] When the server 12 receives gaze information from the user terminal 14, it searches for a video stream 44 in which a viewpoint-specific key frame image 44a is arranged at a timing close in time series relative to the time of reception. Specifically, as described above, the video streams 44 stored in the memory 38 are generated so as to be asynchronous with each other in time series, and therefore the arrangement positions (arrangement timing) of the viewpoint-specific key frame images 44a in the video streams 44 differ from one another in the multiple video streams 44. The CPU 12 of the server 12 calculates the arrangement positions (arrangement timing) of the key frame images based on the key frame period (60 frames) in each video stream 44 and the delay set in each video stream 44, and searches for a video stream 44 that has a viewpoint-specific key frame image 44a at a timing closest to the frame image being distributed at the time the gaze information was received (the frame image being played on the user terminal 14). Then, it is determined whether the viewpoint information corresponding to the searched video stream 44 is closer in position to the viewpoint after the change (viewpoint f) than to the viewpoint before the change (viewpoint a).

[0052] For example, if a search for viewpoint-specific keyframe images 44a determines that viewpoint-specific keyframe image 44a from viewpoint c is close in time series, viewpoint c is closer in position to viewpoint f than viewpoint a, so the video stream 44 from viewpoint c is delivered to the user terminal 14.

[0053] On the other hand, even if it is determined as a result of searching for viewpoint-specific keyframe images 44a that viewpoint-specific keyframe images 44a for viewpoint g are arranged at a close timing in the time series, viewpoint g is farther away from viewpoint f than viewpoint a, so the video stream 44 for viewpoint a is delivered to the user terminal 14.

[0054] In the video distribution system 1 of this embodiment, video streams 44 corresponding to multiple viewpoints are generated in advance, so even if a user's gaze occurs, it is sufficient to distribute the video stream 44 corresponding to the viewpoint determined by the gaze, thereby reducing the increase in load on the server 12 even if there are requests from a large number of user terminals 14.

[0055] Furthermore, even if the display coordinates of the video being displayed are shifted as a result of a change in the posture of the user terminal 14, causing a gradual decrease in image quality, a video stream 44 having a viewpoint-specific keyframe image 44a that is close in time series and position is delivered, thereby preventing a significant decrease in image quality of the image being displayed.

[0056] [Second embodiment]

[0057] In the first embodiment, the server 12 selects the video stream 44 to be distributed based on the line-of-sight information received from the user terminal 14, but in the second embodiment, the user terminal 14 selects the video stream 44 to be received based on the line-of-sight information and requests the server 12 to distribute the selected video stream 44. The following description will focus on configurations and flows that differ from the first embodiment, and will omit configurations and methods that are common to the first embodiment as appropriate.

[0058] In this embodiment, similar to the first embodiment, a video stream 44 generated for each viewpoint is associated with viewpoint identification information and stored in the memory of the server 12, but the viewpoint information is different from the first embodiment in that the viewpoint information is not stored in the memory 24 of the server 12. In this embodiment, each piece of viewpoint information is associated with viewpoint identification information and stored in the memory 24 of the user terminal 14.

[0059] In this embodiment, similarly to the first embodiment, multiple video streams 44 generated for each viewpoint contain viewpoint-specific keyframe images 44a. In these multiple video streams 44, the first viewpoint-specific keyframe image 44a is offset, so that the viewpoint-specific keyframe images 44a are arranged asynchronously with one another in time series. In this embodiment, the arrangement timing of the viewpoint-specific keyframe images 44a in the video stream 44 for each viewpoint is stored in the memory 24 of the user terminal 14. The arrangement timing indicates the timing (frame) at which the viewpoint-specific keyframe image 44a is arranged in the video stream 44 for each viewpoint, and is typically the interval (arrangement period) of the viewpoint-specific keyframe images 44a in each video stream 44 and the offset number (number of frames to delay) of the first viewpoint-specific keyframe image 44a in the video stream 44 for each viewpoint. In this embodiment, as shown in FIG. 4 , a viewpoint-specific keyframe image is arranged every 60 frames, so the interval is "60." 10, the viewpoint-specific key frame image 44a of viewpoint a is not offset, so the offset number for viewpoint a is "0." In addition, the first viewpoint-specific key frame image 44a in the video stream 44 of viewpoint b is offset by four frames, so the offset number for viewpoint b is "4." Similarly, the offset number for viewpoint c is "9," and the offset number for viewpoint d is "14." In this way, the placement timings defined for each viewpoint are stored in association with viewpoint identification information, and each piece of viewpoint identification information is stored in association with viewpoint information.

[0060] As described above, the user terminal 14 of this embodiment determines the video stream 44 of the viewpoint to be received based on the line-of-sight information, and requests the delivery of the video stream 44 of the determined viewpoint from the server 12. Specifically, the user terminal 14 executes line-of-sight information acquisition processing, request processing, and display processing in this order. (1) The gaze information acquisition process is a process in which the CPU 26 of the user terminal 14 acquires gaze information based on the output from the angular velocity sensor 22, and similar to the first embodiment, acquires the roll angle (α), pitch angle (β), and yaw angle (γ). (2) The request process extracts viewpoint information that is similar to the viewpoint information acquired in the above-described viewpoint information acquisition process, and transmits viewpoint identification information corresponding to the extracted viewpoint information to the server 12. Upon receiving the viewpoint identification information from the user terminal 14, the server 12 distributes the video stream 44 corresponding to the viewpoint identification information to the user terminal 14. (3) Display processing is processing for displaying the video stream 44 on the liquid crystal display 18 while receiving the video stream 44 from the server 12. The above flow executes the delivery and display of the video stream 44 in the initial stage.

[0061] As described above, the CPU 26 of the user terminal 14 executes gaze information acquisition processing, determination processing, request processing, and display processing in synchronization with the frame rate of the video stream 44 while displaying the video stream 44, in order to display the video stream 44 in accordance with changes in gaze caused by the user operating the user terminal 14.

[0062] (4) The line-of-sight information acquisition process is similar to the process (1) above, and is a process of acquiring line-of-sight information (roll angle (α), pitch angle (β), and yaw angle (γ)) based on the output of the angular velocity sensor 22.

[0063] (5) The determination process is a process for determining a video stream to be requested from the server 12, and the CPU 26 of the user terminal 14 selects viewpoint identification information in which the viewpoint-specific keyframe images 44a are arranged close in time series. (5-1) Specifically, the frame number in the currently playing video stream 44 (hereinafter referred to as the currently playing frame number) is identified. For example, when the video stream 44 from viewpoint a is being played and the 100th frame image is being displayed, the frame number is identified as "100." (5-2) Next, based on the interval and offset determined as the placement timing, the placement position of the viewpoint-specific keyframe image 44a is calculated for each viewpoint, and the number of the keyframe that is placed after the identified frame number and is close in time series is extracted. For example, in the placement timing of viewpoint b, the interval is defined as "60" and the offset is defined as "4", so the position of the first viewpoint-specific keyframe image 44a of viewpoint b is determined as "5", the position of the second viewpoint-specific keyframe image 44a as "65", the position of the third viewpoint-specific keyframe image 44a as "125", and the position of the fourth viewpoint-specific keyframe image 44a as "185". Then, each time the position of these viewpoint-specific keyframe images 44a is determined, the difference from the identified frame number "100" is calculated, and the position of the viewpoint-specific keyframe image 44a with the smallest difference, specifically the position "124" of the third viewpoint-specific keyframe image, is deemed to approximate the identified frame number "100". Similar calculations are performed for viewpoints c, d, etc., and for viewpoint c, the position "129" of the third viewpoint-specific keyframe image is deemed to be closest to the identified frame number "100." For viewpoint d, the position "74" of the second viewpoint-specific keyframe image is the closest, but since this is located before the identified frame number "100," the next closest position "134" of the third viewpoint-specific keyframe image is deemed to be closest to the identified frame number "100." In this way, when the position of the viewpoint-specific keyframe image 44a that is closest to each viewpoint is calculated, the viewpoint that is closest to the identified frame number is selected from among them. In the above example, viewpoint b that is closest to the identified frame number "100" is selected. (5-3) When the viewpoint with the closest viewpoint-specific keyframe image 44a (viewpoint b in the above example) is selected, the distance between that viewpoint (viewpoint b) and the line-of-sight information is calculated. Also, the distance between the currently played viewpoint (viewpoint a) and the line-of-sight information is calculated. Then, the viewpoint with the shorter distance is determined as the viewpoint to be played, and viewpoint identification information corresponding to that viewpoint is extracted. In other words, if the currently played viewpoint (viewpoint a) is close to the line-of-sight information, the currently played viewpoint (viewpoint a) is continuously requested. On the other hand, if the line-of-sight information of a viewpoint (viewpoint b) in which a viewpoint-specific keyframe is arranged at a timing close to the currently played frame is closer in coordinates than the currently played viewpoint (viewpoint a), a new video stream 44 from that viewpoint (viewpoint b) is requested.

[0064] (6) In the request processing, the viewpoint identification information and the identified frame number are transmitted to the server 12. Upon receiving the viewpoint identification information and the frame number, the server 12 transmits to the user terminal 14 a video stream 44 that corresponds to the viewpoint identification information and that starts from the frame image corresponding to the identified frame number.

[0065] (7) The user terminal 14 displays the received video stream 44 at the center of the liquid crystal display 18 at a position according to the line-of-sight information.

[0066] [Third embodiment] In the first and second embodiments described above, the user terminal 14 plays back the video stream 44 while receiving it from the server 12, but the present invention is not limited to this playback mode. A third embodiment does not include a server 12, and video streams 44 generated for each viewpoint are stored in the memory 24 of the user terminal 14 in association with viewpoint information. As in the first and second embodiments, the first viewpoint-specific keyframe image 44a of each of the multiple video streams 44 is offset, so that the viewpoint-specific keyframe images 44a are arranged asynchronously with one another in time series. The arrangement timing of the viewpoint-specific keyframe images 44a in the video stream 44 for each viewpoint is stored in the memory 24 of the user terminal 14.

[0067] In this embodiment, the user terminal 14 executes the line-of-sight information acquisition process and the playback process in this order. (1) The gaze information acquisition process is a process in which the CPU 26 of the user terminal 14 acquires gaze information based on the output from the angular velocity sensor 22, and similar to the first and second embodiments, acquires the roll angle (α), pitch angle (β), and yaw angle (γ). (2) In the playback process, viewpoint information whose value is close to the viewpoint information acquired in the above-described viewpoint information acquisition process is extracted, and the video stream 44 corresponding to the extracted viewpoint information is played back.

[0068] As described above, while playing back the video stream 44, the CPU 26 of the user terminal 14 executes the gaze information acquisition process and the playback process in synchronization with the frame rate of the video stream 44 in order to play back the video stream 44 in accordance with the change in gaze caused by the user operating the user terminal 14.

[0069] (4) The line-of-sight information acquisition process is similar to the process (1) above, and is a process of acquiring line-of-sight information (roll angle (α), pitch angle (β), and yaw angle (γ)) based on the output of the angular velocity sensor 22.

[0070] (5) The playback process is a process of determining and playing back the video stream 44 to be played back, in which the CPU 26 of the user terminal 14 selects viewpoint information in which the viewpoint-specific keyframe image 44a is located close in time series at the time of playback, and selects the video stream 44 to be played back based on the selected viewpoint information and the line-of-sight information acquired in (4) above. (5-1) Specifically, the frame number in the video stream 44 being played back (hereinafter referred to as the currently played frame number) is identified. (5-2) Next, the arrangement position of the viewpoint-specific key frame image 44a is calculated based on the arrangement timing for each viewpoint-specific video stream 44 stored in the memory 24. Then, the number of frames up to the calculated arrangement position of the key frame image is calculated for each viewpoint-specific video stream 44. That is, the number of frames from the currently played frame number to the arrangement position of the key frame image is counted, the viewpoint-specific video stream 44 with the key frame image with the smallest count value is identified, and viewpoint information corresponding to that viewpoint-specific video stream 44 is extracted. (5-3) Next, the distance between the extracted viewpoint information and the line-of-sight information acquired in (4) above is calculated. Also, the distance between the viewpoint information of the video stream 44 being played and the line-of-sight information is calculated. Then, the viewpoint with the shorter distance between the two is determined as the viewpoint to be played, and the video stream 44 corresponding to that viewpoint is played. That is, if the viewpoint information of the currently played video stream 44 is close to the user's line of sight information, the currently played video stream 44 continues to be played. On the other hand, if viewpoint information in which a viewpoint-specific keyframe image 44a is arranged at a timing close to the currently played frame is closer in coordinate terms to the user's line of sight information than the currently played viewpoint information, the video stream 44 of that viewpoint information is newly played.

[0071] The present invention is not limited to the above-described embodiment, and may be modified as follows.

[0072] <Variation 1> In the above embodiment, the video stream 44 for each viewpoint is generated based on a group of hemispherical images captured by the camera 10. However, the video stream for each viewpoint may be generated based on a group of spherical images captured by the camera 10. Furthermore, the object is not limited to a hemispherical or spherical object, and may be a virtual sphere of infinite radius viewed from the observation point of the camera 10 with a 45-degree angle of view. In this way, the present invention may be configured to generate a video stream for each viewpoint based on a group of spherical images captured by the camera 10.

[0073] <Variation 2> In the above embodiment, the camera 10 that captures images of the real world is used, but a camera that captures images of a virtual world may also be used.

[0074] <Variation 3> In the above embodiment, the acquisition process and generation process of the server 12 are not essential processes, and the video stream 44 may be prepared in advance in association with multiple viewpoints and stored in the memory 38 prior to distribution to the user terminal 14.

[0075] <Variation 4> In the above embodiment, the video stream 44 is arranged so that the viewpoint-specific keyframe images 44a are asynchronous in time series between the viewpoints. However, asynchronous arrangement may also be achieved by consecutively arranging images unrelated to the viewpoint-specific keyframe images or viewpoint-specific difference frames, such as multiple blank images, at the beginning of the video stream. Furthermore, the arrangement intervals of the viewpoint-specific keyframe images may be different in the video stream between the viewpoints. For example, viewpoint-specific keyframe images for viewpoint a may be arranged every 60 frames, while viewpoint-specific keyframe images for viewpoint b may be arranged every 55 frames, viewpoint-specific keyframe images for viewpoint c every 50 frames, and viewpoint-specific keyframe images for viewpoint d every 45 frames, thereby achieving asynchronous arrangement in time series.

[0076] <Variation 5> In the second embodiment, the interval between key frames is constant, so the array list is defined by the interval value and the offset value, but this is not limited to this. For example, if the interval between key frames is random, a list of the positions (numbers) of key frames in each video stream may be stored in association with viewpoint identification information.

[0077] <Variation 6> In the above embodiment, when the user changes his / her line of sight due to operating the user terminal 14, the video stream 44 containing the viewpoint-specific keyframe image 44a at the closest timing in the time series is selected as the target for distribution or playback, but it is not limited to the closest timing, and the video stream 44 containing the viewpoint-specific keyframe image 44a at a nearby timing may also be selected. Here, "close timing" refers to a frame count that is smaller than the reference frame count number from the delivery timing (playback timing) to the keyframe image of the video stream 44 of the closest viewpoint, focusing on the video stream 44 of the viewpoint that is being delivered or played (nearest viewpoint video stream 44). That is, in each of the viewpoint-specific video streams 44, the placement positions (placement timing) of the viewpoint-specific keyframe images 44a are calculated, and if the placement positions (placement timing) of the viewpoint-specific keyframe images 44a before and after the delivery timing (playback timing) are smaller than the reference frame count number, the video stream 44 is selected as the one that includes the viewpoint-specific keyframe image 44a at a nearby timing in the time series, and the viewpoint information of the selected video stream 44a is compared with the line-of-sight information after the change. The reference frame count may be a frame count obtained by subtracting a number of frames that takes into account a delay in delivery caused by the network environment. [Explanation of symbols]

[0078] 1...Video distribution system 10... Camera 12... Server 14...User terminal 28... Image sensor (imaging section) 30...CPU 34...Memory 38...Memory 40...CPU

Claims

1. a storage step of storing a plurality of video streams generated for each of a plurality of viewpoints defined on a celestial sphere with the camera as an observation point; a distribution step of distributing the video stream to a user terminal; Including, each of the video streams generated for each viewpoint includes an image of the entire celestial sphere such that an image of the viewpoint is included in a center thereof; The delivery step includes: A video distribution method, characterized in that a video stream of a viewpoint other than the nearest viewpoint on the celestial sphere corresponding to the line of sight determined on the user's terminal is distributed.

2. the video stream includes key frame images and difference frame images; 2. The video distribution method according to claim 1, wherein the key frame images from one viewpoint and the key frame images from another viewpoint are asynchronous in time series.

3. The delivery step includes: The video distribution method according to claim 2, characterized in that when the line of sight changes, the video stream of a viewpoint in which the key frame images are arranged nearby in chronological order is set as the video stream of a viewpoint other than the nearest viewpoint.

4. 4. The video distribution method according to claim 2, wherein the video stream from another viewpoint is an array of a plurality of consecutive key frame images.

5. The video distribution method according to claim 2 or claim 3, characterized in that in the video stream of the other viewpoint, the first key frame image is arranged so as to be delayed relative to the first key frame image in the video stream of the one viewpoint.

6. 4. The video distribution method according to claim 2, wherein the key frame images in the video stream from the other viewpoint are arranged at intervals different from those in the video stream from the one viewpoint.

7. a storage unit that stores a plurality of video streams generated for each of a plurality of viewpoints defined on a celestial sphere with the camera as an observation point; a distribution unit that distributes the video stream to a user terminal; Equipped with each of the video streams generated for each viewpoint includes an image of the entire celestial sphere such that an image of the viewpoint is included in a center thereof; The distribution unit A video distribution device that distributes a video stream of a viewpoint other than the nearest viewpoint on the celestial sphere corresponding to the line of sight determined on the user's terminal.

8. Storing a plurality of video streams generated for each of a plurality of viewpoints defined on a celestial sphere with the camera as an observation point; a playback step of playing the video stream at a user's terminal; Including, each of the video streams generated for each viewpoint includes an image of the entire celestial sphere such that an image of the viewpoint is included in a center thereof; The regeneration step includes: A video playback method, comprising: playing back a video stream of a viewpoint other than a nearest viewpoint on the celestial sphere corresponding to a line of sight determined on the user's terminal.

Citation Information

Patent Citations

  • Air conditioner for car

    JP1986049967A

  • Receiver, data stream output apparatus, broadcasting system, control method of receiver, control method of data stream output apparatus, control program and recording medium

    JP2008252832A

  • Video transmitter, video transmission system, video transmission method and program

    JP2017135464A

  • Extended scene view

    US20190052805A1