Composite video generation system
The composite image generation system addresses limitations by incorporating depth and three-dimensional image management to create versatile and interactive composite images, enhancing user experience and application versatility.
Patent Information
- Application Number
- JP2024113386
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-16
- Publication Date
- 2026-01-28
AI Technical Summary
Existing composite image generation systems are limited in their applications and do not effectively utilize depth information and three-dimensional images to create realistic and versatile composite images.
A composite image generation system that includes a real image acquisition unit capturing depth information, a measurement unit for user position and orientation, a three-dimensional image management unit, and a synthesis unit that generates composite images using depth and measurement information, displayed on a wearable device.
Enables a wider range of applications by providing realistic, three-dimensional composite images that can be viewed from different angles and perspectives, allowing users to interact with obstructed views and utilize real-time streaming videos.
Smart Images

Figure 2026013156000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a composite image generation system for displaying to a user a composite image formed by combining a plurality of images. [Background technology]
[0002] The applicant has conceived of the inventions disclosed in the following Patent Documents 1 to 3 as systems for displaying to a user a composite image formed by combining a real image of the user's surroundings with raw images. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 6717486 [Patent Document 2] Patent No. 6991494 [Patent Document 3] Japanese Patent Application Laid-Open No. 7157271 Summary of the Invention [Problem to be solved by the invention]
[0004] An object of the present invention is to provide a synthetic image generation system that utilizes the techniques disclosed in the above-mentioned patent documents and that can be used for a wider variety of applications. [Means for solving the problem]
[0005] The present invention, which has been made to solve the above-mentioned problems, is a composite image generation system for displaying a composite image to a user, which is formed by combining multiple images, and is characterized by comprising at least a real image acquisition unit that acquires a real image, which is an image taken in the approximate direction of the user's line of sight, in a manner that includes depth information; a measurement unit that acquires measurement information including the position and orientation of the user; a three-dimensional image management unit that manages a three-dimensional image based on multiple two-dimensional images taken by multiple cameras; a synthesis unit that generates a composite image by synthesizing a processed image in which any portion of the real image has been removed based on the depth information and the three-dimensional image based on the measurement information; and a display unit that can be worn on the user's head and displays the composite image to the user. In addition, the present invention may be configured such that the three-dimensional image management unit stores multiple three-dimensional images of the same subject taken at different times, and the synthesis unit generates a synthetic image using one three-dimensional image selected by the user from the multiple three-dimensional images. In addition, the present invention may be configured so that the three-dimensional image management unit stores multiple three-dimensional images with different subjects, and the synthesis unit uses multiple three-dimensional images selected by the user to generate the synthesized image. The present invention may also be configured such that the synthesis unit uses a plurality of identical three-dimensional images and generates a synthesized image with the orientation of each three-dimensional image changed. In addition, in the present invention, streaming video generated in real time may be used as the three-dimensional video used in the synthesis unit. [Effects of the Invention]
[0006] According to the present invention, it is possible to provide a synthetic image generation system that can be used for a wider variety of applications. [Brief explanation of the drawings]
[0007] [Figure 1] 1 is a functional block diagram of a synthetic video generation system according to the present invention; [Figure 2]An illustration of the removal process for real-world footage. [Figure 3] FIG. 2 is a diagram showing the system configuration on the subject side. [Figure 4] FIG. 1 is a diagram showing a system configuration on the side where real images are acquired. [Figure 5] An illustration of the composite image. [Figure 6] FIG. 10 is an image diagram showing the transition of a synthesized image. DETAILED DESCRIPTION OF THE INVENTION
[0008] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. [Example]
[0009] <1> Overall configuration (Fig. 1) The synthetic image generation system shown in FIG. 1 is configured to include at least a real image acquisition unit 10, a measurement unit 20, a three-dimensional image management unit 30, a synthesis unit 40, and a display unit 50. Each part can be realized by any combination of hardware and software, and can also be configured as a separate device, or as an integrated device incorporating multiple parts. Furthermore, in the present invention, there are no particular limitations on the number of each part, the location of each part, or the number of images used for composition. Each part will be described in detail below.
[0010] <2> Real-world image acquisition unit (Figure 1) The real image acquisition unit 10 has a function of acquiring a real image 11 in the approximate direction of the user's line of sight, with depth information. This "real image 11" is an image captured in real space that includes the approximate line of sight of the user, that is, a real-time image that shows the space around the user. The real image acquisition unit 10 can use a camera built into a head-mounted display (hereinafter also referred to as "HMD") worn by the user, a camera attached externally to the HMD, or any other camera that can be worn by the user. Cameras that can be used as the real image acquisition unit 10 include so-called depth cameras or depth cameras that have a built-in depth sensor. In the present invention, the real image acquisition unit 10 may be a camera that can capture images at least in the approximate direction of the user's line of sight, and does not exclude cameras that can capture images of the entire periphery of the user.
[0011] <3> Measurement section (Fig. 1) The measurement unit 20 has a function of acquiring measurement information 21 of the user. This "measurement information 21" includes at least information that allows the user's orientation and position to be recognized. A synthesis unit 40, which will be described later, generates a synthetic video 42 based on this measurement information 21. The measurement unit 20 can use a group of sensors such as an angle sensor, an acceleration sensor, a gyroscope, a motion capture device, a GPS device, and the like. The measurement unit 20 may be of any type, such as a type installed in the living space where the user is present, a type built into the HMD worn by the user, a type worn separately on the user's body, or a type held by the user.
[0012] <4> 3D video control unit (Fig. 1) The 3D video management unit 30 has a function of managing at least one 3D video 31 . This "three-dimensional video 31" is a stereoscopic video consisting of still images or videos (including slow-motion videos and time-lapse videos) containing a 3D model generated based on multiple two-dimensional images obtained by capturing a subject (including people, objects, structures, and anything else) with multiple cameras. Three-dimensional video 31 is also called "volumetric video."
[0013] In the present invention, the method for generating the 3D video 31 is not particularly limited. In addition, in the present invention, the type of subject from which the three-dimensional video 31 is acquired is not particularly limited, and may include a person, an object, a structure, or a combination of these. Furthermore, in the present invention, when generating the 3D video 31, there is no limitation as to whether or not extraction processing for removing elements other than the subject is performed. Furthermore, in the present invention, the "management" of the 3D video 31 by the 3D video management unit 30 includes the generation of the 3D video 31 based on a plurality of received 2D videos, and the reception and storage of the 3D video 31 that is being generated or has already been generated. Therefore, the 3D video 31 handled by the 3D video management unit 30 may be a video that has already been shot, or may be a streaming video that is generated in real time at the stage when the composite video is displayed to the user.
[0014] <5> Synthesis section (Figure 1) The synthesis unit 40 has the function of generating an image (processed image) in which at least any part of the real image 11 has been removed based on depth information (depth-based removal processing), and the function of synthesizing the processed image with a three-dimensional image 41 acquired and / or stored by the three-dimensional image management unit 30 based on measurement information 21 acquired by the measurement unit 20. The synthesis unit 40 can be configured by controlling hardware such as a CPU, memory, and GPU through software. In the present invention, the synthesis unit 40 may further combine and synthesize other raw images (VR images, CG images, 360-degree images, etc.) in addition to the 3D images 31 and the processed images.
[0015] <5.1> Example of removal process (Figure 1) In the present invention, the removal processing performed on the real image 11 includes at least depth-related removal processing, and may also include chroma-related removal processing as necessary. "Depth-based removal processing" refers to a process of removing a real image 11 having depth information based on the value of the depth information, such as a process of removing parts having depth information beyond a predetermined distance. "Chroma system removal processing" is a process that performs removal based on the value of color space information on the image after depth system removal processing has been performed on the real image 11. More specifically, it is a process that removes the part that corresponds to the homogeneous color of the background material placed in the user's living space and is reflected within the field of view of the real image 11, making it possible to composite other images into the removed part. For details of these processes, please refer to the above-mentioned Patent Documents 1 to 3.
[0016] <5.2> Processed image generation image FIG. 2 shows an image diagram of the real image 11 after each removal process. The real image 11 shown in Fig. 2(a) is an image captured by a real image capture unit 10 provided in an HMD worn by a user, and includes the user's hands and the background in the user's line of sight. By performing depth-based removal processing on this real image 11, for example, to remove areas with depth information of 50 cm or more, a processed image 41a can be generated in which only the user's hands are extracted, as shown in Fig. 2(b). If the user's hands are sufficiently extracted in the processed image 41a after this depth-based removal process, then this processed image 41a can be used as the final processed image 41 to be used for compositing. For example, if more detailed extraction is desired, further chroma-based removal process can be performed on the processed image 41a to remove the background color area remaining around the outline of the user's hands, thereby generating the processed image 41b shown in Figure 2(c), and this processed image 41b can then be used as the final processed image 41.
[0017] <6> Display unit (Figure 1) The display unit 50 is a device having a function of displaying the composite image generated by the composition unit 40 to the user. The display unit 50 can be an HMD, goggles, or the like that can be worn by the user, or alternatively, a mobile terminal such as a smartphone that can be housed in a housing and turned into goggles can also be used.
[0018] <7> System configuration example (Fig. 3, Fig. 4) An example of the configuration of a synthetic video generation system according to the present invention will be described with reference to FIGS. In this embodiment, the configurations of the subject-side base (FIG. 3) and the user-side base (FIG. 4) will be described separately.
[0019] <7.1> Subject's base (Fig. 3) Figure 3 shows the system configuration at the site where the 3D images are acquired. At the base, there is provided a subject C from which a three-dimensional image 31 is to be acquired, and multiple cameras A installed around the subject C. Then, each camera A captures an image of the subject C, and obtains a respective two-dimensional image A1. Each two-dimensional image A1 is sent to an information processing device B connected to each camera A and is used to generate a three-dimensional image 31. Then, the three-dimensional video 31 generated by the information processing device B is sent to an information processing device F provided at the user's base shown in FIG. In this embodiment, a three-dimensional image 31 is generated using a bookshelf and a person standing in front of the bookshelf as subject C. If it is not desired to include the wall behind the bookshelf in the three-dimensional image 31, a background material G for chromakeying may be placed on the wall, and the color portion corresponding to the background material G may be removed from each of the two-dimensional images A1 before the three-dimensional image 31 is generated.
[0020] <7.2> User base (Figure 4) Figure 4 shows the system configuration at the user's site. At this base, a head-mounted display E worn by a user D and an information processing device F connected to the head-mounted display E are provided.
[0021] <7.2.1> Head-mounted display (Fig. 4) The head mounted display E is a device that functions as the real image acquisition unit 10, the measurement unit 20, and the display unit 50 in the synthetic image generation system according to the present invention. In this embodiment, a depth camera attached externally to the front of the head-mounted display E is used as the real image acquisition unit 10, a group of sensors such as an angle sensor, an acceleration sensor, and a gyroscope built into the head-mounted display E is used as the measurement unit 20, and a display built into the head-mounted display E is used as the display unit 50. The image (real image 11) captured by the depth camera functioning as the real image acquisition unit 10 utilizes the parallax of the images captured by the two cameras to provide depth information, and is an image that is linked to the line of sight of user D, so that background material G arranged around user D and user D's hands are appropriately reflected within the field of view of the image. It goes without saying that the background material G is a member that is necessary only when chroma removal processing is performed.
[0022] <7.2.2> Information processing device (Fig. 4) The information processing device F is a device that functions as the three-dimensional video management unit 30 and the synthesis unit 40 in the composite video generation system according to the present invention. The information processing device F stores the three-dimensional image 31 sent from the information processing device B shown in Figure 3, generates a processed image by performing a predetermined removal process on the real image 11 sent from the head-mounted display E, and sends a composite image 42 created by combining this processed image with the three-dimensional image 31 to the head-mounted display E so that the user D can view it. The positioning of the three-dimensional image 31 and the processed image 41 when they are combined may be performed by a known method such as using markers (not shown) provided at each location as a reference.
[0023] <8> Image of synthetic image generation (Figures 5 and 6) The image of how a composite video is generated in the example system configuration shown in FIGS. 3 and 4 will be explained with reference to FIGS. 5 and 6. FIG.
[0024] <8.1> Image before user movement (Fig. 5) FIG. 5 is an image diagram of the composite image before the user moves. The three-dimensional image 31 acquired at the subject's base includes subject C (Figure 5(a)), and the processed image 41 generated based on the real image 11 acquired at the user's base includes both hands of user D and the stick he is holding with both hands (Figure 5(b)). In a composite image 42 obtained by combining these images based on the measurement information of user D, subject C is reflected behind both hands of user D (FIG. 5(c)).
[0025] <8.2> Image after user movement (Fig. 6) FIG. 6 is a diagram showing an image of the transition of the composite image after the user moves. When user D moves diagonally forward to the right on the composite image 42 from the standing position shown in FIG. 4 as shown in FIG. 6(a) and looks at subject C, the composite image 42 shown in FIG. 5 transitions to the image shown in FIG. 6(b) in conjunction with changes in user D's position and orientation based on user D's measurement information. At this time, since the user D's hands are lowered, the hand of the user D is not displayed in the composite image 42. In this way, user D can view subject C, which was photographed at the location shown in FIG. 3, in a three-dimensional manner in conjunction with user D's own position and orientation. 3, which shows the system configuration at the subject's base where the 3D video is acquired, for example, if a camera A is added and installed so as to capture the space between the bookshelf and the person in subject C, and 3D video 31 is generated using 2D video A1 captured by camera A, when user D moves further to the right in FIG. 6 so as to approach the left hand of subject C, the space between the person and the bookshelf in subject C can be viewed in three dimensions. As a result, from the position where user D was originally standing (FIGS. 4 and 5) or the position to which he moved shown in FIG. 6, the part of the bookshelf that was blocked by the 3D model person standing in front of subject C's bookshelf was not visible, but by user D moving within the space and peering between the person and the bookshelf while avoiding the obstructing 3D model person, he can view in three dimensions what is placed in the part of the bookshelf that was blocked.
[0026] <9> Other synthesis examples The composite video generation system according to the present invention may be configured to use a plurality of 3D videos 31 in one composite video 42 to obtain a composite video 42 with a larger amount of information. The plurality of 3D images 31 included in the composite image 42 may be different 3D images 31, or the same 3D image 31 may be used. Furthermore, when the same three-dimensional images 31 are used, by changing the coordinates and orientation of each three-dimensional image 31 and combining them in the composite image 42, it is possible to simultaneously view the same subject C in different angles of view.
[0027] <10> Action and effect The composite video system according to the present invention can achieve the following effects, for example. (1) The composite image contains a three-dimensional image generated from multiple two-dimensional images taken by multiple cameras, allowing the user to experience a more realistic composite image. (2) In the synthetic image, even if there is a place in a particular field of view that the user cannot see because it is obstructed by a 3D model such as a person or object contained in the three-dimensional image, the user can still see the place by moving or turning their head to avoid the obstructing 3D model. (3) By moving so that the user overlaps with the 3D model in the 3D image, the user can experience a view that is as if they are one with the 3D model. (4) By storing multiple 3D images of the same subject taken at different times, and allowing the user to select one of these to be synthesized and switch between them as needed, this system is ideal for applications such as confirmation work performed in chronological order. (5) By using multiple 3D images with different subjects in a single composite image, it contributes to further expanding the scope of applications. (6) By preparing multiple identical 3D images in one composite image and combining each 3D image in a different orientation, the user can simultaneously view the same subject at different angles of view without having to move. (7) By using streaming video generated in real time for the 3D video, it is possible to display to the user a real-time video of the subject from which the 3D video was obtained. For example, when a skilled engineer in a remote location provides guidance to an assembly worker at a construction site, by displaying a composite video in which a 3D video is generated from streaming video captured in real time of the construction site, the user can provide guidance to the assembly worker while recognizing the construction site in three dimensions, even from a remote location, as if they were actually in the space of the construction site.
[0028] <11> Usage examples The synthetic video system according to the present invention can be used for the following purposes, for example.
[0029] (1) Creation of teaching materials and manuals By saving the work of a skilled technician as a three-dimensional image and allowing a trainee, as a user of the composite image system of the present invention, to view the three-dimensional image 31 within the composite image 42 from any angle, it is possible to provide teaching materials with a higher learning effect and easy-to-understand work manuals. For example, <9> As explained above, if multiple 3D images 31 containing the work content of a skilled technician are included in one composite image 42 at different angles of view, it is beneficial in that the user can simultaneously check the work of the skilled technician from multiple angles of view without having to move significantly. In addition, the user can experience the field of view of the skilled technician by moving so that the user is superimposed on the 3D model of the skilled technician in the 3D image. In addition, the skilled engineer can move so that the 3D model of the assembly worker in the 3D image overlaps with the 3D model, allowing the skilled engineer to experience the perspective of the assembly worker. For example, when the composite video system according to the present invention is used for remote supervision of an assembly site of a cable connection, an assembly worker or a surrounding work object is set as the subject, and if the user moves to a position in the composite video space that overlaps with the position of the assembly worker, the work status of the cable connection that the assembly worker is looking at can be confirmed from the perspective of the assembly worker. Also, if the user moves to the position of the cable connection in the composite video space, the user can actually measure the dimensions of the cable connection on the subject side with a tape measure that the user has at hand.
[0030] (2) Creating work records By storing the three-dimensional image 31 of the synthetic image system of the present invention instead of photographs taken and recorded with a normal camera for recording construction work, etc., it is possible to create work records that are easier to review after the fact. In addition, by storing multiple three-dimensional images 31 of the same subject taken at different times and allowing the user to switch between the three-dimensional images 31 as appropriate, it becomes possible to check the progress of work and the time when a problem occurred. Furthermore, if the three-dimensional image 31 is saved as a work record of the construction site, when the user looks back on past assembly states, the user can check the work record with a sense of realism as if they were inside the space of the past construction site. Furthermore, by moving so that the user overlaps the 3D model of the assembly worker in the 3D image, the user can experience the view of the assembly worker. Therefore, for example, a skilled engineer can become the user and point out points he or she notices by experiencing the view of the assembly worker, or the user can move around the space and check the assembly worker's hand movements from different directions to provide guidance on points he or she notices. Furthermore, within the range of 3D image 31, the user can also see the surrounding environment that the assembly worker cannot see during assembly work.
[0031] (3) Consider furniture layout when moving The current entire room is saved as a three-dimensional image 31, and by combining a VR image of the new home with the three-dimensional image 31 and processed image 41, it is possible to confirm an image of how the current furniture layout would look if it were applied to the new home as is. [Explanation of symbols]
[0032] 10: Real image acquisition unit 11: Reality footage 20: Measurement section 30: 3D video management department 31: 3D images 40: Synthesis section 41: Processed video 42: Composite image 50: Display section A: Camera B: Information processing device C: Subject D: User E: Head-mounted display F: Information processing equipment G: Background material
Claims
1. A composite video generation system for displaying a composite video to a user, the composite video being formed by combining a plurality of videos, comprising: a real image acquisition unit that acquires a real image that is an image captured in a general line of sight direction of the user in a manner that includes depth information; a measurement unit that acquires measurement information including the position and orientation of the user; a three-dimensional image management unit that manages three-dimensional images based on a plurality of two-dimensional images captured by a plurality of cameras; a synthesis unit that synthesizes a processed image, in which an arbitrary portion has been removed from the real image based on the depth information, and the 3D image based on the measurement information, to generate a synthesized image; a display unit that can be worn on the user's head and that displays the composite image to the user; characterized by comprising at least Synthetic image generation system.
2. the 3D image management unit stores a plurality of 3D images of the same subject taken at different times, The synthesis unit generates a synthetic image using one 3D image selected by a user from the plurality of 3D images. The synthetic image generation system according to claim 1 .
3. The 3D image management unit stores a plurality of 3D images having different objects, The synthesis unit uses a plurality of 3D images selected by the user to generate the synthetic image. The synthetic image generation system according to claim 1 .
4. The synthesis unit uses a plurality of identical 3D images and generates a synthetic image by changing the orientation of each 3D image. The synthetic image generation system according to claim 1 .
5. The three-dimensional image used in the synthesis unit is a streaming image generated in real time. The synthetic image generation system according to claim 1 .
Citation Information
Patent Citations
Augmented virtual space provision system
JP6717486B1
Head-mounted display with improved visibility
JP6991494B1
JP7157271A