Image generating device, program, and image generating method
The image generating device estimates head pose and gaze direction to generate a gaze image, addressing the challenge of identifying objects of interest in surveillance images, enhancing monitoring and marketing efficiency.
Patent Information
- Application Number
- JP2022132374
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-08-23
- Publication Date
- 2025-10-20
- Estimated Expiration
- 2042-08-23
AI Technical Summary
Existing surveillance systems struggle to identify the objects of interest or concern to individuals in images captured by surveillance cameras, beyond just estimating gaze direction.
An image generating device that acquires a space image, estimates a person's head pose and gaze direction, and generates a gaze image by projecting a range image onto a sphere around the head coordinates, extracting an image in the line of sight direction from an environmental three-dimensional image.
Enables easy identification of objects in which a person is interested or concerned, facilitating effective monitoring and marketing strategies by revealing the focus of attention.
Smart Images

Figure 0007756608000001 
Figure 0007756608000002 
Figure 0007756608000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image generation device, a program, and an image generation method. [Background technology]
[0002] Thanks to improvements in image processing technology, it has become possible to analyze images captured by surveillance cameras in real time and detect skeletal or pose information of people in the images. However, it is difficult to identify suspicious behavior, which is required for surveillance, using only information about people, such as skeletal or pose information.
[0003] Patent Document 1 discloses a method for estimating the gaze direction from posture information, since the gaze direction is particularly effective in monitoring work. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent Publication No. 2020-182146 Summary of the Invention [Problem to be solved by the invention]
[0005] However, simply estimating the gaze direction is not enough to identify the object that a person in an image captured by a surveillance camera is interested in or concerned about.
[0006] Therefore, one or more aspects of the present disclosure aim to make it possible to easily identify objects in which a person in an image is interested or concerned. [Means for solving the problem]
[0007] An image generating device according to a first aspect of the present disclosure includes an acquisition unit that acquires an image of a space, a gaze direction detection unit that estimates a head pose that is a three-dimensional pose of a head of a person included in the image, and detects a gaze direction of the person using the head pose;before From the environmental image, which is a three-dimensional image of the space, a range image projected onto a sphere around the head coordinates, which are the three-dimensional coordinates of the head, and an image of a predetermined range in the line of sight direction is extracted from the range image; and a gaze image generating unit that generates a gaze image by extracting the gaze position.
[0009] A program according to a first aspect of the present disclosure includes a computer including: an acquisition unit that acquires an image of a space; a gaze direction detection unit that estimates a head pose that is a three-dimensional pose of a head of a person included in the image, and detects a gaze direction of the person using the head pose; and before From the environmental image, which is a three-dimensional image of the space, a range image projected onto a sphere around the head coordinates, which are the three-dimensional coordinates of the head, and an image of a predetermined range in the line of sight direction is extracted from the range image; The extraction is characterized by functioning as a gaze image generating unit that generates a gaze image.
[0011] In an image generating method according to a first aspect of the present disclosure, an acquisition unit acquires an image of a space, a gaze direction detection unit estimates a head pose that is a three-dimensional pose of the head of a person included in the image, and detects the gaze direction of the person using the head pose, and a gaze image generation unit: before From the environmental image, which is a three-dimensional image of the space, a range image projected onto a sphere around the head coordinates, which are the three-dimensional coordinates of the head, and an image of a predetermined range in the line of sight direction is extracted from the range image; The method is characterized by generating a gaze image by extracting the gaze direction. [Effects of the Invention]
[0013] According to one or more aspects of the present disclosure, it is possible to easily identify objects in which a person in an image is interested or concerned. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a block diagram schematically showing the configuration of a viewpoint image generating system according to first and second embodiments. [Figure 2] 1 is a block diagram schematically illustrating a configuration of an image generating device according to a first embodiment. [Figure 3] FIG. 10 is a schematic diagram for explaining processing in a three-dimensional posture estimation unit. [Figure 4] FIG. 10 is a schematic diagram for explaining a process of estimating a gaze direction. [Figure 5]10A and 10B are block diagrams showing an example of a hardware configuration. [Figure 6] 1 is a schematic diagram showing an image represented by video data of a suspicious person captured by a camera; [Figure 7] 10 is a flowchart illustrating an example of processing by the image generating device. [Figure 8] 10 is a schematic diagram for explaining processing in a gaze image generating unit. FIG. [Figure 9] FIG. 10 is a block diagram schematically illustrating the configuration of an image generating device according to a second embodiment. [Figure 10] 10 is a schematic diagram for explaining an example of mapping by a gaze image generating unit. FIG. [Figure 11] FIG. 11 is a block diagram schematically showing the configuration of a viewpoint image generating system according to a third embodiment. [Figure 12] FIG. 11 is a block diagram schematically illustrating the configuration of an image generating device according to a third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0015] Embodiment 1 Next, an embodiment will be described with reference to the drawings, in which the same parts are designated by the same reference numerals. Furthermore, the drawings are schematic, and the ratios of the dimensions may differ from those of reality. Therefore, the specific dimensions should be determined in consideration of the following explanation. Furthermore, the drawings may include parts in which the dimensional relationships or ratios differ from one another.
[0016] FIG. 1 is a block diagram showing a schematic configuration of a viewpoint image generating system 100 including an image generating device 110 according to the first embodiment. The viewpoint image generation system 100 includes a camera 101 , a display device 102 , and an image generation device 110 .
[0017] The camera 101 is an imaging device that functions as an imaging unit that captures a video including a plurality of images. The video data of the video captured by the camera 101 is provided to the image generating device 110.
[0018] The display device 102 functions as a display unit that displays the image generated by the image generation device 110 .
[0019] The camera 101 and the display device 102 may be connected to the image generating device 110 via a network such as the Internet, or may be connected via a connection interface such as a USB (Universal Serial Bass).
[0020] FIG. 2 is a block diagram showing a schematic configuration of an image generating device 110 according to the first embodiment. The image generating device 110 includes an input I / F (Interface) unit 111, a camera input unit 112, a three-dimensional pose estimation unit 113, a head pose estimation unit 114, an environmental image generating unit 116, a gaze image generating unit 117, and an output I / F unit 118.
[0021] The input I / F unit 111 is an input interface that receives input of video data from the camera 101. The received input video data is provided to a camera input unit 112.
[0022] Camera input unit 112 acquires video data from input I / F unit 111. Since the video data includes a plurality of images captured of a space, camera input unit 112 functions as an acquisition unit that acquires the images captured of the space.
[0023] Camera input unit 112 then determines whether or not a person is captured in the image represented by the video data, and if a person is captured, identifies the position of the person in the image. Camera input unit 112 then provides the video data and person detection information indicating the presence or absence of a person and, if a person is present, the coordinates of the person to three-dimensional posture estimation unit 113.
[0024] The three-dimensional posture estimation unit 113 refers to the person detection information and analyzes the image shown in the video data to estimate the three-dimensional posture and coordinates of the person in the image, including the area around the person's face. For example, as shown in Fig. 3, the three-dimensional posture estimation unit 113 can estimate the three-dimensional posture and coordinates of a person P1 in an image by estimating the three-dimensional coordinates of the joints and estimating the interconnections of the joints. Here, the three-dimensional coordinates of the person can be the coordinates of the person's feet. Furthermore, the three-dimensional posture estimation unit 113 may estimate the three-dimensional posture and coordinates of a person by fitting a human body model to a person area in an image.
[0025] Three-dimensional posture estimation section 113 provides head posture estimation section 114 with three-dimensional posture information indicating the estimated posture and three-dimensional coordinate information indicating the estimated coordinates. Note that here, three-dimensional posture estimation unit 113 performs estimation by referring to human detection information from camera input unit 112, but embodiment 1 is not limited to this example. For example, three-dimensional posture estimation unit 113 may directly estimate the presence or absence of a person, as well as the coordinates and posture of the person, from an image represented by video data. In such a case, camera input unit 112 does not need to detect the presence or absence of a person and their coordinates.
[0026] Head pose estimation unit 114 estimates the three-dimensional pose and coordinates of the person's head by referring to the three-dimensional pose information and three-dimensional coordinate information, and detects the forward direction of the person's face area as the gaze direction from the pose and coordinates. Head pose estimation unit 114 then provides gaze coordinate information indicating the three-dimensional coordinates of the head and the gaze direction to gaze image generation unit 117.
[0027] The head posture estimation unit 114 can, for example, refer to the three-dimensional posture information and three-dimensional coordinate information to determine the gaze direction D1 as a direction perpendicular to a line L1 connecting the positions of the left and right ears on the person's head FA and a perpendicular line L2 descending from the line L1 to the chin, and also facing forward of the head FA, as shown in Figure 4.
[0028] It should be noted that head pose estimation section 114 may fit a head model from feature amounts around the head in an image represented by video data, and detect the front direction of the face area as the gaze direction. Furthermore, if the image shown by the video data captures the eye area, such as the eyeballs, in detail, the head posture estimation unit 114 may estimate the gaze direction using information such as the position of the eyeballs or the position of the eyelids. Furthermore, head pose estimation section 114 may correct the estimated gaze direction using past data. Here, correction using past data refers to a state space model such as a Kalman filter.
[0029] As described above, the three-dimensional posture estimation unit 113 and the head posture estimation unit 114 estimate the head posture and head coordinates, which are the three-dimensional posture and coordinates of the head of a person included in an image shown by video data, and function as a gaze direction detection unit 115 that detects the forward direction of the person's face as the gaze direction from the head posture and head coordinates.
[0030] In the first embodiment, gaze direction detection unit 115 detects the gaze direction of a person at a position specified by camera input unit 112. For example, gaze direction detection unit 115 estimates a head pose, which is the three-dimensional pose of the head of a person included in an image, and detects the gaze direction of the person using the head pose. Furthermore, gaze direction detection unit 115 may estimate the gaze direction using a state space model as described above.
[0031] The environmental image generation unit 116 generates an environmental image, which is a three-dimensional image of the space in which the camera 101 is installed. The environmental image is, for example, a three-dimensional map as data that reproduces the space in which the camera 101 is installed in three dimensions.
[0032] Specifically, the environmental image can be generated by capturing images of objects and the like placed in the space in which the camera 101 is installed with one or more cameras from within the space in advance, and then stitching together the images. In other words, the environmental image generation unit 116 can generate the environmental image by stitching together images captured from within the space in which the camera 101 is installed. The environmental image may be updated by capturing images of the space from the inside at any time using a mobile robot equipped with a camera or the like.
[0033] The gaze image generation unit 117 generates a gaze image by extracting an image in the gaze direction from the head coordinates from the environmental image. For example, the gaze image generation unit 117 generates a gaze image by using head coordinates, which are the three-dimensional coordinates of a person's head, to extract an image in the gaze direction from the environmental image, which is a three-dimensional image of the captured space.
[0034] Specifically, gaze image generation unit 117 applies the gaze direction and head coordinates estimated by head gaze estimation unit 140 to the coordinates of the environmental image generated by environmental image generation unit 116, and generates a gaze image that is an image in the gaze direction within the environmental image. Then, gaze image generation unit 117 provides gaze image data indicating the gaze image to output I / F unit 118.
[0035] The output I / F unit 118 is an output interface that outputs the gaze image data to the display device 102. This allows the display device 102 to display a gaze image based on the gaze image data.
[0036] Some or all of the above-described camera input unit 112, three-dimensional pose estimation unit 113, head pose estimation unit 114, environmental image generation unit 116, and gaze image generation unit 117 can be configured, for example, as shown in FIG. 5(A), by a memory 10 and a processor 11 such as a CPU (Central Processing Unit) that executes a program stored in the memory 10. In other words, the image generation device 110 can be realized by a computer. Such a program may be provided via a network or may be provided by being recorded on a recording medium. That is, such a program may be provided, for example, as a program product.
[0037] In addition, some or all of the camera input unit 112, three-dimensional posture estimation unit 113, head posture estimation unit 114, environmental image generation unit 116, and gaze image generation unit 117 can also be configured as a processing circuit 12 such as a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), or an FPGA (Field Programmable Gate Array), as shown in Figure 5(B). As described above, the camera input unit 112, the three-dimensional posture estimation unit 113, the head posture estimation unit 114, the environmental image generation unit 116, and the gaze image generation unit 117 can be configured as a processing circuit network.
[0038] Next, an overview of the processing when the image generating device 110 according to the first embodiment is used as a monitoring device will be described.
[0039] Here, the monitoring device performs a process of presenting the attention points of a suspicious person to the monitor, for example, but the present embodiment is not limited to this example.
[0040] A security person checking the surveillance footage sees a suspicious person in the footage, but the useful information is not the location of the suspicious person, but what the suspicious person is looking at.
[0041] FIG. 6 is a schematic diagram showing an image IM1 represented by video data of a suspicious person P2 captured by a camera. Image IM1 shows suspicious person P2, but does not show the object in the direction of suspicious person P2's gaze. With only image IM1, the monitor cannot determine what point suspicious person P2 is looking at, making it difficult to determine whether suspicious person P2 is a suspicious person, and therefore unable to decide whether to continue monitoring.
[0042] In contrast to this, a process of presenting a point of interest for the suspicious individual P2 when the image generating device 110 according to the first embodiment is used as a monitoring device will be described. FIG. 7 is a flowchart showing an example of processing by the image generating device 110.
[0043] First, the environmental image generation unit 116 generates an environmental image, which is the original data of the gaze image (S10). As described above, the environmental image may be generated by stitching together images captured by one or more cameras, or may be generated by stitching together images of the space captured in advance by a mobile robot or the like. Furthermore, the environmental image may be generated and updated successively, or a generated image may continue to be used.
[0044] It should be noted that the environmental image generation unit 116 does not need to generate environmental images for the entire space, and may, for example, generate only the range that is visible from the aisle where people pass, and omit images of parts that are out of sight, such as the back of a shelf. Furthermore, if there is a defect in the data acquired by a camera or the like, the environmental image generation unit 116 may generate an environmental image by omitting the defective part.
[0045] Next, the camera input unit 112 receives video data from the camera 101 installed in the space via the input I / F unit 111 (S11).
[0046] Next, if a person is captured in the image represented by the received video data, camera input unit 112 identifies the position of the person in the image (S12). Then, camera input unit 112 provides the video data and person detection information indicating the presence or absence of a person and, if a person is present, the coordinates of the person to three-dimensional posture estimation unit 113.
[0047] Next, three-dimensional posture estimation unit 113 refers to the person detection information and analyzes the image shown in the video data to estimate the three-dimensional posture and coordinates of the person in the image, including the area around the person's face (S13). Three-dimensional posture estimation unit 113 provides three-dimensional posture information indicating the estimated posture and three-dimensional coordinate information indicating the estimated coordinates to head posture estimation unit 114.
[0048] Next, head pose estimation unit 114 refers to the three-dimensional pose information and three-dimensional coordinate information to estimate the three-dimensional pose and coordinates of the person's head, and detects the forward direction of the person's face area as the gaze direction from the pose and coordinates (S14). Then, head pose estimation unit 114 provides gaze coordinate information indicating the three-dimensional coordinates of the head and the gaze direction to gaze image generation unit 117.
[0049] Gaze image generation unit 117 applies the gaze direction and head coordinates estimated by head gaze estimation unit 140 to the coordinates of the environmental image generated by environmental image generation unit 116, and generates a gaze image that is an image in the gaze direction within the environmental image (S15). Then, gaze image generation unit 117 outputs gaze image data indicating the gaze image from output I / F unit 118 to display device 102. This allows display device 102 to display the gaze image.
[0050] Finally, the processing of the line-of-sight image generating unit 117 will be described. FIG. 8 is a schematic diagram for explaining the processing in the gaze image generating unit 117. As shown in FIG. As shown in Figure 8, the environmental image generation unit 116 stitches together images captured from inside the space in which the camera 101 is installed to generate an environmental image IM2, which is a three-dimensional map that is data that reproduces the space in three dimensions.
[0051] Then, gaze image generation unit 117 generates range image IM3 by spherically projecting the periphery of head position P3 estimated by head posture estimation unit 114 from environment image IM2 generated by environment image generation unit 116. Note that a known technique such as perspective projection may be used to generate range image IM3 from environment image IM2.
[0052] Next, gaze image generation unit 117 generates gaze image IM4 by extracting from range image IM3 an image of a predetermined range in the gaze direction estimated by head posture estimation unit 114. Here, gaze image IM4 is an image extracted from spherical range image IM3, so gaze image generation unit 117 may perform image processing on the extracted image to make it a planar image, thereby generating gaze image IM4.
[0053] According to the image generating device 110 of the first embodiment described above, an observer can check the gaze image IM4 on the display device 102 and thereby understand the object that the suspicious person in the video captured by the camera 101 is looking at.
[0054] According to the image generating device 110 of embodiment 1, the monitor can grasp the objects of interest of the suspicious person, so that the areas where the crime prevention effect needs to be improved can be narrowed down and measures can be taken, thereby reducing the costs involved in crime prevention.
[0055] Embodiment 2 As shown in FIG. 1, a viewpoint image generating system 200 according to the second embodiment includes a camera 101, a display device 102, and an image generating device 210.
[0056] The camera 101 and the display device 102 of the viewpoint image generating system 200 in the second embodiment are similar to the camera 101 and the display device 102 of the viewpoint image generating system 100 in the first embodiment.
[0057] FIG. 9 is a block diagram showing a schematic configuration of an image generating device 210 according to the second embodiment. The image generation device 210 includes an input I / F unit 111, a camera input unit 112, a three-dimensional pose estimation unit 113, a head pose estimation unit 114, an environmental image generation unit 116, a gaze image generation unit 217, an output I / F unit 118, and a data storage unit 219.
[0058] Input I / F unit 111, camera input unit 112, three-dimensional pose estimation unit 113, head pose estimation unit 114, environmental image generation unit 116, and output I / F unit 118 of image generation device 210 according to embodiment 2 are similar to input I / F unit 111, camera input unit 112, three-dimensional pose estimation unit 113, head pose estimation unit 114, environmental image generation unit 116, and output I / F unit 118 of image generation device 110 according to embodiment 1. Gaze direction detection unit 115 according to embodiment 2 is also similar to gaze direction detection unit 115 according to embodiment 1.
[0059] As will be described later, the data accumulation unit 219 accumulates the gaze coordinate information from the gaze image generation unit 217. For example, the data accumulation unit 219 accumulates, as gaze coordinate information, a plurality of head coordinates, which are the three-dimensional coordinates of a person's head, and a plurality of gaze directions, which are the person's gaze directions. The data storage unit 219 can be realized by a storage device such as a hard disk drive (HDD), a solid state drive (SSD), a volatile memory, or a non-volatile memory.
[0060] The gaze image generation unit 217 stores the plurality of gaze directions and the plurality of head coordinates estimated by the head gaze estimation unit 140 in the data storage unit 219 as gaze coordinate information. Then, the gaze image generation unit 217 generates a gaze image by extracting from the environmental image an image of an area where points corresponding to the multiple gaze directions are concentrated based on the multiple head coordinates stored in the data storage unit 219.
[0061] For example, the gaze image generation unit 217 applies the gaze direction and head coordinates indicated by the accumulated gaze coordinate information to the coordinates of the environmental image generated by the environmental image generation unit 116, and identifies an area in the environmental image where points corresponding to the gaze direction from the head coordinates are concentrated.The gaze image generation unit 217 then generates a gaze image, which is an image of the identified area in the environmental image.The gaze image generation unit 217 then provides gaze image data indicating the gaze image to the output I / F unit 118.
[0062] For example, when an amount of gaze coordinate information equal to or greater than a predetermined threshold is recorded in the data storage unit 219, the gaze image generation unit 217 reads the gaze coordinate information from the data storage unit 219. Then, the gaze image generation unit 217 maps a location corresponding to the gaze direction from the head coordinates indicated by the gaze coordinate information onto the environmental image generated by the environmental image generation unit 116.
[0063] FIG. 10 is a schematic diagram for explaining an example of mapping by the line-of-sight image generating unit 217. As shown in FIG. As shown in Figure 10, the environmental image generation unit 116 stitches together images captured from inside the space in which the camera 101 is installed to generate an environmental image IM2, which is a three-dimensional map that is data that reproduces the space in three dimensions.
[0064] The line-of-sight image generating unit 217 maps the points PO1 to PO7 corresponding to the line-of-sight direction from the coordinates of the head indicated by the line-of-sight coordinate information onto the environmental image IM2. Then, the line-of-sight image generating unit 217 specifies an area where the points PO1 to PO7 are concentrated. Here, the line-of-sight image generating unit 217 specifies an area TE of a predetermined size centered on the center of gravity of the three-dimensional positions of the points PO1 to PO7. The line-of-sight image generating unit 217 extracts the image of the area TE and sets the extracted image as the line-of-sight image.
[0065] Note that the gaze image generating unit 217 may specify a plurality of regions using the K-means method, etc. In this case, a plurality of gaze images will be generated.
[0066] As described above, in the second embodiment, the image generating device 210 can be used as a data providing device that accumulates the estimation results of the head poses of people over a certain period of time in, for example, a commercial facility, and presents the locations where many gazes are focused.
[0067] As described above, according to the second embodiment, the image generating device 210 records the gaze direction of customers in a commercial facility or the like where an unspecified number of customers come and go in the data storage unit 219, and generates a gaze image, thereby enabling marketing personnel at the commercial facility or the like to visually grasp the objects of interest to customers.
[0068] Therefore, according to the second embodiment, by accumulating the gaze directions of an unspecified number of people, it is possible to display products in locations where customers are likely to notice them, or to display products in accordance with the gaze directions of each target demographic.
[0069] Embodiment 3 FIG. 11 is a block diagram showing a schematic configuration of a viewpoint image generation system 300 including an image generation device 310 according to the third embodiment. The viewpoint image generation system 300 includes a camera 101 , a display device 102 , an image generation device 310 , and an input device 303 .
[0070] The camera 101 and the display device 102 of the viewpoint image generation system 300 in the third embodiment are similar to the camera 101 and the display device 102 of the viewpoint image generation system 100 in the first embodiment.
[0071] The input device 303 functions as an input unit that receives an input of an instruction to select a person for whom a gaze image is to be generated from an image included in the video data from the camera 101 . For example, the user can display an image included in the video data from the camera 101 on the display device 102 and select a person appearing in the displayed image using the input device 303 .
[0072] FIG. 12 is a block diagram showing a schematic configuration of an image generating device 310 according to the third embodiment. The image generation device 310 includes an input I / F unit 311, a camera input unit 312, a three-dimensional pose estimation unit 113, a head pose estimation unit 114, an environmental image generation unit 116, a gaze image generation unit 117, and an output I / F unit 118.
[0073] Three-dimensional pose estimation unit 113, head pose estimation unit 114, environment image generation unit 116, gaze image generation unit 117, and output I / F unit 118 of image generation device 310 according to embodiment 3 are similar to three-dimensional pose estimation unit 113, head pose estimation unit 114, environment image generation unit 116, gaze image generation unit 117, and output I / F unit 118 of image generation device 110 according to embodiment 1. Gaze direction detection unit 115 according to embodiment 3 is also similar to gaze direction detection unit 115 according to embodiment 1.
[0074] The input I / F unit 311 is an input interface that receives input of video data from the camera 101 and input of instructions from the input device 303. The received input video data and instructions are provided to a camera input unit 312. In the third embodiment, the input I / F unit 311 receives an instruction from the input device 303 to select a person from an image included in the video data from the camera 101, for whom a gaze image is to be generated.
[0075] The camera input unit 312 identifies the position of a person for which a gaze image is to be generated from the image represented by the video data. Then, the camera input unit 312 provides the video data and person detection information indicating the coordinates of the person to the three-dimensional posture estimation unit 113.
[0076] For example, camera input unit 312 may send video data from camera 101 from output I / F unit 118 to display device 102, and receive an instruction from the user via input device 303 to select a person for whom a gaze image is to be generated from an image included in the video data. In other words, camera input unit 312 accepts the selection of a person from the image shown in the video data. Then, gaze direction detection unit 115 detects the gaze direction of the selected person.
[0077] The input device 303 described above can be realized by a keyboard, a mouse, or the like. In the third embodiment, for example, the display device 102 and the input device 303 may be provided in a mobile terminal such as a smartphone. In this case, the display device 102 and the input device 303 may be realized by a touch panel. Furthermore, the mobile terminal may include the camera 101.
[0078] In such a case, the user of the mobile terminal can see a line-of-sight image of the person captured by the camera 101 looking in the direction of the gaze at a tourist spot or landmark.
[0079] In addition, at tourist spots or landmarks, image data that allows the generation of three-dimensional maps may be provided by tourists, etc., and the environmental image generation unit 116 may generate an environmental image using images represented by image data uploaded to the Internet.
[0080] The display device 102 in the third embodiment may be a general monitor, or may be a device for AR (Augumented Reality) / VR (Virtual Reality) such as a mobile terminal or an eyeglass-type device.
[0081] When the gaze image display unit 160 can provide immersive video images such as a glasses-type device, the generated gaze image may be an omnidirectional image, or a display indicating the gaze direction may be provided within the omnidirectional image.
[0082] As described above, according to the third embodiment, it is possible to view the scenery seen by people at a tourist spot or landmark from a remote location, and to relive the experience as if you were actually there. [Explanation of symbols]
[0083] 100,200,300 viewpoint image generation system, 101 camera, 102 display device, 303 input device, 110,210,310 image generation device, 111,311 input I / F unit, 112,312 camera input unit, 113 three-dimensional pose estimation unit, 114 head pose estimation unit, 115 gaze direction detection unit, 116 environmental image generation unit, 117,217 gaze image generation unit, 118 output I / F unit, 219 data storage unit.
Claims
1. an acquisition unit that acquires an image of a space; a gaze direction detection unit that estimates a head pose, which is a three-dimensional pose of a head of a person included in the image, and detects a gaze direction of the person using the head pose; a gaze image generating unit that generates a range image projected spherically around head coordinates, which are the three-dimensional coordinates of the head, from an environmental image, which is a three-dimensional image of the space, and generates a gaze image by extracting an image of a predetermined range in the gaze direction from the range image. An image generating device characterized by:
2. Further comprising an environmental image generating unit that generates the environmental image by stitching together images captured from the inside of the space.
2. The image generating device according to claim 1, wherein:
3. The gaze direction detection unit estimates the gaze direction using a state space model.
2. The image generating device according to claim 1, wherein:
4. The acquisition unit identifies a position of the person included in the image, The gaze direction detection unit detects the gaze direction of the person at the specified position.
2. The image generating device according to claim 1, wherein:
5. the acquisition unit receives a selection of the person from the image; The gaze direction detection unit detects the gaze direction of the selected person.
2. The image generating device according to claim 1, wherein:
6. Computer, an acquisition unit that acquires an image of the space; a gaze direction detection unit that estimates a head pose, which is a three-dimensional pose of a head of a person included in the image, and detects a gaze direction of the person using the head pose; and The system functions as a gaze image generating unit that generates a range image projected spherically around head coordinates, which are the three-dimensional coordinates of the head, from an environmental image, which is a three-dimensional image of the space, and generates a gaze image by extracting an image of a predetermined range in the gaze direction from the range image. A program characterized by.
7. The acquisition unit acquires an image of the space, a gaze direction detection unit estimating a head pose, which is a three-dimensional pose of a head of a person included in the image, and detecting a gaze direction of the person using the head pose; The gaze image generation unit generates a range image projected spherically around head coordinates, which are the three-dimensional coordinates of the head, from an environmental image, which is a three-dimensional image of the space, and extracts an image of a predetermined range in the gaze direction from the range image, thereby generating a gaze image. An image generation method characterized by:
Citation Information
Patent Citations
View point free type image display apparatus and method, storage medium, computer program, broadcast system, and attached information management apparatus and method
JP2003235058A
Monitoring apparatus
JP2008288707A
Image processing unit, method of processing image, and program
JP2010268157A
Monitoring device and monitoring method
JP2020182146A
Method and apparatus for managing information
KR1020150093532A