Projection system

The projection system improves behavior estimation accuracy by projecting a virtual person into a virtual space, correlating their actions with object configurations, addressing the limitations of direct two-dimensional image analysis.

JP2025174368APending Publication Date: 2025-11-28TOYOTA JIDOSHA KK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024080692
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing methods for analyzing human behavior in a target space from camera images lack accuracy, particularly in estimating three-dimensional postures and behaviors of individuals.

Method used

A projection system that reconstructs a three-dimensional environment from captured images, projects a virtual person representing a real person into a virtual space, and estimates behavior based on the relationship between the virtual person and object configurations in the virtual space.

Benefits of technology

Enhances the accuracy of behavior estimation by projecting a virtual person into a virtual space, allowing for precise analysis of real person actions by correlating their behavior with the virtual environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025174368000001_ABST
    Figure 2025174368000001_ABST
Patent Text Reader

Abstract

To improve accuracy of human action estimation based on an image captured by a camera installed in a target space.SOLUTION: A projection system reconstructs a three-dimensional environment of a real target space on the basis of a group of images captured in the real target space. The projection system generates a virtual target space representing the three-dimensional environment. The projection system detects a real human shown in an image captured by a real camera installed in the real target space, and estimates a three-dimensional pose of the real human. The projection system projects a virtual human having a three-dimensional pose and representing the real human onto the virtual target space. The projection system estimates an action of the real human in the real target space by estimating an action of the virtual human in the virtual target space on the basis of a relationship between the virtual human having the three-dimensional pose and a configuration of objects in the virtual target space.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a technique for projecting a person in a real target space into a virtual target space. [Background technology]

[0002] Patent Document 1 discloses a gaze estimation system. The gaze estimation system acquires a series of images showing the face of a subject. Then, the gaze estimation system estimates the gaze position of the subject from the images including the face by using a trained model.

[0003] Non-Patent Document 1 discloses a technique for performing 3D reconstruction of an object space by 3D Gaussian Splatting and recognizing linguistic descriptions of objects in the object space by semantic segmentation. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2022-187547 [Non-patent literature]

[0005] [Non-Patent Document 1] Qin et al., “LangSplat: 3D Language Gaussian Splatting”, arXiv:2312.16084v2 [cs.CV], 31 Mar 2024. Summary of the Invention [Problem to be solved by the invention]

[0006] Images captured by a camera installed in a target space can be used to analyze the target space. For example, the behavior of a person present in the target space can be estimated based on the images captured by the camera installed in the target space. It is desirable to improve the accuracy of human behavior estimation based on images captured by a camera. [Means for solving the problem]

[0007] One aspect of the present disclosure relates to a projection system. The projection system comprises one or more processors. The one or more processors reconstruct a three-dimensional environment of the real object space based on a group of images captured in the real object space, and generate a virtual object space that represents the three-dimensional environment of the real object space. The one or more processors recognize configurations and linguistic descriptions of objects in the virtual object space by applying semantic segmentation to the three-dimensional environment. The one or more processors detect real people appearing in images captured by a real camera installed in a real target space. The one or more processors estimate a three-dimensional pose of the real person based on the images. The one or more processors project a virtual person having a three-dimensional pose and representing a real person into a virtual object space. The one or more processors perform a behavior estimation process to estimate the behavior of the real person in the real target space by estimating the behavior of the virtual person in the virtual target space based on the relationship between the virtual person having a three-dimensional pose and the configuration of objects in the virtual target space. [Effects of the Invention]

[0008] According to the present disclosure, the three-dimensional posture of a real person in a real target space is estimated, and a virtual person representing the real person having the three-dimensional posture is projected into the virtual target space. Then, the behavior of the real person in the real target space is estimated by estimating the behavior of the virtual person in the virtual target space. Therefore, it is possible to estimate the behavior of the real person more accurately than when the behavior of the real person is directly estimated from a two-dimensional image. [Brief explanation of the drawings]

[0009] [Figure 1] This is a conceptual diagram to explain the outline of the real-to-virtual projection system. [Figure 2] FIG. 10 is a conceptual diagram for explaining virtual space configuration information. [Figure 3] FIG. 2 is a conceptual diagram for explaining camera configuration information. [Figure 4] FIG. 1 is a conceptual diagram for explaining an overview of a visualization function and an analysis function. [Figure 5] FIG. 10 is a conceptual diagram illustrating an example of an image analysis module. [Figure 6] FIG. 10 is a conceptual diagram for explaining a technique for improving the accuracy of localization processing. [Figure 7] FIG. 10 is a conceptual diagram for explaining an example of a gaze estimation process. [Figure 8] FIG. 10 is a conceptual diagram for explaining an example of a grip estimation process. [Figure 9] FIG. 10 is a conceptual diagram for explaining an example of a people flow estimation process. [Figure 10] FIG. 10 is a conceptual diagram for explaining an example of camera calibration and matching processing. [Figure 11] FIG. 2 is a block diagram showing a functional configuration related to automatic generation of virtual space configuration information. [Figure 12] FIG. 10 is a conceptual diagram for explaining another example of the gaze estimation process. DETAILED DESCRIPTION OF THE INVENTION

[0010] 1. Overview of Real-to-Virtual Projection System Figure 1 is a conceptual diagram to explain the outline of the real-to-virtual projection system. The real target space SP-R is an actual three-dimensional space that is the subject of various analyses. The virtual target space SP-V is a virtual three-dimensional space that represents the real target space SP-R. In other words, the virtual target space SP-V is a virtual three-dimensional space that imitates the real target space SP-R. The real target space SP-R and the virtual target space SP-V are expressed in the same world coordinate system (X, Y, Z).

[0011] Various physical objects exist in the real target space SP-R. Examples of physical objects include walls, pillars, doors, desks, chairs, shelves, boxes, displays, electronic devices, trees, etc. Hereinafter, physical objects that exist in the real target space SP-R will be referred to as real objects. Virtual objects that correspond to these real objects are defined in the virtual target space SP-V. In other words, virtual objects that mimic real objects are defined in the virtual target space SP-V. The configuration of real objects in the real target space SP-R and the configuration of virtual objects in the virtual target space SP-V match with a certain level of accuracy. Note that the term "configuration" here is a concept that includes position, orientation, shape, size, etc.

[0012] One or more real cameras CAM-R are installed in the real target space SP-R. Each real camera CAM-R is a stationary camera (fixed camera). One or more virtual cameras CAM-V corresponding to the one or more real cameras CAM-R are installed in the virtual target space SP-V. A corresponding pair of one real camera CAM-R and one virtual camera CAM-V has the same camera parameters. Here, the camera parameters include intrinsic parameters and extrinsic parameters. The intrinsic parameters include distortion parameters, focal length, etc. The extrinsic parameters include the position and orientation (rotation) of the camera in the world coordinate system. Camera calibration to determine the camera parameters is performed in advance. In addition, a process to align the virtual camera CAM-V in the virtual target space SP-V with the real camera CAM-R in the real target space SP-R is also performed in advance.

[0013] The real-to-virtual projection system 1 projects a person in a real target space SP-R into a virtual target space SP-V. More specifically, a real person existing in the real target space SP-R is photographed by a real camera CAM-R. The real-to-virtual projection system 1 detects the real person in the image photographed by the real camera CAM-R and estimates the three-dimensional posture of the detected real person. Furthermore, the real-to-virtual projection system 1 generates a virtual person that represents (simulates) the real person and has the estimated three-dimensional posture. The real-to-virtual projection system 1 then projects the virtual person into the virtual target space SP-V. The virtual person is projected into the virtual target space SP-V so that the position of the virtual person in the virtual target space SP-V matches the position of the real person in the real target space SP-R with a certain degree of accuracy or higher. The above projection processing may be performed in real time.

[0014] The real-to-virtual projection system 1 may visualize the virtual target space SP-V and the virtual person projected therein. For example, the real-to-virtual projection system 1 may generate an image of the virtual target space SP-V and the virtual person as seen from a virtual camera CAM-V and display the image on a display device. The visualization process may be performed in real time.

[0015] The real-to-virtual projection system 1 may estimate or analyze the actions of a virtual person projected into the virtual target space SP-V. The actions of a virtual person in the virtual target space SP-V are equivalent to the actions of a real person in the real target space SP-R. In other words, the real-to-virtual projection system 1 can estimate (analyze) the actions of a real person in the real target space SP-R by estimating (analyzing) the actions of a virtual person in the virtual target space SP-V. In this sense, the real-to-virtual projection system 1 can also be called a target space analysis system, a human action estimation system, etc. Hereinafter, the real-to-virtual projection system 1 will be simply referred to as "System 1."

[0016] The system 1 may be configured with a single node or multiple nodes. Fig. 1 also shows an example configuration of the system 1. The system 1 includes one or more real cameras CAM-R, one or more processors 10, one or more storage devices 20, one or more communication devices 30, one or more input devices 40, and one or more display devices 50.

[0017] The processor 10 executes various processes. Examples of the processor 10 include a general-purpose processor, a specific-purpose processor, a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), and a field-programmable gate array (FPGA). The storage device 20 stores various information required for processing. Examples of the storage device 20 include a hard disk drive (HDD), a solid-state drive (SSD), a volatile memory, and a non-volatile memory. The communication device 30 communicates with the outside via a communication network. The input device 40 accepts input of various information from a user of the system 1. Examples of the input device 40 include a keyboard, a mouse, a touch panel, and a microphone. The display device 50 displays various information. Examples of the display device 50 include a liquid crystal display, an organic electroluminescence (EL) display, and a head-up display (HUD).

[0018] The processor 10 may execute a computer program. The computer program is stored in the storage device 20. The computer program may be recorded on a computer-readable recording medium. The functions of the system 1 may be realized by cooperation between the processor 10 executing the computer program and the storage device 20.

[0019] The system 1 will be described in more detail below.

[0020] 2. Various information and functions 2-1. Virtual space configuration information FIG. 2 is a conceptual diagram illustrating virtual space configuration information 150. Virtual space configuration information 150 indicates the configuration in virtual target space SP-V. More specifically, virtual space configuration information 150 indicates the "configuration" of each object defined in virtual target space SP-V. Here, "configuration" is a concept that includes position, orientation, shape, size, etc. in the world coordinate system (X, Y, Z). For example, each object is represented by a three-dimensional bounding box. In this case, virtual space configuration information 150 includes information for defining the position, orientation, size, etc. of the bounding box of each object.

[0021] The objects defined in the virtual target space SP-V include virtual objects corresponding to real objects in the real target space SP-R. Each virtual object may be assigned identification information (see [A] in FIG. 2). Each virtual object may be assigned a color. The virtual space configuration information 150 may indicate the identification information and color of each virtual object. The virtual space configuration information 150 may also indicate the "category" of each virtual object (see [B] in FIG. 2). Here, "category" refers to the type of virtual object (e.g., wall, pillar, door, desk, chair, shelf, box, display, electronic device, tree, etc.). Furthermore, the virtual space configuration information 150 may include a language explanation for each virtual object.

[0022] The objects defined in the virtual target space SP-V may include area definition objects for defining areas within the virtual target space SP-V (see [C] in FIG. 2). The area definition objects may be represented by thin three-dimensional bounding boxes. Identification information may be assigned to each area definition object. A color may be assigned to each area definition object. Furthermore, the virtual space configuration information 150 may include a linguistic description of each area definition object.

[0023] The virtual space configuration information 150 is generated in advance and stored in the storage device 20.

[0024] The system 1 may include a customization module 100. The customization module 100 provides a user with a function for customizing (editing) the virtual space configuration information 150. In other words, the customization module 100 provides a user interface for customizing (editing) the virtual space configuration information 150. The customization module 100 displays the virtual space configuration information 150 being edited on the display device 50. The user can freely edit the virtual space configuration information 150 using the input device 40. In other words, the user can freely define virtual objects and area definition objects using the input device 40. The customization module 100 updates the virtual space configuration information 150 according to input from the user.

[0025] 2-2. Camera configuration information 3 is a conceptual diagram illustrating the camera configuration information 250. The camera configuration information 250 indicates the camera parameters of each real camera CAM-R and each virtual camera CAM-V. The camera parameters include intrinsic parameters and extrinsic parameters. The intrinsic parameters include distortion parameters, focal length, etc. The extrinsic parameters include the position and orientation (rotation) of the camera in the world coordinate system. A corresponding pair of one real camera CAM-R and one virtual camera CAM-V has the same camera parameters.

[0026] The camera configuration information 250 is generated in advance and stored in the storage device 20.

[0027] The system 1 may include a calibration module 200. The calibration module 200 performs "camera calibration" to determine the camera parameters of the real camera CAM-R. The calibration module 200 also performs "camera alignment" to correct the camera parameters so that the real target space SP-R seen from the real camera CAM-R is aligned with the virtual target space SP-V. That is, the calibration module 200 performs "camera calibration and alignment" to determine the camera parameters of the real camera CAM-R so that the real target space SP-R seen from the real camera CAM-R is aligned with the virtual target space SP-V. As a result, camera configuration information 250 indicating the camera parameters is obtained.

[0028] A specific example of the camera calibration and matching process will be explained in Section 7 below.

[0029] 2-3.Various functions 4 is a conceptual diagram for explaining an overview of the visualization function and analysis function of the system 1. The system 1 includes an image analysis module 300, a localization module 400, a visualization module 500, and a human analysis module 600.

[0030] The image analysis module 300 acquires a series of 2D images IMG captured by a real camera CAM-R installed in the real target space SP-R. The image analysis module 300 detects real people appearing in the 2D images IMG. The image analysis module 300 may track the detected real people. The image analysis module 300 may perform human re-identification to recognize the same real person across different real cameras CAM-R. The image analysis module 300 also estimates the 2D pose and 3D pose of the real person based on the 2D images IMG. The processing by the image analysis module 300 may be performed in real time. Details of the processing by the image analysis module 300 will be described later in Section 3.

[0031] The localization module 400 performs localization processing to estimate a person's position in the world coordinate system. The real person's position is the position where a real person exists in the real target space SP-R. The virtual person's position is a position in the virtual target space SP-V that corresponds to the real person's position. In other words, the virtual person's position in the virtual target space SP-V is set to match the real person's position in the real target space SP-R. The localization module 400 receives the analysis results from the image analysis module 300 and estimates the real person's position and the virtual person's position based on the analysis results and the camera configuration information 250. The localization module 400 then projects (places) a virtual person at the virtual person's position in the virtual target space SP-V. The virtual person is a virtual person that represents (simulates) a real person and has a three-dimensional posture estimated by the image analysis module 300. Details of the processing by the localization module 400 will be described later in Section 4.

[0032] The visualization module 500 visualizes the virtual target space SP-V and the virtual person projected therein by displaying them on the display device 50. The object configuration in the virtual target space SP-V is obtained from the virtual space configuration information 150. The virtual person has a three-dimensional pose as described above. For example, the visualization module 500 may generate an image of the virtual target space SP-V and the virtual person as seen from the virtual camera CAM-V based on the camera configuration information 250, and display the generated image on the display device 50. In this case, the generated image corresponding to the two-dimensional image IMG captured by the real camera CAM-R will be displayed on the display device 50. The visualization process may be performed in real time. Details of the processing by the visualization module 500 will be described later in Section 5.

[0033] The person analysis module 600 analyzes the virtual person projected into the virtual target space SP-V. For example, the person analysis module 600 performs an "action estimation process" to estimate the action of the virtual person in the virtual target space SP-V based on the relationship between the virtual person having a three-dimensional pose and the object configuration in the virtual target space SP-V. The object configuration in the virtual target space SP-V is obtained from the virtual space configuration information 150. The action of the virtual person in the virtual target space SP-V is equivalent to the action of the real person in the real target space SP-R. In other words, the person analysis module 600 can estimate the action of the real person in the real target space SP-R by estimating the action of the virtual person in the virtual target space SP-V. Since the person action is estimated based on the relationship between the virtual person having a three-dimensional pose and the object configuration, the estimation accuracy is improved compared to when the person action is directly estimated from the two-dimensional image IMG. The processing by the person analysis module 600 may be performed in real time. The person analysis module 600 may display the analysis results on the display device 50. Details of the processing by the person analysis module 600 are described in Section 6 below.

[0034] 3. Image Analysis Module 5 is a conceptual diagram illustrating an example of the image analysis module 300. The image analysis module 300 includes a human detector 310, a tracker 320, a human re-identification unit 330, and a pose estimator 340.

[0035] A series of 2D images IMG captured by a real camera CAM-R are input to the person detection unit 310. The person detection unit 310 performs a person detection process to detect real people appearing in each 2D image IMG. A bounding box represents the position of the real person detected in the 2D image IMG. Note that the person detection process is a well-known technique, and the method is not particularly limited. For example, YOLOX is used as the person detection unit 310.

[0036] The tracker 320 automatically tracks the same real person in a series of 2D images IMG based on a tracking algorithm. Tracking processing is a well-known technique, and the method is not particularly limited. For example, ByteTrack is used as the tracker 320.

[0037] The person re-identification unit 330 performs human re-identification to recognize the same real person across different real cameras CAM-R. More specifically, the person re-identification unit 330 acquires partial images of the real person captured in each 2D image IMG. A partial image surrounded by a bounding box in the 2D image IMG corresponds to the partial image of the real person. The person re-identification unit 330 extracts features of the real person (hereinafter referred to as "ReID features") based on the partial image of the real person. Typically, the person re-identification unit 330 extracts the ReID features from each partial image by using a ReID model based on machine learning. The ReID model may be a model based on a Transformer. The person re-identification unit 330 then calculates the similarity between the first real person and the second real person based on the ReID features of the first real person and the ReID features of the second real person. If the similarity is equal to or greater than a threshold, the person re-identification unit 330 determines that the first real person and the second real person are the same real person. Unique person identification information is assigned to the same real person.

[0038] Multi-Target Multi-Camera Tracking (MTMC) may be adopted. In MTMC, multiple 2D images IMG captured by multiple real cameras CAM-R are used, and tracking and person re-identification are performed on multiple real people in parallel.

[0039] The pose estimation unit 340 estimates the 2D pose (2D Pose) and 3D pose (3D Pose) of the real person based on each 2D image IMG. More specifically, the pose estimation unit 340 acquires a partial image of the real person shown in each 2D image IMG. The partial image surrounded by a bounding box in the 2D image IMG corresponds to the partial image of the real person. The pose estimation unit 340 extracts key points from the partial image using a pose estimation model based on machine learning, and estimates the 2D pose and 3D pose of the real person. The 2D pose is expressed in the image coordinate system of the 2D image IMG. Meanwhile, the 3D pose is expressed in the camera coordinate system (CX, CY, CZ). Information on the camera coordinate system (CX, CY, CZ) is obtained from the camera configuration information 250. The 2D pose and 3D pose are expressed by lines connecting parts such as joints, head, hands, and feet. Note that the pose estimation process is a well-known technique, and the method is not particularly limited. For example, MeTRAbs, TransPose, etc. are used for posture estimation processing.

[0040] Additionally, the image analysis module 300 may detect attributions of real people by analyzing partial images of the real people, such as gender and age.

[0041] 4. Localization Module The localization module 400 performs localization processing to estimate a person's position in the world coordinate system. The real person's position is the position where a real person exists in the real target space SP-R. The virtual person's position is the position in the virtual target space SP-V that corresponds to the real person's position. In other words, the virtual person's position in the virtual target space SP-V is set to match the real person's position in the real target space SP-R.

[0042] 4-1. First example of localization processing In a first example of localization processing, the localization module 400 receives information on the 3D pose of a real person from the pose estimation unit 340. The 3D pose is expressed in the camera coordinate system (CX, CY, CZ). The position of the 3D pose of the real person in the camera coordinate system is used as the real person position and the virtual person position in the camera coordinate system. Furthermore, the localization module 400 converts the real person position and the virtual person position in the camera coordinate system (CX, CY, CZ) into the real person position and the virtual person position in the world coordinate system (X, Y, Z) by using the camera configuration information 250. Then, the localization module 400 projects (places) the virtual person having the 3D pose at the virtual person position in the virtual target space SP-V.

[0043] In this way, in the first example, the person position is estimated based on the position of the three-dimensional posture in the camera coordinate system and the camera configuration information 250. However, in order to further improve the estimation accuracy of the person position, the second example described below may be adopted.

[0044] 4-2. Second example of localization processing FIG. 6 is a conceptual diagram illustrating a second example of the localization process. First, a depth map of the virtual object space SP-V as viewed from the virtual camera CAM-V is prepared in advance. The depth map provides a depth distribution from the virtual camera CAM-V to each object in the virtual object space SP-V. In particular, the depth map provides at least a depth distribution relative to the floor in the virtual object space SP-V. The depth distribution is provided in an image coordinate system as viewed from the virtual camera CAM-V. Such a depth map is generated, for example, based on virtual space configuration information 150 indicating the configuration of the virtual object space SP-V and camera configuration information 250 related to the virtual camera CAM-V. The depth map is stored in the storage device 20.

[0045] The localization module 400 receives information on the "two-dimensional pose" of the real person from the pose estimation unit 340. The two-dimensional pose is expressed in the image coordinate system. The localization module 400 obtains depth information D_ref for the position in the image of the two-dimensional pose of the real person from the depth map. In other words, the localization module 400 obtains the depth information D_ref for the position in the image of the two-dimensional pose by using the depth map as a lookup table (LUT).

[0046] In particular, the localization module 400 may focus on the position of the real person's "feet." More specifically, the localization module 400 estimates the positions of the real person's "feet" in the 2D image IMG based on information about the real person's 2D posture. For example, the real person's left and right feet are identified based on the real person's 2D posture, and the intermediate position between the positions of the left and right feet is used as the position of the "feet." Then, the localization module 400 obtains depth information D_ref for the positions of the real person's feet in the image from the depth map.

[0047] The 3D pose of the real person estimated by the pose estimation unit 340 is expressed in the camera coordinate system (CX, CY, CZ). The original depth D_org is depth information of the original 3D pose estimated by the pose estimation unit 340. The accuracy of the original depth D_org is not necessarily high. Therefore, the localization module 400 performs localization processing using depth information D_ref obtained from the depth map, rather than the original depth D_org.

[0048] For example, the localization module 400 uses the depth information D_ref and the camera configuration information 250 to project the in-image positions of the real person's feet onto three-dimensional positions in the camera coordinate system. In other words, the localization module 400 projects the in-image positions of the real person's feet onto three-dimensional positions corresponding to the depth information D_ref. At this time, the camera ray direction from the camera to the real person is maintained the same as the original (see the explanatory diagram at the bottom left of FIG. 6). The three-dimensional positions obtained in this way are used as highly accurate real person positions and virtual person positions. It can also be said that the localization module 400 reflects the depth information D_ref obtained from the depth map on the person positions while maintaining the original camera ray direction.

[0049] Furthermore, the localization module 400 converts the real person position and the virtual person position in the camera coordinate system (CX, CY, CZ) into the real person position and the virtual person position in the world coordinate system (X, Y, Z) by using the camera configuration information 250. Then, the localization module 400 projects (places) a virtual person having a three-dimensional pose at the virtual person position in the virtual target space SP-V.

[0050] Thus, according to the second example of the localization process, the accuracy of the localization process can be improved by using a depth map. As a result, the accuracy of the projection of the virtual person into the virtual target space SP-V is also improved, and the sense of incongruity caused by the projection result of the virtual person is suppressed. Furthermore, the improvement in the accuracy of the projection of the virtual person into the virtual target space SP-V leads to an improvement in the accuracy of the analysis process by the person analysis module 600.

[0051] Furthermore, the process of acquiring the depth information D_ref from the depth map (lookup table) is extremely simple, the processing load is light, and high-speed processing is possible. High-speed processing is preferable from the viewpoint of real-time processing. That is, according to the second example of localization processing, it is possible to realize real-time projection processing with high accuracy.

[0052] 5. Visualization Module The visualization module 500 visualizes the virtual target space SP-V and the virtual person projected therein by displaying them on the display device 50. The object configuration in the virtual target space SP-V is obtained from the virtual space configuration information 150. The virtual person is drawn to have an estimated three-dimensional pose. The virtual person may be drawn as an avatar having a three-dimensional pose. Attribute information (e.g., gender, age) obtained by the image analysis module 300 may be reflected in the avatar.

[0053] For example, the visualization module 500 may generate an image of the virtual target space SP-V and the virtual person as seen from the virtual camera CAM-V based on the camera configuration information 250, and display the generated image on the display device 50. In this case, a generated image corresponding to the two-dimensional image IMG captured by the real camera CAM-R will be displayed on the display device 50. The visualization process may be performed in real time.

[0054] In the case of MTMC (Multi-Target Multi-Camera Tracking), multiple 2D images IMG captured by multiple real cameras CAM-R are used, and tracking and person re-identification are performed on multiple real people in parallel. The visualization module 500 simultaneously displays multiple virtual people corresponding to the multiple real people on the display device 50.

[0055] The same real person is assigned unique personal identification information. It is also possible that the same real person appears simultaneously in two or more two-dimensional images IMG captured by two or more real cameras CAM-R. In this case, the accuracy of estimating the position of the same real person is improved because two-dimensional images IMG captured at different angles can be used. On the other hand, to avoid overlapping display of two or more virtual people corresponding to the same real person, the visualization module 500 may display only a single virtual person for the same real person on the display device 50.

[0056] 6. Person Analysis Module The person analysis module 600 analyzes the virtual person projected into the virtual target space SP-V. For example, the person analysis module 600 performs an "action estimation process" to estimate the action of the virtual person in the virtual target space SP-V based on the relationship between the virtual person having a three-dimensional pose and the object configuration in the virtual target space SP-V. The object configuration in the virtual target space SP-V is obtained from the virtual space configuration information 150. The action of the virtual person in the virtual target space SP-V is equivalent to the action of the real person in the real target space SP-R. In other words, the person analysis module 600 can estimate the action of the real person in the real target space SP-R by estimating the action of the virtual person in the virtual target space SP-V. Since the action of the person is estimated based on the relationship between the virtual person having a three-dimensional pose and the object configuration, the estimation accuracy is improved compared to when the action of the person is directly estimated from the two-dimensional image IMG.

[0057] For example, the person analysis module 600 estimates the behavior of a virtual person relative to virtual objects in the virtual target space SP-V based on the relationship between the virtual person having a three-dimensional pose and the configuration of each virtual object. The behavior of a virtual person relative to virtual objects in the virtual target space SP-V is equivalent to the behavior of a real person relative to real objects in the real target space SP-R. In other words, the person analysis module 600 can estimate the behavior of a real person relative to real objects in the real target space SP-R by estimating the behavior of a virtual person relative to virtual objects in the virtual target space SP-V. Because the behavior of a person relative to an object is estimated based on the relationship between the virtual person having a three-dimensional pose and the object configuration, the estimation accuracy is improved compared to when the behavior is estimated directly from the two-dimensional image IMG.

[0058] A specific example of the behavior estimation process performed by the person analysis module 600 will be described below.

[0059] 6-1. Gaze Estimation 7 is a conceptual diagram for explaining an example of gaze estimation processing. Person analysis module 600 includes gaze estimation module 610 that performs gaze estimation processing. Gaze estimation module 610 estimates which virtual object a virtual person is looking at, thereby estimating which real object a real person is looking at.

[0060] More specifically, the gaze estimation module 610 estimates the eye ray of the virtual person based on information about the virtual person having a three-dimensional pose. For example, the virtual person's facial orientation can be determined from the virtual person's three-dimensional pose. The facial orientation is considered to be the virtual person's gaze direction. As another example, the gaze direction may be estimated from the virtual person's three-dimensional pose by using a machine learning model. A line extending from the virtual person's face position in the gaze direction is set as the gaze ray. Then, the gaze estimation module 610 determines whether the virtual person's gaze ray intersects with any virtual object defined in the virtual target space SP-V.

[0061] For example, the gaze estimation module 610 determines whether the gaze ray intersects with any virtual object using a ray-triangle intersection algorithm. For example, when each virtual object is represented by a bounding box, the surface of the bounding box is represented by a combination of 12 triangular planes. In the example shown in FIG. 7, a triangle is defined by three vertices A, B, and C, and the gaze ray is represented by a combination of a starting point O and a direction d_g. The gaze estimation module 610 calculates the intersection point P between the plane on which the triangle exists and the gaze ray. If the intersection point P exists within the triangle, it is determined that the gaze ray intersects with the virtual object that owns that triangle. The gaze estimation module 610 can determine whether the gaze ray intersects with any virtual object by performing the above determination process on all triangles defined in the virtual target space SP-V. If it is determined that the gaze ray intersects with multiple virtual objects, the gaze estimation module 610 selects one virtual object that is closest to the starting point O of the gaze ray. The gaze estimation module 610 then estimates the closest virtual object that intersects with the gaze ray as the virtual object that the virtual person is looking at.

[0062] Note that the ray-triangle intersection algorithm is an example, and the present disclosure is not limited thereto. Other shapes may be used instead of triangles. At the very least, the ray-triangle intersection algorithm is extremely simple, has a light processing load, and is capable of high-speed processing. High-speed processing is preferable from the perspective of real-time processing.

[0063] In this way, it is possible to estimate with high accuracy which virtual object the virtual person is looking at based on the three-dimensional posture of the virtual person and the virtual space configuration information 150. In other words, it is possible to estimate with high accuracy which real object the real person is looking at. If the virtual space configuration information 150 indicates the category of each virtual object, it is possible to estimate with high accuracy which real object in which category the real person is looking at. By estimating which real object the real person is looking at, it is possible to know, for example, what the real person is interested in.

[0064] 6-2. Grasp Estimation When a real person's hand is inside any real object, it is highly likely that the real person is grasping or attempting to grasp an item stored in that real object. From this perspective, a grasp estimation process is performed to estimate whether the real person is grasping or attempting to grasp something.

[0065] 8 is a conceptual diagram for explaining an example of the grasp estimation process. The person analysis module 600 includes a grasp estimation module 620 that performs the grasp estimation process. The grasp estimation module 620 estimates whether the virtual person's hand is located inside any virtual object, thereby estimating whether the real person's hand is located inside any real object.

[0066] More specifically, the grasp estimation module 620 estimates the hand positions of the virtual person based on information about the virtual person having a three-dimensional pose. That is, the grasp estimation module 620 estimates the hand positions of the three-dimensional pose as the hand positions of the virtual person. Then, the grasp estimation module 620 determines whether the virtual person's hands are within any of the virtual objects defined in the virtual target space SP-V.

[0067] For example, the grasp estimation module 620 uses a point-cube detection algorithm to determine whether the virtual person's hand is inside any virtual object. For example, each virtual object is represented by a bounding box. In the example shown in FIG. 8, the bounding box of a virtual object is defined by vertices A to H. The position of the center point I of the bounding box is calculated from the positions of vertices D and F (I=(D+F) / 2). Point P is the position of the virtual person's hand. A vector V is defined that points from the center point I to point P. The three axes that define the bounding box are the x-axis, y-axis, and z-axis. [Vx, Vy, Vz] are the x-axis component, y-axis component, and z-axis component of vector V. Lx, Ly, and Lz are the lengths of the bounding box along the x-axis, y-axis, and z-axis directions. The grasp estimation module 620 determines whether the following condition is met: 2 × Vx≦Lx, 2 × Vy≦Ly, 2 × Vz≦Lz. If the condition is met, point P is determined to be inside the bounding box. That is, the virtual person's hand is estimated to be inside the virtual object represented by the bounding box.

[0068] The point-cube detection algorithm is an example, and the present disclosure is not limited thereto. However, the point-cube detection algorithm is extremely simple, has a light processing load, and is capable of high-speed processing. High-speed processing is preferable from the perspective of real-time processing.

[0069] In this way, based on the three-dimensional posture of the virtual person and the virtual space configuration information 150, it is possible to estimate with high accuracy whether the virtual person's hand is inside any virtual object. In other words, it is possible to estimate with high accuracy whether the real person's hand is inside any real object. If the real person's hand is inside any real object, it can be determined that the real person is at least interested in an item inside that real object. Furthermore, if the real person's hand is inside any real object, it is highly likely that the real person is grasping or attempting to grasp an item inside that real object. Therefore, the grasp estimation process can roughly estimate whether the real person is grasping or attempting to grasp something. If the virtual space configuration information 150 indicates the category of each virtual object, it is also possible to identify in more detail the item the real person is grasping or attempting to grasp.

[0070] 6-3. Human Flow Estimation 9 is a conceptual diagram for explaining an example of people flow estimation processing. The person analysis module 600 includes a people flow estimation module 630 that performs people flow estimation processing. The people flow estimation module 630 estimates the flow of virtual people in the virtual target space SP-V, thereby estimating the flow of real people in the real target space SP-R.

[0071] The people flow estimation process uses area definition objects (see [C] in Figure 2). An area definition object is an object for defining an area within the virtual target space SP-V. The virtual space configuration information 150 indicates the configuration of each area definition object defined in the virtual target space SP-V. The people flow estimation module 630 estimates the flow of virtual people in the virtual target space SP-V based on the relationship between the virtual people having three-dimensional postures and the configuration of the area definition objects.

[0072] More specifically, the people flow estimation module 630 estimates the position of the virtual person's feet based on information about the virtual person's three-dimensional pose. For example, the virtual person's left and right feet are identified based on the virtual person's three-dimensional pose, and the midpoint between the left and right foot positions is used as the "foot" position. The people flow estimation module 630 then determines which area definition object the virtual person's feet are located in. This determination is made, for example, based on the point-cube detection algorithm described in Section 6-2 above. The people flow estimation module 630 identifies the area definition object in which the virtual person's feet are located and determines that the virtual person is located in the area defined by that area definition object. Furthermore, the people flow estimation module 630 estimates the flow of the virtual person in the virtual target space SP-V by detecting changes in the area in which the virtual person is located.

[0073] In this way, the flow of a virtual person in the virtual target space SP-V can be estimated with high accuracy based on the three-dimensional posture of the virtual person and the virtual space configuration information 150. That is, the flow of a real person in the real target space SP-R can be estimated with high accuracy. The flow of a real person in the real target space SP-R can be used for various purposes. For example, based on the flow of a real person in the real target space SP-R, it is possible to investigate the congestion situation in the real target space SP-R and identify the causes of congestion. As another example, it is possible to analyze what a real person is interested in based on the flow of a real person in the real target space SP-R.

[0074] 6-4. Displaying analysis results The person analysis module 600 displays the analysis results on the display device 50. The analysis results may be statistical information or time transition information. Information may be classified by person attributes (e.g., gender, age).

[0075] 7. Camera calibration and matching process example The calibration module 200 performs camera calibration to determine the camera parameters of the real camera CAM-R. The calibration module 200 also performs camera alignment to correct the camera parameters so that the real target space SP-R as seen from the real camera CAM-R is aligned with the virtual target space SP-V. That is, the calibration module 200 performs camera calibration and alignment to determine the camera parameters of the real camera CAM-R so that the real target space SP-R as seen from the real camera CAM-R is aligned with the virtual target space SP-V. The camera calibration and alignment ensures the accuracy of the localization process (see Section 4), visualization process (see Section 5), and person analysis process (see Section 6).

[0076] A specific example of the camera calibration and matching process will be described below.

[0077] FIG. 10 is a conceptual diagram for explaining an example of camera calibration and matching processing. In this example, camera calibration is performed based on the PnP (Perspective-n-Point) method. The PnP method uses a point cloud consisting of n points, where n is an integer equal to or greater than 2. 3D point cloud coordinate information is coordinate information of the point cloud in 3D space (world coordinate system). 2D point cloud coordinate information is coordinate information of the point cloud in the image coordinate system when the point cloud is photographed by a camera. When the 3D point cloud coordinate information and the 2D point cloud coordinate information are given, the camera parameters can be calculated by solving the PnP problem.

[0078] In this example, the point cloud (n points) is acquired from a marker 210, which is a mark placed in space. For example, the marker 210 is a rectangle, and the four vertices M1 to M4 of the rectangle are used as the point cloud. The marker 210 has a predetermined pattern and can be recognized on an image.

[0079] More specifically, a real marker 210-R is placed at a predetermined real position photographed by a real camera CAM-R in the real target space SP-R. Meanwhile, a virtual marker 210-V is placed at a predetermined virtual position photographed by a virtual camera CAM-V in the virtual target space SP-V. Here, the predetermined virtual position in the virtual target space SP-V corresponds to a predetermined real position in the real target space SP-R. The real marker 210-R and the virtual marker 210-V have the same shape, orientation, and pattern. Therefore, the point cloud (vertices M1 to M4) of the virtual marker 210-V corresponds to the point cloud (vertices M1 to M4) of the real marker 210-R. When placing the virtual marker 210-V at a predetermined virtual position in the virtual target space SP-V, the customization module 100 shown in FIG. 2 is used.

[0080] The calibration module 200 acquires, as a query image, a two-dimensional image IMG captured by a real camera CAM-R installed in a real target space SP-R. The query image shows a real marker 210-R placed at a predetermined real position. The calibration module 200 has configuration information of the marker 210, and detects the real marker 210-R and a point cloud (four vertices M1 to M4) in the query image by performing pattern matching. Then, the calibration module 200 acquires the position of the point cloud in the query image as "two-dimensional point cloud coordinate information 211."

[0081] Meanwhile, the calibration module 200 acquires the position of the point cloud (four vertices M1 to M4) of the virtual marker 210-V in the virtual target space SP-V as “three-dimensional point cloud coordinate information 212.” The position of the point cloud of the virtual marker 210-V in the virtual target space SP-V is obtained from the virtual space configuration information 150.

[0082] The calibration module 200 determines the camera parameters by solving the PnP problem based on the thus obtained 2D point cloud coordinate information 211 and 3D point cloud coordinate information 212. What is important here is that, according to this method, camera alignment is achieved at the same time as the camera parameters are determined. Since the 2D point cloud coordinate information 211 obtained from the real target space SP-R and the 3D point cloud coordinate information 212 obtained from the virtual target space SP-V are combined, camera alignment is achieved at the same time as the camera parameters are determined.

[0083] As described above, according to this example, it is possible to realize both camera calibration and camera matching processing in one step, which is preferable from the viewpoint of reducing the processing load.

[0084] 8. Automatic generation of virtual space configuration information As described in Section 2-1 above, the virtual space configuration information 150 can be generated by a user through the customization module 100. As another example, as described below, the virtual space configuration information 150 can be automatically generated. Note that after the virtual space configuration information 150 is automatically generated, the user can also customize the virtual space configuration information 150 through the customization module 100.

[0085] 11 is a block diagram showing a functional configuration related to the automatic generation of virtual space configuration information 150. First, a group of images 110 taken in the real target space SP-R is prepared. For example, a person with a camera takes pictures of the real target space SP-R from various viewpoints while moving around in the real target space SP-R.

[0086] The system 1 includes a 3D reconstruction processor 120. The 3D reconstruction processor 120 receives a set of images 110 captured in a real target space SP-R. The 3D reconstruction processor 120 then reconstructs a 3D environment of the real target space SP-R based on the set of images 110.

[0087] For example, SfM (Structure from Motion) is used. SfM is a well-known technique that receives multiple images captured by a camera in a space, estimates (reconstructs) the three-dimensional environment of the space from the multiple images, and estimates the position and orientation of the camera. More specifically, SfM extracts feature points from the multiple images, performs feature matching between the multiple images, and estimates the position of a point cloud representing the three-dimensional environment and the position and orientation of the camera from the geometric relationship between the corresponding feature points. Note that the point cloud obtained by SfM is generally "sparse" and is often called a sparse point cloud. For example, COLMAP is known as software that performs SfM.

[0088] The SfM processing unit 121 applies SfM to the group of images 110 to obtain sparse point cloud data 122 representing the three-dimensional environment of the real target space SP-R. Next, the reconstruction processing unit 123 applies 3D reconstruction processing to the sparse point cloud data 122. Known 3D reconstruction processing methods include NeRF (Neural Radiance Fields) and 3D Gaussian Splatting.

[0089] 3D Gaussian splatting is described, for example, in Non-Patent Document 1. In 3D Gaussian splatting, "splats" representing a Gaussian distribution are arranged based on sparse point cloud data 122. Each splat is defined by its position, covariance, color, etc. The set of arranged splats is optimized to fit a group of images 110, thereby reconstructing a 3D environment in the real target space SP-R. The reconstructed 3D environment is represented by a set of multiple splats. 3D Gaussian splatting is known to be faster than NeRF. Therefore, using 3D Gaussian splatting to perform 3D reconstruction processing is preferable from the perspective of real-time processing.

[0090] The three-dimensional environment of the real target space SP-R reconstructed by the three-dimensional reconstruction processing corresponds to the three-dimensional environment of the virtual target space SP-V. That is, the three-dimensional reconstruction processing unit 120 reconstructs the three-dimensional environment of the real target space SP-R based on the group of images 110, thereby generating the virtual target space SP-V that represents the three-dimensional environment of the real target space SP-R. The reconstructed space information 124 indicates the three-dimensional environment of the virtual target space SP-V generated in this manner.

[0091] Furthermore, the semantic segmentation processor 125 applies semantic segmentation to the 3D environment of the virtual object space SP-V represented by the reconstructed spatial information 124. More specifically, the semantic segmentation processor 125 generates a 2D image from the 3D environment by rendering based on the reconstructed spatial information 124, and then performs semantic segmentation on the resulting 2D image. Semantic segmentation divides the 2D image into multiple regions by classifying each pixel in the 2D image into categories and grouping pixels of the same category together. Regions of the same category correspond to "objects." Applying semantic segmentation to the 3D environment of the virtual object space SP-V makes it possible to recognize the configuration and language explanation of each object in the virtual object space SP-V. The language explanation of an object includes the object's category, color, etc.

[0092] Non-Patent Document 1 discloses a technology (Language Gaussian Splatting) that performs 3D Gaussian splatting and semantic segmentation at the same time. The technology disclosed in Non-Patent Document 1 may be used in the 3D reconstruction processing unit 120.

[0093] The three-dimensional reconstruction processing unit 120 combines the configuration and linguistic description of each object in the virtual target space SP-V with the reconstructed space information 124. This results in virtual space configuration information 150 indicating the configuration and linguistic description of each object in the virtual target space SP-V.

[0094] The virtual space configuration information 150 thus automatically generated can also be applied to the "behavior estimation process" described in Section 6 above.

[0095] 12 shows an application example in which the "gaze estimation process" described in Section 6-1 above is performed. The gaze estimation module 610 estimates which virtual object the virtual person is looking at, thereby estimating which real object the real person is looking at. More specifically, the gaze estimation module 610 determines whether the gaze ray of the virtual person intersects with any virtual object in the virtual target space SP-V.

[0096] In the example shown in FIG. 12, it is assumed that 3D reconstruction processing was performed using 3D Gaussian splatting. Therefore, each virtual object in the virtual object space SP-V is represented by a set of multiple splats. In this case, the gaze estimation module 610 uses a ray-sphere intersection algorithm to determine whether the gaze ray intersects with any of the splats. In FIG. 12, the gaze ray is represented by a combination of a starting point O and a direction d_g, the center of the splat is indicated by point C, and the intersections of the gaze ray and the surface of the splat are indicated by P1 and P2. If it is determined that the gaze ray intersects with multiple splats, the gaze estimation module 610 selects the splat closest to the starting point O of the gaze ray. The gaze estimation module 610 then estimates the virtual object containing the closest splat intersecting the gaze ray as the virtual object being viewed by the virtual person. The ray-sphere intersection algorithm is also simple, has a light processing load, and is capable of high-speed processing. High-speed processing is preferable from the perspective of real-time processing.

[0097] As explained above, by reconstructing the three-dimensional environment of the real target space SP-R, the virtual target space SP-V that represents the three-dimensional environment of the real target space SP-R is automatically generated. Therefore, there is no need to manually predefine the virtual target space SP-V. The behavior estimation process is possible even if the virtual target space SP-V is not manually predefined.

[0098] When the three-dimensional environment of the real target space SP-R changes, the virtual target space SP-V must also be updated to follow the change. In this case, it is possible to manually edit (redefine) the virtual target space SP-V using the customization module 100. However, the method described in this section makes it easier to automatically update the virtual target space SP-V.

[0099] Furthermore, our method automatically recognizes the linguistic description of each object through semantic segmentation. This allows for more detailed behavior estimation. For example, it is possible to accurately estimate which real objects in which categories a real person is looking. It is also possible to know which objects in which categories a real person is interested.

[0100] Fast 3D Gaussian splatting can be used for 3D reconstruction, which contributes to real-time processing. Also, since each object is represented by a set of multiple small splats, it is possible to estimate human behavior relative to the object in more detail. [Explanation of symbols]

[0101] 1...Real-to-virtual projection system (target space analysis system), 100...Customization module, 150...Virtual space configuration information, 200...Calibration module, 250...Camera configuration information, 300...Image analysis module, 400...Localization module, 500...Visualization module, 600...Person analysis module, CAM-R...Real camera, CAM-V...Virtual camera, SP-R...Real target space, SP-V...Virtual target space

Claims

1. one or more processors; the one or more processors: Reconstructing a three-dimensional environment of the real target space based on a group of images taken in the real target space, and generating a virtual target space that represents the three-dimensional environment of the real target space; Recognizing configurations and linguistic descriptions of objects in the virtual object space by applying semantic segmentation to the three-dimensional environment; Detecting real people in images captured by a real camera installed in the real target space; Estimating a three-dimensional pose of the real person based on the image; projecting a virtual person having the three-dimensional pose and representing the real person into the virtual object space; A behavior estimation process is performed to estimate the behavior of the real person in the real target space by estimating the behavior of the virtual person in the virtual target space based on the relationship between the virtual person having the three-dimensional posture and the configuration of the object in the virtual target space. It was configured as Projection system.

2. 10. The projection system of claim 1, the one or more processors are configured to reconstruct the three-dimensional environment of the real object space by utilizing three-dimensional Gaussian splatting. Projection system.

3. 10. The projection system of claim 1, The behavior estimation process includes a gaze estimation process for estimating which real object the real person is looking at by estimating which object the virtual person is looking at. Projection system.

4. 4. A projection system according to claim 3, The gaze estimation process includes: estimating a gaze ray of the virtual person based on the three-dimensional pose of the virtual person; estimating which object the virtual person is looking at by determining whether the line of sight ray intersects with any object in the virtual object space; and Contains Projection system.

5. 5. A projection system according to claim 4, the one or more processors are configured to reconstruct the three-dimensional environment of the real object space by utilizing three-dimensional Gaussian splatting; each object in the virtual object space is represented by a collection of multiple splats; Determining whether the line of sight ray intersects with any object in the virtual object space includes determining whether the line of sight ray intersects with any splat. Projection system.

Citation Information

Patent Citations

  • Virtual object operating system and virtual object operating method

    JP2021068405A

  • Work estimation apparatus, method and program

    JP2022048017A

  • Gaze estimation system

    JP2022187547A