Projection system
The projection system addresses the challenge of integrating analysis results across varied target spaces by using identical spatial shape and camera parameters to project a real person's three-dimensional posture into a virtual space, enhancing behavior estimation accuracy.
Patent Information
- Application Number
- JP2024080695
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-17
- Publication Date
- 2025-11-28
AI Technical Summary
Existing systems struggle to compare, reuse, or integrate analysis results across multiple target spaces due to variations in shape and camera placement.
A projection system with identical spatial shape and camera parameters across multiple real target spaces, projecting a real person's three-dimensional pose into a virtual space for behavior estimation, allowing for accurate analysis by estimating the behavior of a virtual person in the virtual space.
Enables easy comparison and integration of analysis results across multiple real target spaces, improving the accuracy of behavior estimation by projecting a real person's three-dimensional posture into a virtual space.
Smart Images

Figure 2025174371000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a technique for projecting a person in a real target space into a virtual target space. [Background technology]
[0002] Patent Document 1 is A population data file is provided in which the movement line data for each person, which is obtained by tracking the behavior of each person within the monitored area, is accumulated as a population. The movement line data stored in the population data file is used to calculate statistics related to human behavior analysis.
[0003] Other known techniques include those disclosed in Patent Documents 2, 3, and 4. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2010-002997 [Patent Document 2] International Publication No. 2015 / 025490 [Patent Document 3] Japanese Patent Application Laid-Open No. 2005-309951 [Patent Document 4] Japanese Patent Application Laid-Open No. 2011-232864 Summary of the Invention [Problem to be solved by the invention]
[0005] Images captured by a camera installed in a target space can be used to analyze the target space. For example, the behavior of people present in the target space can be estimated based on the images captured by the camera installed in the target space. However, because the shape of the target space and the camera placement in the target space vary, the analysis results cannot be simply compared, reused, or integrated between multiple target spaces. [Means for solving the problem]
[0006] One aspect of the present disclosure relates to a projection system. The projection system includes one or more processors and one or more storage devices. The plurality of real target spaces have the same shape, and a real camera having the same camera parameters is installed in each of the plurality of real target spaces. The one or more storage devices are configured to store virtual space configuration information indicating the configuration of objects defined in each of a plurality of virtual object spaces that respectively represent a plurality of real object spaces. The one or more processors detect a real person appearing in an image captured by a real camera installed in a first real target space among the plurality of real target spaces. The one or more processors estimate a three-dimensional pose of the real person based on the images. The one or more processors project a virtual person having a three-dimensional pose and representing a real person into a first virtual object space representing a first real object space. The one or more processors perform a behavior estimation process to estimate the behavior of the real person in the first real target space by estimating the behavior of the virtual person in the first virtual target space based on the relationship between the virtual person having a three-dimensional pose and the configuration of objects in the first virtual target space. [Effects of the Invention]
[0007] According to the present disclosure, since the spatial shape and camera parameters are the same among multiple real target spaces, it becomes possible to easily compare, reuse, and integrate analysis results among multiple real target spaces. Furthermore, according to the present disclosure, the three-dimensional posture of a real person in a real target space is estimated, and a virtual person representing the real person having the three-dimensional posture is projected into the virtual target space. Then, the behavior of the real person in the real target space is estimated by estimating the behavior of the virtual person in the virtual target space. Therefore, it is possible to estimate the behavior of the real person more accurately than when the behavior of the real person is directly estimated from a two-dimensional image. [Brief explanation of the drawings]
[0008] [Figure 1] This is a conceptual diagram to explain the outline of the real-to-virtual projection system. [Figure 2] FIG. 10 is a conceptual diagram for explaining virtual space configuration information. [Figure 3] FIG. 2 is a conceptual diagram for explaining camera configuration information. [Figure 4] FIG. 1 is a conceptual diagram for explaining an overview of a visualization function and an analysis function. [Figure 5] FIG. 10 is a conceptual diagram illustrating an example of an image analysis module. [Figure 6] FIG. 10 is a conceptual diagram for explaining a technique for improving the accuracy of localization processing. [Figure 7] FIG. 10 is a conceptual diagram for explaining an example of a gaze estimation process. [Figure 8] FIG. 10 is a conceptual diagram for explaining an example of a grip estimation process. [Figure 9] FIG. 10 is a conceptual diagram for explaining an example of a people flow estimation process. [Figure 10] FIG. 10 is a conceptual diagram for explaining an example of camera calibration and matching processing. [Figure 11] FIG. 1 is a conceptual diagram for explaining an example of a plurality of real target spaces. DETAILED DESCRIPTION OF THE INVENTION
[0009] Embodiments of the present disclosure will be described with reference to the accompanying drawings.
[0010] 1. Overview of Real-to-Virtual Projection System Figure 1 is a conceptual diagram to explain the outline of the real-to-virtual projection system. The real target space SP-R is an actual three-dimensional space that is the subject of various analyses. The virtual target space SP-V is a virtual three-dimensional space that represents the real target space SP-R. In other words, the virtual target space SP-V is a virtual three-dimensional space that imitates the real target space SP-R. The real target space SP-R and the virtual target space SP-V are expressed in the same world coordinate system (X, Y, Z).
[0011] Various physical objects exist in the real target space SP-R. Examples of physical objects include walls, pillars, doors, desks, chairs, shelves, boxes, displays, electronic devices, trees, etc. Hereinafter, physical objects that exist in the real target space SP-R will be referred to as real objects. Virtual objects that correspond to these real objects are defined in the virtual target space SP-V. In other words, virtual objects that mimic real objects are defined in the virtual target space SP-V. The configuration of real objects in the real target space SP-R and the configuration of virtual objects in the virtual target space SP-V match with a certain level of accuracy. Note that the term "configuration" here is a concept that includes position, orientation, shape, size, etc.
[0012] One or more real cameras CAM-R are installed in the real target space SP-R. Each real camera CAM-R is a stationary camera (fixed camera). One or more virtual cameras CAM-V corresponding to the one or more real cameras CAM-R are installed in the virtual target space SP-V. A corresponding pair of one real camera CAM-R and one virtual camera CAM-V has the same camera parameters. Here, the camera parameters include intrinsic parameters and extrinsic parameters. The intrinsic parameters include distortion parameters, focal length, etc. The extrinsic parameters include the position and orientation (rotation) of the camera in the world coordinate system. Camera calibration to determine the camera parameters is performed in advance. In addition, a process to align the virtual camera CAM-V in the virtual target space SP-V with the real camera CAM-R in the real target space SP-R is also performed in advance.
[0013] The real-to-virtual projection system 1 projects a person in a real target space SP-R into a virtual target space SP-V. More specifically, a real person existing in the real target space SP-R is photographed by a real camera CAM-R. The real-to-virtual projection system 1 detects the real person in the image photographed by the real camera CAM-R and estimates the three-dimensional posture of the detected real person. Furthermore, the real-to-virtual projection system 1 generates a virtual person that represents (simulates) the real person and has the estimated three-dimensional posture. The real-to-virtual projection system 1 then projects the virtual person into the virtual target space SP-V. The virtual person is projected into the virtual target space SP-V so that the position of the virtual person in the virtual target space SP-V matches the position of the real person in the real target space SP-R with a certain degree of accuracy or higher. The above projection processing may be performed in real time.
[0014] The real-to-virtual projection system 1 may visualize the virtual target space SP-V and the virtual person projected therein. For example, the real-to-virtual projection system 1 may generate an image of the virtual target space SP-V and the virtual person as seen from a virtual camera CAM-V and display the image on a display device. The visualization process may be performed in real time.
[0015] The real-to-virtual projection system 1 may estimate or analyze the actions of a virtual person projected into the virtual target space SP-V. The actions of a virtual person in the virtual target space SP-V are equivalent to the actions of a real person in the real target space SP-R. In other words, the real-to-virtual projection system 1 can estimate (analyze) the actions of a real person in the real target space SP-R by estimating (analyzing) the actions of a virtual person in the virtual target space SP-V. In this sense, the real-to-virtual projection system 1 can also be called a target space analysis system, a human action estimation system, etc. Hereinafter, the real-to-virtual projection system 1 will be simply referred to as "System 1."
[0016] The system 1 may be configured with a single node or multiple nodes. Fig. 1 also shows an example configuration of the system 1. The system 1 includes one or more real cameras CAM-R, one or more processors 10, one or more storage devices 20, one or more communication devices 30, one or more input devices 40, and one or more display devices 50.
[0017] The processor 10 executes various processes. Examples of the processor 10 include a general-purpose processor, a specific-purpose processor, a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), and a field-programmable gate array (FPGA). The storage device 20 stores various information required for processing. Examples of the storage device 20 include a hard disk drive (HDD), a solid-state drive (SSD), a volatile memory, and a non-volatile memory. The communication device 30 communicates with the outside via a communication network. The input device 40 accepts input of various information from a user of the system 1. Examples of the input device 40 include a keyboard, a mouse, a touch panel, and a microphone. The display device 50 displays various information. Examples of the display device 50 include a liquid crystal display, an organic electroluminescence (EL) display, and a head-up display (HUD).
[0018] The processor 10 may execute a computer program. The computer program is stored in the storage device 20. The computer program may be recorded on a computer-readable recording medium. The functions of the system 1 may be realized by cooperation between the processor 10 executing the computer program and the storage device 20.
[0019] The system 1 will be described in more detail below.
[0020] 2. Various information and functions 2-1. Virtual space configuration information FIG. 2 is a conceptual diagram illustrating virtual space configuration information 150. Virtual space configuration information 150 indicates the configuration in virtual target space SP-V. More specifically, virtual space configuration information 150 indicates the "configuration" of each object defined in virtual target space SP-V. Here, "configuration" is a concept that includes position, orientation, shape, size, etc. in the world coordinate system (X, Y, Z). For example, each object is represented by a three-dimensional bounding box. In this case, virtual space configuration information 150 includes information for defining the position, orientation, size, etc. of the bounding box of each object.
[0021] The objects defined in the virtual target space SP-V include virtual objects corresponding to real objects in the real target space SP-R. Each virtual object may be assigned identification information (see [A] in FIG. 2). Each virtual object may be assigned a color. The virtual space configuration information 150 may indicate the identification information and color of each virtual object. The virtual space configuration information 150 may also indicate the "category" of each virtual object (see [B] in FIG. 2). Here, "category" refers to the type of virtual object (e.g., wall, pillar, door, desk, chair, shelf, box, display, electronic device, tree, etc.). Furthermore, the virtual space configuration information 150 may include a language explanation for each virtual object.
[0022] The objects defined in the virtual target space SP-V may include area definition objects for defining areas within the virtual target space SP-V (see [C] in FIG. 2). The area definition objects may be represented by thin three-dimensional bounding boxes. Identification information may be assigned to each area definition object. A color may be assigned to each area definition object. Furthermore, the virtual space configuration information 150 may include a linguistic description of each area definition object.
[0023] The virtual space configuration information 150 is generated in advance and stored in the storage device 20.
[0024] The system 1 may include a customization module 100. The customization module 100 provides a user with a function for customizing (editing) the virtual space configuration information 150. In other words, the customization module 100 provides a user interface for customizing (editing) the virtual space configuration information 150. The customization module 100 displays the virtual space configuration information 150 being edited on the display device 50. The user can freely edit the virtual space configuration information 150 using the input device 40. In other words, the user can freely define virtual objects and area definition objects using the input device 40. The customization module 100 updates the virtual space configuration information 150 according to input from the user.
[0025] 2-2. Camera configuration information 3 is a conceptual diagram illustrating the camera configuration information 250. The camera configuration information 250 indicates the camera parameters of each real camera CAM-R and each virtual camera CAM-V. The camera parameters include intrinsic parameters and extrinsic parameters. The intrinsic parameters include distortion parameters, focal length, etc. The extrinsic parameters include the position and orientation (rotation) of the camera in the world coordinate system. A corresponding pair of one real camera CAM-R and one virtual camera CAM-V has the same camera parameters.
[0026] The camera configuration information 250 is generated in advance and stored in the storage device 20.
[0027] The system 1 may include a calibration module 200. The calibration module 200 performs "camera calibration" to determine the camera parameters of the real camera CAM-R. The calibration module 200 also performs "camera alignment" to correct the camera parameters so that the real target space SP-R seen from the real camera CAM-R is aligned with the virtual target space SP-V. That is, the calibration module 200 performs "camera calibration and alignment" to determine the camera parameters of the real camera CAM-R so that the real target space SP-R seen from the real camera CAM-R is aligned with the virtual target space SP-V. As a result, camera configuration information 250 indicating the camera parameters is obtained.
[0028] A specific example of the camera calibration and matching process will be explained in Section 7 below.
[0029] 2-3.Various functions 4 is a conceptual diagram for explaining an overview of the visualization function and analysis function of the system 1. The system 1 includes an image analysis module 300, a localization module 400, a visualization module 500, and a human analysis module 600.
[0030] The image analysis module 300 acquires a series of 2D images IMG captured by a real camera CAM-R installed in the real target space SP-R. The image analysis module 300 detects real people appearing in the 2D images IMG. The image analysis module 300 may track the detected real people. The image analysis module 300 may perform human re-identification to recognize the same real person across different real cameras CAM-R. The image analysis module 300 also estimates the 2D pose and 3D pose of the real person based on the 2D images IMG. The processing by the image analysis module 300 may be performed in real time. Details of the processing by the image analysis module 300 will be described later in Section 3.
[0031] The localization module 400 performs localization processing to estimate a person's position in the world coordinate system. The real person's position is the position where a real person exists in the real target space SP-R. The virtual person's position is a position in the virtual target space SP-V that corresponds to the real person's position. In other words, the virtual person's position in the virtual target space SP-V is set to match the real person's position in the real target space SP-R. The localization module 400 receives the analysis results from the image analysis module 300 and estimates the real person's position and the virtual person's position based on the analysis results and the camera configuration information 250. The localization module 400 then projects (places) a virtual person at the virtual person's position in the virtual target space SP-V. The virtual person is a virtual person that represents (simulates) a real person and has a three-dimensional posture estimated by the image analysis module 300. Details of the processing by the localization module 400 will be described later in Section 4.
[0032] The visualization module 500 visualizes the virtual target space SP-V and the virtual person projected therein by displaying them on the display device 50. The object configuration in the virtual target space SP-V is obtained from the virtual space configuration information 150. The virtual person has a three-dimensional pose as described above. For example, the visualization module 500 may generate an image of the virtual target space SP-V and the virtual person as seen from the virtual camera CAM-V based on the camera configuration information 250, and display the generated image on the display device 50. In this case, the generated image corresponding to the two-dimensional image IMG captured by the real camera CAM-R will be displayed on the display device 50. The visualization process may be performed in real time. Details of the processing by the visualization module 500 will be described later in Section 5.
[0033] The person analysis module 600 analyzes the virtual person projected into the virtual target space SP-V. For example, the person analysis module 600 performs an "action estimation process" to estimate the action of the virtual person in the virtual target space SP-V based on the relationship between the virtual person having a three-dimensional pose and the object configuration in the virtual target space SP-V. The object configuration in the virtual target space SP-V is obtained from the virtual space configuration information 150. The action of the virtual person in the virtual target space SP-V is equivalent to the action of the real person in the real target space SP-R. In other words, the person analysis module 600 can estimate the action of the real person in the real target space SP-R by estimating the action of the virtual person in the virtual target space SP-V. Since the person action is estimated based on the relationship between the virtual person having a three-dimensional pose and the object configuration, the estimation accuracy is improved compared to when the person action is directly estimated from the two-dimensional image IMG. The processing by the person analysis module 600 may be performed in real time. The person analysis module 600 may display the analysis results on the display device 50. Details of the processing by the person analysis module 600 are described in Section 6 below.
[0034] 3. Image Analysis Module 5 is a conceptual diagram illustrating an example of the image analysis module 300. The image analysis module 300 includes a human detector 310, a tracker 320, a human re-identification unit 330, and a pose estimator 340.
[0035] A series of 2D images IMG captured by a real camera CAM-R are input to the person detection unit 310. The person detection unit 310 performs a person detection process to detect real people appearing in each 2D image IMG. A bounding box represents the position of the real person detected in the 2D image IMG. Note that the person detection process is a well-known technique, and the method is not particularly limited. For example, YOLOX is used as the person detection unit 310.
[0036] The tracker 320 automatically tracks the same real person in a series of 2D images IMG based on a tracking algorithm. Tracking processing is a well-known technique, and the method is not particularly limited. For example, ByteTrack is used as the tracker 320.
[0037] The person re-identification unit 330 performs human re-identification to recognize the same real person across different real cameras CAM-R. More specifically, the person re-identification unit 330 acquires partial images of the real person captured in each 2D image IMG. A partial image surrounded by a bounding box in the 2D image IMG corresponds to the partial image of the real person. The person re-identification unit 330 extracts features of the real person (hereinafter referred to as "ReID features") based on the partial image of the real person. Typically, the person re-identification unit 330 extracts the ReID features from each partial image by using a ReID model based on machine learning. The ReID model may be a model based on a Transformer. The person re-identification unit 330 then calculates the similarity between the first real person and the second real person based on the ReID features of the first real person and the ReID features of the second real person. If the similarity is equal to or greater than a threshold, the person re-identification unit 330 determines that the first real person and the second real person are the same real person. Unique person identification information is assigned to the same real person.
[0038] Multi-Target Multi-Camera Tracking (MTMC) may be adopted. In MTMC, multiple 2D images IMG captured by multiple real cameras CAM-R are used, and tracking and person re-identification are performed on multiple real people in parallel.
[0039] The pose estimation unit 340 estimates the 2D pose (2D Pose) and 3D pose (3D Pose) of the real person based on each 2D image IMG. More specifically, the pose estimation unit 340 acquires a partial image of the real person shown in each 2D image IMG. The partial image surrounded by a bounding box in the 2D image IMG corresponds to the partial image of the real person. The pose estimation unit 340 extracts key points from the partial image using a pose estimation model based on machine learning, and estimates the 2D pose and 3D pose of the real person. The 2D pose is expressed in the image coordinate system of the 2D image IMG. Meanwhile, the 3D pose is expressed in the camera coordinate system (CX, CY, CZ). Information on the camera coordinate system (CX, CY, CZ) is obtained from the camera configuration information 250. The 2D pose and 3D pose are expressed by lines connecting parts such as joints, head, hands, and feet. Note that the pose estimation process is a well-known technique, and the method is not particularly limited. For example, MeTRAbs, TransPose, etc. are used for posture estimation processing.
[0040] Additionally, the image analysis module 300 may detect attributions of real people by analyzing partial images of the real people, such as gender and age.
[0041] 4. Localization Module The localization module 400 performs localization processing to estimate a person's position in the world coordinate system. The real person's position is the position where a real person exists in the real target space SP-R. The virtual person's position is the position in the virtual target space SP-V that corresponds to the real person's position. In other words, the virtual person's position in the virtual target space SP-V is set to match the real person's position in the real target space SP-R.
[0042] 4-1. First example of localization processing In a first example of localization processing, the localization module 400 receives information on the 3D pose of a real person from the pose estimation unit 340. The 3D pose is expressed in the camera coordinate system (CX, CY, CZ). The position of the 3D pose of the real person in the camera coordinate system is used as the real person position and the virtual person position in the camera coordinate system. Furthermore, the localization module 400 converts the real person position and the virtual person position in the camera coordinate system (CX, CY, CZ) into the real person position and the virtual person position in the world coordinate system (X, Y, Z) by using the camera configuration information 250. Then, the localization module 400 projects (places) the virtual person having the 3D pose at the virtual person position in the virtual target space SP-V.
[0043] In this way, in the first example, the person position is estimated based on the position of the three-dimensional posture in the camera coordinate system and the camera configuration information 250. However, in order to further improve the estimation accuracy of the person position, the second example described below may be adopted.
[0044] 4-2. Second example of localization processing FIG. 6 is a conceptual diagram illustrating a second example of the localization process. First, a depth map of the virtual object space SP-V as viewed from the virtual camera CAM-V is prepared in advance. The depth map provides a depth distribution from the virtual camera CAM-V to each object in the virtual object space SP-V. In particular, the depth map provides at least a depth distribution relative to the floor in the virtual object space SP-V. The depth distribution is provided in an image coordinate system as viewed from the virtual camera CAM-V. Such a depth map is generated, for example, based on virtual space configuration information 150 indicating the configuration of the virtual object space SP-V and camera configuration information 250 related to the virtual camera CAM-V. The depth map is stored in the storage device 20.
[0045] The localization module 400 receives information on the "two-dimensional pose" of the real person from the pose estimation unit 340. The two-dimensional pose is expressed in the image coordinate system. The localization module 400 obtains depth information D_ref for the position in the image of the two-dimensional pose of the real person from the depth map. In other words, the localization module 400 obtains the depth information D_ref for the position in the image of the two-dimensional pose by using the depth map as a lookup table (LUT).
[0046] In particular, the localization module 400 may focus on the position of the real person's "feet." More specifically, the localization module 400 estimates the positions of the real person's "feet" in the 2D image IMG based on information about the real person's 2D posture. For example, the real person's left and right feet are identified based on the real person's 2D posture, and the intermediate position between the positions of the left and right feet is used as the position of the "feet." Then, the localization module 400 obtains depth information D_ref for the positions of the real person's feet in the image from the depth map.
[0047] The 3D pose of the real person estimated by the pose estimation unit 340 is expressed in the camera coordinate system (CX, CY, CZ). The original depth D_org is depth information of the original 3D pose estimated by the pose estimation unit 340. The accuracy of the original depth D_org is not necessarily high. Therefore, the localization module 400 performs localization processing using depth information D_ref obtained from the depth map, rather than the original depth D_org.
[0048] For example, the localization module 400 uses the depth information D_ref and the camera configuration information 250 to project the in-image positions of the real person's feet onto three-dimensional positions in the camera coordinate system. In other words, the localization module 400 projects the in-image positions of the real person's feet onto three-dimensional positions corresponding to the depth information D_ref. At this time, the camera ray direction from the camera to the real person is maintained the same as the original (see the explanatory diagram at the bottom left of FIG. 6). The three-dimensional positions obtained in this way are used as highly accurate real person positions and virtual person positions. It can also be said that the localization module 400 reflects the depth information D_ref obtained from the depth map on the person positions while maintaining the original camera ray direction.
[0049] Furthermore, the localization module 400 converts the real person position and the virtual person position in the camera coordinate system (CX, CY, CZ) into the real person position and the virtual person position in the world coordinate system (X, Y, Z) by using the camera configuration information 250. Then, the localization module 400 projects (places) a virtual person having a three-dimensional pose at the virtual person position in the virtual target space SP-V.
[0050] Thus, according to the second example of the localization process, the accuracy of the localization process can be improved by using a depth map. As a result, the accuracy of the projection of the virtual person into the virtual target space SP-V is also improved, and the sense of incongruity caused by the projection result of the virtual person is suppressed. Furthermore, the improvement in the accuracy of the projection of the virtual person into the virtual target space SP-V leads to an improvement in the accuracy of the analysis process by the person analysis module 600.
[0051] Furthermore, the process of acquiring the depth information D_ref from the depth map (lookup table) is extremely simple, the processing load is light, and high-speed processing is possible. High-speed processing is preferable from the viewpoint of real-time processing. That is, according to the second example of localization processing, it is possible to realize real-time projection processing with high accuracy.
[0052] 5. Visualization Module The visualization module 500 visualizes the virtual target space SP-V and the virtual person projected therein by displaying them on the display device 50. The object configuration in the virtual target space SP-V is obtained from the virtual space configuration information 150. The virtual person is drawn to have an estimated three-dimensional pose. The virtual person may be drawn as an avatar having a three-dimensional pose. Attribute information (e.g., gender, age) obtained by the image analysis module 300 may be reflected in the avatar.
[0053] For example, the visualization module 500 may generate an image of the virtual target space SP-V and the virtual person as seen from the virtual camera CAM-V based on the camera configuration information 250, and display the generated image on the display device 50. In this case, a generated image corresponding to the two-dimensional image IMG captured by the real camera CAM-R will be displayed on the display device 50. The visualization process may be performed in real time.
[0054] In the case of MTMC (Multi-Target Multi-Camera Tracking), multiple 2D images IMG captured by multiple real cameras CAM-R are used, and tracking and person re-identification are performed on multiple real people in parallel. The visualization module 500 simultaneously displays multiple virtual people corresponding to the multiple real people on the display device 50.
[0055] The same real person is assigned unique personal identification information. It is also possible that the same real person appears simultaneously in two or more two-dimensional images IMG captured by two or more real cameras CAM-R. In this case, the accuracy of estimating the position of the same real person is improved because two-dimensional images IMG captured at different angles can be used. On the other hand, to avoid overlapping display of two or more virtual people corresponding to the same real person, the visualization module 500 may display only a single virtual person for the same real person on the display device 50.
[0056] 6. Person Analysis Module The person analysis module 600 analyzes the virtual person projected into the virtual target space SP-V. For example, the person analysis module 600 performs an "action estimation process" to estimate the action of the virtual person in the virtual target space SP-V based on the relationship between the virtual person having a three-dimensional pose and the object configuration in the virtual target space SP-V. The object configuration in the virtual target space SP-V is obtained from the virtual space configuration information 150. The action of the virtual person in the virtual target space SP-V is equivalent to the action of the real person in the real target space SP-R. In other words, the person analysis module 600 can estimate the action of the real person in the real target space SP-R by estimating the action of the virtual person in the virtual target space SP-V. Since the action of the person is estimated based on the relationship between the virtual person having a three-dimensional pose and the object configuration, the estimation accuracy is improved compared to when the action of the person is directly estimated from the two-dimensional image IMG.
[0057] For example, the person analysis module 600 estimates the behavior of a virtual person relative to virtual objects in the virtual target space SP-V based on the relationship between the virtual person having a three-dimensional pose and the configuration of each virtual object. The behavior of a virtual person relative to virtual objects in the virtual target space SP-V is equivalent to the behavior of a real person relative to real objects in the real target space SP-R. In other words, the person analysis module 600 can estimate the behavior of a real person relative to real objects in the real target space SP-R by estimating the behavior of a virtual person relative to virtual objects in the virtual target space SP-V. Because the behavior of a person relative to an object is estimated based on the relationship between the virtual person having a three-dimensional pose and the object configuration, the estimation accuracy is improved compared to when the behavior is estimated directly from the two-dimensional image IMG.
[0058] A specific example of the behavior estimation process performed by the person analysis module 600 will be described below.
[0059] 6-1. Gaze Estimation 7 is a conceptual diagram for explaining an example of gaze estimation processing. Person analysis module 600 includes gaze estimation module 610 that performs gaze estimation processing. Gaze estimation module 610 estimates which virtual object a virtual person is looking at, thereby estimating which real object a real person is looking at.
[0060] More specifically, the gaze estimation module 610 estimates the eye ray of the virtual person based on information about the virtual person having a three-dimensional pose. For example, the virtual person's facial orientation can be determined from the virtual person's three-dimensional pose. The facial orientation is considered to be the virtual person's gaze direction. As another example, the gaze direction may be estimated from the virtual person's three-dimensional pose by using a machine learning model. A line extending from the virtual person's face position in the gaze direction is set as the gaze ray. Then, the gaze estimation module 610 determines whether the virtual person's gaze ray intersects with any virtual object defined in the virtual target space SP-V.
[0061] For example, the gaze estimation module 610 determines whether the gaze ray intersects with any virtual object using a ray-triangle intersection algorithm. For example, when each virtual object is represented by a bounding box, the surface of the bounding box is represented by a combination of 12 triangular planes. In the example shown in FIG. 7, a triangle is defined by three vertices A, B, and C, and the gaze ray is represented by a combination of a starting point O and a direction d_g. The gaze estimation module 610 calculates the intersection point P between the plane on which the triangle exists and the gaze ray. If the intersection point P exists within the triangle, it is determined that the gaze ray intersects with the virtual object that owns that triangle. The gaze estimation module 610 can determine whether the gaze ray intersects with any virtual object by performing the above determination process on all triangles defined in the virtual target space SP-V. If it is determined that the gaze ray intersects with multiple virtual objects, the gaze estimation module 610 selects one virtual object that is closest to the starting point O of the gaze ray. The gaze estimation module 610 then estimates the closest virtual object that intersects with the gaze ray as the virtual object that the virtual person is looking at.
[0062] Note that the ray-triangle intersection algorithm is an example, and the present disclosure is not limited thereto. Other shapes may be used instead of triangles. At the very least, the ray-triangle intersection algorithm is extremely simple, has a light processing load, and is capable of high-speed processing. High-speed processing is preferable from the perspective of real-time processing.
[0063] In this way, it is possible to estimate with high accuracy which virtual object the virtual person is looking at based on the three-dimensional posture of the virtual person and the virtual space configuration information 150. In other words, it is possible to estimate with high accuracy which real object the real person is looking at. If the virtual space configuration information 150 indicates the category of each virtual object, it is possible to estimate with high accuracy which real object in which category the real person is looking at. By estimating which real object the real person is looking at, it is possible to know, for example, what the real person is interested in.
[0064] 6-2. Grasp Estimation When a real person's hand is inside any real object, it is highly likely that the real person is grasping or attempting to grasp an item stored in that real object. From this perspective, a grasp estimation process is performed to estimate whether the real person is grasping or attempting to grasp something.
[0065] 8 is a conceptual diagram for explaining an example of the grasp estimation process. The person analysis module 600 includes a grasp estimation module 620 that performs the grasp estimation process. The grasp estimation module 620 estimates whether the virtual person's hand is located inside any virtual object, thereby estimating whether the real person's hand is located inside any real object.
[0066] More specifically, the grasp estimation module 620 estimates the hand positions of the virtual person based on information about the virtual person having a three-dimensional pose. That is, the grasp estimation module 620 estimates the hand positions of the three-dimensional pose as the hand positions of the virtual person. Then, the grasp estimation module 620 determines whether the virtual person's hands are within any of the virtual objects defined in the virtual target space SP-V.
[0067] For example, the grasp estimation module 620 uses a point-cube detection algorithm to determine whether the virtual person's hand is inside any virtual object. For example, each virtual object is represented by a bounding box. In the example shown in FIG. 8, the bounding box of a virtual object is defined by vertices A to H. The position of the center point I of the bounding box is calculated from the positions of vertices D and F (I=(D+F) / 2). Point P is the position of the virtual person's hand. A vector V is defined that points from the center point I to point P. The three axes that define the bounding box are the x-axis, y-axis, and z-axis. [Vx, Vy, Vz] are the x-axis component, y-axis component, and z-axis component of vector V. Lx, Ly, and Lz are the lengths of the bounding box along the x-axis, y-axis, and z-axis directions. The grasp estimation module 620 determines whether the following condition is met: 2 × Vx≦Lx, 2 × Vy≦Ly, 2 × Vz≦Lz. If the condition is met, point P is determined to be inside the bounding box. That is, the virtual person's hand is estimated to be inside the virtual object represented by the bounding box.
[0068] The point-cube detection algorithm is merely an example, and the present disclosure is not limited thereto. However, the point-cube detection algorithm is extremely simple, has a light processing load, and is capable of high-speed processing. High-speed processing is preferable from the perspective of real-time processing.
[0069] In this way, based on the three-dimensional posture of the virtual person and the virtual space configuration information 150, it is possible to estimate with high accuracy whether the virtual person's hand is inside any virtual object. In other words, it is possible to estimate with high accuracy whether the real person's hand is inside any real object. If the real person's hand is inside any real object, it can be determined that the real person is at least interested in an item inside that real object. Furthermore, if the real person's hand is inside any real object, it is highly likely that the real person is grasping or attempting to grasp an item inside that real object. Therefore, the grasp estimation process can roughly estimate whether the real person is grasping or attempting to grasp something. If the virtual space configuration information 150 indicates the category of each virtual object, it is also possible to identify in more detail the item the real person is grasping or attempting to grasp.
[0070] 6-3. Human Flow Estimation 9 is a conceptual diagram for explaining an example of people flow estimation processing. The person analysis module 600 includes a people flow estimation module 630 that performs people flow estimation processing. The people flow estimation module 630 estimates the flow of virtual people in the virtual target space SP-V, thereby estimating the flow of real people in the real target space SP-R.
[0071] The people flow estimation process uses area definition objects (see [C] in Figure 2). An area definition object is an object for defining an area within the virtual target space SP-V. The virtual space configuration information 150 indicates the configuration of each area definition object defined in the virtual target space SP-V. The people flow estimation module 630 estimates the flow of virtual people in the virtual target space SP-V based on the relationship between the virtual people having three-dimensional postures and the configuration of the area definition objects.
[0072] More specifically, the people flow estimation module 630 estimates the position of the virtual person's feet based on information about the virtual person's three-dimensional pose. For example, the virtual person's left and right feet are identified based on the virtual person's three-dimensional pose, and the midpoint between the left and right foot positions is used as the "foot" position. The people flow estimation module 630 then determines which area definition object the virtual person's feet are located in. This determination is made, for example, based on the point-cube detection algorithm described in Section 6-2 above. The people flow estimation module 630 identifies the area definition object in which the virtual person's feet are located and determines that the virtual person is located in the area defined by that area definition object. Furthermore, the people flow estimation module 630 estimates the flow of the virtual person in the virtual target space SP-V by detecting changes in the area in which the virtual person is located.
[0073] In this way, the flow of a virtual person in the virtual target space SP-V can be estimated with high accuracy based on the three-dimensional posture of the virtual person and the virtual space configuration information 150. That is, the flow of a real person in the real target space SP-R can be estimated with high accuracy. The flow of a real person in the real target space SP-R can be used for various purposes. For example, based on the flow of a real person in the real target space SP-R, it is possible to investigate the congestion situation in the real target space SP-R and identify the causes of congestion. As another example, it is possible to analyze what a real person is interested in based on the flow of a real person in the real target space SP-R.
[0074] 6-4. Displaying analysis results The person analysis module 600 displays the analysis results on the display device 50. The analysis results may be statistical information or time transition information. Information may be classified by person attributes (e.g., gender, age).
[0075] 7. Camera calibration and matching process example The calibration module 200 performs camera calibration to determine the camera parameters of the real camera CAM-R. The calibration module 200 also performs camera alignment to correct the camera parameters so that the real target space SP-R as seen from the real camera CAM-R is aligned with the virtual target space SP-V. That is, the calibration module 200 performs camera calibration and alignment to determine the camera parameters of the real camera CAM-R so that the real target space SP-R as seen from the real camera CAM-R is aligned with the virtual target space SP-V. The camera calibration and alignment ensures the accuracy of the localization process (see Section 4), visualization process (see Section 5), and person analysis process (see Section 6).
[0076] A specific example of the camera calibration and matching process will be described below.
[0077] FIG. 10 is a conceptual diagram for explaining an example of camera calibration and matching processing. In this example, camera calibration is performed based on the PnP (Perspective-n-Point) method. The PnP method uses a point cloud consisting of n points, where n is an integer equal to or greater than 2. 3D point cloud coordinate information is coordinate information of the point cloud in 3D space (world coordinate system). 2D point cloud coordinate information is coordinate information of the point cloud in the image coordinate system when the point cloud is photographed by a camera. When the 3D point cloud coordinate information and the 2D point cloud coordinate information are given, the camera parameters can be calculated by solving the PnP problem.
[0078] In this example, the point cloud (n points) is acquired from a marker 210, which is a mark placed in space. For example, the marker 210 is a rectangle, and the four vertices M1 to M4 of the rectangle are used as the point cloud. The marker 210 has a predetermined pattern and can be recognized on an image.
[0079] More specifically, a real marker 210-R is placed at a predetermined real position photographed by a real camera CAM-R in the real target space SP-R. Meanwhile, a virtual marker 210-V is placed at a predetermined virtual position photographed by a virtual camera CAM-V in the virtual target space SP-V. Here, the predetermined virtual position in the virtual target space SP-V corresponds to a predetermined real position in the real target space SP-R. The real marker 210-R and the virtual marker 210-V have the same shape, orientation, and pattern. Therefore, the point cloud (vertices M1 to M4) of the virtual marker 210-V corresponds to the point cloud (vertices M1 to M4) of the real marker 210-R. When placing the virtual marker 210-V at a predetermined virtual position in the virtual target space SP-V, the customization module 100 shown in FIG. 2 is used.
[0080] The calibration module 200 acquires, as a query image, a two-dimensional image IMG captured by a real camera CAM-R installed in a real target space SP-R. The query image shows a real marker 210-R placed at a predetermined real position. The calibration module 200 has configuration information of the marker 210, and detects the real marker 210-R and a point cloud (four vertices M1 to M4) in the query image by performing pattern matching. Then, the calibration module 200 acquires the position of the point cloud in the query image as "two-dimensional point cloud coordinate information 211."
[0081] Meanwhile, the calibration module 200 acquires the position of the point cloud (four vertices M1 to M4) of the virtual marker 210-V in the virtual target space SP-V as “three-dimensional point cloud coordinate information 212.” The position of the point cloud of the virtual marker 210-V in the virtual target space SP-V is obtained from the virtual space configuration information 150.
[0082] The calibration module 200 determines the camera parameters by solving the PnP problem based on the thus obtained 2D point cloud coordinate information 211 and 3D point cloud coordinate information 212. What is important here is that, according to this method, camera alignment is achieved at the same time as the camera parameters are determined. Since the 2D point cloud coordinate information 211 obtained from the real target space SP-R and the 3D point cloud coordinate information 212 obtained from the virtual target space SP-V are combined, camera alignment is achieved at the same time as the camera parameters are determined.
[0083] As described above, according to this example, it is possible to realize both camera calibration and camera matching processing in one step, which is preferable from the viewpoint of reducing the processing load.
[0084] 8.Analysis of multiple real-world target spaces System 1 may analyze multiple real target spaces SP-R. However, if the spatial shapes or camera arrangements of the multiple real target spaces SP-R are different, the analysis results cannot be simply compared, reused, or integrated between the multiple real target spaces SP-R. Therefore, in the following, we will consider multiple real target spaces SP-R that have the same spatial shape.
[0085] FIG. 11 is a conceptual diagram for explaining an example of multiple real target spaces SP-R. FIG. 11 shows real target spaces SP-R-1 to SP-R-4 as an example. The real target spaces SP-R-1 to SP-R-4 have the same shape (unified shape, common shape). For example, the real target spaces SP-R-1 to SP-R-4 are the interior spaces of the same type of vehicle with the same body shape. Multiple vehicles with such the same shape are used for various industries and business types (e.g., mobile convenience store, mobile bookstore, mobile general store, mobile izakaya). Note that each user can freely design the object configuration in each real target space SP-Ri (i = 1 to 4).
[0086] Furthermore, real cameras CAM-R-1 to CAM-R-4 with the same camera parameters are installed in each of the real target spaces SP-R-1 to SP-R-4, respectively. In other words, the positions, orientations, and internal parameters of the real cameras CAM-R-1 to CAM-R-4 installed in each of the real target spaces SP-R-1 to SP-R-4 are unified.
[0087] The virtual target spaces SP-V-1 to SP-V-4 represent the real target spaces SP-R-1 to SP-R-4, respectively. The object configuration in the virtual target space SP-Vi (i=1 to 4) can be customized by the user through the customization module 100. The category (meaning) of each object in the virtual target space SP-Vi can also be freely set by the user through the customization module 100. The virtual space configuration information 150 indicates the configuration of each object defined in each virtual target space SP-Vi.
[0088] In addition, a virtual camera CAM-Vi corresponding to the real camera CAM-Ri is installed in the virtual target space SP-Vi. A pair of a corresponding real camera CAM-Ri and a corresponding virtual camera CAM-Vi has the same camera parameters.
[0089] The camera calibration and matching process described above is also performed in advance. Because the spatial shapes and camera parameters are identical between the real target spaces SP-R-1 to SP-R-4 and the virtual target spaces SP-V-1 to SP-V-4, it is sufficient to perform the camera calibration and matching process only once. For example, the camera calibration and matching process is performed for the combination of the real target space SP-R-1 and the virtual target space SP-V-1. The camera configuration information 250 obtained thereby can be reused for the other real target spaces SP-R-2 to SP-R-4 and the other virtual target spaces SP-V-2 to SP-V-4. This reduces the processing load required for the camera calibration and matching process.
[0090] The system 1 may analyze the real target spaces SP-R-1 to SP-R-4 one by one, or may analyze two or more of the real target spaces SP-R-1 to SP-R-4 in parallel. Among the real target spaces SP-R-1 to SP-R-4, the first real target space SP-RX is the target of analysis this time. Among the virtual target spaces SP-V-1 to SP-V-4, the first virtual target space SP-VX corresponds to (is equivalent to) the first real target space SP-RX. The system 1 performs the various processes described above for the first real target space SP-RX and the first virtual target space SP-VX (see Sections 3 to 6). For example, the system 1 (person analysis module 600) performs the behavior estimation process described in Section 6 above for the first real target space SP-RX and the first virtual target space SP-VX. The system 1 may switch the first virtual target space SP-RX among the real target spaces SP-R-1 to SP-R-4. The system 1 may select two or more of the real target spaces SP-R-1 to SP-R-4 in parallel as the first virtual target space SP-RX.
[0091] The system 1 (person analysis module 600) displays the analysis results on the display device 50. The analysis results may be statistical information or time transition information. The information may be classified by person attributes (e.g., gender, age). The system 1 (person analysis module 600) may display on the display device 50 multiple types of analysis results obtained by the behavior estimation process for each of the real target spaces SP-R-1 to SP-R-4.
[0092] As explained above, the spatial shape and camera parameters are the same among multiple real target spaces SP-R, so it is possible to easily compare, reuse, and integrate analysis results between multiple real target spaces SP-R.
[0093] Furthermore, since the spatial shape and camera parameters are the same among multiple real target spaces SP-R, it is sufficient to perform the camera calibration and matching process only once, thereby reducing the processing load required for the camera calibration and matching process.
[0094] In addition, the user can flexibly set the configuration and category (meaning) of objects in each virtual target space SP-Vi, making analysis easier. [Explanation of symbols]
[0095] 1...Real-to-virtual projection system (target space analysis system), 100...Customization module, 150...Virtual space configuration information, 200...Calibration module, 250...Camera configuration information, 300...Image analysis module, 400...Localization module, 500...Visualization module, 600...Person analysis module, CAM-R...Real camera, CAM-V...Virtual camera, SP-R...Real target space, SP-V...Virtual target space
Claims
1. one or more processors; one or more storage devices; Equipped with The plurality of real target spaces have the same shape, and a real camera having the same camera parameters is installed in each of the plurality of real target spaces; the one or more storage devices are configured to store virtual space configuration information indicating a configuration of objects defined in each of a plurality of virtual object spaces representing each of the plurality of real object spaces; the one or more processors: Detecting a real person appearing in an image captured by the real camera installed in a first real target space among the plurality of real target spaces; estimating a three-dimensional pose of the real person based on the image; projecting a virtual person having the three-dimensional pose and representing the real person into a first virtual object space representing the first real object space; and performing a behavior estimation process for estimating a behavior of the real person in the first real target space by estimating a behavior of the virtual person in the first virtual target space based on a relationship between the virtual person having the three-dimensional pose and the configuration of the object in the first virtual target space. It was configured as Projection system.
2. 10. The projection system of claim 1, The one or more processors are further configured to switch the first virtual object space among the plurality of real object spaces, or to select two or more of the plurality of real object spaces as the first virtual object space in parallel. Projection system.
3. 3. A projection system according to claim 2, The one or more processors are further configured to display, on a display device, a plurality of types of analysis results obtained by the behavior estimation process for each of the plurality of real target spaces. Projection system.
4. 4. A projection system according to any one of claims 1 to 3, the objects defined in the first virtual object space include virtual objects that correspond to real objects in the first real object space; In the behavior estimation process, the one or more processors are configured to estimate a behavior of the real person relative to the real object in the first real object space by estimating a behavior of the virtual person relative to the virtual object in the first virtual object space based on a relationship between the virtual person having the three-dimensional pose and the configuration of the virtual object in the first virtual object space. Projection system.
5. 4. A projection system according to any one of claims 1 to 3, the objects defined in the first virtual object space include an area definition object for defining an area within the first virtual object space; In the behavior estimation process, the one or more processors are configured to execute a people flow estimation process for estimating a flow of the real person in the first real target space by estimating a flow of the virtual person in the first virtual target space based on a relationship between the virtual person having the three-dimensional pose and the configuration of the area definition object. Projection system.
Citation Information
Patent Citations
Virtual object operating system and virtual object operating method
JP2021068405A
Work estimation apparatus, method and program
JP2022048017A
Sales promotion support system
JP2005309951A
Personal behavior analysis apparatus and personal behavior analysis program
JP2010002997A
Facility information classification system and facility information classification program
JP2011232864A