Projection system

JP7916948B2Active Publication Date: 2026-09-08TOYOTA JIDOSHA KK
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024080695
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2026-09-08
Estimated Expiration
2044-05-17

AI Technical Summary

Benefits of technology

【0007】 本開示によれば、複数のリアル対象空間の間で空間形状及びカメラパラメータは同一であるため、複数のリアル対象空間の間で解析結果を容易に比較したり流用したり統合したりすることが可能となる。 また、本開示によれば、リアル対象空間におけるリアル人物の3次元姿勢が推定され、その3次元姿勢を有しリアル人物を表現したバーチャル人物がバーチャル対象空間に投影される。そして、バーチャル対象空間におけるバーチャル人物の行動を推定することによって、リアル対象空間におけるリアル人物の行動が推定される。従って、リアル人物の行動を2次元画像から直接推定する場合と比較して、リアル人物の行動をより正確に推定することが可能となる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007916948000001
    Figure 0007916948000001
  • Figure 0007916948000002
    Figure 0007916948000002
  • Figure 0007916948000003
    Figure 0007916948000003
Patent Text Reader

Abstract

To facilitate analysis based on an image captured by a camera installed in a target space.SOLUTION: A plurality of real target spaces have the same shape, and a real camera having the same camera parameter is installed in each real target space. A configuration of objects is defined in a virtual target space representing the real target space. A projection system detects a real human shown in an image captured by the real camera installed in the real target space, and estimates a three-dimensional pose of the real human. The projection system projects a virtual human having a three-dimensional pose and representing the real human onto the virtual target space. The projection system estimates an action of the real human in the real target space by estimating an action of the virtual human in the virtual target space on the basis of a relationship between the virtual human having the three-dimensional pose and the configuration of objects in the virtual target space.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technology for projecting a person in a real object space into a virtual object space.

Background Art

[0002] Patent Document 1 discloses that a population data file is provided, in which person-specific flow line data obtained by tracking the behavior of each person in a monitoring area is accumulated as a population. The flow line data stored in the group data file is used to calculate statistics related to human behavior analysis

[0003] As other technologies, Patent Document 2, Patent Document 3, and Patent Document 4 are known.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Patent Document 3

Patent Document 4

Summary of Invention

Problem to be Solved by Invention

[0005] Images captured by a camera installed in an object space can be used to analyze the object space. For example, based on images captured by a camera installed in the object space, the behavior of a person present in the object space can be estimated. However, since the shape of the object space and the arrangement of cameras in the object space vary, analysis results cannot be simply compared, diverted, or integrated between a plurality of object spaces.

Means for Solving the Problem

[0006] One aspect of this disclosure relates to projection systems. The projection system comprises one or more processors and one or more storage devices. Multiple real-world target spaces have the same shape, and each of these real-world target spaces is equipped with a real camera having the same camera parameters. One or more storage devices are configured to store virtual space configuration information that indicates the configuration of objects defined in each of the multiple virtual space which represents each of the multiple real space. One or more processors detect real people in images captured by a real camera installed in a first real space among multiple real space locations. One or more processors estimate the 3D pose of a real person based on an image. One or more processors project a virtual person, which has a three-dimensional posture and represents a real person, onto a first virtual object space that represents a first real object space. One or more processors perform an action estimation process to estimate the actions of a real person in a first real object space by estimating the actions of the virtual person in the first virtual object space based on the relationship between the virtual person having a three-dimensional posture and the configuration of objects in the first virtual object space. [Effects of the Invention]

[0007] According to this disclosure, since the spatial shape and camera parameters are identical across multiple real-world target spaces, it becomes possible to easily compare, reuse, and integrate analysis results across multiple real-world target spaces. Furthermore, according to this disclosure, the 3D pose of a real person in a real object space is estimated, and a virtual person representing the real person with that 3D pose is projected into the virtual object space. Then, by estimating the actions of the virtual person in the virtual object space, the actions of the real person in the real object space are estimated. Therefore, it is possible to estimate the actions of the real person more accurately compared to directly estimating the actions of the real person from a 2D image. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram illustrating the overview of a real-to-virtual projection system. [Figure 2] This is a conceptual diagram used to explain virtual space configuration information. [Figure 3] This is a conceptual diagram to explain the camera configuration information. [Figure 4] This is a conceptual diagram illustrating the overview of the visualization and analysis functions. [Figure 5] This is a conceptual diagram illustrating an example of an image analysis module. [Figure 6] This is a conceptual diagram illustrating a method for improving the accuracy of localization processing. [Figure 7] This is a conceptual diagram illustrating an example of gaze estimation processing. [Figure 8] This is a conceptual diagram illustrating an example of the grip estimation process. [Figure 9] This is a conceptual diagram illustrating an example of pedestrian flow estimation processing. [Figure 10] This is a conceptual diagram illustrating an example of camera calibration and alignment processing. [Figure 11] This is a conceptual diagram illustrating examples of multiple real-world object spaces. [Modes for carrying out the invention]

[0009] Embodiments of this disclosure will be described with reference to the attached drawings.

[0010] 1. Outline of Real-to-Virtual Projection System FIG. 1 is a conceptual diagram for explaining the outline of a real-to-virtual projection system 1. A real target space SP-R is an actual three-dimensional space, and is a three-dimensional space that is the target of various analyses. A virtual target space SP-V is a virtual three-dimensional space representing the real target space SP-R. In other words, the virtual target space SP-V is a virtual three-dimensional space imitating the real target space SP-R. The real target space SP-R and the virtual target space SP-V are represented by the same world coordinate system (X, Y, Z).

[0011] Various physical objects exist in the real target space SP-R. Examples of the physical objects include walls, pillars, doors, desks, chairs, shelves, boxes, displays, electronic devices, trees, and the like. Hereinafter, physical objects existing in the real target space SP-R are referred to as real objects. A virtual object corresponding (corresponding) to the real object is defined in the virtual target space SP-V. In other words, a virtual object imitating a real object is defined in the virtual target space SP-V. The configuration of the real object in the real target space SP-R and the configuration of the virtual object in the virtual target space SP-V match with a certain level of accuracy or higher. Note that "configuration" herein is a concept including position, orientation, shape, size, and the like.

[0012] Further, one or more real cameras CAM-R are installed in the real target space SP-R. Each real camera CAM-R is a stationary camera (fixed camera). Then, one or more virtual cameras CAM-V corresponding (corresponding) to the one or more real cameras CAM-R are installed in the virtual target space SP-V. A pair of one corresponding real camera CAM-R and one corresponding virtual camera CAM-V have the same camera parameters. Here, the camera parameters include intrinsic parameters and extrinsic parameters. The intrinsic parameters include distortion parameters, focal length, and the like. The extrinsic parameters include the position and rotation (orientation) of the camera in the world coordinate system. Camera calibration for determining the camera parameters has been performed in advance. Further, processing for aligning the virtual camera CAM-V in the virtual target space SP-V with the real camera CAM-R in the real target space SP-R has also been performed in advance.

[0013] A real-to-virtual projection system 1 projects a person in a real target space SP-R onto a virtual target space SP-V. More specifically, a real person present in the real target space SP-R is photographed by the real camera CAM-R. The real-to-virtual projection system 1 detects a real person appearing in an image photographed by the real camera CAM-R, and estimates a three-dimensional posture of the detected real person. Furthermore, the real-to-virtual projection system 1 generates a virtual person, which is a virtual person representing (imitating) the real person and having the estimated three-dimensional posture. Then, the real-to-virtual projection system 1 projects the virtual person onto the virtual target space SP-V. At this time, the virtual person is projected onto the virtual target space SP-V such that the position of the virtual person in the virtual target space SP-V and the position of the real person in the real target space SP-R match with a precision equal to or higher than a predetermined level. The above projection processing may be performed in real time.

[0014] The real-to-virtual projection system 1 may visualize the virtual object space SP-V and the virtual person projected onto it. For example, the real-to-virtual projection system 1 may generate an image of the virtual object space SP-V and the virtual person as seen from a virtual camera CAM-V, and display that image on a display device. The visualization process may be performed in real time.

[0015] The Real-to-Virtual Projection System 1 may estimate or analyze the actions of virtual figures projected onto the virtual object space SP-V. The actions of virtual figures in the virtual object space SP-V are equivalent to the actions of real figures in the real object space SP-R. That is, by estimating (analyzing) the actions of virtual figures in the virtual object space SP-V, the Real-to-Virtual Projection System 1 can estimate (analyze) the actions of real figures in the real object space SP-R. In this sense, the Real-to-Virtual Projection System 1 can also be called an object space analysis system, a person action estimation system, etc. Hereafter, the Real-to-Virtual Projection System 1 will simply be referred to as "System 1".

[0016] System 1 may consist of a single node or multiple nodes. Figure 1 also shows an example configuration of System 1. System 1 includes one or more real cameras CAM-R, one or more processors 10, one or more storage devices 20, one or more communication devices 30, one or more input devices 40, and one or more display devices 50.

[0017] Processor 10 performs various processes. Examples of processor 10 include general-purpose processors, application-specific processors, CPUs (Central Processing Units), GPUs (Graphics Processing Units), ASICs (Application Specific Integrated Circuits), FPGAs (Field-Programmable Gate Arrays), etc. Storage device 20 stores various information necessary for processing. Examples of storage device 20 include HDDs (Hard Disk Drives), SSDs (Solid State Drives), volatile memory, non-volatile memory, etc. Communication device 30 communicates with the outside world via a communication network. Input device 40 receives various information input from the user of system 1. Examples of input device 40 include keyboards, mice, touch panels, microphones, etc. Display device 50 displays various information. Examples of display device 50 include liquid crystal displays, organic EL displays, heads-up displays (HUDs), etc.

[0018] The processor 10 may execute a computer program. The computer program is stored in the storage device 20. The computer program may be recorded on a computer-readable recording medium. The functions of System 1 may be realized through the cooperation of the processor 10, which executes the computer program, and the storage device 20.

[0019] The following provides a more detailed explanation of System 1.

[0020] 2. Various Information and Functions 2-1. Virtual Space Configuration Information Figure 2 is a conceptual diagram illustrating the virtual space configuration information 150. The virtual space configuration information 150 shows the configuration within the virtual target space SP-V. More specifically, the virtual space configuration information 150 shows the "configuration" of each object defined in the virtual target space SP-V. Here, "configuration" is a concept that includes position, orientation, shape, size, etc. in the world coordinate system (X, Y, Z). For example, each object is represented by a 3D bounding box. In that case, the virtual space configuration information 150 includes information to define the position, orientation, size, etc. of the bounding box of each object.

[0021] Objects defined in the virtual object space SP-V include virtual objects that correspond to real objects in the real object space SP-R. Each virtual object may be assigned identification information (see [A] in Figure 2). Each virtual object may be assigned a color. The virtual space configuration information 150 may indicate the identification information and color of each virtual object. The virtual space configuration information 150 may also indicate the "category" of each virtual object (see [B] in Figure 2). Here, "category" means the type of virtual object (e.g., wall, pillar, door, desk, chair, shelf, box, display, electronic device, tree, etc.). Furthermore, the virtual space configuration information 150 may include a language explanation for each virtual object.

[0022] Objects defined in the virtual object space SP-V may include area definition objects to define areas within the virtual object space SP-V (see [C] in Figure 2). Area definition objects may be represented by thin three-dimensional bounding boxes. Identification information may be assigned to each area definition object. Color may be assigned to each area definition object. Furthermore, the virtual space configuration information 150 may include a linguistic description of each area definition object.

[0023] The virtual space configuration information 150 is generated in advance and stored in the storage device 20.

[0024] System 1 may include a customization module 100. The customization module 100 provides the user with a function to customize (edit) the virtual space configuration information 150. In other words, the customization module 100 provides a user interface for customizing (editing) the virtual space configuration information 150. The customization module 100 displays the virtual space configuration information 150 being edited on the display device 50. The user can freely edit the virtual space configuration information 150 using the input device 40. That is, the user can freely define virtual objects and area definition objects using the input device 40. The customization module 100 updates the virtual space configuration information 150 according to the input from the user.

[0025] 2-2. Camera Configuration Information Figure 3 is a conceptual diagram illustrating the camera configuration information 250. The camera configuration information 250 shows the camera parameters for each real camera CAM-R and each virtual camera CAM-V. The camera parameters include intrinsic parameters and extrinsic parameters. Intrinsic parameters include distortion parameters, focal length, etc. Extrinsic parameters include the camera's position and orientation (rotation) in the world coordinate system. A corresponding pair of one real camera CAM-R and one virtual camera CAM-V has the same camera parameters.

[0026] The camera configuration information 250 is generated in advance and stored in the storage device 20.

[0027] System 1 may include a calibration module 200. The calibration module 200 performs "camera calibration," which determines the camera parameters of the real camera CAM-R. The calibration module 200 also performs "camera alignment," which corrects the camera parameters so that the real target space SP-R, as seen from the real camera CAM-R, matches the virtual target space SP-V. In other words, the calibration module 200 performs "camera calibration and alignment," which determines the camera parameters of the real camera CAM-R so that the real target space SP-R, as seen from the real camera CAM-R, matches the virtual target space SP-V. As a result, camera configuration information 250 showing the camera parameters is obtained.

[0028] Specific examples of camera calibration and alignment processes will be explained later in Section 7.

[0029] 2-3.Various functions Figure 4 is a conceptual diagram illustrating the visualization and analysis functions of System 1. System 1 includes an image analysis module 300, a localization module 400, a visualization module 500, and a human analysis module 600.

[0030] The image analysis module 300 acquires a series of 2D images (IMG) captured by a real camera (CAM-R) installed in the real target space (SP-R). The image analysis module 300 detects real people in the 2D images (IMG). The image analysis module 300 may also perform tracking of the detected real people. The image analysis module 300 may perform human re-identification to recognize the same real person across different real cameras (CAM-R). Furthermore, the image analysis module 300 estimates the 2D pose (2D Pose) and 3D pose (3D Pose) of the real people based on the 2D images (IMG). The processing by the image analysis module 300 may be performed in real time. Details of the processing by the image analysis module 300 will be described in Section 3 below.

[0031] The localization module 400 performs localization processing to estimate the position of a person in the world coordinate system. The real person position is the position where the real person exists in the real target space SP-R. The virtual person position is the position in the virtual target space SP-V that corresponds to the real person position. In other words, the virtual person position in the virtual target space SP-V is set to be consistent with the real person position in the real target space SP-R. The localization module 400 receives the analysis results from the image analysis module 300 and estimates the real person position and the virtual person position based on the analysis results and the camera configuration information 250. Then, the localization module 400 projects (places) the virtual person at the virtual person position in the virtual target space SP-V. The virtual person is a virtual person that represents (models) a real person and has a 3D pose estimated by the image analysis module 300. Details of the processing by the localization module 400 will be explained in Section 4 below.

[0032] The visualization module 500 visualizes the virtual object space SP-V and the virtual person projected onto it on the display device 50. The object configuration within the virtual object space SP-V is obtained from the virtual space configuration information 150. The virtual person has a three-dimensional pose as described above. For example, the visualization module 500 may generate an image of the virtual object space SP-V and the virtual person as seen from the virtual camera CAM-V, based on the camera configuration information 250, and display the generated image on the display device 50. In this case, the generated image corresponding to the two-dimensional image IMG captured by the real camera CAM-R will be displayed on the display device 50. The visualization process may be performed in real time. Details of the processing by the visualization module 500 will be described in Section 5 below.

[0033] The person analysis module 600 analyzes a virtual person projected onto the virtual target space SP-V. For example, the person analysis module 600 performs an "action estimation process" to estimate the actions of the virtual person in the virtual target space SP-V based on the relationship between the virtual person, which has a three-dimensional posture, and the object configuration in the virtual target space SP-V. The object configuration in the virtual target space SP-V is obtained from the virtual space configuration information 150. The actions of the virtual person in the virtual target space SP-V are equivalent to the actions of the real person in the real target space SP-R. That is, the person analysis module 600 can estimate the actions of the real person in the real target space SP-R by estimating the actions of the virtual person in the virtual target space SP-V. Since the person's actions are estimated based on the relationship between the virtual person, which has a three-dimensional posture, and the object configuration, the estimation accuracy is improved compared to when the person's actions are directly estimated from a two-dimensional image IMG. The processing by the person analysis module 600 may be performed in real time. The person analysis module 600 may display the analysis results on the display device 50. Details of the processing performed by the person analysis module 600 will be described in Section 6 below.

[0034] 3. Image Analysis Module Figure 5 is a conceptual diagram illustrating an example of an image analysis module 300. The image analysis module 300 includes a human detector 310, a tracker 320, a human re-identification unit 330, and a pose estimator 340.

[0035] The person detection unit 310 receives a series of 2D images (IMG) captured by the real camera CAM-R. The person detection unit 310 performs person detection processing to detect real people in each 2D image (IMG). The bounding box represents the position of the real person detected in the 2D image (IMG). The person detection processing is a well-known technique, and the method is not particularly limited. For example, YOLOX can be used as the person detection unit 310.

[0036] Tracker 320 automatically tracks the same real person within a series of 2D images (IMG) based on a tracking algorithm. The tracking process is a well-known technique, and the method is not particularly limited. For example, ByteTrack can be used as Tracker 320.

[0037] The human re-identification unit 330 performs human re-identification to recognize the same real person across different real cameras CAM-R. More specifically, the human re-identification unit 330 acquires partial images of real people captured in each 2D image IMG. The partial images enclosed by the bounding box within the 2D image IMG correspond to the partial images of real people. Based on the partial images of real people, the human re-identification unit 330 extracts the features of those real people (hereinafter referred to as "ReID features"). Typically, the human re-identification unit 330 extracts ReID features from each partial image by using a ReID model based on machine learning. The ReID model may also be a model based on a Transformer. Then, the human re-identification unit 330 calculates the similarity between the first real person and the second real person based on the ReID features of the first real person and the ReID features of the second real person. If the similarity is above a threshold, the person re-identification unit 330 determines that the first real person and the second real person are the same real person. Unique person identification information is assigned to the same real person.

[0038] MTMC (Multi-Target Multi-Camera tracking) may be employed. In the case of MTMC, multiple 2D images (IMG) captured by each of multiple real cameras (CAM-R) are used, and tracking and re-identification of multiple real people are performed in parallel.

[0039] The pose estimation unit 340 estimates the 2D pose (2D Pose) and 3D pose (3D Pose) of a real person based on each 2D image (IMG). More specifically, the pose estimation unit 340 acquires partial images of the real person shown in each 2D image (IMG). The partial images enclosed by the bounding box within the 2D image (IMG) correspond to the partial images of the real person. The pose estimation unit 340 extracts key points from the partial images using a machine learning-based pose estimation model to estimate the 2D and 3D poses of the real person. The 2D pose is represented in the image coordinate system of the 2D image (IMG). On the other hand, the 3D pose is represented in the camera coordinate system (CX, CY, CZ). Information on the camera coordinate system (CX, CY, CZ) is obtained from the camera configuration information 250. The 2D and 3D poses are represented by joints, head, hands, feet, etc., and by lines connecting these parts. Note that the pose estimation process is a well-known technique, and the method is not particularly limited. For example, MeTRAbs, TransPose, etc., are used for pose estimation processing.

[0040] In addition, the image analysis module 300 may detect attributes of a real person by analyzing partial images of the real person. Examples of attributes include gender and age.

[0041] 4. Localization Module The localization module 400 performs localization processing to estimate the position of a person in the world coordinate system. The real person's position is the location where the real person exists in the real target space SP-R. The virtual person's position is the position in the virtual target space SP-V that corresponds to the real person's position. In other words, the virtual person's position in the virtual target space SP-V is set to be consistent with the real person's position in the real target space SP-R.

[0042] 4-1. First example of localization processing In the first example of localization processing, the localization module 400 receives information on the 3D pose of a real person from the pose estimation unit 340. The 3D pose is represented in the camera coordinate system (CX, CY, CZ). The position of the real person's 3D pose in the camera coordinate system is used as the real person's position and virtual person's position in the camera coordinate system. Furthermore, the localization module 400 uses the camera configuration information 250 to convert the real person's position and virtual person's position in the camera coordinate system (CX, CY, CZ) to the real person's position and virtual person's position in the world coordinate system (X, Y, Z). Then, the localization module 400 projects (places) the virtual person with the 3D pose at the virtual person's position in the virtual target space SP-V.

[0043] Thus, in the first example, the person's position is estimated based on the 3D orientation position in the camera coordinate system and the camera configuration information 250. However, in order to further improve the accuracy of the person's position estimation, the second example described below may be adopted.

[0044] 4-2. Second example of localization processing Figure 6 is a conceptual diagram illustrating a second example of localization processing. First, a depth map of the virtual target space SP-V as seen from the virtual camera CAM-V is prepared in advance. The depth map provides the depth distribution from the virtual camera CAM-V to each object in the virtual target space SP-V. In particular, the depth map provides at least the depth distribution for the floor in the virtual target space SP-V. The depth distribution is given in the image coordinate system as seen from the virtual camera CAM-V. Such a depth map is generated, for example, based on virtual space configuration information 150 showing the configuration of the virtual target space SP-V and camera configuration information 250 related to the virtual camera CAM-V. The depth map is stored in the storage device 20.

[0045] The localization module 400 receives information about the "2D pose" of a real person from the pose estimation unit 340. The 2D pose is represented in the image coordinate system. The localization module 400 obtains depth information D_ref for the image position of the real person's 2D pose from the depth map. In other words, the localization module 400 obtains depth information D_ref for the image position of the 2D pose by using the depth map as a lookup table (LUT).

[0046] In particular, the localization module 400 may focus on the position of the real person's "feet". More specifically, the localization module 400 estimates the in-image position of the real person's "feet" in the 2D image IMG based on the real person's 2D pose information. For example, the left and right feet of the real person are identified based on the real person's 2D pose, and the midpoint between the positions of the left and right feet is used as the position of the "feet". Then, the localization module 400 obtains depth information D_ref for the in-image position of the real person's feet from the depth map described above.

[0047] The 3D pose of a real person estimated by the pose estimation unit 340 is represented in the camera coordinate system (CX, CY, CZ). The original depth D_org is the depth information of the original 3D pose estimated by the pose estimation unit 340. The accuracy of the original depth D_org is not necessarily high. Therefore, the localization module 400 performs localization processing using depth information D_ref obtained from the depth map, rather than the original depth D_org.

[0048] For example, the localization module 400 uses depth information D_ref and camera configuration information 250 to project the in-image position of a real person's feet onto a 3D position in the camera coordinate system. In other words, the localization module 400 projects the in-image position of a real person's feet onto a 3D position corresponding to the depth information D_ref. At this time, the camera ray direction from the camera to the real person is kept the same as the original (see the explanatory diagram in the lower left of Figure 6). The 3D position obtained in this way is used as a highly accurate real person position and virtual person position. It can also be said that the localization module 400 reflects the depth information D_ref obtained from the depth map onto the person position while maintaining the original camera ray direction.

[0049] Furthermore, the localization module 400 uses the camera configuration information 250 to convert the real person position and virtual person position in the camera coordinate system (CX, CY, CZ) to the real person position and virtual person position in the world coordinate system (X, Y, Z). Then, the localization module 400 projects (places) a virtual person with a three-dimensional orientation at the virtual person position in the virtual target space SP-V.

[0050] Thus, as shown in this second example of localization processing, it is possible to improve the accuracy of localization processing by using depth maps. As a result, the accuracy of projecting virtual figures onto the virtual target space SP-V is also improved, and the sense of incongruity with the projected virtual figures is suppressed. Furthermore, the improvement in the accuracy of projecting virtual figures onto the virtual target space SP-V ultimately leads to an improvement in the accuracy of the analysis processing by the person analysis module 600.

[0051] Furthermore, the process of obtaining depth information D_ref from the depth map (lookup table) is extremely simple, resulting in a light processing load and enabling high-speed processing. High-speed processing is desirable from the perspective of real-timeness. In other words, according to the second example of localization processing, high-precision real-time projection processing becomes possible.

[0052] 5. Visualization Module The visualization module 500 visualizes the virtual object space SP-V and the virtual person projected into it on the display device 50. The object configuration within the virtual object space SP-V is obtained from the virtual space configuration information 150. The virtual person is depicted to have an estimated three-dimensional pose. The virtual person may also be depicted as an avatar with a three-dimensional pose. Attribute information (e.g., gender, age) obtained by the image analysis module 300 may be reflected in the avatar.

[0053] For example, the visualization module 500 may generate images of the virtual target space SP-V and a virtual person as seen from the virtual camera CAM-V, based on the camera configuration information 250, and display the generated images on the display device 50. In this case, the generated image corresponding to the 2D image IMG captured by the real camera CAM-R will be displayed on the display device 50. The visualization process may be performed in real time.

[0054] In the case of MTMC (Multi-Target Multi-Camera tracking), multiple 2D images (IMG) captured by each of multiple real cameras (CAM-R) are used, and tracking and re-identification of multiple real people are performed in parallel. The visualization module 500 simultaneously displays multiple virtual people, each corresponding to one of the multiple real people, on the display device 50.

[0055] Each real person is assigned unique person identification information. It is also possible that the same real person may appear simultaneously in two or more 2D images (IMG) captured by two or more real cameras (CAM-R). In this case, since 2D images (IMG) captured at different angles are available, the accuracy of position estimation for the same real person is improved. On the other hand, to avoid the overlapping display of two or more virtual people corresponding to the same real person, the visualization module 500 may display only a single virtual person on the display device 50 for each real person.

[0056] 6. Human Analysis Module The person analysis module 600 analyzes a virtual person projected onto the virtual target space SP-V. For example, the person analysis module 600 performs an "action estimation process" that estimates the actions of the virtual person in the virtual target space SP-V based on the relationship between the virtual person, which has a three-dimensional posture, and the object configuration in the virtual target space SP-V. The object configuration in the virtual target space SP-V is obtained from the virtual space configuration information 150. The actions of the virtual person in the virtual target space SP-V are equivalent to the actions of the real person in the real target space SP-R. That is, the person analysis module 600 can estimate the actions of the real person in the real target space SP-R by estimating the actions of the virtual person in the virtual target space SP-V. Since the person's actions are estimated based on the relationship between the virtual person, which has a three-dimensional posture, and the object configuration, the estimation accuracy is improved compared to when the person's actions are directly estimated from a two-dimensional image IMG.

[0057] For example, the person analysis module 600 estimates the behavior of a virtual person towards virtual objects in the virtual object space SP-V based on the relationship between a virtual person with a three-dimensional pose and the configuration of each virtual object. The behavior of a virtual person towards virtual objects in the virtual object space SP-V is equivalent to the behavior of a real person towards real objects in the real object space SP-R. That is, by estimating the behavior of a virtual person towards virtual objects in the virtual object space SP-V, the person analysis module 600 can estimate the behavior of a real person towards real objects in the real object space SP-R. Since the behavior of a person towards objects is estimated based on the relationship between a virtual person with a three-dimensional pose and the object configuration, the estimation accuracy is improved compared to when it is estimated directly from a two-dimensional image IMG.

[0058] The following describes a specific example of behavior estimation processing using the Human Analysis Module 600.

[0059] 6-1. Gaze Estimation Processing Figure 7 is a conceptual diagram illustrating an example of gaze estimation processing. The person analysis module 600 includes a gaze estimation module 610 that performs gaze estimation processing. The gaze estimation module 610 estimates which real objects a real person is looking at by estimating which virtual objects a virtual person is looking at.

[0060] More specifically, the gaze estimation module 610 estimates the eye ray of a virtual person based on information about the virtual person having a 3D pose. For example, the orientation of the virtual person's face can be determined from the virtual person's 3D pose. This face orientation is considered to be the direction of the virtual person's gaze. Alternatively, the direction of the gaze may be estimated from the virtual person's 3D pose by using a machine learning model. A line extending from the position of the virtual person's face in the direction of the gaze is set as the eye ray. The gaze estimation module 610 then determines whether the virtual person's eye ray intersects with any virtual object defined in the virtual object space SP-V.

[0061] For example, the gaze estimation module 610 uses the Ray-triangle intersection algorithm to determine whether a gaze ray intersects with any virtual object. For instance, if each virtual object is represented by a bounding box, the surface of that bounding box is represented by a combination of 12 triangular planes. In the example shown in Figure 7, a triangle is defined by three vertices A, B, and C, and the gaze ray is represented by a combination of a starting point O and a direction d_g. The gaze estimation module 610 calculates the intersection point P between the plane on which the triangle exists and the gaze ray. If the intersection point P is within the triangle, the gaze ray is determined to intersect with the virtual object that has that triangle. By performing the above determination process for all triangles defined in the virtual object space SP-V, the gaze estimation module 610 can determine whether a gaze ray intersects with any virtual object. If it is determined that the gaze ray intersects with multiple virtual objects, the gaze estimation module 610 selects the one virtual object closest to the starting point O of the gaze ray. The gaze estimation module 610 then estimates the virtual object that intersects the gaze ray as the virtual object that the virtual character is looking at.

[0062] The ray-triangle intersection algorithm is merely an example, and this disclosure is not limited to it. Other shapes may be used instead of triangles. At least the ray-triangle intersection algorithm is extremely simple, has a low processing load, and enables high-speed processing. High-speed processing is desirable from the standpoint of real-time processing.

[0063] In this way, based on the 3D pose of the virtual person and the virtual space configuration information 150, it is possible to estimate with high accuracy which virtual object the virtual person is looking at. That is, it is possible to estimate with high accuracy which real object a real person is looking at. If the virtual space configuration information 150 indicates the category of each virtual object, it is possible to estimate with high accuracy which real object of which category the real person is looking at. By estimating which real object a real person is looking at, it is possible to know, for example, what the real person is interested in.

[0064] 6-2. Grasp Estimation If a real person's hand is inside a real object, there is a high probability that the real person is grasping or attempting to grasp an item stored inside that object. From this perspective, a grasping estimation process is performed to estimate whether the real person is grasping or attempting to grasp something.

[0065] Figure 8 is a conceptual diagram illustrating an example of the grasp estimation process. The person analysis module 600 includes a grasp estimation module 620 that performs the grasp estimation process. The grasp estimation module 620 estimates whether a real person's hand is inside a real object by estimating whether a virtual person's hand is inside a virtual object.

[0066] More specifically, the grasping estimation module 620 estimates the position of the virtual person's hands based on information about the virtual person having a three-dimensional posture. In other words, the grasping estimation module 620 estimates the position of the hands in the three-dimensional posture as the position of the virtual person's hands. The grasping estimation module 620 then determines whether the virtual person's hands are within any of the virtual objects defined in the virtual object space SP-V.

[0067] For example, the grasp estimation module 620 uses a point-cube detection algorithm to determine if a virtual person's hand is inside one of the virtual objects. For example, each virtual object is represented by a bounding box. In the example shown in Figure 8, the bounding box of a certain virtual object is defined by vertices A to H. The position of the center point I of the bounding box is calculated, for example, from the positions of vertices D and F (I = (D + F) / 2). Point P is the position of the virtual person's hand. A vector V is defined from the center point I to point P. The three axes that define the bounding box are the x-axis, y-axis, and z-axis. [Vx, Vy, Vz] are the x-axis, y-axis, and z-axis components of vector V. Lx, Ly, Lz are the lengths of the bounding box along the x-axis, y-axis, and z-axis directions. The grip estimation module 620 determines whether the conditions "2×Vx≦Lx, 2×Vy≦Ly, 2×Vz≦Lz" are met. If these conditions are met, point P is determined to be inside the bounding box. In other words, the virtual person's hand is estimated to be inside the virtual object represented by that bounding box.

[0068] The point cube detection algorithm is merely an example, and this disclosure is not limited thereto. However, the point cube detection algorithm is extremely simple, has a low processing load, and enables high-speed processing. High-speed processing is desirable from the standpoint of real-time processing.

[0069] In this way, based on the 3D pose of the virtual person and the virtual space configuration information 150, it is possible to estimate with high accuracy whether the virtual person's hand is inside any virtual object. That is, it is possible to estimate with high accuracy whether the real person's hand is inside any real object. If the real person's hand is inside any real object, it can be determined that the real person is at least interested in the item inside that real object. Furthermore, if the real person's hand is inside any real object, there is a high probability that the real person is grasping or attempting to grasp the item inside that real object. Therefore, the grasping estimation process can roughly estimate whether the real person is grasping or attempting to grasp something. If the virtual space configuration information 150 indicates the category of each virtual object, it becomes possible to further identify the item that the real person is grasping or attempting to grasp.

[0070] 6-3. Human Flow Estimation Process Figure 9 is a conceptual diagram illustrating an example of pedestrian flow estimation processing. The pedestrian analysis module 600 includes a pedestrian flow estimation module 630 that performs pedestrian flow estimation processing. The pedestrian flow estimation module 630 estimates the flow of real people in the real target space SP-R by estimating the flow of virtual people in the virtual target space SP-V.

[0071] Area definition objects are used in the human flow estimation process (see [C] in Figure 2). An area definition object is an object used to define an area within the virtual target space SP-V. The virtual space configuration information 150 shows the configuration of each area definition object defined in the virtual target space SP-V. The human flow estimation module 630 estimates the flow of virtual people in the virtual target space SP-V based on the relationship between virtual people with three-dimensional poses and the configuration of area definition objects.

[0072] More specifically, the pedestrian flow estimation module 630 estimates the position of a virtual person's feet based on information about the virtual person having a three-dimensional posture. For example, the left and right feet of the virtual person are identified based on the virtual person's three-dimensional posture, and the midpoint between the positions of the left and right feet is used as the "foot" position. The pedestrian flow estimation module 630 then determines which area-defining object the virtual person's feet are located within. This determination is made, for example, based on the point-cube detection algorithm described in Section 6-2 above. The pedestrian flow estimation module 630 identifies the area-defining object where the virtual person's feet reside and determines that the virtual person is located in the area defined by that area-defining object. Furthermore, the pedestrian flow estimation module 630 estimates the flow of virtual people in the virtual target space SP-V by detecting changes in the area where the virtual person is located.

[0073] In this way, based on the 3D pose of the virtual person and the virtual space configuration information 150, the flow of virtual people in the virtual target space SP-V can be estimated with high accuracy. That is, the flow of real people in the real target space SP-R can be estimated with high accuracy. The flow of real people in the real target space SP-R can be used for various purposes. For example, based on the flow of real people in the real target space SP-R, it is possible to investigate the congestion situation in the real target space SP-R or identify the causes of congestion. As another example, based on the flow of real people in the real target space SP-R, it is possible to analyze what real people are interested in.

[0074] 6-4. Displaying Analysis Results The person analysis module 600 displays the analysis results on the display device 50. The analysis results may be statistical information or time-series information. The information may be classified according to the person's attributes (e.g., gender, age).

[0075] 7. Example of camera calibration and alignment process The calibration module 200 performs camera calibration to determine the camera parameters of the real camera CAM-R. The calibration module 200 also performs camera alignment to correct the camera parameters so that the real target space SP-R, as seen from the real camera CAM-R, matches the virtual target space SP-V. In other words, the calibration module 200 performs camera calibration and alignment to determine the camera parameters of the real camera CAM-R so that the real target space SP-R, as seen from the real camera CAM-R, matches the virtual target space SP-V. This camera calibration and alignment process ensures the accuracy of the localization process (see Section 4), visualization process (see Section 5), and person analysis process (see Section 6).

[0076] The following describes specific examples of camera calibration and alignment processes.

[0077] Figure 10 is a conceptual diagram illustrating an example of camera calibration and alignment processing. In this example, camera calibration is performed based on the PnP (Perspective-n-Point) method. The PnP method uses a point cloud consisting of n points, where n is an integer greater than or equal to 2. The 3D point cloud coordinate information is the coordinate information of the point cloud in 3D space (world coordinate system). The 2D point cloud coordinate information is the coordinate information of the point cloud in the image coordinate system when the point cloud is captured by the camera. Given the 3D point cloud coordinate information and the 2D point cloud coordinate information, camera parameters can be calculated by solving the PnP problem.

[0078] In this example, the point cloud (n points) is obtained from markers 210, which are markers placed in space. For example, marker 210 is a rectangle, and its four vertices M1 to M4 are used as the point cloud. Marker 210 has a predetermined pattern and is recognizable in the image.

[0079] More specifically, a real marker 210-R is placed at a predetermined real position captured by a real camera CAM-R within the real target space SP-R. On the other hand, a virtual marker 210-V is placed at a predetermined virtual position captured by a virtual camera CAM-V within the virtual target space SP-V. Here, the predetermined virtual position within the virtual target space SP-V corresponds to the predetermined real position within the real target space SP-R. Furthermore, the shape, orientation, and pattern of the real marker 210-R and the virtual marker 210-V are all identical. Therefore, the point cloud (vertices M1-M4) of the virtual marker 210-V corresponds to the point cloud (vertices M1-M4) of the real marker 210-R. When placing the virtual marker 210-V at a predetermined virtual position within the virtual target space SP-V, the customized module 100 shown in Figure 2 is used.

[0080] The calibration module 200 acquires a 2D image IMG, captured by a real camera CAM-R installed in the real target space SP-R, as a query image. The query image shows a real marker 210-R placed at a predetermined real position. The calibration module 200 has the configuration information of the marker 210 and detects the real marker 210-R and point cloud (four vertices M1-M4) in the query image by performing pattern matching. The calibration module 200 then acquires the position of the point cloud in the query image as "2D point cloud coordinate information 211".

[0081] Meanwhile, the calibration module 200 acquires the position of the point cloud (four vertices M1 to M4) of the virtual marker 210-V in the virtual target space SP-V as "3D point cloud coordinate information 212". The position of the point cloud of the virtual marker 210-V in the virtual target space SP-V is obtained from the virtual space configuration information 150.

[0082] The calibration module 200 determines the camera parameters by solving a PnP problem based on the two-dimensional point cloud coordinate information 211 and three-dimensional point cloud coordinate information 212 obtained in this way. What is important here is that, according to this method, camera alignment is achieved at the same time that the camera parameters are determined. Because the two-dimensional point cloud coordinate information 211 obtained from the real object space SP-R and the three-dimensional point cloud coordinate information 212 obtained from the virtual object space SP-V are combined, camera alignment is achieved at the same time that the camera parameters are determined.

[0083] As explained above, this example makes it possible to perform both camera calibration and camera alignment processing in a single step. This is advantageous from the standpoint of reducing processing load.

[0084] 8. Analysis of multiple real-world target spaces System 1 may analyze multiple real-world object spaces SP-R. However, if the spatial shape and camera placement differ among the multiple real-world object spaces SP-R, the analysis results cannot be simply compared, reused, or integrated. Therefore, the following discussion will focus on multiple real-world object spaces SP-R having the same spatial shape.

[0085] Figure 11 is a conceptual diagram illustrating examples of multiple real-world object spaces SP-R. Figure 11 shows, as an example, real-world object spaces SP-R-1 to SP-R-4. Real-world object spaces SP-R-1 to SP-R-4 share the same shape (unified shape, common shape). For example, real-world object spaces SP-R-1 to SP-R-4 are the interior spaces of vehicles of the same type with the same body shape. Multiple vehicles with such identical shapes are used in various industries and business types (e.g., mobile convenience stores, mobile bookstores, mobile general stores, mobile bars). Furthermore, the object configuration within each real-world object space SP-Ri (i=1 to 4) can be freely designed by each user.

[0086] Furthermore, each of the real target spaces SP-R-1 to SP-R-4 is equipped with real cameras CAM-R-1 to CAM-R-4, each having the same camera parameters. In other words, the position, orientation, and internal parameters of the real cameras CAM-R-1 to CAM-R-4 installed in each of the real target spaces SP-R-1 to SP-R-4 are unified.

[0087] The virtual target spaces SP-V-1 to SP-V-4 represent the real target spaces SP-R-1 to SP-R-4, respectively. The object configuration within the virtual target space SP-Vi (i=1 to 4) can be customized by the user through the customization module 100. The category (meaning) of each object within the virtual target space SP-Vi can also be freely set by the user through the customization module 100. The virtual space configuration information 150 shows the configuration of each object defined in each virtual target space SP-Vi.

[0088] Furthermore, a virtual camera CAM-Vi, which corresponds to (is equivalent to) the real camera CAM-Ri, is installed in the virtual target space SP-Vi. A pair of one real camera CAM-Ri and one virtual camera CAM-Vi will have the same camera parameters.

[0089] The camera calibration and alignment process described above is also performed in advance. Since the spatial shape and camera parameters are the same between the real target spaces SP-R-1 to SP-R-4 and the virtual target spaces SP-V-1 to SP-V-4, it is sufficient to perform the camera calibration and alignment process only once. For example, the camera calibration and alignment process is performed for the combination of real target space SP-R-1 and virtual target space SP-V-1. The resulting camera configuration information 250 can then be reused for other real target spaces SP-R-2 to SP-R-4 and other virtual target spaces SP-V-2 to SP-V-4. Therefore, the processing load required for the camera calibration and alignment process is reduced.

[0090] System 1 may analyze the real target spaces SP-R-1 to SP-R-4 one by one in sequence, or it may analyze two or more of the real target spaces SP-R-1 to SP-R-4 in parallel. Of the real target spaces SP-R-1 to SP-R-4, the first real target space SP-RX is the target of this analysis. Of the virtual target spaces SP-V-1 to SP-V-4, the first virtual target space SP-VX corresponds to the first real target space SP-RX. System 1 performs the various processes described above on the first real target space SP-RX and the first virtual target space SP-VX (see sections 3 to 6). For example, System 1 (person analysis module 600) performs the behavior estimation process described in section 6 on the first real target space SP-RX and the first virtual target space SP-VX. System 1 may switch the first virtual target space SP-RX among the real target spaces SP-R-1 to SP-R-4. System 1 may select two or more of the real target spaces SP-R-1 to SP-R-4 in parallel as the first virtual target space SP-RX.

[0091] System 1 (person analysis module 600) displays the analysis results on the display device 50. The analysis results may be statistical information or time transition information. The information may be classified according to the person's attributes (e.g., gender, age). System 1 (person analysis module 600) may display multiple types of analysis results obtained by the behavior estimation process for each of the real target spaces SP-R-1 to SP-R-4 on the display device 50.

[0092] As explained above, since the spatial shape and camera parameters are identical across multiple real-world SP-R environments, it becomes possible to easily compare, reuse, and integrate analysis results across multiple real-world SP-R environments.

[0093] Furthermore, since the spatial shape and camera parameters are identical across multiple real-world SP-R target spaces, camera calibration and alignment processing only needs to be performed once. Therefore, the processing load required for camera calibration and alignment processing is reduced.

[0094] Furthermore, since users can flexibly configure the structure and categories (meaning) of objects within each virtual target space SP-Vi, analysis processing becomes easier. [Explanation of symbols]

[0095] 1…Real-to-Virtual Projection System (Target Space Analysis System), 100…Customization Module, 150…Virtual Space Configuration Information, 200…Calibration Module, 250…Camera Configuration Information, 300…Image Analysis Module, 400…Localization Module, 500…Visualization Module, 600…Person Analysis Module, CAM-R…Real Camera, CAM-V…Virtual Camera, SP-R…Real Target Space, SP-V…Virtual Target Space

Claims

1. One or more processors, One or more storage devices and Equipped with, Multiple real-world object spaces have the same shape, and each of the multiple real-world object spaces is equipped with a real camera having the same camera parameters. The one or more storage devices are configured to store virtual space configuration information that indicates the configuration of objects defined in each of the multiple virtual object spaces that represent each of the multiple real object spaces. The one or more processors described above are: The system detects real people appearing in images captured by the real camera installed in the first real target space among the multiple real target spaces. Based on the aforementioned image, the three-dimensional posture of the real person is estimated. The virtual person having the three-dimensional posture and representing the real person is projected onto the first virtual object space which represents the first real object space. Based on the relationship between the virtual person having the three-dimensional posture and the configuration of the object in the first virtual object space, an action estimation process is performed to estimate the actions of the real person in the first real object space by estimating the actions of the virtual person in the first virtual object space. It is configured in such a way A projection system.

2. A projection system according to claim 1, The one or more processors are further configured to switch between the first virtual object spaces among the plurality of real object spaces, or to select two or more of the plurality of real object spaces in parallel as the first virtual object spaces. A projection system.

3. The projection system according to claim 2, The one or more processors are further configured to display on a display device the multiple types of analysis results obtained by the action estimation process for each of the multiple real target spaces. A projection system.

4. A projection system according to any one of claims 1 to 3, The objects defined in the first virtual object space include virtual objects that correspond to real objects in the first real object space, In the behavior estimation process, the one or more processors are configured to estimate the behavior of the real object in the first real object in the first real object by estimating the behavior of the virtual object in the first virtual object in the first virtual object in the first virtual object in the relationship between the virtual person having a three-dimensional posture and the configuration of the virtual object in the first virtual object in the first virtual object in the first virtual object in the first virtual object in the first virtual object in the first real object in the first real object in the first virtual A projection system.

5. A projection system according to any one of claims 1 to 3, The objects defined in the first virtual object space include area definition objects for defining areas within the first virtual object space. In the behavior estimation process, the one or more processors are configured to perform a human flow estimation process that estimates the flow of real people in a first real object space by estimating the flow of the virtual people in the first virtual object space based on the relationship between the virtual people having a three-dimensional posture and the configuration of the area definition object. A projection system.

Citation Information

Patent Citations

  • Sales promotion support system

    JP2005309951A

  • Personal behavior analysis apparatus and personal behavior analysis program

    JP2010002997A

  • Facility information classification system and facility information classification program

    JP2011232864A

  • Virtual object operating system and virtual object operating method

    JP2021068405A

  • Work estimation apparatus, method and program

    JP2022048017A