Projection system
By projecting virtual characters into a virtual object space and using their relationship with objects to infer the actions of real people, the problem of insufficient inference accuracy of camera images in existing technologies is solved, and higher accuracy inference of real people's actions is achieved.
Patent Information
- Application Number
- CN202510380000.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-17
- Filing Date
- 2025-03-28
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies lack sufficient accuracy when inferring human actions based on camera images, making it difficult to accurately infer the actions of real people.
By detecting the 3D pose of real people and projecting virtual people into virtual object space, the actions of real people can be inferred from the relationship between virtual people and objects. Combined with depth mapping and camera calibration technology, the positioning accuracy can be improved.
It enables more accurate inference of the actions of real people, improving the accuracy and real-time performance of the reality-to-virtual projection system.
Smart Images

Figure CN120976303A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a technique for projecting figures from a real-world object space onto a virtual object space. Background Technology
[0002] Patent Document 1 discloses a gaze inference system. The gaze inference system acquires a series of images showing the face of a person being measured. Furthermore, by using a learned model, the gaze inference system infers the gaze position of the person being measured based on the images including the face.
[0003] As technologies associated with the hypothetical space, patent documents 2, 3 and 4 are known.
[0004] Patent Document 1: Japanese Patent Application Publication No. 2022-187547
[0005] Patent Document 2: Japanese Patent Application Publication No. 2024-003401
[0006] Patent Document 3: Japanese Patent Application Publication No. 2022-061305
[0007] Patent Document 4: Japanese Patent Application Publication No. 2010-176662
[0008] Images captured by a camera positioned in an object space can be used to analyze that object space. For example, based on images captured by a camera positioned in an object space, it is possible to infer the actions of a person existing within that object space. The goal is to improve the accuracy of inferring person actions from images captured by a camera. Summary of the Invention
[0009] One aspect of this disclosure relates to projection systems.
[0010] The projection system has one or more processors and one or more storage devices.
[0011] One or more storage devices are configured to store virtual space composition information representing the composition of objects defined in a virtual object space that represents the real object space.
[0012] One or more processors detect real people appearing in images captured by real cameras set in real object space.
[0013] One or more processors infer the 3D pose of a real person based on an image.
[0014] One or more processors project virtual characters with three-dimensional poses and representing real people into a virtual object space.
[0015] One or more processors perform action inference processing, in which the actions of the virtual character in the virtual object space are inferred based on the relationship between the virtual character with a three-dimensional pose and the objects in the virtual object space, thereby inferring the actions of the real character in the real object space.
[0016] According to this disclosure, the three-dimensional pose of a real person in a real-world object space is inferred, and a virtual person having that three-dimensional pose and representing the real person is projected into the virtual object space. Furthermore, the actions of the real person in the real-world object space are inferred by inferring the actions of the virtual person in the virtual object space. Therefore, compared to directly inferring the actions of a real person from a two-dimensional image, the actions of the real person can be inferred more accurately. Attached Figure Description
[0017] Figure 1 This is a schematic diagram used to illustrate the overview of a real-to-virtual projection system.
[0018] Figure 2 It is a schematic diagram used to illustrate the composition information of virtual space.
[0019] Figure 3 This is a schematic diagram used to illustrate the configuration information of a camera.
[0020] Figure 4 This is a diagram used to illustrate the overview of visualization and parsing functions.
[0021] Figure 5 This is a schematic diagram used to illustrate an example of the image parsing module.
[0022] Figure 6 This is a schematic diagram used to illustrate methods for improving the accuracy of positioning processing.
[0023] Figure 7 This is a schematic diagram used to illustrate an example of line-of-sight inference processing.
[0024] Figure 8 This is a schematic diagram used to illustrate an example of the control inference process.
[0025] Figure 9 This is a schematic diagram used to illustrate an example of pedestrian flow inference processing.
[0026] Figure 10 This is a schematic diagram illustrating an example of camera calibration and matching processing.
[0027] Explanation of reference numerals in the attached figures
[0028] 1…Reality-to-virtual projection system (object space resolution system); 100…Custom module; 150…Virtual space composition information; 200…Calibration module; 250…Camera composition information; 300…Image resolution module; 400…Positioning module; 500…Visualization module; 600…Character resolution module; CAM-R…Reality camera; CAM-V…Virtual camera; SP-R…Reality object space; SP-V…Virtual object space. Detailed Implementation
[0029] The embodiments of this disclosure will be described with reference to the accompanying drawings.
[0030] 1. Overview of Reality-to-Virtual Projection Systems
[0031] Figure 1 This is a schematic diagram used to illustrate the overview of a real-to-virtual projection system. The real object space SP-R is the actual existing three-dimensional space, the three-dimensional space that serves as the basis for various analytical objects. The virtual object space SP-V is an imaginary three-dimensional space that represents the real object space SP-R. In other words, the virtual object space SP-V is an imaginary three-dimensional space that simulates the real object space SP-R. The real object space SP-R and the virtual object space SP-V are represented using the same world coordinate system (X, Y, Z).
[0032] In the Reality Object Space (SP-R), various physical objects exist. Examples of physical objects include walls, pillars, doors, tables, chairs, shelves, boxes, monitors, electronic devices, and trees. Hereinafter, physical objects existing in the Reality Object Space (SP-R) will be referred to as real objects. In the Virtual Object Space (SP-V), virtual objects corresponding to these real objects are defined. In other words, virtual objects that simulate real objects are defined in the Virtual Object Space (SP-V). The composition of real objects in the Reality Object Space (SP-R) matches the composition of virtual objects in the Virtual Object Space (SP-V) with a certain level of precision. Furthermore, "composition" here includes concepts such as position, orientation, shape, and size.
[0033] In addition, one or more real-world cameras (CAM-Rs) are set up in the real-world object space SP-R. Each real-world camera (CAM-R) is a still camera (fixed camera). Furthermore, one or more virtual cameras (CAM-Vs) are set up in the virtual object space SP-V, corresponding to one or more real-world camera (CAM-R). The camera pair of one real-world camera (CAM-R) and one virtual camera (CAM-V) has the same camera parameters. These camera parameters include intrinsic parameters and extrinsic parameters. Intrinsic parameters include deformation parameters, focal length, etc. Extrinsic parameters include the camera's position and orientation (rotation) in the world coordinate system. Camera calibration to determine the camera parameters is performed beforehand. Additionally, a process to align the virtual camera (CAM-V) in the virtual object space SP-V with the real camera (CAM-R) in the real-world object space SP-R is also performed beforehand.
[0034] The reality-to-virtual projection system 1 projects a person from the real object space SP-R onto the virtual object space SP-V. More specifically, a real person existing in the real object space SP-R is captured by a real camera CAM-R. The reality-to-virtual projection system 1 detects the real person appearing in the image captured by the real camera CAM-R and infers the detected real person's 3D pose. Furthermore, the reality-to-virtual projection system 1 generates a virtual person with the inferred 3D pose, representing (simulating) the real person. This virtual person is then projected into the virtual object space SP-V. At this time, the virtual person is projected into the virtual object space SP-V in a manner that matches the position of the virtual person in the virtual object space SP-V with the position of the real person in the real object space SP-R to a certain degree of accuracy. This projection process can also be performed in real time.
[0035] The reality-to-virtual projection system 1 can also visualize the virtual object space SP-V and the virtual character projected thereon. For example, the reality-to-virtual projection system 1 can also generate an image of the virtual object space SP-V and the virtual character observed from the virtual camera CAM-V, and display the image on a display device. Visualization processing can also be performed in real time.
[0036] The reality-to-virtual projection system 1 can also infer and resolve the actions of virtual characters projected onto the virtual object space SP-V. The actions of virtual characters in the virtual object space SP-V are equivalent to the actions of real characters in the reality object space SP-R. That is, by inferring (resolving) the actions of virtual characters in the virtual object space SP-V, the reality-to-virtual projection system 1 can infer (resolve) the actions of real characters in the reality object space SP-R. In this sense, the reality-to-virtual projection system 1 can also be called an object space resolution system, a character action inference system, etc. Hereinafter, the reality-to-virtual projection system 1 will be simply referred to as "System 1".
[0037] System 1 can consist of a single node or multiple nodes. Figure 1 The diagram also shows an example of the configuration of System 1. System 1 includes one or more real-world cameras (CAM-R), one or more processors 10, one or more storage devices 20, one or more communication devices 30, one or more input devices 40, and one or more display devices 50.
[0038] Processor 10 performs various processes. Examples of processor 10 include general-purpose processors, application-specific processors, CPUs (Central Processing Units), GPUs (Graphics Processing Units), ASICs (Application Specific Integrated Circuits), and FPGAs (Field-Programmable Gate Arrays). Storage device 20 stores various information required for processing. Examples of storage device 20 include HDDs (Hard Disk Drives), SSDs (Solid State Drives), volatile memory, and non-volatile memory. Communication device 30 communicates with the outside world via a communication network. Input device 40 accepts various information input from the user of system 1. Examples of input device 40 include keyboards, mice, touch panels, and microphones. Display device 50 displays various information. Examples of display device 50 include liquid crystal displays (LCDs), organic EL displays, and head-up displays (HUDs).
[0039] The processor 10 can also execute computer programs. The computer programs are stored in the storage device 20. The computer programs can also be recorded on a computer-readable recording medium. The functions of System 1 can also be realized through the cooperation of the processor 10 executing the computer programs and the storage device 20.
[0040] The following is a more detailed explanation of System 1.
[0041] 2. Various information and functions
[0042] 2-1. Virtual space constitutes information
[0043] Figure 2 This is a schematic diagram illustrating the virtual space composition information 150. The virtual space composition information 150 represents the composition within the virtual object space SP-V. More specifically, the virtual space composition information 150 represents the "composition" of each object defined in the virtual object space SP-V. Here, "composition" is a concept that includes position, orientation, shape, size, etc., in the world coordinate system (X, Y, Z). For example, each object is represented by a three-dimensional bounding box. In this case, the virtual space composition information 150 contains information specifying the position, orientation, size, etc., of the bounding box for each object.
[0044] Objects defined in the virtual object space SP-V include virtual objects that are equivalent to real-world objects in the real-world object space SP-R. Identification information can also be assigned to each virtual object (see [reference]). Figure 2 [A] in the text). Colors can also be assigned to each virtual object. The virtual space composition information 150 can also represent the identification information and color of each virtual object. Additionally, the virtual space composition information 150 can also represent the "category" of each virtual object (see [A] in the text). Figure 2 [B] Here, "category" refers to the type of virtual object (e.g., wall, pillar, door, table, chair, shelf, box, monitor, electronic device, tree, etc.). Furthermore, the virtual space composition information 150 may also include language explanations for each virtual object.
[0045] In the Virtual Object Space SP-V, objects defined within a region can also contain region-defining objects (see [reference]). Figure 2 [C] in the text. Region-defined objects can also be represented using thin 3D bounding boxes. Identification information can also be assigned to each region-defined object. Colors can also be assigned to each region-defined object. Furthermore, the virtual space composition information 150 can also include linguistic descriptions of each region-defined object.
[0046] Virtual space composition information 150 is pre-generated and stored in storage device 20.
[0047] System 1 may also include a customization module 100. Customization module 100 provides the user with the function of customizing (editing) the virtual space composition information 150. In other words, customization module 100 provides a user interface for customizing (editing) the virtual space composition information 150. Customization module 100 displays the virtual space composition information 150 being edited on display device 50. The user can freely edit the virtual space composition information 150 using input device 40. That is, the user can freely define virtual objects and region definition objects using input device 40. Customization module 100 updates the virtual space composition information 150 based on the user's input.
[0048] 2-2. Camera Configuration Information
[0049] Figure 3 This is a schematic diagram illustrating the camera configuration information 250. The camera configuration information 250 represents the camera parameters of each real-world camera (CAM-R) and each virtual camera (CAM-V). The camera parameters include intrinsic parameters and extrinsic parameters. Intrinsic parameters include deformation parameters, focal length, etc. Extrinsic parameters include the camera's position and orientation in the world coordinate system. A pair of one real-world camera (CAM-R) and one virtual camera (CAM-V) has the same camera parameters.
[0050] Camera configuration information 250 is pre-generated and stored in storage device 20.
[0051] System 1 may also include a calibration module 200. The calibration module 200 performs "camera calibration" to determine the camera parameters of the real-world camera CAM-R. Additionally, the calibration module 200 performs "camera alignment" to correct the camera parameters by matching the real-world object space SP-R observed from the real-world camera CAM-R with the virtual object space SP-V. That is, the calibration module 200 performs "camera calibration and alignment" to determine the camera parameters of the real-world camera CAM-R by matching the real-world object space SP-R observed from the real-world camera CAM-R with the virtual object space SP-V. As a result, camera configuration information 250 representing the camera parameters is obtained.
[0052] Furthermore, specific examples of camera calibration and matching processing will be explained in Chapter 7.
[0053] 2-3. Various functions
[0054] Figure 4This is a schematic diagram illustrating the overview of the visualization and analysis functions of System 1. System 1 includes an image analysis module 300, a localization module 400, a visualization module 500, and a human analysis module 600.
[0055] The image analysis module 300 acquires a series of two-dimensional image (IMG) images captured by the reality camera CAM-R positioned in the reality object space SP-R. The image analysis module 300 detects real-world people displayed in the two-dimensional image IMGs. The image analysis module 300 can also track the detected real-world people. The image analysis module 300 can also perform human re-identification processing to identify the same real-world person across different reality cameras CAM-Rs. Furthermore, the image analysis module 300 infers the two-dimensional pose (2DPose) and three-dimensional pose (3DPose) of the real-world people based on the two-dimensional image IMGs. This processing can also be performed in real-time by the image analysis module 300. Details of the processing performed by the image analysis module 300 will be explained in Chapter 3.
[0056] The positioning module 400 performs positioning processing to infer the position of the character in the world coordinate system. The real-world character position is the location where the real-world character exists in the real-world object space SP-R. The virtual character position is the position in the virtual object space SP-V, which corresponds to the real-world character position. That is, the virtual character position in the virtual object space SP-V is set to match the real-world character position in the real-world object space SP-R. The positioning module 400 receives the analysis result from the image analysis module 300 and infers the real-world character position and the virtual character position based on the analysis result and the camera configuration information 250. Furthermore, the positioning module 400 projects (configures) the virtual character onto the virtual character position in the virtual object space SP-V. The virtual character is an imaginary character that represents (simulates) the real-world character and has a three-dimensional pose inferred by the image analysis module 300. Details of the processing performed by the positioning module 400 will be explained in Chapter 4 later.
[0057] The visualization module 500 visualizes the virtual object space SP-V and a virtual character projected onto the virtual object space SP-V on the display device 50. The object composition in the virtual object space SP-V is obtained based on the virtual space composition information 150. As described above, the virtual character has a three-dimensional pose. For example, the visualization module 500 can also generate an image of the virtual object space SP-V and the virtual character observed from the virtual camera CAM-V based on the camera composition information 250, and display this generated image on the display device 50. In this case, a generated image equivalent to a two-dimensional image IMG captured by the real camera CAM-R is displayed on the display device 50. Visualization processing can also be performed in real time. Details of the processing performed by the visualization module 500 will be explained in Section 5 below.
[0058] The character analysis module 600 analyzes the virtual character projected onto the virtual object space SP-V. For example, the character analysis module 600 performs "action inference processing" to infer the action of the virtual character in the virtual object space SP-V based on the relationship between the virtual character with a 3D pose and the object composition in the virtual object space SP-V. The object composition in the virtual object space SP-V is obtained based on the virtual space composition information 150. The action of the virtual character in the virtual object space SP-V is equivalent to the action of the real character in the real object space SP-R. That is, the character analysis module 600 can infer the action of the real character in the real object space SP-R by inferring the action of the virtual character in the virtual object space SP-V. Since the character action is inferred based on the relationship between the virtual character with a 3D pose and the object composition, the inference accuracy is improved compared to directly inferring the character action from the 2D image IMG. Processing can also be performed in real time by the character analysis module 600. The character analysis module 600 can also display the analysis results on the display device 50. Details of the processing performed by the character analysis module 600 will be explained in Chapter 6 later.
[0059] 3. Image parsing module
[0060] Figure 5 This is a schematic diagram used to illustrate an example of the image parsing module 300. The image parsing module 300 includes a human detector 310, a tracker 320, a human re-identification unit 330, and a pose estimator 340.
[0061] A series of two-dimensional image IMGs captured by a reality camera CAM-R are input into the person detection unit 310. The person detection unit 310 performs person detection processing to detect real people displayed in each two-dimensional image IMG. The bounding box indicates the position of the real person detected in the two-dimensional image IMG. Furthermore, person detection processing is a known technique, and its method is not particularly limited. For example, YOLOX can be used as the person detection unit 310.
[0062] Tracker 320 automatically tracks the same real-world figure within a series of 2D image (IMG) files based on a tracking algorithm. Tracking processing is a well-known technique, and its methods are not particularly limited. For example, ByteTrack can be used as tracker 320.
[0063] The person re-identification unit 330 performs human re-identification to recognize the same real-world person across different real-world cameras (CAM-R). More specifically, the person re-identification unit 330 acquires partial images of the real-world person displayed in each two-dimensional image (IMG). The partial image enclosed by a bounding box in the two-dimensional image (IMG) corresponds to a partial image of the real-world person. Based on the partial images of the real-world person, the person re-identification unit 330 extracts the feature quantity of that real-world person (hereinafter referred to as "ReID feature quantity"). Typically, the person re-identification unit 330 extracts the ReID feature quantity from each partial image using a machine learning-based ReID model. The ReID model can also be a transformer-based model. Furthermore, the person re-identification unit 330 calculates the similarity between the first real-world person and the second real-world person based on the ReID feature quantity of the first real-world person and the ReID feature quantity of the second real-world person. If the similarity is above a threshold, the person re-identification unit 330 determines that the first real-world person and the second real-world person are the same real-world person. Inherent person recognition information is assigned to the same real-world person.
[0064] Alternatively, MTMC (Multi-Target Multi-Camera tracking) can be used. In the case of MTMC, multiple 2D images (IMGs) captured by multiple real-world cameras (CAM-R) are used, and multiple real-world figures are tracked and re-identified side by side.
[0065] The pose inference unit 340 infers the two-dimensional pose (2DPose) and three-dimensional pose (3D Pose) of a real person based on each two-dimensional image IMG. More specifically, the pose inference unit 340 obtains partial images of the real person displayed in each two-dimensional image IMG. The partial image enclosed by the bounding box in the two-dimensional image IMG corresponds to a partial image of the real person. The pose inference unit 340 extracts key points from the partial images using a pose inference model based on machine learning to infer the two-dimensional and three-dimensional poses of the real person. The two-dimensional pose is represented in the image coordinate system of the two-dimensional image IMG. On the other hand, the three-dimensional pose is represented in the camera coordinate system (CX, CY, CZ). The camera coordinate system (CX, CY, CZ) information is obtained from the camera configuration information 250. The two-dimensional pose and three-dimensional pose are represented by lines connecting joints, head, hands, feet, etc., to each other. Furthermore, pose inference processing is a known technique, and its method is not particularly limited. For example, MeTRAbs, TransPose, etc., can be used for pose inference processing.
[0066] In addition, the image parsing module 300 can also detect the attributes of real people by parsing a portion of their images. Examples of attributes include gender and age.
[0067] 4. Positioning Module
[0068] The positioning module 400 performs positioning processing to infer the character's position in the world coordinate system. The real-world character's position is the location where the real character exists in the real-world object space SP-R. The virtual character's position is the position within the virtual object space SP-V, which corresponds to the real-world character's position. That is, the virtual character's position in the virtual object space SP-V is set to match the real-world character's position in the real-world object space SP-R.
[0069] 4-1. The first example of location processing
[0070] In the first example of the positioning process, the positioning module 400 receives information about the three-dimensional pose of a real person from the pose inference unit 340. The three-dimensional pose is represented in the camera coordinate system (CX, CY, CZ). The position of the real person's three-dimensional pose in the camera coordinate system is used as the position of the real person and the position of the virtual person in the camera coordinate system. Furthermore, the positioning module 400 transforms the positions of the real person and the virtual person in the camera coordinate system (CX, CY, CZ) into the positions of the real person and the virtual person in the world coordinate system (X, Y, Z) by using the camera configuration information 250. Moreover, the positioning module 400 projects (configures) the virtual person with the three-dimensional pose onto the virtual person position in the virtual object space SP-V.
[0071] Thus, in the first example, the person's position is inferred based on the three-dimensional pose in the camera coordinate system and the camera configuration information 250. To further improve the accuracy of the person's position inference, the second example described below can also be used.
[0072] 4-2. The second example of location processing
[0073] Figure 6 This is a schematic diagram illustrating the second example of the positioning process. First, a depth map of the virtual object space SP-V, observed from the virtual camera CAM-V, is prepared in advance. The depth map gives the depth distribution of each object within the virtual object space SP-V from the virtual camera CAM-V. In particular, the depth map gives at least the depth distribution relative to the floor within the virtual object space SP-V. The depth distribution is given in the image coordinate system observed from the virtual camera CAM-V. Such a depth map is generated, for example, based on virtual space configuration information 150 representing the configuration of the virtual object space SP-V and camera configuration information 250 related to the virtual camera CAM-V. The depth map is stored in the storage device 20.
[0074] The localization module 400 receives information about the "two-dimensional pose" of the real person from the pose inference unit 340. The two-dimensional pose is represented in an image coordinate system. The localization module 400 obtains depth information D_ref, which is the position within the image relative to the two-dimensional pose of the real person, from the aforementioned depth mapping. In other words, the localization module 400 uses the depth mapping in the form of a lookup table (LUT) to obtain the depth information D_ref, which is the position within the image relative to the two-dimensional pose.
[0075] In particular, the positioning module 400 can also focus on the position of the "feet" of the real person. More specifically, the positioning module 400 infers the in-image position of the "feet" of the real person in the 2D image IMG based on the 2D pose information of the real person. For example, based on the 2D pose of the real person, the left and right feet of the real person are determined, and the midpoint between the positions of the left and right feet is used as the position of the "feet". Moreover, the positioning module 400 obtains the depth information D_ref relative to the in-image position of the real person's feet from the aforementioned depth mapping.
[0076] The 3D pose of the real person inferred by the pose inference unit 340 is represented in the camera coordinate system (CX, CY, CZ). The original depth D_org is the depth information of the original 3D pose inferred by the pose inference unit 340. The accuracy of the original depth D_org is not necessarily high. Therefore, the localization module 400 does not use the original depth D_org, but uses the depth information D_ref obtained from the depth map for localization processing.
[0077] For example, the positioning module 400 projects the position of the real person's feet within the image onto a three-dimensional position in the camera coordinate system using depth information D_ref and camera configuration information 250. In other words, the positioning module 400 projects the position of the real person's feet within the image onto a three-dimensional position equivalent to the depth information D_ref. At this time, the direction of the camera light rays from the camera to the real person remains the same as the original (see reference). Figure 6 (See the lower left illustration). The 3D position obtained in this way is used for high-precision positioning of both real and virtual characters. In other words, the positioning module 400, while maintaining the original camera light direction, reflects the depth information D_ref obtained from the depth mapping into the character's position.
[0078] Furthermore, the positioning module 400 transforms the positions of the real and virtual characters in the camera coordinate system (CX, CY, CZ) into the positions of the real and virtual characters in the world coordinate system (X, Y, Z) by using the camera configuration information 250. Moreover, the positioning module 400 projects (configures) a virtual character with a three-dimensional pose onto the virtual character position within the virtual object space SP-V.
[0079] Thus, according to the second example of positioning processing, the accuracy of positioning processing can be improved by utilizing depth mapping. As a result, the accuracy of the projection of the virtual character into the virtual object space SP-V is also improved, thereby suppressing the sense of inconsistency in the projection results relative to the virtual character. In addition, improving the accuracy of the projection of the virtual character into the virtual object space SP-V further helps to improve the accuracy of the parsing processing performed by the character parsing module 600.
[0080] Furthermore, the processing of obtaining depth information D_ref from the depth map (check table) is extremely simple, has a light processing load, and can be performed at high speed. From the perspective of real-time processing, high-speed processing is preferred. That is, based on the second example of positioning processing, real-time projection processing can be achieved with high accuracy.
[0081] 5. Visualization Module
[0082] The visualization module 500 visualizes the virtual object space SP-V and the virtual character projected onto the virtual object space SP-V on the display device 50. The object composition in the virtual object space SP-V is obtained from the virtual space composition information 150. The virtual character is depicted with the inferred three-dimensional pose. The virtual character can also be depicted as a virtual image with a three-dimensional pose. Attribute information (e.g., gender, age) obtained through the image analysis module 300 can also be reflected in the virtual image.
[0083] For example, the visualization module 500 can also generate images of the virtual object space SP-V and virtual characters observed from the virtual camera CAM-V based on the camera configuration information 250, and display the generated images on the display device 50. In this case, the generated image, equivalent to a two-dimensional image IMG captured by the real camera CAM-R, is displayed on the display device 50. Visualization processing can also be performed in real time.
[0084] In the case of MTMC (Multi-Target Multi-Camera tracking), multiple two-dimensional images (IMGs) captured by multiple real-world cameras (CAM-R) are used. Furthermore, multiple real-world figures are tracked and re-identified side-by-side. The visualization module 500 simultaneously displays multiple virtual figures, each corresponding to a real-world figure, on the display device 50.
[0085] The visualization module 500 assigns inherent person recognition information to the same real-world person. It also considers the case where the same real-world person is simultaneously displayed in two or more 2D image IMGs captured by two or more real-world cameras (CAM-R). In this case, by utilizing 2D image IMGs captured from different angles, the accuracy of position inference for the same real-world person is improved. On the other hand, to avoid repeatedly displaying two or more virtual characters equivalent to the same real-world person, the visualization module 500 can also display only a single virtual character on the display device 50 for the same real-world person.
[0086] 6. Character Analysis Module
[0087] The character analysis module 600 analyzes the virtual character projected onto the virtual object space SP-V. For example, the character analysis module 600 performs "action inference processing," inferring the actions of the virtual character in the virtual object space SP-V based on the relationship between the virtual character with a 3D pose and the object composition in the virtual object space SP-V. The object composition in the virtual object space SP-V is obtained from the virtual space composition information 150. The actions of the virtual character in the virtual object space SP-V are equivalent to the actions of the real character in the real object space SP-R. That is, the character analysis module 600 can infer the actions of the real character in the real object space SP-R by inferring the actions of the virtual character in the virtual object space SP-V. Because the character action is inferred based on the relationship between the virtual character with a 3D pose and the object composition, the inference accuracy is improved compared to directly inferring the character action from the 2D image IMG.
[0088] For example, the character analysis module 600 infers the actions of a virtual character relative to virtual objects in the virtual object space SP-V based on the relationship between a virtual character with a 3D pose and the composition of various virtual objects. The actions of a virtual character relative to virtual objects in the virtual object space SP-V are equivalent to the actions of a real character relative to real objects in the real object space SP-R. That is, by inferring the actions of a virtual character relative to virtual objects in the virtual object space SP-V, the character analysis module 600 can infer the actions of a real character relative to real objects in the real object space SP-R. Because the actions of a character relative to objects are inferred based on the relationship between a virtual character with a 3D pose and the composition of objects, the inference accuracy is improved compared to inference directly from a 2D image IMG.
[0089] The following is a specific example of the action inference processing performed by the character analysis module 600.
[0090] 6-1. Gaze Estimation
[0091] Figure 7 This is a schematic diagram illustrating an example of gaze inference processing. The character analysis module 600 includes a gaze inference module 610 that performs gaze inference processing. The gaze inference module 610 infers which real object a real person is observing by inferring which virtual object the virtual character is observing.
[0092] More specifically, the gaze inference module 610 infers the virtual character's eye ray based on information about the virtual character's three-dimensional pose. For example, the orientation of the virtual character's face can be determined from its three-dimensional pose. This orientation is considered the virtual character's gaze direction. Alternatively, the gaze direction can be inferred from the virtual character's three-dimensional pose using a machine learning model. A line extending from the virtual character's face towards the gaze direction is defined as the eye ray. Furthermore, the gaze inference module 610 determines whether the virtual character's eye ray intersects with any virtual object defined in the virtual object space SP-V.
[0093] For example, the gaze inference module 610 determines whether a gaze ray intersects with any virtual object based on a ray-triangle intersection algorithm. For instance, when virtual objects are represented by bounding boxes, the surface of the bounding box is represented by a combination of 12 triangular planes. Figure 7In the example shown, a triangle is defined by three vertices A, B, and C, and the line of sight is represented by a combination of its origin O and direction d_g. The line of sight inference module 610 calculates the intersection point P between the plane containing the triangle and the line of sight. If the intersection point P exists within the triangle, it is determined that the line of sight intersects with a virtual object containing that triangle. By performing the above determination process on all triangles defined within the virtual object space SP-V, the line of sight inference module 610 can determine whether the line of sight intersects with any virtual object. If it is determined that the line of sight intersects with multiple virtual objects, the line of sight inference module 610 selects the virtual object closest to the origin O of the line of sight. Furthermore, the line of sight inference module 610 infers the virtual object closest to the line of sight as the virtual object that the virtual character is observing.
[0094] Furthermore, the ray-triangle intersection algorithm is one example, and this disclosure is not limited to it. Other shapes can also be used instead of triangles. At least the ray-triangle intersection algorithm is extremely simple, has a light processing load, and can be processed at high speed. From the viewpoint of real-time processing, high-speed processing is preferred.
[0095] In this way, based on the virtual character's 3D pose and the virtual space composition information 150, it is possible to infer with high precision which virtual object the virtual character is observing. That is, it is possible to infer with high precision which real object a real person is observing. When the virtual space composition information 150 represents the categories of each virtual object, it is possible to infer with high precision which category and which real object a real person is observing. By inferring which real object a real person is observing, for example, it is possible to understand what the real person is interested in.
[0096] 6-2. Grasp Estimation
[0097] When a real person's hand is inside any real object, it is highly likely that the real person is grasping or attempting to grasp an item contained within that object. Based on this perspective, we can perform a grasping inference process to deduce what the real person is grasping or attempting to grasp.
[0098] Figure 8 This is a schematic diagram illustrating an example of grip inference processing. The character analysis module 600 includes a grip inference module 620 that performs grip inference processing. The grip inference module 620 infers that the hand of a real person is in any real object by inferring that the hand of a virtual character is in any virtual object.
[0099] More specifically, the grip inference module 620 infers the position of the virtual character's hands based on information about the virtual character's three-dimensional pose. That is, the grip inference module 620 infers the position of the three-dimensional pose hand as the position of the virtual character's hands. Furthermore, the grip inference module 620 determines whether the virtual character's hands are located within any virtual object defined in the virtual object space SP-V.
[0100] For example, the grip inference module 620 uses a point-cube detection algorithm to determine whether the virtual character's hand is within any virtual object. For example, each virtual object is represented by a bounding box. Figure 8 In the example shown, the bounding box of a virtual object is defined by vertices A to H. For example, the position of the center point I of the bounding box is calculated based on the positions of vertices D and F (I = (D + F) / 2). Point P is the position of the virtual character's hand. A vector V is defined from the center point I toward point P. The three axes of the bounding box are defined as the x-axis, y-axis, and z-axis. [Vx, Vy, Vz] are the x-axis, y-axis, and z-axis components of vector V. Lx, Ly, and Lz are the lengths of the bounding box along the x-axis, y-axis, and z-axis directions, respectively. The holding inference module 620 determines whether the condition "2 × Vx ≤ Lx, 2 × Vy ≤ Ly, 2 × Vz ≤ Lz" is satisfied. If this condition is satisfied, it is determined that point P is within the bounding box. That is, it is inferred that the virtual character's hand is within the virtual object represented by this bounding box.
[0101] Furthermore, the point-cube detection algorithm is one example, and this disclosure is not limited to it. The point-cube detection algorithm is extremely simple, has a light processing load, and can perform high-speed processing. From the viewpoint of real-time processing, high-speed processing is preferred.
[0102] In this way, based on the virtual character's 3D pose and virtual space composition information 150, it is possible to infer with high precision that the virtual character's hand is within any virtual object. That is, it is possible to infer with high precision that the real character's hand is within any real object. When the real character's hand is within any real object, it can at least be determined that the real character is interested in an item within that object. Furthermore, when the real character's hand is within any real object, there is a high probability that the real character is grasping or attempting to grasp an item within that object. Therefore, through grasping inference processing, it is possible to roughly infer what the real character is grasping or attempting to grasp. When the virtual space composition information 150 represents the categories of each virtual object, it is also possible to determine in more detail the item that the real character is grasping or attempting to grasp.
[0103] 6-3. Human Flow Estimation
[0104] Figure 9 This is a schematic diagram illustrating an example of people flow inference processing. The person analysis module 600 includes a people flow inference module 630 that performs people flow inference processing. The people flow inference module 630 infers the flow of real people in the real object space SP-R by inferring the flow of virtual people in the virtual object space SP-V.
[0105] In pedestrian flow inference processing, region-defined objects are used (see reference). Figure 2 [C]). Region definition objects are objects used to define regions within the virtual object space SP-V. Virtual space composition information 150 indicates the composition of each region definition object defined in the virtual object space SP-V. The human flow inference module 630 infers the flow of virtual characters in the virtual object space SP-V based on the relationship between virtual characters with three-dimensional poses and the composition of region definition objects.
[0106] More specifically, the pedestrian flow inference module 630 infers the position of the virtual character's feet based on information about the virtual character's three-dimensional pose. For example, it determines the virtual character's left and right feet based on the virtual character's three-dimensional pose and uses the midpoint between the left and right foot positions as the "foot" position. Furthermore, the pedestrian flow inference module 630 determines which area definition object the virtual character's feet are located within. For example, this determination is made based on the point-cube detection algorithm described in section 6-2 above. The pedestrian flow inference module 630 identifies the area definition object where the virtual character's feet are located and determines that the virtual character is located within the area defined by that area definition object. Moreover, the pedestrian flow inference module 630 infers the flow of the virtual character in the virtual object space SP-V by detecting changes in the area where the virtual character is located.
[0107] In this way, based on the three-dimensional pose of the virtual character and the virtual space composition information 150, the flow of virtual characters in the virtual object space SP-V can be inferred with high precision. That is, the flow of real characters in the real object space SP-R can be inferred with high precision. The flow of real characters in the real object space SP-R can be used for various purposes. For example, based on the flow of real characters in the real object space SP-R, the congestion situation within the real object space SP-R can be investigated, and the causes of congestion can be determined. As another example, based on the flow of real characters in the real object space SP-R, it is possible to analyze what real characters are interested in.
[0108] 6-4. Display of Analysis Results
[0109] The character analysis module 600 displays the analysis results on the display device 50. The analysis results can be statistical information or information on changes over time. The information can also be categorized according to each attribute of the character (e.g., gender, age).
[0110] 7. Examples of camera calibration and matching processing
[0111] The calibration module 200 performs camera calibration to determine the camera parameters of the real-world camera CAM-R. Additionally, the calibration module 200 performs "camera alignment" to correct the camera parameters by matching the real-world object space SP-R observed from the real-world camera CAM-R with the virtual object space SP-V. That is, the calibration module 200 performs "camera calibration and alignment" to determine the camera parameters of the real-world camera CAM-R by matching the real-world object space SP-R observed from the real-world camera CAM-R with the virtual object space SP-V. By performing camera calibration and alignment, the accuracy of positioning processing (see section 4), visualization processing (see section 5), and character analysis processing (see section 6) is ensured.
[0112] The following is a specific example of camera calibration and matching processing.
[0113] Figure 10 This is a schematic diagram illustrating an example of camera calibration and matching processing. In this example, camera calibration is performed based on the PnP (Perspective n-Point) method. The PnP method uses a point group consisting of n points. Here, n is two or more integers. The 3D point group coordinate information is the coordinate information of the point group in 3D space (world coordinate system). The 2D point group coordinate information is the coordinate information of the point group in the image coordinate system when the camera captured that point group. Given the 3D and 2D point group coordinate information, the camera parameters can be calculated by solving the PnP problem.
[0114] In this example, a point group (n points) is obtained based on identifier 210, which is a marker configured in space. For example, identifier 210 is a quadrilateral, and the four vertices M1 to M4 of this quadrilateral are used as the point group. Identifier 210 has a defined pattern that can be identified on the image.
[0115] More specifically, a real identifier 210-R is placed at a predetermined real location captured by a real camera CAM-R within the real object space SP-R. Conversely, a virtual identifier 210-V is placed at a predetermined virtual location captured by a virtual camera CAM-V within the virtual object space SP-V. Here, the predetermined virtual location within the virtual object space SP-V is equivalent to the predetermined real location within the real object space SP-R. Furthermore, the real identifier 210-R and the virtual identifier 210-V are identical in shape, orientation, and pattern. Therefore, the point group (vertices M1 to M4) of the virtual identifier 210-V is equivalent to the point group (vertices M1 to M4) of the real identifier 210-R. Moreover, when the virtual identifier 210-V is placed at the predetermined virtual location within the virtual object space SP-V, it is utilized... Figure 2 Custom module 100 is shown.
[0116] The calibration module 200 acquires a two-dimensional image (IMG) captured by a reality camera (CAM-R) positioned in the reality object space (SP-R) as a query image. A reality identifier 210-R placed at a predetermined reality location is captured within the query image. The calibration module 200, possessing the composition information of the identifier 210, detects the reality identifier 210-R and point groups (four vertices M1 to M4) within the query image through pattern matching. Furthermore, the calibration module 200 acquires the position of the point groups in the query image as "two-dimensional point group coordinate information 211".
[0117] On the other hand, the calibration module 200 obtains the positions of the point group (four vertices M1 to M4) of the virtual identifier 210-V within the virtual object space SP-V as "three-dimensional point group coordinate information 212". The positions of the point group of the virtual identifier 210-V within the virtual object space SP-V are obtained from the virtual space composition information 150.
[0118] The calibration module 200 determines the camera parameters by solving the PnP problem based on the obtained two-dimensional point group coordinate information 211 and three-dimensional point group coordinate information 212. Importantly, according to this method, camera alignment is achieved simultaneously with the determination of camera parameters. Because the two-dimensional point group coordinate information 211 obtained from the real object space SP-R and the three-dimensional point group coordinate information 212 obtained from the virtual object space SP-V are combined, camera alignment is achieved while determining the camera parameters.
[0119] As explained above, this example demonstrates that both camera calibration and camera matching can be performed in a single step. From the perspective of reducing processing load, this is preferable.
Claims
1. A projection system, wherein, The projection system includes: One or more processors; and One or more storage devices are configured to store virtual space composition information representing the composition of objects defined in a virtual object space that represents a real object space. The one or more processors are configured as follows: Detecting real-world figures appearing in images captured by a real-world camera positioned in the real-world object space. The three-dimensional pose of the real person is inferred from the image. A virtual character, possessing the aforementioned three-dimensional pose and representing the real-world figure, is projected into the virtual object space. An action inference process is performed, in which the action of the virtual character in the virtual object space is inferred based on the relationship between the virtual character having the three-dimensional pose and the composition of the objects in the virtual object space, thereby inferring the action of the real character in the real object space.
2. The projection system according to claim 1, wherein, The objects defined in the virtual object space include virtual objects that are equivalent to real-world objects in the real-world object space. In the action inference process, the one or more processors are configured to: infer the action of the virtual character in the virtual object space relative to the virtual object based on the relationship between the virtual character having the three-dimensional pose and the virtual object in the virtual object space, thereby inferring the action of the real character in the real object space relative to the real object.
3. The projection system according to claim 2, wherein, The action inference process includes a gaze inference process, which infers which real object the real person is observing by inferring which virtual object the virtual character is observing.
4. The projection system according to claim 3, wherein, The line-of-sight inference process includes: Based on the three-dimensional pose of the virtual character, the ray of the virtual character's gaze is inferred; and By determining whether the line of sight intersects with any virtual object defined in the virtual object space, it can be inferred which virtual object the virtual character is observing.
5. The projection system according to claim 2, wherein, The action inference process includes a grip inference process, which infers whether the hand of the real person is in any real object by inferring whether the hand of the virtual character is in any virtual object.
6. The projection system according to claim 1, wherein, The objects defined in the virtual object space include region definition objects used to define regions within the virtual object space. In the action inference process, the one or more processors are configured to perform a people flow inference process, in which the flow of the virtual character in the virtual object space is inferred based on the relationship between the virtual character with the three-dimensional pose and the composition of the area-defined object, thereby inferring the flow of the real character in the real object space.
Citation Information
Patent Citations
Method for giving virtual universal environment, computer system, and computer readable storage medium
JP2010176662A
Simulation system and picking station layout designing method
JP2022061305A
Gaze estimation system
JP2022187547A
Physical distribution simulation system
JP2024003401A