Full-body tracking system
The system leverages existing cameras to expand tracking space and eliminate the need for additional devices, enabling full-body tracking on various devices by capturing and processing image data for skeletal estimation and tracking.
Patent Information
- Application Number
- JP2024081075
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-17
- Publication Date
- 2025-11-28
AI Technical Summary
Conventional full-body tracking systems require additional devices to expand the tracking space and necessitate the user to wear or install specific equipment, limiting their flexibility and applicability.
A full-body tracking system that utilizes existing cameras in the tracking space, such as surveillance and mobile device cameras, to capture and process image data for skeletal estimation and tracking, enabling expansion of the tracking space without additional devices.
Enables tracking in any space with communicable cameras, allowing users to perform full-body tracking using devices without inherent tracking capabilities, such as smartphones, and simplifies the process by requiring only entry into the tracking space.
Smart Images

Figure 2025174598000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a full-body tracking system that acquires a user's body movements in three dimensions in order to reflect the user's body movements in a virtual avatar in XR applications including VR, AR, and MR. [Background technology]
[0002] Japanese Patent Application Laid-Open Publication No. 2023-057498 discloses a motion capture system that acquires information about the physical movement and posture of a user who is a subject. In this motion capture system, the subject's movement and posture are captured using a general camera (for example, a digital camera, smartphone, or tablet). The captured image, a virtual camera with the same field of view as the capturing camera, and a three-dimensional character that resembles the subject are then placed in three-dimensional space, and the bone angles and position coordinates are acquired across multiple frames when the subject's image as seen from the virtual camera matches the character's posture, thereby three-dimensionally capturing the movement of the subject's movement and posture.
[0003] Conventional full-body tracking systems require a user to wear a device, install the device in a tracking space, and introduce some kind of device for full-body tracking into the tracking space. Even when using a general camera such as a smartphone, as in the technology disclosed in JP 2023-057498 A, the camera must be installed so that it can capture the entire body of the user.
[0004] As described above, conventional full-body tracking systems have a limited tracking space, and in order to expand the tracking space, it is necessary to introduce additional devices for full-body tracking.
[0005] In addition to JP 2023-057498 A, JP 2024-030610 A, JP 10-302070 A, and JP 2023-086471 A can be exemplified as documents showing the technical state of the technical field related to the present disclosure. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Japanese Patent Application Publication No. 2023-057498 [Patent Document 2] Japanese Patent Application Publication No. 2024-030610 [Patent Document 3] Japanese Patent Application Publication No. 10-302070 [Patent Document 4] Japanese Patent Application Publication No. 2023-086471 Summary of the Invention [Problem to be solved by the invention]
[0007] The present disclosure has been made in view of the above-mentioned problems, and an object of the present disclosure is to provide a full-body tracking system that can easily expand the tracking space. [Means for solving the problem]
[0008] According to one embodiment of the present disclosure, a processing circuit of a full-body tracking system communicates with various cameras located in a tracking space and acquires image data captured by the cameras. The image data is acquired using cameras already installed in the tracking space for other purposes, such as surveillance cameras installed in the tracking space, cameras mounted on moving objects moving within the tracking space, and cameras mounted on mobile devices carried by users. The processing circuit performs tracking and skeletal estimation for each user in the tracking space using the acquired image data. The processing circuit also authenticates a target user attempting full-body tracking and acquires position information and skeletal information of the authenticated target user from the results of tracking and skeletal estimation. The processing circuit then uses the acquired position information and skeletal information of the target user to reflect the target user's body movements in a virtual avatar in an XR space. [Effects of the Invention]
[0009] According to the full-body tracking system configured as described above, as long as a communicable camera is present in a space, that space can be used as a tracking space. In other words, the tracking space can be expanded without requiring an additional device for full-body tracking. Furthermore, since image data acquired from a camera located in the tracking space is used, the user does not need to prepare a device for full-body tracking. The user can use an application that uses full-body tracking on a device that does not have full-body tracking functionality, such as a smartphone or a head-mounted display. Furthermore, the user can perform full-body tracking simply by entering the tracking space. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram illustrating an overview of a full-body tracking system according to an embodiment of the present disclosure. [Figure 2]FIG. 1 is a diagram illustrating a mechanism of full-body tracking according to an embodiment of the present disclosure. [Figure 3] FIG. 1 is a diagram illustrating a mechanism of full-body tracking according to an embodiment of the present disclosure. [Figure 4] FIG. 1 is a block diagram illustrating a configuration of a full-body tracking system according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0011] 1. Overview of the Full-Body Tracking System In XR applications, including VR, AR, and MR, the user's movements are reflected in a virtual avatar in the XR space. The full-body tracking system described below acquires three-dimensional data indicating the user's body movements and outputs avatar generation information for generating an avatar for each user based on the acquired three-dimensional data.
[0012] 1 is a diagram illustrating an overview of a full-body tracking system 100. The full-body tracking system 100 is a computer or a group of computers including a processor 101 as a processing circuit and a memory 102 coupled to the processor 101. A program made up of a plurality of instructions 103 is stored in the memory 102. When the processor 101 executes the instructions 103, various processes related to full-body tracking are performed by the processor 101.
[0013] A tracking space 10 is set in the full-body tracking system 100. The tracking space 10 is a space where full-body tracking of a person is performed, and is also a space where conversion to an XR space 30 is performed. In the example shown in FIG. 1 , three users 1, 2, and 3 are present in the tracking space 10. In the XR space 30 corresponding to the tracking space 10, there exist avatars 31, 32, and 33 corresponding to the users 1, 2, and 3, respectively. The full-body tracking system 10 reflects the body movements of the users 1, 2, and 3 in the tracking space 10 in the corresponding avatars 31, 32, and 33.
[0014] The XR space 30 is displayed on a head-mounted display 12 worn by user 2 and mobile devices 11 and 13 carried by users 1 and 3. The XR space 30 can also be displayed on a monitor 16 viewed by persons other than users 1, 2, and 3. The head-mounted display 12 includes XR goggles and XR glasses. The mobile devices 11 and 13 include smartphones and tablet PCs. The monitor 16 includes public viewing monitors installed in town, PC monitors installed in homes, and mobile devices owned by persons other than users 1, 2, and 3. The monitor 16 may be installed outside the tracking space 10. These devices 11, 12, 13, and 16 display the XR space 30 using an XR application based on avatar generation information transmitted from the full-body tracking system 100. The XR application can be downloaded from the full-body tracking system 100.
[0015] The head-mounted display 12 displays a subjective image of the XR space 30 as seen from an avatar 32 corresponding to the user 2. On the mobile devices 11 and 13, the users 1 and 3 can arbitrarily change the viewpoint from which the image of the XR space 30 is displayed. For example, the mobile devices 11 and 13 can display a subjective image of the XR space 30 as seen from their own avatars 31 and 33, or can display an objective image including their own avatars 31 and 33. The monitor 16 displays an objective image of the XR space 30 as seen from a fixed or moving observation point.
[0016] In order to generate avatars 31, 32, and 33 to be displayed in the XR space 30, it is necessary to grasp the positions and postures of users 1, 2, and 3 in the tracking space 10. To grasp the positions and postures of the users, multiple cameras that have already been introduced into the tracking space 10 for other purposes and are capable of real-time communication with the full-body tracking system 100 are used. The tracking space 10 can also be defined as a space in which the positions and postures of users 1, 2, and 3 can be grasped using multiple cameras.
[0017] 2. How full body tracking works Next, the mechanism of full-body tracking by the full-body tracking system 100, specifically the mechanism for reflecting the position and posture of a user captured by multiple cameras on an avatar, will be described with reference to Figures 2 and 3. Note that in this description, user 1 of users 1, 2, and 3 in the tracking space 10 is the target of full-body tracking.
[0018] In the example shown in FIG. 2, four cameras are installed in the tracking space 10. The first camera is a surveillance camera 14 installed on a utility pole 4. The surveillance camera 14 is installed for security purposes. When the surveillance camera 14 is used for full-body tracking of the user 1, the surveillance camera 14 communicates with the full-body tracking system 100 in real time and transmits image data to the full-body tracking system 100. Note that the surveillance camera connected to the full-body tracking system 100 may be a surveillance camera installed in a building or a private residence.
[0019] The second camera is an on-board camera 15 mounted on the automobile 5. The on-board camera 15 is provided for the purpose of recording the surroundings of the automobile 5. When the on-board camera 15 is used for full-body tracking of the user 1, the on-board camera 15 communicates with the full-body tracking system 100 in real time and transmits image data to the full-body tracking system 100. In addition to the on-board camera 15, a camera mounted on a mobile object moving in the tracking space 10, such as a patrol robot or a drone, may also be connected to the full-body tracking system 100.
[0020] The third camera is the camera (mobile terminal camera) 11a of the mobile terminal 11 carried by the user 1. When the mobile terminal 11 is used for full-body tracking of the user 1, an in-camera mounted inside the mobile terminal 11 functions as the mobile terminal camera 11a. The mobile terminal 11 communicates with the full-body tracking system 100 in real time and transmits image data captured by the mobile terminal camera 11a to the full-body tracking system 100. Note that the connection between the mobile terminal 11 and the full-body tracking system 100 may be established only while the XR application is running, or may be established automatically when the mobile terminal 11 enters the tracking space 10.
[0021] The fourth camera is a camera (mobile terminal camera) 13a of a mobile terminal 13 carried by user 3. When the mobile terminal 13 is used for full-body tracking of user 1, an outer camera mounted on the back side of the mobile terminal 11 functions as the mobile terminal camera 13a. The mobile terminal 13 communicates with the full-body tracking system 100 in real time and transmits image data captured by the mobile terminal camera 13a to the full-body tracking system 100. Note that the connection between the mobile terminal 13 and the full-body tracking system 100 may be established only while the XR application is running, or may be established automatically when the mobile terminal 13 enters the tracking space 10.
[0022] The full-body tracking system 100 acquires image data from the multiple cameras 11a, 13a, 14, and 15, and determines the position A1 of user 1 in the tracking space 10 from the image data. If the camera that captured the image of user 1 is a surveillance camera 14, its installation position is fixed. Therefore, if position and orientation calibration has been performed, the position A1 of user 1 in the tracking space 10 can be calculated from the position of user 1 in the image. The in-vehicle camera 15 and the mobile terminal cameras 11a and 13a move along with the automobile 5 and users 1 and 3, but their positions and orientations can be estimated using SLAM. The positions and orientations of the in-vehicle camera 15 and the mobile terminal cameras 11a and 13a can also be estimated using GPS installed in the automobile 5 and the mobile terminals 11 and 13. If the camera that captured the image of user 1 is the in-vehicle camera 15 or another user's mobile terminal camera 13a, the position A1 of user 1 in the tracking space 10 can be calculated from the position of user 1 in the image. If the camera that captured the image of user 1 is user 1's mobile terminal camera 11a, the position of mobile terminal camera 11 can be considered to be user 1's position A1 in tracking space 10. If user 1 is captured by multiple cameras, user 1's position A1 in tracking space 10 may be calculated based on the most reliable camera, or user 1's position A1 may be calculated statistically, such as by averaging the positions of user 1 calculated for each camera.
[0023] The full-body tracking system 100 has a three-dimensional coordinate system 20 corresponding to the tracking space 10, and projects the position A1 of the user 1 onto the three-dimensional coordinate system 20. Then, the movement of the user 1 at the position A1 is grasped from the image data. Specifically, skeletal estimation is performed for each image in which the user 1 appears. Skeletal estimation means estimating the joint postures of a human body having a skeleton composed of multiple joints. For skeletal estimation, for example, a two-dimensional joint posture estimation model is used. The two-dimensional joint posture estimation model is a machine learning model that estimates the position of each joint that makes up the human body. According to the two-dimensional joint posture estimation model, one two-dimensional joint posture of the user 1 can be obtained from one image in which the user 1 appears. When multiple images in which the user 1 appears are obtained using multiple cameras, multiple two-dimensional joint postures of the user 1 viewed from different directions can be obtained.
[0024] As the next step in skeletal estimation, the full-body tracking system 100 estimates the three-dimensional joint posture of the user 1 from the two-dimensional joint posture. A three-dimensional joint posture estimation model is used to estimate the three-dimensional joint posture. The three-dimensional joint posture estimation model is a machine learning model that converts the two-dimensional joint posture into a three-dimensional joint posture by adding depth information of the human body obtained by machine learning. When multiple two-dimensional joint postures viewed from different directions are obtained, they can be input into the three-dimensional joint posture estimation model to obtain a more accurate three-dimensional joint posture. The full-body tracking system 100 projects the three-dimensional joint posture 21 of the user 1 obtained by skeletal estimation onto the position A1 of the user 1 in the three-dimensional coordinate system 20.
[0025] The full-body tracking system 100 performs tracking of the user 1 using image data obtained from multiple cameras. Tracking here refers to tracking of the position coordinates of the user 1 as he moves through the tracking space 10, and is distinct from full-body tracking. Full-body tracking is a process that combines the aforementioned skeletal estimation based on image data with tracking. When the user 1 moves from position A1 to position A2, the full-body tracking system 100 estimates the position A2 of the user 1 after the movement in a three-dimensional coordinate system 20 by tracking. At the same time, the full-body tracking system 100 projects the three-dimensional joint posture 21 of the user 1 after the movement, obtained by skeletal estimation, onto the position A2 of the user 1 in the three-dimensional coordinate system 20.
[0026] The full-body tracking system 100 also performs the above-mentioned tracking and skeletal estimation on other users 2 and 3 in the tracking space 10. The full-body tracking system 100 also authenticates target users among users 1, 2, and 3 whose position information and skeletal information are to be reflected in avatars in the XR space 30, and acquires the position information and skeletal information of the target users from the results of tracking and skeletal estimation for each user. Note that the skeletal information of the target users means information related to the three-dimensional joint posture.
[0027] Here, a description will be given of an authentication process for confirming whether or not a person captured by a camera in the tracking space 10 is a target user. In the following description, it is assumed that user 1 in FIG. 2 is a target user.
[0028] A first example of authentication processing is authentication based on the result of matching the location information of the mobile terminal 11 obtained by a GPS installed in the mobile terminal 11 with the location information of each person obtained from image data. The location information of the mobile terminal 11 is the location information of the user 1 who carries the mobile terminal 11. The full-body tracking system 100 acquires the location information of the mobile terminal 11 by communicating with the mobile terminal 11, and also acquires the location information of each person captured on camera from the image data. Of the people captured on camera, the person whose location information matches the location information of the mobile terminal 11 can be confirmed to be user 1.
[0029] A second example of authentication processing is authentication based on the degree of match between the movement trajectory of mobile terminal 11 obtained by a GPS installed in mobile terminal 11 and the movement trajectory of each person acquired by tracking. The movement trajectory of mobile terminal 11 is the movement trajectory of user 1 who carries mobile terminal 11. Full-body tracking system 100 acquires the movement trajectory of mobile terminal 11 by communicating with mobile terminal 11, and also acquires the movement trajectory of each person captured on camera from the tracking results. Of the people captured on camera, the person whose movement trajectory most closely matches the movement trajectory of mobile terminal 11 can be confirmed to be user 1.
[0030] A third example of authentication processing is face authentication using image data. For example, if the tracking space 10 is a closed space with an entrance, the face of user 1 can be photographed at the entrance, and the face image acquired at the entrance can be compared with the face image of the person included in the camera image. This makes it possible to link the account information of user 1 to the location information and skeletal information obtained from the camera image data. If user 1 moves between cameras, the location information and skeletal information can be transferred between cameras by comparing the face image again.
[0031] The full-body tracking system 100 transmits avatar generation information, including the position information and skeletal information of the target user, to the various devices 11, 12, 13, and 16. The devices 11, 12, 13, and 16 use an XR application to reflect the position information and skeletal information in the avatar in the XR space 30.
[0032] In the example shown in FIG. 3, the position information and skeletal information of user 1 are reflected in avatar 31. Therefore, when user 1 moves from position A1 to position A2 in tracking space 10, avatar 31 also moves from position A1 to position A2 in XR space 30. Furthermore, when user 1 changes his / her posture, the posture of avatar 31 also changes accordingly. Similarly, the position information and skeletal information of user 2 are reflected in avatar 32, and the position information and skeletal information of user 3 are reflected in avatar 33. Furthermore, in XR space 30, a stationary object such as utility pole 4 may be replaced with a plant 34 or an inanimate stationary object having another form, and a moving object such as car 5 may be replaced with an animal 35 or an inanimate moving object having another form.
[0033] 3. Full-body tracking system configuration Next, the configuration of a full-body tracking system 100 that enables the above-described process for reflecting the position and posture of a user captured by multiple cameras on an avatar will be described with reference to FIG.
[0034] The full-body tracking system 100 includes a video management unit 110, a processing unit 120, and a system interface unit 130. These units 110, 120, and 130 may be realized by separate hardware, or may be realized by software running on common hardware.
[0035] A specific example of the image management unit 110 is image management software. The image management unit 110 communicates in real time with each of the cameras 11 a, 13 a, 14, and 15 located in the tracking space 10, and acquires image data from each of the cameras 11 a, 13 a, 14, and 15. The image management unit 110 inputs the acquired image data into the processing unit 120.
[0036] A specific example of the processing unit 120 is a machine learning pipeline. The processing unit 120 executes a skeleton estimation process 121 and a tracking process 122 based on the video data input from the video management unit 110. The skeleton information obtained by the skeleton estimation process 121 and the position information obtained by the tracking process 122 are temporarily stored in a memory 123.
[0037] A specific example of the system interface unit 130 is an API endpoint. The system interface unit 130 acquires, for example, position information or movement trajectories from the mobile terminals 11 and 13 and the head-mounted display 12 located in the tracking space 10, and performs authentication processing 131 for each user based on the acquired position information or movement trajectories. The system interface unit 130 also performs output processing 132 to transmit position information and skeletal information of an avatar to the mobile terminals 11 and 13 and the head-mounted display 12 located in the tracking space 10. When the XR space 30 is displayed on the monitor 16, the position information and skeletal information of the avatar are also transmitted to the monitor 16.
[0038] 4.Effects As described above, with the full-body tracking system 100 according to this embodiment, any space can be used as a tracking space as long as a camera capable of communicating is present in the space. Therefore, the tracking space can be expanded without the need for additional devices. Furthermore, with the full-body tracking system 100, image data acquired from a camera located in the tracking space is used, so the user does not need to prepare their own device for full-body tracking. Users can use applications that use full-body tracking with smartphones, head-mounted displays, and other devices that do not have full-body tracking functionality. Furthermore, with the full-body tracking system 100, user authentication can be performed without the need for additional devices for full-body tracking, and the user can perform full-body tracking simply by entering the tracking space. [Explanation of symbols]
[0039] 1, 2, 3...User, 4...Utility pole, 5...Car, 10...Tracking space, 11, 13...Mobile device, 11a, 13a...Mobile device camera, 12...Head-mounted display, 14...Surveillance camera, 15...In-vehicle camera, 16...Monitor, 20...Three-dimensional coordinate system, 21...Three-dimensional joint posture, 30...XR space, 31, 32, 33...Avatar, 100...Full-body tracking system
Claims
1. processing circuitry; The processing circuitry Communicate with multiple cameras located in the tracking space, Acquire image data captured by the plurality of cameras; Using the image data, perform tracking and skeletal estimation for one or more users in the tracking space; authenticating a target user among the one or more users; acquiring position information and skeletal information of the target user from the results of the tracking and the skeletal estimation; Using the position information and the skeletal information, the body movements of the target user are reflected in a virtual avatar in the XR space. It is configured as follows: A full-body tracking system.
2. 10. The full body tracking system of claim 1, The plurality of cameras includes a surveillance camera installed in the tracking space. A full-body tracking system.
3. 10. The full body tracking system of claim 1, The plurality of cameras includes a camera mounted on a moving object moving in the tracking space. A full-body tracking system.
4. 10. The full body tracking system of claim 1, The plurality of cameras includes cameras mounted on mobile terminals owned by the one or more users. A full-body tracking system.
5. 10. The full body tracking system of claim 1, The processing circuitry The target user is authenticated based on a result of matching between the target user's position acquired through communication with a portable terminal carried by the target user and each of the positions of the one or more users acquired based on the image data, or based on the degree of coincidence between the target user's movement trajectory acquired through communication with a portable terminal carried by the target user and each of the movement trajectories of the one or more users acquired by the tracking. It is configured as follows: A full-body tracking system.
Citation Information
Patent Citations
Object recognition system
JP2016212675A
Image processing device
JP2019009752A
Information terminal device, program, presentation system and server
JP2021033377A
Information processing device, information processing method, and program
JP2021168040A
Techniques for rendering three-dimensional animated graphics from video
US20200193671A1