Dynamic target synchronous tracking and interaction system and method

By combining optical positioning and markerless vision modules, and utilizing data processing through filter switching and fusion processing modules, high-precision and highly interference-resistant dynamic target tracking was achieved. This solved the problems of insufficient accuracy and interference resistance in existing technologies, and improved the accuracy and stability of motion capture.

CN121033947AActive Publication Date: 2025-11-28AI TUER
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511563663.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2025-11-28
Estimated Expiration
2045-10-30

AI Technical Summary

Technical Problem

Existing dynamic target tracking technologies struggle to simultaneously achieve both high precision and strong anti-interference capabilities. Hybrid systems suffer from shallow data fusion layers and fail to effectively address the issues of spatiotemporal synchronization and error compensation.

Method used

By combining an optical positioning module and a markerless vision module, switching between Blob mode and video mode is achieved through filter switching. Data fusion and spatiotemporal synchronization are performed using a fusion processing module, and dynamic weight allocation is implemented to achieve real-time tracking of dynamic targets.

Benefits of technology

It improves the accuracy and anti-interference ability of the motion capture process, ensuring high-precision dynamic target tracking in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033947A_ABST
    Figure CN121033947A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic target synchronous tracking and interaction system and method, and relates to the technical field of motion capture, and the system comprises an optical positioning module which is used for switching an optical filter to a narrowband infrared optical filter, marking the narrowband infrared optical filter in a Blob mode, capturing the infrared image data of an actor in real time, and obtaining the information of a marking point; the unmarked visual module is used for switching an optical filter to full-spectrum transmission, marking the optical filter as a video mode, and capturing RGB picture data and space depth information of the actor in real time; the mode switching module is used for switching a Blob mode and a video mode and uniformly using the Blob mode for calibration; and the fusion processing module is used for obtaining the real-time three-dimensional model of the actor, obtaining the dynamic space position of the actor, and carrying out fusion synchronous mapping on the real-time three-dimensional model after weight distribution and the dynamic space position. The method and the device have the effect of improving the accuracy and the anti-interference capability in the motion capture process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of motion capture, and in particular to a dynamic target synchronous tracking and interaction system and method. Background Technology

[0002] With the rapid development of immersive technologies such as VR, AR, and MR, dynamic target tracking technology is becoming increasingly important in these fields.

[0003] However, existing dynamic target tracking technologies are mainly divided into two types: optically labeled and labelless, each with its own advantages and disadvantages. Optically labeled tracking relies on markers attached to the target object for tracking. While it offers high tracking accuracy, its reliability significantly decreases when markers are occluded, detached, or in complex lighting conditions, thus limiting the system's application scope. Labelless tracking tracks targets by identifying natural feature points on the object (such as human joints, object contours, etc.). However, this method is limited by high computational complexity, insufficient dynamic target tracking accuracy, and difficulty in handling similar feature points in complex scenes, resulting in limited effectiveness in practical applications.

[0004] A single system cannot simultaneously meet the requirements of high precision and strong anti-interference capabilities. Therefore, it is necessary to develop a novel hybrid tracking technology that combines the advantages of optical marking and markerless systems to overcome the limitations of existing technologies.

[0005] In the relevant patents, most existing hybrid systems adopt independent working modes and have shallow data fusion levels (such as simple weighted averages), which do not effectively solve the problems of spatiotemporal synchronization and error compensation. Summary of the Invention

[0006] The purpose of this invention is to provide a dynamic target synchronous tracking and interaction system and method to solve the problems mentioned in the background art.

[0007] In a first aspect, this application provides a dynamic target tracking and interaction system, the system comprising: Optical positioning module: used to switch the filter to a narrowband infrared filter and mark it as Blob mode, capture the infrared image data of the actor in the virtual shooting environment in real time, and obtain the marker point information based on the infrared image data; Unmarked visual module: Used to switch the filter to full-spectrum transmission and mark it as video mode, capturing RGB image data and spatial depth information of the actor in the virtual shooting environment in real time; Mode switching module: used to switch between the Blob mode and the video mode, and to use the Blob mode for unified calibration; Fusion processing module: used to obtain the actor's real-time 3D model based on the marker point information, obtain the actor's dynamic spatial position based on the RGB image data and the spatial depth information, perform spatiotemporal synchronization and dynamic weight allocation on the real-time 3D model and the dynamic spatial position, and perform fusion synchronization mapping on the weighted real-time 3D model and the dynamic spatial position.

[0008] Preferably, the optical positioning module includes: Infrared cameras and optical markers; Optical markers: used to be placed on the costumes or accessories worn by actors to reflect infrared light emitted from the outside back; Infrared camera: used to emit infrared light and receive reflected infrared light, determine the marker point information of the optical marker, and generate infrared image data.

[0009] Preferably, the markerless visual module includes: RGB camera and deep learning processor; RGB camera: Used to capture RGB image data of the actors; Deep learning processor: used to calculate the distance value of the pattern at different positions in the RGB image data, and store the distance value in each pixel of the RGB image data.

[0010] Preferably, the mode switching module includes: A calibration unit and a switching unit, wherein the switching unit includes a dynamic filter control component and a mode switching component; Calibration unit: used to uniformly calibrate the infrared camera and the RGB camera in the Blob mode and establish a spatial mapping relationship; Filter control component: used to manually adjust the filter parameters of a preset number of RGB cameras after calibration, and to switch a preset number of RGB cameras to video mode; Mode switching component: used to detect the loss rate of marker points in real time under the Blob mode, and determine whether the loss rate exceeds a preset loss rate threshold; If it is determined that the loss rate exceeds the loss rate threshold, the RGB camera that has switched to the video mode is invoked to perform supplementary shooting.

[0011] Preferably, the mode switching module further includes: A low-latency switching unit and a camera resource allocation unit, wherein the camera resource allocation unit includes a task allocation component and a computing load balancing component; Low-latency switching unit: used to extract the switching duration and switching time point of the filter, extract the exposure period of the RGB camera, and synchronize the switching time point and the switching duration with the exposure period in time sequence; Task allocation component: used to extract the number of cameras of the RGB cameras, fix a portion of the RGB cameras in the Blob mode according to a preset minimum number, and dynamically switch the remaining RGB cameras in the Blob mode; The computational load balancing component is used to process Blob mode data generated in the Blob mode by a preset FPGA and to process video mode data generated in the video mode by a preset GPU.

[0012] Preferably, the fusion processing module includes: A hybrid data fusion unit, comprising a data fusion acquisition component and a dynamic compensation component; Data fusion acquisition component: When in the Blob mode, the infrared camera acquires the main positioning data of the marker point, and the RGB camera assists in completing the marker point through a filter; When in the video mode, the RGB camera performs skeleton tracking on the actor, and the infrared camera identifies key movement nodes of the actor and provides precise pose based on the key nodes; Dynamic compensation component: used to perform tomographic monitoring on the data collected by the RGB camera when the filter parameters are adjusted, to obtain the data tomography, and to perform motion interpolation on the data tomography.

[0013] Preferably, the fusion processing module further includes: A hybrid sensor spatiotemporal synchronization unit, comprising a hardware spatiotemporal synchronization component and a spatial calibration component; Hardware time-space synchronization component: used to trigger hardware synchronization signal, synchronize the working time of the infrared camera and the RGB camera according to the hardware synchronization signal, and align the time axis of the Blob mode data and the video mode data. Spatial calibration component: used to construct a unified coordinate system and establish an error compensation model based on a preset shared calibration target; Based on the unified coordinate system, the data deviation between the Blob mode data and the video mode data is extracted, and the data deviation is corrected according to the error compensation model.

[0014] Preferably, the fusion processing module further includes: A dynamic weight fusion unit, comprising an occlusion detection component and an adaptive weight allocation component; Occlusion detection component: used to extract unmarked feature points of actors in the RGB image data, and monitor the visibility value of the optical markers and the confidence level of the unmarked feature points in real time; Based on the visibility value and the confidence level, the weight allocation parameters of the Blob mode data and the video mode data are dynamically adjusted in real time. Adaptive weight allocation component: used to determine the occlusion status of the optical marker and the unmarked feature points based on the weight allocation parameters; If it is determined that there is no obstruction, the weight value of the Blob mode data is increased, and the video mode data is used as an auxiliary parameter to correct the slight jitter of the Blob mode data. If it is determined to be partial occlusion, the occlusion area is recorded. The video mode data takes over the occlusion area and performs cross-modal feature association based on the unmarked feature points to complete the data of the occlusion area. If the occlusion is determined to be global, a preset long short-term memory model is used to predict the actor's movement trajectory and obtain the movement trajectory data.

[0015] Secondly, this application provides a dynamic target synchronization tracking and interaction method, the method comprising: Real-time capture of infrared image data, RGB image data, and spatial depth information of actors in the virtual shooting environment; Based on the infrared image data and the RGB image data, the marker point information on the actor's body is obtained, and the loss rate of the marker point information is extracted; Determine whether the loss rate exceeds a preset loss rate threshold. If the loss rate exceeds the loss rate threshold, extract the actor's position data, posture data, and motion information based on the RGB image data and the spatial depth information. Based on the marker point information, a real-time 3D model of the actor in the virtual environment is obtained; based on the position data, the posture data, and the action information, the dynamic spatial position of the actor in the virtual performance environment is obtained. The real-time 3D model is fused and synchronously mapped with the dynamic spatial location to generate real-time dynamic images of the actor in the virtual environment.

[0016] In summary, this application includes at least one of the following beneficial technical effects: By separately capturing infrared and RGB image data and spatial depth information of the actor in the virtual shooting environment using infrared and RGB cameras respectively, marker point information is obtained on the actor based on the infrared image data, and position, posture, and motion information is obtained based on the RGB image data and spatial depth information. The optical positioning module acquires the infrared image data and marks it as Blob mode, while the unmarked vision module acquires the RGB image data and spatial depth information and marks it as video mode. Switching between Blob and video modes is achieved through filters on the RGB camera. The acquired data is fused and processed. A real-time 3D model of the actor is obtained based on the marker point information, and the actor's dynamic spatial position is obtained based on the RGB image data and spatial depth information. Real-time weights are assigned to the real-time 3D model and dynamic spatial position, and any missing data in the corresponding mode is addressed based on the assigned weights, resulting in the final real-time dynamic image of the actor in the virtual environment. This improves the accuracy and anti-interference capability of the motion capture process. Attached Figure Description

[0017] Fig. 1 This is a block diagram of a dynamic target synchronous tracking and interaction system provided in this application.

[0018] Fig. 2 This is a flowchart of the steps of a dynamic target synchronization tracking and interaction method provided in this application.

[0019] Explanation of reference numerals in the attached diagram: 1. Optical positioning module; 2. Markerless vision module; 3. Mode switching module; 4. Fusion processing module. Detailed Implementation

[0020] The following combination Figs. 1-2 This application will be described in further detail, but the embodiments of the present invention are not limited thereto.

[0021] This application discloses a dynamic target synchronization tracking and interaction system and method.

[0022] In this application embodiment, a dynamic target synchronous tracking and interaction system is provided, the system comprising: Optical positioning module 1: Used to switch the filter to a narrowband infrared filter and mark it as Blob mode, capture the infrared image data of the actor in the virtual shooting environment in real time, and obtain the marker point information based on the infrared image data; Unmarked visual module 2: Used to switch the filter to full-spectrum transmission and mark it as video mode, capturing the actor's RGB image data and spatial depth information in the virtual shooting environment in real time; Mode switching module 3: Used to switch between Blob mode and video mode, and uses Blob mode for unified calibration; Fusion processing module 4: It is used to obtain the real-time 3D model of the actor based on the marker point information, obtain the dynamic spatial position of the actor based on the RGB image data and spatial depth information, perform spatiotemporal synchronization and dynamic weight allocation on the real-time 3D model and dynamic spatial position, and perform fusion synchronization mapping on the weighted real-time 3D model and dynamic spatial position.

[0023] The optical positioning module includes: Infrared cameras and optical markers; Optical markers: used to be placed on the costumes or accessories worn by actors to reflect infrared light emitted from the outside back; Infrared camera: Used to emit infrared light and receive the reflected infrared light to determine the marker point information of optical markers and generate infrared image data.

[0024] In application, taking the actor's movements in a virtual shooting scene as an example, the actor wears a special costume with 42 reflective markers (6mm in diameter, 92% reflectivity), ergonomically distributed at key joints and torso positions. An infrared camera array (consisting of eight Vicon Vero v2.2 cameras) emits 850nm wavelength infrared light at a frequency of 120Hz, with each camera equipped with a lens having a 30° field of view. When the actor raises their arm, the infrared light reflected from marker VH-12 (right elbow position) is captured by camera number 3, generating infrared image data containing coordinates X=1253.2px, Y=867.5px. The system calculates the precise position of this marker in three-dimensional space (X=1.2m, Y=0.8m, Z=2.3m) using triangulation, with an error range of ±0.5mm. Simultaneously, it detects that the reflection intensity of marker VH-09 (right shoulder position) drops to 65 lux (below the threshold of 80 lux), triggering a low-reflection alarm. All marker point data is transmitted to the processing host via gigabit Ethernet, generating a marker point information frame containing 42 three-dimensional coordinates, with a timestamp of 2023-07-15T14:22:36.125.

[0025] The unmarked visual module includes: RGB camera and deep learning processor; RGB camera: Used to capture RGB image data of the actors; Deep learning processor: Used to calculate the distance values ​​of patterns at different locations in RGB image data and store the distance values ​​in each pixel of the RGB image data.

[0026] In the application, three Intel RealSense D455 cameras were used to capture RGB images of the actor at 30fps, with a resolution of 1920×1080. When the actor turned around, the deep learning processor (NVIDIA Jetson AGX Orin) detected a feature point (texture hash value #A8D3F2) of the cuff of his right arm, and calculated the distance of this feature point to the camera to be 2.1m using a stereo vision algorithm. In depth information processing, the system identified five joints of the actor's left hand (from fingertip to wrist), and measured the depth difference between the middle finger joint (X=1.5m, Y=1.8m, Z=1.2m) and the background green screen to be 0.3m. At the same time, 68 feature points of the actor's face were extracted from the RGB image, and a 3.2mm rise in the right eyebrow was detected (baseline value is 2.8mm), generating facial expression change data. All spatial data were stored in JSON format, including the timestamp 2023-07-15T14:22:36.130, with the time deviation from the infrared data controlled within ±2ms.

[0027] The mode switching module includes: A calibration unit and a switching unit, wherein the switching unit includes a dynamic filter control component and a mode switching component; Calibration unit: used to uniformly calibrate the infrared camera and the RGB camera in the Blob mode and establish a spatial mapping relationship; Filter control component: used to manually adjust the filter parameters of a preset number of RGB cameras after calibration, and to switch a preset number of RGB cameras to video mode; Mode switching component: used to detect the loss rate of marker points in real time under the Blob mode, and determine whether the loss rate exceeds a preset loss rate threshold; If it is determined that the loss rate exceeds the loss rate threshold, the RGB camera that has switched to the video mode is invoked to perform supplementary shooting.

[0028] During operation, in the calibration phase, a preset calibration board (containing 12×9 circular markers with a diameter of 10mm) was used to simultaneously acquire optical calibration reference points (error 0.3px) and RGB calibration reference points (error 0.5px). The established spatial mapping relationship showed an X-axis offset of +1.2px and a Y-axis offset of -0.8px, which the system automatically wrote into the calibration parameter table. The manual switching component recorded the operator's forced mode switching command at 14:22:37, locking camera 1 in Blob mode for keypoint tracking. When the actor entered a high-light area (ambient light intensity reached 1500 lux), the system detected that the marker loss rate suddenly increased from 5% to 32% (threshold set at 20%). The dynamic filter control component immediately adjusted the electronically controlled filter (Lucent Optics DF-850) of camera 3, reducing the transmittance from 90% to 45%, and simultaneously switched the working mode from Blob mode to video mode.

[0029] The mode switching module also includes: The low-latency switching unit and the camera resource allocation unit include a task allocation component and a computing load balancing component. Low-latency switching unit: used to extract the switching duration and switching time point of the filter, extract the exposure cycle of the RGB camera, and synchronize the switching time point and switching duration with the exposure cycle in sequence; Task allocation component: used to extract the number of RGB cameras, fix some RGB cameras in Blob mode according to the preset minimum number, and dynamically switch the remaining RGB cameras in mode. The computational load balancing component is used to process Blob data generated in Blob mode by a preset FPGA and video data generated in video mode by a preset GPU.

[0030] In operation, the system is configured with six cameras. Cameras C, D, E, and F are pre-set to Blob mode (processing latency 8ms), while cameras A, B, and C dynamically switch between them. When the electronically controlled filter switching takes 12ms, the low-latency unit synchronously adjusts the exposure period of camera B to 14ms to ensure the switching process covers the complete exposure interval. The computational load balancing component distributes the Blob data (1.2GB / s) from the four cameras to the Xilinx Alveo U50 FPGA for processing, while the video data (2.4GB / s) is processed by the NVIDIA RTX 6000 Ada GPU. At 14:22:38, a GPU load of 85% is detected, and the video streams from two cameras are automatically downclocked to 15fps to maintain real-time performance. At this point, the overall processing latency remains below 22ms.

[0031] The fusion processing module includes: The hybrid data fusion unit includes a data fusion acquisition component and a dynamic compensation component. Data fusion acquisition component: When in Blob mode, the infrared camera acquires the main positioning data of the marker point, and the RGB camera assists in the completion of the marker point through a filter; When in video mode, the RGB camera performs skeleton tracking on the actor, and the infrared camera identifies key movement nodes of the actor and provides precise pose based on the key nodes. Dynamic compensation component: used to perform tomographic monitoring on the data collected by the RGB camera when the filter parameters are adjusted, obtain the data tomography, and perform motion interpolation on the data tomography.

[0032] In application, in Blob mode, the infrared camera tracked eight marker points on the actor's back (spatial error ±1.1mm), while the RGB camera, using a filter, assisted in identifying the marker point VH-22 (left scapula) obscured by clothing, completing its three-dimensional coordinates (X=0.9m, Y=1.2m, Z=1.8m). After switching to video mode, the RGB camera's skeleton tracking showed the right arm raised at an angle of 53°, and the infrared camera simultaneously provided an elbow joint marker point (error ±0.8°) for calibration. When the filter switched at 14:22:39, the dynamic compensation component detected a two-frame data gap and generated transition data through motion interpolation (right wrist trajectory smoothness reached 94%), ensuring motion continuity.

[0033] The fusion processing module also includes: The hybrid sensor spatiotemporal synchronization unit includes a hardware spatiotemporal synchronization component and a spatial calibration component. Hardware time-space synchronization component: used to trigger hardware synchronization signal, synchronize the working time of infrared camera and RGB camera according to hardware synchronization signal, and align the time axis of Blob mode data and video mode data. Spatial calibration component: used to construct a unified coordinate system and establish an error compensation model based on a preset shared calibration target; Based on a unified coordinate system, the data deviation between Blob mode data and video mode data is extracted, and the data deviation is corrected according to the error compensation model.

[0034] In application, the hardware synchronization signal is triggered at a frequency of 500Hz, keeping the time deviation between the infrared camera (exposure time 2ms) and the RGB camera (exposure time 3ms) within ±0.1ms. After establishing a unified coordinate system using a shared calibration target (size 0.5m×0.5m), an average offset of +2.3mm was detected in the Blob mode data along the Z-axis. The error compensation model applied second-order polynomial correction (R²=0.98). During data fusion at 14:22:40, the system corrected the 1.5px positional deviation of the right-hand marker point caused by lens distortion, ultimately improving the output coordinate accuracy to ±0.6mm.

[0035] The fusion processing module also includes: The dynamic weight fusion unit includes an occlusion detection component and an adaptive weight allocation component. Occlusion detection component: used to extract unmarked feature points of actors in RGB image data and monitor the visibility value of optical markers and the confidence level of unmarked feature points in real time; Based on visibility and confidence level, dynamically adjust the weight allocation parameters of Blob mode data and video mode data in real time; Adaptive weight allocation component: used to determine the occlusion of optically labeled and unlabeled feature points based on weight allocation parameters; If it is determined that there is no obstruction, the weight value of the Blob mode data is increased, and the video mode data is used as an auxiliary parameter to correct the slight jitter of the Blob mode data. If it is determined to be partial occlusion, the occlusion area is recorded, the video mode data takes over the occlusion area, and cross-modal feature association is performed based on unmarked feature points to complete the data of the occlusion area; If the occlusion is determined to be global, a preset long short-term memory model is used to predict the actor's movement trajectory and obtain the movement trajectory data.

[0036] In application, when the actor's right hand was obscured by a prop, the system detected that the visibility of markers VH-15 to VH-18 decreased from 100% to 12%, while the confidence of unmarked feature points remained at 78%. The weighting parameters were dynamically adjusted to 30% for Blob mode and 70% for video mode. For the unobscured left hand area (95% visibility), the weight of Blob mode remained at 85%, and video mode was only used to eliminate hand tremors at the 0.3mm level. When global occlusion occurred (14:22:41), the Long Short-Term Memory model predicted an 82% probability that the actor would move to the right next, generating motion trajectory data within 3 seconds (maximum error 3.2cm), until the markers reappeared.

[0037] This invention provides a dynamic target synchronization tracking and interaction method, using any one of the dynamic target synchronization tracking and interaction systems described above. The method includes the following: S100: Captures infrared image data, RGB image data and spatial depth information of actors in a virtual shooting environment in real time; S200: Based on infrared image data and RGB image data, obtain the marker point information on the actor's body and extract the loss rate of the marker point information; S300: Determine whether the loss rate exceeds the preset loss rate threshold. If the loss rate exceeds the threshold, extract the actor's position data, posture data, and motion information based on the RGB image data and spatial depth information. S400: Based on the marker point information, obtain the real-time 3D model of the actor in the virtual environment, and based on the position data, posture data and motion information, obtain the dynamic spatial position of the actor in the virtual performance environment. S500: It fuses and synchronously maps real-time 3D models with dynamic spatial positions to generate real-time dynamic images of actors in a virtual environment.

[0038] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A dynamic target synchronous tracking and interaction system, characterized in that, include: Optical positioning module: used to switch the filter to a narrowband infrared filter and mark it as Blob mode, capture the infrared image data of the actor in the virtual shooting environment in real time, and obtain the marker point information based on the infrared image data; Unmarked visual module: Used to switch the filter to full-spectrum transmission and mark it as video mode, capturing RGB image data and spatial depth information of the actor in the virtual shooting environment in real time; Mode switching module: used to switch between the Blob mode and the video mode, and to perform unified calibration in the Blob mode; Fusion processing module: used to obtain the actor's real-time 3D model based on the marker point information, obtain the actor's dynamic spatial position based on the RGB image data and the spatial depth information, perform spatiotemporal synchronization and dynamic weight allocation on the real-time 3D model and the dynamic spatial position, and perform fusion synchronization mapping on the weighted real-time 3D model and the dynamic spatial position.

2. The dynamic target synchronous tracking and interaction system according to claim 1, characterized in that, The optical positioning module includes: Infrared cameras and optical markers; Optical markers: used to be placed on the costumes or accessories worn by actors to reflect infrared light emitted from the outside back; Infrared camera: used to emit infrared light and receive reflected infrared light, determine the marker point information of the optical marker, and generate infrared image data.

3. The dynamic target synchronous tracking and interaction system according to claim 2, characterized in that, The label-free vision module includes: RGB camera and deep learning processor; RGB camera: Used to capture RGB image data of the actors; Deep learning processor: used to calculate the distance value of the pattern at different positions in the RGB image data, and store the distance value in each pixel of the RGB image data.

4. The dynamic target synchronous tracking and interaction system according to claim 3, characterized in that, The mode switching module includes: A calibration unit and a switching unit, wherein the switching unit includes a dynamic filter control component and a mode switching component; Calibration unit: used to uniformly calibrate the infrared camera and the RGB camera in the Blob mode and establish a spatial mapping relationship; Filter control component: used to manually adjust the filter parameters of a preset number of RGB cameras after calibration, and to switch a preset number of RGB cameras to video mode; Mode switching component: used to detect the loss rate of marker points in real time under the Blob mode, and determine whether the loss rate exceeds a preset loss rate threshold; If it is determined that the loss rate exceeds the loss rate threshold, the RGB camera that has switched to the video mode is invoked to perform supplementary shooting.

5. A dynamic target synchronous tracking and interaction system according to claim 4, characterized in that, The mode switching module further includes: A low-latency switching unit and a camera resource allocation unit, wherein the camera resource allocation unit includes a task allocation component and a computing load balancing component; Low-latency switching unit: used to extract the switching duration and switching time point of the filter, extract the exposure period of the RGB camera, and synchronize the switching time point and the switching duration with the exposure period in time sequence; Task allocation component: used to extract the number of cameras of the RGB cameras, fix a portion of the RGB cameras in the Blob mode according to a preset minimum number, and dynamically switch the remaining RGB cameras in the Blob mode; The computational load balancing component is used to process Blob mode data generated in the Blob mode by a preset FPGA and to process video mode data generated in the video mode by a preset GPU.

6. The dynamic target synchronous tracking and interaction system according to claim 5, characterized in that, The fusion processing module includes: A hybrid data fusion unit, comprising a data fusion acquisition component and a dynamic compensation component; Data fusion acquisition component: When in the Blob mode, the infrared camera acquires the main positioning data of the marker point, and the RGB camera assists in completing the marker point through a filter; When in the video mode, the RGB camera performs skeleton tracking on the actor, and the infrared camera identifies key movement nodes of the actor and provides precise pose based on the key nodes; Dynamic compensation component: used to perform tomographic monitoring on the data collected by the RGB camera when the filter parameters are adjusted, to obtain the data tomography, and to perform motion interpolation on the data tomography.

7. A dynamic target synchronous tracking and interaction system according to claim 6, characterized in that, The fusion processing module further includes: A hybrid sensor spatiotemporal synchronization unit, comprising a hardware spatiotemporal synchronization component and a spatial calibration component; Hardware time-space synchronization component: used to trigger hardware synchronization signal, synchronize the working time of the infrared camera and the RGB camera according to the hardware synchronization signal, and align the time axis of the Blob mode data and the video mode data. Spatial calibration component: used to construct a unified coordinate system and establish an error compensation model based on a preset shared calibration target; Based on the unified coordinate system, the data deviation between the Blob mode data and the video mode data is extracted, and the data deviation is corrected according to the error compensation model.

8. A dynamic target synchronization tracking and interaction system according to claim 7, characterized in that, The fusion processing module further includes: A dynamic weight fusion unit, comprising an occlusion detection component and an adaptive weight allocation component; Occlusion detection component: used to extract unmarked feature points of actors in the RGB image data, and monitor the visibility value of the optical markers and the confidence level of the unmarked feature points in real time; Based on the visibility value and the confidence level, the weight allocation parameters of the Blob mode data and the video mode data are dynamically adjusted in real time. Adaptive weight allocation component: used to determine the occlusion status of the optical marker and the unmarked feature points based on the weight allocation parameters; If it is determined that there is no obstruction, the weight value of the Blob mode data is increased, and the video mode data is used as an auxiliary parameter to correct the slight jitter of the Blob mode data. If it is determined to be partial occlusion, the occlusion area is recorded. The video mode data takes over the occlusion area and performs cross-modal feature association based on the unmarked feature points to complete the data of the occlusion area. If the occlusion is determined to be global, a preset long short-term memory model is used to predict the actor's movement trajectory and obtain the movement trajectory data.

9. A dynamic target synchronization tracking and interaction method, wherein the method uses a dynamic target synchronization tracking and interaction system as described in any one of claims 1-8, characterized in that, The method includes: Real-time capture of infrared image data, RGB image data, and spatial depth information of actors in the virtual shooting environment; Based on the infrared image data and the RGB image data, the marker point information on the actor's body is obtained, and the loss rate of the marker point information is extracted; Determine whether the loss rate exceeds a preset loss rate threshold. If the loss rate exceeds the loss rate threshold, extract the actor's position data, posture data, and motion information based on the RGB image data and the spatial depth information. Based on the marker point information, a real-time 3D model of the actor in the virtual environment is obtained; based on the position data, the posture data, and the action information, the dynamic spatial position of the actor in the virtual performance environment is obtained. The real-time 3D model is fused and synchronously mapped with the dynamic spatial location to generate real-time dynamic images of the actor in the virtual environment.

Citation Information

Patent Citations

  • AR-based live broadcast real-time interaction system and method

    CN120434412A

  • Method for monitoring object through image fusion in monitoring system

    KR1020140017222A

  • Computer-generated image processing including volumetric scene reconstruction to replace a designated region

    US11055900B1