A dynamic target synchronous tracking and interaction system and method
By combining optical positioning and markerless vision modules, and utilizing filter switching and data fusion processing, high-precision and anti-interference dynamic tracking and interaction of dynamic targets are achieved, solving the problems of spatiotemporal synchronization and error compensation in existing technologies.
Patent Information
- Application Number
- CN202511563663.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-10-30
AI Technical Summary
Existing dynamic target tracking technologies struggle to simultaneously achieve both high precision and strong anti-interference capabilities. Hybrid systems suffer from shallow data fusion layers, failing to effectively address the issues of spatiotemporal synchronization and error compensation.
By combining an optical positioning module and a markerless vision module, switching between Blob mode and video mode is achieved through filter switching. Data fusion and spatiotemporal synchronization are performed using a fusion processing module, and dynamic weight allocation is implemented to achieve real-time tracking of dynamic targets.
It improves the accuracy and anti-interference ability of the motion capture process, ensuring high-precision tracking and real-time interaction of dynamic targets in complex environments.
Smart Images

Figure CN121033947B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of motion capture, and in particular to a dynamic target synchronous tracking and interaction system and method. BACKGROUND
[0002] With the rapid development of immersive technologies such as VR, AR, MR, dynamic target tracking technology is increasingly important in these fields.
[0003] However, existing dynamic target tracking technologies are mainly divided into two types: optical marker-based and markerless, each with its own advantages and disadvantages. Optical marker: relies on marker points attached to the target object for tracking, which can provide high tracking accuracy, but when the marker points are obscured, fall off, or are in a complex lighting environment, the tracking reliability will be significantly reduced, thereby limiting the application range of the system. Markerless: tracks by recognizing natural feature points on the target object (such as human joint points, object contours, etc.). However, this method is limited by high computational complexity, insufficient dynamic target tracking accuracy, and difficulty in handling similar feature points in complex scenes, resulting in limited effectiveness in practical applications.
[0004] A single system cannot simultaneously meet the needs of high accuracy and strong anti-interference capability. Therefore, it is necessary to develop a new hybrid tracking technology that combines the advantages of optical marker and markerless systems to overcome the limitations of existing technologies.
[0005] In related patents, existing hybrid systems mostly use independent working mode, with shallow data fusion level (such as simple weighted average), and do not effectively solve the problems of time and space synchronization and error compensation. SUMMARY
[0006] The purpose of the present application is to provide a dynamic target synchronous tracking and interaction system and method to solve the problems raised in the background art.
[0007] In a first aspect, the present application provides a dynamic target tracking and interaction system, which comprises:
[0008] An optical positioning module: for switching the filter to a narrow-band infrared filter and marking it as Blob mode, real-time capturing of infrared image data of the actor in the virtual shooting environment, and obtaining marker point information according to the infrared image data;
[0009] A markerless vision module: for switching the filter to full-spectrum transmission and marking it as video mode, real-time capturing of RGB image data and spatial depth information of the actor in the virtual shooting environment;
[0010] A mode switching module is configured to switch between the Blob mode and the video mode, and use the Blob mode for unified calibration.
[0011] A fusion processing module is configured to obtain a real-time three-dimensional model of an actor according to the marker point information, obtain a dynamic spatial position of the actor according to the RGB image data and the spatial depth information, perform spatio-temporal synchronization and dynamic weight distribution on the real-time three-dimensional model and the dynamic spatial position, and perform fusion synchronization mapping on the real-time three-dimensional model and the dynamic spatial position after the weight distribution.
[0012] Preferably, the optical positioning module comprises:
[0013] An infrared camera and an optical marker.
[0014] The optical marker is configured to be arranged on a costume or accessory worn by the actor, and reflect infrared light emitted from outside back.
[0015] The infrared camera is configured to emit infrared light, and receive the reflected infrared light, to determine marker point information of the optical marker, and generate infrared image data.
[0016] Preferably, the markerless vision module comprises:
[0017] An RGB camera and a deep learning processor.
[0018] The RGB camera is configured to capture RGB image data of the actor.
[0019] The deep learning processor is configured to calculate distance values of different positions of a pattern in the RGB image data, and store the distance values in each pixel point in the RGB image data.
[0020] Preferably, the mode switching module comprises:
[0021] A calibration unit and a switching unit, the switching unit comprising a dynamic filter control component and a mode switching component.
[0022] The calibration unit is configured to calibrate the infrared camera and the RGB camera in the Blob mode, and construct a spatial mapping relationship.
[0023] The filter control component is configured to manually adjust filter parameters of a part of the RGB cameras of a preset number after calibration, and switch the part of the RGB cameras of the preset number to the video mode.
[0024] The mode switching component is configured to detect a loss rate of marker points in the Blob mode in real time, and determine whether the loss rate exceeds a preset loss rate threshold.
[0025] If it is judged that the loss rate exceeds the loss rate threshold, the RGB camera in the video mode is called for supplementary shooting.
[0026] Preferably, the mode switching module further comprises:
[0027] A low-latency switching unit and a camera resource allocation unit, the camera resource allocation unit comprising a task allocation component and a computing load balancing component;
[0028] The low-latency switching unit is configured to extract the switching duration and switching time point of the filter, extract the exposure period of the RGB camera, and time-synchronize the switching time point and the switching duration with the exposure period.
[0029] The task allocation component is configured to extract the number of cameras of the RGB camera, and fix part of the RGB camera as the Blob mode according to a preset minimum number, and perform dynamic mode switching on the remaining RGB camera.
[0030] The computing load balancing component is configured to process Blob mode data generated in the Blob mode by a preset FPGA, and process video mode data generated in the video mode by a preset GPU.
[0031] Preferably, the fusion processing module comprises:
[0032] A mixed data fusion unit, the mixed data fusion unit comprising a data fusion acquisition component and a dynamic compensation component.
[0033] The data fusion acquisition component is configured to, when in the Blob mode, acquire main positioning data of the marker point by the infrared camera, and complete the marker point by the RGB camera with the aid of the filter.
[0034] When in the video mode, the RGB camera performs skeleton tracking on the actor, and the infrared camera identifies the action key node of the actor and provides accurate pose according to the key node.
[0035] The dynamic compensation component is configured to, when the filter adjusts the filter parameters, perform tomographic monitoring on the data collected by the RGB camera to obtain a data tomography, and perform motion interpolation on the data tomography.
[0036] Preferably, the fusion processing module further comprises:
[0037] A mixed sensor space-time synchronization unit, the mixed sensor space-time synchronization unit comprising a hardware space-time synchronization component and a space calibration component.
[0038] hardware space-time synchronization component: for triggering a hardware synchronization signal, according to the hardware synchronization signal, synchronizing the working time of the infrared camera and the RGB camera, and aligning the time axis of the Blob mode data and the video mode data;
[0039] space calibration component: for constructing a unified coordinate system according to a preset shared calibration target, and establishing an error compensation model;
[0040] based on the unified coordinate system, extracting a data deviation existing between the Blob mode data and the video mode data, and correcting the data deviation according to the error compensation model.
[0041] Preferably, the fusion processing module further comprises:
[0042] a dynamic weight fusion unit, the dynamic weight fusion unit comprising an occlusion detection component and an adaptive weight distribution component;
[0043] the occlusion detection component: for extracting the unmarked feature points of the actor in the RGB picture data, and monitoring the visible value of the optical marker and the confidence of the unmarked feature points in real time;
[0044] according to the visible value and the confidence, dynamically adjusting the weight distribution parameters of the Blob mode data and the video mode data in real time;
[0045] the adaptive weight distribution component: for judging the occlusion of the optical marker and the unmarked feature points according to the weight distribution parameters;
[0046] if it is judged that there is no occlusion, the weight value of the Blob mode data is increased, and the video mode data is used as a secondary parameter to slightly jitter correct the Blob mode data;
[0047] if it is judged that there is partial occlusion, the occluded area of partial occlusion is recorded, the video mode data takes over the occluded area, and cross-modal feature correlation is performed according to the unmarked feature points to complete the data of the occluded area;
[0048] if it is judged that there is global occlusion, a preset long short-term memory model is enabled to predict the motion trajectory of the actor to obtain motion trajectory data.
[0049] In a second aspect, the application provides a dynamic target synchronous tracking and interaction method, the method comprising:
[0050] real-time capturing of infrared picture data, RGB picture data and spatial depth information of the actor in a virtual shooting environment;
[0051] Based on the infrared picture data and the RGB picture data, marker point information on the actor is obtained, and a loss rate of the marker point information is extracted;
[0052] It is judged whether the loss rate exceeds a preset loss rate threshold. If it is judged that the loss rate exceeds the loss rate threshold, position data, posture data and action information of the actor are extracted according to the RGB picture data and the spatial depth information;
[0053] According to the marker point information, a real-time three-dimensional model of the actor in a virtual environment is obtained, and a dynamic spatial position of the actor in a virtual performance environment is obtained according to the position data, the posture data and the action information;
[0054] The real-time three-dimensional model and the dynamic spatial position are fused and synchronously mapped to generate a real-time dynamic image of the actor in the virtual environment.
[0055] In summary, the present application includes at least one of the following beneficial technical effects:
[0056] By respectively calling infrared cameras and RGB cameras to collect infrared picture data and RGB picture data of the actor in a virtual shooting environment and spatial depth information, according to the infrared picture data, marker point information on the actor is obtained, and according to the RGB picture data and the spatial depth information, position data, posture data and action information of the actor are obtained, wherein the optical positioning module is used to obtain the infrared picture data, which is marked as a Blob mode, and the marker-free vision module is used to obtain the RGB picture data and the spatial depth information, which is marked as a video mode. The Blob mode and the video mode are switched by a filter on the RGB camera. The collected data is fused and processed, the real-time three-dimensional model of the actor is obtained according to the marker point information, the dynamic spatial position of the actor is obtained according to the RGB picture data and the spatial depth information, and the real-time three-dimensional model and the dynamic spatial position are real-time weight allocated, the missing data in the corresponding mode is obtained according to the allocated weight in real time, and the final real-time dynamic image of the actor in the virtual environment is obtained. The accuracy and anti-interference ability in the motion capture process are improved. BRIEF DESCRIPTION OF DRAWINGS
[0057] Fig. 1 is a module block diagram of a dynamic target synchronous tracking and interaction system provided by the present application.
[0058] Fig. 2 is a step flow chart of a dynamic target synchronous tracking and interaction method provided by the present application.
[0059] Explanation of reference numerals: 1, optical positioning module; 2, marker-free vision module; 3, mode switching module; 4, fusion processing module. DETAILED DESCRIPTION
[0060] The following description Figs. 1-2 The application is further described in detail by the following embodiments, but the embodiments of the application are not limited thereto.
[0061] The embodiments of the present application disclose a dynamic target synchronous tracking and interaction system and method.
[0062] In the embodiments of the present application, a dynamic target synchronous tracking and interaction system comprises:
[0063] An optical positioning module 1 is configured to switch the filter to a narrow-band infrared filter, mark as a Blob mode, capture infrared picture data of an actor in a virtual shooting environment in real time, and obtain marker point information according to the infrared picture data;
[0064] A marker-free vision module 2 is configured to switch the filter to full-spectrum transmission, mark as a video mode, capture RGB picture data and spatial depth information of the actor in the virtual shooting environment in real time;
[0065] A mode switching module 3 is configured to switch the Blob mode and the video mode, and perform unified calibration using the Blob mode;
[0066] A fusion processing module 4 is configured to obtain a real-time three-dimensional model of the actor according to the marker point information, obtain a dynamic spatial position of the actor according to the RGB picture data and the spatial depth information, perform space-time synchronization and dynamic weight distribution on the real-time three-dimensional model and the dynamic spatial position, and perform fusion synchronous mapping on the real-time three-dimensional model and the dynamic spatial position after weight distribution.
[0067] The optical positioning module comprises:
[0068] An infrared camera and an optical marker;
[0069] The optical marker is configured to be arranged on a costume or accessory worn by the actor, and reflect infrared light emitted from outside back;
[0070] The infrared camera is configured to emit infrared light, receive the reflected infrared light, determine marker point information of the optical marker, and generate infrared picture data.
[0071] In use, take the actor's movement in a virtual shooting scene as an example. The actor wears a specially designed costume with 42 reflective marker points (6mm in diameter, 92% in reflectivity), which are distributed according to ergonomics at key joints and torso locations. An infrared camera array (consisting of 8 Vicon Vero v2.2) emits infrared light at a wavelength of 850nm at a frequency of 120Hz, and each camera is equipped with a lens with a field of view angle of 30°. When the actor makes a hand-raising action, the infrared light reflected by marker point VH-12 (right elbow position) is captured by camera No. 3, generating infrared image data containing X=1253.2px, Y=867.5px coordinates. The system calculates the exact position of the marker point in three-dimensional space (X=1.2m, Y=0.8m, Z=2.3m) by triangulation method, with an error range of ±0.5mm. At the same time, the reflected intensity of marker point VH-09 (right shoulder position) is detected to be reduced to 65lux (lower than the threshold value of 80lux), triggering a low reflection alarm. All marker point data is transmitted to the processing host through gigabit Ethernet, generating a marker point information frame containing 42 three-dimensional coordinates, with a timestamp of 2023-07-15T14:22:36.125.
[0072] Markerless vision module, comprising:
[0073] RGB camera and deep learning processor;
[0074] RGB camera: for shooting RGB image data of the actor;
[0075] Deep learning processor: for calculating distance values of different positions of patterns in the RGB image data, and storing the distance values in each pixel point in the RGB image data.
[0076] In use, 3 Intel RealSense D455 cameras are used to shoot RGB images of the actor at 30fps, with a resolution of 1920x1080. When the actor turns around, the deep learning processor (NVIDIA Jetson AGX Orin) detects the sleeve wrinkle feature point (texture hash value #A8D3F2) of his right arm, and calculates that the feature point is 2.1m away from the camera through a stereo vision algorithm. In depth information processing, the system identifies 5 joint points of the actor's left hand (from fingertip to wrist), and measures that the middle finger joint (X=1.5m, Y=1.8m, Z=1.2m) has a depth difference of 0.3m with the background green screen. At the same time, 68 feature points of the actor's face are extracted from the RGB image, and it is detected that the right eyebrow is raised by 3.2mm (baseline value is 2.8mm), generating expression change data. All spatial data is stored in JSON format, containing timestamp 2023-07-15T14:22:36.130, with a time deviation of ±2ms from the infrared data.
[0077] A mode switching module, comprising:
[0078] A calibration unit and a switching unit, the switching unit comprising a dynamic filter control component and a mode switching component;
[0079] The calibration unit: for calibrating the infrared camera and the RGB camera in the Blob mode, and constructing a spatial mapping relationship;
[0080] The filter control component: for manually adjusting the filter parameters of a part of the RGB cameras of a preset number after calibration, and switching the part of the RGB cameras to the video mode;
[0081] The mode switching component: for detecting the loss rate of the marker points in real time in the Blob mode, and judging whether the loss rate exceeds a preset loss rate threshold;
[0082] If it is judged that the loss rate exceeds the loss rate threshold, the RGB camera switched to the video mode is called to make up for the shooting.
[0083] In use, in the calibration stage, a preset calibration board (containing 12x9 circular markers with a diameter of 10 mm) is used to simultaneously obtain optical calibration reference points (error 0.3px) and RGB calibration reference points (error 0.5px). The established spatial mapping relationship shows that the X-axis offset is +1.2px and the Y-axis offset is -0.8px, and the system automatically writes the correction parameter table. The manual switching component records the forced mode switching instruction of the operator at 14:22:37, and locks the No. 1 camera in the Blob mode for key point tracking. When the actor enters the strong light area (the ambient light intensity reaches 1500 lux), the system detects that the marker point loss rate suddenly rises from 5% to 32% (the threshold is set to 20%). The dynamic filter control component immediately adjusts the electric control filter (Lucent Optics DF-850) of the No. 3 camera, reduces the light transmittance from 90% to 45%, and at the same time switches the working mode from the Blob mode to the video mode.
[0084] The mode switching module further comprises:
[0085] A low-delay switching unit and a camera resource allocation unit, the camera resource allocation unit comprising a task allocation component and a computing load balancing component;
[0086] The low-delay switching unit: for extracting the switching duration and switching time point of the filter, extracting the exposure period of the RGB camera, and time-synchronously the switching time point and the switching duration with the exposure period;
[0087] Task allocation component: used to extract the number of camera of RGB camera, and fix part of RGB camera as Blob mode according to the preset minimum number, and switch the rest of RGB camera to dynamic mode;
[0088] Computational load balancing component: used to hand over the Blob mode data generated in Blob mode to the preset FPGA for processing, and hand over the video mode data generated in video mode to the preset GPU for processing.
[0089] In use, the system is configured with 6 cameras, and the cameras C, D, E and F are fixed as Blob mode according to the preset (processing delay 8ms), and the cameras A, B and C are switched dynamically. When the switching time of the electric control filter is 12ms, the low delay unit synchronously adjusts the exposure period of camera B to 14ms, ensuring that the switching process covers the complete exposure interval. The computational load balancing component allocates the Blob data (1.2GB per second) of 4 cameras to Xilinx Alveo U50 FPGA for processing, and the video data (2.4GB per second) is processed by NVIDIA RTX 6000 Ada GPU. At 14:22:38, the GPU load reaches 85%, and the video stream of 2 cameras is automatically reduced to 15fps to maintain real-time performance, and the overall processing delay is maintained within 22ms at this time.
[0090] Fusion processing module, comprising:
[0091] Mixed data fusion unit, the mixed data fusion unit comprising a data fusion acquisition component and a dynamic compensation component;
[0092] Data fusion acquisition component: used to collect the main positioning data of the marker point by the infrared camera when in Blob mode, and the RGB camera completes the marker point through the filter assistance;
[0093] When in video mode, the RGB camera performs skeleton tracking on the actor, and the infrared camera recognizes the action key node of the actor, and provides accurate pose according to the key node;
[0094] Dynamic compensation component: used to monitor the data collected by the RGB camera when the filter adjusts the filter parameters, obtain the data fault, and perform motion interpolation on the data fault.
[0095] In the Blob mode, the infrared camera tracked the 8 marker points on the actor's back (spatial error ±1.1mm), while the RGB camera assisted in identifying the marker point VH-22 (left scapula) that was blocked by the clothes through the filter, and completed its three-dimensional coordinates (X=0.9m, Y=1.2m, Z=1.8m). After switching to the video mode, the RGB camera skeleton tracking showed that the right arm was raised at an angle of 53°, and the infrared camera provided the elbow joint marker point (error ±0.8°) for calibration at the same time. When the filter switched at 14:22:39, the dynamic compensation component detected a 2-frame data fault, and generated transition data through motion interpolation (right wrist trajectory smoothness reached 94%), ensuring the continuity of the action.
[0096] The fusion processing module further comprises:
[0097] The mixed sensor space-time synchronization unit comprises a hardware space-time synchronization component and a space calibration component.
[0098] The hardware space-time synchronization component is used to trigger a hardware synchronization signal, synchronize the working time of the infrared camera and the RGB camera according to the hardware synchronization signal, and align the time axis of the Blob mode data and the video mode data.
[0099] The space calibration component is used to construct a unified coordinate system according to a preset shared calibration target, and establish an error compensation model.
[0100] Based on the unified coordinate system, the data deviation between the Blob mode data and the video mode data is extracted, and the data deviation is corrected according to the error compensation model.
[0101] In use, the hardware synchronization signal is triggered at a frequency of 500Hz, so that the time deviation of the infrared camera (exposure time 2ms) and the RGB camera (exposure time 3ms) is controlled within ±0.1ms. After establishing the unified coordinate system using the shared calibration target (size 0.5m×0.5m), it is detected that the Blob mode data has an average +2.3mm deviation in the Z axis, and the error compensation model is corrected by a second-order polynomial (R²=0.98). In the data fusion at 14:22:40, the system corrected the 1.5px position deviation of the right hand marker point caused by lens distortion, and finally the coordinate accuracy was improved to ±0.6mm.
[0102] The fusion processing module further comprises:
[0103] The dynamic weight fusion unit comprises an occlusion detection component and an adaptive weight distribution component.
[0104] The occlusion detection component is used to extract the unmarked feature points of the actor in the RGB picture data and monitor the visibility of the optical marker and the confidence of the unmarked feature points in real time.
[0105] According to the visibility and the confidence, the weight distribution parameters of the Blob mode data and the video mode data are dynamically adjusted in real time.
[0106] The adaptive weight distribution component is used to determine the occlusion of the optical marker and the unmarked feature points according to the weight distribution parameters.
[0107] If it is determined that there is no occlusion, the weight value of the Blob mode data is increased, and the video mode data is used as a secondary parameter to slightly jitter and correct the Blob mode data.
[0108] If it is determined that there is partial occlusion, the occluded area of the partial occlusion is recorded, the video mode data takes over the occluded area, and the data of the occluded area is completed according to the cross-modal feature correlation of the unmarked feature points.
[0109] If it is determined that there is global occlusion, a preset long short-term memory model is used to predict the motion trajectory of the actor to obtain the motion trajectory data.
[0110] When the actor's right hand is occluded by a prop, the system detects that the visibility of the marker points VH-15 to VH-18 decreases from 100% to 12%, while the unmarked feature point confidence remains at 78%. The weight distribution parameters are dynamically adjusted to Blob mode 30% and video mode 70%. For the unoccluded left hand area (visibility 95%), the Blob mode weight remains 85%, and the video mode is only used to eliminate 0.3mm level hand jitter. When global occlusion occurs (14:22:41), the long short-term memory model predicts that the probability of the actor moving to the right in the next step is 82%, and generates the motion trajectory data within 3 seconds (maximum error 3.2cm) until the marker points reappear.
[0111] The embodiment of the present application provides a dynamic target synchronous tracking and interaction method using the dynamic target synchronous tracking and interaction system.
[0112] S100: capturing infrared picture data, RGB picture data and spatial depth information of an actor in a virtual shooting environment in real time;
[0113] S200: obtaining marker point information on the actor based on the infrared picture data and the RGB picture data, and extracting a loss rate of the marker point information.
[0114] S300: judging whether the loss rate exceeds a preset loss rate threshold, and if judging that the loss rate exceeds the loss rate threshold, extracting position data, posture data and action information of the actor according to the RGB picture data and the spatial depth information;
[0115] S400: obtaining a real-time three-dimensional model of the actor in the virtual environment according to the marker point information, and obtaining a dynamic spatial position of the actor in the virtual performance environment according to the position data, the posture data and the action information;
[0116] S500: synchronously mapping the real-time three-dimensional model and the dynamic spatial position to generate a real-time dynamic image of the actor in the virtual environment.
[0117] The above are preferred embodiments of the present application, which do not limit the protection scope of the present application, and thus: any equivalent changes made on the structure, shape and principle of the present application should be covered within the protection scope of the present application.
Claims
1. A dynamic target synchronous tracking and interaction system, characterized in that, include: Optical positioning module: used to switch the filter to a narrowband infrared filter and mark it as Blob mode, capture the infrared image data of the actor in the virtual shooting environment in real time, and obtain the marker point information based on the infrared image data; The optical positioning module includes an infrared camera; Unmarked visual module: Used to switch the filter to full-spectrum transmission and mark it as video mode, capturing RGB image data and spatial depth information of the actor in the virtual shooting environment in real time; The markerless vision module includes an RGB camera; Mode switching module: used to switch between the Blob mode and the video mode, and to perform unified calibration in the Blob mode; Fusion processing module: used to obtain the actor's real-time 3D model based on the marker point information, obtain the actor's dynamic spatial position based on the RGB image data and the spatial depth information, perform spatiotemporal synchronization and dynamic weight allocation on the real-time 3D model and the dynamic spatial position, and perform fusion synchronization mapping on the weighted real-time 3D model and the dynamic spatial position; The mode switching module includes: A calibration unit and a switching unit, wherein the switching unit includes a dynamic filter control component and a mode switching component; Calibration unit: used to uniformly calibrate the infrared camera and the RGB camera in the Blob mode and establish a spatial mapping relationship; Filter control component: used to manually adjust the filter parameters of a preset number of RGB cameras after calibration, and to switch a preset number of RGB cameras to video mode; Mode switching component: used to detect the loss rate of marker points in real time under the Blob mode, and determine whether the loss rate exceeds a preset loss rate threshold; If it is determined that the loss rate exceeds the loss rate threshold, the RGB camera that has switched to the video mode is invoked to perform supplementary shooting.
2. The dynamic target synchronous tracking and interaction system according to claim 1, characterized in that, The optical positioning module further includes: Optical marking; Optical markers: used to be placed on the costumes or accessories worn by actors to reflect infrared light emitted from the outside back; Infrared camera: used to emit infrared light and receive reflected infrared light, determine the marker point information of the optical marker, and generate infrared image data.
3. The dynamic target synchronous tracking and interaction system according to claim 2, characterized in that, The label-free vision module further includes: Deep learning processor; RGB camera: Used to capture RGB image data of the actors; Deep learning processor: used to calculate the distance value of the pattern at different positions in the RGB image data, and store the distance value in each pixel of the RGB image data.
4. The dynamic target synchronous tracking and interaction system according to claim 3, characterized in that, The mode switching module further includes: A low-latency switching unit and a camera resource allocation unit, wherein the camera resource allocation unit includes a task allocation component and a computing load balancing component; Low-latency switching unit: used to extract the switching duration and switching time point of the filter, extract the exposure period of the RGB camera, and synchronize the switching time point and the switching duration with the exposure period in time sequence; Task allocation component: used to extract the number of cameras of the RGB cameras, fix a portion of the RGB cameras in the Blob mode according to a preset minimum number, and dynamically switch the remaining RGB cameras in the mode; The load balancing component is used to process Blob mode data generated in the Blob mode by a preset FPGA and video mode data generated in the video mode by a preset GPU.
5. A dynamic target synchronous tracking and interaction system according to claim 4, characterized in that, The fusion processing module includes: A hybrid data fusion unit, comprising a data fusion acquisition component and a dynamic compensation component; Data fusion acquisition component: When in the Blob mode, the infrared camera acquires the main positioning data of the marker point, and the RGB camera assists in completing the marker point through a filter; When in the video mode, the RGB camera performs skeleton tracking on the actor, and the infrared camera identifies key movement nodes of the actor and provides precise pose based on the key nodes; Dynamic compensation component: used to perform tomographic monitoring on the data collected by the RGB camera when the filter parameters are adjusted, to obtain the data tomography, and to perform motion interpolation on the data tomography.
6. The dynamic target synchronous tracking and interaction system according to claim 5, characterized in that, The fusion processing module further includes: A hybrid sensor spatiotemporal synchronization unit, comprising a hardware spatiotemporal synchronization component and a spatial calibration component; Hardware time-space synchronization component: used to trigger hardware synchronization signal, synchronize the working time of the infrared camera and the RGB camera according to the hardware synchronization signal, and align the time axis of the Blob mode data and the video mode data; Spatial calibration component: used to construct a unified coordinate system and establish an error compensation model based on a preset shared calibration target; Based on the unified coordinate system, the data deviation between the Blob mode data and the video mode data is extracted, and the data deviation is corrected according to the error compensation model.
7. A dynamic target synchronous tracking and interaction system according to claim 6, characterized in that, The fusion processing module further includes: A dynamic weight fusion unit, comprising an occlusion detection component and an adaptive weight allocation component; Occlusion detection component: used to extract unmarked feature points of actors in the RGB image data, and monitor the visibility value of the optical markers and the confidence level of the unmarked feature points in real time; Based on the visibility value and the confidence level, the weight allocation parameters of the Blob mode data and the video mode data are dynamically adjusted in real time. Adaptive weight allocation component: used to determine the occlusion status of the optical marker and the unmarked feature points based on the weight allocation parameters; If it is determined that there is no obstruction, the weight value of the Blob mode data is increased, and the video mode data is used as an auxiliary parameter to correct the slight jitter of the Blob mode data. If it is determined to be partial occlusion, the occlusion area is recorded. The video mode data takes over the occlusion area and performs cross-modal feature association based on the unmarked feature points to complete the data of the occlusion area. If the occlusion is determined to be global, a preset long short-term memory model is used to predict the actor's movement trajectory and obtain the movement trajectory data.
8. A dynamic target synchronization tracking and interaction method, wherein the method uses a dynamic target synchronization tracking and interaction system as described in any one of claims 1-7, characterized in that, The method includes: Real-time capture of infrared image data, RGB image data, and spatial depth information of actors in the virtual shooting environment; Based on the infrared image data and the RGB image data, the marker point information on the actor's body is obtained, and the loss rate of the marker point information is extracted; Determine whether the loss rate exceeds a preset loss rate threshold. If the loss rate exceeds the loss rate threshold, extract the actor's position data, posture data, and motion information based on the RGB image data and the spatial depth information. Based on the marker point information, a real-time 3D model of the actor in the virtual environment is obtained; based on the position data, the posture data, and the action information, the dynamic spatial position of the actor in the virtual performance environment is obtained. The real-time 3D model is fused and synchronously mapped with the dynamic spatial location to generate real-time dynamic images of the actor in the virtual environment.
Citation Information
Patent Citations
AR-based live broadcast real-time interaction system and method
CN120434412A
Method for monitoring object through image fusion in monitoring system
KR1020140017222A