Picture dynamic capturing method and system based on tracking system

Through multi-sensor fusion and deep learning-driven multi-modal analysis, combined with adaptive learning mechanism, high-precision capture and prediction of dynamic object motion states is achieved, solving the problem of insufficient equipment dependence and environmental adaptability in the existing technology, and improving the authenticity and robustness of the capture results.

CN120126047AInactive Publication Date: 2025-06-10WIDELINK TECH CO LTD

Patent Information

Application Number
CN202510189167.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing dynamic capture technology relies on specific hardware devices, limits the applicability and flexibility in complex environments, ignores the interaction between dynamic objects and the environment, resulting in a lack of realism and integrity in the capture results, and has weak processing capabilities for abnormal data.

Method used

The dynamic picture capture method based on the tracking system is adopted, and through multi-sensor fusion, deep learning-driven multi-modal analysis and adaptive learning mechanism, the motion state of dynamic objects is captured and predicted in real time, and a high-precision three-dimensional motion trajectory is generated based on external environment information.

Benefits of technology

It significantly improves the accuracy, consistency and environmental adaptability of dynamic object motion capture, enhances the robustness and adaptability of the system, and provides more realistic and reliable picture capture results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126047A_ABST
    Figure CN120126047A_ABST
Patent Text Reader

Abstract

The invention provides a picture dynamic capturing method and system based on a tracking system, and the method comprises the steps: processing a video stream through the real-time transmission of a multi-view video stream and the collection of external environment information through a space calibration algorithm, so as to obtain a video frame; performing multi-modal analysis on the video frame by using a scene understanding algorithm, automatically identifying key motion points of the dynamic object and interaction characteristics between the key motion points and the surrounding environment so as to obtain position data of the key motion points, and then processing the position data by using a three-dimensional reconstruction technology so as to obtain a three-dimensional reconstruction result; a high-precision three-dimensional motion track is generated, a preset multi-mode motion mode library is combined, the motion state of a dynamic object is intelligently matched by using a self-adaptive learning mechanism, the influence of environmental factors is considered, and finally a precise image capture result is formed. According to the technical scheme provided by the invention, multi-sensor fusion, deep learning and adaptive learning mechanisms are integrated, and the precision, coherence and reliability of dynamic object motion capture are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of computer vision and image processing, and in particular to a method and system for capturing dynamic images based on a tracking system. Background Art

[0002] With the development of virtual reality (VR), augmented reality (AR), sports training, medical rehabilitation and other fields, the demand for accurate capture and real-time feedback of dynamic objects is growing. These applications not only require accurate capture of the motion trajectory of dynamic objects, but also require understanding of the interaction between dynamic objects and the surrounding environment. At present, there are many dynamic capture technologies on the market, such as optical-based dynamic capture systems and inertial sensor-based dynamic capture systems, which capture the motion information of target objects through high-speed cameras or inertial measurement units (IMUs). However, these existing solutions have some obvious shortcomings, such as relying on specific hardware devices to limit their applicability and flexibility in complex environments, ignoring the interaction between dynamic objects and the environment, resulting in a lack of realism and integrity in the capture results, and weak processing capabilities for abnormal data, which easily generates erroneous motion trajectories when external interference or sensor failure occurs. Therefore, it is particularly urgent to develop a new generation of dynamic capture technology that can overcome these defects. Summary of the invention

[0003] The embodiments of the present invention provide a method and system for capturing dynamic images based on a tracking system, so as to solve the problems in the prior art that the method only relies on specific hardware devices, which limits its applicability and flexibility in complex environments, ignores the interaction between dynamic objects and the environment, resulting in a lack of realism and integrity in the capture results, and has a weak ability to process abnormal data, and is prone to produce erroneous motion trajectories when there is external interference or sensor failure.

[0004] In a first aspect, an embodiment of the present invention provides a method for capturing dynamic images based on a tracking system, comprising:

[0005] The multi-view video stream is transmitted in real time by using a multi-sensor fusion tracking system, and external environment information is collected to obtain a multi-view video stream, wherein the external environment information includes: light, sound and temperature;

[0006] Using a preset time synchronization signal and a spatial calibration algorithm to perform frame synchronization and spatial alignment processing on the multi-view video stream, so as to obtain video frames with consistent timing and space;

[0007] Performing multimodal analysis on the video frames using a deep learning-driven scene understanding algorithm to automatically identify key motion points of dynamic objects and interactive features of the environment surrounding the dynamic objects, and obtaining key motion point position data containing interactive information;

[0008] Using a three-dimensional reconstruction technology based on a deep neural network to perform advanced mapping processing on the key motion point position data, to obtain a high-precision three-dimensional motion trajectory of the dynamic object;

[0009] Combining the high-precision three-dimensional motion trajectory with the preset multimodal action pattern library, an adaptive learning mechanism is used to intelligently match and predict the motion state of the dynamic object, while calculating the influence of environmental factors to generate a picture capture result.

[0010] Optionally, the key motion point position data is subjected to advanced mapping processing using a three-dimensional reconstruction technology based on a deep neural network to obtain a high-precision three-dimensional motion trajectory of the dynamic object, including:

[0011] Performing multi-view geometric correction on the key motion point position data using a deep neural network model to obtain corrected key motion point position data;

[0012] Performing environmental perception optimization on the corrected key motion point position data according to the external environment information to obtain enhanced key motion point position data to adjust the three-dimensional reconstruction parameters of the key motion points under different lighting, sound and temperature conditions;

[0013] Using time series analysis technology to smooth the enhanced key motion point position data to obtain a dynamic object motion trajectory;

[0014] The multi-scale feature fusion technology is used to integrate the multi-scale features of the motion trajectory of the dynamic object to generate a high-precision three-dimensional motion trajectory.

[0015] Optionally, it is characterized in that the enhanced key motion point position data is smoothed by using a time series analysis technology to obtain a dynamic object motion trajectory, including:

[0016] Using time series analysis technology to perform dual smoothing processing in the time domain and frequency domain on the enhanced key motion point position data, dynamically adjusting the filter parameters through an adaptive filtering algorithm to obtain an initial motion trajectory;

[0017] The initial motion trajectory is optimized by using a Kalman filter or a particle filter, and an optimized motion trajectory is generated by combining historical motion data of the dynamic object and a preset multimodal motion pattern library;

[0018] Using a machine learning model to perform anomaly detection on the optimized motion trajectory, identify and correct abnormal points caused by external interference or sensor errors, and obtain a clear motion trajectory;

[0019] The clear motion trajectory and external environment information are combined, and the motion parameters of the dynamic object are adjusted in real time using an environment-aware dynamic adjustment algorithm to generate a motion trajectory of the dynamic object.

[0020] Optionally, a machine learning model is used to perform anomaly detection on the optimized motion trajectory, identify and correct abnormal points caused by external interference or sensor errors, and obtain a clear motion trajectory, including:

[0021] Using a preset deep learning model to monitor the optimized motion trajectory in real time, detecting abnormal points in the trajectory, and obtaining preliminary identified abnormal points;

[0022] Extracting time series features, spatial features and environmental features from the optimized motion trajectory through multi-dimensional feature extraction technology, comparing them with historical normal motion data, identifying abnormal points caused by external interference, and obtaining accurately identified abnormal points;

[0023] Using multimodal data fusion technology, combining visual information, sound information and temperature information in multi-view video streams to verify the accurately identified abnormal points from multiple angles, and obtain verified abnormal points;

[0024] Using an adaptive interpolation algorithm to correct the verified abnormal points, restore the normal motion characteristics of the motion trajectory, and obtain a corrected motion trajectory;

[0025] Combining the corrected motion trajectory with the motion state of the dynamic object, an adaptive smoothing algorithm is used to optimize the smoothness and continuity of the corrected motion trajectory to generate the final clear motion trajectory.

[0026] Optionally, it is characterized in that the verified abnormal points are corrected by using an adaptive interpolation algorithm to restore the normal motion characteristics of the motion trajectory to obtain a corrected motion trajectory, including:

[0027] Using an adaptive interpolation algorithm, the interpolation parameters are dynamically adjusted according to the time series data before and after the verified abnormal point, and the abnormal point is accurately interpolated in combination with the local motion trend and the global motion pattern to restore the normal motion characteristics at the abnormal point and obtain a preliminary corrected motion trajectory;

[0028] The motion trajectory of the preliminary correction is optimized again by using a weighted average method based on neighboring points combined with the motion data and environmental information of multiple time points before and after the verified abnormal point, and the multimodal data in the time window is introduced for comprehensive correction to obtain an optimized corrected motion trajectory, wherein the motion data includes: visual information, sound information and temperature information;

[0029] The motion model of the dynamic object and the environment perception algorithm are used to perform consistency check on the optimized corrected motion trajectory, so that the optimized motion trajectory is consistent with the motion state of the dynamic object and adapts to different environmental changes to generate the final target corrected motion trajectory.

[0030] Optionally, a deep learning-driven scene understanding algorithm is used to perform multimodal analysis on the video frames, automatically identify key motion points of dynamic objects and interactive features of the surrounding environment of the dynamic objects, and obtain key motion point position data containing interactive information, including:

[0031] Performing multimodal analysis on the video frames using a deep learning-driven scene understanding algorithm to extract visual features, sound features, temperature features, and motion features from multiple perspectives of dynamic objects to obtain initial multimodal feature data;

[0032] Using a target detection and tracking algorithm in combination with the initial multimodal feature data, high-precision detection and tracking of dynamic objects is performed to identify key motion points of the dynamic objects and obtain initial position data of the key motion points;

[0033] Using an environment perception algorithm combined with interactive features in the video frame to optimize the environmental adaptability of the initial position data of the key motion point, adjusting the position and posture of the key motion point under different environmental conditions, and obtaining the key motion point position data containing interactive information;

[0034] Using context-aware technology to perform semantic understanding and logical reasoning on the key motion point position data containing the interactive information, and combining the historical motion trajectory of the dynamic object and the preset multimodal motion mode library to generate optimized key motion point position data;

[0035] The spatiotemporal correlation analysis technology is used to perform spatiotemporal correlation modeling on the optimized key motion point position data, analyze the motion laws of dynamic objects at different time points and spatial positions, generate a spatiotemporal motion model of the dynamic object, and obtain the target key motion point position data.

[0036] Optionally, combining the high-precision three-dimensional motion trajectory with a preset multimodal motion pattern library, an adaptive learning mechanism is used to intelligently match and predict the motion state of the dynamic object, and the influence of environmental factors is calculated to generate a picture capture result, including:

[0037] The high-precision three-dimensional motion trajectory is combined with a preset multi-modal action pattern library to perform multi-dimensional intelligent matching on the motion state of the dynamic object, calculate the spatiotemporal characteristics, speed and acceleration information of the motion trajectory, and obtain a preliminary motion state classification;

[0038] Using an adaptive learning mechanism to dynamically adjust the parameters of the multi-dimensional intelligent matching algorithm according to the historical motion data and real-time motion trajectory of the dynamic object, and combining the deep learning model and the reinforcement learning algorithm to optimize the preliminary motion state classification to obtain an optimized motion state classification;

[0039] Using the environment perception model in combination with the optimized motion state classification and external environment information, the influence of environmental factors on the motion state of the dynamic object is calculated to generate a motion state prediction result with strong environmental adaptability;

[0040] The final picture capture result is generated by utilizing the motion state prediction result with strong environmental adaptability and combining the real-time motion trajectory of the dynamic object and multimodal environmental data.

[0041] In a second aspect, an embodiment of the present application provides a screen dynamic capture system based on a tracking system, comprising:

[0042] A processing module, which performs frame synchronization and spatial alignment processing on the multi-view video stream using a preset time synchronization signal and a spatial calibration algorithm to obtain video frames with consistent timing and space;

[0043] An analysis module, which uses a deep learning-driven scene understanding algorithm to perform multimodal analysis on the video frames, automatically identifies key motion points of dynamic objects and interactive features of the surrounding environment of the dynamic objects, and obtains key motion point position data containing interactive information;

[0044] A mapping module, which uses a three-dimensional reconstruction technology based on a deep neural network to perform advanced mapping processing on the key motion point position data to obtain a high-precision three-dimensional motion trajectory of the dynamic object;

[0045] The generation module combines the high-precision three-dimensional motion trajectory with the preset multimodal action pattern library, adopts an adaptive learning mechanism to intelligently match and predict the motion state of the dynamic object, and calculates the influence of environmental factors to generate a picture capture result.

[0046] In a third aspect, an embodiment of the present invention provides a computing device, comprising a processor and a memory, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute a method for capturing dynamic images based on a tracking system as described in any one of the first aspects.

[0047] In a fourth aspect, an embodiment of the present invention provides a computer storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement a method for capturing dynamic images based on a tracking system as described in any one of the first aspects.

[0048] In an embodiment of the present invention, a multi-sensor fusion tracking system is first used to transmit the received multi-view video stream in real time, and external environment information is collected to obtain a multi-view video stream, and then a preset time synchronization signal and a spatial calibration algorithm are used to perform frame synchronization and spatial alignment processing on the multi-view video stream to obtain video frames with consistent timing and space; then a deep learning-driven scene understanding algorithm is used to perform multimodal analysis on the video frames, and the key motion points of dynamic objects and the interactive features of the surrounding environment of the dynamic objects are automatically identified to obtain key motion point position data containing interactive information; then a deep neural network-based three-dimensional reconstruction technology is used to perform advanced mapping processing on the key motion point position data to obtain a high-precision three-dimensional motion trajectory of the dynamic object; finally, the high-precision three-dimensional motion trajectory is combined with a preset multimodal action pattern library, and an adaptive learning mechanism is used to intelligently match and predict the motion state of the dynamic object, and the influence of environmental factors is calculated to generate a screen capture result. The present invention realizes real-time capture and prediction of the motion state of dynamic objects with high precision, high robustness and strong environmental adaptability by integrating multi-sensor fusion, deep learning-driven multimodal analysis and adaptive learning mechanism, significantly improving the accuracy, consistency and reliability of image capture, and providing strong technical support for applications in virtual reality, augmented reality, sports training, medical rehabilitation and other fields.

[0049] Furthermore, an adaptive interpolation algorithm is used to dynamically adjust the interpolation parameters according to the time series data before and after the verified outlier point, and the outlier point is accurately interpolated in combination with the local motion trend and the global motion pattern to restore the normal motion characteristics at the outlier point and obtain a preliminary corrected motion trajectory. Then, a weighted average method based on neighboring points is used to combine the motion data and environmental information of multiple time points before and after the outlier point, and multimodal data in the time window is introduced for comprehensive correction to obtain an optimized corrected motion trajectory. Finally, the motion model of the dynamic object and the environmental perception algorithm are used to perform a consistency check on the optimized corrected motion trajectory to generate the final corrected motion trajectory of the target.

[0050] According to the above steps, the present invention significantly improves the accuracy and robustness of the motion trajectory of dynamic objects through multi-sensor fusion, multi-modal analysis driven by deep learning, adaptive interpolation algorithm and environmental perception technology. In particular, in terms of processing abnormal data and adapting to complex environments, the present invention can effectively eliminate external interference and sensor errors, generate high-quality motion trajectories, and provide more reliable and accurate technical support for applications in the fields of virtual reality, augmented reality, sports training and medical rehabilitation.

[0051] These and other aspects of the present invention will become more apparent from the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0053] Figure 1 A flowchart of a method for dynamically capturing a picture based on a tracking system provided by an embodiment of the present invention;

[0054] Figure 2 A schematic structural diagram of a system for dynamically capturing a picture based on a tracking system provided by an embodiment of the present invention;

[0055] Figure 3 A schematic structural diagram of a computing device provided by an embodiment of the present invention. Detailed implementation manners

[0056] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention.

[0057] In some processes described in the specification, claims and the above accompanying drawings of the present invention, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0059] In the prior art, most existing methods rely on a single type of sensor or a combination of a few types of sensors, and cannot comprehensively capture the multi-modal information of dynamic objects, resulting in limitations in the integrity and accuracy of the capture results. Based on this, the present invention provides a method for dynamically capturing a picture based on a tracking system, as Figure 1, the specific steps are as follows:

[0060] Step 101: Use a multi-sensor fusion tracking system to perform real-time transmission on the received multi-view video stream and collect external environment information to obtain a multi-view video stream. Among them, the external environment information includes: light, sound, and temperature;

[0061] In this step, multiple cameras are used to capture dynamic objects from different angles to generate a multi-view video stream. By combining data from multiple sensors (such as cameras, microphones, temperature and humidity sensors, etc.), the integrity and reliability of the data are improved. Through a high-speed network or wireless communication technology, the multi-view video stream and external environment information are transmitted to the central processing unit in real time;

[0062] Step 102: Use a preset time synchronization signal and spatial calibration algorithm to perform frame synchronization and spatial alignment processing on the multi-view video stream to obtain video frames that are consistent in time sequence and space;

[0063] In this step, through a preset time synchronization signal, ensure that each frame of the multi-view video stream is aligned in time to eliminate time deviation. Use a spatial calibration algorithm to perform spatial alignment on the multi-view video stream to ensure that the video frames from different perspectives are aligned in space to eliminate perspective deviation. After processing, video frames that are aligned in both time and space are obtained, providing a basis for subsequent processing;

[0064] Step 103: Use a deep learning-driven scene understanding algorithm to perform multi-modal analysis on the video frames to automatically identify the key motion points of the dynamic object and the interaction features of the environment around the dynamic object, and obtain key motion point position data containing interaction information;

[0065] In this step, use a deep learning model (such as a convolutional neural network CNN) to analyze the video frames, extract multi-modal features of the dynamic object, combine visual features, sound features, temperature features, etc. in the video frames for comprehensive analysis, automatically identify the key motion points of the dynamic object (such as joints, facial features, etc.), identify the interaction features between the dynamic object and the surrounding environment, such as lighting conditions, sound information, temperature changes, object contacts, etc., and generate key motion point position data containing interaction information;

[0066] Step 104: Use a three-dimensional reconstruction technology based on a deep neural network to perform advanced mapping processing on the key motion point position data to obtain a high-precision three-dimensional motion trajectory of the dynamic object;

[0067] In this step, a deep neural network model is used to perform multi-view geometric correction on the key motion point position data to improve the accuracy of the data. According to the external environmental information (such as light, sound, temperature), the corrected key motion point position data is optimized, the 3D reconstruction parameters are adjusted, the optimized key motion point position data is smoothed to generate a preliminary motion trajectory of the dynamic object, and multi-scale feature integration is performed on the preliminary motion trajectory to generate a high-precision 3D motion trajectory;

[0068] Step 105: Combine the high-precision 3D motion trajectory with a preset multi-modal action pattern library, and use an adaptive learning mechanism to perform intelligent matching and prediction on the motion state of the dynamic object, while calculating the influence of environmental factors to generate a frame capture result.

[0069] In this step, the preset multi-modal action pattern library contains the 3D motion trajectories and features of various common actions. Using the high-precision 3D motion trajectory and combining with the multi-modal action pattern library, multi-dimensional intelligent matching of the motion state of the dynamic object is performed, and the spatio-temporal features, speed, and acceleration information of the motion trajectory are calculated to obtain a preliminary motion state classification;

[0070] According to the historical motion data and real-time motion trajectory of the dynamic object, the parameters of the multi-dimensional intelligent matching algorithm are dynamically adjusted, and the preliminary motion state classification is optimized by combining a deep learning model and a reinforcement learning algorithm to obtain an optimized motion state classification;

[0071] Combining the optimized motion state classification with the external environmental information (such as light conditions, sound information, temperature changes), calculate the influence of environmental factors on the motion state of the dynamic object, adjust the parameters of the pre-trained motion state prediction model, and generate a motion state prediction result with strong environmental adaptability;

[0072] Using the motion state prediction result with strong environmental adaptability, combining with the real-time motion trajectory of the dynamic object and multi-modal environmental data, generate a final frame capture result to ensure the accuracy, coherence, and environmental adaptability of frame capture.

[0073] The embodiments of the present invention can achieve the following beneficial effects through the above steps: Through multi-sensor fusion, deep learning-driven multi-modal analysis, high-precision 3D reconstruction, and an adaptive learning mechanism, this method significantly improves the accuracy, coherence, and environmental adaptability of dynamic object motion capture.

[0074] Based on this, the present invention provides a specific embodiment. In step 104, a 3D reconstruction technology based on a deep neural network is used to perform advanced mapping processing on the key motion point position data to obtain a high-precision 3D motion trajectory of the dynamic object, which specifically includes the following steps:

[0075] Step 201: Use a deep neural network model to perform multi-view geometric correction on the key motion point position data to obtain corrected key motion point position data;

[0076] In this step, a deep neural network (such as a convolutional neural network CNN or a deep residual network ResNet) is used to process the key motion point position data. These models learn multi-view geometric relationships through a large amount of training data and can effectively handle data inconsistencies under different views;

[0077] Multi-view geometry refers to shooting the same scene from different angles by multiple cameras and using geometric relationships to correct position differences under different views. Specifically, through a deep neural network model, the positions of each key motion point under different views can be corrected to ensure their position consistency in three-dimensional space. This process involves geometric operations such as triangulation and projective transformation.

[0078] The key motion point position data after multi-view geometric correction eliminates the deviation caused by different camera views, improving the accuracy and reliability of the data. These corrected data provide a basis for subsequent 3D reconstruction.

[0079] Step 202: Perform environmental perception optimization on the corrected key motion point position data according to the external environmental information to obtain enhanced key motion point position data, so as to adjust the 3D reconstruction parameters of key motion points under different lighting, sound, and temperature conditions;

[0080] Among them, the external environmental information includes environmental parameters such as light, sound, and temperature, and these information can be collected by additional sensors (such as light intensity sensors, microphones, temperature and humidity sensors);

[0081] Using an environmental perception model, optimize the corrected key motion point position data according to the external environmental information. The environmental perception model will adjust the brightness and contrast of the image according to the lighting conditions, adjust the accuracy of sound source localization according to the sound information, and adjust the performance parameters of the sensor according to the temperature information. These optimization measures can improve the accuracy and robustness of the key motion point position data under different environmental conditions.

[0082] The key motion point position data after environmental perception optimization can maintain higher accuracy and stability under different environmental conditions. For example, in an environment with large lighting changes, the optimized data can better maintain the visibility and accuracy of key motion points;

[0083] Step 203: Use time series analysis technology to smooth the enhanced key motion point position data to obtain the dynamic object motion trajectory;

[0084] In this step, time series analysis methods (such as autoregressive moving average ARMA model, Kalman filter, or particle filter) are used to process the data, and these methods can effectively handle the noise and discontinuities in time series data.

[0085] By smoothing the time series data to remove short-term high-frequency noise and retain the true motion characteristics of dynamic objects, the autoregressive moving average ARMA model can smooth the data by fitting the statistical characteristics of the time series, the Kalman filter can smooth the data through state estimation, and the particle filter can smooth the data through probability distribution;

[0086] Based on the position data of the key motion points after smoothing, a smooth and continuous dynamic object motion trajectory is generated. These trajectories not only remove noise but also retain the true motion characteristics of dynamic objects, improving the reliability and interpretability of the trajectories;

[0087] Step 204: Perform multi-scale feature integration on the dynamic object motion trajectory through multi-scale feature fusion technology to generate a high-precision three-dimensional motion trajectory;

[0088] In this step, multi-scale feature fusion methods (such as multi-scale convolutional neural network or multi-scale recurrent neural network) are used to extract and integrate features from different scales. Multi-scale feature fusion technology can capture the details and overall structure of the motion trajectory at different scales, improving the richness and diversity of features;

[0089] Integrate features at different scales to improve the details and overall coherence of the motion trajectory. Specifically, the multi-scale convolutional neural network can extract local and global features at different scales, and the multi-scale recurrent neural network can capture long-term and short-term dependencies in the time dimension. Through these technologies, a more refined and coherent motion trajectory can be generated;

[0090] Through multi-scale feature integration, high-precision and high-resolution three-dimensional motion trajectories are generated. These trajectories are not only more accurate in details but also more coherent and stable as a whole, ensuring the high quality and high reliability of motion capture.

[0091] Based on this, the present invention provides a specific embodiment. In step 203, time series analysis technology is used to smooth the enhanced key motion point position data to obtain a dynamic object motion trajectory, which specifically includes the following steps:

[0092] Step 301: Use time series analysis technology to perform double smoothing processing on the enhanced key motion point position data in the time domain and frequency domain, and dynamically adjust the filtering parameters through an adaptive filtering algorithm to obtain an initial motion trajectory;

[0093] In this step, time series analysis technology is used to process the enhanced key motion point position data. This technology is a statistical method for analyzing a sequence of data points arranged in chronological order to extract meaningful information and predict future values. Noise is reduced by applying smoothing techniques simultaneously in the time domain (changes on the time axis) and the frequency domain (frequency components), and an adaptive filtering algorithm is used to dynamically adjust the filtering parameters to ensure optimal filtering performance even when the data characteristics change. As a result, a relatively clean and smooth initial motion trajectory is generated, providing a basis for subsequent processing.

[0094] Step 302: Optimize the initial motion trajectory using a Kalman filter or a particle filter, and generate an optimized motion trajectory by combining the historical motion data of the dynamic object and a preset multi-modal action pattern library.

[0095] In this step, a Kalman filter or a particle filter is used to further optimize the initial motion trajectory obtained in the previous step. The Kalman filter is a recursive solution suitable for optimal estimation problems of linear systems, while the particle filter is based on the Monte Carlo simulation method and is suitable for state estimation problems with non-linear and non-Gaussian distributions. This step combines the historical motion data of the dynamic object and a preset multi-modal action pattern library, enabling the optimized motion trajectory to consider not only the data of the current frame but also historical information and expected action patterns, thereby improving the accuracy and rationality of the trajectory.

[0096] Step 303: Use a machine learning model to perform anomaly detection on the optimized motion trajectory, identify and correct anomaly points caused by external interference or sensor errors, and obtain a clear motion trajectory.

[0097] In this step, a machine learning model is used to monitor the optimized motion trajectory in real time and identify possible anomaly points. Here, the machine learning model refers to a mathematical model that learns patterns from data through training and can be used for tasks such as classification, regression, and clustering. Once anomaly points, i.e., data points that do not conform to the normal pattern, which may be caused by external interference or sensor errors, are detected, the system will attempt to accurately identify these anomaly points by comparing historical normal motion data and other features, and finally correct them through appropriate algorithms to ensure the clarity and accuracy of the motion trajectory.

[0098] Step 304: Combine the clear motion trajectory and external environment information, and use an environment-aware dynamic adjustment algorithm to adjust the motion parameters of the dynamic object in real time to generate a dynamic object motion trajectory.

[0099] In this step, the clear motion trajectory processed as above is combined with external environmental information such as light, sound, temperature, etc., and a dynamic adjustment algorithm for environmental perception is used to adjust the motion parameters of the dynamic object in real time. This algorithm can adjust the algorithm parameters in real time according to the external environmental information, enabling the system to better adapt to different environmental conditions. The purpose of doing this is to make the finally generated motion trajectory of the dynamic object not only reflect the real motion of the object, but also appropriately consider the influence of environmental factors, so as to more realistically reproduce the result of the captured picture.

[0100] Through time-domain and frequency-domain double smoothing processing, Kalman or particle filter optimization, machine learning anomaly detection and correction, and environmental perception dynamic adjustment technology in the embodiments of the present invention, the accuracy, clarity, and environmental adaptability of the motion trajectory of the dynamic object are effectively improved, providing a more realistic and reliable result of the captured picture.

[0101] Based on this, the present invention provides a specific embodiment. Step 303 specifically includes the following steps:

[0102] Step 401: Use a preset deep learning model to monitor the optimized motion trajectory in real time, detect the abnormal points in the trajectory, and obtain the preliminarily identified abnormal points;

[0103] In this step, a pre-trained deep learning model is used to monitor the optimized motion trajectory in real time. The deep learning model is a machine learning method that can automatically learn features from a large amount of data and is used here to detect the abnormal points in the trajectory. By comparing the current trajectory with the normal patterns learned by the model, the abnormal points that may be caused by external interference or sensor errors can be identified, thus obtaining the preliminarily identified abnormal points.

[0104] Step 402: Extract time series features, spatial features, and environmental features from the optimized motion trajectory through multi-dimensional feature extraction technology, compare them with historical normal motion data, identify the abnormal points caused by external interference, and obtain the accurately identified abnormal points;

[0105] In this step, multi-dimensional feature extraction technology is adopted, which is a technology used to extract different types of features (such as time series features, spatial features, and environmental features) from the optimized motion trajectory. These features will then be compared with the historical normal motion data to more accurately identify which abnormal points are caused by external factors, thereby obtaining the accurately identified abnormal points.

[0106] Step 403: Use multi-modal data fusion technology to verify the accurately identified abnormal points from multiple perspectives by combining visual information, sound information, and temperature information in multi-view video streams, and obtain the verified abnormal points;

[0107] In this step, the multi-modal data fusion technology combines data from multiple sources (such as visual information, sound information, and temperature information) to provide a more comprehensive and accurate analysis perspective. By integrating various information in the multi-view video stream, the previously identified abnormal points can be verified from multiple angles to ensure the accuracy of the judgment of abnormal points, and finally the verified abnormal points are obtained.

[0108] Step 404: Use the adaptive interpolation algorithm to correct the verified abnormal points, restore the motion characteristics of the normal motion trajectory, and obtain the corrected motion trajectory;

[0109] In this step, the adaptive interpolation algorithm is applied to dynamically adjust the interpolation parameters according to the time series data before and after the abnormal points, and process the abnormal points by combining the local motion trend and the global motion pattern. The purpose is to restore the normal motion characteristics at the abnormal points, thereby generating the corrected motion trajectory. This method can effectively make up for the data loss or errors caused by abnormal points.

[0110] Step 405: Combine the corrected motion trajectory and the motion state of the dynamic object, and use the adaptive smoothing algorithm to optimize the smoothness and coherence of the corrected motion trajectory to generate the final clear motion trajectory;

[0111] In this step, the adaptive smoothing algorithm is used to further optimize the corrected motion trajectory. The adaptive smoothing algorithm can automatically adjust its parameters according to specific conditions to ensure the smoothness and coherence of the motion trajectory. By combining the corrected motion trajectory and the motion state of the dynamic object, this algorithm helps to eliminate any remaining discontinuities or mutations and generate the final clear motion trajectory.

[0112] In the embodiment of the present invention, through the real-time monitoring of the deep learning model, the extraction of multi-dimensional features and the comparison with historical data, the verification of multi-modal data fusion, the correction of the adaptive interpolation algorithm, and the optimization of the adaptive smoothing algorithm, the system can effectively detect and correct the abnormal points in the motion trajectory, ensuring the accuracy, integrity, and coherence of the motion trajectory, while improving the robustness and adaptability of the system, providing a high-quality data basis for subsequent applications.

[0113] Based on this, the present invention provides a specific embodiment. In step 404, using the adaptive interpolation algorithm to correct the verified abnormal points, restore the motion characteristics of the normal motion trajectory, and obtain the corrected motion trajectory, specifically includes the following steps:

[0114] Step 501: Use the adaptive interpolation algorithm to dynamically adjust the interpolation parameters according to the time series data before and after the verified abnormal points, combine the local motion trend and the global motion pattern, perform precise interpolation processing on the abnormal points, restore the normal motion characteristics at the abnormal points, and obtain the preliminarily corrected motion trajectory;

[0115] In this step, an adaptive interpolation algorithm is applied to process the verified outliers. This algorithm can dynamically adjust its parameters according to the time series data before and after the outliers, and combine the local motion trend (i.e., the motion pattern near the outliers) and the global motion pattern (the overall motion characteristics) to perform precise interpolation processing on the outliers. The purpose is to restore the normal motion characteristics that should have been at the outliers, so as to obtain a preliminarily corrected motion trajectory. This method ensures that even in the presence of anomalies, the authenticity and coherence of the motion trajectory can be maintained as much as possible;

[0116] Since traditional methods usually rely on a single type of data (such as only visual information), they are not sensitive enough to changes in the external environment, ignore the value of historical motion data, and may perform poorly when facing rapidly changing motion scenarios. Therefore, the present invention introduces an adaptive interpolation algorithm to correct the verified outliers, restore the normal motion characteristics of the motion trajectory, correct the outliers and generate a more accurate and smoother motion trajectory. Among them, the expression of the adaptive interpolation algorithm is as follows:

[0117] Pcorrected(t) = α(t)·Plocal_trend(t)+(1-α(t))·Pglobal_model(t)+β(t)·Eenv(t)+γ(t)·Mmulti-modal(t)+

[0118] δ(t)·Fhistory(t)

[0119] Among them, P corrected (t) represents the position of the corrected motion trajectory at time t;

[0120] α(t) is a dynamically adjusted weight parameter, which depends on the characteristics of the time series data at time t, including the local motion trend and environmental information, and is used to dynamically adjust the weight between the local motion trend and the global model prediction, ensuring that the system can flexibly adjust the importance of local and global information in different situations, especially when dealing with rapidly changing or long-term stable motions, and automatically adjusts based on the adaptive learning algorithm according to real-time data;

[0121] P local_trend (t) represents the local motion trend, the estimated position at time t, which can be obtained by analyzing the motion data before and after the outliers, and is used to capture the local motion characteristics, reflect the short-term motion trend, and help to handle sudden changes within a short time, such as sudden actions or posture changes, and is obtained by analyzing the motion data before and after the outliers;

[0122] P global_model(t) represents the position predicted by the global motion model, which reflects the overall pattern or expected behavior of the entire motion trajectory. It is used to provide predictions of the overall motion pattern, ensure the continuity of the trajectory, help maintain motion consistency over a long period of time, and avoid trajectory mutations. It is trained based on a preset multi-modal action pattern library and deep learning model;

[0123] β(t) is a dynamically adjusted weight parameter used to adjust the influence of the environmental perception module on the correction result, calculate the influence of external environmental factors on the motion trajectory, such as light changes, noise interference, etc., and is dynamically adjusted according to environmental sensor data through machine learning algorithms;

[0124] E env E(t) is the output of the environmental perception module, representing the influence of the external environment (such as light, sound, temperature, etc.) on the positions of key motion points. It is used to reflect the influence of the external environment on the positions of key motion points, enhance the robustness and adaptability of the system, and ensure accurate capture of motion in different environments. It comes from the environmental perception module and combines sensor data such as light, sound, and temperature;

[0125] γ(t) is a dynamically adjusted weight parameter used to adjust the influence of the multi-modal data fusion module on the correction result, integrate various sensor data, and provide a more comprehensive and accurate description of motion. Based on multi-modal data fusion technology, it combines various sensor information such as vision, hearing, and temperature;

[0126] M multi-modal M(t) is the output of the multi-modal data fusion module, which combines various sensor data such as visual information, sound information, and temperature information to provide a more comprehensive environmental perception. It is used to combine various sensor data to provide a more comprehensive environmental perception, improve the perception ability and accuracy of the system, especially in complex environments;

[0127] δ(t) is a dynamically adjusted weight parameter used to adjust the influence of historical motion data on the correction result, utilize historical data to optimize the current prediction, and improve the stability and accuracy of the prediction. Based on time series analysis and adaptive learning mechanisms;

[0128] F history F(t) is the output of the historical motion data module, which combines the historical motion trajectories of dynamic objects to improve the accuracy and stability of the prediction, helps the system better understand the long-term behavior patterns of objects, and thus makes more accurate predictions. It is obtained by analyzing historical motion data;

[0129] Furthermore, the local motion trend P local_trend(t) captures the motion characteristics within a short period, reflecting the motion state of the dynamic object near the current time point. By introducing the weight parameter α(t), the importance of the local motion trend can be dynamically adjusted, which enables the system to be more flexible and accurate when dealing with sudden changes (such as sudden actions or pose changes);

[0130] Global model prediction P global_model (t) provides the overall pattern or expected behavior of the entire motion trajectory, ensuring the motion coherence and consistency over a long time period. Using (1 - α(t)) as the weight ensures the balance between local and global information. When α(t) is high, more importance is attached to local features; when α(t) is low, more reliance is placed on the global model. This helps to maintain the smoothness and continuity of the motion trajectory and avoid sudden changes;

[0131] Environmental perception module E env (t) reflects the influence of the external environment (such as light, sound, temperature, etc.) on the positions of key motion points. By introducing the weight parameter β(t), the influence of the environmental perception module can be dynamically adjusted according to real-time environmental conditions, which enhances the robustness and adaptability of the system and ensures accurate motion capture in different environments, especially in complex and changeable environments;

[0132] Multimodal data fusion module M multi-modal (t) combines various sensor data (such as vision, sound, temperature), providing a more comprehensive and accurate motion description. By introducing the weight parameter γ(t), the influence of multimodal data can be dynamically adjusted according to actual needs, which improves the perception ability and accuracy of the system, especially in situations where multiple information sources need to be integrated, such as dynamic capture in complex scenarios;

[0133] Historical motion data module F history (t) combines the historical motion trajectories of the dynamic object, using time series analysis and adaptive learning mechanisms to optimize the current prediction. By introducing the weight parameter δ(t), its influence can be dynamically adjusted according to historical data, ensuring that the current prediction takes into account both short-term changes and long-term trends. This improves the stability and accuracy of the prediction, especially when dealing with periodic or regular motions;

[0134] This algorithm formula aims to overcome the limitations of existing dynamic capture technologies. Through multi-modal data fusion (combining information such as vision, sound, and temperature) and an environmental perception module, it solves the problems of single data source limitation and insufficient environmental adaptability; it uses historical motion data to optimize current predictions, making up for the defect of traditional methods ignoring historical data; and through dynamically adjusting weight parameters and an adaptive learning mechanism, it enhances the flexibility of the system and the ability to handle rapidly changing scenarios. Combining these improvements significantly enhances the robustness, accuracy, and adaptability of the dynamic capture system, ensuring that high-quality and stable motion trajectories can still be provided in complex and changing environments, providing more reliable technical support for fields such as virtual reality, augmented reality, and animation production.

[0135] Step 502: Use the weighted average method based on neighboring points to combine the motion data and environmental information of multiple time points before and after the verified abnormal points, and optimize the preliminarily corrected motion trajectory again. Introduce multi-modal data within the time window for comprehensive correction to obtain the optimized corrected motion trajectory, where the motion data includes: visual information, sound information, and temperature information;

[0136] In this step, the weighted average method based on neighboring points is used to further optimize the preliminarily corrected motion trajectory. This process not only considers the motion data of multiple time points before and after the abnormal points but also introduces environmental information (such as vision, sound, temperature, etc.). Through comprehensive correction, the accuracy of the motion trajectory is improved. The multi-modal data within the time window is used to enhance the correction effect, ensuring that the obtained optimized corrected motion trajectory is closer to the actual motion situation and reducing the influence caused by abnormal points.

[0137] Step 503: Use the motion model of the dynamic object and the environmental perception algorithm to perform consistency verification on the optimized corrected motion trajectory, making the optimized motion trajectory conform to the motion state of the dynamic object and adapt to different environmental changes, and generate the final target corrected motion trajectory;

[0138] In this step, the motion model of the dynamic object and the environmental perception algorithm are used to perform consistency verification on the optimized corrected motion trajectory. This step aims to ensure that the optimized motion trajectory not only conforms to the actual motion state of the dynamic object but also can adapt to different environmental changes. By comparing the optimized motion trajectory with the preset motion model, any inconsistencies can be checked and adjusted to ensure that the finally generated target corrected motion trajectory is both faithful to the original motion and can maintain stability in different environments. In addition, the environmental perception algorithm can also adjust the trajectory in real time to cope with changes in environmental factors and provide a more accurate description of the motion.

[0139] In the embodiments of the present invention, through the precise processing of the adaptive interpolation algorithm, the weighted average optimization based on neighboring points, and the final consistency check and environmental adaptability adjustment, the system can maintain or even improve the quality of the motion trajectory while processing abnormal points.

[0140] Based on this, the present invention provides a specific embodiment. In step 103, a deep learning-driven scene understanding algorithm is used to perform multimodal analysis on the video frame, automatically identify the key motion points of the dynamic object and the interaction features of the environment around the dynamic object, and obtain the key motion point position data containing interaction information, which specifically includes the following steps:

[0141] Step 601: Use a deep learning-driven scene understanding algorithm to perform multimodal analysis on the video frame, extract the visual features, sound features, temperature features, and motion features from multiple perspectives of the dynamic object, and obtain the initial multimodal feature data;

[0142] In this step, a deep learning-driven scene understanding algorithm is used to perform multimodal analysis on the video frame. This algorithm can extract various features of the dynamic object from the video frame, including visual features (such as shape, color, texture), sound features (such as audio spectrum), temperature features (such as thermal imaging data), and motion features obtained from multiple perspectives. By integrating these different types of features, the initial multimodal feature data about the dynamic object and its environment can be obtained, providing rich and detailed information for subsequent processing.

[0143] Step 602: Use the object detection and tracking algorithm in combination with the initial multimodal feature data to perform high-precision detection and tracking on the dynamic object, identify the key motion points of the dynamic object, and obtain the initial position data of the key motion points;

[0144] In this step, the object detection and tracking algorithm is used in combination with the initial multimodal feature data obtained in the previous step to perform high-precision detection and tracking on the dynamic object. The goal of this step is to accurately identify the key motion points of the dynamic object and determine their initial position data. The object detection algorithm is used to locate and identify the dynamic object, while the tracking algorithm ensures continuous tracking of these objects in consecutive video frames, thus achieving accurate capture of the key motion points.

[0145] Step 603: Use the environmental perception algorithm in combination with the interaction features in the video frame to perform environmental adaptability optimization on the initial position data of the key motion points, adjust the position and posture of the key motion points under different environmental conditions, and obtain the key motion point position data containing interaction information;

[0146] In this step, an environment perception algorithm is applied to optimize the position data of key motion points. The environment perception algorithm takes into account interactive features in the video frame, such as lighting conditions, sound information, temperature changes, object contacts, etc., to adjust the position and posture of the key motion points under different environmental conditions. This can ensure that the data of the key motion points not only reflects the motion characteristics of the object itself, but also takes into account the influence of the surrounding environment, and finally obtains the position data of the key motion points containing interactive information, improving the accuracy of the data.

[0147] Step 604: Use context awareness technology to perform semantic understanding and logical reasoning on the position data of the key motion points containing interactive information, and generate optimized position data of the key motion points in combination with the historical motion trajectory of the dynamic object and the preset multi-modal action pattern library;

[0148] In this step, context awareness technology is used to perform semantic understanding and logical reasoning on the position data of the key motion points containing interactive information. The context awareness technology can combine the historical motion trajectory of the dynamic object and the preset multi-modal action pattern library to help the system understand the current motion state and predict the future motion trend. Through semantic analysis and logical reasoning of the data, the position data of the key motion points can be further optimized to make it more in line with the actual motion situation and the expected action pattern.

[0149] Step 605: Use spatio-temporal correlation analysis technology to perform spatio-temporal correlation modeling on the optimized position data of the key motion points, analyze the motion laws of the dynamic object at different time points and spatial positions, generate a spatio-temporal motion model of the dynamic object, and obtain the target position data of the key motion points;

[0150] In this step, the optimized position data of the key motion points is modeled through spatio-temporal correlation analysis technology. This technology aims to analyze the motion laws of the dynamic object at different time points and spatial positions, and construct a spatio-temporal motion model that describes how the dynamic object changes over time and space. This model can help to more deeply understand the behavior pattern of the dynamic object, and finally obtain the target position data of the key motion points, providing strong support for subsequent applications (such as animation production, virtual reality, augmented reality, etc.).

[0151] The embodiments of the present invention enhance the ability to understand and express the motion of dynamic objects through the above steps, and provide high-quality data support for motion analysis and applications in complex scenarios.

[0152] Based on this, the present invention provides a specific embodiment, and step 105 specifically includes the following steps:

[0153] Step 701: Use the high-precision three-dimensional motion trajectory to perform multi-dimensional intelligent matching on the motion state of the dynamic object in combination with a preset multi-modal action pattern library, calculate the spatio-temporal characteristics, speed, and acceleration information of the motion trajectory, and obtain a preliminary motion state classification.

[0154] In this step, the system uses the high-precision three-dimensional motion trajectory to perform multi-dimensional intelligent matching on the motion state of the dynamic object in combination with a preset multi-modal action pattern library. By analyzing the spatio-temporal characteristics (i.e., changes in time and space), speed, and acceleration information of the motion trajectory, the system can identify the motion pattern of the dynamic object and compare it with the pre-stored action patterns, thereby obtaining a preliminary motion state classification.

[0155] Step 702: Use an adaptive learning mechanism to dynamically adjust the parameters of the multi-dimensional intelligent matching algorithm according to the historical motion data and real-time motion trajectory of the dynamic object, and optimize the preliminary motion state classification by combining a deep learning model and a reinforcement learning algorithm to obtain an optimized motion state classification.

[0156] In this step, an adaptive learning mechanism is introduced to dynamically adjust the parameters of the multi-dimensional intelligent matching algorithm according to the historical motion data and real-time motion trajectory of the dynamic object. In this process, a deep learning model and a reinforcement learning algorithm are combined to further optimize the preliminary motion state classification. The deep learning model can capture complex non-linear relationships, while the reinforcement learning helps to make the best decisions in an uncertain environment. In this way, the system can more accurately identify and classify the motion state of the dynamic object, providing a more accurate optimization result.

[0157] Step 703: Use an environment perception model to calculate the influence of environmental factors on the motion state of the dynamic object in combination with the optimized motion state classification and external environmental information, and generate a motion state prediction result with strong environmental adaptability.

[0158] In this step, an environment perception model is used to calculate the influence of environmental factors (such as light, sound, temperature, etc.) on the motion state of the dynamic object in combination with the optimized motion state classification and external environmental information. The model aims to generate a motion state prediction result with strong environmental adaptability, ensuring that the prediction not only takes into account the current motion state but also fully considers the changes in the surrounding environment. This can improve the accuracy of the prediction, making the motion state of the dynamic object consistent and reasonable in different environments.

[0159] Step 704: Use the motion state prediction result with strong environmental adaptability to generate a final picture capture result in combination with the real-time motion trajectory and multi-modal environmental data of the dynamic object.

[0160] In this step, using the motion state prediction results with strong environmental adaptability, combining the real-time motion trajectories of dynamic objects and multi-modal environmental data, the final frame capture result is generated. This step integrates all the information processed previously, including the optimized motion state classification, the prediction results provided by the environmental perception model, and the real-time multi-modal data, ensuring that the final frame capture is both faithful to the original motion and can accurately reflect the behavior of dynamic objects in different environments. The finally output frame capture results can be used in various applications such as virtual reality, augmented reality, animation production, etc.

[0161] Through the above steps, the embodiments of the present invention improve the understanding and prediction ability of the motion state of dynamic objects, and ensure the accuracy, coherence and environmental adaptability of the capture results.

[0162] Figure 2 The following is a schematic structural diagram of a frame dynamic capture system based on a tracking system provided by an embodiment of the present application, as Figure 2 shown. The system includes:

[0163] A transmission module 21, configured to use a multi-sensor fusion tracking system to perform real-time transmission on the received multi-view video stream, and collect external environmental information to obtain a multi-view video stream, where the external environmental information includes: light, sound, and temperature;

[0164] A processing module 22, configured to perform frame synchronization and spatial alignment processing on the multi-view video stream by using a preset time synchronization signal and a spatial calibration algorithm to obtain video frames with consistent time sequence and space;

[0165] An analysis module 23, configured to perform multi-modal analysis on the video frames by using a deep learning-driven scene understanding algorithm, automatically identify the key motion points of dynamic objects and the interaction features of the environment around the dynamic objects, and obtain key motion point position data including interaction information;

[0166] A mapping module 24, configured to perform advanced mapping processing on the key motion point position data by using a three-dimensional reconstruction technology based on a deep neural network to obtain a high-precision three-dimensional motion trajectory of the dynamic object;

[0167] A generation module 25, configured to combine the high-precision three-dimensional motion trajectory with a preset multi-modal action pattern library, adopt an adaptive learning mechanism to perform intelligent matching and prediction on the motion state of the dynamic object, and calculate the influence of environmental factors at the same time to generate a frame capture result.

[0168] Figure 2 The described frame dynamic capture system based on a tracking system can execute Figure 1A method for dynamically capturing images based on a tracking system as described in the illustrated embodiments, the implementation principle and technical effects of which will not be elaborated further. For a system for dynamically capturing images based on a tracking system in the above embodiments, the specific ways in which each module and unit perform operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0169] Figure 2 A system for dynamically capturing images based on a tracking system in the illustrated embodiments can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;

[0170] The storage component 31 stores one or more computer instructions, where the one or more computer instructions are called and executed by the processing component 32.

[0171] The processing component 32 is used to perform real-time transmission on the multi-view video stream received by using a multi-sensor fusion tracking system, and collect external environment information to obtain a multi-view video stream, where the external environment information includes: light, sound, and temperature;

[0172] Perform frame synchronization and spatial alignment processing on the multi-view video stream by using a preset time synchronization signal and spatial calibration algorithm to obtain video frames with consistent timing and space;

[0173] Perform multi-modal analysis on the video frames by using a scene understanding algorithm driven by deep learning, automatically identify the key motion points of dynamic objects and the interaction features of the environment around the dynamic objects to obtain key motion point position data containing interaction information;

[0174] Perform high-level mapping processing on the key motion point position data by using a three-dimensional reconstruction technology based on a deep neural network to obtain a high-precision three-dimensional motion trajectory of the dynamic object;

[0175] Combine the high-precision three-dimensional motion trajectory with a preset multi-modal action pattern library, and adopt an adaptive learning mechanism to perform intelligent matching and prediction on the motion state of the dynamic object, and calculate the influence of environmental factors at the same time to generate an image capture result.

[0176] Among them, the processing component 32 includes one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component can also be implemented by one or more application-specific integrated circuits (AICs), digital signal processors (DPs), digital signal processing devices (DPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.

[0177] The storage component 31 is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc.

[0178] The computing device further includes other components, such as an input / output interface, a display component, and a communication component.

[0179] The input / output interface provides an interface between the processing component and the peripheral interface module, and the peripheral interface module can be an output device or an input device.

[0180] The communication component is configured to facilitate communication between the computing device and other devices in a wired or wireless manner, etc.

[0181] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. At this time, the computing device can refer to a cloud server, and the above-mentioned processing component, storage component, etc. can be basic server resources leased or purchased from a cloud computing platform.

[0182] An embodiment of the present invention further provides a computer storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the above-mentioned Figure 1 A method and system for dynamic capture of a screen based on a tracking system shown in the embodiment.

[0183] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0184] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0185] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0186] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for capturing dynamic images based on a tracking system, characterized in that: include: The multi-view video stream is transmitted in real time by using a multi-sensor fusion tracking system, and external environment information is collected to obtain a multi-view video stream, wherein the external environment information includes: light, sound and temperature; Using a preset time synchronization signal and a spatial calibration algorithm to perform frame synchronization and spatial alignment processing on the multi-view video stream, so as to obtain video frames with consistent timing and space; Performing multimodal analysis on the video frames using a deep learning-driven scene understanding algorithm to automatically identify key motion points of dynamic objects and interactive features of the environment surrounding the dynamic objects, and obtaining key motion point position data containing interactive information; Using a three-dimensional reconstruction technology based on a deep neural network to perform advanced mapping processing on the key motion point position data, to obtain a high-precision three-dimensional motion trajectory of the dynamic object; Combining the high-precision three-dimensional motion trajectory with the preset multimodal action pattern library, an adaptive learning mechanism is used to intelligently match and predict the motion state of the dynamic object, while calculating the influence of environmental factors to generate a picture capture result.

2. The method according to claim 1, characterized in that The key motion point position data is subjected to advanced mapping processing using a 3D reconstruction technology based on a deep neural network to obtain a high-precision 3D motion trajectory of the dynamic object, including: Performing multi-view geometric correction on the key motion point position data using a deep neural network model to obtain corrected key motion point position data; Performing environmental perception optimization on the corrected key motion point position data according to the external environment information to obtain enhanced key motion point position data to adjust the three-dimensional reconstruction parameters of the key motion points under different lighting, sound and temperature conditions; Using time series analysis technology to smooth the enhanced key motion point position data to obtain a dynamic object motion trajectory; The multi-scale feature fusion technology is used to integrate the multi-scale features of the motion trajectory of the dynamic object to generate a high-precision three-dimensional motion trajectory.

3. The method according to claim 2, characterized in that The enhanced key motion point position data is smoothed using time series analysis technology to obtain the motion trajectory of the dynamic object, including: Using time series analysis technology to perform dual smoothing processing in the time domain and frequency domain on the enhanced key motion point position data, dynamically adjusting the filter parameters through an adaptive filtering algorithm to obtain an initial motion trajectory; The initial motion trajectory is optimized by using a Kalman filter or a particle filter, and an optimized motion trajectory is generated by combining historical motion data of the dynamic object and a preset multimodal motion pattern library; Using a machine learning model to perform anomaly detection on the optimized motion trajectory, identify and correct abnormal points caused by external interference or sensor errors, and obtain a clear motion trajectory; The clear motion trajectory and external environment information are combined, and the motion parameters of the dynamic object are adjusted in real time using an environment-aware dynamic adjustment algorithm to generate a motion trajectory of the dynamic object.

4. The method according to claim 3, characterized in that The optimized motion trajectory is detected using a machine learning model to identify and correct abnormal points caused by external interference or sensor errors to obtain a clear motion trajectory, including: Using a preset deep learning model to monitor the optimized motion trajectory in real time, detecting abnormal points in the trajectory, and obtaining preliminary identified abnormal points; Extracting time series features, spatial features and environmental features from the optimized motion trajectory through multi-dimensional feature extraction technology, comparing them with historical normal motion data, identifying abnormal points caused by external interference, and obtaining accurately identified abnormal points; Using multimodal data fusion technology, combining visual information, sound information and temperature information in multi-view video streams to verify the accurately identified abnormal points from multiple angles, and obtain verified abnormal points; Using an adaptive interpolation algorithm to correct the verified abnormal points, restore the normal motion characteristics of the motion trajectory, and obtain a corrected motion trajectory; Combining the corrected motion trajectory with the motion state of the dynamic object, an adaptive smoothing algorithm is used to optimize the smoothness and continuity of the corrected motion trajectory to generate the final clear motion trajectory.

5. The method according to claim 4, characterized in that The verified abnormal points are corrected by using an adaptive interpolation algorithm to restore the normal motion characteristics of the motion trajectory to obtain a corrected motion trajectory, including: Using an adaptive interpolation algorithm, the interpolation parameters are dynamically adjusted according to the time series data before and after the verified abnormal point, and the abnormal point is accurately interpolated in combination with the local motion trend and the global motion pattern to restore the normal motion characteristics at the abnormal point and obtain a preliminary corrected motion trajectory; The motion trajectory of the preliminary correction is optimized again by using a weighted average method based on neighboring points combined with the motion data and environmental information of multiple time points before and after the verified abnormal point, and the multimodal data in the time window is introduced for comprehensive correction to obtain an optimized corrected motion trajectory, wherein the motion data includes: visual information, sound information and temperature information; The motion model of the dynamic object and the environment perception algorithm are used to perform consistency check on the optimized corrected motion trajectory, so that the optimized motion trajectory is consistent with the motion state of the dynamic object and adapts to different environmental changes to generate the final target corrected motion trajectory.

6. The method according to claim 1, characterized in that A deep learning-driven scene understanding algorithm is used to perform multimodal analysis on the video frames, automatically identify key motion points of dynamic objects and interactive features of the surrounding environment of the dynamic objects, and obtain key motion point position data containing interactive information, including: Performing multimodal analysis on the video frames using a deep learning-driven scene understanding algorithm to extract visual features, sound features, temperature features, and motion features from multiple perspectives of dynamic objects to obtain initial multimodal feature data; Using a target detection and tracking algorithm in combination with the initial multimodal feature data, high-precision detection and tracking of dynamic objects is performed to identify key motion points of the dynamic objects and obtain initial position data of the key motion points; Using an environment perception algorithm combined with interactive features in the video frame to optimize the environmental adaptability of the initial position data of the key motion point, adjusting the position and posture of the key motion point under different environmental conditions, and obtaining the key motion point position data containing interactive information; Using context-aware technology to perform semantic understanding and logical reasoning on the key motion point position data containing the interactive information, and combining the historical motion trajectory of the dynamic object and the preset multimodal motion mode library to generate optimized key motion point position data; The spatiotemporal correlation analysis technology is used to perform spatiotemporal correlation modeling on the optimized key motion point position data, analyze the motion laws of dynamic objects at different time points and spatial positions, generate a spatiotemporal motion model of the dynamic object, and obtain the target key motion point position data.

7. The method according to claim 1, characterized in that Combining the high-precision three-dimensional motion trajectory with the preset multi-modal motion pattern library, an adaptive learning mechanism is used to intelligently match and predict the motion state of the dynamic object, while calculating the influence of environmental factors to generate a picture capture result, including: The high-precision three-dimensional motion trajectory is combined with a preset multi-modal action pattern library to perform multi-dimensional intelligent matching on the motion state of the dynamic object, calculate the spatiotemporal characteristics, speed and acceleration information of the motion trajectory, and obtain a preliminary motion state classification; Using an adaptive learning mechanism to dynamically adjust the parameters of the multi-dimensional intelligent matching algorithm according to the historical motion data and real-time motion trajectory of the dynamic object, and combining the deep learning model and the reinforcement learning algorithm to optimize the preliminary motion state classification to obtain an optimized motion state classification; Using the environment perception model in combination with the optimized motion state classification and external environment information, the influence of environmental factors on the motion state of the dynamic object is calculated to generate a motion state prediction result with strong environmental adaptability; The final picture capture result is generated by utilizing the motion state prediction result with strong environmental adaptability and combining the real-time motion trajectory of the dynamic object and multimodal environmental data.

8. A picture dynamic capture system based on a tracking system, characterized in that: include: A transmission module, used to transmit the received multi-view video stream in real time by using a multi-sensor fusion tracking system, and collect external environment information to obtain a multi-view video stream, wherein the external environment information includes: light, sound and temperature; A processing module, used to perform frame synchronization and spatial alignment processing on the multi-view video stream using a preset time synchronization signal and a spatial calibration algorithm to obtain video frames with consistent timing and space; An analysis module, configured to perform multimodal analysis on the video frames using a deep learning-driven scene understanding algorithm, automatically identify key motion points of dynamic objects and interactive features of the surrounding environment of the dynamic objects, and obtain key motion point position data containing interactive information; A mapping module, used to perform advanced mapping processing on the key motion point position data using a three-dimensional reconstruction technology based on a deep neural network to obtain a high-precision three-dimensional motion trajectory of the dynamic object; A generation module is used to combine the high-precision three-dimensional motion trajectory with a preset multimodal action pattern library, use an adaptive learning mechanism to intelligently match and predict the motion state of the dynamic object, and calculate the influence of environmental factors to generate a picture capture result.

9. A computing device, characterized in that It comprises a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a method for capturing dynamic images based on a tracking system as described in any one of claims 1 to 7.

10. A computer storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a computer, a method for capturing dynamic images based on a tracking system as claimed in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Construction personnel warning protection system based on attitude risk degree analysis

    CN117711131A

  • Animation character generation system and method based on virtual reality

    CN118429494A

  • Volleyball motion trail extraction method based on video data

    CN118470068A

  • Athlete throwing action analysis and training method and equipment based on deep learning

    CN119068558A

Cited By

  • Intelligent sports live broadcast data management system and method

    CN120416528A