Video data processing method and system based on digital twinborn scene
The method enhances video data processing by extracting spatiotemporal features and mapping dynamic elements to 3D models, addressing inefficiencies in complex scenes and improving analysis precision and prediction.
Patent Information
- Application Number
- CN202510790217.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-13
AI Technical Summary
The prior art is difficult to accurately identify and predict static information and object motion states in video data in complex scenarios, especially in a combination of static information such as light and object motion states, video data processing efficiency is low.
By using video data processing methods based on digital twin scenarios, video data is collected and decomposed into video frame sequences, spatial and temporal feature extraction is performed, spatial and temporal feature matrix is generated, video elements are segmented, object motion paths are fused, dynamic elements are located in real time, mapped to 3D space, and model parameters are adjusted to generate predicted video frame sequences.
It improves the accuracy and efficiency of video data analysis, enhances dynamic tracking and prediction capabilities, optimizes digital twin applications, and promotes the application of smart cities, industrial maintenance, security monitoring and other fields.
Smart Images

Figure CN120321433A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of digital twin technology, and in particular, to a method and system for video data processing based on a digital twin scenario. Background Art
[0002] In the digital age, "Digital Twin", as an emerging technological concept, is profoundly changing the way data is processed and decisions are made in all walks of life, especially in the field of video data analysis. Digital Twin refers to creating a virtual replica that synchronizes with a physical entity by integrating digital representations of physical entities or systems, combined with sensor data, historical records, real-time monitoring, prediction models, and advanced analytics. In the field of video data processing, the application of this concept focuses on constructing highly accurate virtual environment models to achieve in-depth insight into video content, predictive analysis, and optimized decision support.
[0003] Currently, video data processing solutions usually use AI analysis technology to automatically identify objects and activities in the video frame and create a simplified "dynamic storyboard" in time and space to help quickly understand video information.
[0004] However, it is still difficult to accurately identify and predict in the above-mentioned solutions in complex scenarios, especially in the combined scenario of static information (such as light) and the motion state of objects, it is very difficult to effectively process video data. Summary of the Invention
[0005] Embodiments of the present application provide a method and system for video data processing based on a digital twin scenario to solve the problem of poor video data processing efficiency in the prior art.
[0006] In a first aspect, embodiments of the present application provide a method for video data processing based on a digital twin scenario, which is characterized by including: Collect video data from the actual environment, decompose the video data into a video frame sequence, and extract spatio-temporal features from the video frame sequence to determine a first spatio-temporal feature matrix. Each row of the first spatio-temporal feature matrix includes a video frame in the video frame sequence, and each column includes the object motion path, the main body behavior pattern, and the scene information corresponding to the video frame. The object motion path, the main body behavior pattern, and the scene information corresponding to the video frame are determined by spatio-temporal feature extraction; Based on the first spatio-temporal feature matrix, segment a plurality of video elements from the video frame sequence and generate an element segmentation mask corresponding to each video element; Fuse the element segmentation mask with the object motion paths corresponding to multiple video frames in the first spatio-temporal feature matrix to generate a second spatio-temporal feature matrix, where each row of the second spatio-temporal feature matrix represents a time step, and each column contains the segmentation mask features and motion features of all objects within the time step; Based on the second spatio-temporal feature matrix, dynamically locate and track the dynamic elements in the video frame sequence in real time, and obtain the 2D motion trajectories corresponding to the dynamic elements, where the dynamic elements are video elements with dynamic features in the video elements; Map the 2D motion trajectories of the dynamic elements into 3D space to create or update the target virtual 3D scene model, and input the first spatio-temporal feature matrix and the dynamic elements into the target virtual 3D scene model to generate a predicted video frame sequence through the target virtual 3D scene model; Compare the predicted video frame sequence with the video frame sequence decomposed from the video data frame by frame, and based on the comparison results, adjust the model parameters in the target virtual 3D scene model, where the model parameters at least include camera parameters or rendering parameters.
[0007] Optionally, the spatio-temporal feature extraction of the video frame sequence to determine the first spatio-temporal feature matrix includes: Identify all objects in the video frame sequence, and continuously track the positions of each object in different frames to obtain the object motion paths corresponding to the video frames; Identify the action information of all subjects in the video frame sequence to determine the subject behavior patterns of each subject; Identify the scene change information that changes over time in the video frame sequence, as well as the static information in the video frame sequence, and form scene information according to the scene change information and the static information; Wherein, the static information at least includes: scene structure, object layout, and lighting conditions; Convert the object motion paths, subject behavior patterns, and scene information corresponding to multiple video frames into a preset numerical form to constitute the first spatio-temporal feature matrix.
[0008] Optionally, based on the first spatio-temporal feature matrix, segment multiple video elements from the video frame sequence and generate an element segmentation mask corresponding to each video element, including: Input the first spatio-temporal feature matrix into a pre-established semantic segmentation model to predict each video frame in the first spatio-temporal feature matrix through the semantic segmentation model and output a segmentation result. The segmentation result includes a plurality of video elements and an element label corresponding to each video element. The element label is used to indicate whether the video element belongs to the foreground, background, or a specific object category in the video frame. The semantic segmentation model is trained based on a plurality of video frame samples with element labels; Generate an element segmentation mask corresponding to each video element. The element segmentation mask exists in the form of a binary image, and the binary image form includes black and white. Among them, white represents the area where the video element is located, and black represents the background area outside the video element.
[0009] Optionally, the fusing the element segmentation mask with the object motion paths corresponding to multiple video frames in the first spatio-temporal feature matrix to generate a second spatio-temporal feature matrix includes: Extract the object motion paths corresponding to multiple video frames from the first spatio-temporal feature matrix, and the object motion path corresponding to each video frame has the same time step as the corresponding element segmentation mask; Convert each element segmentation mask into a segmentation mask feature, and convert the object motion path corresponding to each video frame into a motion feature; For each video frame, splice the segmentation mask feature of the corresponding object with the motion feature to form a comprehensive feature vector; Arrange the comprehensive feature vectors corresponding to each video frame in sequence according to the time step and the number of objects to generate a second spatio-temporal feature matrix.
[0010] Optionally, based on the second spatio-temporal feature matrix, real-time locate and track the dynamic elements in the video frame sequence and obtain the 2D motion trajectory corresponding to the dynamic elements. The dynamic elements are the video elements with dynamic features among the video elements, including: Based on the second spatio-temporal feature matrix, identify the area in the video frame sequence that has a significant motion difference from the background area outside the video element, and determine the area as the area where the dynamic element is located; In the area where the dynamic element is located, identify the category and bounding box of the dynamic element, and for each dynamic element, establish the motion trajectory of the dynamic element among the video frames based on the category and bounding box of the dynamic element to generate the 2D motion trajectory corresponding to the dynamic element.
[0011] Optionally, mapping the 2D motion trajectory of the dynamic element into 3D space to create or update the target virtual 3D scene model, and inputting the first spatio-temporal feature matrix and the dynamic element into the target virtual 3D scene model to generate a predicted video frame sequence through the target virtual 3D scene model, includes: Extracting 2D feature points on the 2D motion trajectory corresponding to the dynamic element, and matching the 2D feature points between different video frames to determine the matched 2D feature point pairs; Using the matched 2D feature point pairs to calculate the corresponding points of the dynamic element in 3D space to generate the 3D motion trajectory of the dynamic element; Based on the subject behavior pattern and the scene information in the first spatio-temporal feature matrix, constructing an initial virtual 3D scene model, and integrating the 3D motion trajectory of the dynamic element into the initial virtual 3D scene model to generate a target virtual 3D scene model, wherein the dynamic element can move at the correct time and spatial position in the virtual 3D scene model; According to the target virtual 3D scene model and the position information of the obtained dynamic element in each video frame, rendering new video frames under preset camera parameters to form a predicted video frame sequence.
[0012] Optionally, comparing the predicted video frame sequence frame by frame with the video frame sequence decomposed from the video data, and based on the comparison result, adjusting the model parameters in the target virtual 3D scene model, includes: Calculating a difference loss function between the predicted video frame sequence and the video frame sequence decomposed from the video data, where the difference loss function contains a set of model parameters, and the set of model parameters at least includes: camera parameters or rendering parameters.
[0013] Optionally, calculating the difference loss function between the predicted video frame sequence and the video frame sequence decomposed from the video data, includes: Through the formula group: , calculating the difference loss function between the predicted video frame sequence and the video frame sequence decomposed from the video data; Wherein, represents a pixel-level loss function, represents the predicted video frame corresponding to the frame with the time stamp of , represents the video frame sequence decomposed from the video data, It is expressed as a structural similarity index, which is used to measure the structural similarity between the predicted video frame sequence and the video frame sequence decomposed from the video data. Its value range is between -1 and 1, where 1 indicates complete identity and 0 indicates no correlation, and it is calculated through the SSIM algorithm; It is expressed as a hyperparameter, which is used to balance the weights of the SSIM loss and the L1 norm loss; It is expressed as the L1 norm loss, which is used for the sum of the absolute errors between the predicted video frame sequence and the video frame sequence decomposed from the video data; Among them, It is expressed as a geometric consistency loss function, It is expressed as the number of sampling points in the real 3D scene model, It is expressed as the number of sampling points in the target virtual 3D scene model, It is expressed as the real 3D scene model, It is expressed as the target virtual 3D scene model, It is expressed as calculating the point set for each point in the point set to the nearest point in the point set and the sum of the squares of the distances, for each point in the point set to the nearest point in the point set and the sum of the squares of the distances, It is expressed as a hyperparameter, which is used to adjust the weight of the normal consistency loss in the geometric consistency loss It is expressed as the normal consistency loss, which is used to measure and the difference between them; Among them, It is expressed as a temporal consistency loss function, It is expressed as the predicted video frame corresponding to the frame with the timestamp of It is expressed as the predicted video frame corresponding to the frame with the timestamp of It is expressed as the optical flow error, which is used to measure the consistency of pixel motion between two predicted video frames, It is expressed as the penalty coefficient of the optical flow error, It is expressed as the optical smoothness loss, which is used to control the smooth transition of the video frame in terms of timestamp, It is expressed as a hyperparameter, which is used to balance the influence of the optical smoothness loss and the optical flow error on the temporal consistency loss; Among them, It is expressed as a difference loss function, represented as a set of model parameters, represented as the weights of the pixel-level loss function, represented as the weights of the geometric consistency loss function, represented as the weights of the temporal consistency loss function.
[0014] In a second aspect, an embodiment of the present application provides a video data processing system based on a digital twin scenario, including: An acquisition and processing module, configured to acquire video data from an actual environment, decompose the video data into a video frame sequence, and perform spatio-temporal feature extraction on the video frame sequence to determine a first spatio-temporal feature matrix, where each row of the first spatio-temporal feature matrix includes a video frame in the video frame sequence, and each column includes the object motion path, the main body behavior pattern, and the scene information corresponding to the video frame, and the object motion path, the main body behavior pattern, and the scene information corresponding to the video frame are determined through spatio-temporal feature extraction; A segmentation module, configured to segment a plurality of video elements from the video frame sequence based on the first spatio-temporal feature matrix, and generate an element segmentation mask corresponding to each video element; A fusion module, configured to fuse the element segmentation mask with the object motion paths corresponding to a plurality of the video frames in the first spatio-temporal feature matrix to generate a second spatio-temporal feature matrix, where each row of the second spatio-temporal feature matrix represents a time step, and each column includes the segmentation mask features and motion features of all objects within the time step; A positioning module, configured to based on the second spatio-temporal feature matrix, real-time locate and track dynamic elements in the video frame sequence, and obtain 2D motion trajectories corresponding to the dynamic elements, where the dynamic elements are video elements with dynamic features among the video elements; A generation module, configured to map the 2D motion trajectories of the dynamic elements into a 3D space to create or update a target virtual 3D scene model, and input the first spatio-temporal feature matrix and the dynamic elements into the target virtual 3D scene model to generate a predicted video frame sequence through the target virtual 3D scene model; An adjustment module, configured to compare the predicted video frame sequence with the video frame sequence decomposed from the video data frame by frame, and based on the comparison result, adjust the model parameters in the target virtual 3D scene model, where the model parameters at least include object motion parameters, camera parameters, or rendering parameters.
[0015] In a third aspect, an embodiment of the present application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a video data processing method based on a digital twin scenario as described in the first aspect above.
[0016] In an embodiment of the present application, video data is collected from the actual environment, the video data is decomposed into a video frame sequence, and spatio-temporal features are extracted from the video frame sequence to determine a first spatio-temporal feature matrix. Each row of the first spatio-temporal feature matrix includes a video frame in the video frame sequence, and each column includes the object motion path, the main body behavior pattern, and the scene information corresponding to the video frame. The object motion path, the main body behavior pattern, and the scene information corresponding to the video frame are determined by spatio-temporal feature extraction; based on the first spatio-temporal feature matrix, multiple video elements are segmented from the video frame sequence, and an element segmentation mask corresponding to each video element is generated; the element segmentation mask is fused with the object motion paths corresponding to multiple video frames in the first spatio-temporal feature matrix to generate a second spatio-temporal feature matrix. Each row of the second spatio-temporal feature matrix represents a time step, and each column includes the segmentation mask features and motion features of all objects within the time step; based on the second spatio-temporal feature matrix, the dynamic elements in the video frame sequence are located and tracked in real time, and the 2D motion trajectories corresponding to the dynamic elements are obtained. The dynamic elements are the video elements with dynamic characteristics among the video elements; the 2D motion trajectories of the dynamic elements are mapped into a 3D space to create or update a target virtual 3D scene model, and the first spatio-temporal feature matrix and the dynamic elements are input into the target virtual 3D scene model to generate a predicted video frame sequence through the target virtual 3D scene model; the predicted video frame sequence is compared frame by frame with the video frame sequence decomposed from the video data, and based on the comparison result, the model parameters in the target virtual 3D scene model are adjusted. The model parameters at least include object motion parameters, camera parameters, or rendering parameters.
[0017] The beneficial effects of this method are mainly reflected in the following aspects: Improve analysis accuracy and efficiency: By performing fine spatio-temporal feature extraction on video data, not only can static information in the video be captured, but also dynamic changes such as object motion, behavior patterns, and scene changes can be deeply understood, thereby greatly improving the accuracy and processing speed of data analysis.
[0018] Enhance dynamic tracking and prediction capabilities: Real-time locate and track dynamic elements, generate 2D motion trajectories and map them to 3D space, which helps to construct a more realistic and dynamically responsive virtual environment. This is particularly effective for predicting future states and simulating "what-if" scenarios, enhancing the foresight and practicality of decision support systems.
[0019] Optimize digital twin applications: Continuously iterate and optimize the 3D scene model to ensure the consistency and synchronization between the virtual model and the physical world, promoting the wide application of digital twin technology in various fields such as industrial maintenance, urban management, and security monitoring, and improving the efficiency of simulation, fault prediction, and resource scheduling.
[0020] In summary, this method not only promotes the progress of video data analysis technology, but also shows significant positive effects at multiple levels such as promoting the construction of smart cities, optimizing production management, and enhancing user experience.
[0021] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0023] Figure 1 It is a flowchart of a video data processing method based on a digital twin scenario provided by an embodiment of the present application; Figure 2 It is a schematic structural diagram of a video data processing system based on a digital twin scenario provided by an embodiment of the present application; Figure 3 It is a schematic structural diagram of a computing device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] To enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application.
[0025] In some of the processes described in the specification, claims, and the above-mentioned drawings of this application, a number of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any order of execution. Additionally, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first", "second", etc. in this document are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types.
[0026] The technical solution of this application can be applied to a variety of scenarios that require in-depth analysis and simulation of real-world dynamics, specifically including but not limited to the following aspects: Intelligent security monitoring: In the monitoring environment, it can automatically identify abnormal behaviors and track the movement trajectories of specific people or vehicles, improving security warnings and the speed of event response.
[0027] Smart city management: For example, traffic flow analysis, by analyzing the flow patterns of vehicles and pedestrians, optimizing signal control strategies, reducing congestion, and improving urban traffic efficiency.
[0028] Industrial production and maintenance: In a digital factory, monitor the operating status of equipment, predict maintenance needs, analyze bottlenecks in the production process, and achieve a double improvement in production efficiency and safety.
[0029] Sports event analysis: Real-time capture of the detailed movements of athletes, analysis of game strategies, evaluation of athlete performance, providing instant feedback to the coaching team, and can also be used to enhance the audience experience.
[0030] Film and television special effects and animation production: By capturing the dynamic characteristics of the real world, quickly generate highly realistic virtual scenes and character animations, reducing production costs and improving creative efficiency.
[0031] Disaster emergency response: At the scene of natural disasters or emergencies, quickly identify the affected situation and personnel distribution, assist in formulating rescue plans, and improve the accuracy and timeliness of rescue operations.
[0032] Retail and consumer behavior research: Analyze the walking paths and staying areas of customers in commercial venues, optimize product layouts and promotional strategies, and enhance the customer experience and sales performance.
[0033] The common feature of these scenarios is that they all require in-depth analysis and processing of video data to extract key information and perform predictions or simulations, and the above method just provides a full-process solution from extracting spatio-temporal features from video data, real-time tracking of dynamic elements to constructing and optimizing digital twin models.
[0034] The inventors' research found that in current video data processing solutions, AI analysis technology is usually used to automatically identify objects and activities in the video frame, and create a simplified "dynamic storyboard" in time and space to help quickly understand video information.
[0035] However, it is still difficult to accurately identify and predict in complex scenarios, especially in the combined scenario of static information (such as light) and the object motion state. It is very difficult to effectively process video data, such as the extraction and processing of dynamic elements.
[0036] In view of this, the present application provides a video data processing method based on a digital twin scenario. The method includes: collecting video data from the actual environment, decomposing the video data into a video frame sequence, and extracting spatio-temporal features from the video frame sequence to determine a first spatio-temporal feature matrix. Each row of the first spatio-temporal feature matrix includes a video frame in the video frame sequence, and each column includes the object motion path, the main body behavior pattern, and the scene information corresponding to the video frame. The object motion path, the main body behavior pattern, and the scene information corresponding to the video frame are determined by spatio-temporal feature extraction; based on the first spatio-temporal feature matrix, multiple video elements are segmented from the video frame sequence, and an element segmentation mask corresponding to each video element is generated; the element segmentation mask is fused with the object motion paths corresponding to multiple video frames in the first spatio-temporal feature matrix to generate a second spatio-temporal feature matrix. Each row of the second spatio-temporal feature matrix represents a time step, and each column includes the segmentation mask features and motion features of all objects within the time step; based on the second spatio-temporal feature matrix, the dynamic elements in the video frame sequence are located and tracked in real time, and the 2D motion trajectory corresponding to the dynamic elements is obtained. The dynamic elements are the video elements with dynamic features in the video elements; the 2D motion trajectory of the dynamic elements is mapped into a 3D space to create or update a target virtual 3D scene model, and the first spatio-temporal feature matrix and the dynamic elements are input into the target virtual 3D scene model to generate a predicted video frame sequence through the target virtual 3D scene model; the predicted video frame sequence is compared frame by frame with the video frame sequence decomposed from the video data, and based on the comparison result, the model parameters in the target virtual 3D scene model are adjusted. The model parameters at least include object motion parameters, camera parameters, or rendering parameters.
[0037] By performing fine spatio-temporal feature extraction on video data, the above method can not only capture static information in the video, but also deeply understand dynamic changes such as object movement, behavior patterns, and scene changes, thereby greatly improving the accuracy and processing speed of data analysis. In addition, by real-time positioning and tracking dynamic elements, generating 2D motion trajectories and mapping them to 3D space, the above method helps to construct a more realistic and dynamically responsive virtual environment, which is particularly effective for predicting future states and simulating "what-if" scenarios, enhancing the forward-looking and practicality of the decision support system.
[0038] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0039] Figure 1 The flowchart of a video data processing method based on a digital twin scenario provided by an embodiment of the present application is as follows Figure 1 As shown, the method includes: 101. Collect video data from the actual environment, decompose the video data into a video frame sequence, and perform spatio-temporal feature extraction on the video frame sequence to determine a first spatio-temporal feature matrix; In this step, each row of the first spatio-temporal feature matrix includes a video frame in the video frame sequence, and each column includes the object movement path, the main behavior pattern, and the scene information corresponding to the video frame, which are determined by spatio-temporal feature extraction.
[0040] Specifically, first, continuous video data streams are collected from the actual environment through devices such as high-definition cameras. Subsequently, the continuous video stream is decomposed into a sequence of frame-by-frame images, that is, a video frame sequence, using video processing software or algorithms.
[0041] For spatio-temporal feature extraction, this process involves complex computer vision techniques and deep learning models, aiming to extract meaningful information from each frame of the image. The object movement path is obtained through optical flow estimation or tracking algorithms, analyzing the change of pixel points over time to determine the moving direction and speed of the object. The recognition of the main body behavior pattern uses an action recognition model, classifying or predicting the behavior type by analyzing the human body (or other main bodies) postures and action sequences. The extraction of scene information involves scene understanding, including identifying objects in the background, lighting conditions, and overall layout, which is usually completed by a scene parsing model. These features together constitute a spatio-temporal feature matrix, where each row corresponds to the feature profile of a certain frame, and each column is refined to specific feature dimensions, such as motion, behavior, and scene details.
[0042] In the embodiments of this application, it is assumed that a video processing module is being developed for an intelligent transportation system. Starting from the high-definition video stream collected by the surveillance cameras installed at urban intersections, the video data per second is decomposed into a sequence of 30 video frames. A combined model of Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) is used for spatio-temporal feature extraction. CNN is responsible for extracting visual features such as object contours and textures from each frame, while RNN tracks the changes of these features over time to capture the movement path. For the recognition of the main body behavior pattern, especially for pedestrians and vehicles, a pre-trained behavior recognition model can be used, such as the I3D network, which can predict the driving direction, acceleration or braking behavior of vehicles, and the crossing pattern of pedestrians, etc. based on video segments. The scene information is obtained through a scene classification model, such as using the ResNet architecture to classify the background of each video frame, distinguishing elements such as streets, sidewalks, and buildings. In this way, after each frame of video data is processed, it is integrated into the spatio-temporal feature matrix. Each row records the comprehensive information at a certain moment, and the columns are refined to specific descriptions of object movement (such as the right-turn path of vehicle ID_001), behavior (pedestrian_002 waiting to cross the street), and scene (night, rainy day), providing a rich data basis for subsequent analysis and decision-making.
[0043] 102. Based on the first spatio-temporal feature matrix, segment multiple video elements from the video frame sequence, and generate an element segmentation mask corresponding to each video element; In this step, first, a deep learning model, specifically a Convolutional Neural Network (CNN) or a Recurrent Neural Network (RNN) based on time series analysis, is used to analyze the input video frame sequence. These models can learn and extract spatio-temporal features in the video, forming a high-dimensional feature matrix, namely the first spatio-temporal feature matrix. This process involves multiple layers of convolutional operations and pooling operations (for CNN), or recursive processing of temporal information (for RNN), aiming to capture static visual features and dynamic change patterns in the video. Subsequently, these features are used to identify and locate different video elements in the video, such as objects, people, actions, etc. The generation of the element segmentation mask is achieved through relatively complex post-processing algorithms, such as Conditional Random Field (CRF) or graph cut-based methods, which accurately define the boundaries of each video element according to the response intensity in the feature matrix, thereby realizing the effective segmentation of video content.
[0044] In the embodiment of this application, a pre-trained 3D CNN model is used to process the video frame sequence. This model can simultaneously consider information in both spatial and temporal dimensions. First, the video frames are adjusted to a unified size and normalized, and then input into the 3D CNN for forward propagation to obtain a first spatio-temporal feature matrix with rich spatio-temporal information.
[0045] Next, a Mask R-CNN model is applied to predict the segmentation mask of each video element. It not only outputs the positions of objects in each video frame but also provides pixel-level segmentation masks additionally. For example, in the analysis of a sports event video, the system can accurately segment different elements such as athletes, balls, and fields from a complex scene and generate fine segmentation masks for each type of video element, making subsequent video editing, analysis, or content understanding tasks more efficient and accurate.
[0046] 103. Fuse the element segmentation mask with the object motion paths corresponding to multiple video frames in the first spatio-temporal feature matrix to generate a second spatio-temporal feature matrix. Each row of the second spatio-temporal feature matrix represents a time step, and each column contains the segmentation mask features and motion features of all objects within the time step. In this step, it is first necessary to parse the motion trajectory of the object from the first spatio-temporal feature matrix. This usually involves analyzing the position changes of the object between consecutive frames through tracking algorithms such as optical flow method, Kalman filter or deep learning tracking network to determine the motion path of the object. The motion characteristics of the object can include position change, speed, acceleration, etc. Subsequently, these motion characteristics are combined with the segmentation masks of each video element generated previously. The combination method is to directly splice the segmentation mask features and the motion feature vectors, or design a more complex fusion mechanism, such as using an attention mechanism to emphasize the features of key motion regions. After such processing, a new feature representation - the second spatio-temporal feature matrix is formed, which not only retains the static appearance information of the object at each time step, but also incorporates the description of dynamic behavior, providing a more comprehensive feature basis for subsequent high-level analysis or recognition tasks.
[0047] In the embodiment of the present application, the DeepSORT algorithm is used to track the objects in the video and extract their motion paths. DeepSORT combines appearance feature matching and motion model prediction to ensure the stability and accuracy of tracking. For each object at each time step, we extract its segmentation mask in the first spatio-temporal feature matrix and record its motion parameters. Then, a feature vector is constructed for each object, where the first half is the feature encoding of the segmentation mask and the second half is the motion feature of the object. These vectors are arranged in chronological order to form the second spatio-temporal feature matrix. For example, in a surveillance video analysis scenario, the system not only identifies objects such as pedestrians and vehicles, but also records their movement trajectories and speed changes, and all this information is integrated in the second spatio-temporal feature matrix, providing detailed data support for further analysis of pedestrian behavior patterns or vehicle flow management.
[0048] 104. Based on the second spatio-temporal feature matrix, real-time locate and track the dynamic elements in the video frame sequence, and obtain the 2D motion trajectories corresponding to the dynamic elements, where the dynamic elements are the video elements with dynamic features among the video elements; In this step, the goal is to identify and track the dynamic elements in the video in real time from the already constructed second spatio-temporal feature matrix, and then obtain the two-dimensional motion trajectories of these elements. This step mainly includes feature decoding, motion detection, and trajectory generation. First, using the object segmentation masks and motion features encoded in the matrix, decoded by machine learning or deep learning models, the presence and categories of dynamic elements in each frame are identified. Then, by applying trackers based on correlation filtering, deep learning-based tracking networks (such as Siamese Tracker, DeepSORT, etc.), or optical flow estimation methods, according to the appearance and position changes of objects between consecutive frames, their motion trajectories in the two-dimensional space are calculated and updated. In this way, the motion paths of each dynamic element are gradually constructed, providing intuitive motion information for subsequent analyses such as behavior recognition and interaction detection.
[0049] In the embodiments of this application, an improved deep learning tracking framework, namely the Tracking Transformer (TranSTR) model based on Transformer, is adopted to implement this process. The TranSTR model can effectively integrate temporal and spatial information, improving the robustness of tracking in complex scenarios. Specifically, during operation, first, the second spatio-temporal feature matrix is parsed to extract the feature maps of each time step and input into the TranSTR model. The model uses the self-attention mechanism to correlate object features across temporal and spatial dimensions, thereby accurately locating and predicting the positions of each dynamic element in the next frame. For example, in the scenario of sports event video analysis, the system can track the movements of athletes and balls in real time, accurately capturing their 2D motion trajectories even in cases of fast movement and frequent occlusion, which is of great value for analyzing game strategies, athlete performance, etc. By continuously iterating this process, the system continuously updates the motion trajectories of each dynamic element until the end of the video sequence.
[0050] 105. Map the 2D motion trajectories of the dynamic elements into the 3D space to create or update the target virtual 3D scene model, and input the first spatio-temporal feature matrix and the dynamic elements into the target virtual 3D scene model to generate a predicted video frame sequence through the target virtual 3D scene model; In this step, the motion trajectories of dynamic elements extracted from the 2D video are transformed into the 3D space, which is achieved by a technique called "2D-to-3D lifting". This technique typically involves using depth estimation, multi-view geometry, or pre-established 3D models to infer the three-dimensional positions and poses of objects in the video. Once the 2D trajectories are obtained, they are mapped to the 3D space by combining the depth information of the scene or applying kinematic principles to form the 3D motion trajectories of the objects. Subsequently, these 3D trajectories are used to create or update a virtual 3D scene model, which includes the static background in the video and the dynamic representations of the dynamic elements. The first spatio-temporal feature matrix and the dynamic element information are used as inputs to help the model understand the spatio-temporal context of the video, thereby more accurately predicting the subsequent video frame sequence. This step is crucial for generating realistic video predictions, augmented reality content, or video content editing.
[0051] In the embodiments of this application, a method combining deep learning and geometric constraints is adopted to achieve the 2D-to-3D mapping. First, a monocular depth estimation network (such as Monodepth2) is used to estimate the depth value of each pixel from the video frames, which provides depth cues for the 2D motion trajectories. Then, algorithms such as Bundle Adjustment are used to combine the motion trajectories and depth information to optimize and obtain the 3D trajectories of the objects. For the construction of the target virtual 3D scene model, we choose the Unity 3D engine because it provides powerful physical simulation and rendering capabilities. We import the scene structure information decoded from the first spatio-temporal feature matrix and the 3D positions and poses of the dynamic elements into Unity to create the corresponding 3D objects and animations. Further, using the built-in time series prediction model in Unity or a custom recurrent neural network (RNN), combined with the current 3D scene state and the motion laws of the dynamic elements, a predicted video frame sequence is generated. For example, in a reconstruction project of an urban street scene, not only static elements such as streets and buildings are finely modeled, but also dynamic behaviors such as the walking of pedestrians and the driving of vehicles are accurately reproduced in the virtual 3D scene through the above process, and then the possible street scene changes in the next few frames are predicted.
[0052] 106. Compare the predicted video frame sequence with the video frame sequence decomposed from the video data frame by frame, and based on the comparison results, adjust the model parameters in the target virtual 3D scene model, where the model parameters at least include object motion parameters, camera parameters, or rendering parameters.
[0053] In this step, firstly, the corresponding video frame sequence needs to be extracted from the predicted video frame sequence and the original video data. The predicted video frame sequence is usually rendered by simulating parameters such as lighting, material, and camera motion by the established target virtual 3D scene model. The video frame sequence decomposed by the video data is each frame image directly extracted from the actual recorded or captured video. In order to perform frame-by-frame comparison, an image similarity measurement method, such as the structural similarity index (SSIM), the peak signal-to-noise ratio (PSNR), or a deep learning-driven image difference measurement model, can be used to evaluate the difference between the predicted frame and the actual frame.
[0054] After acquiring these frame sequences, they are preprocessed using image processing techniques, such as alignment, scaling, and color space conversion, to ensure consistency and accuracy of the comparison. Then, through pixel-level or feature-level comparisons, the deviations between the predicted frames and the actual frames in terms of content, color, texture, etc. can be analyzed. Based on these comparison results, it is possible to identify which aspects of the target virtual 3D scene model fail to accurately reflect the real scene, such as inaccurate object positions and motion trajectories, or improper camera viewing angles and exposure settings.
[0055] In the embodiment of this application, it is assumed that a deep learning-based 3D reconstruction system is being developed to restore the virtual scene of an ancient building. First, we use a neural network to predict a video sequence of the building, which includes the changes in light and shadow at different time periods during the day and the dynamics of tourists. Then, we obtain a real video record of the location in the same time period and decompose it into a single-frame image sequence.
[0056] During the implementation, SSIM was selected as the image quality evaluation index, and combined with optical flow estimation to analyze the coherence of object motion. By comparison, it was found that some frames in the predicted sequence had large deviations from the actual light and shadow effects, especially the insufficient expression of the golden glow at dusk. In addition, the virtual tourists in the video were slightly misaligned in their moving paths compared to the actual situation.
[0057] Based on these analysis results, the parameters in the target virtual 3D scene model were adjusted in a targeted manner: the color temperature adjustment of the light source was added to make the light at dusk softer and richer in warm tones; at the same time, the behavior logic and path planning in the crowd simulation algorithm were optimized to better match the actual crowd flow pattern. Finally, the predicted sequence was re-rendered and compared again until a satisfactory visual fit was achieved. This process reflects the practice of iteratively optimizing model parameters to improve the authenticity of video synthesis.
[0058] Optionally, in the embodiments of the present application, first, through deep learning and computer vision technologies, especially based on convolutional neural networks (CNNs) and recurrent neural networks (RNNs), a detailed analysis of the video frame sequence is achieved. This covers key aspects such as object recognition, motion path tracking, behavior patterns, and scene understanding. Therefore, the process of "extracting spatio-temporal features from the video frame sequence to determine the first spatio-temporal feature matrix" in step 101 may include: 1011. Identify all objects in the video frame sequence and continuously track the positions of each object in different frames to obtain the object motion paths corresponding to the video frames; In this step, an object detection model (such as YOLOD or FasterNet) is used to identify all objects in the video, and the movement of each object between video frames is tracked through optical flow estimation or sequence frame matching techniques (such as Kalman-Bucs filtering) to draw its motion path. This requires the model to have real-time recognition and prediction capabilities to ensure the stability of object identities between consecutive frames.
[0059] 1012. Identify the action information of all subjects in the video frame sequence to determine the subject behavior patterns of each subject; In this step, the dynamics of individuals or groups in the video sequence are analyzed through behavior recognition algorithms (such as based on 3D-CNN or LSTM), and their behaviors (walking, running, sitting, waving, etc.) are identified and classified. This requires understanding the action sequence in the time series, extracting the features of the behavior, and combining the temporal context for pattern recognition.
[0060] 1013. Identify the scene change information that changes over time in the video frame sequence, as well as the static information in the video frame sequence, and form scene information based on the scene change information and the static information; In this step, the static information at least includes: scene structure, object layout, and lighting conditions. Scene understanding is achieved through scene segmentation (such as Mask R-CNN) to identify static structures (buildings, roads, furniture layout) and lighting conditions (shadows, color temperature). At the same time, the scene changes are analyzed, and the dynamic elements (such as weather changes, people appearing and disappearing) over time are captured and combined with the static elements to form a comprehensive scene description.
[0061] 1014. Convert the object motion paths, subject behavior patterns, and scene information corresponding to multiple video frames into a preset numerical form to constitute the first spatio-temporal feature matrix.
[0062] In this step, the above information is encoded into numerical values. For example, the object movement path can be represented as velocity and acceleration vectors, the behavior pattern as a class label or behavior probability, scene changes, light intensity, etc., and organized into a matrix. Each row corresponds to a video frame, and the columns represent different feature dimensions (object movement, behavior, scene, etc.).
[0063] In the embodiment of the present application, it is assumed that this process is applied in an intelligent monitoring system. In the monitoring video stream, first, objects such as pedestrians and vehicles are identified using YOLODv3, and the pedestrian trajectories are tracked through the DeepSORT algorithm to record their position changes in different frames. At the same time, the behavior patterns of pedestrians such as "walking", "running", "standing", etc. are identified through 3D-CNN, and the "road", "trees", "buildings", and "daytime" lighting conditions in the monitoring area are identified through scene segmentation. Finally, this information is encoded into a feature matrix, with each frame corresponding to a row, and the features including the trajectory change speed of object A, the behavior label "walking", the scene label "road", and the light intensity. This matrix provides a basis for subsequent analysis and prediction, such as abnormal behavior detection and crowd flow pattern analysis.
[0064] Optionally, in the embodiment of the present application, the process of "segmenting a plurality of video elements from the video frame sequence based on the first spatio-temporal feature matrix and generating an element segmentation mask corresponding to each video element" in step 102 may include: 1021. Input the first spatio-temporal feature matrix into a pre-established semantic segmentation model to predict each video frame in the first spatio-temporal feature matrix through the semantic segmentation model and output a segmentation result. The segmentation result includes a plurality of video elements and an element label corresponding to each video element. The element label is used to indicate whether the video element belongs to the foreground, background, or a specific object category in the video frame. The semantic segmentation model is trained based on a plurality of video frame samples with element labels. In this step, first, by using the first spatio-temporal feature matrix as the input, the composition of the video frame is analyzed using semantic segmentation technology in deep learning. The semantic segmentation model is a deep neural network architecture, such as U-Net, DeepLabV3+, or Mask R-CNN, which is designed for pixel-level classification, that is, assigning a class label to each pixel in the image. During the training process of such a model, it learns how to distinguish the foreground from the background and identify specific object categories within the video frame, such as people, vehicles, buildings, etc. The training dataset contains a large number of labeled video frame samples, and each pixel on each sample image has a corresponding element label to guide the model to learn the features of each category. In the inference stage, based on the rich context information provided by the first spatio-temporal feature matrix, the model predicts the semantic attribution of each part in each video frame and outputs a high-precision segmentation result.
[0065] 1022. Generate an element segmentation mask corresponding to each of the video elements. The element segmentation mask exists in the form of a binary image, and the binary image form includes black and white. Among them, white represents the area where the video element is located, and black represents the background area outside the video element.
[0066] In this step, when actually generating the element segmentation mask, based on the output of the semantic segmentation model, a binary mask is created for each predicted video element. This process is essentially a binarization process of the segmentation result. Among them, the pixel points predicted by the model as video elements are assigned white (value 1), indicating that this part of the image belongs to a specific video element; while the pixel points predicted as the background or other non-element areas are assigned black (value 0). In this way, each video element corresponds to a clear binary image mask, which can intuitively display the spatial distribution of the element in the original video frame, facilitating subsequent video editing, analysis, or content enhancement operations.
[0067] In the embodiment of the present application, consider a sports event broadcast application scenario. From a video of a football game, we have obtained a first spatio-temporal feature matrix containing various spatio-temporal features such as players, balls, and field lines. Next, the DeepLabV3+ semantic segmentation model pre-trained on a large number of sports video samples is used to process the input of this feature matrix. The model performs fine segmentation on each frame of the picture, differentiates different elements such as players (with clothing of different team colors), footballs, spectator stands, and grasslands, and assigns correct labels to each element. Subsequently, according to the output of the model, a series of binary images are automatically generated as element segmentation masks. For example, a mask specifically highlights the positions of all players, with the white area representing the players and the black area being the other background. These masks can not only help achieve real-time game data analysis, such as player running trajectories and ball control area statistics, but also facilitate advanced editing tasks such as adding special effects or inserting advertisements during post-production.
[0068] Optionally, in the embodiment of the present application, the process of "fusing the element segmentation mask with the object motion paths corresponding to multiple video frames in the first spatio-temporal feature matrix to generate a second spatio-temporal feature matrix" in step 103 may include: 1031. Extract the object motion paths corresponding to multiple video frames from the first spatio-temporal feature matrix, and the object motion path corresponding to each video frame has the same time step as the corresponding element segmentation mask; In this step, first by analyzing the first spatio-temporal feature matrix, we extract the motion paths of each object in the video frame sequence over time. These motion paths reflect the dynamic position changes of the objects in the scene, ensuring their consistency with the element segmentation mask in the time dimension, that is, the object motion information at each time step corresponds to the mask at that moment.
[0069] 1032. Convert each of the element segmentation masks into segmentation mask features, and convert the motion paths of the objects corresponding to each video frame into motion features. In this step, the mask information is converted into a numerical feature representation, which usually involves encoding the structural information of the mask into a one-dimensional or higher-dimensional vector to reflect the spatial distribution state of the objects. At the same time, the extraction of the motion features of the objects includes the speed, acceleration or more complex trajectory characteristics of the objects to describe their dynamic behaviors.
[0070] 1033. For each video frame, concatenate the segmentation mask feature of the corresponding object with the motion feature to form a comprehensive feature vector. In this step, a comprehensive feature vector is constructed for each object in each video frame. This step integrates the static position of the object (through the segmentation mask feature) and the dynamic behavior (through the motion feature) to form a comprehensive description of the object in space-time.
[0071] 1034. Arrange the comprehensive feature vectors corresponding to each video frame in sequence according to the time step and the number of objects to generate a second spatio-temporal feature matrix.
[0072] In this step, the comprehensive feature vectors of all video frames are arranged in order of time and objects to construct a new spatio-temporal feature representation - the second spatio-temporal feature matrix. This matrix not only contains the current state of the objects, but also embeds the process of their evolution over time, providing a richer information basis for subsequent advanced analysis or processing. The dimensions of the second spatio-temporal feature matrix are: the number of time steps x (the number of objects x (the dimension of the segmentation mask feature + the dimension of the motion feature)).
[0073] In the embodiments of the present application, it is assumed that a segment of an action movie is being processed, which includes the protagonist and multiple supporting characters chasing him. In step 1031, through previous analysis, we have obtained the precise movement trajectories of the protagonist and each supporting character in consecutive frames. These trajectories record the movement of the characters in the scene in the form of a sequence of coordinates. At the same time, in step 1032, for each frame, the element segmentation masks of the protagonist and each supporting character are converted into feature vectors, representing their position and shape information in the frame; their movement paths are also converted into motion feature vectors, describing their speed and direction in the scene. In step 1033, taking a certain frame as an example, the segmentation mask feature of the protagonist (assumed to be 256-dimensional) is concatenated with the motion feature (assumed to be 6-dimensional) to form a 322-dimensional comprehensive feature vector. Similarly, each supporting character is processed. Finally, in step 1034, the comprehensive feature vectors of all characters in all frames of the entire movie are arranged in chronological order to form a large second spatio-temporal feature matrix, whose dimension is (total number of frames x (number of characters x (dimension of segmentation mask feature + dimension of motion feature))). Such a matrix is helpful for further analyzing the interaction patterns between characters, action strategy analysis, or automated generation of visual effects, etc.
[0074] Optionally, in the embodiments of the present application, the process of "locating and tracking the dynamic elements in the video frame sequence in real time based on the second spatio-temporal feature matrix and obtaining the corresponding 2D movement trajectories of the dynamic elements" in step 104 may include: 1041. Based on the second spatio-temporal feature matrix, identify the regions in the video frame sequence that have significant motion differences from the background regions other than the video elements, and determine these regions as the regions where the dynamic elements are located; In this step, first, utilize the high-dimensional information of the second spatio-temporal feature matrix to identify the significantly different regions by comparing the changes of each pixel or region over time with the changes of the background region. These regions usually correspond to the dynamic elements in the video, such as moving people, vehicles, etc. Technologies such as optical flow estimation, background subtraction, or motion segmentation can be used to effectively distinguish the dynamic elements from the static background.
[0075] 1042. In the regions where the dynamic elements are located, identify the categories and bounding boxes of the dynamic elements, and for each dynamic element, establish the movement trajectory of the dynamic element between the video frames based on the category and bounding box of the dynamic element to generate the corresponding 2D movement trajectory of the dynamic element.
[0076] In this step, for each identified dynamic region, object detection algorithms such as YOLO (You Only Look Once) and SSD (Single Shot MultiBox Detector) are applied to accurately locate the categories of dynamic elements (such as people, vehicles, etc.) and their real-time bounding box positions. By cross-frame matching these bounding boxes and considering the appearance features, positions, and velocity consistency of objects, the motion trajectories of dynamic elements between consecutive frames are constructed, thus generating 2D motion trajectories. This process requires the support of efficient object tracking algorithms, such as Kalman filters and Siamese tracking networks in deep learning, to ensure the accuracy and stability of tracking.
[0077] In the embodiment of the present application, consider a surveillance video analysis scenario where pedestrians in a specific area need to be monitored and tracked in real time. In step 104.1, by analyzing the second spatio-temporal feature matrix, the system quickly identifies the regions where pedestrians move in the frame, which form a sharp contrast with the static background, and these regions are regarded as dynamic element regions. In step 104.2, an advanced deep learning object detection model, such as Faster R-CNN, is used to accurately classify and localize the bounding boxes of pedestrians in these regions. Once a pedestrian is detected for the first time, the system immediately starts the deep learning-based SORT (Simple Online and Realtime Tracking) algorithm, which combines appearance features and motion models to continuously predict and adjust the bounding boxes of pedestrians, and can maintain tracking even in cases of partial occlusion or lighting changes. In this way, the system generates continuous 2D motion trajectories for each pedestrian, and these trajectory data can be used in further applications such as behavior analysis, crowd density estimation, or abnormal event detection.
[0078] Optionally, in the embodiment of the present application, the process of "mapping the 2D motion trajectories of the dynamic elements into the 3D space to create or update the target virtual 3D scene model, and inputting the first spatio-temporal feature matrix and the dynamic elements into the target virtual 3D scene model to generate a predicted video frame sequence through the target virtual 3D scene model" in step 105 may include: 1051. Extract 2D feature points on the 2D motion trajectories corresponding to the dynamic elements, and perform matching of the 2D feature points between different video frames to determine the matching pairs of 2D feature points; In this step, first, through step 1051, key feature points are carefully selected on the 2D motion trajectory of the dynamic element, such as contour inflection points or points with obvious texture changes, and the corresponding matches of these feature points are searched between adjacent video frames to establish cross-frame 2D feature associations. This process involves feature detection algorithms such as SIFT (Scale-Invariant Feature Transform) or ORB (Oriented FAST and Rotated BRIEF), which can effectively identify and match invariant feature points and remain reliable even under perspective changes or scale scaling.
[0079] 1052. Use the pair of matched 2D feature points to calculate the corresponding points of the dynamic element in 3D space to generate the 3D motion trajectory of the dynamic element; Next, in step 1052, using these pairs of well-matched 2D feature points, through principles of stereo vision or multi-view geometry, such as fundamental matrix or essential matrix calculation, the true coordinates of these points in three-dimensional space are deduced, and then the 3D motion trajectory of the dynamic element is reconstructed. This step is crucial for understanding the true motion path of the dynamic element and lays the foundation for subsequent 3D modeling.
[0080] 1053. Based on the subject behavior pattern and the scene information in the first spatio-temporal feature matrix, construct an initial virtual 3D scene model and integrate the 3D motion trajectory of the dynamic element into the initial virtual 3D scene model to generate a target virtual 3D scene model, where the dynamic element can move at the correct time and spatial positions in the virtual 3D scene model; In step 1053, based on the subject behavior pattern and the static structural information of the scene extracted from the first spatio-temporal feature matrix, a preliminary virtual 3D scene model is constructed. This model not only includes the geometric layout of the scene but also incorporates the expected patterns of subject activities, such as the walking paths of pedestrians or the driving routes of vehicles. Then, the 3D motion trajectory of the dynamic element is embedded into this initial model to ensure that each dynamic element can move at the correct spatio-temporal positions according to its actual behavior pattern, thus perfecting it into a target virtual 3D scene model.
[0081] 1054. According to the target virtual 3D scene model and the position information of the obtained dynamic element in each video frame, render new video frames under preset camera parameters to form a predicted video frame sequence.
[0082] Finally, in step 1054, by using the target virtual 3D scene model and the specific positions of each frame of the dynamic elements, combined with the preset camera perspective and lighting conditions, real-time rendering technology (such as using game engines like Unity or Unreal Engine) is adopted to generate new video frames. These new frames are strung together to form a predicted video sequence, which not only reflects the content of the original video but also predicts the possible future movements of the dynamic elements, and is applicable to various application scenarios such as video synthesis, predictive analysis, and virtual reality.
[0083] In the embodiments of the present application, assume that a predicted video is to be made for the preview of an action movie. First, from the existing action clips, for the fighting actions of the protagonist, in step 1051, the SIFT algorithm is used to extract the 2D feature points of key actions such as the protagonist's arm waving and kicking, and the matching pairs are found through feature matching of consecutive frames. In step 1052, binocular vision technology is adopted to calculate the movement trajectory of the protagonist in the three-dimensional space based on these matching point pairs. In step 1053, according to the scene layout set in the script (such as an ancient battlefield environment) and the protagonist's behavior patterns (attacking, dodging), a preliminary 3D scene model is constructed using the Unity game engine, and the calculated 3D movement trajectory is integrated into the model so that the protagonist can move naturally in the scene. Finally, in step 1054, while keeping the virtual camera following the protagonist's actions, according to the set shot language (such as tracking shots, overhead shots), a series of new video frames are rendered to generate an action preview sequence that is not actually shot in the movie but has a consistent style, helping the director and producer to preview and adjust the shooting plan in advance.
[0084] Optionally, in the embodiments of the present application, the process of "comparing the predicted video frame sequence with the video frame sequence decomposed from the video data frame by frame, and based on the comparison result, adjusting the model parameters in the target virtual 3D scene model" in step 106 may include: Calculating the difference loss function between the predicted video frame sequence and the video frame sequence decomposed from the video data, where the difference loss function contains a set of model parameters, and the set of model parameters at least includes: camera parameters or rendering parameters.
[0085] Among them, calculating the difference loss function between the predicted video frame sequence and the video frame sequence decomposed from the video data includes: Through the formula group: , calculate the difference loss function between the predicted video frame sequence and the video frame sequence decomposed from the video data; Among them, represents the pixel-level loss function, represents the timestamp as The predicted video frame corresponding to the frame, Denoted as the sequence of video frames decomposed from the video data, Denoted as the structural similarity index, used to measure the structural similarity between the predicted sequence of video frames and the sequence of video frames decomposed from the video data, with a value range between -1 and 1, where 1 indicates exactly the same and 0 indicates no correlation, calculated by the SSIM algorithm; Denoted as a hyperparameter for balancing the weights of the SSIM loss and the L1 norm loss; Denoted as the L1 norm loss, which is the sum of the absolute errors between the predicted sequence of video frames and the sequence of video frames decomposed from the video data; Among them, Denoted as the geometric consistency loss function, Denoted as the number of sampling points in the real 3D scene model, Denoted as the number of sampling points in the target virtual 3D scene model, Denoted as the real 3D scene model, Denoted as the target virtual 3D scene model, Denoted as the calculated point set For each point In the point set To the nearest point The sum of the squares of the distances, Denoted as the calculated point set For each point In the point set To the nearest point The sum of the squares of the distances, Denoted as a hyperparameter for adjusting the weight of the normal consistency loss In the geometric consistency loss The weight, Denoted as the normal consistency loss, used to measure And The difference between; Among them, Denoted as the temporal consistency loss function, Denoted as the predicted video frame corresponding to the frame with timestamp , Denoted as the predicted video frame corresponding to the frame with timestamp , Denoted as the optical flow error, used to measure the consistency of pixel motion between two predicted video frames, Denoted as the penalty coefficient of the optical flow error, Denoted as the optical smoothness loss, used to control the smooth transition of video frames over timestamps, Denoted as a hyperparameter, used to balance the influence of the optical smoothness loss and the optical flow error on the temporal consistency loss; Wherein, Denoted as a difference loss function, Denoted as a set of model parameters, Denoted as the weight of the pixel-level loss function, Denoted as the weight of the geometric consistency loss function, Denoted as the weight of the temporal consistency loss function.
[0086] In this step, the calculation of the difference loss function comprehensively considers the matching degree between the predicted video frame sequence and the original video frame sequence in multiple dimensions, aiming to minimize this difference by optimizing the model parameters, thereby improving the quality of the predicted video. Specifically, The similarity between the predicted frame and the original frame at the pixel level is measured by the structural similarity index (SSIM) and the L1 norm, emphasizing the preservation of the visual quality and structural consistency of the image content; Focus on the geometric accuracy of the 3D scene model, and evaluate the closeness between the predicted model and the real scene by the distance between point clouds and the normal consistency; Then ensure the smoothness of the video sequence in the temporal dimension, and promote coherent dynamic performance by analyzing the optical flow consistency and optical smoothness between adjacent frames. The weights of these loss terms (such as α, β, γ) need to be adjusted according to the specific application scenario to achieve the best reconstruction effect.
[0087] In the embodiments of the present application, it is assumed that a replay prediction system for a sports event needs to be optimized, and the following numerical examples are specifically implemented: First, set the initial hyperparameters as: , and the weights of the pixel-level loss function, the geometric consistency loss function, and the temporal consistency loss function are respectively , so as to balance the influence of each loss term on the final model adjustment.
[0088] Next, through an iterative optimization algorithm (such as Adam or SGD), for each predicted frame, calculate the value of. For example, for a certain frame, if , the L1 norm loss is 500, the mean squared sum of the nearest neighbor distances between the point sets and is 100, the normal consistency loss is, the optical flow error is 20, and the optical smoothness loss is 0.02, then the calculation of the sub-loss functions (pixel-level loss function, geometric consistency loss function, temporal consistency loss function) for this frame is as follows: =(1 - 0.9)+0.01 500 = 10 + 5 = 15; =100 + 100 + 0.001 0.01 = 200.0001; =0.1 20 + 0.005 0.02 = 2 + 0.001 = 2.001; Therefore, the difference loss function L for this frame is as follows: =1.0 15 + 0.5 200.0001 + 0.3 2.001 ≈ 15 + 100 + 0.6 ≈ 115.6; By continuously iterating this process, the loss function is calculated for each frame in the entire video sequence, and the camera parameters and rendering parameters are adjusted through backpropagation based on the calculation results until the total loss L drops to an acceptable range to ensure that the predicted video frame sequence is highly visually consistent and dynamically smooth with the original video sequence.
[0089] Furthermore, in the process of adjusting the model parameters based on the difference loss function, optimization algorithms such as gradient descent or its variants (such as Adam, RMSprop, etc.) are usually adopted. The following are the basic steps, taking the gradient descent method as an example: First, it is necessary to calculate the difference loss function with respect to the model parameters i.e., . The gradient points in the direction where the loss function increases fastest at the current parameter values, and each component corresponds to the partial derivative of the corresponding parameter in the parameter vector. This step is usually automatically completed by deep learning frameworks such as TensorFlow or PyTorch through the backpropagation algorithm.
[0090] Having obtained the gradient, according to the principle of gradient descent, it is necessary to adjust the parameters along the negative gradient direction in order to expect to reduce the value of the loss function. The general form of the update rule is: ; where is the learning rate, which determines the step size of parameter update. Choosing an appropriate learning rate is crucial. Too large will lead to an unstable training process and may miss the optimal solution; too small will lead to a too slow convergence speed.
[0091] The above process is repeated multiple times over the entire dataset (batch data or mini-batch data), and each iteration is called an epoch. At the end of each epoch, all samples (or subsets of samples) are traversed, and the parameters are updated accordingly. As the number of iterations increases, the model parameters are gradually adjusted to the values that minimize the differential loss function 𝐿, so that the performance of the model on the training data gets better and better.
[0092] For example, a simplified example can be provided to illustrate how this formula is applied in practice. Suppose there is a very simple linear regression model that attempts to fit a set of data points to predict the output . The form of the model can be , where are the camera parameters, and are the rendering parameters.
[0093] The loss function is the mean squared error (MSE).
[0094] Therefore, the following simplified situation can be obtained: Current model parameters: ; Learning rate ( ): 0.01; Suppose for a specific input-output pair , the calculated gradients (partial derivatives) with respect to and are respectively: ; According to the gradient descent formula, the parameters can be updated as follows: For ; For ; Therefore, after one round of gradient descent update, the model parameters are adjusted to and .
[0095] It should be noted that in practical applications, the gradient calculation involves the average of the entire training dataset (or its mini-batch), and the model and loss function may be more complex, but the basic principle is the same.
[0096] Figure 2 FIG. is a schematic structural diagram of a video data processing system based on a digital twin scenario provided by an embodiment of the present application. As Figure 2 shown, the system includes: The acquisition and processing module 21 is configured to acquire video data from the actual environment, decompose the video data into a video frame sequence, and perform spatio-temporal feature extraction on the video frame sequence to determine a first spatio-temporal feature matrix. Each row of the first spatio-temporal feature matrix includes a video frame in the video frame sequence, and each column includes the object motion path, the main body behavior pattern, and the scene information corresponding to the video frame. The object motion path, the main body behavior pattern, and the scene information corresponding to the video frame are determined through spatio-temporal feature extraction; The segmentation module 22 is configured to segment a plurality of video elements from the video frame sequence based on the first spatio-temporal feature matrix, and generate an element segmentation mask corresponding to each video element; The fusion module 23 is configured to fuse the element segmentation mask with the object motion paths corresponding to a plurality of the video frames in the first spatio-temporal feature matrix to generate a second spatio-temporal feature matrix. Each row of the second spatio-temporal feature matrix represents a time step, and each column includes the segmentation mask feature and the motion feature of all objects within the time step; The positioning module 24 is configured to, based on the second spatio-temporal feature matrix, real-time locate and track the dynamic elements in the video frame sequence, and obtain the 2D motion trajectory corresponding to the dynamic elements. The dynamic elements are the video elements with dynamic features among the video elements; The generation module 25 is configured to map the 2D motion trajectory of the dynamic elements into a 3D space to create or update a target virtual 3D scene model, and input the first spatio-temporal feature matrix and the dynamic elements into the target virtual 3D scene model to generate a predicted video frame sequence through the target virtual 3D scene model; The adjustment module 26 is configured to compare the predicted video frame sequence with the video frame sequence decomposed from the video data frame by frame, and based on the comparison result, adjust the model parameters in the target virtual 3D scene model. The model parameters at least include object motion parameters, camera parameters, or rendering parameters.
[0097] Optionally, in the embodiment of the present application, the acquisition and processing module 21 is specifically configured to identify the action information of all the main bodies in the video frame sequence to determine the main body behavior pattern of each main body; identify the scene change information that changes over time in the video frame sequence, and the static information in the video frame sequence, and form scene information according to the scene change information and the static information; where the static information at least includes: scene structure, object layout, and lighting conditions; convert the object motion paths, the main body behavior patterns, and the scene information corresponding to a plurality of the video frames into a preset numerical form to constitute a first spatio-temporal feature matrix.
[0098] Optionally, in the embodiments of the present application, the segmentation module 22 is specifically configured to input the first spatio-temporal feature matrix into a pre-established semantic segmentation model, so as to predict each video frame in the first spatio-temporal feature matrix through the semantic segmentation model and output a segmentation result, where the segmentation result includes a plurality of video elements and an element label corresponding to each video element, and the element label is used to indicate that the video element belongs to the foreground, background, or a specific object category in the video frame. The semantic segmentation model is trained based on a plurality of video frame samples with element labels; generate an element segmentation mask corresponding to each video element, where the element segmentation mask exists in the form of a binary image, and the binary image form includes black and white two colors, where white represents the area where the video element is located, and black represents the background area outside the video element.
[0099] Optionally, in the embodiments of the present application, the fusion module 23 is specifically configured to extract the object motion paths corresponding to a plurality of the video frames from the first spatio-temporal feature matrix, and the object motion path corresponding to each video frame is consistent with the time step of the corresponding element segmentation mask; convert each element segmentation mask into a segmentation mask feature, and convert the object motion path corresponding to each video frame into a motion feature; for each video frame, splice the segmentation mask feature of the corresponding object with the motion feature to form a comprehensive feature vector; arrange the comprehensive feature vectors corresponding to each video frame in sequence according to the time step and the number of objects to generate a second spatio-temporal feature matrix, and the dimension of the second spatio-temporal feature matrix includes: the number of time steps x (the number of objects x (the dimension of the segmentation mask feature + the dimension of the motion feature)).
[0100] Optionally, in the embodiments of the present application, the positioning module 24 is specifically configured to, based on the second spatio-temporal feature matrix, identify an area in the video frame sequence that has a significant motion difference from the background area outside the video element, and determine the area as the area where the dynamic element is located; in the area where the dynamic element is located, identify the category and bounding box of the dynamic element, and for each dynamic element, establish a motion trajectory of the dynamic element between the video frames based on the category and bounding box of the dynamic element to generate a 2D motion trajectory corresponding to the dynamic element.
[0101] Optionally, in the embodiments of the present application, the generation module 25 is specifically configured to extract 2D feature points on the 2D motion trajectory corresponding to the dynamic element, and match the 2D feature points between different video frames to determine the matched 2D feature point pairs; use the matched 2D feature point pairs to calculate the corresponding points of the dynamic element in the 3D space to generate the 3D motion trajectory of the dynamic element; based on the main body behavior pattern and the scene information in the first spatio-temporal feature matrix, construct an initial virtual 3D scene model, and integrate the 3D motion trajectory of the dynamic element into the initial virtual 3D scene model to generate a target virtual 3D scene model, where the dynamic element can move at the correct time and spatial position in the virtual 3D scene model; according to the target virtual 3D scene model and the obtained position information of the dynamic element in each video frame, render new video frames under the preset camera parameters to form a predicted video frame sequence.
[0102] Optionally, in the embodiments of the present application, the adjustment module 26 is specifically configured to calculate the difference loss function between the predicted video frame sequence and the video frame sequence decomposed from the video data, and the difference loss function includes a set of model parameters, and the set of model parameters includes at least: camera parameters or rendering parameters.
[0103] Optionally, in the embodiments of the present application, the adjustment module 26 is further configured to use the formula group: , to calculate the difference loss function between the predicted video frame sequence and the video frame sequence decomposed from the video data; wherein, represents the pixel-level loss function, represents the predicted video frame corresponding to the frame with the time stamp , represents the video frame sequence decomposed from the video data, represents the structural similarity index, which is used to measure the structural similarity between the predicted video frame sequence and the video frame sequence decomposed from the video data, and the value range is between -1 and 1, 1 means exactly the same, 0 means no correlation, and it is calculated by the SSIM algorithm; represents the hyperparameter, which is used to balance the weights of the SSIM loss and the L1 norm loss; represents the L1 norm loss, which is used for the sum of the absolute errors between the predicted video frame sequence and the video frame sequence decomposed from the video data; wherein, represents the geometric consistency loss function, represents the number of sampling points in the real 3D scene model, is expressed as the number of sampling points in the target virtual 3D scene model, is expressed as the real 3D scene model, is expressed as the target virtual 3D scene model, is expressed as the set of calculation points for each point to the point set the sum of the squares of the distances to the nearest point in, is expressed as the set of calculation points for each point to the point set the sum of the squares of the distances to the nearest point in, is expressed as a hyperparameter for adjusting the normal consistency loss in the geometric consistency loss weight, is expressed as the normal consistency loss for measuring and the difference between; Among them, is expressed as the temporal consistency loss function, is expressed as the predicted video frame corresponding to the frame with timestamp , is expressed as the predicted video frame corresponding to the frame with timestamp , is expressed as the optical flow error for measuring the consistency of pixel motion between two predicted video frames, is expressed as the penalty coefficient of the optical flow error, is expressed as the optical smoothness loss for controlling the smooth transition of video frames in timestamps, is expressed as a hyperparameter for balancing the influence of the optical smoothness loss and the optical flow error on the temporal consistency loss; Among them, is expressed as the difference loss function, is expressed as the set of model parameters, is expressed as the weight of the pixel-level loss function, is expressed as the weight of the geometric consistency loss function, is expressed as the weight of the temporal consistency loss function.
[0104] Figure 2 The video data processing system based on the digital twin scenario can execute Figure 1The video data processing method based on the digital twin scenario described in the illustrated embodiment, its implementation principle and technical effects will not be elaborated further. For the video data processing system based on the digital twin scenario in the above embodiment, the specific ways in which each module and unit perform operations have been described in detail in the embodiment related to this method, and will not be elaborated here.
[0105] In a possible design, Figure 2 The video data processing system based on the digital twin scenario of the illustrated embodiment can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32; The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32. The processing component 32 is configured to: collect video data from the actual environment, decompose the video data into a video frame sequence, and perform spatio-temporal feature extraction on the video frame sequence to determine a first spatio-temporal feature matrix, where each row of the first spatio-temporal feature matrix includes a video frame in the video frame sequence, and each column includes the object motion path, the main body behavior pattern, and the scene information corresponding to the video frame, and the object motion path, the main body behavior pattern, and the scene information corresponding to the video frame are determined by spatio-temporal feature extraction; based on the first spatio-temporal feature matrix, segment multiple video elements from the video frame sequence and generate an element segmentation mask corresponding to each video element; fuse the element segmentation mask with the object motion paths corresponding to multiple video frames in the first spatio-temporal feature matrix to generate a second spatio-temporal feature matrix, where each row of the second spatio-temporal feature matrix represents a time step, and each column includes the segmentation mask feature and the motion feature of all objects within the time step; based on the second spatio-temporal feature matrix, real-time locate and track the dynamic elements in the video frame sequence and obtain the 2D motion trajectory corresponding to the dynamic elements, where the dynamic elements are the video elements with dynamic features among the video elements; map the 2D motion trajectory of the dynamic elements into the 3D space to create or update the target virtual 3D scene model, and input the first spatio-temporal feature matrix and the dynamic elements into the target virtual 3D scene model to generate a predicted video frame sequence through the target virtual 3D scene model; compare the predicted video frame sequence with the video frame sequence decomposed from the video data frame by frame, and based on the comparison result, adjust the model parameters in the target virtual 3D scene model, where the model parameters at least include object motion parameters, camera parameters, or rendering parameters.
[0106] Among them, the processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above methods. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above methods.
[0107] The storage component 31 is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks or optical discs.
[0108] The display component 33 can be an electroluminescent (EL) element, a liquid crystal display or a microdisplay with a similar structure, or a retina-direct display or a similar laser scanning display.
[0109] Of course, the computing device may also necessarily include other components, such as input / output interfaces, communication components, etc.
[0110] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above peripheral interface module can be an output device, an input device, etc.
[0111] The communication component is configured to facilitate communication between the computing device and other devices in a wired or wireless manner, etc.
[0112] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. At this time, the computing device can refer to a cloud server, and the above processing component, storage component, etc. can be basic server resources leased or purchased from the cloud computing platform.
[0113] The embodiment of the present application also provides a computer storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the above Figure 1 video data processing method based on the digital twin scenario shown in the embodiment.
[0114] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.
[0115] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0116] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A video data processing method based on a digital twin scenario, characterized in that, Including: Collecting video data from an actual environment, decomposing the video data into a video frame sequence, and performing spatio-temporal feature extraction on the video frame sequence to determine a first spatio-temporal feature matrix. Each row of the first spatio-temporal feature matrix includes a video frame in the video frame sequence, and each column includes the object motion path, the main body behavior pattern, and the scene information corresponding to the video frame. The object motion path, the main body behavior pattern, and the scene information corresponding to the video frame are determined through spatio-temporal feature extraction; Based on the first spatio-temporal feature matrix, splitting a plurality of video elements from the video frame sequence, and generating an element segmentation mask corresponding to each video element; Fusing the element segmentation mask with the object motion paths corresponding to a plurality of the video frames in the first spatio-temporal feature matrix to generate a second spatio-temporal feature matrix. Each row of the second spatio-temporal feature matrix represents a time step, and each column contains the segmentation mask features and motion features of all objects within the time step; Based on the second spatio-temporal feature matrix, real-time positioning and tracking of dynamic elements in the video frame sequence, and obtaining a 2D motion trajectory corresponding to the dynamic elements. The dynamic elements are video elements with dynamic characteristics among the video elements; Mapping the 2D motion trajectory of the dynamic elements into a 3D space to create or update a target virtual 3D scene model, and inputting the first spatio-temporal feature matrix and the dynamic elements into the target virtual 3D scene model to generate a predicted video frame sequence through the target virtual 3D scene model; Comparing the predicted video frame sequence frame by frame with the video frame sequence decomposed from the video data, and based on the comparison result, adjusting the model parameters in the target virtual 3D scene model. The model parameters at least include camera parameters or rendering parameters.
2. The method according to claim 1, wherein The performing spatio-temporal feature extraction on the video frame sequence to determine a first spatio-temporal feature matrix includes: Identifying all objects in the video frame sequence, and continuously tracking the positions of each object in different frames to obtain the object motion path corresponding to the video frame; Identifying the action information of all main bodies in the video frame sequence to determine the main body behavior pattern of each main body; Identifying the scene change information that changes over time in the video frame sequence, as well as the static information in the video frame sequence, and composing scene information according to the scene change information and the static information; Wherein, the static information at least includes: scene structure, object layout, and lighting conditions; Converting the object motion paths, the main body behavior patterns, and the scene information corresponding to a plurality of the video frames into a preset numerical form to constitute a first spatio-temporal feature matrix.
3. The method according to claim 1, wherein The splitting a plurality of video elements from the video frame sequence based on the first spatio-temporal feature matrix, and generating an element segmentation mask corresponding to each video element includes: Input the first spatio-temporal feature matrix into a pre-established semantic segmentation model to predict each video frame in the first spatio-temporal feature matrix through the semantic segmentation model and output a segmentation result. The segmentation result includes a plurality of video elements and an element label corresponding to each video element. The element label is used to indicate whether the video element belongs to the foreground, background, or a specific object category in the video frame. The semantic segmentation model is trained based on a plurality of video frame samples with element labels; Generate an element segmentation mask corresponding to each video element. The element segmentation mask exists in the form of a binary image, and the binary image form includes black and white colors. Among them, white represents the area where the video element is located, and black represents the background area outside the video element.
4. The method according to claim 1, wherein The fusion of the element segmentation mask and the object movement paths corresponding to multiple video frames in the first spatio-temporal feature matrix to generate a second spatio-temporal feature matrix includes: Extract the object movement paths corresponding to multiple video frames from the first spatio-temporal feature matrix, and the object movement path corresponding to each video frame has the same time step as the corresponding element segmentation mask; Convert each element segmentation mask into a segmentation mask feature, and convert the object movement path corresponding to each video frame into a movement feature; For each video frame, splice the segmentation mask feature of the corresponding object with the movement feature to form a comprehensive feature vector; Arrange the comprehensive feature vectors corresponding to each video frame in sequence according to the time step and the number of objects to generate a second spatio-temporal feature matrix.
5. The method according to claim 1, wherein Based on the second spatio-temporal feature matrix, locate and track the dynamic elements in the video frame sequence in real time and obtain the 2D movement trajectory corresponding to the dynamic elements. The dynamic elements are the video elements with dynamic characteristics among the video elements, including: Based on the second spatio-temporal feature matrix, identify the area in the video frame sequence that has a significant movement difference from the background area outside the video element, and determine the area as the area where the dynamic element is located; In the area where the dynamic element is located, identify the category and bounding box of the dynamic element, and for each dynamic element, establish the movement trajectory of the dynamic element based on the category and bounding box of the dynamic element among the video frames to generate the 2D movement trajectory corresponding to the dynamic element.
6. The method according to claim 1, characterized in that, Map the 2D movement trajectory of the dynamic element into the 3D space to create or update the target virtual 3D scene model, and input the first spatio-temporal feature matrix and the dynamic element into the target virtual 3D scene model to generate a predicted video frame sequence through the target virtual 3D scene model, including: Extract 2D feature points on the 2D movement trajectory corresponding to the dynamic element, and perform matching of the 2D feature points between different video frames to determine the matching 2D feature point pairs; Use the matching 2D feature point pairs to calculate the corresponding points of the dynamic element in the 3D space to generate the 3D movement trajectory of the dynamic element. Construct an initial virtual 3D scene model based on the subject behavior pattern and the scene information in the first spatio-temporal feature matrix, and integrate the 3D motion trajectories of the dynamic elements into the initial virtual 3D scene model to generate a target virtual 3D scene model, wherein the dynamic elements can move at the correct time and spatial positions in the virtual 3D scene model; Render new video frames under preset camera parameters according to the target virtual 3D scene model and the position information of the dynamic elements obtained in each video frame to form a predicted video frame sequence.
7. The method according to claim 1, characterized in that Compare the predicted video frame sequence with the video frame sequence decomposed from the video data frame by frame, and based on the comparison result, adjust the model parameters in the target virtual 3D scene model, including: Calculate a difference loss function between the predicted video frame sequence and the video frame sequence decomposed from the video data, and the difference loss function contains a set of model parameters, and the set of model parameters includes at least: camera parameters or rendering parameters.
8. The method according to claim 7, characterized in that, The calculating the difference loss function between the predicted video frame sequence and the video frame sequence decomposed from the video data includes: By a set of formulas: , calculate the difference loss function between the predicted video frame sequence and the video frame sequence decomposed from the video data; Among them, is expressed as a pixel-level loss function, is expressed as the predicted video frame corresponding to the frame with time stamp ; is expressed as the video frame sequence decomposed from the video data, is expressed as the structural similarity index, which is used to measure the structural similarity between the predicted video frame sequence and the video frame sequence decomposed from the video data. The value range is between -1 and 1, where 1 means exactly the same and 0 means no correlation, and it is calculated by the SSIM algorithm; is expressed as a hyperparameter, which is used to balance the weights of the SSIM loss and the L1 norm loss; is expressed as the L1 norm loss, which is used for the sum of the absolute errors between the predicted video frame sequence and the video frame sequence decomposed from the video data; Among them, is expressed as the geometric consistency loss function, is expressed as the number of sampling points in the real 3D scene model, is expressed as the number of sampling points in the target virtual 3D scene model, is expressed as the real 3D scene model, is expressed as the target virtual 3D scene model, is expressed as the calculated point set for each point to the point set the sum of the squares of the distances to the nearest point in it, is expressed as the calculated point set for each point to the point set the sum of the squares of the distances to the nearest point in it, is expressed as a hyperparameter for adjusting the normal consistency loss in the geometric consistency loss weight, is expressed as the normal consistency loss for measuring and the difference between them; Among them, is expressed as a temporal consistency loss function, is expressed as the predicted video frame corresponding to the frame with timestamp ; is expressed as the predicted video frame corresponding to the frame with timestamp ; is expressed as an optical flow error, which is used to measure the consistency of pixel motion between two predicted video frames, is expressed as a penalty coefficient of the optical flow error, is expressed as an optical smoothness loss, which is used to control the smooth transition of video frames in timestamps, is expressed as a hyperparameter, which is used to balance the influence of the optical smoothness loss and the optical flow error on the temporal consistency loss; Among them, is expressed as a difference loss function, is expressed as a set of model parameters, is expressed as the weight of the pixel-level loss function, is expressed as the weight of the geometric consistency loss function, is expressed as the weight of the temporal consistency loss function.
9. A video data processing system based on a digital twin scenario, characterized in that, Including: An acquisition and processing module, configured to acquire video data from an actual environment, decompose the video data into a video frame sequence, and perform spatio-temporal feature extraction on the video frame sequence to determine a first spatio-temporal feature matrix. Each row of the first spatio-temporal feature matrix includes a video frame in the video frame sequence, and each column includes the object motion path, the subject behavior pattern, and the scene information corresponding to the video frame. The object motion path, the subject behavior pattern, and the scene information corresponding to the video frame are determined by spatio-temporal feature extraction; A segmentation module, configured to segment a plurality of video elements from the video frame sequence based on the first spatio-temporal feature matrix, and generate an element segmentation mask corresponding to each video element; A fusion module, configured to fuse the element segmentation mask with the object motion paths corresponding to a plurality of video frames in the first spatio-temporal feature matrix to generate a second spatio-temporal feature matrix. Each row of the second spatio-temporal feature matrix represents a time step, and each column includes the segmentation mask features and motion features of all objects within the time step; A positioning module, configured to real-time locate and track the dynamic elements in the video frame sequence based on the second spatio-temporal feature matrix, and obtain the 2D motion trajectories corresponding to the dynamic elements. The dynamic elements are the video elements with dynamic features in the video elements; A generation module, configured to map the 2D motion trajectories of the dynamic elements into 3D space to create or update a target virtual 3D scene model, and input the first spatio-temporal feature matrix and the dynamic elements into the target virtual 3D scene model to generate a predicted video frame sequence through the target virtual 3D scene model; An adjustment module is configured to compare the predicted video frame sequence frame by frame with the video frame sequence decomposed from the video data, and based on the comparison result, adjust the model parameters in the target virtual 3D scene model, where the model parameters at least include object motion parameters, camera parameters, or rendering parameters.
10. A computing device, characterized in that, It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a video data processing method based on a digital twin scenario as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Optimizations for dynamic object instance detection, segmentation, and structure mapping
CN111670457A
Video style migration attitude detection method based on graph neural network
CN119919458A
Airport crossing safety management method and system based on video capture
CN120071218A
Digital twinning method and system for scene flow based on dynamic trajectory flow
US20250087082A1
Cited By
Dynamic target mapping and track reproduction method and system in video based on digital twinborn scene
CN121458767A
A Method and System for Dynamic Target Mapping and Trajectory Reproduction in Video Based on Digital Twin Scenes
CN121458767B
Bidirectional mapping positioning and interaction method and system for video picture and three-dimensional model
CN121486547A