Video picture and three-dimensional model bidirectional mapping positioning and interaction method and system
By constructing a scene decomposition model that includes an environment layer and a dynamic entity layer, the motion trajectory and deformation of dynamic targets are calculated, achieving precise two-way mapping and interaction between video footage and 3D models. This solves the problem of the disconnect between virtual and reality in existing technologies and improves interaction efficiency and real-time performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-06
- Publication Date
- 2026-03-20
AI Technical Summary
Existing technologies in digital twin scenarios lack the accuracy and real-time performance of two-way interaction between video and 3D models, resulting in a disconnect between the virtual and reality, low interaction efficiency, and an inability to meet the high fidelity and real-time requirements of scenarios such as industrial monitoring and emergency drills.
A scene decomposition model is constructed, including an environment layer and a dynamic entity layer. The deformation module is used to calculate the motion trajectory and deformation of dynamic targets. User operations are responded to to generate editing commands, which drive the model to update the video footage. This achieves accurate separation and 3D reconstruction of dynamic targets and supports users to edit and generate multi-channel video sequences with updated content.
It achieves precise two-way mapping and interaction between video footage and 3D models, ensuring that editing operations are reflected in the video in real time, improving the accuracy and real-time nature of the interaction, and guaranteeing the continuity of video content and the consistency of perspective.
Smart Images

Figure CN121486547B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of digital twin interaction, and particularly relates to a video picture and three-dimensional model bidirectional mapping positioning and interaction method and system. BACKGROUND
[0002] In the fields of digital twin, smart city and industrial automation, fusing real-time video pictures with three-dimensional virtual models is a core technology for realizing scene visualization, dynamic monitoring and interaction simulation. The ideal goal is to establish a seamless bridge between the real world and the digital world (such as presented by three-dimensional models), and to realize accurate and real-time information correspondence and interaction between the two, thereby greatly improving the understanding, management and control capabilities of complex environments.
[0003] However, the current mainstream technology has significant limitations in achieving this goal. Most applications still remain in the primary stage of "one-way mapping", for example, monitoring videos are projected as textures onto three-dimensional building models for display, or vehicles in the video are manually labeled in the three-dimensional model. This one-way information flow causes disconnection between the virtual and the real, and cannot form a true interaction loop. Although some technologies have attempted to separate dynamic targets (such as pedestrians and vehicles) from static backgrounds in videos, due to algorithmic limitations, they perform poorly in accurately calculating the motion trajectories and complex morphological changes of targets, which directly leads to low accuracy and slow response in subsequent three-dimensional reconstruction.
[0004] These fundamental technical deficiencies directly lead to a series of problems in practical applications. First, due to the lack of smooth bidirectional linkage mechanism, user edits to the three-dimensional model (such as adjusting the structure of a bridge in simulation) cannot be reflected in the corresponding video pictures in real time, disrupting the interactive experience. Second, the separation and reconstruction of dynamic targets and static backgrounds are not satisfactory, resulting in low fidelity of the digital twin scene and an inability to accurately reproduce dynamic changes in the real world. These problems collectively result in low operational efficiency of the entire system, making it difficult to meet the stringent requirements of high fidelity and real-time performance in industrial monitoring, emergency drills and other scenarios. Therefore, the core problem of existing technology is the inability to solve the precision and real-time problem of video and three-dimensional model bidirectional interaction in the digital twin scene. SUMMARY
[0005] The present application aims to provide a video picture and three-dimensional model bidirectional mapping positioning and interaction method and system to solve the problem of insufficient precision and real-time performance of video and three-dimensional model bidirectional interaction in the digital twin scene in the prior art.
[0006] To solve the above technical problems, in a first aspect, the present application provides a video picture and three-dimensional model bidirectional mapping positioning and interaction method, comprising:
[0007] acquire multi-channel video picture data of a real scene corresponding to a target digital twin scene, and a static three-dimensional white model corresponding to the target digital twin scene, the multi-channel video picture data including dynamic targets and static backgrounds;
[0008] construct a scene decomposition model based on the multi-channel video picture data and the static three-dimensional white model, the scene decomposition model including an environment layer and at least one dynamic entity layer, the dynamic entity layer having a deformation module built-in, for calculating a motion trajectory and deformation of the dynamic target according to a time parameter;
[0009] in response to a specific time selected by a user in the multi-channel video picture data, query the scene decomposition model, separate the dynamic target from the static background, and map and position the dynamic target from the video picture to the static three-dimensional white model for shape reconstruction to obtain a reconstructed target;
[0010] in response to an editing operation performed by a user on the reconstructed target through a three-dimensional visualization interface, generate an editing instruction, the three-dimensional visualization interface being generated based on the static three-dimensional white model;
[0011] the scene decomposition model adjusts model parameters of the dynamic entity layer corresponding to the reconstructed target according to the editing instruction, and drives the deformation module to update the state, to generate a multi-channel video picture sequence with updated content based on the adjusted scene decomposition model mapping.
[0012] Optionally, the scene decomposition model adjusts model parameters of the dynamic entity layer corresponding to the reconstructed target according to the editing instruction, and drives the deformation module to update the state, to generate a multi-channel video picture sequence with updated content based on the adjusted scene decomposition model mapping, including:
[0013] the scene decomposition model receives and analyzes the editing instruction, and locates to the corresponding dynamic entity layer according to the editing instruction;
[0014] adjust the basic morphology parameters of the deformable mesh in the dynamic entity layer and the parameters of the internal calculation function network of the deformation module in reverse according to the editing instruction;
[0015] drive the deformation module to update the state based on the adjusted model parameters, so that for any input time, the deformation module can calculate the deformable mesh morphology according to the updated parameters to meet the editing intention;
[0016] for each target time in the video picture sequence to be generated, input the target time to each updated deformation module to calculate the three-dimensional morphology of each dynamic target at the target time;
[0017] combining the three-dimensional shape of each dynamic target with the environmental layer in the scene decomposition model to form a complete three-dimensional scene expression corresponding to the target moment;
[0018] According to a preset mapping relationship from the environmental layer to each video frame, the complete three-dimensional scene expression is projected into each video frame view to generate a sequence of multi-channel video frames with updated content.
[0019] Optionally, after constructing the scene decomposition model based on the multi-channel video frame data and the static three-dimensional white model, the method further comprises:
[0020] performing cross-view collaborative optimization on the scene decomposition model, wherein the cross-view collaborative optimization comprises:
[0021] selecting at least two different observation views for the same dynamic target in the multi-channel video frames;
[0022] generating simulated appearances of the dynamic target under the at least two different views based on the scene decomposition model;
[0023] calculating a difference between the simulated appearances and the appearance of the dynamic target in the real video frame under the corresponding view;
[0024] adjusting parameters of the deformation module in the dynamic entity layer corresponding to the dynamic target, so that the difference is simultaneously reduced under multi-view constraints, to optimize the consistency of the motion trajectory and deformation solution of the dynamic target in the three-dimensional space.
[0025] Optionally, the adjusting, according to the editing instruction, of the basic shape parameters of the deformable mesh and the parameters of the internal calculation function network of the deformation module comprises:
[0026] parsing the editing instruction to obtain a target shape to be reached by the reconstruction target at the specific moment;
[0027] using the target shape as a hard constraint condition at the specific moment and using original shapes reflected by the multi-channel video frame data at other moments as soft constraint conditions to construct a joint optimization function;
[0028] solving adjustment values of the basic shape parameters of the deformable mesh and the parameters of the internal calculation function network of the deformation module by minimizing the joint optimization function;
[0029] synchronously updating the basic shape parameters and the parameters of the calculation function network using the adjustment values.
[0030] Optionally, the response to the user selecting a specific time in the multi-channel video picture data, querying the scene decomposition model, separating the dynamic target from the static background, and mapping the dynamic target from the video picture to the static three-dimensional white model for shape reconstruction to obtain the reconstructed target, comprising:
[0031] In response to the user selecting a specific time in the multi-channel video picture data;
[0032] Query the scene decomposition model to obtain the deformable mesh shape and surface appearance information of each dynamic entity layer calculated by the deformation module at the specific time;
[0033] According to the deformable mesh shape and surface appearance information, separate each dynamic target from the corresponding static background of the multi-channel video picture data;
[0034] According to the spatial relationship between the corresponding dynamic entity layer and the environment layer, map the shape of each separated dynamic target to the corresponding spatial region of the static three-dimensional white model, and perform three-dimensional shape reconstruction;
[0035] The reconstructed three-dimensional shape is determined as the reconstructed target.
[0036] Optionally, the scene decomposition model is constructed based on the multi-channel video picture data and the static three-dimensional white model, comprising:
[0037] Identify at least one dynamic target in the multi-channel video picture data;
[0038] Based on the static three-dimensional white model, an environment layer representing a fixed scene structure is established;
[0039] Construct an independent dynamic entity layer for each dynamic target, which contains deformable mesh and surface appearance information, and the basic shape of the deformable mesh is related to the spatial region of the dynamic target in the environment layer;
[0040] Set a deformation module in each dynamic entity layer, which is a calculation function network that receives time input and calculates the changes of deformable mesh shape and surface appearance information;
[0041] Combine the state of the environment layer and each dynamic entity layer under the time drive, and synthesize the picture according to the preset mapping relationship to construct the scene decomposition model.
[0042] Optionally, the response to the user selecting a specific time in the multi-channel video picture data, querying the scene decomposition model, separating the dynamic target from the static background, and mapping the dynamic target from the video picture to the static three-dimensional white model for shape reconstruction to obtain the reconstructed target, comprising:
[0043] generate and display a three-dimensional visualization interface based on the geometry data and the texture data of the static three-dimensional white model;
[0044] render and display the reconstructed target as an interactive three-dimensional entity in the three-dimensional visualization interface;
[0045] capture an interactive event corresponding to the editing operation and three-dimensional space transformation parameters in response to a user-triggered editing operation, the editing operation being a spatial geometry editing operation triggered by an input device for the reconstructed target in the three-dimensional visualization interface;
[0046] generate a structured editing instruction according to the interactive event, the three-dimensional space transformation parameters, the unique identifier of the reconstructed target, and the specific time information, the editing instruction recording the index of the operated vertex set, the transformation type, and the specific transformation amount, the specific transformation amount being represented by a three-dimensional coordinate or a transformation matrix.
[0047] In a second aspect, the present application provides a video picture and three-dimensional model bidirectional mapping positioning and interaction system, comprising:
[0048] an acquisition module configured to acquire multi-channel video picture data of a real scene corresponding to a target digital twin scene, and a static three-dimensional white model corresponding to the target digital twin scene, the multi-channel video picture data including a dynamic target and a static background;
[0049] a construction module configured to construct a scene decomposition model based on the multi-channel video picture data and the static three-dimensional white model, the scene decomposition model including an environment layer and at least one dynamic entity layer, the dynamic entity layer being internally provided with a deformation module configured to calculate a motion trajectory and deformation of the dynamic target according to a time parameter;
[0050] a response module configured to, in response to a specific time selected by a user in the multi-channel video picture data, query the scene decomposition model, separate the dynamic target from the static background, and map and position the dynamic target from the video picture to the static three-dimensional white model for shape reconstruction to obtain a reconstructed target;
[0051] the response module is further configured to, in response to an editing operation performed by a user on the reconstructed target through a three-dimensional visualization interface, generate an editing instruction, the three-dimensional visualization interface being generated based on the static three-dimensional white model;
[0052] a generation module configured to adjust model parameters of a dynamic entity layer corresponding to the reconstructed target according to the editing instruction, and drive the deformation module to update the state, so as to generate a multi-channel video picture sequence with updated content based on the adjusted scene decomposition model.
[0053] In a third aspect, the present application provides an electronic device, comprising:
[0054] a memory for storing a computer program;
[0055] a processor for executing the computer program to implement the steps of the video frame and three-dimensional model bidirectional mapping positioning and interaction method according to the first aspect.
[0056] In a fourth aspect, the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executable by a processor to implement the steps of the video frame and three-dimensional model bidirectional mapping positioning and interaction method according to the first aspect.
[0057] The video frame and three-dimensional model bidirectional mapping positioning and interaction method provided by the present application obtains multi-path video frame data containing dynamic targets and static backgrounds and static three-dimensional white models of a target digital twin scene, thereby providing basic materials for bidirectional mapping interaction; constructs a scene decomposition model containing an environment layer and a dynamic entity layer with a built-in deformation module, thereby realizing scene static and dynamic layering and accurately calculating the motion trajectory and deformation of dynamic targets; in response to a specific time selected by a user, separates dynamic and static targets through the model and maps the dynamic targets from the video frame to the three-dimensional white model to complete shape reconstruction, thereby realizing accurate positioning from the video to the three-dimensional model; supports the user to edit the reconstructed targets through a visual interface based on the static three-dimensional white model and generates editing instructions, thereby providing a convenient interaction entry; the scene decomposition model adjusts the corresponding dynamic entity layer parameters according to the editing instructions and drives the deformation module to update, thereby finally mapping and generating a multi-path video frame sequence with updated content, and realizing synchronous linkage from the three-dimensional model to the video frame.
[0058] Further, after receiving and analyzing the editing instructions, the scene decomposition model is positioned to the corresponding dynamic entity layer, reversely adjusts the basic shape parameters of the deformable mesh and the parameters of the internal calculation function network of the deformation module of the layer, drives the deformation module to update to ensure that the deformable mesh shape conforming to the editing intention can be calculated at any input time, and then inputs each target time of the video frame sequence to be generated into each updated deformation module to calculate the three-dimensional shape of each dynamic target, combines the environment layer to form a complete three-dimensional scene expression, and then projects the complete three-dimensional scene expression to each video perspective according to the preset mapping relationship from the environment layer to each video frame, thereby generating a multi-path video frame sequence with updated content. This scheme accurately adjusts the model parameters and the deformation module state, ensures that the dynamic target shape completely matches the user editing intention, and can adapt to shape calculation at any time dimension, and through complete three-dimensional scene combination and multi-perspective projection, realizes accurate synchronization of the multi-path video frame sequence and three-dimensional editing operation, and guarantees the coherence, integrity and perspective consistency of the updated video content. Attached Figure Description
[0059] To more clearly illustrate the technical solutions of the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0060] Figure 1 A flowchart illustrating a bidirectional mapping positioning and interaction method between video footage and a 3D model, provided in an embodiment of this application;
[0061] Figure 2 A flowchart illustrating another method for bidirectional mapping, positioning, and interaction between video footage and a 3D model, provided in an embodiment of this application;
[0062] Figure 3 This is a schematic diagram of a bidirectional mapping, positioning, and interaction system between video footage and a 3D model, provided as an embodiment of this application. Detailed Implementation
[0063] In current digital twin applications, the fusion technology of video footage and 3D models faces a core bottleneck: the interaction between the two is often unidirectional, lacking real-time bidirectional synchronization. While existing technologies can overlay video information onto 3D scenes or annotate video objects within models, they are insufficient in accurately separating dynamic targets from static backgrounds and in real-time tracking of target motion trajectories and shape changes. A more critical flaw is that when users modify the 3D model, these changes cannot be reverse-engineered and updated in real-time in the video content. This leads to a disconnect between the virtual and real worlds, reducing the realism of scene recreation and severely impacting interaction efficiency, constituting a technical challenge of insufficient interactive accuracy and real-time performance.
[0064] To address the aforementioned issues, this invention proposes a novel bidirectional mapping, positioning, and interaction method between video footage and 3D models. The core of this method lies in constructing a scene decomposition model comprising an environment layer and a dynamic entity layer, capable of accurately separating static backgrounds and dynamic targets from multiple video streams. When a user selects any moment in the video, the system automatically maps the dynamic target to 3D space and performs precise shape reconstruction. Crucially, any edits made by the user to the reconstructed model within the 3D interface are parsed by the system and used to update the scene decomposition model, thereby generating a video sequence with synchronized content updates in real time. This closed-loop interaction mechanism not only overcomes the challenges of inaccurate separation of static and dynamic elements and poor bidirectional interaction but also achieves instantaneous linkage from 3D editing to video updates, fundamentally improving the accuracy and real-time performance of virtual-real interaction in digital twin scenarios.
[0065] For those skilled in the technical field, the present application will be further described in detail below in conjunction with the drawings and specific embodiments. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0066] The core of the present application is to provide a video picture and three-dimensional model bidirectional mapping positioning and interaction method, and a specific embodiment of the method is shown in the flowchart Figure 1 As shown, the method comprises:
[0067] S101, acquiring multi-channel video picture data of a real scene corresponding to a target digital twin scene, and a static three-dimensional white model corresponding to the target digital twin scene.
[0068] Among them, the multi-channel video picture data is the recording of the real physical scene, which contains dynamic targets and static backgrounds. Dynamic targets refer to objects in the real scene that can move, deform or change state, and static backgrounds are the environment part of the scene that always remains unchanged. The static three-dimensional white model is the basic skeleton of the target digital twin scene, which is used to accurately present the fixed spatial structure of the scene and does not contain any motion or change elements.
[0069] In a specific embodiment, in order to fully capture the overall appearance of the real physical scene and avoid visual angle blind area, multiple cameras are arranged around the physical scene at different positions. These cameras will start and work synchronously, real-time collect the video pictures of the real scene, and finally form multi-channel video data that can cover all angles of the scene, and each channel of video contains the motion process of dynamic targets and the complete image of static backgrounds. At the same time, by pre-scanning, modeling and other ways of the real physical scene, its static three-dimensional white model is constructed as the basic model for subsequent bidirectional mapping and interaction.
[0070] S102, constructing a scene decomposition model based on the multi-channel video picture data and the static three-dimensional white model.
[0071] Among them, the scene decomposition model aims to decompose a complex digital twin scene into static "environment layer" and dynamic "entity layer" according to function.
[0072] The environment layer represents the fixed part of the scene, such as buildings, terrain, large fixed equipment, etc. The dynamic entity layer creates an independent layer for each movable or changeable target in the scene (such as vehicles, robots, personnel). It exclusively records and manages the geometry, appearance and motion state of the target.
[0073] Specifically, S102 specifically includes the following processes:
[0074] S1021, identify at least one dynamic target in the multi-channel video picture data.
[0075] For example, analyze the multi-channel video picture by a target detection algorithm (such as YOLO) to automatically identify and lock all targets in an active state.
[0076] S1022, based on the static three-dimensional white model, an environment layer representing the fixed scene structure is established.
[0077] Using an existing static three-dimensional model (white model), an "environment layer" representing the fixed scene structure is constructed to ensure that the base of the digital scene completely corresponds to the static environment of the physical world.
[0078] S1023, for each dynamic target, an independent dynamic entity layer is constructed, which contains a deformable mesh and surface appearance information, and the basic form of the deformable mesh is associated with the spatial area of the dynamic target in the environment layer.
[0079] Among them, the deformable mesh is the basic structure for simulating the shape of the dynamic target in the dynamic entity layer, which can change correspondingly with the movement or deformation of the target. The surface appearance information is the visual features of the surface of the dynamic target, including color, texture, identification and other contents that can reflect the visual attributes of the target.
[0080] S1024, a deformation module is set in each dynamic entity layer, which is a calculation function network for receiving time input and calculating the changes of the deformable mesh form and the surface appearance information.
[0081] Among them, the deformation module is the core calculation unit built-in the dynamic entity layer, which can accurately calculate the shape change and surface visual feature change of the dynamic target after receiving time information. The calculation function network is a set composed of a series of pre-designed calculation rules, which can automatically output the corresponding calculation results according to the input of specific information.
[0082] S1025, combine the state of the environment layer and each dynamic entity layer under time driving, and synthesize the picture according to the preset mapping relationship to construct a scene decomposition model.
[0083] In a specific embodiment, the overall process of constructing the scene decomposition model is to first identify the dynamic target from the multi-channel video, then establish the environment layer relying on the static three-dimensional white model, match the independent dynamic entity layer for each dynamic target and configure the deformation module, and finally synthesize the picture by combining the time state of each layer to form a scene decomposition model that can be disassembled and built.
[0084] As an example, take the digital twin factory scene as an example, which contains 3 robots, 2 conveying lines and other dynamic targets, and the static background is the factory building, fixed equipment, etc.:
[0085] Firstly, step S1021 analyzes multiple video pictures through a target detection algorithm to identify 3 robots, 2 conveying lines and other dynamic targets that will move, and records their appearance positions and initial states in the video. In the application, the target detection algorithm can be YOLO algorithm.
[0086] Secondly, step S1022 extracts the information of the fixed structures such as the factory building framework and equipment base in the static three-dimensional white model of the factory, and directly establishes the environment layer to ensure that the environment layer is completely consistent with the fixed structures of the real factory.
[0087] Then, step S1023 constructs an independent dynamic entity layer for each robot and each conveying line, and configures a deformable mesh for each dynamic entity layer. For example, the deformable mesh of the robot fits its body contour, and the deformable mesh of the conveying line matches its track length. At the same time, the surface appearance information of the dynamic targets is recorded, such as the blue body of the robot and the gray track of the conveying line. The initial shape of the deformable mesh corresponds to the actual installation position of the dynamic target in the environment layer, such as the mesh shape of No. 1 robot fitting its installation area on the east side of the factory building.
[0088] Then, step S1024 sets a deformation module in each dynamic entity layer. The calculation function network of the deformation module needs to be trained first. The specific process is as follows: collect the historical motion data of the robots and conveying lines in the factory, which labels the shape and appearance changes at different time points; take the time parameter as the input and the shape and appearance change data as the output to train the network to learn the corresponding relationship between the two, and adjust the network parameters until the prediction error meets the requirements. The trained deformation module can calculate the shape change through the following formula (1):
[0089] (1)
[0090] In the formula, is the deformable mesh shape at time t, is the basic shape of the deformable mesh, is a time decay function, whose value range is 0-1, used to reflect the proportion of shape change amplitude over time, is the maximum deformation amplitude of the dynamic target.
[0091] For example, the of No. 1 robot is the initial body contour mesh, and its maximum stretching amplitude is 5 cm, when t = 1 second , substituting formula (1) gives: , the state of the robot arm is stretched by 3 cm.
[0092] Finally, the state-fixed environment layer is combined with the state-changing dynamic entity layer, and according to the preset mapping relationship of the environment layer to each video frame, a complete scene frame at each time point is synthesized, and finally a scene decomposition model is constructed. In this example, the mapping relationship is a correspondence rule between the three-dimensional scene area and the camera shooting angle set in advance, such as the west area of the factory building can only be shot by camera 1, the east area of the factory building can only be shot by camera 2, and the central area of the factory building can be shot by camera 3 and camera 4 at the same time. This ensures that each area has a corresponding video perspective and that there is no blind area.
[0093] The above example is only one example of the present application, and in actual application, the recognition algorithm of the dynamic target and the training data amount of the deformation module can be adjusted according to the needs, which is not limited in the present application.
[0094] In another specific embodiment, the recognition of the dynamic target can use the background difference method to extract the dynamic target by comparing the differences between the video frames and the static background; the calculation function network of the deformation module can use a simple linear regression network, which is suitable for scenes with simple dynamic target deformation rules, such as uniform motion conveying lines, and can also achieve accurate correspondence between time and shape change.
[0095] The present application constructs a scene decomposition model through the above steps, realizes the hierarchical and fine management of the digital twin scene, accurately captures the motion trajectory and deformation rule of the dynamic target, and has the ability of scene state splitting and reconstruction, providing reliable structured model support for subsequent bidirectional mapping positioning and interaction of video frames and three-dimensional models.
[0096] After S102, it further includes:
[0097] The scene decomposition model is cross-view collaborative optimized, and the cross-view collaborative optimization includes: selecting at least two different observation angles for the same dynamic target in the multiple video frames; based on the scene decomposition model, respectively generating simulated appearances of the dynamic target under the at least two different angles; calculating the difference between the simulated appearance and the appearance of the dynamic target in the real video frame under the corresponding angle; adjusting the parameters of the deformation module in the dynamic entity layer corresponding to the dynamic target, so that the difference is simultaneously reduced under the multi-angle constraint, to optimize the consistency of the motion trajectory and deformation calculation of the dynamic target in the three-dimensional space.
[0098] The simulated appearance refers to the visual presentation effect of the same dynamic target under a specific observation angle based on the constructed scene decomposition model, which is equivalent to the target appearance predicted by the model, including the shape contour, surface texture and other core visual features of the target. The multi-angle constraint refers to the core condition that needs to be followed in the optimization process, that is, when adjusting the model parameters, the differences of the same dynamic target under at least two different selected angles need to be reduced synchronously, rather than optimizing the differences of a single angle, to ensure that the optimized target state has uniform consistency in the three-dimensional space.
[0099] The overall optimization logic of this part is to verify the simulated appearance generated by the model through the real appearance of the dynamic target under multiple angles, adjust the key parameters of the scene decomposition model in the opposite direction, eliminate the calculation deviation caused by a single angle, and finally improve the consistency of the three-dimensional state calculation of the dynamic target.
[0100] As an example, still taking No. 1 robot in the digital twin factory scene as an example:
[0101] First, two different observation angles are selected for No. 1 robot, camera 1 and camera 3, where camera 1 is a front view angle that can clearly cover the front of the robot body and the head area, and camera 3 is a side view angle that can completely cover the side of the robot body and the arm movement area, ensuring that the two angles can complementarily capture the appearance details and movement posture of the robot from different dimensions, avoiding incomplete optimization caused by visual angle blind area.
[0102] Secondly, based on the constructed scene decomposition model, the dynamic entity layer and the built-in deformation module corresponding to No. 1 robot are called, and the current time parameter is input to generate the front simulated appearance of the robot under the camera 1 view angle and the side simulated appearance under the camera 3 view angle. The generated simulated appearance needs to completely reproduce the contour shape, current posture of the arm and surface texture features of the robot under the corresponding view angle.
[0103] Then, the difference between the simulated appearance and the real appearance is calculated: for each view angle, the key visual features such as the contour and texture of the robot in the real video picture are extracted, and then compared with the simulated appearance features of the corresponding view angle to judge the matching degree. For example, under the front view angle of camera 1, it is found that the head contour of the simulated appearance is narrower than the real appearance; under the side view angle of camera 3, there is a significant deviation between the stretching angle of the simulated appearance and the real appearance of the robot arm, which are all differences that need to be quantified.
[0104] Finally, based on the difference between the two perspectives, adjust the parameters of the deformation module in the dynamic entity layer corresponding to the first robot, such as the time decay function parameters and the maximum deformation amplitude parameters in formula (1). During the adjustment process, strictly follow the multi-perspective constraint to ensure that the difference between the two perspectives is reduced synchronously after adjustment, for example, after adjustment, the robot head profile in the camera 1 perspective completely matches the real appearance, and the robot arm stretching angle in the camera 3 perspective is consistent with the real state, at this time, the difference between the simulation appearance and the real appearance of the same robot under the two perspectives is significantly reduced, and the consistency of the motion trajectory and the deformation calculation in the three-dimensional space is effectively improved.
[0105] The above example is only an example of the present application, and in actual application, more observation perspectives can be selected and more detailed visual features can be used for difference comparison according to requirements, which is not limited in the present application.
[0106] In another specific embodiment, for dynamic targets with complex deformation rules, such as multi-joint robot arms and foldable conveying lines, three or more non-coplanar observation perspectives can be selected for collaborative optimization. At the same time, different weights are given according to the imaging clarity and coverage integrity of different perspectives, and the difference of clear perspectives is preferentially reduced, and other perspectives are simultaneously optimized, which can further improve the optimization accuracy and is suitable for precise manufacturing digital twin scenarios with high consistency requirements.
[0107] Through cross-perspective collaborative optimization, the present application effectively eliminates the deviation of the motion trajectory and the deformation calculation of the dynamic target under a single perspective, significantly improves the consistency of the calculation results of the same dynamic target under multiple perspectives, and strengthens the calculation accuracy and reliability of the scene decomposition model, providing a more optimal and stable model basis for subsequent accurate bidirectional mapping positioning of video pictures and three-dimensional models.
[0108] S103, in response to a specific time selected by a user in the multi-path video picture data, querying the scene decomposition model, separating the dynamic target from the static background, and mapping and positioning the dynamic target from the video picture to the static three-dimensional white model for shape reconstruction to obtain a reconstructed target.
[0109] Wherein, the specific time refers to a certain time point selected by the user from the multi-path video picture data according to the requirements, which is used to accurately extract the state of the dynamic target at that time point. The reconstructed target refers to the three-dimensional model consistent with the real dynamic target shape obtained by mapping and positioning the dynamic target in the video picture to the static three-dimensional white model.
[0110] S103 specifically includes:
[0111] S1031, in response to the selection of a specific time by a user in the multi-path video picture data.
[0112] S1032, query the scene decomposition model to obtain the deformable mesh shape and surface appearance information of each dynamic entity layer calculated by the deformation module at the specific moment.
[0113] S1033, according to the deformable mesh shape and surface appearance information, separate each dynamic target from the static background corresponding to the multi-channel video picture data.
[0114] S1034, map and position the shape of each separated dynamic target to the corresponding spatial region of the static three-dimensional white model according to the spatial correlation of the corresponding dynamic entity layer and the environment layer, and perform three-dimensional shape reconstruction.
[0115] Wherein, three-dimensional shape reconstruction refers to the process of restoring the real three-dimensional shape of the target in the static three-dimensional white model based on the appearance characteristics and spatial position information of the separated dynamic target. The reconstruction result needs to accurately match the size, posture and spatial position of the target in the real scene.
[0116] S1035, determine the reconstructed three-dimensional shape as the reconstruction target.
[0117] In a specific embodiment, first, in response to the user completing the specific moment selection operation through the visual interaction interface, the time index of the selected moment is queried to extract the key data such as the deformable mesh shape and surface appearance information of the dynamic entity layer corresponding to each dynamic target at the moment from the scene decomposition model that has completed cross-view collaborative optimization; based on these extracted key data, the accurate separation of each dynamic target and the static background in the video picture is completed; then according to the preset spatial correlation of the dynamic entity layer and the environment layer, the shape of the separated dynamic target is mapped and positioned to the corresponding spatial region of the static three-dimensional white model, and the three-dimensional shape reconstruction work is carried out synchronously, and finally the reconstructed three-dimensional shape is determined as the reconstruction target.
[0118] As an example, taking the digital twin factory scene as an example:
[0119] First, step S1031 responds to user operation, and the user selects t=8 seconds as a specific moment from the multi-channel factory operation video through the video playing interface. At this moment, the No. 1 robot is in the posture of grabbing materials.
[0120] Secondly, through step S1032, the optimized scene decomposition model is automatically queried, and according to the time parameter t=8 seconds, the deformable mesh shape and surface appearance information of the dynamic entity layer corresponding to the No. 1 robot at this moment are extracted, wherein the deformable mesh shape is the outline of the body and arm in the grabbing posture, and the surface appearance information includes the metal texture of the blue body and the front grabbing claw.
[0121] Then, step S1033 separates the first robot from the static background in the video image according to the extracted deformable mesh shape and surface appearance information, and the static background includes the factory wall and the ground. Specifically, the robot region pixels with high matching degree are retained and the static background region pixels with low matching degree are removed by comparing the pixel features of the video image with the appearance information of the dynamic entity layer.
[0122] Then, step S1034 accurately maps and positions the separated robot shape to the corresponding spatial area on the east side of the static three-dimensional white model according to the spatial correlation between the dynamic entity layer of the first robot and the environment layer, and the correlation is pre-set, i.e., the initial position of the dynamic entity layer corresponds to the equipment area on the east side of the factory. The three-dimensional shape is reconstructed by integrating the multi-view appearance features of the robot at this moment in the multi-channel video to restore the three-dimensional pose and real size of the robot when grabbing the material, wherein the three-dimensional pose includes the arm stretching angle and the opening and closing state of the grabbing claw.
[0123] Finally, step S1035 determines the three-dimensional shape of the robot obtained by the reconstruction as the reconstruction target, and completes the accurate conversion from the specific moment of the video to the three-dimensional model. The above example is only one example of the present application, and in actual application, the separation method, the detail accuracy of the reconstruction, etc. can be adjusted according to the needs, which are not limited in the present application.
[0124] Through the above steps, the present application realizes the accurate mapping and shape reconstruction of the dynamic target from the video image to the static three-dimensional white model at a specific moment, ensures complete separation of dynamic and static, and reconstruction of the target shape, effectively opens the conversion channel of the video visual information to the three-dimensional space information, and provides an accurate and reliable operation object for the subsequent user interaction and editing of the three-dimensional target.
[0125] S104, in response to the editing operation of the reconstruction target implemented by the user through the three-dimensional visualization interface, generating an editing instruction.
[0126] S104 specifically includes:
[0127] S1041, generating and displaying a three-dimensional visualization interface through a three-dimensional graphics rendering engine based on the geometric data and texture data of the static three-dimensional white model.
[0128] The geometry data is the spatial structure core data of the static three-dimensional white model, including vertex coordinates, edge connection relationship, face structure information, etc. of the white model, and is used to define the three-dimensional shape and spatial position of the white model. The texture data is the surface visual texture information of the static three-dimensional white model, including surface color distribution, material texture pattern, etc., and is used to restore the real visual effect of the white model. The three-dimensional graphics rendering engine is a core tool specially used for converting three-dimensional data into visual images. It can calculate and generate three-dimensional images conforming to the visual habits of human eyes according to the input geometry data and texture data and display them. The three-dimensional visualization interface is a visual interactive window for presenting a digital twin scene, which is generated based on the basic data of the static three-dimensional white model and can intuitively display the three-dimensional spatial structure and reconstruction target of the scene.
[0129] S1042, in the three-dimensional visualization interface, the reconstruction target is rendered and displayed as an interactive three-dimensional entity.
[0130] S1043, in response to a user triggered editing operation, an interactive event corresponding to the editing operation and a three-dimensional space transformation parameter are captured.
[0131] The interactive event is a definition of the user editing operation type, including specific operation categories such as moving, rotating, and scaling. The three-dimensional space transformation parameter is data describing the influence of the editing operation on the spatial state of the reconstruction target, including key information such as the direction, angle, and distance of the transformation. The editing operation refers to a spatial geometry editing operation triggered by an input device for the reconstruction target in the three-dimensional visualization interface. The input device is a tool for the user to carry out editing operations, which commonly includes a mouse, a keyboard, a touch screen, and a three-dimensional operation handle.
[0132] S1044, according to the interactive event, the three-dimensional space transformation parameter, the unique identifier of the reconstruction target, and the specific time information, a structured editing instruction is generated.
[0133] The unique identifier is the exclusive identification information assigned to each reconstruction target, which is used to accurately locate the operated reconstruction target and avoid confusion with other targets. The editing instruction records the index of the operated vertex set, the transformation type, and the specific transformation amount; the index of the operated vertex set is the numbering information of the edited part of the reconstruction target three-dimensional model, which is used to accurately locate the model position to be adjusted; the transformation type and the specific transformation amount are the core content of the editing instruction, which respectively clearly adjust the mode and the specific degree of adjustment, wherein the specific transformation amount is represented by three-dimensional coordinates or a transformation matrix; the transformation matrix is a mathematical tool for representing three-dimensional space transformation, which can accurately quantify the specific parameters of complex transformations such as translation and rotation.
[0134] In a specific embodiment, the geometric data and texture data of the static three-dimensional white model are first extracted, rendering calculation is performed on the data by means of a three-dimensional graphics rendering engine, a three-dimensional visualization interface capable of intuitively presenting the spatial structure of the digital twin scene is generated, then in the interface, the reconstruction target of which the three-dimensional shape reconstruction has been completed is rendered and displayed as a three-dimensional entity that can be directly operated, ensuring that the user can clearly see the reconstruction target and the scene environment in which it is located, then the user operation is monitored in real time, when the user triggers an editing operation on the reconstruction target in the interface by means of an input device, the corresponding interactive event and three-dimensional space transformation parameter are captured in response to the operation, finally the captured interactive event, three-dimensional space transformation parameter, reconstruction target unique identifier for accurately positioning the operation object, and specific time information of the associated operation time point are integrated to generate a structured editing instruction with a standard format and complete information, ensuring that the instruction can be accurately parsed by the scene decomposition model.
[0135] As an example, taking the reconstruction target of No. 1 work robot in the digital twin factory scene as an example:
[0136] Firstly, the geometric data and texture data of the static three-dimensional white model are extracted in step S1041, wherein the geometric data includes the vertex coordinates of the factory building, the edge connection relationship of the wall, etc., and the texture data includes the gray paint texture of the factory wall, the anti-skid texture of the ground, etc., rendering calculation is performed on the data by means of a three-dimensional graphics rendering engine, a three-dimensional visualization interface is generated and displayed, and the three-dimensional structure of the factory building is clearly presented in the interface. Optionally, in the application, the three-dimensional graphics rendering engine can be Unity engine.
[0137] Secondly, the reconstruction target (three-dimensional model of the grasping material posture) of No. 1 robot is rendered and displayed as an interactive entity in the three-dimensional visualization interface by step S1042, so that the user can directly see the three-dimensional shape of the robot and its spatial position in the factory building in the interface.
[0138] Then, the user triggers the editing operation of “rotating the arm” on the robot reconstruction target in the interface by means of a mouse as an input device, step S1043 captures the corresponding rotation operation interactive event and three-dimensional space transformation parameter in response to the operation, at this time, the three-dimensional space transformation parameter corresponding to the rotation operation is represented by a rotation matrix, and the formula of the rotation matrix is:
[0139] (2)
[0140] In the formula, R is the rotation matrix around the Z axis, is the rotation angle. Assuming that the user rotates the robot arm by 30°, then , , , and the formula (2) is obtained: .
[0141] At the same time, the vertex set index of the operated arm part is captured, such as the vertex corresponding to index 101-200.
[0142] Finally, all the above information is integrated to generate structured editing instructions, which contain clear and specific content, including the interactive event as rotation, the three-dimensional space transformation parameter as the above formula (2) after the parameter is brought in, the unique identifier of the reconstructed target of the No. 1 robot, that is, ID: Robot-001, the specific time information t=8 seconds, the operated vertex set index 101-200, the transformation type around the Z-axis rotation, and the specific transformation amount 30°.
[0143] The above example is only one example of the present application, and in actual application, different input devices and adjustment of transformation parameter representation methods can be selected according to requirements, and the present application does not limit this.
[0144] In another specific embodiment, for a complex digital twin scene with multiple reconstructed targets, the three-dimensional visualization interface can support simultaneous display and identification of multiple reconstructed targets, such as different color annotations for different targets, and the user can select multiple reconstructed targets at the same time through the frame selection operation to carry out batch editing. The system synchronously captures the editing operation information of each target to generate batch editing instructions containing multiple target adjustment contents, and the instructions distinguish each operated target through different unique identifiers, which can significantly improve the editing efficiency in the multi-target scene, and is suitable for dynamic target intensive industrial production digital twin scene.
[0145] Through the above steps, the present application realizes intuitive interaction between the user and the digital twin scene, accurately captures the user's editing intention for the reconstructed target and converts it into structured editing instructions, ensures that the editing instruction information is complete and the format is standard, and provides accurate and reliable instruction basis for subsequent scene decomposition model parameter adjustment and generation of updated video picture sequence.
[0146] S105, the scene decomposition model adjusts the model parameters of the dynamic entity layer corresponding to the reconstructed target according to the editing instruction, and drives the deformation module to update the state to generate a multi-path video picture sequence with updated content based on the adjusted scene decomposition model mapping.
[0147] Among them, the model parameter is the core data of the dynamic entity layer that defines the dynamic target morphology and motion law, including the basic morphology parameter of the deformable mesh and the internal parameter of the deformation module calculation function network, which directly determines the three-dimensional presentation effect and time variation law of the dynamic target. The state update refers to the update of the calculation logic and output result of the deformation module according to the adjusted model parameters, to ensure that the subsequent calculated dynamic target morphology meets the editing intention.
[0148] S105 specifically comprises:
[0149] S1051, the scene decomposition model receives and parses the editing instruction, and locates to the corresponding dynamic entity layer according to the editing instruction.
[0150] S1052, according to the editing instruction, reversely adjusting the basic morphological parameters of the deformable mesh in the dynamic entity layer and the parameters of the calculation function network inside the deformation module.
[0151] S1052 specifically comprises:
[0152] Parsing the editing instruction to obtain the target morphology that the reconstruction target needs to reach at the specific time; taking the target morphology as a hard constraint condition at the specific time, and taking the original morphology reflected by the multi-path video picture data at other times as a soft constraint condition, constructing a joint optimization function; by minimizing the joint optimization function, the adjustment value of the basic morphological parameters of the deformable mesh and the parameters of the calculation function network inside the deformation module is solved; using the adjustment value, synchronously updating the basic morphological parameters and the parameters of the calculation function network.
[0153] Wherein, the reverse adjustment refers to the process of reverse derivation and correction of the original parameters of the model with the target morphology edited by the user as the final result, and the core is to ensure that the adjusted parameters can make the deformation module calculate the morphology that meets the editing intention. The hard constraint condition is a condition that must be strictly met in the optimization process, which requires that the adjusted model must accurately match the target morphology edited at a specific time without deviation space. The soft constraint condition is a condition that needs to be met as much as possible in the optimization process, which is used to ensure that the morphology of the dynamic target at other times still meets the motion law in the original video, avoiding the disorder of the overall motion logic caused by single-time adjustment. The joint optimization function is a mathematical function integrating the hard constraint and the soft constraint, and the minimum value of the sum of the deviations of the two constraints corresponds to the optimal adjustment value. The adjustment value is the specific magnitude of the model parameter that needs to be corrected, which is used to correct the original parameter to the optimal parameter that meets the constraint condition.
[0154] S1053, based on the adjusted model parameters, driving the deformation module to update the state, so that for any input time, the deformation module can calculate the deformable mesh morphology that meets the editing intention according to the updated parameters.
[0155] S1054, for each target time in the to-be-generated video picture sequence, inputting the target time to each updated deformation module to calculate the three-dimensional morphology of each dynamic target at the target time.
[0156] S1055, combine the three-dimensional shape of each dynamic target with the environmental layer in the scene decomposition model to form a complete three-dimensional scene expression corresponding to the target moment.
[0157] Among them, the complete three-dimensional scene expression refers to the complete digital twin scene data containing static background and all dynamic targets, which can fully reproduce the three-dimensional space state of the scene at a specific moment, and is the basis for generating video pictures.
[0158] S1056, according to the preset mapping relationship from the environmental layer to each video picture, project the complete three-dimensional scene expression to each video picture view respectively to generate a multi-channel video picture sequence with updated content.
[0159] Among them, the multi-channel video picture sequence refers to a new video sequence generated after parameter adjustment corresponding to user editing operation, and the shape and posture of dynamic targets in the sequence are consistent with the reconstructed target after editing.
[0160] In a specific embodiment, as shown in Figure 2 The scene decomposition model first receives the structured editing instructions converted from user operations, analyzes the core information in the instructions, accurately locates the dynamic entity layer corresponding to the reconstructed target according to the unique identifier obtained by analysis, and then starts the reverse optimization process. With the specific moment target shape specified by the editing instruction as a hard constraint and the dynamic target shape reflected by the original video at other moments as a soft constraint, a joint optimization function is constructed, and the optimal adjustment value of the deformable mesh basic shape parameter and the deformation module calculation function network parameter is obtained by solving the minimum value of the function. The adjustment value is used to accurately modify the model parameters of the corresponding dynamic entity layer.
[0161] After parameter adjustment, the state is updated to ensure that the deformable mesh shape of the dynamic target corresponding to the user's editing intention can be calculated when the deformation module receives any time parameter in the future; then the time range of the video picture sequence to be generated is determined, and each target moment in this range is input into all updated deformation modules in turn to calculate the three-dimensional shape of each dynamic target at each target moment.
[0162] Then integrate the three-dimensional shape of all dynamic targets at each target moment with the fixed and unchanging environmental layer in the scene decomposition model to form a complete three-dimensional scene expression corresponding to each target moment; finally, according to the preset mapping relationship from the environmental layer to each video picture, project the complete three-dimensional scene expression of each target moment to the corresponding video picture view respectively, and after rendering processing, finally generate an updated multi-channel video picture sequence with completely matched content and editing intention, and smooth picture.
[0163] As an example, take the editing instruction of the No. 1 robot in the digital twin factory scene: rotate the arm 30° around the Z axis at t = 8 seconds:
[0164] First, the scene decomposition model receives and parses the structured editing instruction in step S1051, and accurately locates the dynamic entity layer corresponding to the No. 1 robot through the unique identifier ID: Robot-001 in the instruction.
[0165] Secondly, through step S1052, reverse parameter adjustment is carried out. First, parse the editing instruction and determine that the target shape of the No. 1 robot at t = 8 seconds is the grasping pose after the arm is rotated 30°, and take this shape as a hard constraint condition, and take the shape of the robot in the original video at other times before and after t = 8 seconds, such as t = 7 seconds and t = 9 seconds, as a soft constraint condition, and construct a joint optimization function. The joint optimization function is designed as:
[0166] (3)
[0167] In the formula, is the value of the joint optimization function; is the set of model parameters to be adjusted, including the basic shape parameters of the deformable mesh and the network parameters of the deformation module calculation function; is the hard constraint weight, which is 1.0 in value. Such a value can ensure that the hard constraint is satisfied first; is the soft constraint weight, which is 0.5 in value. This value can balance the overall motion law; is the hard constraint deviation, which specifically refers to the difference between the adjusted t = 8 seconds shape and the target shape; is the soft constraint deviation, which specifically refers to the difference between the adjusted shape at other times and the original shape.
[0168] In order to solve the optimal solution of the above joint optimization function and determine the adjustment value of the model parameter, the specific calculation process is as follows: assuming that under the original parameters, the value of is 0.8, which represents a large shape deviation; the value of is 0.2, which represents the original motion law. Based on these values, the original function value is calculated as 1.0 multiplied by 0.8 plus 0.5 multiplied by 0.2, and the final original function value is . By minimizing using the gradient descent algorithm, the optimal parameter is obtained. At this time the value of is 0.01, which represents that the adjusted shape basically fits the target shape; 0.25, which represents a slight deviation after adjustment but does not affect the overall motion rule. Based on these new values, the optimized function value is calculated as 1.0 multiplied by 0.01 plus 0.5 multiplied by 0.25, and finally The result meets the optimization requirement. Then, the solved adjustment value is used to synchronously update the base shape parameters of the deformable mesh and the parameters of the deformation module calculation function network, where the update of the base shape parameters of the deformable mesh is to correct the initial contour angle of the arm, and the update of the parameters of the deformation module calculation function network is to adjust the correspondence between time and rotation angle.
[0169] Then, step S1053 drives the deformation module state update based on the adjusted parameters, ensuring that no matter which time parameter is input, the deformation module can solve the robot shape that meets the editing intention of t=8 seconds and 30° rotation.
[0170] Then, step S1054 determines that the time range of the video picture sequence to be generated is t=5 seconds to t=15 seconds, and inputs each target time in the range to the updated deformation module in turn, which specifically includes t=5.0 seconds, t=5.1 seconds, and t=15.0 seconds. The three-dimensional shape of the robot at each time is solved by inputting these target times.
[0171] Then, step S1055 combines the three-dimensional shape of the robot at each target time with the environment layer in the scene decomposition model, which contains elements such as the factory building and fixed equipment, and forms a complete three-dimensional scene expression at each time after combination.
[0172] Finally, step S1056 projects the complete three-dimensional scene expression at each time to the video view of each camera according to the preset mapping relationship of the environment layer to each video picture, where the specific performance of the preset mapping relationship is that the east side area of the factory building corresponds to the view of camera 1, and the north side area of the factory building corresponds to the view of camera 3. Finally, a multi-channel video picture sequence with updated content is generated. In this sequence, the robot arm at t=8 seconds and the time before and after presents the rotated posture, and the motion rule at other times remains consistent with the original video.
[0173] The above example is only one example of the present application, and in actual application, the weights of the joint optimization function, the time range of the video to be generated, etc. can be adjusted according to the needs, which are not limited in the present application.
[0174] In another specific embodiment, for a scene edited simultaneously for multiple dynamic targets, the scene decomposition model can synchronously receive multiple structured editing instructions, locate to the corresponding dynamic entity layer through the unique identifier in the instruction, and carry out reverse parameter adjustment and deformation module state update in parallel. When constructing the joint optimization function, multi-target space interference constraints can be added to ensure that multiple dynamic targets do not overlap after adjustment, and then the three-dimensional shape of each target moment of multiple dynamic targets is solved synchronously to generate a multi-channel video picture sequence after projection of the combined environment layer. This way can adapt to the needs of multi-target collaborative editing, improve the interaction efficiency in complex scenes, and is suitable for large-scale digital twin factory, smart city and other multi-dynamic target scenes.
[0175] The present application realizes accurate reverse mapping of user editing operation to video content through the above steps, ensures that the updated multi-channel video picture sequence completely matches the editing intention, while guaranteeing the coherence of dynamic target motion law and the integrity of the scene, finally completes the bidirectional mapping interaction loop of video picture and three-dimensional model, and improves the accuracy and smoothness of digital twin scene interaction.
[0176] Figure 3 A specific embodiment of a video picture and three-dimensional model bidirectional mapping positioning and interaction system provided by the present application is shown in the structure diagram Figure 3 The system can include:
[0177] The acquisition module 31 is configured to acquire multi-channel video picture data of a real scene corresponding to a target digital twin scene, and a static three-dimensional white model corresponding to the target digital twin scene, wherein the multi-channel video picture data includes dynamic targets and static backgrounds.
[0178] The construction module 32 is configured to construct a scene decomposition model based on the multi-channel video picture data and the static three-dimensional white model, wherein the scene decomposition model includes an environment layer and at least one dynamic entity layer, and the dynamic entity layer is internally provided with a deformation module for calculating the motion trajectory and deformation of the dynamic target according to a time parameter.
[0179] The response module 33 is configured to respond to a specific moment selected by a user in the multi-channel video picture data, query the scene decomposition model, separate the dynamic target from the static background, and map and position the dynamic target from the video picture to the static three-dimensional white model for shape reconstruction to obtain a reconstructed target.
[0180] The response module 33 is further configured to generate an editing instruction in response to an editing operation performed by a user on the reconstructed target through a three-dimensional visualization interface, wherein the three-dimensional visualization interface is generated based on the static three-dimensional white model.
[0181] The generating module 34 is configured to adjust the model parameters of the dynamic entity layer corresponding to the reconstruction target according to the editing instruction, and drive the deforming module to update the state, so as to generate a content-updated multi-path video picture sequence based on the adjusted scene decomposition model.
[0182] The video picture and three-dimensional model bidirectional mapping positioning and interaction system of the embodiments of the present application is used to implement the aforementioned video picture and three-dimensional model bidirectional mapping positioning and interaction method, and therefore the specific embodiments of the video picture and three-dimensional model bidirectional mapping positioning and interaction system can be seen from the aforementioned embodiment part of the video picture and three-dimensional model bidirectional mapping positioning and interaction method, and the specific embodiments can be referred to the description of the corresponding embodiment part, which will not be described herein again.
[0183] The present application further provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the aforementioned video picture and three-dimensional model bidirectional mapping positioning and interaction method.
[0184] The present application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the aforementioned video picture and three-dimensional model bidirectional mapping positioning and interaction method.
[0185] In an exemplary embodiment, the aforementioned computer readable storage medium can include, but is not limited to, a U disk, a read-only memory, a random access memory, a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0186] The embodiments of the present application further provide a computer program product, wherein the aforementioned computer program product comprises a computer program, and the computer program is executed by a processor to implement the steps in the aforementioned video picture and three-dimensional model bidirectional mapping positioning and interaction method embodiments.
[0187] The skilled person can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in a general manner in the foregoing description. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0188] The above describes in detail the video picture and three-dimensional model bidirectional mapping positioning and interaction method and system provided by the present application. The principles and implementation manners of the present application are described by using specific examples, and the above description of the embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that, for those skilled in the art, some improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the present application.
Claims
1. A method for bidirectional mapping, positioning, and interaction between video footage and a 3D model, characterized in that, include: Acquire multiple video feeds of the real scene corresponding to the target digital twin scene, as well as a static 3D white model of the target digital twin scene. The multiple video feeds include dynamic targets and static backgrounds. Based on the multi-channel video image data and the static 3D white model, a scene decomposition model is constructed. The scene decomposition model includes an environment layer and at least one dynamic entity layer. The dynamic entity layer has a built-in deformation module, which is used to calculate the motion trajectory and deformation of the dynamic target according to the time parameter. In response to a specific moment selected by the user from the multi-channel video frame data, the scene decomposition model is queried to separate the dynamic target from the static background, and the dynamic target is mapped and located from the video frame to the static three-dimensional white model for shape reconstruction to obtain the reconstructed target; In response to the user's editing operation on the reconstructed target through a 3D visualization interface, an editing command is generated, wherein the 3D visualization interface is generated based on the static 3D white model; The scene decomposition model adjusts the model parameters of the dynamic entity layer corresponding to the reconstruction target according to the editing instructions, and drives the deformation module to perform state updates, so as to generate a multi-channel video frame sequence with updated content based on the adjusted scene decomposition model mapping. The scene decomposition model adjusts the model parameters of the dynamic entity layer corresponding to the reconstruction target according to the editing instructions, and drives the deformation module to perform state updates, so as to generate a multi-channel video frame sequence with updated content based on the adjusted scene decomposition model mapping, including: The scene decomposition model receives and parses the editing instructions, and locates the corresponding dynamic entity layer based on the editing instructions; Based on the editing instructions, the basic morphological parameters of the deformable mesh in the dynamic entity layer and the parameters of the computational function network inside the deformation module are adjusted in reverse. Based on the adjusted model parameters, the deformation module is driven to update its state, so that for any input time, the deformation module can calculate a deformable mesh shape that meets the editing intention based on the updated parameters. For each target moment in the video frame sequence to be generated, the target moment is input into each updated deformation module to calculate the three-dimensional shape of each dynamic target at the target moment; The three-dimensional form of each dynamic target is combined with the environment layer in the scene decomposition model to form a complete three-dimensional scene representation corresponding to the target at that moment; Based on the preset mapping relationship from the environment layer to each video frame, the complete 3D scene representation is projected onto the viewpoints of each video frame to generate a multi-video frame sequence with updated content.
2. The method according to claim 1, characterized in that, After constructing the scene decomposition model based on the multi-channel video image data and the static 3D white model, the method further includes: Cross-view collaborative optimization is performed on the scene decomposition model. The cross-view collaborative optimization includes: selecting at least two different viewing angles for the same dynamic target in the multiple video feeds; generating simulated appearances of the dynamic target under the at least two different viewing angles based on the scene decomposition model; calculating the difference between the simulated appearance and the appearance of the dynamic target in the real video feed under the corresponding viewing angles; and adjusting the parameters of the deformation module in the dynamic entity layer corresponding to the dynamic target so that the difference is reduced simultaneously under multi-view constraints, thereby optimizing the consistency between the motion trajectory and deformation calculation of the dynamic target in three-dimensional space.
3. The method according to claim 1, characterized in that, The step of reversely adjusting the basic morphological parameters of the deformable mesh in the dynamic solid layer and the parameters of the computational function network inside the deformation module according to the editing instructions includes: The editing instructions are analyzed to obtain the target form that the reconstruction target needs to achieve at the specific time. Using the target shape as the hard constraint at the specific moment and the original shape reflected by the multi-channel video image data at other moments as the soft constraint, a joint optimization function is constructed. By minimizing the joint optimization function, the adjustment values for the basic morphological parameters of the deformable mesh and the parameters of the computational function network inside the deformation module are obtained. Using the adjusted values, the basic morphological parameters and the parameters of the computational function network are updated synchronously.
4. The method according to claim 1, characterized in that, In response to a specific moment selected by the user from the multi-channel video frame data, the scene decomposition model is queried to separate the dynamic target from the static background, and the dynamic target is mapped and located from the video frame to the static 3D white model for shape reconstruction to obtain the reconstructed target, including: In response to the user's selection of a specific moment from the multi-channel video frame data; Query the scene decomposition model to obtain the deformable mesh morphology and surface appearance information of each dynamic entity layer calculated by the deformation module at the specific time. Based on the deformable mesh shape and surface appearance information, each dynamic target is separated from the static background corresponding to the multi-channel video image data; The shapes of the separated dynamic targets are mapped and located to the corresponding spatial regions of the static 3D white model according to the spatial association between the corresponding dynamic entity layer and the environment layer, and 3D shape reconstruction is performed. The reconstructed three-dimensional shape is defined as the reconstruction target.
5. The method according to claim 1, characterized in that, The construction of a scene decomposition model based on the multi-channel video image data and the static 3D white model includes: Identify at least one dynamic target in the multi-channel video frame data; Based on the static 3D white model, an environment layer representing a fixed scene structure is established; An independent dynamic entity layer is constructed for each dynamic target. The dynamic entity layer contains deformable mesh and surface appearance information, and the basic shape of the deformable mesh is associated with the spatial region of the dynamic target in the environment layer. A deformation module is set in each dynamic entity layer. The deformation module is a computational function network that receives time input and calculates the changes in deformable mesh morphology and surface appearance information. The environmental layer and the states of each dynamic entity layer are combined under time-driven conditions, and the images are synthesized according to a preset mapping relationship to construct a scene decomposition model.
6. The method according to claim 1, characterized in that, The method of generating editing instructions in response to the user's editing operation on the reconstructed target through a 3D visualization interface includes: Based on the geometric and texture data of the static 3D white model, a 3D visualization interface is generated and displayed through a 3D graphics rendering engine. In the 3D visualization interface, the reconstructed target is rendered and displayed as an interactive 3D entity; In response to a user-triggered editing operation, the system captures the corresponding interactive event and three-dimensional spatial transformation parameters. The editing operation is a spatial geometric editing operation triggered by an input device on the reconstruction target within the three-dimensional visualization interface. Based on the interaction event, the three-dimensional space transformation parameters, the unique identifier of the reconstructed target, and the specific time information, a structured editing instruction is generated. The editing instruction records the index, transformation type, and specific transformation amount of the set of vertices to be operated on. The specific transformation amount is represented by three-dimensional coordinates or a transformation matrix.
7. A bidirectional mapping positioning and interaction system between video footage and a 3D model, characterized in that, include: The acquisition module is used to acquire multiple video images of the real scene corresponding to the target digital twin scene, as well as a static 3D white model corresponding to the target digital twin scene. The multiple video images include dynamic targets and static backgrounds. The construction module is used to construct a scene decomposition model based on the multi-channel video image data and the static 3D white model. The scene decomposition model includes an environment layer and at least one dynamic entity layer. The dynamic entity layer has a built-in deformation module for calculating the motion trajectory and deformation of the dynamic target according to time parameters. The response module is used to respond to a specific moment selected by the user from the multi-channel video frame data, query the scene decomposition model, separate the dynamic target from the static background, and map and locate the dynamic target from the video frame to the static three-dimensional white model for shape reconstruction to obtain the reconstructed target; The response module is also used to generate editing instructions in response to the editing operation performed by the user on the reconstructed target through the three-dimensional visualization interface, wherein the three-dimensional visualization interface is generated based on the static three-dimensional white model; The generation module is used to adjust the model parameters of the dynamic entity layer corresponding to the reconstruction target according to the editing instructions of the scene decomposition model, and drive the deformation module to perform state updates, so as to generate a multi-channel video frame sequence with updated content based on the adjusted scene decomposition model mapping. The scene decomposition model adjusts the model parameters of the dynamic entity layer corresponding to the reconstruction target according to the editing instructions, and drives the deformation module to perform state updates, so as to generate a multi-channel video frame sequence with updated content based on the adjusted scene decomposition model mapping, including: The scene decomposition model receives and parses the editing instructions, and locates the corresponding dynamic entity layer based on the editing instructions; Based on the editing instructions, the basic morphological parameters of the deformable mesh in the dynamic entity layer and the parameters of the computational function network inside the deformation module are adjusted in reverse. Based on the adjusted model parameters, the deformation module is driven to update its state, so that for any input time, the deformation module can calculate a deformable mesh shape that meets the editing intention based on the updated parameters. For each target moment in the video frame sequence to be generated, the target moment is input into each updated deformation module to calculate the three-dimensional shape of each dynamic target at the target moment; The three-dimensional form of each dynamic target is combined with the environment layer in the scene decomposition model to form a complete three-dimensional scene representation corresponding to the target at that moment; Based on the preset mapping relationship from the environment layer to each video frame, the complete 3D scene representation is projected onto the viewpoints of each video frame to generate a multi-video frame sequence with updated content.
8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the steps of the bidirectional mapping, positioning, and interaction method between video footage and a 3D model as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, enables the bidirectional mapping, positioning, and interaction method between video footage and a 3D model as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Building intelligent teaching virtual simulation training room based on cloud computing system
CN106683502A
Construction method of intelligent assembly system based on digital twinning
CN114580083A