Digital twin video fusion method and device and electronic equipment

Through a single data acquisition device combined with the relative pose data of the digital twin model, the efficient fusion of digital twin videos is achieved, solving the problems of increasing costs and complex processing of multi-view camera systems, and improving the accuracy and reality of the fusion video.

CN120259092APending Publication Date: 2025-07-04GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410008975.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art requires multiple cameras or perspectives in digital twin video fusion, which increases hardware costs and maintenance costs, and the multi-view image data processing is complex, making it difficult to achieve real-time and accurate three-dimensional information fusion.

Method used

Using a single data acquisition device, by obtaining the relative pose data of the data acquisition device relative to the twin, combining the digital twin model, the synthesis of dynamic goals and the fusion of background images is achieved, and the display view is generated to generate a fusion video.

Benefits of technology

It simplifies equipment configuration and management, reduces costs, improves the accuracy and reality of three-dimensional information fusion, and avoids the complexity of multi-view image data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259092A_ABST
    Figure CN120259092A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video image processing, and discloses a digital twin video fusion method and device and electronic equipment. The method comprises the following steps: obtaining twin bodies; acquiring relative pose data of the data acquisition equipment relative to the twin body and first video data acquired by the data acquisition equipment; obtaining a dynamic target in the first video data; according to the relative pose data, synthesizing the dynamic target at a display view angle to generate a synthesized dynamic target view; generating a background picture under the display view angle; and fusing the background picture and the synthesized dynamic target view to obtain a display view so as to generate a fused video through the display view. According to the invention, the first video data is obtained through a single data acquisition device, the configuration and management of the device are simplified, and the cost is reduced; fusion and synthesis of the three-dimensional information of the video content are realized based on one data source, so that the data processing complexity is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video image processing, and particularly to a digital twin video fusion method, apparatus, and electronic device. Background Art

[0002] Digital twin video fusion is a method that combines digital twin technology with video fusion technology and is currently widely used in various fields. To achieve an ideal fusion result of digital twin videos, related technologies fuse the accurate three-dimensional information of the twin model and video content. Among them, a key point is to obtain the three-dimensional information of video content. For this purpose, related technologies use multiple cameras with different angles and installation positions to synchronously image the same video shooting scene to obtain multi-view images of the corresponding scene; then, based on the three-dimensional information implicit between these multi-view images, the video content is fused with the twin model.

[0003] The inventors found that the related technologies have at least the following problems in the process of implementing the embodiments of this application: Using multiple cameras or viewpoints requires additional equipment and facilities to support, increasing the hardware cost and maintenance cost. The image data generated by multiple cameras or viewpoints needs to be synchronized and fused, which involves technical challenges such as image alignment and time synchronization, and processing a large amount of multi-view image data also requires higher computing resources and algorithms to ensure real-time performance and accuracy. Summary of the Invention

[0004] The main technical problem to be solved by the embodiments of this application is to solve the problems of complex data processing and the increased cost caused by a multi-view camera system.

[0005] To solve the above technical problems, a technical solution adopted in an embodiment of the present application is: to provide a digital twin video fusion method, including: obtaining a twin body, where the twin body is obtained based on a digital twin model in a preset scenario; obtaining relative pose data of a data acquisition device in the preset scenario relative to the twin body, and first video data collected by the data acquisition device; obtaining dynamic targets in the first video data; synthesizing the dynamic targets in a display perspective according to the relative pose data to obtain a synthesized dynamic target view; generating a background picture in the display perspective; fusing the background picture and the dynamic target view to obtain a display view, so as to generate a fusion video through the display view. Among them, the first video data is obtained by a single data acquisition device, without the arrangement and synchronization of multiple data acquisition devices, thereby simplifying the device configuration and management and reducing the cost; using a digital twin model and a small number of data acquisition devices, the fusion and synthesis of three-dimensional information of video content are realized. This process avoids the use of multiple data acquisition devices. When using multiple data acquisition devices, it is necessary to ensure that the data collected by these data acquisition devices is at the same time point or within the same time range, which requires data synchronization and time calibration processing to align the data of each data acquisition device in time. In addition, it may also involve processing such as calibration. This process involves complex data processing. In contrast, in the method of using a single data acquisition device, since there is only one data source, the above complex data processing can be avoided or simplified.

[0006] Optionally, the step of obtaining the relative pose data of the data acquisition device in the preset scenario relative to the twin body includes: extracting first feature points from the image of the first video data, and extracting second feature points from the view of the twin body; performing stereo matching on the first feature points and the second feature points to obtain matched video features and twin body image features; extracting three-dimensional coordinates corresponding to the second feature points from the twin body according to the twin body image features; calculating the relative pose data according to the video features and the three-dimensional coordinates. Among them, the perspective of the data acquisition device can be corresponded to the coordinate system of the twin body, realizing the precise alignment of video content and the twin model, thereby helping to improve the accuracy and realism of digital twin video fusion in subsequent processes, making the generated fusion video more vivid and credible.

[0007] Optionally, the step of performing stereo matching on the first feature point and the second feature point to obtain the matched video features and twin image features includes: calculating a first feature descriptor corresponding to the first feature point and a second feature descriptor corresponding to the second feature point; comparing the similarity or distance between the first feature descriptor and the second feature descriptor to obtain the correspondence between the first feature point and the second feature point; performing matching screening on the correspondence and eliminating false matches to obtain the matched video features and twin image features. Among them, through operations such as feature descriptor extraction, similarity comparison, matching screening, and false match elimination, the accuracy and reliability of feature point matching can be improved, providing a reliable basis for subsequent data fusion and pose calculation.

[0008] Optionally, the step of obtaining the dynamic target in the first video data includes: obtaining the dynamic target category in the first video data; constructing an image segmentation data set according to the image data corresponding to the dynamic target category; training a preset image segmentation model through the image segmentation data set; and segmenting the dynamic target from the first video data through the trained image segmentation model. Among them, by obtaining the target category, constructing the image segmentation data set, training the image segmentation model, and applying the model for segmentation, the accurate segmentation of the dynamic target can be realized, providing an accurate target area for subsequent fusion and synthesis steps.

[0009] Optionally, the step of synthesizing the dynamic target in the display perspective according to the relative pose data to obtain the synthesized dynamic target view includes: obtaining image pairs of the dynamic target in different perspectives; calculating the relative pose of the image pairs, where the image pairs include a first image and a second image; adjusting a pre-trained conditional diffusion model, where the first image and the relative pose are used as conditional data to be input into the conditional diffusion model, and the output of the conditional diffusion model is the second image; and obtaining the view of the dynamic target in the display perspective through the adjusted conditional diffusion model. Among them, through operations such as perspective transformation, conditional diffusion model adjustment, and target view generation, the realistic synthesis of the dynamic target in the display perspective can be realized, which provides an accurate target view for the generation of the digital twin video.

[0010] Optionally, the step of obtaining the view of the dynamic target from the adjusted conditional diffusion model at the display perspective includes: obtaining the relative rotation and translation between the display perspective and the acquisition perspective of the data acquisition device; setting the relative rotation and translation, and the dynamic target as the conditional data of the adjusted conditional diffusion model, to obtain the view of the dynamic target at the display perspective output by the adjusted conditional diffusion model. Among them, by obtaining the relative rotation and translation, setting the conditional data, and generating the target view, an accurate and realistic target view can be generated, improving the realism and quality of the synthesis result.

[0011] Optionally, the step of fusing the background picture and the dynamic target view to obtain the display view includes: calculating the view position of the dynamic target at the display perspective according to the relative rotation and translation; fusing the background picture and the synthesized dynamic target view according to the view position to obtain the display view. Among them, by calculating the view position and performing the fusion operation, a realistic display view can be generated, providing a high-quality result for the display of the digital twin video.

[0012] To solve the above technical problems, another technical solution adopted in the embodiments of the present application is: providing a digital twin video fusion device, including: a model acquisition module for acquiring a twin body, which is obtained based on a digital twin model in a preset scenario; a data acquisition module for acquiring the relative pose data of a data acquisition device relative to the twin body in the preset scenario, and the first video data acquired by the data acquisition device; a dynamic target determination module for acquiring the dynamic target in the first video data; a dynamic target synthesis module for synthesizing the dynamic target at the display perspective according to the relative pose data to obtain a synthesized dynamic target view; a background picture determination module for generating a background picture at the display perspective; a video fusion module for fusing the background picture and the dynamic target view to obtain a display view, so as to generate a fusion video through the display view. Among them, the first video data is obtained by a single data acquisition device, without the arrangement and synchronization of multiple data acquisition devices, thus simplifying the device configuration and management and reducing the cost; using the digital twin model and a small number of data acquisition devices, the fusion and synthesis of three-dimensional information of video content are realized. This process avoids the use of multiple data acquisition devices. When using multiple data acquisition devices, it is necessary to ensure that the data collected by these data acquisition devices is at the same time point or within the same time range, which requires data synchronization and time calibration processing to align the data of each data acquisition device in time. Additionally, other processes such as calibration may also be involved. This process involves complex data processing. In contrast, in the method of using a single data acquisition device, since there is only one data source, the above complex data processing can be avoided or simplified.

[0013] To solve the above technical problems, another technical solution adopted in the embodiments of the present application is: to provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the digital twin video fusion method as described above. Among them, the first video data is obtained through a single data acquisition device, without the need for the arrangement and synchronization of multiple data acquisition devices, thus simplifying the device configuration and management and reducing costs; by using the digital twin model and a small number of data acquisition devices, the fusion and synthesis of three-dimensional information of video content are realized. This process avoids the use of multiple data acquisition devices. When using multiple data acquisition devices, it is necessary to ensure that the data collected by these data acquisition devices is at the same time point or within the same time range, which requires data synchronization and time calibration processing to align the data of each data acquisition device in time. Additionally, it may also involve processing such as correction. This process involves complex data processing. In contrast, in the method using a single data acquisition device, since there is only one data source, the above complex data processing can be avoided or simplified.

[0014] To solve the above technical problems, still another technical solution adopted in the embodiments of the present application is: to provide a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by an electronic device, the electronic device executes the digital twin video fusion method as described above. Among them, the first video data is obtained through a single data acquisition device, without the need for the arrangement and synchronization of multiple data acquisition devices, thus simplifying the device configuration and management and reducing costs; by using the digital twin model and a small number of data acquisition devices, the fusion and synthesis of three-dimensional information of video content are realized. This process avoids the use of multiple data acquisition devices. When using multiple data acquisition devices, it is necessary to ensure that the data collected by these data acquisition devices is at the same time point or within the same time range, which requires data synchronization and time calibration processing to align the data of each data acquisition device in time. Additionally, it may also involve processing such as correction. This process involves complex data processing. In contrast, in the method using a single data acquisition device, since there is only one data source, the above complex data processing can be avoided or simplified. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] One or more embodiments are exemplarily illustrated by corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings represent similar elements. Unless otherwise stated, the figures in the drawings do not constitute a scale limitation.

[0016] Figure 1 is a flowchart of a digital twin video fusion method provided by an embodiment of the present application;

[0017] Figure 2 is a flowchart of a method for obtaining relative pose data of a data acquisition device relative to the twin body provided by an embodiment of the present application;

[0018] Figure 3 is a flowchart of a method for performing stereo matching on the first feature point and the second feature point to obtain the video features and twin body image features after matching provided by an embodiment of the present application;

[0019] Figure 4 is a flowchart of a method for obtaining dynamic targets in the first video data provided by an embodiment of the present application;

[0020] Figure 5 is a flowchart of a method for synthesizing the dynamic targets in a display perspective according to the relative pose data to generate a synthesized dynamic target view provided by an embodiment of the present application;

[0021] Figure 6 is a data flow diagram of digital twin video fusion provided by an embodiment of the present application;

[0022] Figure 7 is a schematic structural diagram of a digital twin video fusion device provided by an embodiment of the present application;

[0023] Figure 8 is a schematic hardware structure diagram of an electronic device for executing the digital twin video fusion method provided by an embodiment of the present application. Detailed implementation manners

[0024] In order to make the objectives, technical solutions, and advantages of this application more clear and understandable, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application. It should be noted that if there is no conflict, the various features in the embodiments of this application can be combined with each other, and all are within the protection scope of this application. Additionally, although the functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division from that in the device schematic diagram or a different order from that in the flowchart. Unless otherwise defined, all the technical and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not used to limit this application.

[0025] Digital twin refers to a virtual model created through digital technology that can simulate, imitate, and predict physical entities or systems; it is a digital mirror of the real world and can construct an accurate representation of the actual system's behavior and performance by collecting and integrating various data sources, such as sensor data, physical models, historical data, etc.

[0026] The digital twin model is a virtual model for modeling and simulating the actual system, which can be constructed based on various methods, including physical modeling, data-driven modeling, and machine learning, etc. The digital twin model uses real-time data, physical models, and algorithms to simulate and predict the behavior, performance, and state changes of the actual system. The digital twin model can provide an in-depth understanding of the actual system and is used for optimization, prediction, and decision support. The twin is the real-time connection and synchronization between the digital twin model and the actual system. By integrating the digital twin model with the actual system, the twin can obtain the data of the actual system in real time and input it into the digital twin model for comparison and alignment. The twin can compare and verify the state and behavior of the actual system with the digital twin model, thereby realizing real-time monitoring, simulation, and optimization. The twin can also use the prediction ability of the digital twin model for virtual experiments and scenario simulations to support decision-making and system optimization. Therefore, the digital twin model is a virtual model for modeling and simulating the actual system, while the twin is a real-time instance that integrates and connects the digital twin model with the actual system.

[0027] Digital twin video fusion refers to the combination and integration of digital twin models with video data in the real world to generate more realistic and accurate simulation and analysis results. By aligning and fusing video data with digital twin models, virtual objects or scenes can be visually interacted and integrated with the real world in the real video. This can provide more intuitive and visual simulation and analysis results, enabling the behavior and performance of the model to be more accurately evaluated in the real environment.

[0028] In digital twin video fusion, the twin can be regarded as a real-time connection and synchronization instance between the digital twin model and the actual video data. For example, the twin is responsible for obtaining real-world video data in real time and synchronizing it with the digital twin model; the twin can also preprocess the collected video data to ensure its consistency with the data format and coordinate system of the digital twin model; the twin is also used to fuse and render virtual objects or scenes in the digital twin model with the video data to present the interaction effect between the real world and the virtual world in the video; etc. The twin plays a role of real-time connection and synchronization in digital twin video fusion, responsible for integrating the digital twin model with the actual video data to achieve more realistic and accurate simulation and analysis effects.

[0029] Digital twin video fusion can be applied to multiple fields and scenarios. It can be used for simulation, prediction, and analysis in various fields to provide more accurate and comprehensive data and visual information, helping to optimize decision-making and improve the performance of the system. For example, in the medical field, digital twin video fusion can combine virtual human models with real medical imaging video data, enabling surgical simulation and planning, assisting doctors in disease diagnosis and treatment decision-making, and improving the accuracy and safety of surgeries. In the field of architecture and urban planning, digital twin video fusion can combine virtual building models with video data in the real world, allowing the simulation of the behavior of buildings under different environmental conditions, evaluating the feasibility of design schemes, optimizing energy utilization, and improving performance in aspects such as air flow. In manufacturing, digital twin video fusion can combine virtual product design and production processes with actual production line video data, enabling real-time monitoring and analysis of the operation of the production line, detecting potential problems, performing fault diagnosis and prediction, and optimizing production efficiency and quality.

[0030] It can be seen that digital twin video fusion provides more realistic and accurate simulation and analysis results, enhances visualization and interactivity, improves fault diagnosis and prediction capabilities, and supports optimized decision-making and increased efficiency. These advantages make digital twin video fusion an important tool in various fields, providing strong support for the solution of practical problems and the improvement of systems.

[0031] However, some difficulties have been encountered in the process of implementing digital twin video fusion technology. One of the key steps is video fusion, that is, fusing video information with the twin model. The most intuitive way to implement this step is to establish a homography transformation relationship between the rendering view of the twin model and the video frame through feature matching and directly "overlay" the two frames together. However, due to the parallax between the two views, the fused frame obtained by such direct overlay often shows serious distortion.

[0032] Therefore, related technologies rely on the accurate three-dimensional information of the twin model and video content to achieve accurate three-dimensional structure fusion, use methods such as raster projection to fuse textures, and render the fused frame from the required viewing angle with the help of the graphics rendering process. This method can more intuitively process and display complete scene information. However, to accurately fuse video content with the twin model and display the fusion result with a large viewing angle difference (there is a large difference between the shooting angle and the viewing angle), the three-dimensional information of the content captured by the video needs to be known. In practical applications, since the scenes captured by videos are often dynamically changing, for example, compared with the twin model, dynamic objects such as pedestrians and vehicles may appear in the video, it is difficult to perform three-dimensional modeling on such scenes and directly obtain the accurate three-dimensional information of the video content.

[0033] Based on this, the solution adopted by related technologies is as follows: for the same video shooting scene, multiple cameras with different angles and installation positions are used to synchronously image to obtain multi-view images of the corresponding scene. Then, based on the three-dimensional information implicit in these multi-view images, the video content is fused with the twin model. However, the disadvantages of this method are obvious because in fact, most application scenarios cannot provide enough multi-view images, and even most scenes can only provide single-view images. When the number of shooting angles is limited (in the extreme case, only a single-view image of the scene), due to the lack of complete three-dimensional information of the scene, the directly fused result will show serious distortion and information loss.

[0034] Therefore, the embodiments of the present application provide a digital twin video fusion method and device, which can realize the effective fusion of the twin model and the real-time video only relying on the single-view image captured from a fixed viewing angle of the scene, and can also realize the video fusion scheme for displaying the fusion result with a large viewing angle difference. Among them, the fusion of the twin model and the real-time video is based on the fixed viewing angle of the scene, rather than relying on the input of multiple viewing angles or cameras. The use of such single-view images helps to simplify the complexity of data acquisition and processing and reduces the need for a multi-view camera system. The embodiments of the present application provide a simple and effective way to fuse the digital twin model with the real-time video to produce a satisfactory fusion result.

[0035] The fusion method of single - perspective images provided by the embodiments of this application is applicable to many scenarios, such as monitoring, virtual reality, augmented reality, etc., where only one perspective or camera is used. Specific application scenarios include, for example:

[0036] Virtual fitting: In the e - commerce or fashion industry, single - perspective images taken from a fixed perspective are used, combined with a twin model to simulate the effect of a customer trying on clothes in a virtual environment. By fusing the twin model with real - time video, the customer can see themselves wearing different clothes in the form of a virtual image on the screen, so as to better evaluate the style, color, and fit.

[0037] Virtual home decoration: In the field of interior design and home decoration, a single - perspective fusion method is used to fuse the twin model with real - time video to simulate the effects of different furniture, decorations, or layout schemes in the actual environment. Users can intuitively understand the impact of different decoration choices on the appearance and atmosphere of the room by observing the fused video, thus making better decisions.

[0038] Intelligent traffic management: In an intelligent traffic management system, a single - perspective fusion method is used to fuse the twin model with real - time video to simulate traffic flow, vehicle behavior, and road conditions. This can help traffic managers monitor and predict traffic congestion, accident risks, and road conditions, and support intelligent traffic decision - making and optimization.

[0039] In the above - mentioned application scenarios, the hardware devices that may be involved include: cameras, which are used to capture fixed - perspective images or real - time videos of the scene, and can be ordinary cameras or camera devices specialized for applications such as virtual reality and augmented reality. Computing devices, which are used to perform computing tasks such as image processing, vision algorithms, and model inference, and can be computers, embedded systems, cloud servers, etc. Input devices, which are used to interact with the system. For example, in the virtual fitting scenario, input devices such as touchscreens, mice, or gamepads can be used. Output devices, which are used to display the synthesized video results or virtual scenes, and can be computer monitors, TV screens, projectors, etc. In some scenarios, audio devices may also be included. For example, in virtual home decoration, ambient sound effects or user - interaction sounds can be added. Among them, the computing device can be communicatively connected to the camera, input device, output device, audio device, etc.

[0040] In practical applications, appropriate hardware devices can be selected and configured according to needs to support the implementation of the single - perspective fusion scheme. The digital twin video fusion method in the following embodiments can be specifically executed by the computing device.

[0041] Please refer to Figure 1 , Figure 1 which is a flowchart of a digital twin video fusion method provided by the embodiments of this application. The method includes:

[0042] S11. Obtain a twin, which is obtained based on a digital twin model in a preset scenario.

[0043] A digital twin model is a virtual model that models and simulates an actual system. It can be constructed based on methods such as real-time data, physical models, and machine learning. The digital twin model is an independent virtual entity. By modeling and simulating the actual system, it can provide predictions and analyses of the system's behavior, performance, and state changes. A twin is the real-time connection and synchronization between the digital twin model and the actual system. It is a system that integrates the digital twin model with the actual system to achieve real-time data interaction and information transfer.

[0044] The process of obtaining a twin may include: First, construct a digital twin model according to the characteristics, behaviors, or performance of the actual system in the preset scenario. This model can be based on physical equations, statistical models, machine learning models, or deep learning models, etc. In the preset scenario, collect data of the actual system, which can include sensor data, operating status data, environmental parameters, and other data related to the actual system. Preprocess the collected data, including data cleaning, denoising, calibration, etc., to ensure the quality and accuracy of the data. Next, use the preprocessed data as input and train it with the digital twin model in the preset scenario. The goal of training is to adjust the model's parameters to better fit the behavior of the actual system, which can be achieved through methods such as supervised learning, unsupervised learning, or reinforcement learning. Then, verify and evaluate the trained digital twin model. You can use a part of independent data to verify the model and check its accuracy and reliability in the preset scenario. Adjust and improve the model according to the verification results. When the digital twin model has been trained and verified, it can be connected and synchronized with the actual system in real time to generate a real-time twin. This can be achieved by inputting the data of the actual system into the digital twin model and using the model's prediction results to reflect the behavior and performance of the actual system. Among them, obtaining a twin is an iterative process. By continuously collecting more data, improving the model, and conducting verification, the accuracy and reliability of the twin can be gradually optimized.

[0045] Specifically, for example, the twin model of an outdoor open scene can be constructed by using multi-view stereo reconstruction algorithms through oblique photography technology, or can also be reconstructed based on NeRF (Neural Radiance Fields); indoor scenes can be realized with the help of various 3D scanning technologies, such as the SLAM (Simultaneous Localization and Mapping) mapping solution that combines vision and lidar, structured light 3D scanning, multi-view stereo vision, or NeRF reconstruction and other methods. These methods have their own advantages and disadvantages, as long as the models they construct can be converted into twin representations such as point clouds and textures, or surface patches, they can meet the video fusion requirements. For different scenarios and application requirements, appropriate methods can be selected according to specific situations to construct digital twin models and convert them into twins to ensure that the video fusion requirements can be met.

[0046] The preset scene can be a virtual or real scene set according to specific needs or purposes. These scenes can be the actual environments in the real world, or virtual environments or simulation environments. For example, a preset scene can be a digital twin model of a factory or production line, used for optimizing the production process, predicting equipment failures, conducting virtual on-site training, etc.

[0047] S12. Obtain the relative pose data of the data acquisition device with respect to the twin in the preset scene, and the first video data collected by the data acquisition device.

[0048] The relative pose data refers to the position and attitude information of the data acquisition device with respect to the twin, which describes the translation and rotation relationship of the data acquisition device with respect to the twin. According to the relative pose data, the position of the data acquisition device in the twin coordinate system can be obtained. For example, information such as the horizontal displacement, vertical displacement, and depth displacement of the data acquisition device with respect to the twin. The relative pose data can be obtained through calibration techniques, camera pose estimation algorithms, or other sensor fusion technologies. After obtaining the relative pose data of the data acquisition device with respect to the twin in the preset scene, this relative pose data can be used to convert the video content collected by the data acquisition device into the coordinate space where the twin is located. In this way, the video content can be aligned with the model of the twin, making the simulation and analysis more accurate. According to the relative pose data, the attitude of the data acquisition device in the twin coordinate system can also be obtained. For example, the rotation angle or direction of the data acquisition device with respect to the twin, such as pitch, yaw, and roll.

[0049] In this embodiment, by obtaining the pose of the data acquisition device in the twin coordinate system, the perspective of the video data can be correctly positioned so as to match and fuse it with the model of the twin. In subsequent video fusion and simulation processes, the relative pose data can be used to align and fuse video data from other perspectives with the twin. Thus, video data from different perspectives can be synthesized and simulated in the same scene coordinate system to generate a richer and more realistic display view.

[0050] The first video data can be the video data of a single perspective captured by the data acquisition device. This single perspective can be the perspective that can be captured by the initial position and orientation of the data acquisition device, which is obtained at the beginning of the acquisition process and is used to initialize the twin and establish the correspondence with the actual scene. By acquiring the first video data, the scene information captured by the data acquisition device at the initial position and pose can be obtained. The video data of this single perspective can be used as a reference to establish the correspondence between the digital twin model and the actual scene.

[0051] The data acquisition device refers to a device used to capture video images or video sequences in the real world. It can be a camera, webcam, drone payload, mobile device, etc., depending on the application scenario and requirements. The data acquisition device usually captures optical images or videos through sensors and converts them into digital signals for processing and storage.

[0052] Among them, please refer to Figure 2 , the step of obtaining the relative pose data of the data acquisition device relative to the twin in the preset scene includes:

[0053] S121. Extract first feature points from the image of the first video data, and extract second feature points from the view of the twin;

[0054] S122. Perform stereo matching on the first feature points and the second feature points to obtain the matched video features and twin image features;

[0055] S123. Extract the three-dimensional coordinates corresponding to the second feature points from the twin according to the twin image features;

[0056] S124. Calculate the relative pose data according to the video features and the three-dimensional coordinates.

[0057] In the embodiment of the present application, the Perspective-n-Point (PnP) method is used to solve the pose of the data acquisition device relative to the twin.

[0058] Among them, the first feature points can use SIFT (Scale-Invariant Feature Transform) features or feature extraction methods based on deep learning to extract the coordinate FC of the feature points and calculate the descriptors of these feature points. Similarly, the same type of features FD are extracted from the view of the twin or the constructed twin image, and the descriptors are calculated. When using SIFT or other feature extraction methods, FC represents the position coordinates of the extracted feature points in the image of the first video data. Similar to FC, FD represents the position coordinates of the feature points extracted from the twin image in the image. The descriptor is a vector or feature vector that encodes the local image information in the area around the feature point, and it captures key information such as texture, shape, and gradient in the area around the feature point. In SIFT or deep learning methods, the descriptor is used to represent the feature information of the feature point for matching and recognition. In this embodiment, SIFT or a method based on deep learning is used to extract the coordinate FC of the feature points in the image of the first video data and calculate the descriptors of these feature points; then the coordinate FD of the same type of feature points is extracted from the view of the twin or the constructed twin image, and the descriptors are calculated; by matching these feature points and descriptors, stereo matching can be performed and the pose of the data acquisition device can be further solved. Among them, the data acquisition device can specifically be a camera, etc.

[0059] Next, stereo matching is performed on the above two sets of features FC and FD to obtain the matched video features FCM and the twin image features FDM, that is, by matching the feature points, the corresponding relationship between the image of the first video data and the twin image is found. Among them, FCM and FDM respectively represent the matched video features and the twin image features. Specifically, FCM represents the feature coordinates of the matched feature points in the first video data, and FDM represents the feature coordinates of the matched feature points in the twin image. Among them, please refer to Figure 3 , step S122 includes:

[0060] S1221. Calculate the first feature descriptor corresponding to the first feature point and the second feature descriptor corresponding to the second feature point.

[0061] Among them, for each feature point, according to the selected method (such as SIFT or a method based on deep learning), calculate its corresponding feature descriptor, and the calculated feature descriptor is used to represent the key information in the area around the feature point.

[0062] S1222. Compare the similarity or distance between the first feature descriptor and the second feature descriptor to obtain the corresponding relationship between the first feature point and the second feature point.

[0063] In this embodiment, the similarity or distance between feature descriptors can be compared by calculating the Euclidean distance, Hamming distance, cosine similarity, etc. During the feature matching process, the above calculation methods can help determine the correspondence between two feature points. By comparing the similarity or distance between feature descriptors, the best matching pairs can be selected, and inaccurate and false matches can be excluded.

[0064] S1223. Perform matching screening on the correspondence and eliminate false matches to obtain the matched video features and the twin image features.

[0065] Among them, the best matching pairs can be selected from the matching pairs by using methods such as threshold-based screening, RANSAC (RANdom SAmple Consensus) algorithm, etc. to remove inaccurate and false matches.

[0066] By performing the above operations of feature descriptor extraction, similarity comparison, matching screening, and false match elimination, the accuracy and reliability of feature point matching can be improved. Through the above steps, the matched video features FCM and the twin image features FDM can be obtained, and these features represent the correspondence found between the first video data and the twin image.

[0067] Next, according to the above feature matching results, use the matchable twin image features FDM to extract the 3D coordinates P3D corresponding to these feature points in the twin model. These 3D coordinates represent points in the twin coordinate system. Then, based on the obtained matched video features FCM of the data acquisition device and the corresponding 3D coordinates P3D, combined with the internal parameter matrix of the data acquisition device, use the PnP algorithm to solve the relative pose data of the data acquisition device with respect to the twin. Among them, the PnP algorithm is a method for estimating the pose of the data acquisition device through known 3D-2D point pairs. Here, through the matched video features FCM of the data acquisition device and the corresponding 3D coordinates P3D, combined with the internal parameter matrix of the data acquisition device, 3D-2D point pairs can be constructed.

[0068] It can be known that the data acquisition device and the twin are in different coordinate systems. Therefore, it is necessary to determine the correspondence between them to achieve accurate matching and fusion of data. In digital twin video fusion, the data acquisition device usually records video data in the real world, while the twin is a virtual representation of a real-world object or scene. To align the two, it is necessary to correspond the perspective of the data acquisition device with the coordinate system of the twin model. Therefore, in this embodiment, the perspective of the data acquisition device is corresponded with the coordinate system of the twin to achieve accurate alignment of the video content and the twin model, which helps to improve the accuracy and realism of digital twin video fusion in the subsequent process.

[0069] S13. Obtain the dynamic targets in the first video data.

[0070] In this embodiment, by obtaining the target categories, constructing an image segmentation dataset, training an image segmentation model, and applying the model for segmentation, the accurate segmentation of dynamic targets can be achieved, providing an accurate target area for subsequent fusion and synthesis steps.

[0071] Specifically, please refer to Figure 4 , the segmenting and obtaining the dynamic targets in the first video data includes:

[0072] S131. Obtain the dynamic target categories in the first video data.

[0073] Target detection, motion analysis, or other related methods can be used to detect and identify the dynamic targets in the first video data and assign corresponding category labels to the dynamic targets. For example, in a traffic surveillance video, by using a target detection model, different dynamic targets such as cars, pedestrians, and bicycles can be detected and identified, and corresponding category labels can be assigned to them. For example, in a sports competition video, through motion analysis, the motion trajectories of athletes can be extracted, and the athletes can be classified into different categories according to their motion patterns and behaviors, such as football players, basketball players, etc. For example, in a video surveillance system, by using a deep learning model, pedestrians can be detected and identified, and category labels such as adults and children can be assigned to them.

[0074] There are many methods in the related art that can be used to detect and identify the dynamic targets in the first video data and assign corresponding category labels to them. In detail, they will not be elaborated here.

[0075] S132. Construct an image segmentation dataset according to the image data corresponding to the dynamic target categories.

[0076] According to the obtained dynamic target categories above, extract the image data related to each dynamic target category from the first video data, and these image data will be used to construct the training set for image segmentation.

[0077] S133. Train a preset image segmentation model through the image segmentation dataset.

[0078] For example, use deep learning methods such as convolutional neural networks (CNNs) or semantic segmentation models to construct a model capable of segmenting images. Use the image segmentation dataset constructed in step S132 to train this model so that it can accurately segment the dynamic targets.

[0079] S134. Segment the dynamic target from the first video data using the trained image segmentation model.

[0080] Use the trained image segmentation model to perform image segmentation operations on each frame of the first video data. Among them, through the inference process of the model, the regions corresponding to the dynamic target in the image can be segmented, so as to obtain the segmentation result of the dynamic target in the first video data.

[0081] S14. Synthesize the dynamic target in the display perspective according to the relative pose data to obtain the synthesized dynamic target view.

[0082] Among them, the display perspective refers to the perspective selected when generating the synthesized dynamic target view, which can be multiple different perspectives. When generating the synthesized dynamic target view, different perspectives for observing or presenting the target can be selected to obtain synthesized images from multiple perspectives. The selection of multiple perspectives can be carried out according to needs to meet specific application requirements. For example, in a virtual reality environment, synthesized images from multiple perspectives can be generated to provide a more realistic and immersive viewing experience.

[0083] Obtain the relative pose data according to the above steps. According to the relative pose data, the pose of the data acquisition device in the twin coordinate system can be obtained. In this embodiment, the data acquisition device obtains the first video data based on a fixed perspective and does not directly obtain the three-dimensional information of the target. Therefore, in this embodiment, based on the relative pose data, the pre-trained conditional diffusion model is adjusted, and then based on the adjusted conditional diffusion model, an image of the dynamic target in a new perspective is generated. Specifically, please refer to Figure 5 The synthesizing the dynamic target in the display perspective according to the relative pose data to obtain the synthesized dynamic target view includes:

[0084] S141. Obtain image pairs of the dynamic target from different perspectives and calculate the relative pose of the image pairs, where the image pairs include a first image and a second image.

[0085] An image pair refers to two images with different perspectives captured by the data acquisition device, namely the first image and the second image. The same dynamic target can be photographed or observed multiple times to obtain multiple sets of images from different perspectives or angles. Calculating the relative pose of the image pair specifically means calculating the relative pose between the first image and the second image, that is, the relative position and orientation relationship with respect to the display perspective.

[0086] Among them, the relative pose between the first image and the second image is calculated based on the relative pose data. The relative pose data is the pose of the data acquisition device in the twin coordinate system, which is also the pose of the data acquisition device in the world coordinate system, and this describes the position and orientation of the data acquisition device. Then, the internal parameters of the data acquisition device are obtained, such as the focal length of the camera, the coordinates of the principal point, and the distortion parameters, etc. These parameters are used to convert the pixel coordinates on the image into three-dimensional points in the camera coordinate system. Next, some three-dimensional points that are commonly visible in the first image and the second image are selected and paired with their corresponding pixel coordinates, which can be achieved by using a feature point matching algorithm. Finally, using the three-dimensional to two-dimensional correspondence relationship, combined with the internal parameters of the data acquisition device and the pose of the data acquisition device, the Perspective-n-Point (PnP) algorithm or other pose estimation algorithms can be used to calculate the relative pose between the first image and the second image.

[0087] S142. Adjust the pre-trained conditional diffusion model, where the first image and the relative pose are input into the conditional diffusion model as conditional data, and the output of the conditional diffusion model is the second image.

[0088] The pre-trained conditional diffusion model is obtained by training on a large amount of image data and can generate synthetic images. This conditional diffusion model can learn the characteristics and distributions of the image data during the training process, enabling it to generate images with reasonable characteristics.

[0089] After training the conditional diffusion model, it is adjusted. Adjustment means further adjusting the parameters and weights of the model on the basis of the already trained conditional diffusion model to make it adapt to specific tasks or new conditions. In this embodiment, the stable diffusion model (i.e., the pre-trained conditional diffusion model) has been obtained through training on a large amount of image data; then, by using the first image and the relative pose as conditional data, the stable diffusion model is adjusted to better generate images that meet the requirements of specific perspectives and poses. In this embodiment, by using the first image and the relative pose as conditional data to adjust the stable diffusion model, it can be made to generate images that meet the requirements of specific perspectives and poses.

[0090] S143. Obtain the view of the dynamic target in the display perspective through the adjusted conditional diffusion model.

[0091] Among them, obtaining the view of the dynamic target from the display perspective by the adjusted conditional diffusion model includes: obtaining the relative rotation and translation between the display perspective and the acquisition perspective of the data acquisition device; setting the relative rotation and translation, and the dynamic target as the conditional data of the adjusted conditional diffusion model, and obtaining the view of the dynamic target from the display perspective output by the adjusted conditional diffusion model.

[0092] First, obtain the relative rotation and translation relationship between the display perspective (i.e., the desired viewing angle) and the perspective acquired by the data acquisition device. This relative rotation and translation relationship can be obtained through sensor data or other calibration methods. Then, input the image of the dynamic target and the relative rotation and translation as conditional data into the adjusted conditional diffusion model. The conditional diffusion model will use these conditional data to generate and output the view of the dynamic target from the display perspective. The output view will meet the requirements of the expected display perspective and be able to represent the dynamic target that matches the input conditions.

[0093] For example, the currently acquired video contains a moving car. The car is used as the dynamic target. To obtain the view of the car from the display perspective, where the display perspective is the desired viewing angle of the car and can include one or more, first, measure the rotation and translation of the data acquisition device relative to the display perspective through sensor data or calibration methods. For example, an inertial measurement unit (IMU) or other sensors can be used to measure the attitude of the data acquisition device, that is, the rotation of the data acquisition device relative to gravity and the geomagnetic field. This measurement result is used to calculate the relative rotation and translation relationship between the coordinate system of the data acquisition device and the display perspective. Then, at the input stage, provide the image of the car as the input and the relative rotation and translation relationship as the conditional data. The adjusted conditional diffusion model will generate the view of the car from the display perspective based on these inputs. Finally, through this process, a view of the dynamic target from the display perspective that meets the requirements of the expected display perspective and matches the input conditions can be obtained.

[0094] Among them, at the input stage, providing the image of the car as the input and the relative rotation and translation relationship as the conditional data, and the adjusted conditional diffusion model generating the view of the car from the display perspective based on these inputs can specifically include:

[0095] First, prepare an image dataset containing cars and the relative rotation and translation relationships associated with each image, which can be obtained through sensor data or calibration methods. At the same time, a conditional diffusion model for adjustment needs to be prepared, and this model has been trained in the pre-training stage. Then, according to specific requirements, define different viewing perspectives, such as the front, rear, left, right, etc. Next, select a viewing perspective, select a car image from the dataset as the input. At the same time, according to the selected viewing perspective and the corresponding relative rotation and translation relationships, input the image and conditional data into the adjusted conditional diffusion model. Through the inference process of the model, the adjusted conditional diffusion model will utilize the input car image and the relative rotation and translation relationships corresponding to the viewing perspective to generate a view of the car under the viewing perspective. Next, different viewing perspectives can be selected as needed, and the above steps of inputting data and conditional settings, and generating a view of the car under the viewing perspective can be repeated to generate views of the car under different viewing perspectives.

[0096] For example: Suppose there is a dataset containing car images and the relative rotation and translation relationships associated with each image. We hope to generate views of the car under the front and left viewing perspectives. First, perform data preparation, that is, prepare a dataset containing car images and relative rotation and translation relationships, and an adjusted conditional diffusion model. Then define the viewing perspectives, that is, select the front and left as the viewing perspectives. Next, input data and conditional settings, that is, select a car image from the dataset, and at the same time, according to the relative rotation and translation relationships of the front viewing perspective, input the image and conditional data into the adjusted conditional diffusion model. Through the inference process of the model, the adjusted conditional diffusion model utilizes the input car image and the relative rotation and translation relationships of the front viewing perspective to generate a view of the car under the front viewing perspective.

[0097] Next, select the left viewing perspective and repeat the above process to generate a view of the car under the left viewing perspective.

[0098] Through this example, we can generate views of the car under the corresponding perspectives using the adjusted conditional diffusion model based on different viewing perspectives. Repeating this process can generate multiple views under different viewing perspectives to meet specific requirements. This embodiment can generate accurate and realistic target views, improving the realism and quality of the synthesis results.

[0099] S15. Generate a background picture under the said viewing perspective, wherein, the background picture is rendered based on the digital twin model.

[0100] Among them, the static background is basically the same in appearance when the twin model is modeled and the data acquisition device takes pictures. Therefore, any perspective image of the static background can be directly rendered based on the twin model. The process of rendering the background picture based on the digital twin model may include: using digital twin technology to model the static background, including obtaining the three-dimensional geometric shape, texture information, and other relevant attributes of the background. Using the digital twin model, select the required display perspective, and then use rendering technology to generate an image from this perspective. For example, it can be achieved by setting corresponding camera parameters (such as perspective, position, and orientation) on the digital twin model and performing rendering. According to the image from the rendered perspective, generate an image of the static background. For example, perform post-processing on the rendering result, such as color correction, lighting adjustment, or other image processing techniques, to make the generated image more in line with the requirements of the expected display perspective.

[0101] S16. Merge the background picture and the synthesized dynamic target view to obtain a display view, and generate a fused video through the display view.

[0102] Specifically, this step is to fuse the static background and the dynamic target into the same view. The merging of the background picture and the synthesized dynamic target view to obtain a display view includes: calculating the view position of the dynamic target in the display perspective according to the relative rotation and translation; fusing the background picture and the synthesized dynamic target view according to the view position to obtain the display view. Among them, using the above relative rotation and translation relationships, calculate the view position of the dynamic target in the display perspective according to the requirements of the display perspective. This process involves converting the position and pose of the dynamic target into the coordinate system of the display perspective. Among them, according to the calculated view position of the dynamic target, fuse the background picture and the synthesized dynamic target view. The specific fusion method can be image synthesis technology, such as superimposing, blending, or performing blending mode operations on the two images to make them blend naturally, so as to obtain a display view containing the static background and the dynamic target. Among them, by continuously playing the generated display view at a certain frame rate, a video integrating the static background and the dynamic target can be generated. Through calculating the view position and performing the fusion operation in this embodiment, a realistic display view can be generated, providing high-quality results for the display of digital twin videos.

[0103] Based on the above embodiments of the digital twin video fusion method, please refer to Figure 6 , Figure 6 which is the data flow diagram of digital twin video fusion provided by the embodiments of the present application. As Figure 6As shown in the figure, a digital twin model and a twin body are constructed according to the digital twin modeling module, and the twin body is obtained based on the digital twin model in a preset scenario; first video data captured by a data acquisition device in the preset scenario from a single perspective is obtained according to the single perspective image acquisition module; then, the relative pose calculation module calculates the relative pose data of the data acquisition device relative to the twin body, and at the same time, dynamic targets are obtained by segmenting the first video data; then, the dynamic targets are synthesized in the display perspective according to the relative pose data to generate a synthesized dynamic target view; at the same time, a background image in the display perspective is rendered based on the digital twin model; finally, the background image and the synthesized dynamic target view are fused to obtain a display view, and a fused video is generated through the display view.

[0104] In the embodiment of the present application, a target view with a reasonable appearance can be generated according to input conditions (such as rotation and translation relationships). By fusing the generated target view with a static background, a fusion result can be obtained without three-dimensional information acquired by multiple data acquisition devices. Real-time online fusion can also be achieved. The above-mentioned conditional diffusion model is a model that has been adjusted and can instantaneously generate a target view when requested. This real-time nature enables the video fusion process from different perspectives to be carried out quickly without recalculating complex three-dimensional information for each frame. In addition, it can also be compatible with multi-perspective situations. When adjusting the conditional diffusion model, the number of views and relative poses used as conditional inputs can be changed as needed. For example, by increasing or decreasing the number of views and adjusting the relative pose settings, different multi-perspective fusion requirements can be met.

[0105] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a digital twin video fusion device provided by an embodiment of the present application. The digital twin video fusion device 20 includes:

[0106] A model acquisition module 201 for acquiring a twin body, where the twin body is obtained based on a digital twin model in a preset scenario; a data acquisition module 202 for acquiring the relative pose data of the data acquisition device relative to the twin body in the preset scenario, and the first video data acquired by the data acquisition device; a dynamic target determination module 203 for acquiring dynamic targets in the first video data; a dynamic target synthesis module 204 for synthesizing the dynamic targets in the display perspective according to the relative pose data to obtain a synthesized dynamic target view; a background image determination module 205 for generating a background image in the display perspective, where the background image is rendered based on the digital twin model; and a video fusion module 206 for fusing the background image and the synthesized dynamic target view to obtain a display view, so as to generate a fused video through the display view.

[0107] It should be noted that the above digital twin video fusion device can execute the digital twin video fusion method provided in the embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in the embodiments of the digital twin video fusion device, reference can be made to the digital twin video fusion method provided in the embodiments of the present application.

[0108] Figure 8 It is a schematic hardware structure diagram of an electronic device 30 for executing the digital twin video fusion method provided in the embodiments of the present application, as Figure 8 shown. The electronic device 30 includes:

[0109] One or more processors 301 and a memory 302. Figure 8 Taking one processor 301 as an example. The processor 301 and the memory 302 can be connected through a bus or other means. Figure 8 Taking the connection through the bus as an example.

[0110] The memory 302, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the digital twin video fusion method in the embodiments of the present application. The processor 301 executes various functional applications and data processing of the electronic device by running the non-volatile software programs, instructions, and modules stored in the memory 302, that is, implementing the digital twin video fusion method in the above method embodiments.

[0111] The memory 302 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the digital twin video fusion device, etc. In addition, the memory 302 may include a high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 302 may optionally include a memory remotely set relative to the processor 301, and these remote memories can be connected to the digital twin video fusion device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0112] The one or more modules are stored in the memory 302 and, when executed by the one or more processors 301, execute the digital twin video fusion method in any of the above method embodiments. For example, execute the Figures 1 to 5 method steps described above, and implement the Figure 7 functions of the modules in

[0113] The above-mentioned product can execute the method provided by the embodiment of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiment of the present application.

[0114] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to: ultra-mobile personal computer devices, mobile communication devices, servers, and other electronic devices with data interaction functions.

[0115] The embodiments of the present application provide a non-volatile computer-readable storage medium, which stores computer-executable instructions. These computer-executable instructions are executed by one or more processors, such as Figure 8 one of the processors 301 in [], enabling the above-mentioned one or more processors to execute the digital twin video fusion method in any of the above method embodiments. For example, execute the Figures 1 to 5 method steps described above, and implement the Figure 7 functions of the modules in [].

[0116] The embodiments of the present application provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by the electronic device, the electronic device can execute the digital twin video fusion method in any of the above method embodiments. For example, execute the Figures 1 to 5 method steps described above, and implement the Figure 7 functions of the modules in [].

[0117] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0118] Through the description of the above embodiments, those of ordinary skill in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other changes in different aspects of the present application as described above. For the sake of brevity, they are not provided in detail; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A digital twin video fusion method, characterized in that, Including: Obtain a twin, which is obtained based on a digital twin model in a preset scenario; Obtain the relative pose data of the data acquisition device relative to the twin in the preset scenario, and the first video data collected by the data acquisition device; Obtain the dynamic targets in the first video data; Synthesize the dynamic targets in the display perspective according to the relative pose data to obtain a synthesized dynamic target view; Generate a background image in the display perspective; Fuse the background image and the dynamic target view to obtain a display view, so as to generate a fused video through the display view.

2. The digital twin video fusion method according to claim 1, wherein, The step of obtaining the relative pose data of the data acquisition device relative to the twin in the preset scenario includes: Extract first feature points from the images of the first video data, and extract second feature points from the view of the twin; Perform stereo matching on the first feature points and the second feature points to obtain matched video features and twin image features; According to the twin image features, extract the three-dimensional coordinates corresponding to the second feature points from the twin; Calculate the relative pose data according to the video features and the three-dimensional coordinates.

3. The digital twin video fusion method according to claim 2, wherein The step of performing stereo matching on the first feature points and the second feature points to obtain matched video features and twin image features includes: Calculate a first feature descriptor corresponding to the first feature points, and a second feature descriptor corresponding to the second feature points; Compare the similarity or distance between the first feature descriptor and the second feature descriptor to obtain the corresponding relationship between the first feature points and the second feature points; Perform matching screening on the corresponding relationship and eliminate false matches to obtain the matched video features and twin image features.

4. The digital twin video fusion method according to claim 1, wherein The step of obtaining the dynamic targets in the first video data includes: Obtain the dynamic target categories in the first video data; Construct an image segmentation data set according to the image data corresponding to the dynamic target categories; Train a preset image segmentation model through the image segmentation data set; Segment the dynamic targets from the first video data through the trained image segmentation model.

5. The digital twin video fusion method according to claim 1, wherein The step of synthesizing the dynamic targets in the display perspective according to the relative pose data to obtain a synthesized dynamic target view includes: Obtain image pairs of the dynamic targets in different perspectives; Calculate the relative pose of the image pairs, where the image pairs include a first image and a second image; Adjust a pre-trained conditional diffusion model, where the first image and the relative pose are input into the conditional diffusion model as conditional data, and the output of the conditional diffusion model is the second image; Obtain the view of the dynamic target in the display perspective through the adjusted conditional diffusion model.

6. The digital twin video fusion method according to claim 5, wherein, The step of obtaining the view of the dynamic target in the display perspective through the adjusted conditional diffusion model includes: Obtain the relative rotation and translation between the display perspective and the acquisition perspective of the data acquisition device; Set the relative rotation and translation, and use the dynamic target as the conditional data of the adjusted conditional diffusion model to obtain the view of the dynamic target output by the adjusted conditional diffusion model from the display perspective.

7. The digital twin video fusion method according to claim 6, wherein The step of fusing the background image and the view of the dynamic target to obtain the display view includes: Calculate the view position of the dynamic target from the display perspective according to the relative rotation and translation; Fuse the background image and the synthesized view of the dynamic target according to the view position to obtain the display view.

8. A digital twin video fusion device, characterized in that, It includes: A model acquisition module for acquiring a twin, which is obtained based on a digital twin model in a preset scenario; A data acquisition module for acquiring the relative pose data of a data acquisition device in the preset scenario with respect to the twin, and the first video data collected by the data acquisition device; A dynamic target determination module for acquiring the dynamic target in the first video data; A dynamic target synthesis module for synthesizing the dynamic target from the display perspective according to the relative pose data to obtain a synthesized view of the dynamic target; A background image determination module for generating a background image from the display perspective; A video fusion module for fusing the background image and the view of the dynamic target to obtain a display view, so as to generate a fused video through the display view.

9. An electronic device, characterized in that, It includes: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the digital twin video fusion method according to any one of claims 1-7.

10. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by an electronic device, the electronic device executes the digital twin video fusion method according to any one of claims 1-7.