Video stream and digital twin scene fusion method, terminal equipment and storage medium
By binding the physical camera and specific viewing angles in the digital twin scene, embed the video stream as the surface texture, and dynamically adjusting the texture attributes, the problem of poor splitting and correlation after the video and digital twin scene is solved, and the user experience and scene realism are improved.
Patent Information
- Application Number
- CN202510317970.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-10
AI Technical Summary
In the prior art, video and digital twin scenes have a sense of separation and poor correlation, and are complex in operation, which affects the user's immersion and operation convenience.
By binding a specific viewing angle and physical camera, select a specific viewing angle in the digital twin scene to retrieve the video stream, embed the video stream as a surface texture into the digital twin scene, and parse the video stream decoded frames, extract parameter information and dynamically adjust the texture attributes.
It improves the overall aesthetics and reality of the digital twin scene after video fusion, solves the problem of poor sense of separation and correlation, and improves the user experience.
Smart Images

Figure CN120128679A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital twins, and particularly to a method for integrating video streams with digital twin scenarios, a terminal device, and a storage medium. Background Art
[0002] The application of existing technical solutions can, to a certain extent, enable users to view the video surveillance footage related to the current spatial location in real time. However, it is necessary to click on the camera icon button to open a form to play the video, resulting in adverse effects such as a sense of disconnection, poor relevance, and complex operations after the video is integrated with the digital twin scenario. For example, when viewing the video, it is difficult for users to correspond the content in the video with the actual location in the digital twin scenario, thereby reducing the user's immersion and operational convenience; or, due to obvious differences in visual style, color matching, etc. between the video and the digital twin scenario, the overall aesthetics and realism of the digital twin scenario after video integration are reduced. This greatly affects the user experience. Summary of the Invention
[0003] To address the above problems, this application provides a method for integrating a video stream with a digital twin scenario, including:
[0004] S1: Bind at least one specific perspective and the physical camera corresponding to the specific perspective;
[0005] S2: Select the specific perspective in the digital twin scenario and retrieve the video stream of the physical camera;
[0006] S3: Embed the video stream as a surface texture into the digital twin scenario;
[0007] S4: Analyze the decoded frames of the video stream, extract the parameter information of the current decoded frame, and dynamically adjust the texture attributes of the digital twin scenario.
[0008] Further, step S1 further includes:
[0009] S11: Adjust the virtual camera perspective in the digital twin scenario to the specific perspective in the interaction interface to match the physical camera perspective;
[0010] S12: Store the virtual camera position information, target point information, and video surveillance point information of the physical camera at the specific perspective;
[0011] S13: Bind the specific perspective and the physical camera.
[0012] Further, selecting the specific perspective in step S2 includes selecting the position information and the target point information of the virtual camera.
[0013] Furthermore, step S3 further includes:
[0014] S31: Align the object position and pose of any video frame in the video stream with the digital twin model based on the projection mapping algorithm;
[0015] S32: Render the video frame using a 3D engine;
[0016] S33: Perform distortion correction and fusion registration on the rendered video frame through a script, and generate a surface texture to be embedded into the digital twin scene.
[0017] Furthermore, step S31 further includes using SLAM to perform spatial positioning on the objects in the video frame.
[0018] Furthermore, the object position and pose in step S31 include geometric feature points of the building's rigid structure.
[0019] Furthermore, step S4 extracts the parameter information of the current decoded frame in real time through a script based on the gray-level co-occurrence matrix algorithm.
[0020] Furthermore, the parameter information in step S4 includes the contrast, energy, correlation, and entropy of the texture of the decoded frame.
[0021] Furthermore, in step S4, the OpenGL graphics structure is called to dynamically adjust the texture attributes of the digital twin scene.
[0022] Furthermore, the texture attributes include material reflection and level of detail.
[0023] Furthermore, it further includes:
[0024] S5: Switch the specific perspective, and repeat steps S2 to S5 to observe the digital twin scene from different specific perspectives.
[0025] This application further provides a terminal device, which includes a memory, a processor, and a computer program stored on the processor and executable on the processor. When the computer program is executed by the processor, it implements the steps of the above-mentioned method for fusing a video stream with a digital twin scene.
[0026] This application further provides a computer-readable storage medium, on which a computer program is stored. It is characterized in that when the computer program is executed by a processor, it implements the steps of the above-mentioned method for fusing a video stream with a digital twin scene.
[0027] The present application relates to a method for integrating a video stream and a digital twin scenario, a terminal device, and a storage medium. First, at least one specific perspective and a physical camera corresponding to the specific perspective are bound; the specific perspective is selected in the digital twin scenario, and the video stream of the physical camera is retrieved; then, the video stream is embedded into the digital twin scenario as a surface texture; finally, the decoded frames of the video stream are parsed, the parameter information of the current decoded frame is extracted, and the texture attributes of the digital twin scenario are dynamically adjusted, effectively improving the overall aesthetics and realism of the digital twin scenario after video integration, solving the problem that the overall presentation shows a sense of fragmentation and poor correlation after the video and the digital twin scenario are integrated, and improving the user experience effect. Description of the Drawings
[0028] Figure 1 It is a flowchart of the method for integrating a video stream and a digital twin scenario of the present application;
[0029] Figure 2 It is a flowchart of implementing step S1 of the present application;
[0030] Figure 3 It is a flowchart of implementing step S3 of the present application;
[0031] Figure 4 It is a schematic diagram of a terminal device of the present application.
[0032] Description of the Reference Numerals
[0033] 1. Terminal device; 11. Memory; 12. Processor; 13. Computer program. Detailed Embodiment
[0034] In order to be able to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure will be described in detail below with reference to the drawings. The attached drawings are for reference and illustration only, and are not used to limit the embodiments of the present disclosure. In the following technical description, for the sake of explanation, multiple details are provided to provide a full understanding of the disclosed embodiments. However, one or more embodiments can still be implemented without these details. In other cases, well-known structures and devices can be shown in a simplified manner to simplify the drawings.
[0035] It should be noted that, without conflict, the embodiments in the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0036] To further understand the purpose, structure, features, and functions of the present application, the following is a detailed description in conjunction with the embodiments.
[0037] In view of the above problems, the present application provides a method for integrating a video stream and a digital twin scenario, including:
[0038] S1: Bind at least one specific perspective and the physical camera corresponding to the specific perspective;
[0039] S2: Select the specific perspective in the digital twin scenario and retrieve the video stream of the physical camera;
[0040] S3: Embed the video stream as a surface texture into the digital twin scenario;
[0041] S4: Analyze the decoded frames of the video stream, extract the parameter information of the current decoded frame, and dynamically adjust the texture attributes of the digital twin scenario.
[0042] Select one or more specific perspectives in the digital twin scenario according to actual needs, and bind one or more physical cameras to each specific perspective. The physical camera is a real camera, which can be a fixed monitoring camera or a mobile monitoring device.
[0043] In the digital twin scenario, select a specific perspective by clicking a button, dragging the mouse, or entering an instruction, etc. When a specific perspective is selected, the system immediately obtains the corresponding video stream of the physical camera bound to that perspective.
[0044] Process the retrieved video stream, and apply the processed video stream as a surface texture to the corresponding object or area in the digital twin scenario. To ensure that the mapping method of the texture matches the geometric shape and surface characteristics of the object, methods such as UV mapping and projection mapping are used to accurately fit the texture to the object surface.
[0045] Decode each frame in the video stream to generate decoded frames, analyze the decoded frames, and extract parameter information such as brightness, contrast, color, saturation, sharpness, etc. from the decoded frames through edge detection, motion detection, object recognition, etc. According to the extracted parameter information, dynamically adjust the texture attributes of the corresponding area in the digital twin scenario. For example, if the light in the video stream becomes darker, the system can correspondingly reduce the brightness of the corresponding area in the digital twin scenario to simulate the real light change; if the color in the video stream changes, the system can adjust the color mapping of the texture to maintain the color consistency of the scene.
[0046] In one embodiment, step S1 further includes:
[0047] S11: Adjust the virtual camera perspective in the digital twin scenario to the specific perspective in the interaction interface to match the physical camera perspective;
[0048] S12: Store the position information of the virtual camera, the target point information, and the video monitoring point information of the physical camera when storing the specific perspective;
[0049] S13: Bind the specific perspective and the physical camera.
[0050] By using a mouse, keyboard, or other input devices on the interaction interface of the digital twin scenario, adjust the perspective of the virtual camera. According to the actual installation location and monitoring range of the physical camera, adjust the perspective of the virtual camera to make it as consistent as possible with the perspective of the physical camera. This is achieved by comparing the virtual view in the digital twin scenario and the real-time video stream of the physical camera to ensure the perspective matching between the two.
[0051] When the perspective of the virtual camera is adjusted to the specific perspective, the system records and stores the position information of the virtual camera, including its coordinates and rotation angle in the digital twin scenario, and stores the target point information of the virtual camera, that is, the position information of the object or area directly in front of the current camera, namely, the focus information of the virtual camera. The virtual camera is the camera set in the digital twin scenario corresponding to the physical camera. At the same time, store the video monitoring point information of the physical camera corresponding to the specific perspective, including the installation location, monitoring range, identification information, etc. of the camera, to ensure that the video stream can be accurately mapped to the corresponding position in the digital twin scenario.
[0052] The system establishes a binding relationship between the specific perspective and the physical camera based on the stored virtual camera information and physical camera information. The binding between the specific perspective and the physical camera can be achieved through configuration files, databases, or direct API calls, etc.
[0053] Selecting the specific perspective in step S2 includes selecting the position information and the target point information of the virtual camera.
[0054] In the user interface of the digital twin scenario, find and select the previously stored virtual camera position information by browsing or searching, etc., and at the same time select the information related to the virtual camera target point, which helps to confirm the object or area directly in front of the virtual camera. The position information includes the precise coordinates and rotation angle of the virtual camera in the digital twin scenario, ensuring that the user can accurately locate the specific perspective. The target point information can be used as a reference to verify the correctness of the perspective and ensure that the selected perspective is consistent with the expectation.
[0055] The system searches for and obtains the identification information of the physical camera bound to this perspective according to the selected virtual camera position information and target point information of the user, establishes a connection with the physical camera using an appropriate communication protocol, and retrieves the real-time video stream.
[0056] Step S3 also includes:
[0057] S31: Align any video frame in the video stream with the object position and pose of the digital twin model based on the projection mapping algorithm;
[0058] S32: Render the video frame using a 3D engine;
[0059] S33: Perform distortion correction and fusion registration on the video frame after graphic rendering through a script, and generate a surface texture to be embedded into the digital twin scene.
[0060] In the digital twin scene, the projection mapping algorithm is used to accurately map the two-dimensional image of each frame in the video stream onto the corresponding three-dimensional object surface of the digital twin model. First, identify the target object in the digital twin model and obtain its position information and pose information. According to the perspective and parameters of the virtual camera, perform a projection transformation on each frame image in the video stream to make it perfectly aligned with the surface of the target object. During the alignment process, adjust parameters such as the scaling, rotation, and translation of the video to ensure that the video frame can accurately cover the surface of the target object.
[0061] The system inputs the video frame after projection mapping into the 3D engine. The 3D engine performs rendering processing on the mapped video frame according to parameters such as the material, lighting, and shadow of the digital twin model. During the rendering process, the 3D engine applies operations such as anti-aliasing, texture filtering, and lighting models for graphic rendering to enhance the visual effect and realism of the video frame.
[0062] Due to the limitations of the camera lens and imaging principle, there may be distortion phenomena in the video stream, such as barrel distortion, pincushion distortion, etc. The system performs distortion correction on the video frame after graphic rendering through a script to eliminate these distortion phenomena, and then performs color correction, brightness adjustment, edge fusion, etc. on the video frame through the fusion registration process to generate the final surface texture and embed it onto the corresponding object surface in the digital twin scene. This not only improves the accuracy and authenticity of the video frame but also ensures the consistency of the visual effect between the video frame and the digital twin model, enhancing the visualization effect and user experience of the digital twin scene.
[0063] Step S31 also includes using SLAM to perform spatial positioning on the objects in the video frame. After the projection mapping algorithm, if the alignment effect of the video frame with the object position and pose of the digital twin model is not ideal, perform spatial positioning on the objects in the video through SLAM. SLAM analyzes information such as feature points and edges in the video frame to identify the objects and determine their spatial positions and poses, extracts feature points in the video frame using computer vision algorithms, and matches them with the feature points in the digital twin model. Through feature matching, the transformation matrix of the objects in the video frame relative to the digital twin model can be calculated, including translation and rotation. To achieve seamless integration between the video frame and the digital twin model.
[0064] Among them, the object spatial position includes the x, y, and z axis coordinates, and the object pose includes the rotation angle.
[0065] The position and attitude of the object described in step S31 include the geometric feature points of the building's rigid structure.
[0066] In the video frame, computer vision algorithms are used to identify the geometric feature points of the building's rigid structure. These feature points can be parts with significant geometric features such as the corners, edge intersections, and positions of doors and windows of the building. For the digital twin model, corresponding geometric feature points are also extracted to ensure that these points correspond to the feature points in the video frame. The spatial position of the building's rigid structure in the video frame is determined, and its position and attitude in the three-dimensional space of the digital twin scenario are determined. During the projection mapping process, it is ensured that the feature points in the video frame are precisely aligned with the feature points in the digital twin model, thereby achieving the alignment of the overall image.
[0067] Step S4 extracts the parameter information of the current decoded frame in real time through a script based on the gray-level co-occurrence matrix algorithm.
[0068] Each frame of the video stream is obtained in real time from the video stream and decoded to generate a decoded frame. After parsing the decoded frame, the image data for analysis is obtained. Based on the gray-level co-occurrence matrix, by selecting appropriate distances and directions, the co-occurrence frequency between the pixel gray values of the decoded frame is statistically calculated. Then, parameter information is extracted from the constructed gray-level co-occurrence matrix to quantify the texture features of the image. The script utilizes an image processing library to simplify the process and ensures that each frame in the video stream can be processed in real time, outputting the parameter information of the corresponding decoded frame.
[0069] In some optional embodiments, the parameter information in step S4 includes the contrast, energy, correlation, and entropy of the texture of the decoded frame; among them, the contrast reflects the intensity of the contrast between the light and dark of the image texture, and the contrast is obtained by calculating the difference between different gray-level pixel pairs to reflect the clarity or roughness of the texture; the energy reflects the uniformity of the gray-level distribution of the image, and the energy is obtained by calculating the sum of the squares of the frequencies of the occurrence of pixel pairs to reflect the stability or smoothness of the image; the correlation represents the degree of linear correlation between the gray levels in the image, and the correlation is calculated by calculating the linear correlation of the gray values of pixel pairs to reflect the directionality or regularity of the texture; the entropy measures the complexity of the image texture, and the entropy is obtained by calculating the sum of the logarithms of the frequencies of the occurrence of pixel pairs to reflect the texture information content or chaos degree of the image.
[0070] In step S4, the OpenGL graphics structure is called to dynamically adjust the texture attributes of the digital twin scenario.
[0071] Among them, OpenGL is a cross - language and cross - platform graphics application programming interface used for rendering operations in 2D and 3D graphics environments. In the digital twin scenario, texture properties determine the appearance and details of the surfaces of objects in the scenario. By invoking the graphics structures of OpenGL, the texture properties in the digital twin scenario can be adjusted in real - time and dynamically to achieve a more realistic and vivid scenario effect.
[0072] The texture properties include material reflection and level of detail.
[0073] Material reflection determines how the surface of an object reflects light. By setting material properties to simulate the reflection characteristics of different materials, such as diffuse color, specular color, specular exponent, etc. Generally speaking, metal surfaces usually have higher specular reflection and specular exponent, while wooden or cloth surfaces may have lower specular reflection and softer specular effects. By adjusting the material reflection properties, the objects in the digital twin scenario can be made more realistic.
[0074] The level of detail refers to the resolution and level of detail of the texture at different distances or viewing angles. The dynamic adjustment of the level of detail is achieved by using the multi - level progressive texture technology of the texture. By reasonably setting and using the level of detail, high - quality visual effects can be provided while ensuring rendering performance.
[0075] In other alternative embodiments, the present application further includes:
[0076] S5: Switch the specific viewing angle, and repeat steps S2 to S5 to observe the digital twin scenario from different specific viewing angles.
[0077] The switching of the specific viewing angle can be achieved through keyboard input, mouse click, or touch screen, etc., or can be automatically performed by the system according to the preset viewing - angle switching logic. After switching to a new specific viewing angle, the system re - executes steps S2 to S5 to ensure that the digital twin scenario under the new specific viewing angle can be correctly rendered, and the texture properties can be dynamically adjusted as needed. By switching the specific viewing angle and repeating steps S2 to S5, the digital twin scenario can be observed from multiple different angles to obtain more comprehensive scenario information.
[0078] The present application further provides a terminal device. The terminal device 1 includes a memory 11, a processor 12, and a computer program 13 stored on the processor 12 and executable on the processor 12. When the computer program 13 is executed by the processor 12, the steps of the above - mentioned method for fusing the video stream and the digital twin scenario are implemented.
[0079] The present application further provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the steps of the above method for fusing a video stream with a digital twin scenario are implemented. The computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the above. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0080] The present application relates to a method for fusing a video stream with a digital twin scenario, a terminal device, and a storage medium. First, at least one specific perspective and a physical camera corresponding to the specific perspective are bound; in the digital twin scenario, the specific perspective is selected, and the video stream of the physical camera is retrieved; then, the video stream is embedded into the digital twin scenario as a surface texture; finally, the decoded frames of the video stream are parsed, the parameter information of the current decoded frame is extracted, and the texture attributes of the digital twin scenario are dynamically adjusted, effectively improving the overall aesthetics and realism of the digital twin scenario after video fusion, solving the problem that the overall presentation shows a sense of fragmentation and poor relevance after the video is fused with the digital twin scenario, and improving the user experience effect.
[0081] In the description of this specification, the descriptions with reference to the terms "an embodiment", "some embodiments", "specifically", or "optional embodiments", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0082] The present application has been described by the above related embodiments. However, the above embodiments are only examples for implementing the present application. It must be pointed out that the disclosed embodiments do not limit the scope of the present application. On the contrary, any modifications and refinements made without departing from the spirit and scope of the present application fall within the scope of patent protection of the present application.
Claims
1. A method for fusing video streams with digital twin scenes, characterized in that: include: S1: Binding at least one specific viewing angle and a physical camera corresponding to the specific viewing angle; S2: Select the specific viewing angle in the digital twin scene and retrieve the video stream of the physical camera; S3: embedding the video stream into the digital twin scene as a surface texture; S4: Parse the decoded frame of the video stream, extract parameter information of the current decoded frame and dynamically adjust the texture properties of the digital twin scene.
2. The method for fusing video stream and digital twin scene according to claim 1, characterized in that: Step S1 also includes: S11: adjusting the virtual camera perspective in the digital twin scene to the specific perspective in the interactive interface to match the physical camera perspective; S12: storing the virtual camera position information, target point information and video monitoring point information of the physical camera at the specific viewing angle; S13: Binding the specific viewing angle and the physical camera.
3. The method for fusing video stream and digital twin scene according to claim 2, characterized in that: Selecting the specific viewing angle in step S2 includes selecting the position information and the target point information of the virtual camera.
4. The method for fusing video stream and digital twin scene according to claim 1, characterized in that: Step S3 also includes: S31: aligning any video frame in the video stream with the object position and posture of the digital twin model based on a projection mapping algorithm; S32: Using a three-dimensional engine to perform graphics rendering on the video frame; S33: Perform distortion correction and fusion registration on the video frame after graphics rendering through a script, and generate surface texture to embed into the digital twin scene.
5. The method for fusing video stream and digital twin scene according to claim 4, characterized in that: Step S31 also includes using SLAM to spatially locate the object in the video frame.
6. The method for fusing video stream and digital twin scene according to claim 4, characterized in that: The object position and posture in step S31 include geometric feature points of the building's rigid structure.
7. The method for fusing video stream and digital twin scene according to claim 1, characterized in that: The step S4 extracts parameter information of the current decoded frame in real time based on the gray level co-occurrence matrix algorithm through a script.
8. The method for fusing video stream and digital twin scene according to claim 1, characterized in that: The parameter information in step S4 includes contrast, energy, correlation and entropy of the decoded frame texture.
9. The method for fusing video stream and digital twin scene according to claim 1, characterized in that: In step S4, the OpenGL graphics structure is called to dynamically adjust the texture properties of the digital twin scene.
10. The method for fusing video stream and digital twin scene according to claim 1, characterized in that: The texture properties include material reflectance and level of detail.
11. The method for fusing video stream and digital twin scene according to claim 1, characterized in that: Also includes: S5: Switch the specific perspective and repeat steps S2 to S5 to observe the digital twin scene from different specific perspectives.
12. A terminal device, characterized in that: The terminal device includes a memory, a processor, and a computer program stored on the processor and executable on the processor. When the computer program is executed by the processor, the steps of the method for fusing the video stream and the digital twin scene described in any one of claims 1 to 11 are implemented.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for fusing the video stream and the digital twin scene described in any one of claims 1 to 11 are implemented.
Citation Information
Cited By
Video stitching method and system based on digital twinborn scene, and computing device
CN120378688A
A video splicing method, system and computing device based on digital twin scene
CN120378688B