A video splicing method, system and computing device based on digital twin scene
By constructing spatial and temporal layouts in a digital twin environment and automatically adjusting the perspective and position of video clips, the problems of low efficiency and insufficient precision in traditional video stitching methods are solved, high-precision and automated video stitching is achieved, and the coherence and immersion of video content are improved.
Patent Information
- Application Number
- CN202510854417.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Traditional video stitching methods rely on manual intervention, are inefficient, and have difficulty achieving high-precision and automated processing. Especially when dealing with complex scene changes and multi-perspective transitions, image misalignment and time discontinuity are prone to occur.
By obtaining the real scene coordinates and timestamps of the video clips, a virtual scene is constructed in the digital twin environment, spatial layout and temporal layout are performed, visual features and scene semantic tone are identified, the perspective and position of the video clips are automatically adjusted, and a digital twin display video is generated.
It improves the automation and accuracy of video stitching, enhances visual continuity and immersion, improves the ability and efficiency of handling complex scenes, reduces labor costs, and is suitable for large-scale video content production.
Smart Images

Figure CN120378688B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of digital twin technology, and in particular to a video stitching method, system, and computing device based on a digital twin scene. Background Art
[0002] With the rapid development of information technology, the concept of digital twin, a fusion of advanced technologies such as the Internet of Things, big data, and artificial intelligence, is gaining increasing attention. Essentially a precise digital mirror of the physical world, a digital twin creates virtual replicas of physical objects, systems, or processes, enabling real-time interaction and dynamic simulation of data between the real and virtual worlds. In the field of video stitching, the application of digital twins has opened up new possibilities, particularly demonstrating significant advantages in reconstructing complex scenes and creating immersive experiences. By reproducing every detail of the real environment in a virtual space and combining data from the temporal dimension, digital twins can achieve precise synchronization and integration of video clips, greatly enhancing the coherence and realism of video content and providing viewers with unprecedented visual enjoyment and in-depth understanding.
[0003] Traditional video stitching methods mostly rely on basic image processing technology and video editing software. These methods usually include the following steps: first, manually select video clips shot by multiple cameras; then, use the alignment tools in the video editing software to roughly align them based on identifiable common features in the videos (such as the horizon, building edges, etc.); then, adjust the playback speed of the video clips or insert transition effects to ensure the consistency of the time series as much as possible; finally, perform color correction and image quality optimization to achieve visual coherence.
[0004] However, although these traditional methods meet the basic video stitching needs to a certain extent, their processes are cumbersome, inefficient, and often rely on manual intervention, making it difficult to achieve high-precision and automated processing. Summary of the Invention
[0005] The embodiments of the present application provide a video stitching method, system and computing device based on a digital twin scene to solve the problem of poor video stitching effect in the prior art.
[0006] In a first aspect, an embodiment of the present application provides a video stitching method based on a digital twin scene, comprising:
[0007] Obtaining a set of video clips to be spliced, wherein each video clip in the set of video clips carries corresponding real scene coordinates and a timestamp;
[0008] Constructing a virtual scene in the digital twin environment, and mapping the plurality of video clips into the virtual scene according to the real scene coordinates of the plurality of video clips to construct a spatial layout;
[0009] Adjusting the timing between the plurality of video clips according to the timestamps of the plurality of video clips to obtain an initial video sequence, and mapping the initial video sequence to the virtual scene to construct a time layout;
[0010] Constructing a spatiotemporal layout based on the spatial layout and the temporal layout, identifying visual features and scene semantic tones between the plurality of video clips in the spatiotemporal layout, and adjusting the perspectives and positions of the plurality of video clips based on the visual features and scene semantic tones between the plurality of video clips to generate a video clip adjustment result;
[0011] Based on the video clip adjustment results, multiple video clips are spliced to generate a digital twin display video.
[0012] Optionally, mapping the plurality of video clips into the virtual scene according to the real scene coordinates of the plurality of video clips to construct a spatial layout includes:
[0013] For each of the video clips, determining, according to the real scene coordinates of each of the video clips, the virtual scene coordinates corresponding to the real scene coordinates in the virtual scene, wherein the real scene and the virtual scene have a corresponding coordinate system;
[0014] Determining the position of each video clip in the virtual scene according to the virtual scene coordinates corresponding to each video clip;
[0015] determining a viewing angle of each of the video clips in the virtual scene;
[0016] A spatial layout is generated based on the positions and viewing angles of the plurality of video clips.
[0017] Optionally, determining the viewing angle of each video clip in the virtual scene includes:
[0018] For each of the video clips, calculating a geometric center point of the video clip in the virtual scene according to the virtual scene coordinates corresponding to the video clip and the coverage area of the video clip;
[0019] Determining the video content of the video clip and determining the sight direction vector of the video clip, wherein the video content at least includes: a subject movement direction and a camera pointing direction;
[0020] Determining physical properties of the video clip and determining an effective observation range of the video clip in the virtual scene, wherein the physical properties include at least: lens focal length;
[0021] The viewing angle of the video clip in the virtual scene is determined according to the geometric center point, the sight direction vector, and the focal length of the lens corresponding to the video clip.
[0022] Optionally, adjusting the timing between the plurality of video segments according to the timestamps of the plurality of video segments to obtain an initial video sequence includes:
[0023] For each of the video clips, extracting key frames of each of the video clips, and determining the visual features and scene semantic tone of the video clip based on the key frames of each of the video clips;
[0024] The timestamp, visual features and scene semantic tone of the video clip are taken as input and input into a pre-viewed sequential combing model to adjust the timing between the multiple video clips through the sequential combing model and output an initial video sequence, wherein the sequential combing model is trained based on multiple video clip samples, and the multiple video clip samples contain video sequence data marked with timestamps, visual features and scene semantic tone, as well as context information corresponding to the video sequence data.
[0025] Optionally, it also includes:
[0026] In the process of adjusting the timing between the plurality of video segments by the sequential combing model, if timestamps overlap between the plurality of video segments, determining the sequence priority between the plurality of video segments by determining the continuity and content length of the video content between the plurality of video segments; or
[0027] In the process of adjusting the timing between multiple video clips through the sequential combing model, if it is calculated that the time difference between any two video clips is less than the set time gap, an additional frame is generated between the last frame of the first video clip and the first frame of the next video clip of the two video clips to fill the time gap, and the timing between the two video clips is readjusted based on the additional frame.
[0028] Optionally, in the spatiotemporal layout, identifying visual features and scene semantic tones among the plurality of video clips, and adjusting the perspectives and positions of the plurality of video clips based on the visual features and scene semantic tones among the plurality of video clips to generate a video clip adjustment result, includes:
[0029] Determining the visual features and scene semantic tone of each video clip based on the key frames of the plurality of video clips;
[0030] In the spatiotemporal layout, based on the visual features of each of the video segments, identifying shared visual features between the plurality of video segments;
[0031] Calculating a visual rotation angle and a position translation amount between two adjacent video clips based on shared visual features between the plurality of video clips;
[0032] Determining the adjusted viewing angles and positions of the plurality of video clips based on the visual rotation angle and position translation between two adjacent video clips and the scene semantic tones corresponding to the two adjacent video clips;
[0033] A video segment adjustment result is generated based on the adjusted viewing angles and positions of the plurality of video segments.
[0034] Optionally, calculating the visual rotation angle between two adjacent video segments based on shared visual features between the plurality of video segments includes:
[0035] Determining shared visual features between two adjacent video clips, and constructing a shared visual feature point set corresponding to each video clip, so as to generate a shared visual feature set based on the shared visual feature point sets corresponding to the two adjacent video clips;
[0036] Calculating an initial rotation matrix based on the shared visual feature set, and determining a visual rotation angle between two adjacent video segments according to the initial rotation matrix;
[0037] determining adjusted viewing angles of the plurality of video clips based on a visual rotation angle between two adjacent video clips and a scene semantic tone corresponding to the two adjacent video clips;
[0038] Based on the adjusted rotation matrix, the adjusted viewing angles of the plurality of video clips are determined.
[0039] Optionally, calculating the position translation between two adjacent video clips based on shared visual features between the plurality of video clips includes:
[0040] Determining a shared visual feature between two adjacent video clips, and calculating an average displacement between the two adjacent video clips based on the shared visual feature between the two adjacent video clips, so as to determine a position translation amount between the two adjacent video clips;
[0041] Based on the position translation amount between two adjacent video segments and the scene semantic tones corresponding to the two adjacent video segments, the adjusted positions of the plurality of video segments are determined.
[0042] In a second aspect, an embodiment of the present application provides a video splicing system based on a digital twin scene, comprising:
[0043] An acquisition module, configured to acquire a set of video segments to be spliced, wherein each video segment in the set of video segments carries corresponding real scene coordinates and a timestamp;
[0044] A construction module is configured to construct a virtual scene in the digital twin environment, and map the plurality of video clips into the virtual scene according to the real scene coordinates of the plurality of video clips to construct a spatial layout; adjust the timing between the plurality of video clips according to the timestamps of the plurality of video clips to obtain an initial video sequence, and map the initial video sequence into the virtual scene to construct a temporal layout; and construct a spatiotemporal layout based on the spatial layout and the temporal layout;
[0045] A generation module is used to identify the visual features and scene semantic tone between the multiple video clips in the spatiotemporal layout, and adjust the perspectives and positions of the multiple video clips based on the visual features and scene semantic tone between the multiple video clips to generate video clip adjustment results; based on the video clip adjustment results, the multiple video clips are spliced to generate a digital twin display video.
[0046] In a third aspect, an embodiment of the present application provides a computing device comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the video stitching method based on the digital twin scene as described in the first aspect above.
[0047] In an embodiment of the present application, a set of video clips to be spliced is obtained, each video clip in the set of video clips carries corresponding real scene coordinates and timestamps; a virtual scene is constructed in a digital twin environment, and based on the real scene coordinates of the multiple video clips, the multiple video clips are mapped into the virtual scene to construct a spatial layout; based on the timestamps of the multiple video clips, the timing between the multiple video clips is adjusted to obtain an initial video sequence, and the initial video sequence is mapped into the virtual scene to construct a temporal layout; a spatiotemporal layout is constructed based on the spatial layout and the temporal layout, and in the spatiotemporal layout, the visual features and scene semantic tone between the multiple video clips are identified, and based on the visual features and scene semantic tone between the multiple video clips, the viewing angles and positions of the multiple video clips are adjusted to generate video clip adjustment results; based on the video clip adjustment results, the multiple video clips are spliced to generate a digital twin display video.
[0048] The video stitching method based on digital twin scenes proposed in this application has the following significant benefits compared to traditional video stitching technology:
[0049] High degree of automation and improved accuracy: This method significantly improves the automation of video stitching by automatically acquiring the real-world coordinates and timestamps of video clips and mapping them within the digital twin environment. The construction of spatial and temporal layouts ensures the precise positioning and timing of video clips in virtual space, reducing manual intervention and improving stitching accuracy.
[0050] Enhanced visual coherence and immersion: Based on the construction of spatiotemporal layout, this method can dynamically adjust the perspective and position of video clips through intelligent identification and analysis of the visual features and scene semantic tone between video clips to ensure visual smoothness and consistency of emotional expression, thereby enhancing the overall coherence of the final video and the audience's immersive experience.
[0051] Efficient processing of complex scenes: For video materials containing complex scene changes and multi-perspective conversions, this method can effectively respond to and accurately stitch them together by leveraging the powerful capabilities of digital twin technology. This solves the problems of image dislocation and time discontinuity that are prone to occur in traditional methods when processing such scenes, and improves the ability and efficiency of processing complex scenes.
[0052] Intelligent emotion matching and perspective optimization: By analyzing the scene semantic tone of video clips, this method can intelligently adjust the video display angle and sequence, so that the video content conveys a more expected emotional atmosphere, increasing the artistry and appeal of the video narrative.
[0053] Promoting large-scale content production: Due to its highly automated processing capabilities and effective management of complex scenarios, this method is very suitable for large-scale video content production, which can significantly improve production efficiency, reduce labor costs, and meet the modern media and entertainment industry's demand for high-quality and efficient video content.
[0054] In summary, the present invention not only innovates the technical means of video splicing, but also greatly enriches the creative potential of video content, providing advanced technical support for digital media, film and television production, virtual reality and other fields.
[0055] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0057] Figure 1 A flowchart of an embodiment of a video stitching method based on a digital twin scene provided in an embodiment of the present application;
[0058] Figure 2 A flowchart of another embodiment of a video stitching method based on a digital twin scene provided in an embodiment of the present application;
[0059] Figure 3 A schematic structural diagram of a video splicing system based on a digital twin scene provided in an embodiment of the present application;
[0060] Figure 4 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0061] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0062] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to being different types.
[0063] The technical solution of this application can be widely applied in various scenarios, including but not limited to the following aspects:
[0064] Film and television production and post-editing: In the post-production of movies, TV series, or commercials, this technology can efficiently integrate shots shot at different locations and times, maintain visual and emotional coherence, and accelerate the post-editing process.
[0065] Virtual Reality (VR) and Augmented Reality (AR) content creation: Creating immersive content for VR / AR experiences requires highly accurate spatial and temporal synchronization. This approach allows creators to seamlessly blend live-action video with virtual elements, enhancing the user experience.
[0066] News reporting and documentary production: In rapidly changing news scenes or complex documentary shooting environments, this technology can quickly organize video clips from multiple sources and automatically splice them according to timeline and spatial logic, speeding up news production and ensuring the timeliness and accuracy of reports.
[0067] Live broadcast and review of sports events: For large-scale sporting events, real-time video obtained from multiple camera angles can be quickly integrated through this method, and the perspective and sequence can be adjusted in real time according to the progress of the game, providing viewers with a comprehensive and coherent game viewing experience.
[0068] Urban planning and architectural design presentations: In urban planning or architectural design projects, combining drone footage with ground-based photography and utilizing digital twin technology for spatial layout can generate dynamic urban landscape or architectural tour videos, making it easier for decision makers and the public to understand the design proposals.
[0069] Education and Training: When producing interactive educational video content, this technology can accurately integrate multi-dimensional content such as theoretical explanations and practical operations, historical scenes and modern comparisons, thereby enhancing the appeal of teaching materials and improving teaching effectiveness.
[0070] Tourism promotion: The tourism industry can use this technology to integrate videos of scenic spots in different seasons and time periods to create tourism promotional videos that transcend time boundaries and present a comprehensive and vivid image of the destination to tourists.
[0071] The inventors' research has revealed that most traditional video stitching methods rely on basic image processing techniques and video editing software. These methods typically involve the following steps: first, manually selecting video clips captured by multiple cameras; then, using the alignment tools in the video editing software, roughly aligning the clips based on identifiable common features in the videos (such as the horizon and building edges); then, adjusting the playback speed of the clips or inserting transition effects to ensure temporal consistency; and finally, performing color correction and image quality optimization to achieve visual coherence. However, while these traditional methods meet basic video stitching requirements to a certain extent, they are cumbersome, inefficient, and often rely on manual intervention, making high-precision and automated processing difficult to achieve.
[0072] In other words, the shortcomings of existing video stitching technology are becoming increasingly prominent. On the one hand, due to the lack of intelligent scene understanding and spatiotemporal coordination mechanisms, traditional methods are prone to problems such as image dislocation and temporal discontinuity when dealing with video materials with complex scene changes, rapid perspective changes, or changing lighting conditions, which affect the viewing experience of the final video. On the other hand, the heavy demand for manual participation leads to a huge workload, which is difficult to adapt to the needs of large-scale video content production, especially in the contemporary media environment that pursues efficient and high-quality output. Therefore, exploring a new generation of video stitching methods combined with digital twin technology to achieve higher-precision automatic stitching and intelligent adjustment has become an important direction of current research.
[0073] In view of this, the present application provides a video stitching method based on a digital twin scene, which includes: obtaining a set of video clips to be stitched, each video clip in the set of video clips carries corresponding real scene coordinates and timestamps; constructing a virtual scene in a digital twin environment, and mapping multiple video clips into the virtual scene according to the real scene coordinates of multiple video clips to construct a spatial layout; adjusting the timing between multiple video clips according to the timestamps of multiple video clips to obtain an initial video sequence, and mapping the initial video sequence into the virtual scene to construct a temporal layout; constructing a spatiotemporal layout based on the spatial layout and the temporal layout, and identifying the visual features and scene semantic tone between multiple video clips in the spatiotemporal layout, and adjusting the perspectives and positions of multiple video clips based on the visual features and scene semantic tone between multiple video clips to generate video clip adjustment results; based on the video clip adjustment results, performing splicing processing on multiple video clips to generate a digital twin display video.
[0074] The method of the present application can solve the problems of image misalignment and time discontinuity that are prone to occur in traditional methods when processing such scenes. While improving the ability and efficiency of processing complex scenes, it can also improve video stitching efficiency and reduce labor costs.
[0075] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0076] Figure 1 A flowchart of an embodiment of a video splicing method based on a digital twin scene provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the method includes:
[0077] 101. Obtain a set of video clips to be spliced, where each video clip in the set of video clips carries corresponding real scene coordinates and a timestamp;
[0078] In this step, "obtaining a set of video clips to be stitched" refers to collecting a series of different clips that are intended to be merged into a single continuous video. These video clips come from different sources and may be real-world scenes shot at different times and using different devices. Each "video clip" is an independent video clip that contains a visual record of a portion of a specific scene. In particular, each video clip "carries corresponding real-world scene coordinates and timestamps", which means that in addition to the video content itself, each clip is also attached with geographic location information (real-world scene coordinates) and the exact time point (timestamp) when the clip was shot. Such additional information is crucial for subsequent accurate stitching because it allows the system to understand the relative position of each clip in physical space and time.
[0079] In this example, let's assume we're working on a cityscape documentary project. The project team uses drones and ground-based cameras to capture different areas of the city at different times of day and from multiple locations. Each camera is equipped with a GPS module to record the coordinates of the shooting location, and the cameras are precisely synchronized to ensure that all recorded footage is accurately timestamped.
[0080] The details of the embodiment are as follows:
[0081] Video Clip Collection: From morning to night, the team recorded sunrise scenes, busy streets, sunset riverfront scenes, and night scenes in the city's East District, Downtown, West District, and North District. A total of 20 independent high-definition video clips were collected, each approximately 30 seconds long.
[0082] Real-world coordinates: When recording each video clip, GPS automatically records the latitude and longitude of the shooting location. For example, the coordinates of the sunrise clip are 120.75° East longitude, 31.23° North latitude, while the coordinates of the night scene clip are 120.78° East longitude, 31.20° North latitude.
[0083] Timestamps: All devices are synchronized with a network time server, ensuring that each video clip has the exact time it was recorded. For example, a sunrise clip would be timestamped at 2023-04-01 06:00:00, while a night scene clip would be timestamped at 19:30:00 on the same day.
[0084] With these detailed geographic coordinates and timestamp information, the post-production team can use the method proposed in this application to accurately arrange these clips in spatial and temporal order, and finally splice them into a smooth all-weather landscape documentary showing the city from dawn to dusk.
[0085] 102. Construct a virtual scene in the digital twin environment, and map the multiple video clips into the virtual scene according to the real scene coordinates of the multiple video clips to construct a spatial layout;
[0086] In this step, the "digital twin environment" refers to a highly realistic virtual space that uses digital technology to replicate the real-world environment, including geography, physical properties, buildings, object locations, and even lighting and weather conditions. "Building a virtual scene" means creating a virtual model corresponding to the real world in such a digital twin environment, providing an interactive and operational platform for video stitching. "Mapping multiple video clips into the virtual scene" means using the real-world scene coordinate information carried by the video clips to accurately place each video clip in the corresponding position in the virtual space, achieving alignment and layout in physical space. The purpose of this is to establish spatial connections between video clips, laying the foundation for subsequent time series stitching and visual fluency.
[0087] In this example, let's assume we're producing a documentary about ancient ruins exploration. The team visits three ancient ruins located around a city at different times, capturing a video clip at each location. Each video clip not only captures the visual content of the ruins, but also includes the GPS coordinates and timestamp of the moment the footage was taken.
[0088] The implementation steps are as follows:
[0089] Building a digital twin environment: First, using 3D modeling software (such as Unity or Unreal Engine) along with laser scanning data and satellite map APIs, we create a virtual ancient city ruins environment that is as detailed and realistic as possible. This environment includes terrain, vegetation, roads, the exterior and interior structure of the ruins, and even simulates lighting conditions at different times of day, such as sunrise, afternoon, and dusk.
[0090] Video clips mapping virtual scenes:
[0091] For the first video clip, it was shot in the morning at the entrance of the ruins at the foot of the mountain, with the coordinates (12345.2°N, 11.5°E). This location was found in the digital twin environment, and the video clip was embedded so that its visuals matched the virtual entrance scene.
[0092] The second video was filmed in the afternoon at the central square of the ruins, at the coordinates (1.9°N, 1.8°E). The square was located in the virtual scene and the video clip was integrated to ensure that the light direction and shadows were consistent with the virtual environment.
[0093] For the final segment, the highest point of the ruins at dusk, at coordinates (1.5°N, 1.2°E), the video clip was precisely placed on the top of the virtual mountain, using time-matched dusk lighting to create an atmosphere consistent with that in the real video.
[0094] Through such mapping and layout, all video clips can be accurately positioned in the digital twin environment, forming a coherent spatial sequence, laying the foundation for the next step of splicing in time series and adjusting visual effects.
[0095] 103. Adjust the timing between the multiple video clips according to the timestamps of the multiple video clips to obtain an initial video sequence, and map the initial video sequence to the virtual scene to construct a time layout;
[0096] In this step, "adjusting the timing according to the timestamps of multiple video clips" means using the specific time information recorded when the video clips were recorded to determine their order on the timeline. This step is crucial to ensure that the video clips are spliced in the order in which they actually occurred, so that the narrative or scene presentation is logically coherent. "Obtaining the initial video sequence" means correctly playing the sequence of video clips sorted according to the timestamps. "Mapping the initial video sequence to the virtual scene" means mapping the logic of this time sequence to the previously constructed digital twin environment, so that the video clips are not only spatially positioned, but also temporally ordered, completing the time layout and laying a solid foundation for the final spatiotemporal continuity.
[0097] In this example, let's assume we're producing a time-lapse documentary about daily life in a city. The team captures multiple video clips from morning to night at various locations (e.g., a park, a cafe, an office building, a night market), each with a precise timestamp.
[0098] The implementation steps are as follows:
[0099] Timestamp analysis: We found the following timestamp information in all the collected video clips:
[0100] Clip A: 8:00 AM, morning joggers in the park - Clip B: 10:30 AM, conversations in a cafe - Clip C: 6:00 PM, people leaving the office - Clip D: 8:30 PM, a bustling night market;
[0101] Adjust timing and sorting: Based on these timestamps, determine the natural time sequence as A→B→C→D→D, which is a natural transition from morning activities to nighttime activities.
[0102] Mapping the time layout to the virtual scene: In the constructed digital twin city model, the video clips are placed in this time order: Clip A is placed in the park scene at 8:00 am, where the characters match the ambient lighting; Clip B moves to the cafe at 10:30 am, where the customer interaction is consistent with the indoor lighting; Clip C moves to the office area at 6:00 pm, where the crowd matches the sunset background; Clip D is placed in the night market at 8:30 pm, where the lighting blends in with the lively atmosphere of the market.
[0103] In this way, the video clips are not only accurately geographically located in the virtual space, but also correctly sorted according to the timestamp information to construct a time layout, ensuring the temporal and spatial coherence and logic of the video in the digital twin scene.
[0104] 104. Constructing a spatiotemporal layout based on the spatial layout and the temporal layout, identifying visual features and scene semantic tones among the plurality of video clips in the spatiotemporal layout, and adjusting the viewing angles and positions of the plurality of video clips based on the visual features and scene semantic tones among the plurality of video clips to generate a video clip adjustment result;
[0105] In this step, "spatial layout" refers to a comprehensive presentation framework formed by combining the spatial position information of the video clips (determined by the spatial layout) and the temporal sequence (determined by the temporal layout). This not only considers the three-dimensional coordinate position of the video clips in the virtual scene but also incorporates the temporal dimension, forming a four-dimensional spatiotemporal structure. "Identifying visual features and scene semantic tone" involves using algorithms to analyze the content of the video clips, automatically detecting visual features such as color, movement, and composition, as well as the scene semantic tone conveyed by the video (such as joy, tranquility, and tension). "Adjusting perspective and position" optimizes the video clip's display angle and specific placement within the virtual scene based on these identified video features and scene semantic tone, to enhance narrative coherence or enhance the audience's emotional experience.
[0106] In the embodiment of this application, a production process of a tourism promotion video is envisioned:
[0107] Construction of spatiotemporal layout: A three-dimensional model of a virtual tourist city is used as the basic scene. The video clips of various famous attractions are arranged in the order of playback and spatial position according to their actual shooting time (such as morning, noon, dusk) and geographical location.
[0108] Visual features and emotion recognition: Analyzed by AI algorithms:
[0109] Video clip 1 (beach at sunrise): The colors are warm, the picture is peaceful, and the semantic tone of the scene is tranquility and anticipation.
[0110] Video clip 2 (a bustling market at noon): rich in colors, strong in movement, and the semantic tone of the scene is lively and energetic.
[0111] Video clip 3 (ancient castle at sunset): The colors are soft, the composition is classical, and the semantic tone of the scene is nostalgic and romantic.
[0112] Viewing angle and position adjustment:
[0113] For video clip 1, the viewing angle was adjusted to a low angle to simulate the feeling of standing on the beach and watching the sunrise. The position was set at the coastline, allowing the audience to feel the gentleness of the first ray of sunshine.
[0114] For video clip 2, a bird's-eye view is used to show the overall picture of the market, emphasizing the bustling crowds. The location is set above the center of the market to increase the audience's immersive sense of participation.
[0115] Video clip 3 uses a long lens to slowly zoom in, gradually focusing from the panoramic view of the castle to a beam of light in front of the window, strengthening the semantic tone of the scene. The position is arranged on the dusk side of the castle, matching the surrounding environment.
[0116] Through such adjustments, the video clips are not only coherent and smooth in terms of time and space layout, but also more accurate and attractive in terms of visual and emotional expression. The final video clip adjustment results are more in line with the expected narrative rhythm and emotional expression.
[0117] 105. Based on the video clip adjustment result, multiple video clips are spliced together to generate a digital twin display video.
[0118] In this step, "video clip splicing processing" refers to the process of seamlessly connecting multiple adjusted video clips in a preset order and transition effects. This includes editing, adding transition effects, unifying colors and image quality, and synthesizing sound tracks to ensure that the final output video is smooth and natural, as if it were a work shot as a whole. "Digital twin display video" is a highly realistic and dynamic virtual model display method that uses digital technology to simulate objects, environments or processes in the real world or design concepts, allowing viewers to intuitively understand complex systems or experience scenes that are difficult to observe directly.
[0119] In this embodiment of the present application, it is assumed that a digital twin display video of a factory production line is being created:
[0120] Applying the adjusted video clips: The video clips prepared according to the previous steps cover every stage of the production process, from raw material warehousing, automated assembly line operations, quality inspection, to finished product packaging and shipment. Each stage of the video clip has been fine-tuned to the actual production line layout and timeline, ensuring consistent viewing angles, lighting, and color.
[0121] Stitching process: First, arrange the adjusted video clips in the logical order of the production process. Add smooth transition effects between adjacent clips, such as fade-in and fade-out or sliding switching, to ensure visual fluidity and no abruptness. Unify the color style and brightness of the entire video to ensure visual continuity from one scene to another. Synchronously add environmental sound effects and narration, such as the sound of running machinery, the sound of moving products, and instructions for key steps, to enhance the comprehensiveness of information conveyed. Finally, perform video rendering and output settings to ensure that the video resolution, frame rate, and encoding format are suitable for the target playback platform and device.
[0122] Through the above examples, the completed digital twin display video not only accurately reflects the actual operation of the factory production line, but also enables the audience to deeply understand and feel every detail of the production process through the dual presentation of vision and hearing. It is of great value for scenarios such as training, remote monitoring or customer demonstrations.
[0123] Optionally, in the embodiment of the present application, step 102 may specifically include:
[0124] 1021. For each of the video clips, determine, based on the real scene coordinates of each of the video clips, virtual scene coordinates corresponding to the real scene coordinates in the virtual scene, wherein the real scene and the virtual scene have a corresponding coordinate system;
[0125] 1022. Determine a position of each video clip in the virtual scene according to the virtual scene coordinates corresponding to each video clip;
[0126] 1023. Determine the viewing angle of each video clip in the virtual scene;
[0127] 1024. Generate a spatial layout based on the positions and viewing angles of the plurality of video clips.
[0128] In steps 1021-1024 above, real-world scene coordinates refer to the location information in the actual physical environment captured in the video clip. A three-dimensional coordinate system (X, Y, Z) is typically used to represent the precise position of an object or camera in the real world. A virtual scene is a computer-generated digital representation that mimics a real environment and contains various elements and layouts that correspond to the real world.
[0129] Among them, virtual scene coordinates refer to the corresponding positions assigned to real-world objects or viewpoints in the virtual environment to ensure that they can be correctly reproduced in the digital twin.
[0130] Among them, perspective refers to the direction and angle of the camera viewing the scene in the video clip, which is crucial for reconstructing the appearance of the real environment.
[0131] Among them, spatial layout involves how to reasonably arrange and display these video clips in the virtual scene to maintain the consistency of spatial logic and the viewer's sense of immersion.
[0132] In the embodiment of the present application, assuming that a digital twin model is being created for a large shopping mall, the following process may be involved:
[0133] Determining the correspondence between real-world coordinates and virtual-world coordinates (i.e., step 1021): First, the precise locations (real-world coordinates) of each surveillance camera within the shopping mall are obtained through GPS data and on-site measurements. A virtual scene of the shopping mall is then created in 3D modeling software, and these physical coordinates are converted to corresponding coordinate points in the virtual scene using the same coordinate system.
[0134] Determine the position of the video clip in the virtual scene (i.e., step 1022): Based on the converted virtual scene coordinates, accurately locate the video clip recorded by each surveillance camera to the corresponding monitoring point in the virtual shopping mall to ensure that the video image corresponds to the position in the virtual environment.
[0135] Determine the viewing angle of the video clip (i.e., step 1023): Analyze the shooting direction and field of view of each video clip and convert them into viewing angle parameters of the virtual camera, such as pitch angle, yaw angle, and roll angle, so that the virtual viewing angle matches the actual shooting viewing angle.
[0136] Generate spatial layout (i.e., step 1024): Comprehensively consider the location distribution and respective perspective information of all video clips, optimize the view switching logic in the virtual scene, and design a reasonable navigation path or automatic roaming route to ensure that users can smoothly shuttle between different shopping areas when browsing the digital twin display video, and experience the spatial continuity and realism just like a field trip.
[0137] Through the above embodiments, not only is accurate mapping from physical space to virtual space achieved, but an efficient and intuitive shopping mall digital twin experience platform is also provided for managers, designers or customers, facilitating asset management, layout planning or remote navigation.
[0138] Optionally, as a possible implementation scheme, the process of "determining the viewing angle of each video clip in the virtual scene" in step 1023 may include: for each video clip, according to the virtual scene coordinates corresponding to the video clip and the coverage area of the video clip, calculating the geometric center point of the video clip in the virtual scene; determining the video content of the video clip, determining the sight direction vector of the video clip, the video content at least including: the subject movement direction, the camera pointing; determining the physical properties of the video clip, determining the effective observation range of the video clip in the virtual scene, the physical properties at least including: the focal length of the lens; determining the viewing angle of the video clip in the virtual scene according to the geometric center point, the sight direction vector and the focal length of the lens corresponding to the video clip.
[0139] In this step, the geometric center point refers to the center coordinate point calculated based on the corresponding position of the video clip and its coverage area in the virtual scene. It is used as the basic position for the video perspective reference. The line of sight direction vector represents the direction of the camera in the video clip. It is determined by analyzing the direction of movement of the subject in the video content or the direct camera pointing angle. It is a directional quantity that helps understand the directionality of video observation. Video content analysis includes identifying the main movement direction or the direct direction of the camera in the video, which helps to infer the visual focus that the video clip is trying to capture. Physical properties, especially the focal length of the lens, determine the width of the camera's field of view (angle of view range). The shorter the focal length, the wider the field of view, and vice versa, it is narrower, affecting the range of the virtual scene that the video clip can cover. The effective observation range refers to the area of the virtual scene that the video clip can clearly cover at a given lens focal length. It is affected by the focal length and video quality.
[0140] In this embodiment of the present application, it is assumed that the viewing angle of a video surveillance video of an outdoor park is determined in a digital twin environment:
[0141] Computing the Geometric Center Point: A surveillance video clip shows a main road within a park. By analyzing the ground area covered by the video clip, the location of the surveillance point within the virtual park model is determined, and the geometric center coordinates of this surveillance point within the virtual scene are calculated.
[0142] Determining the gaze direction vector: Video analysis shows that the camera primarily tracks pedestrians along the main road, moving from west to east. Therefore, the gaze direction vector is set to due east, or (1,0,0) (assuming the virtual scene uses a right-handed coordinate system).
[0143] Determine physical properties: It is known that the surveillance camera uses a standard lens with a focal length of 35mm. Based on this focal length, its horizontal and vertical viewing angles can be calculated.
[0144] Determine the effective viewing range: Using trigonometric functions, we combine the geometric center point, the line of sight vector, and the focal length of the lens to calculate the maximum viewing radius and angular range of the camera in the virtual scene. For example, a 35mm lens provides approximately 63 degrees of horizontal viewing angle. This means that in the virtual environment, starting from the center point and following the line of sight vector, the camera can clearly cover a conical area with a specific radius centered at the center point.
[0145] Determine the viewing angle based on the above parameters: determine the viewing angle of the video clip in the virtual scene according to the geometric center point, the sight direction vector, and the focal length of the lens corresponding to the video clip.
[0146] Assume known conditions:
[0147] Geometric center point: , the three-dimensional coordinates of the video clip in the virtual scene.
[0148] View direction vector: , a unit vector, indicating the direction in which the camera is pointing, has been normalized to a unit vector, i.e. its modulus is 1.
[0149] Lens focal length: , the unit is usually in millimeters, which determines the viewing angle of the camera.
[0150] Effective viewing range: W, can be understood as the viewing angle width desired to be covered.
[0151] Calculation steps:
[0152] By formula: , determining the viewing angle of the video clip in the virtual scene according to the geometric center point, the sight direction vector, and the lens focal length corresponding to the video clip;
[0153] in, It is represented by the perspective of the video clip in the virtual scene, Expressed as the desired viewing angle width, is the focal length of the lens, and are the x and y components of the view direction vector, and arctan is the inverse tangent function.
[0154] It should be noted that the The multiplication by 2 is because the conversion from half-width to full-width of the field of view is usually calculated.
[0155] In summary, through these detailed steps, not only the observation point of the video clip in the virtual environment is precisely located, but also its actual observation angle and coverage are accurately simulated, making the digital twin environment closer to reality and helping to improve the accuracy of data analysis and decision-making.
[0156] Optionally, in the embodiment of the present application, the process of “adjusting the timing between the multiple video segments according to the timestamps of the multiple video segments to obtain the initial video sequence” in step 103 may include:
[0157] 1031. For each video segment, extract a key frame of each video segment, and determine the visual features and scene semantic tone of the video segment based on the key frame of each video segment;
[0158] 1032. The timestamp, visual features and scene semantic tone of the video clip are used as input to a previously viewed sequential combing model, so as to adjust the timing between the multiple video clips through the sequential combing model and output an initial video sequence, wherein the sequential combing model is trained based on multiple video clip samples, and the multiple video clip samples contain video sequence data marked with timestamps, visual features and scene semantic tone, as well as context information corresponding to the video sequence data.
[0159] In this step, keyframes are extracted, and visual features and the scene's semantic tone are determined. Keyframes are selected from the video clip to best represent the content of the clip. They are often used to quickly capture the video's visual summary. Visual features may include color distribution, texture, shape, and motion intensity, which help identify video content. The scene's semantic tone is the overall emotional atmosphere conveyed by the video clip, such as joy or sadness. It is typically determined through a comprehensive analysis of elements within the video, such as music, color, expression, and motion.
[0160] Sequential Combining Model: This is a machine learning or deep learning model trained to understand the relationship between contextual information, timestamps, visual features, and the semantic tone of a video clip to predict the optimal playback order. The model inputs include timestamps, visual features, and the semantic tone of the scene. Using this information, the model learns how video clips naturally transition based on emotional and temporal flow, thereby adjusting and sequencing the video clips.
[0161] In this embodiment of the present application, keyframe extraction and scene semantic tone determination are as follows: Assume that a travel vlog video is produced, and each video clip contains different scenery and character activities. First, an algorithm automatically identifies the keyframes in each clip, such as the iconic shot of the Eiffel Tower or the sunset shot at the seaside. Next, the color (such as the warm tones of the Eiffel Tower), the action (the smiles of the crowd), and the music (relaxing melodies) of these keyframes are analyzed to determine that the scene semantic tone is pleasant and relaxing.
[0162] Application of the Sequential Combining Model: Using a trained model, the model takes as input the timestamp of each clip (e.g., the Eiffel Tower at 2 p.m., a sunset on the beach in the evening), visual features (e.g., color distribution, motion intensity), and the semantic tone of the scene (e.g., joy, relaxation). Based on the trained sample database, the model understands how emotions and time transition smoothly across different contexts (e.g., from daytime activities to the tranquility of dusk), adjusts the order of the clips, and outputs an initial video playback sequence. The model database contains a large number of travel video samples, each labeled with time, visual features, scene semantic tone, and context (e.g., the relevance of scenes before and after the activity), helping the model learn how to naturally sequence the clips.
[0163] Through this implementation, video clips are not only reasonably sorted by time, but also intelligently arranged according to the fluency of emotions and visual content, which improves the coherence of the video and the viewing experience.
[0164] Optionally, in the embodiment of the present application, the method further includes:
[0165] In the process of adjusting the timing between the plurality of video segments by the sequential combing model, if timestamps overlap between the plurality of video segments, determining the sequence priority between the plurality of video segments by determining the continuity and content length of the video content between the plurality of video segments; or
[0166] In the process of adjusting the timing between multiple video clips through the sequential combing model, if it is calculated that the time difference between any two video clips is less than the set time gap, an additional frame is generated between the last frame of the first video clip and the first frame of the next video clip of the two video clips to fill the time gap, and the timing between the two video clips is readjusted based on the additional frame.
[0167] In this step, timestamp overlap is handled: When different video clips record overlapping events, to ensure the coherence and logic of the video narrative, the playback order needs to be determined based on the continuity of the video content and the length of each clip. Continuity assessment involves analyzing whether the image and audio between video clips can transition smoothly, while the length of the content may affect which clip's information is more important and should be displayed first.
[0168] Temporal Gap Filling: If the sequential combing model detects a small time interval (less than a preset threshold) between two adjacent video clips, this can cause a perceived jump in playback. To address this issue, the system automatically generates additional frames to insert between the two clips. These additional frames can be interpolated intermediate frames or transition frames synthesized from the content of the previous and next frames to smooth out the time jump and improve the viewing experience.
[0169] In this embodiment of the present application, we address timestamp overlap: Assume that a video editing project contains two video clips recording the same concert from different perspectives, and the timestamps indicate that the two clips overlap. Analysis reveals that the first video focuses on a close-up of the singer, while the second captures audience reactions. Because the first video is more crucial to conveying the core content of the concert (the singer's performance) and provides better visual continuity, we decide to play the first video first, followed by the second, to ensure a clear storyline.
[0170] Filling the Time Gap: In another scenario, the model captured a video clip of a morning sunrise and breakfast preparations. The model calculated a mere two seconds of time difference between the two videos, falling short of the five-second smooth playback time gap. The system then generated several frames of transitional images between the two videos, depicting a gradual shift in sky color and light, simulating the natural light transition from sunrise to breakfast table. This not only filled the time gap but also enhanced the video's coherence and artistic quality.
[0171] Optionally, in the embodiment of the present application, Figure 2 As shown, step 104 may specifically include:
[0172] 1041. Determine the visual features and scene semantic tone of each video segment based on the key frames of the plurality of video segments;
[0173] This step first selects representative keyframes from each video clip. These keyframes often contain the core visual information and emotional expression of the video content. By analyzing these keyframes, we can extract the unique visual features of each video clip, such as color distribution, texture, and shape, and simultaneously assess the semantic tone of the scene, such as whether it is cheerful, calm, or sad.
[0174] 1042. In the spatiotemporal layout, based on the visual features of each of the video segments, identify shared visual features between the multiple video segments;
[0175] In this step, the algorithm further searches for commonalities between different video clips within the framework of spatiotemporal layout. These shared visual features can be similar color themes, recurring patterns, or consistent composition, which help establish visual connections and flow across multiple clips.
[0176] 1043. Calculate a visual rotation angle and a position translation between two adjacent video segments based on shared visual features between the plurality of video segments;
[0177] In this step, the system uses the identified shared visual features to calculate how to smoothly transition between adjacent video segments. This involves adjusting the rotation angle of the viewing angle and the on-screen translation of the video segments to ensure a natural and coherent transition from one segment to the next. This step ensures a smooth visual flow and reduces abruptness.
[0178] 1044. Determine adjusted viewing angles and positions of the plurality of video segments based on the visual rotation angle and position translation between two adjacent video segments and the scene semantic tones corresponding to the two adjacent video segments.
[0179] During this step, the perspective and position adjustments are also considered to ensure the semantic tone of the scenes in adjacent video clips is consistent. If the two clips complement each other emotionally, such as a transition from tranquility to jubilation, the adjustments may include more obvious dynamic changes to enhance this emotional shift. Conversely, if emotional coherence is important, the adjustments will be more subtle to ensure a smooth transition.
[0180] 1045. Generate a video segment adjustment result based on the adjusted viewing angles and positions of the plurality of video segments.
[0181] In this step, the above analysis and calculations are combined to determine the adjusted perspective and position layout of all video clips, forming a harmonious and emotionally coherent video sequence. This process not only optimizes the performance of individual clips but also strengthens the inherent connections between them, enhancing the overall viewing experience.
[0182] In the embodiments of this application, this process can be applied to the production of travel memoirs, music videos, or any form of multi-clip video editing. Through intelligent analysis and adjustment, a visually unified and emotionally rich work can be created. For example, in a video project showing the changing seasons, the soft tones and vibrant scene semantic tone of the spring clips may smoothly transition to the bright colors and vitality of the summer clips. By gradually adjusting the perspective and position, the audience feels as if they have personally experienced a seamless seasonal journey.
[0183] Optionally, in an embodiment of the present application, the process of “calculating the visual rotation angle between two adjacent video clips based on the shared visual features between the plurality of video clips” in step 1043 may include: determining the shared visual features between two adjacent video clips, and constructing a shared visual feature point set corresponding to each video clip, respectively, so as to generate a shared visual feature set based on the shared visual feature point sets corresponding to the two adjacent video clips; calculating an initial rotation matrix based on the shared visual feature set, and determining the visual rotation angle between the two adjacent video clips based on the initial rotation matrix;
[0184] In this step, "shared visual features" refer to visual elements that appear in two adjacent video clips, such as the same objects, colors, line directions, or texture patterns. These features can serve as a bridge connecting the two clips and help achieve a smooth transition. "Shared visual feature point sets" refer to the specific pixels or feature point sets selected from each video clip that can represent these shared features. These point sets can quantify and compare the similarities and differences in the visual content of the two videos.
[0185] The "initial rotation matrix" is a mathematical tool in computer vision and graphics processing that describes the rotational transformation from one coordinate system to another. In this context, it indicates how to rotate the first video clip so that its visual features better match those of the second video clip, ensuring visual continuity and consistency.
[0186] In this example, let's assume we're editing two video clips. Clip A shows a beach at sunset, while Clip B follows, showing the same beach at night under a starry sky. Shared visual features might include the position of the sea level, specific rock formations, or the outline of a distant mountain. We first extract these shared feature points from both A and B, forming two feature point sets.
[0187] Next, an algorithm calculates the best possible match for these point sets, generating an initial rotation matrix. For example, if analysis reveals that segment B needs to be rotated 15 degrees counterclockwise relative to segment A to align the sea level feature points, then this 15-degree rotation is the calculated visual rotation angle. While scaling and translation must also be considered, the focus here is on rotation. Therefore, the perspective of segment A is adjusted to rotate 15 degrees. This ensures that when segment A transitions to segment B, the viewer sees the continuity of the sea level, creating a natural and fluid experience.
[0188] In summary, this embodiment ensures that the overall visual effect of the video remains coherent and harmonious even when there are significant time or lighting changes in the scene by accurately calculating and adjusting the visual rotation angles between video clips.
[0189] Based on the above, the process of "determining the adjusted viewing angles of the plurality of video segments based on the visual rotation angle between the two adjacent video segments and the scene semantic tones corresponding to the two adjacent video segments" in step 1044 may include:
[0190] By formula: , determining the adjusted viewing angles of the plurality of video clips based on the visual rotation angle between the two adjacent video clips and the scene semantic tones corresponding to the two adjacent video clips;
[0191] Among them, R′ represents the adjusted rotation matrix, R represents the initial rotation matrix, represents the coefficient of adjusting the initial rotation matrix by the scene semantic tone, S represents the quantized value of the scene semantic tone, and the exp function represents the matrix exponential function. It is expressed as a diagonal linear transformation of the identity matrix;
[0192] Based on the adjusted rotation matrix, the adjusted viewing angles of the plurality of video clips are determined.
[0193] In this step, the "adjusted rotation matrix R'" refers to the new matrix obtained by modifying the rotation matrix R initially calculated based on visual features after taking into account the semantic tone of the video clip scene. This adjustment is intended to ensure that the transition between video clips is not only visually continuous, but also maintains a consistent emotional atmosphere or produces a desired emotional gradient effect. " is a coefficient that determines the degree of influence of the scene semantic tone on the rotation angle. Its size reflects the importance of adjusting the scene semantic tone. "S" represents the quantitative value of the scene semantic tone, which is usually a numerical indicator that can reflect the intensity of the emotion (such as happiness, sadness, tension, etc.) of the video clip. "exp function" is a matrix exponential function. It is used here to convert the product of the diagonal matrix into the adjustment factor of the rotation matrix, ensuring the mathematical rationality of the rotation operation. And " ” is the stretching transformation of the identity matrix along its diagonal direction, which is used to construct the basic structure of rotation adjustment.
[0194] In this embodiment of the application, suppose there are two video clips. Clip A shows a cheerful party scene, while Clip B shows a quiet sunset scene. Through the previous steps, we have calculated the initial rotation matrix R from A to B as a 30-degree rotation matrix to match the position of the sea level in the scene. The scene semantic tone analysis shows that the scene semantic tone quantization value of Clip A is =8 (representing high activity), the quantified value of the scene semantic tone of segment B =6 (relatively calm), indicating that the semantic tone of the scene transitions from active to calm.
[0195] In order to reflect this emotional change, the =0.1, in order to moderately integrate the emotional factors. Then the adjusted rotation matrix It can be calculated by the formula:
[0196] ;
[0197] Substitute specific values:
[0198] ;
[0199] This means that the adjustment of the scene's semantic tone will slightly reduce the originally calculated rotation angle, reflecting the visual expression of emotions from active to calm. , the video editing software will adjust the perspective of segment B accordingly, making it not only visually smooth, but also emotionally achieving a natural transition from the joy of segment A to the tranquility of segment B.
[0200] It should be noted that the final adjustment matrix represents the rotation angle of each video clip after the perspective is adjusted. In practical applications, this adjusted rotation matrix is applied to the perspective transformation of each video clip, for example, by changing the camera's heading or rotation parameters, so that the video clip is displayed according to the new perspective in the virtual scene. This not only ensures that the video clips are correctly aligned in spatial layout, but also undergoes subtle visual adjustments based on the scene's semantic tone, making the overall video sequence visually and emotionally coherent and smooth.
[0201] In this way, the initial rotation matrix based on visual features is combined with the scene semantic tone to calculate the adjustment matrix It provides precise guidance for adjusting the perspective of video clips, thereby achieving natural transitions and emotional expression in time and space layout.
[0202] Optionally, in the embodiment of the present application, the process of “calculating the positional translation between two adjacent video clips based on the shared visual features between the plurality of video clips” in step 1043 may include: determining the shared visual features between two adjacent video clips, and calculating an average displacement between the two adjacent video clips based on the shared visual features between the two adjacent video clips, so as to determine the calculated positional translation between the two adjacent video clips;
[0203] In this step, "shared visual features" refer to prominent and recognizable visual elements that appear in two adjacent video clips, such as common landmarks, textures, colors, shapes, or specific patterns. These elements are key to connecting video clips and help determine the spatial relationship and position transformation between clips. "Position shift" refers to the relative positional displacement between adjacent video clips that needs to be adjusted to achieve a smooth transition between the spatial layout. This is calculated based on the matching of shared features and visual coherence, ensuring scene coherence and logic.
[0204] In this embodiment, let's assume we're producing a city documentary. Clip A shows a lake view in a city park, followed by a building complex on the other side of the lake. The shared visual feature is the edge of the lake, and image analysis software identifies the pixel set of the outline.
[0205] Determine shared visual features: First, analyze and find the lake edge line shared by segments A and B, and extract the edge contour feature point set of the lake.
[0206] Calculate the average displacement: By aligning the corresponding feature point sets of the lake boundary in the two segments, calculate their average offset difference. For example, it is calculated that segment B needs to be translated 5 meters (in units in virtual space) to the right relative to A to align with the edge of the lake.
[0207] Determine the position shift: Based on the average displacement calculation, we determine that segment B needs to be shifted 5 meters to the right relative to segment A. This is the position shift between the two video segments, ensuring that the lake surface appears smooth and seamless when transitioning from segment A to segment B, maintaining spatial continuity.
[0208] Therefore, the method of calculating position translation based on shared visual features ensures spatial continuity and visually smooth transitions between video clips, enhancing the audience's immersive experience.
[0209] Based on the above, the process of "determining the adjusted positions of the plurality of video segments based on the position translation between the two adjacent video segments and the scene semantic tones corresponding to the two adjacent video segments" in step 1044 may include:
[0210] By formula group: , determining the adjusted positions of the plurality of video segments based on the position translation between two adjacent video segments and the scene semantic tones corresponding to the two adjacent video segments;
[0211] in, Expressed as the x component of the position translation, Expressed as the y component of the position translation, is represented as an adjustment factor based on the semantic tone of the scene, S is represented as a quantized value of the semantic tone of the scene, It is represented as an adjustment factor based on the shared visual feature, and F is represented as the visual feature matching closeness, which is used to measure the consistency of the shared visual features between two adjacent video clips.
[0212] In this step, the position translation (Δx, Δy) refers to the displacement distance between adjacent video segments on the x-axis and y-axis calculated based on the shared visual features.
[0213] Adjustment factor of scene semantic tone : A coefficient based on the scene semantic tone adjustment, used to increase or decrease the displacement to match the emotional coherence between video clips.
[0214] The quantitative value S of the scene semantic tone: reflects the intensity of the emotion in the video clip, such as excitement, calmness, etc.
[0215] Adjustment factors for shared visual features : Weights adjusted based on the closeness of these visual feature matches help refine the displacement and ensure visual consistency.
[0216] Visual feature matching F: an indicator used to measure the similarity of visual features of adjacent segments.
[0217] Substitute these parameters into the formula group: , which can adjust the final position of the video clip on the x-axis and y-axis to adapt to the matching degree of emotional and visual features.
[0218] In this embodiment of the present application, it is assumed that a short film about natural scenery with emotional ups and downs is being produced:
[0219] Clip A: The lake is tranquil in the early morning with light mist, and the keynote quantization value is 𝑆𝐴=5 (calm).
[0220] Clip B: Turbulent waterfall in the afternoon, full of movement, keynote quantization value 𝑆𝐵=8 (active)
[0221] Position translation: Analyze the shared visual feature - the lake edge line, and calculate that segment A to B needs to be shifted right by Δx = 10 meters and Δy = 0 (keeping the horizontal level unchanged) to connect to the lake surface.
[0222] Adjust the scene semantic tone: Set it to 0.5 to blend the emotional changes in a moderate proportion. Set to 1 to emphasize the importance of visual matching.
[0223] Calculate the adjusted position: Δx′=10+0.5×5×1×1×10=15 meters, Δy'=0.
[0224] In the embodiment of the present application, based on the characteristics of the lake surface, the original 10-meter right shift is adjusted to 15 meters, which enhances the visual transition from tranquility to vitality, reflects the dual consideration of emotions and visual characteristics, and makes the transition between video clips natural and emotionally coherent.
[0225] Figure 3 A schematic diagram of the structure of a video splicing system based on a digital twin scene provided in an embodiment of the present application is shown in FIG. Figure 3 As shown, the system includes:
[0226] An acquisition module 31 is configured to acquire a set of video segments to be spliced, wherein each video segment in the set of video segments carries corresponding real scene coordinates and a timestamp;
[0227] A construction module 32 is configured to construct a virtual scene in the digital twin environment, and map the plurality of video clips into the virtual scene according to the real scene coordinates of the plurality of video clips to construct a spatial layout; adjust the timing between the plurality of video clips according to the timestamps of the plurality of video clips to obtain an initial video sequence, and map the initial video sequence into the virtual scene to construct a temporal layout; and construct a spatiotemporal layout based on the spatial layout and the temporal layout;
[0228] The generation module 33 is used to identify the visual features and scene semantic tone between the multiple video clips in the spatiotemporal layout, and adjust the perspectives and positions of the multiple video clips based on the visual features and scene semantic tone between the multiple video clips to generate video clip adjustment results; based on the video clip adjustment results, the multiple video clips are spliced to generate a digital twin display video.
[0229] Optionally, in an embodiment of the present application, the construction module 32 is specifically used to determine, for each of the video clips, the virtual scene coordinates corresponding to the real scene coordinates in the virtual scene according to the real scene coordinates of each of the video clips, wherein the real scene and the virtual scene have a corresponding coordinate system; determine the position of each of the video clips in the virtual scene according to the virtual scene coordinates corresponding to each of the video clips; determine the viewing angle of each of the video clips in the virtual scene; and generate a spatial layout based on the positions and viewing angles of multiple video clips.
[0230] Optionally, in an embodiment of the present application, the construction module 32 is also used to calculate, for each of the video clips, the geometric center point of the video clip in the virtual scene according to the virtual scene coordinates corresponding to the video clip and the coverage area of the video clip; determine the video content of the video clip, determine the sight direction vector of the video clip, and the video content includes at least: the subject movement direction, the camera pointing; determine the physical properties of the video clip, determine the effective observation range of the video clip in the virtual scene, and the physical properties include at least: the focal length of the lens; determine the viewing angle of the video clip in the virtual scene according to the geometric center point, the sight direction vector and the focal length of the lens corresponding to the video clip.
[0231] Optionally, in an embodiment of the present application, the construction module 32 is also used to extract the key frames of each video clip, and determine the visual features and scene semantic tone of the video clip based on the key frames of each video clip; the timestamp, visual features and scene semantic tone of the video clip are used as input, and input into a pre-viewed sequential combing model to adjust the timing between the multiple video clips through the sequential combing model, and output the initial video sequence, wherein the sequential combing model is trained based on multiple video clip samples, and the multiple video clip samples contain video sequence data marked with timestamps, visual features and scene semantic tone, and context information corresponding to the video sequence data.
[0232] Optionally, in the embodiment of the present application, the system further includes a processing module 34;
[0233] The processing module 34 is used to determine the sequence priority between the multiple video clips by determining the continuity and content length of the video content between the multiple video clips during the process of adjusting the timing between the multiple video clips through the sequential combing model. If timestamps overlap between the multiple video clips, the processing module 34 is used to determine the sequence priority between the multiple video clips by determining the continuity and content length of the video content between the multiple video clips; or, in the process of adjusting the timing between the multiple video clips through the sequential combing model, if the time difference between any two video clips is calculated to be less than the set time gap, an additional frame is generated between the last frame of the first video clip and the first frame of the next video clip in the two video clips to fill the time gap, and the timing between the two video clips is readjusted based on the additional frame.
[0234] Optionally, in an embodiment of the present application, the generation module 33 is specifically used to determine the visual features and scene semantic tone of each of the video clips based on the key frames of the multiple video clips; in the spatiotemporal layout, based on the visual features of each of the video clips, identify the shared visual features between the multiple video clips; based on the shared visual features between the multiple video clips, calculate the visual rotation angle and position translation between two adjacent video clips; based on the visual rotation angle and position translation between two adjacent video clips and the scene semantic tone corresponding to the two adjacent video clips, determine the adjusted viewing angle and position of the multiple video clips; based on the adjusted viewing angle and position of the multiple video clips, generate a video clip adjustment result.
[0235] Optionally, in an embodiment of the present application, the generating module 33 is further configured to determine shared visual features between two adjacent video segments, and construct a shared visual feature point set corresponding to each video segment, so as to generate a shared visual feature set based on the shared visual feature point sets corresponding to the two adjacent video segments; calculate an initial rotation matrix based on the shared visual feature set, and determine a visual rotation angle between the two adjacent video segments based on the initial rotation matrix;
[0236] Optionally, in the embodiment of the present application, the generating module 33 is further configured to generate the following equation: , determining the adjusted viewing angles of the plurality of video clips based on the visual rotation angle between the two adjacent video clips and the scene semantic tones corresponding to the two adjacent video clips;
[0237] Among them, R′ represents the adjusted rotation matrix, R represents the initial rotation matrix, represents the coefficient of adjusting the initial rotation matrix by the scene semantic tone, S represents the quantized value of the scene semantic tone, and the exp function represents the matrix exponential function. It is represented as a diagonal linear transformation of the identity matrix;
[0238] Based on the adjusted rotation matrix, the adjusted viewing angles of the plurality of video clips are determined.
[0239] Optionally, in the embodiment of the present application, the generating module 33 is further configured to determine a shared visual feature between two adjacent video segments, and calculate an average displacement between the two adjacent video segments based on the shared visual feature between the two adjacent video segments, so as to determine a position translation amount between the two adjacent video segments;
[0240] Optionally, in the embodiment of the present application, the generating module 33 is further configured to generate the following formula: , determining the adjusted positions of the plurality of video segments based on the position translation between two adjacent video segments and the scene semantic tones corresponding to the two adjacent video segments;
[0241] in, Expressed as the x component of the position translation, Expressed as the y component of the position translation, is represented as an adjustment factor based on the semantic tone of the scene, S is represented as a quantized value of the semantic tone of the scene, It is represented as an adjustment factor based on the shared visual feature, and F is represented as the visual feature matching closeness, which is used to measure the consistency of the shared visual features between two adjacent video clips.
[0242] Figure 3 The video splicing system based on digital twin scenes can perform Figure 1 The implementation principle and technical effects of the video stitching method based on the digital twin scene described in the illustrated embodiment will not be repeated here. The specific manner in which each module and unit performs operations in the video stitching system based on the digital twin scene in the above embodiment has been described in detail in the embodiment of the method and will not be elaborated here.
[0243] In one possible design, Figure 3 The video stitching system based on digital twin scenes of the embodiment shown can be implemented as a computing device, such as Figure 4 As shown, the computing device may include a storage component 41 and a processing component 42;
[0244] The storage component 41 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 42 .
[0245] The processing component 42 is used to: obtain a set of video clips to be spliced, each video clip in the set of video clips carries corresponding real scene coordinates and timestamps; construct a virtual scene in a digital twin environment, and map multiple video clips to the virtual scene based on the real scene coordinates of multiple video clips to construct a spatial layout; adjust the timing between multiple video clips based on the timestamps of multiple video clips to obtain an initial video sequence, and map the initial video sequence to the virtual scene to construct a temporal layout; construct a spatiotemporal layout based on the spatial layout and the temporal layout, and in the spatiotemporal layout, identify the visual features and scene semantic tone between multiple video clips, and adjust the perspectives and positions of multiple video clips based on the visual features and scene semantic tone between multiple video clips to generate video clip adjustment results; based on the video clip adjustment results, perform splicing processing on multiple video clips to generate a digital twin display video.
[0246] The processing component 42 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.
[0247] The storage component 41 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0248] The display component 43 may be an electroluminescent (EL) element, a liquid crystal display or a micro display having a similar structure, or a retinal direct display or a similar laser scanning display.
[0249] Of course, a computing device may also include other components, such as input / output interfaces, communication components, etc.
[0250] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.
[0251] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.
[0252] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.
[0253] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program can achieve the above-mentioned Figure 1 The video stitching method based on the digital twin scene of the illustrated embodiment.
[0254] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0255] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0256] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0257] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A video splicing method based on digital twin scenes, characterized in that: include: Obtaining a set of video clips to be spliced, wherein each video clip in the set of video clips carries corresponding real scene coordinates and a timestamp; Constructing a virtual scene in the digital twin environment, and mapping the plurality of video clips into the virtual scene according to the real scene coordinates of the plurality of video clips to construct a spatial layout; Adjusting the timing between the plurality of video clips according to the timestamps of the plurality of video clips to obtain an initial video sequence, and mapping the initial video sequence to the virtual scene to construct a time layout; Constructing a spatiotemporal layout based on the spatial layout and the temporal layout, identifying visual features and scene semantic tones between the plurality of video clips in the spatiotemporal layout, and adjusting the perspectives and positions of the plurality of video clips based on the visual features and scene semantic tones between the plurality of video clips to generate a video clip adjustment result; Based on the video clip adjustment results, multiple video clips are spliced together to generate a digital twin display video; The step of identifying visual features and scene semantic tones among the plurality of video clips in the spatiotemporal layout, and adjusting the perspectives and positions of the plurality of video clips based on the visual features and scene semantic tones among the plurality of video clips to generate a video clip adjustment result, includes: Determining the visual features and scene semantic tone of each video clip based on the key frames of the plurality of video clips; In the spatiotemporal layout, based on the visual features of each of the video segments, identifying shared visual features between the plurality of video segments; Calculating a visual rotation angle and a position translation amount between two adjacent video clips based on shared visual features between the plurality of video clips; Determining the adjusted viewing angles and positions of the plurality of video clips based on the visual rotation angle and position translation between two adjacent video clips and the scene semantic tones corresponding to the two adjacent video clips; A video segment adjustment result is generated based on the adjusted viewing angles and positions of the plurality of video segments.
2. The method according to claim 1, characterized in that Mapping the plurality of video clips to the virtual scene according to the real scene coordinates of the plurality of video clips to construct a spatial layout includes: For each of the video clips, determining, according to the real scene coordinates of each of the video clips, the virtual scene coordinates corresponding to the real scene coordinates in the virtual scene, wherein the real scene and the virtual scene have a corresponding coordinate system; Determining the position of each video clip in the virtual scene according to the virtual scene coordinates corresponding to each video clip; determining a viewing angle of each of the video clips in the virtual scene; A spatial layout is generated based on the positions and viewing angles of the plurality of video clips.
3. The method according to claim 2, characterized in that Determining the viewing angle of each video clip in the virtual scene includes: For each of the video clips, calculating a geometric center point of the video clip in the virtual scene according to the virtual scene coordinates corresponding to the video clip and the coverage area of the video clip; Determining the video content of the video clip and determining the sight direction vector of the video clip, wherein the video content at least includes: a subject movement direction and a camera pointing direction; Determining physical properties of the video clip and determining an effective observation range of the video clip in the virtual scene, wherein the physical properties include at least: lens focal length; The viewing angle of the video clip in the virtual scene is determined according to the geometric center point, the sight direction vector, and the focal length of the lens corresponding to the video clip.
4. The method according to claim 1, wherein The adjusting the timing between the plurality of video segments according to the timestamps of the plurality of video segments to obtain an initial video sequence includes: For each of the video clips, extracting key frames of each of the video clips, and determining the visual features and scene semantic tone of the video clip based on the key frames of each of the video clips; The timestamp, visual features and scene semantic tone of the video clip are taken as input and input into a pre-viewed sequential combing model to adjust the timing between the multiple video clips through the sequential combing model and output an initial video sequence, wherein the sequential combing model is trained based on multiple video clip samples, and the multiple video clip samples contain video sequence data marked with timestamps, visual features and scene semantic tone, as well as context information corresponding to the video sequence data.
5. The method according to claim 4, characterized in that Also includes: In the process of adjusting the timing between the plurality of video clips by using the sequential combing model, if timestamps overlap between the plurality of video clips, determining the order priority between the plurality of video clips by determining the continuity and content length of the video content between the plurality of video clips; or, In the process of adjusting the timing between multiple video clips through the sequential combing model, if it is calculated that the time difference between any two video clips is less than the set time gap, an additional frame is generated between the last frame of the first video clip and the first frame of the next video clip of the two video clips to fill the time gap, and the timing between the two video clips is readjusted based on the additional frame.
6. The method according to claim 1, characterized in that Calculating a visual rotation angle between two adjacent video clips based on shared visual features between the plurality of video clips includes: Determining shared visual features between two adjacent video clips, and constructing a shared visual feature point set corresponding to each video clip, so as to generate a shared visual feature set based on the shared visual feature point sets corresponding to the two adjacent video clips; Calculating an initial rotation matrix based on the shared visual feature set, and determining a visual rotation angle between two adjacent video segments according to the initial rotation matrix; determining adjusted viewing angles of the plurality of video clips based on a visual rotation angle between two adjacent video clips and a scene semantic tone corresponding to the two adjacent video clips; Based on the adjusted rotation matrix, the adjusted viewing angles of the plurality of video clips are determined.
7. The method according to claim 1, characterized in that Calculating the position translation between two adjacent video clips based on the shared visual features between the plurality of video clips includes: Determining a shared visual feature between two adjacent video clips, and calculating an average displacement between the two adjacent video clips based on the shared visual feature between the two adjacent video clips, so as to determine a position translation amount between the two adjacent video clips; Based on the position translation amount between two adjacent video segments and the scene semantic tones corresponding to the two adjacent video segments, the adjusted positions of the plurality of video segments are determined.
8. A video splicing system based on a digital twin scene, used to execute the video splicing method based on a digital twin scene according to any one of claims 1 to 7, characterized in that: include: An acquisition module, configured to acquire a set of video segments to be spliced, wherein each video segment in the set of video segments carries corresponding real scene coordinates and a timestamp; A construction module is configured to construct a virtual scene in the digital twin environment, and map the plurality of video clips into the virtual scene according to the real scene coordinates of the plurality of video clips to construct a spatial layout; adjust the timing between the plurality of video clips according to the timestamps of the plurality of video clips to obtain an initial video sequence, and map the initial video sequence into the virtual scene to construct a temporal layout; Constructing a spatiotemporal layout based on the spatial layout and the temporal layout; A generation module is used to identify the visual features and scene semantic tone between the multiple video clips in the spatiotemporal layout, and adjust the perspectives and positions of the multiple video clips based on the visual features and scene semantic tone between the multiple video clips to generate video clip adjustment results; based on the video clip adjustment results, the multiple video clips are spliced to generate a digital twin display video.
9. A computing device, characterized in that It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the video stitching method based on the digital twin scene as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Image information processing method for short video production
CN119152096A
Video stream and digital twin scene fusion method, terminal equipment and storage medium
CN120128679A