Video stitching method and system based on digital twinborn scene, and computing device

By using the real scene coordinates and timestamps of video clips in a digital twin environment, the spatial and temporal layout of video clips is automatically adjusted, which solves the problem of inefficiency in traditional video stitching methods, and achieves high-precision and efficient video stitching, which is suitable for a variety of application scenarios.

CN120378688AActive Publication Date: 2025-07-25BEIJING ZHIHUI YUNZHOU TECH CO LTD

Patent Information

Application Number
CN202510854417.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-25
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

Traditional video stitching methods rely on manual intervention, are inefficient and difficult to achieve high-precision and automated processing, especially when complex scene changes and multi-view conversion, they are prone to screen misalignment and time discontinuity problems.

Method used

Based on the video stitching method of digital twin scenes, by obtaining the real scene coordinates and timestamps of the video clips, a virtual scene is constructed in a digital twin environment, the spatial layout and temporal layout of the video clips are adjusted, visual features and scene semantic tone are identified, and the perspective angle and position are dynamically adjusted, and the digital twin display video is finally generated.

Benefits of technology

It improves the automation and accuracy of video stitching, enhances visual coherence and immersion, solves the stitching problem of complex scenes, improves processing efficiency and reduces labor costs, and is suitable for large-scale video content production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378688A_ABST
    Figure CN120378688A_ABST
Patent Text Reader

Abstract

The invention provides a video stitching method and system based on a digital twin scene, and a computing device. The method comprises the following steps: acquiring a to-be-spliced video clip set; mapping the plurality of video clips into a virtual scene according to the real scene coordinates of the plurality of video clips so as to construct a spatial layout; obtaining an initial video sequence according to the timestamps of the plurality of video clips, and mapping the initial video sequence into the virtual scene to construct a time layout; constructing a spatio-temporal layout based on the spatial layout and the temporal layout, identifying visual features and scene semantic tones among the plurality of video clips in the spatio-temporal layout, and adjusting visual angles and positions of the plurality of video clips based on the visual features and the scene semantic tones among the plurality of video clips to generate a video clip adjustment result; and based on the video clip adjustment result, splicing the plurality of video clips to generate a digital twin display video. According to the technical scheme provided by the invention, the video stitching effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of digital twin, and in particular to a video splicing method, system and computing device based on a digital twin scenario. Background Art

[0002] With the rapid development of information technology, Digital Twin, as a concept integrating advanced technologies such as the Internet of Things, big data, and artificial intelligence, has been increasingly emphasized. Digital Twin is essentially an accurate digital mirror of the physical world. By creating virtual copies of physical objects, systems, or processes, it enables real-time data interaction and dynamic simulation between the real and virtual worlds. In the field of video splicing, the application of Digital Twin has opened up new possibilities, especially showing significant advantages in complex scene reconstruction and immersive experience creation. By reproducing every detail of the real environment in the virtual space and combining data in the time dimension, Digital Twin can achieve precise synchronization and integration of video clips, greatly enhancing the coherence and realism of video content, and bringing unprecedented visual enjoyment and in-depth understanding to the audience.

[0003] Most traditional video splicing methods rely on basic image processing technologies and video editing software. These methods generally include the following steps: First, manually select video clips captured by multiple cameras; then, use the alignment tool in the video editing software to perform rough alignment based on recognizable common features in the video (such as the horizon, building edges, etc.); then, try to ensure the consistency of the time sequence by adjusting the playback speed of the video clips or inserting transition effects; finally, perform color correction and image quality optimization in order to achieve visual coherence.

[0004] However, although these traditional methods meet the basic video splicing requirements to a certain extent, their processes are cumbersome, inefficient, and often rely on manual intervention, making it difficult to achieve high-precision and automated processing. Summary of the Invention

[0005] The embodiments of the present application provide a video splicing method, system and computing device based on a digital twin scenario, so as to solve the problem of poor video splicing effect in the prior art.

[0006] In a first aspect, the embodiments of the present application provide a video splicing method based on a digital twin scenario, including: Obtain a set of video clips to be spliced, and each video clip in the set of video clips carries a corresponding real scene coordinate and timestamp; Construct a virtual scene in the digital twin environment, and map multiple video clips into the virtual scene according to the real scene coordinates of the multiple video clips to construct a spatial layout; Adjust the timing sequence among multiple video clips according to the timestamps of the multiple video clips to obtain an initial video order, and map the initial video order into the virtual scene to construct a time layout; Construct a spatio-temporal layout based on the spatial layout and the time layout, and in the spatio-temporal layout, identify the visual features and scene semantic tones among the multiple video clips, and based on the visual features and scene semantic tones among the multiple video clips, adjust the perspectives and positions of the multiple video clips to generate a video clip adjustment result; Based on the video clip adjustment result, perform splicing processing on the multiple video clips to generate a digital twin display video.

[0007] Optionally, the mapping of the multiple video clips into the virtual scene according to the real scene coordinates of the multiple video clips to construct a spatial layout includes: For each video clip, determine the virtual scene coordinates corresponding to the real scene coordinates in the virtual scene according to the real scene coordinates of each video clip, where there is a corresponding coordinate system between the real scene and the virtual scene; Determine the position of each video clip in the virtual scene according to the virtual scene coordinates corresponding to each video clip; Determine the perspective of each video clip in the virtual scene; Generate a spatial layout based on the positions and perspectives of the multiple video clips.

[0008] Optionally, the determining the perspective of each video clip in the virtual scene includes: For each video clip, calculate the geometric center point of the video clip in the virtual scene according to the virtual scene coordinates corresponding to the video clip and the coverage area of the video clip; Determine the video content of the video clip, and determine the line-of-sight direction vector of the video clip, where the video content includes at least: the main body movement direction, the camera pointing direction; Determine the physical attributes of the video clip, and determine the effective observation range of the video clip in the virtual scene, where the physical attributes include at least: the lens focal length; Determine the perspective of the video clip in the virtual scene according to the geometric center point, the line-of-sight direction vector, and the lens focal length corresponding to the video clip.

[0009] Optionally, the adjusting the timing sequence among the multiple video clips according to the timestamps of the multiple video clips to obtain an initial video order includes: For each of the video segments, extract the key frames of each video segment, and based on the key frames of each video segment, determine the visual features and scene semantic tone of the video segment; Use the time stamp, visual features, and scene semantic tone of the video segment as input, and input them into a pre-trained sequence sorting model to adjust the time sequence between multiple video segments through the sequence sorting model, and output an initial video order, where the sequence sorting model is trained based on multiple video segment samples, and the multiple video segment samples include video sequence data with annotated time stamps, visual features, and scene semantic tones, as well as the context information corresponding to the video sequence data.

[0010] Optionally, it further includes: During the process of adjusting the time sequence between multiple video segments through the sequence sorting model, if there is a time stamp overlap between multiple video segments, determine the order priority between multiple video segments by determining the continuity and content length of the video content between multiple video segments; or, During the process of adjusting the time sequence between multiple video segments through the sequence sorting model, if the calculated time difference between any two video segments is less than the set time gap, generate additional frames between the last frame of the first video segment and the first frame of the next video segment among the two video segments to fill the time gap, and re-adjust the time sequence between the two video segments based on the additional frames.

[0011] Optionally, in the spatio-temporal layout, identify the visual features and scene semantic tones between multiple video segments, and based on the visual features and scene semantic tones between multiple video segments, adjust the perspectives and positions of multiple video segments to generate a video segment adjustment result, including: Based on the key frames of multiple video segments, determine the visual features and scene semantic tone of each video segment; In the spatio-temporal layout, based on the visual features of each video segment, identify the shared visual features between multiple video segments; According to the shared visual features between multiple video segments, calculate the visual rotation angle and position translation amount between adjacent two video segments; Based on the visual rotation angle and position translation amount between adjacent two video segments and the corresponding scene semantic tones of adjacent two video segments, determine the perspectives and positions of the adjusted multiple video segments; Based on the perspectives and positions of the adjusted multiple video segments, generate a video segment adjustment result.

[0012] Optionally, calculating a visual rotation angle between two adjacent video segments according to shared visual features between multiple video segments includes: Determining shared visual features between two adjacent video segments, and respectively constructing a set of shared visual feature points corresponding to each video segment, so as to generate a shared visual feature set according to the sets of shared visual feature points corresponding to two adjacent video segments; Calculating an initial rotation matrix based on the shared visual feature set, and determining a visual rotation angle between two adjacent video segments according to the initial rotation matrix; Determining perspectives of the adjusted multiple video segments based on the visual rotation angle between two adjacent video segments and the scene semantic tones corresponding to the two adjacent video segments; Determining perspectives of the adjusted multiple video segments based on the adjusted rotation matrix.

[0013] Optionally, calculating a position translation amount between two adjacent video segments according to shared visual features between multiple video segments includes: Determining shared visual features between two adjacent video segments, and calculating an average displacement between the two adjacent video segments based on the shared visual features between the two adjacent video segments, so as to determine a position translation amount between the two adjacent video segments; Determining positions of the adjusted multiple video segments based on the position translation amount between two adjacent video segments and the scene semantic tones corresponding to the two adjacent video segments.

[0014] In a second aspect, an embodiment of the present application provides a video splicing system based on a digital twin scenario, including: An acquisition module, configured to acquire a set of video segments to be spliced, and each video segment in the set of video segments carries a corresponding real scene coordinate and a time stamp; A construction module, configured to construct a virtual scene in a digital twin environment, map multiple video segments into the virtual scene according to the real scene coordinates of the multiple video segments, so as to construct a spatial layout; adjust the time sequence between the multiple video segments according to the time stamps of the multiple video segments to obtain an initial video sequence, and map the initial video sequence into the virtual scene, so as to construct a time layout; construct a spatio-temporal layout based on the spatial layout and the time layout; A generation module, configured to identify visual features and scene semantic tones among multiple video segments in the spatio-temporal layout, and based on the visual features and scene semantic tones among multiple video segments, adjust the perspectives and positions of multiple video segments to generate a video segment adjustment result; and based on the video segment adjustment result, perform splicing processing on multiple video segments to generate a digital twin display video.

[0015] In a third aspect, an embodiment of the present application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the video splicing method based on a digital twin scene as described in the first aspect above.

[0016] In an embodiment of the present application, a set of video segments to be spliced is obtained, and each video segment in the set of video segments carries a corresponding real scene coordinate and timestamp; a virtual scene is constructed in a digital twin environment, and multiple video segments are mapped into the virtual scene according to the real scene coordinates of the multiple video segments to construct a spatial layout; according to the timestamps of the multiple video segments, the time sequence between the multiple video segments is adjusted to obtain an initial video order, and the initial video order is mapped into the virtual scene to construct a time layout; a spatio-temporal layout is constructed based on the spatial layout and the time layout, and in the spatio-temporal layout, visual features and scene semantic tones among multiple video segments are identified, and based on the visual features and scene semantic tones among multiple video segments, the perspectives and positions of multiple video segments are adjusted to generate a video segment adjustment result; and based on the video segment adjustment result, splicing processing is performed on multiple video segments to generate a digital twin display video.

[0017] The video splicing method based on a digital twin scene proposed in the present application shows the following remarkable beneficial effects compared with traditional video splicing technologies: Improved high automation and accuracy: By automatically obtaining the real scene coordinates and timestamp information of video segments and mapping them in a digital twin environment, this method significantly improves the automation degree of video splicing. The construction of the spatial and time layouts ensures the precise positioning and time sequence arrangement of video segments in the virtual space, reduces manual intervention, and improves the splicing accuracy.

[0018] Enhanced visual coherence and immersion: Based on the construction of the spatio-temporal layout, by intelligently identifying and analyzing the visual features and scene semantic tones between video segments, this method can dynamically adjust the perspectives and positions of video segments to ensure visual fluency and consistency of emotional expression, enhancing the overall coherence of the final video and the immersive experience of the audience.

[0019] Efficiently handle complex scenarios: For video materials containing complex scene changes and multi-perspective conversions, this method, relying on the powerful capabilities of digital twin technology, can effectively respond to and accurately splice them, solving the problems of picture misalignment and time discontinuity that easily occur in traditional methods when dealing with such scenarios, and improving the ability and efficiency in handling complex scenarios.

[0020] Intelligent emotion matching and perspective optimization: By analyzing the scene semantic tone of video clips, this method can intelligently adjust the video display angles and order, making the conveyed video content more in line with the expected emotional atmosphere, and increasing the artistry and appeal of video narration.

[0021] Facilitate large-scale content production: Due to its high degree of automation processing ability and effective management of complex scenarios, this method is very suitable for application in the production of large-scale video content, can significantly improve production efficiency, reduce labor costs, and meet the needs of the modern media and entertainment industries for high-quality and high-efficiency video content.

[0022] In summary, the present invention not only revolutionizes the technical means of video splicing, but also greatly enriches the creative potential of video content, providing advanced technical support for fields such as digital media, film and television production, and virtual reality.

[0023] These aspects or other aspects of this application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0025] Figure 1 It is a flowchart of an embodiment of a video splicing method based on a digital twin scene provided by an embodiment of this application; Figure 2 It is a flowchart of another embodiment of a video splicing method based on a digital twin scene provided by an embodiment of this application; Figure 3 It is a schematic structural diagram of a video splicing system based on a digital twin scene provided by an embodiment of this application; Figure 4 It is a schematic structural diagram of a computing device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application.

[0027] In some processes described in the specification, claims and above-mentioned accompanying drawings of this application, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear herein or in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations can be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0028] The technical solution of this application can be widely applied to a variety of scenarios, specifically including but not limited to the following aspects: Film and television production and post-production editing: In the post-production of movies, TV series or advertisements, this technology can efficiently integrate shots taken at different shooting locations and times, maintain visual and emotional coherence, and accelerate the post-production editing process.

[0029] Virtual reality (VR) and augmented reality (AR) content creation: When creating immersive content for VR / AR experiences, highly precise spatial and temporal synchronization is required. This method can help creators seamlessly integrate real-shot videos and virtual elements to enhance the user experience.

[0030] News reporting and documentary production: In a rapidly changing news scene or a complex documentary shooting environment, this technology can quickly organize video clips from multiple sources, automatically splice them according to the timeline and spatial logic, speed up the news production speed, and ensure the timeliness and accuracy of reports.

[0031] Sports event live broadcast and review: For large-scale sports events, real-time videos obtained from multiple camera angles can be quickly integrated through this method, and the perspective and order can be adjusted in real time according to the progress of the game, providing the audience with a comprehensive and coherent game viewing experience.

[0032] Urban planning and architectural design display: In urban planning or architectural design projects, by combining video clips of drone shooting and ground photography and using digital twin technology for spatial layout, dynamic urban landscapes or architectural roaming videos can be generated, facilitating decision-makers and the public to understand the design scheme.

[0033] Education and training: When producing interactive educational video content, this technology can precisely integrate multi-dimensional content such as theoretical explanations and practical operations, historical scenes and modern comparisons, etc., enhancing the attractiveness of teaching materials and teaching effects.

[0034] Tourism promotion: The tourism industry can use this technology to integrate scenic spot videos of different seasons and different time periods, creating a tourism promotional video that crosses time boundaries and presenting a comprehensive and vivid destination image to tourists.

[0035] The inventors' research found that most current traditional video splicing methods rely on basic image processing technologies and video editing software. These methods generally include the following steps: First, manually select video segments captured by multiple cameras; then, use the alignment tools in the video editing software to roughly align according to the recognizable common features in the video (such as the horizon, building edges, etc.); then, try to ensure the consistency of the time sequence by adjusting the playback speed of the video segments or inserting transition effects; finally, perform color correction and picture quality optimization in order to achieve visual coherence. However, although these traditional methods meet the basic video splicing requirements to a certain extent, their processes are cumbersome, inefficient, and often rely on manual intervention, making it difficult to achieve high-precision and automated processing.

[0036] That is to say, the defects of existing video splicing technologies are becoming increasingly prominent. On the one hand, due to the lack of intelligent scene understanding and spatio-temporal coordination mechanisms, traditional methods are extremely prone to problems such as picture misalignment and time discontinuity when processing video materials with complex scene changes, rapid perspective conversions, or variable lighting conditions, affecting the viewing experience of the final video. On the other hand, the large amount of manual participation leads to a huge workload and is difficult to meet the needs of large-scale video content production, especially in the contemporary media environment that pursues high efficiency and high-quality output. Therefore, exploring a new generation of video splicing methods that combine digital twin technology to achieve higher-precision automatic splicing and intelligent adjustment has become an important research direction.

[0037] In view of this, the present application provides a video splicing method based on a digital twin scenario. The method includes: obtaining a set of video segments to be spliced, where each video segment in the set of video segments carries a corresponding real-scene coordinate and timestamp; constructing a virtual scenario in a digital twin environment, and mapping the multiple video segments into the virtual scenario according to the real-scene coordinates of the multiple video segments to construct a spatial layout; adjusting the time sequence between the multiple video segments according to the timestamps of the multiple video segments to obtain an initial video order, and mapping the initial video order into the virtual scenario to construct a time layout; constructing a spatio-temporal layout based on the spatial layout and the time layout, and in the spatio-temporal layout, identifying the visual features and scene semantic tones between the multiple video segments, and adjusting the perspectives and positions of the multiple video segments based on the visual features and scene semantic tones between the multiple video segments to generate a video segment adjustment result; and performing splicing processing on the multiple video segments based on the video segment adjustment result to generate a digital twin display video.

[0038] The method of the present application can solve the problems of picture misalignment and time discontinuity that are prone to occur in traditional methods when processing such scenarios. While improving the ability and efficiency of processing complex scenarios, it can also improve the video splicing efficiency and reduce the labor cost.

[0039] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0040] Figure 1 It is a flowchart of an embodiment of a video splicing method based on a digital twin scenario provided by an embodiment of the present application. As Figure 1 shown, the method includes: 101. Obtain a set of video segments to be spliced, where each video segment in the set of video segments carries a corresponding real-scene coordinate and timestamp; In this step, "obtaining a set of video segments to be spliced" refers to collecting a series of different segments that are intended to be combined into a single continuous video. These video segments come from different sources and may be real-world scenes filmed at different times and with different devices. Each "video segment" is an independent video clip that contains a visual record of a specific part of a scene. In particular, each video segment "carries corresponding real-scene coordinates and timestamps", which means that in addition to the video content itself, each segment is attached with geographical location information (real-scene coordinates) and the exact time point when the segment was filmed (timestamp). Such additional information is crucial for subsequent precise splicing because it allows the system to understand the relative positions of each segment in physical space and on the time axis.

[0041] In an embodiment of this application, it is assumed that work is being done on a cityscape documentary project. The project team used drones and ground cameras to film different areas of the city at different times of the day and at multiple locations. Each camera device was equipped with a GPS module to record the shooting position coordinates, and the time of the cameras was precisely synchronized to ensure that all video segments carried accurate timestamps.

[0042] The details of the embodiment are as follows: Video segment collection: The team recorded the sunrise view, busy streets, riverside at sunset, and night view in the east, downtown, west, and north areas of the city from morning to night. A total of 20 independent high-definition video segments were collected, and each segment was about 30 seconds long.

[0043] Real-scene coordinates: When each video segment was recorded, the GPS automatically recorded the longitude and latitude information of the shooting position. For example, the coordinates of the sunrise segment were 120.75° east longitude and 31.23° north latitude, while the coordinates of the night view segment were 120.78° east longitude and 31.20° north latitude.

[0044] Timestamp: All devices were synchronized with the network time server to ensure that there was an accurate time record at the moment when each video segment started recording. For example, the timestamp of the sunrise segment was 2023-04-01 06:00:00, and the timestamp of the night view segment was 19:30:00 on the same day.

[0045] With this detailed geographical coordinate and timestamp information, the post-production team can use the method proposed in this application to precisely arrange these segments in spatial and temporal order and finally splice them into a smooth all-weather cityscape documentary that shows the city from dawn to dusk.

[0046] 102. Construct a virtual scene in a digital twin environment, and map multiple said video segments into the virtual scene according to the real-scene coordinates of the multiple said video segments to construct a spatial layout; In this step, the "digital twin environment" refers to a highly realistic virtual space that replicates and maps the real-world environment through digital technology, including geographical and physical attributes, buildings, object locations, etc., and even lighting and weather conditions. "Building a virtual scene" means creating a virtual model corresponding to the real world in such a digital twin environment, providing an interactive and operational platform for video stitching. "Mapping multiple video clips to the virtual scene" means using the real-scene coordinate information carried by the video clips to accurately place each video clip in the corresponding position in the virtual space, achieving alignment and layout in the physical space. The purpose of this is to establish a spatial connection between the video clips, laying the foundation for subsequent time-series stitching and visual fluency.

[0047] In the embodiments of this application, it is assumed that a documentary about ancient ruins exploration is being produced. The team visited three ancient ruins distributed around the city at different times, and a video clip was taken at each ruin. Each video clip not only records the visual content of the ruin but also the GPS coordinate information and timestamp at the time of shooting.

[0048] The implementation steps are as follows: Building the digital twin environment: First, use 3D modeling software (such as Unity or Unreal-time Engine) and laser scanning data, satellite map APIs, etc. to create a virtual ancient city ruins environment as close to the real one as possible. The environment includes terrain undulations, vegetation, roads, the appearance and internal structure of architectural ruins, etc., and even simulates the lighting conditions at different times, such as sunrise, afternoon, and dusk.

[0049] Mapping video clips to the virtual scene: For the first video clip, which was taken at the entrance of the foot of the mountain in the morning with coordinates (12345.2°N, 11.5°E), find this location in the digital twin environment and embed the video clip so that its vision matches the virtual entrance scene.

[0050] The second video, taken in the afternoon at the central square of the ruin with coordinates (1.9°N, 1.8°E), locate the square in the virtual scene and integrate the video clip to ensure that the light direction and shadows are consistent with the virtual environment.

[0051] For the last one, at the highest point of the ruin at dusk with coordinates (1.5°N, 1.2°E), the video clip is accurately placed on the virtual mountaintop, using the dusk lighting that matches the time to create an atmosphere consistent with the real video.

[0052] Through such mapping and layout, all video clips can be accurately positioned in the digital twin environment, forming a coherent spatial sequence, laying the foundation for the next step of splicing in time series and adjusting visual effects.

[0053] 103. Adjust the timing between the multiple video clips according to the timestamps of the multiple video clips to obtain an initial video sequence, and map the initial video sequence to the virtual scene to construct a time layout; In this step, "adjusting the timing according to the timestamps of multiple video clips" means using the specific time information recorded when the video clips were recorded to determine their order on the timeline. This step is crucial to ensure that the video clips are spliced in the order in which they actually occurred, so that the narrative or scene presentation is logically coherent. "Obtaining the initial video sequence" means correctly playing the sequence of video clips sorted according to the timestamps. "Mapping the initial video sequence to the virtual scene" means mapping the logic of this time sequence to the previously constructed digital twin environment, so that the video clips are not only spatially positioned, but also temporally ordered, completing the time layout and laying a solid foundation for the final spatiotemporal continuity.

[0054] In the present embodiment, it is assumed that a time-lapse documentary about daily life in a city is being produced. The team collects multiple video clips from morning to night in different locations (such as parks, cafes, office areas, night markets) on the same day, and each clip carries an accurate timestamp.

[0055] The implementation steps are as follows: Timestamp analysis: Among all the collected video clips, we found the following timestamp information: Clip A: 8:00am, morning joggers in the park - Clip B: 10:30am, conversations in a cafe - Clip C: 6:00pm, people leaving get off work in an office - Clip D: 8:30pm, a busy night market; Adjust timing and sequencing: Based on these timestamps, determine the natural time sequence as A→B→C→D→D, which is a natural transition from morning activities to night.

[0056] Mapping the time layout to the virtual scene: In the constructed digital twin city model, the video clips are placed in this time order: Clip A is in the park scene at 8:00 am, and the characters match the ambient light - Clip B moves to the cafe at 10:30 am, and the customer interaction is consistent with the indoor lighting - Clip C goes to the office area at 6:00 pm, and the crowd matches the sunset background - Clip D is placed in the night market at 8:30 pm, and the lights blend with the lively atmosphere of the market.

[0057] In this way, the video clips not only have accurate geographical positioning in the virtual space, but also are sorted correctly according to the timestamp information, constructing a temporal layout, ensuring the spatio-temporal coherence and logic of the video in the digital twin scenario.

[0058] 104. Construct a spatio-temporal layout based on the spatial layout and the temporal layout, and in the spatio-temporal layout, identify the visual features and scene semantic tones among multiple video clips, and based on the visual features and scene semantic tones among multiple video clips, adjust the perspectives and positions of multiple video clips to generate a video clip adjustment result; In this step, the "spatio-temporal layout" refers to an integrated display framework that combines the spatial position information of video clips (determined by the spatial layout) and the time sequence (determined by the temporal layout). It not only considers the three-dimensional coordinate positions of video clips in the virtual scene, but also incorporates the time dimension, forming a four-dimensional spatio-temporal structure. "Identifying visual features and scene semantic tones" means automatically detecting visual features such as color, motion, composition, etc. and the scene semantic tones (such as joy, tranquility, tension) conveyed by the video through algorithm analysis of the content of the video clips. "Adjusting perspectives and positions" is to optimize the display angles of video clips and their specific placement points in the virtual scene according to these identified video features and scene semantic tones, so as to enhance the narrative coherence or improve the emotional experience of the audience.

[0059] In the embodiment of the present application, imagine the production process of a tourism promotion video: Construction of spatio-temporal layout: There is already a three-dimensional model of a virtual tourist city as the basic scene, and the video clips of each famous scenic spot are arranged in the playing order and spatial positions according to their actual shooting times (such as morning, noon, dusk) and geographical locations.

[0060] Visual feature and emotion recognition: Analyze through AI algorithms: Video clip 1 (beach at sunrise): Warm colors, peaceful picture, and the scene semantic tone is tranquility and expectation.

[0061] Video clip 2 (busy market at noon): Rich colors, strong dynamics, and the scene semantic tone is lively and energetic.

[0062] Video clip 3 (ancient castle at sunset): Soft tone, classical composition, and the scene semantic tone is nostalgic and romantic.

[0063] Adjustment of perspectives and positions: For video clip 1, adjust the perspective to a low angle to simulate the feeling of standing on the beach watching the sunrise, and set the position on the coastline to let the audience feel the tenderness of the first ray of sunlight.

[0064] For video clip 2, a bird's-eye view is used to show the overall picture of the market, emphasizing the bustle of people coming and going. The location is set above the center of the market to increase the audience's immersive sense of participation.

[0065] Video clip 3 uses a long lens to slowly zoom in, gradually focusing from the panoramic view of the castle to a beam of light in front of the window, strengthening the semantic tone of the scene. The location is arranged on the dusk side of the castle to match the surrounding environment.

[0066] Through such adjustments, the video clips are not only coherent and smooth in terms of time and space layout, but also more precise and attractive in terms of visual and emotional expression. The final video clip adjustment results are more in line with the expected narrative rhythm and emotional expression.

[0067] 105. Based on the video clip adjustment result, multiple video clips are spliced to generate a digital twin display video.

[0068] In this step, "video clip splicing processing" refers to the process of seamlessly connecting multiple adjusted video clips in a preset order and transition effects. This includes editing, adding transition effects, unifying colors and picture quality, and synthesizing sound tracks to ensure that the final output video is smooth and natural, as if it were a whole shot. "Digital twin display video" is a highly realistic and dynamic virtual model display method that uses digital technology to simulate objects, environments or processes in the real world or design concepts, allowing viewers to intuitively understand complex systems or experience scenes that are difficult to observe directly.

[0069] In this embodiment of the application, it is assumed that a digital twin display video of a factory production line is being created: Application of video clip adjustment results: The video clips adjusted according to the previous steps include all links from raw material storage, automated assembly line operation, quality inspection to finished product packaging and delivery. The video clips of each link have been finely adjusted according to the actual production line layout and time flow, and the perspective, light, and color have all met the consistency and coherence requirements.

[0070] Stitching process: First, arrange the adjusted video clips in the logical order of the production process. Add smooth transition effects between adjacent clips, such as fade in and fade out or slide switching, to ensure visual smoothness. Unify the color style and brightness of the entire video to ensure visual continuity from one scene to another. Synchronously add environmental sound effects and narration, such as the sound of machinery running, the sound of product movement, and the description of key steps, to enhance the comprehensiveness of information communication. Finally, perform video rendering and output settings to ensure that the video resolution, frame rate, and encoding format are suitable for the target playback platform and device.

[0071] Through the above embodiments, the completed digital twin display video not only accurately reflects the actual operation of the factory production line, but also enables the audience to deeply understand and feel every detail of the production process through the dual presentation of vision and hearing, which has extremely high value for scenarios such as training, remote monitoring, or customer demonstrations.

[0072] Optionally, in the embodiments of the present application, step 102 may specifically include: 1021. For each of the video clips, determine the corresponding virtual scene coordinates in the virtual scene according to the real scene coordinates of each video clip, where there is a corresponding coordinate system between the real scene and the virtual scene; 1022. Determine the position of each video clip in the virtual scene according to the virtual scene coordinates corresponding to each video clip; 1023. Determine the viewing angle of each video clip in the virtual scene; 1024. Generate a spatial layout based on the positions and viewing angles of multiple video clips.

[0073] In the above steps 1021 to 1024, the real scene coordinates refer to the position information in the actual physical environment recorded in the video clip, and usually a three-dimensional coordinate system (X, Y, Z) is used to represent the precise position of an object or a camera in the real world. The virtual scene is a computer-generated digital representation that mimics the real environment and contains various elements and layouts corresponding to the real world.

[0074] Among them, the virtual scene coordinates refer to the corresponding positions assigned to real-world objects or viewpoints in the virtual environment to ensure their correct reproduction in the digital twin.

[0075] Among them, the viewing angle refers to the direction and angle at which the camera in the video clip views the scene, which is crucial for reconstructing the visual experience of the real environment.

[0076] Among them, the spatial layout involves how to reasonably arrange and display these video clips in the virtual scene to maintain the consistency of spatial logic and the immersion of the viewer.

[0077] In the embodiments of the present application, assuming that a digital twin model is being created for a large shopping mall, it may include the following process: Determine the correspondence between real scene coordinates and virtual scene coordinates (i.e., step 1021): First, obtain the precise positions (real scene coordinates) of each surveillance camera in the shopping mall through GPS data and on-site measurements. Then, establish the virtual scene of the shopping mall in 3D modeling software and convert these physical coordinates into the corresponding coordinate points in the virtual scene according to the same coordinate system.

[0078] Determining the position of the video clip in the virtual scene (i.e., step 1022): Based on the transformed virtual scene coordinates, accurately position the video clip recorded by each surveillance camera at the corresponding surveillance point in the virtual shopping mall, ensuring a one-to-one correspondence between the video image and the position in the virtual environment.

[0079] Determining the viewing angle of the video clip (i.e., step 1023): Analyze the shooting direction and field of view of each video clip, and convert them into the viewing angle parameters of the virtual camera, such as pitch angle, yaw angle, and roll angle, so that the virtual viewing angle matches the actual shooting viewing angle.

[0080] Generating the spatial layout (i.e., step 1024): Considering the position distribution of all video clips and their respective viewing angle information, optimize the view switching logic in the virtual scene, design a reasonable navigation path or automatic roaming route, ensuring that users can smoothly shuttle between different shopping areas when browsing the digital twin display video, and feel the spatial continuity and realism like on-site inspection.

[0081] Through the above embodiments, not only is the accurate mapping from the physical space to the virtual space achieved, but also an efficient and intuitive digital twin experience platform for the shopping mall is provided for managers, designers or customers, facilitating asset management, layout planning or remote guided tours.

[0082] Optionally, as a possible implementation, the process of "determining the viewing angle of each video clip in the virtual scene" in step 1023 may include: for each video clip, calculate the geometric center point of the video clip in the virtual scene according to the virtual scene coordinates corresponding to the video clip and the coverage area of the video clip; determine the video content of the video clip, determine the line-of-sight direction vector of the video clip, and the video content at least includes: the moving direction of the main body, the camera pointing; determine the physical attributes of the video clip, determine the effective observation range of the video clip in the virtual scene, and the physical attributes at least include: the lens focal length; based on the geometric center point, line-of-sight direction vector, and lens focal length corresponding to the video clip, determine the viewing angle of the video clip in the virtual scene.

[0083] In this step, the geometric center point refers to the central coordinate point calculated in the virtual scene based on the corresponding position of the video clip and its covered area, which is used as the basic position for reference of the video perspective. The line-of-sight direction vector represents the direction pointed by the camera in the video clip, which is determined by analyzing the moving direction of the main body in the video content or directly the camera pointing angle. It is a directional quantity that helps to understand the directionality of video observation. Video content analysis includes identifying the main moving direction in the video or the direct pointing of the camera, which helps to infer the visual focus that the video clip wants to capture. Physical properties, especially the lens focal length, determine the field of view width (angle of view range) of the camera. The shorter the focal length, the wider the field of view, and vice versa, which affects the range of the virtual scene that the video clip can cover. The effective observation range refers to the virtual scene area that the video clip can clearly cover under a given lens focal length, which is affected by the focal length and video quality.

[0084] In the embodiment of the present application, it is assumed that it is necessary to determine the perspective of an outdoor park surveillance video in a digital twin environment: Calculate the geometric center point: The surveillance video clip shows a main road in the park. By analyzing the ground area covered by the video clip, the position of the surveillance point in the virtual park model is determined, and the geometric center coordinates of this surveillance point in the virtual scene are calculated.

[0085] Determine the line-of-sight direction vector: Video analysis shows that the camera mainly tracks the movement of pedestrians along the main road, and the direction is from west to east. Therefore, the line-of-sight direction vector is set to the due east direction, that is, (1, 0, 0) (assuming the virtual scene uses a right-handed coordinate system).

[0086] Determine the physical properties: It is known that the surveillance camera uses a standard lens with a focal length of 35 mm. Based on this focal length, its horizontal and vertical angle of view ranges can be calculated.

[0087] Determine the effective observation range: Combining the geometric center point, the line-of-sight direction vector, and the lens focal length, the maximum observation radius and angle range of the camera in the virtual scene are calculated using trigonometric functions. For example, a 35 mm lens provides a horizontal angle of view of approximately 63 degrees. This means that in the virtual environment, starting from the center point and along the line-of-sight direction vector, the camera can clearly cover a conical area with the center point as the center and a specific radius.

[0088] Determine the perspective based on the above parameters: According to the geometric center point, the line-of-sight direction vector, and the lens focal length corresponding to the video clip, the perspective of the video clip in the virtual scene is determined.

[0089] Assume known conditions: Geometric center point: , the three-dimensional coordinates of the video clip in the virtual scene.

[0090] Line-of-sight direction vector: , a unit vector, representing the direction pointed by the camera, which has been normalized to a unit vector, i.e., its magnitude is 1.

[0091] Lens focal length: , usually measured in millimeters, which determines the viewing angle range of the camera.

[0092] Effective viewing range: W, which can be understood as the expected viewing angle width to be covered.

[0093] Calculation steps: Through the formula: , based on the geometric center point, line-of-sight direction vector, and lens focal length corresponding to the video clip, determine the viewing angle of the video clip in the virtual scene; Wherein, represents the viewing angle of the video clip in the virtual scene, represents the expected viewing angle width to be covered, represents the lens focal length, and are the x and y components of the line-of-sight direction vector, and arctan is the arctangent function.

[0094] It should be noted that multiplying " " by 2 is because usually the conversion from the half-angle to the full-angle of the field of view is calculated.

[0095] In summary, through these detailed steps, not only the observation point of the video clip in the virtual environment is accurately located, but also its actual observation angle and coverage range are accurately simulated, making the digital twin environment closer to reality and helping to improve the accuracy of data analysis and decision-making.

[0096] Optionally, in the embodiment of the present application, the process of "adjusting the time sequence between multiple video clips according to the timestamps of multiple video clips to obtain the initial video order" in step 103 may include: 1031. For each video clip, extract the key frame of each video clip, and based on the key frame of each video clip, determine the visual feature and scene semantic tone of the video clip; 1032. Use the timestamp, visual feature, and scene semantic tone of the video clip as inputs, and input them into a pre-trained order sorting model to adjust the time sequence between multiple video clips through the order sorting model and output the initial video order, where the order sorting model is trained based on multiple video clip samples, and the multiple video clip samples include video sequence data with annotated timestamps, visual features, and scene semantic tones and the context information corresponding to the video sequence data.

[0097] In this step, key frame extraction and determination of visual features and scene semantic tone: Key frames are the frames selected from a video that best represent the content of that segment and are often used to quickly capture the visual summary meaning of the video. Visual features may include color distribution, texture, shape, action intensity, etc., which help to identify the video content. The scene semantic tone is the overall emotional atmosphere conveyed by the video segment, such as cheerful, sad, etc., and is usually obtained through comprehensive analysis of elements such as music, color, expressions, and actions within the video.

[0098] Sequential sorting model: This is a machine learning or deep learning model that, through training, understands the relationship between the context information, timestamps, visual features, and scene semantic tone of video segments to predict the best playback order. The model inputs include timestamps, visual features, and scene semantic tone. Based on this information, the model can learn how video segments naturally transition according to the emotional flow and time flow, and thus adjust and sort them.

[0099] In the embodiment of the present application, key frame extraction and determination of scene semantic tone: Suppose a travel vlog video is produced, and each video segment contains different landscapes and human activities. First, the algorithm automatically identifies the key frames in each segment, such as the iconic shot of the Eiffel Tower, the sunset photo by the sea, etc. Then, analyze the colors (such as the warm tones of the Eiffel Tower), actions (smiles of the crowd), and music (relaxing melody) of these key frames to determine that the scene semantic tone is pleasant and relaxing.

[0100] Application of the sequential sorting model: Use the trained model. The inputs include the timestamps of each segment (such as the Eiffel Tower at 2 pm, the beach sunset at dusk), visual features (such as color distribution, action intensity), and scene semantic tone (pleasant, relaxing). Based on the trained sample database, the model understands how emotions and time smoothly transition in different situations (such as from daytime activities to the tranquility of dusk), adjusts the segment order, and outputs an initial video playback sequence. The model database contains a large number of travel video samples, each marked with time, visual features, scene semantic tone, and context (such as the scene relevance before and after the activity), which helps the model learn how to sort naturally.

[0101] Through this implementation, the video segments are not only sorted reasonably by time but also intelligently arranged according to the fluency of emotions and visual content, improving the coherence and viewing experience of the video.

[0102] Optionally, in the embodiment of the present application, the method further includes: In the process of adjusting the timing between multiple video segments through the sequential sorting model, if there is a timestamp overlap between multiple video segments, the sequential priority between multiple video segments is determined by determining the continuity and content length of the video content between multiple video segments; or, In the process of adjusting the timing between multiple video segments through the sequential sorting model, if the time difference calculated between any two video segments is less than the set time gap, additional frames are generated between the last frame of the first video segment and the first frame of the next video segment among the two video segments to fill the time gap, and the timing between the two video segments is readjusted based on the additional frames.

[0103] In this step, timestamp overlap processing: When the event occurrence times recorded by different video segments partially overlap, in order to ensure the coherence and logic of video narration, it is necessary to determine their playback order based on the continuity of the video content and the content length of each segment. Continuity assessment involves analyzing whether the pictures and audio between video segments can transition smoothly, while the content length may affect which segment's information is more important and should be presented first.

[0104] Time gap filling: If the sequential sorting model finds that there is an overly small time interval (less than the preset threshold) between two adjacent video segments, this may cause a jump cut feeling during playback. To solve this problem, the system automatically generates additional frames and inserts them between these two segments. These additional frames may be intermediate frames generated through interpolation techniques or transitional frames synthesized based on the content of the previous and subsequent frames to smooth the time jump and enhance the viewing experience.

[0105] In the embodiment of the present application, processing timestamp overlap: Suppose a video editing project contains two video segments that record the same concert but from different perspectives, and the timestamps show that there is partial overlap between them. Through analysis, the first video segment focuses on close-ups of the singer, and the second is the audience reaction. Since the first video segment is more important for expressing the core content of the concert (the singer's performance) and has better visual continuity, it is decided to play the first video segment first and then connect the second one to ensure the clarity of the story line.

[0106] Filling the time gap: In another scenario, video segments of the morning sunrise and breakfast preparation are taken. The model calculates that there is only a two-second time difference between the two videos, which is lower than the set five-second smooth playback time gap. Therefore, the system generates several frames of transitional pictures with gradual sky colors and light changes between these two videos, simulating the process of natural light change from sunrise to the breakfast table, not only filling the time gap but also enhancing the coherence and artistic sense of the video.

[0107] Optionally, in the embodiment of the present application, as Figure 2As shown, step 104 may specifically include: 1041. Determine the visual features and scene semantic tones of each of the multiple video segments based on the key frames of the multiple video segments; In this step, first, representative key frames are selected from each video segment. These key frames often contain the core visual information and emotional expressions of the video content. Through the analysis of these key frames, we can extract the unique visual features of each video segment, such as color distribution, texture, shape, etc., and at the same time evaluate its scene semantic tone, such as cheerful, calm or sad.

[0108] 1042. In the spatio-temporal layout, identify the shared visual features among the multiple video segments based on the visual features of each video segment; In this step, within the framework of the spatio-temporal layout, the algorithm further searches for commonalities among different video segments. These shared visual features can be similar color themes, recurring patterns or compositional consistencies, which help to establish visual connections and fluency among multiple segments.

[0109] 1043. Calculate the visual rotation angle and position translation amount between two adjacent video segments according to the shared visual features among the multiple video segments; In this step, using the identified shared visual features, the system calculates how to smoothly transition between adjacent video segments, that is, by adjusting the rotation angle of the perspective and the position translation amount of the video segment on the screen, so that the transition from one segment to the next appears natural and coherent. This step ensures the smoothness of the visual flow and reduces the sense of abruptness.

[0110] 1044. Determine the perspectives and positions of the adjusted multiple video segments based on the visual rotation angle and position translation amount between two adjacent video segments and the corresponding scene semantic tones of the two adjacent video segments; In this step, when adjusting the perspective and position, the coordination of the corresponding scene semantic tones of adjacent video segments is also considered. If the emotions of two segments complement each other, such as changing from tranquility to exuberance, more obvious dynamic changes may be designed during the adjustment to strengthen this emotional transformation; conversely, if emotional coherence needs to be maintained, the adjustment will be more subtle to ensure a smooth transition of emotions.

[0111] 1045. Generate the adjusted result of the video segments based on the perspectives and positions of the adjusted multiple video segments.

[0112] In this step, by synthesizing the above analysis and calculations, the adjusted perspectives and position layouts of all video segments are finally determined to form an overall harmonious and emotionally coherent video sequence. This process not only optimizes the performance of individual segments but also strengthens the internal connections between segments, enhancing the viewing experience of the entire video.

[0113] In the embodiments of the present application, this process can be applied to the production of travel memoirs, music videos, or any form of multi-segment video editing. Through intelligent analysis and adjustment, works that are both visually unified and emotionally rich can be created. For example, in a video project showing the changing seasons, the soft color tones and vibrant scene semantic tones of the spring segment may smoothly transition to the bright colors and high energy of the summer segment. By gradually adjusting the perspective and position, the audience can feel as if they have experienced a seamless seasonal journey.

[0114] Optionally, in the embodiments of the present application, the process of "calculating the visual rotation angle between two adjacent video segments according to the shared visual features between the multiple video segments" in step 1043 may include: determining the shared visual features between two adjacent video segments, and respectively constructing a set of shared visual feature points corresponding to each video segment to generate a set of shared visual features based on the sets of shared visual feature points corresponding to the two adjacent video segments; calculating an initial rotation matrix based on the set of shared visual features, and determining the visual rotation angle between the two adjacent video segments according to the initial rotation matrix. In this step, "shared visual features" refer to visual elements that commonly appear in two adjacent video segments, such as the same object, color, line direction, or texture pattern. These features can serve as a bridge connecting the two segments to help achieve a smooth transition. The "set of shared visual feature points" refers to a set of specific pixel points or feature points selected from each video segment that can represent these shared features. These sets of points can quantify and compare the similarity and differences in the visual content of the two videos.

[0115] Among them, the "initial rotation matrix" is a mathematical tool in computer vision and graphics processing used to describe the rotational transformation from one coordinate system to another. In this context, it is used to represent how to rotate the first video segment so that its visual features can better match the corresponding features of the second video segment, thereby ensuring visual continuity and consistency.

[0116] In the embodiments of the present application, assume that we are editing two video segments. Segment A shows a beach scene at sunset, and segment B immediately follows with the same beach under the night sky. The shared visual features may include the position of the sea level, a specific rock formation, or the outline of a certain mountain in the distance. We first extract these shared feature points from A and B respectively to form two sets of feature points.

[0117] Next, the best matching method of these point sets is calculated by an algorithm to obtain an initial rotation matrix. For example, if it is found through analysis that segment B needs to be rotated counterclockwise by 15 degrees relative to segment A to align the sea level feature points, then this 15 degrees is the calculated visual rotation angle. In addition, scaling and translation also need to be considered, but the focus here is on rotation. Therefore, the perspective of segment A is adjusted to rotate it by 15 degrees. In this way, when segment A transitions to segment B, the viewer sees the continuity of the sea level, feeling natural and smooth.

[0118] In summary, this embodiment ensures that the overall visual effect of the video remains coherent and harmonious even in the case of significant time or light changes in the scene by precisely calculating and adjusting the visual rotation angle between video segments.

[0119] Based on the above, the process of "determining the perspectives of the adjusted multiple video segments based on the visual rotation angle between two adjacent video segments and the scene semantic tone corresponding to the two adjacent video segments" in step 1044 may include: Through the formula: , determine the perspectives of the adjusted multiple video segments based on the visual rotation angle between two adjacent video segments and the scene semantic tone corresponding to the two adjacent video segments; Among them, R′ represents the adjusted rotation matrix, R represents the initial rotation matrix, represents the coefficient for adjusting the initial rotation matrix by the scene semantic tone, S represents the quantization value of the scene semantic tone, the exp function represents the matrix exponential function, represents the diagonal linear transformation of the identity matrix; Based on the adjusted rotation matrix, determine the perspectives of the adjusted multiple video segments.

[0120] In this step, the "adjusted rotation matrix R′" refers to a new matrix obtained by modifying the rotation matrix R initially calculated based on visual features after considering the scene semantic tone of the video segment. This adjustment aims to make the transition of the video segment not only visually continuous but also maintain consistency or produce the desired emotional gradient effect in the emotional atmosphere. " " is a coefficient that determines the degree of influence of the scene semantic tone on the rotation angle, and its magnitude reflects the importance of the adjustment of the scene semantic tone. "S" represents the quantization value of the scene semantic tone, which is usually a numerical index that can reflect the intensity of the emotion (such as happiness, sadness, tension, etc.) of the video segment. The "exp function" as the matrix exponential function is applied here to convert the product of the diagonal matrices into an adjustment factor of the rotation matrix, ensuring the mathematical rationality of the rotation operation. And " is the stretching transformation of the identity matrix along its diagonal direction, which is used to construct the basic structure for rotation adjustment.

[0121] In the embodiments of the present application, assume there are two video clips. Clip A shows a lively party scene, and Clip B is a quiet sunset landscape. Through the previous steps, we have calculated that the initial rotation matrix R from A to B is a matrix that rotates 30 degrees to match the position of the horizon in the scene. The analysis of the scene semantic tone shows that the quantified value of the scene semantic tone of Clip A = 8 (representing highly active), and the quantified value of the scene semantic tone of Clip B = 6 (relatively calm), indicating that the scene semantic tone transitions from active to calm.

[0122] To reflect this emotional change, set = 0.1 to moderately incorporate emotional factors. Then the adjusted rotation matrix can be obtained by formula calculation: ; Substitute specific values: ; This means that the adjustment of the scene semantic tone will slightly reduce the originally calculated rotation angle, reflecting the visual expression of the transition from active to calm emotions. Finally, based on the adjusted rotation matrix , the video editing software will correspondingly adjust the perspective of Clip B, making it not only visually smooth but also achieving a natural transition in emotion from the cheerfulness of Clip A to the tranquility of Clip B.

[0123] It should be noted that the final adjustment matrix represents the rotation angle after the perspective adjustment of each video clip. In practical applications, this adjusted rotation matrix will be applied to the perspective transformation of each video clip, such as by changing the orientation vector or rotation parameters of the camera, so that the video clip is displayed from a new perspective in the virtual scene. In this way, the video clips are not only correctly aligned in spatial layout but also finely adjusted visually according to the scene semantic tone, making the overall video sequence coherent and smooth both visually and emotionally.

[0124] In this way, the combination of the initial rotation matrix based on visual features and the scene semantic tone provides precise guidance for the calculated adjustment matrix to adjust the perspective of the video clip, thus achieving natural transitions and emotional expressions in the spatio-temporal layout.

[0125] Optionally, in the embodiments of the present application, the process of "calculating the position translation amount between two adjacent video segments according to the shared visual features between multiple video segments" in step 1043 may include: determining the shared visual features between two adjacent video segments, and calculating the average displacement between two adjacent video segments based on the shared visual features between two adjacent video segments to determine the position translation amount between two adjacent video segments; In this step, "shared visual features" refer to significant and recognizable visual elements that co - appear in two adjacent video segments, such as the same landmarks, textures, colors, shapes, or specific patterns. They are the key to connecting video segments and help to judge the spatial relationship and position transformation between segments. "Position translation amount" refers to the displacement distance of the relative position between segments that needs to be adjusted to make the adjacent video segments transition smoothly in the spatial layout. This is calculated based on the matching of shared features and visual coherence to ensure scene coherence and logic.

[0126] In the embodiments of the present application, assume that a documentary of urban scenery is being produced. Segment A shows the lake view in a city park, and segment B follows with the buildings on the other side of the lake. The shared visual feature is the contour line of the lake edge, and the pixel point set of the contour is identified through image analysis software.

[0127] Determine the shared visual features: First, through analysis, it is found that the lake shoreline shared by segment A and B, and the contour feature point set of the lake edge is extracted.

[0128] Calculate the average displacement: By aligning the corresponding feature point sets of the lake boundaries in the two segments, calculate their average offset difference value. For example, it is calculated that segment B needs to be translated 5 meters to the right (in the unit of the virtual space) relative to A to align the lake edge.

[0129] Determine the position translation amount: According to the calculation of the average displacement, it is determined that segment B needs to be translated 5 meters to the right relative to A. This is the position translation amount between the two video segments, ensuring that when transitioning from A to B, the lake looks smooth visually without any abruptness, and the spatial continuity is maintained.

[0130] Therefore, the method of calculating the position translation amount based on shared visual features ensures the spatial continuity and smooth visual transition between video segments, enhancing the immersive experience of the audience.

[0131] Based on the above content, the process of "determining the positions of the adjusted multiple video segments based on the position translation amount between two adjacent video segments and the scene semantic tones corresponding to the two adjacent video segments" in step 1044 may include: Through the formula group: , determine the positions of the adjusted multiple video segments based on the position translation amount between two adjacent video segments and the scene semantic tone corresponding to the two adjacent video segments; Among them, represents the x component in the position translation amount, represents the y component in the position translation amount, represents the adjustment factor based on the scene semantic tone, and S represents the quantization value of the scene semantic tone, represents the adjustment factor based on the shared visual features, and F represents the visual feature matching tightness, which is used to measure the consistency of the shared visual features between two adjacent video segments.

[0132] In this step, the position translation amount (Δx, Δy) refers to the displacement distance between adjacent video segments calculated based on the shared visual features on the x-axis and y-axis.

[0133] Adjustment factor of scene semantic tone : A coefficient adjusted based on the scene semantic tone, used to enhance or weaken the displacement amount to conform to the emotional coherence between video segments.

[0134] Quantization value S of scene semantic tone: Reflects the intensity of the emotion of the video segment, such as excitement, calmness, etc.

[0135] Adjustment factor of shared visual features : The weight adjusted based on the visual feature matching tightness, which helps to refine the displacement amount and ensure visual consistency.

[0136] Visual feature matching degree F: An index used to measure the similarity of visual features between adjacent segments.

[0137] By substituting these parameters into the formula group: , the final positions of the video segments on the x-axis and y-axis can be adjusted to adapt to the emotion and visual feature matching degree.

[0138] In the embodiment of the present application, assume that a short natural scenery film with emotional ups and downs is being produced: Segment A: The lake is calm in the early morning, with light fog, and the tone quantization value 𝑆𝐴 = 5 (calm) Segment B: The waterfall is turbulent in the afternoon, full of dynamics, and the tone quantization value 𝑆𝐵 = 8 (active) Position translation amount: Analyze the shared visual feature - the lakeside line, and calculate that segment A needs to move 10 meters to the right and Δy = 0 (keep horizontal unchanged) to connect to the lake surface when moving from segment A to segment B.

[0139] Scene semantic tone adjustment: Set to 0.5 to moderately fuse the emotional changes, Set to 1 to emphasize the importance of visual matching.

[0140] Calculate the adjusted position: Δx′ = 10 + 0.5 × 5 × 1 × 1 × 10 = 15 meters, Δy' = 0.

[0141] In the embodiment of the present application, based on the lake surface characteristics, the original 10-meter right shift is adjusted to 15 meters, enhancing the visual transition from tranquility to vitality, reflecting the dual considerations of emotion and visual characteristics, and making the transition between video segments natural and emotionally coherent.

[0142] Figure 3 It is a schematic structural diagram of a video splicing system based on a digital twin scenario provided by an embodiment of the present application. As Figure 3 shown, the system includes: An acquisition module 31, configured to acquire a set of video segments to be spliced, and each video segment in the set of video segments carries a corresponding real-scene coordinate and timestamp; A construction module 32, configured to construct a virtual scene in a digital twin environment, and map multiple video segments into the virtual scene according to the real-scene coordinates of the multiple video segments to construct a spatial layout; adjust the time sequence between the multiple video segments according to the timestamps of the multiple video segments to obtain an initial video order, and map the initial video order into the virtual scene to construct a time layout; construct a spatio-temporal layout based on the spatial layout and the time layout; A generation module 33, configured to identify visual features and scene semantic tones between multiple video segments in the spatio-temporal layout, and adjust the perspectives and positions of the multiple video segments based on the visual features and scene semantic tones between the multiple video segments to generate a video segment adjustment result; perform splicing processing on the multiple video segments based on the video segment adjustment result to generate a digital twin display video.

[0143] Optionally, in the embodiment of the present application, the construction module 32 is specifically configured to, for each video segment, determine the virtual scene coordinates corresponding to the real-scene coordinates in the virtual scene according to the real-scene coordinates of each video segment, where there is a corresponding coordinate system between the real scene and the virtual scene; determine the position of each video segment in the virtual scene according to the virtual scene coordinates corresponding to each video segment; determine the perspective of each video segment in the virtual scene; generate a spatial layout based on the positions and perspectives of the multiple video segments.

[0144] Optionally, in the embodiments of the present application, the building block 32 is further configured to, for each of the video segments, calculate the geometric center point of the video segment in the virtual scene according to the virtual scene coordinates corresponding to the video segment and the coverage area of the video segment; determine the video content of the video segment, determine the line-of-sight direction vector of the video segment, where the video content at least includes: the moving direction of the subject, the camera pointing direction; determine the physical attributes of the video segment, determine the effective viewing range of the video segment in the virtual scene, where the physical attributes at least include: the lens focal length; and determine the viewing angle of the video segment in the virtual scene according to the geometric center point, the line-of-sight direction vector, and the lens focal length corresponding to the video segment.

[0145] Optionally, in the embodiments of the present application, the building block 32 is further configured to, for each of the video segments, extract the key frames of each of the video segments, and determine the visual features and scene semantic tones of the video segment based on the key frames of each of the video segments; use the time stamp, visual features, and scene semantic tones of the video segment as inputs, and input them into a pre-trained sequential sorting model, so as to adjust the time sequence between multiple video segments through the sequential sorting model, and output an initial video sequence, where the sequential sorting model is trained based on multiple video segment samples, and the multiple video segment samples include video sequence data with labeled time stamps, visual features, and scene semantic tones, as well as the context information corresponding to the video sequence data.

[0146] Optionally, in the embodiments of the present application, the system further includes a processing module 34; When the processing module 34 adjusts the time sequence between multiple video segments through the sequential sorting model, if there is a time stamp overlap between multiple video segments, the order priority between multiple video segments is determined by determining the continuity and content length of the video content between multiple video segments; or, when the processing module 34 adjusts the time sequence between multiple video segments through the sequential sorting model, if the calculated time difference between any two video segments is less than a set time gap, an additional frame is generated between the last frame of the first video segment and the first frame of the next video segment among the two video segments to fill the time gap, and the time sequence between the two video segments is re-adjusted based on the additional frame.

[0147] Optionally, in the embodiments of the present application, the generating module 33 is specifically configured to determine the visual features and scene semantic tones of each video segment based on the key frames of multiple video segments; in the spatio-temporal layout, identify the shared visual features among multiple video segments based on the visual features of each video segment; calculate the visual rotation angle and position translation amount between two adjacent video segments according to the shared visual features among multiple video segments; determine the perspectives and positions of the adjusted multiple video segments based on the visual rotation angle and position translation amount between two adjacent video segments and the corresponding scene semantic tones of two adjacent video segments; and generate a video segment adjustment result based on the perspectives and positions of the adjusted multiple video segments.

[0148] Optionally, in the embodiments of the present application, the generating module 33 is further configured to determine the shared visual features between two adjacent video segments, and respectively construct a shared visual feature point set corresponding to each video segment, so as to generate a shared visual feature set according to the shared visual feature point sets corresponding to two adjacent video segments; calculate an initial rotation matrix based on the shared visual feature set, and determine the visual rotation angle between two adjacent video segments according to the initial rotation matrix. Optionally, in the embodiments of the present application, the generating module 33 is further configured to use the formula: , to determine the perspectives of the adjusted multiple video segments based on the visual rotation angle between two adjacent video segments and the corresponding scene semantic tones of two adjacent video segments; wherein, R′ represents the adjusted rotation matrix, R represents the initial rotation matrix, represents the coefficient for adjusting the initial rotation matrix by the scene semantic tone, S represents the quantization value of the scene semantic tone, the exp function represents the matrix exponential function, represents the diagonal linear transformation of the identity matrix; Determine the perspectives of the adjusted multiple video segments based on the adjusted rotation matrix.

[0149] Optionally, in the embodiments of the present application, the generating module 33 is further configured to determine the shared visual features between two adjacent video segments, and calculate the average displacement between two adjacent video segments based on the shared visual features between two adjacent video segments, so as to determine the position translation amount between two adjacent video segments. Optionally, in the embodiments of the present application, the generating module 33 is further configured to use the formula group: , determine the positions of the adjusted multiple video segments based on the position translation amount between two adjacent video segments and the scene semantic tones corresponding to the two adjacent video segments; wherein, represents the x component in the position translation amount, represents the y component in the position translation amount, represents the adjustment factor based on the scene semantic tone, S represents the quantization value of the scene semantic tone, represents the adjustment factor based on the shared visual feature, F represents the visual feature matching tightness, which is used to measure the consistency of the shared visual features between two adjacent video segments.

[0150] Figure 3 The video stitching system based on the digital twin scene can execute Figure 1 the video stitching method based on the digital twin scene described in the embodiments shown, and its implementation principle and technical effects will not be elaborated. For the video stitching system based on the digital twin scene in the above embodiments, the specific ways for each module and unit to perform operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0151] In a possible design, Figure 3 the video stitching system based on the digital twin scene in the embodiments shown can be implemented as a computing device, such as Figure 4 shown, the computing device may include a storage component 41 and a processing component 42; The storage component 41 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 42.

[0152] The processing component 42 is used to: obtain a set of video segments to be stitched, and each video segment in the set of video segments carries a corresponding real - scene coordinate and timestamp; construct a virtual scene in the digital twin environment, and map the multiple video segments into the virtual scene according to the real - scene coordinates of the multiple video segments to construct a spatial layout; adjust the time sequence between the multiple video segments according to the timestamps of the multiple video segments to obtain an initial video order, and map the initial video order into the virtual scene to construct a time layout; construct a spatio - temporal layout based on the spatial layout and the time layout, and in the spatio - temporal layout, identify the visual features and scene semantic tones between the multiple video segments, and adjust the perspectives and positions of the multiple video segments based on the visual features and scene semantic tones between the multiple video segments to generate a video segment adjustment result; perform stitching processing on the multiple video segments based on the video segment adjustment result to generate a digital twin display video.

[0153] Among them, the processing component 42 may include one or more processors to execute computer instructions to complete all or part of the steps in the above methods. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above methods.

[0154] The storage component 41 is configured to store various types of data to support the operation of the terminal. The storage component may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc.

[0155] The display component 43 may be an electroluminescent (EL) element, a liquid crystal display or a microdisplay with a similar structure, or a retina-direct display or a similar laser scanning display.

[0156] Of course, the computing device may also necessarily include other components, such as input / output interfaces, communication components, etc.

[0157] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above peripheral interface module may be an output device, an input device, etc.

[0158] The communication component is configured to facilitate wired or wireless communication between the computing device and other devices, etc.

[0159] Among them, the computing device may be a physical device or an elastic computing host provided by a cloud computing platform, etc. At this time, the computing device may refer to a cloud server, and the above processing component, storage component, etc. may be basic server resources leased or purchased from the cloud computing platform.

[0160] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the above Figure 1 video stitching method based on the digital twin scenario shown in the embodiment.

[0161] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0162] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0163] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A video splicing method based on a digital twin scenario, characterized in that, Including: Obtain a set of video segments to be spliced, where each video segment in the set of video segments carries a corresponding real - world scene coordinate and a timestamp; Construct a virtual scene in a digital twin environment, and map multiple video segments into the virtual scene according to the real - world scene coordinates of the multiple video segments to construct a spatial layout; Adjust the time sequence between multiple video segments according to the timestamps of the multiple video segments to obtain an initial video order, and map the initial video order into the virtual scene to construct a time layout; Construct a spatio - temporal layout based on the spatial layout and the time layout, and in the spatio - temporal layout, identify the visual features and scene semantic tones between multiple video segments, and adjust the perspectives and positions of multiple video segments based on the visual features and scene semantic tones between multiple video segments to generate a video segment adjustment result; Based on the video segment adjustment result, perform splicing processing on multiple video segments to generate a digital twin display video.

2. The method according to claim 1, wherein The mapping of multiple video segments into the virtual scene according to the real - world scene coordinates of the multiple video segments to construct a spatial layout includes: For each video segment, determine the virtual scene coordinate corresponding to the real - world scene coordinate in the virtual scene according to the real - world scene coordinate of each video segment, where there is a corresponding coordinate system between the real - world scene and the virtual scene; Determine the position of each video segment in the virtual scene according to the virtual scene coordinate corresponding to each video segment; Determine the perspective of each video segment in the virtual scene; Generate a spatial layout based on the positions and perspectives of multiple video segments.

3. The method according to claim 2, wherein The determination of the perspective of each video segment in the virtual scene includes: For each video segment, calculate the geometric center point of the video segment in the virtual scene according to the virtual scene coordinate corresponding to the video segment and the coverage area of the video segment; Determine the video content of the video segment, and determine the line - of - sight direction vector of the video segment. The video content at least includes: the moving direction of the main body and the camera pointing direction; Determine the physical attributes of the video segment, and determine the effective observation range of the video segment in the virtual scene. The physical attributes at least include: the lens focal length; Determine the perspective of the video segment in the virtual scene according to the geometric center point, line - of - sight direction vector, and lens focal length corresponding to the video segment.

4. The method according to claim 1, wherein The adjustment of the time sequence between multiple video segments according to the timestamps of the multiple video segments to obtain an initial video order includes: For each video segment, extract the key frame of each video segment, and determine the visual features and scene semantic tones of the video segment based on the key frame of each video segment; Taking the timestamps, visual features, and scene semantic tones of the video clips as inputs, input them into a pre-trained sequential sorting model to adjust the temporal order among multiple video clips through the sequential sorting model and output an initial video order, where the sequential sorting model is trained based on multiple video clip samples, and the multiple video clip samples include video sequence data with annotated timestamps, visual features, and scene semantic tones, as well as the context information corresponding to the video sequence data.

5. The method according to claim 4, characterized in that It further includes: During the process of adjusting the temporal order among multiple video clips through the sequential sorting model, if there is a timestamp overlap among multiple video clips, determine the order priority among multiple video clips by determining the continuity and content length of the video content among multiple video clips; Or, During the process of adjusting the temporal order among multiple video clips through the sequential sorting model, if the calculated time difference between any two video clips is less than a set time gap, generate additional frames between the last frame of the first video clip and the first frame of the next video clip among the two video clips to fill the time gap, and re-adjust the temporal order between the two video clips based on the additional frames.

6. The method according to claim 1, wherein In the spatio-temporal layout, identifying the visual features and scene semantic tones among multiple video clips, and based on the visual features and scene semantic tones among multiple video clips, adjusting the perspectives and positions of multiple video clips to generate a video clip adjustment result, including: Based on the key frames of multiple video clips, determining the visual features and scene semantic tones of each video clip; In the spatio-temporal layout, based on the visual features of each video clip, identifying the shared visual features among multiple video clips; According to the shared visual features among multiple video clips, calculating the visual rotation angle and position translation amount between two adjacent video clips; Based on the visual rotation angle and position translation amount between two adjacent video clips and the scene semantic tones corresponding to two adjacent video clips, determining the perspectives and positions of the adjusted multiple video clips; Based on the perspectives and positions of the adjusted multiple video clips, generating a video clip adjustment result.

7. The method according to claim 6, characterized in that, According to the shared visual features among multiple video clips, calculating the visual rotation angle between two adjacent video clips, including: Determining the shared visual features between two adjacent video clips, and respectively constructing a set of shared visual feature points corresponding to each video clip to generate a shared visual feature set according to the sets of shared visual feature points corresponding to two adjacent video clips; Calculating an initial rotation matrix based on the shared visual feature set, and determining the visual rotation angle between two adjacent video clips according to the initial rotation matrix; Based on the visual rotation angle between two adjacent video clips and the scene semantic tones corresponding to two adjacent video clips, determining the perspectives of the adjusted multiple video clips; Based on the adjusted rotation matrix, determine the perspectives of the adjusted multiple video segments.

8. The method according to claim 6, wherein According to the shared visual features among the multiple video segments, calculate the position translation amount between two adjacent video segments, including: Determine the shared visual features between two adjacent video segments, and based on the shared visual features between two adjacent video segments, calculate the average displacement between two adjacent video segments to determine the position translation amount between two adjacent video segments; Based on the position translation amount between two adjacent video segments and the scene semantic tone corresponding to two adjacent video segments, determine the positions of the adjusted multiple video segments.

9. A video splicing system based on a digital twin scenario, characterized in that, Including: An acquisition module, configured to acquire a set of video segments to be spliced, where each video segment in the set of video segments carries a corresponding real-scene coordinate and timestamp; A construction module, configured to construct a virtual scene in a digital twin environment, and map the multiple video segments into the virtual scene according to the real-scene coordinates of the multiple video segments to construct a spatial layout; according to the timestamps of the multiple video segments, adjust the time sequence between the multiple video segments to obtain an initial video order, and map the initial video order into the virtual scene to construct a time layout; Construct a spatio-temporal layout based on the spatial layout and the time layout; A generation module, configured to identify the visual features and scene semantic tones among the multiple video segments in the spatio-temporal layout, and based on the visual features and scene semantic tones among the multiple video segments, adjust the perspectives and positions of the multiple video segments to generate a video segment adjustment result; based on the video segment adjustment result, perform splicing processing on the multiple video segments to generate a digital twin display video.

10. A computing device, characterized in that, Including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the video splicing method based on a digital twin scene according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multimedia fragment splicing method and device, mobile terminal and storage medium

    CN110691276A

  • Method and system for generating virtual world scene in meta-cosmic space

    CN115661417A

  • Splicing method and system for monitoring videos of fully mechanized coal mining face

    CN116320779A

  • Digital twinborn scene intelligent generation method based on multi-modal visual identification

    CN117456136A

  • Image information processing method for short video production

    CN119152096A

Cited By

  • Simulation model construction method and device, electronic equipment and readable storage medium

    CN121179408A