Multi-video 3D view generation method, apparatus, and electronic device

CN122741679APending Publication Date: 2026-09-11BEIJING RUIJIA HESHUO TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610242121.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-28
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0004]然而,现有技术多集中于对于2D图像的处理,在实现多视频流融合并生成连续、一致的3D多维动态视角方面,仍缺乏一个高效的、端到端的、高度协同优化的系统来实时地完成从多个视频流输入到高质量3D多维视角视频流的全自动生成

Benefits of technology

[0017] The beneficial effects of the technical solution provided by the embodiments of the present invention are as follows: The multi-video 3D perspective generation method, apparatus, and electronic device of the present invention first acquire multiple video streams of the same scene, different angles, and time-aligned, and then process the multiple video streams to obtain the motion trajectory of the moving target and the segmentation result of the target region, thereby accurately separating the foreground moving target from the background, thus laying the foundation for subsequent perspective fusion; then, based on the attention mechanism, the segmentation results of different angles at the same timestamp are spatiotemporally aligned and fused, and perspective interpolation is used to generate an image sequence with 3D multi-dimensional perspective effect. In this way, the attention mechanism ensures the feature consistency of the same target, and perspective interpolation can synthesize any virtual perspective among a limited number of real perspectives, thus laying the foundation for subsequent encoding based on the image sequence to generate a smooth and natural 3D multi-dimensional perspective effect video stream. The present invention can efficiently and automatically integrate multiple synchronous video streams to generate spatiotemporally consistent and smooth 3D multi-dimensional perspective videos, achieving smooth and natural 3D multi-dimensional perspective transitions and meeting the real-time requirements of live streaming scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122741679A_ABST
    Figure CN122741679A_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, and electronic device for generating multi-video 3D perspectives, relating to the field of video processing technology. The method includes: acquiring multiple video streams from the same scene but at different angles and time-aligned; processing the multiple video streams to obtain the motion trajectory of a moving target and the segmentation results of the target region; performing spatiotemporal feature alignment and fusion of the segmentation results at different angles under the same timestamp based on an attention mechanism, and generating an image sequence with 3D multi-dimensional perspective effects using perspective interpolation; encoding the image sequence to obtain a video stream with 3D multi-dimensional perspective effects. This invention can efficiently and automatically integrate multiple synchronous video streams to generate spatiotemporally consistent and smooth 3D multi-dimensional perspective videos, achieving smooth and natural 3D multi-dimensional perspective transitions and meeting the real-time requirements of live streaming scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of video processing technology, and specifically relates to a method, apparatus and electronic device for generating multi-video 3D perspectives. Background Technology

[0002] During live sports broadcasts, 3D multi-dimensional perspective rotation is a visually striking effect. It creates an immersive experience by freezing a moment in time and simultaneously changing the perspective around the subject in space, allowing for a holistic observation and a sense of timelessness.

[0003] In recent years, with the development of computer vision and artificial intelligence technologies, especially the breakthroughs in deep learning in object segmentation, 3D reconstruction, and view synthesis, these technological advancements have provided potential underlying capabilities for 3D video generation. For example, instance segmentation models can accurately separate foreground objects from the background.

[0004] However, existing technologies mostly focus on processing 2D images. In terms of fusing multiple video streams and generating continuous, consistent 3D multi-dimensional dynamic viewpoints, there is still a lack of an efficient, end-to-end, highly collaboratively optimized system to automatically generate high-quality 3D multi-dimensional viewpoint video streams from multiple video stream inputs in real time. Therefore, there is an urgent need for an integrated technical solution capable of efficiently and automatically generating 3D multi-dimensional viewpoint video effects. Summary of the Invention

[0005] To address the aforementioned issues, this invention provides a method, apparatus, and electronic device for generating multi-video 3D perspectives, which can efficiently and automatically integrate multiple synchronous video streams to generate spatiotemporally consistent and smooth 3D multi-dimensional perspective videos, achieving smooth and natural 3D multi-dimensional perspective transitions and meeting the real-time requirements of live streaming scenarios.

[0006] In a first aspect, the present invention provides a method for generating multi-video 3D perspectives, comprising: Acquire multiple video streams from the same scene but different angles and with time alignment; The multiple video streams are processed to obtain the motion trajectory of the moving target and the segmentation result of the target region; Based on the attention mechanism, the segmentation results from different angles at the same time stamp are aligned and fused in terms of spatiotemporal features, and the image sequence with 3D multi-dimensional view effect is generated by perspective interpolation. The image sequence is encoded to obtain a 3D multi-dimensional viewpoint video stream.

[0007] In an optional implementation, processing the plurality of video streams to obtain the motion trajectory of the moving target and the segmentation result of the target region includes: The motion targets in each video stream are detected using a pre-established target detection module to obtain the target detection results. Based on the target detection results and the pre-established target tracking module, the moving target is tracked to obtain the motion trajectory of the moving target; The target region is segmented based on the pre-established segmentation module and the target detection results to obtain the segmentation results.

[0008] In an optional implementation, the target detection module, the target tracking module, and the segmentation module share a backbone feature extraction network, which is used to extract features from the multiple video streams. The target detection module, the target tracking module, and the segmentation module are each implemented through a task header network connected after the backbone feature extraction network.

[0009] In an optional implementation, the backbone feature extraction network is a lightweight network, and / or the method further includes lightweighting the backbone feature extraction network and the task head network.

[0010] In an optional implementation, the step of aligning and fusing spatiotemporal features of segmentation results from different angles at the same timestamp based on an attention mechanism, and generating an image sequence with 3D multi-dimensional viewpoint effects using viewpoint interpolation, includes: The Transformer or cross-attention mechanism is used to determine the associated feature data of the same moving target under different viewpoints and timestamps; Based on the associated feature data, viewpoint interpolation, texture fusion, and background processing are performed to obtain an image sequence with a 3D multi-dimensional viewpoint effect.

[0011] In an optional implementation, the target tracking module includes an active trajectory library submodule, a Kalman prediction submodule, a data association submodule, and a trajectory recovery submodule; The active trajectory library submodule is used to store the state information of moving targets; The Kalman prediction submodule is used to predict the expected position information of the moving target in the current frame based on the state of the moving target in the previous moment. The data association submodule is used to perform target matching based on the target detection results of the current frame and the expected location information, and using Mahalanobis distance and intersection-union ratio to obtain the matching result; wherein, the target matching is achieved by solving the optimal association using the Hungarian algorithm; The trajectory recovery submodule is used to reassociate and recover the interrupted trajectory based on the matching result, so as to obtain the continuous motion trajectory of the moving target.

[0012] In an optional implementation, the segmentation module incorporates a temporal constraint loss function and optical flow guidance during training.

[0013] In an optional implementation, the target detection module includes a feature pyramid submodule and an attention mechanism submodule.

[0014] Secondly, the present invention provides a multi-video 3D perspective generation device, comprising: The acquisition module is used to acquire multiple video streams from the same scene but different angles and with time alignment. The processing module is used to process the multiple video streams to obtain the motion trajectory of the moving target and the segmentation result of the target region; The generation module is used to perform spatiotemporal feature alignment and fusion of segmentation results from different angles at the same time stamp based on an attention mechanism, and to generate image sequences with 3D multi-dimensional view effects using viewpoint interpolation. The encoding module is used to encode the image sequence to obtain a 3D multi-dimensional viewpoint video stream.

[0015] Thirdly, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described in any of the foregoing embodiments.

[0016] Fourthly, the present invention provides a computer-readable medium having processor-executable non-volatile program code, the program code causing the processor to perform the method described in any of the foregoing embodiments.

[0017] The beneficial effects of the technical solution provided by the embodiments of the present invention are as follows: The multi-video 3D perspective generation method, apparatus, and electronic device of the present invention first acquire multiple video streams of the same scene, different angles, and time-aligned, and then process the multiple video streams to obtain the motion trajectory of the moving target and the segmentation result of the target region, thereby accurately separating the foreground moving target from the background, thus laying the foundation for subsequent perspective fusion; then, based on the attention mechanism, the segmentation results of different angles at the same timestamp are spatiotemporally aligned and fused, and perspective interpolation is used to generate an image sequence with 3D multi-dimensional perspective effect. In this way, the attention mechanism ensures the feature consistency of the same target, and perspective interpolation can synthesize any virtual perspective among a limited number of real perspectives, thus laying the foundation for subsequent encoding based on the image sequence to generate a smooth and natural 3D multi-dimensional perspective effect video stream. The present invention can efficiently and automatically integrate multiple synchronous video streams to generate spatiotemporally consistent and smooth 3D multi-dimensional perspective videos, achieving smooth and natural 3D multi-dimensional perspective transitions and meeting the real-time requirements of live streaming scenarios. Attached Figure Description

[0018] Figure 1 A flowchart illustrating the multi-video 3D perspective generation method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the target detection module provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the target detection module provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the system principle of the multi-video 3D perspective generation device provided in an embodiment of the present invention; Figure 5 A schematic diagram of the system principle of an electronic device provided in an embodiment of the present invention.

[0019] In the diagram: 100 - Acquisition module; 200 - Processing module; 300 - Generation module; 400 - Encoding module; 1000 - Electronic device; 1001 - Communication interface; 1002 - Processor; 1003 - Memory; 1004 - Bus. Detailed Implementation

[0020] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0021] Reference Figure 1 A method for generating multi-video 3D perspectives includes the following steps S100 to S400.

[0022] Step S100: Acquire multiple video streams from the same scene but different angles and time-aligned.

[0023] Specifically, the method in this embodiment is applied to a server. Optionally, the server is equipped with an NVIDIA RTX4090 GPU, 8-channel memory, and a high-performance AMD CPU, supporting high-speed parallel computing; it uses a RAID0 SSD array storage system, supporting high-speed data read and write, ensuring high efficiency in video data processing and storage.

[0024] The server acquires N video streams with strictly synchronized timestamps. These video streams are based on different angles of the same scene, ensuring that all video streams are strictly aligned in time. This embodiment is an artificial intelligence-based method that can perform target segmentation and viewpoint fusion on multi-angle videos of the same scene in real time using only a limited number of video streams, automatically generating high-quality, 3D multi-dimensional viewpoint video streams.

[0025] Step S200: Process multiple video streams to obtain the motion trajectory of the moving target and the segmentation result of the target region.

[0026] In an optional embodiment, step S200 includes the following steps S201 to S203.

[0027] Step S201: Use the pre-established target detection module to perform target detection on the moving targets in each video stream to obtain the target detection results.

[0028] Step S202: Based on the target detection results and the pre-established target tracking module, target tracking is performed on the moving target to obtain the motion trajectory of the moving target.

[0029] Specifically, in this embodiment, the moving target in each frame of the video stream is detected in real time through step S201, and the target tracking module is used to continuously track the specified moving target to ensure that the ID remains unchanged even when it is occluded.

[0030] Here, "keeping the ID unchanged" is a key technical indicator. It ensures the continuity of the tracked target's identity over time. Furthermore, subsequent instance segmentation and dynamic rotation viewpoint generation modules can operate based on a stable and reliable target trajectory, thus outputting smooth, realistic, and consistent 3D multi-dimensional viewpoint video effects. Therefore, this ID is the algorithm's internal "identity card" used to distinguish different tracked targets, uniquely and consistently identifying the same object across consecutive frames, and is unrelated to user identity information.

[0031] This embodiment employs a lightweight deep neural network to perform real-time target detection (such as athletes, vehicles, etc.) on the input video stream and utilizes a multi-target tracking algorithm to associate targets across frames and form motion trajectories. Here, ResNet is used as the backbone network (i.e., the backbone feature extraction network), and residual connections (Shortcut Connections) and BatchNormalization layers are introduced to solve the gradient vanishing problem of deep networks. In addition, multi-scale feature fusion is performed through a Feature Pyramid Network (FPN), combined with an attention mechanism to focus on important regions, significantly improving the accuracy and robustness of target detection. This structure can accurately identify targets (such as athletes) and their postures (such as jumping and throwing) in video frames.

[0032] Step S203: The target region is segmented according to the pre-established segmentation module and the target detection results to obtain the segmentation results.

[0033] Specifically, this embodiment uses a segmentation module to finely segment the athletes within the target area, outputting precise pixel-level masks to form accurate matting. After optimization, this model can be accelerated on a GPU, ensuring that the processing time for each frame of matting remains within 15 milliseconds. Here, by controlling the matting time, the stringent requirements of real-time video streams for processing speed are met, ensuring that the entire "dynamic freeze-frame rotation perspective" system can function as a whole, completing all calculations within the extremely short time specified for each frame. This achieves low latency and high frame rate output, ultimately reaching the "real-time" effect required for applications such as live sports broadcasts. After each frame of video is processed in parallel by the AI ​​algorithm, the matting result (i.e., the segmentation result) is saved. The frame timestamps of different modules are strictly aligned and correspond, ensuring that the layers of different frames are accurately aligned.

[0034] Step S300: Based on the attention mechanism, the segmentation results from different angles at the same time stamp are aligned and fused in terms of spatiotemporal features, and the image sequence with 3D multi-dimensional perspective effect is generated by perspective interpolation.

[0035] In an optional embodiment, step S300 includes the following steps S301 to S302.

[0036] Step S301: Use Transformer or cross-attention mechanism to determine the associated feature data of the same moving target under different viewpoints and timestamps.

[0037] Here, Transformer is a deep learning model architecture built on a self-attention mechanism. Its core design abandons traditional recurrent or convolutional structures, relying entirely on modeling the global dependencies between all elements in the input sequence. Through "query-key-value" computation, it allows any position in the sequence to directly focus on and fuse information from all other positions, making it particularly adept at handling long-distance dependencies. It has become the mainstream infrastructure in fields such as natural language processing and computer vision.

[0038] Cross-attention is an important variant of attention mechanisms, specifically designed to handle relationships between two different sequences or feature sets. One sequence provides the "query," and the other provides the "key" and "value." Through computation, this mechanism allows each element in the "query" sequence to focus on the most relevant part of the "key-value" sequence for information fusion. In this embodiment, cross-attention is used to align and match features from different perspectives and at different times, achieving multi-perspective spatiotemporal information fusion.

[0039] Step S200 performs pixel-level segmentation of the target area to generate a high-precision binary mask, achieving "masking" of the foreground target. Step S301 follows the processing order of the video stream, sequentially fusing the matting results of the same timestamp in the video to the historical background layer from different perspectives while strictly maintaining the relative position of the moving target, until all matting is processed, forming a video stream with dynamic freeze-frame and rotating perspective effects of the same scene from different angles.

[0040] Step S302: Perform viewpoint interpolation, texture fusion, and background processing based on the associated feature data to obtain an image sequence with a 3D multi-dimensional viewpoint effect.

[0041] The perspective interpolation here utilizes cutting-edge AI technology, employing optical flow and deep learning techniques to calculate the target's appearance from a virtual perspective and dynamically adjust its rotation speed smoothly. First, optical flow is used to estimate the motion field (pixel-level displacement) between images. This process is accomplished by calculating the optical flow field between adjacent frames. Optical flow calculates the motion vector for each pixel, describing the object's movement between two perspectives or two points in time. Then, based on the optical flow field and the input image, a deep learning model generates images from intermediate frames or new perspectives. These models not only consider optical flow information but also leverage contextual information (such as scene depth and texture) to improve interpolation quality. Finally, an interpolation algorithm (such as weighted averaging based on optical flow or deep learning-generated interpolation methods) is used to generate new perspective images between two known perspectives. Throughout this process, depth and optical flow information are comprehensively considered to ensure the generated images are coherent both dynamically and spatially.

[0042] In optional embodiments, the latest AI technologies can be embedded, such as the DAIN model and Super SloMo technology. DAIN (Depth-Aware Video Frame Interpolation) is a deep learning-based video interpolation model specifically designed to generate smooth intermediate frames. This model not only relies on traditional optical flow estimation but also incorporates depth information to optimize the interpolation process, thereby improving image quality, especially in complex scenes and object motion. Super SloMo is a deep neural network designed to generate high-quality slow-motion video from low frame rate videos. By utilizing optical flow and convolutional neural networks, this model is able to generate smooth frame interpolation, suitable for improving video smoothness.

[0043] Texture fusion refers to the merging of texture information from multiple real-world perspectives to generate a complete and realistic target image. Background processing refers to applying dynamic blurring, staticization, or other techniques to the background to make it consistent with the rotating foreground subject. Texture fusion reconstructs the complete appearance of a target from a virtual perspective by integrating texture information from multiple real-world perspectives. Its core is based on multi-view image registration and weighted fusion techniques, utilizing the complementarity of textures between perspectives to eliminate occlusion and distortion, and achieving seamless texture transitions through deep learning or optimization algorithms (such as Poisson fusion), ultimately generating a high-quality synthetic image with consistent details and no artifacts.

[0044] Background processing separates the background from the rotating subject in the foreground using motion blur and static techniques. It applies motion blur to the background area to simulate camera tracking, or staticizes it to enhance the sense of time frozen, thereby reducing the visual conflict between the moving background and foreground, highlighting the subject, and enhancing the immersiveness and dynamic expression of the image.

[0045] Step S400: Encode the image sequence to obtain a 3D multi-dimensional viewpoint video stream.

[0046] Specifically, the cutout results are cached in a queue of timestamp-layer pairs. In this embodiment, cutouts with the same timestamp in the video are sequentially processed, keeping the relative position of the cutout targets strictly unchanged, and iteratively merged into historical background layers from different perspectives until all cutouts are processed, forming a dynamic freeze-frame rotating viewpoint video stream of the same scene from different angles. After dynamic blurring optimization using AI frame technology, a 3D multi-dimensional viewpoint video stream is output through the audio and video processing module.

[0047] This embodiment assembles a series of rendered image frames in a time sequence and compresses and encodes them using video coding standards (such as H.264 / H.265) to output the final video stream.

[0048] This embodiment describes a method for processing multiple video streams and outputting a 3D multi-view video stream based on artificial intelligence algorithms, involving fields such as artificial intelligence, computer vision, and video processing. This embodiment organically combines object detection, object tracking, object segmentation, real-time matting, and multi-view frame interpolation fusion technology, and is widely used in sports live streaming and event analysis. Its advantage lies in its ability to generate video streams with highly immersive dynamic freeze-frame and rotating perspective effects in real time, significantly improving the efficiency and expressiveness of video generation.

[0049] Compared with the prior art, the present invention has the following significant advantages: This embodiment is highly flexible and can achieve the same effect using only a few video streams; This embodiment boasts high efficiency and real-time performance. Through optimized algorithms and hardware acceleration (GPU), the entire process achieves near real-time processing with extremely low latency, enabling applications such as live streaming. Here, hardware acceleration utilizes the massively parallel architecture of GPUs to decompose computationally intensive tasks (such as neural network inference, image rendering, and matrix operations) into thousands of micro-tasks that are executed concurrently. It executes highly unified computational instructions through dedicated stream processors and utilizes high-bandwidth video memory to achieve rapid data throughput, thereby achieving throughput efficiency tens of times that of CPUs and extremely low latency when processing pixel-level calculations in video. This embodiment is highly automated, with the entire process requiring little or no human intervention and being completed automatically by AI algorithms, greatly improving production efficiency; This embodiment produces a lifelike effect. By combining the latest AI optimization algorithms and video frame interpolation techniques, the generated dynamic rotating perspective video images are rich in detail and have smooth transitions, with visual effects comparable to high-end hardware solutions.

[0050] Preferably, the target detection module, target tracking module, and segmentation module share a backbone feature extraction network. The backbone feature extraction network is used to extract features from multiple video streams. The target detection module, target tracking module, and segmentation module are implemented through task head networks connected after the backbone feature extraction network. The general feature map extracted by this shared backbone network is provided to the three task heads for detection, tracking, and segmentation.

[0051] In specific implementation, such as Figure 2 The input video frames (H×W×3) are used, and a ResNet-50 / 101 backbone is employed as the main feature extraction network. The network consists of an initial convolutional layer (conv1) and multiple residual stages (res2 to res5). Through residual connections and batch normalization (BN) layers, the vanishing gradient problem in deep networks is effectively mitigated, gradually generating feature maps with rich semantic information. The spatial size of the output feature maps decreases sequentially at each stage, while the number of channels increases from 64 to 2048, achieving multi-level feature encoding of the input image.

[0052] exist Figure 2 In this embodiment, the backbone feature extraction network is a lightweight network, and the method further lightweights the backbone feature extraction network and the task head network.

[0053] In an optional embodiment, the object detection module includes a feature pyramid submodule and an attention mechanism submodule.

[0054] like Figure 2The Feature Pyramid (FPN) module is introduced in this architecture to fully utilize the multi-scale features extracted by the backbone network. This module fuses deep, high-semantic features with shallow, high-resolution features through a top-down path and lateral connections, generating a set of multi-scale feature maps (P2 to P5) with a uniform number of channels (e.g., 256). Furthermore, the scale range of the feature pyramid can be further expanded through additional convolutional layers (e.g., P6 and P7) to better detect targets of different sizes, especially distant or small targets.

[0055] Following the FPN module, a squeeze-encouragement attention mechanism module (SE Block) is integrated. This module first compresses each feature channel using global average pooling (GAP), then learns the importance weights of each channel through two fully connected layers (containing ReLU and sigmoid activation functions), and finally multiplies these weights with the original feature map channel by channel. This process achieves channel-level feature recalibration, enabling the network to focus more on feature channels relevant to the target, thereby improving detection accuracy and robustness.

[0056] The enhanced multi-scale feature maps are fed into the detection head. The detection head in this embodiment includes two parallel branches: one for object classification, which outputs the probability of each anchor box belonging to each category through convolutional layers (shape K×A, where K is the number of categories and A is the number of anchor boxes); the other for bounding box regression, which outputs fine-grained positional adjustment parameters for each anchor box (such as center point coordinates and width). The two branches are combined to finally output the category labels and precise bounding box coordinates for all detected objects in the image.

[0057] This embodiment's detection architecture integrates deep residual networks, multi-scale feature pyramids, and channel attention mechanisms, enabling efficient and accurate localization and recognition of moving targets (such as athletes) in video frames, providing a reliable foundation for subsequent tracking and segmentation tasks.

[0058] In an optional embodiment, the target tracking module includes an active trajectory library submodule, a Kalman prediction submodule, a data association submodule, and a trajectory recovery submodule; The active trajectory library submodule is used to store the state information of moving targets; The Kalman prediction submodule is used to predict the expected position information of a moving target in the current frame based on the state of the moving target in the previous time step. The data association submodule is used to perform target matching based on the target detection results and expected location information of the current frame, and by using Mahalanobis distance and intersection-union ratio to obtain the matching result; among which, target matching is achieved by solving the optimal association using the Hungarian algorithm; The trajectory recovery submodule is used to reassociate and recover the interrupted trajectory based on the matching results, so as to obtain the continuous motion trajectory of the moving target.

[0059] Specifically, the target tracking module connects with the detection module to perform pixel-level segmentation of the target region, generating a high-precision binary mask to achieve "masking" of the foreground target.

[0060] like Figure 3 As shown, this embodiment uses the OC-SORT (Observation-CentricSORT) multi-target tracking algorithm. First, the original video frames are input, and each frame first passes through a target detector (such as...). Figure 2 The network shown obtains the detection results for the current frame, with each result including bounding boxes, category labels, and confidence scores. Simultaneously, the server maintains an active trajectory library submodule, storing the state information of each existing trajectory, including its position, velocity, timestamp, and consecutive loss count. For each active trajectory in the library, the Kalman filter in the Kalman prediction submodule predicts its expected position (predicted box) in the current frame based on its historical motion state, and performs motion consistency verification.

[0061] The data association submodule first calculates the cost matrix, primarily based on the Mahalanobis distance and intersection-union ratio (IOU) between the predicted bounding box and the current frame's detection bounding box. Impossible association pairs can be filtered out by setting a gating threshold. Subsequently, the Hungarian algorithm is used to find the optimal match, assigning the detection bounding boxes to existing trajectories.

[0062] If a match is successful, the Kalman trajectories of the trajectory are updated using the corresponding detection results, the missing count is reset, and the updated trajectory is output. If a match fails, the missing count of the trajectory is incremented and sent to the trajectory recovery module. This module will re-match the trajectory and use a more lenient matching strategy (such as using matching based solely on appearance features or IOU) in subsequent frames for short-term re-association to effectively handle situations where the target is lost due to brief occlusion or rapid movement.

[0063] The target tracking module in this embodiment combines the Kalman filter algorithm with extracted target detection boxes and motion information. By maintaining the continuous state of the target, this module can handle rapid motion and attitude changes, thereby temporally associating key action frames and providing stable tracking data for subsequent extraction of key actions.

[0064] This embodiment also includes lifecycle management. For new detection boxes that do not match any existing trajectories, they are initialized as new trajectories after certain conditions are met (such as being detected in multiple consecutive frames, i.e., satisfying min_hits). For old trajectories whose loss count exceeds the preset maximum lifespan (max_age), they are terminated and removed from the active trajectory library. Finally, the process outputs the tracking results [target ID, bounding box] for each frame, thereby generating continuous and stable motion trajectories for each target.

[0065] This embodiment significantly improves the ability to maintain the consistency of target identity in complex scenarios (such as target intersection and occlusion), providing stable and reliable target motion trajectory data for subsequent multi-view video fusion.

[0066] This embodiment proposes a spatiotemporal feature alignment and fusion algorithm for multiple video streams based on an attention mechanism. Specifically, it proposes a neural network matting model based on Transformer or Cross-Attention mechanisms to handle frame-by-frame target detection, tracking, and segmentation from multiple video streams at different angles. This model can automatically learn and associate features of the same target from different viewpoints and timestamps, achieving cross-viewpoint spatiotemporal semantic alignment without relying on traditional 3D geometric calculations. This provides highly consistent and rich feature representations for subsequent dynamic rotation viewpoint synthesis, ensuring the smoothness and consistency of the generated video. This is the key algorithmic foundation for achieving high-quality fusion results.

[0067] This embodiment is based on a multi-view target consistency tracking algorithm using motion trajectory prediction, specifically a joint learning framework for multi-modal motion trajectory prediction. This algorithm not only tracks within a single video stream, but more importantly, it can associate targets across video streams with different viewpoints. This ensures that even if a target is severely occluded or temporarily disappears from one viewpoint, the system can maintain the target's ID consistency using information from other viewpoints. This embodiment solves the challenge of multi-view target tracking in complex scenes, providing stable and continuous target labels for accurate frame-by-frame segmentation and subsequent viewpoint fusion, thus avoiding the problem of subject identity jumps in the final video.

[0068] This embodiment also innovates the segmentation algorithm, introducing a high-precision segmentation model for temporal consistency in video sequences. Specifically, it builds upon existing instance segmentation models (such as Mask2Former) by introducing a temporal constraint loss function and optical flow guidance. During training, the model not only learns accurate segmentation of single frames but also learns to ensure edge smoothness and temporal consistency of the segmentation masks between adjacent frames.

[0069] Specifically, this scheme, building upon static image segmentation models such as Mask2Former, introduces temporal consistency by constructing a multi-frame input encoder-decoder architecture. Specifically, during training and inference, the model no longer processes only a single frame, but instead takes a sequence of consecutive frames within a short time window as input. The encoder part extracts and aggregates spatiotemporal features across frames by introducing lightweight temporal fusion modules, such as 3D convolutions or temporal self-attention layers. This allows the model to simultaneously perceive the spatial details of the current frame and the motion information between adjacent frames, enabling it to learn and output segmentation results that are stable in the temporal dimension.

[0070] The loss function used to train this model consists of two parts. The first part is the basic segmentation loss inherited from the original model, such as the binary cross-entropy loss of the mask and the Dice loss, which is mainly responsible for ensuring the pixel-level accuracy of the segmentation result in each frame. The second part is the temporal consistency loss, which is a composite loss that includes mask alignment loss and edge smoothing loss. The alignment loss uses optical flow information to constrain the segmentation result of the current frame to be consistent with the result of the previous frame obtained based on motion distortion; the smoothing loss directly minimizes the change amplitude of the segmentation boundary between adjacent frames, thereby effectively suppressing edge flickering and jitter.

[0071] In this embodiment, optical flow guidance primarily serves as a powerful temporal supervision signal. During the training phase, a pre-trained optical flow estimation network computes a dense pixel-level motion vector field between adjacent frames. This optical flow field is then used for two key purposes: first, to construct the aforementioned mask alignment loss, providing the model with explicit motion consistency constraints; second, the optical flow field itself can serve as an additional set of input feature channels, fed into the segmentation network along with the original RGB image, providing the model with explicit motion cues and helping it better distinguish moving subjects from static backgrounds.

[0072] When training the entire network model, it is first trained on a dataset of static scenes (such as an athlete's ready pose) to allow the model to learn basic 3D geometry and appearance. Then, datasets with larger dynamic ranges are gradually introduced, and finally, fine-tuning is performed on a complete dataset containing high-speed motion. This improves the stability of training and the final result.

[0073] In practical applications, a three-frame sliding window centered on the current frame and extending one frame before and after it is used as the standard input. For optical flow computation, models like RAFT, which offer a good balance between accuracy and efficiency, are preferred. To meet the stringent requirement of processing time less than 15 milliseconds per frame in live streaming scenarios, all temporal fusion modules have been carefully designed to ensure lightweight operation. Supervised fine-tuning on large video datasets ensures that the model ultimately possesses both high-precision segmentation capabilities and excellent temporal stability.

[0074] The target mask generated in this embodiment is not only highly accurate, but also continuous and stable in time, effectively eliminating the flickering and jitter that may be caused by frame-by-frame segmentation, resulting in an extremely smooth and natural final rotating video.

[0075] This embodiment also features architectural innovation, employing a real-time multi-task joint learning approach and lightweight model design. Specifically, it utilizes a multi-task learning framework, designing a shared backbone network to simultaneously perform preliminary feature extraction for object detection, tracking, and segmentation. Then, specific computations are performed using different task heads, significantly reducing the overall computational load.

[0076] This embodiment combines lightweight techniques such as model pruning and quantization, enabling a series of complex AI algorithms to be integrated on terminal devices, meeting the stringent requirements for extremely low latency in scenarios such as live streaming, and achieving the best balance between algorithm performance and efficiency.

[0077] This embodiment employs a multi-task learning architecture with a single forward propagation, based on ResNet as the shared backbone network. Unlike traditional serial architectures, ResNet can preserve multi-scale features throughout the entire network, which is crucial for detection, tracking, and segmentation tasks that require precise localization and detail preservation.

[0078] This embodiment first acquires multiple video streams and uses them as input sources; then, it performs real-time target detection and cross-frame target tracking on the video streams to obtain the trajectory of the moving target; it performs pixel-level segmentation on the target region and generates a pixel-level binary mask; then, it fuses the multi-angle matting results and generates a 3D multi-dimensional viewpoint effect video stream through viewpoint interpolation, texture fusion and background processing; finally, it encodes the generated image frame sequence and outputs it as a video stream.

[0079] In this embodiment, the target detection and target tracking modules employ lightweight deep neural networks for real-time detection and utilize multi-target tracking algorithms to associate targets across frames. A cross-view consistency tracking algorithm is also employed to ensure target consistency across multiple viewpoints. The segmentation module uses an attention-based neural network model to achieve cross-view spatiotemporal feature alignment and segmentation mask generation. Optical flow and deep learning techniques are used for viewpoint interpolation during dynamic rotation viewpoint video generation. This embodiment introduces temporal constraints and optical flow guidance to ensure the temporal continuity of the segmentation results. In the fusion and generation steps, deep learning techniques are used to synthesize virtual viewpoints, and the background is dynamically blurred or statically processed.

[0080] Improvements and innovations in this solution: 1. Multi-tasking adapted architecture extension (1) Task-specific FPN design: After sharing the backbone, a dedicated feature pyramid network is provided for different tasks; (2) Dynamic feature routing: Adaptively allocate computing resources to branches of different resolutions according to task requirements; (3) Gradient balancing mechanism: to prevent a certain task from dominating feature learning in multi-task training.

[0081] 2. Enhanced cross-resolution fusion (1) Bidirectional dense connection: Add reverse information flow on the basis of the original horizontal connection; (2) SE attention weighting: Channel-level importance recalibration of cross-resolution features; (3) Multi-scale context aggregation: Introducing void space pyramid pooling to enhance the receptive field.

[0082] 3. Time-aware feature extraction (1) Inter-frame feature propagation: Temporal information is transmitted between adjacent frames through 3D convolution; (2) Motion consistency constraint: Introduce a temporal smoothness regularization term during training; (3) Dynamic background modeling: Differentiated feature processing strategies are adopted for static background and moving foreground.

[0083] 4. Real-time performance optimization and improvement (1) Selective feature calculation: dynamically adjust the calculation intensity of each branch according to the complexity of the scene; (2) Progressive resolution processing: Use low-resolution branches for simple regions and high-resolution branches for complex regions; (3) Memory-efficient design: Reduce memory usage through feature sharing and early downsampling.

[0084] This embodiment significantly improves the smoothness and realism of the generated video. Thanks to the temporal consistency segmentation model and multi-view feature fusion algorithm, the generated target subject exhibits precise edges and no flickering or jitter during rotation, while maintaining highly intact texture details. The transitions in the virtual viewpoint are natural and smooth, with visual effects comparable to movie-level special effects, greatly enhancing the visual expressiveness and professionalism of the final work.

[0085] This embodiment pioneers a new paradigm for real-time interactive video creation. Its extremely low processing latency enables its application in live broadcasts and sports broadcasts. Directors or users can select highlights and generate rotating replays within a very short time after an event occurs, providing viewers with an unprecedented immersive viewing experience and expanding the application boundaries of video technology.

[0086] This embodiment maintains stable output quality in complex scenes. The multi-view consistency tracking algorithm effectively addresses challenges such as occlusion, lighting changes, and motion blur, ensuring the continuous stability of the target identity. The attention mechanism's feature alignment guarantees that even with significant differences in camera viewpoints, the system can find the correct correspondence and perform effective fusion, improving the system's practicality and reliability in real-world complex environments.

[0087] This embodiment performs real-time target detection, cross-viewpoint tracking, and high-precision segmentation on multi-angle videos of the same scene, and integrates viewpoint interpolation technology to automatically generate a smooth, dynamically rotating viewpoint video stream. This embodiment can significantly reduce hardware costs, achieve efficient real-time video processing, and is widely used in fields such as live sports broadcasts.

[0088] This embodiment uses multiple video streams as input sources for object detection and cross-view tracking. It generates pixel-level masks using a temporal consistency segmentation model, fuses the multi-view segmentation results, and synthesizes a 3D multi-dimensional view video stream. Finally, it encodes and outputs the video. This embodiment employs an attention mechanism for feature alignment and combines deep learning algorithms to automatically select keyframes and rotation paths, offering advantages such as high efficiency, real-time performance, and a high degree of automation. The generated video is smooth and realistic.

[0089] This embodiment utilizes a lightweight neural network to achieve real-time detection, tracking, and segmentation of targets in multi-view videos, and synthesizes videos with rotating perspective effects through video frame interpolation and background processing techniques. The system supports end-to-end processing and features low latency, high consistency, and strong robustness. It can be applied to scenarios such as live streaming, sports replay, and content creation, effectively improving the visual appeal and production efficiency of videos.

[0090] This embodiment integrates target detection, tracking and segmentation modules through a multi-task joint learning framework, and introduces cross-view attention mechanism and reinforcement learning decision-making to achieve fully automatic generation of 3D multi-dimensional view effect video streams from multiple input videos.

[0091] See Figure 4 This invention provides a multi-video 3D perspective generation device, comprising an acquisition module 100, a processing module 200, a generation module 300, and an encoding module 400. The acquisition module 100 acquires multiple video streams from the same scene, at different angles, and time-aligned. The processing module 200 processes the multiple video streams to obtain the motion trajectory of a moving target and the segmentation results of the target region. The generation module 300 performs spatiotemporal feature alignment and fusion of the segmentation results from different angles at the same time stamp based on an attention mechanism, and uses perspective interpolation to generate an image sequence with a 3D multi-perspective effect. The encoding module 400 encodes the image sequence to obtain a video stream with a 3D multi-perspective effect.

[0092] In an optional embodiment, the processing module 200 includes: The first processing module is used to perform target detection on moving targets in each video stream using a pre-established target detection module, and obtain the target detection results. The second processing module is used to track the moving target based on the target detection results and the pre-established target tracking module, and obtain the motion trajectory of the moving target. The third processing module is used to segment the target region based on the pre-established segmentation module and the target detection results to obtain the segmentation results.

[0093] In an optional embodiment, the target detection module, the target tracking module, and the segmentation module share a backbone feature extraction network. The backbone feature extraction network is used to extract features from multiple video streams. The target detection module, the target tracking module, and the segmentation module are implemented through a task head network connected after the backbone feature extraction network.

[0094] In optional embodiments, the backbone feature extraction network is a lightweight network, and / or the method further includes lightweighting the backbone feature extraction network and the task head network.

[0095] In an optional embodiment, the generation module 300 includes: The association module is used to determine the associated feature data of the same moving target under different viewpoints and timestamps using Transformer or cross-attention mechanism; The image sequence module is used to perform viewpoint interpolation, texture fusion, and background processing based on associated feature data to obtain image sequences with 3D multi-dimensional viewpoint effects.

[0096] In an optional embodiment, the target tracking module includes an active trajectory library submodule, a Kalman prediction submodule, a data association submodule, and a trajectory recovery submodule; The active trajectory library submodule is used to store the state information of moving targets; The Kalman prediction submodule is used to predict the expected position information of a moving target in the current frame based on the state of the moving target in the previous time step. The data association submodule is used to perform target matching based on the target detection results and expected location information of the current frame, and by using Mahalanobis distance and intersection-union ratio to obtain the matching result; among which, target matching is achieved by solving the optimal association using the Hungarian algorithm; The trajectory recovery submodule is used to reassociate and recover the interrupted trajectory based on the matching results, so as to obtain the continuous motion trajectory of the moving target.

[0097] In an optional embodiment, the segmentation module incorporates a temporal constraint loss function and optical flow guidance during training.

[0098] In an optional embodiment, the object detection module includes a feature pyramid submodule and an attention mechanism submodule.

[0099] The apparatus provided in the embodiments of this application has the same inventive concept as the method provided in the embodiments of this application. As long as the method can solve the technical problem, the apparatus can also solve the technical problem. This will not be elaborated here.

[0100] Reference Figure 5 The present invention also provides an electronic device 1000, including a communication interface 1001, a processor 1002, a memory 1003, and a bus 1004. The processor 1002, the communication interface 1001, and the memory 1003 are connected via the bus 1004. The memory 1003 is used to store a computer program that supports the processor 1002 in executing the multi-video 3D perspective generation method. The processor 1002 is configured to execute the program stored in the memory 1003.

[0101] Optionally, embodiments of the present invention also provide a computer-readable medium having non-volatile program code executable by a processor 1002, the program code causing the processor 1002 to perform the multi-video 3D perspective generation method as described in the above embodiments.

[0102] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0103] As is known from common technical knowledge, this invention can be implemented through other embodiments that do not depart from its spirit or essential characteristics. Therefore, the disclosed embodiments described above are merely illustrative in all respects and are not the only ones. All modifications within the scope of this invention or its equivalents are included in this invention.

Claims

1. A method for generating multi-video 3D perspectives, characterized in that, include: Acquire multiple video streams from the same scene but different angles and with time alignment; The multiple video streams are processed to obtain the motion trajectory of the moving target and the segmentation result of the target region; Based on the attention mechanism, the segmentation results from different angles at the same time stamp are aligned and fused in terms of spatiotemporal features, and the image sequence with 3D multi-dimensional view effect is generated by perspective interpolation. The image sequence is encoded to obtain a 3D multi-dimensional viewpoint video stream.

2. The multi-video 3D perspective generation method according to claim 1, characterized in that, The process of processing the multiple video streams to obtain the motion trajectory of the moving target and the segmentation result of the target region includes: The motion targets in each video stream are detected using a pre-established target detection module to obtain the target detection results. Based on the target detection results and the pre-established target tracking module, the moving target is tracked to obtain the motion trajectory of the moving target; The target region is segmented based on the pre-established segmentation module and the target detection results to obtain the segmentation results.

3. The multi-video 3D perspective generation method according to claim 2, characterized in that, The target detection module, the target tracking module, and the segmentation module share a backbone feature extraction network. The backbone feature extraction network is used to extract features from the multiple video streams. The target detection module, the target tracking module, and the segmentation module are implemented through a task header network connected after the backbone feature extraction network.

4. The multi-video 3D perspective generation method according to claim 3, characterized in that, The backbone feature extraction network is a lightweight network, and / or the method further includes lightweighting the backbone feature extraction network and the task head network.

5. The multi-video 3D perspective generation method according to claim 1, characterized in that, The process of aligning and fusing spatiotemporal features of segmentation results from different angles at the same time stamp based on an attention mechanism, and generating an image sequence with 3D multi-dimensional viewpoint effects using viewpoint interpolation, includes: The Transformer or cross-attention mechanism is used to determine the associated feature data of the same moving target under different viewpoints and timestamps; Based on the associated feature data, viewpoint interpolation, texture fusion, and background processing are performed to obtain an image sequence with a 3D multi-dimensional viewpoint effect.

6. The multi-video 3D perspective generation method according to claim 2 or 3, characterized in that, The target tracking module includes an active trajectory library submodule, a Kalman prediction submodule, a data association submodule, and a trajectory recovery submodule. The active trajectory library submodule is used to store the state information of moving targets; The Kalman prediction submodule is used to predict the expected position information of the moving target in the current frame based on the state of the moving target in the previous moment. The data association submodule is used to perform target matching based on the target detection results of the current frame and the expected location information, and using Mahalanobis distance and intersection-union ratio to obtain the matching result; wherein, the target matching is achieved by solving the optimal association using the Hungarian algorithm; The trajectory recovery submodule is used to reassociate and recover the interrupted trajectory based on the matching result, so as to obtain the continuous motion trajectory of the moving target.

7. The multi-video 3D perspective generation method according to claim 2 or 3, characterized in that, The segmentation module incorporates a temporal constraint loss function and optical flow guidance during training.

8. The multi-video 3D perspective generation method according to claim 2 or 3, characterized in that, The target detection module includes a feature pyramid submodule and an attention mechanism submodule.

9. A multi-video 3D perspective generation device, characterized in that, include: The acquisition module is used to acquire multiple video streams from the same scene but different angles and with time alignment. The processing module is used to process the multiple video streams to obtain the motion trajectory of the moving target and the segmentation result of the target region; The generation module is used to perform spatiotemporal feature alignment and fusion of segmentation results from different angles at the same time stamp based on an attention mechanism, and to generate image sequences with 3D multi-dimensional view effects using viewpoint interpolation. The encoding module is used to encode the image sequence to obtain a 3D multi-dimensional viewpoint video stream.

10. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method according to any one of claims 1-8.