Multi-camera video stream intelligent editing and real-time special effect synthesis method
Through the intelligent video processing system, the space-time synchronization and model optimization of multi-camera video streams are solved, and the problems of low manual editing efficiency and unnatural implantation of virtual scenes in multi-camera video processing are improved, improving the viewing and visual effects of the video.
Patent Information
- Application Number
- CN202510582378.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-25
AI Technical Summary
In the existing multi-camera video processing, manual editing efficiency is low, natural virtual scene implantation and real-time special effects synthesis are difficult to meet the diverse and personalized needs. The traditional methods lack intelligence, resulting in unsmooth lens switching and unnaturally implanting virtual scenes.
Through the intelligent video processing system, multi-modal information synchronization algorithm is used to ensure the time and space synchronization of videos of various cameras, and a multi-lens switching and virtual scene implantation model is built, combining pre-training weights and machine learning optimization algorithms to realize intelligent lens switching and real-time special effects synthesis.
It realizes efficient and intelligent lens switching and natural virtual scene implantation, improving the viewing and visual effects of videos, and meeting diverse creative needs.
Smart Images

Figure CN120378555A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and particularly to a method for intelligent editing and real-time special effect synthesis of multi-camera video streams. Background Art
[0002] In the current era of booming digital media, multi-camera video shooting has been widely used in many fields, such as film and television production, large-scale live events (including live sports events, live art performances, etc.), video conferencing, and online education course recording. By shooting the same event or scene from different angles and different shot sizes with multiple cameras, richer and more comprehensive video materials can be obtained, providing the possibility for presenting high-quality visual content.
[0003] However, with the continuous increase in the amount of multi-camera video data and the increasing requirements for video quality and visual effects, traditional video editing and special effect synthesis methods face many challenges. On the one hand, manually editing multi-camera videos often requires a large amount of human and time costs. Editors need to repeatedly watch the materials of different cameras and judge the switching timing of the shots and select appropriate shot sizes based on experience. This process is not only inefficient but also prone to problems such as unsmooth shot switching and failure to accurately capture key pictures due to human negligence or subjective factors, affecting the final presentation effect of the video. On the other hand, in terms of special effect synthesis, especially real-time special effect synthesis, traditional methods mostly rely on fixed templates preset in professional software and manual operations, lacking intelligence, and it is difficult to automatically match appropriate virtual scenes according to the video content and add context-compliant special effects in real time, unable to meet the current needs of diversification, personalization, and fast video production. Summary of the Invention
[0004] To make up for the above deficiencies, the present invention provides a method for intelligent editing and real-time special effect synthesis of multi-camera video streams, aiming to improve the problems of low efficiency in manual editing and the implantation of natural virtual scenes and real-time special effect synthesis in existing multi-camera video processing.
[0005] In the first aspect, the present invention provides the following technical solution, a method for intelligent editing and real-time special effect synthesis of multi-camera video streams, and the intelligent video processing software system includes: S1. Video acquisition and synchronization: The intelligent video processing system connects to the shooting devices through various interfaces to acquire multi-camera video streams, and uses a multi-modal information synchronization algorithm to ensure the accurate synchronization of the videos of each camera in space and time; S2. Model pre-training and optimization: Construct a multi-shot switching and virtual scene implantation model, initialize it with pre-trained weights, and adjust the parameters through self-supervised learning, fine-tuning, and optimization algorithms to improve the performance of the model; S3, Feature extraction and analysis: Extract visual, motion, and semantic features related to multi-lens switching and virtual scene implantation from synchronized video streams to comprehensively analyze video content; S4, lens switching decision: Combine preset rules with machine learning methods to determine the lens switching requirements, determine the best switching time and position, and realize intelligent lens switching; S5, scene embedding and fusion: select and adjust virtual scene elements according to video features, and use fusion technology to naturally integrate them into the video frame to enrich the video picture; S6. Real-time special effects synthesis: Add color, blur, and particle real-time special effects to the processed video frames to enhance the video visual effects and complete the overall processing flow.
[0006] By adopting the above technical solutions, precise synchronization lays the foundation for subsequent links, model performance improvement ensures processing accuracy, multi-dimensional feature extraction helps accurate decision-making, intelligent lens switching makes the video smoother and more natural, the scene is naturally integrated into the rich picture, and real-time special effects synthesis enhances the visual impact, ultimately outputting high-quality and highly ornamental video content to meet diverse creative needs.
[0007] Preferably, the video acquisition and synchronization includes: Hardware interface, used to connect a variety of shooting devices to realize video stream acquisition, covering Ethernet, SDI, HDMI, Wi-Fi 6, Bluetooth 5.2 types, connecting different devices; The multimodal information synchronization algorithm is used to synchronize the video streams collected by each camera using timestamps, audio features, and visual features. It first performs preliminary synchronization with timestamps, and then uses audio fingerprints, DTW algorithm, deep learning, and homography matrix calculation to accurately align the audio and video levels to ensure the temporal and spatial consistency of the video frames.
[0008] Preferably, the model pre-training optimization includes: The multi-lens automatic switching model is used to analyze the characteristic information in the multi-camera video stream, and based on the extracted spatial and temporal characteristics, through preset rules and machine learning strategies, intelligently judge and decide the best time and corresponding camera position for lens switching, and automatically switch the video lens; The virtual scene embedding model is used to generate virtual scene elements that are adapted to the video content based on the video frame features with the help of a generative adversarial network, and accurately locate the fusion area with the help of a semantic segmentation network, so as to naturally integrate the virtual scene elements into the video frame. A general strategy for model pre-training and optimization is used to use pre-trained weights and self-supervised learning for initial training, and then use labeled data combined with multiple loss functions and optimization algorithms to adjust parameters. Regularization and Dropout are also used to prevent overfitting and dynamically adjust hyperparameters.
[0009] Preferably, the feature extraction analysis includes: Feature extraction related to multi - camera automatic switching, which is used to mine visual, motion, and semantic features from multi - camera video streams, provide a basis for multi - camera automatic switching decisions, and assist in judging when to switch the camera and which camera to switch to; Feature analysis related to virtual scene implantation, which is used to analyze the scene, lighting, and style features of video frames, provide a reference for accurately selecting, adapting, and integrating virtual scene elements, and help natural integration of virtual scenes into videos.
[0010] Preferably, the above - mentioned camera switching decisions include: Rule - based decision - making method, which is used to make decisions on multi - camera video lens switching according to the rules set in terms of the type, rhythm, and composition of video content, ensuring appropriate lens switching timing and reasonable picture presentation; Machine - learning - based decision - making method, which is used to mine lens switching rules by learning multi - camera video data and intelligently determine the lens switching timing and camera position according to the output of classification or reinforcement learning models.
[0011] Preferably, the above - mentioned scene implantation and fusion include: Feature matching and virtual scene selection, which is used to screen out suitable virtual scene elements from the virtual scene material library according to video frame features and make corresponding adjustments to prepare for subsequent natural integration into video frames; Application of image fusion technology, which is used to seamlessly fuse the selected virtual scene elements with video frames, eliminate fusion traces, make virtual scenes naturally integrate into video pictures, and enhance the overall vision of videos.
[0012] Preferably, the above - mentioned real - time special effect synthesis includes: Color adjustment special effects, which are used to change the color attributes of video frames, enhance contrast, correct color deviation, or convert color styles; Blur special effects, which are used to blur video frames, create a specific atmosphere, highlight the main body, reduce noise, or create a hazy feeling; Particle special effects, which are used to simulate diverse natural or fantasy visual effects such as snowflakes, raindrops, and fireworks in videos, enhancing the visual impact and interest of videos.
[0013] In a second aspect, the present invention provides the following technical solution: an intelligent video processing system for the intelligent editing of multi - camera video streams and real - time special effect synthesis method, which is used for the intelligent editing of multi - camera video streams and real - time special effect synthesis method. The system includes: Video stream acquisition and synchronization unit, which is configured with multiple interfaces for connecting shooting devices of different cameras and has a built - in synchronization algorithm module to achieve high - precision synchronization of video streams; Deep - learning model unit, which stores and runs pre - trained deep - learning models related to multi - lens switching and virtual scene implantation and has a model optimization and update function; A feature extraction and analysis unit that can extract multi-dimensional features from the input video stream, perform in-depth analysis, and provide data support for subsequent processing; A shot transition decision unit that comprehensively uses preset rules and machine learning algorithms to output accurate multi-shot automatic transition instructions; A virtual scene implantation and fusion unit that matches elements according to features and uses various fusion technologies to integrate the virtual scene into the video frame; A real-time special effect synthesis unit that adds various real-time special effects to the video as needed and has the ability to adaptively adjust special effect parameters; An output management unit that is responsible for outputting the processed video according to the set format and resolution parameters to meet the requirements of different application scenarios.
[0014] Thirdly, the present invention provides the following technical solution. A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the intelligent multi-camera video stream editing and real-time special effect synthesis method as described above is implemented.
[0015] Fourthly, the present invention provides the following technical solution. A readable storage medium stores a computer program, and when the computer program is executed by a processor, the intelligent multi-camera video stream editing and real-time special effect synthesis method as described above is implemented.
[0016] The present invention has the following beneficial effects: 1. In the present invention, through the intelligent and automatic processing of multi-camera video streams, high-precision spatio-temporal synchronization of video streams of each camera is achieved. The constructed and optimized model helps to accurately extract features, achieve intelligent and smooth shot transitions, and can screen and integrate virtual scene elements according to features, improving the visual appeal of the picture. It solves the problems of low efficiency of manual editing in existing multi-camera video processing, natural virtual scene implantation, and real-time special effect synthesis.
[0017] 2. In the present invention, through the multi-shot automatic transition model, the features of multi-camera video streams can be intelligently analyzed, and the shots can be accurately switched according to regulations and machine learning strategies, making the video shot transitions smooth and natural, improving visual coherence and appeal; the virtual scene implantation model generates adapted elements according to the features of video frames, accurately locates and naturally integrates them, enriching the content, enhancing the visual effect and immersion; the model pre-training and optimization general strategy improves the model performance and generalization ability, and improves the processing accuracy and efficiency. It solves the problems of low efficiency of traditional manual shot transitions, unrealistic element generation and unnatural fusion during virtual scene implantation.
[0018] 3. In the present invention, by extracting multi-faceted features, the multi-camera video stream information is comprehensively mined, which helps the shot transition decision-making to be more intelligent and smooth, improves the video's appreciation and coherence. At the same time, the relevant features of the video frames are analyzed to provide an accurate reference for the integration of virtual scene elements, ensuring their matching and integration with all aspects of the video frames and naturally integrating into the video to enhance the visual effect and immersion, thus solving the problems of inaccurate and unintelligent shot transition decision-making and the mismatch between virtual scene elements and video frames.
[0019] 4. In the present invention, through feature matching and virtual scene selection, the virtual scene elements are accurately screened and adjusted according to the video frame features, improving their adaptability to the video frames and preparing for integration. Then, with the help of image fusion technology, seamless fusion is achieved, eliminating traces, enhancing the richness, realism and appreciation of the picture, and improving the overall visual expressiveness and viewing experience. This solves the problems of poor adaptability between virtual scene elements and video frames and obvious boundaries in the fusion of virtual scenes and video pictures.
[0020] 5. In the present invention, through the division of labor and cooperation of each unit of the intelligent video processing system, the video stream acquisition and synchronization unit lays the foundation, the deep learning model unit realizes intelligent processing, the feature extraction and analysis unit helps with accurate operation, the shot transition decision-making unit ensures intelligent and smooth shot transition, the virtual scene implantation and fusion unit enables the natural integration of virtual scenes, the real-time special effect synthesis unit improves the video quality, and the output management unit ensures multi-scene playback and display. This solves the problems of asynchronous multi-camera video stream acquisition and unnatural virtual scene implantation. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is the method flow chart of the multi-camera video stream intelligent editing and real-time special effect synthesis method proposed by the present invention; Figure 2 is the system architecture diagram of the intelligent video processing system of the multi-camera video stream intelligent editing and real-time special effect synthesis method proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] Embodiment 1 Referring to Figure 1 , in the first embodiment of the present invention, the present invention provides a multi-camera video stream intelligent editing and real-time special effect synthesis method. The intelligent video processing software system includes: S1. Video acquisition and synchronization: The intelligent video processing system connects to shooting devices through multiple interfaces, acquires multi-camera video streams, and uses multi-modal information synchronization algorithms to ensure that the videos of each camera are accurately synchronized in time and space; S2. Model pre-training optimization: Build a multi-camera switching and virtual scene implantation model, initialize it with pre-trained weights, adjust parameters through self-supervised learning, fine-tuning and optimization algorithms to improve model performance; S3, Feature extraction and analysis: Extract visual, motion, and semantic features related to multi-lens switching and virtual scene implantation from synchronized video streams to comprehensively analyze video content; S4, lens switching decision: Combine preset rules with machine learning methods to determine the lens switching requirements, determine the best switching time and position, and realize intelligent lens switching; S5, scene embedding and fusion: select and adjust virtual scene elements according to video features, and use fusion technology to naturally integrate them into the video frame to enrich the video picture; S6. Real-time special effects synthesis: Add color, blur, and particle real-time special effects to the processed video frames to enhance the video visual effects and complete the overall processing flow.
[0024] Specifically, the intelligent video processing system uses Ethernet, SDI, HDMI, Wi-Fi 6, Bluetooth 5.2 and other interfaces to connect shooting equipment to collect multi-camera video streams. Using the multimodal information synchronization algorithm, the videos of each camera are first preliminarily aligned with microsecond timestamps, and then the audio features are extracted. The time synchronization is finely adjusted through short-time Fourier transform and DTW algorithm. Visual features are extracted for key frames, and the homography matrix is calculated by matching methods and RANSAC algorithm to achieve high-precision spatiotemporal synchronization at the frame level, ensuring the accuracy of the foundation for subsequent processing, solving the problem of asynchrony of multi-camera videos and affecting the subsequent processing effect, and achieving orderly and precise coordination of videos from each camera, providing a reliable premise for subsequent editing and special effects synthesis.
[0025] Construct multi-camera switching and virtual scene implantation models. In the multi-camera switching model, CNN uses deep separable convolution and other methods to extract features, and LSTM uses a bidirectional structure combined with a multi-head attention mechanism to grasp the temporal features; in the virtual scene implantation model, the GAN generator, discriminator, and semantic segmentation network each have a specific architecture to ensure the effect. During pre-training, the multi-camera switching model is first initialized with ImageNet weights, and then adjusted through multi-stage training combined with loss functions and optimization algorithms; the virtual scene implantation model is trained separately and jointly in stages, and regularization, Dropout, etc. are used to prevent overfitting. These improve the performance of the models so that they can better serve subsequent video processing work.
[0026] Extract relevant features from the synchronous video stream for video processing. When extracting features related to multi-shot automatic switching, visual features are obtained by CNN, color histograms, Gabor filters, etc.; motion features are calculated by the optical flow algorithm to determine pixel displacement; semantic features are analyzed by language models such as BERT to analyze text information. In the analysis of relevant features for virtual scene implantation, scene features are used to judge the scene type and layout based on a semantic segmentation network; lighting features are clarified through image analysis algorithms such as lighting intensity; style features are used to master the picture style through style transfer algorithms. Comprehensively extract features to provide accurate basis for subsequent shot switching decisions, scene implantation, etc., and help create high-quality videos.
[0027] This step combines two methods to determine shot switching. In rule-based decision-making, content rules switch shots at key points according to video types and themes. For different scenarios such as sports events and interviews, there are corresponding switching timings; rhythm rules select appropriate switching frequencies and methods according to the speed of the video rhythm; composition rules consider the aesthetic feeling of the picture and select camera positions. In machine learning-based decision-making, classification models determine decisions based on the predicted switching probabilities of trained models; reinforcement learning models use agents to interact with the environment and adjust strategies according to reward feedback. Combining the two can accurately judge the shot switching requirements, improve the rationality and accuracy of shot switching, and optimize the video presentation effect.
[0028] First, feature matching and virtual scene selection are carried out. Based on the scene, lighting, and style features of the video frame, compare the elements in the material library to select suitable virtual scene elements. For example, for a video of a seaside scene, select corresponding seaside virtual elements, and then adjust the elements according to the lighting and style characteristics. Then, image fusion technology is used. Alpha fusion achieves natural transitions by controlling the Alpha values of the pixels of the elements; Poisson fusion uses algorithms to seamlessly connect the elements with the background, making the virtual scene naturally integrate into the video frame, enriching the content of the video picture, enhancing the video's appreciation, making the video more attractive, and avoiding the abruptness of virtual scene implantation.
[0029] Add real-time special effects to the processed video frames. Color adjustment special effects use histogram equalization to enhance contrast, color correction to ensure color unity, and color lookup tables to transform color styles; blur special effects use Gaussian blur to simulate depth of field to highlight the main body, average blur to reduce noise or create an atmosphere, and adjust parameters as needed to change the effect; particle special effects use particle system algorithms to define particle attributes to simulate natural and fantasy effects such as snowflakes and fireworks. By adding these special effects, the visual impact and interest of the video are enhanced, meeting different creative needs, further improving the overall quality of the video, and making it more suitable for diverse application scenarios.
[0030] Through the intelligent and automated processing of multi-camera video streams, the video streams of each camera can be highly accurately synchronized in space and time, laying a precise foundation for subsequent processing. The constructed and optimized model can accurately extract the video content features, accurately judge the camera switching timing, and select the appropriate camera position to achieve intelligent and smooth camera switching. It can screen and naturally integrate virtual scene elements according to the video frame features, enhancing the richness and ornamental value of the video picture. It can also add various special effects to the video frames in real time, enhancing the visual presentation effect of the video, and finally outputting high-quality multi-camera video content that meets diverse requirements. This solves the problems of low efficiency in manual editing, natural virtual scene implantation, and real-time special effect synthesis in existing multi-camera video processing.
[0031] Video acquisition and synchronization include: Hardware interfaces for connecting various shooting devices to achieve the acquisition of video streams, covering types such as Ethernet, SDI, HDMI, Wi-Fi 6, and Bluetooth 5.2 to connect different devices; Multi-modal information synchronization algorithms for synchronizing the video streams collected by each camera using timestamps, audio features, and visual features. First, perform preliminary synchronization with timestamps, and then accurately align at the audio-visual level through audio fingerprints, DTW algorithms, deep learning, and homography matrix calculations to ensure the spatio-temporal consistency of video frames.
[0032] Specifically, the Ethernet interface uses the TCP / IP protocol for data transmission and can stably connect devices such as professional cameras and webcams with a relatively high bandwidth (e.g., Gigabit Ethernet can reach 1Gbps), and transmit the captured video stream to the intelligent video processing system through the network line. The SDI interface follows the serial digital interface standard and transmits uncompressed high-definition video signals through coaxial cables, ensuring the high quality and real-time nature of the video, and is suitable for broadcast-level professional shooting devices. The HDMI interface uses its high-definition multimedia interface protocol to simultaneously transmit video and audio signals, and can conveniently connect consumer-grade cameras, cameras, and some portable shooting devices to quickly transmit video data into the system. The Wi-Fi 6 module is based on the 802.11ax standard, supports multiple users to connect simultaneously, has a higher data transmission rate (theoretical peak can reach 9.6Gbps) and lower latency, and can enable wireless shooting devices (such as drones, action cameras, etc.) to transmit video streams to the system in real time within a certain range (usually dozens of meters indoors and hundreds of meters outdoors). The Bluetooth 5.2 module uses low-power Bluetooth technology. Although the transmission rate is relatively low, it can be used for quick pairing between devices and the transmission of some simple control instructions, such as obtaining the basic status information of the shooting device and controlling some simple parameter settings of the shooting device, to assist the video stream acquisition work.
[0033] When each shooting device records a video, it attaches a timestamp accurate to the microsecond level to each frame of video data. The intelligent video processing system will select the video stream of a host position (for example, the position with the best picture stability and the most critical shooting angle as the host position) as the reference. Then, it compares the timestamps of the video streams of other positions with that of the host position. If the difference in timestamps is small, the system will make adjustments by inserting or deleting a small number of video frames. For example, when it is found that the video stream of a certain position is a few frames faster than that of the host position, the corresponding number of frames of that position will be deleted; if it is a few frames slower, blank frames will be inserted for compensation. If the difference in timestamps is large, the system will adjust by playing the video stream at a variable speed in proportion to make the video streams initially aligned on the macroscopic time scale.
[0034] Based on the initial synchronization, the system will extract the audio signals from the video streams of each position. The audio signals are framed at a duration of 20 - 50 milliseconds per frame, and then the short-time Fourier transform (STFT) is performed on each frame of the audio signal to convert the audio signal in the time domain into the spectral characteristics in the frequency domain, thereby generating audio fingerprints. Through the dynamic time warping (DTW) algorithm, the optimal matching path between the audio fingerprints of the video streams of different positions is calculated, and the time of the video streams is further finely adjusted according to the matching result, so that the audio and video are more accurately aligned in time, and the synchronization accuracy can be improved to the millisecond level.
[0035] The system will select key frames from the video streams of each position according to certain rules (for example, every 1 second or according to the dynamic change law of the video content, selecting the area with large picture changes as the key frame). Then, using the feature extraction network in deep learning (such as the SIFT algorithm improved based on the convolutional neural network), the key frames are subjected to feature extraction to obtain a series of feature points and descriptors. Through the matching method based on the ratio of the nearest neighbor distances, matching feature point pairs are found between the key frames of different positions. Then, the random sample consensus (RANSAC) algorithm is used to estimate the homography matrix according to the matching feature point pairs. Finally, based on this homography matrix, geometric transformation adjustment is performed on each frame of the images of the video streams of each position, so that the video frames of each position can also be accurately aligned in space, thereby achieving the accurate spatio-temporal synchronization of the videos of each position.
[0036] Through the above-mentioned hardware interfaces and multi-modal information synchronization algorithms, the intelligent video processing system can stably and efficiently collect video streams from various types of shooting devices, and can achieve high-precision synchronization in terms of time and space. This enables the video frames from different camera positions to be accurately aligned in time and accurately corresponding in space, providing precise and unified basic data for subsequent processing tasks such as intelligent editing of multi-camera video streams, virtual scene implantation, and real-time special effect synthesis, ensuring the accuracy and consistency of each step in the subsequent processing process. It solves the problems of unstable video stream acquisition and inconsistent transmission rate in traditional multi-camera video processing.
[0037] Model pre-training optimization includes: A multi-shot automatic switching model, which is used to analyze the feature information in the multi-camera video stream. Based on the extracted spatial and temporal features, through preset rules and machine learning strategies, it can intelligently judge and decide the best timing and corresponding camera positions for shot switching, and perform automatic switching of video shots; A virtual scene implantation model, which is used to generate virtual scene elements adapted to the video content by means of a generative adversarial network according to the video frame features, and rely on a semantic segmentation network to accurately locate the fusion area, and naturally integrate the virtual scene elements into the video frame; A general strategy for model pre-training and optimization, which is used to perform preliminary training using pre-trained weights and self-supervised learning, then adjust the parameters using labeled data combined with multi-loss functions and optimization algorithms, and also adopt regularization and Dropout to prevent overfitting, and dynamically adjust hyperparameters.
[0038] Specifically, the model first extracts spatial features through a convolutional neural network (CNN). The convolutional layer of the CNN performs convolutional operations on the video frames to extract local features such as object shapes, colors, and textures. For example, in a sports event video, features of objects such as athletes and balls can be recognized. The pooling layer then downsamples the output of the convolutional layer, reducing the data volume while retaining key features. Convolution kernels of different sizes can extract spatial features of different scales. Subsequently, the features extracted by the CNN are input into a long short-term memory network (LSTM) or its variant (such as a bidirectional LSTM) to capture temporal features. The gating mechanism (input gate, forget gate, output gate) of the LSTM can selectively retain or forget information and learn the temporal dependence between video frames. For example, in a movie scene, temporal information such as the sequence of character actions and dialogues can be captured.
[0039] Make decisions based on the extracted spatial and temporal features, combined with preset rules and machine learning strategies. The preset rules include content-based rules, such as switching to a close-up of the guest when the guest is speaking in an interview program; and rhythm-based rules, such as quickly switching shots in an action movie to create a tense atmosphere. In terms of machine learning strategies, classification models (such as support vector machines, random forests, etc.) or reinforcement learning models (such as deep Q networks) are used. Taking reinforcement learning as an example, regard the shot-switching problem as a Markov decision process. The agent takes actions (selects a switching shot) according to the current video state (feature information) and obtains rewards based on the audience feedback (such as click-through rate, viewing duration, etc.). Through continuous trial and error learning, optimize the decision-making strategy to determine the best shot-switching timing and corresponding camera positions.
[0040] Use a generative adversarial network (GAN) to generate virtual scene elements. A GAN consists of a generator and a discriminator. The generator takes random noise or partial video features as input and generates virtual scene elements (such as virtual buildings, props, etc.) through a multi-layer neural network. The discriminator then discriminates between the generated elements and real scene elements and outputs the probabilities of true and false. During the training process, the generator is continuously optimized to make the generated elements more realistic, and the discriminator continuously improves its discrimination ability. For example, when shooting a period drama, generate virtual scene elements such as ancient-style buildings and decorations.
[0041] Perform semantic segmentation on video frames through a semantic segmentation network (such as U-Net, DeepLab, etc.), and divide the video frames into different semantic regions (such as people, background, sky, etc.). According to the semantic segmentation results, accurately locate the regions suitable for fusing virtual scene elements. For example, in an outdoor scenery video, virtual elements can be fused into the background region. Then, use image fusion techniques (such as Alpha fusion, Poisson fusion, etc.) to naturally integrate the generated virtual scene elements into the video frames. Alpha fusion realizes the fusion by adjusting the transparency (Alpha value) of the elements, and Poisson fusion makes the fused image more natural in terms of color, texture, etc. by solving the Poisson equation.
[0042] Initialize the model parameters with the weights pre-trained on large-scale unlabeled data, such as the CNN weights pre-trained on the ImageNet dataset. Then, use self-supervised learning methods to further train the model. Self-supervised learning designs proxy tasks (such as video frame sorting, image colorization, etc.) to let the model automatically learn feature representations from the data, mine the internal structure and laws of the data, so that the model can also learn useful features without manual annotation.
[0043] Fine-tune the pre-trained model using labeled data. Calculate the difference between the model prediction and the actual annotation by combining multiple loss functions (such as classification loss, generative adversarial loss, semantic segmentation loss, etc.). Update the model parameters by back-propagation through optimization algorithms (such as stochastic gradient descent, Adam, etc.) to minimize the loss function. At the same time, use the L2 regularization method to constrain the parameter size to prevent the model from overfitting; use the Dropout technique to randomly discard some neurons to avoid the model from over-reliance on certain features and improve the model's generalization ability. In addition, dynamically adjust hyperparameters (such as learning rate, batch size, etc.) and adjust in real time according to performance indicators (such as accuracy, loss value, etc.) during training to achieve the best training effect.
[0044] The multi-camera automatic switching model can intelligently analyze the characteristics of multi-camera video streams, automatically and accurately switch lenses according to rules and machine learning strategies, achieve smooth and natural transitions of video lenses, and improve visual coherence and viewing quality; the virtual scene implantation model can generate adaptive virtual scene elements based on video frame features, accurately locate the fusion area and naturally integrate it into the video frame, enrich the video content, and enhance the visual effect and immersion; the general strategy for model pre-training and optimization uses pre-training, self-supervised learning, fine-tuning, and parameter optimization to improve model performance and generalization capabilities, so that the model performs well in multi-camera switching and virtual scene implantation tasks, and improves processing accuracy and efficiency. It solves the problems of low efficiency of traditional manual lens switching, unrealistic element generation, and unnatural fusion when implanting virtual scenes.
[0045] Feature extraction analysis includes: Extraction of related features for multi-lens automatic switching, which is used to mine visual, motion, and semantic features from multi-camera video streams, provide a basis for multi-lens automatic switching decisions, and assist in determining when to switch lenses and which camera to switch to; Virtual scene implantation related feature analysis is used to analyze the scene, lighting, and style characteristics of video frames, providing a reference for the accurate selection, adaptation, and fusion of virtual scene elements, and helping to naturally integrate virtual scenes into videos.
[0046] Specifically, the multi-camera video stream is input into the pre-trained CNN model, such as VGGNet, ResNet, etc., frame by frame. The convolution layer in CNN performs convolution operations by sliding the convolution kernel on the image to extract local features of objects in the image, such as edges and textures. As the number of network layers increases, high-level convolution layers can extract more abstract semantic features, such as the overall features of objects such as faces and vehicles. The pooling layer downsamples the output of the convolution layer to reduce the amount of data while retaining key features. For example, the maximum pooling layer selects the maximum value in each pooling window as the output, and the average pooling layer calculates the average value within the window.
[0047] Convert the video frames from the RGB color space to other color spaces (such as HSV, Lab, etc.), and then divide the color space into several intervals according to certain quantization rules. Count the number of pixels in each interval to form a color histogram. The color histogram can reflect the color distribution in the image. For example, in a video frame with a blue dominant color, the number of pixels in the blue interval will be relatively large. By comparing the color histograms of different video frames, the color similarity or difference between them can be judged.
[0048] Perform convolution operations on the video frames using Gabor filters with different directions and scales. Gabor filters are linear filters used for texture analysis, which can extract texture information with different directions and frequencies in the image. For example, a Gabor filter in the horizontal direction can highlight the texture features in the horizontal direction of the image, and Gabor filters with different scales can capture texture details of different sizes. By analyzing the output of the Gabor filters, the texture features of the video frames can be obtained.
[0049] Common optical flow algorithms include the Lucas - Kanade algorithm and the Farneback algorithm. Taking the Lucas - Kanade algorithm as an example, it is based on the assumption of local pixel gray - level consistency, and determines the motion speed and direction of the object by calculating the displacement of corresponding pixels in adjacent video frames. Specifically, within a small window, it is assumed that the pixels within the window have the same motion vector, and the motion vector is solved by minimizing the difference in pixel gray - levels within the window. The Farneback algorithm is a multi - scale - based optical flow calculation method. It constructs an image pyramid, calculates the optical flow at different scales, and then fuses this optical flow information to obtain a more accurate motion estimate. Through optical flow algorithms, the motion trajectory and speed of the objects in the video frames can be obtained, providing a basis for shot - switching decisions.
[0050] For the text information in the video (such as subtitles, commentaries, etc.), first perform pre - processing, including operations such as word segmentation and stop - word removal. Then input the processed text information into a pre - trained language model. The BERT model is based on the Transformer architecture and learns the semantic relationships between words in the text through the self - attention mechanism. When processing video text information, BERT can extract semantic features such as the theme, sentiment, and keywords of the text. For example, in a news video, by analyzing the subtitle text, BERT can identify whether the news theme is about sports events, political events, or entertainment news, etc.
[0051] Semantic segmentation networks (such as U-Net, DeepLab, etc.) are used to perform semantic segmentation on video frames. Through operations such as convolution, pooling, and upsampling, the semantic segmentation network divides the video frame into different semantic regions, such as people, background, sky, ground, etc. By analyzing the results of semantic segmentation, the scene type of the video frame can be determined, for example, whether it is an indoor scene or an outdoor scene, whether it is a city street or a natural landscape, etc. At the same time, the positional relationship of objects in the scene can also be analyzed, such as the relative position of people and the background, providing a reference for the implantation of virtual scene elements.
[0052] Estimate the light intensity by calculating the average gray value of pixels in the image. In a grayscale image, there is a certain relationship between the gray value of pixels and the light intensity. Generally speaking, the higher the gray value, the stronger the light intensity. In addition, more complex lighting models, such as physically based lighting models, can also be used to calculate the light intensity more accurately. Judge the color temperature by analyzing the proportion and hue of the color components in the image. For example, in the RGB color space, an image with more blue components usually has a lower color temperature, while an image with more red components has a higher color temperature. The judgment of color temperature can help select virtual scene elements that match the lighting conditions of the video frame. Infer the direction of light based on the highlights, shadows of objects in the image and the light and dark contrast of the objects, combined with the structural information of the scene.
[0053] Adopt a style transfer algorithm to analyze the style features of video frames. The style transfer algorithm separates and fuses the content features of the video frame and the style features of the reference image, thereby learning the style characteristics such as color, texture, and composition of the video frame. Specifically, first use a pre-trained CNN model to extract the content features of the video frame and the style features of the reference image, and then adjust the pixel values of the video frame through an optimization algorithm, so that the adjusted video frame is similar to the original video frame in content and similar to the reference image in style. In this way, the unique style features of the video frame can be analyzed, providing a basis for the style adaptation of virtual scene elements.
[0054] By extracting visual, motion, and semantic features, comprehensively and deeply excavate the multi-camera video stream information, provide rich and accurate basis for shot transition decisions, achieve more intelligent, smooth, and natural shot transitions, and enhance the video's appreciation and coherence; in the analysis of relevant features for virtual scene implantation, analyze the scene, lighting, and style features of the video frame, provide precise references for the selection, adaptation, and fusion of virtual scene elements, ensure that the virtual scene elements match and blend with the video frame in terms of scene type, lighting effect, and style characteristics, and naturally integrate into the video, enhancing the visual effect and immersion. Solve the problems of inaccurate and non-intelligent shot transition decisions and the mismatch between virtual scene elements and video frames.
[0055] Shot transition decisions include: The rule-based decision-making method is used to make decisions on multi-camera video lens switching based on the rules set in terms of video content type, rhythm, and composition, to ensure that the lens switching timing is appropriate and the picture presentation is reasonable; The decision-making method of machine learning is used to learn from multi-camera video data, explore the rules of shot switching, and intelligently determine the timing and position of shot switching based on classification or reinforcement learning model output.
[0056] Specifically, for different types of video content, such as sports events, talk shows, drama films, documentaries, etc., specific camera switching rules are set. Taking sports events as an example, in football matches, when players perform key passes, shots, steals, etc., the rules require switching to a camera position that can clearly show the details of the action, such as switching to a camera position that closely follows the player to ensure that the audience does not miss important moments; for basketball games, at key moments such as players taking off to shoot and rebounding, it is necessary to switch to a camera that can fully present the entire action process and the positions of the surrounding players. In talk shows, when the guest begins to speak and expounds on important points, the rules require switching to a close-up of the guest to highlight his expression and demeanor, so that the audience can better understand the content of the speech; and when the host asks questions, it can switch to a medium shot to show the interactive scene between the two parties. For drama films, at key plot points such as plot turning points and character emotional outbursts, switch to the corresponding camera position that can set off the atmosphere and show the emotional changes of the characters according to the logic of the plot development. In documentaries, when introducing important historical relics, natural landscapes and other subjects, switch to a camera position that can clearly show the full picture and details of the subject.
[0057] The rules for switching shots are formulated based on the speed of the overall rhythm of the video and the trend of emotional changes. In scenes with a fast rhythm and tense and exciting emotions, such as the chase and fighting scenes in action movies and the climax of sports events, the frequency of shot switching will be increased accordingly. By using a fast switching method, by frequently switching between different camera positions and different scenes, a tense and dynamic atmosphere will be created, allowing the audience to feel the tension and excitement more immersively. In the soothing and emotionally stable plot parts, such as the slow display of natural scenery in documentaries and the daily dialogue scenes of characters in feature films, the frequency of shot switching will be reduced, and a soft and slow switching method will be used, so that the audience can be more immersed in the current picture atmosphere, feel calm and peaceful, and avoid disrupting the overall soothing rhythm due to frequent shot switching.
[0058] From the perspective of the aesthetic feeling of the camera composition, analyze factors such as the layout of objects in the picture, the integrity of the perspective, and the balance of the picture to determine the camera switching rules. When the layout of objects in the picture is unreasonable, for example, the main character is biased to one side, affecting the overall aesthetic feeling, switch to a camera position that can place the main body in the center of the picture or better conform to the composition principles such as the golden ratio according to the rules; if the current camera position perspective cannot fully present the scene or the whole picture of the object to be shown, switch to a camera position that can provide a more comprehensive perspective; for the balance of the picture, if there are too many elements on one side, making the picture unbalanced, switch to a camera position that can adjust the distribution of the picture elements to achieve visual balance, so as to ensure that the picture presentation is more reasonable and beautiful after each camera switch, meeting the requirements of visual aesthetics.
[0059] First, collect a large number of multi-camera video data with annotations. These annotation information includes key contents such as the timing of camera switching and which camera position to switch to. Divide these video data into a training set, a validation set, and a test set according to a certain proportion. Select a suitable classification model, such as Support Vector Machine (SVM), Random Forest, or Deep Neural Network (DNN), etc. Take video frames as units, extract various features of video frames (such as the visual features, motion features, semantic features mentioned above), and input these features as input data into the classification model for training. During the training process of the model, by continuously adjusting the internal parameters (such as the support vectors of SVM, the neuron weights of DNN, etc.), learn the mapping relationship between the features and camera switching. In the validation and testing phases, for the features of newly input video frames, the model will output a value representing the probability of camera switching. When this probability value exceeds a pre-set threshold, it is determined that a camera switch is required, and the specific camera position to switch to is determined according to the rules learned by the model, so as to achieve the camera switching decision based on the classification model.
[0060] The lens switching problem is constructed as a reinforcement learning problem, and the intelligent video processing system is regarded as an agent, and the environment in which the video is located (including the current state of the video, audience feedback, etc.) is regarded as the environment. The agent takes actions in different video frame states (i.e., decides whether to switch the lens and which camera to switch to), and learns the optimal decision-making strategy based on the feedback obtained after taking the action. For example, using a deep Q network (DQN) or a policy gradient algorithm (such as A2C, A3C, etc.), the agent selects an action based on the characteristics of the current video frame (such as the content of the picture, the previous lens switching history, etc.), and then the environment will give a reward signal. This reward signal comprehensively considers many factors, such as the audience's attention (such as the click-through rate of the video, the viewing time, etc.), visual comfort (whether there is screen jitter, abrupt switching, etc.), plot coherence (whether the lens switching conforms to the logic of the plot development), etc. Based on this reward signal, the agent gradually learns the best shot switching action under different video frame states by continuously updating its value function (in DQN) or policy function (in the policy gradient algorithm). That is, it determines the most appropriate shot switching time and position, thereby realizing intelligent shot switching decisions.
[0061] The decision-making method based on rules makes lens switching decisions based on the video content type, rhythm and composition rules, so that the lens switching fits the key plot, atmosphere and aesthetics, making the video rhythm smooth and the picture beautiful, improving the audience's viewing experience and immersion; the decision-making method of machine learning uses the learning and mining rules of multi-camera video data, and makes intelligent decisions based on classification or reinforcement learning models, which can accurately determine the best switching time and camera position, reduce human errors, improve the accuracy and adaptability of lens switching, and make the video more natural, smooth, and more enjoyable and coherent. It solves the problem that lens switching relies on manual experience and lens switching decisions are not accurate enough in complex environments.
[0062] Scene embedding fusion includes: Feature matching and virtual scene selection, which is used to select suitable virtual scene elements from the virtual scene material library based on the video frame features, and make corresponding adjustments to them in preparation for subsequent natural integration into the video frame; Image fusion technology is used to seamlessly fuse the selected virtual scene elements with the video frames, eliminate fusion traces, make the virtual scene naturally integrated into the video screen, and enhance the overall visual effect of the video.
[0063] Specifically, analyze the video frame to extract its scene features. For example, use a semantic segmentation network to divide the video frame into different semantic regions, such as distinguishing different parts like the sky, ground, people, buildings, etc., and clarify the specific scene type presented by the video frame, whether it is an indoor scene, an outdoor natural scene, or an urban street scene. Then, search for virtual scene elements in the virtual scene material library that match the scene type. For instance, when the video frame shows a seaside beach scene, search for virtual scenes in the material library that contain relevant elements such as beaches, waves, blue sky, and white clouds. At the same time, analyze the positional relationships and scene layouts of various objects in the video frame and other detailed information to further screen out virtual scene elements that can be adaptively matched in terms of layout structure, ensuring that there are no logical conflicts in the scene after the virtual scene elements are incorporated.
[0064] Calculate the lighting characteristics of the video frame. Estimate the lighting intensity by analyzing the gray value distribution of image pixels, judge the color temperature based on the proportion of color components, and then infer the lighting direction in combination with the highlights and shadows of objects. For the virtual scene elements preliminarily screened from the material library, adjust the lighting intensity of the virtual scene elements according to the lighting intensity of the video frame using a lighting model (such as the Phong lighting model or the Blinn - Phong lighting model) to make it match the overall lighting intensity of the video frame.
[0065] Adopt a style transfer algorithm to analyze the style features of the video frame. For example, by performing feature decomposition and fusion on the video frame and a reference style image, learn the style characteristics of the video frame in terms of color combination, texture performance, composition method, etc. Then, make corresponding style adjustments to the virtual scene elements. If the video frame has a retro color tone and an oil painting texture style, use image processing techniques to change the color channel parameters of the virtual scene elements, add texture effects similar to oil painting brushstrokes, and adjust its composition layout to make it closer to the style of the video frame, so that the virtual scene elements can be coordinated with the video frame in terms of style, making full preparations for subsequent natural integration into the video frame.
[0066] Alpha blending is a blending method based on transparency control. For the selected and adjusted virtual scene elements, an Alpha value (ranging from 0 to 1, where 0 represents completely transparent and 1 represents completely opaque) is set for each pixel. When blending the virtual scene elements with the video frame, the display degree of each pixel in the blended video frame is determined according to this Alpha value. For example, for the parts in the virtual scene elements that want to create a hazy effect (such as the smoke effect in the distance), the corresponding pixel Alpha values can be set lower, so that they present a semi-transparent effect in the blended video frame, and naturally transition with the original picture in the video frame, avoiding obvious boundary senses. By reasonably setting the Alpha values of pixels in different regions, the smooth visual blending of virtual scene elements and video frames is achieved.
[0067] Poisson blending mainly achieves the blending effect based on the Poisson equation. First, analyze the pixel characteristics at the boundary between the video frame background and the virtual scene elements to determine the boundary conditions. Then, construct the Poisson equation according to these boundary conditions, and adjust the values of the boundary pixels and their surrounding pixels of the virtual scene elements by solving this equation, so that after the virtual scene elements are integrated into the video frame, the color, texture, etc. at the boundary can be naturally connected. For example, when integrating a virtual flower into a video picture of a grassland, using Poisson blending technology, accurately calculate the pixel changes at the junction of the flower boundary and the grassland background, making the flower look as if it originally grew on the grassland, eliminating the blending traces, and enabling the virtual scene elements to be naturally and seamlessly integrated into the video picture, thereby enhancing the overall visual effect of the video.
[0068] Based on the multi-faceted features of the video frame, the virtual scene elements are accurately screened and adjusted through feature matching and virtual scene selection, making them highly adaptable to the video frame in all dimensions, laying a foundation for natural integration, and making the implantation of the virtual scene fit the video context and style; the application of image blending technology uses Alpha blending, Poisson blending, etc. to achieve seamless blending of virtual scene elements and video frames, eliminating traces, enhancing the richness, realism and ornamental value of the video picture, and improving the overall visual expressiveness and viewing experience. It solves the problems of poor adaptability between virtual scene elements and video frames and obvious boundaries easily appearing in the blending of virtual scenes and video pictures.
[0069] Real-time special effect synthesis includes: Color adjustment special effects, used to change the color attributes of the video frame, enhance contrast, correct color deviation or convert color styles; Blur special effects, used to blur the video frame, create a specific atmosphere, highlight the main body, reduce noise or create a hazy feeling; Particle effects are used to simulate a variety of natural or fantasy visual effects such as snowflakes, raindrops, and fireworks in videos, enhancing the visual impact and interest of the video.
[0070] Specifically, the histogram equalization algorithm is used to achieve contrast enhancement. The algorithm first counts the grayscale histogram of the video frame image, that is, calculates the number of pixels that appear at each grayscale level in the image. Then, by equalizing the histogram, the grayscale distribution of the original image is remapped so that the pixel grayscale values are more evenly distributed throughout the grayscale range. For example, the grayscale difference between the darker and brighter areas in the original image is small. After histogram equalization, the dark area will become darker and the bright area will become brighter, thereby increasing the overall contrast of the image and making the details in the image more clearly visible. This is especially suitable for video frames that were shot in dim light or have low contrast.
[0071] Use color correction algorithms, such as those based on the principle of white balance. First, select a reference point that is considered to be white in the video frame (usually an object in the picture that is known to be white, such as white paper, walls, etc.), and then calculate the deviation value of each color channel (such as RGB channel) based on the actual color of this reference point. By adjusting the gain coefficients of these color channels, the color of the reference point is restored to true white, and then the color of the entire video frame is corrected according to this adjustment ratio to ensure that videos shot from different camera positions or video frames with color deviations due to the shooting environment can achieve relatively standard and unified color effects, avoiding the problem of color inconsistency in the same video.
[0072] The color lookup table (LUT) technology is used to achieve the conversion of color styles. In advance, a corresponding color lookup table is made according to different color style requirements (such as retro tones, Japanese fresh tones, film tones, etc.). It is essentially a mapping relationship that defines the conversion rules from the original image color value to the target color style color value. When processing video frames, you only need to replace the original color value of each pixel in the video frame with the corresponding target color value according to this pre-set lookup table, and you can quickly achieve the transformation of color style and give the video frame a different visual color experience.
[0073] A weight matrix is generated based on the Gaussian function. The size and standard deviation of this matrix can be set as needed. When blurring the video frame, each pixel is taken as the center and the weighted average of the surrounding pixels is calculated according to the weight matrix. For example, for a certain pixel in the image, the closer the pixel is to it, the greater the weight it has in the weighted average, and the farther the pixel is, the smaller the weight is. Through this calculation method, the image as a whole produces a soft and uniform blur effect. It is often used to simulate the depth of field effect. For example, by Gaussian blurring the background, the subject in the focus position can be highlighted, guiding the audience's attention to focus on important elements of the picture.
[0074] The blur effect is achieved by calculating the average value of pixels in a local area of the image. First, determine the size of a blur window (such as a 3×3, 5×5, etc. pixel area), then take each pixel as the center on the video frame, take the average value of all pixels in the window, and use this average value to replace the value of the center pixel. In this way, all pixels of the entire video frame are continuously traversed to blur the entire image. This method can effectively reduce the noise in the image. For example, in video frames shot in low light and with high ISO values, there are often more noise points. Mean blur can make the picture look smoother and cleaner. At the same time, according to the creative needs, the size of the blur window and other parameters can be adjusted to create different degrees of haziness and add a specific atmosphere to the video.
[0075] First, we need to define various properties of particles, including basic properties such as initial position, initial velocity, color, life cycle, size, etc. For example, when simulating the snowflake effect, the initial position of the particles can be randomly distributed in the upper area of the video screen to simulate the starting state of falling from the sky; the initial velocity is set to a smaller value and the direction is downward, which conforms to the natural falling characteristics of snowflakes; the color can be set to white or with a hint of blue to reflect the true color of snowflakes; the life cycle can be determined based on the approximate length of time it takes for snowflakes to melt from the appearance to the ground, so that the particles disappear naturally after a certain period of time. At the same time, it is also necessary to define the movement and change laws of particles, such as the change in acceleration when affected by external forces such as gravity and wind, and the gradual change of properties such as color and transparency during the life cycle.
[0076] In each frame of the video, according to the defined particle attributes and motion change rules, the position, state, etc. of each particle are updated and calculated. The new position, new color, new transparency, etc. of the particles in the current frame are determined through physical simulation or a preset mathematical model, and then these particles are rendered onto the video frame according to their respective states to form the corresponding visual effects. For example, when simulating the fireworks effect, the particles will quickly rise at a set speed and direction in the initial stage of emission, explode and disperse after reaching a certain height, and the color will also change from the initial single color to a riot of colors. By continuously updating and rendering the particle states, a magnificent fireworks blooming scene is presented in the video, thus simulating various natural or fantasy visual effects such as snowflakes, raindrops, fireworks, etc., enhancing the visual impact and interest of the video.
[0077] Through color adjustment special effects, by changing color attributes, the picture quality is improved, the contrast is enhanced to make details clear, the color is corrected to maintain coordination, and the style is transformed to add artistic sense and appreciation; through blur special effects, the video frame is blurred to create an atmosphere. Gaussian blur highlights the main body and enhances the sense of hierarchy, while average blur reduces noise and creates a hazy feeling, enhancing the appeal; particle special effects simulate diverse visual effects, enhance the visual impact and interest, making the video more attractive and stand out. It solves the problems of poor picture color quality and poor special visual effects.
[0078] Embodiment 2: Refer to Figure 2 , in the second embodiment of the present invention, the present invention provides an intelligent video processing system for the intelligent editing and real-time special effect synthesis method of multi-camera video streams, which is used for the intelligent editing and real-time special effect synthesis method of multi-camera video streams. The system includes: A video stream acquisition and synchronization unit, configured with multiple interfaces for connecting shooting devices of different cameras, and built-in synchronization algorithm modules to achieve high-precision synchronization of video streams; A deep learning model unit, which stores and runs pre-trained deep learning models related to multi-shot switching and virtual scene implantation, and has the function of model optimization and update; A feature extraction and analysis unit, which can extract multi-dimensional features from the input video stream and conduct in-depth analysis to provide data support for subsequent processing; A shot switching decision-making unit, which comprehensively uses preset rules and machine learning algorithms to output accurate multi-shot automatic switching instructions; A virtual scene implantation and fusion unit, which matches elements according to features and uses a variety of fusion technologies to integrate virtual scenes into video frames; A real-time special effect synthesis unit, which adds various real-time special effects to the video as needed and has the ability to adaptively adjust special effect parameters; An output management unit, which is responsible for outputting the processed video according to the set format and resolution parameters to meet the requirements of different application scenarios.
[0079] Specifically, each unit of the intelligent video processing system works collaboratively. The video stream acquisition and synchronization unit acquires video streams through multiple interfaces and synchronizes them with high precision using a synchronization algorithm; the deep learning model unit stores and runs relevant models and can optimize and update them; the feature extraction and analysis unit extracts multi-dimensional features and conducts in-depth analysis; the shot switching decision unit combines preset rules and machine learning algorithms to output switching instructions; the virtual scene implantation and fusion unit matches elements and then integrates them into the scene using fusion technology; the real-time special effect synthesis unit adds special effects as needed and adaptively adjusts parameters; the output management unit outputs videos according to set parameters to meet the requirements of different scenarios.
[0080] The video stream acquisition and synchronization unit provides reliable basic data for subsequent processing, improving accuracy and quality; the deep learning model unit realizes intelligent processing, enhancing the viewing experience and professionalism; the feature extraction and analysis unit helps with precise operations in each link, improving the level of intelligence; the shot switching decision unit makes the shot switching intelligent and smooth, enhancing the coherence in line with the content; the virtual scene implantation and fusion unit enables the natural integration of virtual scenes, enhancing the visual expressiveness; the real-time special effect synthesis unit adds special effects in line with the content, improving the video quality; the output management unit ensures that the video is suitable for normal playback and display in multiple scenarios. This solves the problems of asynchronous acquisition of multi-camera video streams and unnatural implantation of virtual scenes.
[0081] Embodiment III In the third embodiment of the present invention, based on the same inventive concept, a computer-readable storage medium is proposed. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the multi-camera video stream intelligent editing and real-time special effect synthesis method of the above embodiment.
[0082] Embodiment IV In the fourth embodiment of the present invention, based on the same inventive concept, a computer device is proposed, including: a processor and a memory; the processor and the memory communicate with each other; the memory is used to store instructions; the processor is used to execute the instructions in the memory to implement the multi-camera video stream intelligent editing and real-time special effect synthesis method of the above embodiment.
[0083] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0084] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. Method for intelligent editing and real-time special effect synthesis of multi-camera video streams, characterized in that, The intelligent video processing software system comprises: S1. Video acquisition and synchronization: The intelligent video processing system connects to shooting devices through multiple interfaces, acquires multi-camera video streams, and uses multi-modal information synchronization algorithms to ensure that the videos of each camera are accurately synchronized in time and space; S2. Model pre-training optimization: Build a multi-camera switching and virtual scene implantation model, initialize it with pre-trained weights, adjust parameters through self-supervised learning, fine-tuning and optimization algorithms to improve model performance; S3, Feature extraction and analysis: Extract visual, motion, and semantic features related to multi-lens switching and virtual scene implantation from synchronized video streams to comprehensively analyze video content; S4, lens switching decision: Combine preset rules with machine learning methods to determine the lens switching requirements, determine the best switching time and position, and realize intelligent lens switching; S5, scene embedding and fusion: select and adjust virtual scene elements according to video features, and use fusion technology to naturally integrate them into the video frame to enrich the video picture; S6. Real-time special effects synthesis: Add color, blur, and particle real-time special effects to the processed video frames to enhance the video visual effects and complete the overall processing flow.
2. The multi-camera video stream intelligent editing and real-time special effect synthesis method according to claim 1, wherein The video acquisition and synchronization include: Hardware interface, used to connect a variety of shooting devices to realize video stream acquisition, covering Ethernet, SDI, HDMI, Wi-Fi6, Bluetooth 5.2 types, connecting different devices; The multimodal information synchronization algorithm is used to synchronize the video streams collected by each camera using timestamps, audio features, and visual features. It first performs preliminary synchronization with timestamps, and then uses audio fingerprints, DTW algorithm, deep learning, and homography matrix calculation to accurately align the audio and video levels to ensure the temporal and spatial consistency of the video frames.
3. The intelligent editing and real-time special effect synthesis method for multi-camera video streams according to claim 1, characterized in that, The model pre-training optimization includes: The multi-lens automatic switching model is used to analyze the characteristic information in the multi-camera video stream, and based on the extracted spatial and temporal characteristics, through preset rules and machine learning strategies, intelligently judge and decide the best time and corresponding camera position for lens switching, and automatically switch the video lens; The virtual scene embedding model is used to generate virtual scene elements that are adapted to the video content based on the video frame features with the help of a generative adversarial network, and accurately locate the fusion area with the help of a semantic segmentation network, so as to naturally integrate the virtual scene elements into the video frame. A general strategy for model pre-training and optimization is used to use pre-trained weights and self-supervised learning for initial training, and then use labeled data combined with multiple loss functions and optimization algorithms to adjust parameters. Regularization and Dropout are also used to prevent overfitting and dynamically adjust hyperparameters.
4. The intelligent multi-camera video stream editing and real-time special effect synthesis method according to claim 1, wherein, The feature extraction analysis includes: Extraction of related features for multi-lens automatic switching, which is used to mine visual, motion, and semantic features from multi-camera video streams, provide a basis for multi-lens automatic switching decisions, and assist in determining when to switch lenses and which camera to switch to; Virtual scene implantation related feature analysis is used to analyze the scene, lighting, and style characteristics of video frames, providing a reference for the accurate selection, adaptation, and fusion of virtual scene elements, and helping to naturally integrate virtual scenes into videos.
5. The intelligent multi-camera video stream editing and real-time special effect synthesis method according to claim 1, wherein The shot switching decision includes: The rule-based decision-making method is used to make decisions on multi-camera video lens switching based on the rules set in terms of video content type, rhythm, and composition, to ensure that the lens switching timing is appropriate and the picture presentation is reasonable; The decision-making method of machine learning is used to learn from multi-camera video data, explore the rules of shot switching, and intelligently determine the timing and position of shot switching based on classification or reinforcement learning model output.
6. The multi-camera video stream intelligent editing and real-time special effect synthesis method according to claim 1, wherein, The scene implantation fusion includes: Feature matching and virtual scene selection, which is used to select suitable virtual scene elements from the virtual scene material library based on the video frame features, and make corresponding adjustments to prepare for the subsequent natural integration into the video frame; Image fusion technology is used to seamlessly fuse the selected virtual scene elements with the video frames, eliminate fusion traces, make the virtual scene naturally integrated into the video screen, and enhance the overall visual effect of the video.
7. The intelligent multi-camera video stream editing and real-time special effect synthesis method according to claim 1, characterized in that The real-time special effects synthesis includes: Color adjustment effects, used to change the color properties of video frames, enhance contrast, correct color deviation or convert color style; Blur effect, used to blur video frames to create a specific atmosphere, highlight the subject, reduce noise or create a hazy feeling; Particle effects are used to simulate a variety of natural or fantasy visual effects such as snowflakes, raindrops, and fireworks in videos, enhancing the visual impact and interest of the video.
8. An intelligent video processing system for the intelligent editing of multi-camera video streams and real-time special effect synthesis, characterized in that, The method for intelligently editing and synthesizing multi-camera video streams according to any one of claims 1 to 7, the system comprising: Video stream acquisition and synchronization unit, equipped with multiple interfaces for connecting different camera shooting devices, built-in synchronization algorithm module to achieve high-precision synchronization of video streams; The deep learning model unit stores and runs pre-trained deep learning models related to multi-camera switching and virtual scene implantation, and has the function of model optimization and update; The feature extraction and analysis unit can extract multi-dimensional features from the input video stream and perform in-depth analysis to provide data support for subsequent processing; The lens switching decision unit uses preset rules and machine learning algorithms to output accurate multi-lens automatic switching instructions; The virtual scene embedding and fusion unit integrates the virtual scene into the video frame based on feature matching elements and adopts a variety of fusion technologies; Real-time special effects synthesis unit, which can add various real-time special effects to the video as needed, and has the ability to adaptively adjust special effects parameters; The output management unit is responsible for outputting the processed video according to the set format and resolution parameters to meet the needs of different application scenarios.
9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the multi-camera video stream intelligent editing and real-time special effects synthesis method as described in any one of claims 1-7 is implemented.
10. A readable storage medium, characterized in that, The readable storage medium stores a computer program, and when the computer program is executed by the processor, the method for intelligent editing and real-time special effects synthesis of multi-camera video streams as described in any one of claims 1-7 is implemented.