Video generation method and electronic device
Patent Information
- Application Number
- CN202610809944.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2046-06-05
AI Technical Summary
虚拟拍摄则能真实捕捉演员表演并与虚拟背景实时结合,但依赖高质量数字资产,制作周期长且成本高
[0045]借由上述技术方案,本发明提供的一种视频生成方法及电子设备,获得人工智能生成的参考视频;提取参考视频中的视觉艺术元素,其中,视觉艺术元素包括摄像机运动参数;利用摄像机运动参数,生成目标运动轨迹;利用目标运动轨迹控制物理拍摄系统捕获前景表演,生成前景画面;判断目标运动轨迹是否超出物理拍摄系统的物理能力极限,如果是,则执行虚实协同拍摄模式,计算目标运动轨迹与物理拍摄系统实际运动轨迹之间的运动差值;利用运动差值驱动虚拟摄像机执行补偿运动,并实时渲染生成动态补偿背景层;将前景画面、动态补偿背景层和参考视频中的背景画面进行智能融合,输出合成画面。本发明通过从人工智能生成的参考视频中提取视觉艺术元素,智能解析并转化为可执行的摄像机运动参数和目标运动轨迹,结合物理拍摄系统与虚拟补偿机制,实现虚实协同拍摄与动态智能合成,有效打通AIGC内容与虚拟拍摄的深度融合流程,从而提升视频内容的视觉一致性和表现力,显著提高视频内容质量。
Smart Images

Figure CN122340289B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a video generation method and electronic device. Background Technology
[0002] Artificial Intelligence Generated Content (AIGC) and virtual filming are two cutting-edge technologies in current video production. AIGC can quickly and cost-effectively generate imaginative scenes and visual effects, but it has limitations in the precise control of character movements and camera motion. Virtual filming, on the other hand, can realistically capture actors' performances and integrate them with virtual backgrounds in real time, but it relies on high-quality digital assets and has a long production cycle and high cost.
[0003] Existing methods of combining the two are mostly sequential, i.e., first generating the background, then filming the actors, and finally compositing in post-production. This has significant drawbacks: 1. Poor visual consistency, AIGC lighting is difficult to map onto the real environment, requiring a lot of manual lighting adjustment; 2. It is difficult to accurately replicate or achieve complex camera movements, resulting in limited cinematic language; 3. The production process is fragmented, relying on manual intervention, and is inefficient.
[0004] Therefore, how to intelligently integrate AIGC content with virtual shooting technology to improve the quality of video content has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of the above problems, the present invention provides a video generation method and electronic device that overcomes or at least partially solves the above problems, the technical solution of which is as follows:
[0006] A video generation method, comprising:
[0007] Obtain reference videos generated by artificial intelligence;
[0008] Extract visual art elements from the reference video, wherein the visual art elements include camera motion parameters;
[0009] Using the camera motion parameters, the target motion trajectory is generated;
[0010] The physical shooting system is used to control the target's motion trajectory to capture the foreground performance and generate a foreground image.
[0011] Determine whether the target's motion trajectory exceeds the physical capability limit of the physical shooting system. If so, execute the virtual-real collaborative shooting mode and calculate the motion difference between the target's motion trajectory and the actual motion trajectory of the physical shooting system.
[0012] The motion difference is used to drive the virtual camera to perform compensated motion, and a dynamic compensated background layer is generated in real time.
[0013] The foreground image, the dynamically compensated background layer, and the background image in the reference video are intelligently fused together to output a composite image.
[0014] Optionally, the visual art elements also include lighting information, and before capturing the foreground performance using the target motion trajectory control physical shooting system and generating the foreground image, the method further includes:
[0015] Using the illumination information, control commands for the physical lighting field are generated;
[0016] The control command is sent to the physical lighting field to drive the physical lighting field to illuminate the shooting environment of the physical shooting system.
[0017] Optionally, extracting visual art elements from the reference video includes:
[0018] Key visual regions in the reference video are identified, including areas directly illuminated by a light source and shadow areas; lighting information is fitted using a physical lighting model based on the areas directly illuminated by a light source and the shadow areas.
[0019] And / or, in the video sequence of the reference video, the pose parameters of the virtual camera in each frame are calculated by using feature point tracking and motion recovery structures to obtain the camera motion parameters.
[0020] Optionally, the physical shooting system includes a slide rail and an omnidirectional gimbal camera mounted on the slide rail. The step of using the target motion trajectory to control the physical shooting system to capture the foreground performance and generate a foreground image includes:
[0021] The target motion trajectory is decomposed and planned to obtain the translational motion command sequence of the slide rail and the rotational zoom motion command sequence of the camera device;
[0022] The physical shooting system is driven to perform coordinated movements according to the translational motion command sequence and the rotational zoom motion command sequence in order to capture the foreground performance;
[0023] The foreground performance is keyed to obtain the foreground image.
[0024] Optionally, the step of performing motion decomposition and planning on the target motion trajectory to obtain the translational motion command sequence of the slide rail and the rotational zoom motion command sequence of the camera device includes:
[0025] The camera pose of each frame in the target motion trajectory is decomposed into the translation transformation of the slide rail in the world coordinate system, and the rotation transformation and focal length transformation of the camera device in its own base coordinate system.
[0026] Based on minimizing the difference between the actual pose of the physical shooting system and the target pose, and satisfying the physical motion constraints of the slide rail and the camera device, a translational motion command sequence and a rotational zoom motion command sequence that are aligned in time are obtained.
[0027] Optionally, obtaining the time-aligned translational motion command sequence and rotational zoom motion command sequence includes:
[0028] Project the target position in the target motion trajectory onto the slide rail motion plane to obtain the desired position sequence of the slide rail;
[0029] Motion planning is performed on the desired position sequence to generate a translational motion command sequence that satisfies velocity and acceleration constraints;
[0030] As the slide rail moves sequentially to each desired position in the desired position sequence, the horizontal rotation angle, pitch rotation angle, and zoom value required by the camera device to align the camera optical axis with the scene target are determined, and a rotation zoom motion command sequence is generated.
[0031] Optionally, the step of using the motion difference to drive the virtual camera to perform compensated motion and rendering a dynamically compensated background layer in real time includes:
[0032] Establish a unified transformation relationship between the world coordinate system, the physical camera coordinate system, and the virtual camera coordinate system;
[0033] The pose parameters of the virtual camera are calculated in real time using the motion difference.
[0034] Based on the timestamp synchronized with the physical shooting frames, the virtual camera is driven to move according to the pose parameters, and a dynamically compensated background layer is rendered in real time.
[0035] Optionally, the step of intelligently fusing the foreground image, the dynamic compensation background layer, and the background image in the reference video to output a composite image includes:
[0036] The color statistical characteristics and light and shadow distribution information of the background image in the reference video were analyzed.
[0037] Using the aforementioned color statistical characteristics, a color transformation matrix and a color offset vector are calculated to perform color correction on the foreground image;
[0038] Using the light and shadow distribution information, add highlights and shadows to the foreground image that are consistent with the direction of the background light source;
[0039] The foreground image after dynamic compensation, color correction, and lighting rendering is merged with the background image in the reference video to output a composite image.
[0040] Optionally, the method further includes:
[0041] The encoder feedback of the slide rail and the actual angle and zoom feedback of the camera device are read in real time to obtain the feedback value;
[0042] The feedback value is compared with the corresponding instruction value to obtain the real-time tracking error;
[0043] Based on the real-time tracking error, a feedforward-feedback control algorithm is used to dynamically adjust the control commands sent to the physical imaging system.
[0044] An electronic device includes at least one processor, at least one memory connected to the processor, and a bus; wherein the processor and the memory communicate with each other via the bus; the processor is used to call program instructions in the memory to execute the video generation method.
[0045] By employing the above technical solution, this invention provides a video generation method and electronic device that obtains an AI-generated reference video; extracts visual art elements from the reference video, wherein the visual art elements include camera motion parameters; generates a target motion trajectory using the camera motion parameters; uses the target motion trajectory to control a physical shooting system to capture foreground performance and generate a foreground image; determines whether the target motion trajectory exceeds the physical capability limit of the physical shooting system; if so, executes a virtual-real collaborative shooting mode and calculates the motion difference between the target motion trajectory and the actual motion trajectory of the physical shooting system; uses the motion difference to drive a virtual camera to perform compensation motion and renders a dynamic compensation background layer in real time; and intelligently merges the foreground image, the dynamic compensation background layer, and the background image in the reference video to output a composite image. This invention extracts visual art elements from an AI-generated reference video, intelligently analyzes and transforms them into executable camera motion parameters and target motion trajectories, and combines a physical shooting system with a virtual compensation mechanism to achieve virtual-real collaborative shooting and dynamic intelligent synthesis. This effectively streamlines the deep integration process between AIGC content and virtual shooting, thereby improving the visual consistency and expressiveness of video content and significantly enhancing video content quality.
[0046] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0047] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0048] Figure 1 A flowchart illustrating one embodiment of the video generation method provided by this invention is shown.
[0049] Figure 2 A schematic diagram of the physical imaging system provided in an embodiment of the present invention is shown;
[0050] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present invention is shown. Detailed Implementation
[0051] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.
[0052] With the rapid development of artificial intelligence (AI) technology, AI-generated content (AIGC) and virtual shooting have become two cutting-edge technologies in the video production field. AIGC excels at generating creative and imaginative scenes and visual effects at extremely high speed and low cost, greatly enriching the possibilities of content production. However, AIGC still has certain limitations in controlling specific elements, especially in the precise control of character movements and camera motion, making it difficult to meet the stringent requirements for detail and expressiveness in high-quality film and television production. In contrast, virtual shooting technology can perfectly capture the performances of real actors and achieve real-time integration with virtual backgrounds, enhancing the realism and immersion of the images. However, the core of virtual shooting relies on pre-built high-quality digital assets, a process that is usually time-consuming and costly, limiting its application in rapid content production scenarios.
[0053] Currently, the industry's attempts to integrate AIGC (AI-generated content) with virtual shooting often employ sequential processing or offline replacement methods. For example, background content is first generated using AIGC, then actors are filmed against a green screen, and finally, the foreground and background are stitched together using post-production compositing software. However, this traditional workflow has several fundamental flaws: 1. Visual consistency is difficult to guarantee. The lighting characteristics of the scene generated by AIGC cannot be automatically mapped to the real shooting environment, resulting in inconsistent lighting effects between actors and the background. This heavily relies on manual lighting adjustments and post-production corrections, leading to a large workload and low efficiency. 2. Disjointed camera movement. Camera movements in AIGC videos are mostly based on preset trajectories, which are difficult for physical cameras to accurately replicate. Furthermore, certain complex or physically limit-breaking movements cannot be achieved, resulting in a monotonous camera language or requiring expensive and complex motion control equipment to compensate. 3. Fragmented production process. The current workflow lacks closed-loop connections between different stages, heavily relying on manual intervention. This hinders efficient automation and rapid iteration, limiting overall production efficiency.
[0054] Based on this, this invention provides a video generation method. It acquires an AI-generated reference video, extracts visual art elements including camera motion parameters, and generates a target motion trajectory based on these parameters. This trajectory is used to control a physical shooting system to capture foreground performance and generate a foreground image. If the target motion trajectory exceeds the capabilities of the physical system, a virtual-real collaborative shooting mode is activated. Motion differences are calculated, and a virtual camera is driven to perform compensating motion, rendering a dynamically compensated background layer in real time. The foreground image, the compensated background layer, and the reference video background are intelligently integrated to output a high-quality composite image. This achieves deep integration of AIGC content and virtual shooting, improving the video's visual consistency and expressiveness, and significantly enhancing content quality.
[0055] like Figure 1 The diagram shows a flowchart of one embodiment of the video generation method provided by this invention. The method may include:
[0056] S100, Obtain reference videos generated by artificial intelligence.
[0057] The reference video refers to a creative video generated by artificial intelligence, showcasing initial visual concepts and artistic expressions, serving as a digital director's script for the entire virtual shooting and compositing process.
[0058] Specifically, embodiments of the present invention can receive creative reference videos automatically generated by AIGC tools. These videos not only showcase scene content but also contain complete visual artistic intent, serving as a digital director's script for subsequent shooting and compositing.
[0059] As examples, embodiments of the present invention can receive creative reference videos generated by AIGC tools (such as Sora, Stable Video Diffusion, etc.) via API interfaces or file systems, and use them as a "digital director's script" for the entire production process. This video not only defines the scene content but also embeds a complete visual artistic intent, including camera movement, lighting atmosphere, color style, and composition design. The system loads it as the core processing data source, providing a unique visual reference standard for all subsequent automated processing steps.
[0060] S110. Extract visual art elements from the reference video, including camera motion parameters.
[0061] Visual art elements refer to key visual information in the video, including lighting atmosphere, composition design, color style, and camera movement language, which are used to guide physical shooting and compositing.
[0062] Among them, camera motion parameters refer to numerical data describing changes in camera position, rotation angle, and focal length, reflecting the dynamic changes of the camera in space.
[0063] Specifically, embodiments of the present invention can automatically identify and quantify key visual elements in a reference video, such as lighting atmosphere, composition, color style, and camera motion language, and in particular, use feature point tracking and Perspective-n-Point (PnP) algorithms to inversely calculate the motion parameters of the virtual camera.
[0064] As examples, embodiments of the present invention can extract and track salient feature points in video across frames; then, by solving the perspective n-point problem and combining the RANSAC robust algorithm, the extrinsic parameter matrix (position t and rotation R) of the virtual camera in the world coordinate system and the focal length f are calculated for each frame. For difficult frames with missing textures, users can manually mark feature points as strong constraints to improve accuracy. The original pose sequence is smoothed using Kalman filters to remove high-frequency jitter, thereby obtaining a smooth and continuous sequence of target camera motion parameters, accurately quantifying the camera movement language of the AIGC video.
[0065] S120. Using camera motion parameters, generate the target motion trajectory.
[0066] The target motion trajectory refers to the ideal camera motion path generated based on the camera motion parameters, serving as the motion target for both the physical shooting system and virtual compensation.
[0067] Specifically, in this embodiment of the invention, the camera parameter sequence can be filtered to generate a smooth target motion trajectory, providing a standard trajectory for the execution of the physical shooting system.
[0068] As examples, embodiments of the present invention can arrange the sequence of camera motion parameters sequentially on a time axis to form a complete, digitized target motion trajectory. This target motion trajectory accurately describes the motion path, attitude changes, and viewpoint scaling of the AIGC virtual camera in three-dimensional space, serving as the absolute benchmark for subsequently driving the physical system and calculating virtual compensation.
[0069] S130: The physical shooting system, which controls the target motion trajectory, captures the foreground performance and generates a foreground image.
[0070] Among them, the physical shooting system refers to the actual shooting equipment consisting of a camera motion device and a camera shooting device, which is used to capture the actors' performances and some camera movements.
[0071] In this context, foreground performance refers to the performance content of actors or real-life objects in front of a green screen, which serves as the main element of the video composite.
[0072] The foreground footage refers to the video footage captured by the physical shooting system that includes foreground performances.
[0073] Specifically, embodiments of the present invention can intelligently decompose motion commands according to the target motion trajectory and allocate them to the corresponding modules or devices of the physical shooting system. Through trajectory planning and synchronous control, the physical equipment is driven to capture the actors' performance in front of the green screen in real time and generate foreground images.
[0074] As examples, embodiments of the present invention can decompose the 6-DOF target motion trajectory into control commands for corresponding modules or devices of the physical shooting system through a constrained optimization problem. These commands are then sent to the corresponding modules or devices of the physical shooting system in a high-precision time-synchronized manner via a corresponding protocol. The corresponding modules or devices of the physical shooting system then coordinate their movements to film the actors in front of a green screen under a lighting environment automatically configured by the light sensing and mapping module according to the AIGC blueprint, capturing the original foreground image containing the performance in real time.
[0075] S140. Determine whether the target's motion trajectory exceeds the physical capability limit of the physical shooting system. If so, proceed to step S150.
[0076] Among them, the physical capability limit refers to the maximum load-bearing capacity limit of the physical shooting system in terms of range of motion, speed and acceleration.
[0077] Specifically, embodiments of the present invention can evaluate in real time whether the planned trajectory exceeds the device's range of motion, speed, or acceleration limits, or whether the predicted tracking error exceeds the visual tolerance threshold. If the limit is exceeded, it is determined that it cannot be fully realized physically, and the virtual-real collaborative shooting mode is entered.
[0078] As examples, later embodiments of the present invention can check whether the corresponding modules or devices of the decomposed physical shooting system exceed the corresponding physical constraints. If all physical constraints cannot be met, the target trajectory is determined to be "physically unrealizable", and a flag is triggered to switch from "full physical drive mode" to "virtual and real collaborative shooting mode".
[0079] S150, Execute the virtual and real collaborative shooting mode and calculate the motion difference between the target's motion trajectory and the actual motion trajectory of the physical shooting system.
[0080] Among them, the virtual-real collaborative shooting mode refers to a shooting method that combines physical shooting with virtual camera compensation to achieve complete motion performance when the target's motion trajectory exceeds the capabilities of physical equipment.
[0081] Among them, motion difference refers to the deviation between the target's motion trajectory and the actual motion trajectory of the physical shooting system, which is used to drive the virtual camera to make compensation.
[0082] Specifically, embodiments of the present invention can calculate the three-dimensional position difference, rotation difference, and focal length difference between the target motion trajectory and the actual motion trajectory of the physical shooting system, and use these as inputs for virtual camera motion compensation to ensure the complete expression of visual motion language.
[0083] As some examples, positional differences that need to be compensated by virtual cameras The calculation formula is:
[0084] ,
[0085] in, The reference is the camera position in the target motion trajectory corresponding to the t-th frame in the video; Let t be the actual camera position of the physical imaging system in frame t.
[0086] Rotation difference that needs to be compensated by virtual camera The calculation formula is:
[0087] ,
[0088] in, The reference is the rotation matrix corresponding to the target's motion trajectory in the t-th frame of the video. Let be the actual rotation matrix of the physical imaging system at frame t.
[0089] Focal length difference that needs to be compensated by virtual camera The calculation formula is:
[0090] ,
[0091] in, The focal length corresponding to the target's motion trajectory in the t-th frame of the reference video; Let be the actual focal length of the physical imaging system in frame t.
[0092] S160: Use motion difference to drive the virtual camera to perform compensated motion and render and generate a dynamic compensated background layer in real time.
[0093] Among them, the dynamic compensation background layer refers to the background image layer rendered in real time by the virtual camera based on the motion difference, which is used to compensate for the visual loss caused by physical shooting limitations.
[0094] Specifically, in this embodiment of the invention, a virtual camera can perform a compensation trajectory in Unreal Engine based on motion difference, dynamically render the background layer to achieve parallax changes, and the foreground captured by the physical camera is synchronized with the dynamic background in real time to ensure the spatial and temporal consistency of the composite image.
[0095] As examples, later embodiments of the invention can use motion difference as control commands to drive a virtual camera running in Unreal Engine to perform corresponding compensating motion: the actor foreground captured by the physical camera is relatively stationary (in the physical camera coordinate system), while the virtual camera moves according to the difference trajectory, causing its rendered background layer to produce a corresponding, inverse parallax change. Through precise coordinate transformation and millisecond-level time synchronization, it is ensured that each frame of compensating motion by the virtual camera is strictly aligned with the actual shooting frame of the physical camera. The virtual camera renders and outputs a dynamically changing compensated background layer synchronized with the physical motion in real time.
[0096] S170: Intelligently fuse the foreground image, the dynamic compensation background layer, and the background image in the reference video to output a composite image.
[0097] The background in the reference video refers to a virtual scene background generated by artificial intelligence, which displays the overall visual atmosphere and environmental information.
[0098] Among them, the composite image refers to the final output image after intelligent fusion of the foreground image, the dynamically compensated background layer and the reference video background, which has high-quality visual consistency and artistic expression.
[0099] Specifically, embodiments of the present invention can automatically separate the foreground through deep learning keying, combine intelligent color correction and light and shadow matching algorithms, and fuse physically captured foreground images, virtual camera movement compensation background layers and AIGC backgrounds to achieve seamless compositing and output high-quality composite images.
[0100] As examples, embodiments of the present invention can employ a deep learning keying model to perform high-quality keying of the foreground image, obtaining an alpha mask. Then, intelligent matching is performed: 1. Color and lighting matching: Analyzing the color statistics and lighting distribution of the original AIGC background image, automatically correcting the foreground color by calculating the color transformation matrix, and adding highlights / shadows consistent with the background light source direction. 2. Multi-layer fusion: Precisely aligning and overlaying the processed foreground, the original AIGC static background image, and the dynamically compensated background layer. If virtual camera movement compensation is enabled, the dynamically compensated background layer will be used as the final background; if fully physically driven, the AIGC static background is used directly. Finally, through edge optimization and color fusion, a composite image with cinematic quality that is visually highly consistent with the AIGC reference video is output and streamed in real time as the final video.
[0101] This invention provides a video generation method, which includes: obtaining an AI-generated reference video; extracting visual art elements from the reference video, wherein the visual art elements include camera motion parameters; generating a target motion trajectory using the camera motion parameters; controlling a physical shooting system to capture foreground performance using the target motion trajectory to generate a foreground image; determining whether the target motion trajectory exceeds the physical capability limit of the physical shooting system; if so, executing a virtual-real collaborative shooting mode and calculating the motion difference between the target motion trajectory and the actual motion trajectory of the physical shooting system; using the motion difference to drive a virtual camera to perform compensation motion and rendering a dynamic compensation background layer in real time; and intelligently fusing the foreground image, the dynamic compensation background layer, and the background image from the reference video to output a composite image. This invention extracts visual art elements from an AI-generated reference video, intelligently analyzes and transforms them into executable camera motion parameters and a target motion trajectory, and combines a physical shooting system with a virtual compensation mechanism to achieve virtual-real collaborative shooting and dynamic intelligent compositing. This effectively streamlines the deep integration process between AIGC content and virtual shooting, thereby improving the visual consistency and expressiveness of video content and significantly enhancing video content quality.
[0102] Optionally, in the above Figure 1 Based on one or more corresponding embodiments, in another optional embodiment provided by the present invention, the visual art elements further include lighting information, and before step S130, the method may further include:
[0103] Using illumination information, control commands for the physical lighting field are generated; these commands are then sent to the physical lighting field to drive it to illuminate the shooting environment of the physical shooting system.
[0104] Lighting information refers to detailed data about scene lighting extracted from AI-generated reference videos, including visual characteristics such as the direction, intensity, color, hardness, and shadow distribution of light sources. This information reflects the lighting atmosphere and effects in the video and is an important basis for guiding the physical reproduction of lighting.
[0105] The physical lighting field refers to the actual lighting system consisting of various programmable lights (such as LED lights, spotlights, and soft lights) installed on the shooting location. The physical lighting field adjusts the position, brightness, color temperature, and projection angle of the light sources according to control commands, simulating the ambient lighting effects consistent with those in the reference video in real time, thereby creating a realistic shooting atmosphere.
[0106] This invention enables intelligent visual feature extraction from reference videos, focusing on identifying key lighting features such as direct light sources (highlights), shadow areas, and the softness, direction, and contrast of light in the scene. If precise depth information is unavailable, a preset normal map or geometric assumption is selected based on the scene style (e.g., cartoon or realistic), and combined with brightness gradient field calculations, the direction of the main light source and the intensity and color of the ambient light are estimated, achieving accurate quantification of the overall lighting atmosphere.
[0107] This invention can employ a simplified physical lighting model (e.g., the Phong model) to fit the extracted lighting information, obtaining the main light source direction vector, ambient light intensity, and color temperature parameters. Using a pre-calibrated mapping function, the highlight areas and overall brightness and color information in the video are converted into the absolute brightness and color temperature values required by the physical light source, ensuring a high degree of consistency between the physical light source output and the lighting effect of the reference video, thus avoiding visual deviations caused by simple linear mapping.
[0108] This invention can automatically match the type and location of physical lighting fixtures around the camera by combining the fitted lighting parameters, and convert the lighting parameters into specific lighting control data packets through a built-in cross-device control protocol library. These data packets contain instructions such as lighting fixture on / off, brightness, color temperature, color, and projection angle, forming an executable lighting control command sequence. The lighting control command sequence is sent in real time to each lighting control unit in the physical lighting field via a network or wired communication channel, achieving precise driving of the lighting equipment. This process is automated and requires no manual intervention. The physical lighting fixtures dynamically adjust their light intensity, color temperature, and direction according to the instructions, quickly reproducing the lighting atmosphere set in the reference video, ensuring a high degree of synchronization between the physical shooting environment's lighting and the virtual visual blueprint.
[0109] This invention incorporates lighting information and generates and issues physical lighting field control commands accordingly, thereby achieving precise lighting of the shooting environment. This can highly reproduce the light and shadow atmosphere in the reference video, enhance the realism and visual consistency of the foreground performance, and thus improve the artistic expression and immersiveness of the final composite image.
[0110] Optionally, in the above Figure 1 Based on one or more corresponding embodiments, in another optional embodiment provided by the present invention, step S110 may specifically include:
[0111] Key visual regions in the reference video are identified, including areas directly illuminated by the light source and shadowed areas.
[0112] The area directly illuminated by the light source, also known as the highlight area, refers to the bright area in an image or picture formed by the direct illumination of the light source onto the surface of the object and the reflection. It is the part that directly reflects the light source.
[0113] The shadow area refers to the area in an image or picture that cannot be directly illuminated by the light emitted by the light source because it is blocked by the foreground object; that is, the area that cannot be seen from the perspective of the light source.
[0114] Specifically, embodiments of the present invention can detect areas in an image with significantly higher brightness than the surrounding environment as directly illuminated areas by a light source through brightness threshold segmentation and local contrast analysis; simultaneously, areas with extremely low brightness and smooth edge gradients are used as shadow areas. To enhance the accuracy of recognition, texture features, color distribution, and morphological operations can be combined to remove noise and non-illumination-related bright or dark areas. Furthermore, temporal information is used to track the illumination change trend in consecutive frames to ensure the stability and consistency of key visual regions.
[0115] Illumination information is obtained by fitting a physical lighting model based on the areas directly illuminated by the light source and the shadow areas.
[0116] The lighting information includes the direction of the main light source, the intensity of the light source, and color parameters.
[0117] Specifically, embodiments of the present invention can fit the lighting characteristics of these key areas based on classic physical lighting models (such as the Phong model or the Lambertian reflection model): using the brightness gradient information of the area directly illuminated by the light source in the image and the assumed scene surface normals (obtained through a preset normal map or geometric inference), the direction vector of the main light source is solved. Combining the average brightness of the area directly illuminated by the light source and the ambient brightness of the overall image, the intensity value and color temperature of the light source are estimated through a pre-calibrated mapping function. The RGB color parameters of the light source are inferred based on the average color value of the area directly illuminated by the light source. This fitting process is solved using optimization algorithms (such as the least squares method), so that the physical lighting model, while maintaining simplicity, approximates the lighting and shadow effects presented in the reference video to the greatest extent. The final output lighting information provides a precise physical quantitative basis for subsequent physical lighting field setup and lighting and shadow consistency synthesis.
[0118] As examples, embodiments of the present invention can identify key lighting regions (such as highlight areas and shadow areas) in video frames of a reference video, and preliminarily determine the softness, hardness, direction, and contrast of the light source. Then, it proceeds to a physically-based lighting model fitting stage: assuming the scene satisfies the Lambertian reflection model, it solves an optimization problem: ,in, Indicates the image at pixel points Brightness gradient at that location Represents pixels The surface normal vector at that location, This represents the normalized main light source direction vector to be determined. The normalized main light source direction vector is calculated using the image brightness gradient field and estimated surface normals (a preset normal map can be selected based on the AIGC style). Simultaneously, based on the average brightness and average color of the area directly illuminated by the light source and the overall image, the required light source intensity for the physical lighting field is calculated using a pre-calibrated, non-linear mapping function. and color mapping The calculation formula is:
[0119] ,
[0120] ,
[0121] in, This represents the average brightness of the area in the image directly illuminated by the light source. This is expressed as the overall average brightness of the image; It is represented as the average color (RGB value) of the area directly illuminated by the light source. and It is represented as a pre-calibrated mapping function, which establishes a correspondence between the AIGC image brightness value and the physical lamp field absolute brightness / color temperature value, avoiding the effect distortion caused by simple linear mapping.
[0122] This invention can utilize the driver protocol library (such as DMX512, Art-Net, and sACN) of common lighting equipment (such as smart LED lights, DMX dimming cabinets, and LED walls) to calculate quantitative illumination parameters. , and The system automatically converts the preset three-dimensional position, adjustable angle, and channel definition of each luminaire in the physical lighting field into standard data packets corresponding to the control protocol. These data packets are sent in real time to the controller of the physical lighting field via a network or dedicated line. The controller drives the luminaires to adjust their direction, brightness, and color, thereby accurately reproducing the light and shadow atmosphere in the reference video on the green screen shooting site, providing a lighting environment for the foreground performance that is completely consistent with the target's visual perception, and realizing "one-click automated lighting".
[0123] This invention identifies key visual areas of direct illumination and shadow in a reference video and accurately fits the direction, intensity, and color parameters of the main light source based on a physical lighting model. This enables the quantitative extraction of lighting information, providing a scientific basis for precise lighting of physical lighting fields and significantly improving the light and shadow reproduction of the shooting environment and the visual consistency and immersiveness of the final composite image.
[0124] Optionally, in the above Figure 1 Based on one or more corresponding embodiments, in another optional embodiment provided by the present invention, step S110 may specifically include:
[0125] In the video sequence of the reference video, the pose parameters of the virtual camera in each frame are obtained by using feature point tracking and motion reconstruction structure, thus obtaining the camera motion parameters.
[0126] Specifically, in this embodiment of the invention, significant feature points in each frame of a reference video sequence are automatically extracted using a feature point detection algorithm. Feature matching technology is then used to track feature points between adjacent frames, forming continuous feature trajectories. Combined with Structure from Motion (SfM) technology (or a rough 3D geometry, such as a cube or plane, specified by the user), multi-view geometry and triangulation are used to sparsely reconstruct the 3D spatial positions corresponding to the matched 2D feature points, reconstructing the sparse point cloud of the scene and the camera's position and pose in the world coordinate system. Then, based on known or estimated camera intrinsic parameters, a perspective n-point algorithm is used to solve for the camera extrinsic parameter matrix (rotation matrix and translation vector) in each frame to obtain the virtual camera pose for that frame. To ensure robustness of the solution, the RANSAC algorithm can be introduced to eliminate incorrect matches. Furthermore, user interaction allows for manual labeling of feature points and their 3D positions in keyframes, further improving the accuracy and stability of pose estimation. The estimated camera pose sequence can be smoothed by applying Kalman filtering or Savitzky-Golay filtering to reduce jitter and noise, and obtain coherent and smooth camera motion parameters, providing a high-quality data foundation for subsequent physical motion decomposition and virtual compensation.
[0127] As examples, embodiments of the present invention can process a reference video sequence frame by frame, first extracting and tracking salient feature points across frames. Using structure-of-motion (SOG) techniques, the coordinates of these feature points in the 3D world coordinate system are sparsely reconstructed from their 2D trajectories. Then, for each frame, a perspective n-point problem is solved: using the 2D pixel coordinates of multiple feature points in that frame and their known 3D world coordinates, combined with the camera intrinsic matrix obtained or estimated from the AIGC generation settings, the extrinsic parameter matrix describing the transformation from the world coordinate system to the camera coordinate system of that frame, i.e., the pose parameters, includes a 3x3 rotation matrix and translation vector. The calculation formula is:
[0128] ;
[0129] in, Indicates the first The 2D pixel coordinates of each feature point in the current frame image; Indicates the first The known or reconstructed 3D world coordinates corresponding to each feature point; This represents the camera intrinsic parameter matrix, including focal length and principal point parameters. It can be obtained from the AIGC generation settings or estimated using calibration. Let represent the 3x3 rotation matrix and 3x1 translation vector to be determined, describing the transformation from the world coordinate system to the camera coordinate system; This represents the scaling factor.
[0130] To improve robustness, embodiments of this invention can employ the RANSAC framework combined with efficient algorithms (such as EPNP) to minimize reprojection errors and eliminate the influence of outlier matching points. The calculation formula is as follows:
[0131] ,
[0132] in, For projection functions; This represents the set of interior points selected through RANSAC iteration. This ensures that a stable solution can still be obtained even with outlier matching points.
[0133] For challenging frames with missing textures or motion blur, this invention supports manual labeling assistance, using user-specified feature points as strong constraints in optimization, and outputting a smoothed camera pose sequence for each frame. ,in, Indicates the first The three-dimensional position vector of the frame virtual camera; Indicates the first The rotation matrix of the frame virtual camera; Indicates the first Focal length parameters of the frame virtual camera.
[0134] This invention employs feature point tracking and motion recovery structures in the video sequence of a reference video to inversely calculate the pose parameters of the virtual camera in each frame. This allows for the accurate reconstruction and quantification of the real camera motion in the reference video, providing a reliable basis for high-fidelity motion control and virtual compensation in the subsequent physical shooting system. Consequently, it significantly improves the spatial consistency and immersiveness of the synthesized image.
[0135] Optionally, in the above Figure 1 Based on one or more corresponding embodiments, in another optional embodiment provided by the present invention, the physical shooting system provided by the present invention may include a slide rail and a camera device mounted on the slide rail.
[0136] A slide rail is a mechanical guiding system used to support camera equipment in linear or curved motion along a preset track. A slide rail can consist of a track body, a sliding platform, and a drive system. It uses a motor to drive the precise translational movement of the camera equipment and can be used with a motion control system to achieve programmable composite shooting trajectories. Types of slide rails include linear slide rails, circular slide rails, and multi-axis slide rails, with different ranges and modes of motion selectable according to shooting requirements.
[0137] Specifically, the camera device can be an omnidirectional pan-tilt camera capable of both left-right and up-down rotation, enabling wide-area scanning coverage, such as a PTZ (Pan-Tilt-Zoom) camera. The camera device provided in this embodiment is mounted on a slide rail and moves in coordination with the slide rail to capture foreground performance, achieving integrated control of translation, rotation, and zoom.
[0138] As some examples: Reference Figure 2 The physical shooting system provided in this embodiment of the invention can specifically consist of a PTZ pan-tilt camera, a vertical slide rail, a vertical servo motor, a horizontal slide rail, and a horizontal servo motor.
[0139] Optionally, in the above Figure 1 Based on one or more corresponding embodiments, another optional embodiment provided by the present invention may specifically include step S130 as follows:
[0140] The motion trajectory of the target is decomposed and planned to obtain the translational motion command sequence of the slide rail and the rotation and zoom motion command sequence of the camera device.
[0141] Specifically, this embodiment of the invention can decouple the target motion trajectory in the world coordinate system based on the smooth and high-precision 6-DOF target motion trajectory of the virtual camera obtained from the reference video. This decomposes the trajectory into two-dimensional linear translation motion handled by the sliding rail system and three-DOF rotation (horizontal rotation Pan and vertical rotation Tilt) and zoom motion handled by the PTZ pan-tilt camera. The X and Z axis components of the target camera position are mapped to the motion commands of the sliding rail. A trajectory planning algorithm (such as trapezoidal velocity planning or S-curve planning) is used to generate a sliding rail position sequence that conforms to the speed and acceleration limits of the physical equipment. Simultaneously, based on the expected position of the sliding rail, the required rotation angle and zoom parameters of the PTZ pan-tilt camera are calculated. By solving the inverse kinematics problem, the camera's optical axis is always aligned with the predetermined target point, ensuring that the image composition is consistent with the reference video settings. This process incorporates the physical limits of the equipment; if the trajectory exceeds the limits, virtual compensation space is reserved. Finally, the sliding rail translation trajectory and the PTZ pan-tilt camera rotation and zoom trajectory are transformed into a specific control command sequence, preparing for subsequent hardware execution.
[0142] As examples, embodiments of the present invention can use the target motion trajectory as input and perform cooperative trajectory planning through a constrained optimization problem. The goal is to minimize the difference between the actual reachable pose and the target pose. The process employs a simplified solution strategy: the target position is projected onto the two motion axes of the slide rail to obtain the target position sequence of the slide rail, and a smooth and feasible sequence of slide rail translation motion commands is generated by a trajectory planner that considers rate and acceleration limitations (such as S-curve planning). Assuming that the slide rail is accurately positioned, the sequence of PTZ pan-tilt camera rotation and zoom motion commands required to align the camera optical axis with the scene target and match the object size in the reference video is calculated by solving the inverse kinematics problem. Simultaneously, a physical reachability judgment is performed; if the planned motion exceeds the physical limits of the device or the prediction error is too large, it is marked in advance and virtual compensation is activated.
[0143] The physical shooting system is driven to move in coordination with a sequence of translational motion commands and a sequence of rotational zoom commands to capture the foreground performance.
[0144] Specifically, in this embodiment of the invention, a planned sequence of translational motion commands can be sent to the slide rail servo controller, driving it to move precisely along a predetermined trajectory. Simultaneously, the camera device's controller receives a sequence of rotation and zoom motion commands and adjusts the camera device's horizontal rotation angle, pitch angle, and lens focal length. To ensure time synchronization between the slide rail and the camera device's movements, a unified timestamp and interpolation mechanism is used to simultaneously update both commands within a high-frequency master control cycle. Encoder and sensor data are monitored in real-time and feedback is provided in a closed loop. PID feedback control is used to correct tracking errors. During physical filming, actors perform in front of a green screen according to the director's instructions, while the physical camera continuously films along a coordinated motion trajectory, ensuring that the captured foreground image is highly consistent with the camera motion language of the reference video.
[0145] As examples, embodiments of the present invention send translational motion command sequences to the slide rail servo motor controller via protocols such as Modbus TCP or EtherCAT, and rotational zoom motion command sequences to the PTZ pan-tilt camera via VISCA or ONVIF PTZ protocols. A unified high-frequency master control clock (e.g., 100Hz) is used to interpolate the command trajectory in each control cycle, and the interpolated commands are simultaneously sent to the controllers of both the slide rail and the PTZ pan-tilt camera. The angle / zoom feedback from the slide rail encoder and the PTZ pan-tilt camera is read in real time to form a closed-loop control, and tracking errors are corrected through feedforward-feedback. When actors perform in front of a precisely lit green screen, the slide rail and the PTZ pan-tilt camera move in coordination in this manner, capturing the foreground performance footage in real time that is consistent with the perspective of the AIGC camera.
[0146] The foreground performance is keyed to obtain the foreground image.
[0147] Specifically, in this embodiment of the invention, a continuous green screen video stream of a foreground performance, captured by camera, can be input into an AI-driven keying module. This module employs a deep learning semantic segmentation model (such as a network structure based on U-Net or DeepLab) to automatically and accurately separate the actor from the green screen background, generating a high-quality transparent foreground image. The keying algorithm combines background information for edge refinement, reducing feathering and color overflow, and improving the natural transition effect of edges. Furthermore, based on the lighting conditions and color characteristics of the reference video background, the color and lighting of the foreground image are automatically adjusted, ensuring that the keying result matches the target environment in terms of hue and brightness, enhancing the immersiveness and realism of subsequent compositing. The output is a foreground image with an alpha channel, providing a material basis for subsequent intelligent multi-layer fusion.
[0148] As examples, embodiments of the present invention can receive raw footage containing actors and a green screen captured by a physical shooting system. First, a deep learning chroma keying model (such as a segmentation-based model) is used to automatically and accurately separate the actor's foreground from the green background. Unlike traditional chroma keying, embodiments of the present invention utilize lighting information (such as whether it is backlit) parsed from the reference video background to intelligently optimize the foreground edges. For example, if the background is backlit, the edges of the foreground actor are automatically subjected to appropriate semi-transparency and brightness enhancement to simulate the "overflow" or "transparency" effect of real light on the character's outline, thereby obtaining a more natural foreground image with a higher degree of integration with the background lighting atmosphere, preparing for subsequent intelligent compositing.
[0149] This invention, through precise motion decomposition and planning based on the target motion trajectory, generates coordinated motion commands for the sliding rail and camera equipment, and drives the physical shooting system to execute synchronously. This achieves high-precision capture and intelligent keying of the foreground performance, significantly improving the seamless integration of physical shooting and virtual creativity, and effectively ensuring the consistency of the foreground image with the artistic intent of the reference video and the natural fusion of visual effects.
[0150] Optionally, in the above Figure 1 Based on one or more corresponding embodiments, another optional embodiment provided by the present invention involves performing motion decomposition and planning on the target motion trajectory to obtain a translational motion command sequence for the slide rail and a rotational zoom motion command sequence for the camera device. Specifically, this may include:
[0151] The camera pose of each frame in the target's motion trajectory is decomposed into the translation transformation of the slide rail in the world coordinate system, and the rotation transformation and focal length transformation of the camera device in its own base coordinate system.
[0152] Specifically, embodiments of the present invention can acquire complete camera pose information for each frame of the target motion trajectory, including position vectors, translation matrices, rotation matrices, and focal length parameters. Based on the slider-PTZ physical structure model, the camera pose is decoupled into two parts in the world coordinate system: the two-dimensional or three-dimensional translation transformation of the slider base, mainly corresponding to the linear movement of the camera in the world coordinate system, especially the X and Z axis displacements along the track direction; and the rotation transformation of the camera device in the base coordinate system (including horizontal rotation Pan and vertical rotation Tilt) and lens focal length transformation (Zoom). This process is achieved through coordinate transformation matrix decomposition and inverse kinematics calculation, ensuring that the motion parameters of the slider and camera device can be accurately mapped to the physical device control variables, laying the foundation for cooperative control.
[0153] As some examples, embodiments of the present invention can process smoothed target camera pose sequences. Perform mathematical modeling. Let the pose of the k-th frame be... The actual pose that the PTZ pan-tilt camera mounted on the slide rail can achieve. Defined as the composition of two transformations, the formula is:
[0154] ,
[0155] in, Indicates the position reached by the slide rail. The transformation from the world coordinate system to the gimbal base coordinate system is defined when (the displacement vectors are along the two axes of the track). This indicates that the PTZ camera is rotating horizontally within its own base coordinate system. Pitch and rotation and zoom The subsequent transformation.
[0156] Based on minimizing the difference between the actual pose of the physical shooting system and the target pose, and satisfying the physical motion constraints of the slide rail and the camera device, a translational motion command sequence and a rotational zoom motion command sequence that are aligned in time are obtained.
[0157] Specifically, embodiments of the present invention can construct a nonlinear optimization model based on the decomposed motion parameters, using the slide rail position, camera rotation, and focal length as variables to be optimized. The objective function is defined as the error measure between the actual pose of the physical system and the pose of the target camera, which can be achieved by weighting the rotation matrix and the position vector using the Frobenius norm to balance position and rotation errors. Constraints include the maximum displacement range, velocity, and acceleration limits of the slide rail and camera, as well as the physical limits of the mechanical structure, ensuring that the planned trajectory is within the executable range of the equipment. This optimization problem can also consider temporal continuity and motion smoothness to ensure the temporal coordination of the trajectory, forming a spatiotemporal co-optimization problem, providing executable instructions for subsequent precise drive.
[0158] As examples, the objective of the cooperative trajectory planning optimization problem provided in this embodiment of the invention is to minimize the difference between the actual pose and the target pose. This can be formalized as an optimization problem of solving for the slide rail position for each frame (or keyframe). PTZ camera parameters The formula is:
[0159] ,
[0160] Simultaneously satisfying physical constraints: 1. Slide rail travel limit: 2. Slide rail speed limit: 3. Horizontal angle limit of the gimbal: 4. Other PTZ limits and speed limits.
[0161] in, These are the weighting coefficients that balance position and rotation errors; F represents the Frobenius norm of the matrix. Minimum slide rail travel limit; Maximum slide rail travel limit; Minimum rail speed limit; Maximum rail speed limit; The time interval between two consecutive time points; Minimum horizontal angle limit for the gimbal; This is the maximum horizontal angle limit for the gimbal.
[0162] Specifically, embodiments of the present invention can employ numerical optimization algorithms (such as Sequential Quadratic Programming (SQP), Model Predictive Control (MPC), or Gradient Projection Method) to solve the constructed cooperative trajectory planning problem online. During the solution process, device feedback and dynamic constraints are considered in real time to ensure the feasibility and real-time performance of the optimization results. The output includes a sequence of slider translation position commands strictly aligned with the target trajectory timestamp and a sequence of camera rotation and zoom parameters. These command sequences, after interpolation and filtering, possess smooth velocity and acceleration curves, enabling synchronous driving of the slider and camera, achieving high-precision physical reproduction of the target trajectory, and providing a solid foundation for subsequent shooting and virtual compensation.
[0163] This invention decomposes the camera pose in the target motion trajectory into the translation transformation of the sliding rail and the rotation and zoom transformation of the gimbal, and combines physical motion constraints to construct a collaborative trajectory optimization model for solving, thereby obtaining a precise and time-synchronized sequence of sliding rail and camera device motion commands. This enables the physical shooting system to reproduce the target trajectory with high precision, significantly improving the motion consistency and shooting quality of foreground capture.
[0164] Optionally, in the above Figure 1 Based on one or more corresponding embodiments, in another optional embodiment provided by the present invention, obtaining a translational motion command sequence and a rotational zoom motion command sequence aligned in time may specifically include:
[0165] Project the target position in the target motion trajectory onto the slide rail motion plane to obtain the desired position sequence of the slide rail; perform motion planning on the desired position sequence to generate a translational motion command sequence that satisfies the speed and acceleration constraints; when the slide rail moves to each desired position in the desired position sequence in sequence, determine the horizontal rotation angle, pitch rotation angle and zoom value required by the camera device to align the camera optical axis with the scene target, and generate a rotation zoom motion command sequence.
[0166] Specifically, the present invention can be used to target the location The target position of the slide rail is obtained by projecting the main projection onto the direction of movement of the slide rail (e.g., the X-axis and Z-axis). Then, through a trajectory planner with speed and acceleration limitations (such as trapezoidal velocity planning or S-curve planning), smooth and feasible guide rail position commands are generated. Assuming the slide rail will move precisely to... Under the premise of [specific conditions], calculate the camera equipment parameters required to align the camera's optical axis with the target. This can be accomplished by solving an inverse kinematics problem. Given the target point's coordinates in the world coordinate system and the expected position of the slide rail, calculate the rotation required for the PTZ pan-tilt camera. and To center the target point in the image. Zoom. The desired size of the target in the image (consistent with the object's apparent size in the reference video) is then determined. This involves the planned sequence of slide rail positions. PTZ camera parameter sequence These commands are then converted into specific motor control instructions. In this embodiment of the invention, servo motors can be controlled via Modbus TCP, EtherCAT, or pulse commands, with instructions including target position, speed, and acceleration. PTZ cameras can be controlled via VISCA, PELCO-D / P, or ONVIF PTZ protocols, with instructions including absolute or relative motion angles and zoom values. Finally, through a unified high-precision timestamp and interpolation mechanism, all device clocks are synchronized using Network Time Protocol (NTP) or Precision Time Protocol (PTP), ensuring that these two sets of instruction sequences are strictly aligned in time, forming a synchronized instruction stream that can directly drive the coordinated movement of physical devices.
[0167] This invention, through projecting the target position in the target motion trajectory onto the sliding rail motion plane and performing motion planning under velocity and acceleration constraints, and simultaneously utilizing inverse kinematics to calculate the rotation and zoom commands of the camera device, achieves high-precision, time-aligned reproduction of the target motion trajectory under the coordination of the sliding rail and the pan-tilt unit, thereby effectively improving the motion tracking accuracy of the physical shooting system and the shooting consistency of the foreground performance.
[0168] Optionally, in the above Figure 1 Based on one or more corresponding embodiments, in another optional embodiment provided by the present invention, step S160 may specifically include:
[0169] Establish a unified transformation relationship between the world coordinate system, the physical camera coordinate system, and the virtual camera coordinate system.
[0170] Specifically, in this invention, a unified spatial reference frame, namely the world coordinate system, can be defined, and the base coordinate system of the physical shooting system and the coordinate system of the virtual camera can be clearly defined. For each frame, the target motion trajectory and the actual motion trajectory of the physical shooting system are used to represent its position, rotation, focal length, and other parameters in the world coordinate system, respectively. Through matrix operations (such as Euclidean transformation matrices or quaternion representation), the motion trajectories of the physical camera coordinate system and the virtual camera coordinate system are uniformly mapped to the world coordinate system, achieving consistency in spatial position and orientation. This unified transformation relationship provides a benchmark reference for subsequent motion difference calculation and virtual-real collaborative control.
[0171] The pose parameters of the virtual camera are calculated in real time using motion difference.
[0172] Specifically, in this embodiment of the invention, the motion difference of each frame can be used as the driving signal for virtual camera motion compensation. An interpolation smoothing algorithm is used to filter the difference trajectory to eliminate abrupt changes and discontinuities, ensuring that the virtual camera motion is natural and smooth, generating the virtual camera pose corresponding to each frame, and strictly aligning it with the physical shooting system's actions.
[0173] Based on the timestamp synchronized with the physical shooting frames, the virtual camera is driven to move according to the pose parameters, and a dynamically compensated background layer is rendered in real time.
[0174] Specifically, embodiments of the present invention can maintain a unified timestamp queue to ensure that the virtual compensation motion of each frame is strictly aligned with the physical shooting frame. For each frame, the Unreal Engine interface is invoked to pass the real-time calculated virtual camera pose parameters to the camera entity in the virtual scene, driving it to perform spatial motion and lens transformation. Based on the compensation motion trajectory, the corresponding background image is rendered in real time and output as a dynamic compensation background layer.
[0175] For ease of understanding, an example is provided here: The calculation formula for coordinate transformation provided in this embodiment of the invention can be:
[0176] ,
[0177] in, This is the position compensation vector calculated in a unified world coordinate system; The transformation matrix represents the rotational transformation from the world coordinate system to the local coordinate system of the virtual camera in Unreal Engine; This is the position compensation command in the Unreal Engine virtual camera's local coordinate system after conversion; This is the difference function used for smooth transitions.
[0178] To prevent abrupt changes in the compensation motion, the calculated difference trajectory is smoothed using the following formula:
[0179] ,
[0180] in, This represents the original motion difference calculated at the current time t; The smoothing coefficient is dynamically adjusted based on the velocity and acceleration of the motion. This represents the smoothed motion difference at the previous time t-1; This represents the final motion difference after smoothing at the current time t.
[0181] The compensation motion provided in this invention is particularly suitable for achieving complex camera movements such as large-amplitude push-pull shots (e.g., the propulsion effect when the physical slider travel is insufficient), compound curve motion (e.g., complex camera movements that simultaneously rotate and move in arcs), rapid vibration and shaking (e.g., special camera effects that exceed the response speed of a PTZ pan-tilt camera), and extreme perspective changes (e.g., impact shots that require instantaneous changes in focal length and position). Through intelligent compensation technology, complex camera language that can only be accomplished by traditionally expensive motion control systems is realized at extremely low cost, greatly expanding the possibilities for creative expression.
[0182] This invention establishes a unified coordinate system and uses motion difference to calculate the virtual camera pose in real time, synchronously driving the virtual camera to render a dynamically compensated background layer. This achieves high-precision fusion of physical and virtual lens motion, effectively compensating for the limitations of the physical shooting system's motion capabilities, and ensuring accurate reproduction of the final composite image in terms of spatial consistency and artistic expression.
[0183] Optionally, in the above Figure 1 Based on one or more corresponding embodiments, in another optional embodiment provided by the present invention, step S170 may specifically include:
[0184] The color statistics and light and shadow distribution information of the background image in the reference video were analyzed.
[0185] Specifically, embodiments of the present invention can perform pixel-level analysis on the background of a reference video, calculate its color histogram, color mean, and covariance matrix, and quantify the overall color distribution characteristics; at the same time, by extracting the brightness channel and combining it with image gradient and other methods, the light and shadow distribution pattern, lighting direction, and shadow areas of the background are evaluated to form an accurate description of the lighting atmosphere of the background environment.
[0186] As examples, embodiments of the present invention can calculate global statistics of the background image in a specific color space (such as Lab or RGB) for color statistical characteristics, including the color mean vector and color covariance matrix of all pixels. These statistics describe the overall hue, saturation, and brightness distribution of the background. For light and shadow distribution information, by analyzing the brightness channels of the background image, algorithms (such as brightness gradient analysis) are used to estimate the direction, intensity, and contrast of the main light source. These analysis results collectively constitute the visual context of the background, providing a precise and quantifiable data foundation for subsequent intelligent adaptation of the foreground.
[0187] By utilizing the statistical properties of color, the color transformation matrix and color offset vector are calculated to perform color correction on the foreground image.
[0188] Specifically, in this embodiment of the invention, a 3×3 color transformation matrix and color offset vector can be calculated using an affine transformation method based on the mean and covariance of the background color, so that the color of the foreground pixel matches the background in statistical characteristics, thereby automatically adjusting the hue, saturation and brightness of the foreground image to achieve color harmony and unity between the two.
[0189] As examples, the color statistics of the foreground image are aligned with corresponding statistics obtained from the background analysis. Based on the formula:
[0190] ,
[0191] in, The original color vector of the foreground pixel; The 3x3 color transformation matrix is calculated by matching the color mean and covariance matrix of the foreground and background. It is a color offset vector; This is the corrected foreground color. This process achieves global color transfer, ensuring that the foreground's hue, saturation, and brightness distribution are consistent with the influence of ambient light, avoiding the visual disjointedness caused by simple color overlay.
[0192] By utilizing the information on light and shadow distribution, highlights and shadows are added to the foreground image in the same direction as the background light source.
[0193] Specifically, embodiments of the present invention can analyze the direction and intensity of light sources in the background to dynamically generate corresponding highlight and shadow effects in the foreground image, and use a lighting model to render the light and shadow on the actor's surface, thereby enhancing the consistency of the lighting environment between the foreground and the background and improving the realism and sense of depth of the composite image.
[0194] As examples, embodiments of the present invention can perform physically consistent lighting and shadow rendering on a keyed foreground actor image based on the direction, intensity, and softness / hardness of the main light source estimated from background analysis. First, based on the actor's rough geometric model or normal map, and combined with the light source direction, the illumination intensity that each pixel should receive is calculated. Then, based on the original foreground image, simulated highlights (enhancing brightness and saturation in areas facing the light source) and soft shadows (reducing brightness and blending ambient light colors in backlit areas) are dynamically added. Simultaneously, the edges of foreground objects are specially treated, for example, by appropriately semi-transparentizing and brightening the edges against a backlit background to simulate realistic light scattering (subsurface scattering) effects, making the foreground figure appear truly "immersed" in the background's lighting environment.
[0195] The foreground image after dynamic compensation, color correction, and lighting rendering is combined with the background image in the reference video to output a composite image.
[0196] Specifically, embodiments of the present invention can employ multi-layer compositing technology to fuse a dynamically compensated background layer rendered by a virtual camera, a foreground image that has undergone color correction and lighting processing, and the original reference video background, so that the three are visually seamlessly connected, outputting a high-quality composite image with cinematic visual effects.
[0197] As examples, embodiments of the present invention can receive three input layers: 1. the original / static background image from the reference video; 2. a dynamically compensated background layer rendered in real time by Unreal Engine (this layer contains parallax changes caused by virtual camera movement when virtual camera movement compensation is enabled); 3. a foreground actor image (with an alpha channel) after intelligent color correction and lighting rendering. The blending process employs alpha blending technology. First, the dynamically compensated background layer is overlaid or blended with the original background layer (the specific method depends on the compensation mode). Then, the processed foreground layer is precisely composited onto the blended background based on its alpha channel (which defines the opacity), ensuring that the edges of the foreground are anti-aliased and seamlessly blended with the lighting and color of the background. Finally, a composite image with cinematic visual effects is generated and output in real time, fully reproducing the artistic intent of the AIGC reference video, including complex camera movements, consistent lighting, and harmonious color tones.
[0198] This invention, through intelligent analysis of the color statistics and light and shadow distribution of the reference video background, performs precise color correction and light and shadow rendering on the foreground image, and combines dynamic compensation of the background layer to achieve multi-layer seamless alpha fusion, significantly improving the color consistency and lighting naturalness of the synthesized image, thereby effectively enhancing the visual realism and artistic expression of the virtual-real fusion.
[0199] Optionally, in the above Figure 1 Based on one or more corresponding embodiments, in another optional embodiment provided by the present invention, the method may further include:
[0200] The encoder feedback of the slide rail and the actual angle and zoom feedback of the camera device are read in real time to obtain the feedback value.
[0201] Specifically, in this embodiment of the invention, the current position and speed information can be read in real time from the encoder of the slide rail servo motor through a high-speed data acquisition interface. At the same time, the current horizontal rotation angle, pitch angle and zoom value can be obtained through the built-in sensor of the camera device or an external position feedback module. These feedback data are updated at a high frequency to ensure that the current state of the physical device is accurately reflected.
[0202] The feedback value is compared with the corresponding instruction value to obtain the real-time tracking error.
[0203] Specifically, in each control cycle, the present invention can calculate the difference between the currently acquired encoder position and gimbal angle, zoom feedback and the target control command sent at that moment to obtain the slide rail position error, camera equipment rotation error and zoom error, which are used as input parameters for closed-loop control to measure the deviation and response quality of the device's execution command.
[0204] Based on real-time tracking error, a feedforward-feedback control algorithm is used to dynamically adjust the control commands sent to the physical imaging system.
[0205] Specifically, in this embodiment of the invention, error data can be input into a PID or a more advanced feedforward-feedback control model. The feedforward part predicts motion compensation based on a preset trajectory, while the feedback part adjusts the gain based on the error amount, calculates the correction amount in real time, and resends the adjusted control command to the slide rail servo driver and PTZ pan-tilt camera controller to reduce errors, improve tracking accuracy and system stability, and ensure that the physical equipment accurately executes the planned trajectory.
[0206] This invention achieves high-precision closed-loop control of the physical shooting system by real-time acquisition of motion feedback from the slide rail encoder and camera equipment, precise comparison with command values, and dynamic adjustment of control commands based on tracking errors. This significantly improves the accuracy and stability of motion execution, thereby ensuring that the foreground performance capture is highly consistent with the target motion trajectory, and enhancing the overall visual effect of virtual-real fusion and system reliability.
[0207] Optionally, in the above Figure 1 In addition to one or more corresponding embodiments, another optional embodiment provided by the present invention may further include:
[0208] When the target motion trajectory exceeds the physical capabilities of the physical shooting system, the full physical drive shooting mode is executed: during the shooting execution phase, the camera motion trajectory defined in the reference video is realized entirely by the physical shooting equipment, without any motion compensation from the virtual camera.
[0209] Specifically, in the fully physical drive shooting mode, the target motion trajectory can be directly decomposed and planned to obtain the translational motion command sequence of the slide rail and the rotational zoom motion command sequence of the camera device. The physical shooting system is then driven to perform coordinated motion based on the translational motion command sequence and the rotational zoom motion command sequence to capture the foreground performance. Throughout the shooting process, the virtual camera remains stationary in the Unreal Engine (such as UE) or is only used to render a static background, without participating in any dynamic compensation to match the lens movement.
[0210] The embodiments of the present invention, through the full physical drive shooting mode, can achieve the optimal and most efficient shooting under physical conditions, realizing the lossless and automated conversion from AIGC blueprint to physical execution.
[0211] Although the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous.
[0212] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0213] This invention provides a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements the video generation method.
[0214] This invention provides a processor for running a program, wherein the program executes the video generation method during runtime.
[0215] like Figure 3 As shown, this embodiment of the invention provides an electronic device 1000, which includes at least one processor 1001, at least one memory 1002 connected to the processor 1001, and a bus 1003. The processor 1001 and the memory 1002 communicate with each other via the bus 1003. The processor 1001 is used to call program instructions in the memory 1002 to execute the aforementioned video generation method. The electronic device in this document can be a server, PC, PAD, mobile phone, etc.
[0216] The present invention also provides a computer program product that, when executed on an electronic device, is suitable for executing a program that initializes a video generation method step.
[0217] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, electronic devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0218] In a typical configuration, an electronic device includes one or more processors (CPUs), memory, and a bus. The electronic device may also include input / output interfaces, network interfaces, etc.
[0219] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM, and memory includes at least one memory chip. Memory is an example of computer-readable media.
[0220] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0221] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0222] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0223] In the description of this invention, it should be understood that if the terms "upper", "lower", "front", "rear", "left" and "right" are used to indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the position or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention.
[0224] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0225] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0226] The above are merely embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the present invention.
Claims
1. A video generation method, characterized in that, include: Obtain reference videos generated by artificial intelligence; Extract visual art elements from the reference video, wherein the visual art elements include camera motion parameters; Using the camera motion parameters, the target motion trajectory is generated; The physical shooting system is used to control the target's motion trajectory to capture the foreground performance and generate a foreground image. Determine whether the target's motion trajectory exceeds the physical capability limit of the physical shooting system. If so, execute the virtual-real collaborative shooting mode and calculate the motion difference between the target's motion trajectory and the actual motion trajectory of the physical shooting system. The motion difference is used to drive the virtual camera to perform compensated motion, and a dynamic compensated background layer is generated in real time. The foreground image, the dynamically compensated background layer, and the background image in the reference video are intelligently fused together to output a composite image.
2. The method according to claim 1, characterized in that, The visual art elements also include lighting information. Before capturing the foreground performance using the target motion trajectory-controlled physical shooting system and generating the foreground image, the method further includes: Using the illumination information, control commands for the physical lighting field are generated; The control command is sent to the physical lighting field to drive the physical lighting field to illuminate the shooting environment of the physical shooting system.
3. The method according to claim 2, characterized in that, The extraction of visual art elements from the reference video includes: Key visual regions in the reference video are identified, including areas directly illuminated by a light source and shadow areas; lighting information is fitted using a physical lighting model based on the areas directly illuminated by a light source and the shadow areas. And / or, in the video sequence of the reference video, the pose parameters of the virtual camera in each frame are calculated by using feature point tracking and motion recovery structures to obtain the camera motion parameters.
4. The method according to claim 1, characterized in that, The physical shooting system includes a slide rail and a camera mounted on the slide rail. The step of using the target motion trajectory to control the physical shooting system to capture foreground performance and generate foreground images includes: The target motion trajectory is decomposed and planned to obtain the translational motion command sequence of the slide rail and the rotational zoom motion command sequence of the camera device; The physical shooting system is driven to perform coordinated movements according to the translational motion command sequence and the rotational zoom motion command sequence in order to capture the foreground performance; The foreground performance is keyed to obtain the foreground image.
5. The method according to claim 4, characterized in that, The step of performing motion decomposition and planning on the target motion trajectory to obtain the translational motion command sequence of the slide rail and the rotational zoom motion command sequence of the camera device includes: The camera pose of each frame in the target motion trajectory is decomposed into the translation transformation of the slide rail in the world coordinate system, and the rotation transformation and focal length transformation of the camera device in its own base coordinate system. Based on minimizing the difference between the actual pose of the physical shooting system and the target pose, and satisfying the physical motion constraints of the slide rail and the camera device, a translational motion command sequence and a rotational zoom motion command sequence that are aligned in time are obtained.
6. The method according to claim 5, characterized in that, The process of obtaining the time-aligned translational motion command sequence and rotational zoom motion command sequence includes: Project the target position in the target motion trajectory onto the slide rail motion plane to obtain the desired position sequence of the slide rail; Motion planning is performed on the desired position sequence to generate a translational motion command sequence that satisfies velocity and acceleration constraints; As the slide rail moves sequentially to each desired position in the desired position sequence, the horizontal rotation angle, pitch rotation angle, and zoom value required by the camera device to align the camera optical axis with the scene target are determined, and a rotation zoom motion command sequence is generated.
7. The method according to claim 1, characterized in that, The step of using the motion difference to drive the virtual camera to perform compensated motion and rendering a dynamically compensated background layer in real time includes: Establish a unified transformation relationship between the world coordinate system, the physical camera coordinate system, and the virtual camera coordinate system; The pose parameters of the virtual camera are calculated in real time using the motion difference. Based on the timestamp synchronized with the physical shooting frames, the virtual camera is driven to move according to the pose parameters, and a dynamically compensated background layer is rendered in real time.
8. The method according to claim 1, characterized in that, The step of intelligently fusing the foreground image, the dynamically compensated background layer, and the background image in the reference video to output a composite image includes: The color statistical characteristics and light and shadow distribution information of the background image in the reference video were analyzed. Using the aforementioned color statistical characteristics, a color transformation matrix and a color offset vector are calculated to perform color correction on the foreground image; Using the light and shadow distribution information, add highlights and shadows to the foreground image that are consistent with the direction of the background light source; The foreground image after dynamic compensation, color correction, and lighting rendering is merged with the background image in the reference video to output a composite image.
9. The method according to claim 4, characterized in that, Also includes: The encoder feedback of the slide rail and the actual angle and zoom feedback of the camera device are read in real time to obtain the feedback value; The feedback value is compared with the corresponding instruction value to obtain the real-time tracking error; Based on the real-time tracking error, a feedforward-feedback control algorithm is used to dynamically adjust the control commands sent to the physical imaging system.
10. An electronic device, characterized in that, The electronic device includes at least one processor, at least one memory connected to the processor, and a bus; wherein the processor and the memory communicate with each other through the bus; the processor is used to call program instructions in the memory to execute the video generation method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Virtual camera and real camera switching system and method
CN104349020A
Camera correction method and system for space simulation shooting
CN117527992A