Image processing-based ai film shot three-dimensional accurate control method and system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-13
- Publication Date
- 2026-08-11
AI Technical Summary
在实际影视创作中,导演往往需要通过构图方式、镜头运动节奏以及视角变化等多种因素来表达特定情绪,但现有方法通常仅对构图或运镜单一因素进行优化,缺乏对情绪表达、构图合理性和运镜流畅性的统一建模与协同优化,难以在复杂场景下保证镜头情绪表达的一致性与准确性
通过本发明的基于图像处理的AI影视镜头三维精准控制方法及系统,可以实现从二维参考图像和导演输入的镜头情绪表达要求、构图要求、运镜要求,到多机位三维场景建模、机位配置、AI优化及渲染的全流程自动化处理。在三维场景构建阶段,本发明能够准确解析图像特征数据,生成包含场景结构、物体位置及层次关系的三维场景模型,并对人物资产进行骨骼绑定及动作参数设置,同时对物体资产和摄像机参数进行统一编码,从而为后续的多机位镜头控制提供高精度的场景基础数据。
Smart Images

Figure CN122554614A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and film and television production, and in particular to an AI-based method and system for precise 3D control of film and television shots based on image processing. Background Technology
[0002] In film and television production, shot composition, camera movement, and multi-camera collaboration are crucial for realizing the director's creative intentions and achieving desired cinematic effects. In traditional filmmaking, directors and cinematographers typically rely on experience to design shots and adjust shooting parameters, including angle of view, focal length, framing, subject position, and camera movement. However, with the increasing complexity of shooting scenarios and the growing scale of productions, existing technologies have gradually revealed several shortcomings in practical applications. In multi-camera shooting scenarios, there is a problem of poor consistency in the footage from multiple cameras. Because each camera typically sets and adjusts its parameters independently, lacking a unified linkage mechanism, when the composition or camera movement parameters of one camera change, other cameras struggle to adjust synchronously in a timely manner. This can easily lead to subject shifts, unbalanced compositional proportions, and unnatural shot transitions, affecting the overall visual impact. Existing AI-based shot control methods primarily focus on optimizing single-camera positions or fixed trajectories, lacking the ability to model and collaboratively optimize multi-camera systems as a whole. Existing AI models typically predict parameters only for the composition or camera movement of a single shot, failing to fully consider the spatial relationships and temporal correlations between different camera positions, making it difficult to achieve real-time collaborative control and overall optimization among multiple camera positions.
[0003] In the process of 3D scene construction, the import efficiency of external 3D models is low. Current technologies typically require manual intervention when integrating external 3D models into shooting scenes, including model analysis, coordinate alignment, scale matching, and spatial position adjustment. This low level of automation makes it difficult to quickly construct 3D scene models that meet shooting requirements, impacting overall production efficiency. Existing shot design methods generally lack comprehensive consideration of emotional expression in shots. In actual film and television creation, directors often need to express specific emotions through various factors such as composition, camera movement rhythm, and perspective changes. However, existing methods usually only optimize single factors like composition or camera movement, lacking unified modeling and collaborative optimization of emotional expression, compositional rationality, and camera movement smoothness. This makes it difficult to ensure the consistency and accuracy of emotional expression in complex scenes. How to achieve collaborative control and parameter linkage of multi-camera shots in 3D scenes to improve the consistency of multi-camera footage; how to construct AI shot control methods that support multi-camera joint optimization; how to improve the automation level of importing external 3D models and scene construction; and how to comprehensively consider multiple dimensions such as emotional expression, composition, and camera movement during shot parameter optimization are urgent technical problems that need to be solved in this field. Summary of the Invention
[0004] One objective of this invention is to propose an AI-based method and system for precise 3D control of film and television shots based on image processing. This invention fully utilizes image processing technology, 3D scene modeling technology, and deep learning algorithms to achieve precise optimization and intelligent control of the shooting angle, focal length, framing, composition parameters, and camera movement parameters for each camera position. This invention boasts advantages such as high compositional rationality, smooth camera movement, consistent emotional expression in shots, and high efficiency in multi-camera collaboration. Simultaneously, it can generate structured parameter data that can be directly used for rendering and film and television production, improving the efficiency and intelligence level of film and television creation.
[0005] The AI-based three-dimensional precision control method for film and television shots based on image processing according to an embodiment of the present invention includes the following steps: Acquire reference images and the director's input requirements for shot emotion expression, composition, and camera movement; preprocess the reference images to obtain image feature data. Based on image feature data, a 3D scene is constructed by calling the film and television 3D asset library to form a 3D scene model; Skeletal binding and motion parameter settings are performed on object assets and character assets in the 3D scene model, and the camera in the 3D scene model is parameterized to obtain scene parameters, asset motion parameters and camera parameters. Based on the 3D scene model, camera positions are configured, parameter linkage relationships between camera positions are established, and the composition parameters and camera movement parameters of the associated camera positions are updated synchronously according to the parameter linkage relationships to obtain multi-camera parameters; Image feature data, scene parameters, asset motion parameters, lens parameters, multi-camera parameters, as well as lens emotion expression requirements, composition requirements, and camera movement requirements are input into the AI lens control model. The AI lens control model then optimizes the shooting angle, focal length, framing, composition parameters, and camera movement parameters of each camera position to obtain optimized composition parameters and camera movement parameters. Based on the optimized composition and camera movement parameters, the lens images and camera movement effects corresponding to each camera position are rendered, and the rendering results are verified by multi-camera collaborative verification. The verified 3D scene model, scene parameters, asset motion parameters, camera parameters, multi-camera parameters, composition parameters, and camera movement parameters are stored in a structured manner, and the camera composition previews of each camera position are exported.
[0006] Optionally, the acquisition process involves receiving a reference image uploaded by the user via an input device. The reference image is a still image or a video frame image. The process also involves receiving the director's input requirements for shot emotion expression, composition, and camera movement via a human-computer interaction interface. The shot emotion expression requirements include an emotion category identifier. The composition requirements include the position of the main subject in the frame and the proportion of the frame. The camera movement requirements include the type of camera movement, the direction of movement, and the speed of movement. The image preprocessing includes noise reduction, enhancement, feature point extraction, and normalization.
[0007] Optionally, the formation of the three-dimensional scene model specifically includes: Generate 3D scene construction data based on scene structure features, object position features, and image hierarchy features in image feature data; Search the film and television 3D asset library for scene assets, object assets and character assets that correspond to the 3D scene construction data, determine the spatial position relationship of each asset in the 3D scene, and form the initial 3D scene; Receive external 3D model upload instructions, parse the format of the external 3D model, and import the parsed external 3D model into the initial 3D scene; The imported external 3D model is subjected to coordinate alignment, scale matching, and spatial position configuration processing to obtain fused scene data. Based on the 3D scene construction data and fused scene data, the spatial layout of scene assets, object assets, character assets and external 3D models in the initial 3D scene is adjusted to form a 3D scene model.
[0008] Optionally, obtaining the scene parameters, asset motion parameters, and camera parameters specifically includes: Skeletal binding is performed on the character assets in the 3D scene model to establish the skeletal hierarchy of the character assets, and initial position parameters are set for each joint node in the character assets. Based on the skeletal hierarchy, displacement, rotation, and scale parameters are set for each joint node in the character asset, and the motion parameters of the character asset are obtained based on the initial position parameters, displacement parameters, rotation parameters, and scale parameters of each joint node. Set initial position parameters, displacement velocity parameters, and rotation parameters for object assets in the 3D scene model, and determine the motion parameters of the object assets; Set the spatial position parameters, orientation parameters, focal length parameters, and imaging parameters for the camera in the 3D scene model, and determine the lens parameters; Based on the spatial distribution of scene assets, the motion parameters of character assets, the motion parameters of object assets, and the lens parameters of the camera in the 3D scene model, scene parameters, asset motion parameters, and lens parameters are formed.
[0009] Optionally, obtaining the multi-camera parameters specifically includes: Each camera position is determined based on a 3D scene model, and a unique identifier is assigned to each camera position. Initial position parameters, shooting angle parameters, and camera movement trajectory parameters are set for each camera position to obtain the initial camera position parameters for each camera position; Based on the initial camera position parameters, composition requirements, and camera movement requirements, establish the parameter linkage relationship between each camera position; When the parameters of the target camera position are adjusted, the changes in the initial position parameters, shooting angle parameters, and camera movement trajectory parameters after the adjustment of the target camera position are obtained, and the update amounts of the initial position parameters, shooting angle parameters, and camera movement trajectory parameters corresponding to the associated camera position are determined according to the parameter linkage relationship. Based on the update amounts of initial position parameters, shooting angle parameters, and camera movement trajectory parameters, the composition parameters and camera movement parameters of the associated camera positions are updated synchronously, and multi-camera parameters are generated based on the updated parameters of each camera position.
[0010] Optionally, the optimized composition parameters and camera movement parameters are obtained by means of: Image feature data, scene parameters, asset motion parameters, lens parameters, multi-camera parameters, as well as lens emotion expression requirements, composition requirements, and camera movement requirements are processed by feature encoding to generate fused feature data; The fused feature data is input into the AI camera control model, which adopts a hybrid architecture combining convolutional neural networks and Transformer networks to extract features and model feature associations from the fused feature data, thereby obtaining scene structure features, asset motion features, and camera features corresponding to each camera position. Based on scene structure features, asset action features, and lens features, combined with lens emotion expression requirements, composition requirements, and camera movement requirements, the shooting angle, focal length, framing, composition parameters, and camera movement parameters of each camera position are initially estimated to generate an initial parameter set for each camera position. Using the initial parameter set of each camera position as the optimization input, a composition evaluation function is constructed based on the requirements for emotional expression and composition of the shot, and a camera movement evaluation function is constructed based on the camera movement requirements, forming a comprehensive optimization objective function; Based on the comprehensive optimization objective function, the shooting angle, focal length, and framing of each camera position are iteratively optimized to generate composition schemes. The composition schemes are then screened to determine the composition schemes that meet the requirements of the emotional expression of the shot, and the corresponding composition parameters and camera movement parameters are output. When receiving instructions to adjust composition or camera movement parameters, the AI lens control model re-inputs the adjusted camera parameters to optimize them, resulting in optimized composition and camera movement parameters.
[0011] Optionally, the multi-camera collaborative verification specifically includes: Based on the optimized composition and camera movement parameters, the shot images corresponding to each camera position are rendered to generate a visual image sequence for each camera position. Multi-camera collaborative verification is performed on the visualized image sequence. The multi-camera collaborative verification includes image continuity verification, composition rationality verification, camera movement smoothness verification, and shot emotion expression consistency verification. The image continuity verification judges the smoothness of the image transition by analyzing the changes in the subject position of consecutive frames in the image sequence of each camera. The composition rationality verification determines the composition rationality by calculating the deviation of the subject position, image ratio, and layer relationship in each camera's image from the preset composition requirements. The camera movement smoothness verification judges the smoothness of the camera movement by analyzing the deviation of the changes in the camera movement trajectory and speed of each camera from the optimized camera movement parameters. The shot emotion expression consistency verification judges the consistency of emotion expression by comparing the matching degree of each camera's image features with the director's input shot emotion expression requirements. The verification result status of each camera is generated, indicating whether each verification item meets the preset threshold. Based on the verification results of each camera position, when a verification item fails to meet the preset threshold, the composition parameters and camera movement parameters of the corresponding camera position are adjusted, the updated parameters are reused to render the visualization image sequence of the corresponding camera position, and multi-camera collaborative verification is performed again until all verification items of all cameras meet the preset threshold.
[0012] Optionally, the structured storage involves organizing and storing the verified 3D scene model, scene parameters, asset motion parameters, lens parameters, multi-camera parameters, composition parameters, and camera movement parameters. Specifically, it involves associating and saving the scene assets, object assets, and character assets in the 3D scene model with the scene parameters and asset motion parameters, storing the lens parameters, multi-camera parameters, composition parameters, and camera movement parameters corresponding to each camera position according to the camera position number, and calling the rendering module to render the image content in the 3D scene model from the shooting perspective of the corresponding camera position to generate a lens composition preview image for each camera position.
[0013] An AI-based 3D precision control system for film and television shots, according to an embodiment of the present invention, includes: The image acquisition module is used to acquire reference images and the director's input requirements for shot emotion expression, composition, and camera movement. The reference images are preprocessed to obtain image feature data. The 3D scene construction module is used to construct 3D scenes based on image feature data and call the film and television 3D asset library to form 3D scene models; The parameter setting module is used to perform skeletal binding and motion parameter settings for object assets and character assets in the 3D scene model, and to parameterize the camera in the 3D scene model to obtain scene parameters, asset motion parameters, and camera parameters. The camera position management module is used to configure camera positions based on the 3D scene model, establish parameter linkage relationships between camera positions, and synchronously update the composition parameters and camera movement parameters of the associated camera positions according to the parameter linkage relationships to obtain multi-camera parameters. The AI lens control module is used to input image feature data, scene parameters, asset action parameters, lens parameters, multi-camera parameters, as well as lens emotion expression requirements, composition requirements and camera movement requirements into the AI lens control model. The AI lens control model optimizes the shooting angle, focal length, framing, composition parameters and camera movement parameters of each camera position to obtain optimized composition parameters and camera movement parameters. The rendering and verification module is used to render the shot images and camera movement effects corresponding to each camera position based on the optimized composition parameters and camera movement parameters, and to perform multi-camera collaborative verification of the rendering results. The storage and preview module is used to structurally store the verified 3D scene model, scene parameters, asset motion parameters, camera parameters, multi-camera parameters, composition parameters, and camera movement parameters, and to export the camera composition preview screens for each camera position.
[0014] The beneficial effects of this invention are: The AI-based 3D precision control method and system for film and television shots based on image processing of this invention enables fully automated processing from 2D reference images and director-input requirements for shot emotion expression, composition, and camera movement, to multi-camera 3D scene modeling, camera position configuration, AI optimization, and rendering. In the 3D scene construction stage, this invention can accurately analyze image feature data, generate a 3D scene model including scene structure, object positions, and hierarchical relationships, and perform skeletal binding and motion parameter settings for character assets. Simultaneously, it uniformly encodes object assets and camera parameters, thereby providing high-precision scene foundation data for subsequent multi-camera shot control.
[0015] In the multi-camera control and optimization stage, this invention establishes a parameter linkage relationship between camera positions, enabling other camera positions to synchronously update their composition and camera movement parameters when the parameters of one camera position are adjusted. This ensures consistency and coordination in composition and camera movement across multiple camera positions. Combined with an AI lens control model, this invention optimizes the shooting angle, focal length, framing, composition parameters, and camera movement parameters of each camera position across all dimensions. Composition and camera movement evaluation functions are introduced during the optimization process, and a weighted summation is used to form an optimization objective function. This ensures that the parameters of each camera position meet the director's emotional expression needs while also considering the rationality of the composition and the smoothness of the camera movement. Furthermore, it supports the generation of multiple composition variations, improving the flexibility of shot creation.
[0016] This invention utilizes a multi-camera collaborative verification module to perform real-time verification of rendered shot footage. It automatically adjusts unsuitable composition and camera movement parameters, iteratively optimizing until all verification indicators meet the standards. This ensures that the final output shot composition and camera movement parameters conform to film and television production standards while accurately conveying the director's intent. Through structured storage and visual preview, this invention allows optimized 3D scene models and multi-camera shot parameters to be directly used for rendering and film and television production, improving production efficiency, reducing the complexity of manual adjustments, and significantly enhancing the accuracy and consistency of shot effects, thus achieving intelligent and efficient film and television shot design. Attached Figure Description
[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is an overall flowchart of the AI-based three-dimensional precision control method and system for film and television shots based on image processing proposed in this invention; Figure 2 This is a schematic diagram illustrating the construction of scene parameters, asset motion parameters, and lens parameters for the AI-based film and television shot 3D precision control method and system based on image processing proposed in this invention. Figure 3 This is a schematic diagram illustrating the construction of a multi-camera collaborative verification system for an AI-based three-dimensional precision control method and system for film and television shots based on image processing, as proposed in this invention. Detailed Implementation
[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0019] refer to Figures 1-3 A method for precise 3D control of film and television shots based on image processing includes the following steps: Acquire reference images and the director's input requirements for shot emotion expression, composition, and camera movement; preprocess the reference images to obtain image feature data. Based on image feature data, a 3D scene is constructed by calling the film and television 3D asset library to form a 3D scene model; Skeletal binding and motion parameter settings are performed on object assets and character assets in the 3D scene model, and the camera in the 3D scene model is parameterized to obtain scene parameters, asset motion parameters and camera parameters. Based on the 3D scene model, camera positions are configured, parameter linkage relationships between camera positions are established, and the composition parameters and camera movement parameters of the associated camera positions are updated synchronously according to the parameter linkage relationships to obtain multi-camera parameters; Image feature data, scene parameters, asset motion parameters, lens parameters, multi-camera parameters, as well as lens emotion expression requirements, composition requirements, and camera movement requirements are input into the AI lens control model. The AI lens control model optimizes the shooting angle, focal length, framing, composition parameters, and camera movement parameters of each camera position to obtain optimized composition parameters and camera movement parameters. Based on the optimized composition and camera movement parameters, the lens images and camera movement effects corresponding to each camera position are rendered, and the rendering results are verified by multi-camera collaborative verification. The verified 3D scene model, scene parameters, asset motion parameters, camera parameters, multi-camera parameters, composition parameters, and camera movement parameters are stored in a structured manner, and the camera composition previews of each camera position are exported.
[0020] In this embodiment, the acquisition process involves receiving a reference image uploaded by the user through an input device. The reference image can be a still image or a video frame image. The director's input requirements for shot emotion expression, composition, and camera movement are received through a human-computer interaction interface. Shot emotion expression requirements include emotion category identifiers. Composition requirements include the position of the main subject in the frame and the proportion of the frame. Camera movement requirements include the type of camera movement, direction of movement, and speed parameters. Image preprocessing includes noise reduction, enhancement, feature point extraction, and normalization.
[0021] In this embodiment, the formation of the 3D scene model specifically includes: Generate 3D scene construction data based on scene structure features, object position features, and image hierarchy features in image feature data; The specific process of generating 3D scene construction data is as follows: Analyzing the scene structure features in the image feature data to extract boundary information and regional distribution information representing the spatial structure; analyzing the object position features in the image feature data to determine the position coordinates and relative positional relationships of each object in the 2D image; analyzing the image layer features in the image feature data to determine the hierarchical relationship between objects; establishing a mapping relationship from the 2D image coordinate system to the 3D spatial coordinate system based on the scene structure features, object position features, and image layer features; performing 3D coordinate transformation on the 2D position coordinates of each object to obtain the position parameters of each object in 3D space; and integrating the spatial structure information, object position parameters, and hierarchical relationships to generate 3D scene construction data. Search the film and television 3D asset library for scene assets, object assets and character assets that correspond to the 3D scene construction data, determine the spatial position relationship of each asset in the 3D scene, and form the initial 3D scene; The initial 3D scene formation process is as follows: Based on the scene structure features, object position features, and image hierarchy features in the 3D scene construction data, scene assets, object assets, and character assets are matched and retrieved in the film and television 3D asset library; the retrieved scene assets are loaded as the basic structure of the 3D scene; based on the object position features in the 3D scene construction data, object assets and character assets are configured to their corresponding spatial positions in the scene assets; based on the image hierarchy features in the 3D scene construction data, the spatial relationship between object assets and character assets is determined; and the scene assets, object assets, and character assets are spatially arranged in a unified coordinate system to generate the initial 3D scene. Receive external 3D model upload instructions, parse the format of the external 3D model, and import the parsed external 3D model into the initial 3D scene; The specific process of format parsing for external 3D models involves: performing data parsing processing on the model files of external 3D models to extract geometric structure data, material data, and hierarchical structure data from the model files; among which, geometric structure data includes vertex coordinates, face information, and normal vector information, material data includes texture information and material properties, and hierarchical structure data includes the structural relationships between various parts in the model; and converting the parsed data into a unified data format within the system. The imported external 3D model is subjected to coordinate alignment, scale matching, and spatial position configuration processing to obtain fused scene data. The process of obtaining fused scene data is as follows: Coordinate alignment is performed on the imported external 3D model to transform its coordinate system to match that of the initial 3D scene; scale matching is performed on the aligned external 3D model to scale it according to the size information of the corresponding objects in the initial 3D scene; spatial positioning is performed on the external 3D model based on the object position features in the 3D scene construction data to position it at the target location in the initial 3D scene; the external 3D model, after coordinate alignment, scale matching, and spatial positioning, is then integrated with the scene assets, object assets, and character assets in the initial 3D scene to obtain fused scene data. Based on the 3D scene construction data and fused scene data, the spatial layout of scene assets, object assets, character assets and external 3D models in the initial 3D scene is adjusted to form a 3D scene model; The spatial layout adjustment process is as follows: Based on the scene structure features in the 3D scene construction data, the overall spatial structure of the scene assets in the initial 3D scene is adjusted to determine the spatial boundaries and regional distribution of the 3D scene; based on the object position features in the 3D scene construction data, the position parameters of object assets, character assets, and external 3D models in 3D space are adjusted; based on the image hierarchy features in the 3D scene construction data, the spatial hierarchy relationship between object assets, character assets, and external 3D models is adjusted; and the scene assets, object assets, character assets, and external 3D models after the spatial structure adjustment, position parameter adjustment, and hierarchy relationship adjustment are uniformly spatially integrated to generate a 3D scene model.
[0022] In this embodiment, the acquisition of scene parameters, asset motion parameters, and camera parameters specifically includes: Skeletal binding is performed on the character assets in the 3D scene model to establish the skeletal hierarchy of the character assets and set the initial position parameters for each joint node in the character assets. The process of establishing the skeletal hierarchy of character assets is as follows: perform structural analysis on the 3D model of the character assets to identify the joint nodes corresponding to each component in the character assets; determine the parent-child hierarchical relationship between each joint node based on the connection relationship of each joint node in 3D space; and establish the hierarchical association between each joint node step by step, starting from the root joint node, to form a skeletal hierarchy structure composed of joint nodes and parent-child relationships. The initial position parameter setting process includes: determining the spatial position of each joint node in the character asset model based on the geometric structure data of the character asset in the 3D scene model; and calibrating the position of the joint node in the 3D space by combining the object position features and image layer features in the 3D scene construction data to obtain the initial position parameters of each joint node. Based on the skeletal hierarchy, displacement, rotation, and scale parameters are set for each joint node in the character asset, and the motion parameters of the character asset are obtained based on the initial position parameters, displacement parameters, rotation parameters, and scale parameters of each joint node. The settings for displacement, rotation, and scale parameters are as follows: Based on the object position features and image hierarchy features in the 3D scene construction data, the target spatial position and target pose of the character asset in the 3D scene model are determined; the displacement parameter is represented by the position offset of each joint node in 3D space, which is determined by the difference between the target spatial position and the initial position parameter; the rotation parameter is represented by the angular change of each joint node relative to the initial pose, which is determined based on the target orientation of the character asset in the 3D scene model; the scale parameter is represented by the scaling ratio of each joint node, which is determined based on the spatial proportions in the 3D scene construction data. Set initial position parameters, displacement velocity parameters, and rotation parameters for object assets in the 3D scene model, and determine the motion parameters of the object assets; The process of setting the initial position parameters, displacement velocity parameters, and rotation parameters is as follows: Based on the geometric structure data of the object assets in the 3D scene model, determine the initial position parameters of the object assets in 3D space; based on the object position features and image layer features in the 3D scene construction data, determine the target position and motion direction of the object assets in the 3D scene; calculate the displacement velocity parameters of the object assets according to the spatial distance between the target position and the initial position parameters and the preset time parameters; determine the rotation parameters of the object assets according to the motion direction. Set the spatial position parameters, orientation parameters, focal length parameters, and imaging parameters for the camera in the 3D scene model, and determine the lens parameters; The process of setting spatial position parameters, orientation parameters, focal length parameters, and imaging parameters is as follows: Based on the object position features and image layer features in the 3D scene construction data, determine the target shooting position of the camera in the 3D scene model, and set the target shooting position as the camera's spatial position parameter; Based on composition requirements and shot emotional expression requirements, determine the camera's target shooting direction, and set the target shooting direction as the camera's orientation parameter; Based on the shot size requirements in the composition requirements and the spatial distance between the shooting object and the camera in the 3D scene model, determine the camera's focal length parameter; Based on the spatial distribution of scene assets, the motion parameters of character assets, the motion parameters of object assets, and the lens parameters of the camera in the 3D scene model, scene parameters, asset motion parameters, and lens parameters are formed.
[0023] In this embodiment, obtaining the multi-camera parameters specifically includes: Each camera position is determined based on a 3D scene model, and a unique identifier is assigned to each camera position. The process of determining each camera position is as follows: Based on the scene parameters and lens parameters in the 3D scene model, combined with composition requirements and shot emotional expression requirements, determine the number of camera positions that need to be set; based on the spatial distribution of scene assets in the 3D scene model and the positional relationship between character assets and object assets, determine the candidate shooting positions of each camera position in the 3D scene; based on composition requirements, screen the candidate shooting positions to determine the camera positions that meet the requirements of the subject position, the proportion of the image, and the layering of the image; adjust the camera positions according to the shot emotional expression requirements to obtain the final determined camera positions; and assign a unique identifier to each camera position. Initial position parameters, shooting angle parameters, and camera movement trajectory parameters are set for each camera position to obtain the initial camera position parameters for each camera position; The process of setting initial position parameters, shooting angle parameters, and camera movement trajectory parameters is as follows: Based on the spatial distribution of scene assets in the 3D scene model and the positional relationship between character assets and object assets, determine the target shooting position of each camera position in 3D space and set the target shooting position as the initial position parameter; Based on composition requirements and shot emotional expression requirements, determine the target shooting direction of each camera position and set the target shooting direction as the shooting angle parameter; Based on camera movement requirements, determine the movement path and movement mode of each camera position in the time dimension and set the movement path and movement mode as the camera movement trajectory parameter; Based on the initial position parameters, shooting angle parameters, and camera movement trajectory parameters, generate the initial camera position parameters for each camera position. Based on the initial camera position parameters, composition requirements, and camera movement requirements, establish the parameter linkage relationship between each camera position; The process of establishing parameter linkage is as follows: The initial camera position parameters of each camera position are analyzed to obtain the initial position parameters, shooting angle parameters, and camera movement trajectory parameters of each camera position; one camera position is selected as the reference camera position, and the initial position parameters, shooting angle parameters, and camera movement trajectory parameters are used as the reference parameters; for each associated camera position, the spatial position difference between the initial position parameters and the initial position parameters of the reference camera position is calculated, the angular difference between the shooting angle parameters and the shooting angle parameters of the reference camera position is calculated, and the trajectory difference between the camera movement trajectory parameters and the camera movement trajectory parameters of the reference camera position is calculated; the spatial position difference, angular difference, and trajectory difference are used as the parameter offsets of the associated camera position relative to the reference camera position; based on the parameter offsets, a correspondence is established between the initial camera position parameters of each associated camera position and the initial camera position parameters of the reference camera position, forming the parameter linkage relationship between each camera position; When the parameters of the target camera position are adjusted, the changes in the initial position parameters, shooting angle parameters, and camera movement trajectory parameters after the adjustment of the target camera position are obtained, and the update amounts of the initial position parameters, shooting angle parameters, and camera movement trajectory parameters corresponding to the associated camera position are determined according to the parameter linkage relationship. Based on the update amounts of initial position parameters, shooting angle parameters, and camera movement trajectory parameters, the composition parameters and camera movement parameters of the associated camera positions are updated synchronously, and multi-camera parameters are generated based on the updated parameters of each camera position. The multi-camera parameter generation process is as follows: First, obtain the initial position parameters, shooting angle parameters, and camera movement trajectory parameters of the associated cameras. Second, based on the update amount of the initial position parameters, overlay the initial position parameters of the associated cameras to obtain updated spatial position parameters. Third, based on the update amount of the shooting angle parameters, overlay the shooting angle parameters of the associated cameras to obtain updated shooting angle parameters. Fourth, based on the update amount of the camera movement trajectory parameters, overlay the camera movement trajectory parameters of the associated cameras to obtain updated camera movement trajectory parameters. Fifth, recalculate the composition parameters of the associated cameras based on the updated spatial position parameters and updated shooting angle parameters. Sixth, based on the updated camera movement trajectory parameters, redetermine the camera movement parameters of the associated cameras, including movement path and movement speed. Finally, summarize the updated composition parameters and camera movement parameters of each associated camera to form the multi-camera parameters.
[0024] In this embodiment, the optimized composition parameters and camera movement parameters are obtained specifically through: Image feature data, scene parameters, asset motion parameters, lens parameters, multi-camera parameters, as well as lens emotion expression requirements, composition requirements, and camera movement requirements are processed by feature encoding to generate fused feature data; The feature encoding process is as follows: convolutional feature extraction is performed on image feature data to obtain key features of scene texture, edges, and target objects; scene parameters, asset motion parameters, and camera parameters are vectorized and encoded to map 3D spatial position, skeleton motion parameters, and camera parameters to a unified feature vector space; multi-camera parameters are serialized, arranging the initial position, shooting angle, and camera movement trajectory parameters of each camera in chronological order; image feature vectors, scene feature vectors, asset motion feature vectors, camera feature vectors, and multi-camera feature vectors are concatenated and fused, and the consistency of each feature dimension is ensured through standardization to generate fused feature data; The fused feature data is input into the AI camera control model, which adopts a hybrid architecture combining convolutional neural networks and Transformer networks to extract features and model feature associations from the fused feature data, thereby obtaining scene structure features, asset motion features, and camera features corresponding to each camera position. The convolutional neural network module comprises a cascaded structure of multiple convolutional layers and nonlinear activation layers. Each convolutional layer uses convolutional kernels of different sizes to extract features from the fused feature data. The kernel sizes include 3×3 and 5×5. Normalization and downsampling layers are set between adjacent convolutional layers. The Transformer module comprises a multi-layer encoding structure. Each layer contains a multi-head self-attention submodule and a feedforward neural network submodule. The multi-head self-attention submodule is used to perform parallel modeling of camera sequence features from multiple feature subspaces. The feedforward neural network submodule is used to perform nonlinear mapping and feature reconstruction on the attention output features. During model training, the AI lens control model is trained based on data samples containing multi-camera shot parameters and corresponding composition and movement annotation information. The model parameters are iteratively updated by minimizing the objective function composed of composition deviation, movement deviation, and multi-camera consistency deviation. The process of obtaining scene structure features, asset motion features, and camera features involves the following steps: The fused feature data is input into a CNN module, where multi-layer convolutional operations are performed to extract local spatial features, including scene textures, object boundaries, character joint keypoints, and local camera position features. The spatial features output by the CNN are mapped to a Transformer module, where a self-attention mechanism is used to globally model the features of different time steps or camera sequences, capturing the spatial-temporal dependencies between camera positions and the correlation between composition and camera movement. At the Transformer output layer, the features of each camera position are branched to extract scene structure features to represent the spatial position and distribution of scene assets in the 3D scene, asset motion features to represent the joint motion parameters of character assets and the displacement and rotation parameters of object assets, and camera features to represent the camera's spatial position, orientation, focal length, and camera trajectory. Finally, the processed feature vectors are normalized and encoded to form the scene structure features, asset motion features, and camera features corresponding to each camera position. Based on scene structure features, asset action features, and lens features, combined with lens emotion expression requirements, composition requirements, and camera movement requirements, the shooting angle, focal length, framing, composition parameters, and camera movement parameters of each camera position are initially estimated to generate an initial parameter set for each camera position. The process of generating the initial parameter set for each camera position is as follows: Spatial analysis of scene structure features is performed to determine the shooting area and shooting angle range for each camera position in the 3D scene; asset motion features are analyzed to obtain the joint motion range of character assets and the displacement and rotation range of object assets; lens features are analyzed to obtain the initial position, orientation, focal length, and camera movement trajectory information of each camera position; based on scene structure features, asset motion features, and lens features, combined with the director's input requirements for emotional expression, composition, and camera movement, the initial shooting angle, initial focal length, initial shot size, initial composition parameters, and initial camera movement parameters for each camera position are calculated. The initial shooting angle is determined based on the position of the target subject in the frame and the emotional expression requirements; the initial focal length is determined based on the shot size and the distance to the subject; the initial shot size is determined based on the frame coverage and composition requirements; the initial composition parameters are determined based on the subject position, frame proportion, and layering; and the initial camera movement parameters are determined based on the movement path and rhythm requirements. All parameters are then summarized to generate the initial parameter set for each camera position. Using the initial parameter set of each camera position as the optimization input, a composition evaluation function is constructed based on the requirements of shot emotion expression and composition, and a camera movement evaluation function is constructed based on the requirements of camera movement, forming a comprehensive optimization objective function; The formation process of the comprehensive optimization objective function is as follows: Using the initial parameter set of each camera position as input, a composition evaluation function and a camera movement evaluation function are calculated for each camera position. The composition evaluation function analyzes the shooting angle, focal length, framing, and composition parameters of each camera position to calculate the deviation of the subject position, aspect ratio, layering, and compactness of the image. Based on the director's input of the emotional expression requirements for the shot, each deviation is weighted and summed to obtain the composition evaluation score for each camera position. The camera movement evaluation function analyzes the camera movement trajectory parameters, movement path, movement speed, and camera movement rhythm of each camera position to quantify the degree of deviation between the camera movement and the camera movement requirements. Based on the director's input of the emotional expression requirements and the rhythm requirements for the shot, each deviation is weighted and summed to obtain the camera movement evaluation score for each camera position. Finally, the composition evaluation scores and camera movement evaluation scores of all camera positions are weighted and fused according to preset weights to obtain the comprehensive optimization objective function for the entire multi-camera system. Based on the comprehensive optimization objective function, the shooting angle, focal length, and framing of each camera position are iteratively optimized to generate composition schemes. The composition schemes are then selected to determine those that meet the requirements of the emotional expression of the shot, and the corresponding composition parameters and camera movement parameters are output. The specific process for outputting composition parameters and camera movement parameters is as follows: Based on the comprehensive optimization objective function, the initial shooting angle, focal length, and framing parameters of each camera position are iteratively adjusted. In each iteration, the composition evaluation score and camera movement evaluation score under the current parameter combination are calculated according to the comprehensive optimization objective function. In each iteration, the parameter combination that meets the preset composition accuracy threshold and camera movement error threshold is recorded as a composition scheme. The iteration is repeated until the set number of iterations or optimization convergence conditions are reached to form different feasible composition schemes. Each composition scheme includes the shooting angle, focal length, framing, composition parameters, and camera movement parameters of the corresponding camera position. The composition schemes are screened, and based on the requirements of shot emotional expression, composition requirements, and camera movement requirements, the composition scheme that meets the director's intention and film and television picture specifications is selected, and the composition parameters and camera movement parameters corresponding to the composition scheme are output. When receiving instructions to adjust composition or camera movement parameters, the AI lens control model re-inputs the adjusted camera parameters to optimize them, resulting in optimized composition and camera movement parameters.
[0025] In this embodiment, multi-camera collaborative verification specifically includes: Based on the optimized composition and camera movement parameters, the shot images corresponding to each camera position are rendered to generate a visual image sequence for each camera position. Multi-camera collaborative verification is performed on the visualized image sequence. The multi-camera collaborative verification includes image continuity verification, composition rationality verification, camera movement smoothness verification, and shot emotion expression consistency verification. Image continuity verification judges the smoothness of the image transition by analyzing the changes in the subject position of consecutive frames in the image sequence of each camera. Composition rationality verification determines the rationality of the composition by calculating the deviation of the subject position, image ratio, and layer relationship in each camera's image from the preset composition requirements. Camera movement smoothness verification judges the smoothness of the camera movement by analyzing the deviation of the changes in camera movement trajectory and speed in each camera from the optimized camera movement parameters. The preset threshold values for the deviation of the preset composition requirements in multi-camera collaborative verification are ≤1° and the deviation of the changes in camera movement trajectory and speed from the optimized camera movement parameters are ≤0.06m / s. Shot emotion expression consistency verification judges the consistency of emotion expression by comparing the matching degree of the image features of each camera with the shot emotion expression requirements input by the director. The verification result status of each camera is generated to indicate whether each verification item meets the preset threshold. Based on the verification results of each camera position, when a verification item fails to meet the preset threshold, the composition parameters and camera movement parameters of the corresponding camera position are adjusted, the updated parameters are reused to render the visualization image sequence of the corresponding camera position, and multi-camera collaborative verification is performed again until all verification items of all cameras meet the preset threshold. The process of adjusting the composition and camera movement parameters of each camera position is as follows: Based on the verification results of each camera position, identify composition and camera movement verification items that do not meet the preset thresholds; for each camera position that fails verification, calculate the required parameter correction amount. The composition parameter correction amount includes adjustments to the subject's position, aspect ratio, and layering; the camera movement parameter correction amount includes adjustments to the motion path, speed, and rhythm. Apply the correction amounts to the corresponding camera position's composition and camera movement parameters, obtaining the adjusted parameters through parameter overlay or update operations. Input the adjusted composition and camera movement parameters into the rendering module to regenerate the camera position's visual image sequence. Perform multi-camera collaborative verification again in the generated image sequence to determine if the updated parameters meet the preset thresholds. If the thresholds are not met, repeat the calculation of correction amounts and parameter updates until all camera positions' composition and camera movement parameters meet the preset threshold requirements.
[0026] In this embodiment, structured storage involves organizing and storing the verified 3D scene model, scene parameters, asset motion parameters, lens parameters, multi-camera parameters, composition parameters, and camera movement parameters. Specifically, it involves associating and saving the scene assets, object assets, and character assets in the 3D scene model with the scene parameters and asset motion parameters. The lens parameters, multi-camera parameters, composition parameters, and camera movement parameters corresponding to each camera position are stored according to the camera position number. The rendering module is then called to render the image content in the 3D scene model from the shooting perspective of the corresponding camera position, generating a lens composition preview screen for each camera position.
[0027] An AI-based 3D precision control system for film and television lenses, based on image processing, includes: The image acquisition module is used to acquire reference images and the director's input requirements for shot emotion expression, composition, and camera movement. The reference images are preprocessed to obtain image feature data. The 3D scene construction module is used to construct 3D scenes based on image feature data and call the film and television 3D asset library to form 3D scene models; The parameter setting module is used to perform skeletal binding and motion parameter settings for object assets and character assets in the 3D scene model, and to parameterize the camera in the 3D scene model to obtain scene parameters, asset motion parameters, and camera parameters. The camera position management module is used to configure camera positions based on the 3D scene model, establish parameter linkage relationships between camera positions, and synchronously update the composition parameters and camera movement parameters of the associated camera positions according to the parameter linkage relationships to obtain multi-camera parameters. The AI lens control module is used to input image feature data, scene parameters, asset action parameters, lens parameters, multi-camera parameters, as well as lens emotion expression requirements, composition requirements and camera movement requirements into the AI lens control model. The AI lens control model optimizes the shooting angle, focal length, framing, composition parameters and camera movement parameters of each camera position to obtain optimized composition parameters and camera movement parameters. The rendering and verification module is used to render the shot images and camera movement effects corresponding to each camera position based on the optimized composition parameters and camera movement parameters, and to perform multi-camera collaborative verification of the rendering results. The storage and preview module is used to structurally store the verified 3D scene model, scene parameters, asset motion parameters, camera parameters, multi-camera parameters, composition parameters, and camera movement parameters, and to export the camera composition preview screens for each camera position.
[0028] Example 1: In a typical film shooting scenario, filming an indoor chase scene in a suspenseful drama, the director wants to capture the tense expressions, movement trajectories, and the sense of depth of the surrounding environment using three cameras simultaneously. Traditional shooting methods require manual setup of camera positions, adjustment of shooting angles and camera movements, which can easily lead to inconsistent composition, mismatched camera rhythms, and deviations in the emotional expression of the shots, resulting in insufficient tension or choppy transitions in the final footage. In this scenario, the AI-based three-dimensional precision control method and system for film and television shots based on image processing of this invention is applied to the entire process from initial reference images to multi-camera shot optimization.
[0029] The system uploads reference images, consisting of static storyboards and partial video clips of the scene, to the user via an input device. The director inputs requirements for emotional expression through the human-computer interface, such as "increased tension and suspense," compositional requirements (e.g., the main subject in the upper left corner, a 3:2 aspect ratio), and camera movement requirements (e.g., push-in, forward-facing, speed between 0.2 and 0.4 meters per second). The system preprocesses the reference images, including denoising, enhancement, feature point extraction, and normalization, to obtain scene texture features, key points of characters and objects, and spatial hierarchy information. Subsequently, the system accesses scene assets, character assets, and object assets from the Xcine film and television 3D asset library and receives external 3D models uploaded by the director. All models are imported into the 3D scene for coordinate alignment, scale matching, and spatial layout integration to generate a complete 3D scene model. The system rigs the character assets with skeletons and sets motion parameters; it initializes the position, rotation, and camera movement parameters for object assets and the camera, forming scene parameters, asset motion parameters, and initial shot parameters.
[0030] During the multi-camera setup phase, the system automatically generates candidate shooting positions for the three cameras based on scene structure, character positions, and object layout, and then selects the final camera positions based on the director's input requirements for shot emotional expression. The system generates initial camera parameters for each camera's initial position, shooting angle, and camera movement trajectory, establishing parameter linkage between cameras to ensure that when one camera is adjusted, the composition and camera movement parameters of other cameras are simultaneously optimized. The system inputs image feature data, scene parameters, asset motion parameters, initial shot parameters, multi-camera parameters, and director's input requirements into the AI shot control model. This model uses a hybrid architecture of convolutional neural networks and Transformers for feature extraction and spatiotemporal correlation modeling, outputting scene structure features, asset motion features, and shot features for each camera, and generating initial estimates of shooting angle, focal length, framing, composition parameters, and camera movement parameters. The system constructs a comprehensive optimization objective function, taking shot emotional expression, composition rationality, camera movement smoothness, and multi-camera consistency as objective terms, and iteratively optimizes to obtain the final optimized parameters.
[0031] Based on the optimized parameters, the system renders footage from three cameras for multi-camera collaborative verification. Verification includes shot continuity, compositional rationality, camera movement smoothness, and consistency in emotional expression. If the composition or camera movement parameters of any camera fail to meet the threshold, the system automatically calculates the parameter correction amount and re-renders until all cameras pass verification. Finally, the system structurally stores the verified 3D scene model, scene parameters, asset motion parameters, camera parameters, multi-camera parameters, composition parameters, and camera movement parameters, and generates a preview of the shot composition for each camera for subsequent film production and rendering.
[0032] As can be seen from the above embodiments, the present invention can automatically optimize the composition parameters, shooting angle, focal length, framing and camera movement parameters of each camera position in a multi-camera 3D scene, ensure the consistency of multi-camera footage in composition, camera movement and emotional expression, improve film and television shooting efficiency, reduce reliance on manual adjustments, and provide structured parameter data that can be directly used for rendering and production.
[0033] Table 1 Performance Comparison Before and After Multi-Camera Optimization
[0034] As shown in Table 1, the method of this invention significantly improves the performance of three-camera shots before and after optimization. Before optimization, the average compositional deviation of the three cameras was approximately 3.2°, the camera movement speed deviation was approximately 0.18 m / s, and the emotional consistency score for multi-camera shots was only 68-70 points, indicating problems with inaccurate composition, unstable camera movement rhythm, and inaccurate emotional expression. After optimization using this invention, the compositional deviation was significantly reduced to between 0.7° and 0.9°, the camera movement speed deviation was reduced to between 0.04 and 0.06 m / s, and the emotional consistency score for multi-camera shots improved to between 95 and 97 points. The continuity of the footage and the smoothness of the camera movement both meet the requirements of film and television production. This demonstrates that this invention can effectively solve the problems of inconsistent footage, unsmooth camera movement, and inconsistent emotional expression in traditional multi-camera shooting, significantly improving the accuracy of camera control and the efficiency of film and television production.
[0035] Example 2: An outdoor chase scene from an urban action film is used as an example. This scene is set in a city street environment, including vehicles, buildings, and multiple characters and assets. The director wanted to use multiple cameras simultaneously to depict characters running, vehicles moving, and a sense of environmental pressure, thus conveying a tense and urgent emotional atmosphere. This scene involves three cameras: one camera follows the main character's movement, another camera captures the spatial relationship between the character and vehicles from the side, and a third camera provides a distant, overhead view showcasing the overall environmental layout.
[0036] In traditional shooting methods, each camera position is typically set up and adjusted separately by the photographer. When one camera changes its focal length or angle to emphasize a person's expression, other cameras often cannot adjust synchronously, leading to inconsistent composition, subject misalignment, and unnatural transitions between shots. Furthermore, in complex outdoor environments, the dynamic movement of vehicles and people makes it difficult to maintain stable camera movement manually, resulting in uneven camera speeds and inconsistent shot transitions, thus affecting the overall image quality. Additionally, such scenes often require importing numerous external 3D models, which traditionally require manual alignment and adjustment, resulting in low efficiency.
[0037] In this embodiment, the reference image of the scene and the director's input requirements for shot emotion expression, composition, and camera movement are first obtained through the input device. The reference image consists of environmental photographs and storyboard sketches taken on set. The director's input requirements include the emotional expression goal of "increased tension and pressure," the composition requirement is that the main character be kept slightly to the right of the frame and occupy about one-third of the frame, and the camera movement requirement is to use medium-to-high speed following motion as the main technique, with slight acceleration at key positions.
[0038] The system processes the reference image, extracting information such as road structure, building distribution, character positions, and vehicle positions from the scene, and constructs a corresponding 3D scene model based on this information. During the construction process, external vehicle models and road facility models are imported into the scene, and their positions are matched and spatially arranged to ensure consistency with the environment in the reference image. Subsequently, the character assets are skeletally rigged and motion parameters are set, enabling the character to perform continuous running movements in the 3D scene. Simultaneously, the initial positions, shooting angles, and camera trajectories of each camera position are set.
[0039] After the camera positions are set up, the system automatically generates initial shooting parameters for the three camera positions based on the spatial relationship between the people and vehicles in the 3D scene, and establishes a parameter linkage relationship between the camera positions. When the main camera is tracking the person, if its shooting angle or focal length changes, the system can automatically calculate the impact of this change on other camera positions and synchronously adjust the composition and camera movement parameters of the side camera positions and the distant camera positions, thereby ensuring the consistency of the images between each camera position.
[0040] Image feature data, scene parameters, asset motion parameters, lens parameters, and multi-camera parameters are input into the AI lens control model. This model jointly optimizes the shooting angle, focal length, framing, composition parameters, and camera movement parameters of each camera position. During the optimization process, it comprehensively considers the rationality of composition, the smoothness of camera movement, and the consistency of emotional expression, so that the generated lens scheme can simultaneously meet the director's creative intentions and film and television production standards.
[0041] After obtaining the optimized parameters, the system renders the shots for each camera position and generates a continuous image sequence. Then, the shots from each camera position are collaboratively verified. By analyzing changes in subject position, camera movement trajectory, and consistency of shot rhythm, parameters that do not meet the requirements are automatically adjusted and re-rendered until all camera positions meet the requirements. Finally, the system outputs preview shots of the lens composition for each camera position and uses the relevant parameters in subsequent production processes.
[0042] By comparing the method of this invention with the traditional manual adjustment method, multiple sets of experiments were conducted under the same scenario conditions, and the following performance comparison results were obtained; Table 2 Performance Comparison of Outdoor Multi-Camera Lenses Before and After Optimization
[0043] As shown in Table 2, in complex and dynamic outdoor scenes, the performance indicators of multi-camera shots optimized using the method of this invention are significantly improved. Regarding composition, the compositional deviation of the first three camera positions was over 3°, indicating a significant image shift between different camera positions. However, after applying the method of this invention, the compositional deviation was reduced to less than 1°, an overall reduction of over 75%. This demonstrates that through the camera position parameter linkage mechanism, the compositional deviation between different camera positions is significantly improved. Figure 1 Consistency has been effectively improved.
[0044] Regarding camera movement control, before optimization, due to the difficulty in precisely controlling the movement rhythm through manual adjustments, the deviation of camera movement speed at each camera position was generally between 0.17m / s and 0.22m / s. After optimization, it decreased to about 0.05m / s, a reduction of about 70%. This indicates that by using an AI model to jointly optimize camera movement trajectory and speed, the stability and smoothness of camera movement can be significantly improved.
[0045] In terms of emotional expression, the emotional consistency scores of the first three camera positions were between 66 and 70 points, showing significant differences and making it difficult to uniformly express the tense atmosphere. However, after optimization, all three positions improved to over 95 points, indicating that by introducing composition evaluation functions, camera movement evaluation functions, and multi-camera collaborative optimization mechanisms, it is possible to effectively achieve consistency in emotional expression among multiple camera positions.
[0046] This invention can effectively solve the problems of inconsistent multi-camera shots, unsmooth camera movements, and inconsistent emotional expression in complex outdoor dynamic scenes, significantly improving the accuracy of camera control and the efficiency of film and television production, and has good practical application effects.
[0047] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. An AI-based method for precise 3D control of film and television shots based on image processing, characterized in that, Includes the following steps: Acquire reference images and the director's input requirements for shot emotion expression, composition, and camera movement; preprocess the reference images to obtain image feature data. Based on image feature data, a 3D scene is constructed by calling the film and television 3D asset library to form a 3D scene model; Skeletal binding and motion parameter settings are performed on object assets and character assets in the 3D scene model, and the camera in the 3D scene model is parameterized to obtain scene parameters, asset motion parameters and camera parameters. Based on the 3D scene model, camera positions are configured, parameter linkage relationships between camera positions are established, and the composition parameters and camera movement parameters of the associated camera positions are updated synchronously according to the parameter linkage relationships to obtain multi-camera parameters; Image feature data, scene parameters, asset motion parameters, lens parameters, multi-camera parameters, as well as lens emotion expression requirements, composition requirements, and camera movement requirements are input into the AI lens control model. The AI lens control model then optimizes the shooting angle, focal length, framing, composition parameters, and camera movement parameters of each camera position to obtain optimized composition parameters and camera movement parameters. Based on the optimized composition and camera movement parameters, the lens images and camera movement effects corresponding to each camera position are rendered, and the rendering results are verified by multi-camera collaborative verification. The verified 3D scene model, scene parameters, asset motion parameters, camera parameters, multi-camera parameters, composition parameters, and camera movement parameters are stored in a structured manner, and the camera composition previews of each camera position are exported.
2. The AI-based three-dimensional precision control method for film and television shots based on image processing according to claim 1, characterized in that, The acquisition process involves receiving a reference image uploaded by the user via an input device. The reference image can be a still image or a video frame image. The process also involves receiving the director's input requirements for shot emotion expression, composition, and camera movement via a human-computer interaction interface. The shot emotion expression requirements include an emotion category identifier. The composition requirements include the position of the main subject in the frame and the proportion of the frame. The camera movement requirements include the type, direction, and speed parameters of the camera movement. The image preprocessing includes noise reduction, enhancement, feature point extraction, and normalization.
3. The AI video shot three-dimensional accurate control method based on image processing according to claim 1, characterized in that, The formation of the three-dimensional scene model specifically includes: Generate 3D scene construction data based on scene structure features, object position features, and image hierarchy features in image feature data; Search the film and television 3D asset library for scene assets, object assets and character assets that correspond to the 3D scene construction data, determine the spatial position relationship of each asset in the 3D scene, and form the initial 3D scene; Receive external 3D model upload instructions, parse the format of the external 3D model, and import the parsed external 3D model into the initial 3D scene; The imported external 3D model is subjected to coordinate alignment, scale matching, and spatial position configuration processing to obtain fused scene data. Based on the 3D scene construction data and fused scene data, the spatial layout of scene assets, object assets, character assets and external 3D models in the initial 3D scene is adjusted to form a 3D scene model.
4. The AI video shot three-dimensional accurate control method based on image processing according to claim 1, characterized in that, The specific methods for obtaining the scene parameters, asset motion parameters, and camera parameters include: Skeletal binding is performed on the character assets in the 3D scene model to establish the skeletal hierarchy of the character assets, and initial position parameters are set for each joint node in the character assets. Based on the skeletal hierarchy, displacement, rotation, and scale parameters are set for each joint node in the character asset, and the motion parameters of the character asset are obtained based on the initial position parameters, displacement parameters, rotation parameters, and scale parameters of each joint node. Set initial position parameters, displacement velocity parameters, and rotation parameters for object assets in the 3D scene model, and determine the motion parameters of the object assets; Set the spatial position parameters, orientation parameters, focal length parameters, and imaging parameters for the camera in the 3D scene model, and determine the lens parameters; Based on the spatial distribution of scene assets, the motion parameters of character assets, the motion parameters of object assets, and the lens parameters of the camera in the 3D scene model, scene parameters, asset motion parameters, and lens parameters are formed.
5. The AI video shot three-dimensional accurate control method based on image processing according to claim 1, characterized in that, The acquisition of the multi-camera parameters specifically includes: Each camera position is determined based on a 3D scene model, and a unique identifier is assigned to each camera position. Initial position parameters, shooting angle parameters, and camera movement trajectory parameters are set for each camera position to obtain the initial camera position parameters for each camera position; Based on the initial camera position parameters, composition requirements, and camera movement requirements, establish the parameter linkage relationship between each camera position; When the parameters of the target camera position are adjusted, the changes in the initial position parameters, shooting angle parameters, and camera movement trajectory parameters after the adjustment of the target camera position are obtained, and the update amounts of the initial position parameters, shooting angle parameters, and camera movement trajectory parameters corresponding to the associated camera position are determined according to the parameter linkage relationship. Based on the update amounts of initial position parameters, shooting angle parameters, and camera movement trajectory parameters, the composition parameters and camera movement parameters of the associated camera positions are updated synchronously, and multi-camera parameters are generated based on the updated parameters of each camera position.
6. The AI video shot three-dimensional accurate control method based on image processing according to claim 1, characterized in that, The optimized composition parameters and camera movement parameters are obtained specifically through: Image feature data, scene parameters, asset motion parameters, lens parameters, multi-camera parameters, as well as lens emotion expression requirements, composition requirements, and camera movement requirements are processed by feature encoding to generate fused feature data; The fused feature data is input into the AI camera control model, which adopts a hybrid architecture combining convolutional neural networks and Transformer networks to extract features and model feature associations from the fused feature data, thereby obtaining scene structure features, asset motion features, and camera features corresponding to each camera position. Based on scene structure features, asset action features, and lens features, combined with lens emotion expression requirements, composition requirements, and camera movement requirements, the shooting angle, focal length, framing, composition parameters, and camera movement parameters of each camera position are initially estimated to generate an initial parameter set for each camera position. Using the initial parameter set of each camera position as the optimization input, a composition evaluation function is constructed based on the requirements for emotional expression and composition of the shot, and a camera movement evaluation function is constructed based on the camera movement requirements, forming a comprehensive optimization objective function; Based on the comprehensive optimization objective function, the shooting angle, focal length, and framing of each camera position are iteratively optimized to generate composition schemes. The composition schemes are then screened to determine the composition schemes that meet the requirements of the emotional expression of the shot, and the corresponding composition parameters and camera movement parameters are output. When receiving instructions to adjust composition or camera movement parameters, the AI lens control model re-inputs the adjusted camera parameters to optimize them, resulting in optimized composition and camera movement parameters.
7. The AI movie shot three-dimensional accurate control method based on image processing according to claim 1, characterized in that, The multi-camera collaborative verification specifically includes: Based on the optimized composition and camera movement parameters, the shot images corresponding to each camera position are rendered to generate a visual image sequence for each camera position. Multi-camera collaborative verification is performed on the visualized image sequence. The multi-camera collaborative verification includes image continuity verification, composition rationality verification, camera movement smoothness verification, and shot emotion expression consistency verification. The image continuity verification judges the smoothness of the image transition by analyzing the changes in the subject position of consecutive frames in the image sequence of each camera. The composition rationality verification determines the composition rationality by calculating the deviation of the subject position, image ratio, and layer relationship in each camera's image from the preset composition requirements. The camera movement smoothness verification judges the smoothness of the camera movement by analyzing the deviation of the changes in the camera movement trajectory and speed of each camera from the optimized camera movement parameters. The shot emotion expression consistency verification judges the consistency of emotion expression by comparing the matching degree of each camera's image features with the director's input shot emotion expression requirements. The verification result status of each camera is generated, indicating whether each verification item meets the preset threshold. Based on the verification results of each camera position, when a verification item fails to meet the preset threshold, the composition parameters and camera movement parameters of the corresponding camera position are adjusted, the updated parameters are reused to render the visualization image sequence of the corresponding camera position, and multi-camera collaborative verification is performed again until all verification items of all cameras meet the preset threshold.
8. The AI video shot three-dimensional accurate control method based on image processing according to claim 1, characterized in that, The structured storage involves organizing and storing the verified 3D scene model, scene parameters, asset motion parameters, camera parameters, multi-camera parameters, composition parameters, and camera movement parameters. Specifically, it involves associating and saving scene assets, object assets, and character assets in the 3D scene model with scene parameters and asset motion parameters. It also involves storing the camera parameters, multi-camera parameters, composition parameters, and camera movement parameters corresponding to each camera position according to the camera position number. Finally, it calls the rendering module to render the image content in the 3D scene model from the shooting perspective of the corresponding camera position, generating a preview of the camera composition for each position.
9. An AI-based three-dimensional precision control system for film and television lenses based on image processing, executing the AI-based three-dimensional precision control method for film and television lenses based on image processing as described in any one of claims 1 to 8, characterized in that, include: The image acquisition module is used to acquire reference images and the director's input requirements for shot emotion expression, composition, and camera movement. The reference images are preprocessed to obtain image feature data. The 3D scene construction module is used to construct 3D scenes based on image feature data and call the film and television 3D asset library to form 3D scene models; The parameter setting module is used to perform skeletal binding and motion parameter settings for object assets and character assets in the 3D scene model, and to parameterize the camera in the 3D scene model to obtain scene parameters, asset motion parameters, and camera parameters. The camera position management module is used to configure camera positions based on the 3D scene model, establish parameter linkage relationships between camera positions, and synchronously update the composition parameters and camera movement parameters of the associated camera positions according to the parameter linkage relationships to obtain multi-camera parameters. The AI lens control module is used to input image feature data, scene parameters, asset action parameters, lens parameters, multi-camera parameters, as well as lens emotion expression requirements, composition requirements and camera movement requirements into the AI lens control model. The AI lens control model optimizes the shooting angle, focal length, framing, composition parameters and camera movement parameters of each camera position to obtain optimized composition parameters and camera movement parameters. The rendering and verification module is used to render the shot images and camera movement effects corresponding to each camera position based on the optimized composition parameters and camera movement parameters, and to perform multi-camera collaborative verification of the rendering results. The storage and preview module is used to structurally store the verified 3D scene model, scene parameters, asset motion parameters, camera parameters, multi-camera parameters, composition parameters, and camera movement parameters, and to export the camera composition preview screens for each camera position.