A three-dimensional animation real-time director system fusing virtual production and AI previsualization

CN122597596APending Publication Date: 2026-08-18QINGDAO HENGXING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610973334.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0004]本发明的目的在于提出一种融合虚拟制片与AI预演的三维动画实时导演系统,通过导演风格行为模型驱动的智能预演与基于场景语义偏差感知的闭环决策机制,解决传统虚拟制片系统中导演艺术风格无法融入自动化推演,以及虚实融合偏差无法自主感知与调控的问题

Benefits of technology

通过空间锚点生成模块从场景语义分布图谱中聚类生成的导演引导锚点,并非单纯的场景几何标记,而是承载了场景内部事件发生可能性密度的语义控制点。AI预演推演模块获取这些锚点后,采用预训练的导演风格行为模型对其进行时序状态编码和场景上下文关联匹配,该模型在马氏决策过程的框架内内化了特定导演在类似场景下的镜头调度偏好、节奏控制习惯和叙事展开逻辑。由此产出的动态预演路径序列,其情节推进节奏和视觉呈现风格与目标导演的艺术个性高度一致,克服了传统预演系统输出路径风格同质化、缺乏叙事张力的缺陷。虚实融合场景状态生成后,智能导演决策模块不是简单地比对几何配准精度或色彩一致性,而是连续追踪该状态在时间维度上的场景语义漂移量,构建场景状态的实际演化轨迹。将该轨迹与动态预演路径序列中预设的期望语义特征向量进行逐关键帧偏离度量,经滑动窗口累计求和与归一化映射后生成导演干预紧迫度指标。该指标反映了当前实拍演绎偏离预期叙事路线的严重程度,系统据此从预设导演动作空间内自动化选择匹配的调控操作,形成导演决策控制信号并反馈至多模态解析模块,调节下一时刻实时场景数据采集的敏感度。这种以语义级叙事偏差为感知对象、以调节前端解析参数为执行手段的闭环链路,使系统在导演介入前即可自主完成对微小失控趋势的纠正,维系了创意意图从预演到实拍的连贯传递。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122597596A_ABST
    Figure CN122597596A_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional animation real-time director system fusing virtual production and AI pre-performance, and belongs to the technical field of three-dimensional animation production. The system comprises a multi-modal analysis module, which is used for multi-modal feature extraction on real-time scene data of a three-dimensional animation virtual film set to obtain a scene semantic distribution atlas; an AI pre-performance deduction module, which is used for calling a pre-trained director style behavior model to perform plot deduction and generate a dynamic pre-performance path sequence; a virtual production integration module, which is used for receiving external live-action signals and performing three-dimensional space fusion calibration according to the dynamic pre-performance path sequence to obtain a virtual-real fusion scene state; and an intelligent director decision module, which is used for performing time sequence change trend analysis on the virtual-real fusion scene state, generating a director decision control signal and feeding back to the multi-modal analysis module to complete closed-loop regulation and control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of 3D animation production technology, specifically to a real-time 3D animation directing system that integrates virtual production and AI pre-visualization. Background Technology

[0002] 3D animation directors have long faced the dilemma of a broken feedback chain between artistic conception and the actual finished product during the creative process. In the traditional animation production workflow, directors plan the camera language and narrative rhythm through storyboards and static pre-visualization. These pre-visualization methods can only provide discrete, non-real-time reference frameworks and cannot dynamically respond to the real-time changes in actor performances, virtual scene adjustments, or camera movements. When the director wants to adjust narrative details on set, the production process must be interrupted, the technical team must manually modify the animation data and re-render, and then another round of review must be organized. The whole process is time-consuming and costly. Although the introduction of virtual production technology has enabled the real-time synthesis of live-action footage and computer-generated images, its core function is limited to visual spatial alignment and image overlay, and does not touch the most core aspects of the director's work: plot construction and rhythm control. Rule-based traditional automatic pre-visualization systems can only trigger fixed animation sequences according to preset conditions, lacking the ability to understand and imitate the director's personal style. The camera movements and editing rhythms they generate are monotonous and far removed from the director's artistic pursuits. The more critical contradiction lies in the fact that when the existing system deviates from the live-action signal and the virtual pre-rehearsal result, it cannot detect whether this deviation has affected the communication of the narrative intention, nor can it trigger corrective measures on its own. This forces the director to continuously monitor the virtual-real fusion state and frequently intervene manually, which seriously restricts creative efficiency.

[0003] This application aims to address two core issues. First, how to transform a director's personal stylistic experience into a computable deductive model, enabling the system to automatically generate a rehearsal path that possesses artistic individuality and matches the current scene context. Second, how to allow the system to automatically perceive the degree of deviation between the real-time fusion state and the expected plot direction during the virtual-real fusion process, and autonomously adjust the sensitivity of scene analysis accordingly, achieving a closed-loop adjustment driven by the director's intent. Summary of the Invention

[0004] The purpose of this invention is to propose a real-time 3D animation director system that integrates virtual production and AI pre-visualization. Through intelligent pre-visualization driven by a director style behavior model and a closed-loop decision-making mechanism based on scene semantic deviation perception, it solves the problems in traditional virtual production systems where the director's artistic style cannot be integrated into automated deduction, and where the deviation between virtual and real fusion cannot be autonomously perceived and controlled.

[0005] To achieve the above objectives, the present invention provides the following technical solution: The present invention provides a real-time 3D animation director system that integrates virtual production and AI pre-visualization. The system forms a closed loop integrating scene understanding, director style deduction, virtual-real fusion, and intelligent decision-making, specifically including: The multimodal parsing module is used to extract multimodal features from real-time scene data of 3D animation virtual studios to obtain scene semantic distribution maps. This module extracts common semantics from geometric meshes, texture images and motion capture data, unifying heterogeneous information into the same spatial coordinate system, so that scene understanding no longer depends on a single modality, and fully captures the relationship between object shape, surface attributes and character action intentions in the studio.

[0006] The spatial anchor point generation module is used to perform spatial event sensitivity clustering in the virtual film set 3D grid based on the scene semantic distribution map to generate editable director guidance anchor points; by detecting areas with drastic semantic changes and adaptively adjusting the grid resolution, it extracts peak positions in semantically dense areas and eliminates neighborhood conflicts, automatically generating compact and narratively valuable director guidance anchor points, providing the director with intuitive and adjustable plot intervention nodes.

[0007] The AI ​​pre-visualization and deduction module is used to invoke a pre-trained director style behavior model to perform plot deduction based on the editable director-guided anchor points, generating a dynamic pre-visualization path sequence. This module performs temporal encoding on the anchor point state sequence, and the director style behavior model outputs the plot development probability distribution based on scene context matching. On this basis, it searches and selects the transition path with the best smoothness, so that the deduction result not only conforms to a specific director style, but also ensures the continuity of camera movement and visual comfort.

[0008] The virtual production integration module is used to receive external live-action signals and perform three-dimensional spatial fusion calibration on the external live-action signals according to the dynamic pre-playing path sequence to obtain the virtual-real fusion scene state; by analyzing the intrinsic and extrinsic parameters of the live-action camera pose, it spatially aligns the live-action camera with the expected trajectory of the virtual camera determined by the pre-playing path, iteratively solves the rigid body transformation to eliminate registration residuals, and performs inverse distortion correction, depth comparison and occlusion weight calculation on the live-action image frames and virtual scene frames, finally achieving pixel-level virtual-real overlay, so that the live-action material and the three-dimensional animation scene are seamlessly integrated in space and time.

[0009] The intelligent director decision-making module is used to analyze the temporal change trend of the virtual-real fusion scene state, generate director decision control signals and feed them back to the multimodal analysis module to complete closed-loop control. This module tracks the drift trajectory of scene semantics within the time window, compares it with the expected semantics of the dynamic pre-playing path sequence, calculates the urgency of director intervention, and then selects control operations from the preset action space and sends them to the multimodal analysis module to adjust the data acquisition sensitivity of the next moment. This allows the director's intention to be injected into the perception link in real time, forming a continuous optimization loop of perception-pre-playing-fusion-feedback.

[0010] As a technical solution of this invention, the multimodal parsing module specifically performs the following: vertex normal deviation statistics are performed on the geometric mesh data in the real-time scene data to obtain a surface geometric feature field; color space distribution statistics are performed on the texture data to obtain a texture semantic label map; joint trajectory fitting is performed on the motion capture data to obtain a motion intent encoding vector; the surface geometric feature field, texture semantic label map, and motion intent encoding vector are mapped to a unified spatial coordinate system and aggregated element-by-element to generate the scene semantic distribution map. Thus, the originally scattered geometric, texture, and motion information is compactly represented as a unified cross-modal semantic description, effectively improving the completeness and consistency of semantic details in complex scene environments.

[0011] As another technical solution of the present invention, the spatial anchor point generation module specifically performs the following steps: spatial gradient calculation is performed on the scene semantic distribution map to obtain a scene semantic change gradient field; the virtual set 3D mesh is non-uniformly divided according to the gradient field to obtain an adaptive resolution mesh region; semantic density peak detection is performed within each adaptive resolution mesh region to obtain initial anchor point candidate positions; spatial neighborhood conflict resolution and merging are performed on the initial anchor point candidate positions to generate the editable director guidance anchor points. This process ensures that anchor points are concentrated in areas where scene semantics change significantly, and the anchor point density is adaptively adjusted according to the set complexity, avoiding redundant anchor points while highlighting key director entry points.

[0012] Preferably, the AI ​​pre-visualization and deduction module specifically performs the following: temporal state encoding of editable director-guided anchor points to obtain anchor point state feature sequences; inputting these anchor point state feature sequences into a director style behavior model, which performs scene context association matching on the anchor point state feature sequences and outputs a plot development probability distribution; performing path search on the editable director-guided anchor points based on the plot development probability distribution to obtain a set of candidate plot transition paths; evaluating the smoothness of the candidate plot transition paths and sorting them according to the evaluation results, determining the top-ranked candidate plot transition path as the dynamic pre-visualization path sequence. This solution utilizes a style model to replace manually setting keyframes, enabling plot deduction to automatically follow a preset director style and ensuring that the deduced shot paths are smooth and without jumps in spatial movement, significantly reducing the reliance on human experience in pre-visualization production.

[0013] Preferably, the virtual production integration module specifically performs the following steps: receiving external live-action signals and analyzing their camera intrinsic parameter matrix and camera extrinsic pose to obtain live-action spatial positioning parameters; extracting the desired trajectory of the virtual camera based on the dynamic pre-play path sequence to obtain virtual spatial positioning parameters; calculating the coordinate transformation residual between the live-action spatial positioning parameters and the virtual spatial positioning parameters, and updating the spatial mapping matrix based on the coordinate transformation residual to obtain fusion spatial alignment parameters; and performing pixel-level blending and overlay of external live-action image frames and 3D virtual scene frames according to the fusion spatial alignment parameters to generate a virtual-real fusion scene state. When calculating the coordinate transformation residual, an iterative nearest-point algorithm can be used to solve the rigid body transformation matrix of the matching feature point pairs, and the root mean square of the registration error can be used as the coordinate transformation residual, thereby quickly converging to obtain an accurate spatial mapping relationship, ensuring that the alignment accuracy between the live-action and virtual scenes meets the requirements of real-time compositing.

[0014] Furthermore, when the virtual production integration module performs pixel-level blending and overlay of external live-action image frames and 3D virtual scene frames according to the fusion space alignment parameters, it specifically includes: performing inverse distortion correction and color space conversion on the external live-action image frames according to the fusion space alignment parameters to obtain a pre-processed live-action image; extracting the corresponding frame from the 3D virtual scene frame cache according to the time stamp of the dynamic pre-playback path sequence and performing viewport cropping to obtain an aligned virtual scene frame; performing a depth comparison test on the pre-processed live-action image and the aligned virtual scene frame to generate a pixel-by-pixel occlusion weight map; and using the pixel-by-pixel occlusion weight map to perform weighted fusion output on the pre-processed live-action image and the aligned virtual scene frame to generate a virtual-real fusion scene state. In the depth comparison test, the scene depth value corresponding to each pixel in the pre-processed live-action image is compared with the rendering depth value of the corresponding pixel in the aligned virtual scene frame, and the occlusion weight is assigned according to the relationship between the depth values, thereby accurately solving the occlusion relationship between the live-action foreground and the virtual object and eliminating the sense of penetration and misalignment during compositing.

[0015] Preferably, the intelligent director decision-making module specifically performs the following: statistically tracking the semantic drift of the virtual-real fusion scene state within a continuous time window to obtain the scene state evolution trajectory; measuring the deviation between the scene state evolution trajectory and the dynamic pre-playing path sequence to obtain a director intervention urgency index; selecting matching control operations from the preset director action space based on the director intervention urgency index to generate a director decision control signal; and sending the director decision control signal to the multimodal parsing module to adjust the sensitivity of real-time scene data acquisition at the next moment. During the deviation measurement process, multiple pre-playing keyframes can be set on the dynamic pre-playing path sequence, and a desired semantic feature vector can be configured for each pre-playing keyframe. The actual semantic feature vector corresponding to the time sequence is extracted from the scene state evolution trajectory, and the semantic deviation distance between the two is calculated to obtain a single-frame deviation sequence. Then, a sliding window summation is performed on the single-frame deviation sequence, followed by normalization mapping, to obtain the director intervention urgency index. In this way, the system can promptly perceive when the virtual-real fusion result deviates from the preset style expectation and automatically apply control operations matching the degree of deviation, making the on-set data acquisition focus more on semantic regions with larger deviations, thus achieving precise control of the director's intentions.

[0016] This invention deeply integrates virtual production with AI pre-visualization through multimodal semantic parsing, spatially sensitive anchor point generation, style-driven plot deduction, virtual-real fusion spatial calibration, and closed-loop decision-making based on deviation metric. In the real-time filming environment, it endows the system with the ability to perceive scene semantics, automatically generate director-level pre-visualization paths, seamlessly integrate live-action elements, and dynamically adjust data acquisition based on fusion feedback. This not only ensures the spatial and semantic consistency of virtual-real synthesis but also significantly improves the director's real-time control response speed and creative freedom.

[0017] The technical effects and advantages provided by the present invention in the above technical solution are as follows: The director-guided anchors generated by the spatial anchor generation module from the scene semantic distribution map are not merely geometric markers of the scene, but semantic control points that carry the probability density of events occurring within the scene. After acquiring these anchors, the AI ​​pre-visualization and deduction module uses a pre-trained director style behavior model to encode their temporal states and match them with scene context. This model internalizes a specific director's camera scheduling preferences, pacing habits, and narrative development logic in similar scenes within the framework of the Mahalanobis decision process. The resulting dynamic pre-visualization path sequence has a plot progression rhythm and visual presentation style that is highly consistent with the artistic personality of the target director, overcoming the shortcomings of traditional pre-visualization systems that output path styles that are homogeneous and lack narrative tension. After the virtual-real fusion scene state is generated, the intelligent director decision module does not simply compare geometric registration accuracy or color consistency, but continuously tracks the scene semantic drift of the state in the time dimension, constructing the actual evolution trajectory of the scene state. The trajectory is then compared with the preset expected semantic feature vector in the dynamic pre-visualization path sequence using a keyframe-by-keyframe deviation measurement. After cumulative summation and normalization mapping through a sliding window, a director intervention urgency index is generated. This indicator reflects the severity of the deviation between the current live-action performance and the intended narrative path. Based on this, the system automatically selects matching control operations from the preset director's action space, forming a director decision control signal and feeding it back to the multimodal analysis module to adjust the sensitivity of real-time scene data acquisition for the next moment. This closed-loop link, which uses semantic-level narrative deviation as the object of perception and adjusts front-end analysis parameters as the means of execution, allows the system to autonomously correct minor trends of loss of control before the director intervenes, maintaining the coherent transmission of creative intent from rehearsal to live-action shooting. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0019] Figure 1 This is a schematic diagram of the structure of a real-time 3D animation director system that integrates virtual production and AI pre-visualization; Figure 2 This is a schematic diagram of the process for generating a semantic distribution map of a multimodal scene; Figure 3 This is a flowchart of the spatial anchor point generation method; Figure 4 It is an AI pre-visualization and deduction flowchart based on director-guided anchor points; Figure 5 This is a flowchart of the virtual-real fusion processing of the virtual production integration module; Figure 6 This is a flowchart of the closed-loop control process of the intelligent director decision-making module. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] See Figure 1 This invention provides a real-time 3D animation directing system integrating virtual production and AI pre-visualization. The system includes a multimodal analysis module, a spatial anchor point generation module, an AI pre-visualization and deduction module, a virtual production integration module, and an intelligent director decision-making module. The multimodal analysis module extracts multimodal features from real-time scene data of a 3D animation virtual studio to obtain a scene semantic distribution map. The spatial anchor point generation module performs spatial event sensitivity clustering in the 3D grid of the virtual studio based on the scene semantic distribution map, generating editable director guidance anchor points. The AI ​​pre-visualization and deduction module uses the editable director guidance anchor points to invoke a pre-trained director style behavior model to perform plot deduction, generating a dynamic pre-visualization path sequence. The virtual production integration module receives external live-action signals and performs 3D spatial fusion calibration on the external live-action signals according to the dynamic pre-visualization path sequence to obtain a virtual-real fusion scene state. The intelligent director decision-making module analyzes the temporal change trend of the virtual-real fusion scene state, generates director decision control signals, and feeds them back to the multimodal analysis module to complete closed-loop control.

[0022] Example 1: In specific implementation, please refer to Figure 2 The multimodal parsing module extracts multimodal features from the real-time scene data of the 3D animation virtual studio to obtain a scene semantic distribution map. The real-time scene data of the 3D animation virtual studio includes geometric mesh data, texture data, and motion capture data. These three types of data are synchronously acquired and input into the multimodal parsing module in each frame.

[0023] For geometric mesh data, the multimodal analysis module performs vertex normal deviation statistics to obtain the surface geometric feature field. The surface geometric feature field represents a scalar field of the local shape undulations of the 3D model surface, with each vertex position corresponding to a geometric eigenvalue. For any vertex in the geometric mesh data, a set of neighboring vertices directly connected to that vertex via edges is obtained. The dot product of the vertex's normal vector and the normal vectors of each neighboring vertex in the set is calculated, and the geometric eigenvalue of the vertex is calculated according to the following formula: ; in, This represents the vertex with index 'a' in the geometric mesh data. Represents vertices The set of neighboring vertices, This represents the number of neighboring vertices in the set of neighboring vertices. Represents vertices The unit normal vector, Let represent the unit normal vector of the vertex with index b in the set of neighboring vertices. This represents the dot product of two unit normal vectors. Represents vertices The geometric eigenvalues ​​are obtained by performing the above calculations on all vertices in the geometric mesh data. The geometric eigenvalues ​​of each vertex are then used to construct the surface geometric feature field.

[0024] For texture data, the multimodal parsing module performs color space distribution statistics to obtain a texture semantic label map. In specific implementation, the color value of each pixel in the texture data is converted from the RGB color space to the Lab color space. Cluster analysis is then performed on the color values ​​of all pixels in the Lab color space using the K-means clustering algorithm, with the number of clusters set to 5. The reason for setting the number of clusters to 5 is that the surface materials of the 3D models carrying textures in 3D animation virtual studios typically include five categories: skin, fabric, metal, wood, and background. Clustering automatically divides pixels into clusters corresponding to the material categories. After clustering, a semantic label is assigned to each cluster. The semantic label is an integer identifier from 1 to 5, with different clusters corresponding to different semantic labels. The semantic labels of each pixel in the texture data are arranged according to the original pixel position to form a texture semantic label map. The texture semantic label map has the same width and height resolution as the texture data.

[0025] For motion capture data, the multimodal parsing module performs joint trajectory fitting to obtain motion intent encoding vectors. Motion capture data contains the three-dimensional spatial coordinates of multiple joints in the human skeleton for each frame. For each joint, the three-dimensional coordinate sequence of that joint over several consecutive frames is extracted, and a cubic Bézier curve is used to fit this sequence, resulting in a spatial trajectory curve for that joint. Cubic Bézier curve fitting is performed on all joints to obtain a set of spatial trajectory curves. This set of fitted spatial trajectory curves is then discretized and sampled to obtain a fixed-length joint trajectory feature vector. This feature vector contains the three-dimensional coordinate values ​​of all joints resampled at fixed frame intervals.

[0026] The joint trajectory feature vector is input into a pre-trained autoencoder, which outputs a motion intent encoded vector. The autoencoder consists of an encoder and a decoder. The encoder is composed of three fully connected layers connected sequentially, with 256, 128, and 64 neurons respectively. Each layer is followed by a batch normalization layer and a linear rectified activation function. The decoder is also composed of three fully connected layers connected sequentially, with 128, 256 neurons respectively, and the same number of neurons as the joint trajectory feature vector. The first two layers are followed by a batch normalization layer and a linear rectified activation function, while the last layer has no activation function. The autoencoder is trained through the following process: A large number of human motion capture sequences are collected as training samples. Each training sample is fitted with the joint trajectory and discretized to obtain a joint trajectory feature vector. This joint trajectory feature vector is then input into the encoder to obtain a motion intent encoding vector. The decoder then reconstructs the motion intent encoding vector into a reconstructed trajectory feature vector. The mean squared error between the joint trajectory feature vector and the reconstructed trajectory feature vector is used as the loss function. An adaptive moment estimation optimizer updates the weight parameters of the fully connected layers in both the encoder and decoder. The initial learning rate is set to 0.001, the batch size to 64, and training is iteratively continued until the loss function converges. After training, the encoder is retained. When processing real-time scene data, the real-time joint trajectory feature vector is input into the encoder, and the encoder outputs the motion intent encoding vector.

[0027] After obtaining the surface geometric feature field, texture semantic label map, and action intent encoding vector, the multimodal parsing module maps these three modal features to a unified spatial coordinate system for element-by-element association and aggregation, generating a scene semantic distribution map. In specific implementation, the unified spatial coordinate system is based on the world coordinate system of the virtual set. The surface geometric feature field is originally defined at the vertex positions of the 3D model, with each vertex having its own world coordinates; by aligning the world coordinates of each vertex in the surface geometric feature field to the grid nodes of the unified spatial coordinate system, the mapping of geometric features is completed. The texture semantic label map is originally a 2D image; through texture mapping relationships, the semantic label of each pixel in the texture semantic label map is back-projected onto the surface of the 3D model to obtain the semantic label value of each point on the 3D model surface, and then the semantic label value is stored in the corresponding grid node of the unified spatial coordinate system according to the world coordinates of the surface point. The motion intent encoding vector is determined by the joint positions of the human skeleton. This vector is then diffused from the joint positions to the surrounding spatial region using radial basis functions. For any spatial point in the unified spatial coordinate system, the Euclidean distance from that point to all joint positions is taken. The reciprocal of each Euclidean distance is used as the weight to weight the motion intent encoding vectors corresponding to all joint positions, resulting in a weighted sum for that spatial point. At each grid node in the unified spatial coordinate system, the surface geometric feature values, texture semantic label values, and components of the diffusion encoding vector are concatenated to form the local semantic feature vector of that grid node. The local semantic feature vectors of all grid nodes constitute a scene semantic distribution map.

[0028] Example 2: In specific implementation, please refer to Figure 3 After the spatial anchor point generation module obtains the scene semantic distribution map output by the multimodal parsing module, it performs spatial event sensitivity clustering in the virtual studio 3D mesh based on the scene semantic distribution map to generate editable director guidance anchor points.

[0029] In practical implementation, the spatial anchor point generation module performs spatial gradient calculations on the scene semantic distribution map to obtain the scene semantic change gradient field. The scene semantic distribution map is composed of the local semantic feature vectors of all grid nodes in a unified spatial coordinate system. For any grid node in the unified spatial coordinate system, the adjacent grid nodes of that grid node in the three coordinate axis directions are taken, and the partial derivative vectors of the local semantic feature vector along the X-axis, Y-axis, and Z-axis are calculated respectively using the central difference method. The L2 norm of the three partial derivative vectors is summed to obtain the gradient magnitude at that grid node. The above central difference calculation is performed on all grid nodes in the unified spatial coordinate system one by one to obtain the gradient magnitude corresponding to each grid node. The gradient magnitudes of all grid nodes constitute the scene semantic change gradient field.

[0030] In practice, the spatial anchor point generation module performs non-uniform subdivision of the virtual scene's 3D mesh based on the scene's semantic change gradient field, obtaining an adaptive resolution mesh region. The adaptive resolution mesh region subdivision employs a multi-level mesh subdivision strategy, with a subdivision level set to three layers. The basis for setting the subdivision level to three layers is as follows: the gradient magnitude of all mesh nodes in the scene's semantic change gradient field is divided into three intervals at equal intervals from minimum to maximum. The low gradient interval corresponds to the background area with gentle semantic changes in the scene, the medium gradient interval corresponds to the edge area of ​​character movement or interaction, and the high gradient interval corresponds to the core area where events occur. These three intervals correspond to three subdivision levels, from coarse to fine. The meshing process is as follows: Initialize the virtual scene's 3D mesh as a basic cube mesh, with a side length equal to half the side length of the virtual scene's maximum bounding box. Traverse each basic cube mesh region and calculate the average gradient magnitude of all mesh nodes within that region. If the average value falls within a low gradient range, the basic cube mesh region maintains its current resolution. If the average value falls within a medium gradient range, divide the basic cube mesh region into 8 equal-sized sub-mesh regions along the midpoints of its edges. If the average value falls within a high gradient range, recursively divide the basic cube mesh region twice along the midpoints of its edges, resulting in 64 equal-sized sub-mesh regions. All sub-mesh regions constitute the adaptive resolution mesh region.

[0031] The threshold for determining the segmentation level is determined by the statistical distribution of gradient magnitudes in the scene semantic change gradient field. The gradient magnitudes of all grid nodes in the scene semantic change gradient field are sorted, and the 33rd percentile of the gradient magnitude is taken as the first threshold γ1, and the 67th percentile as the second threshold γ2. The two percentiles are chosen based on roughly dividing the gradient magnitudes into three cumulative intervals. For any basic cube grid region, the arithmetic mean of the gradient magnitudes of all grid nodes within that region is calculated, and this arithmetic mean is denoted as... The subdivision level of the basic cubic mesh region is determined according to the following formula. : ; in, This indicates the partitioning level assigned to the basic cubic mesh region R. γ1 represents the arithmetic mean of the gradient magnitudes of all grid nodes within the basic cubic grid region R; γ2 represents the first threshold obtained by taking the 33rd percentile after sorting the gradient magnitudes of all grid nodes in the scene semantic change gradient field; and γ3 represents the second threshold obtained by taking the 67th percentile after sorting the gradient magnitudes of all grid nodes in the scene semantic change gradient field. This represents the first subdivision level, corresponding to the original resolution without further subdivision. This indicates the second subdivision level, which divides the basic cube mesh region into 8 equal-sized sub-mesh regions along the midpoint of each edge. This represents the third subdivision level, which involves recursively subdividing the basic cube mesh region twice along the midpoints of each edge to obtain 64 equally sized sub-mesh regions. The above subdivision level determination formula is applied to each basic cube mesh region in the virtual scene 3D mesh to complete the non-uniform subdivision of the virtual scene 3D mesh, resulting in an adaptive resolution mesh region.

[0032] In practical implementation, the spatial anchor point generation module performs semantic density peak detection within each adaptive resolution grid region to obtain initial anchor point candidate positions. Semantic density peak detection employs a mean-shift algorithm to search for local density maxima in the semantic feature space. For a single adaptive resolution grid region, local semantic feature vectors of all grid nodes within that region are obtained, constructing a semantic feature vector set. A Gaussian kernel function is defined as the density estimation kernel function, with its bandwidth parameter set to the average Euclidean distance from all grid nodes within the adaptive resolution grid region to the region center. In the semantic feature space, the position of each local semantic feature vector in the semantic feature vector set is used as the starting search point for the mean-shift algorithm. Iterative calculation of the shift vector and position updates are performed: in each iteration, with the current search point as the center, the weighted average of all local semantic feature vectors within the bandwidth radius is calculated, with the weights determined by the Gaussian kernel function value. The current search point is then updated to this weighted average. This iteration is repeated until the position offset of the search point between two adjacent iterations is less than the convergence threshold, which is set to one-thousandth of the bandwidth parameter. When multiple initial search points in the semantic feature vector set converge to the same feature space location, the convergence location is taken as the semantic density peak location. The semantic density peak location is then mapped back to the unified spatial coordinate system to obtain the three-dimensional spatial coordinates of the initial anchor point candidate locations within the adaptive resolution grid region.

[0033] In practice, the spatial anchor point generation module resolves and merges spatial neighborhood conflicts of initial anchor point candidate positions to generate editable director guidance anchor points. The specific method for resolving and merging spatial neighborhood conflicts is as follows: Initial anchor point candidate positions output from all adaptive resolution grid regions are aggregated into a single candidate position set, with a spatial merging threshold set to one-quarter of the side length of the virtual studio's 3D grid base cube. Each pair of initial anchor point candidate positions in the candidate position set is traversed, and the Euclidean distance between the two initial anchor point candidate positions is calculated. When the Euclidean distance is less than the spatial merging threshold, it is determined that the two initial anchor point candidate positions have a spatial conflict. For two initial anchor point candidate positions with a spatial conflict, the density peak response values ​​recorded in the semantic density peak detection step are compared, i.e., the Gaussian kernel density estimate at the convergence position of the mean-shift algorithm. The initial anchor point candidate position with the larger density peak response value is retained, while the initial anchor point candidate position with the smaller density peak response value is discarded. After comparing and discarding all initial anchor point candidates with spatial conflicts, the remaining initial anchor point candidates constitute an editable director-guided anchor point set. Each editable director-guided anchor point contains two attributes: three-dimensional spatial coordinates and a density peak response value. The editable director-guided anchor point set is provided to the director through the system interface for viewing, selecting, moving, and parameter adjustment.

[0034] Example 3: In specific implementation, please refer to Figure 4 After the AI ​​pre-show simulation module obtains the set of editable director guidance anchors output by the spatial anchor generation module, it calls the pre-trained director style behavior model based on the editable director guidance anchors to perform plot simulation and generate a dynamic pre-show path sequence.

[0035] In practical implementation, the AI ​​pre-visualization and deduction module performs temporal state encoding on the editable director-guided anchor points to obtain an anchor point state feature sequence. Each editable director-guided anchor point in the set contains three-dimensional spatial coordinates and a density peak response value. For the current frame, all editable director-guided anchor points at the current frame and several historical frames backward are obtained and arranged in chronological order according to timestamps to form an anchor point time sequence. Graph structure construction is performed on the set of editable director-guided anchor points at each moment in the anchor point time sequence: using each anchor point in the set as a graph node, the Euclidean distance between any two graph nodes is calculated. When the Euclidean distance is less than a connection threshold, the two graph nodes are connected by an edge. The connection threshold is set to half the side length of the virtual studio's three-dimensional grid base cube. The basis for setting the connection threshold to half is that the probability of contextual association between anchor points within this distance range meets the pre-visualization requirements. The anchor point spatial relationship graph is constructed from all graph nodes and all edges. The anchor point spatial relationship graph is input into a graph neural network encoder, which contains three graph convolutional layers. Each graph convolutional layer is followed by a linear rectified activation function. The message passing method of the graph convolutional layers is to aggregate the features of the neighboring nodes of each graph node by averaging them and then concatenating them with the graph node's own features, followed by a linear transformation through a learnable weight matrix. The output feature vectors of the three graph convolutional layers have dimensions of 128, 64, and 32, respectively. The 32-dimensional feature vector output by the three graph convolutional layers is used as the global graph representation vector of the anchor point spatial relationship graph, and then mapped to a 16-dimensional temporal state vector through a fully connected layer. The above graph neural network encoding process is performed on the anchor spatial relationship graph at each time step in the anchor time series, resulting in a set of 16-dimensional temporal state vectors arranged in chronological order. This set of 16-dimensional temporal state vectors is input into a bidirectional long short-term memory (LSTM) network, which consists of a forward LSM layer and a backward LSM layer, each with 32 hidden units. The forward LSM layer processes the temporal state vector sequence in ascending chronological order, while the backward LSM layer processes it in descending chronological order. The hidden state vector of the forward LSM layer at the last time step is concatenated with the hidden state vector of the backward LSM layer at the first time step, resulting in a 64-dimensional anchor state feature sequence. The learnable weight matrix of the graph neural network encoder and the weight parameters of the bidirectional LSM network are obtained through subsequent joint end-to-end training with the director's style behavior model.

[0036] In practical implementation, the AI ​​pre-visualization and deduction module inputs the anchor point state feature sequence into the director style behavior model. The director style behavior model performs scene context association matching on the anchor point state feature sequence and outputs the probability distribution of plot development. The core architecture of the director style behavior model is a conditional variational autoencoder, which includes a conditional encoder, a latent variable sampling layer, and a plot decoder. The conditional encoder consists of a multilayer perceptron, which contains three fully connected layers with 128, 64, and 32 neurons respectively. It receives the anchor point state feature sequence as conditional input and outputs an anchor point conditional feature vector. The latent variable sampling layer uses the anchor point conditional feature vector as a condition to generate a Gaussian distribution mean vector and standard deviation vector, both with a dimension of 16. A 16-dimensional plot latent variable vector is sampled from the Gaussian distribution. The plot decoder consists of a multilayer perceptron, which contains three fully connected layers with 64, 128 neurons, and the number of categories in the plot development probability distribution, respectively. The plot decoder receives the spliced ​​plot latent variable vector and anchor condition feature vector, and outputs the plot development probability distribution, which contains several categories. Each category corresponds to a plot development type, including camera zoom in, camera zoom out, camera pan, character movement, character interaction, and scene switching.

[0037] The training process of the conditional variational autoencoder is as follows: Camera and character scheduling sequences from multiple directors in different animation projects are collected as training data. Each scheduling sequence is extracted as an anchor point time series, and anchor point state feature sequence samples are generated according to the aforementioned temporal state encoding steps. The anchor point state feature sequence samples correspond to the actual plot development type as labels. During training, the anchor point state feature sequence samples are input into the conditional encoder to obtain the anchor point conditional feature vector. The latent variable sampling layer generates a plot latent variable vector based on the anchor point conditional feature vector. The plot latent variable vector is concatenated with the anchor point conditional feature vector and input into the plot decoder to output the plot development probability distribution. The loss function consists of two parts: the cross-entropy classification loss between the plot development probability distribution and the actual plot development type label, and the KL divergence loss between the Gaussian distribution generated by the latent variable sampling layer and the standard normal distribution. The weight coefficient of the KL divergence loss is set to 0.1. The basis for setting the weight coefficient to 0.1 is that the regularity of the latent variable space and the plot classification accuracy are balanced under this coefficient. An adaptive moment estimation optimizer is used to update all parameters of the conditional encoder, plot decoder, and latent variable sampling layer. Simultaneously, the parameters of the graph neural network encoder and bidirectional long short-term memory network in the temporal state encoding step are also updated. The initial learning rate is set to 0.0005, and the batch size is 32. Iterative training is performed until the loss function converges. After training, the parameters of the conditional encoder, latent variable sampling layer, and plot decoder of the director's style behavior model are frozen and retained, as are the parameters of the graph neural network encoder and bidirectional long short-term memory network in the temporal state encoding step.

[0038] In practice, the AI ​​pre-visualization and deduction module searches for paths to editable director-guided anchor points based on the probability distribution of plot development, obtaining a set of candidate plot transition paths. The path search employs a beam search algorithm with a beamwidth set to 5. This beamwidth setting of 5 represents a trade-off between search efficiency and the diversity of candidate paths. The search process is as follows: Using the set of editable director-guided anchor points corresponding to the current frame as the initial state node, the five plot development types with the highest probability values ​​are selected from the plot development probability distribution, serving as five path branches starting from the initial state node. For each path branch, the corresponding virtual camera motion parameter template and character displacement parameter template are called according to the selected plot development type to generate the next state anchor point configuration, which is then used as the next-level state node. At the next-level state node, the next state anchor point configuration is input into the temporal state encoding step to obtain a new anchor point state feature sequence, which is then input again into the director style behavior model to obtain a new plot development probability distribution. The five plot development types with the highest probability values ​​are selected to continue expanding the path. This expansion process is repeated until the preset plot deduction depth is reached, which is set to 8 steps. Each complete path from the initial state node to the final state node constitutes a candidate plot transition path, and all candidate plot transition paths form a candidate plot transition path set.

[0039] In practical implementation, the AI ​​pre-show simulation module performs path smoothness evaluation on the candidate plot transition path set and sorts them according to the evaluation results, determining the top-ranked candidate plot transition path as the dynamic pre-show path sequence. The path smoothness evaluation comprehensively considers three dimensions of smoothness cost: camera motion acceleration cost, viewpoint switching frequency cost, and character movement coherence cost. Camera motion acceleration cost is the cumulative sum of the absolute values ​​of the instantaneous acceleration of the virtual camera between adjacent state nodes in the candidate plot transition path. Viewpoint switching frequency cost is the normalized value of the number of switching times between adjacent state nodes in the candidate plot transition path where the plot development type is camera zoom in, camera zoom out, or camera pan. Character movement coherence cost is the frame-by-frame displacement variance of the character displacement in the candidate plot transition path. For each candidate plot transition path in the candidate plot transition path set, the path smoothness score is calculated according to the following formula: ; in, This represents the path smoothness score of the candidate plot transition path p. A smaller value indicates a smoother path; This represents the camera motion acceleration cost of the candidate plot transition path p; The cost of perspective switching frequency represents the candidate plot transition path p. This represents the cost of character movement coherence in candidate plot transition path p. The weighting coefficient representing the cost of camera motion acceleration is set to 0.5; The weighting coefficient representing the cost of viewpoint switching frequency is set to 0.3; The weighting coefficient representing the cost of character movement coherence is set to 0.2. The three weighting coefficients are set based on the principle that camera movement smoothness has the highest impact on the viewing experience, followed by the frequency of perspective switching, while character movement coherence has a relatively low impact. A path smoothness score is calculated for each candidate plot transition path in the candidate plot transition path set. All candidate plot transition paths are then sorted in ascending order of path smoothness score, and the candidate plot transition path with the highest score is selected as the dynamic pre-draft path sequence. The dynamic pre-draft path sequence includes the timestamp of each deduction step, virtual camera pose parameters, character position parameters, and the corresponding plot development type label.

[0040] Example 4: In specific implementation, please refer to Figure 5 The virtual production integration module receives external live-action signals and performs three-dimensional spatial fusion calibration on the external live-action signals according to the dynamic pre-play path sequence to obtain the virtual-real fusion scene state.

[0041] In practical implementation, the virtual production integration module receives external live-action signals and parses the camera intrinsic parameter matrix and camera extrinsic pose of the signals to obtain the live-action spatial positioning parameters. The external live-action signals are input to the virtual production integration module from the live-action camera via a video capture card. These signals contain a sequence of live-action image frames and corresponding camera metadata. The camera metadata is acquired in real-time by the live-action camera tracking system, which continuously outputs the camera's position coordinates and attitude quaternions in the world coordinate system using infrared optical tracking or inertial sensing tracking. For each frame of the external live-action signal, the virtual production integration module extracts the camera intrinsic parameter matrix from the camera metadata. This matrix includes focal length parameters, principal point coordinate parameters, and pixel aspect ratio parameters. The focal length parameter is expressed in pixels, and the principal point coordinate parameters represent the pixel coordinates of the projection point of the image plane's optical axis. Simultaneously, the virtual production integration module extracts the camera extrinsic pose from the camera metadata. This pose includes a translation vector and a rotation matrix. The translation vector represents the three-dimensional position of the camera's optical center in the world coordinate system, and the rotation matrix represents the rotation transformation of the camera coordinate system relative to the world coordinate system. The camera intrinsic parameter matrix and the camera extrinsic parameter pose together constitute the spatial positioning parameters for real-world shooting.

[0042] In practice, the virtual production integration module extracts the desired trajectory of the virtual camera based on the dynamic pre-visualization path sequence to obtain virtual space positioning parameters. The dynamic pre-visualization path sequence is output by the AI ​​pre-visualization and deduction module, and includes a timestamp for each deduction step, virtual camera pose parameters, character position parameters, and a corresponding plot development type label. The virtual production integration module aligns the deduction step timestamps in the dynamic pre-visualization path sequence with the frame timestamps of the external live-action signal, using a frame sequence number synchronization mapping method: the frame rate of the external live-action signal is 24 frames per second, and the time interval between deduction steps in the dynamic pre-visualization path sequence is also set to one twenty-fourth of a second, ensuring that each live-action frame corresponds to one deduction step. For each aligned deduction step, the virtual production integration module extracts the virtual camera pose parameters, which include the virtual camera's three-dimensional position coordinates in the virtual scene's world coordinate system and the virtual camera's orientation quaternion. Simultaneously, the virtual camera intrinsic parameter matrix is ​​read from the virtual camera settings of the 3D animation software. This matrix includes virtual focal length parameters and virtual principal point coordinate parameters, which use the same normalization reference as the focal length and principal point coordinate parameters of the live camera. The virtual camera intrinsic parameter matrix and the virtual camera pose parameters together constitute the virtual spatial positioning parameters.

[0043] In practical implementation, the virtual production integration module calculates the coordinate transformation residual between the real-world spatial positioning parameters and the virtual spatial positioning parameters, and updates the spatial mapping matrix based on the coordinate transformation residual to obtain the fused spatial alignment parameters. The calculation of the coordinate transformation residual uses an iterative nearest-neighbor algorithm to solve for the rigid body transformation matrix of the matching feature point pairs in the real-world and virtual spatial positioning parameters, and the root mean square of the registration error after solving is used as the coordinate transformation residual. The matching feature point pairs are obtained as follows: scale-invariant feature transformation (SMT) points are extracted from the real-world image frames corresponding to the real-world spatial positioning parameters, and similarly, scale-invariant SMT points are extracted from the virtual scene frames corresponding to the virtual spatial positioning parameters. Feature descriptor matching is performed on the real-world and virtual scale-invariant SMT points, using Euclidean distance nearest neighbor matching combined with a ratio test to screen reliable matching point pairs. The threshold for the ratio test is set to 0.75, based on the premise that this threshold can filter out most mismatches while retaining a sufficient number of correct matching point pairs. The screened reliable matching point pairs are then used as the set of matching feature point pairs.

[0044] The solution process of the iterative nearest point algorithm is as follows: Four pairs of matching feature points are randomly selected from the set of matching feature point pairs. A rigid body transformation matrix is ​​calculated using singular value decomposition. The rigid body transformation matrix includes rotation matrix components and translation vector components. This rigid body transformation matrix is ​​then applied to the coordinates of all matching feature point pairs in the real-world spatial positioning parameters. The root mean square (RMS) Euclidean distance between the transformed real-world feature point coordinates and the corresponding virtual feature point coordinates is calculated. This random selection and rigid body transformation matrix calculation process is repeated 200 times. The rigid body transformation matrix corresponding to the iteration with the smallest RMS Euclidean distance is selected as the optimal rigid body transformation matrix, and the RMS Euclidean distance corresponding to the optimal rigid body transformation matrix is ​​used as the coordinate transformation residual. The iteration count is set to 200 because the coordinate transformation residual converges to a stable value within this number of iterations, and the computation time is controllable. The spatial mapping matrix is ​​updated by multiplying the current spatial mapping matrix from the real-world coordinate system to the virtual coordinate system with the optimal rigid body transformation matrix to obtain the updated spatial mapping matrix. This updated spatial mapping matrix is ​​the fused spatial alignment parameter.

[0045] In practice, the virtual production integration module performs pixel-level blending and overlay of external live-action image frames and 3D virtual scene frames according to the fusion space alignment parameters to generate a virtual-real fusion scene state. The pixel-level blending and overlay includes four sub-steps.

[0046] The first sub-step involves performing inverse distortion correction and color space conversion on the external real-shot image frame based on the fusion spatial alignment parameters to obtain a preprocessed real-shot image. Inverse distortion correction utilizes the radial and tangential distortion coefficients in the camera's intrinsic parameter matrix to perform inverse distortion mapping on the coordinates of each pixel in the real-shot image frame, obtaining the distorted pixel position. The color value of the distorted pixel is then obtained through bilinear interpolation. Color space conversion transforms the distorted real-shot image frame from the camera's original color space to a linear standard red-green-blue color space. The camera's original color space is determined by the real-shot camera's color profile, and the linear standard red-green-blue color space adopts the sRGB linear component definition established by the International Commission on Illumination (ICI).

[0047] The second sub-step is as follows: Extract the corresponding frame from the 3D virtual scene frame cache based on the time stamp of the dynamic pre-rendering path sequence and perform viewport clipping to obtain the aligned virtual scene frame. The 3D virtual scene frame cache stores the sequence of virtual scene frames pre-rendered by the animation rendering engine, with each frame accompanied by a time stamp tag. The virtual production integration module uses the time stamp of the external live-action image frame as an index to search for a virtual scene frame with a completely equal time stamp in the 3D virtual scene frame cache, and reads this virtual scene frame as the virtual scene frame to be aligned. Based on the updated spatial mapping matrix included in the fusion spatial alignment parameters, the virtual camera frustum corresponding to the virtual scene frame to be aligned is projected onto the planar coordinate system of the live-action image frame. The projection polygon region of the virtual camera frustum on the plane of the live-action image frame is calculated. The virtual scene frame to be aligned is clipped based on the projection polygon region, retaining only the pixel content within the projection polygon region to obtain the aligned virtual scene frame.

[0048] The third sub-step is to perform a depth comparison test on the preprocessed live-action image and the aligned virtual scene frame to generate a pixel-by-pixel occlusion weight map. The specific method of the depth comparison test is as follows: For each pixel in the preprocessed live-action image, obtain the scene depth value of that pixel in the live-action scene. The scene depth value is provided by the depth sensor of the live-action camera or calculated from multi-view live-action image frames using a stereo matching algorithm. For pixels in the aligned virtual scene frame that have the same pixel coordinates as those in the preprocessed live-action image, obtain the rendering depth value of that pixel in the virtual scene. The rendering depth value is read from the depth buffer of the animation rendering engine. The occlusion weight of each pixel is determined and assigned according to the following formula: ; in, Indicates pixel coordinates The pixel-by-pixel occlusion weight value at the location, A value of 1.0 indicates that pixels from the preprocessed real-shot image receive full weight during fusion. A value of 0.0 indicates that pixels aligned with the virtual scene frame receive full weight during fusion. This indicates the pixel coordinates of the preprocessed real-shot image. The scene depth value at that location; Indicates the alignment of virtual scene frames in pixel coordinates The rendering depth value at that location; This represents the depth tolerance threshold, which is set to 0.05 meters. The reason for setting the depth tolerance threshold to 0.05 meters is that this value can avoid flickering at the occlusion boundary caused by depth measurement noise. The above occlusion weight assignment is applied to the coordinates of all pixels covered by the preprocessed real-shot image to determine the occlusion, resulting in a pixel-by-pixel occlusion weight map with the same resolution as the preprocessed real-shot image.

[0049] The fourth sub-step involves using a pixel-by-pixel occlusion weight map to perform weighted fusion of the preprocessed real-world image and the aligned virtual scene frame, generating a virtual-real fusion scene state. For each pixel coordinate, the color vectors of the preprocessed real-world image and the aligned virtual scene frame at that pixel coordinate are extracted. A weighted fusion formula is then used to calculate the fused color vector: the output color vector equals the pixel-by-pixel occlusion weight multiplied by the preprocessed real-world image color vector plus 1, minus the pixel-by-pixel occlusion weight multiplied by the aligned virtual scene frame color vector. After pixel-by-pixel weighted fusion of all pixel coordinates, a fused image frame is obtained. This fused image frame, along with the real-world spatial positioning parameters and the updated spatial mapping matrix, constitutes the virtual-real fusion scene state. This state is used by the intelligent director decision-making module for subsequent temporal trend analysis.

[0050] Example 5: In specific implementation, please refer to Figure 6 After the intelligent director decision module obtains the virtual-real scene status output by the virtual production integration module, it performs time-series change trend analysis on the virtual-real scene status, generates director decision control signals, and feeds them back to the multimodal analysis module to complete closed-loop control.

[0051] In practical implementation, the intelligent director decision-making module statistically tracks the semantic drift of the virtual-real fusion scene state within a continuous time window to obtain the scene state evolution trajectory. The virtual-real fusion scene state includes fused image frames, live-action spatial positioning parameters, and an updated spatial mapping matrix. For each fused image frame, the intelligent director decision-making module invokes the scene semantic distribution map extraction process deployed in the multimodal analysis module. Taking the fused image frame as input, it re-extracts geometric mesh data, texture data, and motion capture data from the fused image frame, and obtains the fused scene semantic distribution map through the same vertex normal deviation statistics, color space distribution statistics, and keypoint trajectory fitting. The continuous time window span is set to 90 frames. The basis for setting the continuous time window span to 90 frames is that at a frame rate of 24 frames per second, 90 frames correspond to a time range of approximately 3.75 seconds, which can cover the complete plot segment while maintaining timely response to the evolution trend. The fused scene semantic distribution map corresponding to the first fused image frame within a continuous time window is used as the baseline semantic distribution map, and the fused scene semantic distribution map corresponding to each subsequent fused image frame within the continuous time window is used as the comparison semantic distribution map. For each frame of the comparison semantic distribution map, the scene semantic drift between the comparison semantic distribution map and the baseline semantic distribution map is calculated. The scene semantic drift is calculated by performing element-wise difference calculations between the local semantic feature vector of each grid node in the baseline semantic distribution map and the corresponding local semantic feature vector in the comparison semantic distribution map, and taking the average of the L2 norms of all grid node difference vectors as the scene semantic drift for that frame. The scene semantic drift of each frame within the continuous time window is arranged in chronological order to form a one-dimensional temporal sequence, which represents the scene state evolution trajectory.

[0052] In practical implementation, the intelligent director decision-making module measures the deviation between the scene state evolution trajectory and the dynamic pre-show path sequence to obtain a director intervention urgency index. The deviation measurement process is as follows: multiple pre-show keyframes are set on the dynamic pre-show path sequence, and a desired semantic feature vector is configured for each pre-show keyframe. The pre-show keyframes are set as follows: starting from the starting frame of the dynamic pre-show path sequence, one frame is selected every 15 frames as a pre-show keyframe. The selection interval of 15 frames is based on the fact that at a frame rate of 24 frames per second, 15 frames correspond to approximately 0.625 seconds, which can capture key turning points in the plot evolution. For each pre-show keyframe, the desired semantic feature vector corresponding to the pre-show keyframe is extracted from the intermediate results stored by the AI ​​pre-show deduction module when generating the dynamic pre-show path sequence. The desired semantic feature vector is the ideal semantic representation calculated by the AI ​​pre-show deduction module for the state node of the pre-show keyframe during the director style behavior model inference stage.

[0053] The actual semantic feature vector corresponding to the time sequence of the pre-playing keyframes is extracted from the scene state evolution trajectory. Each time point in the scene state evolution trajectory corresponds to a fused image frame. For each pre-playing keyframe, the actual frame with the same timestamp as the pre-playing keyframe is found in the scene state evolution trajectory. The fused scene semantic distribution map corresponding to the actual frame is extracted. The local semantic feature vectors of all grid nodes in the fused scene semantic distribution map are globally averaged and pooled to obtain a fixed-dimensional vector. The fixed-dimensional vector is the actual semantic feature vector corresponding to the pre-playing keyframe.

[0054] Calculate the semantic deviation distance between the actual semantic feature vector and the expected semantic feature vector to obtain a single-frame deviation sequence. For each pair of actual and expected semantic feature vectors corresponding to the same pre-visualization keyframe, calculate the cosine distance between the two vectors. The cosine distance is set to 1 minus the cosine similarity. Arrange the cosine distances corresponding to all pre-visualization keyframes in chronological order to obtain a single-frame deviation sequence.

[0055] A sliding window summation and normalization mapping are performed on the single-frame deviation sequence to obtain the director intervention urgency index. The sliding window size is set to 4 pre-rehearsal keyframes. The reason for setting the sliding window size to 4 is that 4 pre-rehearsal keyframes cover an approximately 2.5-second scene segment, which can smooth short-term fluctuations while maintaining sensitivity to continuous deviation trends. The sliding window summation method is as follows: starting from the first element of the single-frame deviation sequence, the window slides with a step size of 1. At each window position, the sum of all single-frame deviation values ​​within the window is calculated to obtain a cumulative deviation value sequence. The normalization mapping adopts the min-max normalization method. For all elements in the cumulative deviation value sequence, the director intervention urgency index is obtained by normalizing and mapping according to the following formula: ; in, This represents the urgency of director intervention when the sliding window covers the t-th preview keyframe in the single-frame deviation sequence. The value range is between 0.0 and 1.0; This represents the cumulative deviation value calculated when the sliding window ends at the t-th preview keyframe. This represents the minimum value in the sequence of cumulative deviations calculated from all sliding window positions. This represents the maximum value in the cumulative deviation sequence calculated across all sliding window positions. The upper and lower bounds of the normalization mapping are taken from the actual minimum and maximum values ​​in the cumulative deviation sequence to enable the urgency index to adapt to the range of deviation values ​​under different semantic complexities of different scenarios.

[0056] In practice, the intelligent director decision-making module selects matching control operations from the preset director action space based on the urgency index of director intervention, generating director decision control signals. The preset director action space is a predefined set of director control operations, containing six categories: the first is adjusting the scene lighting direction, the second is modifying the character behavior weights, the third is resetting the camera motion curve, the fourth is triggering specific plot branches, the fifth is adjusting the virtual camera depth-of-field parameters, and the sixth is changing the character animation blending weights. Each type of control operation corresponds to an operation code and a parameter vector template.

[0057] The matching of the director intervention urgency index with control operations employs a tiered threshold determination mechanism. The value range of the director intervention urgency index [0.0, 1.0] is divided into three sub-intervals: low urgency interval [0.0, 0.3), medium urgency interval [0.3, 0.7), and high urgency interval [0.7, 1.0]. The division of these three sub-intervals is based on the fact that the magnitude of the director intervention urgency index reflects the severity of the scene state deviating from the pre-rehearsal expectation: the low urgency interval corresponds to slight deviation, the medium urgency interval to significant deviation, and the high urgency interval to severe deviation. When the director intervention urgency index falls into the low urgency interval, the first type of control operation is selected from the preset director action space, generating a director decision control signal containing operation codes and default parameters. When the director intervention urgency index falls into the medium urgency interval, both the first and second type of control operations are selected from the preset director action space simultaneously, generating a director decision control signal containing two operation codes and corresponding parameters. When the director intervention urgency index falls into the high urgency range, the third, fourth and fifth types of control operations are selected simultaneously from the preset director action space to generate a director decision control signal containing three operation codes and corresponding parameters. The parameters of the fourth type of control operation are selected from the preset plot branch library based on the semantic tags of the current scene.

[0058] In practice, the intelligent director decision-making module sends director decision control signals to the multimodal analysis module to adjust the sensitivity of real-time scene data acquisition for the next moment. The director decision control signals are sent asynchronously to the message receiving interface of the multimodal analysis module via the system's internal message bus. Upon receiving the director decision control signals, the multimodal analysis module parses the operation codes and parameter vectors within them. For the first type of adjustment operation, the multimodal analysis module adjusts the vertex sampling density of the geometric mesh data during the next frame of real-time scene data acquisition, increasing the number of sampling points for vertex normal deviation statistics in the region corresponding to the illumination direction, with the increase proportional to the director intervention urgency index value. For the second type of adjustment operation, the multimodal analysis module adjusts the length of the joint trajectory fitting window in the motion capture data, reducing the fitting window length from its original size, with the reduction ratio inversely proportional to the director intervention urgency index value. For the third type of adjustment operation, the multimodal analysis module resets the virtual camera trajectory history data in the real-time scene data acquisition cache and reloads the updated camera motion curve parameters carried in the director decision control signals. For the fourth type of adjustment operation, the multimodal parsing module injects new semantic tags into the target region of the scene semantic distribution map, triggering the subsequent spatial anchor point generation module to generate new editable director-guided anchor points in the corresponding region. For the fifth type of adjustment operation, the multimodal parsing module adjusts the depth sampling range corresponding to the virtual camera's depth parameters and adds texture detail suppression markers outside the depth range to the texture data acquisition channel. After completing the above acquisition sensitivity adjustment, the multimodal parsing module applies the adjusted acquisition parameters in the next real-time scene data acquisition cycle, making the newly acquired real-time scene data closer to the director's creative intent.

[0059] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A real-time 3D animation directing system integrating virtual production and AI pre-visualization, characterized in that, The system includes: The multimodal parsing module is used to extract multimodal features from real-time scene data of a 3D animation virtual studio to obtain a scene semantic distribution map; The spatial anchor point generation module is used to perform spatial event sensitivity clustering in the virtual film set 3D grid based on the scene semantic distribution map to generate editable director guidance anchor points. The AI ​​pre-show deduction module is used to call a pre-trained director style behavior model to perform plot deduction based on the editable director guidance anchor points and generate a dynamic pre-show path sequence. The virtual production integration module is used to receive external live-action signals and perform three-dimensional spatial fusion calibration on the external live-action signals according to the dynamic pre-playing path sequence to obtain the virtual-real fusion scene state. The intelligent director decision-making module is used to analyze the temporal change trend of the virtual-real fusion scene state, generate director decision control signals, and feed them back to the multimodal analysis module to complete closed-loop control.

2. The 3D animation real-time director system integrating virtual production and AI pre-visualization as described in claim 1, characterized in that, The multimodal parsing module extracts multimodal features from real-time scene data of a 3D animation virtual studio to obtain a scene semantic distribution map, specifically including: Perform vertex normal deviation statistics on the geometric mesh data in the real-time scene data to obtain the surface geometric feature field; Perform color space distribution statistics on the texture data in the real-time scene data to obtain a texture semantic label map; Perform joint trajectory fitting on the motion capture data in the real-time scene data to obtain the motion intent encoding vector; The surface geometric feature field, the texture semantic label map, and the action intent encoding vector are mapped to a unified spatial coordinate system and then associated and aggregated element by element to generate the scene semantic distribution map.

3. The 3D animation real-time director system integrating virtual production and AI pre-visualization as described in claim 1, characterized in that, In the spatial anchor point generation module, spatial event sensitivity clustering is performed in the virtual film set 3D mesh based on the scene semantic distribution map to generate editable director guidance anchor points, specifically including: Spatial gradient calculation is performed on the scene semantic distribution map to obtain the scene semantic change gradient field; Based on the scene semantic change gradient field, the virtual scene 3D mesh is non-uniformly subdivided to obtain an adaptive resolution mesh region; Semantic density peak detection is performed within each adaptive resolution grid region to obtain initial anchor point candidate positions; spatial neighborhood conflict resolution and merging are performed on the initial anchor point candidate positions to generate the editable director guidance anchor points.

4. The 3D animation real-time director system integrating virtual production and AI pre-visualization as described in claim 1, characterized in that, In the AI ​​pre-visualization and deduction module, a pre-trained director style behavior model is invoked based on the editable director guidance anchor points to perform plot deduction and generate a dynamic pre-visualization path sequence, specifically including: The editable director guidance anchor points are subjected to temporal state encoding to obtain anchor point state feature sequences; The anchor point state feature sequence is input into the director style behavior model, and the director style behavior model performs scene context association matching on the anchor point state feature sequence and outputs the plot development probability distribution. Based on the probability distribution of plot development, a path search is performed on the editable director guidance anchor point to obtain a set of candidate plot transition paths; The path smoothness of the candidate plot transition path set is evaluated and sorted according to the evaluation results. The candidate plot transition path ranked first is determined as the dynamic pre-playing path sequence.

5. A real-time 3D animation directing system integrating virtual production and AI pre-visualization as described in claim 1, characterized in that, The virtual production integration module receives external live-action signals and performs three-dimensional spatial fusion calibration on the external live-action signals according to the dynamic pre-playback path sequence to obtain the virtual-real fusion scene state, specifically including: Receive external real-shot signals and analyze the camera intrinsic parameter matrix and camera extrinsic pose of the external real-shot signals to obtain real-shot spatial positioning parameters; The desired trajectory of the virtual camera is extracted based on the dynamic pre-simulation path sequence to obtain virtual space positioning parameters; Calculate the coordinate transformation residual between the real-world spatial positioning parameters and the virtual spatial positioning parameters, and perform spatial mapping matrix update based on the coordinate transformation residual to obtain fused spatial alignment parameters; According to the fusion spatial alignment parameters, the external real-shot image frames and the three-dimensional virtual scene frames are pixel-level mixed and superimposed to generate the virtual-real fusion scene state.

6. A real-time 3D animation directing system integrating virtual production and AI pre-visualization as described in claim 5, characterized in that, In the virtual production integration module, the external real-shot image frames and the 3D virtual scene frames are pixel-level mixed and superimposed according to the fusion space alignment parameters to generate the virtual-real fusion scene state, specifically including: Based on the fusion spatial alignment parameters, inverse distortion correction and color space conversion are performed on the external real-shot image frame to obtain a pre-processed real-shot image; Based on the time stamp of the dynamic pre-show path sequence, the corresponding frame is extracted from the 3D virtual scene frame buffer and viewport clipping is performed to obtain the aligned virtual scene frame; A depth comparison test is performed on the preprocessed real-shot image and the aligned virtual scene frame to generate a pixel-by-pixel occlusion weight map. The preprocessed real-shot image and the aligned virtual scene frame are weighted and fused using the pixel-by-pixel occlusion weight map to generate the virtual-real fusion scene state.

7. A real-time 3D animation directing system integrating virtual production and AI pre-visualization as described in claim 1, characterized in that, In the intelligent director decision-making module, the temporal change trend of the virtual-real fusion scene state is analyzed, a director decision control signal is generated, and fed back to the multimodal analysis module to complete closed-loop control, specifically including: The scene semantic drift of the virtual-real fusion scene state within a continuous time window is statistically tracked to obtain the scene state evolution trajectory. The deviation between the scene state evolution trajectory and the dynamic pre-play path sequence is measured to obtain the director intervention urgency index. Based on the director intervention urgency index, a matching control operation is selected from the preset director action space to generate the director decision control signal; The director's decision control signal is sent to the multimodal analysis module to adjust the sensitivity of real-time scene data acquisition at the next moment.

8. A real-time 3D animation directing system integrating virtual production and AI pre-visualization as described in claim 7, characterized in that, In the intelligent director decision-making module, the deviation between the scene state evolution trajectory and the dynamic pre-playback path sequence is measured to obtain a director intervention urgency index, specifically including: Multiple pre-simulation keyframes are set on the dynamic pre-simulation path sequence, and a desired semantic feature vector is configured for each pre-simulation keyframe; Extract the actual semantic feature vector corresponding to the timing of the pre-rehearsal keyframes from the scene state evolution trajectory; Calculate the semantic deviation distance between the actual semantic feature vector and the expected semantic feature vector to obtain a single-frame deviation sequence; The single-frame deviation sequence is summed using a sliding window and then normalized to obtain the director intervention urgency index.

9. A real-time 3D animation directing system integrating virtual production and AI pre-visualization as described in claim 5, characterized in that, When calculating the coordinate transformation residual between the real-world spatial positioning parameters and the virtual spatial positioning parameters, the iterative nearest point algorithm is used to solve the rigid body transformation matrix of the matching feature points in the real-world spatial positioning parameters and the virtual spatial positioning parameters, and the root mean square of the registration error after solving is used as the coordinate transformation residual.

10. A real-time 3D animation directing system integrating virtual production and AI pre-visualization as described in claim 6, characterized in that, When performing a depth comparison test on the preprocessed real-shot image and the aligned virtual scene frame, the scene depth value corresponding to each pixel in the preprocessed real-shot image is compared with the rendering depth value of the corresponding pixel in the aligned virtual scene frame, and the occlusion weight is determined and assigned according to the relationship between the depth values.