AI visual effects dynamic generation system integrating multimodal perception
Through the multimodal perception of AI visual special effects dynamic generation system, the problems of cumbersome manual adjustment and difficult content integration in visual special effects creation are solved, and efficient and high-quality special effects and content are deeply integrated, which enhances the expressiveness and immersion of audio-visual content.
Patent Information
- Application Number
- CN202510780313.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing visual special effects creation relies on manual adjustments and lacks scientific basis, which leads to the overwhelming and misplaced application of special effects, and cannot be deeply integrated with the content, the production cycle and cost are uncontrollable, and it cannot meet the needs of high-quality audio-visual content.
The AI visual special effects dynamic generation system that integrates multimodal perception can achieve deep fusion and intelligent adaptation of special effects and content through multimodal input processing, importance analysis, parameter configuration, parameterized rendering and rendering and output modules.
It improves the efficiency and quality of visual special effects creation, harmoniously unified special effects and content, lowers the creative threshold, shortens the production cycle, reduces costs, and enhances narrative and emotional expression capabilities.
Smart Images

Figure CN120318379B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual special effects technology, and more specifically, to an AI visual special effects dynamic generation system integrating multimodal perception. Background Art
[0002] As a core component of digital content creation, visual effects technology has undergone a remarkable evolution, evolving from traditional manual synthesis to computer graphics-driven production. Early visual effects relied primarily on physical props and optical techniques. With the rapid development of computer graphics, 3D modeling, particle systems, and fluid simulation have gradually become mainstream methods for creating special effects. In recent years, technologies like physically based rendering (PBR) and real-time ray tracing have further enhanced the realism and immersiveness of special effects, revolutionizing fields like film, gaming, and virtual reality.
[0003] However, the current field of visual effects creation still faces technical bottlenecks and creative difficulties, significantly hindering the expressiveness of digital content. Film and television post-production teams often spend a significant amount of time on tedious manual effects adjustments. A simple lighting effect often requires technical artists to repeatedly adjust parameters dozens of times to achieve the director's expectations. These adjustments lack scientific basis and rely primarily on empirical judgment. In actual production, special effects artists struggle to accurately perceive the semantic structure of the content, resulting in special effects often overshadowing the main theme or being misplaced. For example, in a film, overly flashy background effects can distract the audience from key dialogue scenes, weakening the emotional connection. Livestreaming and short video creators face even more severe challenges. They must produce content quickly and are limited to using preset templates, preventing them from achieving customized special effects that deeply integrate with the content. In game development, special effects often exhibit physical inconsistencies when viewed from different perspectives, such as magic effects that suddenly change shape when the perspective changes, disrupting immersion. Even more troubling is the timing instability of special effects in dynamic scenes, such as the flickering and jittering of particle effects in VR experiences, which can cause user discomfort. Special effects creation has long focused on isolated visual processing, neglecting the synergy between audio content and emotional expression. This has reduced many exquisite special effects to mere decoration rather than narrative enhancement. Industry evaluation standards are subjective and vague, making it difficult for creative teams to quantify special effects quality. Project cycles and costs are uncontrollable due to repeated revisions, severely restricting the large-scale production of innovative content and failing to meet the contemporary market's exploding demand for high-quality audiovisual content.
[0004] In view of this, the present invention proposes an AI visual effects dynamic generation system that integrates multimodal perception to solve the above problems. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art and achieve the above-mentioned objectives, the present invention provides the following technical solution: an AI visual effects dynamic generation system integrating multimodal perception, comprising:
[0006] A multimodal input processing module is used to obtain multimodal input data including image sequences, audio data and scene parameters, perform spatiotemporal alignment and normalization on the multimodal input data, and generate a multimodal feature dataset;
[0007] An importance analysis module is used to extract features from the multimodal feature dataset, construct a multimodal semantic feature map, identify key visual areas of the multimodal semantic feature map, and then generate a scene visual importance distribution map;
[0008] A parameter configuration module, configured to select an adapted special effect template and generate a special effect parameter configuration scheme based on the scene visual importance distribution map and the multimodal semantic feature map;
[0009] A parametric rendering module, configured to apply the special effect parameter configuration scheme to the multimodal feature dataset and generate an initial special effect rendered image using an improved parametric rendering algorithm;
[0010] a parameter optimization module, configured to perform a visual quality evaluation on the initial special effect rendered image, optimize the special effect parameter configuration scheme based on the evaluation result, and generate an optimized special effect parameter set;
[0011] A rendering and output module is used to perform final rendering processing on the multimodal feature dataset using the optimized special effects parameter set to generate an AI visual effects image sequence that integrates multimodal perception.
[0012] The technical effects and advantages of the AI visual effects dynamic generation system integrating multimodal perception of the present invention are as follows:
[0013] The present invention improves the efficiency and quality of visual special effects creation, changes the traditional special effects production process that relies on manual trial and error and experience-based judgment, and enables creators to devote more energy to artistic conception rather than tedious technical adjustments. By deeply understanding the semantics and emotional expression of the content, the system makes special effects no longer just decorative elements, but a powerful tool to enhance narrative, strengthen emotions and guide the audience's attention, greatly improving the expressiveness and immersiveness of audio-visual content. The significant reduction in the threshold for creation allows ordinary users to easily create professional-level visual effects, while film and television production organizations can achieve standardization and scale of special effects creation processes while maintaining artistic personalized expression. The harmonious unity of special effects and content eliminates the problem of traditional special effects being too dominant or split in style. By accurately enhancing key visual areas, special effects can enhance rather than interfere with the audience's attention to the core content. The system's intelligent adaptability enables it to automatically adjust the style and intensity of special effects according to different content types, emotional tones and creative intentions, significantly shortening the production cycle and reducing production costs while ensuring high-quality output. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 Schematic diagram of the AI visual effects dynamic generation system integrating multimodal perception of the present invention;
[0015] Figure 2 A logical diagram for identifying key special effect application areas of the present invention;
[0016] Figure 3 A logical diagram of the present invention for performing perspective-consistent special effects rendering. DETAILED DESCRIPTION
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0018] The present application example provides an AI visual effects dynamic generation system that integrates multimodal perception. The execution subjects of the AI visual effects dynamic generation system that integrates multimodal perception include but are not limited to: digital content creation platforms, video processing systems, film and television post-production tools, real-time rendering engines, augmented reality applications, etc. that are equipped with the system can be regarded as general computing nodes of the present application, and the data processing platform includes but is not limited to: multimodal feature extraction systems, visual importance analysis systems, and at least one parametric rendering system.
[0019] See also Figure 1 The present invention provides an AI visual effects dynamic generation system integrating multimodal perception, including the following modules:
[0020] The multimodal input processing module is used to obtain multimodal input data including image sequences, audio data and scene parameters, perform spatiotemporal alignment and standardization on the multimodal input data, and generate a multimodal feature dataset;
[0021] Importance analysis module, which is used to extract features from multimodal feature datasets, construct multimodal semantic feature maps, identify key visual areas of the multimodal semantic feature maps, and then generate scene visual importance distribution maps;
[0022] The parameter configuration module is used to select an adaptive special effect template and generate a special effect parameter configuration scheme based on the scene visual importance distribution map and the multimodal semantic feature map;
[0023] The parametric rendering module is used to apply the special effects parameter configuration scheme to the multimodal feature dataset and generate the initial special effects rendered image through the improved parametric rendering algorithm;
[0024] The parameter optimization module is used to evaluate the visual quality of the initial special effects rendering image, optimize the special effects parameter configuration scheme based on the evaluation results, and generate an optimized special effects parameter set;
[0025] The rendering and output module is used to perform final rendering processing on the multimodal feature dataset using the optimized special effects parameter set to generate an AI visual effects image sequence that integrates multimodal perception. The modules are connected via wired and / or wireless means to achieve data transmission between modules.
[0026] The present invention realizes comprehensive perception of image sequences, audio and scene parameters through multimodal input processing, constructs a multimodal feature data set to provide a data basis for subsequent analysis, performs importance analysis based on multimodal features to enable special effects generation to have semantic perception capabilities, and the scene visual importance distribution map can accurately guide the application location and intensity of special effects. The selection of adaptive special effects templates based on importance distribution and semantic features improves the situational adaptability of special effects. The special effects parameter configuration scheme takes into account multiple perceptual factors to ensure the artistic expression of special effects. High-quality preliminary rendering of special effects is achieved through an improved parametric rendering algorithm. The initial special effects rendered image maintains the visual integrity of the original content. Visual quality evaluation and parameter optimization of the rendering results improve the visual performance of the special effects. The optimized special effects parameter set ensures the visual coherence and perceptual harmony of the special effects. The special effects image sequence generated by the final rendering process has multimodal coordination and high visual quality.
[0027] In an embodiment of the present invention, the multimodal input processing module is used to obtain multimodal input data including image sequences, audio data, and scene parameters, perform spatiotemporal alignment and normalization on the multimodal input data, and generate a multimodal feature dataset. Specifically, it is used to:
[0028] Performing key frame extraction and resolution unification processing on the image sequence to obtain a standardized image frame sequence, and performing color space conversion on the standardized image frame sequence to obtain a color-corrected image sequence;
[0029] Perform sampling rate resampling and spectrum analysis on the audio data to obtain a time-frequency characteristic map, and perform emotional feature extraction on the time-frequency characteristic map to obtain an audio emotional feature vector;
[0030] Perform semantic parsing and hierarchical organization on scene parameters to obtain structured parameter representation, and build a scene semantic network based on the structured parameter representation to obtain a scene semantic vector;
[0031] The color-corrected image sequence, audio emotion feature vector, and scene semantic vector are timestamp-aligned to obtain synchronized multimodal data, and the synchronized multimodal data is feature-normalized to obtain a standardized feature set.
[0032] The standardized feature set is fused and encoded to construct a unified representation space and obtain a multimodal feature dataset.
[0033] In this embodiment, an input image sequence is first collected, and visual features, timestamps, resolution, and other information of each frame of the image are obtained through a computer vision algorithm to form an image sequence set containing complete visual information. Key frame extraction is performed on the obtained image sequence set, and key frames are identified using methods such as scene switch detection, content change measurement, and visual attention model to ensure that significant change points of the visual content are captured. The extracted key frames are subjected to resolution unification processing, including operations such as scaling, cropping, and pixel resampling, so that all frames have consistent resolution and aspect ratio to form a standardized image frame sequence. The standardized image frame sequence is subjected to color space conversion to convert the RGB color space into a color space more suitable for special effects processing (such as HSL, LAB, etc.). Color correction is performed, including white balance adjustment, tone mapping, and color enhancement, to ultimately obtain a color-corrected image sequence. At the same time, audio data is extracted from the input data, including various sound elements such as music, dialogues, ambient sounds, and sound effects. The audio data is resampled, and audio with different sampling rates is unified to a standard sampling rate (such as 44.1kHz or 48kHz) to ensure processing consistency. The resampled audio is subjected to spectral analysis, and methods such as short-time Fourier transform (STFT), wavelet transform, or Mel-frequency cepstral coefficients (MFCC) are used to extract the time-frequency characteristics of the audio, and generate a time-frequency characteristic map reflecting the distribution and changes of audio energy. The emotional feature extraction algorithm is applied to the time-frequency characteristic map to analyze the audio's rhythm, emotional color, energy changes, and other characteristics. Deep learning models (such as CNN, LSTM, etc.) are used to extract high-order emotional features to form an audio emotional feature vector representing the audio's emotional and rhythmic characteristics. In addition, the algorithm receives scene parameter input, including high-level semantic information such as scene type, narrative stage, emotional tone, and special effect intent. It then performs semantic parsing on the scene parameters, identifying keywords, emotional tags, and intent instructions, understanding the relationships between parameters, and hierarchically organizing the parsing results into a structured semantic tree or knowledge graph. This structured parameter representation is then constructed, and a scene semantic network is constructed based on the structured parameter representation. A graph neural network or attention mechanism model is used to capture the complex relationships between parameters, generating an abstract semantic representation of the scene and forming a scene semantic vector. The color-corrected image sequence, audio emotion feature vector, and scene semantic vector are then aligned according to timestamps to ensure temporal synchronization of the different modal data. The differences in delay and duration between the different modal data are addressed, and a temporal correlation map is constructed to form synchronized multimodal data. The different features in the synchronized multimodal data are then normalized, converting features of different scales and dimensions to a unified numerical range. Methods such as normalization, maximum-minimum scaling, or rank normalization are used to balance the weights of the different modal features to prevent a single modality from dominating the feature space, resulting in a standardized feature set.Finally, the standardized feature set is input into the multimodal feature fusion network, and early fusion, mid-term fusion or late fusion strategies are adopted to perform cross-attention calculation or tensor fusion on the features of different modalities to construct a feature space that can uniformly express multimodal information. The Transformer architecture or cross-modal autoencoder is used to achieve semantic alignment and complementary enhancement of features, and finally generate a multimodal feature dataset containing rich audiovisual semantic information. This dataset not only retains the key features of each modality, but also captures the interactive relationship between modalities, providing a comprehensive data foundation for subsequent importance analysis.
[0034] In an embodiment of the present invention, key frame extraction and resolution unification processing are performed on an image sequence to obtain a standardized image frame sequence, and color space conversion is performed on the standardized image frame sequence to obtain a color-corrected image sequence, including:
[0035] A visual information entropy measurement model is applied to image sequences to calculate the information richness of each frame. Combined with inter-frame visual difference metrics, a dual evaluation mechanism is constructed to identify frames with peak visual information and obtain a set of keyframe candidates. A semantic importance filtering algorithm is then applied to the keyframe candidates to retain narrative key points and visual turning points, resulting in a streamlined set of keyframes.
[0036] The content composition and subject distribution of each frame in the streamlined keyframe set are analyzed, and an adaptive content-aware grid is constructed to guide the resolution adjustment process. A content-preserving map is obtained, and a non-uniform scaling operation is performed based on the content-preserving map to achieve resolution standardization while maintaining the integrity of the visual subject, resulting in a uniform resolution frame with subject enhancement.
[0037] A super-resolution neural network is applied to the uniform-resolution frames of the subject enhancement to intelligently repair detail loss caused by scaling, reconstructing texture and edge details to obtain detail-restored image frames. These detail-restored image frames are then optimized for temporal consistency to eliminate fluctuations in detail representation between frames, resulting in a temporally coherent standardized image frame sequence.
[0038] Establish a scene lighting condition analysis model, infer ambient light characteristics from standardized image frame sequences, and select the color space that best suits the target special effect, achieving perceptually consistent color space conversion and obtaining a color gamut-optimized image;
[0039] Semantic segmented color mapping is performed on the color gamut optimized image, and customized color enhancement strategies are applied to different semantic areas while maintaining global color harmony. A viewer visual adaptability compensation mechanism is introduced to optimize color performance under different viewing environments, ultimately obtaining a perceptually enhanced color-corrected image sequence.
[0040] In this embodiment, a visual information entropy measurement model is first applied to the input image sequence, and the Shannon entropy or relative entropy value of each frame image is calculated to evaluate the information content and complexity contained therein. At the same time, the visual differences between adjacent frames are calculated, and the degree of change between frames is quantified using methods such as structural similarity (SSIM), feature point matching difference or depth feature distance. The two indicators of information entropy and inter-frame difference are combined to construct a dual evaluation mechanism, and an adaptive threshold is set to identify information mutation points and frames with significant visual changes. Frames containing rich information and with significant changes are screened out as key frame candidate sets. A semantic importance filtering algorithm is applied to the key frame candidate sets, and a pre-trained visual semantic model (such as CLIP or ViT) is used to evaluate the semantic importance and narrative value of each frame. Key points with story turning significance (such as scene switching, emotional change points, etc.) are identified, frames with high semantic importance and large content differences are retained, and frames with semantic redundancy or visual similarity are filtered out to obtain a streamlined but information-rich key frame set. Then, the content composition and subject distribution of each frame in the streamlined key frame set are analyzed, and target detection and semantic segmentation techniques are used to identify the main objects, background elements and their spatial distribution in the image. An adaptive content-aware grid is constructed based on visual saliency and semantic importance, and the image is divided into regions of different importance levels to form a grid structure with variable density. The grid density is high in important content areas and low in secondary areas. A content protection map reflecting the distribution of content importance is generated, and a non-uniform scaling operation is performed based on the content protection map. The strategy of maintaining the original proportion or slightly scaling is adopted for important content areas, and a more aggressive scaling strategy is adopted for secondary areas. Content-aware scaling algorithms (such as seam carving or grid deformation) are used to achieve resolution standardization, while maximally retaining the integrity and details of the visual subject, to obtain a uniform resolution frame with subject enhancement. Next, the enhanced uniform resolution frame is input into a super-resolution neural network (such as SRGAN, ESRGAN or Real-ESRGAN), which intelligently repairs and enhances the texture details and edge features lost during the scaling process. The network learns the feature distribution of high-resolution images, reconstructs the blurred or jagged details caused by scaling, restores texture richness and edge clarity, and generates detail-restored image frames. The temporal consistency optimization algorithm is applied to the detail-restored image frame sequence to analyze the coherence of detail performance between adjacent frames, eliminate the inter-frame detail fluctuations and flickering caused by single-frame super-resolution processing, and ensure a smooth transition of texture, color and detail in the temporal dimension. Finally, a standardized image frame sequence that maintains high detail quality and temporal coherence is obtained.Then, a scene lighting condition analysis model is established to extract lighting features from standardized image frame sequences, analyze the light source direction, intensity, color temperature and scattering characteristics, and infer the ambient lighting conditions (such as indoor, outdoor, daylight, artificial light, etc.). According to the lighting conditions and the target special effect type (such as flame, water, light effect, etc.), the color space most suitable for special effect rendering is selected. For example, HDR special effects are suitable for ACES color space, and transparency-related special effects are suitable for linear RGB color space. Maintain perceptual consistency when performing color space conversion to ensure that the visual experience before and after the conversion does not change significantly, and obtain a color gamut optimized image. Finally, semantic segmentation color mapping is performed on the color gamut optimized image. First, a semantic segmentation algorithm is used to divide the image into regions with different semantics (such as people, background, objects, etc.), and a customized color enhancement strategy is designed for each semantic region, such as emphasizing the naturalness of people's skin color, enhancing the color saturation of key objects, etc. Local color enhancement is applied while maintaining global color harmony to ensure that the color processing of different regions remains beautiful and unified as a whole. The audience's visual adaptability compensation mechanism is introduced to consider the differences in color perception in different viewing environments (such as theaters, home TVs, mobile devices, etc.), and optimize the color performance so that it can maintain the best visual effect under various display devices and ambient light conditions. Finally, a color-corrected image sequence that conforms to artistic intentions and has perception enhancement characteristics is obtained, providing a high-quality visual foundation for subsequent special effects generation.
[0041] In an embodiment of the present invention, the importance analysis module is used to extract features from a multimodal feature dataset, construct a multimodal semantic feature map, identify key visual areas of the multimodal semantic feature map, and then generate a scene visual importance distribution map. Specifically, it is used to:
[0042] The multimodal feature dataset is input into the pre-trained multimodal Transformer network to extract deep semantic features and obtain a multi-level feature tensor. The multi-level feature tensor is then processed by the channel attention mechanism to obtain an enhanced feature map.
[0043] Use graph convolutional networks to model the spatial relationships of enhanced feature maps, construct scene graph structure representations, and obtain scene element relationship networks. Then, perform semantic segmentation based on the scene element relationship networks and obtain semantic region partition maps.
[0044] The semantic region division map is cross-modally fused with the audio emotion feature vector to obtain an emotion-enhanced semantic feature map. Multi-scale information is then integrated through an adaptive feature pyramid network to obtain a multimodal semantic feature map.
[0045] Applying a visual saliency detection algorithm to the multimodal semantic feature map, calculating the attention weight of each region to obtain an initial saliency map, and then calibrating the initial saliency map with the human visual perception model to obtain a perceptually corrected saliency map.
[0046] The perceptually corrected saliency map is weightedly fused with the semantic region partition map to construct a saliency heat map that considers semantic importance, and obtains the scene visual importance distribution map.
[0047] In this embodiment, the multimodal feature dataset is first input into a pre-trained multimodal Transformer network. The network has been pre-trained on large-scale multimodal data and has the ability to understand cross-modal semantics. It uses a multi-head self-attention mechanism to process the relationship between different modalities, extract deep semantic features of images, audio and scene parameters, and generate a feature tensor containing multi-level semantic information. The channel attention mechanism is applied to the multi-level feature tensor to automatically identify and enhance feature channels related to the current content, suppress irrelevant channels, and achieve adaptive enhancement of features. The weight of each channel is dynamically adjusted through the attention gating function to highlight semantic key information and form an enhanced feature map. Next, the enhanced feature map is input into the graph convolutional network (GCN), which regards the visual elements in the scene as nodes in the graph structure, and the spatial relationships and semantic associations between elements as edges. Complex scene relationships are modeled through multi-layer graph convolution operations, capturing the interactive relationships and contextual dependencies between entities, and constructing a scene graph structure representation containing rich relationship information to form a scene element relationship network. Semantic segmentation tasks are performed based on the scene element relationship network, and the relationship information learned by the network is used to guide the segmentation process, improve the semantic consistency of the segmentation, and divide the image into regions with clear semantic labels, such as people, objects, background, etc., to generate a semantic region partition map. Then, the semantic region division map is cross-modally fused with the audio emotion feature vector, and a cross-modal attention mechanism is designed to enable the audio emotion features to modulate the expression intensity of the visual semantic region. For example, emotionally intense audio will enhance the emotional representation of the corresponding visual region, thereby achieving collaborative enhancement of audio and video semantics. Multimodal co-occurrence patterns are captured through interactive learning between modalities to obtain an emotion-enhanced semantic feature map. The emotion-enhanced semantic feature map is input into an adaptive feature pyramid network, which extracts and integrates features at multiple scales, capturing everything from fine-grained local features to coarse-grained global semantics. Features of different scales are integrated through top-down and bottom-up bidirectional information flows to achieve a balanced expression of details and semantics, and finally a multimodal semantic feature map is generated, which integrates the semantic information of visual, audio, and scene parameters. Next, a visual saliency detection algorithm is applied to the multimodal semantic feature map to calculate the degree to which each area in the image attracts visual attention. A biologically inspired attention model or a deep learning saliency detection network (such as SAM, SalGAN, etc.) is used to generate an initial saliency map reflecting the distribution of visual attention. The initial saliency map is combined with the human visual perception model, which is based on psychophysical experiments and visual cognition theory. It simulates the attention allocation mechanism of the human visual system to the scene, takes into account visual characteristics such as center preference, face priority, and motion sensitivity, and calibrates the saliency map to make it more consistent with real human visual perception, thus obtaining a perceptually corrected saliency map.Finally, the perceptually corrected saliency map is weightedly fused with the semantic region partition map to assign importance weights to different semantic regions. For example, higher weights are given to narratively important characters or key objects, and lower weights are given to secondary background elements. This constructs a heat map that considers both visual saliency and semantic importance. This heat map can accurately indicate the comprehensive importance of each region in the scene, and ultimately generates a scene visual importance distribution map. This distribution map will guide the subsequent special effects parameter configuration and rendering process, ensuring that special effects are applied to the most appropriate areas and enhancing the visual narrative effect.
[0048] In the embodiment of the present invention, a graph convolutional network is used to model the spatial relationship of the enhanced feature map, construct a scene graph structure representation, and obtain a scene element relationship network, including:
[0049] Generate region candidates for the enhanced feature map, extract potential scene entity regions, obtain a region candidate set, and perform attribute prediction and classification on the region candidate set to obtain a semantically labeled region set;
[0050] Construct feature nodes for each region in the semantically labeled region set, calculate the spatial and semantic similarities between regions, generate a weighted adjacency matrix, and form the initial scene graph;
[0051] Apply multi-layer graph convolution operations to transfer and update information on the initial scene graph, achieve context enhancement of node features, obtain context-aware node representations, and predict semantic relationships between nodes based on the context-aware node representations to obtain a relationship labeling graph;
[0052] Interactively integrate the relationship labeling graph with global scene features to capture long-range dependencies and obtain a global correlation scene graph. Then, perform attention-guided information screening on the global correlation scene graph to highlight key relationships.
[0053] A hierarchical scene understanding model is constructed based on the global associated scene graph, which supports multi-granularity scene parsing and ultimately obtains a scene element relationship network.
[0054] In this embodiment, a region candidate generation algorithm, such as a region proposal network (RPN), selective search, or DeepMask, is first applied to the enhanced feature map to identify image regions that may contain meaningful entities. A set of region candidates containing bounding box coordinates is generated. Feature extraction is performed on each candidate region. A convolutional neural network is used to extract visual feature representations of the region. Attribute prediction and classification are performed on the extracted features. The semantic category (such as person, animal, object, etc.) and attributes (such as color, texture, state, etc.) to which the region belongs are identified. A multi-label classifier is used to assign semantic labels and attribute descriptions to each region, forming a set of semantically annotated semantically labeled regions. Then, a graph structure node is created for each region in the semantically labeled region set. The node features contain the region's visual features, semantic labels, and attribute information. The spatial relationships between regions, such as relative position (up, down, left, right), inclusion relationship, overlap, etc., are calculated. The semantic similarity between regions is calculated. Based on the feature cosine distance or the learned semantic distance metric, a weighted adjacency matrix is constructed based on the spatial relationship and semantic similarity. The matrix elements represent the strength and type of connections between regions. The initial scene relationship graph structure is constructed to form an initial scene graph. Next, the initial scene graph is fed into a multi-layer graph convolutional network, where a message passing mechanism is implemented. This allows each node to aggregate information about its neighbors. Multi-hop message passing captures a wider range of contextual relationships, achieving context-aware enhancement of node features. This results in a context-aware node representation that incorporates information about the surrounding environment. Semantic relationships between nodes, such as "located in," "contains," "supports," and "holds," are predicted based on the context-aware node representation. A relationship classifier is used to predict the relationship type for each pair of connected nodes, generating a relationship labeling graph containing rich semantic relationship annotations. Global scene features are then extracted to represent the overall scene context and atmosphere. The relationship labeling graph is interactively integrated with the global scene features. A global-local attention mechanism is used to enable node features to focus on the global context while enabling global features to focus on locally important entities. Long-range dependencies across regions are captured through bidirectional information flow, forming a globally contextualized scene graph. An attention-guided information filtering mechanism is then applied to the global contextualized scene graph. This automatically focuses on key nodes and relationships based on task objectives and content importance, suppressing irrelevant secondary connections and highlighting the core relationship structures that are crucial for scene understanding, achieving effective information simplification.Finally, a hierarchical scene understanding model is constructed based on the global contextual scene graph, organizing scene elements according to different levels of abstraction, from bottom-level pixels and regions to mid-level objects and entities, and then to high-level events and contextual understanding. It supports multi-level scene parsing from fine-grained to coarse-grained, allowing the system to understand and operate scene content at different levels of abstraction according to task requirements, integrate information at each level, and construct a unified representation that contains both details and high-level semantics, ultimately forming a scene element relationship network. This network comprehensively captures the spatial layout, semantic associations, and interaction methods between entities in the scene, laying a structured foundation for subsequent visual importance analysis and special effects application.
[0055] In an embodiment of the present invention, the parameter configuration module is used to select an adaptive special effect template based on the scene visual importance distribution map and the multimodal semantic feature map, and generate a special effect parameter configuration scheme, specifically for:
[0056] Build a special effects template knowledge base, including parameterized templates for multiple categories of visual effects, and create semantic labels and applicable scenario descriptions for each template to obtain a semantic special effects library;
[0057] Perform regional clustering on the scene visual importance distribution map to identify key special effect application areas and obtain a set of special effect target areas. Furthermore, extract the scene emotional tone and style features based on the multimodal semantic feature map to obtain the scene style representation.
[0058] The features of the special effect target area set are combined with the scene style representation to construct a special effect demand vector. The special effect demand vector is then matched with the templates in the semantic special effect library to obtain a set of candidate special effect templates.
[0059] A multi-objective evaluation mechanism is applied to the candidate special effects template set, evaluating it from multiple dimensions such as visual coordination, emotional matching, and computational complexity to obtain a template score list. Based on the template score list, the optimal special effects template combination is selected to obtain a special effects application plan.
[0060] According to the special effects application plan, the initial parameter values are calculated for each selected special effects template, and the parameter intensity and application range are adjusted in combination with the scene visual importance distribution map to finally form a special effects parameter configuration plan.
[0061] In this embodiment, a special effects template knowledge base is first constructed, and various categories of visual special effects templates are collected and designed, including light effects (such as halos, light beams, and flashes), particle effects (such as flames, smoke, and water), atmosphere effects (such as fog and light scattering), deformation effects (such as distortion and ripples), etc., and a parametric description is implemented for each special effects template, including shape parameters, dynamic parameters, physical parameters, and visual style parameters, etc., so that special effects can be customized through parameter adjustment. Semantic labels are added to each special effects template to describe its visual characteristics, emotional expression, and applicable scenarios, such as "warm halo - creates a warm atmosphere - suitable for warm scenes", and a mapping relationship between special effects and narrative intentions is established to form a structured semantic special effects library. Then, a spatial clustering algorithm, such as DBSCAN, Mean Shift or K-means, is applied to the scene visual importance distribution map. Clusters of regions with similar importance are identified based on the distribution pattern of importance values. Regions where special effects should be applied are screened out considering spatial continuity and importance thresholds. A set of special effects target regions containing spatial positions and importance weights is generated. The emotional tone of the scene, such as tension, cheerfulness, sadness, mystery, etc., is analyzed based on the multimodal semantic feature map. The visual style features of the scene, such as tonal tendencies, visual rhythm, compositional characteristics, etc., are extracted. The emotional tone and visual style information are integrated into a vector representation that represents the artistic style of the scene to form a scene style representation. Next, the visual features, location information and importance weights of the special effects target area set are fused with the scene style representation to construct a multidimensional vector expressing the special effects requirements. This vector encodes special effects application requirements such as "where", "what style" and "how strong". The feature similarity between the special effects requirement vector and each template in the semantic special effects library is calculated. Using cosine similarity, semantic matching or learned similarity measurement, special effects templates with high matching degree with the requirement vector are screened out to form a candidate special effects template set. Then, a multi-dimensional evaluation is conducted on the special effects in the candidate special effects template set. The style compatibility and harmony between the special effects and the original picture are evaluated from the dimension of visual coordination. The enhancement or supplementary effect of the special effects on the emotional expression of the scene is evaluated from the dimension of emotional matching. The resource consumption and rendering time cost of the special effects are evaluated from the dimension of computational complexity. Each dimension is quantitatively scored and a comprehensive score is calculated. A template scoring list containing the scores of each indicator is generated. Based on the template scoring list, the optimal special effects template combination is selected, taking into account the complementarity and synergy between templates to optimize the overall visual experience. It may include multiple special effects of different types or the same type with different parameters to form a special effects application plan.Finally, according to the special effects application plan, the initial parameter values are determined for each selected special effects template. Reasonable initial parameter values are set based on the special effects type and application scenario, such as light effect intensity, number of particles, animation speed, etc., and the parameters are adjusted in combination with the scene visual importance distribution map. The details and intensity of the special effects are enhanced for high-importance areas, and the presence of special effects is reduced for low-importance areas to avoid distraction. The parameter coordination and superposition effects between special effects are considered to prevent visual overload or conflict, and finally a special effects parameter configuration plan is formed. The plan includes detailed parameter settings and application area definitions for each special effects template, providing precise guidance for subsequent parametric rendering.
[0062] In the embodiment of the present invention, please refer to Figure 2 , a logical diagram for identifying key special effect application areas. Specifically, the scene visual importance distribution map is clustered to identify key special effect application areas, and a set of special effect target areas is obtained, including:
[0063] An adaptive threshold segmentation algorithm is applied to the scene visual importance distribution map. The optimal segmentation threshold is automatically determined according to the global importance distribution characteristics to obtain a preliminary region partition map. The preliminary region partition map is then subjected to morphological operations to remove noise and smooth the boundaries to obtain a fine region mask.
[0064] A multi-level image segmentation tree is constructed based on fine-grained region masks. Combining the visual Gestalt law and the principle of semantic consistency, a hierarchical region representation from pixel to object level is achieved, and a multi-scale region candidate pool is obtained. A region merging strategy based on visual perceptual correlation is applied to the multi-scale region candidate pool to fuse regions that visually belong to the same perceptual unit to obtain a set of perceptually consistent regions.
[0065] A dynamic density peak clustering algorithm is designed to adaptively cluster areas with concentrated perceptual consistency, automatically identify important density centers, and consider the semantic relevance between regions to achieve collaborative clustering of semantically related regions, thereby obtaining semantically enhanced clustering results.
[0066] A temporal stability evaluation mechanism is introduced to analyze semantics across frames to enhance the consistency of clustering results. This identifies stable and important regions with temporal continuity and filters out transient unstable regions to obtain a set of temporally stable and important regions. Furthermore, a time-varying importance weight is calculated for each region to construct a dynamic importance curve.
[0067] Combining the temporally stable important region set with the dynamic importance curve, a multi-target region value assessment system is designed. Taking into account factors such as region size, importance intensity, perceptual complexity, narrative importance, and special effects adaptability, regional priority is sorted to obtain a graded special effects target region set. The region boundaries are optimized through the regional automatic diffusion algorithm to ensure the visual consistency of special effects application, and finally a special effects target region set with precise boundary definition and importance weights is generated.
[0068] In this embodiment, an adaptive threshold segmentation algorithm, such as the Otsu method, adaptive threshold or multi-level threshold segmentation, is first applied to the scene visual importance distribution map, the distribution histogram of the global importance value is analyzed, the optimal threshold that can maximize the between-class variance is automatically calculated, the importance map is divided into high-importance areas and low-importance areas, and a preliminary area partition map is generated. Morphological operations, including opening and closing operations, median filtering, and area connection, are applied to the preliminary area partition map to remove isolated noise pixels and small irrelevant areas, smooth the area boundaries, fill small internal holes, improve the coherence and integrity of the area, and obtain a fine area mask. Then, a multi-level image segmentation tree is constructed based on the fine region mask, and a hierarchical segmentation algorithm (such as hierarchical superpixel or layered graph segmentation) is used. Starting from the pixel level, similar regions are gradually merged to form a multi-level segmentation result from fine-grained to coarse-grained. The visual Gestalt rule (such as proximity, continuity, closure, etc.) is combined to guide the region merging process to ensure that the segmentation result conforms to the laws of human visual perception. At the same time, the principle of semantic consistency is followed, so that parts of the same semantic entity tend to be grouped together, and a multi-scale region candidate pool containing segmentation results of different scales is generated. The region merging strategy based on visual perceptual correlation is applied to the regions in the multi-scale region candidate pool, and the visual correlation between regions, such as texture similarity, color continuity and edge coherence, is analyzed to identify regions that belong to the same entity or scene unit in visual perception. These regions are merged to form a complete unit that is more in line with human perceptual habits, and a set of perceptually consistent regions that maintain consistency in visual perception is obtained. Next, a dynamic density peak clustering algorithm is designed. This algorithm can adapt to the importance distribution characteristics of different scenarios, calculate the local density and relative distance of regions in the importance space, and automatically identify the importance density center, that is, the region with high importance value and the surrounding areas also have high importance. Considering the semantic correlation between regions, semantically related regions (such as regions belonging to the same object or action sequence) tend to be divided into the same cluster, even if they are not completely adjacent in space, realizing collaborative clustering with semantic perception ability, and obtaining semantically enhanced clustering results that take into account the semantic relationship of regions. Then, a temporal stability evaluation mechanism is introduced to analyze the temporal consistency of the semantic enhancement clustering results in consecutive multi-frame images, track the changing trajectory of the region in the temporal dimension, identify important regions that persist in multiple frames and have relatively stable positions, such as protagonists, key objects, etc., filter out temporary or random regions that only appear in a few frames, ensure the temporal consistency of special effects application, form a set of temporally stable important regions, and at the same time calculate the time-varying importance weight function for each region to reflect the changing trend of regional importance over time, such as the importance that gradually increases or decreases with the development of the narrative, and construct a dynamic importance curve representing the temporal change of importance.Finally, combining the temporally stable important region set with the dynamic importance curve, a comprehensive multi-factor regional value evaluation system is designed. The evaluation factors include: region size (area and proportion), importance intensity (average and peak), perceptual complexity (texture complexity and edge density), narrative importance (relevance to the storyline), and special effects adaptability (compatibility of the region with special effects). A multidimensional value score is calculated for each region and weighted aggregated. The regions are prioritized according to the comprehensive score to form a hierarchical set of graded special effects target regions with a clear hierarchy. The region boundaries are optimized through the regional automatic diffusion algorithm to ensure that the edges of special effects applications have a natural transition and attenuation, avoiding a harsh sense of boundaries. At the same time, the region range is adaptively adjusted according to the type of special effects. For example, light effects may require a larger impact range, and particle effects may require more precise boundary control. Ultimately, a set of special effects target regions with precise boundary definitions (contour point sets or parameterized curves) and importance weights (numerical values or functions) is generated. This set precisely specifies the location, range, and intensity of the special effects application, providing clear guidance for parameter configuration and rendering.
[0069] In an embodiment of the present invention, a parametric rendering module is used to apply a special effects parameter configuration scheme to a multimodal feature dataset and generate an initial special effects rendered image using an improved parametric rendering algorithm. Specifically, it is used to:
[0070] The special effect parameter configuration scheme is parsed into a specific rendering instruction sequence, a rendering calculation graph is constructed, and the special effect evolution trajectory is pre-calculated based on the spatiotemporal characteristics of the multimodal feature dataset to obtain a dynamic effect prediction graph;
[0071] A depth estimation model is established for the image sequence in the multimodal feature dataset to generate a scene depth map, and the scene depth map is used to construct a three-dimensional scene structure representation to obtain a stereoscopic scene model;
[0072] Combine the special effects parameter configuration scheme with the three-dimensional scene model to perform physical simulation calculations, simulate the interaction between special effects and scene elements, and obtain a physical constraint special effects behavior model;
[0073] Based on a physically constrained special effects behavior model, we used improved neural radiation field technology to perform perspective-consistent special effects rendering, obtaining preliminary rendering results. We also applied a temporal smoothing algorithm to eliminate inter-frame jitter, improve temporal coherence, and obtain a smooth rendering sequence.
[0074] The smoothed rendering sequence is adaptively mixed with the original image sequence, and the mixing intensity is controlled according to the scene visual importance distribution map to generate a realistic initial special effect rendering image.
[0075] In this embodiment, the special effect parameter configuration scheme is first converted into a specific instruction sequence executable by the rendering engine, including special effect type selection, parameter setting, application area and time control, etc. A directed acyclic graph (DAG) structure is used to organize the rendering operation process, optimize computing dependencies and resource utilization, and build an efficient rendering calculation graph. The temporal and spatial characteristics of the multimodal feature dataset, such as motion trajectory, scene changes and audio rhythm, are analyzed. The evolution trajectory of the special effect over time, including position change, shape change and intensity change, is pre-calculated in combination with the physical characteristics of the special effect. A prediction graph reflecting the dynamic behavior of the special effect is generated to provide a reference for subsequent rendering. Then, a monocular depth estimation network (such as MiDaS, DPT or AdaBins, etc.) is applied to the image sequence in the multimodal feature dataset to infer the depth information of the scene from the single-view image, generate a relative distance representation of each pixel to the camera, and form a scene depth map. The depth map is used in combination with the camera parameters to infer the point cloud data or mesh model in the three-dimensional space, restore the geometric structure of the scene, and construct a stereo scene representation containing depth information to obtain a stereo scene model that can be used for three-dimensional space rendering calculations. Next, the special effect properties in the special effect parameter configuration scheme are combined with the spatial structure in the three-dimensional scene model, and a physical simulation environment is set up, including environmental factors such as gravity field, wind field, and lighting, to perform physics-based special effect behavior simulation, such as fluid dynamics calculation (for water and smoke effects), particle system simulation (for sparks, rain and snow effects), light propagation simulation (for light effects and halo effects), etc., to simulate the interaction between special effects and scene elements, such as particle collision, light reflection, shadow projection, etc., to generate a special effect behavior model that conforms to the laws of physics to ensure the natural and realistic performance of special effects. Then, based on the physical constraint special effects behavior model, the improved neural radiation field (NeRF) technology is used for special effects rendering. The neural radiation field can represent the volume density and directional radiation in three-dimensional space, and is particularly suitable for rendering special effects with translucent characteristics, such as smoke, fog, light, etc. The traditional NeRF is improved by adding the time dimension and physical constraints to achieve a consistent rendering effect of special effects with changes in viewing angle, generate preliminary rendering results, and apply time domain smoothing algorithms to adjacent frames in the rendering sequence, such as temporal convolutional networks or optical flow-guided temporal filtering, analyze and correct the mutations and jitter phenomena between frames, ensure the smooth transition and coherent performance of special effects in the time dimension, especially maintain the stability of special effects in complex dynamic scenes, and obtain a temporally coherent smooth rendering sequence.Finally, the smooth rendering sequence is adaptively mixed with the original image sequence, and an intelligent mixing algorithm is designed to automatically select the most appropriate mixing mode according to the special effect type and scene characteristics, such as alpha mixing, additive mixing, multiplicative mixing, etc. The scene visual importance distribution map is used to control the mixing process, maintaining the integrity of the original content in important areas while allowing special effects to enhance visual performance. Special effects can be allowed to have a stronger presence in secondary areas. Adaptive color correction is applied to ensure that the mixed colors and lighting remain natural and coordinated, generating an initial special effect rendering image that retains the integrity of the original content and enhances visual expression, providing a basis for subsequent parameter optimization.
[0076] In the embodiment of the present invention, please refer to Figure 3 , a logical diagram for perspective-consistent special effects rendering. Specifically, a physical constraint-based special effects behavior model is used to perform perspective-consistent special effects rendering using improved neural radiation field technology, including:
[0077] Construct a parametric neural radiation field network, which uses spatial coordinates and viewing direction as input to predict voxel density and directional radiation characteristics, and obtain a basic radiation field model. Parameters in the special effect parameter configuration scheme are then embedded into the basic radiation field model to achieve controllable adjustment of special effect properties and obtain a parametric special effect radiation field.
[0078] Designing special effects-specific density and emission functions based on parameterized special effects radiation fields, simulating complex special effects physical properties including light scattering, refraction, and luminescence, and deriving special effects physical rendering functions. Combining these special effects physical rendering functions with traditional volume rendering equations, we construct a hybrid rendering pipeline capable of handling special effects physical properties.
[0079] In the hybrid rendering pipeline, adaptive ray sampling is performed for each viewpoint. Based on the scene's visual importance distribution map, the sampling density is increased in visually important areas to improve detail representation, resulting in a differentiated sampling distribution. A gradient-guided sampling strategy is then applied based on the differentiated sampling distribution to reduce computational redundancy, improve rendering efficiency, and form an adaptive sampling ray set.
[0080] Inputting the adaptive sampling ray set into the hybrid rendering pipeline implements perspective-based continuity constraints, ensuring consistent performance of special effects at different perspectives and avoiding jumps when the viewpoint changes. This results in perspective-continuous rendering intermediate results, and adopts a multi-resolution rendering strategy for these perspective-continuous rendering intermediate results. This dynamically allocates computing resources based on regional importance and generates special effects rendering data with multiple levels of sophistication.
[0081] The special effects rendering data with multiple levels of precision are input into the optimized hybrid rendering pipeline through hardware acceleration technology, making full use of the GPU parallel computing capabilities to achieve efficient neural radiation field special effects rendering calculations. At the same time, post-processing optimization algorithms are applied to enhance edge details and depth consistency, ultimately obtaining special effects rendering results that are consistent in perspective and physically reasonable.
[0082] In this embodiment, a parameterized neural radiation field network is first designed and trained. The network architecture is based on a multi-layer perceptron (MLP) with three-dimensional spatial coordinates (x, y, z) and viewing directions (θ, ) as input, and output the voxel density σ and directional radiation color c at the corresponding position. The network includes a position encoding layer, a feature extraction layer, and an output prediction layer. Regularization and feature decoupling are used to improve the generalization ability, form a basic radiation field model, and design a parameter embedding module. The parameters in the special effect parameter configuration scheme (such as transparency, scattering coefficient, luminous intensity, etc.) are embedded in the neural network in the form of conditional encoding. Conditional feature modulation (FiLM) or hypernetwork technology is used to achieve dynamic control of network behavior, so that the special effect properties can be directly adjusted through parameters without retraining the network, and a parametric special effect radiation field is constructed that can dynamically adjust the special effect performance according to the input parameters. Then, based on the parameterized special effects radiation field, special density functions suitable for different types of special effects are designed, such as the volume density model for smoke, the temperature gradient density model for flames, the energy attenuation density model for light effects, etc. Special effects-specific emission functions are designed to simulate the radiation characteristics of special effects, such as the flame luminescence model based on blackbody radiation, the atmospheric light scattering model based on Rayleigh scattering, the water surface reflection model based on the Fresnel effect, etc., combined with the principles of physical optics to simulate complex light interaction effects, including scattering (multi-directional dispersion of light in the medium), refraction (change of direction of light when passing through different media) and luminescence (light generated by the material itself) and other phenomena, integrated to form a special effects physical rendering function, combined with the traditional volume rendering integral equation, extended the standard volume rendering equation to support more complex light propagation behavior, and designed a hybrid rendering pipeline to support the seamless integration of multiple rendering technologies, such as ray tracing, path tracing, volume rendering, etc., which can handle the complex physical characteristics of special effects and maintain computational efficiency. Next, an adaptive ray sampling strategy is implemented in the hybrid rendering pipeline. The sampling density requirements of different areas are determined based on the scene visual importance distribution map analysis. A higher sampling density is assigned to visually important areas (such as main objects and key special effects areas), and the number of rays is increased to achieve detail enhancement. A lower sampling density is used for visually secondary areas to form a differentiated sampling distribution for scene content. Gradient-guided sampling optimization is applied based on the differentiated sampling distribution. The gradient distribution of the intermediate rendering results is analyzed to identify the areas where the rendering results are most sensitive to the sampling points. The sampling density is increased in these areas, and the sampling points in areas with less impact on the results are reduced. Importance sampling is used to reduce computational redundancy and improve the efficiency of computing resource utilization, forming an adaptive sampling ray set that focuses on both visual quality and computational efficiency.Then, the adaptive sampling ray set is input into the hybrid rendering pipeline to implement perspective-based continuity constraints, design a perspective smooth transition mechanism, ensure that special effects maintain a smooth transition during viewpoint changes, avoid the jump phenomenon caused by discrete perspective sampling, and achieve perspective continuity of special effects performance through feature interpolation and constraint optimization, generate rendering intermediate results that remain consistent when observed from any perspective, adopt a multi-resolution rendering strategy for the rendering intermediate results, divide the image into areas of different importance, use high-resolution rendering for important areas, and low-resolution rendering for secondary areas, and then up-sample and synthesize through super-resolution or interpolation technology, dynamically allocate computing resources, and improve the overall rendering efficiency while ensuring the quality of key visual areas, and generate special effects rendering data with different levels of fineness. Finally, the hybrid rendering pipeline is optimized to support GPU acceleration, and CUDA, OpenCL or specific hardware acceleration libraries are used to achieve highly parallel rendering calculations. Rendering data with multiple levels of precision are input into the optimized rendering pipeline for final processing, making full use of the parallel computing capabilities of the GPU to achieve efficient neural radiation field special effects calculations. A series of post-processing optimization algorithms are applied, including edge enhancement (to improve the clarity of special effects boundaries), depth consistency correction (to ensure the rationality of the relationship between special effects and scene depth), detail sharpening (to enhance texture and microstructure), etc., to ultimately generate special effects rendering results that are both perspective consistent (maintaining a coherent visual performance from any angle) and in line with physical laws (following optical and physical principles), providing a high-quality visual foundation for subsequent parameter optimization and final rendering.
[0083] In an embodiment of the present invention, the parameter optimization module is used to perform visual quality evaluation on the initial special effect rendering image, optimize the special effect parameter configuration scheme based on the evaluation results, and generate an optimized special effect parameter set, specifically including:
[0084] A multi-level visual quality assessment framework was constructed, including pixel-level fidelity assessment, structural similarity analysis, and perceptual quality assessment modules to obtain a comprehensive quality score. A / B testing simulations were then conducted to compare the visual effects under different special effects parameters and obtain parameter sensitivity analysis results.
[0085] A computational aesthetic evaluation model was introduced to evaluate the artistic value of rendering effects based on compositional balance, color harmony, and visual rhythm, generating an aesthetic score. This score was then combined with a scene visual importance distribution map to test the targetedness of special effects enhancements and generate an enhancement accuracy index.
[0086] Based on the comprehensive quality score, aesthetic score and enhanced accuracy index, a special effect parameter optimization objective function is constructed, and the Bayesian optimization algorithm is used to explore the parameter space and obtain the parameter optimization direction;
[0087] Generate improved parameter candidate sets in the parameter space, perform rapid rendering evaluation, and apply reinforcement learning methods to automatically select the optimal parameter adjustment strategy to achieve fine-tuning of special effect parameters and obtain improved special effect parameter solutions;
[0088] Verify the stability and rendering consistency of the improved special effects parameter scheme, eliminate possible visual artifacts and time inconsistency problems, and finally generate an optimized special effects parameter set.
[0089] In this embodiment, a multi-level visual quality assessment framework is first constructed, and fidelity assessment is performed at the pixel level. The pixel difference between the special effect rendered image and the original image is calculated, and indicators such as mean square error (MSE) and peak signal-to-noise ratio (PSNR) are used to measure the degree to which the special effects retain the original content. Similarity analysis is performed at the structural level, and the structural similarity index (SSIM) or feature similarity index (FSIM) is used to evaluate whether the special effects rendering retains the key structural information of the original image. Quality assessment is performed at the perceptual level, and a perceptual quality assessment model based on deep learning (such as LPIPS or PieAPP) is used to simulate the human visual system's perceptual evaluation of image quality. The evaluation results of the three levels are integrated, and a weighted average score is calculated to form a comprehensive quality score. At the same time, A / B test simulation is performed to generate rendering result samples under different special effect parameter configurations. A computer vision model is used to simulate human preference judgment, and the degree of influence of different parameter settings on visual effects is compared. The sensitivity of parameter changes to rendering quality is analyzed, and the key parameters that have the greatest impact on the final effect are identified to form parameter sensitivity analysis results. Then, a computational aesthetic evaluation model is introduced. This model is constructed based on art theory and aesthetic principles. It evaluates the balance of composition, analyzes whether the spatial distribution of special effects elements conforms to the principles of visual balance, such as the golden section and the rule of thirds, evaluates color harmony, analyzes whether the color scheme after the introduction of special effects is harmonious and unified, and whether it conforms to the harmonious relationship in color theory, evaluates the sense of visual rhythm, analyzes whether the dynamic changes of special effects create a rhythmic visual experience, and whether it matches the narrative rhythm of the content, calculates the aesthetic score based on comprehensive aesthetic indicators, reflects the artistic expression value of special effects, combines the scene visual importance distribution map to test the targetedness of special effects enhancement, calculates the degree of matching between the special effects intensity distribution and the visual importance distribution, evaluates whether the special effects accurately enhance the key areas that should be highlighted, and avoids over-emphasis on secondary areas, forming an enhancement accuracy indicator to quantify the accuracy of special effects application. Next, based on the comprehensive quality score, aesthetic score and enhanced accuracy index, a multi-objective function for special effects parameter optimization is constructed, and the weight coefficient of each indicator is set to reflect the priority of different quality dimensions to form a comprehensive optimization goal. The Bayesian optimization algorithm is used to explore the high-dimensional parameter space. Bayesian optimization efficiently explores possible optimal parameter combinations by constructing a Gaussian process model of the relationship between parameters and scores, continuously updates the prior probability distribution, quickly converges to the optimal solution area, analyzes the parameter sensitivity results, and identifies the optimization direction, that is, whether the parameters should be increased or decreased to improve the overall score, forming an optimization direction to guide parameter adjustment.Then, according to the parameter optimization direction, multiple groups of improved parameter candidate sets are generated in the parameter space. The parameter change range is based on the parameter sensitivity, and the sensitive parameters are finely changed in small steps. A rapid rendering evaluation is performed on each group of candidate parameters, and a simplified rendering pipeline or proxy model is used for rapid preview to evaluate the effectiveness of the parameter adjustment. Deep reinforcement learning methods are applied to construct a parameter adjustment strategy network, and the rendering effect score is used as a reward signal. Through multiple rounds of iterations, the optimal parameter adjustment strategy is learned, and refined parameter tuning is automatically performed, such as fine-tuning the scattering coefficient, adjusting the light intensity attenuation curve, etc., to generate an improved special effects parameter solution. Finally, the stability of the improved special effects parameter scheme is verified, and the performance stability of the parameters under different scene conditions is tested, such as adaptability under different lighting and different movement speeds. A rendering consistency check is performed, and the temporal continuity of the special effects performance is tested on continuous frame sequences to ensure that there is no flickering or mutation. Potential visual artifacts such as moiré, edge jaggedness, unnatural halos, etc. are analyzed and eliminated. Temporal continuity is optimized to ensure that the special effects transition smoothly with scene changes without abrupt jumps. After completing the comprehensive verification, the optimized special effects parameter set is finally generated. This parameter set has high computational efficiency and cross-scene adaptability while ensuring visual quality, artistic expression and application accuracy, providing optimized parameter configuration for the final rendering.
[0090] In an embodiment of the present invention, the rendering and output module is used to perform final rendering processing on the multimodal feature dataset using the optimized special effects parameter set to generate an AI visual effects image sequence that integrates multimodal perception, specifically for:
[0091] Apply the optimized special effects parameter set to a high-precision rendering engine, perform full-resolution processing on the multimodal feature dataset to obtain the original special effects rendering sequence, and perform high dynamic range processing on the original special effects rendering sequence to enhance the details of highlights and shadows to obtain an HDR-enhanced sequence.
[0092] Based on the rhythm and emotional characteristics of the audio data, the timing characteristics of the dynamic changes of special effects are adjusted to achieve special effects rhythm synchronized with audio and video, resulting in a rhythm-synchronized special effects sequence. Furthermore, based on the narrative information in the scene parameters, the fluctuations in the intensity of the special effects are adjusted to strengthen the narrative expression, resulting in a narrative-enhancing special effects sequence.
[0093] Apply post-production color grading technology to unify the color style of special effects and the original image, ensuring consistency in visual style and obtaining a color-unified special effects sequence. Finally, perform final visual optimization processing, including sharpening, noise reduction, and detail enhancement, to obtain a refined special effects sequence.
[0094] Conduct final quality inspection on the refined special effects sequence to ensure that the rendering quality, emotional expression, and narrative enhancement effect meet the expected goals. The verified special effects sequence is then encoded and packaged according to the target output format to obtain a media file that meets the playback specifications.
[0095] Packaging media files that meet playback specifications with interactive control metadata supports the end user's ability to make limited adjustments to the intensity of special effects, ultimately generating AI visual effects image sequences that integrate multimodal perception.
[0096] In this embodiment, the optimized special effects parameter set is first imported into a high-precision rendering engine, such as a physically based renderer (PBR) or a real-time ray tracing engine, and high-quality rendering settings are configured, including high sampling rate, fine ray tracing, and high-precision physical simulation. The multimodal feature data set is processed at full resolution, and proxy models or low-resolution previews are no longer used. Instead, complete and accurate special effects calculations are performed to generate an original special effects rendering sequence. High dynamic range (HDR) processing technology is applied to the original special effects rendering sequence to expand the dynamic range of the image. A local tone mapping algorithm is used to enhance the detail level and brightness changes in the highlight area, improve the visibility and texture performance of the dark area, ensure the visibility of special effects details in high-contrast scenes, and obtain an HDR enhanced sequence with a wider dynamic range and richer details. Then, the rhythmic features of the audio data are analyzed, the beats, rhythm and energy change points in the audio are extracted, the emotional features of the audio are analyzed, and emotional climaxes, tension or relaxation passages are identified. The audio features are associated with the dynamic parameters of the special effects, so that the intensity, speed, density or color of the special effects and other attributes change synchronously with the audio rhythm. For example, particle bursts are enhanced at the beat points of the music, and the intensity of light effects is enhanced at the emotional climax. This creates an immersive experience of audio-visual collaboration and generates a rhythm-synchronized special effects sequence that is perfectly synchronized with the audio. The narrative information in the scene parameters is further analyzed to understand the emotional intention and narrative function of the current scene in the story. The temporal change curve of the special effects intensity is adjusted according to narrative needs, and the fluctuations in the special effects intensity are designed to strengthen the narrative expression. For example, special effects are enhanced at plot turning points, subtle special effects are maintained in the foreshadowing stage, and the strongest special effects are achieved in the climax. This makes special effects an organic part of the narrative means and generates a narrative-enhancing special effects sequence that can strengthen the emotion and meaning of the story. Next, professional post-production color grading technology is applied to analyze the color style and tonal tendency of the original picture, and the color parameters of the special effects elements are adjusted to keep them consistent with the style of the original picture, ensuring that the special effects will not destroy the established color emotions and visual style of the picture, while maintaining the unique visual characteristics of the special effects, balancing the presence of special effects and the harmony of the overall picture, generating a special effects sequence with a unified color style, and performing the final visual optimization processing. An adaptive sharpening algorithm is applied to enhance the clarity of the special effects edges and textures without introducing over-sharpening artifacts. An advanced noise reduction algorithm is used to eliminate noise and graininess that may be generated during the rendering process, especially at the edges of transparency and in areas of gradual lighting changes. Detail enhancement technology is applied to enhance the microstructure and texture performance of the special effects, enhancing the sense of reality and three-dimensionality. Combining these optimization processes, a refined special effects sequence with the ultimate visual quality is generated.Then, the refined special effects sequence is subjected to a final quality inspection, and objective quality measurement tools are used to evaluate technical indicators such as resolution, bit depth, color accuracy, etc. Through visual comparison, it is verified whether the special effects accurately express the expected emotions and atmosphere, and the enhancing effect of special effects on the narrative is evaluated to confirm whether the story expression and emotional transmission are strengthened. The comprehensive evaluation results ensure that all quality indicators meet or exceed the expected goals, and a verified special effects sequence is obtained. The verified special effects sequence is encoded and packaged according to the target output format, and appropriate codecs (such as H.264, H.265, AV1, etc.) and container formats (such as MP4, MOV, MKV, etc.) are selected, and appropriate bit rates and encoding parameters are set to ensure that a reasonable file size is achieved while maintaining visual quality, and high-quality media files that meet industry playback specifications are generated. Finally, the encoded media files are packaged with interactive control metadata, and a lightweight special effects parameter control interface is designed, allowing end users to adjust certain properties of special effects within a preset range, such as overall intensity, color tendency, or animation speed. An interactive playback control component is developed so that terminal applications can read and apply the user's special effects adjustment instructions, providing a moderately personalized experience while maintaining artistic integrity. Ultimately, an AI visual effects image sequence that integrates multimodal perception and has both professional visual performance and a certain degree of interactive flexibility is generated, providing users with an immersive visual experience.
[0097] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art will be able to modify the technical solutions described in the foregoing embodiments or to substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
[0098] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0099] In the description of the present invention, it should be understood that the terms "first", "second", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0100] In the description of the present invention, unless otherwise specified, "plurality" means two or more.
[0101] In the description of the present invention, “several” means one or more, and “a large number” means two or more.
[0102] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0103] The preset parameters and thresholds in this specification are selected by those skilled in the art according to actual conditions.
[0104] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
Claims
1. An AI visual effects dynamic generation system integrating multimodal perception, characterized by: include: A multimodal input processing module is used to obtain multimodal input data including image sequences, audio data and scene parameters, perform spatiotemporal alignment and normalization on the multimodal input data, and generate a multimodal feature dataset; An importance analysis module is used to extract features from the multimodal feature dataset, construct a multimodal semantic feature map, identify key visual areas of the multimodal semantic feature map, and then generate a scene visual importance distribution map; A parameter configuration module, configured to select an adapted special effect template and generate a special effect parameter configuration scheme based on the scene visual importance distribution map and the multimodal semantic feature map; A parametric rendering module, configured to apply the special effects parameter configuration scheme to the multimodal feature dataset and generate an initial special effects rendered image using an improved parametric rendering algorithm; the module comprises: parsing the special effects parameter configuration scheme into a specific rendering instruction sequence, constructing a rendering calculation graph, and precalculating the special effects evolution trajectory based on the spatiotemporal characteristics of the multimodal feature dataset to obtain a dynamic effect prediction graph; Establishing a depth estimation model for the image sequence in the multimodal feature dataset to generate a scene depth map, and constructing a three-dimensional scene structure representation using the scene depth map to obtain a stereoscopic scene model; Combining the special effect parameter configuration scheme with the three-dimensional scene model, performing physical simulation calculations, simulating the interaction between the special effect and the scene elements, and obtaining a physical-constrained special effect behavior model; Based on the physical constraints of the special effects behavior model, an improved neural radiation field technology is used to perform perspective-consistent special effects rendering to obtain preliminary rendering results. A temporal smoothing algorithm is then applied to eliminate inter-frame jitter, improve temporal coherence, and obtain a smooth rendering sequence. Adaptively blending the smoothed rendering sequence with the original image sequence, controlling the blending intensity according to the scene visual importance distribution map, and generating an initial special effects rendered image with a realistic feel; The improved neural radiation field technology improves the traditional neural radiation field technology by adding time dimension and physical constraints to achieve a rendering effect that remains consistent with the change of viewing angle, and generates a preliminary rendering result; a parameter optimization module, configured to perform a visual quality evaluation on the initial special effect rendered image, optimize the special effect parameter configuration scheme based on the evaluation result, and generate an optimized special effect parameter set; A rendering and output module is used to perform final rendering processing on the multimodal feature dataset using the optimized special effects parameter set to generate an AI visual effects image sequence that integrates multimodal perception.
2. The AI visual effects dynamic generation system integrating multimodal perception according to claim 1 is characterized in that: The step of acquiring multimodal input data including image sequences, audio data, and scene parameters, performing spatiotemporal alignment and normalization processing on the multimodal input data, and generating a multimodal feature dataset comprises: performing key frame extraction and resolution unification processing on the image sequence to obtain a standardized image frame sequence, and performing color space conversion on the standardized image frame sequence to obtain a color-corrected image sequence; Performing sampling rate resampling and spectrum analysis on the audio data to obtain a time-frequency characteristic graph, and performing emotion feature extraction on the time-frequency characteristic graph to obtain an audio emotion feature vector; Performing semantic parsing and hierarchical organization on the scene parameters to obtain a structured parameter representation, and constructing a scene semantic network based on the structured parameter representation to obtain a scene semantic vector; Performing time stamp alignment on the color-corrected image sequence, the audio emotion feature vector, and the scene semantic vector to obtain synchronized multimodal data, and performing feature normalization on the synchronized multimodal data to obtain a standardized feature set; The standardized feature set is fused and encoded to construct a unified representation space to obtain a multimodal feature dataset.
3. The AI visual effects dynamic generation system integrating multimodal perception according to claim 2 is characterized in that: The step of performing key frame extraction and resolution unification processing on the image sequence to obtain a standardized image frame sequence, and performing color space conversion on the standardized image frame sequence to obtain a color-corrected image sequence includes: Applying a visual information entropy measurement model to the image sequence to calculate the information richness of each frame, combining inter-frame visual difference metrics to construct a dual evaluation mechanism, identifying visual information peak frames, and obtaining a key frame candidate set; and applying a semantic importance filtering algorithm to the key frame candidate set to obtain a streamlined key frame set; Analyzing the content composition and subject distribution of each frame in the reduced key frame set, constructing an adaptive content-aware grid, obtaining a content protection map, and performing a non-uniform scaling operation based on the content protection map to obtain a uniform resolution frame with subject enhancement; Applying a super-resolution neural network to the subject-enhanced uniform-resolution frames to reconstruct texture and edge details to obtain detail-restored image frames, and performing temporal consistency optimization on the detail-restored image frames to obtain a temporally coherent standardized image frame sequence; Establishing a scene lighting condition analysis model, inferring ambient light characteristics from the standardized image frame sequence, and selecting a color space that best suits the target special effect to obtain a color gamut optimized image; Semantic segmentation color mapping is performed on the color gamut optimized image, and a viewer visual adaptation compensation mechanism is introduced to ultimately obtain a color-corrected image sequence with enhanced perception.
4. The AI visual effects dynamic generation system integrating multimodal perception according to claim 3 is characterized in that: The step of extracting features from the multimodal feature dataset, constructing a multimodal semantic feature map, identifying key visual areas of the multimodal semantic feature map, and then generating a scene visual importance distribution map includes: Inputting the multimodal feature dataset into a pre-trained multimodal Transformer network, extracting deep semantic features to obtain a multi-level feature tensor, and performing a channel attention mechanism on the multi-level feature tensor to obtain an enhanced feature map; Using a graph convolutional network to model the spatial relationship of the enhanced feature map, construct a scene graph structure representation, obtain a scene element relationship network, and perform semantic segmentation based on the scene element relationship network to obtain a semantic region partition map; Cross-modally fusing the semantic region division map with the audio emotion feature vector to obtain an emotion-enhanced semantic feature map, and integrating multi-scale information through an adaptive feature pyramid network to obtain a multimodal semantic feature map; Applying a visual saliency detection algorithm to the multimodal semantic feature map, calculating the attention weight of each region to obtain an initial saliency map, and calibrating the initial saliency map in combination with a human visual perception model to obtain a perceptually corrected saliency map; The perceptually corrected saliency map is weightedly fused with the semantic region partition map to construct a saliency heat map that considers semantic importance, thereby obtaining a scene visual importance distribution map.
5. The AI visual effects dynamic generation system integrating multimodal perception according to claim 4 is characterized in that: The method of using a graph convolutional network to perform spatial relationship modeling on the enhanced feature map, constructing a scene graph structure representation, and obtaining a scene element relationship network includes: Generating region candidates on the enhanced feature map, extracting potential scene entity regions to obtain a region candidate set, and performing attribute prediction and classification on the region candidate set to obtain a semantically labeled region set; Constructing a feature node for each region in the semantically labeled region set, calculating the spatial and semantic similarities between regions, generating a weighted adjacency matrix, and forming an initial scene graph; Applying multi-layer graph convolution operations to transfer and update information on the initial scene graph to obtain a context-aware node representation, and predicting semantic relationships between nodes based on the context-aware node representation to obtain a relationship labeling graph; Interactively integrating the relationship labeling graph with global scene features to obtain a global correlation scene graph, and performing attention-guided information screening on the global correlation scene graph to highlight key relationships; A hierarchical scene understanding model is constructed based on the global association scene graph, which supports multi-granularity scene parsing and ultimately obtains a scene element relationship network.
6. The AI visual effects dynamic generation system integrating multimodal perception according to claim 5 is characterized in that: The step of selecting an adapted special effect template based on the scene visual importance distribution map and the multimodal semantic feature map and generating a special effect parameter configuration scheme includes: Build a special effects template knowledge base, including parameterized templates for multiple categories of visual effects, and create semantic labels and applicable scenario descriptions for each template to obtain a semantic special effects library; Performing regional clustering on the scene visual importance distribution map to obtain a set of special effect target regions, and extracting scene emotional tone and style features based on the multimodal semantic feature map to obtain a scene style representation; Combining the features of the special effect target area set with the scene style representation to construct a special effect demand vector, and performing similarity matching between the special effect demand vector and the templates in the semantic special effect library to obtain a candidate special effect template set; Applying a multi-objective evaluation mechanism to the candidate special effect template set, evaluating it from multiple dimensions such as visual coordination, emotional matching, and computational complexity, obtaining a template score list, and selecting the optimal special effect template combination based on the template score list to obtain a special effect application solution; According to the special effect application plan, initial parameter values are calculated for each selected special effect template, and the parameter strength and application range are adjusted in combination with the scene visual importance distribution map, and finally a special effect parameter configuration plan is formed.
7. The AI visual effects dynamic generation system integrating multimodal perception according to claim 6 is characterized in that: The performing regional clustering on the scene visual importance distribution map to obtain a set of special effect target regions includes: Applying an adaptive threshold segmentation algorithm to the scene visual importance distribution map, automatically determining an optimal segmentation threshold according to global importance distribution characteristics, obtaining a preliminary region partition map, and performing noise removal and boundary smoothing on the preliminary region partition map through morphological operations to obtain a refined region mask; A multi-level image segmentation tree is constructed based on the fine region mask, and a multi-scale region candidate pool is obtained by combining the visual Gestalt law and the semantic consistency principle. A region merging strategy based on visual perceptual correlation is applied to the multi-scale region candidate pool to obtain a perceptually consistent region set. A dynamic density peak clustering algorithm is designed to adaptively cluster the areas where the perceptual consistency areas are concentrated, and obtain semantically enhanced clustering results; A temporal stability evaluation mechanism is introduced to analyze the consistency of the semantic enhancement clustering results across frames to obtain a set of temporally stable important regions. At the same time, a time-varying importance weight is calculated for each region to construct a dynamic importance curve. Combining the temporally stable important region set with the dynamic importance curve, a multi-target region value assessment system is designed. Taking into account the region size, importance intensity, perceptual complexity, narrative importance and special effects adaptability, regional priority is sorted to obtain a graded special effects target region set. The region boundaries are optimized through a regional automatic diffusion algorithm to finally generate a special effects target region set.
8. The AI visual effects dynamic generation system integrating multimodal perception according to claim 1 is characterized in that: The special effects behavior model based on the physical constraints uses an improved neural radiation field technology to perform perspective-consistent special effects rendering, including: Constructing a parameterized neural radiation field network, taking spatial coordinates and viewing direction as input, predicting voxel density and directional radiation characteristics, obtaining a basic radiation field model, and embedding parameters in the special effect parameter configuration scheme into the basic radiation field model to obtain a parameterized special effect radiation field; Designing a density function and an emission function dedicated to special effects based on the parameterized special effect radiation field to obtain a special effect physical rendering function, and combining the special effect physical rendering function with a traditional volume rendering equation to construct a hybrid rendering pipeline; Adaptively sampling rays for each viewpoint in the hybrid rendering pipeline, increasing sampling density in visually important areas according to the scene visual importance distribution map to obtain a differentiated sampling distribution, and applying a gradient-guided sampling strategy based on the differentiated sampling distribution to form an adaptive sampling ray set; Inputting the adaptive sampling ray set into the hybrid rendering pipeline to obtain a perspective-continuous rendering intermediate result, and applying a multi-resolution rendering strategy to the perspective-continuous rendering intermediate result, dynamically allocating computing resources according to regional importance, and generating special effect rendering data; The special effects rendering data is input into the optimized hybrid rendering pipeline through hardware acceleration technology, and a post-processing optimization algorithm is applied to enhance edge details and depth consistency, ultimately obtaining a special effects rendering result.
9. The AI visual effects dynamic generation system integrating multimodal perception according to claim 1 is characterized in that: The performing visual quality evaluation on the initial special effect rendering image, optimizing the special effect parameter configuration scheme based on the evaluation result, and generating an optimized special effect parameter set includes: Construct a multi-level visual quality assessment framework, including pixel-level fidelity assessment, structural similarity analysis, and perceptual quality assessment modules to obtain a comprehensive quality score. A / B test simulations were also conducted to obtain parameter sensitivity analysis results. A computational aesthetic evaluation model is introduced to evaluate the artistic value of rendering effects from the aspects of compositional balance, color harmony, and visual rhythm to obtain an aesthetic score. This score is then combined with the scene visual importance distribution map to test the pertinence of special effects enhancement and obtain an enhancement accuracy index. Based on the comprehensive quality score, the aesthetic score, and the enhanced accuracy index, a special effect parameter optimization objective function is constructed, and a Bayesian optimization algorithm is used to explore the parameter space to obtain a parameter optimization direction; Generate improved parameter candidate sets in the parameter space, perform fast rendering evaluation, and apply reinforcement learning methods to automatically select the optimal parameter adjustment strategy to obtain improved special effects parameter solutions; Verify the stability and rendering consistency of the improved special effects parameter scheme, and finally generate an optimized special effects parameter set.
10. The AI visual effects dynamic generation system integrating multimodal perception according to claim 9 is characterized in that: The method of performing final rendering processing on the multimodal feature dataset using the optimized special effects parameter set to generate an AI visual special effects image sequence integrating multimodal perception includes: Applying the optimized special effects parameter set to a high-precision rendering engine, performing full-resolution processing on the multimodal feature dataset to obtain an original special effects rendering sequence, and performing high dynamic range processing on the original special effects rendering sequence to obtain an HDR enhanced sequence; Adjusting the timing characteristics of the dynamic changes of the special effects based on the rhythm and emotional characteristics of the audio data to obtain a rhythm-synchronized special effects sequence, and adjusting the fluctuations in the intensity of the special effects based on the narrative information in the scene parameters to obtain a narrative-enhancing special effects sequence; Apply post-production color grading technology to unify the color style of special effects and the original image to obtain a color-unified special effects sequence. Then perform final visual optimization processing, including sharpening, noise reduction, and detail enhancement, to obtain a refined special effects sequence. Performing a final quality inspection on the refined special effects sequence to obtain a verified special effects sequence, and encoding and packaging the verified special effects sequence according to a target output format to obtain a media file that meets playback specifications; The media files that meet the playback specifications are packaged with interactive control metadata to support the end user to make limited adjustments to the special effects intensity, and ultimately generate an AI visual effects image sequence that integrates multimodal perception.
Citation Information
Patent Citations
AR / VR real-time scene construction and rendering system for multi-dimensional scene matching
CN118071933A
Special effect video generation method and device, electronic equipment and readable storage medium
CN118138830A