AI visual special effect dynamic generation system fused with multi-modal perception

The system addresses inefficiencies in visual effects creation by using AI to process multi-modal inputs, analyze importance, and optimize parameters, resulting in efficient, high-quality, and artistically expressive visual effects that align with content semantics and emotional expression.

CN120318379AActive Publication Date: 2025-07-15SHENZHEN XINGHUO MUTUAL ENTERTAINMENT DIGITAL TECH CO LTD

Patent Information

Application Number
CN202510780313.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-15
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The existing visual special effects creation relies on manual adjustments, lacks scientific basis, and it is difficult to accurately perceive the semantic structure of the content, resulting in the special effects being overwhelming or misplaced applications, making it impossible to achieve customized integration with the content, and the production cycle and cost are uncontrollable.

Method used

The AI visual effects dynamic generation system that integrates multimodal perception generates high-quality visual effects that are harmonious and unified with the content through multimodal input processing, importance analysis, parameter configuration and parameterized rendering.

Benefits of technology

It improves the efficiency and quality of visual special effects creation, lowers the creative threshold, and allows ordinary users to create professional-level effects, reduces production cycle and cost, and enhances the harmonious unity of special effects and content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318379A_ABST
    Figure CN120318379A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of visual special effects, and discloses an AI visual special effect dynamic generation system fusing multi-modal perception. Through cooperative work of six core modules of multi-modal input processing, importance analysis, parameter configuration, parameterized rendering, parameter optimization and rendering and output, deep fusion of special effects and contents is realized. An image sequence, audio data and scene parameters can be processed at the same time, a multi-modal feature data set is constructed, a scene key visual area is scientifically recognized, a visual importance distribution map is generated, and the special effect parameter configuration and rendering process is guided. An improved neural radiation field technology and a physical constraint model are adopted to ensure the consistency and reality of the special effect at different visual angles; and a multi-dimensional quality evaluation and parameter automatic optimization mechanism is introduced, so that the visual expressive force and the artistic value of the special effect are guaranteed. The creation threshold is remarkably reduced, the cooperative expression ability of the special effect and the content is improved, and the special effect becomes a powerful tool for enhancing the narration and enhancing the emotion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of visual effects, and more specifically, to an AI visual effects dynamic generation system that integrates multi-modal perception. Background Art

[0002] As a core component of digital content creation, visual effects technology has experienced a leap from traditional manual synthesis to computer graphics-driven development. Early visual effects mainly relied on physical props and optical technologies. With the rapid development of computer graphics, technologies such as 3D modeling, particle systems, and fluid simulation have gradually become the mainstream methods for special effects creation. In recent years, technologies such as physically based rendering (PBR) and real-time ray tracing have further enhanced the realism and immersive experience of special effects, bringing revolutionary changes to fields such as movies, games, and virtual reality.

[0003] However, the current field of visual effects creation still faces technical bottlenecks and creative dilemmas, significantly restricting the improvement of the expressiveness of digital content. In post-production teams of film and television, a large amount of time is often spent on tedious manual special effects adjustments. For a simple light effect, technical artists often need to repeatedly adjust dozens of parameters to meet the director's expectations, and these adjustments lack scientific basis and mainly rely on empirical judgment. In actual production, it is difficult for special effects artists to accurately perceive the semantic structure of the content, resulting in special effects often overshadowing the main content or being misapplied. For example, in a movie, a key dialogue scene is distracted by overly elaborate background special effects, weakening the emotional transmission. Live broadcast and short video creators face more severe challenges. They lack professional special effects knowledge and need to produce content quickly, so they can only use preset templates and cannot achieve customized special effects that are deeply integrated with the content. In game development, special effects often show physically unreasonable phenomena from different perspectives. For example, a magic effect suddenly changes its form when the perspective changes, destroying the immersion. What is more troublesome is the problem of unstable timing of special effects in dynamic scenes. For example, the flickering and jittering of particle effects in VR experiences cause discomfort to users. Special effects creation has long remained at the level of isolated visual processing, ignoring the coordination of audio content and emotional expression, making a large number of exquisite special effects become pure decorations rather than narrative enhancement tools. The industry evaluation criteria are subjectively vague, and it is difficult for creative teams to quantify the quality of special effects. The project cycle and cost are uncontrollable due to repeated modifications, severely restricting the large-scale production capacity of innovative content and unable to meet the explosive demand for high-quality audio-visual content in the contemporary market.

[0004] In view of this, the present invention proposes an AI visual effects dynamic generation system that integrates multi-modal perception to solve the above problems. Summary of the Invention

[0005] To overcome the above-mentioned defects of the prior art and to achieve the above object, the present invention provides the following technical solution: An AI visual effects dynamic generation system that integrates multi-modal perception, including: A multi-modal input processing module, which is used to obtain multi-modal input data including an image sequence, audio data, and scene parameters, perform spatio-temporal alignment and normalization processing on the multi-modal input data, and generate a multi-modal feature data set; An importance analysis module, which is used to extract features from the multi-modal feature data set, construct a multi-modal semantic feature map, identify key visual regions of the multi-modal semantic feature map, and then generate a scene visual importance distribution map; A parameter configuration module, which is used to select a suitable special effect template according to the scene visual importance distribution map and the multi-modal semantic feature map, and generate a special effect parameter configuration scheme; A parameterized rendering module, which is used to apply the special effect parameter configuration scheme to the multi-modal feature data set, and generate an initial special effect rendered image through an improved parameterized rendering algorithm; A parameter optimization module, which is used to perform a visual quality assessment on the initial special effect rendered image, optimize the special effect parameter configuration scheme based on the assessment result, and generate an optimized special effect parameter set; A rendering and output module, which is used to perform final rendering processing on the multi-modal feature data set using the optimized special effect parameter set, and generate an AI visual special effect image sequence integrating multi-modal perception.

[0006] The technical effects and advantages of the AI visual special effect dynamic generation system integrating multi-modal perception of the present invention: The present invention improves the efficiency and quality of visual special effect creation, changes the traditional special effect production process that relies on manual trial and error and empirical judgment, and enables creators to devote more energy to artistic conception rather than cumbersome technical adjustments. By deeply understanding content semantics and emotional expression, the special effects are no longer just decorative elements, but become a powerful tool to enhance narrative, strengthen emotions, and guide the audience's attention, greatly improving the expressiveness and immersion of audio-visual content. The significant reduction of the creation threshold enables ordinary users to easily create professional-level visual effects, while film and television production institutions can standardize and scale the special effect creation process while maintaining artistic personalized expression. The harmonious unity of special effects and content eliminates the problems of traditional special effects being overwhelming or having a fragmented style. By precisely enhancing key visual regions, the special effects can strengthen rather than interfere with the audience's attention to the core content. The intelligent adaptability of the system enables it to automatically adjust the special effect style and intensity according to different content types, emotional tones, and creative intentions, significantly shortening the production cycle and reducing the production cost while ensuring high-quality output. Description of the Drawings

[0007] Figure 1 It is a schematic diagram of the AI visual special effect dynamic generation system integrating multi-modal perception of the present invention; Figure 2Schematic logic diagram for identifying key special effect application areas in the present invention; Figure 3 Schematic logic diagram for performing special effect rendering with consistent perspectives in the present invention. Detailed implementation manners

[0008] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0009] This application example provides an AI visual special effect dynamic generation system integrating multi-modal perception. The execution entities of the AI visual special effect dynamic generation system integrating multi-modal perception include, but are not limited to, the following systems that carry this system: digital content creation platforms, video processing systems, film and television post-production tools, real-time rendering engines, augmented reality applications, etc., which can be regarded as general computing nodes of this application. The data processing platform includes, but is not limited to, at least one of the following: multi-modal feature extraction systems, visual importance analysis systems, and parameterized rendering systems.

[0010] Please refer to Figure 1 , the present invention provides an AI visual special effect dynamic generation system integrating multi-modal perception, including the following modules: A multi-modal input processing module, which is used to obtain multi-modal input data including image sequences, audio data, and scene parameters, perform spatio-temporal alignment and normalization processing on the multi-modal input data, and generate a multi-modal feature data set; An importance analysis module, which is used to extract features from the multi-modal feature data set, construct a multi-modal semantic feature map, identify the key visual areas of the multi-modal semantic feature map, and then generate a scene visual importance distribution map; A parameter configuration module, which is used to select an appropriate special effect template according to the scene visual importance distribution map and the multi-modal semantic feature map, and generate a special effect parameter configuration scheme; A parameterized rendering module, which is used to apply the special effect parameter configuration scheme to the multi-modal feature data set, and generate an initial special effect rendering image through an improved parameterized rendering algorithm; A parameter optimization module, which is used to evaluate the visual quality of the initial special effect rendering image, optimize the special effect parameter configuration scheme based on the evaluation results, and generate an optimized special effect parameter set; A rendering and output module, which is used to perform final rendering processing on the multi-modal feature data set using the optimized special effect parameter set, generate an AI visual special effect image sequence integrating multi-modal perception, and the various modules are connected by wired and / or wireless means to realize data transmission between the modules.

[0011] The present invention realizes the comprehensive perception of image sequences, audio, and scene parameters through multi-modal input processing, constructs a multi-modal feature data set to provide a data basis for subsequent analysis, conducts importance analysis based on multi-modal features to enable semantic perception ability in special effect generation, the scene visual importance distribution map can accurately guide the application position and intensity of special effects, selects an appropriate special effect template according to the importance distribution and semantic features to improve the context adaptability of special effects, the special effect parameter configuration scheme considers various perception factors to ensure the artistic expressiveness of special effects, realizes the preliminary rendering of high-quality special effects through an improved parametric rendering algorithm, the initial special effect rendering image maintains the visual integrity of the original content, conducts visual quality evaluation and parameter optimization on the rendering result to enhance the visual performance of special effects, and the optimized special effect parameter set ensures the visual coherence and perceptual harmony of special effects. The finally rendered special effect image sequence generated has multi-modal coordination and high visual quality.

[0012] In an embodiment of the present invention, a multi-modal input processing module is used to obtain multi-modal input data including image sequences, audio data, and scene parameters, perform spatio-temporal alignment and normalization processing on the multi-modal input data, and generate a multi-modal feature data set, specifically for: Extract key frames from the image sequence and perform resolution unification processing to obtain a standardized image frame sequence, and perform color space conversion on the standardized image frame sequence to obtain a color-corrected image sequence; Perform sampling rate resampling and spectral analysis on the audio data to obtain a time-frequency characteristic map, and perform emotional feature extraction on the time-frequency characteristic map to obtain an audio emotional feature vector; Perform semantic parsing and hierarchical organization on the scene parameters to obtain a structured parameter representation, and construct a scene semantic network based on the structured parameter representation to obtain a scene semantic vector; Align the time stamps of the color-corrected image sequence, the audio emotional feature vector, and the scene semantic vector to obtain synchronized multi-modal data, and perform feature normalization on the synchronized multi-modal data to obtain a standardized feature set; Fuse and encode the standardized feature set to construct a unified representation space to obtain a multi-modal feature data set.

[0013] In this embodiment, first, the input image sequence is collected, and information such as the visual features, timestamps, and resolutions of each frame of the image are obtained through computer vision algorithms to form an image sequence set containing complete visual information. Key frame extraction is performed on the obtained image sequence set, and methods such as scene change detection, content change measurement, and visual attention models are used to identify key frames to ensure the capture of significant change points in the visual content. Resolution unification processing is performed on the extracted key frames, including operations such as scaling, cropping, and pixel resampling, so that all frames have consistent resolutions and aspect ratios, forming a standardized image frame sequence. Color space conversion is performed on the standardized image frame sequence, converting the RGB color space to a color space more suitable for special effects processing (such as HSL, LAB, etc.), and color correction is performed, including white balance adjustment, tone mapping, and color enhancement, etc., finally obtaining a color-corrected image sequence. At the same time, audio data is extracted from the input data, including various sound elements such as music, dialogue, ambient sound, and sound effects. The audio data is resampled to unify the audio with different sampling rates to a standard sampling rate (such as 44.1 kHz or 48 kHz) to ensure processing consistency. Spectral analysis is performed on the resampled audio, and methods such as short-time Fourier transform (STFT), wavelet transform, or mel-frequency cepstral coefficients (MFCC) are used to extract the time-frequency characteristics of the audio, generating a time-frequency characteristic map reflecting the audio energy distribution and changes. An emotion feature extraction algorithm is applied to the time-frequency characteristic map to analyze features such as the rhythm, emotional color, and energy change of the audio, and a deep learning model (such as CNN, LSTM, etc.) is used to extract high-order emotion features, forming an audio emotion feature vector representing the audio emotion and rhythm features. In addition, scene parameter inputs are received, including high-level semantic information such as scene type, narrative stage, emotional tone, and special effects intention. Semantic parsing is performed on the scene parameters to identify keywords, emotion labels, and intention instructions, understand the association relationships between the parameters, hierarchically organize the parsing results to form a structured semantic tree or knowledge graph, obtain a structured parameter representation, construct a scene semantic network based on the structured parameter representation, and use a graph neural network or attention mechanism model to capture the complex relationships between the parameters, generating an abstract semantic representation of the scene to form a scene semantic vector. Then, the color-corrected image sequence, audio emotion feature vector, and scene semantic vector are aligned according to the timestamps to ensure the synchronization of different modality data in the time dimension, process the time delay and duration differences between different modality data, construct a time association mapping, form synchronized multi-modal data, and perform normalization processing on different features in the synchronized multi-modal data, converting features with different scales and dimensions to a unified numerical range, using methods such as standardization, min-max scaling, or rank normalization to balance the weights of different modality features and prevent a certain modality from dominating the feature space, obtaining a standardized feature set.Finally, the standardized feature set is input into the multi-modal feature fusion network. Using early fusion, mid-term fusion, or late fusion strategies, cross-attention calculation or tensor fusion is performed on features of different modalities to construct a feature space that can uniformly represent multi-modal information. The Transformer architecture or cross-modal autoencoder is used to achieve semantic alignment and complementary enhancement of features, and finally a multi-modal feature data set containing rich audio-visual semantic information is generated. This data set not only retains the key features of each modality but also captures the interaction relationships between modalities, providing a comprehensive data basis for subsequent importance analysis.

[0014] In the embodiments of the present invention, key frame extraction and resolution unification processing are performed on the image sequence to obtain a standardized image frame sequence, and color space conversion is performed on the standardized image frame sequence to obtain a color-corrected image sequence, including: Apply a visual information entropy measurement model to the image sequence, calculate the information richness of each frame, combine the inter-frame visual difference metric, construct a dual evaluation mechanism, identify the visual information peak frames to obtain a candidate key frame set, and apply a semantic importance filtering algorithm to the candidate key frame set to retain the narrative key points and visual turning points to obtain a refined key frame set; Analyze the content composition and subject distribution of each frame in the refined key frame set, construct an adaptive content-aware grid to guide the resolution adjustment process to obtain a content protection mapping graph, and perform a non-uniform scaling operation based on the content protection mapping graph to achieve resolution standardization while maintaining the integrity of the visual subject, obtaining a unified resolution frame with enhanced subject; Apply a super-resolution neural network to the unified resolution frame with enhanced subject to intelligently repair the detail loss caused by scaling, reconstruct the texture and edge details to obtain a detail-restored image frame, and perform time-domain consistency optimization on the detail-restored image frame to eliminate the fluctuations in the inter-frame detail performance, obtaining a temporally coherent standardized image frame sequence; Establish a scene illumination condition analysis model, infer the environmental light characteristics from the standardized image frame sequence, and select the color space most suitable for the target special effects to achieve a perceptually consistent color space conversion to obtain a gamut-optimized image; Perform semantic segmentation color mapping on the gamut-optimized image, apply customized color enhancement strategies to different semantic regions while maintaining global color harmony, and introduce a viewer visual adaptability compensation mechanism to optimize the color performance in different viewing environments, finally obtaining a perceptually enhanced color-corrected image sequence.

[0015] In this embodiment, first, apply the visual information entropy measurement model to the input image sequence, calculate the Shannon entropy or relative entropy value of each frame of the image, evaluate the information content and complexity it contains, and at the same time calculate the visual differences between adjacent frames. Use methods such as structural similarity (SSIM), feature point matching differences, or depth feature distances to quantify the degree of change between frames. Combine the two indicators of information entropy and inter-frame differences to construct a dual evaluation mechanism, set an adaptive threshold to identify the mutation points of information content and the frames with significant visual changes, and screen out the frames with rich information and significant changes as the candidate key frame set. Apply the semantic importance filtering algorithm to the candidate key frame set, and use a pre-trained visual semantic model (such as CLIP or ViT) to evaluate the semantic importance and narrative value of each frame, identify the key points with story-turning significance (such as scene switching, emotional change points, etc.), retain the frames with high semantic importance and large content differences from each other, and filter out the frames with semantic redundancy or visual similarity to obtain a refined but information-rich key frame set. Then, analyze the content composition and object distribution of each frame in the refined key frame set, use object detection and semantic segmentation techniques to identify the main objects, background elements, and their spatial distribution in the image, construct an adaptive content-aware grid based on visual saliency and semantic importance, divide the image into regions of different importance levels, form a grid structure with variable density, where the grid density of important content regions is high and that of secondary regions is low, generate a content protection mapping diagram reflecting the importance distribution of the content, perform a non-uniform scaling operation based on the content protection mapping diagram, adopt a strategy of maintaining the original ratio or slightly scaling for important content regions, and adopt a more aggressive scaling strategy for secondary regions. Use a content-aware scaling algorithm (such as Seam Carving or grid deformation) to achieve resolution normalization, while maximizing the integrity and details of the visual object, to obtain uniformly resolved frames with enhanced objects. Next, input the uniformly resolved frames with enhanced objects into a super-resolution neural network (such as SRGAN, ESRGAN, or Real-ESRGAN, etc.), intelligently repair and enhance the texture details and edge features lost during the scaling process. The network learns the feature distribution of high-resolution images, reconstructs the blurred or jagged details caused by scaling, restores the texture richness and edge sharpness, and generates a frame sequence of detail-restored images. Apply the temporal domain consistency optimization algorithm to the frame sequence of detail-restored images, analyze the coherence of detail performance between adjacent frames, eliminate the inter-frame detail fluctuations and flickering phenomena caused by single-frame super-resolution processing, and ensure the smooth transition of texture, color, and details in the time dimension. Finally, obtain a standardized image frame sequence that not only maintains high detail quality but also has temporal coherence.Then, establish a scene lighting condition analysis model, extract lighting features from the standardized image frame sequence, analyze the light source direction, intensity, color temperature, and scattering characteristics, etc., infer the environmental lighting conditions (such as indoor, outdoor, sunlight, artificial light, etc.), and select the most suitable color space for special effect rendering according to the lighting conditions and the type of target special effect (such as flame, water, light effect, etc.). For example, HDR special effects are suitable for the ACES color space, and transparency-related special effects are suitable for the linear RGB color space. Maintain perceptual consistency during color space conversion to ensure that the visual perception before and after conversion does not change significantly, and obtain a gamut-optimized image. Finally, perform semantic segmentation color mapping on the gamut-optimized image. First, use the semantic segmentation algorithm to divide the image into regions with different semantics (such as people, background, objects, etc.), design customized color enhancement strategies for each semantic region, such as emphasizing the naturalness of the skin color of people and enhancing the color saturation of key objects, apply local color enhancement while maintaining global color harmony, ensure that the color processing of different regions is aesthetically unified as a whole, introduce an audience visual adaptation compensation mechanism, consider the color perception differences in different viewing environments (such as cinemas, home TVs, mobile devices, etc.), optimize the color performance so that it can maintain the best visual effect under various display devices and ambient light conditions, and finally obtain a color correction image sequence that conforms to the artistic intention and has perceptual enhancement characteristics, providing a high-quality visual basis for subsequent special effect generation.

[0016] In the embodiment of the present invention, the importance analysis module is used to extract features from the multi-modal feature dataset, construct a multi-modal semantic feature map, identify the key visual regions of the multi-modal semantic feature map, and then generate a scene visual importance distribution map, specifically for: Input the multi-modal feature dataset into the pre-trained multi-modal Transformer network to extract deep semantic features, obtain a multi-level feature tensor, and perform channel attention mechanism processing on the multi-level feature tensor to obtain an enhanced feature map; Use the graph convolutional network to model the spatial relationship of the enhanced feature map, construct a scene graph structure representation, obtain a scene element relationship network, and perform semantic segmentation based on the scene element relationship network to obtain a semantic region division map; Perform cross-modal fusion on the semantic region division map and the audio emotion feature vector to obtain an emotion-enhanced semantic feature map, and integrate multi-scale information through an adaptive feature pyramid network to obtain a multi-modal semantic feature map; Apply a visual saliency detection algorithm to the multi-modal semantic feature map, calculate the attention weight of each region to obtain an initial saliency map, and calibrate the initial saliency map in combination with the human visual perception model to obtain a perceptually corrected saliency map; Perform weighted fusion on the perceptually corrected saliency map and the semantic region division map to construct a saliency heat map considering semantic importance, and obtain a scene visual importance distribution map.

[0017] In this embodiment, first, the multi-modal feature dataset is input into a pre-trained multi-modal Transformer network. This network has been pre-trained with a large amount of multi-modal data and has the ability of cross-modal semantic understanding. It uses the multi-head self-attention mechanism to process the relationships between different modalities, extracts the deep semantic features of image, audio, and scene parameters, generates a feature tensor containing multi-level semantic information, applies the channel attention mechanism to the multi-level feature tensor, automatically identifies and enhances the feature channels relevant to the current content, suppresses the irrelevant channels, realizes the adaptive enhancement of features, dynamically adjusts the weights of each channel through the attention gating function, highlights the semantic key information, and forms an enhanced feature map. Then, the enhanced feature map is input into a graph convolutional network (GCN). The visual elements in the scene are regarded as nodes in the graph structure, and the spatial relationships and semantic associations between the elements are regarded as edges. The complex scene relationships are modeled through multi-layer graph convolutional operations, the interaction relationships and context dependencies between entities are captured, a scene graph structure representation containing rich relationship information is constructed, a scene element relationship network is formed, and a semantic segmentation task is performed based on the scene element relationship network. The relationship information learned by the network is used to guide the segmentation process, improve the semantic consistency of the segmentation, divide the image into regions with clear semantic labels, such as people, objects, backgrounds, etc., and generate a semantic region segmentation map. Then, the semantic region segmentation map is fused with the audio emotion feature vector cross-modally. A cross-modal attention mechanism is designed so that the audio emotion feature can modulate the expression intensity of the visual semantic region. For example, intense audio will enhance the emotional representation of the corresponding visual region, realizing the collaborative enhancement of audio-visual semantics. The multi-modal co-occurrence pattern is captured through the interaction learning between modalities, and a sentiment-enhanced semantic feature map is obtained. The sentiment-enhanced semantic feature map is input into an adaptive feature pyramid network. This network extracts and integrates features at multiple scales, captures both fine-grained local features and coarse-grained global semantics, integrates features at different scales through the top-down and bottom-up two-way information flow, realizes the balanced expression of details and semantics, and finally generates a multi-modal semantic feature map. This feature map synthesizes the semantic information of visual, audio, and scene parameters. Then, a visual saliency detection algorithm is applied to the multi-modal semantic feature map to calculate the degree to which each region in the image attracts visual attention. A biologically inspired attention model or a deep learning saliency detection network (such as SAM, SalGAN, etc.) is used to generate an initial saliency map reflecting the visual attention distribution. The initial saliency map is combined with a human visual perception model. This model is based on psychophysical experiments and visual cognitive theories, simulates the attention allocation mechanism of the human visual system to the scene, considers visual characteristics such as central preference, face priority, and motion sensitivity, calibrates the saliency map to make it more in line with real human visual perception, and obtains a perception-corrected saliency map.Finally, the perceptual correction saliency map and the semantic region partition map are weighted and fused to assign importance weights to different semantic regions. For example, higher weights are assigned to characters or key objects that are important for the narrative, and lower weights are assigned to secondary background elements, to construct a heat map that takes into account both visual saliency and semantic importance. This heat map can accurately indicate the comprehensive importance degree of each region in the scene, and finally generate a scene visual importance distribution map, which will guide the subsequent special effect parameter configuration and rendering process to ensure that the special effects are applied to the most appropriate regions and enhance the visual narrative effect.

[0018] In the embodiment of the present invention, a graph convolutional network is used to model the spatial relationship of the enhanced feature map, construct a scene graph structure representation, and obtain a scene element relationship network, including: Generate region candidates for the enhanced feature map, extract potential scene entity regions to obtain a region candidate set, and perform attribute prediction and classification on the region candidate set to obtain a semantic labeled region set; Construct feature nodes for each region in the semantic labeled region set, and at the same time calculate the spatial and semantic similarities between regions to generate a weighted adjacency matrix, forming an initial scene graph; Apply multi-layer graph convolutional operations to the initial scene graph for information transmission and update to realize context enhancement of node features, obtain context-aware node representations, and predict the semantic relationships between nodes based on the context-aware node representations to obtain a relationship labeled graph; Interactively integrate the relationship labeled graph with the global scene features to capture long-range dependencies, obtain a globally associated scene graph, and perform attention-guided information screening on the globally associated scene graph to highlight key relationships; Construct a hierarchical scene understanding model based on the globally associated scene graph to support multi-granularity scene parsing, and finally obtain a scene element relationship network.

[0019] In this embodiment, first, an algorithm for generating candidates for the enhanced feature mapping application area, such as the Region Proposal Network (RPN), Selective Search, or DeepMask, etc., is applied to identify image regions that may contain meaningful entities, generating a set of region candidates containing bounding box coordinates. Feature extraction is performed on each candidate region, and the convolutional neural network is used to extract the visual feature representation of the region. The extracted features are used for attribute prediction and classification to identify the semantic categories (such as people, animals, objects, etc.) and attributes (such as color, texture, status, etc.) to which the region belongs. A multi-label classifier is used to assign semantic labels and attribute descriptions to each region, forming a set of semantically marked regions with semantic annotations. Then, nodes with graph structures are created for each region in the set of semantically marked regions. The node features include the visual features, semantic labels, and attribute information of the region. The spatial relationships between regions, such as relative positions (up, down, left, right), inclusion relationships, overlap degrees, etc., are calculated, and the semantic similarity between regions is calculated based on the feature cosine distance or the learned semantic distance metric. Based on the spatial relationships and semantic similarities, a weighted adjacency matrix is constructed. The matrix elements represent the strength and type of the connections between regions, forming an initial scene relationship graph structure and an initial scene graph. Next, the initial scene graph is input into a multi-layer graph convolutional network to execute the message passing mechanism, enabling each node to aggregate the information of its neighbor nodes. Through multi-hop message passing, a wider range of context relationships are captured, realizing the context-aware enhancement of node features and obtaining a context-aware node representation that integrates the surrounding environment information. Based on the context-aware node representation, the semantic relationships between nodes, such as specific relationship types like "is located in", "contains", "supports", "holds", etc., are predicted. A relationship classifier is used to predict the relationship type for each pair of connected nodes, generating a relationship marked graph with rich semantic relationship annotations. Then, the global scene features are extracted to represent the context and atmosphere of the overall scene. The relationship marked graph and the global scene features are interactively integrated using the global-local attention mechanism, enabling node features to focus on the global context while allowing the global features to focus on local important entities. Through the two-way information flow, the long-distance dependencies across regions are captured, forming a globally associated scene graph considering the global context. An attention-guided information screening mechanism is applied to the globally associated scene graph, automatically focusing on key nodes and relationships according to the task objectives and content importance, suppressing irrelevant secondary connections, and highlighting the core relationship structure crucial for scene understanding, thus realizing the effective reduction of information.Finally, a hierarchical scene understanding model is constructed based on the global correlation scene graph, organizing scene elements at different abstraction levels, from underlying pixels and regions to middle-level objects and entities, and then to high-level event and situation understanding, supporting multi-level scene parsing from fine-grained to coarse-grained, allowing the system to understand and operate on scene content at different abstraction levels according to task requirements, integrating information at each level, constructing a unified representation that includes both details and high-level semantics, and finally forming a scene element relationship network, which comprehensively captures the spatial layout, semantic associations, and interaction methods among entities in the scene, laying a structured foundation for subsequent visual importance analysis and special effect applications.

[0020] In the embodiment of the present invention, a parameter configuration module is used to select an appropriate special effect template according to the scene visual importance distribution map and the multi-modal semantic feature map, and generate a special effect parameter configuration scheme, specifically for: Construct a special effect template knowledge base, including parameterized templates for multiple categories of visual special effects, and establish semantic labels and applicable scene descriptions for each template to obtain a semantic special effect library; Perform regional clustering on the scene visual importance distribution map to identify key special effect application regions, obtain a set of special effect target regions, and extract the scene emotional tone and style features based on the multi-modal semantic feature map to obtain a scene style representation; Combine the features of the special effect target region set with the scene style representation to construct a special effect requirement vector, and perform similarity matching between the special effect requirement vector and the templates in the semantic special effect library to obtain a candidate special effect template set; Apply a multi-objective evaluation mechanism to the candidate special effect template set, evaluate it from multiple dimensions such as visual coordination, emotional matching degree, and computational complexity, obtain a template score list, and select the optimal special effect template combination based on the template score list to obtain a special effect application scheme; According to the special effect application scheme, calculate the initial parameter values for each selected special effect template, and adjust the parameter intensity and application range in combination with the scene visual importance distribution map to finally form a special effect parameter configuration scheme.

[0021] In this embodiment, first, a special effect template knowledge base is constructed. A variety of visual special effect templates are collected and designed, including light effects (such as halos, light columns, flashes), particle effects (such as flames, smoke, water), atmosphere effects (such as fog, light scattering), deformation effects (such as distortion, ripples), etc. Parametric descriptions are implemented for each special effect template, including shape parameters, dynamic parameters, physical parameters, and visual style parameters, etc., so that the special effects can achieve customized performance through parameter adjustment. Semantic tags are added to each special effect template to describe its visual characteristics, emotional expression, and applicable scenarios, such as "warm halo - creating a warm atmosphere - applicable to tender scenes", establishing a mapping relationship between the special effects and the narrative intention, and forming a structured semantic special effect library. Then, a spatial clustering algorithm, such as DBSCAN, Mean Shift, or K-means, etc., is applied to the scene visual importance distribution map. Based on the distribution pattern of the importance values, regional clusters with similar importance are recognized. Considering spatial continuity and importance thresholds, the regions that should be preferentially applied with special effects are selected, generating a set of special effect target regions containing spatial positions and importance weights. Based on the multi-modal semantic feature map, the emotional tone of the scene, such as tense, cheerful, sad, mysterious, etc., is analyzed, and the visual style characteristics of the scene, such as color tone tendency, visual rhythm, composition features, etc., are extracted. The emotional tone and visual style information are integrated into a vector representation characterizing the artistic style of the scene, forming a scene style representation. Next, the visual features, position information, and importance weights of the special effect target region set are feature-fused with the scene style representation to construct a multi-dimensional vector expressing the special effect requirements. This vector encodes special effect application requirements such as "where", "what style", and "how intense", calculating the feature similarity between the special effect requirement vector and each template in the semantic special effect library, using cosine similarity, semantic matching degree, or learned similarity metrics, and screening out the special effect templates with high matching degrees to the requirement vector, forming a candidate special effect template set. Then, multi-dimensional evaluations are performed on the special effects in the candidate special effect template set. From the dimension of visual coordination, the style compatibility and harmony degree between the special effect and the original picture are evaluated. From the dimension of emotional matching degree, the enhancement or supplementary effect of the special effect on the scene emotional expression is evaluated. From the dimension of computational complexity, the resource consumption and rendering time cost of the special effect implementation are evaluated. Quantitative scores are given to each dimension and the comprehensive score is calculated, generating a template score list containing the scores of each index. Based on the template score list, the optimal combination of special effect templates is selected, considering the complementarity and synergy effects between the templates to optimize the overall visual experience, which may include multiple special effects of different types or the same type with different parameters, forming a special effect application plan.Finally, according to the special effect application plan, determine the initial parameter values for each selected special effect template. Set reasonable initial parameter values based on the special effect type and application scenario, such as light effect intensity, particle quantity, animation speed, etc. Adjust the parameters in combination with the scene visual importance distribution map. Enhance the details and intensity of the special effects in high-importance areas, and reduce the presence of special effects in low-importance areas to avoid distraction. Consider the parameter coordination and superposition effects between special effects to prevent visual overload or conflict. Finally, form a special effect parameter configuration plan, which includes the detailed parameter settings and application area definitions for each special effect template, providing precise guidance for subsequent parametric rendering.

[0022] In the embodiment of the present invention, please refer to Figure 2 , which is a logical schematic diagram for identifying key special effect application areas. Specifically, perform regional clustering on the scene visual importance distribution map to identify key special effect application areas, and obtain a set of special effect target areas, including: Apply an adaptive threshold segmentation algorithm to the scene visual importance distribution map, automatically determine the optimal segmentation threshold according to the global importance distribution characteristics, obtain a preliminary regional division map, and perform noise removal and boundary smoothing on the preliminary regional division map through morphological operations to obtain a fine regional mask; Construct a multi-level image segmentation tree based on the fine regional mask, and combine the visual Gestalt law and semantic consistency principle to achieve a hierarchical regional representation from pixels to object level, obtain a multi-scale regional candidate pool, and apply a regional merging strategy based on visual perception correlation to the multi-scale regional candidate pool to fuse regions that visually belong to the same perceptual unit to obtain a set of perceptually consistent regions; Design a dynamic density peak clustering algorithm to adaptively cluster the regions in the set of perceptually consistent regions, automatically identify the importance density centers, and consider the semantic association degree between regions to achieve collaborative clustering of semantically related regions to obtain a semantically enhanced clustering result; Introduce a temporal stability evaluation mechanism to analyze the consistency of the semantically enhanced clustering result across frames, identify stable important regions with temporal continuity, and filter out unstable regions that appear briefly to obtain a set of temporally stable important regions. At the same time, calculate the time-varying importance weights for each region to construct a dynamic importance curve; Combine the set of temporally stable important regions and the dynamic importance curve, design a multi-objective regional value evaluation system, comprehensively consider factors such as region size, importance intensity, perceptual complexity, narrative importance, and special effect adaptability, perform regional priority ranking to obtain a hierarchical set of special effect target regions, and optimize the region boundaries through a regional automatic diffusion algorithm to ensure the visual coherence of special effect application. Finally, generate a set of special effect target regions that includes precise boundary definitions and importance weights.

[0023] In this embodiment, first, an adaptive threshold segmentation algorithm, such as the Otsu method, adaptive threshold, or multi-level threshold segmentation, is applied to the scene visual importance distribution map. The distribution histogram of the global importance values is analyzed to automatically calculate the optimal threshold that can maximize the between-class variance. The importance map is divided into high-importance regions and low-importance regions to generate a preliminary region division map. Morphological operations, including opening and closing operations, median filtering, and region connection, are applied to the preliminary region division map to remove isolated noisy pixels and small irrelevant regions, smooth the region boundaries, fill small internal holes, and improve the coherence and integrity of the regions, resulting in a refined region mask. Then, a multi-level image segmentation tree is constructed based on the refined region mask. A hierarchical segmentation algorithm (such as hierarchical superpixels or hierarchical graph segmentation) is used to start from the pixel level and gradually merge similar regions to form a multi-level segmentation result from fine-grained to coarse-grained. The visual Gestalt laws (such as proximity, continuity, closure, etc.) are combined to guide the region merging process to ensure that the segmentation result conforms to the laws of human visual perception. At the same time, the semantic consistency principle is followed, making parts of the same semantic entity tend to be grouped together, generating a multi-scale region candidate pool containing segmentation results at different scales. A region merging strategy based on visual perception correlation is applied to the regions in the multi-scale region candidate pool to analyze the visual relevance between regions, such as texture similarity, color continuity, and edge coherence, identify regions that belong to the same entity or scene unit in visual perception, and merge these regions to form a more complete unit that conforms to human perception habits, obtaining a set of perceptual consistency regions that maintain consistency in visual perception. Next, a dynamic density peak clustering algorithm is designed. This algorithm can adapt to the importance distribution characteristics of different scenes, calculate the local density and relative distance of regions in the importance space, automatically identify the importance density centers, that is, regions with high importance values and relatively high importance in the surrounding regions, and consider the semantic correlation degree between regions, making semantically related regions (such as regions belonging to the same object or action sequence) tend to be assigned to the same cluster, even if they are not completely adjacent in space, to achieve collaborative clustering with semantic perception ability, obtaining a semantic-enhanced clustering result considering the semantic relationship of regions. Then, a temporal stability evaluation mechanism is introduced to analyze the temporal consistency of the semantic-enhanced clustering results in consecutive multiple frames of images, track the change trajectory of regions in the time dimension, identify important regions that persist in multiple frames and are relatively stable in position, such as the protagonist character, key objects, etc., filter out temporary or random regions that only appear in a few frames, ensure the temporal coherence of special effect applications, form a set of temporally stable important regions, and at the same time calculate a time-varying importance weight function for each region to reflect the changing trend of region importance over time, such as importance that gradually increases or decreases with the development of the narrative, and construct a dynamic importance curve representing the temporal change of importance.Finally, by combining the temporally stable important region set with the dynamic importance curve, a region value evaluation system that synthesizes multiple factors is designed. The evaluation factors include: region size (area and proportion), importance intensity (average value and peak value), perceived complexity (texture complexity and edge density), narrative importance (degree of association with the storyline), and special effect adaptability (compatibility of the region with special effects). Multidimensional value scores are calculated for each region and weighted and aggregated. The regions are prioritized according to the comprehensive scores, forming a hierarchical special effect target region set with clear levels. The region boundaries are optimized through a region automatic diffusion algorithm to make the edges of special effect applications have natural transitions and attenuations, avoiding a rigid boundary feeling. At the same time, the region range is adaptively adjusted according to the special effect type. For example, light effect special effects may require a larger influence range, and particle effects may require more precise boundary control. Finally, a special effect target region set containing precise boundary definitions (contour point sets or parametric curves) and importance weights (numerical values or functions) is generated. This set precisely specifies the positions, ranges, and intensities where special effects should be applied, providing clear guidance for parameter configuration and rendering.

[0024] In an embodiment of the present invention, a parametric rendering module is configured to apply a special effect parameter configuration scheme to a multi-modal feature data set and generate an initial special effect rendering image through an improved parametric rendering algorithm. Specifically, it is used for: Parse the special effect parameter configuration scheme into a specific rendering instruction sequence, construct a rendering computation graph, and pre-compute the special effect evolution trajectory according to the spatio-temporal characteristics of the multi-modal feature data set to obtain a dynamic effect prediction map; Establish a depth estimation model for the image sequence in the multi-modal feature data set to generate a scene depth map, and use the scene depth map to construct a three-dimensional scene structure representation to obtain a stereo scene model; Combine the special effect parameter configuration scheme with the stereo scene model to perform physical simulation calculations, simulate the interaction between the special effects and the scene elements, and obtain a physically constrained special effect behavior model; Based on the physically constrained special effect behavior model, use an improved neural radiance field technology to perform view-consistent special effect rendering to obtain a preliminary rendering result, and apply a temporal smoothing algorithm to eliminate inter-frame jitter and improve temporal coherence to obtain a smooth rendering sequence; Adaptively blend the smooth rendering sequence with the original image sequence, control the blending intensity according to the scene visual importance distribution map, and generate a realistic initial special effect rendering image.

[0025] In this embodiment, first, the special effect parameter configuration scheme is converted into a specific instruction sequence executable by the rendering engine, including special effect type selection, parameter setting, application area, and time control, etc. The rendering operation process is organized using a directed acyclic graph (DAG) structure to optimize the computational dependencies and resource utilization, constructing an efficient rendering computation graph. Analyze the temporal and spatial characteristics in the multi-modal feature dataset, such as motion trajectories, scene changes, and audio rhythms, etc. Combine the physical characteristics of the special effects to pre-compute the evolution trajectory of the special effects over time, including position changes, morphological changes, and intensity changes, etc., generating a prediction graph reflecting the dynamic behavior of the special effects to provide a reference for subsequent rendering. Then, apply a monocular depth estimation network (such as MiDaS, DPT, or AdaBins, etc.) to the image sequence in the multi-modal feature dataset to infer the depth information of the scene from a single-view image, generating a relative distance representation from each pixel to the camera to form a scene depth map. Use the depth map combined with the camera parameters to reverse-infer the point cloud data or mesh model in three-dimensional space, restoring the geometric structure of the scene, constructing a three-dimensional scene representation containing depth information, and obtaining a three-dimensional scene model that can be used for three-dimensional space rendering calculations. Next, combine the special effect attributes in the special effect parameter configuration scheme with the spatial structure in the three-dimensional scene model, set up a physical simulation environment, including environmental factors such as gravitational fields, wind fields, and lighting, etc., and perform physical-based special effect behavior simulations, such as hydrodynamic calculations (for water and smoke effects), particle system simulations (for spark and rain / snow effects), light propagation simulations (for light effects and halo effects), etc., simulating the interactions between special effects and scene elements, such as particle collisions, light reflections, and shadow projections, etc., generating a special effect behavior model that conforms to physical laws to ensure the natural and realistic performance of the special effects. Then, based on the physical constraint-based special effect behavior model, use an improved neural radiance field (NeRF) technology for special effect rendering. The neural radiance field can represent the volume density and directional radiation in three-dimensional space, which is particularly suitable for rendering special effects with translucent characteristics, such as smoke, fog, light, etc. Improve the traditional NeRF by adding a time dimension and physical constraint conditions to achieve a consistent rendering effect of the special effects as the viewing angle changes, generating a preliminary rendering result. Apply a temporal smoothing algorithm, such as a temporal convolutional network or an optical flow-guided temporal filter, to adjacent frames in the rendering sequence to analyze and correct the abrupt changes and jitters between frames, ensuring the smooth transition and coherent performance of the special effects in the time dimension, especially maintaining the stability of the special effects in complex dynamic scenes, and obtaining a temporally coherent smooth rendering sequence.Finally, adaptively blend the smoothed rendering sequence with the original image sequence, design an intelligent blending algorithm, automatically select the most suitable blending mode according to the special effect type and scene characteristics, such as Alpha blending, additive blending, multiplicative blending, etc., use the scene visual importance distribution map to control the blending process, maintain the integrity of the original content in important areas, while allowing the special effects to enhance the visual performance, and allow the special effects to have a stronger presence in secondary areas. Apply adaptive color correction to ensure that the colors and lighting after blending remain natural and coordinated, and generate an initial special effect rendering image that not only retains the integrity of the original content but also has enhanced visual expressiveness, providing a basis for subsequent parameter optimization.

[0026] In the embodiment of the present invention, please refer to Figure 3 , which is a logical schematic diagram for perspective-consistent special effect rendering. Specifically, based on the physical constraint special effect behavior model, an improved neural radiance field technology is used for perspective-consistent special effect rendering, including: Construct a parameterized neural radiance field network, take the spatial coordinates and viewing direction as inputs, predict the voxel density and directional radiation characteristics, obtain the basic radiance field model, and embed the parameters in the special effect parameter configuration scheme into the basic radiance field model to achieve controllable adjustment of special effect attributes, and obtain a parameterized special effect radiance field; Design a special density function and emission function for special effects based on the parameterized special effect radiance field, simulate complex special effect physical characteristics including light scattering, refraction, and luminescence, obtain the special effect physical rendering function, and combine the special effect physical rendering function with the traditional volume rendering equation to construct a hybrid rendering pipeline capable of handling special effect physical characteristics; Perform adaptive ray sampling for each viewpoint in the hybrid rendering pipeline, increase the sampling density in the visually important areas according to the scene visual importance distribution map to improve the detail performance, obtain a differential sampling distribution, and apply a gradient-guided sampling strategy based on the differential sampling distribution to reduce computational redundancy and improve the rendering efficiency, forming an adaptive sampling ray set; Input the adaptive sampling ray set into the hybrid rendering pipeline to achieve continuity constraints based on the viewing angle, ensure the consistent performance of special effects under different viewing angles, avoid jumping phenomena when the viewpoint changes, obtain an intermediate rendering result with continuous viewing angles, and adopt a multi-resolution rendering strategy for the intermediate rendering result with continuous viewing angles, dynamically allocate computing resources according to the regional importance, and generate special effect rendering data with multiple levels of fineness; Input the special effect rendering data with multiple levels of fineness into the optimized hybrid rendering pipeline through hardware acceleration technology, make full use of the parallel computing ability of the GPU to achieve efficient neural radiance field special effect rendering calculation, and at the same time apply post-processing optimization algorithms to enhance edge details and depth consistency, and finally obtain a special effect rendering result with consistent viewing angles and reasonable physics.

[0027] In this embodiment, a parametric neural radiance field network is first designed and trained. The network architecture is based on a multi-layer perceptron (MLP), with three-dimensional spatial coordinates (x, y, z) and viewing directions (θ, ) takes [the input] and outputs the voxel density σ and the directional radiation color c at the corresponding positions. The network includes a position encoding layer, a feature extraction layer, and an output prediction layer. It improves the generalization ability through regularization and feature decoupling to form a basic radiation field model. A parameter embedding module is designed to embed the parameters (such as transparency, scattering coefficient, luminous intensity, etc.) in the special effect parameter configuration scheme into the neural network in the form of conditional encoding. Conditional feature modulation (FiLM) or hypernetwork technology is used to achieve dynamic control of the network behavior, enabling the special effect attributes to be directly adjusted by parameters without retraining the network. A parameterized special effect radiation field that can dynamically adjust the special effect performance according to the input parameters is constructed. Then, based on the parameterized special effect radiation field, special density functions suitable for different special effect types are designed, such as the volume density model for smoke, the temperature gradient density model for flame, the energy attenuation density model for light effects, etc. Special emission functions for special effects are designed to simulate the radiation characteristics of special effects, such as the flame emission model based on blackbody radiation, the atmospheric light scattering model based on Rayleigh scattering, the water surface reflection model based on the Fresnel effect, etc. Combining the principles of physical optics to simulate complex light interaction effects, including scattering (the multi-directional dispersion of light in a medium), refraction (the change in the direction of light when passing through different media), and emission (the light generated by the material itself), etc. The special effect physical rendering function is integrated. The special effect physical rendering function is combined with the traditional volume rendering integral equation to extend the standard volume rendering equation to support more complex light propagation behaviors. A hybrid rendering pipeline is designed to support the seamless integration of multiple rendering technologies, such as ray tracing, path tracing, volume rendering, etc., which can handle the complex physical characteristics in special effects and maintain computational efficiency. Then, an adaptive ray sampling strategy is implemented in the hybrid rendering pipeline. According to the analysis of the scene visual importance distribution map, the sampling density requirements of different regions are determined. A higher sampling density is assigned to the visually important regions (such as the main object, the key special effect regions), increasing the number of rays to achieve detail enhancement. A lower sampling density is used for the visually less important regions to form a differential sampling distribution for the scene content. Based on the differential sampling distribution, gradient-guided sampling optimization is applied. The gradient distribution of the rendering intermediate result is analyzed to identify the regions in the rendering result that are most sensitive to the sampling points. The sampling density is increased in these regions, and the sampling points in the regions with less impact on the result are reduced. Computational redundancy is reduced through importance sampling, and the utilization efficiency of computational resources is improved to form an adaptive sampling ray set that both pays attention to visual quality and takes computational efficiency into account.Then, input the adaptive sampling light set into the hybrid rendering pipeline to achieve view-based continuity constraints, design a view smoothing transition mechanism to ensure smooth transitions of special effects during viewpoint changes, avoid jumping phenomena caused by discrete view sampling, achieve view continuity of special effect performance through feature interpolation and constraint optimization, generate rendering intermediate results that are consistent when observed from any viewpoint, adopt a multi-resolution rendering strategy for the rendering intermediate results, divide the image into regions of different importance, use high-resolution rendering for important regions and low-resolution rendering for secondary regions, and then perform upsampling synthesis through super-resolution or interpolation techniques, dynamically allocate computing resources, improve the overall rendering efficiency while ensuring the quality of visually critical regions, and generate special effect rendering data containing different levels of fineness. Finally, optimize the hybrid rendering pipeline to support GPU acceleration, utilize CUDA, OpenCL or specific hardware acceleration libraries to implement highly parallel rendering calculations, input the rendering data of multiple levels of fineness into the optimized rendering pipeline for final processing, make full use of the parallel computing power of the GPU to achieve efficient neural radiance field special effect calculations, apply a series of post-processing optimization algorithms, including edge enhancement (to improve the clarity of special effect boundaries), depth consistency correction (to ensure the rationality of the depth relationship between special effects and the scene), detail sharpening (to enhance textures and microstructures), etc., and finally generate special effect rendering results that have both view consistency (maintaining a coherent visual performance when observed from any angle) and conform to physical laws (following optical and physical principles), providing a high-quality visual basis for subsequent parameter optimization and final rendering.

[0028] In the embodiment of the present invention, the parameter optimization module is used to perform visual quality evaluation on the initial special effect rendering image, optimize the special effect parameter configuration scheme based on the evaluation result, and generate an optimized special effect parameter set, specifically including: Construct a multi-level visual quality evaluation framework, including pixel-level fidelity evaluation, structural similarity analysis, and perceptual quality evaluation modules, obtain a comprehensive quality score, and conduct A / B test simulations to compare the visual effects under different special effect parameters, and obtain the parameter sensitivity analysis result; Introduce a computational aesthetics evaluation model to evaluate the artistic value of the rendering effect from aspects such as composition balance, color harmony, and visual rhythm, obtain an aesthetics score, and combine it with the scene visual importance distribution map to examine the pertinence of special effect enhancement, and obtain the enhancement accuracy index; Based on the comprehensive quality score, aesthetics score, and enhancement accuracy index, construct an objective function for special effect parameter optimization, and use the Bayesian optimization algorithm to explore the parameter space to obtain the parameter optimization direction; Generate a candidate set of improved parameters in the parameter space, conduct fast rendering evaluation, and apply reinforcement learning methods to automatically select the optimal parameter adjustment strategy to achieve fine-tuning of special effect parameters and obtain an improved special effect parameter scheme; Verify the stability and rendering consistency of the improved special effect parameter scheme, eliminate possible visual artifacts and temporal incoherence issues, and finally generate an optimized special effect parameter set.

[0029] In this embodiment, first, a multi-level visual quality assessment framework is constructed. Fidelity assessment is carried out at the pixel level by calculating the pixel differences between the special effect rendered image and the original image, and metrics such as mean squared error (MSE) and peak signal-to-noise ratio (PSNR) are used to measure the degree to which the special effects retain the original content. Similarity analysis is carried out at the structural level, and the structural similarity index (SSIM) or feature similarity index (FSIM) is used to evaluate whether the special effect rendering retains the key structural information of the original image. Quality assessment is carried out at the perceptual level using a deep learning-based perceptual quality assessment model (such as LPIPS or PieAPP) to simulate the perceptual evaluation of image quality by the human visual system. The evaluation results of the three levels are integrated, and a weighted average score is calculated to form a comprehensive quality score. At the same time, A / B test simulation is carried out to generate samples of rendering results under different special effect parameter configurations, and a computer vision model is used to simulate human preference judgment to compare the influence degree of different parameter settings on the visual effect, analyze the sensitivity of parameter changes to the rendering quality, identify the key parameters that have the greatest impact on the final effect, and form the parameter sensitivity analysis results. Then, a computational aesthetics evaluation model is introduced. This model is constructed based on art theory and aesthetic principles to evaluate the composition balance, analyze whether the spatial distribution of special effect elements conforms to the visual balance principle, such as the golden ratio and the rule of thirds, evaluate the color harmony, analyze whether the color scheme after the introduction of special effects is harmonious and unified and whether it conforms to the harmonious relationship in color theory, evaluate the visual rhythm, analyze whether the dynamic changes of special effects create a rhythmic visual experience and whether it matches the content narrative rhythm, calculate the computational aesthetics score by integrating aesthetic indicators to reflect the artistic expression value of special effects, combine with the scene visual importance distribution map to test the pertinence of special effect enhancement, calculate the matching degree between the special effect intensity distribution and the visual importance distribution, evaluate whether the special effects accurately enhance the key areas that should be highlighted while avoiding overemphasis on secondary areas, and form an enhancement accuracy index that quantifies the precision of special effect application. Next, based on the comprehensive quality score, aesthetic score, and enhancement accuracy index, a multi-objective function for optimizing special effect parameters is constructed, and the weight coefficients of each index are set to reflect the priorities of different quality dimensions, forming a comprehensive optimization goal. The Bayesian optimization algorithm is used to explore the high-dimensional parameter space. Bayesian optimization efficiently explores possible optimal parameter combinations by constructing a Gaussian process model of the relationship between parameters and scores, continuously updates the prior probability distribution, and quickly converges to the optimal solution region. Analyze the parameter sensitivity results to identify the optimization direction, that is, whether the parameters should be increased or decreased to improve the overall score, and form the optimization direction for guiding parameter adjustment.Then, according to the parameter optimization direction, multiple sets of improved parameter candidate sets are generated in the parameter space. The parameter variation range is based on parameter sensitivity. For sensitive parameters, fine-grained changes are made with small step sizes. For each set of candidate parameters, rapid rendering evaluation is performed. A simplified rendering pipeline or proxy model is used for quick preview to evaluate the effectiveness of parameter adjustment. The deep reinforcement learning method is applied to construct a parameter adjustment policy network. The rendering effect score is used as the reward signal. Through multiple rounds of iterative learning, the optimal parameter adjustment policy is obtained, and refined parameter tuning is automatically executed, such as fine-tuning the scattering coefficient, adjusting the light intensity attenuation curve, etc., to generate an improved special effect parameter scheme. Finally, the stability of the improved special effect parameter scheme is verified. The performance stability of the parameters under different scene conditions is tested, such as the adaptability under different lighting and different action speeds. Rendering consistency checks are carried out. The temporal coherence of the special effect performance is tested on a continuous frame sequence to ensure there are no flickering or sudden changes. Potential visual artifacts are analyzed and eliminated, such as moiré patterns, edge jagging, unnatural halos, etc. The temporal coherence is optimized to ensure that the special effects smoothly transition with scene changes without abrupt jumps. After comprehensive verification, an optimized special effect parameter set is finally generated. This parameter set has high computational efficiency and cross-scene adaptability while ensuring visual quality, artistic expression, and application accuracy, providing an optimized parameter configuration for the final rendering.

[0030] In an embodiment of the present invention, the rendering and output module is used to perform final rendering processing on the multi-modal feature dataset using the optimized special effect parameter set to generate an AI visual special effect image sequence integrating multi-modal perception, specifically used for: Apply the optimized special effect parameter set to a high-precision rendering engine to perform full-resolution processing on the multi-modal feature dataset to obtain an original special effect rendering sequence, and perform high dynamic range processing on the original special effect rendering sequence to enhance the details of the bright and dark parts to obtain an HDR enhanced sequence; According to the rhythm and emotional characteristics of the audio data, adjust the temporal characteristics of the dynamic changes of the special effects to achieve the special effect rhythm synchronized with the audio and video to obtain a rhythm-synchronized special effect sequence, and according to the narrative information in the scene parameters, adjust the undulating changes of the special effect intensity to strengthen the narrative expression to obtain a narrative-enhanced special effect sequence; Apply post-processing color grading technology to unify the color styles of the special effects and the original picture to ensure the consistency of the visual style to obtain a color-unified special effect sequence, and perform final visual optimization processing, including sharpening, noise reduction, and detail enhancement, to obtain a refined special effect sequence; Perform final quality inspection on the refined special effect sequence to ensure that the rendering quality, emotional expression, and narrative enhancement effects meet the expected goals to obtain a verified special effect sequence, and encode and package the verified special effect sequence according to the target output format to obtain a media file that meets the playback specifications; Package media files that comply with playback specifications with interactive control metadata, support the ability of end-users to make limited adjustments to the intensity of special effects, and finally generate an AI visual effect image sequence that integrates multi-modal perception.

[0031] In this embodiment, first, the optimized special effect parameter set is imported into a high-precision rendering engine, such as a physically based renderer (PBR) or a real-time ray tracing engine, and high-quality rendering settings are configured, including high sampling rate, fine ray tracing, and high-precision physical simulation, etc. The multi-modal feature dataset is processed at full resolution. Instead of using a proxy model or low-resolution preview, complete and accurate special effect calculations are performed to generate the original special effect rendering sequence. The high dynamic range (HDR) processing technology is applied to the original special effect rendering sequence to expand the dynamic range of the image. The local tone mapping algorithm is used to enhance the detail level and brightness change of the highlight area, improve the visibility and texture performance of the dark area, and ensure the visibility of special effect details in high-contrast scenes, obtaining an HDR-enhanced sequence with a wider dynamic range and richer details. Then, analyze the rhythm characteristics of the audio data, extract the beats, rhythms, and energy change points in the audio, analyze the emotional characteristics of the audio, identify the emotional climax, tense or soothing passages, and associate the audio characteristics with the dynamic parameters of the special effects, so that the attributes such as the intensity, speed, density, or color of the special effects change synchronously with the audio rhythm. For example, enhance the particle burst at the music beat point and enhance the light effect intensity at the emotional climax to create an immersive audio-visual collaborative experience and generate a rhythm-synchronized special effect sequence that is perfectly synchronized with the audio. Further analyze the narrative information in the scene parameters, understand the emotional intention and narrative function of the current scene in the story, adjust the temporal change curve of the special effect intensity according to the narrative requirements, and design the undulating change of the special effect intensity to strengthen the narrative expression. For example, enhance the special effects at the plot turning point, maintain subtle special effects in the foreshadowing stage, and reach the strongest special effects at the climax part, making the special effects an organic part of the narrative means and generating a narrative-enhanced special effect sequence that can strengthen the conveyance of the story's emotion and meaning. Then, apply professional post-production color grading technology, analyze the color style and tone tendency of the original image, adjust the color parameters of the special effect elements to make them consistent with the original image style, ensure that the special effects do not damage the established color emotion and visual style of the image, while maintaining the unique visual characteristics of the special effects, balance the presence of the special effects and the harmony of the overall image, and generate a special effect sequence with a unified color style. Perform the final visual optimization process, apply an adaptive sharpening algorithm to enhance the clarity of the special effect edges and textures without introducing artifacts of oversharpening, use an advanced noise reduction algorithm to eliminate the noise and graininess that may be generated during the rendering process, especially in the transparency edges and light gradient areas, apply detail enhancement technology to improve the microscopic structure and texture performance of the special effects, enhance the realism and three-dimensional sense. Combining these optimization processes, a refined special effect sequence with extremely high visual quality is generated.Then, conduct a final quality inspection on the refined special effects sequence, evaluate technical indicators using objective quality measurement tools such as resolution, bit depth, color accuracy, etc., verify whether the special effects accurately express the expected emotions and atmosphere through visual comparison, evaluate the enhancement effect of the special effects on the narrative, confirm whether the story expression and emotional transmission are strengthened, and ensure that all quality indicators meet or exceed the expected goals based on the comprehensive evaluation results to obtain a verified special effects sequence. Encode and package the verified special effects sequence according to the target output format, select appropriate codecs (such as H.264, H.265, AV1, etc.) and container formats (such as MP4, MOV, MKV, etc.), set appropriate bitrates and encoding parameters to ensure a reasonable file size while maintaining visual quality, and generate a high-quality media file that complies with industry playback specifications. Finally, package the encoded media file with interactive control metadata, design a lightweight special effects parameter control interface that allows end-users to adjust certain attributes of the special effects within a preset range, such as overall intensity, color tendency, or animation speed, develop an interactive playback control component that enables the terminal application to read and apply the user's special effects adjustment instructions, provide a moderate personalized experience while maintaining artistic integrity, and ultimately generate an AI visual special effects image sequence that integrates multi-modal perception with both professional-level visual performance and a certain degree of interactive flexibility, providing an immersive visual experience for users.

[0032] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0033] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0034] In the description of the present invention, it should be understood that the terms "first", "second", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.

[0035] In the description of the present invention, unless otherwise specified, the meaning of "a plurality of" is two or more.

[0036] In the description of the present invention, "several" means one or more, and "a large number" means two or more.

[0037] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0038] The preset parameters and threshold selections for this specification are set by those skilled in the art according to the actual situation.

[0039] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

Claims

1. An AI visual effect dynamic generation system integrating multi-modal perception, characterized in that, including: A multimodal input processing module, configured to obtain multimodal input data including an image sequence, audio data, and scene parameters, perform spatio-temporal alignment and normalization processing on the multimodal input data, and generate a multimodal feature dataset; An importance analysis module, configured to extract features from the multimodal feature dataset, construct a multimodal semantic feature map, identify key visual regions of the multimodal semantic feature map, and then generate a scene visual importance distribution map; A parameter configuration module, configured to select an appropriate special effect template according to the scene visual importance distribution map and the multimodal semantic feature map, and generate a special effect parameter configuration scheme; A parameterized rendering module, configured to apply the special effect parameter configuration scheme to the multimodal feature dataset, and generate an initial special effect rendered image through an improved parameterized rendering algorithm; A parameter optimization module, configured to perform a visual quality assessment on the initial special effect rendered image, optimize the special effect parameter configuration scheme based on the assessment result, and generate an optimized special effect parameter set; A rendering and output module, configured to perform a final rendering process on the multimodal feature dataset using the optimized special effect parameter set, and generate an AI visual special effect image sequence integrating multimodal perception.

2. The AI visual effect dynamic generation system integrating multi-modal perception according to claim 1, characterized in that The obtaining of the multimodal input data including an image sequence, audio data, and scene parameters, and performing spatio-temporal alignment and normalization processing on the multimodal input data to generate a multimodal feature dataset includes: Performing key frame extraction and resolution unification processing on the image sequence to obtain a standardized image frame sequence, and performing color space conversion on the standardized image frame sequence to obtain a color-corrected image sequence; Performing sampling rate resampling and spectrum analysis on the audio data to obtain a time-frequency characteristic spectrum, and performing emotional feature extraction on the time-frequency characteristic spectrum to obtain an audio emotional feature vector; Performing semantic parsing and hierarchical organization on the scene parameters to obtain a structured parameter representation, and constructing a scene semantic network based on the structured parameter representation to obtain a scene semantic vector; Performing timestamp alignment on the color-corrected image sequence, the audio emotional feature vector, and the scene semantic vector to obtain synchronized multimodal data, and performing feature normalization on the synchronized multimodal data to obtain a standardized feature set; Performing fusion coding on the standardized feature set to construct a unified representation space, and obtaining a multimodal feature dataset.

3. The AI visual effect dynamic generation system integrating multi-modal perception according to claim 2, characterized in that, The performing of key frame extraction and resolution unification processing on the image sequence to obtain a standardized image frame sequence, and performing color space conversion on the standardized image frame sequence to obtain a color-corrected image sequence includes: Applying a visual information entropy measurement model to the image sequence, calculating the information richness of each frame, combining the inter-frame visual difference metric, constructing a dual evaluation mechanism, identifying visual information peak frames to obtain a key frame candidate set, and applying a semantic importance filtering algorithm to the key frame candidate set to obtain a refined key frame set; Analyze the content composition and main body distribution of each frame in the refined key frame set, construct an adaptive content-aware grid, obtain a content protection mapping graph, and perform a non-uniform scaling operation based on the content protection mapping graph to obtain a unified resolution frame with enhanced main body; Apply a super-resolution neural network to the unified resolution frame with enhanced main body to reconstruct textures and edge details, obtain a detail-restored image frame, and perform temporal domain consistency optimization on the detail-restored image frame to obtain a temporally coherent and standardized image frame sequence; Establish a scene lighting condition analysis model, infer the ambient light characteristics from the standardized image frame sequence, and select the color space most suitable for the target special effect to obtain a gamut-optimized image; Perform semantic segmentation color mapping on the gamut-optimized image and introduce an audience visual adaptability compensation mechanism to finally obtain a perceptually enhanced color correction image sequence.

4. The AI visual effect dynamic generation system integrating multi-modal perception according to claim 3, characterized in that, The feature extraction of the multi-modal feature dataset, the construction of a multi-modal semantic feature map, the identification of the key visual regions of the multi-modal semantic feature map, and then the generation of a scene visual importance distribution map, including: Input the multi-modal feature dataset into a pre-trained multi-modal Transformer network to extract deep semantic features, obtain a multi-level feature tensor, and perform channel attention mechanism processing on the multi-level feature tensor to obtain an enhanced feature map; Use a graph convolutional network to model the spatial relationship of the enhanced feature map, construct a scene graph structure representation, obtain a scene element relationship network, and perform semantic segmentation based on the scene element relationship network to obtain a semantic region division map; Perform cross-modal fusion of the semantic region division map and the audio emotion feature vector to obtain an emotion-enhanced semantic feature map, and integrate multi-scale information through an adaptive feature pyramid network to obtain a multi-modal semantic feature map; Apply a visual saliency detection algorithm to the multi-modal semantic feature map to calculate the attention weight of each region, obtain an initial saliency map, and calibrate the initial saliency map in combination with the human visual perception model to obtain a perceptually corrected saliency map; Perform weighted fusion of the perceptually corrected saliency map and the semantic region division map to construct a saliency heat map considering semantic importance to obtain a scene visual importance distribution map.

5. The AI visual effect dynamic generation system integrating multi-modal perception according to claim 4, wherein The use of a graph convolutional network to model the spatial relationship of the enhanced feature map, construct a scene graph structure representation, and obtain a scene element relationship network, including: Generate region candidates for the enhanced feature map, extract potential scene entity regions, obtain a region candidate set, and perform attribute prediction and classification on the region candidate set to obtain a semantic labeled region set; Construct feature nodes for each region in the semantic labeled region set, and at the same time calculate the spatial and semantic similarity between regions, generate a weighted adjacency matrix, and form an initial scene graph; Apply multi-layer graph convolutional operations to the initial scene graph for information transmission and update to obtain context-aware node representations, and predict the semantic relationships between nodes based on the context-aware node representations to obtain a relationship labeled graph; Interactively integrate the relationship labeled graph with the global scene features to obtain a global associated scene graph, and perform attention-guided information screening on the global associated scene graph to highlight key relationships; Construct a hierarchical scene understanding model based on the global associated scene graph to support multi-granularity scene parsing, and finally obtain a scene element relationship network.

6. The AI visual effect dynamic generation system integrating multi-modal perception according to claim 5, characterized in that, According to the scene visual importance distribution map and the multi-modal semantic feature map, select an appropriate special effect template to generate a special effect parameter configuration scheme, including: Construct a special effect template knowledge base, including parameterized templates for multi-category visual special effects, and establish semantic labels and applicable scene descriptions for each template to obtain a semantic special effect library; Perform regional clustering on the scene visual importance distribution map to obtain a special effect target area set, and extract the scene emotional tone and style features based on the multi-modal semantic feature map to obtain a scene style representation; Combine the features of the special effect target area set with the scene style representation to construct a special effect requirement vector, and perform similarity matching between the special effect requirement vector and the templates in the semantic special effect library to obtain a candidate special effect template set; Apply a multi-objective evaluation mechanism to the candidate special effect template set, evaluate it from multiple dimensions such as visual coordination, emotional matching degree, and computational complexity, obtain a template score list, and select the optimal special effect template combination based on the template score list to obtain a special effect application scheme; According to the special effect application scheme, calculate the initial parameter values for each selected special effect template, and adjust the parameter intensity and application range in combination with the scene visual importance distribution map, and finally form a special effect parameter configuration scheme.

7. The AI visual effect dynamic generation system integrating multi-modal perception according to claim 6, characterized in that, The performing regional clustering on the scene visual importance distribution map to obtain a special effect target area set includes: Apply an adaptive threshold segmentation algorithm to the scene visual importance distribution map, automatically determine the optimal segmentation threshold according to the global importance distribution characteristics to obtain a preliminary regional division map, and perform noise removal and boundary smoothing on the preliminary regional division map through morphological operations to obtain a fine regional mask; Construct a multi-level image segmentation tree based on the fine regional mask, combine the visual Gestalt law and the semantic consistency principle to obtain a multi-scale regional candidate pool, and apply a regional merging strategy based on visual perception correlation to the multi-scale regional candidate pool to obtain a set of perceptually consistent regions; Design a dynamic density peak clustering algorithm to perform adaptive clustering on the regions in the set of perceptually consistent regions to obtain a semantically enhanced clustering result; Introduce a temporal stability evaluation mechanism to analyze the consistency of the semantically enhanced clustering result across frames to obtain a temporally stable important region set, and calculate the time-varying importance weight for each region at the same time to construct a dynamic importance curve; Combine the temporally stable important region set with the dynamic importance curve, design a multi-objective region value evaluation system, comprehensively consider region size, importance intensity, perceptual complexity, narrative importance, and special effect adaptability, perform region priority ranking to obtain a hierarchical special effect target area set, and optimize the region boundary through a region automatic diffusion algorithm to finally generate a special effect target area set.

8. The AI visual effect dynamic generation system integrating multi-modal perception according to claim 1, characterized in that, Applying the special effect parameter configuration scheme to the multi-modal feature dataset and generating an initial special effect rendering image through an improved parametric rendering algorithm includes: Parsing the special effect parameter configuration scheme into a specific rendering instruction sequence, constructing a rendering computation graph, and pre-computing the special effect evolution trajectory according to the spatio-temporal characteristics of the multi-modal feature dataset to obtain a dynamic effect prediction graph; Establishing a depth estimation model for the image sequence in the multi-modal feature dataset, generating a scene depth map, and using the scene depth map to construct a three-dimensional scene structure representation to obtain a stereoscopic scene model; Combining the special effect parameter configuration scheme with the stereoscopic scene model, performing physical simulation calculations, simulating the interaction between the special effects and scene elements, and obtaining a physically constrained special effect behavior model; Based on the physically constrained special effect behavior model, using an improved neural radiance field technology to perform view-consistent special effect rendering, obtaining a preliminary rendering result, and applying a temporal smoothing algorithm to eliminate inter-frame jitter and improve temporal coherence to obtain a smooth rendering sequence; Adaptive mixing of the smooth rendering sequence and the original image sequence, controlling the mixing intensity according to the scene visual importance distribution map, and generating a realistic initial special effect rendering image; The above-mentioned using an improved neural radiance field technology to perform view-consistent special effect rendering based on the physically constrained special effect behavior model includes: Constructing a parametric neural radiance field network, taking spatial coordinates and view directions as inputs, predicting voxel density and directional radiation characteristics to obtain a basic radiation field model, and embedding the parameters in the special effect parameter configuration scheme into the basic radiation field model to obtain a parametric special effect radiation field; Designing a special effect-specific density function and emission function based on the parametric special effect radiation field to obtain a special effect physical rendering function, and combining the special effect physical rendering function with the traditional volume rendering equation to construct a hybrid rendering pipeline; Performing adaptive ray sampling for each view point in the hybrid rendering pipeline, increasing the sampling density in the visually important area according to the scene visual importance distribution map to obtain a differential sampling distribution, and applying a gradient-guided sampling strategy based on the differential sampling distribution to form an adaptive sampling ray set; Inputting the adaptive sampling ray set into the hybrid rendering pipeline to obtain a view-continuous rendering intermediate result, and adopting a multi-resolution rendering strategy for the view-continuous rendering intermediate result, dynamically allocating computing resources according to the regional importance to generate special effect rendering data; Inputting the special effect rendering data into the optimized hybrid rendering pipeline through hardware acceleration technology, and at the same time applying a post-processing optimization algorithm to enhance edge details and depth consistency to finally obtain a special effect rendering result.

9. The AI visual effect dynamic generation system integrating multi-modal perception according to claim 1, characterized in that, The above-mentioned visually evaluating the initial special effect rendering image and optimizing the special effect parameter configuration scheme based on the evaluation result to generate an optimized special effect parameter set includes: Constructing a multi-level visual quality evaluation framework, including a pixel-level fidelity evaluation, a structural similarity analysis, and a perceptual quality evaluation module, obtaining a comprehensive quality score, and performing an A / B test simulation to obtain a parameter sensitivity analysis result; An aesthetic evaluation model is introduced to evaluate the artistic value of the rendering effect in terms of composition balance, color harmony, and visual rhythm, obtaining an aesthetic score. The pertinence of special effect enhancement is verified by combining with the visual importance distribution map of the scene, obtaining an enhancement accuracy index; Based on the comprehensive quality score, the aesthetic score, and the enhancement accuracy index, an optimization objective function for special effect parameters is constructed, and the Bayesian optimization algorithm is used to explore the parameter space to obtain the parameter optimization direction; A candidate set of improved parameters is generated in the parameter space, rapid rendering evaluation is carried out, and the reinforcement learning method is applied to automatically select the optimal parameter adjustment strategy to obtain an improved special effect parameter scheme; Verify the stability and rendering consistency of the improved special effect parameter scheme, and finally generate an optimized special effect parameter set.

10. The AI visual effect dynamic generation system integrating multi-modal perception according to claim 9, characterized in that, Using the optimized special effect parameter set to perform final rendering processing on the multi-modal feature dataset to generate an AI visual special effect image sequence integrating multi-modal perception, including: Apply the optimized special effect parameter set to a high-precision rendering engine to perform full-resolution processing on the multi-modal feature dataset to obtain an original special effect rendering sequence, and perform high-dynamic range processing on the original special effect rendering sequence to obtain an HDR enhanced sequence; According to the rhythm and emotional characteristics of the audio data, adjust the temporal characteristics of the dynamic changes of the special effects to obtain a rhythm-synchronized special effect sequence, and adjust the undulating changes of the special effect intensity according to the narrative information in the scene parameters to obtain a narrative-enhanced special effect sequence; Apply post-production color grading technology to unify the color style of the special effects and the original picture to obtain a color-unified special effect sequence, and perform final visual optimization processing, including sharpening, noise reduction, and detail enhancement, to obtain a refined special effect sequence; Perform a final quality inspection on the refined special effect sequence to obtain a verified special effect sequence, and encode and package the verified special effect sequence in the target output format to obtain a media file that meets the playback specifications; Package the media file that meets the playback specifications with interactive control metadata to support the ability of the end user to make limited adjustments to the special effect intensity, and finally generate an AI visual special effect image sequence integrating multi-modal perception.

Citation Information

Patent Citations

  • AR / VR real-time scene construction and rendering system for multi-dimensional scene matching

    CN118071933A

  • Special effect video generation method and device, electronic equipment and readable storage medium

    CN118138830A

  • Generalized neural radiation field reconstruction method based on multi-modal information fusion

    CN119359934A

  • KR20240049098A

Cited By

  • Popularization method and system for ice powder marketing

    CN120525616A

  • Game scene optimization method and system based on dynamic rendering

    CN120586387A

  • Visual special effect video generation method and system

    CN120602747A

  • Method and system for generating visual effects video

    CN120602747B

  • Lens label extraction method and device based on full-modal understanding and medium

    CN120635790A