Stage visual content real-time rendering system and method based on ai multi-modal generation
By constructing a real-time rendering index and introducing a standard rendering terminology system, the rendering error problem caused by the semantic gap in AI multimodal generation was solved, achieving efficient and accurate rendering of stage visual content and reducing the waste of time and human resources.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG UNIV OF TECH
- Filing Date
- 2026-04-30
- Publication Date
- 2026-07-24
AI Technical Summary
In existing AI-generated multimodal stage visual content rendering systems, semantic gaps and professional issues lead to audio-visual misalignment and character feature confusion in the generated effects, requiring extensive corrections with inaccurate results, resulting in wasted time and human resources.
A real-time rendering index is constructed, rendering semantic feature vectors are extracted through the semantic acquisition module, and cosine similarity calculation and cross-modal similarity function are used for classification and matching. A standard rendering terminology system and priority adjustment are introduced to ensure that the rendering scheme is highly consistent with user needs.
Improve rendering accuracy, reduce post-processing costs, ensure rendering results match user intent, reduce errors such as character identity confusion, and improve rendering efficiency and accuracy.
Smart Images

Figure CN122453971A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal stage visual content rendering technology, specifically to a real-time rendering system and method for stage visual content based on AI multimodal generation. Background Technology
[0002] With the development of AI technology, many stage effects can be processed and rendered using AI's multimodal models, achieving the same stage effects as those traditionally produced manually, while also saving a significant amount of manpower, material resources, and financial resources.
[0003] However, in the actual generation process, it is necessary to generate corresponding instructions based on a large amount of professional natural language input. During this process, due to semantic gaps and personnel professionalism issues, the generated effect may have problems such as audio-visual misalignment, confusion of role characteristics or identity, which leads to a large number of corrections in the later stage. Moreover, the correction results may not be accurate, resulting in a waste of a lot of time and human resources.
[0004] To address this, a real-time rendering system and method for stage visual content based on AI multimodal generation is proposed. Summary of the Invention
[0005] The purpose of this invention is to provide a real-time rendering system and method for stage visual content based on AI multimodal generation. By generating rendering indexes and optimizing, the system aims to reduce corrections, improve rendering accuracy, and reduce the waste of time and human resources.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A real-time rendering system for stage visual content based on AI multimodal generation includes:
[0008] The semantic acquisition module is used to acquire rendering semantic information of stage visuals and establish a real-time rendering index based on the rendering semantic information.
[0009] Optionally, multiple stage rendering semantic feature vectors input by the user are extracted based on the trained semantic feature extraction model, and the multiple stage rendering semantic feature information is input into the rendering semantic library and classified by cosine similarity calculation to obtain multiple rendering semantic category vectors. The rendering semantic category includes at least the rendering object, the rendering method and the rendering result.
[0010] Cosine similarity is calculated for each of the rendering semantic category vectors to obtain rendering relevance values for multiple rendering semantic category vectors. These rendering relevance values are then used as real-time index weights to connect the rendering semantic categories and generate a real-time rendering index. This method enables the structured organization of user-input rendering semantic information, transforming potentially fragmented and ambiguous natural language instructions into an index system with clear categories and associated weights. This constructs a hierarchical and clearly correlated real-time rendering index, laying a data foundation for subsequent multimodal feature matching and scheme optimization. It effectively avoids the problem of unclear rendering targets caused by semantic information chaos, and also establishes a basic rendering workflow framework, improving rendering accuracy.
[0011] The multimodal index matching module is used to obtain the rendering model features of the multimodal model, and obtain a preliminary rendering scheme by matching and verifying the real-time rendering index based on the rendering model features.
[0012] Optionally, a multimodal feature library is constructed by acquiring the multimodal rendering model features of the multimodal model; the rendering semantic category vector in the real-time rendering index is used to calculate the cross-modal similarity between the rendering semantic category vector and the modal rendering model features in the multimodal feature library to obtain the rendering matching degree; the rendering model features with the rendering matching degree greater than the verification matching degree are associated with the real-time rendering index to generate a preliminary rendering scheme; the verification matching degree is set based on historical data and expert experience.
[0013] Cross-modal similarity is calculated based on a cross-modal similarity function, the expression of which is:
[0014] ;
[0015] in, For the j-th semantic category vector and the features of the i-th modal rendering model Rendering matching degree Render the maximum similarity function for the categories to the model. for The weight, For model rendering accuracy function, for The method assigns weights to different modalities of rendering models, such as images, audio, and text, and accurately matches them with the real-time rendering index. By constructing a multimodal feature library and using a cross-modal similarity function to calculate the rendering matching degree, the method comprehensively considers the maximum similarity of model rendering categories and model rendering accuracy, ensuring that the matched rendering model features highly match the user's semantic needs. Simultaneously, the matching degree verification process effectively filters out irrelevant or low-matching model features, providing a high-quality foundation for subsequent semantic optimization and reducing rendering deviations caused by improper feature matching.
[0016] The semantic rendering optimization module is used to perform semantic classification based on the real-time rendering index of the preliminary rendering scheme, and to perform semantic optimization on the real-time rendering index after semantic classification to generate a semantically optimized rendering scheme.
[0017] Optionally, a term vector library of standard rendering terms is obtained, and all the term vector libraries are normalized and then dimensionality is reduced using the PCA algorithm; DBSCAN clustering is performed on the dimensionality-reduced standard rendering terms to obtain multiple term vector clusters, and the standard rendering term closest to the cluster center of the term vector cluster is mapped to the cluster label of the term vector cluster.
[0018] After normalizing the real-time rendering index of the preliminary rendering scheme, dimensionality reduction is performed using the PCA algorithm. The dimensionality-reduced real-time rendering index is then clustered in multiple term vector clusters using the DBSCAN clustering algorithm. The term vector cluster closest to the cluster center of the term vector cluster is obtained, and the cluster label of the term vector cluster is optimized into a new index for the real-time rendering index.
[0019] A semantically optimized rendering scheme is generated based on the new index optimized from the real-time rendering index. This method standardizes the real-time rendering index by introducing a standard rendering terminology system, effectively solving the semantic ambiguity caused by the fuzziness and lack of professionalism in natural language input. This process not only unifies terminology and eliminates semantic gaps, making the index more accurately reflect rendering requirements, but also improves the consistency of index understanding among various modules in the subsequent rendering process. This provides a strong guarantee for generating a more accurate rendering scheme that better matches user intent, further reducing rendering errors and post-correction costs caused by semantic misunderstanding biases.
[0020] The rendering priority adjustment module is used to obtain the real-time rendering index of the semantic optimization rendering scheme to obtain the rendering key object. The rendering key object includes at least the rendering target and the rendering object. The module establishes the key object priority based on the rendering key object and performs a weighted calculation on the index weight of the real-time rendering index of the semantic optimization rendering scheme based on the key object priority to generate the adjusted and optimized rendering scheme.
[0021] Optionally, the priority of the key objects is established by calculating semantic weights using a semantic weight function based on the rendered key objects. The expression of the semantic weight function is as follows:
[0022] ;
[0023] ;
[0024] in, Let t be the semantic weight. The semantic weights are calculated for the model, and N is the number of real-time rendering index layers. The index weight for real-time rendering of the k-th layer. Let cosine similarity be the activation function. For the k-th semantic, For the t-th semantic to be weighted, The method uses cosine similarity as the activation threshold. By assigning different semantic weights to key rendering objects, a priority system is established, enabling the rendering system to clearly distinguish between core and secondary elements in the rendering process. This approach ensures that the needs of core rendering targets and key objects are prioritized during stage visual rendering, avoiding unreasonable resource allocation or a lack of emphasis on rendering focus due to equal weights for all elements. It also reduces issues such as role confusion during AI rendering, further improving the alignment between rendered content and user needs, and reducing rendering effect deviations caused by priority confusion.
[0025] The real-time rendering module is used to call up materials for stage visual rendering based on the adjusted and optimized rendering scheme.
[0026] This application also provides a method for real-time rendering of stage visual content based on AI multimodal generation, including:
[0027] Obtain the rendering model features of the multimodal model, and obtain a preliminary rendering scheme based on the real-time rendering index after matching and verification of the rendering model features;
[0028] Based on the real-time rendering index of the preliminary rendering scheme, semantic classification is performed, and the semantically classified real-time rendering index is semantically optimized to generate a semantically optimized rendering scheme.
[0029] The real-time rendering index of the semantic optimization rendering scheme is obtained to obtain key rendering objects, which include at least rendering targets and rendering objects. Key object priorities are established based on the key rendering objects. The index weights of the real-time rendering index of the semantic optimization rendering scheme are weighted according to the key object priorities to generate an adjusted and optimized rendering scheme.
[0030] Based on the aforementioned adjusted and optimized rendering scheme, materials are used for stage visual rendering.
[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0032] 1. By constructing a real-time rendering index, the user-input natural language rendering semantic information is transformed into a structured and relational index system, effectively solving the semantic gap problem. Simultaneously, by comprehensively considering the maximum similarity and accuracy of model rendering categories through a cross-modal similarity function, rendering model features are accurately matched, and low-matching features are filtered out, ensuring a high degree of consistency between the initial rendering scheme and the user's semantic needs. A standard rendering terminology system is introduced to standardize the real-time rendering index, unify terminology, eliminate semantic ambiguity, and improve the consistency of index understanding across modules, providing a guarantee for generating accurate rendering schemes. Furthermore, a priority calculation mechanism is introduced to clearly distinguish between core and secondary elements, ensuring that core rendering needs are prioritized, avoiding unreasonable resource allocation and a lack of focus in rendering, reducing errors such as role confusion, and improving the alignment between rendered content and user needs. This reduces rendering errors caused by semantic misunderstanding, improper feature matching, and priority confusion, thereby significantly reducing the time and manpower required for later corrections and improving the efficiency and accuracy of stage visual content rendering. Attached Figure Description
[0033] Figure 1 This is a schematic diagram of the system flow of the present invention.
[0034] Figure 2 This is a schematic diagram of the method flow of the present invention. Detailed Implementation
[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0036] Example 1:
[0037] This invention provides a real-time rendering system for stage visual content based on AI multimodal generation, the technical solution of which is as follows:
[0038] A real-time rendering system for stage visual content based on AI multimodal generation, referencing Figure 1 As shown, it includes:
[0039] The semantic acquisition module is used to acquire rendering semantic information of stage visuals and establish a real-time rendering index based on the rendering semantic information.
[0040] Optionally, multiple stage rendering semantic feature vectors input by the user are extracted based on a trained semantic feature extraction model (e.g., but not limited to the BERT model, etc.). These multiple stage rendering semantic feature information are input into a rendering semantic library and classified by cosine similarity calculation to obtain multiple rendering semantic category vectors. The rendering semantic categories at least include the rendering object, the rendering method, and the rendering result. The stage rendering semantic feature vectors are dynamically expanded based on dynamic parameters of the actual stage scene, such as lighting intensity, stage area division, and performance rhythm, to supplement scene association information, so that the real-time rendering index can more accurately reflect the current stage rendering needs. For example, when a user inputs "to present a rotating sphere in the center of the stage that changes color with the rhythm of the music, with a flowing nebula effect superimposed on the background," the semantic acquisition module extracts semantic feature vectors such as "center of the stage," "rotating sphere," "color change," "music rhythm," "background," and "flowing nebula." It then uses cosine similarity matching with categories in the rendering semantic library, such as "rendering object" (sphere, nebula), "rendering method" (rotation, color change, superimposition), and "rendering result" (changing with the music rhythm), to determine the category vector to which each semantic feature belongs. Subsequently, it calculates the rendering correlation values between these category vectors; for example, "rotating sphere" has a high correlation with "color change," while "music rhythm" has a higher trigger correlation weight with "color change."
[0041] Cosine similarity is calculated for each of the rendering semantic category vectors to obtain rendering relevance values for multiple rendering semantic category vectors. These rendering relevance values are then used as real-time index weights to connect the rendering semantic categories and generate a real-time rendering index. This method enables the structured organization of user-input rendering semantic information, transforming potentially fragmented and ambiguous natural language instructions into an index system with clear categories and associated weights. This constructs a hierarchical and clearly correlated real-time rendering index, laying a data foundation for subsequent multimodal feature matching and scheme optimization. It effectively avoids the problem of unclear rendering targets caused by semantic information chaos, and also establishes a basic rendering workflow framework, improving rendering accuracy.
[0042] The multimodal index matching module is used to obtain the rendering model features of the multimodal model, and obtain a preliminary rendering scheme by matching and verifying the real-time rendering index based on the rendering model features.
[0043] Optionally, a multimodal feature library is constructed by acquiring the multimodal rendering model features of the multimodal model. The multimodal feature library mainly covers the domains that each model is good at processing, such as images, audio, and text. The rendering semantic category vector in the real-time rendering index is compared with the modal rendering model features in the multimodal feature library to calculate the cross-modal similarity and obtain the rendering matching degree. The rendering model features with the rendering matching degree greater than the verification matching degree are associated with the real-time rendering index to generate a preliminary rendering scheme.
[0044] Cross-modal similarity is calculated based on a cross-modal similarity function, the expression of which is:
[0045] ;
[0046] in, For the j-th semantic category vector and the features of the i-th modal rendering model Rendering matching degree This function renders the maximum similarity between categories for the model. The default is the cosine similarity function, but it can be set according to the actual situation. for The weight, This function sets the accuracy for model rendering. By default, it uses the maximum accuracy value from the historical model training process, but it can be set according to the actual situation. for The weighting of the data is determined. This method enables accurate matching of rendering model features from different modalities, such as images, audio, and text, with the real-time rendering index in a multimodal model. By constructing a multimodal feature library and using a cross-modal similarity function to calculate the rendering matching degree, and comprehensively considering the maximum similarity of model rendering categories and model rendering accuracy, it ensures that the matched rendering model features are highly consistent with the user's semantic needs.
[0047] The verification process includes checking whether the features of the rendering model conform to the basic technical specifications of stage visual rendering (such as resolution compatibility, color space matching degree, real-time rendering frame rate requirements, etc.) and whether they are consistent with the core rendering semantics input by the user (such as specific style, emotional tone). The matching degree verification process effectively filters out irrelevant or low-matching model features, providing a high-quality basic solution for subsequent semantic optimization and reducing rendering deviations caused by improper feature matching.
[0048] The semantic rendering optimization module is used to perform semantic classification based on the real-time rendering index of the preliminary rendering scheme, and to perform semantic optimization on the real-time rendering index after semantic classification to generate a semantically optimized rendering scheme.
[0049] Optionally, a term vector library of standard rendering terms is obtained, and all the term vector libraries are normalized and then dimensionality is reduced using the PCA algorithm; DBSCAN clustering is performed on the dimensionality-reduced standard rendering terms to obtain multiple term vector clusters, and the standard rendering term closest to the cluster center of the term vector cluster is mapped to the cluster label of the term vector cluster.
[0050] After normalizing the real-time rendering index of the preliminary rendering scheme, dimensionality reduction is performed using the PCA algorithm. The dimensionality-reduced real-time rendering index is then clustered in multiple term vector clusters using the DBSCAN clustering algorithm. The term vector cluster closest to the cluster center of the term vector cluster is obtained, and the cluster label of the term vector cluster is optimized into a new index for the real-time rendering index. The distance calculation to the cluster includes, but is not limited to, algorithms such as Euclidean distance or cosine similarity.
[0051] For example, when the initial rendering scheme's real-time rendering index includes the phrase "background lighting effect gradient," the semantic rendering optimization module first normalizes it, converting the text information into a standardized vector form. Then, it uses the PCA algorithm to reduce the vector dimension while retaining core features. At this point, the standard rendering term vector library has already formed term vector clusters such as "dynamic lighting effect," "static lighting effect," and "color transition" through DBSCAN clustering. The cluster center of the "color transition" cluster corresponds to the standard term "color gradient effect." The system calculates the distance between the dimensionality-reduced "background lighting effect gradient" index vector and each term vector cluster, finding that it is closest to the cluster center of the "color transition" cluster. Therefore, the real-time rendering index is optimized to the standard term "background color gradient effect" as the new index. Through this process, the potentially ambiguous term "gradient lighting effect" is accurately mapped into the standard terminology system, ensuring that the subsequent rendering module can accurately understand the user's specific requirements for the background lighting effect change method, and avoid rendering deviations caused by inconsistent terminology, such as mistakenly implementing "gradient" as "flickering" or "fixed color value change" and other effects that do not conform to the user's intention.
[0052] A semantically optimized rendering scheme is generated based on the new index optimized from the real-time rendering index. This method standardizes the real-time rendering index by introducing a standard rendering terminology system, effectively solving the semantic ambiguity caused by the fuzziness and lack of professionalism in natural language input. This process not only unifies terminology and eliminates semantic gaps, making the index more accurately reflect rendering requirements, but also improves the consistency of index understanding among various modules in the subsequent rendering process. This provides a strong guarantee for generating a more accurate rendering scheme that better matches user intent, further reducing rendering errors and post-correction costs caused by semantic misunderstanding biases.
[0053] The rendering priority adjustment module is used to obtain the real-time rendering index of the semantic optimization rendering scheme to obtain the rendering key object. The rendering key object includes at least the rendering target and the rendering object. The module establishes the key object priority based on the rendering key object and performs a weighted calculation on the index weight of the real-time rendering index of the semantic optimization rendering scheme based on the key object priority to generate the adjusted and optimized rendering scheme.
[0054] Optionally, the priority of the key objects is established by calculating semantic weights using a semantic weight function based on the rendered key objects. The expression of the semantic weight function is as follows:
[0055] ;
[0056] ;
[0057] in, Let t be the semantic weight. The semantic weights are calculated for the model, and N is the number of real-time rendering index layers. The index weight for real-time rendering of the k-th layer. Let cosine similarity be the activation function. For the k-th semantic, For the t-th semantic to be weighted, The activation threshold for cosine similarity is set by relevant technical personnel.
[0058] The index weight of the real-time rendering index of the semantically optimized rendering scheme based on the priority of the key objects. The expression for generating the weighted calculation and optimized rendering scheme is: ;in, The index weights of the rendering scheme have been optimized for the updated adjustments.
[0059] This method establishes a priority system by assigning different semantic weights to key rendering objects, enabling the rendering system to clearly distinguish between core and secondary elements in the rendering process. This approach ensures that the needs of core rendering targets and key objects are prioritized during stage visual rendering, avoiding problems such as unreasonable resource allocation or lack of emphasis in rendering due to equal weights for each element. It also reduces issues such as role confusion during AI rendering, further improving the fit between the rendered content and user needs, and reducing rendering effect deviations caused by priority confusion.
[0060] The real-time rendering module is used to call up materials for stage visual rendering based on the adjusted and optimized rendering scheme.
[0061] Example 2:
[0062] Based on the content of Embodiment 1, this application also provides a real-time rendering method for stage visual content generated by AI multimodal generation, see reference. Figure 2 As shown, it includes:
[0063] Obtain the rendering model features of the multimodal model, and obtain a preliminary rendering scheme based on the real-time rendering index after matching and verification of the rendering model features;
[0064] Based on the real-time rendering index of the preliminary rendering scheme, semantic classification is performed, and the semantically classified real-time rendering index is semantically optimized to generate a semantically optimized rendering scheme.
[0065] The real-time rendering index of the semantic optimization rendering scheme is obtained to obtain key rendering objects, which include at least rendering targets and rendering objects. Key object priorities are established based on the key rendering objects. The index weights of the real-time rendering index of the semantic optimization rendering scheme are weighted according to the key object priorities to generate an adjusted and optimized rendering scheme.
[0066] Based on the aforementioned adjusted and optimized rendering scheme, materials are used for stage visual rendering.
[0067] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A real-time rendering system for stage visual content based on AI multimodal generation, characterized in that, include: The semantic acquisition module is used to acquire rendering semantic information of stage visuals and establish a real-time rendering index based on the rendering semantic information. The multimodal index matching module is used to obtain the rendering model features of the multimodal model, and obtain a preliminary rendering scheme by matching and verifying the real-time rendering index based on the rendering model features. The semantic rendering optimization module is used to perform semantic classification based on the real-time rendering index of the preliminary rendering scheme, and to perform semantic optimization on the real-time rendering index after semantic classification to generate a semantically optimized rendering scheme. The rendering priority adjustment module is used to obtain the real-time rendering index of the semantic optimization rendering scheme to obtain the rendering key object. The rendering key object includes at least the rendering target and the rendering object. The module establishes the key object priority based on the rendering key object and performs a weighted calculation on the index weight of the real-time rendering index of the semantic optimization rendering scheme based on the key object priority to generate the adjusted and optimized rendering scheme. The real-time rendering module is used to call up materials for stage visual rendering based on the adjusted and optimized rendering scheme.
2. The real-time rendering system for stage visual content based on AI multimodal generation according to claim 1, characterized in that, Obtaining stage visual rendering semantic information and establishing a real-time rendering index based on the rendering semantic information includes: Based on the trained semantic feature extraction model, multiple stage rendering semantic feature vectors input by the user are extracted. The multiple stage rendering semantic feature information is input into the rendering semantic library and classified by cosine similarity calculation to obtain multiple rendering semantic category vectors. The rendering semantic category includes at least the rendering object, the rendering method and the rendering result. Cosine similarity is calculated for each of the rendering semantic category vectors to obtain rendering relevance values for multiple rendering semantic category vectors. The rendering relevance values are then used as real-time index weights to connect the rendering semantic categories and generate a real-time rendering index.
3. The real-time rendering system for stage visual content based on AI multimodal generation according to claim 1, characterized in that, Obtaining the rendering model features of the multimodal model, and obtaining a preliminary rendering scheme based on the real-time rendering index after matching and verification of the rendering model features, includes: A multimodal feature library is constructed by acquiring multimodal rendering model features of the multimodal model; cross-modal similarity calculation is performed between the rendering semantic category vector in the real-time rendering index and the modal rendering model features in the multimodal feature library to obtain the rendering matching degree; the rendering model features with rendering matching degrees greater than the verification matching degree are associated with the real-time rendering index to generate a preliminary rendering scheme. Cross-modal similarity is calculated based on a cross-modal similarity function, the expression of which is: ; in, For the j-th semantic category vector and the features of the i-th modal rendering model Rendering matching degree Render the maximum similarity function for the categories to the model. for The weight, For model rendering accuracy function, for The weight.
4. The real-time rendering system for stage visual content based on AI multimodal generation according to claim 1, characterized in that, Based on the real-time rendering index of the preliminary rendering scheme, semantic classification is performed, and semantic optimization is performed on the semantically classified real-time rendering index to generate a semantically optimized rendering scheme, including: Obtain a term vector library of standard rendering terms, normalize all the term vector libraries and then reduce their dimensionality using the PCA algorithm; perform DBSCAN clustering calculation on the dimensionality-reduced standard rendering terms to obtain multiple term vector clusters, and map the standard rendering term closest to the cluster center of the term vector cluster to the cluster label of the term vector cluster. After normalizing the real-time rendering index of the preliminary rendering scheme, dimensionality reduction is performed using the PCA algorithm. The dimensionality-reduced real-time rendering index is then clustered in multiple term vector clusters using the DBSCAN clustering algorithm. The term vector cluster closest to the cluster center of the term vector cluster is obtained, and the cluster label of the term vector cluster is optimized into a new index for the real-time rendering index. A semantically optimized rendering scheme is generated based on the new index optimized from the real-time rendering index.
5. The real-time rendering system for stage visual content based on AI multimodal generation according to claim 1, characterized in that, The process of obtaining the real-time rendering index of the semantically optimized rendering scheme to obtain key rendering objects, wherein the key rendering objects include at least a rendering target and a rendering object, establishing key object priorities based on the key rendering objects, and generating an adjusted and optimized rendering scheme by weighting the index weights of the real-time rendering index of the semantically optimized rendering scheme based on the key object priorities, includes: Based on the rendering key objects, semantic weights are calculated using a semantic weight function to establish the priority of the key objects. The expression of the semantic weight function is as follows: ; ; in, Let t be the semantic weight. The semantic weights calculated for the model, where N is the actual weight. Render the index level at any time. The index weight for the real-time rendering index of the k-th layer. Let cosine similarity be the activation function. For the k-th semantic, For the t-th semantic to be weighted, The cosine similarity activation threshold is used.
6. A real-time rendering method for stage visual content based on AI multimodal generation, characterized in that, include: Obtain the rendering semantic information of the stage visuals, and establish a real-time rendering index based on the rendering semantic information; Obtain the rendering model features of the multimodal model, and obtain a preliminary rendering scheme based on the real-time rendering index after matching and verification of the rendering model features; Based on the real-time rendering index of the preliminary rendering scheme, semantic classification is performed, and the semantically classified real-time rendering index is semantically optimized to generate a semantically optimized rendering scheme. The real-time rendering index of the semantic optimization rendering scheme is obtained to obtain key rendering objects, which include at least rendering targets and rendering objects. Key object priorities are established based on the key rendering objects. The index weights of the real-time rendering index of the semantic optimization rendering scheme are weighted according to the key object priorities to generate an adjusted and optimized rendering scheme. Based on the aforementioned adjusted and optimized rendering scheme, materials are used for stage visual rendering.