Method and platform for generating metaverse content based on multi-modal large model

By parsing multimodal inputs using a multimodal large model, an initial structured content list is generated. Combined with a historical task experience base, a heterogeneous common knowledge graph is constructed, solving the adaptation problem of multimodal fusion input in existing technologies and achieving efficient and accurate metaverse content generation.

CN121809696BActive Publication Date: 2026-06-02BEIJING MIAOYIN ANIMATION CULTURE CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING MIAOYIN ANIMATION CULTURE CO LTD
Filing Date
2026-03-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing metaverse content generation technologies struggle to adapt to multimodal fusion inputs and lack the systematic reuse of cross-scene, cross-modal historical task experience, resulting in low matching degree between generated content and user needs, low efficiency, and poor quality.

Method used

The system analyzes requirements using a multimodal large model, generates an initial structured content list, performs retrieval and matching based on a historical task experience base, extracts a heterogeneous set of similar tasks, and constructs a heterogeneous common knowledge graph to achieve cross-modal semantic alignment and accurate extraction and fusion of common information.

Benefits of technology

It improves the alignment between generated content and user needs, enhances generation efficiency and quality stability, and ensures the consistency and reliability of generated results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809696B_ABST
    Figure CN121809696B_ABST
Patent Text Reader

Abstract

The application provides a method and platform for generating metaverse content based on a multi-modal large model, and relates to the technical field of metaverse content generation. The method comprises: in response to original multi-modal input of a target user, performing demand analysis through a multi-modal large model to generate an initial structured content list; based on the initial structured content list, performing retrieval matching in a historical task experience library to obtain a heterogeneous approximate task set, wherein the heterogeneous approximate task set is configured differently from the generation mode path of the initial structured content list; analyzing the heterogeneous approximate task set through the multi-modal large model, extracting heterogeneous common information, and constructing the heterogeneous common information into a structured heterogeneous common knowledge graph; and generating and merging output of metaverse content based on the initial structured content list and the heterogeneous common knowledge graph in combination with the multi-modal large model. The method improves the generation efficiency and quality of metaverse content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of metaverse content generation technology, and in particular to a metaverse content generation method and platform based on a multimodal large model. Background Technology

[0002] With the rapid development of metaverse technology, users' demand for metaverse content generation is becoming more multimodal, personalized, and complex.

[0003] However, existing metaverse content generation technologies mostly focus on parsing and generating single-modal inputs, making it difficult to effectively adapt to complex input formats such as text, images, voice, and video, resulting in insufficient alignment between generated content and users' real needs. At the same time, existing methods lack a systematic reuse mechanism for historical task experience across scenarios and modalities, and a large amount of repetitive development not only restricts the efficiency of content generation but also makes it difficult to ensure the consistency and reliability of the generated results.

[0004] Therefore, there is an urgent need for a metaverse content generation method and platform based on a multimodal large model to break through the existing technological bottlenecks. Summary of the Invention

[0005] This invention addresses the technical problems in existing technologies where metaverse content generation struggles to adapt to multimodal fusion inputs and lacks systematic reuse of historical task experience, resulting in low matching between generated content and user needs, low generation efficiency, and poor quality. It provides a metaverse content generation method and platform based on a multimodal large model.

[0006] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:

[0007] In a first aspect, the present invention provides a method for generating metaverse content based on a multimodal large model, including:

[0008] In response to the target user's original multimodal input, the requirements are analyzed through a large multimodal model to generate an initial structured content list;

[0009] Based on the initial structured content list, a search and matching process is performed in the historical task experience base to obtain a heterogeneous set of similar tasks, wherein the heterogeneous set of similar tasks is configured differently from the generation modal path of the initial structured content list.

[0010] The heterogeneous approximate task set is analyzed by multimodal large model, heterogeneous common information is extracted, and the heterogeneous common information is constructed into a structured heterogeneous common knowledge graph.

[0011] Based on the initial structured content list and the heterogeneous common knowledge graph, the metaverse content is generated and merged and output using a multimodal large model.

[0012] Secondly, this invention provides a metaverse content generation platform based on a multimodal large model, including:

[0013] The initial inventory generation module is used to respond to the target user's original multimodal input, analyze the requirements through a multimodal big model, and generate an initial structured content inventory.

[0014] The retrieval and matching module is used to perform retrieval and matching in the historical task experience base based on the initial structured content list to obtain a heterogeneous set of similar tasks, wherein the heterogeneous set of similar tasks is configured differently from the generation modal path of the initial structured content list.

[0015] The graph construction module is used to parse the heterogeneous approximate task set through a multimodal large model, extract heterogeneous common information, and construct the heterogeneous common information into a structured heterogeneous common knowledge graph.

[0016] The merged output module is used to generate and merge metaverse content based on the initial structured content list and the heterogeneous common knowledge graph, combined with the multimodal large model.

[0017] The beneficial effects of this invention are:

[0018] Compared to existing technologies, this application first responds to the original multimodal input of the target user, using a large multimodal model to analyze requirements and generate an initial structured content list, achieving cross-modal semantic alignment and accurate extraction of core requirements from the original multimodal input. Secondly, based on the initial structured content list, it performs retrieval and matching in a historical task experience database to obtain a set of heterogeneous similar tasks, broadening the scope of reuse of historical task experience and effectively mining cross-modal common knowledge. Thirdly, it analyzes the set of heterogeneous similar tasks using a large multimodal model, extracting heterogeneous common information and constructing it into a structured heterogeneous common knowledge graph. This achieves accurate extraction of cross-modal common information from the set of heterogeneous similar tasks, eliminating knowledge reuse barriers caused by heterogeneous modal differences and providing efficient and callable knowledge support for the generation of metaverse content. Finally, based on the initial structured content list and the heterogeneous common knowledge graph, the metaverse content is generated and merged using a multimodal big model. This deeply integrates the personalized requirements of the initial structured content list with the common knowledge of the heterogeneous common knowledge graph. The multimodal big model efficiently completes the generation of metaverse content and the merging of multimodal outputs, ensuring accurate matching between the generated content and user needs, while also improving the quality, efficiency, and coherence and consistency of content generation and modal fusion.

[0019] Through the above technical solutions, this application achieves efficient adaptation to multimodal fusion input, improves the consistency between generated content and users' real needs, and reduces repetitive development costs by effectively reusing cross-modal historical task experience, thereby improving the generation efficiency and quality stability of metaverse content, while ensuring the consistency and reliability of the generated results. Attached Figure Description

[0020] Figure 1 A flowchart illustrating the metaverse content generation method based on a multimodal large model provided by this invention;

[0021] Figure 2 This is a schematic diagram of the structure of the metaverse content generation platform based on a multimodal large model provided by the present invention.

[0022] In the attached diagram, the components represented by each number are as follows:

[0023] Initial list generation module 11, retrieval and matching module 12, map construction module 13, and merged output module 14. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0026] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.

[0027] Example 1, as Figure 1 As shown, embodiments of the present invention provide a method for generating metaverse content based on a multimodal large model, including:

[0028] S10: In response to the target user's original multimodal input, perform requirement analysis through a multimodal large model to generate an initial structured content list.

[0029] As the demand for metaverse content evolves towards multimodal fusion of text, images, and voice, traditional content generation methods suffer from limitations such as single-modal input, insufficient multimodal semantic alignment, and a lack of structured output in demand parsing. This results in a low degree of matching between users' multimodal demands and generation goals, and makes it difficult to support subsequent accurate generation processes.

[0030] To address the aforementioned issues, this application responds to the target user's original multimodal input by performing requirement analysis through a multimodal large model to generate an initial structured content list.

[0031] Specifically, step S10 in the method includes:

[0032] The original multimodal input is semantically parsed using multimodal aligned coding, and the semantic parsing results are mapped to a unified multimodal semantic space to generate a multimodal semantic representation.

[0033] Based on the content complexity of multimodal semantic representation, we perform modal path planning and modal standard initialization to obtain the initial modal path and initial modal standard.

[0034] By combining the preset data paradigm, the multimodal semantic representation, the initial generated modal path, and the initial generated standard are structured and output to obtain the initial structured content list.

[0035] In this embodiment, the original multimodal input is first semantically parsed using multimodal alignment coding, and the semantic parsing result is mapped to a unified multimodal semantic space to generate a multimodal semantic representation. Multimodal alignment coding refers to a coding technique that aligns the semantic information of data from different modalities. Semantic parsing of the original multimodal input using multimodal alignment coding aims to eliminate semantic differences between different modalities. A unified multimodal semantic space is a pre-defined high-dimensional vector space capable of accommodating the semantic information of data from different modalities, making the semantic representations of different modalities comparable and fusionable. A multimodal semantic representation refers to the high-dimensional vector form of the semantic expression formed in the unified multimodal semantic space after the original multimodal input has undergone multimodal alignment coding.

[0036] Specifically, for each modal data in the original multimodal input, a corresponding modal encoder is used to extract features. Then, multimodal alignment algorithms, such as cross-modal attention mechanisms and modal adaptive weight allocation algorithms, are used to semantically align the features of each modality. Finally, the aligned features are mapped to a unified multimodal semantic space to form a multimodal semantic representation that can comprehensively reflect the semantics of all modalities.

[0037] For example, if the original multimodal input is text: red circular object, image: red circular picture, and speech: red circle, then the text encoder extracts text semantic features, the image encoder extracts image visual features, and the speech encoder extracts speech acoustic features. Then, the semantic bias of the three types of features is eliminated by multimodal alignment encoding, and mapped to a unified multimodal semantic space to generate a multimodal semantic representation containing core semantics such as red, circle, and object.

[0038] Secondly, based on the content complexity of the multimodal semantic representation, we plan the generation modality path and initialize the generation criteria to obtain the initial generation modality path and initial generation criteria. Here, content complexity refers to the complexity of the metaverse content corresponding to the multimodal semantic representation, and its evaluation indicators can include the number of content elements, modality fusion difficulty, and detail refinement requirements. The initial generation modality path refers to the initial transformation path from the input modality to the output modality planned according to the content complexity. The initial generation criteria refer to the quality standards and constraints set based on the content complexity for the generation of metaverse content.

[0039] Specifically, by analyzing parameters such as the dimension, sparsity, and feature correlation of multimodal semantic representations, the complexity of the content is evaluated, such as low complexity, medium complexity, and high complexity. Then, corresponding generation modal paths are planned for different complexities. For example, low-complexity content adopts a single-modal direct generation path, medium-complexity content adopts a dual-modal fusion generation path, and high-complexity content adopts a multimodal step-by-step generation path. At the same time, corresponding initial generation standards are set. For example, the accuracy requirement for low-complexity content is 80%, the accuracy requirement for medium-complexity content is 85%, and the accuracy requirement for high-complexity content is 90%.

[0040] For example, if the multimodal semantic representation is: generate a simple blue cube 3D model, based on the dimension, sparsity, and feature correlation of the multimodal semantic representation, it is determined to be low complexity, and the planned initial generation modal path for it is: 3D model generation, and the initial generation criteria for it include cube side length error ≤ 0.1, color deviation ≤ 5%, and model face number ≥ 100.

[0041] Finally, based on a predefined data paradigm, the multimodal semantic representation, the initial generation modality path, and the initial generation criteria are output in a structured manner to obtain an initial structured content list. The predefined data paradigm refers to a predefined structured output format used to standardize the fields, data types, and organization of the initial structured content list. For example, the predefined data paradigm may include fields such as list number, requirement description, input modality type, core multimodal semantic features, initial generation modality path steps, generation criteria for each step, time constraints, and resource constraints. The multimodal semantic representation, initial generation modality path, and initial generation criteria are filled and organized according to this paradigm to form a structured initial structured content list.

[0042] For example, if the field for the initial generated modal path in the preset data paradigm is required to be filled in the following format: step number, modal processing method, output result, if the initial generated modal path is: text, image semantics, character generation, scene generation, animation fusion, then the field should be filled in as: 1. text and image semantic alignment, character generation input, 2. character model generation, 3D character model, 3. scene model generation, 3D scene model, 4. character and scene fusion, animation rendering, 3D animation.

[0043] In summary, compared to existing technologies, this application responds to the target user's original multimodal input by performing requirement analysis through a large multimodal model to generate an initial structured content list. This achieves cross-modal semantic alignment and accurate extraction of core requirements from the original multimodal input. The generated initial structured content list provides a clear and unified basis for subsequent metaverse content generation, improving the accuracy of requirement analysis and the relevance of the subsequent generation process.

[0044] S20: Based on the initial structured content list, perform a search and matching in the historical task experience base to obtain a heterogeneous approximate task set, wherein the heterogeneous approximate task set is configured differently from the generation modal path of the initial structured content list.

[0045] Existing methods for generating metaverse content often limit themselves to similar tasks that align with the current modal path when reusing historical task experience. This neglects the common knowledge hidden in heterogeneous modal paths, resulting in a narrow scope of historical experience reuse, low content generation efficiency, and a tendency to homogenize content. Furthermore, redundant development leads to resource waste.

[0046] To address the aforementioned issues, this application, based on the initial structured content list, performs a search and matching process in a historical task experience database to obtain a heterogeneous set of approximate tasks, wherein the heterogeneous set of approximate tasks is configured differently from the generation modal path of the initial structured content list.

[0047] Specifically, step S20 in the method includes:

[0048] Define the equivalent vacancy rate rule and the corresponding equivalent vacancy rate constraint;

[0049] Based on the vacancy rate constraint, the multimodal semantic representation of the initial structured content list is randomly pruned to obtain several multimodal semantic index vectors that satisfy the vacancy rate constraint.

[0050] Based on several multimodal semantic index vectors, iterative vector similarity retrieval is performed by traversing the historical task experience base to extract historical task records that meet preset similarity constraints, and the output is the heterogeneous approximate task set.

[0051] In this embodiment, the equivalent vacancy rate rule and its corresponding equivalent vacancy rate constraint are first defined. The equivalent vacancy rate rule refers to the rule used to measure the impact of missing dimensions in a multimodal semantic index vector on semantic expression; the equivalent vacancy rate constraint refers to the threshold range set for the equivalent vacancy rate of the multimodal semantic index vector, used to ensure that the index vector retains core semantic information even after missing dimensions. This is because multimodal semantic vectors of different dimensions contribute differently to the core semantics, and simple random vacancy can lead to semantic distortion. Therefore, it is necessary to transform vacancy in different dimensions into a unified equivalent vacancy index through the equivalent vacancy rate rule, and then set the equivalent vacancy rate constraint to ensure the semantic integrity of the index vector.

[0052] Secondly, based on the vacancy rate constraint, the multimodal semantic representation of the initial structured content list is randomly pruned in dimensions to obtain several multimodal semantic index vectors that satisfy the vacancy rate constraint. Random dimension pruning refers to randomly deleting some dimensions from the multimodal semantic representation; the multimodal semantic index vector is the vector used for retrieval in the historical task experience database after random dimension pruning. Specifically, according to the equivalent vacancy rate rule, the range of pruning dimensions is determined, and the multimodal semantic representation is randomly pruned multiple times. After each pruning, the equivalent vacancy rate is calculated, and vectors with equivalent vacancy rates within the constraint range are selected as multimodal semantic index vectors. By generating multimodal semantic index vectors, the missed detection problem caused by single-vector retrieval can be avoided, improving the comprehensiveness of the retrieval.

[0053] For example, if the multimodal semantic representation is 100-dimensional and the vacancy rate constraint is ≤20%, then the number of dimensions for each random dimension pruning is ≤20. After 10 random dimension prunings, 8 vectors with equivalent vacancy rates between 10% and 20% are selected as multimodal semantic index vectors.

[0054] Finally, based on several multimodal semantic index vectors, iterative vector similarity retrieval is performed by traversing the historical task experience base to extract historical task records that meet preset similarity constraints, and the output is a heterogeneous set of approximate tasks. Specifically, iterative vector similarity retrieval refers to the process of sequentially searching the historical task experience base using multiple multimodal semantic index vectors and iteratively optimizing the retrieval results; preset similarity constraints refer to pre-set vector similarity thresholds used to determine the degree of approximation between historical tasks and current requirements.

[0055] Furthermore, the step of "iteratively searching vector similarity by traversing the historical task experience base based on several multimodal semantic index vectors, extracting historical task records that satisfy preset similarity constraints, and outputting the heterogeneous approximate task set" includes:

[0056] Extract the multimodal semantic history vector of each historical task record;

[0057] Based on the multimodal semantic history vector and the multimodal semantic index vector, a vector similarity retrieval is performed to obtain the first approximate task set;

[0058] The historical task experience base is traversed, and the cross-modal semantic history vector of each historical task record is obtained through multimodal large model analysis.

[0059] Cross-modal association retrieval is performed based on the cross-modal semantic history vector and the multimodal semantic index vector to obtain a second approximate task set;

[0060] Based on the first approximate task set and the second approximate task set, heterogeneity screening for generating modal paths is performed to obtain the heterogeneous approximate task set.

[0061] In this embodiment, the multimodal semantic history vector of each historical task record is first extracted. The multimodal semantic history vector refers to the semantic representation vector formed in a unified multimodal semantic space after multimodal alignment encoding of historical input multimodal data in the historical task record. Its generation method is consistent with the multimodal semantic representation required by the current requirement. Specifically, each historical task record in the historical task experience base is traversed, and the same multimodal alignment encoding method as the current requirement is called to perform semantic parsing and mapping on the input multimodal data of the historical task, generating the corresponding multimodal semantic history vector, ensuring that this vector is comparable to the multimodal semantic index vector required by the current requirement.

[0062] Secondly, vector similarity retrieval is performed based on the multimodal semantic history vector and the multimodal semantic index vector to obtain the first approximate task set. The first approximate task set refers to the set of approximate historical tasks selected through direct similarity matching between the multimodal semantic history vector and the multimodal semantic index vector.

[0063] Specifically, the cosine similarity between each multimodal semantic index vector and all multimodal semantic history vectors is calculated, and a preset similarity threshold, such as 0.75, is set. Historical task records with a cosine similarity greater than or equal to the preset similarity threshold are included in the first approximate task set. In this way, preliminary screening based on semantic similarity ensures that the first approximate task set is highly semantically relevant to the current requirement.

[0064] Secondly, the historical task experience base is traversed, and cross-modal semantic history vectors for each historical task record are obtained through multimodal large-scale model analysis. These cross-modal semantic history vectors are cross-modal vectors that reflect the core semantics of the historical task, derived from the output modality results of the historical task through reverse inference using the multimodal large-scale model. Specifically, for each historical task record, its output modality results are extracted, and the multimodal large-scale model performs semantic parsing on the output results, mapping them back to a unified multimodal semantic space to generate cross-modal semantic history vectors. These cross-modal semantic history vectors reflect the core semantics of the historical task from the perspective of the output results, complementing the multimodal semantic history vectors.

[0065] Furthermore, cross-modal association retrieval is performed based on the cross-modal semantic history vector and the multimodal semantic index vector to obtain a second approximate task set. Here, cross-modal association retrieval refers to retrieval based on the semantic association between the cross-modal semantic history vector and the multimodal semantic index vector; the second approximate task set refers to the set of approximate historical tasks selected through cross-modal association retrieval.

[0066] For example, the semantic relevance of each multimodal semantic index vector to all cross-modal semantic history vectors is calculated. This semantic relevance can be calculated using cosine similarity, and a preset relevance threshold, such as 0.7, is set. Historical task records with a semantic relevance ≥ the preset threshold are included in the second approximate task set. This overcomes the limitation of the first approximate task set, which only retrieves data from the input perspective, and filters out historical tasks with different input modalities but similar output semantics.

[0067] Finally, a heterogeneous task set is obtained by performing heterogeneous filtering based on the first and second approximate task sets to select the first and second approximate task sets according to the degree of difference in their generated modal paths. For example, the generated modal paths of each historical task in the first and second approximate task sets are extracted and compared with the initial generated modal paths of the current initial structured content list. Path heterogeneity indices are calculated, such as the number of path step differences, the number of modality type differences, and the number of processing method differences. A preset heterogeneity threshold is set, such as a path difference count ≥ 2. Historical tasks whose path heterogeneity indices meet the preset threshold are selected, merged, and deduplicated to obtain the heterogeneous task set. This ensures that the final heterogeneous task set differs significantly from the currently required generated modal paths, providing a foundation for subsequent extraction of heterogeneous common information.

[0068] Further, the step of "performing heterogeneity screening for generative modal paths based on the first approximate task set and the second approximate task set to obtain the heterogeneous approximate task set" includes:

[0069] Traverse the first approximate task set and the second approximate task set to extract contrast modal information, wherein the contrast modal information includes input modal combination, output modal combination and modal sequence;

[0070] The comparison modal information is parameterized and encoded by an encoder to obtain the comparison generation modal path;

[0071] Based on the initial structured content list, the initial generated modal path is extracted and parameterized accordingly to obtain the reference generated modal path;

[0072] The parameterized residuals of multiple contrast generated modal paths and the reference generated modal path are calculated iteratively, and the selection probabilities of the first approximate task set and the second approximate task set are defined based on the parameterized residuals, wherein the parameterized residuals are negatively correlated with the selection probabilities;

[0073] Based on the selection probability, the first approximate task set and the second approximate task set are traversed to perform heterogeneity screening based on iterative probability selection, thereby obtaining the heterogeneous approximate task set.

[0074] In this embodiment, the first and second approximate task sets are traversed to extract contrast modal information. This contrast modal information includes input modal combination, output modal combination, and modal sequence: input modal combination refers to the set of multiple input modalities provided by the user in the historical task, such as [text, image, speech, video, etc.]; output modal combination refers to the set of multiple output modalities generated by the historical task, such as [3D model, animation, speech, text, etc.]; modal sequence refers to the processing and fusion order of each modality during the generation process of the historical task, such as text → image → 3D model, speech → text → animation, etc.

[0075] For example, each historical task record in the first approximate task set and the second approximate task set is traversed, and the input mode combination, output mode combination and mode sequence are extracted from the task description and generation path field to obtain the comparative mode information of each historical task.

[0076] Secondly, the contrast modality information is parametrically encoded using an encoder to obtain the contrast generation modality path. Here, the encoder refers to an encoding model that can convert contrast modality information into parametric vectors, such as the Transformer encoder or CNN encoder; the contrast generation modality path refers to the vector form that, after parametric encoding, can quantify the generation modality path of the historical task.

[0077] For example, the input modality combinations, output modality combinations, and modality sequences in the contrast modality information are mapped to numerical codes, such as text modality code 1, image code 2, speech code 3, video code 4, and 3D model code 5. Then, an encoder extracts and fuses features from the numerical codes to generate a fixed-dimensional parameterized vector, which serves as the contrast generated modality path. The contrast generated modality path can accurately quantify the generated modality path features of historical tasks, facilitating subsequent calculation of path differences.

[0078] Next, based on the initial structured content list, the initial generated modal path is extracted and parameterized, and a reference generated modal path is obtained. The reference generated modal path refers to the vector form of the initial generated modal path of the current requirement after parameterization and encoding. Its encoding method is consistent with that of the comparison generated modal path, ensuring comparability between the two.

[0079] For example, the input mode combination, output mode combination, and mode sequence corresponding to the initial generated modal path are extracted from the initial structured content list, and converted into parameterized vectors according to the same numerical encoding rules and encoder as the comparison generated modal path to obtain the reference generated modal path.

[0080] Furthermore, the parameterized residuals of multiple contrast-generated modal paths and reference-generated modal paths are iteratively calculated, and the selection probabilities of the first and second approximate task sets are defined based on these parameterized residuals. The parameterized residuals are negatively correlated with the selection probability because a larger parameterized residual indicates a greater path difference, thus a higher selection probability; conversely, a smaller parameterized residual indicates a smaller path difference, thus a lower selection probability. The parameterized residual refers to the degree of difference between the contrast-generated modal paths and the reference-generated modal paths, which can be obtained by calculating Euclidean distance, Manhattan distance, etc.; the selection probability refers to the probability that a historical task is included in the heterogeneous approximate task set.

[0081] For example, for each contrast-generated modal path, its Euclidean distance to the reference-generated modal path is calculated as a parameterized residual. This residual is then mapped to the [0,1] interval using the Sigmoid function to obtain the selection probability. Following the same method, by iteratively calculating the residuals and selection probabilities of multiple contrast-generated modal paths, the selection priority of each historical task can be dynamically adjusted.

[0082] Finally, based on the selection probability, the first and second approximate task sets are traversed to perform heterogeneity filtering based on iterative probability selection, resulting in a heterogeneous approximate task set. For example, a selection probability threshold, such as 0.6, can be set. Each historical task in the first and second approximate task sets is traversed; if the selection probability of a historical task is greater than or equal to the selection probability threshold, it is included in the heterogeneous approximate task set. This iterative probability selection process ensures that the selected historical tasks possess both high path heterogeneity and cover common knowledge across different dimensions, improving the comprehensiveness of subsequent common information extraction.

[0083] Furthermore, the "defining equivalent vacancy rate rules and corresponding equivalent vacancy rate constraints" includes:

[0084] Knowledge is generated by combining prior content, and the intrinsic importance index of each dimension in the multimodal semantic index vector is defined.

[0085] Normalize the intrinsic importance index of multiple dimensions to obtain the dimensional void cost for each dimension;

[0086] The sum of the dimensional vacancy costs of all vacant dimensions is defined as the equivalent vacancy rate, and an equivalent vacancy rate less than or equal to a preset equivalent vacancy rate threshold is defined as the equivalent vacancy rate constraint.

[0087] The equivalent vacancy rate threshold is a significance threshold determined based on statistical analysis methods.

[0088] In this embodiment, the intrinsic importance index of each dimension in the multimodal semantic index vector is first defined by combining prior content generation knowledge. Prior content generation knowledge refers to the knowledge accumulated in the metaverse content generation field regarding the influence of each dimension of the multimodal semantic vector on the generation result. The intrinsic importance index is an indicator that quantifies the importance of each dimension in the multimodal semantic vector to the core semantic expression, with a value range of [0,1]. The closer the intrinsic importance index is to 1, the greater the influence of that dimension on the core semantics.

[0089] For example, knowledge is generated based on prior content, and the semantic features corresponding to each dimension of the multimodal semantic vector, such as color features, shape features, and structural features, are analyzed. The intrinsic importance index of each dimension is determined by the analytic hierarchy process (AHP) or the entropy weight method. For instance, the intrinsic importance index of the dimension corresponding to the core shape in the multimodal semantic vector is 0.9, the intrinsic importance index of the dimension corresponding to the secondary color is 0.3, and the intrinsic importance index of the dimension corresponding to the irrelevant background is 0.1.

[0090] Secondly, the intrinsic importance indices of multiple dimensions are normalized to obtain the dimensionality cost for each dimension. Here, the dimensionality cost refers to the loss cost to the core semantic expression caused by deleting a certain dimension from the multimodal semantic vector; normalization means mapping the intrinsic importance indices of multiple dimensions to the same order of magnitude to ensure the fairness of the calculation of dimensionality cost.

[0091] For example, the min-max normalization method can be used to process the intrinsic importance index of all dimensions, mapping it to the [0,1] interval. The normalized value is the dimensional vacancy cost of that dimension. For instance, if the intrinsic importance index of a certain dimension is 0.9, the dimensional vacancy cost after min-max normalization is 0.9.

[0092] Finally, the sum of the dimensional void costs of all void dimensions is defined as the equivalent void rate, and an equivalent void rate constraint is defined as an equivalent void rate that is less than or equal to a preset equivalent void rate threshold. Here, the equivalent void rate is a comprehensive indicator that measures the degree of core semantic loss after dimensional clipping of multimodal semantic vectors; the equivalent void rate threshold is a significance threshold determined based on statistical analysis methods, i.e., a pre-set, allowed maximum core semantic loss threshold.

[0093] For example, statistical analysis methods refer to methods that determine the significance threshold that can guarantee retrieval accuracy by analyzing a large amount of historical task semantic vector dimension-trimmed data, such as t-tests and analysis of variance. If, after dimension trimming, the sum of the dimension gap costs of all missing dimensions in the multimodal semantic vector is 0.15, then the equivalent gap rate is 0.15. By statistically analyzing the relationship between the equivalent gap rate and retrieval accuracy in historical data, it is determined that when the equivalent gap rate is ≤0.2, the retrieval accuracy is ≥90%. Therefore, the equivalent gap rate threshold can be set to 0.2, and the equivalent gap rate constraint is ≤0.2. This constraint ensures that the multimodal semantic index vector can still retain core semantic information after dimension trimming, avoiding retrieval bias caused by missing dimensions.

[0094] In summary, compared to existing technologies, this application, based on the initial structured content list, performs retrieval and matching within a historical task experience database to obtain a heterogeneous set of approximate tasks. The heterogeneous set of approximate tasks differs from the initial structured content list in its generation modal path configuration. Thus, by obtaining a set of heterogeneous approximate tasks with a different configuration from the initial generation modal path through retrieval and matching, the scope of reuse of historical task experience is broadened, cross-modal common knowledge is effectively mined, and rich data support is provided for subsequent heterogeneous common information extraction and graph construction. Simultaneously, content generation homogenization is avoided, improving generation efficiency and quality.

[0095] S30: The heterogeneous approximate task set is analyzed by multimodal large model, heterogeneous common information is extracted, and the heterogeneous common information is constructed into a structured heterogeneous common knowledge graph.

[0096] In the existing metaverse content generation, the heterogeneous approximate task set contains historical experience under different modal paths. However, these experiences are mostly in a scattered and unstructured form. The differences between heterogeneous modalities make it difficult to extract common information, and the lack of a systematic organization makes it difficult to reuse directly, thus failing to provide accurate knowledge support for subsequent content generation.

[0097] To address the aforementioned issues, this application uses a multimodal large model to analyze the heterogeneous approximate task set, extract heterogeneous common information, and construct a structured heterogeneous common knowledge graph from the heterogeneous common information.

[0098] Specifically, step S30 in the method includes:

[0099] Extract a structured list of heterogeneous approximation content and a summary of heterogeneous approximation results from the set of heterogeneous approximation tasks;

[0100] The list of heterogeneous approximations and the summary of heterogeneous approximations are input into the multimodal large model, and the structured heterogeneous common information is obtained by combining the preset common prompts.

[0101] The heterogeneous common information is constructed into a micro-knowledge graph oriented towards the needs of target users, and the output is the heterogeneous common knowledge graph;

[0102] The heterogeneous common knowledge graph uses content entities, attribute entities, and constraint entities as nodes and entity relationships as edges.

[0103] In this embodiment, a structured list of heterogeneous approximation content and a summary of heterogeneous approximation results are first extracted from the heterogeneous approximation task set. The list of heterogeneous approximation content refers to a structured requirement list for each historical task in the heterogeneous approximation task set, containing information such as the generation objective, modal path, and generation criteria of the historical task. The summary of heterogeneous approximation results is a summary of the generation results of each historical task, which may include, for example, the core features of the generation results, quality assessment results, and key optimization points.

[0104] For example, each historical task record in the heterogeneous approximation task set is traversed, and a list of heterogeneous approximation content is directly extracted from the structured storage fields. Simultaneously, the generation results of the historical tasks are analyzed using a multimodal large model to extract core features and key information, generating a summary of the heterogeneous approximation results. For instance, for a historical task that generates a 3D character model based on speech and text input, its list of heterogeneous approximation content includes: generation target (3D character model), input modality combination (speech, text), generation path (speech, text semantics, character model generation), and generation criteria (character similarity ≥ 85%). The summary of the heterogeneous approximation results includes: the character is a male human figure, 180cm tall, with a casual clothing style, a model accuracy of 88%, and optimization points including facial feature detail enhancement.

[0105] Secondly, the list of heterogeneous approximation content and the summary of heterogeneous approximation results are input into the multimodal large model. Combined with preset common prompts, structured heterogeneous common information is obtained. The preset common prompts refer to pre-defined instructions used to guide the multimodal large model in extracting heterogeneous common information; for example, they may include extraction targets, extraction dimensions, and structured format requirements. The structured heterogeneous common information refers to core information shared by different historical tasks, organized according to a preset format.

[0106] For example, the preset common prompt instruction is: Please extract the common generation target features, modality processing core logic, generation standard key indicators, and result optimization general methods from the input list of heterogeneous approximation content and the summary of heterogeneous approximation results, and output the structured results according to the category and specific content format.

[0107] For example, the list of heterogeneous approximate content, the summary of heterogeneous approximate results, and the preset common prompts are input into a multimodal large model. By analyzing the common features of different historical tasks, structured heterogeneous common information is output. For example, the heterogeneous common information extracted from the list of heterogeneous approximate content and the summary of heterogeneous approximate results may include: the generated target features are all 3D visualized content, and the core logic of modal processing is mainly based on text semantics, etc.

[0108] Finally, the heterogeneous common information is constructed into a micro-knowledge graph oriented towards the needs of target users, and the output is a heterogeneous common knowledge graph. This heterogeneous common knowledge graph uses content entities, attribute entities, and constraint entities as nodes, and entity relationships as edges. Specifically, content entities refer to the content objects involved in the heterogeneous common information, such as 3D models, generation targets, and optimization methods; attribute entities refer to information describing the attributes of content entities, such as accuracy ≥85% and text semantics as the core; constraint entities refer to the constraints on the content generation process, such as modal fusion constraints and quality constraints; and entity relationships refer to the associations between content entities, attribute entities, and constraint entities, such as inclusion and constraint.

[0109] For example, a graph database, such as Neo4j, can be used to construct a heterogeneous common knowledge graph. Content entities, attribute entities, and constraint entities are used as nodes, and the type and direction of edges are defined according to their inherent logical relationships to form a structured knowledge graph. For instance, 3D visualization content can be used as content entities, and attributes with an accuracy of ≥85% can be used as attribute entities, with the two connected by an attribute relationship; modal fusion constraints can be used as constraint entities, connected to the 3D visualization content by constraint relationships; and multimodal feature fusion optimization algorithms can be used as content entities, connected to the 3D visualization content by optimization relationships, ultimately forming a complete heterogeneous common knowledge graph.

[0110] Furthermore, step S30 of the method also includes:

[0111] The heterogeneous common knowledge graph is hierarchically and hierarchically labeled, wherein the hierarchical and hierarchical labeling results include:

[0112] A default content knowledge class, wherein the default content knowledge class has a soft target label, and the soft target label is determined according to the intrinsic importance index of the default content knowledge class;

[0113] Implicit logical knowledge class, wherein the implicit logical knowledge class has a hard target label.

[0114] In this embodiment of the application, the heterogeneous common knowledge graph is first hierarchically labeled. The hierarchical labeling results include: default content knowledge class, which has a soft target label, and the soft target label is determined according to the intrinsic importance index of the default content knowledge class; and implicit logic knowledge class, which has a hard target label.

[0115] Among them, the default content knowledge class refers to the knowledge category in the heterogeneous common knowledge graph used to supplement the content information that may be missing in the current requirement, and its impact on the generation result has a certain degree of flexibility; the implicit logic knowledge class refers to the core logic and rule-based knowledge hidden behind the generation process in the heterogeneous common knowledge graph, and it plays a decisive role in the correctness and rationality of the generation result; the soft target label refers to the label that marks the importance of the default content knowledge class, and is used to indicate the reuse priority of this type of knowledge; the hard target label refers to the label that forcibly marks the implicit logic knowledge class, indicating that this type of knowledge must be reused in the generation process; the intrinsic importance index is a quantitative indicator that measures the importance of the default content knowledge class to the current requirement, and its calculation basis includes the correlation between the knowledge and the current requirement, the frequency of knowledge reuse in historical tasks, and the degree of influence of the knowledge on the generation quality.

[0116] Specifically, by analyzing each node and edge in the heterogeneous common knowledge graph using a multimodal large model, the knowledge category to which it belongs is determined: if a knowledge node is used to supplement content details, such as the choice of character clothing color or the type of scene decoration elements, it is classified as the default content knowledge category, its intrinsic importance index is calculated, such as 0.3, and a corresponding soft target label is assigned, such as an intrinsic importance index of 0.7-0.8 is marked as high priority reuse, 0.5-0.6 is marked as medium priority reuse, and 0.3-0.4 is marked as low priority reuse; if a knowledge node is the core logic of generation, such as multimodal semantic alignment rules, generation path optimization algorithms, core formulas for quality assessment, etc., it is classified as the implicit logic knowledge category and a hard target label is assigned: must be reused.

[0117] For example, nodes of the scene decoration element type in the heterogeneous common knowledge graph are classified as default content knowledge class, and the intrinsic importance index is calculated to be 0.5. Soft targets are marked as medium priority reuse. At the same time, nodes of multimodal semantic alignment rules are classified as implicit logic knowledge class, and hard targets are marked as: must be reused.

[0118] In summary, compared to existing technologies, this application uses a multimodal large model to analyze the heterogeneous approximation task set, extracts heterogeneous common information, and constructs this heterogeneous common information into a structured heterogeneous common knowledge graph. This achieves accurate extraction of cross-modal common information from the heterogeneous approximation task set, and the structured heterogeneous common knowledge graph eliminates the knowledge reuse barriers caused by heterogeneous modal differences, providing efficient and readily available knowledge support for subsequent metaverse content generation, and also improving the rationality and consistency of content generation.

[0119] S40: Based on the initial structured content list and the heterogeneous common knowledge graph, combine the multimodal large model to generate and merge metaverse content and output it.

[0120] In the current metaverse content generation, relying solely on the initial structured content list can easily lead to a lack of historical common experience to support the generated results and insufficient quality stability. Relying solely on scattered heterogeneous common information makes it difficult to accurately match current personalized needs. Furthermore, after multimodal content generation, there are often problems such as inconsistent modal fusion and inconsistent output formats.

[0121] To address the aforementioned issues, this application uses the initial structured content list and the heterogeneous common knowledge graph, combined with a multimodal large model, to generate and merge metaverse content for output.

[0122] Specifically, step S40 in the method includes:

[0123] By comparing the initial structured content list with the heterogeneous common knowledge graph, content differences and missing content items are identified;

[0124] Based on the heterogeneous common information of the heterogeneous common knowledge graph, the missing content items are filled in, and based on the heterogeneous common information of the heterogeneous common knowledge graph and combined with the large model parameters preset by the target user, the content difference items are randomly corrected to generate an enhanced structured content list.

[0125] Based on the enhanced structured content list, the multimodal large model is scheduled to generate content and output metaverse content.

[0126] In this embodiment, the initial structured content list is first compared with the heterogeneous common knowledge graph to identify content discrepancies and missing content items. Content discrepancies refer to inconsistencies between the requirement information in the initial structured content list and the common information in the heterogeneous common knowledge graph; missing content items refer to common information not explicitly mentioned in the initial structured content list but present in the heterogeneous common knowledge graph and significantly impacting the generated results.

[0127] Specifically, the core requirements, generation standards, and modal paths of the initial structured content list are compared one by one with the content entities, attribute entities, and constraint entities in the heterogeneous common knowledge graph through a multimodal large model. If the generation standards set in the initial structured content list are inconsistent with the common generation standards in the heterogeneous common knowledge graph, then the generation standard is a content difference item. If the initial structured content list does not mention the multimodal feature fusion optimization algorithm, but the algorithm is an implicit logical knowledge class in the heterogeneous common knowledge graph (i.e., it must be reused), then the relevant information of the algorithm is a content missing item.

[0128] Secondly, based on the heterogeneous common information of the heterogeneous common knowledge graph, missing content items are filled in. Then, based on this information and the target user's pre-set large model parameters, content discrepancies are randomly corrected to generate an enhanced structured content list. The target user's pre-set large model parameters refer to the generation parameters of the multimodal large model, such as generation temperature, iteration count, and feature weights, set by the target user according to their needs. The enhanced structured content list refers to the optimized structured requirement list after filling in missing content items and correcting content discrepancies.

[0129] Specifically, for missing content items, corresponding heterogeneous common information is directly extracted from the heterogeneous common knowledge graph to fill in the missing items; for content differences, random corrections are made by combining heterogeneous common information and the target user's preset large model parameters. During the correction process, it is ensured that the corrected information conforms to common rules and meets the user's personalized needs.

[0130] For example, if the accuracy of the 3D model in the initial structured content list is 80% and the common accuracy of the heterogeneous common knowledge graph is 85%, the accuracy of the 3D model can be identified as a content difference item. If the user presets the large model generation temperature to 0.7, the accuracy of the randomly corrected 3D model may be 85% or 83%.

[0131] Finally, based on the enhanced structured content list, the multimodal large model is scheduled to perform content generation, outputting metaverse content. Specifically, according to information such as the initial generation modal path, generation criteria, filled missing items, and corrected discrepancies in the enhanced structured content list, the multimodal large model is scheduled to perform content generation. For example, each modal data is processed sequentially according to the modal path, implicit logical knowledge classes in the heterogeneous common knowledge graph are reused, and soft target tags of default content knowledge classes are referenced for priority reuse. During the generation process, quality is monitored in real time according to the generation criteria. If the generation result does not meet the criteria, the parameters of the large model are adjusted, such as increasing the number of iterations or adjusting feature weights. Finally, the multimodal content generated in each step is fused, spliced, and optimized to ensure the consistency and coherence of the content, ultimately outputting complete metaverse content.

[0132] In summary, compared to existing technologies, this application utilizes the initial structured content list and the heterogeneous common knowledge graph, combined with a multimodal large model, to generate and merge metaverse content. This deeply integrates the personalized needs of the initial structured content list with the common knowledge of the heterogeneous common knowledge graph, efficiently completing metaverse content generation and multimodal output merging through the multimodal large model. This ensures accurate matching between generated content and user needs while improving the quality, efficiency, and coherence and consistency of content generation and modal fusion.

[0133] In summary, the embodiments of this application have at least the following technical effects:

[0134] Compared to existing technologies, this application first responds to the target user's original multimodal input, and then uses a large multimodal model to analyze requirements and generate an initial structured content list. This achieves cross-modal semantic alignment and accurate extraction of core requirements from the original multimodal input. The generated initial structured content list provides a clear and unified basis for subsequent metaverse content generation, improving the accuracy of requirements analysis and the relevance of the subsequent generation process.

[0135] Secondly, based on the initial structured content list, this application performs a search and matching operation in the historical task experience database to obtain a heterogeneous set of approximate tasks. The heterogeneous set of approximate tasks has a different generation modal path configuration from the initial structured content list. Thus, by obtaining a set of heterogeneous approximate tasks with a different generation modal path configuration through search and matching, the scope of reuse of historical task experience is broadened, cross-modal common knowledge is effectively mined, and rich data support is provided for subsequent heterogeneous common information extraction and graph construction. At the same time, homogenization of content generation is avoided, improving generation efficiency and quality.

[0136] Furthermore, this application uses a multimodal large model to analyze the heterogeneous approximation task set, extracts heterogeneous common information, and constructs this heterogeneous common information into a structured heterogeneous common knowledge graph. In this way, accurate extraction of cross-modal common information from the heterogeneous approximation task set is achieved. The structured heterogeneous common knowledge graph eliminates the knowledge reuse barriers caused by heterogeneous modal differences, providing efficient and callable knowledge support for subsequent metaverse content generation, and also improving the rationality and consistency of content generation.

[0137] Finally, based on the initial structured content list and the heterogeneous common knowledge graph, this application combines a multimodal large model to generate and merge metaverse content. In this way, the personalized requirements of the initial structured content list are deeply integrated with the common knowledge of the heterogeneous common knowledge graph. The multimodal large model efficiently completes the generation of metaverse content and the merging of multimodal outputs, ensuring both accurate matching of generated content with user needs and improving the quality, efficiency, and coherence and consistency of content generation and modal fusion.

[0138] Through the above technical solutions, this application achieves efficient adaptation to multimodal fusion input, improves the consistency between generated content and users' real needs, and reduces repetitive development costs by effectively reusing cross-modal historical task experience, thereby improving the generation efficiency and quality stability of metaverse content, while ensuring the consistency and reliability of the generated results.

[0139] Example 2, as Figure 2 As shown, based on the same inventive concept as the metaverse content generation method based on a multimodal large model provided in Embodiment 1, this embodiment of the invention also provides a metaverse content generation platform based on a multimodal large model, including:

[0140] The initial list generation module 11 is used to respond to the target user's original multimodal input, analyze the requirements through the multimodal big model, and generate an initial structured content list.

[0141] The retrieval and matching module 12 is used to perform retrieval and matching in the historical task experience base based on the initial structured content list to obtain a heterogeneous approximate task set, wherein the heterogeneous approximate task set and the generation modal path of the initial structured content list are configured differently.

[0142] The graph construction module 13 is used to parse the heterogeneous approximate task set through a multimodal large model, extract heterogeneous common information, and construct the heterogeneous common information into a structured heterogeneous common knowledge graph.

[0143] The merged output module 14 is used to generate and merge metaverse content based on the initial structured content list and the heterogeneous common knowledge graph, combined with the multimodal large model.

[0144] The initial inventory generation module 11 is specifically used for:

[0145] The original multimodal input is semantically parsed using multimodal aligned coding, and the semantic parsing results are mapped to a unified multimodal semantic space to generate a multimodal semantic representation.

[0146] Based on the content complexity of multimodal semantic representation, we perform modal path planning and modal standard initialization to obtain the initial modal path and initial modal standard.

[0147] By combining the preset data paradigm, the multimodal semantic representation, the initial generated modal path, and the initial generated standard are structured and output to obtain the initial structured content list.

[0148] The retrieval and matching module 12 is specifically used for:

[0149] Define the equivalent vacancy rate rule and the corresponding equivalent vacancy rate constraint;

[0150] Based on the vacancy rate constraint, the multimodal semantic representation of the initial structured content list is randomly pruned to obtain several multimodal semantic index vectors that satisfy the vacancy rate constraint.

[0151] Based on several multimodal semantic index vectors, iterative vector similarity retrieval is performed by traversing the historical task experience base to extract historical task records that meet preset similarity constraints, and the output is the heterogeneous approximate task set.

[0152] Specifically, the step of "iteratively searching vector similarity based on several multimodal semantic index vectors, traversing the historical task experience base, extracting historical task records that satisfy preset similarity constraints, and outputting the heterogeneous approximate task set" includes:

[0153] Extract the multimodal semantic history vector of each historical task record;

[0154] Based on the multimodal semantic history vector and the multimodal semantic index vector, a vector similarity retrieval is performed to obtain the first approximate task set;

[0155] The historical task experience base is traversed, and the cross-modal semantic history vector of each historical task record is obtained through multimodal large model analysis.

[0156] Cross-modal association retrieval is performed based on the cross-modal semantic history vector and the multimodal semantic index vector to obtain a second approximate task set;

[0157] Based on the first approximate task set and the second approximate task set, heterogeneity screening for generating modal paths is performed to obtain the heterogeneous approximate task set.

[0158] Specifically, the step of "performing heterogeneity screening for generative modal paths based on the first approximate task set and the second approximate task set to obtain the heterogeneous approximate task set" includes:

[0159] Traverse the first approximate task set and the second approximate task set to extract contrast modal information, wherein the contrast modal information includes input modal combination, output modal combination and modal sequence;

[0160] The comparison modal information is parameterized and encoded by an encoder to obtain the comparison generation modal path;

[0161] Based on the initial structured content list, the initial generated modal path is extracted and parameterized accordingly to obtain the reference generated modal path;

[0162] The parameterized residuals of multiple contrast generated modal paths and the reference generated modal path are calculated iteratively, and the selection probabilities of the first approximate task set and the second approximate task set are defined based on the parameterized residuals, wherein the parameterized residuals are negatively correlated with the selection probabilities;

[0163] Based on the selection probability, the first approximate task set and the second approximate task set are traversed to perform heterogeneity screening based on iterative probability selection, thereby obtaining the heterogeneous approximate task set.

[0164] Furthermore, the "defining equivalent vacancy rate rules and corresponding equivalent vacancy rate constraints" includes:

[0165] Knowledge is generated by combining prior content, and the intrinsic importance index of each dimension in the multimodal semantic index vector is defined.

[0166] Normalize the intrinsic importance index of multiple dimensions to obtain the dimensional void cost for each dimension;

[0167] The sum of the dimensional vacancy costs of all vacant dimensions is defined as the equivalent vacancy rate, and an equivalent vacancy rate less than or equal to a preset equivalent vacancy rate threshold is defined as the equivalent vacancy rate constraint.

[0168] The equivalent vacancy rate threshold is a significance threshold determined based on statistical analysis methods.

[0169] Specifically, the map construction module 13 is used for:

[0170] Extract a structured list of heterogeneous approximation content and a summary of heterogeneous approximation results from the set of heterogeneous approximation tasks;

[0171] The list of heterogeneous approximations and the summary of heterogeneous approximations are input into the multimodal large model, and the structured heterogeneous common information is obtained by combining the preset common prompts.

[0172] The heterogeneous common information is constructed into a micro-knowledge graph oriented towards the needs of target users, and the output is the heterogeneous common knowledge graph;

[0173] The heterogeneous common knowledge graph uses content entities, attribute entities, and constraint entities as nodes and entity relationships as edges.

[0174] Furthermore, the map construction module 13 is also specifically used for:

[0175] The heterogeneous common knowledge graph is hierarchically and hierarchically labeled, wherein the hierarchical and hierarchical labeling results include:

[0176] A default content knowledge class, wherein the default content knowledge class has a soft target label, and the soft target label is determined according to the intrinsic importance index of the default content knowledge class;

[0177] Implicit logical knowledge class, wherein the implicit logical knowledge class has a hard target label.

[0178] Specifically, the merging output module 14 is used for:

[0179] By comparing the initial structured content list with the heterogeneous common knowledge graph, content differences and missing content items are identified;

[0180] Based on the heterogeneous common information of the heterogeneous common knowledge graph, the missing content items are filled in, and based on the heterogeneous common information of the heterogeneous common knowledge graph and combined with the large model parameters preset by the target user, the content difference items are randomly corrected to generate an enhanced structured content list.

[0181] Based on the enhanced structured content list, the multimodal large model is scheduled to generate content and output metaverse content.

[0182] In summary, the embodiments of this application have at least the following technical effects:

[0183] Compared to existing technologies, this application firstly utilizes an initial list generation module. Responding to the target user's original multimodal input, it analyzes requirements using a large multimodal model to generate an initial structured content list, achieving cross-modal semantic alignment and accurate extraction of core requirements from the original multimodal input. Secondly, through a retrieval and matching module, based on the initial structured content list, it performs retrieval and matching in a historical task experience database to obtain a set of heterogeneous similar tasks, broadening the scope of reuse of historical task experience and effectively mining cross-modal common knowledge. Thirdly, through a graph construction module, it analyzes the set of heterogeneous similar tasks using a large multimodal model, extracts heterogeneous common information, and constructs this heterogeneous common information into a structured heterogeneous common knowledge graph. This achieves accurate extraction of cross-modal common information from the set of heterogeneous similar tasks, eliminating knowledge reuse barriers caused by heterogeneous modal differences and providing efficient and callable knowledge support for metaverse content generation. Finally, through the output merging module, based on the initial structured content list and heterogeneous common knowledge graph, and combined with the multimodal large model, metaverse content is generated and merged for output. This ensures accurate matching between the generated content and user needs, while also improving the quality, efficiency, and coherence and consistency of content generation and modal fusion. In this way, efficient adaptation to multimodal fusion input is achieved, improving the relevance of generated content to real user needs. Furthermore, by effectively reusing cross-modal historical task experience, repetitive development costs are reduced, improving the generation efficiency and quality stability of metaverse content, while ensuring the consistency and reliability of the generated results.

[0184] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0185] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0186] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0187] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0188] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0189] Although preferred embodiments of the invention have been described, those skilled in the art, once they have learned the basic inventive concept, can make other changes and modifications to these embodiments.

[0190] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for generating metaverse content based on a multimodal large model, characterized in that, include: In response to the target user's original multimodal input, the requirements are analyzed through a large multimodal model to generate an initial structured content list; Based on the initial structured content list, a search and matching process is performed in the historical task experience base to obtain a heterogeneous set of similar tasks, wherein the heterogeneous set of similar tasks is configured differently from the generation modal path of the initial structured content list. The heterogeneous approximate task set is analyzed by multimodal large model, heterogeneous common information is extracted, and the heterogeneous common information is constructed into a structured heterogeneous common knowledge graph. Based on the initial structured content list and the heterogeneous common knowledge graph, the metaverse content is generated and merged and output using a multimodal large model; In response to the target user's original multimodal input, the first large-scale language model is used to parse the requirements and generate an initial structured content list, including: The original multimodal input is semantically parsed using multimodal aligned coding, and the semantic parsing result is mapped to a unified multimodal semantic space to generate a multimodal semantic representation. Based on the content complexity of multimodal semantic representation, we perform modal path planning and modal standard initialization to obtain the initial modal path and initial modal standard. By combining the preset data paradigm, the multimodal semantic representation, the initial generated modal path, and the initial generated standard are output in a structured manner to obtain the initial structured content list; Specifically, based on the initial structured content list, a search and matching process is performed in the historical task experience base to obtain a heterogeneous set of similar tasks. The heterogeneous set of similar tasks and the generation modal path of the initial structured content list are configured differently, including: Define the equivalent vacancy rate rule and the corresponding equivalent vacancy rate constraint; Based on the vacancy rate constraint, the multimodal semantic representation of the initial structured content list is randomly pruned to obtain several multimodal semantic index vectors that satisfy the vacancy rate constraint. Based on several multimodal semantic index vectors, iterative vector similarity retrieval is performed by traversing the historical task experience base to extract historical task records that meet preset similarity constraints, and the output is the heterogeneous approximate task set.

2. The method for generating metaverse content based on a multimodal large model as described in claim 1, characterized in that, Based on several multimodal semantic index vectors, iterative vector similarity retrieval is performed by traversing the historical task experience base to extract historical task records that satisfy preset similarity constraints, and the output is the heterogeneous approximate task set, including: Extract the multimodal semantic history vector of each historical task record; Based on the multimodal semantic history vector and the multimodal semantic index vector, a vector similarity retrieval is performed to obtain the first approximate task set; The historical task experience base is traversed, and the cross-modal semantic history vector of each historical task record is obtained through multimodal large model analysis. Cross-modal association retrieval is performed based on the cross-modal semantic history vector and the multimodal semantic index vector to obtain a second approximate task set; Based on the first approximate task set and the second approximate task set, heterogeneity screening for generating modal paths is performed to obtain the heterogeneous approximate task set.

3. The method for generating metaverse content based on a multimodal large model as described in claim 2, characterized in that, Based on the first approximate task set and the second approximate task set, heterogeneity filtering for generating modal paths is performed to obtain the heterogeneous approximate task set, including: Traverse the first approximate task set and the second approximate task set to extract contrast modal information, wherein the contrast modal information includes input modal combination, output modal combination and modal sequence; The comparison modal information is parameterized and encoded by an encoder to obtain the comparison generation modal path; Based on the initial structured content list, the initial generated modal path is extracted and parameterized accordingly to obtain the reference generated modal path; The parameterized residuals of multiple contrast generated modal paths and the reference generated modal path are calculated iteratively, and the selection probabilities of the first approximate task set and the second approximate task set are defined based on the parameterized residuals, wherein the parameterized residuals are negatively correlated with the selection probabilities; Based on the selection probability, the first approximate task set and the second approximate task set are traversed to perform heterogeneity screening based on iterative probability selection, thereby obtaining the heterogeneous approximate task set.

4. The metaverse content generation method based on a multimodal large model as described in claim 1, characterized in that, The heterogeneous approximation task set is analyzed using a multimodal large model to extract heterogeneous common information, and this heterogeneous common information is constructed into a structured heterogeneous common knowledge graph, including: Extract a structured list of heterogeneous approximation content and a summary of heterogeneous approximation results from the set of heterogeneous approximation tasks; The list of heterogeneous approximations and the summary of heterogeneous approximations are input into the multimodal large model, and the structured heterogeneous common information is obtained by combining the preset common prompts. The heterogeneous common information is constructed into a micro-knowledge graph oriented towards the needs of target users, and the output is the heterogeneous common knowledge graph; The heterogeneous common knowledge graph uses content entities, attribute entities, and constraint entities as nodes and entity relationships as edges.

5. The metaverse content generation method based on a multimodal large model as described in claim 1, characterized in that, The heterogeneous approximation task set is analyzed using a multimodal large model to extract heterogeneous common information, and this heterogeneous common information is then constructed into a structured heterogeneous common knowledge graph. The method also includes: The heterogeneous common knowledge graph is hierarchically and hierarchically labeled, wherein the hierarchical and hierarchical labeling results include: A default content knowledge class, wherein the default content knowledge class has a soft target label, and the soft target label is determined according to the intrinsic importance index of the default content knowledge class; Implicit logical knowledge class, wherein the implicit logical knowledge class has a hard target label.

6. The method for generating metaverse content based on a multimodal large model as described in claim 1, characterized in that, Based on the initial structured content list and the heterogeneous common knowledge graph, a metaverse content generation and merging output is performed using a multimodal large model, including: By comparing the initial structured content list with the heterogeneous common knowledge graph, content differences and missing content items are identified; Based on the heterogeneous common information of the heterogeneous common knowledge graph, the missing content items are filled in, and based on the heterogeneous common information of the heterogeneous common knowledge graph and combined with the large model parameters preset by the target user, the content difference items are randomly corrected to generate an enhanced structured content list. Based on the enhanced structured content list, the multimodal large model is scheduled to generate content and output metaverse content.

7. The method for generating metaverse content based on a multimodal large model as described in claim 1, characterized in that, Define the equivalent vacancy rate rules and corresponding equivalent vacancy rate constraints, including: Knowledge is generated by combining prior content, and the intrinsic importance index of each dimension in the multimodal semantic index vector is defined. Normalize the intrinsic importance index of multiple dimensions to obtain the dimensional void cost for each dimension; The sum of the dimensional vacancy costs of all vacant dimensions is defined as the equivalent vacancy rate, and an equivalent vacancy rate less than or equal to a preset equivalent vacancy rate threshold is defined as the equivalent vacancy rate constraint. The equivalent vacancy rate threshold is a significance threshold determined based on statistical analysis methods.

8. A metaverse content generation platform based on a multimodal large model, characterized in that: The method for generating metaverse content based on a multimodal large model as described in any one of claims 1-7 includes: The initial inventory generation module is used to respond to the target user's original multimodal input, analyze the requirements through a multimodal big model, and generate an initial structured content inventory. The retrieval and matching module is used to perform retrieval and matching in the historical task experience base based on the initial structured content list to obtain a heterogeneous set of similar tasks, wherein the heterogeneous set of similar tasks is configured differently from the generation modal path of the initial structured content list. The graph construction module is used to parse the heterogeneous approximate task set through a multimodal large model, extract heterogeneous common information, and construct the heterogeneous common information into a structured heterogeneous common knowledge graph. The merged output module is used to generate and merge metaverse content based on the initial structured content list and the heterogeneous common knowledge graph, combined with the multimodal large model.