Functional scene modeling and large model cross-dimension content generation method and system
By abstracting scene entities into programmable function units and constructing a 256-dimensional shared embedding space, the problems of low modeling efficiency, modal semantic fragmentation and insufficient dynamic response in existing technologies are solved, real-time dynamic adjustment and strict alignment of multimodal outputs are achieved, and modeling efficiency and consistency are improved.
Patent Information
- Application Number
- CN202510816286.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies have low modeling efficiency in multimodal scene modeling. Users need to spend a lot of time manually configuring parameters, the cost of modifying interaction logic is high, modal semantics are fragmented, cross-modal generation results are inconsistent, and the lack of dynamic response mechanism leads to scene rigidity.
The scene entities are abstracted into programmable function units, and real-time dynamic adjustment is supported through the internal parameter interface. A 256-dimensional shared embedding space is constructed, and a cross-modal feature alignment algorithm and an attention-weighted fusion layer are used to integrate multimodal feature vectors. The output is optimized in combination with a reinforcement learning strategy.
Real-time dynamic adjustment of the scene model is achieved, the cross-modal error rate is reduced, and multi-modal outputs are strictly aligned, thereby improving modeling efficiency and accuracy.
Smart Images

Figure CN120676222A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video generation technology, and in particular to a method and system for functional scene modeling and large-model cross-dimensional content generation. Background Art
[0002] The rapid development of the metaverse, virtual reality, and gaming industries has created an urgent demand for multimodal scene modeling and content generation technologies. Mainstream solutions in the industry rely primarily on two technical systems. Traditional 3D modeling tools like Blender and Unity, while capable of detailed scene construction, rely heavily on manual work by professional artists, and dynamic interaction logic still requires hard-coding via C++ or Python scripts. Cross-modal generation models like CLIP and DALL·E, while achieving breakthroughs in text-to-image generation, perform poorly in the collaborative control of multiple objects in complex scenes and lack a response mechanism for dynamic parameters.
[0003] Existing technologies face the following problems: low modeling efficiency. When users describe the scene using natural language, they still need to spend several hours manually configuring fluid simulation parameters. Modifying the interactive logic is costly due to reliance on hard coding; modal semantics are fragmented, cross-modal generation results are seriously inconsistent, and multimodal outputs are difficult to synchronize on the timeline, requiring additional manual alignment; the lack of a dynamic response mechanism leads to scene rigidity, and user feedback data cannot be fed back into parameter optimization in real time, forcing the system to rely on offline rendering iterations.
[0004] In response to the above problems, the present invention proposes a method and system for functional scene modeling and large-scale model cross-dimensional content generation to make up for the shortcomings of the existing technology. Summary of the Invention
[0005] In response to the above-mentioned shortcomings of the existing technology, the present invention provides a method and system for functional scene modeling and large-model cross-dimensional content generation, which can effectively solve the problems mentioned in the existing technology.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions:
[0007] The present invention provides a method for functional scene modeling and large-scale model cross-dimensional content generation, which specifically includes the following steps:
[0008] S100: Receive multimodal scene description data input by a user, wherein the scene description data includes at least one modality of text data, image data, and audio data, wherein the text data includes a semantic description field, a style preference field, and an interaction logic constraint field;
[0009] S101: Parse the multimodal scene description data to extract scene object entities, object attribute parameters, and object interaction logic, wherein:
[0010] The scene object entity is located by an object type identifier and is associated with a geometric feature parameter set, a material attribute parameter set, and a behavior rule set;
[0011] The object attribute parameters are divided into a static parameter set and a dynamic parameter set. The static parameter set includes color value and size value, and the dynamic parameter set includes light intensity time function and ambient noise spectrum function.
[0012] The object interaction logic is defined as condition triggering rules, loop execution rules and event response rules;
[0013] S102: Construct a dynamically adjustable scene model, instantiate the scene object entity into functional units, each functional unit including an input interface, an output interface, and an internal parameter configuration interface;
[0014] Generate a dependency topology graph between functional units based on the interaction logic, wherein the topology graph is a directed acyclic graph structure, in which nodes are functional units and edges are data flow paths;
[0015] According to the real-time adjustment instruction input by the user, the scene model state is updated by modifying the value of the internal parameter configuration interface;
[0016] S103: Calling a pre-trained large language model or a multimodal generative model to perform cross-dimensional content generation, mapping the functionalized unit parameters in the scene model to the generation request of the target modality, wherein the mapping is achieved by a cross-modal feature alignment algorithm that calculates the cosine similarity between the function parameter set and the input features of the generative model;
[0017] An attention-weighted fusion layer is used to integrate multimodal feature vectors and output collaborative content of at least two modalities among text, image, audio or 3D model;
[0018] S104: Dynamically optimize output based on user feedback data, where the feedback data includes explicit feedback data and implicit feedback data. The explicit feedback data includes parameter modification instructions and quality ratings, and the implicit feedback data includes user operation timestamps and interface interaction traces.
[0019] Adjust the internal parameter configuration interface value of the functional unit through reinforcement learning strategy, and update the temperature coefficient and sampling probability of the large language model;
[0020] S105: Output the optimized cross-modal content to the application terminal.
[0021] Furthermore, the specific steps of parsing the multimodal scene description data and extracting scene object entities, object attribute parameters and interaction logic between objects include:
[0022] S200, performing semantic dependency analysis and named entity recognition on text data to generate a scene object entity set and associated attribute keywords;
[0023] S201, using an instance segmentation model to extract visual object boundaries from image data, and using an acoustic event detection model to identify sound source objects from audio data, and aligning them with the scene object entity set;
[0024] S202, parsing static parameters into a key-value pair data structure, and parsing dynamic parameters into a time function expression;
[0025] S203: Encode the interaction logic into a triple structure, where the triple includes an event trigger condition, an execution constraint condition, and an action instruction;
[0026] S204: Output a standardized scene description structure, including:
[0027] The geometric feature parameter set and material attribute parameter set associated with the object type identifier;
[0028] Static parameter sets and dynamic parameter sets;
[0029] A behavioral rule set consisting of conditional triggering rules, loop execution rules, and event response rules.
[0030] Furthermore, the time function expression in step S202 includes a light intensity time function and an ambient noise spectrum function;
[0031] The formula of the light intensity time function is:
[0032]
[0033] in, is the light intensity value at time t, The maximum intensity value set by the user, is a preset periodic function, is the angular frequency parameter, is the phase offset, is the environmental attenuation factor, Is the basic light intensity constant;
[0034] The formula of the ambient noise spectrum function is:
[0035]
[0036] in, is the noise energy density at frequency f, is the number of sound sources, is the amplitude weight of the kth sound source, is the center frequency of the kth sound source, is the frequency bandwidth parameter.
[0037] Furthermore, the updating of the dependency topology graph in step S102 includes the following operations:
[0038] When a user operation triggers an event response rule, the user operation is parsed to determine the data flow path to be modified and the modification type. The adjacency matrix storing the dependency relationships is updated. If a new path is added, the value of the corresponding matrix element is set to the preset default dependency weight; if a path is deleted, the value of the corresponding matrix element is set to zero. Finally, the updated topology is verified to ensure that it meets the directed acyclic property.
[0039] When parameters of multiple functional units conflict, the conflicting parameters and the associated functional unit identifiers are identified, and the final parameter value is selected according to the priority rule, wherein the priority rule is: user explicit instructions take precedence over real-time operation behavior; real-time operation behavior takes precedence over scene default values; finally, the selected parameter value is written to the internal parameter configuration interface of the target functional unit;
[0040] According to the updated content of the adjacency matrix, the range of the affected functionalized units is determined. The specific formula is:
[0041]
[0042] in, is the set of affected functionalized units, is a functional unit that is directly modified. For all functional units involved in this update, is the graph reachability determination function, is the updated adjacency matrix;
[0043] Then recalculate the state value of the affected functionalized unit. The specific formula is:
[0044]
[0045] in, Functionalized Unit The new state value of Functionalized Unit arrive The dependency weight of For upstream functional units The current state, is the activation function;
[0046] Record update logs to the version management database.
[0047] Furthermore, the cross-modal feature alignment algorithm execution process in step S103 includes:
[0048] Establish a shared embedding space from the function parameter set to the generative model input, with an embedding space dimension of 256;
[0049] The contrast loss function is used to constrain the Euclidean distance between the text description vector and the image feature vector in the shared space to be less than the threshold. .
[0050] Furthermore, the reinforcement learning strategy in step S104 specifically includes:
[0051] Define the reward function R, the specific formula is:
[0052]
[0053] in, Rate the user, To generate a content consistency score, 、 is the weight.
[0054] Functional scene modeling and large-scale cross-dimensional content generation system, including:
[0055] A multimodal data receiving module, configured to receive multimodal scene description data input by a user, wherein the scene description data includes at least one modality of text data, image data, and audio data, wherein the text data includes a semantic description field, a style preference field, and an interaction logic constraint field;
[0056] A scene data parsing module, configured to parse the multimodal scene description data and extract scene object entities, object attribute parameters, and interaction logic between objects;
[0057] A dynamic scene modeling module is used to construct a dynamically adjustable scene model, instantiating the scene object entity into functional units, each functional unit including an input interface, an output interface, and an internal parameter configuration interface; generating a dependency topology diagram between the functional units based on the interaction logic, wherein the topology diagram is a directed acyclic graph structure; and updating the scene model state by modifying the value of the internal parameter configuration interface according to the real-time adjustment instructions input by the user;
[0058] A cross-dimensional content generation module is used to call a pre-trained large language model or a multimodal generation model to perform cross-dimensional content generation, mapping the functionalized unit parameters in the scene model to the generation request of the target modality through a cross-modal feature alignment algorithm; using an attention-weighted fusion layer to integrate the multimodal feature vectors and output collaborative content of at least two modalities: text, image, audio, or a three-dimensional model;
[0059] A feedback optimization module is used to dynamically optimize the output based on user feedback data, including explicit feedback data and implicit feedback data; adjust the internal parameter configuration interface values of the functional unit through a reinforcement learning strategy, and update the temperature coefficient and sampling probability of the large language model;
[0060] The content output module is used to output optimized cross-modal content to the application terminal.
[0061] Furthermore, the scene data parsing module includes:
[0062] A text parsing unit, used to perform semantic dependency analysis and named entity recognition on text data to generate a set of scene object entities and associated attribute keywords;
[0063] An audio and video parsing unit, configured to extract visual object boundaries from image data using an instance segmentation model, and identify sound source objects from audio data using an acoustic event detection model, and align the sound source objects with the scene object entity set;
[0064] A parameter encoding unit is used to parse static parameters into a key-value pair data structure and parse dynamic parameters into a time function expression; encode the interaction logic into a triple structure, wherein the triple includes an event trigger condition, an execution constraint condition, and an action instruction;
[0065] The structure output unit is used to output a standardized scene description structure, including a geometric feature parameter set associated with an object type identifier, a material attribute parameter set, a static parameter set and a dynamic parameter set, as well as a behavioral rule set consisting of conditional triggering rules, loop execution rules and event response rules.
[0066] Furthermore, the system also includes a visual debugging module for displaying the node status and data flow path of the dependency topology diagram in real time, recording the parameter adjustment history and highlighting the conflict resolution process, and providing a side-by-side comparison preview window for multi-modal generation results.
[0067] Compared with the prior art, the technical solution provided by the present invention has the following beneficial effects:
[0068] 1. The present invention abstracts scene entities into programmable function units and supports real-time dynamic adjustment through internal parameter interfaces. When the user modifies the angular frequency parameter ω, the dependency topology graph can link and update the status of all associated units in a short period of time, avoiding the lag of traditional script coding.
[0069] 2. The present invention constructs a 256-dimensional shared embedding space and uses cosine similarity to calculate the matching degree between function parameters and generative model features, reducing the cross-modal error rate from 30% to below 8%. The attention-weighted fusion layer synchronously coordinates text and audio feature vectors to achieve strict alignment of multimodal outputs.
[0070] 3. The present invention integrates text, image, and audio data through semantic dependency analysis, instance segmentation and other technologies, extracts object entities, attribute parameters and interaction logic, forms a standardized scene description structure, and improves modeling efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.
[0072] Figure 1 Schematic diagram of the method flow of the present invention;
[0073] Figure 2 Schematic diagram of the system structure of the present invention;
[0074] Figure 3 Schematic diagram of the scene data analysis module structure of the present invention. DETAILED DESCRIPTION
[0075] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0076] The present invention will be further described below with reference to the embodiments.
[0077] Example 1: Reference Figure 1 , functional scene modeling and large model cross-dimensional content generation method, the specific steps include:
[0078] S100: Receive multimodal scene description data input by a user, where the scene description data includes at least one modality of text data, image data, and audio data, wherein the text data includes a semantic description field, a style preference field, and an interaction logic constraint field;
[0079] S101. Parse multimodal scene description data to extract scene object entities, object attribute parameters, and object interaction logic, where:
[0080] The scene object entity is located by the object type identifier and is associated with the geometric feature parameter set, material attribute parameter set and behavior rule set;
[0081] Object attribute parameters are divided into static parameter sets and dynamic parameter sets. The static parameter set includes color values and size values, and the dynamic parameter set includes light intensity time function and ambient noise spectrum function.
[0082] The interaction logic between objects is defined as condition triggering rules, loop execution rules and event response rules;
[0083] S102: Build a dynamically adjustable scene model, instantiate scene object entities into functional units, each functional unit including an input interface, an output interface, and an internal parameter configuration interface;
[0084] Generate a topological graph of dependencies between functional units based on interaction logic. The topological graph is a directed acyclic graph structure, where nodes are functional units and edges are data flow paths.
[0085] According to the real-time adjustment instructions input by the user, the scene model state is updated by modifying the value of the internal parameter configuration interface;
[0086] S103: Calling a pre-trained large language model or a multimodal generative model to perform cross-dimensional content generation, mapping the functionalized unit parameters in the scene model to the generation request of the target modality, and implementing the mapping through a cross-modal feature alignment algorithm that calculates the cosine similarity between the function parameter set and the input features of the generative model;
[0087] An attention-weighted fusion layer is used to integrate multimodal feature vectors and output collaborative content of at least two modalities among text, image, audio or 3D model;
[0088] S104: Dynamically optimize output based on user feedback data, where the feedback data includes explicit feedback data and implicit feedback data. The explicit feedback data includes parameter modification instructions and quality ratings, and the implicit feedback data includes user operation timestamps and interface interaction trajectories.
[0089] Adjust the internal parameter configuration interface value of the functional unit through reinforcement learning strategy, and update the temperature coefficient and sampling probability of the large language model;
[0090] S105: Output the optimized cross-modal content to the application terminal.
[0091] In a specific embodiment, assuming that the implementation scenario is virtual home design scene generation, the specific generation steps are:
[0092] Multimodal data reception and parsing: The user enters a text description of "modern minimalist living room, white sofa, floor-to-ceiling windows, natural light changes over time" and uploads a sketch image of the living room;
[0093] The multimodal data receiving module obtains text, including the semantic description "living room", the style preference "modern simplicity", the interaction logic "natural light changes over time", and image data;
[0094] Scene data parsing module: performs semantic dependency analysis on the text, identifies the object entities "sofa" and "floor-length window", and the attribute keywords "white" and "natural light". It uses instance segmentation to extract the visual boundaries of the sofa and window from the image and aligns them with the text object entities. It also parses the dynamic parameter "natural light changes over time" as a function of light intensity over time. ,in , , .
[0095] Dynamic scene model construction:
[0096] The dynamic scene modeling module instantiates objects such as "sofa" and "floor-length window" into functional units. Each unit contains a material parameter interface, such as the color value of the sofa, and a lighting parameter interface, such as the transmittance of the window.
[0097] Based on the interaction logic "natural light changes over time", a dependency topology diagram between functional units is generated: the light intensity function unit and the window transmittance unit are associated through the data flow path to form a directed acyclic graph.
[0098] The user adjusts the "sofa color" to "beige" in real time, updates the scene model by modifying the internal parameter configuration interface, synchronously adjusts the topology adjacency matrix, and verifies the acyclicity.
[0099] Cross-dimensional content generation:
[0100] The cross-dimensional content generation module calls a pre-trained multimodal generation model, such as the Stable Diffusion model, and maps function parameters, such as the RGB value of the sofa color and the light intensity function, to the generation request through a 256-dimensional shared embedding space.
[0101] The cross-modal feature alignment algorithm calculates the cosine similarity between the parameter vector and the model input features to ensure that the color and lighting parameters match the generated image; the attention weighted fusion layer integrates the text style features and the image visual features to output a three-dimensional model of the living room and the corresponding text description "beige sofa with floor-to-ceiling windows, the afternoon light changes softly over time."
[0102] Users rated the generated 3D model 4 stars. , and proposed a modification instruction to "increase the environmental noise effect".
[0103] Based on the reinforcement learning strategy, the feedback optimization module adds the ambient noise spectrum function to the scene model and adjusts the temperature coefficient of the large model to improve the generation diversity. The optimized cross-modal content is pushed to the user terminal through the content output module.
[0104] Furthermore, the specific steps of parsing the multimodal scene description data and extracting scene object entities, object attribute parameters, and the interaction logic between objects include:
[0105] S200, performing semantic dependency analysis and named entity recognition on text data to generate a scene object entity set and associated attribute keywords;
[0106] S201, using an instance segmentation model to extract visual object boundaries from image data, and using an acoustic event detection model to identify sound source objects from audio data, and aligning them with a scene object entity set;
[0107] S202, parsing static parameters into a key-value pair data structure, and parsing dynamic parameters into a time function expression;
[0108] S203, encoding the interaction logic into a triple structure, where the triple includes an event trigger condition, an execution constraint condition, and an action instruction;
[0109] S204: Output a standardized scene description structure, including:
[0110] The geometric feature parameter set and material attribute parameter set associated with the object type identifier;
[0111] Static parameter sets and dynamic parameter sets;
[0112] A behavioral rule set consisting of conditional triggering rules, loop execution rules, and event response rules.
[0113] Furthermore, the time function expression in step S202 includes a light intensity time function and an ambient noise spectrum function;
[0114] The formula for the light intensity-time function is:
[0115]
[0116] in, is the light intensity value at time t, The maximum intensity value set by the user, is a preset periodic function, is the angular frequency parameter, is the phase offset, is the environmental attenuation factor, Is the basic light intensity constant;
[0117] The formula for the ambient noise spectrum function is:
[0118]
[0119] in, is the noise energy density at frequency f, is the number of sound sources, is the amplitude weight of the kth sound source, is the center frequency of the kth sound source, is the frequency bandwidth parameter.
[0120] Furthermore, the updating of the dependency topology graph in step S102 includes the following operations:
[0121] When a user operation triggers an event response rule, the user operation is parsed to determine the data flow path to be modified and the modification type. The adjacency matrix storing the dependency relationships is updated. If a new path is added, the value of the corresponding matrix element is set to the preset default dependency weight; if a path is deleted, the value of the corresponding matrix element is set to zero. Finally, the updated topology is verified to ensure that it meets the directed acyclic property.
[0122] When the parameters of multiple functional units conflict, identify the conflicting parameters and the associated functional unit identifiers, and select the final parameter value according to the priority rule. The priority rule is: user explicit instructions take precedence over real-time operation behavior; real-time operation behavior takes precedence over scenario default values; finally, write the selected parameter value into the internal parameter configuration interface of the target functional unit;
[0123] According to the updated content of the adjacency matrix, the range of the affected functionalized units is determined. The specific formula is:
[0124]
[0125] in, is the set of affected functionalized units, is a functional unit that is directly modified. For all functional units involved in this update, is the graph reachability determination function, is the updated adjacency matrix;
[0126] Then recalculate the state value of the affected functionalized unit. The specific formula is:
[0127]
[0128] in, Functionalized Unit The new state value of Functionalized Unit arrive The dependency weight of For upstream functional units The current state, is the activation function;
[0129] Record update logs to the version management database.
[0130] Furthermore, the cross-modal feature alignment algorithm execution process in step S103 includes:
[0131] Establish a shared embedding space from the function parameter set to the generative model input, with an embedding space dimension of 256. Specifically, the embedding layer converts the feature vectors of different modalities into 256-dimensional vector representations, so that the features of different modalities can be compared and fused in the same space.
[0132] Encode the text description to obtain a text description vector, and extract the image features to obtain an image feature vector. The text description vector can be encoded using a pre-trained language model, and the image feature vector can be extracted using a pre-trained convolutional neural network.
[0133] Contrastive loss function constraint: Construct a contrastive loss function to calculate the Euclidean distance between the text description vector and the image feature vector in the shared space. By optimizing the contrastive loss function, the Euclidean distance between the text description vector and the image feature vector is made less than a preset threshold δ. The threshold δ can be set according to the actual application scenario and requirements to ensure the alignment effect between different modal features.
[0134] Furthermore, the reinforcement learning strategy in step S104 specifically includes:
[0135] Define the reward function R, the specific formula is:
[0136]
[0137] in, Rate the user, To generate a content consistency score, 、 is the weight.
[0138] Specifically, user ratings directly reflect user satisfaction with the generated content; the generated content consistency score measures the consistency between the generated content and the scene model. and It is used to balance the contribution of user rating and consistency score in the reward function and can be adjusted according to actual application requirements.
[0139] Example 2: Reference Figures 2 to 3 , functional scene modeling and large model cross-dimensional content generation system, including:
[0140] A multimodal data receiving module, configured to receive multimodal scene description data input by a user, the scene description data including at least one modality of text data, image data, and audio data, wherein the text data includes a semantic description field, a style preference field, and an interaction logic constraint field;
[0141] The scene data parsing module is used to parse multimodal scene description data and extract scene object entities, object attribute parameters and interaction logic between objects;
[0142] The dynamic scene modeling module is used to build a dynamically adjustable scene model. It instantiates scene object entities into functional units. Each functional unit contains an input interface, an output interface, and an internal parameter configuration interface. Based on the interaction logic, it generates a topological diagram of the dependency relationships between functional units. The topological diagram is a directed acyclic graph structure. Based on the real-time adjustment instructions input by the user, the scene model state is updated by modifying the value of the internal parameter configuration interface.
[0143] A cross-dimensional content generation module is used to call a pre-trained large language model or a multimodal generation model to perform cross-dimensional content generation. The module maps the functionalized unit parameters in the scene model to the generation request of the target modality. The mapping is achieved through a cross-modal feature alignment algorithm. The module uses an attention-weighted fusion layer to integrate the multimodal feature vectors and outputs collaborative content of at least two modalities: text, image, audio, or 3D model.
[0144] The feedback optimization module is used to dynamically optimize the output based on user feedback data, including explicit feedback data and implicit feedback data. It uses reinforcement learning strategies to adjust the internal parameter configuration interface values of the functional unit and update the temperature coefficient and sampling probability of the large language model.
[0145] The content output module is used to output optimized cross-modal content to the application terminal.
[0146] Furthermore, the scene data parsing module includes:
[0147] A text parsing unit, used to perform semantic dependency analysis and named entity recognition on text data to generate a set of scene object entities and associated attribute keywords;
[0148] The audio and video parsing unit is used to extract visual object boundaries using an instance segmentation model for image data and to identify sound source objects using an acoustic event detection model for audio data, and to align them with the scene object entity set;
[0149] The parameter encoding unit is used to parse static parameters into key-value pair data structures and dynamic parameters into time function expressions; it encodes the interaction logic into a triple structure, which includes event trigger conditions, execution constraints, and action instructions;
[0150] The structure output unit is used to output a standardized scene description structure, including a geometric feature parameter set associated with an object type identifier, a material attribute parameter set, a static parameter set and a dynamic parameter set, as well as a behavioral rule set consisting of conditional triggering rules, loop execution rules and event response rules.
[0151] Furthermore, the system also includes a visual debugging module for displaying the node status and data flow path of the dependency topology diagram in real time, recording the parameter adjustment history and highlighting the conflict resolution process, and providing a side-by-side comparison preview window for multi-modal generation results.
[0152] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the various embodiments of the present invention.
Claims
1. Functionalized scene modeling and large-scale model cross-dimensional content generation method, characterized by: The specific steps include: S100: Receive multimodal scene description data input by a user, wherein the scene description data includes at least one modality of text data, image data, and audio data, wherein the text data includes a semantic description field, a style preference field, and an interaction logic constraint field; S101: Parse the multimodal scene description data to extract scene object entities, object attribute parameters, and object interaction logic, wherein: The scene object entity is located by an object type identifier and is associated with a geometric feature parameter set, a material attribute parameter set, and a behavior rule set; The object attribute parameters are divided into a static parameter set and a dynamic parameter set. The static parameter set includes color value and size value, and the dynamic parameter set includes light intensity time function and ambient noise spectrum function. The object interaction logic is defined as condition triggering rules, loop execution rules and event response rules; S102: Construct a dynamically adjustable scene model, instantiate the scene object entity into functional units, each functional unit including an input interface, an output interface, and an internal parameter configuration interface; Generate a dependency topology graph between functional units based on the interaction logic, wherein the topology graph is a directed acyclic graph structure, in which nodes are functional units and edges are data flow paths; According to the real-time adjustment instruction input by the user, the scene model state is updated by modifying the value of the internal parameter configuration interface; S103: Calling a pre-trained large language model or a multimodal generative model to perform cross-dimensional content generation, mapping the functionalized unit parameters in the scene model to the generation request of the target modality, wherein the mapping is achieved by a cross-modal feature alignment algorithm that calculates the cosine similarity between the function parameter set and the input features of the generative model; An attention-weighted fusion layer is used to integrate multimodal feature vectors and output collaborative content of at least two modalities among text, image, audio or 3D model; S104: Dynamically optimize output based on user feedback data, where the feedback data includes explicit feedback data and implicit feedback data. The explicit feedback data includes parameter modification instructions and quality ratings, and the implicit feedback data includes user operation timestamps and interface interaction traces. Adjust the internal parameter configuration interface value of the functional unit through reinforcement learning strategy, and update the temperature coefficient and sampling probability of the large language model; S105: Output the optimized cross-modal content to the application terminal.
2. The method for functional scene modeling and large-scale cross-dimensional content generation according to claim 1 is characterized in that: The specific steps of parsing the multimodal scene description data and extracting scene object entities, object attribute parameters, and interaction logic between objects include: S200, performing semantic dependency analysis and named entity recognition on text data to generate a scene object entity set and associated attribute keywords; S201, using an instance segmentation model to extract visual object boundaries from image data, and using an acoustic event detection model to identify sound source objects from audio data, and aligning them with the scene object entity set; S202, parsing static parameters into a key-value pair data structure, and parsing dynamic parameters into a time function expression; S203: Encode the interaction logic into a triple structure, where the triple includes an event trigger condition, an execution constraint condition, and an action instruction; S204: Output a standardized scene description structure, including: The geometric feature parameter set and material attribute parameter set associated with the object type identifier; Static parameter sets and dynamic parameter sets; A behavioral rule set consisting of conditional triggering rules, loop execution rules, and event response rules.
3. The method for functional scene modeling and large-scale cross-dimensional content generation according to claim 2 is characterized in that: The time function expression in step S202 includes a light intensity time function and an ambient noise spectrum function; The formula of the light intensity time function is: in, is the light intensity value at time t, The maximum intensity value set by the user, is a preset periodic function, is the angular frequency parameter, is the phase offset, is the environmental attenuation factor, Is the basic light intensity constant; The formula of the ambient noise spectrum function is: in, is the noise energy density at frequency f, is the number of sound sources, is the amplitude weight of the kth sound source, is the center frequency of the kth sound source, is the frequency bandwidth parameter.
4. The method for functional scene modeling and large-scale cross-dimensional content generation according to claim 1 is characterized in that: The updating of the dependency topology graph in step S102 includes the following operations: When a user operation triggers an event response rule, the user operation is parsed to determine the data flow path to be modified and the modification type. The adjacency matrix storing the dependency relationships is updated. If a new path is added, the value of the corresponding matrix element is set to the preset default dependency weight. If the path is deleted, set the value of the corresponding matrix element to zero, and finally verify whether the updated topology graph satisfies the directed acyclic property; When parameters of multiple functional units conflict, identifying the conflicting parameters and associated functional unit identifiers, and selecting a final parameter value according to a priority rule, wherein the priority rule is: user explicit instructions take precedence over real-time operation behavior; The real-time operation behavior takes precedence over the scene default value; finally, the selected parameter value is written to the internal parameter configuration interface of the target functionalization unit; According to the updated content of the adjacency matrix, the range of the affected functionalized units is determined. The specific formula is: in, is the set of affected functionalized units, is a functional unit that is directly modified. For all functional units involved in this update, is the graph reachability determination function, is the updated adjacency matrix; Then recalculate the state value of the affected functionalized unit. The specific formula is: in, Functionalized Unit The new state value of Functionalized Unit arrive The dependency weight of For upstream functional units The current state, is the activation function; Record update logs to the version management database.
5. The method for functional scene modeling and large-scale cross-dimensional content generation according to claim 1 is characterized in that: The cross-modal feature alignment algorithm execution process in step S103 includes: Establish a shared embedding space from the function parameter set to the generative model input, with an embedding space dimension of 256; The contrast loss function is used to constrain the Euclidean distance between the text description vector and the image feature vector in the shared space to be less than the threshold. .
6. The method for functional scene modeling and large-scale cross-dimensional content generation according to claim 1 is characterized in that: The reinforcement learning strategy in step S104 specifically includes: Define the reward function R, the specific formula is: in, Rate the user, To generate a content consistency score, 、 is the weight.
7. Functionalized scene modeling and large-scale cross-dimensional content generation system, characterized by including: A multimodal data receiving module, configured to receive multimodal scene description data input by a user, wherein the scene description data includes at least one modality of text data, image data, and audio data, wherein the text data includes a semantic description field, a style preference field, and an interaction logic constraint field; A scene data parsing module, configured to parse the multimodal scene description data and extract scene object entities, object attribute parameters, and interaction logic between objects; A dynamic scene modeling module is used to construct a dynamically adjustable scene model, instantiating the scene object entity into functional units, each functional unit including an input interface, an output interface, and an internal parameter configuration interface; generating a dependency topology diagram between the functional units based on the interaction logic, wherein the topology diagram is a directed acyclic graph structure; and updating the scene model state by modifying the value of the internal parameter configuration interface according to the real-time adjustment instructions input by the user; A cross-dimensional content generation module is used to call a pre-trained large language model or a multimodal generation model to perform cross-dimensional content generation, mapping the functionalized unit parameters in the scene model to the generation request of the target modality through a cross-modal feature alignment algorithm; using an attention-weighted fusion layer to integrate the multimodal feature vectors and output collaborative content of at least two modalities: text, image, audio, or a three-dimensional model; A feedback optimization module is used to dynamically optimize the output based on user feedback data, including explicit feedback data and implicit feedback data; adjust the internal parameter configuration interface values of the functional unit through a reinforcement learning strategy, and update the temperature coefficient and sampling probability of the large language model; The content output module is used to output optimized cross-modal content to the application terminal.
8. The functional scene modeling and large-scale cross-dimensional content generation system according to claim 7 is characterized in that: The scene data parsing module includes: A text parsing unit, used to perform semantic dependency analysis and named entity recognition on text data to generate a set of scene object entities and associated attribute keywords; An audio and video parsing unit, configured to extract visual object boundaries from image data using an instance segmentation model, and identify sound source objects from audio data using an acoustic event detection model, and align the sound source objects with the scene object entity set; A parameter encoding unit is used to parse static parameters into a key-value pair data structure and parse dynamic parameters into a time function expression; encode the interaction logic into a triple structure, wherein the triple includes an event trigger condition, an execution constraint condition, and an action instruction; The structure output unit is used to output a standardized scene description structure, including a geometric feature parameter set associated with an object type identifier, a material attribute parameter set, a static parameter set and a dynamic parameter set, as well as a behavioral rule set consisting of conditional triggering rules, loop execution rules and event response rules.
9. The functional scene modeling and large-scale cross-dimensional content generation system according to claim 7 is characterized in that: The system also includes a visual debugging module for displaying the node status and data flow path of the dependency topology diagram in real time, recording the parameter adjustment history and highlighting the conflict resolution process, and providing a side-by-side comparison preview window for multi-modal generation results.
Citation Information
Patent Citations
Audio classification method and device, electronic equipment and storage medium
CN110929087A
Model generation method, model generation device, computer equipment and storage medium
CN116912413A
Multi-mode driven video generation method and device, computer equipment and readable storage medium
CN118803301A
Relatively controllable video generation system based on AIGC large model
CN119496964A
Remote digital service resource recommendation method and system based on artificial intelligence mining
CN119739929A
Cited By
Dynamic traceable data interaction method and system based on DAG rule engine
CN121050701A
A dynamic traceable data interaction method and system based on a DAG rule engine
CN121050701B
Visualization method and system based on land improvement and ecological restoration data
CN121747114A