A controllable virtual scene generation method and system fusing physical constraints

By introducing a physical constraint loss term and a multimodal feature fusion generation method, the problems of physical inconsistencies and insufficient adaptability in virtual scene generation are solved, achieving high-fidelity and controllable virtual scene generation, and improving the credibility and transfer performance of the simulation system.

CN122312929APending Publication Date: 2026-06-30NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing virtual scene generation technologies lack physical constraints, resulting in physically unreasonable generated results. This makes it difficult to generate low-probability, high-risk scenes as needed, and the generated scenes differ greatly from the real environment, affecting the simulation effect.

Method used

By introducing physical constraint loss terms, reconstruction loss terms, and adversarial loss terms, adversarial generator networks and conditional generator networks are constructed. Feature extraction and fusion are performed by combining multimodal input data, and a domain discriminator is used to generate scenes with physical rationality and controllability.

Benefits of technology

The generated virtual scenes are physically reasonable and controllable, and highly adaptable to the real environment, which improves the credibility and transfer performance of the simulation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122312929A_ABST
    Figure CN122312929A_ABST
Patent Text Reader

Abstract

This invention discloses a controllable virtual scene generation method and system that integrates physical constraints, belonging to the field of virtual scene generation technology. The method includes: acquiring multimodal input data for generating a virtual scene, and preprocessing the multimodal input data to obtain standard data; extracting features from the multimodal input data to obtain feature data, and fusing the feature data to obtain fused data; inputting the fused data and a first conditional signal into an initial generation network in a preset virtual scene generation network to obtain an initial virtual scene, wherein the optimization function of the initial generation network includes a physical constraint loss term, a reconstruction loss term, and an adversarial loss term; and inputting the initial virtual scene and a second conditional signal into a refined generation network in a preset virtual scene to obtain a refined virtual scene, wherein the optimization function of the refined generation network includes a physical constraint loss term, a cross-modal consistency loss term, and a conditional adversarial loss term.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual scene generation technology, and in particular to a controllable virtual scene generation method and system that integrates physical constraints. Background Technology

[0002] In industrial digital twin systems (including autonomous driving simulation testing scenarios), high-fidelity virtual scene data rich in edge cases is crucial. In particular, the on-demand generation of low-probability, high-risk scenarios such as vehicle loss of control, extreme weather, and rare traffic participation behaviors can significantly improve the safety and robustness of intelligent systems in the real world. An ideal virtual scene generation technology should be able to efficiently and controllably generate datasets that conform to physical laws and contain rich edge cases, providing a reliable testing environment for digital twin systems and autonomous driving algorithms.

[0003] As the requirements for high fidelity in simulation scenarios continue to increase, current virtual scene generation technologies used for simulation testing cannot meet existing needs and still suffer from the following problems. Existing mainstream generation models lack a fundamental understanding of the physical world, and their generated results often exhibit physically unreasonable phenomena in terms of geometric relationships, motion trajectories, and material properties, such as object penetration and violations of dynamic laws, making the generated scenes unsuitable for high-fidelity simulations. Furthermore, existing methods have insufficient control over the generated content, making it difficult to generate scenes with specific semantics and low-probability scenarios, such as rare machine malfunctions and traffic behaviors. In addition, the domain differences between the generated scene and the real environment also severely affect the transfer effect from the simulation scene to the real object.

[0004] Currently, there are four main technical approaches to virtual scene generation: Traditional rendering engine-based generation schemes, while ensuring physical plausibility, heavily rely on manual design, resulting in low efficiency and insufficient flexibility; purely data-driven generative adversarial networks (GANs) can quickly generate diverse scenes, but the lack of physical constraints may lead to results that violate physical laws; conditional variational autoencoder-based methods can control the general category of generated content through input labels, but have limited fine-grained control over complex dynamic scenes; and unconstrained diffusion models can generate high-quality images, but cannot guarantee the physical plausibility and simulation usability of the generated scenes. Existing solutions have limitations in physical realism, generation controllability, and domain adaptability, making it difficult to meet the stringent requirements of high-fidelity virtual scenes in industrial digital twins (including autonomous driving simulation).

[0005] In terms of generation quality, most generative models lack the constraints and guidance of a physics engine, leading to physical discrepancies such as unreasonable object trajectories and abnormal material properties in the generated scenes. Secondly, regarding case coverage, traditional generation methods are insufficient in modeling low-probability edge cases, making it difficult to systematically generate targeted and controllable data for specific needs. Finally, in terms of domain adaptability, the distributional differences between the generated scenes and the real environment cause performance degradation when converting simulations to physical objects, limiting the practical value of the generated data. Current technologies struggle to simultaneously meet the multiple requirements of high physical realism, high controllability, and strong domain adaptability, thus hindering the further development of industrial digital twin simulation systems. Summary of the Invention

[0006] The purpose of this invention is to provide a controllable virtual scene generation method and system that integrates physical constraints, so as to solve the problems of the limitations of the prior art.

[0007] This invention provides a controllable virtual scene generation method that integrates physical constraints, comprising the following steps:

[0008] Acquire multimodal input data for generating virtual scenes, and preprocess the multimodal input data to obtain standard data; Feature extraction is performed on the multimodal input data to obtain feature data, and feature fusion is performed on the feature data to obtain fused data; The fused data and the first condition signal are input into the initial generation network in the preset virtual scene generation network to obtain the initial virtual scene. The initial generation network includes an adversarial generator network, and the optimization function of the adversarial generator network includes a physical constraint loss term, a reconstruction loss term, and an adversarial loss term. The initial virtual scene and the second condition signal are input into the refinement generation network in the preset virtual scene to obtain the refined virtual scene. The refinement generation network includes a condition generator network, and the optimization function of the refinement generation network includes a physical constraint loss term, a cross-modal consistency loss term, and a conditional adversarial loss term.

[0009] Optionally, before the step of inputting the fused data and the first condition signal into the initial generation network of the preset virtual scene generation network, the method further includes: The physical constraint loss term, which represents the dynamics of motion and spatial geometric constraints, is embedded in the preset virtual scene generation network in the form of a loss. Based on the optimization function, which includes the physical constraint loss term, reconstruction loss term, and adversarial loss term, the original adversarial network, which includes the original generator network, is trained through backpropagation to obtain the secondary generator network. Based on the optimization function, which includes physical constraint loss, cross-modal consistency loss and conditional adversarial loss, the original conditional adversarial network, which includes the original conditional generator network, is trained by backpropagation to obtain the secondary conditional generator network. The secondary generator network and the secondary conditional generator network are trained using a domain discriminator based on an adversarial domain discriminant loss function to obtain an initial generator network and a conditional generator network. A preset virtual scene generation network is then constructed based on the initial generator network and the conditional generator network.

[0010] Optionally, the step of embedding the physical constraint loss term, which characterizes motion dynamics and spatial geometric constraints, into a preset virtual scene generation network in the form of a loss includes: Through the formula:

[0011] The physical constraint loss term, representing the dynamics of motion and spatial geometric constraints, is embedded in a pre-defined virtual scene generation network in the form of a loss function. To iterate through intermediate states, The gradient of the physical loss. For physical loss items, To rebuild the lost items, Let t be the learning rate and t be the number of iterations.

[0012] Optionally, the step of training the primary adversarial network, including the primary generator network, through backpropagation based on an optimization function comprising a physical constraint loss term, a reconstruction loss term, and an adversarial loss term to obtain the secondary generator network includes: By maximizing the adversarial loss term of the discriminator corresponding to the original generator network and minimizing the optimization function of the original generator network, which includes the physical constraint loss term, reconstruction loss term, and adversarial loss term, and by training the original adversarial network including the original generator network through backpropagation according to the gradient penalty or spectral normalization strategy, a secondary generator network is obtained.

[0013] Optionally, the mathematical representation of the optimization function, which includes the physical constraint loss term, the reconstruction loss term, and the adversarial loss term, is as follows:

[0014] in, For physical constraint loss terms, To rebuild the lost items, For the initial generator network and the discriminator corresponding to the initial generator network The counter-loss item.

[0015] Optionally, the mathematical representation of the optimization function, which includes the physical constraint loss term, the cross-modal consistency loss term, and the conditional adversarial loss term, is as follows:

[0016] in, For physical constraint loss terms, For cross-modal consistency loss term, To counteract the loss term, and These are the balancing weighting coefficients.

[0017] Optionally, the adversarial domain discriminant loss function is characterized by the following mathematical representation:

[0018] in, For the domain discriminator, To refine the virtual scene, These are real samples.

[0019] On the other hand, this application provides a controllable virtual scene generation system that integrates physical constraints, including: The data acquisition module is used to acquire multimodal input data for generating virtual scenes and to preprocess the multimodal input data to obtain standard data. The feature fusion module is used to extract features from the multimodal input data to obtain feature data, and to fuse the feature data to obtain fused data. The first generation module is used to input the fused data and the first condition signal into the initial generation network in the preset virtual scene generation network to obtain the initial virtual scene. The initial generation network includes an adversarial generator network, and the optimization function of the adversarial generator network includes a physical constraint loss term, a reconstruction loss term, and an adversarial loss term. The second generation module is used to input the initial virtual scene and the second condition signal into the refinement generation network in the preset virtual scene to obtain the refined virtual scene. The refinement generation network includes a condition generator network, and the condition adversarial network optimization function includes a physical constraint loss term, a cross-modal consistency loss term, and a condition adversarial loss term.

[0020] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the controllable virtual scene generation method that integrates physical constraints as described above.

[0021] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the controllable virtual scene generation method that integrates physical constraints as described above.

[0022] Beneficial effects: By transforming the dynamics and kinematic constraints, such as motion trajectories and collision detection calculated by the physics engine, into differentiable loss functions and embedding them into the sampling and optimization process of the generative model, the physical plausibility of the generated scenes is achieved. Based on a conditional generative adversarial training mechanism, multimodal conditional signals such as semantic labels, bounding boxes, and natural language descriptions are introduced to guide the generative model to generate virtual scenes with specific targets, layouts, and attributes on demand, significantly improving the controllability of generating complex scenes. By introducing a domain discriminator into the generative model, the generated scene is aligned with the real scene in the feature space, effectively enhancing the practical value of the generated data and ensuring the transfer performance from simulation to reality. Under the premise of ensuring physical realism, efficient and controllable generation of virtual scenes is achieved, and its adaptability and credibility in different simulation platforms and real environments are significantly improved. Attached Figure Description

[0023] Figure 1 This is a flowchart illustrating a controllable virtual scene generation method that integrates physical constraints according to the present invention. Figure 2 This is a flowchart illustrating another controllable virtual scene generation method that integrates physical constraints according to the present invention. Figure 3 This is a schematic diagram illustrating the process of physical constraint modeling and differentiable embedding generation in a controllable virtual scene generation method that integrates physical constraints according to the present invention. Figure 4 This is a schematic diagram of the initial generation network in a controllable virtual scene generation method that integrates physical constraints according to the present invention. Figure 5 This is a flowchart illustrating the process of refining the generation network in a controllable virtual scene generation method that integrates physical constraints according to the present invention. Figure 6 This is a schematic diagram of the structure of a controllable virtual scene generation system that integrates physical constraints, provided in an embodiment of this application. Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation

[0024] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments. Figure 1 As shown in the figure, this embodiment of a controllable virtual scene generation method that integrates physical constraints includes: S110. Obtain multimodal input data for generating virtual scenes, and preprocess the multimodal input data to obtain standard data; S120. Perform feature extraction on the multimodal input data to obtain feature data, and perform feature fusion on the feature data to obtain fused data; S130. Input the fused data and the first condition signal into the initial generation network in the preset virtual scene generation network to obtain the initial virtual scene. The initial generation network includes an adversarial generator network. The optimization function of the adversarial generator network includes a physical constraint loss term, a reconstruction loss term, and an adversarial loss term. S140. Input the initial virtual scene and the second condition signal into the refinement generation network in the preset virtual scene to obtain the refined virtual scene. The refinement generation network includes a condition generator network, and the optimization function of the refinement generation network includes a physical constraint loss term, a cross-modal consistency loss term, and a conditional adversarial loss term.

[0025] For example, such as Figure 2 As shown, firstly, the multimodal input data required for generating the virtual scene is obtained, and then uniformly formatted and semantically aligned to ensure that the multi-source input data can be represented within a unified feature space. The input dataset can be represented as:

[0026] in, Indicates the first A semantic segmentation map of each sample is used to reflect the spatial distribution of objects of each category in the scene; This represents the corresponding scene structure sketch or layout diagram, used to describe the geometric topological relationships of the scene; Represents natural language descriptive text, used to provide semantic information describing scene distribution and content; This represents a vector of physical parameters corresponding to the scene, including but not limited to gravity coefficient, friction factor, wind field distribution, light intensity, temperature gradient, etc. It can be represented as:

[0027] in, The coefficient of gravitational acceleration. The coefficient of friction, The wind field velocity vector, Indicates fluid density, Indicates light intensity.

[0028] Through data standardization, data from different modalities are unified in spatial dimension and semantic level, thus solving the problem that multi-source input data cannot be effectively integrated in the same coordinate and semantic space in existing technologies.

[0029] To transform standardized data from different modalities into a unified feature representation and achieve deep fusion of semantic, spatial structure, and physical information, for each set of inputs... The system employs different feature encoders:

[0030] in, For image encoders, extract spatial texture and layout information; It is a structural encoder that captures geometric and topological features; For text encoders, extract semantic information; It is a physical feature mapper that embeds physical parameters into a learnable feature space.

[0031] To achieve multimodal fusion, a cross-modal attention mechanism is introduced, and the output fused data is represented as follows:

[0032] in, This represents a multi-head attention operation, used to adaptively adjust feature weights based on inter-modal correlations. The output is fused data. As input for subsequent generation modules.

[0033] By fusing features, textual semantics, geometric layout, and physical constraints can be perceived simultaneously, making the generation process multi-dimensionally controllable. This effectively solves the problems of "fuzzy semantic description, poor geometric consistency, and lack of physical context" in traditional methods.

[0034] In one possible implementation, prior to the step of inputting the fused data and the first condition signal into the initial generation network of the preset virtual scene generation network, the method further includes: The physical constraint loss term, which represents the dynamics of motion and spatial geometric constraints, is embedded in the preset virtual scene generation network in the form of a loss. Based on the optimization function, which includes the physical constraint loss term, reconstruction loss term, and adversarial loss term, the original adversarial network, which includes the original generator network, is trained through backpropagation to obtain the secondary generator network. Based on the optimization function, which includes physical constraint loss, cross-modal consistency loss and conditional adversarial loss, the original conditional adversarial network, which includes the original conditional generator network, is trained by backpropagation to obtain the secondary conditional generator network. The secondary generator network and the secondary conditional generator network are trained using a domain discriminator based on an adversarial domain discriminant loss function to obtain an initial generator network and a conditional generator network. A preset virtual scene generation network is then constructed based on the initial generator network and the conditional generator network.

[0035] For example, such as Figures 2-3 As shown, by constructing a differentiable composite physical constraint function (physical constraint loss term), motion dynamics and spatial geometric constraints are embedded in the optimization objective of the generative network in the form of loss, thereby ensuring that the real-time constraint model output in the generation stage conforms to physical laws. The input of the composite physical constraint function is fused data. With physical parameters The output is the physical constraint loss. The loss function is a weighted sum of motion dynamics constraints and collision constraints, and its expression is:

[0036] in, and It is a generative network Object velocity and geometric fields generated based on fused data and physical parameters; Derived from physical parameter vector The gravitational acceleration component in the equation; coefficient The proportionality coefficient is calculated based on physical parameters and is used to control the acceleration balance term; This is a collision constraint term based on geometry detection, used to penalize spatial overlap between different objects; These are the weighting coefficients for each loss term.

[0037] The first term in the loss function reflects the degree of deviation between the velocity field of the generated object and the direction of gravity. Minimizing it can constrain the generated result to conform to the basic dynamic laws and constrain the consistency between the direction of motion and the direction of gravity. The second term is used to detect and punish spatial penetration and unreasonable contact behavior between geometric objects in the scene, thereby ensuring that the generated result meets the continuity of the spatial structure and physical feasibility.

[0038] Optionally, the step of embedding the physical constraint loss term, which characterizes motion dynamics and spatial geometric constraints, into a preset virtual scene generation network in the form of a loss includes: Through the formula:

[0039] The physical constraint loss term, representing the dynamics of motion and spatial geometric constraints, is embedded in a pre-defined virtual scene generation network in the form of a loss function. To iterate through intermediate states, The gradient of the physical loss. For physical loss items, To rebuild the lost items, Let t be the learning rate and t be the number of iterations.

[0040] For example, with the physical constraints being differentiable embedded in the feature space, each generation iteration of the system requires the computation of intermediate states. physical loss and utilize its gradient The sampling direction is guided. Then, through backpropagation, the gradient is passed along the generation path to the fusion module and the generator network, thereby realizing the differentiable embedding of physical constraints in the feature space.

[0041] This process can be formally represented as:

[0042] in, The learning rate or sampling step size, To reconstruct the loss, and to prevent the model from generating visually meaningless or out-of-distribution results simply to satisfy physical constraints.

[0043] By introducing the aforementioned differentiable physical constraint modeling during the sampling and optimization process, dynamic and geometric constraints can be learned simultaneously at the pixel-level learning stage, and the output direction can be automatically adjusted. This significantly reduces physical distortion in the generated results, making the generated scene visually realistic and physically interpretable, thereby improving the system's simulation credibility and transferability.

[0044] For example, based on fused data With physical constraints A foundational generative network is constructed to generate physically plausible scenes. The main task of this network is to learn the spatial distribution, geometric structure, and fundamental physical relationships of objects in the scene, generating a macroscopically plausible but not yet refined "rough" scene image. This step lays the foundation for subsequent refined generation, ensuring physical consistency and semantic correctness.

[0045] like Figure 2 and Figure 4 As shown. The system first constructs a generator network. With discriminator network The generator uses fused data and coarse-grained conditional signals As input, output a base scene image. Input conditions This includes semantic layout diagrams, partition labels, or scene-level descriptive information, used to guide the generator in constructing a reasonable overall scene structure. The generator can be formally represented as:

[0046] in, This represents the parameterized function of the generator. This is the set of learnable parameters. Output The initial scene results are represented.

[0047] To ensure that the generated results are visually realistic, semantically accurate, and conform to physical laws, a comprehensive optimization objective function for the basic generative network is designed here:

[0048] Among them, the counter-loss item The distribution used to constrain the generated image to match the distribution of the real scene is defined as:

[0049] Here is the discriminator The generator learns to distinguish between real-world and generated scenes, thereby guiding it to continuously improve the visual realism of the generated results.

[0050] This is the reconstruction loss term, used to maintain the semantic and spatial structure consistency between the generated result and the input conditions. It is defined as follows: in, An optional reference coarse truth image. Indicates the feature extraction operator, These are weighting coefficients. The purpose of this reconstruction loss term is to ensure that the generator reconstructs a scene structure consistent with the input description in terms of global semantics.

[0051] Physical constraints This ensures that the generated results conform to physical laws in terms of dynamics and geometry. During training, the gradient of the physical constraint term directly affects the generator's parameter updates through backpropagation, enabling the model to follow basic physical constraints while learning visual features.

[0052] The optimization of the generator and discriminator follows an adversarial strategy of alternating training, as follows:

[0053] During training, the discriminator By maximizing adversarial loss, the generator improves its ability to distinguish between real and generated samples. By minimizing Optimize network parameters Achieving a comprehensive balance between physical consistency and visual realism.

[0054] To improve training stability, gradient penalty or spectral normalization strategies can be introduced to prevent discriminator overfitting or gradient vanishing problems. Specific training update rules are as follows: in, These are the learning rates for the generator and the discriminator, respectively. These are the parameters for the discriminator.

[0055] Through the above optimization process, the generator It can learn the spatial layout, shape distribution, and physical relationships of objects in a scene, thereby generating a structurally coherent and physically plausible basic scene. Compared with traditional generation methods that rely solely on visual discrimination, by introducing physical constraints, it effectively reduces distortions such as object overlap, floating, and inconsistent gravity directions in virtual scenes, making the generated results both visually realistic and dynamically interpretable. The system's final output is a physically plausible basic scene image that has not yet undergone fine-tuning of textures and lighting. This output serves as input for subsequent steps, providing a stable and reliable physical basis for the refined generation of conditional adversarial processes.

[0056] For example, such as Figure 2 and Figure 5 As shown, by introducing multimodal conditional constraints and adversarial training mechanisms, the generated results not only possess macroscopic physical consistency but also conform to specified semantic descriptions and perceptual features at the detailed level. The input is the basic scene. With high precision conditions The output is a high-fidelity scene. Among them, high-precision conditions Information including natural language descriptions, texture attributes, lighting parameters, and material types is used to perform semantic-level control over scene details.

[0057] via condition generator With discriminator This constitutes a controllable generative adversarial network (GAN) structure. The generator's function is to generate a refined scene that meets the requirements based on the input conditions, while the discriminator is used to determine whether the generated scene is realistic and meets the input conditions.

[0058] A generator can be formally represented as:

[0059] in, Function mappings representing condition generators. The base scene image output from step four. To refine the conditional signals, These are the multimodal high-level features extracted by the semantic enhancement module. The generator output is... This refers to the refined scenario results of the target.

[0060] To enable the generator to produce high-fidelity and semantically consistent scene results, a conditional adversarial loss function is used as one of the main optimization objectives, defined as follows:

[0061] in, This indicates that the discriminator operates under given conditions. Lower the judgment sample The probability of whether an image is realistic is determined by minimizing this loss function, which guides the generator to produce images that are both realistic and conform to the input conditions.

[0062] To further ensure the correspondence between semantic descriptions and generated image content, a cross-modal consistency loss term is introduced. This is used to constrain the degree of matching between text features and visual features, and is defined as follows:

[0063] Cross-modal consistency loss term can effectively prevent semantic drift problem and ensure that the model generates scene content with clear semantic correspondence based on the input description.

[0064] Meanwhile, to maintain the physical plausibility of the generated results, a physical constraint loss term from step 3 is introduced again during the refinement stage. This ensures that the generated image adheres to the principle of physical consistency during high-precision texture generation. Finally, the comprehensive optimization objective function is defined as:

[0065] in, and To balance the weighting coefficients, the influence of semantic consistency and physical constraints on the overall optimization is controlled separately.

[0066] Through a conditional adversarial refined generation process, detailed semantic control and texture generation can be achieved while ensuring basic physical plausibility. Compared with traditional generation methods, this step can effectively control visual details such as the material, lighting direction, and reflection characteristics of objects in the scene, achieving a balance in realism, semantic consistency, and physical consistency. The final output high-fidelity scene image has the visual quality of a real scene and can be directly applied to autonomous driving simulation, virtual reality environment rendering, and digital twin system construction.

[0067] Optionally, the step of training the primary adversarial network, including the primary generator network, through backpropagation based on an optimization function comprising a physical constraint loss term, a reconstruction loss term, and an adversarial loss term to obtain the secondary generator network includes: By maximizing the adversarial loss term of the discriminator corresponding to the original generator network and minimizing the optimization function of the original generator network, which includes the physical constraint loss term, reconstruction loss term, and adversarial loss term, and by training the original adversarial network including the original generator network through backpropagation according to the gradient penalty or spectral normalization strategy, a secondary generator network is obtained.

[0068] Optionally, the mathematical representation of the optimization function, which includes the physical constraint loss term, the reconstruction loss term, and the adversarial loss term, is as follows:

[0069] in, For physical constraint loss terms, To rebuild the lost items, For the initial generator network and the discriminator corresponding to the initial generator network The counter-loss item.

[0070] Optionally, the mathematical representation of the optimization function, which includes the physical constraint loss term, the cross-modal consistency loss term, and the conditional adversarial loss term, is as follows:

[0071] in, For physical constraint loss terms, For cross-modal consistency loss term, To counteract the loss term, and These are the balancing weighting coefficients.

[0072] Optionally, the adversarial domain discriminant loss function is characterized by the following mathematical representation:

[0073] in, For the domain discriminator, To refine the virtual scene, These are real samples.

[0074] For example, to address the discrepancy in feature distribution between the generated domain and the real domain, and to enhance the cross-domain generalization and domain adaptation capabilities of the generative model, this invention employs a domain adaptation strategy based on feature alignment. The input is the generated scene. Real-world sample The output is a feature representation after domain alignment, and the aligned features will serve as the input basis for the quality assessment module.

[0075] To ensure consistency between the generated domain and the real domain distribution, a domain discriminator is introduced. Its task is to determine whether the input features belong to the generated domain or the real domain.

[0076] The optimization objective of the domain discriminator is defined as the adversarial domain discrimination loss function:

[0077] By maximizing this loss, the discriminator is trained to effectively distinguish between generated and real domain samples; while the generator minimizes this loss so that the generated features gradually approximate the real domain features in terms of distribution.

[0078] By introducing a gradient inversion layer (GRL) for inverse optimization, the generator actively "deceives" the discriminator through adversarial learning, thereby achieving inter-domain distribution alignment in the feature space and significantly reducing the gap between simulation and reality.

[0079] In one possible implementation, the overall quality and physical consistency of the generated scene are quantitatively evaluated. The input is the generated scene. Real samples Semantic information With physical reference index The output is the overall quality score Q.

[0080] The comprehensive evaluation indicators are defined as follows:

[0081] in Measuring distributional differences Reflecting the physical rationality of the generated results, Measure the degree of semantic alignment.

[0082] This module can be deployed independently as a verification stage, and the output quality indicators can be directly used for model performance analysis and feedback adjustment.

[0083] When the scene quality score Below the set threshold At this point, the system automatically triggers a closed-loop feedback mechanism to adaptively update the parameters of the generator and physical constraint modules. The system adaptively adjusts the parameters of the generator and physical constraint modules based on quality differences, achieving continuous learning and performance improvement.

[0084] The update rule is expressed as follows:

[0085] in For learning rate, This is the feedback intensity coefficient.

[0086] Through this mechanism, the system forms a self-circulating process of "generation-evaluation-feedback-regeneration", ensuring the long-term stability and self-evolutionary performance of the model.

[0087] On the other hand, such as Figure 6 As shown, this application provides a controllable virtual scene generation system that integrates physical constraints, including: The data acquisition module 201 is used to acquire multimodal input data for generating virtual scenes, and to preprocess the multimodal input data to obtain standard data; The feature fusion module 202 is used to extract features from the multimodal input data to obtain feature data, and to fuse the feature data to obtain fused data. The first generation module 203 is used to input the fused data and the first condition signal into the initial generation network in the preset virtual scene generation network to obtain an initial virtual scene. The initial generation network includes an adversarial generator network, and the optimization function of the adversarial generator network includes a physical constraint loss term, a reconstruction loss term, and an adversarial loss term. The second generation module 204 is used to input the initial virtual scene and the second condition signal into the refinement generation network in the preset virtual scene to obtain the refined virtual scene. The refinement generation network includes a condition generator network, and the condition adversarial network optimization function includes a physical constraint loss term, a cross-modal consistency loss term, and a condition adversarial loss term.

[0088] In one possible implementation, such as Figure 7 As shown, this application embodiment provides a terminal device 300, including: a memory 310, a processor 320, and a first computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the first computer program 311, it acquires multimodal input data for generating a virtual scene and preprocesses the multimodal input data to obtain standard data. Feature extraction is performed on the multimodal input data to obtain feature data, and feature fusion is performed on the feature data to obtain fused data; The fused data and the first condition signal are input into the initial generation network in the preset virtual scene generation network to obtain the initial virtual scene. The initial generation network includes an adversarial generator network, and the optimization function of the adversarial generator network includes a physical constraint loss term, a reconstruction loss term, and an adversarial loss term. The initial virtual scene and the second condition signal are input into the refinement generation network in the preset virtual scene to obtain the refined virtual scene. The refinement generation network includes a condition generator network, and the optimization function of the refinement generation network includes a physical constraint loss term, a cross-modal consistency loss term, and a conditional adversarial loss term.

[0089] In one possible implementation, such as Figure 8 As shown, this application embodiment provides a computer-readable storage medium 400, on which a second computer program 411 is stored. When the second computer program 411 is executed by a processor, it acquires multimodal input data for generating a virtual scene and preprocesses the multimodal input data to obtain standard data. Feature extraction is performed on the multimodal input data to obtain feature data, and feature fusion is performed on the feature data to obtain fused data; The fused data and the first condition signal are input into the initial generation network in the preset virtual scene generation network to obtain the initial virtual scene. The initial generation network includes an adversarial generator network, and the optimization function of the adversarial generator network includes a physical constraint loss term, a reconstruction loss term, and an adversarial loss term. The initial virtual scene and the second condition signal are input into the refinement generation network in the preset virtual scene to obtain the refined virtual scene. The refinement generation network includes a condition generator network, and the optimization function of the refinement generation network includes a physical constraint loss term, a cross-modal consistency loss term, and a conditional adversarial loss term.

[0090] The methods and framework disclosed in this invention can be implemented in other ways. The embodiments described above are merely illustrative, and the division of modules is only a logical functional division; in actual implementation, there may be other division methods. It will be apparent to those skilled in the art that this invention is not limited to the details of the above exemplary embodiments, and that this invention can be implemented in other specific forms without departing from its essential characteristics.

[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for generating controllable virtual scenes that integrates physical constraints, characterized in that, Includes the following steps: Acquire multimodal input data for generating virtual scenes, and preprocess the multimodal input data to obtain standard data; Feature extraction is performed on the multimodal input data to obtain feature data, and feature fusion is performed on the feature data to obtain fused data; The fused data and the first condition signal are input into the initial generation network in the preset virtual scene generation network to obtain the initial virtual scene. The initial generation network includes an adversarial generator network, and the optimization function of the adversarial generator network includes a physical constraint loss term, a reconstruction loss term, and an adversarial loss term. The initial virtual scene and the second condition signal are input into the refinement generation network in the preset virtual scene to obtain the refined virtual scene. The refinement generation network includes a condition generator network, and the optimization function of the refinement generation network includes a physical constraint loss term, a cross-modal consistency loss term, and a conditional adversarial loss term.

2. The method for generating a controllable virtual scene incorporating physical constraints according to claim 1, characterized in that, Before the step of inputting the fused data and the first condition signal into the initial generation network of the preset virtual scene generation network, the method further includes: The physical constraint loss term, which represents the dynamics of motion and spatial geometric constraints, is embedded in the preset virtual scene generation network in the form of a loss. Based on the optimization function, which includes the physical constraint loss term, reconstruction loss term, and adversarial loss term, the original adversarial network, which includes the original generator network, is trained through backpropagation to obtain the secondary generator network. Based on the optimization function, which includes physical constraint loss, cross-modal consistency loss and conditional adversarial loss, the original conditional adversarial network, which includes the original conditional generator network, is trained by backpropagation to obtain the secondary conditional generator network. The secondary generator network and the secondary conditional generator network are trained using a domain discriminator based on an adversarial domain discriminant loss function to obtain an initial generator network and a conditional generator network. A preset virtual scene generation network is then constructed based on the initial generator network and the conditional generator network.

3. The method for generating a controllable virtual scene incorporating physical constraints according to claim 2, characterized in that, The step of embedding the physical constraint loss term, which characterizes motion dynamics and spatial geometric constraints, into a preset virtual scene generation network in the form of a loss includes: Through the formula: The physical constraint loss term, representing the dynamics of motion and spatial geometric constraints, is embedded in a pre-defined virtual scene generation network in the form of a loss function. To iterate through intermediate states, The gradient of the physical loss. For physical loss items, To rebuild the lost items, Let t be the learning rate and t be the number of iterations.

4. The method for generating a controllable virtual scene incorporating physical constraints according to claim 2, characterized in that, The step of training the primary adversarial network, including the primary generator network, through backpropagation based on an optimization function that includes a physical constraint loss term, a reconstruction loss term, and an adversarial loss term to obtain a secondary generator network includes: By maximizing the adversarial loss term of the discriminator corresponding to the original generator network and minimizing the optimization function of the original generator network, which includes the physical constraint loss term, reconstruction loss term, and adversarial loss term, and by training the original adversarial network including the original generator network through backpropagation according to the gradient penalty or spectral normalization strategy, a secondary generator network is obtained.

5. The method for generating a controllable virtual scene incorporating physical constraints according to claim 1, characterized in that, The mathematical representation of the optimization function, which includes the physical constraint loss term, the reconstruction loss term, and the adversarial loss term, is as follows: in, For physical constraint loss terms, To rebuild the lost items, For the initial generator network and the discriminator corresponding to the initial generator network The counter-loss item.

6. The method for generating a controllable virtual scene incorporating physical constraints according to claim 1, characterized in that, The mathematical representation of the optimization function, which includes the physical constraint loss term, the cross-modal consistency loss term, and the conditional adversarial loss term, is as follows: in, For physical constraint loss terms, For cross-modal consistency loss term, To counteract the loss term, and These are the balancing weighting coefficients.

7. The method for generating a controllable virtual scene incorporating physical constraints according to claim 2, characterized in that, The mathematical representation of the adversarial domain discriminant loss function is as follows: in, For the domain discriminator, To refine the virtual scene, These are real samples.

8. A controllable virtual scene generation system integrating physical constraints, characterized in that, include: The data acquisition module is used to acquire multimodal input data for generating virtual scenes and to preprocess the multimodal input data to obtain standard data. The feature fusion module is used to extract features from the multimodal input data to obtain feature data, and to fuse the feature data to obtain fused data. The first generation module is used to input the fused data and the first condition signal into the initial generation network in the preset virtual scene generation network to obtain the initial virtual scene. The initial generation network includes an adversarial generator network, and the optimization function of the adversarial generator network includes a physical constraint loss term, a reconstruction loss term, and an adversarial loss term. The second generation module is used to input the initial virtual scene and the second condition signal into the refinement generation network in the preset virtual scene to obtain the refined virtual scene. The refinement generation network includes a condition generator network, and the condition adversarial network optimization function includes a physical constraint loss term, a cross-modal consistency loss term, and a condition adversarial loss term.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the controllable virtual scene generation method that integrates physical constraints as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the controllable virtual scene generation method that integrates physical constraints as described in any one of claims 1 to 7.