Virtual environment data processing method, electronic equipment and storage medium

By combining visual language models with multi-constraint iterative optimization, the problem of unreliable physics in AI-generated virtual environments is solved, enabling efficient human-computer collaborative creation and ensuring the physical realism of the virtual environment and user experience.

CN120848743AActive Publication Date: 2025-10-28THE HONG KONG POLYTECHNIC UNIV SHENZHEN RES INST
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511377422.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2025-10-28
Estimated Expiration
2045-09-25

AI Technical Summary

Technical Problem

Existing AI-driven methods for generating virtual environment scenes cannot guarantee physical realism and semantic rationality, which can lead to problems such as objects clipping through each other, being suspended in mid-air, or having positions that do not conform to physical common sense and functional logic. This reduces the reliability of virtual environment data and the user's immersive experience.

Method used

By combining intelligent prediction and multi-constraint iterative optimization using a visual language model, information about the objects to be placed and scene features are acquired. The visual language model is used to generate optimized layout data, and pose correction is performed based on multiple physical placement constraints. The optimization is iteratively performed until a stable state is reached, ensuring that the generated virtual environment data conforms to physical laws and user intent.

Benefits of technology

It significantly improves the reliability and physical realism of virtual environment data, realizes an efficient and transparent human-computer collaborative creation process, fixes physical errors in AI layout, and enhances the user's immersive creation experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120848743A_ABST
    Figure CN120848743A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a virtual environment data processing method, electronic equipment and a storage medium. The method comprises the steps of responding to a virtual data control instruction of a user in a virtual environment scene; when the virtual data control instruction represents updating of the virtual environment scene, obtaining placement object information of the to-be-placed object and structural feature information and visual information of the virtual environment scene, and inputting the structural feature information, the structural feature information and the visual information into a visual language model for data processing; obtaining optimized arrangement data of the to-be-placed object; generating a to-be-placed object in the virtual environment scene based on the optimized arrangement data to obtain an adjusted virtual scene; and on the basis of the multiple physical placement constraints, at least one time of pose iteration correction is performed on the to-be-placed object in the adjusted virtual scene to generate the target virtual scene which can be adjusted by the user, so that the reliability and the physical reality sense of scene data are remarkably improved, and the requirements of the user are better met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to virtual environment data processing methods, electronic devices, and storage media. Background Technology

[0002] Virtual reality (VR) technology provides an immersive interactive experience, creating a physically realistic and believable experience for users within a virtual environment. Currently, the creation of virtual environment scenes primarily relies on manual operation using specialized software, or automated layout through procedural content generation (PCG) and artificial intelligence (AI) driven methods. Among these, some emerging AI generation methods can directly output a complete 3D scene layout scheme using large models based on text or image prompts. These methods are typically a "one-click" black-box process; that is, after receiving high-level instructions, the AI ​​predicts and generates what it considers the optimal, idealized object arrangement, thus achieving a degree of automation in scene generation.

[0003] However, the layout results generated by the aforementioned AI-driven scene generation methods are merely idealized predictions based on data and logic, which cannot fully guarantee the physical realism and semantic rationality of the final scene. This can lead to problems such as objects clipping through each other, floating in mid-air, or being placed in positions that do not conform to physical common sense and functional logic. As a result, the reliability of the generated virtual environment data is low, which seriously damages the physical realism of the virtual scene and the user's immersive experience. Summary of the Invention

[0004] This application provides a virtual environment data processing method, electronic device, and storage medium, which can improve the reliability of the generated virtual environment data.

[0005] To achieve the above objectives, a first aspect of this application proposes a virtual environment data processing method, the method comprising: Responding to user commands for controlling virtual data in a virtual environment; When the virtual data control instruction characterizes the update of the virtual environment scene, it acquires the placement information of the object to be placed and the structural feature information and visual information of the virtual environment scene, and inputs the structural feature information, the structural feature information and the visual information into the visual language model for data processing to obtain the optimized arrangement data of the object to be placed in the virtual environment scene; Based on the optimized layout data, the object to be placed is generated in the virtual environment scene to obtain an adjusted virtual scene. Based on multiple physical placement constraints, the pose adjustment amount of the object to be placed in the adjusted virtual scene is calculated, and the pose of the object to be placed in the adjusted virtual scene is corrected based on the pose adjustment amount to obtain an updated virtual scene. The updated virtual scene is used as the new adjusted virtual scene, and the pose of the object to be placed in the adjusted virtual scene is corrected at least once. The updated virtual scene obtained after the last iteration of pose correction is used as the target virtual scene, and the target virtual scene is displayed or responded to in response to the user's modification operation.

[0006] In some embodiments, obtaining the placement information of the object to be placed includes: The object identifier of the object to be placed is determined from the virtual data control instructions; Obtain the scene style information of the virtual environment scene; Style object retrieval information is generated based on the object identifier and the scene style information; Based on the style object retrieval information, a match is performed in the 3D model library to obtain the placement object information.

[0007] In some embodiments, the step of inputting the structural feature information, the visual information, and the visual information into a visual language model for data processing to obtain optimized placement data of the object to be placed in the virtual environment scene includes: When the virtual data control command includes reference position information of the object to be placed, the reference position information, the structural feature information, the structural feature information, and the visual information are input into the visual language model for data processing to obtain the optimized arrangement data near the reference position information; When the virtual data control command does not include the reference position information of the object to be placed, the structural feature information, the structural feature information and the visual information are input into the visual language model for data processing to obtain the optimized arrangement data that conforms to the position logic in the virtual environment scene.

[0008] In some embodiments, calculating the pose adjustment amount of the object to be placed in the adjusted virtual scene based on multiple physical placement constraints includes: Based on each of the physical placement constraints, the placement state of the objects to be placed in the adjusted virtual scene is matched one by one to obtain the state matching result. When the state matching result indicates that the object to be placed has a position error, an adjustment thrust is generated based on the position error information. When the state matching result indicates that the object to be placed has an orientation error, an adjustment rotation is generated based on the orientation error information; The pose adjustment amount is obtained based on the adjusted thrust and / or the adjusted rotation.

[0009] In some embodiments, the method further includes: The target virtual scene is analyzed and processed to obtain scene summary information of the target virtual scene; Responding to user inquiries; Based on the query information and the scenario summary information, a search and matching operation is performed in the layout rule knowledge base to obtain design rule information; The scenario summary information, the query information, and the design rule information are input into the question-answering model to obtain the response result data corresponding to the query information.

[0010] In some embodiments, the step of inputting the scene summary information, the query information, and the design rule information into the question-answering model to obtain the response result data corresponding to the query information includes: Confirm the target structural feature information and target visual information of the target virtual scene; The scene summary information, the query information, the design rule information, the target structural feature information, and the target visual information are input into the question-answering model to obtain the response result data corresponding to the query information.

[0011] In some embodiments, the method further includes: When the response result data includes information about newly placed objects, the new placement data of the newly placed objects in the target virtual scene is determined based on the information about newly placed objects; The newly placed object is generated in the target virtual scene based on the newly added layout data.

[0012] In some embodiments, the method further includes: Responding to a user's direct modification command to a target object in the target virtual scene; The target object in the target virtual scene is adjusted based on the direct modification command; Based on the adjusted target virtual scene, updated scene information is generated and saved.

[0013] To achieve the above objectives, a second aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the virtual environment data processing method as described in the first aspect.

[0014] To achieve the above objectives, a third aspect of this application provides a storage medium, which is a computer-readable storage medium storing a computer program that, when executed by a processor, implements the virtual environment data processing method described in the first aspect.

[0015] The virtual environment data processing method, electronic device, and storage medium proposed in this application include: First, responding to a user's virtual data control command in a virtual environment scene; then, when the virtual data control command indicates an update to the virtual environment scene, acquiring the placement information of the object to be placed, the structural feature information, and the visual information of the virtual environment scene, and inputting the structural feature information, the structural feature information, and the visual information into a visual language model for data processing to obtain optimized arrangement data of the object to be placed in the virtual environment scene; next, generating the object to be placed in the virtual environment scene based on the optimized arrangement data to obtain an adjusted virtual scene; second, calculating the pose adjustment amount of the object to be placed in the adjusted virtual scene based on multiple physical placement constraints, and performing pose correction on the object to be placed in the adjusted virtual scene based on the pose adjustment amount to obtain an updated virtual scene, using the updated virtual scene as a new adjusted virtual scene, and then performing at least one pose iteration correction on the object to be placed in the adjusted virtual scene; finally, using the updated virtual scene obtained after the last iteration of pose correction as the target virtual scene. This application's embodiments effectively solve the two major technical problems of physical unreliability and uncontrollable interaction caused by the 'black box' mode in existing AI scene generation technologies by combining intelligent prediction of visual language models with subsequent multi-constraint iterative optimization. The solution provided by this application does not take the idealized layout output by AI as the final result, but on this basis, it combines intelligent prediction of visual language models, multi-constraint physical iterative optimization, and user-led control in the loop to ensure that the final generated virtual environment data not only strictly follows physical laws in micro-details, but also achieves an efficient and transparent human-computer collaboration process in the macro-creation process. In the iteration process, through an iterative physical optimization engine, based on various physical placement constraints such as horizontal collision, vertical support, and room boundaries, the pose of objects in the scene is continuously corrected until a stable state is reached. This process can automatically detect and repair physical errors such as clipping, suspension, and insufficient support in the AI ​​layout, and achieve efficient and transparent human-computer collaboration in the macro-creation process, transforming the creation process from a rigid 'instruction-execution' mode to a flexible 'dialogue-collaboration' mode, thereby significantly improving the reliability of scene data, physical realism, and the user's immersive creation experience.

[0016] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description

[0017] Figure 1 This is a flowchart of a virtual environment data processing method provided in an embodiment of this application.

[0018] Figure 2 This is a flowchart of obtaining the placement information of an object to be placed, provided in another embodiment of this application.

[0019] Figure 3 This is a flowchart of another embodiment of the present application, which provides a method for obtaining optimized layout data of objects to be placed in a virtual environment scene.

[0020] Figure 4 This is a flowchart of calculating the pose adjustment amount of an object to be placed in an adjusted virtual scene, provided in another embodiment of this application.

[0021] Figure 5 This is a flowchart of the adjustment and optimization of a multi-physical placement constraint optimization algorithm provided in another embodiment of this application.

[0022] Figure 6 This is a schematic diagram of preparing and optimizing data according to another embodiment of this application.

[0023] Figure 7 This is a flowchart of responding to user queries provided in another embodiment of this application.

[0024] Figure 8 yes Figure 7 The flowchart for step 704.

[0025] Figure 9 This is a flowchart of generating a newly placed object, provided in another embodiment of this application.

[0026] Figure 10 This is a flowchart of a response to a direct modification instruction provided in another embodiment of this application.

[0027] Figure 11 This is a schematic diagram of the overall architecture of a virtual environment data processing system provided in another embodiment of this application.

[0028] Figure 12 This is a flowchart of object placement position prediction and optimization in a virtual environment data method provided in another embodiment of this application.

[0029] Figure 13This is a flowchart of a virtual environment data method for generating a scene from scratch, provided in another embodiment of this application.

[0030] Figure 14 This is a schematic diagram of the hardware structure of an electronic device provided in another embodiment of this application. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0032] It should be noted that although functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0034] Virtual reality (VR) technology provides an immersive interactive experience, creating a physically realistic and believable experience for users within a virtual environment. Currently, the creation of virtual environment scenes primarily relies on manual operation using specialized software, or automated layout through procedural content generation (PCG) and artificial intelligence (AI) driven methods. Among these, some emerging AI generation methods can directly output a complete 3D scene layout scheme using large models based on text or image prompts. These methods are typically a "one-click" black-box process; that is, after receiving high-level instructions, the AI ​​predicts and generates what it considers the optimal, idealized object arrangement, thus achieving a degree of automation in scene generation.

[0035] However, the layout results generated by the aforementioned AI-driven scene generation methods are merely idealized predictions based on data and logic, which cannot fully guarantee the physical realism and semantic rationality of the final scene. This can lead to problems such as objects clipping through each other, floating in mid-air, or being placed in positions that do not conform to physical common sense and functional logic. As a result, the reliability of the generated virtual environment data is low, which seriously damages the physical realism of the virtual scene and the user's immersive experience.

[0036] To improve the reliability of the generated virtual environment data, this application combines intelligent prediction of a visual language model with subsequent multi-constraint iterative optimization. This effectively solves the two major technical problems of physical unreliability and uncontrollable interaction caused by the 'black box' mode in existing AI scene generation technologies. The solution provided in this application does not take the idealized layout output by AI as the final result. Instead, it combines intelligent prediction of a visual language model, multi-constraint physical iterative optimization, and user-led control within the loop to ensure that the final generated virtual environment data not only strictly adheres to physical laws in microscopic details but also achieves [the desired effect] in the macroscopic creation process. It achieves an efficient and transparent human-machine collaboration process. During the iteration process, an iterative physics optimization engine continuously corrects the pose of objects in the scene based on various physical placement constraints such as horizontal collision, vertical support, and room boundaries until a stable state is reached. This process can automatically detect and repair physical errors in AI layout such as clipping, suspension, and insufficient support. Furthermore, it achieves efficient and transparent human-machine collaboration in the macro-creation process, transforming the creation process from a rigid 'instruction-execution' mode to a flexible 'dialogue-collaboration' mode, thereby significantly improving the reliability of scene data, physical realism, and the user's immersive creation experience.

[0037] The following describes the virtual environment data processing method, electronic device, and storage medium provided in the embodiments of this application. The virtual environment data processing method provided in the embodiments of this application is based on a client-server architecture. The client (VR application) is responsible for user interaction and scene rendering, while the server is responsible for core computing. This method can be applied to any server or computing processor with computing resources.

[0038] The virtual environment data processing method in the embodiments of this application will be described in detail below. (Refer to...) Figure 1 This is an optional flowchart of the virtual environment data processing method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps 101 to 105. It is also understood that this embodiment... Figure 1 The order of steps 101 to 105 is not specifically limited. The order of steps can be adjusted or some steps can be reduced or added according to actual needs.

[0039] Step 101: Respond to the user's virtual data control commands in the virtual environment scene.

[0040] Step 101 will be described in detail below.

[0041] In some embodiments, the starting point of the entire data processing flow is responding to virtual data control commands issued by the user in the virtual environment (i.e., the VR environment). Here, virtual data control commands refer to input signals generated by the user within the immersive virtual environment through natural interaction methods, such as voice commands or spatial pointing and grasping operations using VR controllers (e.g., controller rays), to express their creative intent. These commands serve as the trigger for all subsequent scene updates and layout generation operations.

[0042] The system first identifies and classifies the user's virtual data control commands by calling a large language model that supports audio input. The commands are mainly divided into four types: Retrieve (the user wants to retrieve a new object model), Place (the user confirms that they want to place the retrieved object), Generate (the user wants to generate a complete scene from scratch or based on a theme), and Reply (the user asks a question to the system and engages in conversational interaction).

[0043] When the system recognizes the "search" intent, it uses the object name described by the user's voice (e.g., "a wooden table") and the style information of the current scene to search for the most matching model in the 3D model library using a text-image matching model (e.g., CLIP). The specific process is described below.

[0044] 3D Model Library Preparation: The system uses 3D-Front as its 3D model library, which contains more than 16,000 objects. Each object model has a corresponding front preview image. Before the system starts running, the vector encoded from the preview image of each model and the storage location of the model are stored in the Faiss vector database and a JSON file containing the location information for subsequent image-text matching and retrieval.

[0045] Style cues enhancement: Combine the object name in the object description (e.g., "a wooden table") with the style information of the current scene to form an object description with style information (e.g., "a Chinese-style wooden table"), which is used to retrieve object models with the same style as the existing scene.

[0046] Image-text matching retrieval: The style-enhanced text information is transmitted to the model retrieval server. The model retrieval server then uses the CLIP model to vectorize the style-enhanced text information, and uses the Faiss vector database to calculate the similarity and find the model preview image with the highest vector similarity to the current text information, thus obtaining its model-related information.

[0047] Store to Model Cache List: After obtaining the model-related information, including the URL of the model's compressed file, the model name, the model's top view, and the pointing position provided by the user using the controller's raycast when issuing commands, are all stored in a temporary model cache list. This step can be repeated to batch collect multiple objects.

[0048] Step 102: When the virtual data control command characterizes the update of the virtual environment scene, the placement information of the object to be placed and the structural feature information and visual information of the virtual environment scene are obtained. The structural feature information, structural feature information and visual information are input into the visual language model for data processing to obtain the optimized arrangement data of the object to be placed in the virtual environment scene.

[0049] Step 102 is described in detail below.

[0050] In some embodiments, when the system parses the user's virtual data control instructions and determines that the intention is to update the virtual environment scene (e.g., "place" or "generate" a new object), the system enters a multimodal information preparation and processing phase. First, the system acquires the object's own placement information (such as 3D model data). This 3D model data primarily includes the object's .obj model, the object's name, and a scaled top view of the object in the empty scene, where the scale is used to determine the object's proportion within the scene.

[0051] Simultaneously, the system collects two key types of information about the current virtual environment scene: the first is structural feature information, which typically refers to a structured data file (such as Scene.json) used to accurately describe the identifiers, parent-child relationships, and physical parameters of all existing objects in the scene; the second is visual information, namely, real-time rendered images of the scene, such as a real-time bird's-eye view. This bird's-eye view includes planar coordinate scales to indicate the planar coordinates of each location within the view.

[0052] Subsequently, the system inputs these three types of information into a visual language model trained for a specific task. This model can comprehensively understand the structure and visual context of the scene, thereby processing the data and outputting optimized placement data. This data is a structured set of parameters that defines the initial position, orientation, and spatial relationship constraints of the objects to be placed in the current scene that best conform to logical and aesthetic considerations.

[0053] Understandably, if the identified virtual data control command is an intent to "generate," the system invokes a large language model (i.e., a visual language model) to perform overall planning for scene generation. This model is assigned the role of a "VR scene architect," combining Retrieval Augmented Generation (RAG) technology to retrieve relevant design knowledge from a pre-set layout rule knowledge base. Then, based on the user's generation prompts (such as "a modern bedroom"), it outputs a detailed scene plan strictly following JSON format, including a list of objects to be retrieved and text describing how these objects are placed.

[0054] The following describes how to obtain the placement information for the object to be placed.

[0055] Reference Figure 2 Obtain the placement information of the object to be placed, including the following steps 201 to 204.

[0056] Step 201: Determine the object identifier of the object to be placed from the virtual data control instructions.

[0057] Step 202: Obtain scene style information of the virtual environment scene.

[0058] Step 203: Generate style object retrieval information based on object identifiers and scene style information.

[0059] Step 204: Match the style object retrieval information in the 3D model library to obtain the placement object information.

[0060] Steps 201 to 204 are described in detail below.

[0061] In some embodiments, when the system parses the user's virtual data control commands and determines that the user's intention is to update the virtual environment scene (e.g., "place" or "generate" a new object), the system first parses and extracts key information from the user's virtual data control commands. The core of this step is determining the object identifier of the object to be placed. This identifier specifically refers to the object name described by the user in natural language, such as "a wooden table" in the voice command "give me a wooden table." This object identifier is the most original and fundamental query basis for subsequent model retrieval processes.

[0062] Next, similar to the "retrieve" command mentioned above, to ensure that newly added objects are aesthetically consistent with the existing environment, the system needs to obtain the scene style information of the current virtual environment. This information is a variable describing the overall aesthetic characteristics of the scene, such as "modern minimalism" or "Chinese style." It can be derived by analyzing the current scene using a visual language model and serves as an important contextual parameter to guide subsequent model retrieval.

[0063] The system then integrates the information obtained in the first two steps to generate a more precise and context-aware retrieval command. Specifically, the system combines object identifiers with scene style information to generate style object retrieval information. For example, if the object identifier is "a wooden table" and the scene style information is "Chinese style," then the generated style object retrieval information would be "a Chinese-style wooden table," which will serve as the final query text input into the retrieval model.

[0064] Then, the system uses the style object retrieval information generated in the previous step to perform a matching operation in a pre-prepared 3D model library. This matching process uses a text-image matching model (such as the CLIP model) to vectorize the text information, and then performs similarity calculations in a vector database (such as Faiss) to find the model preview image that best matches the text description, thereby obtaining the specific data of the 3D model. The final object placement information is structured data containing the URL of the model compressed package, the model name, etc., which provides all the necessary data support for the subsequent instantiation of the object in the scene.

[0065] Through steps 201 to 204 above, these steps collectively constitute an intelligent, context-aware model retrieval method. The beneficial effect of this method lies in its significant improvement in the accuracy and relevance of model retrieval by combining the user's direct intent (object identifiers) with the global environment of the scene (scene style information). Compared to simple keyword matching, this method ensures that the retrieved objects are stylistically highly coordinated and unified with the existing scene, thus guaranteeing the overall aesthetic quality of the virtual environment from the creative source, and enhancing the intelligence of automated layout and the harmony of the final effect.

[0066] Next, the client sends the scene information (JSON and top view) along with the information of the objects to be placed in the object cache (or "generating" the list and description of objects to be planned) to the backend visual language model. This model is given the role of a "professional 3D room layout assistant," and processes the data in different working modes depending on whether there is a "prior location" provided by the user, as described below.

[0067] Reference Figure 3 The structural feature information, structural feature information and visual information are input into the visual language model for data processing to obtain the optimized layout data of the object to be placed in the virtual environment scene, including the following steps 301 to 302.

[0068] Step 301: When the virtual data control command includes reference position information of the object to be placed, the reference position information, structural feature information, and visual information are input into the visual language model for data processing to obtain optimized arrangement data near the reference position information.

[0069] Step 302: When the virtual data control command does not include the reference position information of the object to be placed, the structural feature information, structural feature information and visual information are input into the visual language model for data processing to obtain optimized layout data that conforms to the position logic in the virtual environment scene.

[0070] Steps 301 to 302 are described in detail below.

[0071] In some embodiments, when the user provides explicit spatial guidance (the user provides reference location information), the reference location information here refers to a specific coordinate or approximate area specified by the user in the virtual environment scene through the ray pointing of the VR controller, etc., as the a priori position for placing the object. The model is instructed that its primary goal is to "place the object at or near the reference position" and to make intelligent fine adjustments based on this, such as automatically adhering to the support surface and avoiding collisions.

[0072] When the user's virtual data control commands include this reference position information, the system provides this information, along with the structural features and visual information of the current scene, as input to the visual language model for data processing. In this mode, the visual language model is explicitly instructed that its primary task is to place the object to be placed at the location indicated by the reference position information or within a reasonable range in its immediate vicinity, and then perform intelligent fine-tuning based on this, such as automatically attaching it to the supporting surface or avoiding collisions with other objects, ultimately obtaining optimized placement data near the reference position information.

[0073] When users do not provide specific spatial guidance and instead leave the layout decisions entirely to the AI, the model is instructed to "determine the most logical and aesthetically pleasing location for new objects without any user prompts," thus granting it greater creative freedom.

[0074] When the user's virtual data control commands do not include reference position information for the objects to be placed, the system only inputs the structural features and visual information of the scene into the visual language model for data processing. In this reference-free working mode, the visual language model is given greater creative freedom. It is instructed to autonomously determine the most logical and aesthetically pleasing position for the new object by comprehensively analyzing the existing layout, functional zoning, and aesthetic style of the scene without any user prompts. Therefore, the optimized placement data ultimately output by the model is a result generated entirely by its internal knowledge within the virtual environment, conforming to positional logic.

[0075] Through steps 301 and 302 above, a highly efficient human-computer collaborative creation paradigm is constructed by providing two different working modes, embodying the core concept of "human in the loop." When the user has a clear intention, it ensures that the AI ​​accurately executes the user's spatial guidance, preserving the user's ultimate control. Conversely, when the user desires creative suggestions or rapid layout, the AI ​​can act as an automated designer, fully leveraging its intelligent planning advantages. This dual-mode mechanism enables this method to adapt to diverse creative needs, achieving an ideal balance between user-led fine-grained control and AI-driven automated generation.

[0076] All the AI ​​models described above are guided by carefully designed system prompts. These prompts define clear roles, tasks, input / output formats, and logical rules that the models must follow (such as "large furniture must be placed on the floor" and "parent object ID is only used for physical support relationships"). Through this prompting engineering, general-purpose AI models are transformed into professional scene design tools, ensuring the stability and professionalism of their output. The model ultimately predicts the optimal parameters for each object to be placed in the scene and returns them in a structured JSON format. This JSON format defines the AI ​​model's prediction output, which includes the following function instructions: batch_predictions, a list containing prediction information for all objects to be placed; predicted_object, containing the core physical parameters of an object; objectId, a unique string identifier for the object; parentId, the objectId of the parent object supporting this object (e.g., "floor" represents the floor); position, the object's two-dimensional coordinates [x, z] on the XZ plane; rotation, the object's rotation angle [x, y, z]; scale, the object's scaling ratio [x, y, z]; object_name, a natural language description of the object; point_towards, spatial relationship constraints, defining the objectId of another object that the object needs to face; against_wall, spatial relationship constraints, defining whether the object needs to be against a wall; and adjacent, spatial relationship constraints, defining the distance that the object needs to maintain from another object (target).

[0077] Step 103: Generate objects to be placed in the virtual environment scene based on the optimized layout data, and obtain the adjusted virtual scene.

[0078] Step 103 will be described in detail below.

[0079] In some embodiments, the system instantiates and initially places the object to be placed in a virtual environment scene based on the optimized layout data generated by the visual language model in the previous step. The scene obtained after this process is called the adjusted virtual scene. It should be noted that this adjusted virtual scene is a temporary, intermediate scene. Although it reflects the intelligent prediction of the object layout by the artificial intelligence model, it has not been rigorously verified by physical laws, and therefore may have potential physical inconsistencies such as clipping between objects or objects suspended in mid-air.

[0080] During the generation process, the system performs topological sorting based on the parent-child dependencies in the prediction results, ensuring that the supporting object (parent object) is loaded before the supported object (child object). Subsequently, all object models are loaded asynchronously in sequence to the initial position predicted by the AI, and a data structure containing its physical and constraint information is created for each object to be optimized.

[0081] Step 104: Based on multiple physical placement constraints, calculate the pose adjustment amount of the object to be placed in the adjustment virtual scene, and perform pose correction on the object to be placed in the adjustment virtual scene based on the pose adjustment amount to obtain an updated virtual scene. Use the updated virtual scene as the new adjustment virtual scene, and then perform at least one pose iteration correction on the object to be placed in the adjustment virtual scene.

[0082] Step 104 is described in detail below.

[0083] In some embodiments, after generating the initial adjusted virtual scene, the system initiates a core, iterative physics optimization process. This process is based on multiple physical placement constraints, which are a set of preset rules used to simulate the physical laws and spatial logic of the real world, such as seven categories of constraints including horizontal collision avoidance, vertical support relationships, room boundary restrictions, object proximity and orientation relationships, etc. In each iteration, the system traverses the objects to be placed, checking whether they violate any of the above constraints. If a violation is found, the system calculates a pose adjustment amount to correct the error, which is specifically manifested as a "push" or a rotation operation applied to the object. The system then corrects the pose of the object based on this pose adjustment amount, thereby obtaining an updated virtual scene with improved physical plausibility. Crucially, this process is not completed in one step. The system uses this updated virtual scene as input for a new iteration (i.e., a new adjusted virtual scene), repeatedly performing constraint checks and pose corrections—this is iterative pose correction. This cycle continues until there are no more constraint conflicts in the scene, or the preset maximum number of iterations is reached.

[0084] Before performing pose iteration correction, it is necessary to prepare optimization data for the objects to be optimized in the scene, including 3D bounding boxes, 2D bounding boxes, etc. After preparation, a state iteration loop is entered (for example, up to 400 rounds).

[0085] Next, in each iteration, all objects to be optimized (i.e., objects to be placed) are traversed, and a resultant force is calculated for each object based on seven different physical placement constraints. A violation of each constraint is quantified as a pose adjustment amount (i.e., a thrust or a rotation operation) applied to the object, as described below.

[0086] Reference Figure 4Based on multiple physical placement constraints, the pose adjustment amount of the object to be placed in the virtual scene is calculated, including the following steps 401 to 404.

[0087] Step 401: Based on each physical placement constraint, perform placement state matching on the objects to be placed in the virtual scene to obtain the state matching result.

[0088] Step 402: When the state matching result indicates that the object to be placed has a position error, an adjustment thrust is generated based on the position error information.

[0089] Step 403: When the state matching result indicates that the object to be placed has an orientation error, generate an adjustment rotation based on the orientation error information.

[0090] Step 404: Obtain the pose adjustment amount based on adjusting the thrust and / or adjusting the rotation.

[0091] Steps 401 to 404 are described in detail below.

[0092] In some embodiments, at the beginning of each round of iterative optimization, the system performs a comprehensive state assessment of the objects to be placed in the virtual scene. Here, physical placement constraints refer to a series of predefined rules used to ensure the physical realism and semantic rationality of the scene, such as horizontal collision, vertical support, room boundaries, and proximity relationships. The system measures the current state of the object against each physical placement constraint; this process is called placement state matching. The result of placement state matching, i.e., the state matching result, clearly indicates whether the object currently violates a specific constraint rule and quantifies the degree of violation (such as collision depth, distance beyond the boundary, etc.), as described below.

[0093] Horizontal collision optimization: Detects whether the 2D bounding boxes (Rect) of objects on the same supporting plane overlap. This 2D bounding box is a rectangle obtained by orthogonally projecting the complete 3D bounding boxes of the objects onto the horizontal (XZ) plane. If they overlap, a repulsive pushing force is calculated along the centroid direction of the two objects' bounding boxes based on the overlap depth, pushing the objects apart until the minimum separation distance is met.

[0094] Vertical collision optimization: Detect whether the bottom of an object's 3D bounding box penetrates the top of its support (parent object). If it does, calculate an upward thrust based on the penetration depth to push the object back to a very small safe distance above the support surface.

[0095] Room boundary optimization: Detects whether the bounding box of an object exceeds the preset room boundary. If it does, a pushing force is calculated based on the excess distance to push the object back into the boundary, preventing the object from passing through walls or being placed outside the scene.

[0096] Support optimization: Calculate the overlap ratio of the projected areas of the supported object and the support on the XZ plane. If this ratio is lower than a preset threshold (e.g., 90%), calculate a centripetal thrust pointing towards the center of the support to prevent the object from "hanging" on the edge and ensure its physical stability.

[0097] Orientation Optimization: This is a rotation operation. Based on the AI-predicted point_towards constraint, the required rotation angle to orient the object toward the target is calculated, and a small portion of the rotation is applied in each iteration to achieve a smooth turn. This rotation is centered on the bounding box of the object, avoiding displacement caused by inconsistent model anchor points.

[0098] Wall-alignment optimization: This is a combined operation of position and rotation. The rotation is completed in one step during object creation, using raycasting to detect the wall normal and position the object so that it faces away from the wall. The position is calculated iteratively to generate a force that pushes the object towards the wall until the distance to the wall is within a preset wall-alignment tolerance range.

[0099] Proximity optimization: Based on the proximity constraint predicted by AI, the difference between the current distance between objects and the target distance is calculated. If the distance is too far, a horizontal force of mutual attraction is generated; if the distance is too close, a horizontal force of repulsion is generated, until the distance between objects is within the tolerance range of the target distance.

[0100] When the state matching result obtained in the previous step indicates that the object to be placed has a positional error, the system will initiate a correction mechanism. Here, a positional error refers to any violation of physical placement constraints related to spatial location, such as overlapping objects, objects exceeding room boundaries, or objects failing to be stably placed on their supports. Based on the positional error information (such as overlap depth, penetration distance, etc.) contained in the state matching result, the system will calculate and generate a vector, i.e., an adjustment thrust, using a preset algorithm. The direction and magnitude of this adjustment thrust are carefully designed to push the object in a direction that can reduce or eliminate the positional error.

[0101] Similar to handling positional errors, the system also processes objects with incorrect orientations when state matching indicates they are facing the wrong direction. Here, an orientation error specifically refers to an object's rotational posture failing to meet certain semantic constraints; for example, the AI-predicted constraint that "a chair needs to face a table" is not satisfied. Based on this orientation error information (such as the angular difference between the current orientation and the target orientation), the system calculates and generates an adjustment rotation to correct the orientation. This adjustment rotation is a specific rotational operation parameter used to guide the object smoothly towards its preset target direction.

[0102] Then, the system integrates all the corrections calculated for a single object in a single iteration. Specifically, the system accumulates all the adjustment thrusts generated due to different positional errors to form a resultant force, and combines this with the adjustment rotations generated due to orientation errors to finally obtain a comprehensive pose adjustment. This pose adjustment is the sum of all pose transformations that the object needs to perform in this iteration; it includes both translation in position and possibly rotation in attitude, providing a precise execution basis for subsequent unified pose correction of the object.

[0103] Through steps 401 to 404 above, a systematic and modular method for calculating pose adjustment is established. This method decomposes the complex physical optimization problem into a clear process of "state matching" and "type-specific error handling" (position and orientation), enabling the system to systematically identify and quantify all constraint violations. By generating adjustment thrust and adjustment rotation for different types of errors and ultimately unifying them into a comprehensive pose adjustment, this method ensures the accuracy, stability, and controllability of the optimization process. It can simultaneously handle multiple complex constraint conflicts, thereby efficiently guiding the virtual scene to converge to a physically reasonable stable state.

[0104] In each iteration, all calculated horizontal thrusts are summed and then applied uniformly to the object to update its position in the XZ plane. Vertical thrusts and rotation operations are handled independently.

[0105] Finally, update the boundary information of all objects and check if the scene has reached a stable state (i.e., no object is generating thrust or rotation). If the scene is stable or the maximum number of iterations has been reached, the loop terminates.

[0106] Step 105: Use the updated virtual scene obtained after the last iteration of pose correction as the target virtual scene, and display the target virtual scene or respond to the user's modification operation.

[0107] Step 105 is described in detail below.

[0108] In some embodiments, after the above pose iteration correction process terminates, the system performs a final physical rationality check and formally determines the updated virtual scene obtained after the last iteration, which is stable and conforms to all physical constraints, as the target virtual scene. This target virtual scene is the final output of the entire data processing method; it is a high-quality, highly reliable virtual environment that embodies the intelligent layout of the artificial intelligence model and has undergone rigorous physical law verification.

[0109] Then the target virtual scene is displayed to the user for viewing, or it is adaptively modified in response to the user's modification operation.

[0110] In addition, after the optimization process is complete, the system performs a final physical constraint verification on each object in the scene to check for serious horizontal collisions, insufficient support, and other issues. Any object that does not meet basic physical requirements will be automatically removed from the scene.

[0111] Then, the scene state is updated and recorded: After a final physical plausibility check and removal of unreasonable objects, the scene reaches a final stable state. At this point, the final physical parameters of all valid objects in the system are formatted and updated to the Scene.json file. This file has a dual key role: serving as structured input for the AI ​​model, and recording the precise state of all objects in the scene, which will be used as structured data alongside the top-down view of the scene. Figure 1 Both are input into the visual language model. This allows the AI ​​model to better combine visual and spatial logic, providing accurate contextual judgments for subsequent scene summarization generation, intelligent question answering, and new object layout prediction. As a dynamic object list for the client, this file clearly identifies all objects added by the user using this method. This enables client applications (such as systems built in Unity) to distinguish these "added objects" from the game objects inherent in the scene, thus allowing them to accurately operate only on user-added objects when performing interactive operations such as "delete" or "delete all."

[0112] The Scene.json file is a JSON array. Each object in the array represents a valid object in the scene, including: objectId, a unique string identifier for the object; parentId, the objectId of the parent object supporting this object; position, the final 2D coordinates of the object on the XZ plane after optimization [x, z]; rotation, the final rotation angle of the object after optimization [x, y, z]; and scale, the scaling ratio of the object [x, y, z].

[0113] Reference Figure 5 This is a flowchart illustrating the adjustment and optimization process of a multi-physics placement constraint optimization algorithm provided in an embodiment of this application. Figure 5The diagram shows a detailed flowchart of the multi-constraint optimization engine provided in this application, illustrating the core algorithm for transforming AI-predicted idealized layouts into physically realistic layouts. The process includes: preparing optimization data, starting with receiving the AI-predicted layout results and creating an optimization object containing the physical and constraint information for each object to be placed, while simultaneously calculating the bounding box information for each object for subsequent optimization calculations; initiating optimization and iteration loops, entering an iteration loop. At the beginning of each iteration, all thrusts are first reset; all thrusts are calculated in parallel, which is the core of the optimization. The system calculates thrusts generated by seven different constraints in parallel, including horizontal collision thrusts, vertical support thrusts, wall thrusts, proximity thrusts, and boundary thrusts. All these thrusts are accumulated to form a resultant force acting on the object; application and updating, applying the calculated resultant force to all thrusts and updating its position. Simultaneously, rotations caused by orientation optimization are handled independently. After the transformation is completed, object information (such as bounding boxes) is updated to prepare for the next iteration; convergence or reaching the maximum loop, checking whether the scene has converged (i.e., all thrusts approach zero). If stability is achieved, or the number of iterations reaches the preset limit, the optimization is complete. Physical constraint check and termination: After optimization, the system performs a final physical constraint check. If all objects meet the constraints, the process ends; if any objects do not meet the conditions, these unreasonable objects are deleted to ensure the stability and realism of the final output scene.

[0114] Reference Figure 6 This is a schematic diagram illustrating the preparation of optimized data provided in an embodiment of this application. Figure 6 As shown in the attached diagram, this figure visually illustrates how a 3D object is simplified into a bounding box for various physics calculations. First, for each 3D object, the system obtains its complete 3D bounding box by rendering a mesh, as shown in the middle image. This 3D bounding box will be used to calculate vertical collisions and constraints requiring height information for support relationships. Second, by orthogonally projecting the 3D bounding box onto the horizontal (XZ) plane, a 2D bounding box (Rect) is obtained, as shown in the right image. This 2D bounding box is specifically used for efficiently calculating planar constraints such as horizontal collisions and proximity distances. This data preparation process from 3D to 2D is a core method for balancing complex physics simulations with computational efficiency.

[0115] Reference Figure 7 The virtual environment data processing method also includes the following steps 701 to 704.

[0116] Step 701: Analyze and process the target virtual scene to obtain scene summary information of the target virtual scene.

[0117] Step 702: Respond to the user's query information.

[0118] Step 703: Based on the query information and scenario summary information, perform retrieval and matching in the layout rule knowledge base to obtain design rule information.

[0119] Step 704: Input the scenario summary information, query information, and design rule information into the question-answering model to obtain the response result data corresponding to the query information.

[0120] Steps 701 to 704 are described in detail below.

[0121] In some embodiments, after a scene is updated—that is, after a target virtual scene is generated—the system initiates a scene understanding process. The system calls a multimodal visual language model for analysis to comprehensively analyze the visual information (such as a top-down view) and structured data (such as Scene.json) of the current scene, resulting in a "style" variable and a "scene summary." The "style" variable (selected from a predefined style list, such as "modern minimalist") is used for subsequent model retrieval to ensure style consistency; the "scene summary" serves as core contextual information, input into the next step of the intelligent question-answering system (e.g., "This room is a restaurant scene. In the center of the restaurant is a black dining table, surrounded by four silver metal chairs. A table lamp illuminates the central area, including the table and chairs. In the corner of the room is a white sofa, next to which is a coffee table"). This summary information is a descriptive natural language text that accurately summarizes the layout, style, and spatial relationships of the core objects in the current scene, providing crucial contextual information for subsequent intelligent interaction.

[0122] The system then continuously listens for and parses the user's virtual data control commands. When it recognizes the user's intent as "Reply," indicating that the user is asking a question, the system responds to the user's inquiry. This inquiry is a specific question posed by the user in natural language, such as "This room looks a bit empty, what should I add?" This step is the starting point for triggering the entire context-aware question-and-answer and execution process.

[0123] To make the model's responses more professional and targeted, the system employs a Retrieval Augmentation (RAG) technique. The system combines the user's query information and the generated scene summary information as query conditions, performing semantic retrieval matching within a pre-defined layout rule knowledge base (vector database). This knowledge base is a vector database storing a large amount of professional interior design principles, spatial layout examples, and other knowledge. The most relevant knowledge fragments obtained after the retrieval and matching process constitute the design rule information, which will be injected into the language model as important external knowledge in the next step. All of this retrieved content will be provided as context to the AI ​​model, making its responses more professional and accurate.

[0124] The system then integrates all contextual information and invokes the question-answering model to generate the final response. Specifically, the system provides the question-answering model with scenario summary information, user query information, and retrieved design rule information as input. After receiving this comprehensive context, the model generates a structured response result. This data is typically a JSON object, which not only contains the text answering the user's question in natural language but may also contain directly executable scenario modification instructions parsed from the dialogue content.

[0125] Among them, reference Figure 8 The process involves inputting scenario summary information, query information, and design rule information into the question-answering model to obtain the response result data corresponding to the query information, including the following steps 801 to 802.

[0126] Step 801: Confirm the target structural feature information and target visual information of the target virtual scene.

[0127] Step 802: Input the scene summary information, query information, design rule information, target structural feature information, and target visual information into the question answering model to obtain the response result data corresponding to the query information.

[0128] Steps 801 to 802 are described in detail below.

[0129] To provide the question-answering model with the most comprehensive and accurate scene context, the system verifies the underlying data of the target virtual scene before generating the final response. This step aims to obtain two core types of raw scene data: first, target structural feature information, which specifically refers to a complete structured data file (i.e., Scene.json) that records the precise physical parameters and parent-child relationships of all valid objects in the scene; second, target visual information, which typically refers to a real-time image that can intuitively reflect the overall layout of the current scene, such as a top-down view of the scene. These two pieces of information together constitute the "ground reality data" regarding the state of the target virtual scene.

[0130] Next, all the information prepared in the preceding process, including high-level scene summary information, direct user queries, external design rule information retrieved from the layout rule knowledge base, and confirmed underlying target structural feature information and target visual information, are packaged together to form a comprehensive information package. This complete contextual information package, containing multi-level and multi-modal data, is then input into the question-answering model, which performs the final inference and generation to obtain the response result data corresponding to the query information.

[0131] Through steps 801 to 802 above, a comprehensive input with rich information layers and diverse data modalities is constructed, providing the question-answering model with an unprecedented depth of context. This ensures that the model's decisions are not only based on high-level text summarization and external knowledge, but also on the underlying structured data of the scene and real-time visual information as factual basis for dual anchoring. This multi-level, multi-modal context aggregation strategy greatly improves the model's understanding of the scene state in depth and accuracy, effectively reducing the risk of the model generating "illusionary" responses that do not conform to the actual scene situation, thereby ensuring that the generated response results data have high reliability, accuracy and scene relevance.

[0132] The system uses sophisticated prompting engineering to instruct the question-answering model to output a strictly formatted JSON object, rather than simple text. This JSON object contains three key fields: "answer" (a Chinese answer in natural language), "add_objects" (a boolean value indicating whether the user's question implies an intention to add objects), and "objects" (a list of names of objects to be added).

[0133] The client then parses the returned JSON. The system then plays the content of the "answer" field to the user using the text-to-speech (TTS) module.

[0134] Reference Figure 9 The virtual environment data processing method also includes the following steps 901 to 902.

[0135] Step 901: When the response result data includes information about newly placed objects, determine the new placement data of the newly placed objects in the target virtual scene based on the information about newly placed objects.

[0136] Step 902: Generate newly placed objects in the target virtual scene based on the newly added layout data.

[0137] Steps 901 to 902 are described in detail below.

[0138] In some embodiments, this step addresses the specific case where the question-answering model's response contains executable creation instructions. The response result data is structured information output by the question-answering model. When this data includes information about new objects—typically referring to an explicit Boolean flag (i.e., "add_objects" is true) and a list of object names to be added—the system triggers an automated scene addition process. Based on this information, the system invokes its core intelligent layout prediction module within the context of the current target virtual scene. This module calculates the optimal placement data for each new object to be added, including precise physical parameters such as the new object's required position and orientation.

[0139] After obtaining the new layout data calculated for all newly placed objects, the system will perform scene instantiation update. This step loads and generates the corresponding 3D models of the newly placed objects in the target virtual scene based on the new layout data. These newly generated objects will be initially placed in the positions specified by the new layout data and then seamlessly integrated into the scene. They will usually undergo a subsequent multi-constraint physics optimization process to ensure that their relationship with existing objects in the scene also conforms to physical realism requirements.

[0140] Through steps 901 to 902 above, the two stages of intelligent question answering and scene creation are connected, constructing a complete interactive closed loop from "dialogue understanding" to "scene execution". This makes the question answering system no longer just a passive information query tool, but an intelligent assistant that can actively participate in and execute creative tasks. Users can seamlessly propose and implement scene modifications and content additions during natural language dialogue with the system, which greatly improves the smoothness, intuitiveness and efficiency of the creation process, thereby realizing a more advanced and in-depth human-computer collaborative interaction mode.

[0141] Through steps 701 to 704 above, an intelligent question-answering system with deep scene understanding and professional knowledge enhancement is constructed, making the interaction no longer a simple command response, but upgraded to an intelligent dialogue with context awareness. By combining real-time scene summaries and external professional design knowledge bases, the system can not only "understand" the current scene and accurately answer user questions, but also provide professional suggestions that conform to design principles. Furthermore, it can understand and execute the user's creative instructions from the dialogue, forming a complete intelligent interaction closed loop from analysis, understanding to execution, which greatly improves the intelligence level and interaction depth of human-computer collaborative creation.

[0142] In addition, refer to Figure 10 The virtual environment data processing method also includes the following steps 1001 to 1003.

[0143] Step 1001: Respond to the user's direct modification command on the target object in the target virtual scene.

[0144] Step 1002: Adjust the target objects in the target virtual scene based on the direct modification command.

[0145] Step 1003: Generate updated scene information based on the adjusted target virtual scene and save it.

[0146] Steps 1001 to 1003 are described in detail below.

[0147] In some embodiments, after modifying the scene using a large language model and a visual language model, the user can use the grab button on the VR controller to modify objects in the existing scene. When the user makes more refined and subjective creative adjustments based on the AI-generated target virtual scene, the system will respond accordingly. This direct modification command is a low-level interactive command distinct from high-level voice commands. It typically refers to the user directly selecting and applying the command to a specific target object in the scene using physical operations such as the grab button on the VR controller, thus initiating a manual modification intention. This ensures that the user retains the final editing rights over any element in the scene after the automated layout process.

[0148] After receiving a direct modification command from the user, the system will adjust the selected target object accordingly. These adjustments are entirely user-controlled and may include: changing the spatial position of the target object in the scene by physically moving the controller; adjusting the orientation of the target object by rotating the controller; or executing a delete command to completely remove the target object from the scene. This step enables manual fine-tuning or correction of the AI ​​layout results.

[0149] In step 903 of some embodiments, after the user completes the manual adjustment of the target object, the system initiates a background scene information synchronization process. This process actively collects the latest state of the modified object in the scene, including its final position, orientation, and support relationships (i.e., using rays to extend downwards from the center of each object to collect information about the parent object supporting it), and generates updated scene information based on this adjusted information. This updated scene information is then written to and saved in the scene's structured data file (Scene.json), ensuring that any manual modifications by the user are persistently recorded, facilitating further modifications such as AI prediction and placement based on the user's changes to the scene.

[0150] Through steps 1001 to 1003 above, a truly meaningful "human-in-the-loop" human-machine collaborative creation mode is constructed. By giving users the ability to directly intervene manually after AI-automated generation, the user's final creative leadership and control are ensured. Users can use AI to complete tedious batch layout work, and then make personalized artistic fine-tuning through manual adjustments. More importantly, the automatic update and saving mechanism of scene information ensures that the user's "manual modification" and AI's "automated creation" can be seamlessly connected and alternated in the same scene, forming an efficient, flexible, and continuously iterative hybrid interactive process for creative results.

[0151] Reference Figure 11 This is a schematic diagram of the overall architecture of a virtual environment data processing system provided in an embodiment of this application. Figure 11The diagram illustrates an intelligent agent system with intent recognition as its core scheduler. The system employs a typical client-server architecture and integrates multiple AI models to form an agentic system. This includes: a user input layer, where users input information in the VR environment via voice audio commands and controller ray pointing; an intent routing layer, where the intent recognition module acts as the system's overall scheduler, calling a large language model that supports audio input to parse the user's speech and route their intent to one of four main workflows: Retrieve, Place, Generate, or Reply; and a scene generation workflow, including: Retrieve, where commands trigger a model retrieval service in the model retrieval server, which uses the CLIP model for image-text matching to find the object described by the user from the model library. The retrieved OBJ model compressed files are stored in the model cache list; Generation: The command first triggers the planning module, which performs scene planning by calling a large language model and outputs an object list and layout description; Placement: Whether it is a user-triggered "Placement" or a subsequent step in the "Generate" workflow, it will eventually call the Placement or Generation module. These modules receive cached object information and the current scene snapshot as input by calling a visual language model, perform multimodal inference, and output accurate layout predictions; Question-and-answer interaction workflow: Reply: The command triggers the Reply module, which receives user questions and scene summary information generated by the scene analysis module by calling a visual language model, and provides context-aware answers. If the answer contains a command to add objects, it will further call the Add Objects module to add new objects to the scene. The architecture of the Add Objects module is the same as that of the Placement module; Feedback: The text output of the Reply module is converted into audio and played to the user through a TTS (Text-to-Speech) service; Scene analysis and closure: Whenever the scene changes, the scene analysis module calls a visual language model to analyze the new scene and extract style variables and scene summary information. Style information is fed back to the model retrieval server to ensure style consistency in subsequent retrievals; summary information provides context for the question answering system, forming a closed loop of continuous optimization.

[0152] Reference Figure 12 This is a flowchart illustrating the prediction and optimization of object placement positions in a virtual environment data method provided in this application embodiment. For example... Figure 12The diagram illustrates the complete object position prediction and optimization algorithm flow of this application. This flow is the core of the invention's achievement of physical realism and semantic rationality in scene layout generation. The flow includes: Input: The flow receives two core pieces of information: object model cache information, including the name of the object to be placed, its top view, and the user-selectable prior position; and existing scene information, including the scene's JSON file and top view. AI Prediction: After receiving these multimodal inputs, the visual language model batch-predicts the position and spatial constraints of each object, forming a formatted object layout representation output. Instantiation and Optimization: After the object is initially placed in the new scene, it immediately enters the physical and semantic optimization stage. This is a local computation process based on an iterative push model, applying seven constraints (such as collision, support, proximity, etc.) to gradually adjust the object's position and orientation in a loop of up to 400 rounds. Final Verification: After optimization, the system performs a final physical rationality check, verifying through the algorithm whether each object meets the most basic physical laws (such as no collision and sufficient support). Any object that does not meet the conditions will be automatically removed to ensure the stability and realism of the final output scene.

[0153] Reference Figure 13 This is a flowchart illustrating the process of generating a scene from scratch in a virtual environment data method provided in this application embodiment. For example... Figure 13 The diagram illustrates the scene generation process from scratch in this application, demonstrating the application of Retrieval-Enhanced Generation (RAG) technology. The process includes: In the planning phase, the user inputs a high-level generation instruction (e.g., "a restaurant"). The Large Language Model (LLM) receives this instruction, first vectorizing the instruction text information using a vector embedding model, and then performing a semantic search in a vector database to retrieve relevant layout rules. This vector database contains a complete layout design rule document. Some of the retrieved rules are dynamically injected into the LLM's system prompts, enabling it to generate a more professional and reasonable scene plan, including a list of object models and a target scene description. In the execution phase, the system automatically retrieves all models based on the generated model list and inputs their information, along with the scene description and an image of the current empty scene, into the visual language model. The visual language model, acting as the "layout executor," predicts the positions and constraints of all objects based on the detailed text description. The subsequent processes are optimized and output. Figure 12 The optimization and physical inspection process shown is the same, ultimately generating a complete new scene that conforms to the user's high-level intent and is physically realistic.

[0154] The virtual environment data processing method proposed in this application constructs a user-centric "human-in-the-loop" VR creation mode: This application changes the one-way relationship between humans and tools in traditional VR content creation, establishing an efficient human-machine collaborative mode. In the VR environment, users provide high-level creative intentions and key spatial references through natural voice commands and controller spatial pointing, while AI is responsible for performing tedious batch retrieval, layout calculations, and detail optimization. Simultaneously, during the automated layout creation process, users can still interactively move and delete objects. This "human-in-the-loop" mode retains the creator's leadership and ultimate control while seamlessly combining human creativity with AI's execution capabilities, allowing users to focus on the design itself rather than tool operation.

[0155] The virtual environment data processing method proposed in this application improves the efficiency and quality of VR scene layout: through AI-driven intelligent planning and automated layout, this application reduces the manual modeling and adjustment work that may take several hours in traditional methods to minutes. At the same time, with the help of AI models enhanced by professional knowledge (through RAG technology), the layout generated by the system is not only fast, but also highly reasonable in terms of spatial relationships, functional zoning, and aesthetic principles, effectively avoiding the randomness and irrationality of traditional programmatically generated content.

[0156] The virtual environment data processing method proposed in this application ensures the physical realism and immersion of the virtual layout: the original multi-constraint optimization engine based on iterative thrust is the key to ensuring the realism of the VR experience. This engine can automatically correct physical errors such as collisions, clipping, and suspension between objects based on the AI-generated layout, ensuring that every object in the scene follows the physical laws of the real world. This guarantee of physical realism greatly enhances the user's immersion and credibility in the VR environment.

[0157] The virtual environment data processing method proposed in this application realizes dynamic intelligent interaction with contextual understanding: this invention is not just a static layout tool, but also an intelligent assistant capable of "observing" and "understanding" the scene. Through an innovative question-and-answer and execution loop, the system can combine the scene's visual information (top view) and structured data (Scene.json) to understand the user's dialogue and dynamically and intelligently modify the scene based on the dialogue content. This provides a deeper level of interaction for creating complex and dynamic content in a VR environment.

[0158] The solution provided in this application utilizes the "human-in-the-loop" approach, which allows users to improve and correct AI-generated data, achieving a true human-machine collaborative interaction mode rather than a rigid, one-way human-machine interaction mode. At the same time, it maintains physical semantic rationality and real-time performance during the generation process.

[0159] This application also provides an electronic device, including: At least one memory; At least one processor; At least one program; The program is stored in a memory, and the processor executes the at least one program to implement the virtual environment data processing method described above. The electronic device can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.

[0160] See also Figure 14 , Figure 14 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 1401 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 1402 can be implemented in the form of ROM (Read-Only Memory), static storage device, dynamic storage device, or RAM (Random Access Memory). The memory 1402 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1402 and is called and executed by the processor 1401 using the virtual environment data processing method of the embodiments of this application. The input / output interface 1403 is used to implement information input and output; The communication interface 1404 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1405 transmits information between various components of the device (e.g., processor 1401, memory 1402, input / output interface 1403, and communication interface 1404); The processor 1401, memory 1402, input / output interface 1403 and communication interface 1404 are connected to each other within the device via bus 1405.

[0161] This application embodiment also provides a storage medium, which is a computer-readable storage medium, storing a computer program that, when executed by a processor, implements the above-described virtual environment data processing method.

[0162] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0163] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0164] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0165] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0166] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0167] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0168] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0169] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The coupling or direct coupling or communication connection between the shown or discussed units may be through some interfaces, or indirect coupling or communication connection between the apparatus or units, and may be electrical, mechanical, or other forms.

[0170] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0171] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0172] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0173] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A virtual environment data processing method, characterized in that, The method includes: Responding to user commands for controlling virtual data in a virtual environment; When the virtual data control instruction characterizes the update of the virtual environment scene, it acquires the placement information of the object to be placed and the structural feature information and visual information of the virtual environment scene, and inputs the structural feature information, the structural feature information and the visual information into the visual language model for data processing to obtain the optimized arrangement data of the object to be placed in the virtual environment scene; Based on the optimized layout data, the object to be placed is generated in the virtual environment scene to obtain an adjusted virtual scene. Based on multiple physical placement constraints, the pose adjustment amount of the object to be placed in the adjusted virtual scene is calculated, and the pose of the object to be placed in the adjusted virtual scene is corrected based on the pose adjustment amount to obtain an updated virtual scene. The updated virtual scene is used as the new adjusted virtual scene, and the pose of the object to be placed in the adjusted virtual scene is corrected at least once. The updated virtual scene obtained after the last iteration of pose correction is used as the target virtual scene, and the target virtual scene is displayed or responded to in response to the user's modification operation.

2. The virtual environment data processing method according to claim 1, characterized in that, The process of obtaining the placement information of the object to be placed includes: The object identifier of the object to be placed is determined from the virtual data control instructions; Obtain the scene style information of the virtual environment scene; Style object retrieval information is generated based on the object identifier and the scene style information; Based on the style object retrieval information, a match is performed in the 3D model library to obtain the placement object information.

3. The virtual environment data processing method according to claim 1, characterized in that, The step of inputting the structural feature information, the visual information, and the visual information into a visual language model for data processing to obtain optimized placement data of the object to be placed in the virtual environment scene includes: When the virtual data control command includes reference position information of the object to be placed, the reference position information, the structural feature information, the structural feature information, and the visual information are input into the visual language model for data processing to obtain the optimized arrangement data near the reference position information; When the virtual data control command does not include the reference position information of the object to be placed, the structural feature information, the structural feature information and the visual information are input into the visual language model for data processing to obtain the optimized arrangement data that conforms to the position logic in the virtual environment scene.

4. The virtual environment data processing method according to claim 1, characterized in that, The calculation of the pose adjustment amount of the object to be placed in the adjusted virtual scene based on multiple physical placement constraints includes: Based on each of the physical placement constraints, the placement state of the objects to be placed in the adjusted virtual scene is matched one by one to obtain the state matching result. When the state matching result indicates that the object to be placed has a position error, an adjustment thrust is generated based on the position error information. When the state matching result indicates that the object to be placed has an orientation error, an adjustment rotation is generated based on the orientation error information; The pose adjustment amount is obtained based on the adjusted thrust and / or the adjusted rotation.

5. The virtual environment data processing method according to claim 1, characterized in that, The method further includes: The target virtual scene is analyzed and processed to obtain scene summary information of the target virtual scene; Responding to user inquiries; Based on the query information and the scenario summary information, a search and matching operation is performed in the layout rule knowledge base to obtain design rule information; The scenario summary information, the query information, and the design rule information are input into the question-answering model to obtain the response result data corresponding to the query information.

6. The virtual environment data processing method according to claim 5, characterized in that, The step of inputting the scene summary information, the query information, and the design rule information into the question-answering model to obtain the response result data corresponding to the query information includes: Confirm the target structural feature information and target visual information of the target virtual scene; The scene summary information, the query information, the design rule information, the target structural feature information, and the target visual information are input into the question-answering model to obtain the response result data corresponding to the query information.

7. The virtual environment data processing method according to claim 5, characterized in that, The method further includes: When the response result data includes information about newly placed objects, the new placement data of the newly placed objects in the target virtual scene is determined based on the information about newly placed objects; The newly placed object is generated in the target virtual scene based on the newly added layout data.

8. The virtual environment data processing method according to claim 1, characterized in that, The method further includes: Responding to a user's direct modification command to a target object in the target virtual scene; The target object in the target virtual scene is adjusted based on the direct modification command; Based on the adjusted target virtual scene, updated scene information is generated and saved.

9. An electronic device, characterized in that, The system includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the virtual environment data processing method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the virtual environment data processing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Object display method and device, storage medium and intelligent glasses

    CN117252968A

  • Virtual scene generation method and device, electronic equipment and storage medium

    CN117745987A

  • Large model-based scene retrieval method and terminal

    CN120508614A