Space-anchored virtual environment transformation method, system, and related devices

By reconstructing the 3D point cloud structure of the physical environment and generating stylized assets, the problem of integrating the virtual environment with the real space in virtual reality is solved, improving user experience and safety, and supporting user body interaction.

CN122510451APending Publication Date: 2026-08-04TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2026-05-15
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In existing virtual reality and mixed reality technologies, the integration of virtual environments with real spaces presents problems such as security risks, insufficient immersion and spatial perception, inconsistent styles, and inconvenient user interaction.

Method used

By acquiring visual data of the physical environment, a 3D point cloud structure is reconstructed, spatial structural elements are identified to construct a spatial framework, stylized assets are generated by combining user-input themes, and spatial alignment is performed to ensure that the virtual environment is consistent with the real structure.

Benefits of technology

It achieves a precise integration of the virtual environment and the real space, enhances the user's immersion and security, supports the user's physical interaction, and ensures style consistency and functional rationality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122510451A_ABST
    Figure CN122510451A_ABST
Patent Text Reader

Abstract

The application provides a space anchor-based virtual environment transformation method, system and related equipment, acquires visual data of a physical environment, reconstructs a three-dimensional point cloud structure, identifies space structure elements in the point cloud, constructs a space framework, extracts global style features according to a target theme input by a user, and generates local stylization constraints for furniture elements, generates corresponding stylization assets based on the space framework, the global style features and the local stylization constraints, spatially aligns the stylization assets according to pose information of elements in the space framework, and constructs a virtual environment that is visually uniform, spatially consistent and theme distinctive. The mechanism ensures that virtual content corresponds to real structures, supports coarse-grained body interaction by users, avoids mispositioning or suspension of models, and ensures style uniformity and functional rationality among different objects. A new efficient, natural and controllable technical path is provided, which breaks through the limitations of existing technologies in space perception, style consistency and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and artificial intelligence, and in particular to a method, system and related equipment for virtual environment transformation based on spatial anchoring. Background Technology

[0002] With the rapid development of Virtual Reality (VR) and Mixed Reality (MR) technologies, their applications in scenarios such as home entertainment, immersive narrative experiences, virtual exhibitions, and design visualization are becoming increasingly widespread. Users expect to maintain awareness of the physical room structure while enjoying highly immersive virtual themed spaces to ensure safety and functional operation.

[0003] However, existing virtual environment generation technologies suffer from several major problems that severely impact user experience. First, many virtual environments are entirely virtualized, making it difficult for users to accurately perceive the location of real-world furniture and boundaries, creating safety hazards. While some solutions use edge hints or perspective views to alert users, these methods often undermine immersion. Second, current themed scene generation solutions typically rely on pre-set models, lacking an understanding of the actual furniture's volume, orientation, and function, resulting in misalignment and inconsistency between generated virtual and real objects. Furthermore, the use of uniform textures often leads to inconsistent styles, affecting the visual experience. Finally, existing systems are either completely automated, with no user intervention, or entirely reliant on manual construction, requiring a huge workload and limiting users' ability to reasonably adjust object retention and deformation, thus impacting user freedom.

[0004] Therefore, these problems limit the application effect of VR / MR technology, and there is an urgent need to develop new solutions to achieve effective integration of virtual environment and real space in order to enhance user immersion and safety. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method, system and related equipment for virtual environment transformation based on spatial anchoring, in order to solve the problems that a completely virtual scene provides immersion but damages spatial perception, and that video perspective preserves physical structure but weakens the virtual experience.

[0006] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:

[0007] The first aspect of this invention discloses a virtual environment transformation method based on spatial anchoring, the method comprising:

[0008] Acquire visual data of the physical environment and reconstruct the three-dimensional point cloud structure of the physical environment based on the visual data;

[0009] Identify the spatial structural elements in the three-dimensional point cloud structure, and construct a spatial framework using the geometric and semantic information of the spatial structural elements;

[0010] In response to the target theme input by the user, the target theme is parsed to extract global style features, and local stylistic constraints for the furniture elements are generated based on the functional attributes of the furniture elements in the spatial framework.

[0011] Based on the spatial framework, the global style features, and the local stylization constraints, stylized assets corresponding to each spatial structural element in the spatial framework are generated.

[0012] The stylized assets are spatially aligned according to the pose information of the corresponding spatial structural elements in the spatial framework, so that the stylized assets visually overlap with the real structures in the physical environment.

[0013] Preferably, the step of acquiring visual data of the physical environment and reconstructing the three-dimensional point cloud structure of the physical environment based on the visual data includes:

[0014] Continuously acquire video frame sequences of the physical environment within a preset duration;

[0015] The camera pose of each frame in the video frame sequence is recovered using localization and mapping methods, and dense point cloud data of the physical environment is generated based on the recovered camera pose.

[0016] Based on the indoor structural features of the physical environment, the dense point cloud data is realigned to obtain a three-dimensional point cloud structure of the physical environment aligned in a unified world coordinate system.

[0017] Preferably, the step of identifying spatial structural elements in the 3D point cloud structure and constructing a spatial framework using the geometric and semantic information of the spatial structural elements includes:

[0018] Semantic decomposition is performed on the three-dimensional point cloud structure and the acquired video frame sequence to identify the spatial structural elements in the physical environment, and a three-dimensional bounding box with semantic labels is generated for each spatial structural element.

[0019] The three-dimensional bounding box serves as a spatial container for the physical environment, constituting the basic geometric and semantic constraints of the spatial framework.

[0020] For each target furniture element in the spatial structural elements, excluding walls and the ground, the observation quality of the target furniture element in each frame is evaluated from the video frame sequence to obtain the evaluation result;

[0021] Based on the evaluation results, a reference frame is selected for each target furniture element, and the camera pose of the reference frame is recorded as the reference pose of the target furniture element.

[0022] Each of the reference poses is associated with the spatial frame to construct a three-dimensional bounding box containing each spatial structural element and its reference pose.

[0023] Preferably, the step of responding to a target theme input by the user, parsing the target theme to extract global style features, and generating local stylistic constraints for the furniture elements based on the functional attributes of the furniture elements in the spatial framework includes:

[0024] Obtain the target theme input by the user through text or example image, and parse the target theme to extract core style keywords as global style features;

[0025] The global style features, the 3D bounding boxes of each spatial structural element in the spatial framework, and their semantic labels are input into the mapping module to generate a style mapping table. The style mapping table includes at least the local stylization constraints corresponding to each furniture element. The local stylization constraints include the functional semantics retained by the furniture element, the stylized object type and potential collision risk markers, and the generated prompts.

[0026] Preferably, the step of generating stylized assets corresponding to each spatial structural element in the spatial framework based on the spatial framework, the global style features, and the local stylization constraints includes:

[0027] For each furniture element in the spatial frame, a two-dimensional stylized image of the furniture element is generated based on the reference frame corresponding to the furniture element and the generation prompt in the local stylization constraint;

[0028] A three-dimensional stylized model of the furniture element is reconstructed based on the two-dimensional stylized image;

[0029] While reconstructing the three-dimensional stylized model of the furniture elements, surface texture assets corresponding to the wall elements and ground elements in the spatial frame are generated according to the global style features and the local stylization constraints, and environmental background assets corresponding to the environmental background in the spatial frame are generated according to the global style features.

[0030] The three-dimensional stylized models, surface texture assets, and environmental background assets of each furniture element are used as stylized assets corresponding to each spatial structural element in the spatial framework.

[0031] Preferably, the step of spatially aligning the stylized asset according to the pose information of the corresponding spatial structural elements in the spatial frame, so that the stylized asset visually overlaps with the real structure in the physical environment, includes:

[0032] For the three-dimensional stylized model in the stylized asset, the three-dimensional stylized model is transformed according to the three-dimensional bounding box of the corresponding spatial structural element in the spatial frame, so that the size of the three-dimensional stylized model is adapted to the three-dimensional bounding box and the orientation is matched with the actual orientation of the spatial structural element.

[0033] The position of the transformed 3D stylized model is adjusted to align with the spatial reference of the 3D bounding box;

[0034] For the surface texture assets in the stylized assets, the surface texture assets are mapped to the surface areas of the corresponding wall and ground elements in the spatial frame; for the environmental background assets in the stylized assets, the environmental background assets are loaded and presented as the panoramic background of the virtual scene.

[0035] All stylized assets, after alignment, are integrated into a complete virtual scene, so that each stylized asset in the virtual scene visually overlaps with the corresponding real structure in the physical environment.

[0036] Preferably, after spatially aligning the stylized assets according to the pose information of the corresponding spatial structural elements in the spatial frame, so that the stylized assets visually overlap with the real structures in the physical environment, the method further includes:

[0037] In response to a user's adjustment command for the three-dimensional bounding box corresponding to any spatial structural element in the spatial frame in mixed reality mode, the size, position, or orientation of the three-dimensional bounding box is adjusted.

[0038] In response to the user's command to add or delete bounding boxes of the spatial frame in mixed reality mode, add or delete three-dimensional bounding boxes and their associated spatial structural elements;

[0039] In response to a user's instruction to modify the semantic label of any spatial structural element, update the semantic label of that spatial structural element and use the updated semantic label when regenerating the stylized asset;

[0040] In response to the user's command to regenerate any stylized asset, the stylized asset is regenerated and replaced with the original stylized asset. The step of spatially aligning the stylized asset according to the pose information of the corresponding spatial structural element in the spatial frame is re-executed so that the stylized asset visually overlaps with the real structure in the physical environment.

[0041] During the process of adjusting the 3D bounding box, modifying semantic tags, or regenerating the stylized assets, the update status of the stylized assets corresponding to each spatial structural element is displayed in real time.

[0042] In response to a user's mode switching command, the system switches from the mixed reality mode to the virtual reality mode to display the adjusted or regenerated complete virtual scene.

[0043] A second aspect of this invention discloses a virtual environment transformation system based on spatial anchoring, the system comprising:

[0044] The spatial understanding module is used to acquire visual data of the physical environment and reconstruct the three-dimensional point cloud structure of the physical environment based on the visual data.

[0045] A construction module is used to identify spatial structural elements in the three-dimensional point cloud structure and to construct a spatial framework using the geometric and semantic information of the spatial structural elements.

[0046] The style extraction and mapping module is used to respond to the target theme input by the user, parse the target theme to extract global style features, and generate local stylization constraints for the furniture elements based on the functional attributes of the furniture elements in the spatial framework.

[0047] The content generation module is used to generate stylized assets corresponding to each spatial structural element in the spatial framework based on the spatial framework, the global style features, and the local stylization constraints.

[0048] The scene integration and presentation module is used to spatially align the stylized assets according to the pose information of the corresponding spatial structural elements in the spatial framework, so that the stylized assets visually overlap with the real structures in the physical environment.

[0049] The third aspect of this invention discloses a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements a virtual environment transformation method based on spatial anchoring disclosed in the first aspect of this invention.

[0050] The fourth aspect of this invention discloses an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement a spatially anchored virtual environment transformation method disclosed in the first aspect of this invention.

[0051] Based on the spatially anchored virtual environment transformation method, system, and related equipment provided in the above embodiments of the present invention, visual data of the physical environment is acquired, and a 3D point cloud structure is reconstructed; spatial structural elements in the point cloud are identified, and a spatial framework is constructed; global style features are extracted according to the target theme input by the user, and local stylistic constraints are generated for furniture elements; corresponding stylized assets are generated based on the spatial framework, global style features, and local stylistic constraints; the stylized assets are spatially aligned according to the pose information of the elements in the spatial framework, constructing a visually unified, spatially consistent, and thematically distinct virtual environment. This mechanism ensures that virtual content corresponds to the real structure, supports users to perform coarse-grained physical interactions such as touching, walking, and sitting, avoids model misalignment or floating, and ensures stylistic consistency and functional rationality among different objects. It provides an efficient, natural, and controllable new technological path, breaking through the limitations of existing technologies in spatial perception, style consistency, and user experience. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0053] Figure 1 A flowchart illustrating a virtual environment transformation method based on spatial anchoring, provided in an embodiment of the present invention;

[0054] Figure 2 A schematic diagram illustrating a virtual environment transformation method based on spatial anchoring provided in an embodiment of the present invention;

[0055] Figure 3 This is a schematic diagram illustrating the generation of a style mapping table based on user input, provided in an embodiment of the present invention.

[0056] Figure 4 A schematic diagram of the interaction process and interface provided in the embodiments of the present invention;

[0057] Figure 5 A flowchart of a virtual environment transformation system based on spatial anchoring provided for an embodiment of the present invention;

[0058] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0060] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0061] As the background technology shows, existing fully virtual scenes lack geometric correspondence with real rooms, resulting in unnatural interactions. Furthermore, perspective-based physical environment overlays lack a thematic and stylized experience, and traditional image style transfer cannot maintain the logic of the three-dimensional structure. Therefore, the core objective of this invention is to reconstruct the thematic appearance of elements such as walls, furniture, and furnishings through generative models without altering the actual position, volume, functional semantics, and some appearance features of objects in a real room. This allows the generated virtual environment to possess both aesthetic consistency and precise matching with the structure and functional logic of the physical space, thereby simultaneously satisfying the needs for immersive experience, spatial safety, and creative controllability.

[0062] Therefore, this invention provides a method, system, and related equipment for virtual environment transformation based on spatial anchoring. It reconstructs the 3D point cloud structure of the physical environment using visual data, analyzes and identifies spatial structural elements, and generates a 3D bounding box. Combining large-model-driven style analysis and generative image and 3D asset generation methods, it stylizes the spatial container under geometric and semantic constraints. By maintaining the geometric boundaries and functional attributes of physical objects, it generates stylized assets that are precisely aligned with the real room structure, constructing a visually unified, spatially consistent, and thematically distinct virtual environment. This mechanism ensures that virtual content corresponds to the real structure, supports coarse-grained physical interactions such as touching, walking, and sitting, avoids model misalignment or floating, and ensures stylistic consistency and functional rationality among different objects. In summary, it provides an efficient, natural, and controllable new technological path, breaking through the limitations of existing VR / MR technologies in spatial perception, style consistency, and user experience.

[0063] See Figure 1The diagram shows a flowchart of a virtual environment transformation method based on spatial anchoring provided by an embodiment of the present invention. The method is explained below:

[0064] The method is implemented by one or more computing devices, which include at least a processor, memory, and an image acquisition interface. In a specific implementation scenario, the computing device can be integrated into a head-mounted display device or a standalone host device communicatively connected to the head-mounted display. The camera on the head-mounted display is used to acquire visual data of the physical environment, the processor is used to execute the reconstruction, recognition, mapping, generation, and alignment steps of this embodiment of the invention, and the memory is used to store intermediate data and the finally generated stylized assets.

[0065] Understandably, the proposed spatially anchored virtual environment stylization transformation method comprises four core components: a spatial understanding component, a style parsing and mapping component, a stylized content generation component, and a scene integration and presentation component. These four components work together to enable the system to construct a thematic, stylized, and functionally preserved virtual environment while fully adhering to the geometric structure of a real room.

[0066] It should be noted that the function of the spatial understanding component is to ensure that the generated virtual environment always accurately corresponds to the real physical room in three-dimensional space, avoiding collisions or misjudgments of position for the user in VR. Specifically, the spatial understanding component is responsible for parsing the geometric structure and semantic information of the physical environment, constructing the spatial framework upon which all subsequent stylization generation depends. Its processing steps include steps S101 and S102.

[0067] Step S101: Obtain visual data of the physical environment and reconstruct the three-dimensional point cloud structure of the physical environment based on the visual data.

[0068] In the specific implementation step S101, the head-mounted display camera collects a monocular video frame sequence of the user's physical environment, and a reconstruction algorithm is used to generate a three-dimensional point cloud structure of the physical environment.

[0069] Specifically, in combination Figure 2 The content shown in “a. Spatial Understanding Module” first involves continuously collecting video frame sequences of the physical environment within a preset duration.

[0070] Understandably, users continuously capture images of their physical environment using a camera on a head-mounted display. The camera is a monocular camera that captures a preset duration of indoor video (e.g., 30 to 60 seconds), generating a continuous sequence of video frames. This preset duration can be adaptively adjusted based on room size and complexity to cover sufficient environmental structural information.

[0071] Then, the camera pose of each frame in the video frame sequence is recovered using localization and mapping methods, and dense 3D point cloud data of the physical environment is generated based on the recovered camera pose.

[0072] It should be noted that each frame of the acquired video frame sequence is processed using localization and mapping methods (such as the SLAM3R algorithm): First, feature points are extracted from the image, and the camera pose (i.e., the camera's position and orientation in 3D space) for each frame is recovered through inter-frame feature matching; simultaneously, based on the principle of multi-view stereo vision, the image information and camera pose of all frames are fused to generate dense 3D point cloud data of the physical environment. This point cloud data contains the spatial coordinates of various surface points in the environment and can preliminarily represent the geometry of the room.

[0073] Finally, based on the indoor structural features of the physical environment, the dense 3D point cloud data is realigned to obtain the 3D point cloud structure of the physical environment aligned in a unified world coordinate system.

[0074] In some specific embodiments, due to the potential accumulation of errors in camera pose estimation during the acquisition process, and the arbitrary coordinate system of the point cloud reconstructed by the positioning and mapping methods, the system performs realignment processing on the generated dense 3D point cloud data to ensure the accuracy of subsequent spatial semantic parsing. Specifically, the system uses an uncertainty-aware rotation estimation method to infer the principal axis direction of the room. Based on this, the entire point cloud is rotated and translated to align its coordinate axes with the actual horizontal and vertical directions of the room, ultimately obtaining a 3D point cloud structure aligned in a unified world coordinate system with a clear gravity direction, serving as the geometric basis for all subsequent spatial understanding and stylization operations.

[0075] Step S102: Identify spatial structural elements in the 3D point cloud structure and construct a spatial framework using the geometric and semantic information of the spatial structural elements.

[0076] In the specific implementation step S102, the 3D point cloud structure and the acquired video frame sequence are jointly input into an intelligent language model (e.g., a Spatial Language Model, SpatialLM) oriented towards the 3D scene. This model can jointly infer the geometric distribution in the point cloud and the appearance features in the video frames, automatically identifying key spatial structural elements in the physical environment, including but not limited to walls, doors, windows, and furniture. A spatial framework is constructed using the geometric and semantic information of these spatial structural elements. This framework serves as the spatial boundary for all subsequent generated content, ensuring that the virtual environment is always geometrically aligned with reality.

[0077] It is understandable that the specific implementation process of step S102 is as follows (processes A1 to A5):

[0078] Process A1: Semantic decomposition is performed on the 3D point cloud structure and the acquired video frame sequence to identify spatial structural elements in the physical environment and generate a 3D bounding box with semantic labels for each spatial structural element.

[0079] In process A1, the 3D point cloud structure and the video frame sequence are input into the intelligent language model. This model uses the geometric distribution information provided by the point cloud and the appearance texture information provided by the video frames to reason about and classify functional structures in the physical environment. After processing, the model identifies spatial structural elements in the physical environment. For each identified spatial structural element, the model further outputs a 3D bounding box.

[0080] It is understandable that this 3D bounding box contains attribute information in the following four dimensions:

[0081] Dimensions: The length along the three axes of the world coordinate system, reflecting the actual physical size of the element;

[0082] Location: The three-dimensional coordinates of the center point of the bounding box, which determines the specific orientation of the element in the room space;

[0083] Orientation: The direction of the main axis of the enclosure box, used to characterize the posture of the element (such as the opening and closing direction of a door, or the front of a sofa) in space;

[0084] Category tags: Semantic category identifiers for elements, such as "walls", "windows", "sofas", "desks", etc.

[0085] Process A2: Using the 3D bounding box as a spatial container for the physical environment, it forms the basic geometric and semantic constraints of the spatial framework.

[0086] Understandably, these 3D bounding boxes, as spatial containers of the physical environment, constitute the basic geometric and semantic constraints of the spatial framework, providing a precise spatial benchmark for the subsequent construction of the spatial framework and the generation and alignment of stylized content.

[0087] Process A3: For each target furniture element in the spatial structure elements other than the walls and the ground, evaluate the observation quality of the target furniture element in each frame from the video frame sequence to obtain the evaluation results.

[0088] In implementing process A3, for each target furniture element in the spatial structure elements other than the walls and the ground, the entire video frame sequence is evaluated frame by frame based on the 3D bounding box information obtained from spatial semantic parsing to obtain the evaluation result.

[0089] Specifically, the observation quality of an object in each frame is comprehensively judged from three key dimensions: visibility, centering, and visible area size. Occluded viewpoints are filtered before being included in the scoring. Visibility measures whether sufficient surface area of ​​the object is captured in the frame; centering reflects the proximity of the object's projection to the image center, emphasizing visual salience; and visible area size excludes significantly occluded viewpoints. To ensure robustness, a lexicographical priority is used: visibility is compared first, followed by centering, and finally, visible area size.

[0090] In actual calculations, the semantic bounding boxes are first transformed to the SLAM coordinate system through Sim(3) registration, which means aligning the semantic bounding boxes in the world coordinate system back to the SLAM coordinate system so that the corner points of the bounding boxes can be projected into the camera view of each frame. For each frame With objects Define each candidate frame The evaluation vector is:

[0091] s(f) = (v_cnt(f), -d_center(f), A_vis(f))(1).

[0092] In formula (1), v_cnt(f) represents the number of visible points of the target in the frame, d_center(f) represents the distance between the target center and the image center, and A_vis(f) represents the visible area of ​​the target.

[0093] Process A4: Based on the evaluation results, select a reference frame for each target furniture element and record the camera pose of the reference frame as the reference pose of the target furniture element.

[0094] Understandably, the best-view frame in the evaluation results, which best represents the complete geometry of the object, is selected as the best-view frame for the object by lexicographical maximization, and the camera pose of the reference frame is recorded as the reference pose of the target furniture element.

[0095] Specifically, by negativening the distance term, each component satisfies the principle that "the larger the value, the better the quality." The optimal frame is defined as:

[0096] f* = argmax_{f ∈ F} s(f)(2).

[0097] In formula (2), argmax takes the maximum value in the sense of lexicographic order, that is, v_cnt is compared first, then -d_center (i.e., the smaller the distance, the better), and finally A_vis is compared.

[0098] In practical applications, the selected reference frame is bounded by a green box (e.g., Figure 2 The "Select Best Frame" step in the "a. Spatial Understanding Module" (corresponding to the green bounding box in the example image) is recorded along with object label annotations. Simultaneously, the camera pose of the reference frame is written into the scene's JSON file to guide subsequent stylization generation and 3D model registration. This mechanism ensures that each object enters the generation process from a clear, unobstructed perspective with the most representative geometric structure, thereby significantly improving the stability and spatial alignment accuracy of the stylized content.

[0099] Process A5: Associate each reference pose with the spatial frame, and construct a spatial frame containing each spatial structural element and its reference pose.

[0100] It should be noted that the spatial framework includes: each spatial structural element (at least including walls, doors, windows, and furniture) identified from the 3D point cloud structure of the physical environment; the 3D bounding box corresponding to each spatial structural element and its size, position, orientation, and category semantic labels; and the reference frame camera pose recorded for each furniture element as a reference pose. The spatial framework constitutes the spatial benchmark and geometric semantic constraints for subsequent stylized asset generation and alignment.

[0101] Step S103: In response to the target theme input by the user, parse the target theme to extract global style features, and generate local stylization constraints for the furniture elements based on the functional attributes of the furniture elements in the spatial framework.

[0102] It should be noted that step S103 is implemented through the style parsing and mapping component. This component is responsible for transforming the user's thematic requirements into executable style generation rules, aiming to ensure consistency between style and function, while guaranteeing that the style of the object matches the size and purpose of the real object. Furthermore, the user can further edit or lock key objects. These goals collectively promote the overall coordination and flexibility of the system.

[0103] Specifically, in combination Figure 2 The "b. Style Extraction and Mapping Module" shows that users input their desired target theme through text descriptions, voice descriptions, or sample images. The system uses a language model to parse the user's input and extract style keywords related to the target theme. Style keywords include at least dimensions such as material, color tone, and atmosphere.

[0104] Building upon this foundation, the language model further generates stylized descriptions for each furniture element within the spatial framework, adapting them to the target theme. For instance, for objects with a "sitting" function (such as sofas and chairs), the language model maps them to alternative representations within the same stylistic framework that are visually recognizable and retain the basic functional logic of seating. For example, a modern sofa might be mapped to a wood-grain rope-woven style bench. Through this process, the system ensures that the basic functional semantics of each furniture element are not compromised while maintaining overall visual style consistency.

[0105] Combination Figure 3 The specific implementation process of step S103 shown is as follows (process B1 and process B2):

[0106] Process B1: Obtain the target topic input by the user through text or sample images, parse the target topic to extract core style keywords as global style features.

[0107] In implementing process B1, users can input their desired theme via text or sample images, such as "forest cabin," "cyberpunk bar," or "Hogwarts School of Witchcraft and Wizardry." The system automatically extracts 4 to 8 core style keywords (i.e., ...) using a large language model. Figure 3 The "style keywords" shown include color preferences, main materials, atmospheric characteristics, and design language. These core style keywords are used as the overall style characteristics.

[0108] Process B2: Input the global style features, the 3D bounding boxes of each spatial structural element in the spatial frame, and their semantic labels into the mapping module to generate a style mapping table.

[0109] In implementing process B2, the global style features and the 3D bounding boxes and semantic labels of each spatial structural element in the spatial framework are input into the mapping module, and the language model generates a style mapping table (e.g., ...). Figure 3 The “MappingTable” shown.

[0110] The style mapping table includes at least the local stylization constraints corresponding to each furniture element. The local stylization constraints include the functional semantics retained by the furniture element, the stylized object type and potential collision risk markers, as well as the generated prompts.

[0111] Specifically, the style mapping table integrates user style intent, spatial semantic information, and generated prompts and various constraints obtained from language model inference. For each type of spatial structural element in the scene, the style mapping table defines the following in sequence:

[0112] Label;

[0113] Replace object function;

[0114] Preserved functional semantics (such as Figure 3 The "object functions" shown include, for example, "storing items", "protecting privacy", "placing items", etc.

[0115] Stylized object types (such as) Figure 3 The "replacement object" shown can be used to map a "sofa" to a "wooden rope-woven bench".

[0116] Generate a prompt (e.g.) Figure 3 The "appearance cues" shown are used to guide the subsequent generation of stylized assets;

[0117] Potential collision risk marker (collision_risk) (e.g.) Figure 3 The "Clash Risk" indicator is used to mark the potential spatial interaction risks of the element after stylization.

[0118] In addition, for fixed structural elements such as walls, floors, and ceilings, the style mapping table also includes their corresponding material rules.

[0119] In the subsequent stylization generation stage, the style mapping table is used uniformly in a structured form. On the one hand, since this mapping table is generated once based on the entire spatial framework, it can constrain the generation process at the global level, thereby ensuring stylistic consistency and semantic rationality among different objects; on the other hand, it enables centralized management of key attributes of each object, facilitating subsequent editing or replacement by users.

[0120] Understandably, step S103 establishes a cross-modal, multi-level style control system that allows users to specify themes through text, images, or instance styles. It also utilizes a large language model to construct a style mapping table, generating structured style rules for walls, floors, furniture, and background environments to maintain consistency in the overall scene's materials, tones, and lighting atmosphere. Simultaneously, templated prompts designed for different object categories constrain the behavior of the generation model, reduce illusions, improve style consistency, and ensure a stable, controllable generation process with professional-grade style expression capabilities.

[0121] Step S104: Based on the spatial framework, global style features, and local stylization constraints, generate stylized assets corresponding to each spatial structural element in the spatial framework.

[0122] It should be noted that step S104 is implemented through a stylized generation component. This component generates textures, images, and 3D models for constructing the virtual environment based on a style map, ensuring that the functional semantics of objects remain consistent with geometric constraints. This component includes three types of generation processes: stylized 3D generation of objects within the scene, generation of wall and ground textures, and generation of environmental backgrounds and skyboxes. These processes work together to ensure that the visual effects and functionality of the virtual environment meet the expected standards.

[0123] Combination Figure 2 The content shown in "c. Content Generation Module" is used in step S104 to generate, based on the spatial framework and style mapping results, thematic textures for walls and floors, stylized images and corresponding 3D models of furniture, and a skybox environment consistent with the overall theme. The generated objects maintain the same volume, orientation, and placement as real furniture, allowing users to naturally maintain consistency with the real world when performing coarse-grained body interactions such as walking and sitting in the virtual environment.

[0124] It is understandable that the specific implementation process of step S104 is as follows (processes C1 to C4):

[0125] Process C1: For each furniture element in the spatial frame, generate a two-dimensional stylized image of the furniture element based on the reference frame corresponding to the furniture element and the generation prompts in the local stylization constraints.

[0126] In the implementation of process C1, for each furniture element in the spatial frame, the reference frame corresponding to the furniture element and the generation prompt in the local stylization constraint are input into the image generation model to generate a two-dimensional stylized image of the furniture element that has a consistent appearance style and retains the basic structure.

[0127] Process C2: Reconstruct a three-dimensional stylized model of the furniture elements based on the two-dimensional stylized image.

[0128] In process C2, a two-dimensional stylized image is input into a three-dimensional generative model for reconstruction, generating a lightweight three-dimensional stylized model of the furniture element (such as a rope chair, magic door, forest wardrobe, etc.).

[0129] Understandably, the stylized 3D generation process of objects within the scene corresponds to processes C1 and C2, which ensures that the generated objects have a strong thematic style while retaining the approximate volume, posture, and footprint of the original objects, thus ensuring that users can still have normal physical interactions with real objects.

[0130] Process C3: While reconstructing the 3D stylized model of the furniture elements, surface texture assets corresponding to the wall elements and ground elements in the spatial frame are generated according to the global style features and local stylization constraints. Environmental background assets corresponding to the environmental background in the spatial frame are also generated according to the global style features.

[0131] It is understandable that the process of generating wall and ground textures and the process of generating environmental background and skybox correspond to process C3.

[0132] In the implementation of process C3, for wall and ground elements within the spatial framework, global style features and local stylization constraints are input into the image generation model to generate tileable material textures that do not require stitching, and normal maps and metallic maps are automatically generated. Simultaneously, static skyboxes are generated based on the theme to provide an immersive atmosphere, such as the night sky of a magic academy, bioluminescent corals in the deep sea, and the skyline of a cyberpunk city. Furthermore, a large model is used to seamlessly loop the forward and reverse playback of the skybox video, achieving a seamless playback effect.

[0133] For example, surface texture assets include: tileable material textures that do not require stitching (i.e., diffuse maps), as well as automatically generated normal maps and metallic maps. Through the combination of these maps, walls and floors can present realistic lighting effects and material textures in subsequent rendering.

[0134] Process C4: Use the 3D stylized models, surface texture assets, and environmental background assets of each furniture element as stylized assets corresponding to the spatial structural elements in the spatial framework.

[0135] Step S105: Spatially align the stylized assets according to the pose information of the corresponding spatial structural elements in the spatial frame, so that the stylized assets visually overlap with the real structures in the physical environment.

[0136] It should be noted that step S105 is specifically implemented through the scene integration and presentation component. The scene integration component is responsible for strictly aligning all generated content into the spatial anchoring frame, so that the virtual environment is completely transformed visually, but corresponds one-to-one with the real physical environment in terms of spatial structure.

[0137] Combination Figure 2As shown in the "d. Immersive Integration and Presentation Module", in step S105, the system automatically aligns all generated stylized assets into the corresponding spatial frame. Through scale matching, Intersection over Union (IoU) rotation correction, and position optimization mechanisms, it achieves precise overlap between the virtual content and the real physical environment structure. The final presented theme environment can be directly entered in VR, allowing users to obtain an immersive visual experience while maintaining accuracy in movement and spatial perception.

[0138] The specific implementation process of step S105 is as follows (processes D1 to D4):

[0139] Process D1: For the 3D stylized model in the stylized asset, transform the 3D stylized model according to the 3D bounding box of the corresponding spatial structural element in the spatial frame, so that the size of the 3D stylized model fits the 3D bounding box and the orientation matches the actual orientation of the spatial structural element.

[0140] Understandably, the system executes process D1 to ensure that all stylized content falls precisely within its corresponding bounding box.

[0141] First, for each furniture element's 3D stylized model, a uniform scaling and normalization process is performed: based on the length of the longest side of the furniture element's bounding box, the 3D stylized model is scaled proportionally so that the longest side of the model is aligned with the longest side of the bounding box, thereby making the overall size of the model fit its spatial container while maintaining the original shape proportions.

[0142] Secondly, based on the camera pose of the best viewpoint frame corresponding to the furniture element, a set of candidate rotation directions are generated around the vertical direction (Y-axis) of that camera pose (e.g., sampling within ±45 degrees). For each candidate rotation direction, the 3D overlap (e.g., 3D intersection-over-union ratio) between the 3D stylized model and the spatial bounding box is calculated, and the direction with the highest overlap is selected as the final orientation of the model. For planar objects such as carpets and picture frames, since their front orientation may be distributed along any coordinate axis (e.g., picture frames against a wall, carpets on the floor), in addition to the conventional vertical axis rotation optimization, the system performs additional positive and negative 90° axis flip tests along the three coordinate axes, and selects the orientation that best matches the stylized reference image to ensure that the front and back orientations of the model are correct.

[0143] Process D2: Adjust the position of the transformed 3D stylized model to align with the spatial reference of the 3D bounding box.

[0144] It should be noted that after coarse alignment, the system applies fine-grained scale constraints. Specifically, it limits the dimensions of the 3D stylized model in each axis to a preset multiple (e.g., 1.3 times) of the corresponding axial dimension of the bounding box, to ensure the stability of spatial occupancy while maintaining the freedom of stylized appearance. Finally, the system performs ground reference alignment: aligning the bottom surface of the 3D stylized model with the lower surface of the bounding box (i.e., the ground reference), thereby preventing the model from floating above the ground or partially embedded below it.

[0145] Process D3: For surface texture assets in stylized assets, map the surface texture assets to the surface areas of the corresponding wall and ground elements in the spatial frame; for environmental background assets in stylized assets, load and present the environmental background assets as the panoramic background of the virtual scene.

[0146] Understandably, through the above-mentioned scaling, rotation, flipping, bottom face alignment, and panoramic background loading operations, the system enables each stylized 3D model to maintain the appearance of the theme while achieving a stable and correct spatial mapping with the corresponding object structure in the real room, and makes the virtual scene present an environmental atmosphere consistent with the target theme as a whole.

[0147] Process D4: Integrate all the stylized assets after alignment into a complete virtual scene, so that the stylized assets in the virtual scene visually overlap with the corresponding real structures in the physical environment.

[0148] It should be noted that after all objects have completed geometric alignment and position correction, the system integrates the wall textures, floor materials, skybox environment background, and the 3D stylized models of each piece of furniture into a complete virtual reality scene. Users can immediately enter this themed environment and interact with virtual objects in a natural way (such as touching, sitting, and walking) without worrying about misalignment or spatial perception conflicts with the actual furniture in the real room.

[0149] Understandably, step S105 introduces an object-level spatial constraint and optimal viewpoint inference mechanism. Through a per-object optimal viewpoint selection algorithm, the system automatically selects the synchronously located and mapped video frame that best represents the object's geometry as the reference input for stylization generation, ensuring that the generated result is consistent with the shape, orientation, and semantic usage of the real object. Combined with the generated content registration method, the system achieves high-precision mapping between virtual content and physical space through steps such as scaling normalization, IoU pose optimization, axial flip checking, and ground alignment, avoiding problems such as misalignment, clipping, or floating. This refined object-level alignment scheme has a significant leading advantage over similar technologies.

[0150] In some specific embodiments, the present invention operates in two modes: mixed reality mode and virtual reality mode.

[0151] In response to a user's adjustment command for the 3D bounding box corresponding to any spatial structural element in the spatial frame in mixed reality mode, the size, position, or orientation of the 3D bounding box is adjusted.

[0152] In response to user commands to add or delete bounding boxes of the spatial frame in mixed reality mode, add or delete 3D bounding boxes and their associated spatial structural elements.

[0153] In response to a user's instruction to modify the semantic label of any spatial structural element, the semantic label of that spatial structural element is updated, and the updated semantic label is used when regenerating the stylized asset.

[0154] In response to a user's regeneration instruction for any stylized asset (including a 3D stylized model, surface texture asset, or environmental background asset), the stylized asset is regenerated and replaced with the original stylized asset, and step S105 is re-executed.

[0155] During the process of adjusting the 3D bounding box, modifying semantic tags, or regenerating stylized assets, the update status of the stylized assets corresponding to each spatial structural element is displayed in real time.

[0156] In response to the user's mode switching command, it switches from mixed reality mode to virtual reality mode to display the adjusted or regenerated complete virtual scene.

[0157] It should be added that after generating the stylized assets corresponding to each spatial structural element, the system executes step S105, which involves spatially aligning the stylized assets according to the pose information of the corresponding spatial structural elements in the spatial frame, so that the stylized assets visually overlap with the real structures in the physical environment. In this embodiment, this step also integrates a user-oriented interactive interface and operation flow, such as... Figure 4 As shown.

[0158] like Figure 4 As shown in "(a) Container Visualization", the system displays a 3D bounding box (i.e., a spatial container) obtained through spatial semantic parsing in the reconstructed physical scene. Users can intuitively view the spatial location and extent of various objects in the room through a head-mounted display and enter mixed reality mode. This visualization interface helps users understand the spatial correspondence between stylized assets and real structures.

[0159] like Figure 4 As shown in "(b) Extracting Style from Multimodal Information", the user specifies the desired target style by uploading a reference image or entering text prompts. The system extracts style features from the multimodal input as conditional inputs for subsequent generative asset creation.

[0160] like Figure 4 As shown in "(c) Manipulating the Container," in a virtual reality environment, users can directly manipulate the displayed 3D bounding box using interactive devices such as controllers, including selecting, rotating, moving, and scaling. The system responds to user adjustment commands, updating the pose information of the corresponding spatial structural elements in real time, and accordingly re-performs spatial alignment of the stylized assets. This interactive mechanism allows users to have fine-grained control over the scene layout, meeting personalized editing needs.

[0161] like Figure 4 As shown in "(d) Generation Information Monitoring," during the stylized asset generation and alignment process, the system displays the real-time status of the generation process, including generation progress, information on the currently processed object, generation prompts used, and a preview of the generation result. Users can confirm, regenerate, or edit the generated assets via voice commands or interface buttons. Finally, the system accurately places the confirmed 3D model or texture asset back into its corresponding spatial location, completing the closed loop from stylized generation to spatial alignment.

[0162] Through the above interactive process, users can intuitively and flexibly participate in various stages of stylization transformation while maintaining the geometric consistency of physical space, thus achieving efficient and controllable virtual environment creation.

[0163] In this embodiment of the invention, a full-link approach combining spatial understanding, semantic parsing, and geometric alignment is employed, using the real-world room structure as the spatial benchmark for the virtual environment. This allows themed virtual scenes to strictly adhere to the size, location, and accessibility requirements of real-world spaces while undergoing complete style changes. This technical approach, combining structural preservation with stylistic freedom, significantly enhances immersive experience, security, and spatial predictability, filling a gap in current VR / MR systems. This invention possesses excellent scalability and model replaceability. The system framework is completely decoupled from specific model implementations; the currently used image generation model, 3D reconstruction model, semantic parsing model, and SLAM reconstruction model can all be replaced with other models with similar functions or future more advanced algorithms without affecting the overall process and final effect. This loosely coupled design allows the invention to evolve over the long term and flexibly adapt to different hardware platforms and application scenarios.

[0164] It should be added that the virtual environment transformation method based on spatial anchoring provided in this embodiment of the invention has been verified in the following two typical scenarios:

[0165] Scene 1: Immersive narrative experience.

[0166] Ordinary rooms are transformed into virtual environments with specific story themes through stylized transformations, allowing users to experience themed virtual content and engage in immersive interactive entertainment within real-world rooms. In this scenario, the system uses real-world space as a narrative vehicle, enabling users to experience the atmosphere and plot of the stylized environment while moving safely.

[0167] Scene 2: Space Design Preview.

[0168] This method allows designers to quickly generate various interior design styles (such as modern, classical, and sci-fi) while maintaining the actual geometric structure of the room, for creative exploration and spatial effect preview. This scenario significantly improves design iteration efficiency and reduces the cost of physical set design.

[0169] Understandably, compared to traditional video perspective solutions, the user's subjective immersion in the stylized virtual environment is significantly enhanced; spatial cognition is maintained: users can accurately judge the position of real objects, effectively reducing the risk of collisions caused by misalignment between virtual content and physical space; and creative expression efficiency is improved: in design applications, users can quickly generate multiple stylized environments and compare and filter them, significantly improving the efficiency and flexibility of creative exploration. These experimental data demonstrate that, while maintaining spatial structural consistency, this invention successfully achieves highly immersive, highly secure, and highly controllable stylized transformations of virtual environments.

[0170] Corresponding to the virtual environment transformation method based on spatial anchoring provided in the above embodiments of the present invention, see also... Figure 5 The diagram shows a structural block diagram of a virtual environment transformation system based on spatial anchoring provided by an embodiment of the present invention.

[0171] The system includes: a spatial understanding module 501, a construction module 502, a style extraction and mapping module 503, a content generation module 504, and a scene integration and presentation module 505.

[0172] The spatial understanding module 501 is used to acquire visual data of the physical environment and reconstruct the three-dimensional point cloud structure of the physical environment based on the visual data.

[0173] The spatial understanding module 501 is specifically used for: continuously acquiring video frame sequences of the physical environment within a preset time period; recovering the camera pose of each frame in the video frame sequence using localization and mapping methods, and generating dense point cloud data of the physical environment based on the recovered camera pose; and realigning the dense point cloud data based on the indoor structural features of the physical environment to obtain a three-dimensional point cloud structure of the physical environment aligned in a unified world coordinate system.

[0174] Module 502 is used to identify spatial structural elements in a 3D point cloud structure and to construct a spatial framework using the geometric and semantic information of the spatial structural elements.

[0175] The style extraction and mapping module 503 is used to respond to the target theme input by the user, parse the target theme to extract global style features, and generate local stylization constraints for furniture elements based on the functional attributes of furniture elements in the spatial framework.

[0176] The style extraction and mapping module 503 is specifically used for: obtaining the target theme input by the user through text or example images, parsing the target theme to extract core style keywords as global style features; inputting the global style features, the 3D bounding boxes of each spatial structural element in the spatial frame and their semantic labels into the mapping module to generate a style mapping table, wherein the style mapping table includes at least the local stylization constraints corresponding to each furniture element, the local stylization constraints include the functional semantics retained by the furniture element, the stylized object type and potential collision risk markers, and the generated prompts.

[0177] The content generation module 504 is used to generate stylized assets corresponding to each spatial structural element in the spatial framework based on the spatial framework, global style features, and local stylization constraints.

[0178] The scene integration and presentation module 505 is used to spatially align stylized assets according to the pose information of the corresponding spatial structural elements in the spatial framework, so that the stylized assets visually overlap with the real structures in the physical environment.

[0179] In this embodiment of the invention, the system framework is completely decoupled from specific model implementations. The currently used image generation model, 3D reconstruction model, semantic parsing model, and SLAM reconstruction model can all be replaced with other models with similar functions or future more advanced algorithms without affecting the overall process and final result. This loosely coupled design allows the invention to evolve over the long term and flexibly adapt to different hardware platforms and application scenarios.

[0180] Combination Figure 5 As shown, the construction module 502 includes: semantic decomposition unit, composition unit, evaluation unit, selection unit, and construction unit.

[0181] The semantic decomposition unit is used to perform semantic decomposition on the 3D point cloud structure and the acquired video frame sequence, identify the spatial structural elements in the physical environment, and generate a 3D bounding box with semantic labels for each spatial structural element.

[0182] The constituent unit is used to treat the 3D bounding box as a spatial container for the physical environment, forming the basic geometric and semantic constraints of the spatial framework.

[0183] The evaluation unit is used to evaluate the observation quality of each target furniture element in the video frame sequence, excluding walls and the ground, and obtain the evaluation result.

[0184] The selection unit is used to select a reference frame for each target furniture element based on the evaluation results, and record the camera pose of the reference frame as the reference pose of the target furniture element.

[0185] The building unit is used to associate each reference pose with the spatial frame, and to build a spatial frame containing each spatial structural element and its reference pose.

[0186] Combination Figure 5 The content generation module 504, as shown, includes: a first generation unit, a reconstruction unit, a second generation unit, and a determination unit.

[0187] The first generation unit is used to generate a two-dimensional stylized image of each furniture element in the spatial frame, based on the reference frame corresponding to the furniture element and the generation prompts in the local stylization constraints.

[0188] The reconstruction unit is used to reconstruct a three-dimensional stylized model of furniture elements from a two-dimensional stylized image.

[0189] The second generation unit is used to generate surface texture assets corresponding to wall elements and ground elements in the spatial frame, based on global style features and local stylization constraints, while reconstructing the three-dimensional stylized model of furniture elements, and generating environmental background assets corresponding to the environmental background in the spatial frame based on global style features.

[0190] The defined unit is used to take the 3D stylized model, surface texture assets, and environmental background assets of each furniture element as stylized assets corresponding to each spatial structural element in the spatial framework.

[0191] Combination Figure 5 The scene integration and presentation module 505, as shown, includes: a transformation unit, an adjustment unit, a mapping unit, and an integration unit.

[0192] The transformation unit is used to transform the 3D stylized model in the stylized asset according to the 3D bounding box of the corresponding spatial structural element in the spatial frame, so that the size of the 3D stylized model fits the 3D bounding box and the orientation matches the actual orientation of the spatial structural element.

[0193] The adjustment unit is used to adjust the position of the transformed 3D stylized model to align with the spatial reference of the 3D bounding box.

[0194] The mapping unit is used to map surface texture assets in stylized assets to the surface areas of corresponding wall and ground elements in the spatial frame; and to load and present environmental background assets as the panoramic background of the virtual scene.

[0195] The integration unit is used to integrate all stylized assets after alignment into a complete virtual scene, so that the stylized assets in the virtual scene visually overlap with the corresponding real structures in the physical environment.

[0196] Combination Figure 5 The system also includes the following modules: adjustment module, bounding box editing module, semantic tag editing module, regeneration module, real-time display module, and switching display module.

[0197] The adjustment module is used to adjust the size, position, or orientation of the 3D bounding box corresponding to any spatial structural element in the spatial frame in response to the user's adjustment command in the mixed reality mode.

[0198] The bounding box editing module is used to add or delete 3D bounding boxes and their associated spatial structural elements in response to user commands to add or delete bounding boxes of a spatial frame in mixed reality mode.

[0199] The semantic tag editing module is used to update the semantic tag of any spatial structural element in response to the user's instruction to modify the semantic tag of that spatial structural element, and to use the updated semantic tag when regenerating the stylized asset.

[0200] The regeneration module is used to respond to the user's regeneration command for any stylized asset, regenerate the stylized asset and replace the original stylized asset, and re-execute the scene integration and presentation module 505.

[0201] The real-time display module is used to display the update status of the stylized assets corresponding to each spatial structural element in real time during the process of adjusting the 3D bounding box, modifying semantic tags, or regenerating stylized assets.

[0202] The switching display module is used to respond to the user's mode switching command, switching from mixed reality mode to virtual reality mode to display the adjusted or regenerated complete virtual scene.

[0203] Another embodiment of this application provides an electronic device, such as... Figure 6 As shown, it includes: memory 601 and processor 602.

[0204] The memory 601 is used to store computer programs.

[0205] The processor 602 is used to execute a computer program, which, when executed, is specifically used to implement the spatially anchored virtual environment transformation method provided in any of the above embodiments.

[0206] The electronic devices mentioned in this article can be servers, PCs, tablets, mobile phones, ECUs (Electronic Control Units), VCUs (Vehicle Control Units), MCUs (Micro Controller Units), HCUs (Hybrid Control Units), etc.

[0207] Another embodiment of this application provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, is used to implement the spatially anchored virtual environment transformation method provided in any of the above embodiments.

[0208] Computer-readable storage media include both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0209] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0210] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0211] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A virtual environment transformation method based on spatial anchoring, characterized in that, The method includes: Acquire visual data of the physical environment and reconstruct the three-dimensional point cloud structure of the physical environment based on the visual data; Identify the spatial structural elements in the three-dimensional point cloud structure, and construct a spatial framework using the geometric and semantic information of the spatial structural elements; In response to the target theme input by the user, the target theme is parsed to extract global style features, and local stylistic constraints for the furniture elements are generated based on the functional attributes of the furniture elements in the spatial framework. Based on the spatial framework, the global style features, and the local stylization constraints, stylized assets corresponding to each spatial structural element in the spatial framework are generated. The stylized assets are spatially aligned according to the pose information of the corresponding spatial structural elements in the spatial framework, so that the stylized assets visually overlap with the real structures in the physical environment.

2. The method according to claim 1, characterized in that, The step of acquiring visual data of the physical environment and reconstructing the three-dimensional point cloud structure of the physical environment based on the visual data includes: Continuously acquire video frame sequences of the physical environment within a preset duration; The camera pose of each frame in the video frame sequence is recovered using localization and mapping methods, and dense point cloud data of the physical environment is generated based on the recovered camera pose. Based on the indoor structural features of the physical environment, the dense point cloud data is realigned to obtain a three-dimensional point cloud structure of the physical environment aligned in a unified world coordinate system.

3. The method according to claim 1, characterized in that, The process of identifying spatial structural elements in the 3D point cloud structure and constructing a spatial framework using the geometric and semantic information of these spatial structural elements includes: Semantic decomposition is performed on the three-dimensional point cloud structure and the acquired video frame sequence to identify the spatial structural elements in the physical environment, and a three-dimensional bounding box with semantic labels is generated for each spatial structural element. The three-dimensional bounding box serves as a spatial container for the physical environment, constituting the basic geometric and semantic constraints of the spatial framework. For each target furniture element in the spatial structural elements, excluding walls and the ground, the observation quality of the target furniture element in each frame is evaluated from the video frame sequence to obtain the evaluation result; Based on the evaluation results, a reference frame is selected for each target furniture element, and the camera pose of the reference frame is recorded as the reference pose of the target furniture element. Each of the reference poses is associated with the spatial frame to construct a three-dimensional bounding box containing each spatial structural element and its reference pose.

4. The method according to claim 1, characterized in that, The process of responding to a target theme input by the user, parsing the target theme to extract global style features, and generating local stylistic constraints for the furniture elements based on the functional attributes of the furniture elements in the spatial framework includes: Obtain the target theme input by the user through text or example image, and parse the target theme to extract core style keywords as global style features; The global style features, the 3D bounding boxes of each spatial structural element in the spatial framework, and their semantic labels are input into the mapping module to generate a style mapping table. The style mapping table includes at least the local stylization constraints corresponding to each furniture element. The local stylization constraints include the functional semantics retained by the furniture element, the stylized object type and potential collision risk markers, and the generated prompts.

5. The method according to claim 3, characterized in that, The process of generating stylized assets corresponding to each spatial structural element in the spatial framework, based on the spatial framework, the global style features, and the local stylization constraints, includes: For each furniture element in the spatial frame, a two-dimensional stylized image of the furniture element is generated based on the reference frame corresponding to the furniture element and the generation prompt in the local stylization constraint; A three-dimensional stylized model of the furniture element is reconstructed based on the two-dimensional stylized image; While reconstructing the three-dimensional stylized model of the furniture elements, surface texture assets corresponding to the wall elements and ground elements in the spatial frame are generated according to the global style features and the local stylization constraints, and environmental background assets corresponding to the environmental background in the spatial frame are generated according to the global style features. The three-dimensional stylized models, surface texture assets, and environmental background assets of each furniture element are used as stylized assets corresponding to each spatial structural element in the spatial framework.

6. The method according to claim 1, characterized in that, The step of spatially aligning the stylized assets according to the pose information of the corresponding spatial structural elements in the spatial frame, so that the stylized assets visually overlap with the real structures in the physical environment, includes: For the three-dimensional stylized model in the stylized asset, the three-dimensional stylized model is transformed according to the three-dimensional bounding box of the corresponding spatial structural element in the spatial frame, so that the size of the three-dimensional stylized model is adapted to the three-dimensional bounding box and the orientation is matched with the actual orientation of the spatial structural element. The position of the transformed 3D stylized model is adjusted to align with the spatial reference of the 3D bounding box; For the surface texture assets in the stylized assets, the surface texture assets are mapped to the surface areas of the corresponding wall and ground elements in the spatial frame; for the environmental background assets in the stylized assets, the environmental background assets are loaded and presented as the panoramic background of the virtual scene. All stylized assets, after alignment, are integrated into a complete virtual scene, so that each stylized asset in the virtual scene visually overlaps with the corresponding real structure in the physical environment.

7. The method according to claim 1, characterized in that, After spatially aligning the stylized assets according to the pose information of the corresponding spatial structural elements in the spatial framework, so that the stylized assets visually overlap with the real structures in the physical environment, the process further includes: In response to a user's adjustment command for the three-dimensional bounding box corresponding to any spatial structural element in the spatial frame in mixed reality mode, the size, position, or orientation of the three-dimensional bounding box is adjusted. In response to the user's command to add or delete bounding boxes of the spatial frame in mixed reality mode, add or delete three-dimensional bounding boxes and their associated spatial structural elements; In response to a user's instruction to modify the semantic label of any spatial structural element, update the semantic label of that spatial structural element and use the updated semantic label when regenerating the stylized asset; In response to the user's command to regenerate any stylized asset, the stylized asset is regenerated and replaced with the original stylized asset. The step of spatially aligning the stylized asset according to the pose information of the corresponding spatial structural element in the spatial frame is re-executed so that the stylized asset visually overlaps with the real structure in the physical environment. During the process of adjusting the 3D bounding box, modifying semantic tags, or regenerating the stylized assets, the update status of the stylized assets corresponding to each spatial structural element is displayed in real time. In response to a user's mode switching command, the system switches from the mixed reality mode to the virtual reality mode to display the adjusted or regenerated complete virtual scene.

8. A virtual environment transformation system based on spatial anchoring, characterized in that, The system includes: The spatial understanding module is used to acquire visual data of the physical environment and reconstruct the three-dimensional point cloud structure of the physical environment based on the visual data. A construction module is used to identify spatial structural elements in the three-dimensional point cloud structure and to construct a spatial framework using the geometric and semantic information of the spatial structural elements. The style extraction and mapping module is used to respond to the target theme input by the user, parse the target theme to extract global style features, and generate local stylization constraints for the furniture elements based on the functional attributes of the furniture elements in the spatial framework. The content generation module is used to generate stylized assets corresponding to each spatial structural element in the spatial framework based on the spatial framework, the global style features, and the local stylization constraints. The scene integration and presentation module is used to spatially align the stylized assets according to the pose information of the corresponding spatial structural elements in the spatial framework, so that the stylized assets visually overlap with the real structures in the physical environment.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements a spatially anchored virtual environment transformation method as described in any one of claims 1 to 7.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement a spatially anchored virtual environment transformation method as described in any one of claims 1 to 7.