Cross-platform virtual-real fusion scene construction method and system based on AI space calculation
By employing a cross-platform virtual-real fusion method based on AI spatial computing, and utilizing generative adversarial networks and 3D reconstruction algorithms for multimodal modeling, combined with GPU rendering and deep learning distribution, this approach addresses the issues of low efficiency, high latency, and poor adaptability in existing virtual-real fusion systems, enabling the construction of efficient and low-latency cross-platform virtual-real fusion scenarios.
Patent Information
- Application Number
- CN202511501268.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-01-23
AI Technical Summary
Existing virtual-real fusion systems suffer from low content generation efficiency, high real-time interaction latency, high deployment costs, and poor adaptability in cross-platform live streaming or dynamic interactive scenarios. In particular, they lack the ability to uniformly model multimodal inputs, and the rendering and distribution processes are fragmented, making it difficult to dynamically optimize visual quality based on network conditions.
A cross-platform virtual-real fusion method based on AI spatial computing is adopted. Multi-scale feature fusion is performed through generative adversarial networks, implicit field coding is performed by combining 3D reconstruction algorithms, a two-way data channel between virtual scenes and physical sensors is established, hybrid rendering is performed using GPU clusters, and distribution is performed by combining deep learning bitrate adaptive algorithms.
It achieves end-to-end automated construction from user input to virtual-real fusion scenarios, improves the flexibility of content generation and modeling accuracy, enhances the real-time interactivity and immersion, and ensures the real-time output of high-fidelity visual effects and simultaneous streaming across multiple platforms.
Smart Images

Figure CN121392201A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and spatial computing technology, and in particular to a cross-platform virtual-real fusion scene construction method and system based on AI spatial computing. Background Technology
[0002] Existing virtual-real fusion systems typically employ a step-by-step process, relying on manual design of 3D models and streaming through independent engines, resulting in low content generation efficiency, high real-time interaction latency, and difficulty in meeting real-time requirements.
[0003] Especially in cross-platform live streaming or dynamic interactive scenarios, traditional methods lack the ability to uniformly model multimodal inputs (such as text and image commands), and the synchronization of virtual content with physical space relies on high-precision external positioning devices, resulting in high deployment costs and poor adaptability. In addition, the rendering and distribution processes are disconnected, making it difficult to dynamically optimize visual quality based on network conditions, thus affecting user experience.
[0004] Therefore, there is an urgent need to provide an end-to-end collaborative optimization method for constructing virtual-real fusion. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this application provides an end-to-end collaborative optimization method for constructing virtual-real fusion scenarios applicable to cross-platform live streaming or dynamic interactive scenarios, achieving high efficiency and low latency from input to output. This application also provides a cross-platform virtual-real fusion scene construction method and system based on AI spatial computing.
[0006] Firstly, the objective of this invention is achieved through the following technical solution: A cross-platform virtual-real fusion scene construction method based on AI spatial computing includes: Based on the received cross-modal conversion instruction set, multi-scale feature fusion is performed through generative adversarial network to generate an initial image sequence; and the target object is implicitly field-coded by a 3D reconstruction algorithm to obtain an initial parameterized model. Spatiotemporal alignment is performed on the multi-view video streams acquired by XR devices. Digital scene asset packages are loaded by combining physical sensor data and a preset spatial index structure to establish a two-way data channel and mapping interaction rules between virtual scenes and physical sensors. A hybrid rendering pipeline that integrates rasterization and ray tracing is run on a GPU cluster to render and adjust the lighting parameters of the initial parametric model, resulting in an optimized parametric model with high-fidelity visual flow. The optimized parameter model is processed by differentiated encoding according to regional characteristics to obtain the target virtual-real scene fusion model; the bitrate resource is allocated by combining deep learning bitrate adaptive algorithm to distribute the target virtual-real scene fusion model to the preset terminal.
[0007] By adopting the above technical solutions, the cross-modal conversion instruction set includes cross-modal conversion instructions composed of text, images, and line drawings. This application achieves end-to-end automated construction from user input to a virtual-real fusion scene through the fusion of multimodal generation and spatial computing. Utilizing generative adversarial networks (such as AI models of Transformer-GAN) and 3D implicit reconstruction techniques (such as the NeRF 3D reconstruction algorithm), it supports unified processing of cross-modal instructions such as text, images, and line drawings, significantly improving the flexibility of content generation and modeling accuracy. It performs spatiotemporal alignment of multi-view video streams through edge computing nodes, achieving high-precision spatiotemporal alignment of virtual and real spaces through edge computing and sensor fusion, enhancing real-time interaction and immersion. It employs a hybrid rendering pipeline combined with GPU parallel acceleration to ensure real-time output of high-fidelity visual effects. Combining intelligent encoding and deep learning bitrate control, it achieves adaptive distribution of high-quality content on multiple platform terminals, enabling simultaneous streaming across multiple platforms. The overall solution outperforms traditional methods in terms of generation efficiency, rendering quality, and system response latency, making it suitable for high-concurrency, low-latency application scenarios such as live e-commerce, remote collaboration, and digital twins.
[0008] In a preferred example, this application further includes the following method before implicit field encoding of the target object using a 3D reconstruction algorithm to obtain an initial parametric model: Acquire environmental perception data of the target physical space, including depth images, RGB images, inertial measurement data, and environmental sound field information; Based on the environmental perception data, a three-dimensional spatial digital model containing geometric structure and semantic labels is generated using a pre-trained spatial semantic reconstruction model. During the construction of the three-dimensional spatial digital model, the pose alignment deviation between the current reconstructed region and the global spatial reference is calculated; Spatial correction processing is performed on the pose alignment deviation to obtain spatial correction regression parameters; Based on the spatial correction regression parameters, the spatial deformation error under the current environmental conditions is calculated, and the point cloud fusion strategy and texture mapping rate are simultaneously optimized based on the spatial deformation error.
[0009] By employing the aforementioned technical solutions, multimodal environmental perception data, including depth images, RGB images, inertial measurement data, and environmental sound field information, were acquired, enabling comprehensive perception of the physical space. Multi-source data fusion effectively overcomes the limitations of single-modality approaches such as illumination variations, texture loss, and motion blur. A pre-trained spatial semantic reconstruction model generates a 3D spatial digital model containing geometric structures and semantic labels. This not only reconstructs the geometric shape of the space but also endows it with semantic information (such as "desktop," "door," and "corridor"), providing semantic context support for the intelligent layout and interactive logic design of subsequent virtual content. Furthermore, during the modeling process, the pose alignment deviation between the current reconstructed region and the global spatial reference is calculated in real time, and spatial correction regression parameters are generated through spatial correction processing. This mechanism effectively suppresses the "drift" phenomenon caused by accumulated errors in the SLAM system during long-term operation, ensuring the long-term consistency of spatial modeling.
[0010] In a preferred example of this application: the step of calculating the spatial deformation error under the current environmental conditions based on the spatial correction regression parameters, and simultaneously optimizing the point cloud fusion strategy and texture mapping rate based on the spatial deformation error, includes: obtaining the static environmental geometric stability coefficient corresponding to the current spatial region and the dynamic pose disturbance intensity corresponding to the pose alignment deviation based on the spatial correction regression parameters; Based on the geometric stability coefficient and the dynamic pose disturbance intensity, calculate the spatial deformation error of the three-dimensional spatial digital model in the current environment; Adjust the voxel filtering resolution of the point cloud data according to the spatial deformation error, and control the number of point cloud registrations and feature extraction density within the same time window according to the voxel filtering resolution to form a point cloud fusion optimization strategy. Based on the spatial consistency enhancement results generated by the point cloud fusion optimization strategy, the sampling frequency and UV unwrapping accuracy in the texture mapping process are dynamically adjusted to optimize the texture mapping rate.
[0011] By adopting the above technical solution, the static environment geometric stability coefficient reflects whether the environment itself is conducive to stable reconstruction. For example, areas with rich textures have high stability, while areas with white walls have low stability. The dynamic pose disturbance intensity reflects whether the device movement is violent, such as rapid head rotation causing pose jitter. This facilitates the analysis of the source of spatial errors. The spatial morphological error calculated based on the comprehensive analysis of these two parameters is more consistent with the complexity of actual reconstruction scenarios. Then, the spatial consistency enhancement result generated by the point cloud fusion optimization strategy guides the adjustment of the sampling frequency and UV unwrapping accuracy in the texture mapping process, realizing the synergistic optimization of geometry and texture processing and avoiding the inconsistency problem of "accurate geometry but blurry texture" or "clear texture but incorrect geometry".
[0012] In a preferred embodiment of this application, the dynamic adjustment of the sampling frequency and UV unwrapping accuracy during the texture mapping process further includes: Obtain the visual distortion angle of the completed texture mapping block and the mesh reconstruction resolution of the three-dimensional spatial digital model; Calculate the illumination compensation parameters for the next mapping block based on the visual distortion angle, and perform visual consistency compensation for the brightness difference of the previous mapping block based on the illumination compensation parameters. Based on the mesh reconstruction resolution and the pixel density of the corresponding texture atlas, adjust the amount of texture mapping data and the multi-level detail hierarchy of the next mapping block; The texture mapping amount T is calculated using formula (1): T=α×β×γ×S (1) Where α is the texture compression coefficient, ranging from 0.3 to 1.0, β is the LOD adjustment factor, γ is the illumination compensation weight, and S is the geometric area of the surface to be mapped. Based on the spatial anchor point positions of the constructed region in the three-dimensional spatial digital model, the edge alignment parameters of each spatial anchor point are obtained, and the spatial edge change value of the visual distortion angle is calculated after visual consistency compensation. The UV coordinate alignment parameters of the next mapping block are adjusted based on the spatial edge change value, and the overall texture mapping process is optimized by combining the adjusted UV coordinate alignment parameters and the amount of texture mapping data.
[0013] By employing the above technical solution, the system obtains the "visual distortion angle" and "mesh reconstruction resolution" of the completed mapping blocks, enabling it to quantify the quality defects of the current texture mapping. The visual distortion angle reflects the stretching or compression caused by geometric curvature during UV unwrapping, directly reflecting texture distortion problems. Based on the visual distortion angle, the system calculates the lighting compensation parameters for the next mapping block and compensates for the brightness differences of the previous mapping block, achieving a smooth transition with visual consistency across blocks. This effectively alleviates problems such as uneven brightness and color jumps at seams caused by block mapping.
[0014] In a preferred embodiment of this application, the spatial correction processing of the pose alignment deviation to obtain spatial correction regression parameters specifically includes: During the construction of the three-dimensional spatial digital model, the actual pose angle of the current frame point cloud data in the target physical space relative to the global coordinate system is obtained; Obtain the ideal spatial axis corresponding to the current construction position, and calculate the pose alignment deviation value between the actual pose angle and the ideal spatial axis. Based on the pose alignment deviation value, the construction trend of the three-dimensional spatial digital model is analyzed to obtain the spatial drift trend analysis results in the current construction process; Based on the spatial drift trend analysis results, the corresponding continuous drift intervals are identified, and the cumulative error within the continuous drift intervals is compensated in a segmented manner to obtain the spatial correction regression parameters of the three-dimensional spatial digital model under the current drift state.
[0015] By employing the above technical solution, the actual pose angle of the current frame point cloud relative to the global coordinate system is obtained. Combined with the ideal axis, the pose alignment deviation value is calculated, achieving precise quantification of pose error. Based on the deviation value analysis, a trend is constructed to identify the direction and magnitude of spatial drift, providing trend prediction capabilities. Furthermore, continuous drift intervals are identified, and cumulative errors are compensated in segments to achieve gradual correction. This mechanism avoids visual abrupt changes caused by large-scale corrections at once, maintaining user experience smoothness while ensuring registration accuracy. It effectively suppresses cumulative drift during long-term operation of SLAM (Simultaneous Localization and Mapping) systems.
[0016] In a preferred embodiment of this application: the step of identifying the corresponding continuous drift interval based on the spatial drift trend analysis results, and performing piecewise reverse compensation on the cumulative error within the continuous drift interval to obtain the spatial correction regression parameters of the three-dimensional spatial digital model in the current drift state includes: Obtain the environmental texture feature density under the current drift state and the current reconstructed pose of the three-dimensional spatial digital model; Based on the environmental texture feature density and the current reconstruction pose, analyze the optimal spatial regression path that fits the pose deviation back to the global coordinate system reference range during the spatial reconstruction process. The optimal spatial regression path is segmented into multiple scales, and the pose compensation angle of the current drift interval is calculated based on the spatial alignment accuracy of the previous segment. Based on the pose compensation angle of the previous paragraph, the spatial interpolation step size and ICP registration threshold of the next paragraph are dynamically adjusted to obtain the multi-level spatial correction regression parameters of the three-dimensional spatial digital model in the current drift state.
[0017] By employing the aforementioned technical solution, based on the environmental texture feature density and the current reconstructed pose, the optimal spatial regression path for returning pose deviation to the global baseline is analyzed, ensuring a smooth and natural correction process. Next, the optimal spatial regression path is segmented into multiple scales to achieve phased, progressive correction. The spatial alignment accuracy of the previous segment guides the pose compensation angle of the current segment, forming a feedforward adjustment mechanism. The spatial interpolation step size and ICP registration threshold of the next segment are dynamically adjusted to achieve an adaptive correction strategy that combines high efficiency and high accuracy, which is beneficial for improving the robustness and visual comfort of the spatial correction process.
[0018] Secondly, the objective of this invention is achieved through the following technical solution: A cross-platform virtual-real fusion scene construction system based on AI spatial computing is used to execute the cross-platform virtual-real fusion scene construction method based on AI spatial computing as described above. The system includes: The multimodal instruction parsing module is used to receive cross-modal conversion instruction sets and parse them into executable semantic operation sequences; the feature fusion and modeling module is used to perform multi-scale fusion of the visual features corresponding to the semantic operation sequences through a generative adversarial network to generate an initial image sequence, and combine it with a 3D reconstruction algorithm to perform implicit field encoding on the target object and output an initial parameterized model. The data processing and scene loading module is used to connect to the multi-view video stream input interface of the XR device, perform spatiotemporal synchronization processing on the video stream, integrate physical sensor data, load digital scene asset packages in combination with a preset spatial index structure, and establish a two-way data channel and mapping interaction rules between the virtual scene and the physical sensor. The hybrid rendering computing module, deployed on a GPU cluster, is used to run a hybrid rendering pipeline that integrates rasterization and ray tracing. It performs high-fidelity rendering processing on the initial parameterized model and dynamically adjusts the material reflectivity, shadow intensity, and global illumination parameters based on the ambient lighting estimation results to generate an optimized parameter model with high-fidelity visual flow. The multi-platform streaming module is used to perform differentiated encoding processing on the optimized parameter model according to regional characteristics to obtain the target virtual-real scene fusion model; and to allocate bitrate resources by combining deep learning bitrate adaptive algorithm to distribute the target virtual-real scene fusion model to preset terminals.
[0019] Thirdly, the objective of this invention is achieved through the following technical solution: A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method for constructing cross-platform virtual-real fusion scenes based on AI spatial computing.
[0020] Fourthly, the objective of this invention is achieved through the following technical solution: A computer program product includes a computer program / instructions that, when executed by a processor, implement the steps of the cross-platform virtual-real fusion scene construction method based on AI spatial computing as described above.
[0021] In summary, this application includes at least one of the following beneficial technical effects: 1. Through the collaborative design of multimodal AI generation, implicit 3D modeling, multi-source sensor fusion, hybrid rendering and intelligent distribution, the end-to-end automated construction from user intent to cross-platform virtual-real fusion presentation is realized, which significantly improves the intelligence level, visual realism and system compatibility of virtual-real fusion scene generation; 2. This application further introduces a high-precision digital modeling and spatial error correction mechanism for physical space, which focuses on solving the registration error problem caused by "inaccurate spatial perception" in virtual-real fusion. This application effectively solves key technical problems such as inaccurate spatial registration, model drift, and texture distortion in virtual-real fusion by using multimodal environmental perception data, spatial semantic reconstruction model, spatial correction regression parameters, calculation of spatial deformation error, and full-process error closed-loop compensation mechanism. Attached Figure Description
[0022] Figure 1 This is a flowchart of a cross-platform virtual-real fusion scene construction method based on AI spatial computing in one embodiment of this application. Detailed Implementation
[0023] The present application will be further described in detail below with reference to the accompanying drawings.
[0024] In one embodiment, such as Figure 1 As shown, this application discloses a cross-platform virtual-real fusion scene construction method based on AI spatial computing, including the following steps: S1: Based on the received cross-modal conversion instruction set, multi-scale feature fusion is performed through generative adversarial network to generate an initial image sequence; and the target object is implicitly field-coded by a 3D reconstruction algorithm to obtain an initial parameterized model.
[0025] In this embodiment, the cross-modal transformation instruction set includes cross-modal transformation instructions composed of text, images, and line drawings, representing a combination of multi-source information input by the user. The cross-modal transformation instruction set extracts semantic and visual features through a natural language processing module and an image encoder, and inputs them into a Transformer-based generative adversarial network (Transformer-GAN). The GAN's generator adopts a U-Net architecture, combining an attention mechanism to perform feature fusion at different scales (e.g., 128×128, 256×256, 512×512), outputting a semantically consistent and detailed initial image sequence.
[0026] Subsequently, a NeRF (Neural Radiance Fields) type 3D reconstruction algorithm is used to implicitly encode the target object. Specifically, the initial image sequence and its corresponding camera pose (estimated by the SfM algorithm) are used as input, and a neural network learns from the spatial coordinates (x, y, z) and viewpoint direction. A continuous mapping function to color (r, g, b) and density σ Equivalent to The training loss function is: L = λ1||I render -I gt ||+λ2L reg Where λ1 = 1.0 is the weighting coefficient of the data fidelity term; λ2 = 0.1 is the weighting coefficient of the regularization term; I render Images generated by rendering from a NeRF model; I gt A real image, i.e., an actual photograph or a known image of the target object; L reg This is the TV regularization term. After training, the implicit field is decoded into a parametric mesh model containing vertex, normal, and texture coordinates. The implicit field encoding does not rely on traditional mesh representations, but stores 3D geometry and appearance information in the form of neural network weights. Finally, it is decoded into an initial parametric model containing topology and material properties, in USDZ or GLB format, for easy use in subsequent rendering.
[0027] Specifically, in practical applications, a comprehensive technology platform integrating AI spatial computing, virtual live streaming, digital twins, and 3D modeling is provided, called the AI spatial computing platform. This platform extracts text semantic feature vectors using the BERT model, extracts image visual features using ResNet-50, and converts line drawings into binary image features after edge detection. These three types of features are fused by a feature alignment module (using a cross-attention mechanism) and then input into the generator of a Generative Adversarial Network (GAN). The GAN generator introduces multi-head self-attention modules (8 heads) at each level of the encoder and decoder. The discriminator uses a PatchGAN structure, combined with the CLIP (Contrastive Language-Image Pre-training) model to calculate cross-modal consistency loss. During training, L1+perceptual loss+adversarial loss are jointly optimized, outputting an initial image sequence with a resolution of 512×512.
[0028] S2: Perform spatiotemporal alignment on the multi-view video streams acquired by the XR device, load the digital scene asset package by combining physical sensor data and preset spatial index structure, and establish a two-way data channel and mapping interaction rules between the virtual scene and the physical sensor.
[0029] In this embodiment, the XR device includes, but is not limited to: AR glasses, or a camera array formed by at least three wide-angle cameras at different mounting angles, at least one depth camera, and an IMU (Inertial Measurement Unit) sensor on the moving target object. Multi-view synchronous video acquisition is performed through the arrangement of cameras in the actual physical space to obtain a real-time multi-view video stream. To eliminate frame misalignment caused by transmission delay, a timestamp alignment mechanism based on PTP (Precise Time Protocol) is adopted, combined with optical flow methods (such as...). The algorithm performs sub-pixel-level spatial registration to achieve spatiotemporal synchronization. Physical sensor data includes point cloud data from acceleration, angular velocity, and depth cameras. The AI spatial computing platform's application system pre-configures an Octree structure as the spatial index, dividing the physical space into multi-level cubic voxels, supporting rapid retrieval and collision detection. Based on the current user pose and gaze direction, it dynamically loads the corresponding digital scene asset package (such as a prefab from Unity or Unreal Engine).
[0030] Specifically, physical sensor data is collected uniformly through the ROS (Robot Operating System) middleware, including IMU acceleration and angular velocity data, and point cloud data from depth cameras. Point cloud data requires ≥30,000 points per frame. An octree structure is used to model the physical space, with the root node covering a 10m × 10m × 3m area and a maximum depth of 6 levels. Dynamic insertion and deletion of voxel nodes are supported for rapid collision detection and occlusion assessment.
[0031] Furthermore, a bidirectional data channel is established between the virtual scene engine and physical sensors via the WebRTC protocol: on the one hand, sensor data is used to drive the movement of virtual characters and update lighting conditions; on the other hand, interactive events in the virtual scene (such as clicking virtual buttons) can trigger feedback from physical devices (such as vibration cues). The signaling server is deployed on edge nodes, using the SDP protocol to negotiate media and data channels; the data channel uses DataChannel to transmit control commands, with a Protobuf encapsulation format and a transmission frequency of 60Hz. Mapping interaction rules are defined by the event-driven engine, such as "when a user stares at a virtual product for more than 3 seconds, an automatic purchase link pops up."
[0032] S3: Runs a hybrid rendering pipeline that combines rasterization and ray tracing on a GPU cluster to render and adjust the lighting parameters of the initial parametric model, resulting in an optimized parametric model with high-fidelity visual flow.
[0033] In this embodiment, the GPU cluster consists of multiple servers (e.g., four) equipped with NVIDIA A100 graphics cards, deploying a hybrid rendering pipeline: rasterization is used for distant or static backgrounds to ensure high frame rates; ray tracing is enabled for key objects in the foreground (e.g., the main body of the product) to simulate realistic lighting effects, including specular reflection, ambient occlusion (SSAO), and global illumination. For example, the hybrid rendering pipeline employs a dynamic switching strategy: when a virtual object is more than 3 meters away from the camera or occupies less than 15% of the screen area, rasterization mode (forward rendering) is enabled; otherwise, it switches to ray tracing mode. In ray tracing mode, the path tracing has a maximum bounce count of 4. The switching process uses a gradual transition (e.g., fade-in / out time of 100ms) to avoid abrupt screen changes.
[0034] Specifically, the system integrates an ambient lighting estimation module. Utilizing HDR images of real-world scenes captured by a camera, it extracts the direction (azimuth and elevation) and color temperature (K-value) of the main light source through a CNN network (based on the HDRNet architecture). This data drives the dynamic changes in the PBR material parameters (such as metallic and roughness) of virtual objects, adjusting their material reflectivity and shadow intensity. For example, when the ambient light dims, the edge lighting effects of virtual goods are automatically enhanced to maintain visibility. The final output is a high-fidelity visual stream at 60 frames per second with a 4K resolution, encapsulated as an optimized parameter model containing texture maps, normal maps, PBR material parameters, and animation skeletal information.
[0035] S4: Perform differentiated encoding processing on the optimized parameter model according to regional characteristics to obtain the target virtual-real scene fusion model; combine deep learning bitrate adaptive algorithm to allocate bitrate resources to distribute the target virtual-real scene fusion model to the preset terminal.
[0036] In this embodiment, to adapt to different network conditions and terminal performance, the system performs intelligent encoding on the rendered output. A lightweight SegFormer model is used for semantic segmentation based on the characteristics of the image regions, identifying areas such as background, foreground tasks, virtual products, and UI elements: static background areas are encoded using WebP or AV1, resulting in high compression ratios and good subjective quality; dynamic foreground areas (such as anchor actions and virtual effects) are encoded using H.265 to preserve motion details, and transparent layers are extracted separately into WebP format. The encoding parameters are jointly determined by the region motion intensity, texture complexity, and saliency detection results.
[0037] Bitrate allocation is controlled by the DeepQoS deep learning model, which is based on an LSTM network architecture. It monitors network bandwidth, latency, and packet loss rate in real time (e.g., over the past 10 seconds) and combines this with the terminal device type (mobile phone, PC, VR headset) to output the target bitrate for each area. For example, in a 5G MEC (Multi-access Edge Computing) environment, priority is given to ensuring the bitrate of critical areas; in a weak 4G network environment, the resolution of non-critical areas is reduced to maintain smoothness. The final generated target virtual-real scene fusion model is distributed through CDN edge nodes, supporting RTMP for live streaming, HLS for web playback, and DASH for adaptive streaming media, enabling cross-platform synchronous transmission to e-commerce platforms, social media, or enterprise self-built systems.
[0038] In one embodiment, before implicitly encoding the target object using a 3D reconstruction algorithm to obtain an initial parameterized model, the cross-platform virtual-real fusion scene construction method based on AI spatial computing further includes: S10: Acquire environmental perception data of the target physical space, including depth images, RGB images, inertial measurement data, and environmental sound field information.
[0039] In this embodiment, the target physical space refers to the user's real environment, such as a live broadcast room, exhibition hall, or film shooting location. Depth images are acquired by a Time-of-Flight (ToF) or structured light depth camera, with a resolution of 640×480 and a frame rate of 30fps. RGB images are simultaneously acquired by a wide-angle color camera, with a resolution of 1920×1080. Inertial measurement data (IMU data) includes measurement data output from a three-axis accelerometer and a three-axis gyroscope. Ambient sound field information is obtained by acquiring spatial audio signals through a microphone array (no fewer than four channels), with a sampling rate of 48kHz.
[0040] S20: Based on environmental perception data, a three-dimensional spatial digital model containing geometric structure and semantic labels is generated using a pre-trained spatial semantic reconstruction model.
[0041] In this embodiment, the spatial semantic reconstruction model is a multimodal fusion neural network based on the Transformer architecture. Its structure includes a visual encoding branch, an IMU encoding branch, an audio encoding branch, and a cross-modal fusion module. The visual encoding branch uses ResNet-34 to extract RGB image features and PointNet++ to process the point cloud features generated from the depth map. The IMU encoding branch uses a 1D-CNN+BiLSTM network to extract motion temporal features. The audio encoding branch uses a VGGish network to extract sound field spectral features (Mel-spectrogram). The cross-modal fusion module aligns and fuses visual, motion, and audio features through a cross-attention mechanism, outputting a unified feature tensor. The spatial semantic reconstruction model is pre-trained on a large-scale indoor scene dataset, outputting a voxel grid with semantic labels or a mesh model with category labels, such as a wall (label=1, color=gray) and a floor (label=2, color=brown). The three-dimensional spatial digital model in this embodiment not only includes precise geometric boundaries but also assigns semantic attributes to each region.
[0042] S30: During the construction of the three-dimensional spatial digital model, calculate the pose alignment deviation between the current reconstructed area and the global spatial reference.
[0043] In this embodiment, during continuous reconstruction, due to IMU drift, missing visual features, or dynamic occlusion, cumulative errors may occur in local reconstruction areas, leading to a shift from the initial coordinate system (i.e., the "global spatial reference"). The global spatial reference refers to a fixed coordinate system established with the user's initial standing position as the origin, geographic north as the Y-axis, and the gravity direction as the -Z-axis. The global spatial reference is a fixed reference coordinate system. The system uses the SLAM (Simultaneous Localization and Mapping) algorithm to estimate the camera pose T of the current frame in real time. local ∈SE(3), and the desired pose T in the global coordinate system. global Compare them.
[0044] Pose alignment deviation is defined as: Where ΔT is the 6-DOF error, including 3 translations and 3 rotations, and is represented by the Lie algebra SE(3) as a 6-dimensional vector ξ={δ x ,δ y ,δ z ,δθ x ,δθ y ,δθ z} T The system calculates ΔT every 5 frames. When the rotation deviation exceeds 1.5° or the translation deviation exceeds 3cm, the subsequent correction mechanism is triggered.
[0045] S40: Perform spatial correction processing on the pose alignment deviation to obtain spatial correction regression parameters.
[0046] In this embodiment, to eliminate accumulated errors, the system performs spatial correction processing, and the spatial correction regression parameters are a set of transformation matrices used to correct drift.
[0047] Specifically, step S40 includes: S401: During the construction of a three-dimensional digital model, obtain the actual pose angle of the current frame point cloud data in the target physical space relative to the global coordinate system.
[0048] In this embodiment, the actual pose angle refers to the attitude rotation component of the sensor coordinate system in the current acquisition frame relative to the global coordinate system, usually expressed in Euler angle form as (θ). x ,θ y ,θ z The units are degrees (°), which correspond to the rotation angles around the X-axis (pitch), Y-axis (yaw), and Z-axis (roll), respectively.
[0049] Specifically, a visual-inertial SLAM algorithm (such as ORB-SLAM3) is used to estimate the camera pose in real time; after feature extraction (ORB keypoints) of each frame of depth image, it is matched with the previous frame, and nonlinear optimization (using g2o) is performed in combination with the IMU pre-integration results; the optimized output is the 6-DOF pose matrix T of the current frame. current ∈SE(3), its rotational part R∈SO(3) can be converted into Euler angles using the Rodrigues formula; the actual pose angle is this Euler angle. The actual pose angle reflects the instantaneous orientation of the device in real space and is the basis for judging whether spatial misalignment has occurred. For example, in a standard live broadcast room environment, if the system is initially set to point the Y-axis towards the long side of the room (ideally north), then when the user holds the XR device and turns 90° to the right, the yaw angle θ in the actual pose angle is... y It will change from 0° to approximately 88.5°, indicating a slight drift.
[0050] S402: Obtain the ideal spatial axis corresponding to the current construction position, and calculate the pose alignment deviation between the actual pose angle and the ideal spatial axis.
[0051] In this embodiment, the ideal spatial axis refers to the geometric reference direction that the currently reconstructed area should follow under error-free conditions. Its determination depends on the scene's semantic structure: for regular indoor environments (such as rectangular rooms), the ideal axis is determined by the direction of the main walls: the walls are detected through RANSAC plane fitting, and the normal direction of the wall with the largest area is taken as the X / Z principal axis, with the ground normal as the -Z axis; for industrial equipment layouts, the ideal axis can be pre-imported into the design coordinate system in the CAD drawings; if no recognizable structure is found, a local orthogonal coordinate system is established by default along the initial movement direction.
[0052] Specifically, the dominant geometric features of the current reconstruction area, such as corner lines and ceiling edges, are extracted from the three-dimensional spatial digital model generated in step S20, and the spatial axis direction vectors are fitted. Then, the main orientation of the current frame (such as the forward direction) is determined. )and Compare the two and calculate the angle between them: This included angle is the pose alignment deviation value. It represents the degree to which the current reconstruction direction deviates from the ideal structure. The pose alignment deviation value quantifies the inconsistency between local modeling and the overall spatial structure.
[0053] For example, in a corridor scene, if the ideal axis is along the direction of the corridor's extension... The actual forward direction of the equipment changes to [0.96, 0.28, 0] due to drift, resulting in a pose alignment deviation of approximately 16.3°. S403: Based on the pose alignment deviation value, analyze the construction trend of the three-dimensional spatial digital model and obtain the spatial drift trend analysis results in the current construction process.
[0054] In this embodiment, to avoid misjudging instantaneous noise as continuous drift, trend analysis of spatial drift is required. Specifically, a sliding window of length N = 10 frames is maintained to store the pose alignment deviation values {Δθ1, Δθ2, ..., Δθ} for the most recent 10 moments. 10}; Calculate the first difference δ of the sequence. i =Δθ i+1 -Δθ i If 5 consecutive δ i A slope greater than 0.5° is considered a positive drift trend. The linear regression slope is also calculated. Among them, t i Here, i represents the timestamp (frame number) and i represents the sampling point identifier. For the average time; Δθ i The value represents the pose alignment deviation; the linear regression slope k is the quantification of the drift velocity. The Pearson correlation coefficient is calculated. in, The average deviation angle represents the average deviation level within the current window. The value of r ranges from -1 to 1, with units of ° / frame. If k = +1.2° / frame, it means that the deviation increases by 1.2° with each frame, indicating rapid positive drift. If |k| > 0.8 / frame and the correlation coefficient r > 0.7, the spatial drift is considered to have a significant trend. Combining this with Kalman filtering to predict the deviation growth curve for the next 5 frames forms the "spatial drift trend analysis result." Furthermore, an environmental confidence weight is introduced: trend sensitivity is reduced in textured regions with many feature points; the response threshold is increased in regions with fewer features.
[0055] S404: Based on the spatial drift trend analysis results, identify the corresponding continuous drift intervals, and perform segmented reverse compensation on the cumulative error within the continuous drift intervals to obtain the spatial correction regression parameters of the three-dimensional spatial digital model under the current drift state.
[0056] In this embodiment, once a significant drift trend is confirmed, the system performs segmented reverse compensation. In actual execution, continuous drift intervals are first identified. In this embodiment, a continuous drift interval is defined as: starting from the first deviation exceeding the threshold Δθ. th The frame with a deviation of 2.0° ends when the deviation falls back to less than 1.5° in the subsequent three consecutive frames; all point cloud frames within the continuous drift interval are marked as the set to be corrected, F = {F...} i ,F i+1 ,…,F j Arranged in chronological order, F i ,F i+1 ,...,F j For each point cloud frame, there is a set of 3D points and their poses obtained from the depth image transformation; then, the pose matrix T of each frame within the interval is... k Pose matrix T k 4×4 homogeneous transformation matrices belonging to the special Euclidean group E(3): Where R k ∈SO(3) is a rotation matrix. This is the translation vector. Next, its position relative to the first frame T is calculated. i Relative transformation ΔT k =T i -1 T k , where k∈[i,j]; T i -1 The inverse pose matrix is given; the cumulative error is defined as the relative rotation angle of the final frame ||log(Rot(ΔT)| ... j ))||, in radians; where log(·) is a Lie algebra mapping that transforms matrix operations into measurable rotation vectors; Rot(ΔT) j Extraction of the rotated portion from ΔTj Extract its 3×3 rotation submatrix R∈SO(3).
[0057] Specifically, the segmented reverse compensation strategy includes: dividing the drift interval into three sub-segments: the front segment, the middle segment, and the final segment, each segment having an equal length; and applying a reverse rigid body transformation to each segment. Where ξ m ∈SE(3) is the Lie algebra of the average error of this segment; m is the number of sub-segments; α m As compensation coefficients, α1 = 0.3, α2 = 0.3, and α3 = 0.3 are chosen to gradually increase the correction strength and avoid abrupt changes. After compensation, point cloud registration and mesh fusion are re-performed. Finally, a set of spatial correction regression parameters is generated. Each parameter is a 4×4 homogeneous transformation matrix, recorded in the metadata.
[0058] Further, step S404 includes: S4041: Obtain the environmental texture feature density and the current reconstructed pose of the 3D spatial digital model under the current drift state.
[0059] In this embodiment, to achieve adaptive and robust spatial regression control, the current environmental state and the pose of the 3D spatial digital model are perceived in real time. Environmental texture feature density refers to the number of effective visual feature points that can be extracted per unit area. This is achieved by dividing the current frame image into a grid (e.g., a 5×5 region); detecting keypoints in each grid and removing duplicate or blurry points, and then calculating the global average density (i.e., environmental texture feature density). Environmental texture feature density = number of effective keypoints / estimated projected area of the field of view, unit: m². 2 Current reconstruction posture T curr This is the pose matrix of the point cloud in the current frame in the global coordinate system. It is output by the front-end VO (Visual Odometry) or VIO (Visual-Inertial Odometry) module.
[0060] S4042: Based on the environmental texture feature density and the current reconstruction pose, analyze the optimal spatial regression path that fits the pose deviation back to the global coordinate system reference range during the spatial reconstruction process.
[0061] In this embodiment, the global coordinate system reference range refers to the ideal reference frame established during the initial modeling of the system, such as the pose of the first frame. The regression to the reference range represents the relative rotation angle ||log(Rot(ΔT)) that causes the current pose deviation. j If the drift angle is less than a preset threshold, such as 5°, and the translation error is less than 0.1m, the optimal spatial regression path is from the current drift state T. curr To the target reference state T targetA series of intermediate pose sequences {T0,T1,...,T n The optimal spatial regression path satisfies geometric continuity (no abrupt changes), visual stability (no user dizziness), and convergence (eventually approaching the baseline). Geodesic interpolation on a Lie group can be used to find the optimal spatial regression path. s∈(0,1), or B-spline curve fitting is used to optimize path smoothness and energy consumption in SE(3) space. The path feasibility evaluation coefficient can be obtained by weighting the dynamic pose disturbance intensity and the environmental texture feature density correlation weight value. Compensation is started only when the path feasibility evaluation coefficient is greater than the preset feasibility threshold (such as 0.4); otherwise, regression is delayed and waiting for better conditions.
[0062] S4043: Perform multi-scale segmentation on the optimal spatial regression path, and calculate the pose compensation angle of the current drift interval based on the spatial alignment accuracy of the previous segment.
[0063] In this embodiment, the spatial alignment accuracy is the residual error after the previous segment has been compensated, expressed as the mean square error (RMSE) after ICP registration. The pose compensation angle = maximum allowable single-segment compensation angle × (1 - (spatial alignment difference of the previous segment / maximum tolerable registration error)), where the maximum tolerable registration error is, for example, 0.05m. If the spatial alignment difference of the previous segment is large, the compensation for this segment is weakened. For example, a multi-scale segmented progressive compensation strategy includes a segmented granularity adaptive mechanism: fine-grained compensation in high-texture areas is performed every 0.5 seconds; coarse-grained compensation in low-texture areas is performed every 2 seconds; each segment corresponds to a set of point cloud frames to be compensated.
[0064] S4044: Based on the pose compensation angle of the previous paragraph, dynamically adjust the spatial interpolation step size and ICP registration threshold of the next paragraph to obtain the multi-level spatial correction regression parameters of the three-dimensional spatial digital model in the current drift state.
[0065] In this embodiment, the spatial interpolation step size adjustment rule is as follows: if the previous compensation angle is large, the current step size is reduced; conversely, the compensation can be appropriately increased. Specific adjustment parameters can be defined by the user. The ICP (Iterative Closest Point) registration threshold includes a distance threshold (for removing outliers) and a convergence threshold (for stopping conditions). The dynamic adjustment strategy is as follows: if the previous compensation angle is large (e.g., >2°), the distance threshold is increased, for example, from 0.1m to 0.15m. If the previous compensation angle is small (e.g., <1°), the distance threshold is decreased to 0.08m. If the texture density is low, the maximum number of iterations is increased, for example, from 50 to 100. Multi-level spatial correction regression parameters. in For the m-th segment of the reverse rigid body transformation; Δs m This is the spatial interpolation step size; For ICP registration parameters; ε m denoted as , where m is the expected residual; m is the compensation segment number; and M is the total number of segments, dividing the entire continuous drift interval into M segments. This scheme significantly reduces the jitter rate of the 3D spatial digital model through segmentation, feedback, and adaptive parameter adjustment. The 3D spatial digital model in this embodiment is an important foundation and guiding condition for generating the initial parameterized model. In the constructed 3D spatial digital model, after performing region extraction, data alignment and clipping, implicit field encoding guidance, and parameterized modeling in sequence, a differentiable and continuous initial parameterized model is output.
[0066] S50: Calculate the spatial deformation error under the current environmental conditions based on the spatial correction regression parameters, and simultaneously optimize the point cloud fusion strategy and texture mapping rate based on the spatial deformation error.
[0067] In this embodiment, spatial deformation error is a geometric distortion measure that integrates stability and deviation.
[0068] Specifically, step S50 includes: S501: Based on the spatial correction regression parameters, obtain the static environmental geometric stability coefficient and the dynamic pose disturbance intensity corresponding to the pose alignment deviation of the current spatial region.
[0069] In this embodiment, two key indicators are extracted from the spatial correction regression parameters to quantify the modeling difficulty of the current region. The static environment geometric stability coefficient α stable This reflects the structural reconstructability of the current spatial region under static conditions, with a value range of [0.0, 1.0]. A higher value indicates greater stability. The calculation formula is: α stable =ω1×D texture +ω2×N feature +ω3×C contrast D texture Texture richness is calculated using grayscale variance; the larger the variance, the richer the texture. N feature C represents the number of ORB feature points per unit area. contrast For local contrast, the average gradient magnitude of the Sobel operator was calculated, with ω1 = 0.4, ω2 = 0.4, and ω3 = 0.2, and calibrated experimentally. For example, low-texture areas such as white walls and glass windows are used. stable An α value <0.3 can easily lead to SLAM inaccuracies; high-texture areas such as furniture and decorative walls have an α value of <0.3. stable >0.7, modeling is stable.
[0070] Dynamic pose disturbance intensity β drift This reflects the severity of pose drift caused by equipment movement and is positively correlated with the rate of change of pose alignment deviation, ranging from [0.0, 1.0]. Calculation method: Where ||ΔT corr ||=||log(Rot(T corr ))||,T corr This represents the principal compensation change in the current spatially corrected regression parameters; θ max =10°, serving as the normalization baseline to ensure β drift ≤1.0. For example, β occurs when the user moves or rotates the device rapidly. drift >0.6; <0.2 when stationary or moving slowly.
[0071] S502: Calculate the spatial deformation error of the three-dimensional digital model under the current environment based on the geometric stability coefficient and the intensity of dynamic pose disturbance.
[0072] In this embodiment, the spatial deformation error E deform =(1-α) stable )×β drift +γ'×||ΔT corr ||, among which iTable ft This illustrates the combined error caused by "low stability + high disturbance"; γ'×||ΔT corr The rotational error (in radians) is directly introduced from the correction parameters, with γ' = 0.1 as the weighting coefficient. Spatial deformation error E deform ∈[0.0, 1.0], a larger value indicates that the region is more prone to stretching, twisting, and other deformations. For example, if E deform >0.7: High-risk area, point cloud density needs to be reduced and texture updates temporarily suspended; if E deform <0.3: Low-risk zone, which can improve reconstruction accuracy.
[0073] S503: Adjust the voxel filtering resolution of the point cloud data according to the spatial deformation error, and control the number of point cloud registrations and feature extraction density within the same time window according to the voxel filtering resolution to form a point cloud fusion optimization strategy.
[0074] In this embodiment, to address the risk of spatial deformation, the point cloud fusion optimization strategy needs to be dynamically adjusted. Specifically, the voxel filtering resolution is the spatial granularity of point cloud downsampling; the voxel filtering resolution r voxel Adjust the parameters used for downsampling point clouds to reduce computational load. voxe and l Spatial deformation error E deform Positive correlation: r voxel =r min +(r max -r min )×E deform , where r min =0.02cm; r max =0.1m, for example, if E deformWhen r = 0.8, voxel =9.6cm, significantly reducing point cloud density. The number of point cloud registrations is limited by restricting the number of frames participating in ICP (Iterative Closest Point) registration within a fixed time window (e.g., 1 second): E deform <0.3: Registration 10 frames per second; 0.3≤E deform <0.7: 6 frames per second; E deform ≥0.7: Only 3 frames per second, avoiding the accumulation of mismatches. Feature extraction density adjustment includes retaining 1000 keypoints in high stability regions and reducing them to 300 in high deformation regions.
[0075] S504: Based on the spatial consistency enhancement result generated by the point cloud fusion optimization strategy, the sampling frequency and UV unwrapping accuracy in the texture mapping process are dynamically adjusted to optimize the texture mapping rate.
[0076] In this embodiment, the sampling frequency refers to the image acquisition frequency during texture mapping. UV unwrapping is the process of mapping a 3D mesh to a 2D texture atlas. The accuracy is determined by the resolution of the parameterized mesh: at high accuracy, 1000 UV triangles are generated per unit area; at low accuracy, this is reduced to 300 to reduce computational overhead. The texture mapping rate is optimized to the surface area (m²) of texture mapping completed per unit time. 2 / s), the system uses progressive texture loading: first transmit the low-resolution version (512×512), then wait for E deform Once the image quality drops to a safe threshold, the high-resolution version (2048×2048) is loaded. Texture updates in non-critical areas (such as background walls) are paused during periods of high deformation to prioritize foreground objects.
[0077] Furthermore, step S504 also includes: S5041: Obtain the visual distortion angle of the completed texture mapping block and the mesh reconstruction resolution of the 3D spatial digital model.
[0078] In this embodiment, the visual distortion angle θ distort This refers to the angle between the pixel stretching direction and the ideal orthogonal direction when the current texture block is projected onto the 3D mesh surface due to camera tilt or non-rigid deformation. It is obtained by extracting the normals of the corresponding triangular facets after mapping a texture block. relative to the camera's line of sight The angle between the points is calculated, and the local affine transformation Jacobian matrix J of each pixel is calculated. Its singular value decomposition (SVD) yields the maximum stretching direction. The visual distortion angle is defined as the average deviation between the maximum stretching direction and the principal axis of the patch. in, In the direction of stretching, For the ideal alignment direction (e.g., the texture U-axis). For example, θ distort A value >15° indicates significant stretching, requiring compensation in the next block. Mesh reconstruction resolution r mesh Extract the mesh topology of the completed region from the current 3D spatial digital model and count the number of patches per square meter; for example, the coarse-grained reconstruction threshold is r. mesh <500 faces / m 2 The high-precision reconstruction threshold is r. mesh >2000 faces / m 2 .
[0079] S5042: Calculate the illumination compensation parameters for the next mapped block based on the visual distortion angle, and perform visual consistency compensation for the brightness difference of the previous mapped block based on the illumination compensation parameters.
[0080] In this embodiment, to eliminate the inconsistency in brightness across blocks caused by changes in lighting conditions (such as shadows and highlights) during shooting, lighting compensation needs to be performed: lighting compensation parameter λ light Based on the visual distortion angle θ of the current block distort and ambient light intensity I env Calculate: λ light =f(θ) diatort The function, ΔI), takes the form of a lookup table or neural network regression, and outputs include: brightness offset Δb, contrast gain g, and color temperature correction. Visual consistency compensation applies gamma correction and histogram matching to the texture map of the previous mapped region: T' prev (x,y)=clamp(g×(T prev (x,y)+Δb), where T prev (x, y) represents the original texture pixel value of the previous block, and the pixel brightness (or RGB value) at coordinates (x, y) is in the range of [0, 255]; T' prev (x,y) represents the corrected texture pixel values; Δb is the brightness offset; g is the contrast gain factor, which controls the scaling factor of the contrast. If the current contrast is stronger (e.g., sunlight), then g>1.0, enhancing the contrast of the previous block; clamp(·) is the saturation cutoff function. If color difference exists, a white balance algorithm (e.g., the gray world assumption) is used for color unification.
[0081] S5043: Adjust the amount of texture mapping data and the multi-level detail hierarchy of the next mapping block according to the mesh reconstruction resolution and the pixel density of the corresponding texture atlas.
[0082] Specifically, the texture mapping amount T is calculated using formula (1): T=α×β×γ×S (1) Where α is the texture compression coefficient, ranging from 0.3 to 1.0; β is the LOD adjustment factor, corresponding to multiple levels of detail; β = 1.0 loads the highest resolution Mipmap level (1024×1024), β = 0.4 uses a low resolution level (256×256); γ is the lighting compensation weight, representing the proportion of metadata that needs to be additionally stored after compensation: γ = 1.0: fully saves compensation parameters, γ = 0.7: only saves brightness offset; S is the geometric area of the surface to be mapped, referring to the actual surface area (m²) of the current target region on the 3D model. 2 The system defaults to T_max = 8MB as the upper limit for a single block. If T > T_max is calculated according to formula (1), β or α will be automatically reduced to prioritize real-time performance. LOD level switching is based on the viewpoint distance d:
[0083] S5044: Based on the spatial anchor point positions of the constructed region in the 3D spatial digital model, obtain the edge alignment parameters of each spatial anchor point, and calculate the spatial edge change value obtained after visual distortion angle compensation.
[0084] In this embodiment, the spatial anchor point refers to a fixed reference point in three-dimensional space, typically located in prominent positions such as corners or device edges, used for spatial binding of virtual content. Each anchor point records its three-dimensional coordinates and normal vector. The edge alignment parameter δ... i This represents the alignment error between the current texture boundary and adjacent blocks or physical boundaries, expressed in pixels or millimeters. It is calculated by ICP registration or feature matching, taking the mean vertex offset at the intersection of two blocks. The spatial edge variation value Δe represents the change in edge alignment error after applying illumination compensation and geometric correction, calculated as Δe = ||δ after -δ bef ore||. If Δe < -0.5mm, the compensation is effective; if > 0.3mm, a new error is introduced.
[0085] S5045: Adjust the UV coordinate alignment parameters of the next mapping block based on the spatial edge change value, and optimize the overall texture mapping process by combining the adjusted UV coordinate alignment parameters and the amount of texture mapping data.
[0086] In this embodiment, the adjustment of the UV coordinate alignment parameter φ includes: if Δe>0.2mm, it indicates that the current alignment is not good and the UV unwrapping strategy of the next block needs to be adjusted: enable boundary-constrained parameterization; force strict alignment of the UV coordinates of the boundary edge; and adjust the Laplacian weight to reduce distortion.
[0087] Specifically, the overall optimization process includes: based on a preset optimization objective function: Where E visual Energy representing visual distortion, including illumination and color errors; E geo For edge misalignment and UV distortion; C reso Let T be the amount of texture data; and ω be the weight. i Configurable (e.g., live streaming scenarios emphasize ω3, while film and television rendering emphasizes ω1).
[0088] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0089] In one embodiment, a cross-platform virtual-real fusion scene construction system based on AI spatial computing is provided. This system corresponds to the cross-platform virtual-real fusion scene construction method based on AI spatial computing described above. The cross-platform virtual-real fusion scene construction system based on AI spatial computing includes a multimodal instruction parsing module, a feature fusion and modeling module, a data processing and scene loading module, a hybrid rendering computing module, and a multi-platform streaming module. Detailed descriptions of each functional module are as follows: The multimodal instruction parsing module is used to receive cross-modal conversion instruction sets and parse them into executable semantic operation sequences; the feature fusion and modeling module is used to perform multi-scale fusion of visual features corresponding to semantic operation sequences through generative adversarial networks to generate initial image sequences, and combine 3D reconstruction algorithms to perform implicit field encoding on target objects to output initial parameterized models. The data processing and scene loading module is used to connect to the multi-view video stream input interface of XR devices, perform spatiotemporal synchronization processing on the video stream, integrate physical sensor data, load digital scene asset packages in combination with preset spatial index structure, and establish a two-way data channel and mapping interaction rules between virtual scenes and physical sensors. The hybrid rendering computing module, deployed on a GPU cluster, is used to run a hybrid rendering pipeline that combines rasterization and ray tracing. It performs high-fidelity rendering on the initial parametric model and dynamically adjusts the material reflectivity, shadow intensity, and global illumination parameters based on the ambient lighting estimation results to generate an optimized parametric model with high-fidelity visual flow. The multi-platform streaming module is used to perform differentiated encoding processing on the optimized parameter model according to regional characteristics to obtain the target virtual-real scene fusion model; combined with the deep learning bitrate adaptive algorithm, bitrate resources are allocated to distribute the target virtual-real scene fusion model to the preset terminal.
[0090] Optionally, a cross-platform virtual-real fusion scene construction system based on AI spatial computing also includes: The environmental perception module is used to acquire environmental perception data of the target physical space; The digital modeling module is used to input environmental perception data into a pre-trained spatial semantic reconstruction model to generate a three-dimensional spatial digital model containing geometric structure and semantic labels. The pose calibration and error compensation module is used to calculate the pose alignment deviation between the current reconstructed area and the global spatial reference in real time during the construction of the 3D spatial digital model; to perform spatial correction processing on the pose alignment deviation, generate spatial correction regression parameters, and calculate the spatial deformation error under the current environmental conditions based on the spatial correction regression parameters. The point cloud and texture optimization unit, integrated into the pose calibration and error compensation module, is used to dynamically adjust the voxel filtering resolution and feature matching threshold during the point cloud fusion process based on spatial deformation errors, and simultaneously adjust the sampling frequency and UV unwrapping accuracy of texture mapping to optimize the point cloud fusion strategy and texture mapping rate.
[0091] For specific limitations regarding the cross-platform virtual-real fusion scene construction system based on AI spatial computing, please refer to the limitations of the cross-platform virtual-real fusion scene construction method based on AI spatial computing mentioned above, which will not be repeated here. Each module in the above-mentioned cross-platform virtual-real fusion scene construction system based on AI spatial computing can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in the processor of the computer device in hardware form or independent of the processor, or it can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0092] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements steps such as a method for constructing a cross-platform virtual-real fusion scene based on AI spatial computing.
[0093] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Furthermore, any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory.
[0094] In one embodiment, particularly according to embodiments of the present invention, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, embodiments of the present invention include a computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the cross-platform virtual-real scene construction method based on AI spatial computing. In such embodiments, when the computer program is executed by a central processing unit (CPU), it performs the various functions defined in the present invention.
[0095] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0096] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A cross-platform virtual-real fusion scene construction method based on AI spatial computing, characterized in that, include: Based on the received cross-modal conversion instruction set, multi-scale feature fusion is performed through a generative adversarial network to generate an initial image sequence; The target object is implicitly field-coded using a 3D reconstruction algorithm to obtain an initial parameterized model. Spatiotemporal alignment is performed on the multi-view video streams acquired by XR devices. Digital scene asset packages are loaded by combining physical sensor data and a preset spatial index structure to establish a two-way data channel and mapping interaction rules between virtual scenes and physical sensors. A hybrid rendering pipeline that integrates rasterization and ray tracing is run on a GPU cluster to render and adjust the lighting parameters of the initial parametric model, resulting in an optimized parametric model with high-fidelity visual flow. The optimized parameter model is processed by differentiated encoding according to regional characteristics to obtain the target virtual-real scene fusion model; the bitrate resource is allocated by combining deep learning bitrate adaptive algorithm to distribute the target virtual-real scene fusion model to the preset terminal.
2. The cross-platform virtual-real fusion scene construction method based on AI spatial computing according to claim 1, characterized in that, Before combining 3D reconstruction algorithms to perform implicit field encoding on the target object and obtain an initial parametric model, the method also includes: Acquire environmental perception data of the target physical space, including depth images, RGB images, inertial measurement data, and environmental sound field information; Based on the environmental perception data, a three-dimensional spatial digital model containing geometric structure and semantic labels is generated using a pre-trained spatial semantic reconstruction model. During the construction of the three-dimensional spatial digital model, the pose alignment deviation between the current reconstructed region and the global spatial reference is calculated; Spatial correction processing is performed on the pose alignment deviation to obtain spatial correction regression parameters; Based on the spatial correction regression parameters, the spatial deformation error under the current environmental conditions is calculated, and the point cloud fusion strategy and texture mapping rate are simultaneously optimized based on the spatial deformation error.
3. The cross-platform virtual-real fusion scene construction method based on AI spatial computing according to claim 2, characterized in that, The step of calculating the spatial deformation error under the current environmental conditions based on the spatial correction regression parameters, and simultaneously optimizing the point cloud fusion strategy and texture mapping rate based on the spatial deformation error, includes: Based on the spatial correction regression parameters, obtain the static environmental geometric stability coefficient corresponding to the current spatial region and the dynamic pose disturbance intensity corresponding to the pose alignment deviation; Based on the geometric stability coefficient and the dynamic pose disturbance intensity, calculate the spatial deformation error of the three-dimensional spatial digital model in the current environment; Adjust the voxel filtering resolution of the point cloud data according to the spatial deformation error, and control the number of point cloud registrations and feature extraction density within the same time window according to the voxel filtering resolution to form a point cloud fusion optimization strategy. Based on the spatial consistency enhancement results generated by the point cloud fusion optimization strategy, the sampling frequency and UV unwrapping accuracy in the texture mapping process are dynamically adjusted to optimize the texture mapping rate.
4. The cross-platform virtual-real fusion scene construction method based on AI spatial computing according to claim 3, characterized in that, The dynamic adjustment of the sampling frequency and UV unwrapping accuracy during the texture mapping process also includes: Obtain the visual distortion angle of the completed texture mapping block and the mesh reconstruction resolution of the three-dimensional spatial digital model; Calculate the illumination compensation parameters for the next mapping block based on the visual distortion angle, and perform visual consistency compensation for the brightness difference of the previous mapping block based on the illumination compensation parameters. Based on the mesh reconstruction resolution and the pixel density of the corresponding texture atlas, adjust the amount of texture mapping data and the multi-level detail hierarchy of the next mapping block; The texture mapping amount T is calculated using formula (1): T = α × β × γ × S (1) Where α is the texture compression coefficient, ranging from 0.3 to 1.0, β is the LOD adjustment factor, γ is the illumination compensation weight, and S is the geometric area of the surface to be mapped. Based on the spatial anchor point positions of the constructed region in the three-dimensional spatial digital model, the edge alignment parameters of each spatial anchor point are obtained, and the spatial edge change value of the visual distortion angle is calculated after visual consistency compensation. The UV coordinate alignment parameters of the next mapping block are adjusted based on the spatial edge change value, and the overall texture mapping process is optimized by combining the adjusted UV coordinate alignment parameters and the amount of texture mapping data.
5. The cross-platform virtual-real fusion scene construction method based on AI spatial computing according to claim 2, characterized in that, The spatial correction processing of the pose alignment deviation to obtain spatial correction regression parameters specifically includes: During the construction of the three-dimensional spatial digital model, the actual pose angle of the current frame point cloud data in the target physical space relative to the global coordinate system is obtained; Obtain the ideal spatial axis corresponding to the current construction position, and calculate the pose alignment deviation value between the actual pose angle and the ideal spatial axis. Based on the pose alignment deviation value, the construction trend of the three-dimensional spatial digital model is analyzed to obtain the spatial drift trend analysis results in the current construction process; Based on the spatial drift trend analysis results, the corresponding continuous drift intervals are identified, and the cumulative error within the continuous drift intervals is compensated in a segmented manner to obtain the spatial correction regression parameters of the three-dimensional spatial digital model under the current drift state.
6. The cross-platform virtual-real fusion scene construction method based on AI spatial computing according to claim 5, characterized in that, The step involves identifying corresponding continuous drift intervals based on the spatial drift trend analysis results, and performing piecewise reverse compensation on the accumulated errors within the continuous drift intervals to obtain the spatial correction regression parameters of the three-dimensional spatial digital model under the current drift state, including: Obtain the environmental texture feature density under the current drift state and the current reconstructed pose of the three-dimensional spatial digital model; Based on the environmental texture feature density and the current reconstruction pose, analyze the optimal spatial regression path that fits the pose deviation back to the global coordinate system reference range during the spatial reconstruction process. The optimal spatial regression path is segmented into multiple scales, and the pose compensation angle of the current drift interval is calculated based on the spatial alignment accuracy of the previous segment. Based on the pose compensation angle of the previous paragraph, the spatial interpolation step size and ICP registration threshold of the next paragraph are dynamically adjusted to obtain the multi-level spatial correction regression parameters of the three-dimensional spatial digital model in the current drift state.
7. A cross-platform virtual-real fusion scene construction system based on AI spatial computing, characterized in that, The system is used to execute the cross-platform virtual-real fusion scene construction method based on AI spatial computing as described in any one of claims 1 to 6, the system comprising: The multimodal instruction parsing module is used to receive cross-modal conversion instruction sets and parse them into executable semantic operation sequences; The feature fusion and modeling module is used to perform multi-scale fusion of visual features corresponding to the semantic operation sequence through a generative adversarial network to generate an initial image sequence, and to perform implicit field coding on the target object in combination with a 3D reconstruction algorithm to output an initial parameterized model. The data processing and scene loading module is used to connect to the multi-view video stream input interface of the XR device, perform spatiotemporal synchronization processing on the video stream, integrate physical sensor data, load digital scene asset packages in combination with a preset spatial index structure, and establish a two-way data channel and mapping interaction rules between the virtual scene and the physical sensor. The hybrid rendering computing module, deployed on a GPU cluster, is used to run a hybrid rendering pipeline that integrates rasterization and ray tracing. It performs high-fidelity rendering processing on the initial parameterized model and dynamically adjusts the material reflectivity, shadow intensity, and global illumination parameters based on the ambient lighting estimation results to generate an optimized parameter model with high-fidelity visual flow. The multi-platform streaming module is used to perform differentiated encoding processing on the optimized parameter model according to regional characteristics to obtain the target virtual-real scene fusion model; and to allocate bitrate resources by combining deep learning bitrate adaptive algorithm to distribute the target virtual-real scene fusion model to preset terminals.
8. The cross-platform virtual-real fusion scene construction system based on AI spatial computing according to claim 7, characterized in that, The system also includes: The environmental perception module is used to acquire environmental perception data of the target physical space; The digital modeling module is used to input the environmental perception data into a pre-trained spatial semantic reconstruction model to generate a three-dimensional spatial digital model containing geometric structure and semantic labels. The pose calibration and error compensation module is used to calculate the pose alignment deviation between the current reconstructed area and the global spatial reference in real time during the construction of the three-dimensional spatial digital model; to perform spatial correction processing on the pose alignment deviation, generate spatial correction regression parameters, and calculate the spatial deformation error under the current environmental conditions based on the spatial correction regression parameters. The point cloud and texture optimization unit, integrated into the pose calibration and error compensation module, is used to dynamically adjust the voxel filtering resolution and feature matching threshold during the point cloud fusion process according to the spatial deformation error, and simultaneously adjust the sampling frequency and UV unwrapping accuracy of texture mapping to optimize the point cloud fusion strategy and texture mapping rate.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the cross-platform virtual-real fusion scene construction method based on AI spatial computing as described in any one of claims 1 to 6.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the cross-platform virtual-real fusion scene construction method based on AI spatial computing as described in any one of claims 1 to 6.
Citation Information
Cited By
SLAM system optimization method and device based on layered anchor diagram structure and storage medium
CN121740063A
SLAM system optimization method, device, and storage medium based on hierarchical anchor point graph structure
CN121740063B
Aerospace equipment data-real fusion test cross-domain scene-task generation and evaluation method
CN121766417A
Cross-domain scenario for data-real fusion testing of aerospace equipment - mission generation and evaluation methods
CN121766417B
Cross-border area PM2.5 three-dimensional reconstruction method and device
CN122134950A