3D modeling algorithm based on large and small model collaboration

Through the collaborative architecture of large cloud models and lightweight edge models, combined with multi-source data fusion and dynamic task allocation, the contradiction between high precision and real-time performance is resolved, efficient reconstruction of complex scenes is achieved, and the robustness of the model and resource utilization efficiency are improved.

CN120635309APending Publication Date: 2025-09-12DONGGUAN NEW GENERATION ARTIFICIAL INTELLIGENCE IND TECH RES INST
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510715925.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing 3D modeling technology relies on a single computing architecture, making it difficult to balance high-precision modeling and real-time performance. There are deviations in multimodal data fusion, dynamic scene processing lacks a real-time repair mechanism, and cloud and end-side models lack adaptive collaboration, resulting in resource waste and performance bottlenecks.

Method used

By integrating large cloud models and lightweight edge models through heterogeneous computing collaborative architecture, and adopting multi-source data fusion, dynamic task allocation and cross-modal conflict resolution mechanisms, efficient and high-precision reconstruction of complex scenes can be achieved.

Benefits of technology

While achieving high-precision 3D modeling, it significantly improves the real-time reconstruction efficiency and model robustness in dynamic scenes, reduces the end-side computing power requirements, expands the application scope of the algorithm, and supports real-time processing of complex dynamic scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635309A_ABST
    Figure CN120635309A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D modeling algorithm based on large and small model collaboration, and relates to the technical field of computer vision and three-dimensional modeling, the algorithm comprises the following steps: preprocessing multi-source data such as an RGB image sequence, and generating labels and task priorities through semantic analysis; the cloud-side large model splits tasks, and distributes the tasks to the end-side lightweight model or the cloud-side large model according to calculation complexity and real-time performance; task execution is dynamically adjusted according to feedback in collaborative planning; the heterogeneous models cooperatively complete geometric reconstruction, texture optimization and the like; and cross-modal conflict resolution is combined with multi-source data and dynamic correction is carried out. According to the algorithm, efficient collaboration, balance performance and real-time performance of computing power resources are achieved, and the contradiction between high precision and low delay of a traditional method is solved through cross-modal data fusion and end-cloud collaboration continuous learning. Meanwhile, the large and small model cooperation normal form improves the resource efficiency and the system reliability, can be applied to the fields of industrial detection, AR / VR and the like, reduces the end side hardware threshold, and supports the real-time processing of complex dynamic scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and three-dimensional modeling, and specifically to a 3D modeling algorithm based on collaboration between large and small models. Background Art

[0002] Existing 3D modeling technology has significant defects due to its reliance on a single computing architecture (cloud / end): high-precision modeling is difficult to balance due to the conflict between computing power requirements and real-time performance, pure cloud processing is limited by network latency, and the end side cannot support complex algorithms; multimodal data (RGB, depth map, etc.) fusion deviations lead to distortion of reflective / low-texture areas, dynamic scene processing lacks a real-time repair mechanism, causing model holes, and there is a lack of adaptive collaboration between large cloud models and lightweight end-side models, resulting in resource waste or performance bottlenecks. In response to these problems, the present invention proposes an innovative solution - integrating cloud and edge computing power through a heterogeneous computing collaborative architecture, introducing dynamic programming tokens to achieve intelligent scheduling of subtasks, and combining a multimodal confidence fusion mechanism to optimize data conflict resolution. Ultimately, while ensuring global modeling accuracy, it significantly improves real-time reconstruction efficiency and model robustness in dynamic scenes. Summary of the Invention

[0003] To address the above problems, the present invention relates to a 3D modeling algorithm and system based on the collaboration of large cloud models and lightweight edge models. Through multi-source data fusion, dynamic task allocation, heterogeneous model collaboration and cross-modal conflict resolution mechanism, efficient and high-precision reconstruction of complex scenes can be achieved.

[0004] To achieve the above objectives, the present invention is implemented through the following technical solutions: a 3D modeling algorithm based on collaboration between large and small models, comprising the following steps:

[0005] Step S1: Multi-source input data preprocessing

[0006] Input data preprocessing:

[0007] The cloud-side large model receives multi-source input data, including RGB image sequences, depth maps (such as ToF or LiDAR data), point cloud data (such as structured light scanning results), and possible IMU sensor data.

[0008] Align and calibrate the input data, for example by calibrating the camera's internal and external parameters, to ensure temporal synchronization and spatial consistency of multi-view data.

[0009] Based on the Transformer architecture, the input data is semantically parsed to identify the object categories, material properties, and dynamic object areas in the scene, and generate semantic labels and task priority lists.

[0010] Step S2: Task decomposition and dynamic resource allocation:

[0011] 1. The large model analyzes task requirements based on the Transformer architecture and decomposes the 3D reconstruction process into key subtasks:

[0012] ①, Geometric reconstruction (sparse / dense point cloud generation, surface reconstruction);

[0013] ② Texture mapping (texture mapping optimization based on multi-view images);

[0014] ③ Dynamic completion (moving object removal or interpolation repair);

[0015] ④Detail enhancement (high-resolution geometry refinement or normal optimization);

[0016] Dynamically allocate subtasks to the cloud or device side based on their computational complexity (e.g., dense reconstruction requires more computing power) and real-time requirements (e.g., dynamic scenes require low latency).

[0017] 2. Resource Allocation

[0018] The end device runs a lightweight model (such as Mobile-UNet or Tiny-YOLO) to handle tasks with high real-time requirements (such as feature point extraction or preliminary point cloud generation).

[0019] Large cloud-side models (such as PointNet++ or NeRF-based architectures) handle computationally intensive tasks (such as global optimization or high-precision texture synthesis);

[0020] Step S3: Collaborative planning and knowledge transfer

[0021] 1. Global plan generation:

[0022] The cloud side constructs a task dependency graph (DAG) to clarify the temporal relationship between subtasks (for example, point cloud registration must be completed before surface reconstruction can be performed). A lightweight planning token is generated, encoding the following information:

[0023] ①, Task type (geometry / texture / dynamic);

[0024] ② Input / output data specifications (such as point cloud resolution requirements);

[0025] ③, QoS constraints (maximum delay, energy budget);

[0026] 2. Client-side execution and feedback:

[0027] After the client parses the Planning Token, it calls the corresponding lightweight model to execute the task.

[0028] For example:

[0029] ①. Use SIFT or SuperPoint to extract feature points;

[0030] ②. Perform local point cloud registration based on ICP or Bundle Adjustment;

[0031] ③. Real-time feedback of execution status (such as feature matching success rate and local reconstruction error) to the cloud side.

[0032] 3. Dynamic adjustment:

[0033] If the client-side feedback confidence level is lower than the threshold (for example, the point cloud alignment error is greater than 5%), the cloud-side triggers one of the following policies:

[0034] ① Computation offloading: Migrate tasks to the cloud for execution;

[0035] ② Parameter adjustment: Send a new LoRA adapter to fine-tune the end-side model.

[0036] Step S4: Collaborative execution of heterogeneous models

[0037] 1. Geometric reconstruction collaboration:

[0038] ①. Cloud side: Run large Transformer-based models (such as Point-BERT) to complete global point cloud denoising and topology optimization.

[0039] ②. On the device side: Use a lightweight DGCNN to process the local point cloud and upload the geometric features to the cloud side through a cross-modal residual module.

[0040] 2. Texture optimization collaboration:

[0041] ①. Cloud side: Generate high-resolution texture maps (e.g., through StyleGAN-ADA) and distill them into low-dimensional feature vectors.

[0042] ②. On the device side: Receive texture features through the multimodal knowledge injection interface and perform real-time texture mapping in combination with local images.

[0043] 3. Dynamic scene processing:

[0044] The device detects moving objects (such as YOLO-NAS), and the cloud side completes the 3D structure of the dynamic area through optical flow estimation and spatiotemporal consistency analysis.

[0045] Step S5: Cross-modal conflict resolution and semantic enhancement

[0046] 1. Multimodal fusion:

[0047] The system integrates device-side RGB-D data, IMU pose estimation, and cloud-side prior knowledge (such as the object's CAD model library), and uses an attention mechanism to weight the contributions of different modalities. For example, when reconstructing metal parts, the weight of depth data is increased (because RGB in reflective areas is unreliable).

[0048] 2. Conflict detection and arbitration:

[0049] ① Geometric conflict: When the IoU between the reconstruction results on the client and cloud sides is less than 0.7, RANSAC-based verification is started.

[0050] ② Texture conflict: The CLIP model is used to calculate the image-text alignment score (such as "metal surface should be glossy") and correct unreasonable textures.

[0051] 3. Continuous semantic learning:

[0052] The device side collects special scene data (transparent objects, low-texture areas) and uploads it to the cloud side; the cloud side updates the large model parameters through federated learning and sends incremental semantic knowledge (adapter parameters, material reflection model) to the device side.

[0053] Furthermore, in step S1, semantic analysis specifically includes:

[0054] Based on the Transformer architecture, it parses the input data, identifies object categories, material properties, and dynamic areas, and outputs semantic labels and task priority lists.

[0055] Furthermore, in step S2, resource allocation specifically includes using Mobile-UNet or Tiny-YOLO on the client side to process real-time tasks; and using PointNet++ or NeRF-based architecture on the cloud side to process computationally intensive tasks.

[0056] Furthermore, in step S3, if the geometric alignment error fed back by the client side exceeds 5%, the cloud side triggers one of the following operations:

[0057] ① Migrate tasks to the cloud for execution;

[0058] ② Send a LoRA adapter to fine-tune the device-side model.

[0059] Furthermore, in step S4, the geometric reconstruction collaboration specifically includes:

[0060] The cloud side uses Point-BERT to complete global point cloud denoising, the client side uses DGCNN to process local point clouds, and uploads geometric features to the cloud side through the cross-modal residual module.

[0061] Furthermore, in step S4, the texture optimization collaboration specifically includes:

[0062] The cloud side generates a 512×512 resolution texture map and compresses it into a low-dimensional vector through VAE. The client side adapts to the local lighting based on GAN and integrates AO parameters for rendering.

[0063] Furthermore, in step S5, RANSAC is used to verify the geometric conflicts, and CLIP model is used to correct the texture conflicts. The fusion formula is:

[0064] Score = w1·Pgeo+w2·Simtexture+w3·IMUstability, where the weights w1, w2, and w3 are dynamically adjusted according to the scenario.

[0065] When the Score is less than the threshold, the manual annotation intervention process is triggered.

[0066] The present invention provides a 3D modeling algorithm based on the collaboration of large and small models. Compared with the existing technology, it has the following advantages:

[0067] 1. Efficient coordination of computing resources, balancing performance and real-time performance

[0068] Through dynamic task decomposition and resource allocation, real-time-critical tasks (such as feature extraction and local point cloud generation) are assigned to lightweight models on the device (e.g., Mobile-UNet and Tiny-YOLO), while computationally intensive tasks (e.g., global optimization and high-resolution texture synthesis) are handled by large cloud-side models (e.g., NeRF and Point-BERT). This improvement significantly reduces the computing power required on the device while ensuring high-precision reconstruction through cloud-side collaboration, meeting the low-latency requirements of dynamic scenarios, such as enabling a smooth experience in real-time AR / VR interactions.

[0069] 2. Cross-modal data fusion and intelligent conflict resolution

[0070] The attention mechanism dynamically weights multimodal contributions (e.g., prioritizing depth data in reflective metal areas), verifies geometric conflicts (triggered when Intersection over Union (IoU) < 0.7) through RANSAC, and corrects texture conflicts (e.g., semantic alignment scores) through the CLIP model. This improvement reduces the impact of sensor limitations (e.g., RGB reflection interference and IMU drift) on the results, ensuring the physical plausibility and consistency of the reconstruction results.

[0071] 3. Continuous semantic learning and long-tail scenario adaptation through end-cloud collaboration

[0072] By collecting special scenario data on the edge and uploading it to the cloud, federated learning is used to update large model parameters and distribute incremental semantic knowledge (such as adapter parameters and material reflection models). This mechanism enables continuous optimization of the model at the edge, reduces manual labeling costs, improves modeling capabilities for long-tail scenarios (such as glassware and foggy environments), and expands the application range of the algorithm.

[0073] 4. Collaborative paradigm for large model disassembly tasks and small model execution

[0074] The present invention uses a large cloud-side model (based on the Transformer architecture) to intelligently decompose global tasks into subtasks (such as geometric reconstruction and texture mapping), and distributes them to small end-side models (such as DGCNN and SIFT) or large cloud-side models according to computational complexity and real-time requirements. Feature interaction is achieved through cross-modal residual modules and knowledge distillation (such as texture feature compression). This collaborative mechanism maximizes the efficiency of end-cloud resources, ensuring real-time performance on the end side (such as AR interaction) and high precision on the cloud side (such as industrial part reconstruction). At the same time, the dynamic fault-tolerance mechanism (such as cloud-side takeover when errors exceed the limit) further improves system reliability.

[0075] 5. Comprehensive advantages and application value

[0076] This invention addresses the conflict between high precision and low latency in traditional approaches through a collaborative architecture that combines a large cloud-based model with a small terminal model, combined with dynamic task allocation, cross-modal fusion, continuous learning, and flexible adjustment mechanisms. This significantly lowers the hardware requirements on the terminal side and supports real-time processing of complex dynamic scenes. Its technological breakthroughs cover areas such as industrial inspection, autonomous driving, and AR / VR, providing efficient and reliable solutions for high-precision part reconstruction, dynamic obstacle modeling, and real-time interactive rendering. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 It is a schematic diagram of the process of the present invention;

[0078] Figure 2 This is the system architecture diagram of the present invention DETAILED DESCRIPTION

[0079] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0080] See also Figure 1-2 The present invention provides a technical solution: a 3D modeling algorithm based on collaboration between large and small models

[0081] 1. Multi-source input data preprocessing (step S1):

[0082] (1) Data collection

[0083] Use an industrial-grade RGB-D camera (such as Intel RealSense D455) and LiDAR (such as Velodyne VLP-16) to obtain high-resolution RGB images and depth maps of the mold surface.

[0084] High-precision point cloud data is acquired through a structured light scanner (such as Microsoft Azure Kinect), covering the complex curved surfaces and tiny structures of the mold.

[0085] Synchronously collect IMU sensor data (such as inertial measurement unit) to record device posture changes and ensure spatial consistency of multi-view data.

[0086] (2) Data alignment and calibration

[0087] The RGB image and depth map are synchronized in time and space based on the calibrated camera internal and external parameters (through the checkerboard calibration method) to eliminate misalignment caused by device latency.

[0088] The point cloud data is registered using ICP (Iterative Closest Point) to generate a global point cloud set in a unified coordinate system.

[0089] (3) Semantic analysis and priority allocation:

[0090] The cloud-side large model (Transformer architecture) parses the input data and identifies key mold components (such as cavity, core, cooling system) and material properties (metal, plastic, transparent areas).

[0091] Dynamic area detection: Locate mold moving parts (such as sliders and ejectors) based on IMU data and generate a priority list of dynamic subtasks.

[0092] Output semantic labels (such as “high reflective area - metal cavity”, “low texture area - cooling channel”) for subsequent processing reference.

[0093] 2. Task decomposition and resource allocation (S2)

[0094] (1) Subtask division

[0095] Geometric reconstruction: through sparse point cloud generation (using Open3D's RANSAC plane detection) and dense point cloud optimization (multi-view reconstruction based on NeRF).

[0096] Texture Mapping: Texture map optimization based on multi-view RGB images (using COLMAP for SfM reconstruction).

[0097] Dynamic completion: Remove occluded areas of moving objects (such as a transport robot) and repair missing structures through spatiotemporal interpolation.

[0098] Detail enhancement: Utilize high-resolution point clouds (0.1mm accuracy) for normal optimization and surface smoothing (based on Poisson reconstruction).

[0099] (2) Resource allocation strategy

[0100] Devices on the client side: Deploy lightweight models (such as Mobile-UNet) to handle real-time tasks.

[0101] ① Feature point extraction (SuperPoint)

[0102] ②Preliminary point cloud generation (local registration based on ICP)

[0103] Cloud-side large model: handles computationally intensive tasks:

[0104] ①Global point cloud optimization (PointNet+++Transformer)

[0105] ② High-precision texture synthesis (StyleGAN-ADA generates 512×512 texture maps)

[0106] Resource scheduling logic:

[0107] The client prioritizes real-time tasks (such as feature point extraction and preliminary point cloud registration) to reduce cloud communication latency.

[0108] The cloud side undertakes global optimization tasks (such as dense point cloud reconstruction and high-precision texture synthesis) and takes advantage of its computing power.

[0109] Dynamic adjustment mechanism: If the end side detects that the point cloud registration error in the mold cavity area is greater than 5% (such as caused by vibration), it triggers the task to be offloaded to the cloud side, or fine-tunes the end-side model parameters through the LoRA adapter.

[0110] 3. Collaborative Planning and Knowledge Transfer (S3)

[0111] (1) Global plan generation:

[0112] The cloud side builds a task dependency graph (DAG) to clarify the execution order:

[0113] Feature point extraction (end-side) → 2. Local point cloud registration (end-side) → 3. Global denoising and topology optimization (cloud-side) → 4. Texture mapping (cloud-side → end-side) → 5. Dynamic area restoration (cloud-side collaborative end-side).

[0114] Generate a lightweight Planning Token (JSON format), including the task type, data specifications (such as point cloud resolution must be ≥ 0.1mm), and QoS constraints (such as total delay ≤ 2s).

[0115] (2) End-side execution and feedback:

[0116] After the client parses the Planning Token, it calls the corresponding model:

[0117] ①Use SuperPoint to extract key feature points on the mold surface (such as parting line and gate contour);

[0118] ②Based on the ICP algorithm, the local point cloud is quickly registered and the local reconstruction error is calculated.

[0119] Real-time feedback to the cloud:

[0120] ① If the feature matching success rate is less than 85% (e.g., the transparent plastic area features are missing), the cloud will be triggered to send the updated material reflection model of federated learning.

[0121] ② If the IMU detects that the mold posture is unstable, the cloud adjusts the geometric weight (for example, increasing the depth data weight to 0.7).

[0122] 4. Heterogeneous Model Collaborative Execution (S4)

[0123] (1) Geometric reconstruction collaboration

[0124] ① Cloud side: Run Point-BERT to perform global point cloud denoising to eliminate abnormal points caused by mold vibration.

[0125] ② On the device side: Use DGCNN to process local point clouds (such as the pinhole area) and upload local geometric features to the cloud side through the cross-modal residual module.

[0126] (2) Texture optimization collaboration

[0127] ① Cloud side: Generate high-resolution texture maps (StyleGAN-ADA) and compress them into low-dimensional feature vectors (e.g., 128 dimensions) through VAE.

[0128] ② On the device side: Receives texture feature vectors and combines them with local images and lighting parameters (AO maps) for real-time texture mapping.

[0129] (3) Dynamic scene processing

[0130] YOLO-NAS is used on the client side to detect moving parts of the mold (such as the mold opening and closing mechanism), and the cloud side completes the 3D structure of the dynamic area through optical flow estimation and spatiotemporal consistency analysis.

[0131] 5. Cross-modal conflict resolution and semantic enhancement (S5)

[0132] (1) Multimodal fusion

[0133] Fusion of end-side RGB-D data, IMU pose estimation, and cloud-side prior knowledge (such as mold CAD model library).

[0134] Increase the depth data weight (w1=0.7w1=0.7) in reflective areas (such as the polished surface of the mold), and increase the RGB weight (w2=0.8w2=0.8) in low-texture areas (such as cooling water channels).

[0135] (2) Conflict Detection and Arbitration

[0136] Geometric conflict: If the IoU between the client-side and cloud-side reconstruction results is less than 0.7, RANSAC verification is initiated to remove abnormal point clouds.

[0137] Texture conflict: The CLIP model is used to calculate the alignment score between the "metal surface highlight" and the actual texture, and to correct unreasonable textures.

[0138] (3) Continuous semantic learning

[0139] The end side collects special scenario data (such as transparent coolant channels) and uploads it to the cloud side.

[0140] The cloud side updates the large model parameters through federated learning and sends incremental semantic knowledge (such as the material reflection model adapter) to the end side.

[0141] Implementation Effect

[0142] (1) Accuracy: Achieve 0.1mm-level mold geometry reconstruction, with texture mapping resolution reaching 512×512.

[0143] (2) Efficiency: Real-time processing on the end (<200ms / frame), global optimization on the cloud (<10s / task).

[0144] (3) Robustness: Dynamically adjust through the LoRA adapter to adapt to complex scenes such as reflections and low textures.

[0145] The above-described embodiments merely represent several implementation methods of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A 3D modeling algorithm based on collaboration between large and small models, characterized in that: include: Multi-source input data preprocessing, including receiving multi-source input data, aligning and calibrating the input data, and semantically parsing the input data based on the Transformer architecture to identify object categories, material properties, and dynamic object areas in the scene, and generate semantic labels and task priority lists; Task decomposition and dynamic resource allocation, including analyzing task requirements based on the Transformer architecture and breaking down the 3D reconstruction process into subtasks such as geometry reconstruction, texture mapping, dynamic completion, and detail enhancement. Subtasks are then dynamically allocated to the cloud or device based on their computational complexity and real-time requirements. Devices on the device side run lightweight models to handle tasks with high real-time requirements, while large models on the cloud side handle computationally intensive tasks. Collaborative planning and knowledge transfer: The cloud side builds a task dependency graph and generates a lightweight PlanningToken that encodes information such as the task type, input / output data specifications, and QoS constraints. The client side parses the PlanningToken, calls the corresponding lightweight model to execute the task, and provides real-time feedback on the execution status to the cloud side. If the client-side feedback confidence level falls below a threshold, the cloud side triggers computation offloading or adjusts parameters. Collaborative execution of heterogeneous models, including collaborative geometry reconstruction, collaborative texture optimization, and dynamic scene processing; Cross-modal conflict resolution and semantic enhancement, including multimodal fusion, conflict detection and arbitration, and continuous semantic learning.

2. A 3D modeling algorithm based on large and small model collaboration according to claim 1, characterized in that: The multi-source input data includes RGB image sequences, depth maps, point cloud data, and IMU sensor data.

3. The 3D modeling algorithm based on large and small model collaboration according to claim 1, characterized in that: The semantic parsing is specifically to parse the input data based on the Transformer architecture, identify object categories, material properties and dynamic areas, and output semantic labels and task priority lists.

4. The 3D modeling algorithm based on large and small model collaboration according to claim 1, characterized in that: The lightweight model includes Mobile-UNet or Tiny-YOLO, and the cloud-side large model includes PointNet++ or NeRF-based architecture.

5. The 3D modeling algorithm based on large and small model collaboration according to claim 1, characterized in that: If the geometric alignment error reported by the device side exceeds 5%, the cloud side triggers one of the following operations: migrating the task to the cloud side for execution; sending a LoRA adapter to fine-tune the device side model.

6. The 3D modeling algorithm based on large and small model collaboration according to claim 1, characterized in that: The geometric reconstruction collaboration specifically involves completing global point cloud denoising through Point-BERT on the cloud side, processing local point clouds through DGCNN on the client side, and uploading geometric features to the cloud side through a cross-modal residual module.

7. The 3D modeling algorithm based on large and small model collaboration according to claim 1, characterized in that: The texture optimization collaboration specifically generates a 512×512 resolution texture map on the cloud side and compresses it into a low-dimensional vector through VAE. The client side adapts the local lighting based on GAN and fuses AO parameters for rendering.

8. The 3D modeling algorithm based on large and small model collaboration according to claim 1, characterized in that: In the conflict detection and arbitration, geometric conflicts are verified by RANSAC, and texture conflicts are corrected by CLIP model. The fusion formula is: Score=w1・Pgeo+w2・Simtexture+w3・IMUstability The weights w1, w2, and w3 are dynamically adjusted according to the scenario. When the score is less than the threshold, the manual labeling intervention process is triggered.

9. The 3D modeling algorithm based on large and small model collaboration according to claim 1, characterized in that: The multimodal fusion specifically involves fusing the device-side RGB-D data, IMU pose estimation, and cloud-side prior knowledge, and weighting the contributions of different modalities through an attention mechanism. Continuous semantic learning specifically involves collecting special scenario data on the device side and uploading it to the cloud side. The cloud side updates the large model parameters through federated learning and sends incremental semantic knowledge to the device side.

10. The 3D modeling algorithm based on large and small model collaboration according to claim 1, characterized in that: Dynamic scene processing involves detecting moving objects on the device side and completing the 3D structure of the dynamic area on the cloud side through optical flow estimation and spatiotemporal consistency analysis. The geometric reconstruction includes sparse / dense point cloud generation and surface reconstruction. The texture mapping includes texture map optimization based on multi-view images. The dynamic completion includes moving object removal or interpolation repair. The detail enhancement includes high-resolution geometric refinement or normal optimization.

Citation Information

Cited By

  • End cloud multi-mode sensing method and system driven by external light route, and storage medium

    CN122247921A