A terminal-side lightweight semantic segmentation three-dimensional positioning and tracking method and system

CN122714501APending Publication Date: 2026-09-08BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610568346.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-27
Publication Date
2026-09-08

AI Technical Summary

Technical Problem

[0005]本发明旨在解决现有基于特征点的三维定位追踪技术在移动端环境中存在的依赖离线地图、计算量高、网络延迟大以及在光照变化或遮挡情况下稳定性不足等问题

Benefits of technology

离线采集包含目标识别物的视频,并通过切帧、自动标注与人工校正规则生成图像—掩码数据集;基于所述数据集对轻量级语义分割网络进行微调训练,并将训练完成的模型转换为终端侧可运行的格式。终端在三维渲染环境中构建虚拟坐标系与投影映射规则,加载所述模型并对实时摄像头图像执行语义分割推理,获取目标识别物的掩码及其统计量。系统结合掩码信息、平面法向量、深度估计及IMU传感器数据,通过加权投影与深度还原计算目标识别物在相机坐标系下的三维质心,并据此更新虚拟模型的位姿,实现对目标识别物的三维定位与追踪。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122714501A_ABST
    Figure CN122714501A_ABST
Patent Text Reader

Abstract

The application belongs to the field of computer vision and augmented reality, and discloses a terminal-side lightweight semantic segmentation three-dimensional positioning and tracking method and system. The method comprises offline and online processing stages: offline collection of target video, construction of image-mask data set, fine-tuning training of a lightweight semantic segmentation network and conversion into a terminal-side compatible inference model; online construction of a three-dimensional scene on the terminal side and agreement of coordinate mapping rules, loading of the model to perform semantic segmentation on the real-time image of the camera to obtain a target mask, fusion of the mask, IMU attitude and depth estimation data, calculation of the three-dimensional centroid of the target object through weighted projection and depth restoration, estimation of the pose after filtering optimization, completion of virtual model superposition rendering and interaction. The system matches offline and online processing modules to respectively realize model training and terminal-side positioning and tracking. The application is free from dependence on an offline feature point map, can be independently operated on the terminal side, has high positioning accuracy, stability and real-time performance, and is flexible in deployment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and augmented reality (AR), specifically to a lightweight semantic segmentation 3D localization and tracking method and system for terminal side. Background Technology

[0002] With the rapid development of 3D registration, positioning, and tracking technologies, these capabilities have been widely applied in scenarios such as cultural tourism, education, industrial inspection, and commercial displays. Taking digital cultural tourism as an example, through spatial understanding of the real environment, mobile terminals can overlay 3D reconstruction models, historical figure animations, or interactive explanations onto real-world scenic scenes, thereby creating an immersive experience that blends the virtual and real worlds. These applications require high levels of spatial positioning accuracy, stability, and real-time performance, forming the foundation for augmented reality content presentation.

[0003] Existing positioning and tracking systems mostly employ Spatial Mapping and Localization (SLAM) technology based on feature points. This type of technology typically includes two stages: offline mapping and online tracking. In the offline stage, scene images are acquired and feature point extraction, such as SIFT, ORB, and AKAZE matching, is performed to generate an offline map for localization. In the online stage, the terminal uploads keyframes of the current scene to the server, which then performs feature matching and pose calculation in the offline map and returns the pose to the terminal to achieve the overlay rendering of the 3D model in the real scene.

[0004] While the aforementioned methods can theoretically provide high positioning accuracy, they have several limitations: First, offline mapping is costly and heavily reliant on scene textures; second, online matching involves large computational loads and is sensitive to network bandwidth and communication latency, making it difficult to meet the real-time experience requirements of mobile devices; third, feature points are significantly affected by factors such as lighting changes, dynamic objects, and partial occlusion, which may lead to unstable positioning or even tracking failure. Especially in scenarios where the device operates independently, how to achieve lightweight, efficient, and robust 3D positioning and tracking without the need for feature point maps has become a pressing technical problem to be solved. Summary of the Invention

[0005] This invention aims to address the problems of existing feature-point-based 3D localization and tracking technologies in mobile environments, such as reliance on offline maps, high computational cost, large network latency, and insufficient stability under varying lighting conditions or occlusion. Traditional methods estimate pose through feature point extraction and matching, which is sensitive to scene texture, prone to localization jitter, and the online matching process is time-consuming, making it difficult to meet the requirements of lightweight, efficient, and robust augmented reality applications on the terminal side. Therefore, there is an urgent need for a new method that can independently complete 3D localization on the terminal without the need for feature point maps.

[0006] To achieve the above objectives, the present invention provides the following technical solution: A lightweight semantic segmentation 3D localization and tracking method on the terminal side, the specific steps of which are as follows: S1. Construct a server-terminal architecture. The terminal includes various sensors and image acquisition modules. The terminal acquires and uploads raw video containing the target object and its surrounding scene. S2. Perform frame segmentation and normalization on the original video to construct an image-mask dataset of the target object; S3. Input the image-mask dataset into the lightweight semantic segmentation network for training, and convert the trained model into a mobile-compatible inference model. S4. Construct a virtual 3D scene on the terminal, establish a coordinate system, and establish mapping rules between the camera image plane and the virtual 3D space; S5. The terminal downloads and loads the inference model from the server, runs it locally through the semantic segmentation inference framework, and completes the model preheating and inference backend configuration. S6. The terminal collects real-time image frames from the camera frame by frame, and after preprocessing, it is input into the inference model to perform semantic segmentation, and obtains the mask of the target object and the mask statistics. S7. The terminal fuses the mask and mask statistics with the output data of multiple sensors to establish a planar coordinate model of the camera image. It then calculates the three-dimensional centroid coordinates of the target object in the camera coordinate system through weighted projection, depth adaptive compensation, and filtering smoothing. Specifically: The three-dimensional centroid projection coordinates are calculated using weighted projection, and the formula is as follows: Where n is the normal vector of the tangent plane, pi=(x i ,y i (z0) represents the planar coordinates of the foreground pixel, z0 represents the depth of the camera texture plane, and N represents the total number of foreground pixels in the mask; Using data from multiple sensors, the centroid projection coordinates are restored from two-dimensional planar coordinates to three-dimensional centroid coordinates in the camera coordinate system. When depth is missing, adaptive compensation is performed according to the mask area ratio. The compensation formula is as follows: in The most recently recorded valid depth, and These are the mask areas of the current frame and the most recent frame, respectively; Then, the three-dimensional centroid coordinates are reconstructed using the formula, and a stable three-dimensional target position is obtained by combining it with sliding window weighted filtering. The formula is as follows: in The distance from the camera texture plane to the camera. Projected coordinates The coordinates in the camera coordinate system, where R is the camera rotation matrix calculated from the sensor output data; S8. Estimate the target's position and pose based on the three-dimensional centroid coordinates and data from multiple sensors, and complete the overlay rendering and interaction of the virtual three-dimensional model in the real scene.

[0007] Specifically, step S2 includes the following sub-steps: S21. Divide the original video uploaded in step S1 into independent image frames according to the preset sampling rate and record the timestamps. Standardize the size, format and color space of the image frames and organize and store them according to the directory structure and index files to form a data source library. S22. The SAM model is called to generate candidate masks for the image frames in step S21. Candidate masks are selected by area, edge sharpness, and confidence score. Low-confidence image frames are manually reviewed and corrected to obtain image-mask pairs. S23. Perform data augmentation operations on the image-mask pair, including random horizontal or vertical flipping, random small-angle rotation, scaling and cropping, brightness and contrast perturbation, local occlusion synthesis and background replacement. The augmented image-mask pair is divided into training set, validation set and test set according to a preset ratio and an index file is generated. Step S3 includes the following sub-steps: S31. To meet the real-time inference requirements of the terminal, a lightweight semantic segmentation network is selected as the base network, and an edge enhancement module is introduced at its output or intermediate feature layer to clarify the network input and output constraints and preprocessing specifications. S32. Using the dataset constructed in step S2, the improved network is fine-tuned through transfer learning. The loss function and optimization strategy are configured, and retraining is performed for edge targets and small target false detection types to obtain the trained model. S33. Convert the trained model into a terminal-compatible model, with preprocessing instructions and verification of the consistency between the terminal output and the training output.

[0008] Specifically, step S4 includes the following sub-steps: S41. Construct a 3D scene on the terminal that includes virtual camera nodes, a 2D texture plane of the camera image, and a root node for mounting the virtual model. Map the camera image to a dynamic video texture of the 2D texture plane to achieve coordinate system alignment. S42. Fix the virtual camera projection parameters and establish a mapping rule from pixel coordinates to world coordinates or camera coordinates. The mapping rule includes the construction method of the virtual coordinate system, the planar position and orientation of the camera image in the virtual coordinate system, and subsequent centroid and pose calculations shall follow this convention. S5 includes the following steps: S51. The terminal requests the inference model exported in step S33 from the server via the network, loads it, completes the inference resource binding, and updates the interface status indicating that the model has been loaded. S52. The terminal selects the inference backend based on the device performance and performs parameter optimization and memory pre-allocation for the inference framework.

[0009] Specifically, step S6 includes the following sub-steps: S61. Acquire camera image frames at a preset frame rate and record timestamps to achieve time alignment with data from multiple sensors; S62. After performing resolution scaling, tensor quantization, and normalization preprocessing on the image frame, input it into the model for inference, output a pixel-level probability map and binarize it to generate a mask map, calculate the mask statistics and visualization bitmap, and the mask statistics include the number of foreground pixels, centroid candidate coordinates and the vertical range of the mask. S63. Perform connected component analysis, noise filtering, and morphological processing on the mask image obtained in step S62, calculate the required statistics, including the target mask area, the two-dimensional centroid of the target pixel set, and the height and width range of the target pixels in the image, and provide the above statistics to the subsequent three-dimensional centroid and depth estimation algorithms.

[0010] Specifically, the basic lightweight semantic segmentation network of S31 is the MobileSeg network of the PaddlePaddle platform or its equivalent lightweight semantic segmentation network.

[0011] Specifically, the loss function of S32 adopts a weighted combination of pixel-level cross-entropy and region consistency loss, and the training process adopts a quantization strategy from FP32 to FP16 to reduce model size and improve inference efficiency.

[0012] The system also includes a lightweight semantic segmentation 3D localization and tracking system on the terminal side, comprising an offline processing module and an online processing module. The offline processing module is deployed on a server or offline processing node and includes a video acquisition unit, a dataset construction unit, a network training and format conversion unit, and a model storage and distribution unit. The video acquisition unit is used to acquire and upload the original target video. The dataset construction unit is used to construct an image-mask dataset. The network training and format conversion unit is used to train and convert the lightweight semantic segmentation network with an edge enhancement module. The model storage and distribution unit is used to store and distribute the inference model to the terminal. The online processing module is deployed on the terminal and includes a 3D scene construction unit, a model loading unit, an image acquisition and inference unit, a multi-sensor fusion unit, and a pose rendering and interaction unit. The 3D scene construction unit is used to construct a virtual 3D scene and define coordinate mapping rules. The model loading unit is used to download and load the inference model and configure the inference framework. The image acquisition and inference unit is used to acquire real-time images and output masks and statistics. The multi-sensor fusion unit is used to fuse IMU and depth estimation data to calculate the 3D centroid. The pose rendering and interaction unit is used for pose estimation, virtual model rendering, and user interaction.

[0013] Specifically, the multi-sensor fusion unit incorporates a weighted projection calculation module, a depth adaptive compensation module, and a sliding window filtering module, which are used to perform the three-dimensional centroid calculation and smoothing optimization as described in claim 5.

[0014] Specifically, the pose rendering interaction unit is based on the three.js, LayaAir, and WebGL rendering framework to realize virtual model overlay and supports rotation and scaling interaction; the system is suitable for augmented reality scenarios such as cultural tourism, education, industrial inspection, and commercial display.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: Offline video capture of target objects is performed, and an image-mask dataset is generated through frame segmentation, automatic annotation, and manual correction rules. A lightweight semantic segmentation network is fine-tuned and trained based on this dataset, and the trained model is converted into a format executable on the terminal. The terminal constructs a virtual coordinate system and projection mapping rules in a 3D rendering environment, loads the model, and performs semantic segmentation inference on real-time camera images to obtain the mask and statistics of the target objects. The system combines mask information, plane normal vectors, depth estimation, and IMU sensor data to calculate the 3D centroid of the target object in the camera coordinate system through weighted projection and depth reconstruction, and updates the pose of the virtual model accordingly, achieving 3D localization and tracking of the target object.

[0016] This invention has the advantages of lightweight model, strong real-time performance on the terminal side, and high positioning stability, and is suitable for augmented reality applications in mobile and browser environments. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the overall process of the present invention. Detailed Implementation

[0018] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0019] A method for implementing a lightweight semantic segmentation 3D localization and tracking system on the terminal side includes the following steps: S1: Offline video capture and upload.

[0020] Offline trainers use cameras, mobile phones, or other acquisition devices to record videos of the target object and its surrounding environment, obtaining raw video files containing the target object and the scene environment. These video files are then uploaded to a server or offline processing node for storage and subsequent processing.

[0021] S2: Dataset Construction.

[0022] Offline trainers use the video file as input and perform frame segmentation and image normalization on the server or offline processing node to obtain an image and mask dataset for target recognition training.

[0023] S3: Target network training and format conversion.

[0024] The target recognition training set obtained in S2 is input into the lightweight semantic segmentation network, the network is fine-tuned and trained, and after training is completed, the network is exported and its format is converted to obtain a target semantic segmentation network model that can run on mobile devices or browsers.

[0025] S4: Rules for 3D scene construction and coordinate mapping.

[0026] The terminal constructs a virtual 3D scene in the 3D renderer, establishes a virtual coordinate system and the corresponding world coordinate system, and establishes mapping rules between the camera image plane and the virtual 3D space, providing unified coordinate constraints for subsequent centroid and pose calculations.

[0027] S5: Target network delivery and loading.

[0028] The terminal obtains the target network model file generated by S3 from the server, for example, by downloading it via the HTTP protocol, and loads and runs it locally using a semantic segmentation inference framework, such as ONNXRuntimeWeb or TensorFlow.js in a web environment, to complete model warm-up and inference backend configuration.

[0029] S6: Real-time image acquisition and semantic segmentation inference.

[0030] The terminal turns on its camera in a scene environment that is the same as or similar to offline acquisition, and captures camera images frame by frame. Each frame of image is preprocessed according to the target network input specification and then input into the target network for inference to obtain the mask or pixel-level category probability map of the target object.

[0031] S7: 3D centroid estimation using multi-sensor fusion.

[0032] The terminal fuses the target identification mask or mask statistics obtained by S6 with the output data of various sensors (such as IMU, depth estimation module, etc.) to calculate the approximate three-dimensional centroid coordinates of the target identification object in the terminal camera coordinate system.

[0033] S8: 3D pose estimation and subsequent task execution.

[0034] Based on the 3D centroid position obtained by S7, the terminal estimates the pose of the target object by combining IMU or other sensor data, and can then perform subsequent tasks based on this pose. For example, the terminal can use a WebGL renderer to overlay an AR 3D model onto the 3D centroid position of the target object, thereby achieving an augmented reality effect.

[0035] According to claim 1, a lightweight semantic segmentation 3D localization and tracking system for terminal side is preferably provided, wherein step S2 includes the following sub-steps: S21: Video frame cutting and image normalization.

[0036] The original video file uploaded to S1 is segmented into independent image frames on the server or offline processing node according to a preset sampling rate (e.g., several frames per second, configurable), with the timestamp of each frame recorded during the segmentation process. The segmented images are then normalized in terms of size, format, and color space. All image frames are organized and stored according to a directory structure and index files, forming a data source library for subsequent automatic annotation and training.

[0037] S22: Mask generation and annotation.

[0038] For each image obtained in S21, a target identification mask is generated using a combination of automatic and manual annotation. During automatic annotation, an automatic annotation script based on SAM (Segment Anything Model) or an equivalent general segmenter is invoked to generate a candidate mask set. The candidate masks are scored based on area, edge sharpness, and confidence level, and a master mask or composite mask is selected as the initial label according to preset rules. Image frames with low confidence levels or multiple target ambiguities are marked in a manual review queue. Annotators correct and confirm the mask boundaries and categories in a dedicated annotation interface. The manually corrected results are then written back to the training sample database.

[0039] S23: Sample augmentation and dataset splitting.

[0040] Systematic data augmentation is performed on the image-mask pairs obtained from S22 to improve the model's robustness to deformation, scaling, brightness changes, and partial occlusion. The augmentation operations include, but are not limited to: random horizontal / vertical flipping, random small-angle rotation, scaling and cropping, brightness and contrast perturbation, local occlusion synthesis, and background replacement. The augmented samples are divided into training, validation, and test sets according to a preset ratio, and an index file is generated to record the sample paths and augmentation parameters of each subset. The storage format and path conventions of the dataset are clearly defined to ensure unambiguous data loading by the training scripts, improving experimental reproducibility.

[0041] Preferably, step S3 includes the following sub-steps: S31: Lightweight network selection and architecture adaptation.

[0042] For real-time inference requirements on mobile or browser-based devices, a lightweight semantic segmentation network architecture is selected. This architecture can be a variant of an existing lightweight network, with edge enhancement modules introduced at the output or intermediate feature layers to improve segmentation accuracy at target boundaries. The network design clearly defines input / output size constraints and preprocessing specifications, such as input pixel normalization methods, channel order, and input resolution.

[0043] S32: Targeted fine-tuning training.

[0044] The selected lightweight network is trained using a dataset constructed with S2 for transfer learning or fine-tuning. The training process includes: specifying the data input pipeline for the training and validation sets, configuring the loss function (e.g., a weighted combination of pixel-level cross-entropy and region consistency loss), selecting the optimizer and implementing learning rate scheduling strategies, and employing early stopping or optimal checkpoint saving strategies. During fine-tuning, evaluation metrics such as IoU, precision, and recall are recorded at each stage. For false positives in edge regions or small targets, targeted sample selection and retraining are performed to improve the model's recognition stability in real-world deployment scenarios.

[0045] S33: Model export, graph optimization and terminal-side format conversion.

[0046] The trained model is converted to a terminal-side inference framework-compatible format, such as ONNX and / or TensorFlow.js, using an export toolchain. The export process includes input layer preprocessing instructions (including normalization scale, channel order, etc.) and performs consistency verification on the terminal-side loading and inference results to confirm that the difference between the terminal-side output probability distribution and the training result is within an acceptable range. The exported model includes the model file, local preprocessing instructions, and version information for terminal loading, upgrading, and rollback management.

[0047] Preferably, step S4 includes the following sub-steps: S41: 3D spatial context and scene benchmarking construction.

[0048] The terminal establishes a 3D rendering context and constructs a 3D scene within a browser or mobile operating environment. The 3D scene includes at least: a virtual camera node, a 2D texture plane for displaying the camera image, and a scene root node for mounting virtual 3D models. The normal vector of the 2D texture plane preferably faces the virtual camera; the camera image can be dynamically updated and mapped to this plane as a video texture, or it can be superimposed on the image as a 2D element, serving as a rendering base, thereby ensuring that the 3D space and the camera image are aligned in the same world coordinate system.

[0049] S42: Projection parameters and coordinate mapping rules.

[0050] During the rendering initialization phase, the system fixes the projection parameters of the virtual camera and establishes mapping rules from pixel coordinates to world coordinates (or camera coordinates). These rules include at least: the method of constructing the virtual coordinate system, the planar position and orientation of the camera image within the virtual coordinate system, and the projection relationships used for subsequent pose calculations. Subsequent centroid and pose calculations are performed under these constraints, thereby ensuring coordinate consistency between different modules.

[0051] Preferably, step S5 includes the following sub-steps: S51: Model distribution and loading strategy.

[0052] During startup, the terminal downloads the target network model file exported by S33 from the server via a network request (e.g., HTTP) and loads the model and weight parameters in the local execution environment. After loading, the system performs warm-up inference (at least one forward propagation) to initialize backend resources (e.g., WebGL or WebGPU resource allocation on the web client) and completes resource binding and UI state updates for inference in the background.

[0053] S52: Terminal-side inference framework compatibility with backend configuration.

[0054] The terminal selects a suitable inference backend (e.g., WebGL, WebGPU, WASM, etc.) based on device performance and browser capabilities, and optimizes backend parameters and pre-allocates memory for the inference framework. The inference framework provides a unified inference interface: accepting standardized image data (conforming to the input specifications of the exported model) and returning pixel-level probability maps or mask outputs. Developers explicitly configure the inference input size, data arrangement format, and preprocessing flow on the terminal to ensure consistency between the inference input and the training input.

[0055] Preferably, step S6 includes the following sub-steps: S61: Real-time camera frame acquisition.

[0056] During operation, the terminal continuously acquires video streams from the camera and captures image frames from the video stream at a preset frame rate; a timestamp is recorded for each captured image frame for subsequent alignment with sensor data.

[0057] S62: Inference execution and mask probability mapping.

[0058] The worker thread or inference thread receives image frame data obtained by S61 and performs preprocessing operations such as tensor quantization and normalization according to the model input specifications before performing forward inference. The inference result is a probability distribution of each pixel category or a binary mask image. The system can binarize the probability map to generate a mask image according to a preset threshold, and can optionally output mask statistics (such as the number of foreground pixels, centroid candidate coordinates, mask vertical range, etc.) and a mask visualization bitmap.

[0059] S63: Mask post-processing and statistical calculation.

[0060] Post-processing operations such as connected component analysis, noise filtering, and morphological processing are performed on the mask image obtained by S62. The required statistics are calculated, including but not limited to the target mask area, the two-dimensional centroid of the target pixel set, and the height and width range of the target pixels in the image. These statistics are then provided to the subsequent three-dimensional centroid and depth estimation algorithms.

[0061] Preferably, step S7 includes the following sub-steps: S71: Calculation of three-dimensional centroid projection.

[0062] Based on the target mask or its statistics obtained by S6, combined with various sensor data (such as plane normal vectors, IMU attitude information, etc.) or prior knowledge, a plane coordinate model of the camera image is established. The foreground pixels in the mask are mapped from two-dimensional pixel coordinates to plane coordinates in the camera coordinate system, and the projection coordinates of the three-dimensional centroid of the target object on the plane are calculated accordingly.

[0063] S72: Three-dimensional centroid reconstruction.

[0064] By utilizing depth and attitude information output from multiple sensors, such as depth estimation modules, IMUs, and accelerometers, the centroid projection points obtained by S71 are restored from two-dimensional plane coordinates to three-dimensional centroid coordinates in the camera coordinate system. The depth values ​​are then filtered and smoothed to obtain a stable three-dimensional position of the target.

[0065] Example 1 like Figure 1As shown, this invention provides a lightweight semantic segmentation 3D localization and tracking method and system for terminal-side applications. The system includes an offline component and an online component, wherein: 1) Offline part: responsible for collecting videos of the target object and its environment, generating semantic segmentation datasets, fine-tuning and training the lightweight network and converting its format, and finally obtaining the target network that can run on the terminal. 2) Online component: Responsible for loading the target network on the terminal, performing semantic segmentation inference on real-time camera images, and combining multi-sensor information to estimate the three-dimensional centroid and pose of the target object in the camera coordinate system, and completing the rendering and interaction of the three-dimensional model.

[0066] The system implementation method includes steps S1 to S8, and each step is specifically implemented as follows in this embodiment. The invention proposes a lightweight localization and tracking system based on semantic segmentation. This system includes: S1: Offline trainers record and upload videos. Offline trainers use cameras, mobile phones, or other video capture devices to record videos of the target object and its surrounding environment in the scene where the target object is located, and obtain the original video file containing the target object, such as in mp4 format.

[0067] The original video files are uploaded to a server or offline processing node for storage, serving as the foundational data for subsequent dataset construction. In a preferred embodiment, a script is triggered immediately after uploading to prepare for subsequent frame segmentation and annotation.

[0068] S2: Frame segmentation and mask generation to construct the target recognition dataset. Offline trainers use the video recorded in step S1 as input, segment the video into frames, obtain the mask for each frame, and construct a dataset for target recognition. This step can be further broken down into sub-steps S21 to S23: S21: Video Frame Segmentation and Image Normalization The uploaded original video file is processed on the server or offline processing node according to a preset sampling rate, such as several frames per second or once every 5 frames, and divided into independent image frames, and the timestamp of each image frame is recorded.

[0069] In one specific implementation, the processing flow can be as follows: the offline acquisition personnel provide the approximate location of the target object in the initial frame of the video through the annotation interface (e.g., selecting the target object with a closed rectangle), the script cuts the video into frames at fixed intervals (e.g., every 5 frames), and stores the obtained image frames in the original image dataset folder in sequence.

[0070] The images obtained from frame segmentation are standardized in terms of size, format, and color gamut, for example, by unifying them to a fixed resolution and RGB color space. All frame segmentation results are organized according to the directory structure and index files to form a data source library for subsequent automatic annotation and training.

[0071] S22: Mask Generation and Labeling The present invention can use a combination of automated annotation and manual annotation to generate a target mask for the image obtained in step S21.

[0072] In a preferred embodiment, a general segmenter based on SAM (SegmentAnythingModel) or an equivalent is used to call an automatic annotation script for each original image to generate a set of candidate masks. The automatic annotator scores the candidate masks based on indicators such as mask area, edge sharpness, overlap with the initial rectangle position, and confidence level, and selects the master mask or composite mask as the initial label according to preset rules.

[0073] For images with low confidence, multiple target ambiguities, or severe occlusion, the system automatically places the image into a manual review queue. Annotators use a dedicated annotation interface to correct the mask boundaries or confirm the category. The corrected mask results are written back to the training sample library, forming a one-to-one image-mask pair with the original image.

[0074] In another implementation, segmentation tools such as SAMURAI can be used to track the positions of manually labeled target objects in the initial frame, automatically segmenting and generating mask images with the same names as the original image in subsequent frames. After generation, the mask folder is manually reviewed to remove or correct obviously erroneous samples.

[0075] S23: Sample Augmentation and Dataset Construction A systematic data augmentation process is performed on the image-mask pair obtained from S22 to improve the model's robustness to deformation, scaling, brightness changes, and partial occlusion. Augmentation operations may include, but are not limited to: random horizontal / vertical flipping, random small-angle rotation, random scaling and cropping, brightness, contrast and color perturbation, local occlusion synthesis, and background replacement.

[0076] The augmented samples are divided into training, validation, and test sets according to a preset ratio, and an index file is generated to record the sample paths and augmentation parameters of each subset. The storage format and path conventions of the dataset are clearly defined so that the training scripts can load the data unambiguously, improving the reproducibility of the experiment.

[0077] Through the above steps S21 to S23, the semantic segmentation dataset of the target object is constructed, thus realizing step S2.

[0078] S3: Lightweight Semantic Segmentation Network Fine-tuning and Terminal-Side Format Conversion The target recognition training set obtained in step S2 is used as input, and fine-tuned training is performed through a lightweight semantic segmentation network. The network format is then converted to obtain a target network that can run on mobile devices or browsers. This step may specifically include sub-steps S31 to S33: S31: Lightweight Network Selection and Structure Adaptation To meet the real-time inference requirements of mobile or web-based applications, a lightweight semantic segmentation network structure is selected in step S3, preferably the MobileSeg network based on the PaddlePaddle platform or other equivalent lightweight networks. An edge enhancement module can be further introduced into the network structure to improve the segmentation accuracy of target boundary regions.

[0079] When designing the network, clearly define the preprocessing specifications such as input size, channel order, and normalization method to ensure consistency between the training and inference ends.

[0080] S32: Targeted Fine-Tuning Training The selected lightweight network is trained using the dataset constructed in step S2 for transfer learning or fine-tuning. The training process may include: 1) data input PIPELINE specification (including data augmentation, random shuffling, batch size, etc.); 2) loss function configuration, such as a weighted combination of pixel-level cross-entropy loss and region consistency loss; 3) optimizer selection and learning rate scheduling strategy; 4) early stopping strategy or best checkpoint saving strategy; 5) recording evaluation metrics such as IoU, precision, and recall during training.

[0081] To further reduce model size and improve inference efficiency, a quantization strategy from FP32 to FP16 can be used during training. After training, a lightweight .paddle format model file is obtained.

[0082] For errors such as false detections of edge regions or small targets, the stability of the model in real deployment scenarios can be improved by selecting typical samples and fine-tuning the training again.

[0083] S33: Model Export, Graph Optimization, and Terminal-Side Format Conversion The trained model is converted to ONNX format sequentially using a toolchain, then to TensorFlow format using the onnx2tf tool, and finally to tfjs format using TensorFlow.js, resulting in a target network that can run in a browser environment or a lightweight terminal.

[0084] During the export process, input preprocessing instructions (including normalization scale, channel order, input resolution, etc.) and model version information are saved. After export, the output probability maps of the training end and the terminal side inference are compared to verify that they are consistent within an acceptable error range.

[0085] Through the above steps S31 to S33, the fine-tuning training of the lightweight semantic segmentation network and the terminal-side format conversion are completed, thus realizing step S3.

[0086] S4: 3D Scene Construction and Mapping of Virtual Coordinate System to World Coordinate System The terminal constructs a 3D scene within the 3D renderer and defines the mapping rules between the virtual coordinate system and the world coordinate system. This step may include sub-steps S41 and S42: S41: 3D Spatial Context and Scene Benchmarking Construction The terminal creates a 3D rendering context (e.g., based on WebGL or three.js) in a browser environment or mobile operating environment to construct a 3D scene. The 3D scene includes at least: 1) a virtual camera node, whose position is fixed at the origin (0,0,0) of the virtual coordinate system; 2) a 2D texture plane for displaying the camera image, the normal vector of which preferably faces the virtual camera; and 3) a scene root node for mounting virtual 3D models.

[0087] The camera image can be dynamically mapped onto the two-dimensional plane as a video texture, or it can be superimposed on the image as a 2D element as a rendering base, thereby ensuring that the three-dimensional space and the camera image are aligned under a unified coordinate system.

[0088] S42: Camera projection parameters and pixel-to-world coordinate mapping rules During the rendering initialization phase, the system fixes the projection parameters of the virtual camera and agrees on the mapping rules from pixel coordinates to the virtual coordinate system / world coordinate system.

[0089] In a preferred embodiment, the camera texture plane is defined as being located at: Among them The height of the camera view. This is the camera's field of view; when camera intrinsics are lacking, this can be used as the default value. It is 90°. For the camera image... Pixels in coordinate system Its three-dimensional coordinates in the aforementioned plane It can be represented as: in, These are the horizontal coordinates in the camera coordinate system. These are the vertical coordinates in the camera coordinate system. Here, W represents the depth coordinate in the camera coordinate system, and W represents the width of the camera view.

[0090] The above three-dimensional mapping relationship provides the basis for subsequent calculation of the in-plane centroid projection and three-dimensional centroid through the mask pixel set.

[0091] Step S4 is achieved through steps S41 and S42.

[0092] Step S5: Terminal target network acquisition and inference framework loading The terminal obtains the target network trained in step S3 (e.g., downloads the model file from the server via HTTP protocol) and loads and runs it using a semantic segmentation network framework. This step may specifically include sub-steps S51 and S52: S51: Model Distribution and Loading Strategy During the startup phase, the terminal downloads the target network model file from the server via protocols such as HTTP, loads it into the execution environment, and performs pre-inference to initialize underlying computing resources (such as textures and buffers in the WebGL or WebGPU backend). After the model is loaded, the system updates the UI state and prepares to share the inference results with the rendering module.

[0093] S52: Terminal-side inference framework compatibility and backend configuration The terminal selects a suitable inference backend based on device performance; for example, in a web environment, it can choose a WebGL, WebGPU, or WASM backend. After the inference framework is loaded, parameters such as input size, channel order, and normalization method are configured to achieve the same input specifications as the training end, and a unified inference interface is provided: receiving standardized image data and outputting pixel-level probability maps or mask maps.

[0094] Step S5 is achieved through steps S51 and S52.

[0095] Step S6: Real-time camera acquisition and target mask inference When the terminal is running, it opens the device's camera in the scene environment or a similar environment where the video was recorded in step S1. Each frame of camera image data is transmitted to the target network for inference to obtain the target object mask. This step can be further broken down into sub-steps S61 to S63: S61: Real-time camera frame acquisition The terminal continuously captures image frames from the camera video stream at a preset frame rate and records the timestamp of each frame for alignment with sensor data such as IMU.

[0096] S62: Inference Execution and Mask Probability Mapping The inference thread receives the image frame obtained in step S61, performs preprocessing operations such as resolution scaling, tensor quantization and normalization on it, and inputs the processed data into the target network loaded in step S5 for forward inference to obtain the probability distribution of each pixel category or the binary segmentation result.

[0097] The probability map is binarized according to a preset threshold to generate a target identification mask map, and statistical quantities such as the number of foreground pixels, the vertical range of the mask, and the two-dimensional centroid can be calculated.

[0098] S63: Mask Post-processing and Statistical Calculation Post-processing operations such as connected component analysis, noise filtering, and morphological processing are performed on the mask obtained by S62 to stabilize the mask region boundary and output key statistics for subsequent 3D reconstruction, including mask area, mask pixel set and its two-dimensional centroid coordinates.

[0099] Step S6 is achieved through the above steps S61 to S63.

[0100] Step S7: 3D centroid estimation via multi-sensor fusion During terminal operation, an approximate three-dimensional centroid of the target object in the terminal camera coordinate system is calculated by combining multiple sensor fusion methods with the target object mask. This step can be further refined into sub-steps S71 and S72: S71: Calculation of 3D centroid projection The system establishes a geometric model on the camera texture plane, and marks the three-dimensional coordinates of the plane corresponding to all foreground pixels in the mask as follows: x i y i For p i The coordinates on the x and y axes, where For the aforementioned plane depth .

[0101] Assume the normal vector of the tangent plane containing the target object is The projection of the three-dimensional centroid of the target object onto the plane can then be calculated based on the following weighted form. : in express and The inner product of the two sides. This weighted projection method can achieve more stable centroid estimation even when image distortion is caused by camera perspective.

[0102] S72: Three-dimensional centroid reconstruction The system acquires the depth information of the target object and the camera pose information through a depth estimation module and sensors such as an IMU.

[0103] When the current frame has a valid depth estimation result, record the depth. If the current frame depth is missing, adaptive compensation can be performed on the nearest valid depth based on the mask area ratio. in, The most recently recorded valid depth, and These are the mask areas of the current frame and the most recent frame, respectively.

[0104] make The distance from the plane to the camera. for In the camera coordinate system, where R is the camera rotation matrix calculated from the IMU output, the three-dimensional centroid of the target object is... It can be represented as: The system uses a sliding window to record IMU data and centroid estimation results in recent frames. Through linear acceleration calculation and weighted filtering mechanism, it fuses and smooths the three-dimensional coordinates at different times to obtain a stable current three-dimensional centroid position.

[0105] Step S7 is achieved through the above steps S71 and S72.

[0106] Step S8: Pose-based 3D model rendering and interaction Based on the centroid position obtained in step S7 and the pose jointly acquired by sensors such as the IMU, the terminal renders the virtual 3D model onto the target object location in the 3D rendering environment and supports interactive operations, thereby achieving augmented reality effects.

[0107] In the specific implementation, a virtual 3D space is constructed using three.js, and the virtual camera position is fixed at (0,0,0). When the plane detection and centroid estimation are completed for the first time, the 3D model is made visible, and its initial orientation is set to the quaternion direction of the corresponding plane normal vector.

[0108] During system operation, the virtual camera orientation is updated to the quaternion data obtained by the current IMU in each frame, and the position of the 3D model is updated to the target 3D centroid coordinates obtained in step S7, so as to maintain the stable superposition of the model in the real scene.

[0109] In addition, the webpage can listen for user touch events to achieve interactive functions: when a single-finger swipe operation is detected, the 3D model is rotated around the Y-axis by a corresponding angle according to the swipe distance, thus rotating the model; when a two-finger zoom operation is detected, the scaling factor of the 3D model is adjusted according to the change ratio of the distance between the two fingers, thus enlarging and shrinking the model.

[0110] Through the above step S8, the rendering and interaction of the 3D model driven by the object's pose are completed, thereby realizing the augmented reality presentation of a lightweight positioning and tracking system based on semantic segmentation on mobile devices or browsers.

[0111] Glossary Semantic segmentation: Semantic segmentation is a computer vision task that uses deep learning algorithms to assign class labels to pixels. It is one of three subcategories in the overall image segmentation process, helping computers understand visual information. Semantic segmentation identifies sets of pixels and classifies them based on various features. The other two subcategories of image segmentation are instance segmentation and panoptic segmentation.

[0112] Tracking: The process by which a system acquires sensor pose in real time based on changes in the target's position in a real scene, re-establishes a spatial coordinate system according to the user's current perspective, and renders the virtual scene onto the accurate position in the real environment is called tracking.

[0113] Feature points: Feature points are points in image processing where grayscale values ​​change drastically or where edge curvature is large. They are used to identify targets, recognize objects, and support image matching algorithms. They include corner points and other types. Corner points are defined as points in a 2D grayscale image where grayscale brightness information changes drastically in different directions, or points with maximum curvature on edge contours. Feature points consist of keypoints (containing location, size, and orientation information) and descriptors (recording information about surrounding pixels), playing a crucial role in 3D reconstruction, autonomous driving localization, and image matching.

[0114] Simultaneous Localization and Mapping (SLAM) is a core technology that enables robots to locate themselves in unknown environments and build environmental maps in real time using sensors. Its core principle is to fuse measurement data from sensors (such as LiDAR, cameras, and millimeter-wave radar) to simultaneously estimate the robot's own trajectory and the coordinates of environmental features.

[0115] Terminal: refers to mobile devices or browser devices equipped with cameras, IMU sensors, and 3D rendering capabilities, including mobile phones, tablets, and web clients, used for local execution of semantic segmentation inference, 3D localization, and tracking.

Claims

1. A lightweight semantic segmentation 3D localization and tracking method on the terminal side, characterized in that, The specific steps are as follows: S1. Construct a server-terminal architecture. The terminal includes sensors and an image acquisition module. The terminal acquires and uploads the original video of the target object and its surrounding scene. S2. Normalize the original video and construct an image-mask dataset for target recognition, specifically: The original video uploaded in step S1 is divided into independent image frames according to a preset sampling rate and timestamps are recorded. The image frames are standardized in terms of size, format, and color space. They are then organized and stored according to a directory structure and index files to form a data source library. The SAM model is called to generate candidate masks for image frames. Candidate masks are then selected based on area, edge sharpness, and confidence score. Low-confidence image frames are manually reviewed and corrected to obtain image-mask pairs. Data augmentation operations are performed on the image-mask pairs, including random horizontal or vertical flipping, random small-angle rotation, scaling and cropping, brightness and contrast perturbation, local occlusion synthesis and background replacement. The augmented image-mask pairs are divided into training set, validation set and test set according to a preset ratio and an index file is generated. S3. Input the image-mask dataset into the lightweight semantic segmentation network for training, and convert the trained model into a mobile-compatible inference model. S4. Construct a virtual 3D scene on the terminal, building a scene structure that includes virtual camera nodes, a 2D texture plane for the camera view, and a root node for the virtual model. Map the camera view to a dynamic video texture on the 2D texture plane to achieve coordinate system alignment. Fix the virtual camera projection parameters and establish mapping rules from pixel coordinates to world coordinates or camera coordinates. The mapping rules include the virtual coordinate system construction method, the planar position and orientation of the camera view in the virtual coordinate system, and subsequent centroid and pose calculations will follow this mapping convention. S5. The terminal downloads and loads the inference model from the server, runs it locally through the semantic segmentation inference framework, and completes the model preheating and inference backend configuration. S6. The terminal collects real-time image frames from the camera frame by frame, inputs them into the inference model to perform semantic segmentation, and obtains the mask of the target object and the mask statistics. S7. The terminal fuses the mask and mask statistics with the sensor output data to establish a planar coordinate model of the camera image. It then calculates the three-dimensional centroid coordinates of the target object in the camera coordinate system through weighted projection, depth adaptive compensation, and filtering smoothing. Specifically: The three-dimensional centroid projection coordinates are calculated using weighted projection, and the formula is as follows: Where n is the normal vector of the tangent plane, p i =(x i ,y i (x, z0) represents the planar coordinates of the foreground pixel, x... i y i For p i In the x and y coordinates, z0 represents the depth of the camera texture plane, and N represents the total number of foreground pixels in the mask. Indicate n and The inner product; Using data from multiple sensors, the centroid projection coordinates are restored from two-dimensional planar coordinates to three-dimensional centroid coordinates in the camera coordinate system. When depth is missing, adaptive compensation is performed according to the mask area ratio. The compensation formula is as follows: in The most recently recorded valid depth, and These are the mask areas of the current frame and the most recent frame, respectively; Then, the three-dimensional centroid coordinates are reconstructed using the formula, and a stable three-dimensional target position is obtained by combining it with sliding window weighted filtering. The formula is as follows: in The distance from the camera texture plane to the camera. Projected coordinates The coordinates in the camera coordinate system, where R is the camera rotation matrix calculated from the sensor output data; S8. Estimate the target's position and pose based on the three-dimensional centroid coordinates and data from multiple sensors, and complete the overlay rendering and interaction of the virtual three-dimensional model in the real scene.

2. The terminal-side lightweight semantic segmentation 3D localization and tracking method according to claim 1, characterized in that, S3 includes the following steps: Step S3 includes the following sub-steps: S31. To meet the real-time inference requirements of the terminal, a lightweight semantic segmentation network is selected, and the network input and output constraints and preprocessing specifications are defined. S32. Use the dataset constructed in step S2 to perform transfer learning or fine-tuning training on the improved network to obtain the trained model. S33. Convert the trained model into a terminal-compatible model, with preprocessing instructions and verify the consistency between the terminal output and the training output.

3. The terminal-side lightweight semantic segmentation 3D localization and tracking method according to claim 1, characterized in that, S5 includes the following steps: S51. The terminal requests the terminal-compatible model exported in step S33 from the server via the network, loads it, completes the inference resource binding, and updates the interface status after the model is loaded. S52. The terminal selects the inference backend based on the device performance and performs parameter optimization and memory pre-allocation for the inference framework.

4. The terminal-side lightweight semantic segmentation 3D localization and tracking method according to claim 1, characterized in that, S6 includes the following sub-steps: S61. Acquire camera image frames at a preset frame rate and record timestamps to achieve time alignment with data from multiple sensors; S62. After performing resolution scaling, tensor quantization, and normalization preprocessing on the image frame, perform forward inference, output a pixel-level probability map, and generate a mask map by binarization. Calculate the mask statistics and visualization bitmap. The mask statistics include the number of foreground pixels, centroid candidate coordinates, and vertical range of the mask. S63. Perform connected component analysis, noise filtering, and morphological processing on the mask image obtained in step S62, calculate the required statistics, including the target mask area, the two-dimensional centroid of the target pixel set, and the height and width range of the target pixels in the image, and provide the above statistics to the subsequent three-dimensional centroid and depth estimation algorithms.

5. A lightweight semantic segmentation 3D localization and tracking method for the terminal side according to claim 2, characterized in that, The basic lightweight semantic segmentation network of S31 is the MobileSeg network of the PaddlePaddle platform or its equivalent lightweight semantic segmentation network.

6. A lightweight semantic segmentation 3D localization and tracking system for terminal-side applications, characterized in that, It includes an offline processing module and an online processing module; the offline processing module is deployed on a server and includes a video acquisition unit, a dataset construction unit, a network training and format conversion unit, and a model storage and distribution unit; the video acquisition unit is used to acquire and upload the target original video; the dataset construction unit is used to construct an image-mask dataset; the network training and format conversion unit is used to train and convert the format of the lightweight semantic segmentation network with edge enhancement module; The model storage and distribution unit is used to store and distribute inference models to the terminal; The online processing module is deployed on the terminal and includes a 3D scene construction unit, a model loading unit, an image acquisition and inference unit, a multi-sensor fusion unit, and a pose rendering and interaction unit. The 3D scene construction unit is used to construct a virtual 3D scene and define coordinate mapping rules. The model loading unit is used to download and load the inference model and configure the inference framework. The image acquisition and inference unit is used to acquire real-time images and output masks and statistics. The multi-sensor fusion unit is used to fuse sensor data and depth estimation data to calculate the three-dimensional centroid; The pose rendering interaction unit is used for pose estimation, virtual model rendering, and user interaction.

7. A lightweight semantic segmentation 3D localization and tracking system for the terminal side according to claim 6, characterized in that, The multi-sensor fusion unit incorporates a weighted projection calculation module, a depth adaptive compensation module, and a sliding window filtering module, which are used to perform the three-dimensional centroid calculation and smoothing optimization as described in claim 5.

8. A lightweight semantic segmentation 3D localization and tracking system for the terminal side according to claim 7, characterized in that, The pose rendering interaction unit uses the three.js, LayaAir, and WebGL rendering frameworks to implement virtual model overlay and supports rotation and scaling interactions. The system is applicable to augmented reality scenarios such as cultural tourism, education, industrial inspection, and commercial display.