AI-based open world three-dimensional object detection and segmentation method and electronic equipment
Through an AI-based open-world 3D object detection and segmentation method, it receives images and user interaction instructions, generates alpha channel masks and optimizes edges, solving the problem of missing 3D geometry in existing technologies, achieving direct conversion from 2D to 3D and improving segmentation stability.
Patent Information
- Application Number
- CN202510730937.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-12
AI Technical Summary
Existing 2D segmentation systems cannot directly obtain quantitative data on the physical size, spatial pose, and occlusion relationship of objects. They have the problem of missing three-dimensional geometry and lack segmentation stability and cross-domain generalization capabilities in dynamic scenes.
An AI-based open-world 3D object detection and segmentation method is adopted. It receives images and user interaction instructions through a visual basic model, generates an alpha channel mask and applies an edge optimization algorithm, and outputs a PNG format image with a transparent channel to achieve end-to-end conversion from 2D segmentation to 3D detection.
It achieves direct conversion from 2D images to 3D detection, improves segmentation stability and accuracy, solves the problem of missing 3D geometry, and optimizes the accuracy and visual effects of object segmentation.
Smart Images

Figure CN120635880A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to an AI-based open-world three-dimensional object detection and segmentation method, electronic device, and storage medium. Background Art
[0002] In the field of computer vision, image object segmentation technology has evolved through three generations: the traditional image processing era, the deep learning explosion, and the open-world exploration era. Throughout these generations, the following technical flaws have emerged: The 3D geometry missing issue: Existing 2D segmentation systems cannot directly obtain quantitative data on an object's physical dimensions (length / width / height), spatial pose (6DoF parameters), and occlusion relationships. Summary of the Invention
[0003] The embodiments of the present invention aim to provide an AI-based open-world three-dimensional object detection and segmentation method, electronic device, and storage medium, aiming to solve the problem of three-dimensional geometry loss in existing image object segmentation technologies.
[0004] To solve the above technical problems, a first embodiment of the present invention provides an AI-based open-world 3D object detection and segmentation method, comprising:
[0005] Receive external input images and user input operation interaction instructions;
[0006] Based on the visual basic model, the user's input operation interaction instructions are processed to the external input image, and the preliminary segmented object area information is output;
[0007] Generate alpha channel mask based on the object area information of the preliminary segmentation;
[0008] Apply an edge optimization algorithm to adjust the edge pixel values of the alpha channel mask and output a PNG format image with a transparent channel.
[0009] Optionally, the visual base model includes an input layer, and the input layer includes a multimodal input module;
[0010] The receiving of an external input image and an operation interaction instruction input by a user includes:
[0011] Calling the multimodal input module to receive external input images and user input operation interaction instructions;
[0012] Spatially align the external input RGB images and adjust the input RGB images to a uniform size.
[0013] Optionally, the visual basic model further includes a feature extraction layer, and the feature extraction layer includes an ROI feature extraction device;
[0014] The method processes the external input image based on the user's input operation interaction instruction based on the visual basic model and outputs the preliminary segmented object area information, including:
[0015] Based on the operation interaction instruction input by the user, the ROI feature extraction device is called to obtain key features related to the target object from the spatially aligned RGB image input externally;
[0016] Based on the key features related to the target object, the region growing algorithm is used to expand the growing region and output the preliminary segmented object region information.
[0017] Optionally, the ROI feature extraction device includes a SAM encoder and a DINOv2 encoder;
[0018] The operation interaction instruction based on the user input calls the ROI feature extraction device to obtain key features related to the target object from the spatially aligned RGB image input externally, including:
[0019] Call the SAM encoder to encode the input spatially aligned RGB image, extract the spatial features of the RGB image, and output a 1027-dimensional feature vector;
[0020] Call the DINOv2 encoder to further extract the geometric features of the RGB image and output a 768-dimensional feature vector;
[0021] The dynamic feature fusion module fuses the 1027-dimensional feature vector output by the SAM encoder and the 768-dimensional feature vector output by the DINOv2 encoder to generate a fused feature including geometric primitives that describe the structure of the object.
[0022] Optionally, the visual basic model further includes an innovative processing layer, and the innovative processing layer includes a dynamic feature fusion module;
[0023] The method processes an external input image based on the user's input operation interaction instruction based on the visual basic model and outputs preliminary segmented object area information, and further includes:
[0024] The dynamic feature fusion module is called to fuse the 1027-dimensional feature vector output by the SAM encoder and the 768-dimensional feature vector output by the DINOv2 encoder to generate a fused feature including geometric primitives describing the structure of the object.
[0025] Optionally, the innovative processing layer further includes a geometric decoupling module;
[0026] The method processes an external input image based on the user's input operation interaction instruction based on the visual basic model and outputs preliminary segmented object area information, and further includes:
[0027] The geometric decoupling module is called to parse the geometric primitive information of the object according to the fusion features.
[0028] Optionally, the innovation processing layer further comprises a 3D primitive generator;
[0029] The method processes an external input image based on the user's input operation interaction instruction based on the visual basic model and outputs preliminary segmented object area information, and further includes:
[0030] The 3D primitive generator is called to generate a 3D primitive model representing the 3D structure of the object based on the geometric primitive information provided by the geometric decoupling module and in combination with the spatial position and shape characteristics of the object.
[0031] Optionally, the innovation processing layer further includes a pose estimator;
[0032] The method processes an external input image based on the user's input operation interaction instruction based on the visual basic model and outputs preliminary segmented object area information, and further includes:
[0033] The pose estimator is called to calculate the pose information of the object in 3D space according to the 3D primitive model and fusion features, obtain the rotation matrix and cube parameters of the object through feature matching of basic geometric space transformation, and output 6DoF parameters.
[0034] Optionally, the innovative treatment layer further comprises a ZEM stabilizer;
[0035] The method processes an external input image based on the user's input operation interaction instruction based on the visual basic model and outputs preliminary segmented object area information, and further includes:
[0036] The ZEM stabilizer is called to stabilize the processing of the pose estimator based on the zero embedding mapping mechanism and output standardized 6DoF parameters.
[0037] Optionally, the visual basic model further includes an output layer, and the output layer includes a feature generation module and an output interface;
[0038] The method processes an external input image based on the user's input operation interaction instruction based on the visual basic model and outputs preliminary segmented object area information, and further includes:
[0039] Calling the feature generation module to generate a 3D bounding box of the object in 3D space based on the 6DoF parameters, and generating an accurate instance mask based on the projection constraint and the fused features;
[0040] The output interface is called to output the 6DoF parameters and the instance mask as final output, wherein the 6DoF parameters and the instance mask become key features related to the target object obtained from the spatially aligned RGB image.
[0041] Optionally, the method of processing an external input image based on the user's input operation interaction instruction based on the visual basic model and outputting preliminary segmented object region information further includes:
[0042] Based on the key features related to the target object output by the output layer, a region growing algorithm is used to expand the growing region and output preliminary segmented object region information.
[0043] Optionally, generating an alpha channel mask based on the initially segmented object region information includes:
[0044] The object area that is initially segmented is binarized to separate the object area from the background and obtain a binary mask:
[0045] Smoothing the binary mask based on the 3D bounding box of the object area;
[0046] According to the depth information of the object area or according to a certain weight distribution rule, different transparency values are assigned to the smoothed binary mask to generate an alpha channel mask.
[0047] Accordingly, an embodiment of the second aspect of the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and running on the processor. When the computer program is executed by the processor, it implements the AI-based open-world three-dimensional object detection and segmentation method described in the embodiment of the first aspect of the present invention.
[0048] Accordingly, an embodiment of the third aspect of the present invention provides a storage medium, on which a program of an AI-based open-world three-dimensional object detection and segmentation method is stored. When the program of the AI-based open-world three-dimensional object detection and segmentation method is executed by a processor, the AI-based open-world three-dimensional object detection and segmentation method described in the embodiment of the first aspect of the present invention is implemented.
[0049] Compared with the prior art, embodiments of the present invention provide an AI-based open-world 3D object detection and segmentation method, electronic device, and storage medium. The AI-based open-world 3D object detection and segmentation method includes: receiving an externally input image and an operation interaction instruction input by a user; processing the externally input image based on the operation interaction instruction input by the user based on a visual basic model, and outputting preliminary segmented object region information; generating an alpha channel mask based on the preliminary segmented object region information; applying an edge optimization algorithm to adjust the edge pixel values of the alpha channel mask, and outputting a PNG format image with a transparent channel. By receiving external input images and user-input interactive commands, the system integrates multiple input methods to achieve multimodal input. By processing the external input images based on user-input interactive commands based on a visual foundational model and outputting preliminary segmented object region information, the system achieves end-to-end conversion from 2D segmentation to 3D detection, directly obtaining 3D object detection results from the input 2D image. This breakthrough in geometric perception and improved segmentation stability are achieved. By generating an alpha channel mask based on the preliminary segmented object region information, the object region can be separated from the background region and smoothed, resulting in a more natural boundary transition and better fusion of the generated mask with the original image. By applying an edge optimization algorithm to adjust the edge pixel values of the alpha channel mask and outputting a PNG format image with a transparent channel, the alpha channel mask edges can be optimized, improving the accuracy and visual quality of object segmentation. This solves the problem of 3D geometry loss in existing image object segmentation technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0051] Figure 1 This is a flowchart of an AI-based open-world 3D object detection and segmentation method provided by the present invention;
[0052] Figure 2 Schematic diagram of the structure of a visual basic model used in an AI-based open-world three-dimensional object detection and segmentation method provided by the present invention;
[0053] Figure 3 1 is a flow chart of S2 in an AI-based open-world three-dimensional object detection and segmentation method provided by the present invention;
[0054] Figure 4 It is a structural schematic diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION
[0055] In order to facilitate the understanding of the present invention, the present invention will be described in more detail below with reference to the accompanying drawings and specific embodiments. It should be noted that when an element is described as being "fixed to" another element, it can be directly on the other element, or there can be one or more centered elements therebetween. When an element is described as being "electrically connected" to another element, it can be directly connected to the other element, or there can be one or more centered elements therebetween. The orientation or positional relationship indicated by the terms "upper", "lower", "inner", "outer", "bottom" etc. used in this specification is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", "third", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0056] Unless otherwise defined, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which this invention pertains. The terms used in this specification are intended only to describe specific embodiments and are not intended to limit the invention. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items.
[0057] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0058] First, here are some explanations of the terms that appear below:
[0059] Geometric Perception Breakthrough: In computer vision, this refers to the ability to accurately extract 3D geometric information (such as an object's physical size, spatial position, and occlusion relationships) from 2D image features through technological innovation. Specifically, by establishing a direct link between 2D features and 3D space, this overcomes the inability of traditional 2D segmentation techniques to capture an object's 3D geometric properties, enabling the perception of an object's position, form, and structure in real space.
[0060] Explicit mapping: This refers to establishing a direct correspondence between 2D image features and 3D spatial coordinates through a clear mathematical formula or model. This relationship can be quantitatively described using interpretable parameters (such as rotation matrices and translation vectors), rather than relying on implicit black-box mappings.
[0061] End-to-end conversion: This refers to the complete process from inputting 2D images to outputting 3D detection results (such as 3D bounding boxes and pose parameters) without manual intervention or step-by-step processing, and is completed directly by a single model or system. This process integrates steps such as feature extraction, geometric mapping, and 3D reconstruction to form an integrated processing chain.
[0062] Temporal consistency constraint: In dynamic scenes (such as video sequences), the consistency of segmentation results in adjacent frames is constrained to ensure the stability of object segmentation in the temporal dimension. Specifically, the difference between the segmentation mask of the current frame and the mask of the previous frame after motion alignment is calculated and used as the loss function to optimize the model.
[0063] Dynamic scene: In the field of computer vision, a dynamic scene refers to a complex visual environment that contains elements or conditions that change over time, causing real-time changes in image content, imaging environment, or object status.
[0064] In the field of computer vision, image object segmentation technology has evolved through three generations: the traditional image processing era, the deep learning explosion, and the open world exploration era. During these generations, the following technical flaws exist:
[0065] The first generation: the era of traditional image processing (2000-2012). This generational evolution of technology focused on edge detection (such as the Canny operator) and region growing algorithms (such as the GrabCut algorithm). Its typical technical limitations included relying on manually set thresholds and a recall rate of less than 60% in complex backgrounds.
[0066] Second Generation: The Deep Learning Explosion Period (2012-2020). This generational evolution saw breakthroughs in key technologies, including FCN (2015) achieving end-to-end pixel-level prediction, U-Net (2015) establishing an encoder-decoder architecture, and Mask R-CNN (2017) establishing the instance segmentation paradigm. However, this generation also had certain drawbacks, such as the need for large amounts of labeled data (the COCO dataset has >2 million labeled instances) and closed-category limitations (recognizing only 80 predefined object categories).
[0067] The third generation: the open-world exploration phase (2020-present). During this generational evolution, the following emerging technologies emerged: visual foundational models (SAM / DINO) provide zero-shot capabilities, and multimodal fusion (CLIP, etc.) enable text-guided segmentation. These emerging technologies present unresolved challenges: maintaining geometric consistency from 2D to 3D (existing methods have projection errors >15%) and real-time segmentation latency for dynamic scenes (>300ms at 4K resolution).
[0068] In the above generational evolution of technology, the following technical defects exist:
[0069] 1. Missing 3D geometry: Existing 2D segmentation systems cannot directly obtain quantitative data on the object’s physical size (length / width / height), spatial pose (6DoF parameters), and occlusion relationship.
[0070] 2. Dynamic scene adaptability issues: The performance is poor in the following scenarios: sudden changes in illumination (accuracy drops by 40% when lux changes > 1000), motion blur (edge breakage rate > 25% when speed > 60 km / h), and transparent objects (glassware segmentation error rate > 65%).
[0071] 3. Cross-domain generalization bottleneck: There is a performance degradation problem during model migration.
[0072] In view of this, if Figure 1 As shown, the present invention provides an open-world three-dimensional object detection and segmentation method based on AI (Artificial Intelligence), comprising:
[0073] S1, receiving an external input image and an operation interaction instruction input by a user;
[0074] S2, processing the external input image based on the user's input operation interaction instructions based on the visual basic model, and outputting the preliminary segmented object area information;
[0075] S3, generating an alpha channel mask based on the object area information of the preliminary segmentation;
[0076] S4. Apply an edge optimization algorithm to adjust the edge pixel values of the alpha channel mask and output a PNG format image with a transparent channel.
[0077] In this embodiment, an AI-based open-world 3D object detection and segmentation method is provided, including: receiving an external input image and an operation interaction instruction input by a user; processing the external input image based on the operation interaction instruction input by the user based on a visual basic model, and outputting preliminary segmented object region information; generating an alpha channel mask based on the preliminary segmented object region information; applying an edge optimization algorithm to adjust edge pixel values of the alpha channel mask, and outputting a PNG format image with a transparent channel. Therefore, an AI-based open-world 3D object detection and segmentation method is provided. This method integrates multiple input methods and implements multimodal input by receiving external input images and user-input interactive commands. The external input image is processed based on a visual foundational model, and preliminary segmented object region information is output. This method achieves end-to-end conversion from 2D segmentation to 3D detection, directly obtaining 3D object detection results from the input 2D image, achieving a breakthrough in geometric perception and improving segmentation stability. An alpha channel mask is generated based on the preliminary segmented object region information, separating the object region from the background region and smoothing the object region to achieve a more natural boundary transition and better fusion of the generated mask with the original image. An edge optimization algorithm is applied to adjust the edge pixel values of the alpha channel mask, and a PNG format image with a transparent channel is output. This optimizes the alpha channel mask edge, improving the accuracy and visual quality of object segmentation. This method thus addresses the 3D geometry loss problem of existing image object segmentation techniques.
[0078] In this paper, visual base models (SAM / DINO) are used to provide zero-shot capability and multimodal fusion (CLIP, etc.) is used to achieve text-guided segmentation.
[0079] SAM (Segment Anything Model) is a segmentation-anything model proposed by Meta. It breaks through the boundaries of segmentation and greatly promotes the development of basic computer vision models. SAM is a hint-based model that has been trained on over 1 billion masks on 11 million images, achieving strong zero-shot generalization. The SAM model architecture mainly consists of three parts: an image encoder, a hint encoder, and a mask decoder. DINO (Distillation with No Labels) is a self-supervised learning framework proposed by the Facebook AI team in 2021, mainly for vision tasks. DINO generates high-quality feature representations by learning representations from unlabeled image data. Its design is inspired by knowledge distillation, but unlike traditional methods, it does not require a pre-trained teacher model. Instead, it uses "self-distillation" to enable the model to self-optimize during training. DINO has the characteristics of no label dependency, high-performance representation, high computational efficiency, strong transferability, and advantages in interpretability and visualization. It has a wide range of applications in vision pre-training, data-scarce fields, unsupervised clustering, and multimodal learning.
[0080] like Figure 2 As shown, the visual basic model adopted by the present invention includes an input layer, a feature extraction layer, an innovation processing layer and an output layer.
[0081] The input layer receives external input images and user input operation interaction instructions; the input layer includes a multimodal input module 10, which receives external input images and user input operation interaction instructions, and includes an optical lens 11 and an interaction unit 12.
[0082] The feature extraction layer processes the external input image based on the operation interaction instructions input by the user and extracts key features related to the target object; the feature extraction layer includes a ROI feature extraction device 20, and the ROI feature extraction device 20 includes a SAM encoder 21 and a DINOv2 encoder 22.
[0083] The innovative processing layer generates 6DoF (six degrees of freedom) parameters of the target object based on the extracted key features related to the target object; the innovative processing layer includes a dynamic feature fusion module 31, a geometric decoupling module 32, a 3D primitive generator 33, a pose estimator 34 and a ZEM stabilizer 35.
[0084] The output layer generates a 3D bounding box of the object in 3D space based on the 6DoF parameters, and generates an accurate instance mask based on the projection constraint and the fusion feature F, and outputs the 3D bounding box and instance mask; the output layer includes a feature generation module 41 and an output interface 42.
[0085] In one embodiment, in step S1, receiving an external input image and an operation interaction instruction input by a user specifically includes:
[0086] S11. Call the multimodal input module of the input layer of the visual basic model to receive external input images and user input operation interaction instructions.
[0087] The external input image is an RGB image, which is raw color image data and serves as the fundamental source of visual information. The external input image can be an RGB image uploaded by the user, or an RGB image of the external environment captured by the optical lens 11. The optical lens 11 captures image data of the external environment and provides raw image information for subsequent processing.
[0088] Receive user input of operation interaction instructions, determine the location or range of the target object, and provide semantic guidance for subsequent processing. The operation interaction instructions can reflect the user's intention information, including: box selection, click selection, or text description. By receiving user input of operation interaction instructions including box selection, click selection, or text description, the user's intention information can be obtained, thereby determining the location or range of the target object and providing semantic guidance for subsequent processing. For example, the user's intention information can be obtained by receiving the user input of operation interaction instructions including box selection, click selection, or text description through the interaction unit 12.
[0089] By calling the multimodal input module 10 to receive external input images and user input operation interaction instructions, multiple input methods can be integrated to achieve multimodal input.
[0090] S12. Spatially align the external input RGB image and adjust the input RGB image to a uniform size to ensure consistency in subsequent processing and facilitate feature extraction and model calculation.
[0091] For example, the image size is unified to 512*512, and the external input RGB image is spatially aligned and resized to 512*512. This ensures consistency in subsequent processing and facilitates feature extraction and model calculation.
[0092] S13: Rapidly transmit the spatially aligned RGB image to the ROI feature extraction device 20 through the optical fiber channel.
[0093] Specifically, the fiber channel is responsible for high-speed data transmission, which quickly transmits the collected spatially aligned image data and user operation instructions to the ROI feature extraction device 20, with a transmission delay of less than 5 milliseconds (ms), thereby ensuring the real-time processing of the system.
[0094] In this embodiment, by receiving external input images and user input operation interaction instructions, multiple input methods can be integrated to achieve multimodal input; the external input images and user input operation interaction instructions can be quickly transmitted to the subsequent ROI feature extraction device through the optical fiber channel, ensuring the real-time processing of the system, wherein the optical fiber channel can be responsible for high-speed data transmission, and its transmission delay is less than 5 milliseconds (ms), thereby ensuring the real-time processing of the system.
[0095] In one embodiment, Figure 3 As shown, in step S2, the external input image is processed based on the user's input operation interaction instruction of the visual basic model, and the preliminary segmented object area information is output. The object area information includes the pixel set of the object area and the feature vector of the object area. Specifically, it includes:
[0096] S21. Based on the operation interaction instruction input by the user, the ROI feature extraction device 20 of the feature extraction layer is called to obtain key features related to the target object from the spatially aligned RGB image input from the external input, wherein the operation interaction instruction includes: box selection, point selection or text description.
[0097] ROI, also known as Region of Interest in English or Region of Interest in Chinese, is a feature extraction device 20 for obtaining key features related to a target object from an external input spatially aligned RGB image based on an interactive operation instruction (box selection, point selection, or text description) input by a user.
[0098] Specifically, if the operation interaction instruction is a frame selection, the ROI feature extraction device 20 accurately locates the frame selection area in the image according to the coordinates of the frame selection area, and then uses a convolutional neural network to extract the feature vector of the image of the frame selection area, and extracts the feature vector including rich information such as object shape and object texture.
[0099] Convolutional Neural Networks (CNNs) are a type of feedforward neural network with a deep structure that incorporates convolutional computations. They are a representative algorithm for deep learning. Convolutional neural networks possess representation learning capabilities and can perform shift-invariant classification of input information based on their hierarchical structure. Hence, they are also known as shift-invariant artificial neural networks (SIANNs).
[0100] If the operation interaction instruction is a point selection, the ROI feature extraction device 20 selects an image area based on a pre-set neighborhood range with the point selected as the center, and then uses a convolutional neural network to extract the feature vector of the image in the image area, extracting a feature vector including rich information such as object shape and object texture.
[0101] If the operation interaction instruction is a text description, the ROI feature extraction device 20 first converts the text into a semantic vector through natural language processing technology, and then combines the image features to find the image area that matches the semantics, and then uses a convolutional neural network to extract the feature vector of the image in the image area, extracting a feature vector that includes rich information such as object shape and object texture.
[0102] Through the above operations based on the operation interaction instructions input by the user, the ROI feature extraction device 20 can obtain key features related to the target object from the user's operation interaction instructions and images.
[0103] Furthermore, the ROI feature extraction device 20 includes a SAM encoder 2 and a DINOv2 encoder 22. In step S21, the ROI feature extraction device 20 of the feature extraction layer is called based on the operation interaction instruction input by the user to obtain key features related to the target object from the spatially aligned RGB image input externally, including:
[0104] The SAM encoder 21 is called to encode the input spatially aligned RGB image, extract the spatial features of the RGB image, and output a 1027-dimensional feature vector. These 1027-dimensional feature vectors contain rich image spatial features, such as object contours, object position relationships, and other information.
[0105] The DINOv2 encoder 22 is called to further extract the geometric features of the RGB image and output a 768-dimensional feature vector. These 768-dimensional feature vectors focus on the geometric features of RGB, such as object edges, object angles, object textures, etc.
[0106] In the present invention, the SAM encoder and the DINOv2 encoder extract features from the RGB image from different angles respectively, providing multi-dimensional feature vector information for subsequent processing.
[0107] S22, call the dynamic feature fusion module 31 of the innovative processing layer to fuse the feature vectors output by the SAM encoder 21 and the DINOv2 encoder 22, and generate a fusion feature F∈R including geometric primitives describing the object structure 2048 .
[0108] Specifically, the dynamic feature fusion module 31 uses the attention gating mechanism and feature alignment technology to fuse the feature vectors output by the SAM encoder and the DINOv2 encoder to generate a fusion feature F∈R including geometric primitives describing the object structure. 2048 .
[0109] Among them, the attention gating mechanism is to weight different feature vectors according to their importance, highlighting the key features related to the target object.
[0110] The feature alignment technology is to align the feature vectors output by the SAM encoder 21 and the DINOv2 encoder 22 in space and dimension to ensure that the feature information can be effectively integrated.
[0111] In this embodiment, after the 1027-dimensional feature vector output by the SAM encoder and the 768-dimensional feature vector output by the DINOv2 encoder are fused through the above-mentioned attention gating mechanism and feature alignment technology, a fused feature F is generated. The fused feature F combines the advantages of the two encoders, the SAM encoder and the DINOv2 encoder. The fused feature F includes a 2048-dimensional comprehensive feature vector that fuses spatial and geometric features. The 2048-dimensional comprehensive feature vector includes generating geometric primitives that describe the structure of the object, thereby combining the semantic features of the text description input by the user, and dynamically adjusting the feature weights through the attention gating mechanism and feature alignment technology, so as to fuse the spatial features, geometric features and semantic feature information to generate a fused feature F including geometric primitives that describe the structure of the object, including more comprehensive image information, and enhancing the representation ability of the target object.
[0112] S23. Call the geometric decoupling module 32 of the innovative processing layer to parse the geometric primitive information of the object based on the fusion feature F. The geometric primitive information is used to describe the shape characteristics and structure of the object, providing a basis for subsequent pose estimation and 3D model generation.
[0113] S24: Invoke the 3D primitive generator 33 of the innovative processing layer to generate a 3D primitive model representing the object's 3D structure based on the geometric primitive information provided by the geometric decoupling module 32 and the object's spatial position and shape characteristics. The 3D primitive model is a preliminary representation of the object in 3D space and provides important basic data for subsequent generation of an accurate 3D model and pose estimation.
[0114] S25. Call the pose estimator 34 of the innovative processing layer to calculate the pose information of the object in the 3D space according to the 3D primitive model and the fusion feature F, and obtain the rotation matrix and cube parameters of the object through feature matching of basic geometric space transformations, so as to determine the position and posture of the object in the 3D space, including translation and rotation information, and output 6DoF (six degrees of freedom) parameters, where the 6DoF parameters are used to describe the complete pose state of the object in the 3D space.
[0115] 6DoF (Six degrees of freedom tracking) allows users to freely view content from any position and orientation within physical space. User movement can be captured by sensors or input controllers, supporting both spatial displacement and head posture changes.
[0116] In this embodiment, the pose information (such as translation information and rotation information) of the object in 3D space is separated by the pose estimator based on the 3D primitive model and the fusion feature F, and the accuracy and standardization of the above pose information are ensured through pose regularization.
[0117] Pose regularization refers to a technique for constraining or optimizing the position and orientation of an object in fields such as computer vision, robotics, and augmented reality (AR) / virtual reality (VR). Its purpose is to improve the robustness, smoothness, and physical rationality of pose estimation and avoid unreasonable pose output (such as jitter, mutation, or non-compliance with kinematic constraints).
[0118] Among them, position (Position) is the coordinate (x, y, z) in 3D space.
[0119] Orientation is represented by rotation matrix R∈SO(3), quaternion or Euler angles.
[0120] S26. Call the ZEM stabilizer 35 of the innovative processing layer to stabilize the processing process of the pose estimator 34 based on the zero embedding mapping mechanism ZEM(G), reduce fluctuations in training or inference, and output standardized 6DoF (six degrees of freedom) parameters, so that the position and posture of the object in 3D space can be accurately described based on the 6DoF parameters.
[0121] The ZEM stabilizer (Zero Effort Moment Stabilizer) is an active control device used to improve platform stability. Its core goal is to offset external disturbances (such as vibration and shaking) through real-time adjustments to ensure that the device remains horizontal or pointing stably.
[0122] Zero Embedding Mapping (ZEM), or Zero Embedding Mapping Mechanism (ZEM(G)), is a dimensionality reduction or feature embedding technique used in mathematics and computer science (particularly machine learning and graph representation learning) for processing high-dimensional or graph-structured data. Its core goal is to embed data into a low-dimensional space while preserving key structural information (such as adjacency relationships and topological properties) through a specific mapping method. It can also optimize the embedding process by introducing the concept of "zero" (such as the zero vector or null space).
[0123] The zero embedding mapping mechanism ZEM(G) is expressed as: ZEM(G) = δ(W0*G), where:
[0124] ZEM(G) is the zero embedding mapping function, which takes graph G as input and outputs an embedded representation.
[0125] W0 is a learnable weight matrix (or linear transformation matrix).
[0126] G is the input graph data.
[0127] δ(·) is a nonlinear activation function (e.g., ReLU, Sigmoid), and may also represent a sparsification function (e.g., threshold truncation, which makes some outputs zero).
[0128] ZEM(G)=δ(W0*G) is used to describe a graph data processing mechanism based on zero embedding mapping (ZEM), which generates an embedded representation of the graph through linear transformation (W0*G) and nonlinear mapping (δ), and may induce some embedding values to be zero through the sparsity of δ.
[0129] In this embodiment, the processing of the pose estimator is stabilized by a ZEM stabilizer based on a zero embedding mapping mechanism ZEM(G), thereby achieving stable training to reduce fluctuations in training or inference.
[0130] S27 , calling the feature generation module 41 of the output layer to generate a 3D bounding box of the object in 3D space based on the 6DoF parameters, and to generate an accurate instance mask based on the projection constraint and the fusion feature F.
[0131] Specifically, the feature generation module 41 generates a 3D bounding box of the object in the 3D space based on the 6DoF parameters, including: calculating and generating the 3D bounding box of the object in the 3D space based on the 6DoF parameters, and clarifying the spatial range and size of the object.
[0132] Specifically, the DINO model typically establishes an explicit mapping between the DINO features and 3D space, using a deep neural network to learn the complex relationship between DINO features and 3D spatial coordinates. DINO (DIstillation with NOlabels) is a visual feature extraction method based on self-supervised learning, proposed by the Meta AI team. Its core idea is to use self-distillation and contrastive learning to enable the model to learn highly semantic visual features without manually annotated data.
[0133] In this invention, the 6DoF parameters output by the pose estimator are used to convert 2D image features into 3D space coordinates using an explicit mapping relationship. The explicit mapping relationship represents the accurate projection of 2D features into 3D space, which can be expressed by the following formula:
[0134] P 3D =R·φ 2D +t, where R∈SO(3),
[0135] In the above formula:
[0136] P 3D is the 3D bounding box of the object in 3D space;
[0137] R∈SO(3) is the rotation matrix in 3D space, which is used to describe the rotation of an object in space;
[0138] t∈R3 is the translation vector, which determines the position of the object in space;
[0139] φ 2D is the 2D feature vector obtained from the DINO model. In the present invention, φ 2D 6DoF parameters.
[0140] By training the DINO model with a large amount of 3D annotated data, the model learns the corresponding positions of different 2D features in 3D space, thus achieving end-to-end conversion from 2D segmentation to 3D detection. This means that the 3D detection results of objects, including 3D bounding boxes and pose information, can be directly obtained from the input 2D image. This enables end-to-end conversion from 2D segmentation to 3D detection, achieving a breakthrough in geometric perception.
[0141] The feature generation module is also used to generate accurate instance masks based on the projection constraint and the fusion feature F, including: projecting the information of the 3D bounding box onto the 2D image plane to ensure the spatial consistency of the 3D information and the 2D information; and generating accurate instance masks based on the projection constraint and the fusion feature F, accurately segmenting the pixel area of the target object in the image, and realizing instance-level segmentation of objects on the 2D image.
[0142] Specifically, the projection constraint is the temporal consistency constraint L temporal , as shown below:
[0143] L temporal =∑‖α t -warp(α t-1 )‖ 2
[0144] Among them, α t is the mask of the object in the current frame image, α t-1 is the mask of the object in the previous frame image, warp(α t-1 ) is the mask α of the object in the previous frame image t-1 Position in the current frame.
[0145] In dynamic scenes, there is a certain degree of motion continuity between objects in adjacent frames. The temporal consistency constraint L temporal Using this feature, the mask α of the object in the previous frame image is calculated by the optical flow algorithm t-1 The position warp(α t-1 ), and compared with the mask α predicted by the model in the current frame t A comparison is made and the sum of the squares of the differences between the two is calculated as the loss function. During the model training process, the model parameters are continuously adjusted to minimize the loss function, thereby ensuring that the masks predicted by the model in dynamic scenes have temporal consistency, improving the segmentation stability in dynamic scenes, and improving the segmentation stability in dynamic scenes by 3.2 times, thereby achieving dynamic scene optimization.
[0146] In the present invention, the feature generation module projects the information of the 3D bounding box onto the 2D image plane to ensure the spatial consistency of the 3D information and the 2D information, and based on the projection constraint (i.e., the temporal consistency constraint L temporal ) is combined with the fusion feature F to generate an accurate instance mask, accurately segmenting the pixel area of the target object in the image and realizing object instance-level segmentation on the 2D image.
[0147] S28. Call the output interface 42 of the output layer to take the 6DoF parameters obtained by the pose estimator 34 and the instance mask generated by the feature generation module 41 as the final output. The 6DoF parameters and instance mask also become the key features related to the target object from the spatially aligned RGB image. The instance mask is used to accurately segment the object instance in the image and clarify the boundary between the object and the background. The 6DoF message provides an accurate description of the positioning and posture of the object in the 3D space. The 6DoF parameters and instance mask as the final output can be directly applied to actual scenarios such as AR content merging, 3D printing model generation, and autonomous driving scene reconstruction.
[0148] S29. Based on the key features related to the target object output by the output layer, the region growing algorithm is used to expand the growing region and output the preliminary segmented object region information.
[0149] The region growth algorithm is a segmentation algorithm based on image region features. It uses a seed point (a frame selection area, a point selection location, or the center of a semantic matching area) determined by the ROI feature extraction device as the starting point. Based on pre-set similarity criteria (such as color similarity, texture similarity, etc.), adjacent pixels with similar features to the seed point are gradually merged into the growth region. For example, the color difference between the adjacent pixel and the seed point is calculated. If the difference is less than a preset threshold, the adjacent pixel is included in the growth region.
[0150] When the output layer outputs the key features related to the target object, the region growing algorithm is automatically activated according to the output key feature information. The region growing algorithm will continue to expand the growing area until the stopping condition is met (for example, the area no longer grows, the feature change of the growing area is less than the preset threshold, etc.), and output the preliminary segmented object area information.
[0151] In this embodiment, after processing by the feature extraction layer and the innovative processing layer and / or the region growing algorithm, preliminary segmented object region information is output, including the pixel set of the object region, the feature vector of the object region, etc. This preliminary segmented object region information will be used for subsequent generation of the alpha channel mask and further processing. In this process, the output preliminary segmented object region information not only includes the object's position information and object shape information, but also includes a description of the object's features, providing a basis for subsequent precise segmentation and processing. This can achieve end-to-end conversion from 2D segmentation to 3D detection, that is, directly obtaining the 3D detection result of the object from the input 2D image, achieving a breakthrough in geometric perception, and improving segmentation stability.
[0152] In one embodiment, in step S3, an alpha channel mask is generated based on the initially segmented object region information, specifically comprising:
[0153] S31. Binarize the initially segmented object region to separate the object region from the background, and obtain a binary mask.
[0154] S32. Smoothing the binary mask according to the 3D bounding box of the object region to make the boundary transition of the object region more natural. For example, a Gaussian filter method can be used to smooth the binary mask to obtain a boundary transition of the object region with a more natural transition.
[0155] S33. In order to achieve a better fusion effect between the generated mask and the original image, different transparency values are assigned to the smoothed binary mask according to the depth information of the object area or according to a preset weight distribution rule to generate an alpha channel mask, wherein each transparency value in the alpha channel mask represents the transparency of the object at the pixel.
[0156] In this embodiment, by binarizing the initially segmented object area, the object area can be separated from the background area, and a binary mask can be obtained. The binary mask is then smoothed to make the boundary transition of the object area more natural. Different transparency values are assigned to the smoothed binary mask based on the depth information of the object area or based on a preset weight distribution rule to generate an alpha channel mask, so that the generated mask can be better fused with the original image.
[0157] In one embodiment, in step S4, an edge optimization algorithm is applied to adjust the edge pixel values of the alpha channel mask, and a PNG format image with a transparent channel is output.
[0158] PNG (Portable Network Graphics) is a bitmap format that uses a lossless compression algorithm and supports indexed, grayscale, and RGB color schemes as well as alpha channels.
[0159] The edge optimization algorithm operates based on the boundary energy function E(α), which is as follows:
[0160]
[0161] Among them, λ1 is the smoothing weight, which is used to control the smoothness of the alpha channel mask edge to avoid excessive jagged edges; λ2 is the data fidelity term, which ensures that the generated alpha channel mask is as close as possible to the original segmentation result; λ3 is the edge constraint term, which makes the alpha channel mask edge more closely fit the actual edge of the object.
[0162] The edge optimization algorithm iteratively adjusts the edge pixel values of the alpha channel mask to minimize the edge energy function E(α). In each iteration, the gradient of the edge energy function is calculated based on the edge of the current alpha channel mask. The alpha channel mask edge pixel values are then adjusted in the direction of gradient descent. After multiple iterations, the alpha channel mask edge is optimized, improving the accuracy and visual quality of object segmentation. Based on the optimized alpha channel mask edge, a PNG image with a transparent channel is output. This PNG image with a transparent channel includes 6DoF parameters and instance masks.
[0163] In the present invention, by setting λ1=0.3, λ2=0.7, and λ3=1.2, an optimized alpha channel mask edge can be obtained, thereby improving the accuracy and visual effect of object segmentation; and based on the optimized alpha channel mask edge, a PNG format image with a transparent channel is output, wherein the PNG format image with a transparent channel includes 6DoF parameters and an instance mask.
[0164] The present invention provides an AI-based open-world 3D object detection and segmentation method and / or an AI-based open-world 3D object detection system, which can be applied to practical scenarios such as AR content merging, 3D printing model generation, and autonomous driving scene reconstruction.
[0165] The following are several application scenarios of the AI-based open-world 3D object detection and segmentation method and / or AI-based open-world 3D object detection system of the present invention.
[0166] Scenario 1: In-vehicle scenario.
[0167] In a vehicle-mounted scenario, new obstacles are detected. A dashcam video frame (resolution 1920×1080, focal length f=6mm) is input, and the user selects an irregularly shaped obstacle on the right side of the image (frame coordinates x1y1=1240x360, x2y2=1480x720).
[0168] The present invention provides an AI-based open-world 3D object detection and segmentation method and / or an AI-based open-world 3D object detection system, the output of which is:
[0169] 3D bounding box parameters:
[0170] Center coordinates: x = 4.2m, y = 1.8m, z = 32.7m (vehicle coordinate system);
[0171] Dimensions: w = 0.8m, h = 1.2m, l = 0.5m;
[0172] Rotation angle: pitch = 5°, yaw = 12°.
[0173] Processing time: 48ms (NVIDIA Orin platform).
[0174] Based on the output results, the present invention provides an AI-based open-world 3D object detection and segmentation method and / or an AI-based open-world 3D object detection system, which are compared with MonoFlex. The comparison results are shown in the following table:
[0175] index The present invention MonoFlex Improvement Depth error (m) 0.32 1.15 72%↓ Yaw angle error (degrees) 3.1 8.7 64%↓ Unknown object detection rate 89% 41% 117%↑
[0176] Scenario 2: AR application scenario.
[0177] In an AR application, furniture dimensions are measured. A camera captures a wide-angle photo of the living room (4032×3024, no internal parameters provided) as input. Multiple prompts are then used for interaction: the user taps the sofa armrest (screen coordinates (1560, 2200)) and voice input is: "measure this sectional sofa."
[0178] The present invention provides an AI-based open-world 3D object detection and segmentation method and / or an AI-based open-world 3D object detection system, the output of which is shown as follows:
[0179] 3D wireframe superimposed on AR view (error < 2cm);
[0180] Dimensions: total length 2.8m / seat depth 0.9m / back height 0.75m;
[0181] Material identification: Textile (92% confidence).
[0182] In terms of performance data: end-to-end latency: 1.2 seconds on iPhone 14Pro; memory usage: peak 1.4GB (including DINOv2 feature cache).
[0183] Scenario 3: Industrial inspection scenario.
[0184] In industrial inspection scenarios, locating irregular-shaped parts is done using an assembly line monitoring camera (global shutter, 2 megapixels) to identify non-standard gear parts and implement improvements to address metal reflections.
[0185] The present invention provides an AI-based open-world 3D object detection and segmentation method and / or an AI-based open-world 3D object detection system, the output verification of which is:
[0186] Pose estimation results: translation error: ±0.12mm (RMS), rotation error: ±0.5° (compared with laser scanning).
[0187] The anti-interference performance is shown in the following table:
[0188] Interference Type Successful detection rate Oil stain blocking 30% 87% Strong reflective areas 79% Motion Blur 92%
[0189] Scenario 4: Medical scenario.
[0190] In a medical scenario, surgical instruments are tracked. The following configuration is input: Data: Laparoscopic video (1280×720@30fps, f=3.4mm), Constraint: Latency must be <50ms (real-time requirement). Configure a lightweight solution and perform temporal consistency processing.
[0191] The present invention provides an AI-based open-world 3D object detection and segmentation method and / or an AI-based open-world 3D object detection system, which outputs clinical validation data as shown in the following table:
[0192] Device type End-to-end delay (ms) Position error (mm) Electric coagulation hook 43 0.8±0.3 Ultrasonic scalpel 39 1.2±0.5 needle holder 47 0.6±0.2
[0193] Based on the same concept, the present invention also provides an electronic device, such as Figure 4 As shown, the electronic device 900 includes: a memory 902, a processor 901, and one or more computer programs stored in the memory 902 and executable on the processor 901. The memory 902 and the processor 901 are coupled together via a bus system 903. When the one or more computer programs are executed by the processor 901, the steps of an AI-based open-world three-dimensional object detection and segmentation method provided in an embodiment of the present invention are implemented.
[0194] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 901. Processor 901 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be performed by hardware integrated logic circuits or software instructions within processor 901. Processor 901 can be a general-purpose processor, a DSP, or other programmable logic device, a discrete gate or transistor logic device, or discrete hardware components. Processor 901 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present invention can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules within the decoding processor. The software modules can be located in a storage medium located in memory 902. Processor 901 reads information from memory 902 and, in conjunction with its hardware, completes the steps of the above method.
[0195] It can be understood that the memory 902 in the embodiment of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device; the volatile memory can be random access memory (RAM), by way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), etc. Memory), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), Direct Rambus Random Access Memory (DRRAM).The memories described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memories.
[0196] In the present invention, the electronic device 900 may be any device equipped with a processor and having processing capabilities, such as a smart phone, a tablet computer, a PDA, a laptop computer, a server, a workstation, or other electronic device equipped with a processor.
[0197] It should be noted that the above-mentioned electronic device embodiment and method embodiment belong to the same concept, and their specific implementation process is detailed in the method embodiment, and the technical features in the method embodiment are applicable to the electronic device embodiment, which will not be repeated here.
[0198] In addition, in an exemplary embodiment, an embodiment of the present invention further provides a storage medium, specifically a computer-readable storage medium, for example, including a memory 902 for storing computer programs, and the computer storage medium stores one or more programs of an AI-based open-world three-dimensional object detection and segmentation method. When the one or more programs of the AI-based open-world three-dimensional object detection and segmentation method are executed by the processor 901, the steps of an AI-based open-world three-dimensional object detection and segmentation method provided in an embodiment of the present invention are implemented.
[0199] It should be noted that the AI-based open-world three-dimensional object detection and segmentation method program embodiment on the above-mentioned computer-readable storage medium and the method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment, and the technical features in the method embodiment are correspondingly applicable in the embodiment of the above-mentioned computer-readable storage medium, which will not be repeated here.
[0200] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Under the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present invention as described above. For the sake of simplicity, they are not provided in detail. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in this field should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An AI-based open-world 3D object detection and segmentation method, characterized in that: include: Receive external input images and user input operation interaction instructions; Based on the visual basic model, the user's input operation interaction instructions are processed to the external input image, and the preliminary segmented object area information is output; Generate alpha channel mask based on the object area information of the preliminary segmentation; Apply an edge optimization algorithm to adjust the edge pixel values of the alpha channel mask and output a PNG format image with a transparent channel.
2. The AI-based open-world 3D object detection and segmentation method according to claim 1, characterized in that: The visual base model includes an input layer, and the input layer includes a multimodal input module; The receiving of an external input image and an operation interaction instruction input by a user includes: Calling the multimodal input module to receive external input images and user input operation interaction instructions; Spatially align the external input RGB images and adjust the input RGB images to a uniform size.
3. The AI-based open-world 3D object detection and segmentation method according to claim 1, characterized in that: The visual basic model further comprises a feature extraction layer, wherein the feature extraction layer comprises an ROI feature extraction device; The method processes the external input image based on the user's input operation interaction instruction based on the visual basic model and outputs the preliminary segmented object area information, including: Based on the operation interaction instruction input by the user, the ROI feature extraction device is called to obtain key features related to the target object from the spatially aligned RGB image input externally; Based on the key features related to the target object, the region growing algorithm is used to expand the growing region and output the preliminary segmented object region information.
4. The AI-based open-world 3D object detection and segmentation method according to claim 3, characterized in that: The ROI feature extraction device includes a SAM encoder and a DINOv2 encoder; The operation interaction instruction based on the user input calls the ROI feature extraction device to obtain key features related to the target object from the spatially aligned RGB image input externally, including: Call the SAM encoder to encode the input spatially aligned RGB image, extract the spatial features of the RGB image, and output a 1027-dimensional feature vector; Call the DINOv2 encoder to further extract the geometric features of the RGB image and output a 768-dimensional feature vector; The dynamic feature fusion module fuses the 1027-dimensional feature vector output by the SAM encoder and the 768-dimensional feature vector output by the DINOv2 encoder to generate a fused feature including geometric primitives that describe the structure of the object.
5. The AI-based open-world 3D object detection and segmentation method according to claim 4, characterized in that: The visual basic model also includes an innovative processing layer, which includes a dynamic feature fusion module; The method processes an external input image based on the user's input operation interaction instruction based on the visual basic model and outputs preliminary segmented object area information, and further includes: The dynamic feature fusion module is called to fuse the 1027-dimensional feature vector output by the SAM encoder and the 768-dimensional feature vector output by the DINOv2 encoder to generate a fused feature including geometric primitives describing the structure of the object.
6. The AI-based open-world 3D object detection and segmentation method according to claim 5, characterized in that: The innovative processing layer also includes a geometric decoupling module; The method processes an external input image based on the user's input operation interaction instruction based on the visual basic model and outputs preliminary segmented object area information, and further includes: The geometric decoupling module is called to parse the geometric primitive information of the object according to the fusion features.
7. The AI-based open-world 3D object detection and segmentation method according to claim 6, characterized in that: The innovative processing layer also includes a 3D primitive generator; The method processes an external input image based on the user's input operation interaction instruction based on the visual basic model and outputs preliminary segmented object area information, and further includes: The 3D primitive generator is called to generate a 3D primitive model representing the 3D structure of the object based on the geometric primitive information provided by the geometric decoupling module and in combination with the spatial position and shape characteristics of the object.
8. The AI-based open-world 3D object detection and segmentation method according to claim 7, characterized in that: The innovative processing layer also includes a pose estimator; The method processes an external input image based on the user's input operation interaction instruction based on the visual basic model and outputs preliminary segmented object area information, and further includes: The pose estimator is called to calculate the pose information of the object in 3D space according to the 3D primitive model and fusion features, obtain the rotation matrix and cube parameters of the object through feature matching of basic geometric space transformation, and output 6DoF parameters.
9. The AI-based open-world 3D object detection and segmentation method according to claim 8, characterized in that: The innovative treatment layer also includes a ZEM stabilizer; The method processes an external input image based on the user's input operation interaction instruction based on the visual basic model and outputs preliminary segmented object area information, and further includes: The ZEM stabilizer is called to stabilize the processing of the pose estimator based on the zero embedding mapping mechanism and output standardized 6DoF parameters.
10. The AI-based open-world 3D object detection and segmentation method according to claim 8 or 9, characterized in that: The visual basic model also includes an output layer, which includes a feature generation module and an output interface; The method processes an external input image based on the user's input operation interaction instruction based on the visual basic model and outputs preliminary segmented object area information, and further includes: Calling the feature generation module to generate a 3D bounding box of the object in 3D space based on the 6DoF parameters, and generating an accurate instance mask based on the projection constraint and the fused features; The output interface is called to output the 6DoF parameters and the instance mask as final output, wherein the 6DoF parameters and the instance mask become key features related to the target object obtained from the spatially aligned RGB image.
11. The AI-based open-world 3D object detection and segmentation method according to claim 10, characterized in that: The method processes an external input image based on the user's input operation interaction instruction based on the visual basic model and outputs preliminary segmented object area information, and further includes: Based on the key features related to the target object output by the output layer, a region growing algorithm is used to expand the growing region and output preliminary segmented object region information.
12. The AI-based open-world 3D object detection and segmentation method according to claim 11, characterized in that: The generating of the alpha channel mask based on the object region information of the preliminary segmentation includes: The object area that is initially segmented is binarized to separate the object area from the background and obtain a binary mask: Smoothing the binary mask based on the 3D bounding box of the object area; According to the depth information of the object area or according to a certain weight distribution rule, different transparency values are assigned to the smoothed binary mask to generate an alpha channel mask.
13. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the computer program is executed by the processor, the open world three-dimensional object detection and segmentation method based on AI according to any one of claims 1 to 12 is implemented.
14. A storage medium, characterized in that The storage medium stores a program for an AI-based open-world three-dimensional object detection and segmentation method. When the program for the AI-based open-world three-dimensional object detection and segmentation method is executed by a processor, the AI-based open-world three-dimensional object detection and segmentation method according to any one of claims 1 to 12 is implemented.