A method and system for detecting and identifying an unknown object, a terminal device and a medium
Patent Information
- Application Number
- CN202611097548.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-23
AI Technical Summary
[0005]本发明要解决的技术问题在于,在开放环境下的目标识别领域,现有方法依赖单帧RGB图像且分割、识别与增量学习彼此割裂,无法综合利用深度几何信息、时序信息及大模型语义反馈,导致难以形成从发现未知目标、理解目标语义到更新识别能力的完整流程
[0016]有益效果:本发明公开一种未知物体的检测与识别方法、系统、终端设备及介质,涉及目标识别技术领域。方法首先获取与待处理RGB视频片段同步的深度图像,对所述深度图像进行预处理和聚类,基于聚类结果生成空间提示掩码。其后,将所述空间提示掩码与所述RGB视频片段共同输入视频分割模型,得到时序分割结果,其中,所述空间提示掩码用于引导关注区域,所述RGB视频片段的时序信息用于跨帧传播。随后,从所述RGB视频片段中选取样本帧,在所述时序分割结果中所述样本帧对应的分割掩码区域内进行实例分割,得到包含多个目标实例的目标实例集合。接着,提取所述目标实例集合中各目标实例的视觉特征,将所述视觉特征输入可训练分类器进行类别识别,并将满足预设未知判定条件的目标实例确定为未知类别目标。然后,通过多模态大模型对所述未知类别目标进行结构化语义识别,得到结构化目标属性信息,并将所述结构化目标属性信息与所述目标实例集合进行匹配,构建带语义标签的训练样本集合。最后,基于所述训练样本集合对所述可训练分类器进行增量训练,获得更新后的分类器,以使所述可训练分类器具备对所述未知类别目标的再识别能力。
Smart Images

Figure CN122618356B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target recognition technology, and in particular to a method, system, terminal device and medium for detecting and recognizing unknown objects. Background Technology
[0002] Existing object detection and recognition technologies typically rely on predefined sets of categories to locate and identify targets in input images using object detection, image classification, or segmentation networks. While these approaches achieve good recognition results in closed-category tasks with uniform training data distribution, they struggle in open environments such as robot perception, unmanned system environmental understanding, intelligent monitoring, and industrial recognition. These systems frequently encounter objects not present in the training set, making it difficult for existing solutions to effectively address these challenges.
[0003] The current methods for handling the problem of unknown object recognition mainly include: determining the unknown category based on a low confidence threshold, performing open vocabulary reasoning based on a visual language model, retraining the classification model based on incremental learning, and extracting the potential target region based on segmentation before recognition. However, the above methods have the following shortcomings: (1) They mainly rely on single-frame RGB images and lack the use of scene geometric depth information, making it difficult to accurately locate potential targets in complex backgrounds, occlusions, or low contrast conditions; (2) They can only complete the unknown determination and cannot output structured semantic attribute information, making it difficult to form new category samples that can be used for subsequent learning; (3) The segmentation, recognition, and incremental learning stages are separated from each other, making it impossible to form a complete process from discovering the unknown, understanding the unknown to learning the unknown; (4) Although large models can output rich semantics, their results lack a stable automatic correspondence with pixel-level target regions, making it difficult to convert them into supervised training samples; (5) For video scenes, processing only a single frame can easily lead to problems such as incomplete target regions and large instantaneous noise interference; (6) New category learning relies on large-scale model overall retraining, which is costly and inefficient.
[0004] Therefore, there is an urgent need for a method that comprehensively utilizes depth information, RGB temporal information, segmentation models, and semantic feedback from large models to continuously update the ability to detect and recognize unknown objects, in order to fill the gaps in existing technologies. Summary of the Invention
[0005] The technical problem this invention aims to solve is that, in the field of target recognition in open environments, existing methods rely on single-frame RGB images, and segmentation, recognition, and incremental learning are isolated from each other. This makes it impossible to comprehensively utilize depth geometric information, temporal information, and semantic feedback from large models, resulting in a difficulty in forming a complete process from discovering unknown targets, understanding target semantics, to updating recognition capabilities. Therefore, an effective solution is urgently needed to address the aforementioned technical problems.
[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a method for detecting and identifying unknown objects, the method comprising: Acquire a depth image synchronized with the RGB video clip to be processed, preprocess and cluster the depth image, and generate a spatial cue mask based on the clustering results; The spatial cue mask and the RGB video clip are input into the video segmentation model to obtain the temporal segmentation result. The spatial cue mask is used to guide the region of interest, and the temporal information of the RGB video clip is used for cross-frame propagation. Sample frames are selected from the RGB video segments, and instance segmentation is performed within the segmentation mask region corresponding to the sample frames in the temporal segmentation result to obtain a target instance set containing multiple target instances. Visual features of each target instance in the target instance set are extracted, and the visual features are input into a trainable classifier for category recognition. Target instances that meet the preset unknown determination conditions are identified as unknown category targets. The unknown category target is subjected to structured semantic recognition by a multimodal large model to obtain structured target attribute information, and the structured target attribute information is matched with the target instance set to construct a training sample set with semantic labels; The trainable classifier is incrementally trained based on the training sample set to obtain an updated classifier, so that the trainable classifier has the ability to re-identify the unknown category target.
[0007] In one implementation, the preprocessing and clustering of the depth image, and the generation of a spatial cue mask based on the clustering results, includes: The depth image is denoised, and the denoised depth data is filtered by distance range to remove pixels that exceed a preset distance threshold, thereby obtaining an effective depth region. The pixels within the effective depth region are clustered based on depth values to divide the scene of the depth image into different depth levels; Select the pixel region corresponding to the preset depth level of the depth image to generate the spatial cue mask.
[0008] In one implementation, the step of inputting the spatial cue mask and the RGB video segment into a video segmentation model to obtain a temporal segmentation result includes: The spatial cue mask is used as the initial frame spatial cue for the video segmentation model. The video segmentation model is then propagated through each frame of the RGB video segment to optimize the target segmentation region, resulting in a temporal segmentation result. The video segmentation model utilizes historical frame information from adjacent video segments to assist in segmentation prediction for the first few frames of the current segment when segmenting the current RGB video segment, so as to achieve a smooth transition between segmentation results between adjacent segments.
[0009] In one implementation, selecting sample frames from the RGB video segment and performing instance segmentation within the segmentation mask region corresponding to the sample frames in the temporal segmentation result to obtain a target instance set containing multiple target instances includes: Based on preset image quality evaluation indicators, the highest quality frame is selected from the RGB video clip as the sample frame; Using the segmentation mask region corresponding to the sample frame in the temporal segmentation result as a constraint range, full segmentation is performed within the constraint range to obtain the full segmentation result; The target instances that intersect with the constraint range in the full segmentation result are included in the target instance set.
[0010] In one implementation, the step of extracting visual features from each target instance in the target instance set, inputting the visual features into a trainable classifier for category recognition, and identifying target instances that meet preset unknown determination conditions as unknown category targets includes: The image regions corresponding to each target instance are input into the feature extraction model to obtain the visual feature vectors of the target instances; The visual feature vector is input into the trainable classifier to obtain the category prediction result; When the category prediction result meets the preset unknown determination condition, the corresponding target instance is determined as an unknown category target.
[0011] In one implementation, the step of performing structured semantic recognition on the unknown category target using a multimodal large model to obtain structured target attribute information, and matching the structured target attribute information with the target instance set to construct a training sample set with semantic labels, includes: The image region containing the unknown category target and the corresponding effective segmentation range information are sent to the multimodal large model; Based on a preset structured data format, the multimodal large model returns attribute information of the unknown category target, wherein the attribute information includes at least category information, location information, and confidence information; Based on the location information in the structured target attribute information and the location information of each target instance in the target instance set, the spatial matching degree between the two is calculated; Establish a correspondence between target pairs whose spatial matching degree meets the preset matching conditions; The training sample set is constructed based on the target instances with which the correspondence is established and their corresponding semantic category labels.
[0012] In one implementation, incrementally training the trainable classifier based on the training sample set to obtain an updated classifier, so that the trainable classifier has the ability to re-identify the unknown category target, includes: Using the output distribution of the trainable classifier on old category samples as a constraint, the trainable classifier is incrementally trained based on samples in the training sample set to update the parameters of the trainable classifier and obtain the updated classifier. The weights of each sample during training are determined based on the confidence level of the target output corresponding to that sample in the multimodal large model.
[0013] Secondly, embodiments of the present invention also provide a system for detecting and identifying unknown objects, the system comprising: The depth cue generation module is used to acquire a depth image synchronized with the RGB video clip to be processed, preprocess and cluster the depth image, and generate a spatial cue mask based on the clustering results; The temporal segmentation module is used to input the spatial cue mask and the RGB video clip into the video segmentation model to obtain the temporal segmentation result. The spatial cue mask is used to guide the region of interest, and the temporal information of the RGB video clip is used for cross-frame propagation. The instance splitting module is used to select sample frames from the RGB video segment and perform instance splitting within the segmentation mask area corresponding to the sample frames in the temporal segmentation result to obtain a target instance set containing multiple target instances. The classification and recognition module is used to extract the visual features of each target instance in the target instance set, input the visual features into a trainable classifier for category recognition, and determine the target instance that meets the preset unknown judgment condition as an unknown category target; The sample construction module is used to perform structured semantic recognition on the unknown category target through a multimodal large model, obtain structured target attribute information, and match the structured target attribute information with the target instance set to construct a training sample set with semantic labels; The incremental training module is used to incrementally train the trainable classifier based on the training sample set to obtain an updated classifier, so that the trainable classifier has the ability to re-identify the unknown category target.
[0014] Thirdly, embodiments of the present invention also provide a terminal device, the terminal device including a memory, a processor, and an unknown object detection and identification program stored in the memory and executable on the processor, wherein when the processor executes the unknown object detection and identification program, it implements the steps of the unknown object detection and identification method described in any of the above schemes.
[0015] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a detection and identification program for an unknown object. When the unknown object detection and identification program is executed by a processor, it implements the steps of the unknown object detection and identification method described in any of the above schemes.
[0016] Beneficial Effects: This invention discloses a method, system, terminal device, and medium for detecting and recognizing unknown objects, relating to the field of target recognition technology. The method first acquires a depth image synchronized with an RGB video clip to be processed. The depth image is preprocessed and clustered, and a spatial cue mask is generated based on the clustering results. Subsequently, the spatial cue mask and the RGB video clip are input into a video segmentation model to obtain a temporal segmentation result. The spatial cue mask is used to guide the region of interest, and the temporal information of the RGB video clip is used for cross-frame propagation. Then, sample frames are selected from the RGB video clip, and instance segmentation is performed within the segmentation mask region corresponding to the sample frames in the temporal segmentation result to obtain a target instance set containing multiple target instances. Next, the visual features of each target instance in the target instance set are extracted, and the visual features are input into a trainable classifier for category recognition. Target instances that meet preset unknown determination conditions are identified as unknown category targets. Then, a multimodal large model is used to perform structured semantic recognition on the unknown category target to obtain structured target attribute information. This structured target attribute information is then matched with the target instance set to construct a training sample set with semantic labels. Finally, the trainable classifier is incrementally trained based on the training sample set to obtain an updated classifier, enabling the trainable classifier to re-recognize the unknown category target.
[0017] This invention constructs an automated closed loop from discovering the unknown to learning to recognize it by using deep prompting-guided segmentation, temporal information-optimized masking, and automatic alignment of large model semantics and pixel segmentation. This enables the system to autonomously complete continuous learning from unknown object detection and semantic understanding to classifier updates in open environments without human intervention, thereby enhancing the system's adaptability to unfamiliar scenes and its recognition robustness. Attached Figure Description
[0018] Figure 1 A flowchart illustrating a specific implementation of the method for detecting and identifying unknown objects provided in this invention.
[0019] Figure 2 This is a schematic diagram of the overall process of the method for detecting and identifying unknown objects provided in an embodiment of the present invention.
[0020] Figure 3 These are the front and side views of the visual data acquisition module.
[0021] Figure 4 These are the acquired depth images and RGB video image frames.
[0022] Figure 5 This is a schematic diagram showing the original depth image and the depth image after clustering to generate a spatial cue mask.
[0023] Figure 6 This is a schematic diagram of a relatively smooth mask region obtained after temporal image segmentation.
[0024] Figure 7 This is a schematic diagram of the target instance set obtained after intersection processing.
[0025] Figure 8 This is a schematic diagram of the category prediction results of the target instance obtained after processing by a trainable classifier.
[0026] Figure 9 This is a schematic diagram of the results after annotation of a multimodal large model.
[0027] Figure 10 This is a schematic diagram showing the class prediction results of the target instance after processing by the classifier trained with small sample increments.
[0028] Figure 11 This is a principle block diagram of the unknown object detection and identification system provided in the embodiments of the present invention.
[0029] Figure 12 This is a block diagram illustrating the internal structure of the terminal device provided in an embodiment of the present invention. Detailed Implementation
[0030] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0031] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content, operations, or steps, nor does it require execution in the described order. For example, some operations or steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0032] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0033] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, the terms "first" and "second" are used in the embodiments of the present invention to distinguish identical or similar items with essentially the same function and effect. For example, "first control information" and "second control information" are only used to distinguish different control information and do not limit their order.
[0034] Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or the order of execution, and that the words "first" and "second" do not necessarily imply that they are different.
[0035] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0036] Existing object detection and recognition technologies typically rely on predefined sets of categories and use object detection networks, image classification networks, or segmentation networks to locate and identify targets in input images. For common closed-category tasks, these approaches can achieve good recognition results when the training data distribution is consistent. However, in complex and open environments, especially in robot perception, unmanned system environmental understanding, intelligent monitoring, and diverse industrial site recognition scenarios, systems often encounter objects that have not appeared in the training set. Therefore, enabling these systems to autonomously learn to discover and recognize unknown objects can significantly enhance their robustness in complex scenarios and improve their adaptability to various situations.
[0037] Existing technologies typically address this type of problem from the following aspects: (1) Threshold-based methods, which directly classify low-confidence samples as unknown categories; (2) Open-vocabulary recognition-based methods, which use visual language models to perform open-category reasoning on targets; (3) Incremental learning-based methods, which retrain or fine-tune classification models after obtaining new category data; (4) Segmentation-based methods, which first separate potential target regions in the scene and then perform subsequent recognition.
[0038] However, these methods still have certain shortcomings: (1) Most methods mainly rely on RGB single-frame images for target recognition, lacking effective utilization of scene geometric depth information, resulting in difficulty in accurately locating potential target areas under complex backgrounds, occlusion, and low contrast conditions; (2) Existing unknown category processing methods often can only complete unknown determination, but cannot further provide more structured semantic attribute information, thus making it difficult to form new category samples that can be directly used for subsequent learning; (3) Many schemes process segmentation, recognition, and incremental learning separately, failing to form a process from discovering unknown targets to obtaining new category knowledge and then updating their own recognition capabilities. Complete closed loop; (4) Although open recognition based on large models can output rich semantic information, its output often lacks a stable correspondence with the pixel-level target region, making it difficult to automatically align with existing segmentation results, and thus difficult to use for supervised learning of subsequent trainable models; (5) For video scene recognition tasks, if only single frame is used for target discovery, problems such as incomplete target region, large instantaneous noise interference, and unstable candidate targets are likely to occur, affecting subsequent classification and learning effects; (6) For learning new categories, if large-scale model retraining is relied upon each time, it is not only costly and inefficient, but also not conducive to real-time deployment in engineering.
[0039] The main reasons for the above problems are: (1) Existing methods do not use depth information to generate high-quality spatial cues and lack a front-end cues mechanism that can quickly narrow down the candidate region from a geometrical perspective; (2) Existing methods do not make full use of temporal information and cannot make full use of the stability of the continuous appearance of targets in video clips to obtain more reliable initial target masks; (3) Existing open category recognition results usually remain at the text level and lack an automatic alignment mechanism with pixel segmentation results; (4) Most existing incremental learning mechanisms rely directly on manually labeled data and lack a technical path that uses a large model to automatically construct small sample supervision data and update a lightweight classification head; (5) Closed category classification heads lack adaptive expansion capabilities for unknown targets, which means that even if the system can discover unknown targets, it cannot gradually learn to recognize the targets.
[0040] This embodiment provides a method for detecting and identifying unknown objects, such as... Figure 1 As shown, the specific steps include the following: Step S100: Obtain a depth image synchronized with the RGB video clip to be processed, preprocess and cluster the depth image, and generate a spatial cue mask based on the clustering results.
[0041] In this embodiment, a depth image refers to a two-dimensional image synchronously acquired by a depth camera, where each pixel value represents the distance from the corresponding real-scene point to the camera's imaging plane. The depth image can have a resolution of 1280×720 and a data type of uint16 (unsigned 16-bit integer). Sources for acquiring depth images include, but are not limited to, structured light depth cameras, time-of-flight depth cameras, binocular parallax calculation, LiDAR range map projection, or multi-view reconstruction. Simultaneously acquired with the depth image is an RGB video segment to be processed. The RGB video segment contains multiple consecutive frames of RGB images, for example, 30 frames, with the starting frame timed with the depth image, and a frame rate of 30 frames per second. The RGB image resolution can also be 1280×720, and the data type is uint8. The RGB video segment can also be replaced with video input from other imaging modalities, such as infrared video, thermal imaging video, grayscale video, multispectral video, or video fused from RGB and infrared multimodal modes. The number of frames in the video segment can be adjusted to an arbitrary length of continuous frame sequence according to scene complexity and real-time requirements, or only a few keyframes can be used as input.
[0042] Preprocessing refers to a set of operations that suppress noise and remove invalid regions from the original depth image. Due to factors such as changes in ambient light, differences in surface reflection from objects, and the sensor's own thermal noise during acquisition, the original depth image often contains random noise and local holes, thus requiring initial denoising. After denoising, the depth data needs to be filtered by distance range, retaining only pixels with depth values less than or equal to a preset distance threshold as valid depth regions. Pixels exceeding this threshold are considered background or invalid regions and are not included in subsequent processing. Objects that are too far away have a small imaging area and blurred details in RGB images, and are usually not the focus of current operations in practical applications.
[0043] Clustering refers to the operation of grouping pixels within an effective depth region according to the similarity of their depth values. Through clustering, a scene is divided into different levels along the depth dimension, such as near-field, mid-field, and far-field regions. In practice, the K-Means clustering method can be used, setting the number of cluster centers to 3, corresponding to the near, mid, and far depth categories respectively.
[0044] A spatial cue mask is a binary image of the same size as the original depth image, generated based on clustering results. Pixel values within the cue region are marked as valid, such as 1 or white, while the rest of the region is marked as invalid, such as 0 or black.
[0045] The specific selection strategies for the prompt area include: (1) selecting only the nearby clusters as the prompt area; and (2) selecting a combination of nearby and mid-range clusters as the prompt area. In this embodiment, the strategy of selecting only the pixel area corresponding to the nearby clusters is adopted.
[0046] The role of spatial cue masks is to constrain the initial focus of subsequent video segmentation models, prioritizing segmentation in areas with closer geometric distances and clearer object outlines, thereby shielding distant, blurred background areas and reducing interference from complex backgrounds in the segmentation process. This mechanism transforms depth geometric information into spatial attention guidance, automatically extracting prior information highly correlated with the spatial distribution of potential objects from underlying sensor data, improving the targeting and efficiency of target discovery.
[0047] In one implementation, the preprocessing and clustering of the depth image, and the generation of a spatial cue mask based on the clustering results, specifically includes the following steps: Step S110: Denoise the depth image and perform distance range filtering on the denoised depth data to remove pixels that exceed a preset distance threshold, thereby obtaining an effective depth region. Step S120: Cluster the pixels in the effective depth region based on the depth value to divide the scene of the depth image into different depth levels; Step S130: Select the pixel region corresponding to the preset depth level of the depth image and generate the spatial cue mask.
[0048] In this embodiment, the depth image can be acquired through... Figure 3 The visual data acquisition module shown is used to achieve this. This module consists of a RealSense d435i depth camera and a LubanCat-RK3588 embedded processing unit, and is equipped with a 5G-Wi-Fi communication module. The entire module weighs less than 500 grams. During data acquisition, the demonstrator carries the device while walking naturally, simulating the movement of a robot in an indoor environment, capturing real-time first-person perspective depth maps and RGB images. The acquired data is as follows... Figure 4 As shown, Figure 4 The left side is the depth image. Figure 4 The right side shows synchronized RGB video image frames.
[0049] After acquiring the depth image, denoising is performed first. The original depth image often contains random noise and local holes. The purpose of denoising is to suppress this noise while preserving object edge information. Denoising methods include, but are not limited to, one or more combinations of median filtering, mean filtering, bilateral filtering, morphological filtering, hole filling, or temporal smoothing. Median filtering has a good suppression effect on salt-and-pepper noise; bilateral filtering can preserve depth edges well while denoising; for local holes in the depth image caused by occlusion or specular reflection, a hole filling algorithm based on neighbor pixel interpolation can be used for repair; in continuous video acquisition scenarios, temporal smoothing can be combined to further stabilize the depth estimation of the current frame by utilizing the temporal continuity of depth values from adjacent frames.
[0050] After denoising, distance range filtering is performed on the depth data. Depth values are evaluated pixel-by-pixel, and only pixels with depth values less than or equal to a preset distance threshold are retained as valid depth regions. The selection of the preset distance threshold depends on the depth camera's range, calibration parameters, and specific application scenario. In one example, using the effective range of the RealSense d435i camera as a reference, the preset distance threshold is set to 5000 mm. Pixels exceeding this threshold are considered background or invalid regions and do not participate in subsequent clustering and segmentation processing.
[0051] Subsequently, pixels within the effective depth region are clustered based on their depth values. Let the set of effective depth pixels after distance filtering be... Represented as:
[0052] in, This represents the first dimension of the original depth map after it has been stretched into a one-dimensional vector. The depth value of each pixel. The clustering goal is to... The pixels in the image are divided into several non-overlapping subsets, such that the pixel depth values within the same subset are as close as possible, and the depth values between different subsets are as far apart as possible.
[0053] In one example of this embodiment, the K-Means clustering method is used. The optimization objective function of the K-Means method is:
[0054] in, Indicates the first A cluster, Indicates the corresponding cluster center, Represents pixel depth value and its cluster center The squared Euclidean distance between them. The objective function gradually converges to a local minimum by iteratively updating cluster assignments and cluster centers. In this application scenario, to balance clustering speed and subsequent prompting effects, three cluster centers are used, representing the depth data of the near, medium, and far categories, respectively. Figure 5 As shown, the left side is the original depth image, and the right side is the depth image after K-Means clustering. The red area is the cue mask area generated based on the clustering results, that is, the area of the spatial cue mask.
[0055] After clustering, a spatial cue mask is generated by selecting pixel regions corresponding to a preset depth level from multiple depth levels. Since objects closest to the camera typically have the clearest images and most complete outlines, this embodiment can select only the nearest clusters as cue regions. Alternatively, a combination of near-range and mid-range clusters can be selected as cue regions, suitable for scenarios requiring simultaneous attention to objects at multiple distance levels. The spatial cue mask is a binary mask, with cue regions marked as 1 and non-cue regions marked as 0, and its size is consistent with the original depth image.
[0056] Spatial cue masks provide explicit spatial priors for video segmentation models, guiding them to the near-field regions most likely containing clear objects instead of blindly performing dense searches across the entire image. Especially in complex backgrounds, scenes with multiple stacked objects, or partial occlusion, cue masks generated by deep clustering can effectively filter distant distractions and blurred backgrounds, improving the targeting and accuracy of object segmentation.
[0057] Step S200: Input the spatial cue mask and the RGB video segment into the video segmentation model to obtain the temporal segmentation result. The spatial cue mask is used to guide the region of interest, and the temporal information of the RGB video segment is used for cross-frame propagation.
[0058] In this embodiment, the video segmentation model refers to a neural network model that can receive spatial cue information and video frame sequences as joint input, and continuously track and optimize the target segmentation region in each frame of the video. Specifically, the video segmentation model can be the SAM2 model, or it can be replaced with other models capable of video target segmentation, video instance segmentation models, interactive segmentation models, or integrated detection and segmentation models.
[0059] During operation, the generated spatial cue mask is used as the initial frame spatial cue for the video segmentation model. The spatial cue mask guides the region of interest, informing the model where to focus its search and segmentation efforts on the target within the image, avoiding wasted computational resources or missegmentation in background areas outside the cue region. After reading the spatial cue mask, the video segmentation model generates the initial target segmentation region based on the cue region in the starting frame synchronized with the depth image. Subsequently, it uses the temporal information of the RGB video clip for cross-frame propagation, gradually spreading the segmentation result of the initial frame to subsequent frames. Cross-frame propagation refers to the model continuously adjusting and optimizing the boundary of the target segmentation region in the temporal dimension based on the motion and appearance continuity of pixels between adjacent frames, ensuring the segmentation result maintains temporal continuity and consistency throughout the entire video clip.
[0060] Temporal segmentation results refer to the set of segmentation masks output by the video segmentation model for each frame of an RGB video clip, which is a mask sequence with a frame-length dimension. Through this method, the spatial cues provided by a single-frame depth image are successfully extended to temporally consistent segmentation results across the entire RGB video clip. This mechanism comprehensively utilizes the spatial localization advantages of depth information and the continuity advantages of video temporal information. Specifically, depth information quickly locates candidate regions and filters background interference at the geometric level, while temporal information improves the stability, integrity, and robustness of the segmentation mask through the complementary and smoothing effects of multi-frame observations. Compared to segmentation methods that rely solely on a single-frame RGB image, it performs more reliably in scenarios with occlusion, motion blur, or transient noise.
[0061] In one implementation, the step of inputting the spatial cue mask and the RGB video segment into a video segmentation model to obtain a temporal segmentation result specifically includes the following steps: Step S210: Use the spatial cue mask as the initial frame spatial cue of the video segmentation model, and propagate and optimize the target segmentation region in each frame of the RGB video segment through the video segmentation model to obtain the temporal segmentation result; The video segmentation model utilizes historical frame information from adjacent video segments to assist in segmentation prediction for the first few frames of the current segment when segmenting the current RGB video segment, so as to achieve a smooth transition between segmentation results between adjacent segments.
[0062] In this embodiment, a specific implementation of the video segmentation model is the SAM2 model. SAM2 is a deep learning model that can receive spatial cues and continuously segment targets in a video sequence. Internally, it maintains a memory module for storing and retrieving historical frame feature information in the temporal dimension of the video.
[0063] To balance the real-time performance and stability of the model inference, 30 frames of RGB video images were treated as a set of data for continuous semantic segmentation. Each set of data is denoted as [data type]. V :
[0064] in, to This represents the RGB image of each frame out of 30 frames.
[0065] Spatial cue mask and RGB video clip The common input to the SAM2 model, where the spatial cue mask is derived from and The results of the depth map acquired at the same time. SAM2 receives spatial cue mask and video temporal information. Perform temporal segmentation on the target region in the video clip and output the initial segmentation result of the video clip:
[0066] in, , Indicates the first The segmentation mask result corresponding to the frame.
[0067] After the initial frame is segmented, SAM2 leverages the motion features and appearance continuity learned between consecutive RGB frames to propagate the segmentation results from the initial frame to subsequent frames. In each frame, SAM2 fine-tunes and optimizes the boundaries of the propagated segmented region based on the RGB image information of the current frame to adapt to changes in the actual position and shape of the target within that frame; this process is called cross-frame propagation. Through cross-frame propagation, the spatial cue information that originally existed only in the depth map of a single frame is expanded into a temporal segmentation result covering the entire RGB video segment. For example... Figure 6 As shown, the target mask region obtained after temporal image segmentation maintains good temporal smoothness and spatial consistency throughout the entire video clip.
[0068] SAM2 also employs a memory-assisted mechanism during cross-frame propagation. When video is segmented into continuous fixed-length segments, such as 30 frames per segment, there are temporal transitions between adjacent video segments. For the currently processed video segment, the first 6 frames have not yet accumulated sufficient historical information in the SAM2 memory. In this case, SAM2 utilizes the memory features of the last 6 frames of the previous video segment to assist in the segmentation prediction of the first 6 frames of the current segment. This mechanism allows for a natural and smooth transition between the segmentation results of adjacent video segments at the boundaries, preventing sudden jumps in segmentation results or target loss due to the video being divided into processing units.
[0069] Specifically, the above formula utilizes the timing information of the previous video when t < 6. In A total of 6 frames of data are used as a reference, which allows for an automatic and smooth transition between two adjacent frames. The memory data in the video clips makes the prediction results smoother.
[0070] The advantages of using SAM2 for temporal segmentation are twofold: firstly, it extends the spatial guidance capability of depth cues from a single frame to the entire video segment, allowing subsequent frames to inherit spatial prior information from the first frame without requiring independent depth map input. Secondly, SAM2 utilizes multi-frame temporal information for joint inference, effectively mitigating issues such as transient occlusion, motion blur, or local noise that may occur in a single frame through complementary information from adjacent frames, resulting in a more stable and complete target mask than single-frame segmentation. Furthermore, the adjacent segment memory assistance mechanism ensures the coherence of segmentation results in continuous video stream processing scenarios, which is particularly important for robot real-time environmental perception and continuous target tracking.
[0071] It is understood that SAM2 is merely an example of the video segmentation model in this embodiment, and the video segmentation model can also be replaced by other models with video object segmentation capabilities, video instance segmentation models, interactive segmentation models, or integrated detection and segmentation models. The information used for spatial cues is not limited to spatial cue masks; it can also be point cues, bounding box cues, or center position cues generated from depth maps. Furthermore, during temporal segmentation, the initial region obtained from depth cues can be combined with motion information, optical flow information, or background modeling results to jointly constrain the segmentation model's scope of interest and cross-frame propagation path, thereby further enhancing the segmentation robustness in dynamic scenes.
[0072] Step S300: Select sample frames from the RGB video segments, and perform instance segmentation within the segmentation mask region corresponding to the sample frames in the temporal segmentation result to obtain a target instance set containing multiple target instances.
[0073] In this embodiment, a sample frame refers to the highest quality RGB image selected from subsequent frames of an RGB video segment according to a preset image quality evaluation metric. The purpose of selecting a sample frame is to provide a reference image with the highest clarity and richest target detail for subsequent instance segmentation and semantic recognition. The preset image quality evaluation metric can be a sharpness metric, calculated using methods such as the Laplacian variance method, the Tenengrad gradient method, or the Fourier high-frequency energy method; it can also be a segmentation stability metric, a target region area metric, or an image information entropy metric, etc. The reason for selecting a sample frame from subsequent frames rather than the starting frame is that subsequent frames have already undergone temporal propagation and optimization, and the target region boundaries are usually more accurate and stable.
[0074] Instance segmentation refers to not only separating the foreground region from the background in an image, but also further distinguishing pixels belonging to different individual objects within the foreground region, generating a unique instance mask for each independent object. Instance segmentation is not performed on the entire sample frame image, but rather within the segmentation mask region corresponding to the sample frame in the temporal segmentation result. Specifically, the segmentation mask region corresponding to the sample frame in the temporal segmentation result is used as the constraint range, and full segmentation is performed only within this constraint range. Full segmentation means that the segmentation model automatically detects and segments all object instances appearing in the image. Instances in the full segmentation result that intersect with the constraint range are retained, while instances that do not intersect are discarded. Here, intersection refers to two regions overlapping at the pixel level.
[0075] The target instance set refers to the collection of multiple independent target instances retained after filtering within the aforementioned constraints. Each target instance corresponds to an independent object region and has its own instance mask. The initial temporal segmentation result typically provides a relatively coarse-grained foreground region, which may contain multiple adjacent or partially occluded objects. Through instance segmentation within the constraints, these objects can be separated into independent target instances one by one. The two-stage segmentation architecture, which extracts fine-grained instances from the coarse-grained region, achieves fine-grained analysis of multiple objects in the scene while ensuring segmentation robustness.
[0076] It should be noted that the method in this embodiment is applicable to the discovery and recognition of a single object in a single-object scenario, as well as to the discovery and classification of multiple new categories of objects in a multi-object scenario.
[0077] In one implementation, the step of selecting sample frames from the RGB video segment and performing instance segmentation within the segmentation mask region corresponding to the sample frames in the temporal segmentation result to obtain a target instance set containing multiple target instances specifically includes the following steps: Step S310: Based on the preset image quality evaluation index, select the frame with the best quality from the RGB video clip as the sample frame; Step S320: Using the segmentation mask region corresponding to the sample frame in the temporal segmentation result as the constraint range, perform full segmentation within the constraint range to obtain the full segmentation result; Step S330: Include the target instances in the full segmentation result that intersect with the constraint range into the target instance set.
[0078] In this embodiment, the selection of sample frames is based on a preset image quality evaluation index. The image quality evaluation index is a calculation standard used to quantify whether an image is suitable as input for subsequent instance segmentation. The quality score is calculated for each frame after the starting frame of the RGB video clip, and the frame with the highest quality score is selected as the sample frame. Sharpness is one such index, reflecting the richness of high-frequency details in the image. Higher sharpness indicates sharper object edges and clearer texture details, enabling the segmentation model to obtain more accurate instance boundaries on such images.
[0079] Specifically, regarding from The 29 frames, according to the resolution operator Filter:
[0080] This means selecting the frame with the highest resolution within a given time window. The resolution operator... The calculation methods for image sharpness typically include, but are not limited to, the following: One is the Laplacian variance method, which uses the Laplacian operator to perform convolution operations on the image to extract edge information, and then calculates the variance of the edge image as a sharpness score; the larger the variance, the sharper the image. Another is the Tenengrad gradient method, which calculates the horizontal and vertical gradients of the image using the Sobel operator, and uses the sum of the squares of the gradient magnitudes as a measure of sharpness. Yet another is the Fourier high-frequency energy method, which transforms the image to the frequency domain and then calculates the proportion of energy occupied by high-frequency components; the higher the high-frequency energy, the sharper the image. Besides sharpness metrics, image quality evaluation metrics can also include segmentation stability metrics, which measure the consistency of the target region between adjacent frames, target region area metrics, or image information entropy metrics, etc.
[0081] Specifically, frames selected from the 29 frames following the starting frame. ,in Selected sample frame as The corresponding initial segmentation result is .
[0082] After selecting a sample frame, its corresponding segmentation mask region in the temporal segmentation result is used as the constraint range. The constraint range defines the spatial boundary of subsequent instance segmentation operations. The segmentation model performs full segmentation only within this spatial range, and pixels outside the range are not considered. The temporal segmentation result already provides the approximate foreground region where the target object is located. Limiting the full segmentation operation to this region can avoid unnecessary intensive computation on irrelevant background regions, and at the same time reduce the probability of background interference objects being incorrectly segmented as target instances.
[0083] Within the constraints, a full segmentation operation is performed using a segmentation model, automatically detecting and segmenting all object instances present. In one example, full segmentation can also use the SAM2 model, but in this case, the model operates in full image segmentation mode rather than guided mode. The resulting full segmentation output contains masks of all object instances detected within the constraints.
[0084] Specifically, construct sample pairs Then, it is fed back into the SAM2 model to obtain a set of masked targets for the region:
[0085] in, Each Represents a target instance mask. This indicates the total number of targets generated.
[0086] Subsequently, the intersection of each instance in the full segmentation result with the constraint range is processed. If the instance mask obtained from the full segmentation has pixel-level overlap with the constraint range, that is, the two regions share at least one pixel, it is included in the target instance set. If the instance mask is completely outside the constraint range, it is discarded.
[0087] Specifically, these targets are sequentially connected with... The pixels in the image are checked for intersection, and all intersecting target objects are selected as the final target mask.
[0088] like Figure 7 As shown, the final set of target instances consists of all target instances that intersect with the constraint range. Through the above operations, multiple objects originally marked as the same foreground region in the temporal segmentation result are successfully decomposed into their own independent target instances. For example, a water bottle, a phone stand, and a pen placed side by side on a table, each instance has a precise pixel-level mask. The two-stage segmentation strategy of first obtaining a stable coarse-grained foreground region using depth cues and temporal information, and then performing fine-grained instance decomposition within that region, achieves fine-grained analysis of multiple objects in the scene while ensuring segmentation robustness and efficiency.
[0089] Step S400: Extract the visual features of each target instance in the target instance set, input the visual features into a trainable classifier for category recognition, and determine the target instances that meet the preset unknown determination conditions as unknown category targets.
[0090] In this embodiment, visual features refer to high-dimensional vector representations extracted from the image regions corresponding to target instances that characterize the appearance attributes of the object. The extraction method involves inputting the image regions corresponding to each target instance into a feature extraction model, which then outputs a fixed-dimensional feature vector. The specific implementation of the feature extraction model can be the DINOv2 model, or it can be replaced with other self-supervised visual models, visual Transformer models, convolutional neural networks, multimodal visual encoders, CLIP-type models, or lightweight backbone networks. The input method for the image regions can be to directly extract the foreground region using an instance mask, to crop the region using the bounding box corresponding to the mask, or to simultaneously retain both the foreground mask image and the bounding box image as joint input.
[0091] A trainable classifier is a classification module whose parameters can be updated through training data. It maps the input visual feature vector to the predicted probability distribution of each known category. Trainable classifiers can take the form of a linear classifier head, which is a linear transformation layer consisting of a weight matrix and a bias vector. The output is normalized to obtain the category probability. Linear classifier heads are characterized by small parameter count, fast training speed, and ease of incremental updates. Other forms include multilayer perceptrons, prototype networks, metric learning heads, cosine classifier heads, parameter-efficient adapters, or LoRA modules.
[0092] After obtaining the category prediction results, the target instance is determined to belong to a known category or an unknown category based on preset unknown criteria. Preset unknown criteria refer to a set of rules used to define whether the classifier's prediction results are sufficiently confident; meeting these criteria indicates that the classifier cannot reliably classify the target into any known category. Preset unknown criteria may include at least one of the following: the maximum category probability is lower than a preset probability threshold, the maximum logits value is lower than a preset logits threshold, the classification entropy is higher than a preset entropy threshold, the difference between the scores of the first two categories is lower than a preset interval threshold, or a combination of the above conditions. In one example, when the maximum probability of the target belonging to any known category is lower than 0.5, it is determined to be an unknown category target.
[0093] In this way, the trainable classifier prioritizes recognition tasks based on existing category knowledge, triggering subsequent large-scale model semantic recognition and incremental learning processes only when encountering uncertain targets. This design constructs a lightweight, trainable, and continuously updatable target recognition front-end that can quickly complete recognition in most common scenarios, only calling external large-scale model resources when necessary, thus balancing recognition efficiency and privacy protection.
[0094] In one implementation, the step of extracting the visual features of each target instance in the target instance set, inputting the visual features into a trainable classifier for category recognition, and determining the target instance that meets the preset unknown determination condition as an unknown category target specifically includes the following steps: Step S410: Input the image region corresponding to each target instance into the feature extraction model to obtain the visual feature vector of the target instance; Step S420: Input the visual feature vector into the trainable classifier to obtain the category prediction result; Step S430: When the category prediction result meets the preset unknown determination condition, the corresponding target instance is determined as an unknown category target.
[0095] In this embodiment, a specific implementation of the feature extraction model is the DINOv2 model. DINOv2 is a large-scale visual pre-trained model based on self-supervised learning. It learns visual representations with strong generalization ability through self-distillation training on massive amounts of unlabeled image data. The DINOv2 model takes an image region as input and outputs a fixed-dimensional feature vector, which highly abstractly encodes key information such as the appearance, texture, shape, and semantics of objects in the image region.
[0096] for Each target mask in From the corresponding RGB frame Extract the target region. Extraction methods may include: (1) directly extracting the foreground region using a mask; (2) cropping the region using the bounding box (BBox) corresponding to the mask; and (3) retaining both the mask foreground image and the bounding box image as joint input.
[0097] Specifically, the foreground region can be extracted directly using the mask of the target instance, retaining only the object's own pixels and setting the background pixels to zero or replacing them with uniform grayscale values. Alternatively, the region can be cropped using the bounding rectangle corresponding to the mask, retaining the original RGB information without pixel-level separation. Or, both the foreground mask image and the bounding box image can be retained simultaneously, using them as joint input to provide both fine contour information of the object and overall appearance context.
[0098] Let the first Each target region is represented as Then Input the DINOv2 model to obtain tokens or feature vectors of the target region. , is represented as:
[0099] In another implementation, the RGB foreground image, bounding box region and depth local image are used as joint inputs to fuse appearance texture information and geometric distance information into a unified feature representation, enabling the feature extraction model to use both the visual appearance and spatial location information of the target for more robust representation learning.
[0100] The visual feature vectors output by DINOv2 are discriminative. The feature vectors of objects of the same class are close to each other in the vector space, while the feature vectors of objects of different classes are far apart. This allows a single lightweight classifier to effectively divide the decision boundaries of each class in the feature space.
[0101] A trainable classifier takes a visual feature vector as input and outputs a class prediction. In one example, the trainable classifier uses a linear classification head. , is represented as:
[0102] in, and These are the trainable parameters of the linear classifier head.
[0103] Linear classification heads offer advantages such as fewer parameters, faster forward inference speed, and ease of incremental updates, making them suitable for deployment on edge devices with limited computing resources. Besides linear classification heads, trainable classifiers can also employ multilayer perceptrons to increase non-linear classification capacity, prototype networks to classify based on the distance between feature vectors and prototypes of each category, metric learning heads to calculate similarity using distance metric functions, cosine classification heads to classify based on cosine similarity, and parameter-efficient adapters or LoRA modules to achieve classification functionality with minimal trainable parameters. The structure of trainable classifiers can also adopt a hierarchical design, where the first layer predicts the broad category to which the target belongs, and the second layer predicts the finer category within that broad category. This coarse-to-fine classification strategy enhances fine-grained recognition capabilities.
[0104] After obtaining the category prediction results, known category targets and unknown category targets are distinguished according to preset unknown judgment conditions. When the trainable classifier shows significant uncertainty in its prediction result for a target instance, it means that the target may not belong to any learned known category.
[0105] Pre-defined unknown determination conditions can employ one or more of the following strategies: (1) Maximum category probability is lower than the preset threshold: If the maximum predicted probability of a target instance in all known categories is lower than the preset probability threshold, it is determined to be an unknown category target; (2) Maximum logits below the preset threshold: If the unnormalized maximum logits value is lower than the preset logits threshold, it is determined to be an unknown category target; (3) Classification entropy is higher than the preset threshold: If the Shannon entropy of the category prediction distribution is higher than the preset entropy threshold, the higher the entropy, the more even the distribution and the more uncertain the classifier is, and it is determined to be an unknown category target; (4) The difference between the scores of the top two categories is lower than the preset threshold: If the difference between the predicted probabilities of the top two categories is lower than the preset interval threshold, it is determined to be an unknown category target; (5) Simultaneously satisfying both low confidence and low interval conditions: It is only determined to be unknown when both low confidence and low interval conditions are satisfied.
[0106] In one example, to demonstrate the effectiveness of the method, the simplest threshold determination strategy is used:
[0107] in, Indicate target Belongs to the The probability of a known category. Threshold for determining unknown categories.
[0108] In one example, such as Figure 8 As shown, the trainable classifier outputs category prediction results for each input target instance. It can give a high-confidence correct prediction for the learned categories, but for objects that have not appeared in the training set, the prediction probability is low, thus being identified as unknown category targets, triggering the large model semantic recognition process.
[0109] Through the aforementioned feature extraction and classification mechanism, a lightweight, trainable, and continuously updated known category recognition process is constructed, which can quickly handle known object recognition tasks in most common scenarios. The subsequent large model recognition and incremental learning process is only activated when encountering unknown objects with high prediction uncertainty. This balances recognition speed and adaptability to open environments, and also reduces the frequency of dependence on large cloud models.
[0110] Step S500: Perform structured semantic recognition on the unknown category target using a multimodal large model to obtain structured target attribute information, and match the structured target attribute information with the target instance set to construct a training sample set with semantic labels.
[0111] In this embodiment, a multimodal large model refers to a general-purpose artificial intelligence model capable of simultaneously receiving image and text input and generating structured output, such as a large language model supporting visual understanding like GPT-4V. Multimodal large models can be deployed in the cloud or utilize locally deployed large model resources. Structured semantic recognition requires that the multimodal large model not output recognition results in the form of natural language paragraphs, but instead return target attribute information according to a preset structured data format. The preset structured data format can be YAML, JSON, XML, or a predefined key-value structure. The structured target attribute information includes at least category information, location information, and confidence information. In a specific example, the structured target attribute information includes five attributes: color, major category, minor category, confidence, and target bounding box.
[0112] The sample frame image region containing the unknown category target, along with its corresponding effective segmentation range information, is sent to the multimodal large model. The large model is then restricted to returning the attribute information of all objects within that region in a preset format. The effective segmentation range information defines the image space range that the large model should focus on, avoiding unnecessary recognition of background objects outside the range.
[0113] After obtaining the structured target attribute information, it is matched with the target instance set to determine the semantic category corresponding to each unknown category segmentation target. The matching criterion is the spatial consistency between the two, specifically, the spatial matching degree is calculated based on the positional information in the structured target attribute information and the positional information of each target instance in the target instance set. The spatial matching degree is a quantitative indicator that measures the degree of spatial correspondence between the object recognized by the large model and the instance obtained by pixel-level segmentation. It can be calculated based on the intersection-union ratio between bounding rectangles, the distance between mask center points, or a weighted combination of the above factors. Target pairs whose spatial matching degree meets the preset matching conditions are established to form a correspondence. The matching condition can be that the spatial matching degree is higher than the preset matching threshold. Finally, a training sample set is constructed based on the established correspondence between the target instances and their corresponding semantic category labels. Each sample in the training sample set contains at least the image region of the target instance, the instance mask, the positional information, and the semantic category label.
[0114] In traditional approaches, the semantic recognition results of large-scale models typically remain at the text output level, failing to establish a precise correspondence with pixel-level segmentation results. By automatically calculating and pairing spatial matching degrees, the open semantic understanding capabilities of large-scale models are successfully transformed into supervisory signals that can be used to train local classification models. This constructs a conversion channel from general semantic knowledge to specialized recognition capabilities. The entire process requires no manual annotation, achieving fully automated sample construction.
[0115] In one implementation, the step of performing structured semantic recognition on the unknown category target using a multimodal large model to obtain structured target attribute information, and matching the structured target attribute information with the target instance set to construct a training sample set with semantic labels, specifically includes the following steps: Step S510: Send the image region where the unknown category target is located and the corresponding effective segmentation range information to the multimodal large model; Step S520: Based on a preset structured data format, the multimodal large model returns the attribute information of the unknown category target, wherein the attribute information includes at least category information, location information, and confidence information; Step S530: Calculate the spatial matching degree between the location information in the structured target attribute information and the location information of each target instance in the target instance set; Step S540: Establish a correspondence between target pairs whose spatial matching degree meets the preset matching conditions; Step S550: Construct the training sample set based on the target instances with established correspondences and their corresponding semantic category labels.
[0116] In this embodiment, targets classified as unknown categories are not directly discarded; instead, a multimodal large model is invoked to process the sample frames. In the middle Structured identification of targets within the effective range.
[0117] The specific implementation of a multimodal large-scale model can be a large language model that supports image understanding, such as ChatGPT, or other multimodal models with visual question answering or image description capabilities. The deployment location of the multimodal large-scale model can be a cloud server or a local computing node with sufficient computing power. A predefined structured data format refers to a predefined data organization form with fixed fields and hierarchical relationships. The reason for using a structured format is that while natural language output is easy to read, it is difficult for programs to automatically parse and extract key information, while structured formats can be directly processed automatically by downstream matching and sample construction modules. The predefined structured data format can adopt YAML, JSON, XML, or a predefined key-value structure.
[0118] In a specific example, the RGB image of the sample frame and the effective segmentation range information are sent to the large model, which is then requested to return target information objects in YAML format. Each target object includes at least the following attributes: color, major category, minor category, confidence level, and the corresponding bounding box. The request format sent to the large model is shown in Table 1, and the format returned by the large model is shown in Table 2.
[0119] Table 1
[0120] Table 2
[0121] Examples of structured target attribute information returned after processing a multimodal large model are as follows: Figure 9 As shown.
[0122] After obtaining the structured target attribute information, each object is matched with each target instance in the target instance set to determine the correspondence between them. The matching is based on spatial location consistency. The position of the object identified by the large model in the image should be close to or coincide with the position of the target instance obtained by pixel-level segmentation. Specifically, the matching criteria may include at least one of the following: (1) the intersection-union ratio (IoU) between the bounding box of the BBox and the target mask; (2) the distance between the center point of the mask and the center point of the BBox; (3) the compatibility between the area of the mask and the area of the BBox; (4) the consistency between the color description and the color statistics of the target region; and (5) multi-condition weighted matching score.
[0123] Spatial matching degree is an indicator that quantifies the degree of spatial correspondence between two objects. To balance the hole effect and effective pixel ratio between the target region (BBox) and the mask region, a weighted score combining Intersection over Union (IoU) and center distance is used, expressed as:
[0124] in, Indicates mask The outer frame, Indicates the first in the target A box, Indicates the coordinates of the center point. and These are the weighting coefficients.
[0125] For target pairs with matching scores higher than a threshold, a correspondence between the mask and semantic attributes is established, thereby constructing a training sample set, represented as:
[0126] After establishing the correspondence, each successfully matched segmented target instance obtains a semantic category label and other auxiliary attributes such as color and confidence score from the multimodal large model. Based on this information, a training sample set is constructed, where each sample contains at least the image region corresponding to the target instance, instance mask, location information, semantic category label, and the confidence score output by the multimodal large model. For cases where multiple target instances correspond to the same sample frame image, storage space can be saved by sharing the image index. Thus, the semantic output of the multimodal large model, originally existing only in text form, is successfully converted into supervised training data that can be directly used by a trainable classifier, achieving an end-to-end automated flow from open semantic understanding to model training.
[0127] Step S600: Incrementally train the trainable classifier based on the training sample set to obtain an updated classifier, so that the trainable classifier has the ability to re-identify the unknown category target.
[0128] In this embodiment, incremental training refers to updating the parameters of a trainable classifier using only a small number of newly acquired training samples without retraining the entire model. This allows the classifier to retain its ability to recognize old categories while gaining the ability to recognize new categories. Specific methods for incremental training can include gradient descent-based parameter updates, prototype-based parameter updates, sample replay mechanisms, meta-learning methods, or other incremental learning strategies.
[0129] The training process is based on the constructed training sample set. The weight of each sample in training can be determined based on the confidence level of the target output corresponding to that sample in the multimodal large model. Samples with higher confidence levels contribute more to the model update, making the model more inclined to learn the feature distribution of high-quality labeled samples. At the same time, to prevent catastrophic forgetting during incremental training, i.e., the learning of new categories leading to a significant degradation in the recognition performance of old categories, a constraint on the output distribution of the trainable classifier on old category samples is introduced during training. This is achieved by keeping the prediction distribution of the classifier for old category samples basically consistent before and after the update.
[0130] After incremental training, an updated classifier is obtained. This classifier can directly output the corresponding known category prediction results for objects that were previously judged as unknown categories and objects with similar appearances that will reappear in the future, without having to call the multimodal large model again. This enables the trainable classifier to re-identify objects that were originally unknown categories.
[0131] The new knowledge provided by the multimodal large model is only accumulated into a local lightweight classifier through incremental training with small samples, allowing the recognition capability to gradually accumulate and expand with interaction in unfamiliar environments. Compared to solutions that rely on a large model for recognition every time a new object is encountered, or that retrain a large-scale model entirely every time new category data is obtained, this embodiment balances recognition efficiency, computational cost, and privacy protection, enabling the system to have sustainable autonomous evolution capabilities in open environments.
[0132] In one implementation, the incremental training of the trainable classifier based on the training sample set to obtain an updated classifier, so that the trainable classifier has the ability to re-identify the unknown category target, specifically includes the following steps: Step S610: Using the output distribution of the trainable classifier on old category samples as a constraint, incrementally train the trainable classifier based on the samples in the training sample set, update the parameters of the trainable classifier, and obtain the updated classifier. The weights of each sample during training are determined based on the confidence level of the target output corresponding to that sample in the multimodal large model.
[0133] In this embodiment, the goal of incremental training is to use the constructed training sample set to train the classifier. Update the parameters to obtain the updated classification header. This allows it to recognize newly added, unknown categories of targets while maintaining its ability to identify existing old categories.
[0134] Each sample in the training sample set carries a confidence score provided by the multimodal large model, reflecting the model's degree of certainty regarding the semantic annotation of that sample. During training, samples with higher confidence scores should have a greater impact on model parameter updates because their annotation quality is more reliable, while samples with lower confidence scores may contain annotation noise and their contribution to model training should be appropriately reduced. Based on this, a weighted classification loss function is constructed by assigning training weights related to each sample's confidence score.
[0135] For a batch of A dataset consisting of data samples, using target region features As input, the fine-grained categories serve as supervisory labels, and the confidence scores provided by the large model are introduced as sample weights to construct a weighted classification loss function:
[0136] in, For the first The weights of each sample, For label indication value, To predict probabilities, This indicates the number of categories after the update.
[0137] Besides using confidence scores directly as sample weights, confidence scores can also be used to set a pseudo-label screening threshold. Only samples with confidence scores higher than the preset threshold are retained for incremental training, while samples with confidence scores lower than the threshold are considered unreliable and are discarded. Confidence scores can also be used as a temperature parameter in the loss function to adjust the differences in gradient contributions among different samples during training. Furthermore, the update intensity of this incremental training can be dynamically controlled based on the average confidence score of the current batch of samples. A higher average confidence score allows for a larger parameter update magnitude, while a lower average confidence score employs a more conservative update strategy.
[0138] In one embodiment, sample weights With the confidence of the large model output Positive correlation:
[0139] in, This is the adjustment coefficient.
[0140] During incremental training, training a trainable classifier using only new category samples can easily lead to a catastrophic forgetting problem, where the model learns new categories but forgets its previously acquired ability to recognize older categories. To mitigate this issue, a constraint is introduced during training on the output distribution of the trainable classifier on old category samples. This constraint is achieved through a distillation constraint term. The distillation constraint term is based on the classifier before the update. With the updated classifier Constructing the output distribution differences of old category samples, to The predicted probability distribution is used as a soft label constraint. The output distributions are made as close as possible to the target class. To prevent incremental training from causing performance degradation in older classes, an old classifier head distillation constraint term is added, denoted as:
[0141] The total loss function can be written as:
[0142] in, The distillation term weights are used to balance the trade-off between learning new categories and preserving old categories.
[0143] Through the above small-sample incremental training, a new classification head is obtained. This enables it to re-identify new categories. For example... Figure 10 As shown, after incremental training with small samples, the trainable classifier has successfully learned to recognize the tablet computer and the white mat that were originally classified as unknown categories, while maintaining the ability to correctly recognize objects of known categories.
[0144] The incremental training mechanism is effective in several ways. First, by using confidence weighting, high-quality samples contribute more to model updates, reducing the negative impact of low-quality pseudo-labels on classifier performance and enabling faster and more stable convergence of the incremental learning process. Second, the distillation constraint effectively mitigates the catastrophic forgetting problem in incremental training, ensuring that the ability to identify both new and old classes coexists in the same lightweight classifier. Furthermore, the sample size for the entire incremental training process is relatively small, involving only newly discovered unknown targets and their multimodal large-scale model annotation results. It eliminates the need for overall retraining of the large-scale model, resulting in low training costs, fast convergence, and suitability for periodic execution on computationally limited edge devices.
[0145] Figure 2This demonstrates the complete workflow of this embodiment. Specifically, synchronously acquired depth maps and RGB video clips are used as input. First, the depth maps are denoised, and valid depth data within a defined distance range is selected. K-Means clustering is used to divide the scene into near, mid, and far depth layers, generating a Mask-Prompt to guide the segmentation model. Then, this Mask-Prompt, along with the RGB temporal clips, is input into the SAM2 model to obtain the initial temporal segmentation results for the video clips. Next, sample frames are selected from subsequent frames, and full segmentation is performed again within the valid area of the initial segmentation results, resulting in multiple target instance mask sets. For each target instance, DINOv2 is used to extract the target region representation, and a trainable linear classification head is used to complete known category identification. For unknown category targets with low classification scores, large models such as ChatGPT are further invoked, requiring them to return structured target attribute information in YAML format, including color, major category, fine category, confidence score, and bounding box. Next, these structured semantic results are automatically matched with the existing target mask set to construct a new sample set containing RGB, Mask, BBox, color, category, and confidence. Finally, a weighted loss function is designed by combining fine-grained categories and confidence to perform incremental training on the linear classifier head with a few samples, resulting in an updated classifier head that enables the system to continuously learn and re-recognize new categories.
[0146] After testing, the overall task detection accuracy reached 95.1% (simulating 10 common categories of objects in daily life, repeated 40 times), and the model inference speed on an RTX 5070 reached 30 frames per second, meeting real-time requirements. Thanks to its excellent communication interface, this invention does not require additional computing power from mobile platforms.
[0147] In summary, the core of this embodiment lies in four aspects: First, using deep clustering results as front-end spatial cues to guide the temporal segmentation model to more accurately focus on potential object regions; second, using a two-stage mechanism of initial temporal segmentation and local full segmentation to further extract independent target instances from coarse-grained regions; third, introducing a large-model structured semantic output + pixel-level automatic target alignment mechanism to achieve automatic conversion from open semantic recognition to supervised training sample construction; and fourth, constructing a closed-loop incremental learning framework that integrates new knowledge from a large model with a lightweight classification head to complete knowledge accumulation, enabling the system to gradually learn unknown object categories.
[0148] Compared with existing technologies, this invention can not only discover unknown objects, but also automatically obtain their category semantics and convert them into trainable samples, further completing model updates, thereby achieving continuous learning in an open environment. This solution has strong engineering feasibility and scalability, and can be applied to scenarios such as robot perception, intelligent monitoring, industrial recognition, unmanned system environmental understanding, and intelligent annotation.
[0149] like Figure 11 As shown in the figure, this embodiment of the invention provides a detection and recognition system for unknown objects. The system includes: a depth cue generation module 10, a temporal segmentation module 20, an instance splitting module 30, a classification and recognition module 40, a sample construction module 50, and an incremental training module 60.
[0150] Specifically, the depth cue generation module 10 is used to acquire a depth image synchronized with the RGB video clip to be processed, preprocess and cluster the depth image, and generate a spatial cue mask based on the clustering result; the temporal segmentation module 20 is used to input the spatial cue mask and the RGB video clip into a video segmentation model to obtain a temporal segmentation result, wherein the spatial cue mask is used to guide the region of interest, and the temporal information of the RGB video clip is used for cross-frame propagation; the instance splitting module 30 is used to select sample frames from the RGB video clip, and perform instance segmentation within the segmentation mask region corresponding to the sample frames in the temporal segmentation result to obtain a target instance set containing multiple target instances; the classification and identification... The identification module 40 is used to extract the visual features of each target instance in the target instance set, input the visual features into a trainable classifier for category recognition, and identify target instances that meet preset unknown judgment conditions as unknown category targets; the sample construction module 50 is used to perform structured semantic recognition on the unknown category targets through a multimodal large model to obtain structured target attribute information, and match the structured target attribute information with the target instance set to construct a training sample set with semantic labels; the incremental training module 60 is used to perform incremental training on the trainable classifier based on the training sample set to obtain an updated classifier, so that the trainable classifier has the ability to re-identify the unknown category targets.
[0151] Based on the above embodiments, the present invention also provides a terminal device, the principle block diagram of which can be as follows: Figure 12 As shown, the terminal device includes a processor, memory, network interface, display screen, and temperature sensor connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for detecting and identifying unknown objects. The display screen can be an LCD screen or an e-ink screen. The temperature sensor is pre-installed inside the terminal device to detect the operating temperature of the internal components.
[0152] Those skilled in the art will understand that Figure 12 The schematic diagram shown is only a partial structural diagram related to the present invention and does not constitute a limitation on the terminal device to which the present invention is applied. The specific terminal device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0153] In one embodiment, a terminal device is provided, including a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors, the one or more programs including instructions for performing operations as described in the embodiments of the methods above.
[0154] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0155] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0156] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for detecting and identifying unknown objects, characterized in that, The method includes: Acquire a depth image synchronized with the RGB video clip to be processed, preprocess and cluster the depth image, and generate a spatial cue mask based on the clustering results; The spatial cue mask and the RGB video clip are input together into the video segmentation model to obtain the temporal segmentation result. The spatial cue mask is used to guide the region of interest, and the temporal information of the RGB video clip is used for cross-frame propagation. Sample frames are selected from the RGB video segments, and instance segmentation is performed within the segmentation mask region corresponding to the sample frames in the temporal segmentation result to obtain a target instance set containing multiple target instances. Visual features of each target instance in the target instance set are extracted, and the visual features are input into a trainable classifier for category recognition. Target instances that meet the preset unknown determination conditions are identified as unknown category targets. The unknown category target is subjected to structured semantic recognition by a multimodal large model to obtain structured target attribute information, and the structured target attribute information is matched with the target instance set to construct a training sample set with semantic labels; The trainable classifier is incrementally trained based on the training sample set to obtain an updated classifier, so that the trainable classifier has the ability to re-identify the unknown category target. The step involves selecting sample frames from the RGB video segment and performing instance segmentation within the segmentation mask region corresponding to the sample frames in the temporal segmentation result to obtain a target instance set containing multiple target instances, including: Based on preset image quality evaluation indicators, the highest quality frame is selected from the RGB video clip as the sample frame; Using the segmentation mask region corresponding to the sample frame in the temporal segmentation result as a constraint range, full segmentation is performed within the constraint range to obtain the full segmentation result; The target instances in the full segmentation result that intersect with the constraint range are included in the target instance set; The step of performing structured semantic recognition on the unknown category target using a multimodal large model to obtain structured target attribute information, and matching the structured target attribute information with the target instance set to construct a training sample set with semantic labels includes: The image region containing the unknown category target and the corresponding effective segmentation range information are sent to the multimodal large model; Based on a preset structured data format, the multimodal large model returns attribute information of the unknown category target, wherein the attribute information includes at least category information, location information, and confidence information; Based on the location information in the structured target attribute information and the location information of each target instance in the target instance set, the spatial matching degree between the two is calculated; Establish a correspondence between target pairs whose spatial matching degree meets the preset matching conditions; The training sample set is constructed based on the target instances with which the correspondence is established and their corresponding semantic category labels.
2. The method for detecting and identifying unknown objects according to claim 1, characterized in that, The step of preprocessing and clustering the depth image, and generating a spatial cue mask based on the clustering results, includes: The depth image is denoised, and the denoised depth data is filtered by distance range to remove pixels that exceed a preset distance threshold, thereby obtaining an effective depth region. The pixels within the effective depth region are clustered based on depth values to divide the scene of the depth image into different depth levels; Select the pixel region corresponding to the preset depth level of the depth image to generate the spatial cue mask.
3. The method for detecting and identifying unknown objects according to claim 1, characterized in that, The step of inputting the spatial cue mask and the RGB video segment into the video segmentation model to obtain the temporal segmentation result includes: The spatial cue mask is used as the initial frame spatial cue for the video segmentation model. The video segmentation model is then propagated through each frame of the RGB video segment to optimize the target segmentation region, resulting in a temporal segmentation result. The video segmentation model utilizes historical frame information from adjacent video segments to assist in segmentation prediction for the first few frames of the current segment when segmenting the current RGB video segment, so as to achieve a smooth transition between segmentation results between adjacent segments.
4. The method for detecting and identifying unknown objects according to claim 1, characterized in that, The step of extracting visual features from each target instance in the target instance set, inputting the visual features into a trainable classifier for category recognition, and identifying target instances that meet preset unknown determination conditions as unknown category targets includes: The image regions corresponding to each target instance are input into the feature extraction model to obtain the visual feature vectors of the target instances; The visual feature vector is input into the trainable classifier to obtain the category prediction result; When the category prediction result meets the preset unknown determination condition, the corresponding target instance is determined as an unknown category target.
5. The method for detecting and identifying unknown objects according to claim 1, characterized in that, The step of incrementally training the trainable classifier based on the training sample set to obtain an updated classifier, so that the trainable classifier has the ability to re-identify the unknown category target, includes: Using the output distribution of the trainable classifier on old category samples as a constraint, the trainable classifier is incrementally trained based on samples in the training sample set to update the parameters of the trainable classifier and obtain the updated classifier. The weights of each sample during training are determined based on the confidence level of the target output corresponding to that sample in the multimodal large model.
6. A system for detecting and identifying unknown objects, characterized in that, The system, applied to the steps of implementing the method for detecting and identifying unknown objects as described in any one of claims 1-5, comprises: The depth cue generation module is used to acquire a depth image synchronized with the RGB video clip to be processed, preprocess and cluster the depth image, and generate a spatial cue mask based on the clustering results; The temporal segmentation module is used to input the spatial cue mask and the RGB video clip into the video segmentation model to obtain the temporal segmentation result. The spatial cue mask is used to guide the region of interest, and the temporal information of the RGB video clip is used for cross-frame propagation. The instance splitting module is used to select sample frames from the RGB video segment and perform instance splitting within the segmentation mask area corresponding to the sample frames in the temporal segmentation result to obtain a target instance set containing multiple target instances. The classification and recognition module is used to extract the visual features of each target instance in the target instance set, input the visual features into a trainable classifier for category recognition, and determine the target instance that meets the preset unknown judgment condition as an unknown category target; The sample construction module is used to perform structured semantic recognition on the unknown category target through a multimodal large model, obtain structured target attribute information, and match the structured target attribute information with the target instance set to construct a training sample set with semantic labels; The incremental training module is used to incrementally train the trainable classifier based on the training sample set to obtain an updated classifier, so that the trainable classifier has the ability to re-identify the unknown category target.
7. A terminal device, characterized in that, The terminal device includes a memory, a processor, and an unknown object detection and identification program stored in the memory and executable on the processor. When the processor executes the unknown object detection and identification program, it implements the steps of the unknown object detection and identification method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program for detecting and identifying unknown objects. When the program is executed by a processor, it implements the steps of the method for detecting and identifying unknown objects as described in any one of claims 1-5.
Citation Information
Patent Citations
SAM2 visual basis model assisted vehicle-mounted laser point cloud semantic segmentation method and system
CN120543855A
Target detection method based on unsupervised feature clustering and multi-modal large model collaborative iteration
CN121883895A