Target detection method and device of perception model based on incremental learning

By adopting an perception model based on incremental learning in autonomous vehicles, combining the regional proposal network and Transformer coding layer, and using prototype vectors for feature fusion, the problems of uncertainty and low efficiency of target information updates in dynamic environments are solved, and efficient target detection and recognition are achieved.

CN119942479AActive Publication Date: 2025-05-06BEIJING UNIV OF CHEM TECH

Patent Information

Application Number
CN202411715425.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-05-06
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

When existing autonomous driving vehicles deal with the update of target information in dynamic environments, there is uncertainty in the training of the perception model, the training is time-consuming and inefficient, and when training on new tasks, the performance of the old tasks will usually decrease significantly, and there are problems of catastrophic forgetting, resulting in low detection accuracy and limited recognition capabilities.

Method used

Using an perception model based on incremental learning, the RGB images on the bicycle are acquired, and the pre-trained regional proposal network and the Transformer encoding layer are processed, and combined with the prototype vector of the target type is fused to achieve object detection and classification.

Benefits of technology

It improves the accuracy of object detection in unknown scenarios, reduces the time and resource consumption of model training, avoids catastrophic forgetting, and ensures that the system can maintain efficient detection and recognition capabilities when facing new environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942479A_ABST
    Figure CN119942479A_ABST
Patent Text Reader

Abstract

The invention provides a target detection method and device of a perception model based on incremental learning. The method comprises the following steps: acquiring an RGB image at the current moment acquired by a camera on a vehicle; processing the RGB image at the current moment by using a regional proposal network to obtain a plurality of proposal boxes; processing the RGB image at the current moment by using a Transform coding layer to obtain a deep semantic feature map; fusing each proposal box, the deep semantic feature map and the prototype vector of each target type to obtain a fusion feature of each proposal box; wherein the prototype vector of each target type is obtained by performing incremental learning according to the target type in the current scene of the vehicle; processing the fusion feature of each proposal box by using a region projection layer to obtain a target positioning result; and processing the fusion feature of each proposal box by using the classification layer to obtain a target classification result. According to the invention, the target detection precision of the RGB image in an unknown scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, and in particular to a target detection method and device based on an incremental learning perception model. Background Art

[0002] When autonomous vehicles are processing target information updates in dynamic environments, the visual perception models they use often rely on a large amount of labeled data. However, in actual applications, there are some problems: in some sensitive areas, there is little prior information and the training is difficult, which leads to uncertainty in the training of the perception model; retraining / fine-tuning the entire model is time-consuming and inefficient; in addition, when training on new tasks, the performance on old tasks usually drops significantly, and there is a common defect - catastrophic forgetting. These problems result in low detection accuracy and limited recognition capabilities of visual perception models when facing complex dynamic scenes, which cannot meet actual needs. Summary of the invention

[0003] In view of this, the present application provides a target detection method and device based on an incremental learning perception model to solve the above-mentioned technical problems.

[0004] In a first aspect, an embodiment of the present application provides a target detection method based on an incremental learning perception model, which is applied to a vehicle, including:

[0005] Get the RGB image at the current moment captured by the camera on the vehicle;

[0006] Use the pre-trained region proposal network to process the current RGB image and obtain multiple proposal boxes;

[0007] Use the pre-trained Transformer encoding layer to process the current RGB image to obtain a deep semantic feature map;

[0008] Each proposal frame, deep semantic feature map and prototype vector of each target type are fused to obtain the fusion feature of each proposal frame; wherein the prototype vector of each target type is obtained by incremental learning according to the target type in the current scene of the vehicle;

[0009] The regional projection layer is used to process the fusion features of each proposal frame to obtain the target positioning result;

[0010] The classification layer is used to process the fusion features of each proposal box to obtain the target classification result.

[0011] In one possible implementation, the method further includes:

[0012] The prototype file recording the prototype vectors of each target type sent by the operating end is received, and the original prototype vectors of each target type stored in the vehicle end are updated.

[0013] In one possible implementation, the method further includes:

[0014] The target detection result of the current RGB image collected by the vehicle end is sent to the operation end, and the target detection result includes: target positioning result and target classification result.

[0015] In a possible implementation, the prototype file is generated by the operating end, and the generation steps include:

[0016] Get k RGB images containing the new target type from all RGB images at the current moment and before;

[0017] Use the SAM image segmentation algorithm to process k RGB images containing new target types and obtain target masks of k new target types;

[0018] Use the prototype network to process the RGB images of k new target types and their corresponding target masks to obtain the prototype vectors of the new target types;

[0019] Add the new target type and its prototype vector to the prototype file and send the updated prototype file to the ego vehicle.

[0020] In one possible implementation, each proposal box, deep semantic feature map and prototype vector of each target type are fused to obtain the fused features of each proposal box, including:

[0021] Align the deep semantic feature map with each proposal box to obtain the proposal features of each proposal box;

[0022] The proposed features of each proposal box are dot-producted with the prototype vector of each target type to obtain the fused features of each proposal box.

[0023] In one possible implementation, the method further includes: a step of jointly training the region proposal network, the Transformer encoding layer, and the prototype network.

[0024] In one possible implementation, the steps of jointly training the region proposal network, the Transformer encoding layer, and the prototype network include:

[0025] Obtain a training set, wherein the training set includes multiple RGB image samples annotated with real target frames and target categories;

[0026] Use the region proposal network to process the RGB image samples and obtain multiple proposal box samples;

[0027] Use the pre-trained Transformer encoding layer to process the RGB image samples to obtain deep semantic feature map samples;

[0028] Use the SAM image segmentation network to process RGB image samples and obtain multiple object masks;

[0029] Using a prototype network to process multiple RGB images and their corresponding target masks to obtain a prototype vector of a basic target type; wherein the basic target type is a target category in the RGB image sample;

[0030] The proposal box samples, deep semantic feature map samples and prototype vectors of basic target types are fused to obtain the fusion features of each proposal box sample;

[0031] The regional projection layer is used to process the fusion features of each proposal box sample to obtain multiple predicted target boxes;

[0032] Using multiple predicted target frames and true target frames, calculate a first loss value;

[0033] The classification layer is used to process the fusion features of each proposal box sample to obtain the predicted classification result of the target;

[0034] Using the predicted classification result of the target and the actual target category, calculate the second loss value;

[0035] The first loss value and the second loss value are used to update the parameters of the region proposal network, the Transformer encoding layer, and the prototype network.

[0036] In a second aspect, an embodiment of the present application provides a target detection device based on an incremental learning perception model, which is applied to a vehicle, including:

[0037] An acquisition unit, used to acquire the RGB image at the current moment captured by the camera on the vehicle;

[0038] A first processing unit is used to process the RGB image at the current moment using a pre-trained region proposal network to obtain multiple proposal boxes;

[0039] The second processing unit is used to process the RGB image at the current moment using the pre-trained Transformer encoding layer to obtain a deep semantic feature map;

[0040] A fusion unit is used to fuse each proposal frame, the deep semantic feature map and the prototype vector of each target type to obtain a fusion feature of each proposal frame; wherein the prototype vector of each target type is obtained by incremental learning according to the target type in the current scene of the vehicle;

[0041] The positioning unit is used to process the fusion features of each proposal box using the regional projection layer to obtain the target positioning result;

[0042] The classification unit is used to process the fusion features of each proposal box using the classification layer to obtain the target classification result.

[0043] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the target detection method based on the incremental learning perception model of the embodiment of the present application is implemented.

[0044] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the target detection method based on the incremental learning perception model of the embodiment of the present application is implemented.

[0045] This application improves the object detection accuracy of RGB images collected by the vehicle in unknown scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0047] Figure 1 A flowchart of a target detection method based on an incremental learning perception model provided in an embodiment of the present application;

[0048] Figure 2 A schematic diagram of an incremental perception system provided in an embodiment of the present application;

[0049] Figure 3 A functional structure diagram of a target detection device based on an incremental learning perception model provided in an embodiment of the present application;

[0050] Figure 4 A functional structure diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.

[0052] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for which protection is sought, but merely represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.

[0053] First, the design concept of the embodiments of the present application is briefly introduced.

[0054] Existing visual perception model learning methods are mostly targeted at situations with a large amount of prior information or training samples. There is still a lack of effective perception model learning strategies for a small amount of prior target information, which cannot fully meet the needs of real-time perception and recognition in complex environments. The scarcity and imbalance of data pose great challenges to model training and optimization, and effective methods need to be found to enhance the detection performance of the model. Therefore, how to incrementally learn the model based on a small amount of prior target information, maintain model performance and avoid catastrophic forgetting is a technical problem that needs to be solved urgently.

[0055] To this end, the present application proposes a target detection method based on an incremental learning perception model. First, real-time data is obtained, including an image to be detected and a small number of prompt images, wherein a small number of prompt images refer to key frame images containing new target categories. The image to be detected is regarded as a query image, and a small number of prompt images are regarded as support images. Then, the query image is used to generate an initial proposal box through a region proposal network (RPN). Next, the feature map of the query image and the support image is extracted using the Transformer encoding layer. By calculating the average feature of the support image, the Prototype feature vector of the new category is generated, and these prototype feature vectors represent the central feature of each category. Then, the query image feature map, the class prototype feature vector and the proposal box are fused so that the fused feature contains more semantic information and spatial information. Subsequently, the regional projection layer uses the fused features to convert the initial proposal into accurate bounding box coordinates to ensure the accuracy of target positioning. Finally, in the classification layer, the target is classified according to the fused features to generate the target's category label and confidence score. Through these steps, the incremental detection algorithm can achieve efficient and accurate target detection and classification in complex environments.

[0056] The advantages of this application are:

[0057] 1. The adopted SAM image segmentation algorithm can obtain the mask and text annotation information of new targets in the image in real time through human-computer interaction. It performs well in the face of data scarcity and category updates, and provides accurate annotation data for the model learning process. During the testing process, the image segmentation accuracy and speed were tested through on-site testing. After testing, the image segmentation accuracy reached more than 95% for both large and small targets; the time cost of the segmentation speed was less than 1s, which fully met the requirements of instant detection performance.

[0058] 2. The prototype feature extraction algorithm adopts the idea of ​​metric learning, and generates representative prototype features by learning the feature distribution of a small number of prior samples; it not only performs well in the face of data scarcity and category updates, but also captures the commonality between tasks through the multi-task training process of meta-learning, further enhancing the generalization ability of the model. During the test process, the speed of prototype feature file generation was tested. When faced with small sample image data, the actual test time cost was less than 1s, which fully met the requirements of instant detection performance.

[0059] 3. The incremental learning algorithm of the perception model with a small amount of prior target information is adopted. Based on the pre-trained model, the incremental learning of the model under few samples is realized by adopting the prototype network method based on meta-learning. By introducing the prototype module on the basic model, the detection performance is quickly improved without training and fine-tuning the model, while still having a high recall rate, avoiding the catastrophic forgetting problem of full fine-tuning; ensuring that the system has the ability to continuously learn and adapt to new environments, and can respond quickly when new targets and data continue to appear.

[0060] After introducing the application scenarios and design concepts of the embodiments of the present application, the technical solutions provided by the embodiments of the present application are described below.

[0061] like Figure 1 As shown, the present application implements a target detection method based on an incremental learning perception model, which is applied to the vehicle side, including:

[0062] Step 101: Obtain the RGB image at the current moment captured by the camera on the vehicle;

[0063] Step 102: Use the pre-trained region proposal network to process the RGB image at the current moment to obtain multiple proposal boxes;

[0064] For example, the Region Proposal Network (RPN) can generate an initial proposal box for a target in an RGB image. Specifically, the RGB image is regarded as a query image, and the RPN is responsible for generating multiple candidate regions in the image, which may contain the target object. The RPN generates a series of anchor boxes on the image through a sliding window mechanism. These anchor boxes have different scales and aspect ratios to accommodate various possible target sizes and shapes. For each anchor box, the RPN calculates the probability that it contains the target and adjusts it to locate the target more accurately.

[0065] Step 103: Use the pre-trained Transformer encoding layer to process the RGB image at the current moment to obtain a deep semantic feature map;

[0066] Exemplarily, the Transformer Encoder Layer is used to extract the feature map of the RGB image. The Transformer Encoder Layer can capture the deep semantic features of the image and extract the feature map of the query image, which plays an important role in the subsequent fusion processing. At the same time, the same Transformer Encoder Layer is used to extract the feature map of the support image, and the feature vector is obtained through the feature map for the calculation of the prototype feature vector of the new category. Here, the backbone network of the Transformer-based pre-trained Dinov2 ViT model is used as the Transformer Encoder Layer. Specifically, the Transformer Encoder Layer processes the RGB image in blocks, and calculates the relationship between each image block through a multi-head self-attention mechanism to generate a high-dimensional feature representation. These feature representations not only contain the global information of the image, but also retain fine-grained local features, thereby improving the expressive power of the model.

[0067] Step 104: fusing each proposal frame, the deep semantic feature map and the prototype vector of each target type to obtain a fusion feature of each proposal frame; wherein the prototype vector of each target type is obtained by incremental learning according to the target type in the current scene of the vehicle;

[0068] Step 105: Use the region projection layer to process the fusion features of each proposal box to obtain the target positioning result;

[0069] Exemplary, Region Projection Layer: used to further correct the proposal box proposed by the region proposal network to generate accurate bounding box coordinates. For the target object in the query image, the detection box generated by the region proposal network will have a certain deviation, which affects the precise positioning of the target. By expanding the proposal box, and then using the rich semantic information and spatial information contained in the fusion features, the initial proposal is diffused to the target area at the pixel level in the form of a heat map conversion, and finally the target area is projected to the accurate bounding box through the learnable spatial integration layer. Specifically, the learnable spatial integration layer calculates the center position and size of the bounding box by weighted averaging the probability distribution of the target area. This method avoids the gradient explosion and vanishing problems that may exist in traditional regression methods, thereby achieving more accurate target positioning.

[0070] Step 106: Use the classification layer to process the fusion features of each proposal box to obtain the target classification result.

[0071] Exemplary, the classification layer is used to classify targets based on the fused features obtained by fusion. Since in this incremental detection system, it is necessary to continuously introduce prototype feature vectors of new categories to achieve the incremental effect, a multi-layer binary network is used in the classification layer to calculate the classification score of each category relative to other categories. These scores are converted into probability distributions through the softmax function, indicating the probability that the proposal box belongs to each category. Finally, the system classifies each proposal box according to these probability distributions. In this way, the system can accurately identify and classify targets of new categories under few sample conditions, and achieve efficient target detection and classification.

[0072] Exemplary target types include: pedestrians, bicycles, electric vehicles, trucks, trailers, cars, and animals.

[0073] In some embodiments, the method further comprises:

[0074] The prototype file recording the prototype vectors of each target type sent by the operating end is received, and the original prototype vectors of each target type stored in the vehicle end are updated.

[0075] In some embodiments, the method further comprises:

[0076] The target detection result of the current RGB image collected by the vehicle end is sent to the operation end, and the target detection result includes: target positioning result and target classification result.

[0077] In some embodiments, the prototype file is generated by the operating end, and the generation steps include:

[0078] Get k RGB images containing the new target type from all RGB images at the current moment and before;

[0079] Use the SAM image segmentation algorithm to process k RGB images containing new target types and obtain target masks of k new target types;

[0080] Use the prototype network to process the RGB images of k new target types and their corresponding target masks to obtain the prototype vectors of the new target types;

[0081] Add the new target type and its prototype vector to the prototype file and send the updated prototype file to the ego vehicle.

[0082] In this embodiment, a small amount of image data containing the new target is collected as a prompt image for the perception model, and the SAM image segmentation algorithm is used to perform mask extraction on the new target in the image through human-computer interaction. The label data of the new target object is formed by combining the category text label, which provides input data for the subsequent extraction of prototype features.

[0083] SAM can achieve image segmentation of any target based on instruction information and image data. It draws on the Prompt strategy in the NLP field and provides prompts for image segmentation tasks to complete the rapid segmentation of any target. The prompt type can be a foreground / background point, a rough box, a mask, or any form of text. The input of the model is the original image and prompt information, and the goal is to output a "valid" segmentation. The so-called valid means that when the prompt is vague, the model can output at least one mask of the target.

[0084] The SAM model consists of three core components: Image Encoder, Prompt Encoder, and Mask Decoder. The image is encoded by the Image Encoder, the prompt is encoded by the Prompt Encoder, and the two embeddings are then passed through a lightweight Mask Decoder to obtain the fused features. The Encoder part uses an existing model, and the Decoder part uses Transformer.

[0085] Component 1: Image Encoder

[0086] The role of the Image Encoder is to map the image to the feature space. Essentially, this Encoder can be any network structure. Here we use the fine-tuned ViT model, or we can use the traditional convolutional structure.

[0087] Based on the vit_b version of the ViT model, the input original image is supplemented into a square according to the maximum side length, and then the size is resized to 3x1024x1024. The process through the ViT structure is as follows:

[0088] Patch Embedding. The input image passes through a convolution base, which divides the image into 16x16 patches with a stride of 16. This reduces the size of the feature map by 16 times, and maps the channel from 3 to 768.

[0089] Positional Embedding. After Patch Embedding, the output tokens need to be encoded with positional encoding to preserve the spatial information of the image. Positional encoding can be understood as a map. The number of rows in the map is the same as the number of input sequences. Each row represents a vector. The dimension of the vector is the same as the dimension of the input sequence tokens. The operation of positional encoding is sum, so the dimension remains unchanged.

[0090] Transformer Encoder. The feature map passes through 16 Transformer Blocks, 12 of which use the attention mechanism based on Window Partition (dividing the feature map into 14*14 windows for local attention) to process local information. The other 4 blocks are global attention modules, which are interspersed between WindowPartition modules to capture the global context of the image.

[0091] Partition operation. In the non-global attention block, in order to adapt to the 14x14 window size, the input feature map needs to be padded and split. The specific process is as follows:

[0092] ① Input feature map: The initial size of the input feature map is 1x64x64x768.

[0093] ② Determine the minimum divisible size: The window size is 14*14, and we need to find the minimum feature map size that is divisible by 14. For width and height, we need to find the minimum number that is greater than or equal to 64 and divisible by 14. These two numbers are 70(64+6) and 70(64+6), so the minimum divisible feature map size is 1x70x70x768.

[0094] ③Padding: In order to expand the feature map size from 64x64 to 70x70, a 6x6 area needs to be filled in the lower right corner, because 70-64=6. This padding method ensures that the window can be correctly divided at the edge of the feature map.

[0095] ④ Split feature map: Split the features after padding Figure 1 x70x70x768 is split into 14x14 windows. Since 70 / 14=5, the feature map can be split into 5x5 windows of 14x14, for a total of 5x5=25 windows. The size of each window is 14x14x768.

[0096] Unpartition operation. In the non-global attention block, the features output by the attention layer are Figure 1 The conversion of x70x70x768 into a 1x64x64x768 feature map is actually done by taking out the 1x64x64x768 part in the upper left corner from the 1x70x70x768 feature map through the slicing operation x=x[:1,:64,:64,:].

[0097] Neck Convolution: Through two layers of convolution (Neck), the channel is reduced to 256 and the final ImageEmbedding is generated.

[0098] Component 2: Prompt Encoder

[0099] Prompt Encoder maps the input prompt information to the corresponding feature space. The channel of its feature is consistent with the channel of image embedding. The input data type of this component can be point, box, mask or text. The present invention uses box data as prompt input to generate the sparse_embedding required by the subsequent MaskEncoder.

[0100] The bounding box generally has two points, and the encoding process is as follows:

[0101] Transfer the input marker box corner coordinates from the upper right corner of the pixel to the center of the pixel.

[0102] The adjusted bounding box corner coordinates and the size of the input image are taken as input, and a high-dimensional embedding representation corner_Embedding is obtained using a random Gaussian matrix and sine and cosine encoding.

[0103] After adding corner_Embedding to a learnable embedding vector, we get sparseembedding as the output of Prompt Encoder.

[0104] Component 3: Mask Decoder

[0105] The role of Mask Encoder is to use the output of Image Encoder and Prompt Encoder as information source, and predict the segmentation mask after deep fusion.

[0106] The key idea of ​​Mask Encoder lies in the self-attention mechanism and cross-attention mechanism of Transformer. Specifically, in a Transformer block, one self-attention and three cross-attention are used to repeatedly fuse and update the previously acquired input information image_embedding and sparse_embedding. At the same time, in order to ensure that the position information is never lost, position encoding is introduced multiple times to ensure that accurate segmentation results are obtained in the end.

[0107] In this embodiment, a picture containing a new target is used as a prompt picture, and the SAM image segmentation algorithm is used to extract the mask of the target through manual interactive input of box data, which is then combined with the category text label to form a small amount of annotation data of the new target object in the prompt image, providing a data basis for the subsequent extraction of the prototype of the new target.

[0108] The workflow of the Prototype network includes: generating a prototype feature vector for each basic category in the learning phase. This vector is obtained by calculating the mean of the feature vectors of all samples of the category and represents the central feature of the category. In the detection phase, after the model receives a new category sample, it first extracts the feature vector of the sample through the pre-trained model. Then, the prototype feature vector of the new category sample is calculated in the same way, and the prototype feature vectors of the new category and the basic category are combined for subsequent target prediction.

[0109] In the small sample classification task, suppose there are N labeled samples S = (x1, y1),..., (x N ,y N ),in is a sample feature vector of D dimensions, and y∈1,...,K is the corresponding label. K Represents the set of samples of the kth class. The prototype network calculates the M-dimensional prototype feature vector of each class The calculated function is Where φ is a learnable parameter. The prototype feature vector c k The calculation formula is as follows:

[0110]

[0111] The Prototype network method under the meta-learning framework solves the few-shot learning problem through metric learning, which can significantly improve the model's adaptability to new tasks.

[0112] In some embodiments, each proposal box, the deep semantic feature map and the prototype vector of each target type are fused to obtain the fusion features of each proposal box, including:

[0113] Align the deep semantic feature map with each proposal box to obtain the proposal features of each proposal box;

[0114] The proposed features of each proposal box are dot-producted with the prototype vector of each target type to obtain the fused features of each proposal box.

[0115] In some embodiments, the method further includes: a step of jointly training the region proposal network, the Transformer encoding layer, and the prototype network; specifically including:

[0116] Obtain a training set, wherein the training set includes multiple RGB image samples annotated with real target frames and target categories;

[0117] Use the region proposal network to process the RGB image samples and obtain multiple proposal box samples;

[0118] Use the pre-trained Transformer encoding layer to process the RGB image samples to obtain deep semantic feature map samples;

[0119] Use the SAM image segmentation network to process RGB image samples and obtain multiple object masks;

[0120] Using a prototype network to process multiple RGB images and their corresponding target masks to obtain a prototype vector of a basic target type; wherein the basic target type is a target category in the RGB image sample;

[0121] The proposal box samples, deep semantic feature map samples and prototype vectors of basic target types are fused to obtain the fusion features of each proposal box sample;

[0122] The regional projection layer is used to process the fusion features of each proposal box sample to obtain multiple predicted target boxes;

[0123] Using multiple predicted target frames and true target frames, calculate a first loss value;

[0124] The classification layer is used to process the fusion features of each proposal box sample to obtain the predicted classification result of the target;

[0125] Using the predicted classification result of the target and the actual target category, calculate the second loss value;

[0126] The first loss value and the second loss value are used to update the parameters of the region proposal network, the Transformer encoding layer, and the prototype network.

[0127] The specific implementation process of this application is described below in conjunction with a specific application scenario.

[0128] like Figure 2 As shown, the incremental perception system includes an operation end, a server end, and a vehicle end, wherein:

[0129] The vehicle side uses the camera to read image data in the real environment, covering complex scenes and environmental conditions. These image data are processed by the vehicle side's detection algorithm to identify the target in the image and generate detection results. These detection results are transmitted to the server side through the network. After receiving these images, the server side forwards them to the operation side for further processing. At the same time, the vehicle side also receives the prototype file from the server side and restarts the algorithm to adapt to the new target category.

[0130] The operator receives the vehicle images from the server and displays them on the QT interactive interface. The operator selects the target in the image using the segmentation algorithm and generates a mask for the target. Using these masks and image information, the operator generates a new prototype file and sends it back to the server via the network. After receiving the file, the server forwards it to the vehicle, which receives the new prototype file and restarts the detection algorithm to ensure that the system can adapt to the dynamic changes in the environment in real time.

[0131] The server is responsible for data transfer and transmission throughout the entire process. First, it receives the detection result images transmitted by the vehicle and forwards them to the operator. At the same time, the server also receives the prototype file generated by the operator and forwards it to the vehicle. To ensure that the new algorithm and prototype file are applied, the server sends a restart command to the vehicle after confirming that the file transfer is successful.

[0132] The specific workflow is as follows:

[0133] The on-board camera captures real-time images and performs preliminary processing through the vehicle-side detection algorithm to generate detection results. These detection results are transmitted to the server through the network. The server receives and stores these image data and then forwards them to the operator. The operator displays the received image on the QT interactive interface, generates the mask of the target using the prompt box selected on the image through the segmentation algorithm, and generates a new prototype file based on the mask and text category annotation. The operator sends the generated prototype file back to the server through the network, and the server receives and forwards it to the vehicle. After confirming that the file transfer is successful, the server sends a restart command to the vehicle. After receiving the new prototype file, the vehicle restarts the detection algorithm and uses the model after updating the prototype file to perform target detection to adapt to the new target category.

[0134] The above steps use the SAM segmentation algorithm to obtain the mask of the new target in the prompt image as annotation information. Based on the pre-trained model, an incremental perception system that adapts to the recognition of new targets is developed through a fast fine-tuning method that does not require training, which improves the system's recognition ability for new targets and ensures that target detection and recognition can still be effectively performed when data is limited. The system achieves efficient and accurate target detection and classification in complex real-world environments. The incremental detection algorithm combines real-time image acquisition, dynamic target detection, and model updates to ensure that the system can quickly adapt to dynamic changes in the real environment, providing solid technical support for modern computer vision detection and recognition.

[0135] Based on the above embodiments, the present application embodiment provides a target detection device based on an incremental learning perception model. Figure 3 As shown, the target detection device 200 based on the perception model of incremental learning provided in the embodiment of the present application at least includes:

[0136] An acquisition unit 201 is used to acquire an RGB image at the current moment captured by a camera on the vehicle;

[0137] A first processing unit 202 is used to process the RGB image at the current moment using a pre-trained region proposal network to obtain multiple proposal boxes;

[0138] The second processing unit 203 is used to process the RGB image at the current moment by using the pre-trained Transformer encoding layer to obtain a deep semantic feature map;

[0139] A fusion unit 204 is used to fuse each proposal frame, the deep semantic feature map and the prototype vector of each target type to obtain a fusion feature of each proposal frame; wherein the prototype vector of each target type is obtained by incremental learning according to the target type in the current scene of the vehicle;

[0140] A positioning unit 205 is used to process the fusion features of each proposal frame using the regional projection layer to obtain a target positioning result;

[0141] The classification unit 206 is used to process the fusion features of each proposal box using the classification layer to obtain a target classification result.

[0142] It should be noted that the principle of solving the technical problem by the target detection device 200 based on the incremental learning perception model provided in the embodiment of the present application is similar to the method provided in the embodiment of the present application. Therefore, the implementation of the target detection device 200 based on the incremental learning perception model provided in the embodiment of the present application can refer to the implementation of the method provided in the embodiment of the present application, and the repeated parts will not be repeated.

[0143] Based on the above embodiments, the present application also provides an electronic device, referring to Figure 4 As shown, the electronic device 300 provided in the embodiment of the present application includes at least: a processor 301, a memory 302, and a computer program stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program, the target detection method based on the incremental learning perception model provided in the embodiment of the present application is implemented.

[0144] The electronic device 300 provided in the embodiment of the present application may further include a bus 303 connecting different components (including the processor 301 and the memory 302). The bus 303 represents one or more of several types of bus structures, including a memory bus, a peripheral bus, a local bus, and the like.

[0145] The memory 302 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 3021 and / or a cache memory 3022 , and may further include a read-only memory (ROM) 3023 .

[0146] The memory 302 may also include a program tool 3025 having a set (at least one) of program modules 3024, the program modules 3024 including but not limited to: an operating subsystem, one or more applications, other program modules and program data, each of which or some combination may include the implementation of a network environment.

[0147] The electronic device 300 may also communicate with one or more external devices 304 (e.g., keyboards, remote controls, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device 300 (e.g., mobile phones, computers, etc.), and / or communicate with any device that enables the electronic device 300 to communicate with one or more other electronic devices 300 (e.g., routers, modems, etc.). Such communication may be performed through an input / output (I / O) interface 305. Furthermore, the electronic device 300 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 306. Figure 4 As shown, the network adapter 306 communicates with other modules of the electronic device 300 via the bus 303. It should be understood that although Figure 4 Not shown, other hardware and / or software modules may be used in conjunction with the electronic device 300, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, disk arrays (Redundant Arrays of Independent Disks, RAID) subsystems, tape drives, and data backup storage subsystems.

[0148] It should be noted that Figure 4 The electronic device 300 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0149] The embodiment of the present application also provides a computer-readable storage medium, which stores computer instructions, and when the computer instructions are executed by the processor, the target detection method based on the incremental learning perception model provided in the embodiment of the present application is implemented. Specifically, the executable program can be built-in or installed in the electronic device 300, so that the electronic device 300 can implement the target detection method based on the incremental learning perception model provided in the embodiment of the present application by executing the built-in or installed executable program.

[0150] The target detection method based on the incremental learning perception model provided in the embodiment of the present application can also be implemented as a program product, which includes a program code. When the program product can be run on the electronic device 300, the program code is used to enable the electronic device 300 to execute the target detection method based on the incremental learning perception model provided in the embodiment of the present application.

[0151] The program product provided in the embodiments of the present application may adopt any combination of one or more readable media, wherein the readable medium may be a readable signal medium or a readable storage medium, and the readable storage medium may be but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. Specifically, more specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, RAM, ROM, Erasable Programmable Read Only Memory (EPROM), optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0152] The program product provided in the embodiment of the present application may adopt a CD-ROM and include program code, and may also be run on a computing device. However, the program product provided in the embodiment of the present application is not limited thereto. In the embodiment of the present application, the readable storage medium may be any tangible medium containing or storing a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.

[0153] It should be noted that, although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided into multiple units to be embodied.

[0154] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present application and are not intended to limit it. Although the present application is described in detail with reference to the embodiments, a person skilled in the art should understand that any modification or equivalent replacement of the technical solution of the present application does not depart from the spirit and scope of the technical solution of the present application and should be included in the scope of the claims of the present application.

Claims

1. A target detection method based on an incremental learning perception model, applied to a vehicle, characterized in that: include: Get the RGB image at the current moment captured by the camera on the vehicle; Use the pre-trained region proposal network to process the current RGB image and obtain multiple proposal boxes; Use the pre-trained Transformer encoding layer to process the current RGB image to obtain a deep semantic feature map; Each proposal frame, deep semantic feature map and prototype vector of each target type are fused to obtain the fusion feature of each proposal frame; wherein the prototype vector of each target type is obtained by incremental learning according to the target type in the current scene of the vehicle; The regional projection layer is used to process the fusion features of each proposal frame to obtain the target positioning result; The classification layer is used to process the fusion features of each proposal box to obtain the target classification result.

2. The target detection method based on the perception model of incremental learning according to claim 1 is characterized in that: The method further comprises: The prototype file recording the prototype vectors of each target type sent by the operating end is received, and the original prototype vectors of each target type stored in the vehicle end are updated.

3. The target detection method based on the perception model of incremental learning according to claim 2 is characterized in that: The method further comprises: The target detection result of the current RGB image collected by the vehicle end is sent to the operation end, and the target detection result includes: target positioning result and target classification result.

4. The target detection method based on the perception model of incremental learning according to claim 3 is characterized in that: The prototype file is generated by the operation end, and the generation steps include: Get k RGB images containing the new target type from all RGB images at the current moment and before; Use the SAM image segmentation algorithm to process k RGB images containing new target types and obtain target masks of k new target types; Use the prototype network to process the RGB images of k new target types and their corresponding target masks to obtain the prototype vectors of the new target types; Add the new target type and its prototype vector to the prototype file and send the updated prototype file to the ego vehicle.

5. The target detection method based on the perception model of incremental learning according to claim 4 is characterized in that: Each proposal box, deep semantic feature map and prototype vector of each target type are fused to obtain the fusion features of each proposal box, including: Align the deep semantic feature map with each proposal box to obtain the proposal features of each proposal box; The proposed features of each proposal box are dot-producted with the prototype vector of each target type to obtain the fused features of each proposal box.

6. The target detection method based on the perception model of incremental learning according to claim 5 is characterized in that: The method further comprises: Steps for jointly training the region proposal network, Transformer encoding layer, and prototypical network.

7. The target detection method based on the incremental learning perception model according to claim 6, characterized in that: The steps for jointly training the region proposal network, Transformer encoding layer, and prototype network include: Obtain a training set, wherein the training set includes multiple RGB image samples annotated with real target frames and target categories; Use the region proposal network to process the RGB image samples and obtain multiple proposal box samples; Use the pre-trained Transformer encoding layer to process the RGB image samples to obtain deep semantic feature map samples; Use the SAM image segmentation network to process RGB image samples and obtain multiple object masks; Using a prototype network to process multiple RGB images and their corresponding target masks to obtain a prototype vector of a basic target type; wherein the basic target type is a target category in the RGB image sample; The proposal box samples, deep semantic feature map samples and prototype vectors of basic target types are fused to obtain the fusion features of each proposal box sample; The regional projection layer is used to process the fusion features of each proposal box sample to obtain multiple predicted target boxes; Using multiple predicted target frames and true target frames, calculate a first loss value; The classification layer is used to process the fusion features of each proposal box sample to obtain the predicted classification result of the target; Using the predicted classification result of the target and the actual target category, calculate the second loss value; The first loss value and the second loss value are used to update the parameters of the region proposal network, the Transformer encoding layer, and the prototype network.

8. A target detection device based on an incremental learning perception model, applied to a vehicle, characterized in that: include: An acquisition unit, used to acquire the RGB image at the current moment captured by the camera on the vehicle; A first processing unit is used to process the RGB image at the current moment using a pre-trained region proposal network to obtain multiple proposal boxes; The second processing unit is used to process the RGB image at the current moment using the pre-trained Transformer encoding layer to obtain a deep semantic feature map; A fusion unit is used to fuse each proposal frame, the deep semantic feature map and the prototype vector of each target type to obtain a fusion feature of each proposal frame; wherein the prototype vector of each target type is obtained by incremental learning according to the target type in the current scene of the vehicle; The positioning unit is used to process the fusion features of each proposal box using the regional projection layer to obtain the target positioning result; The classification unit is used to process the fusion features of each proposal box using the classification layer to obtain the target classification result.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • A RGBD image target recognition method based on depth learning

    CN109543697A

  • Incremental information classification method based on prototype

    CN111667016A

  • Multi-mode and adversarial learning-based multi-task target detection and identification method and device

    CN114821014A

  • Variable illumination traffic target detection method based on RGB and event camera fusion

    CN117274945A

  • Low-computing-power automatic driving real-time multi-task sensing method and device

    CN117372983A

Cited By

  • Multi-view target detection method and device, computer equipment and storage medium

    CN120496126A