Object Detection Method and System Based on Prompt Engineering and Regional Text Description

By introducing prompt engineering propt and image text generation decoder into the object detection model, fine-grained features of the target area and generating attribute text descriptions, the feature conflict problems caused by subdivision classification and area attribute judgment problems are solved, and the model performance and event judgment capabilities are improved.

CN118609149BActive Publication Date: 2025-05-27CHENGDU ZHIHUI HENENG CITY TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410648143.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-23
Publication Date
2025-05-27
Estimated Expiration
2044-05-23

AI Technical Summary

Technical Problem

The existing target detection methods based on subcategorization lead to feature conflicts, resulting in reduced model performance, and cannot provide a description of target area attributes, which cannot effectively solve the two types of problems of category attributes of smart cities.

Method used

A target detection model with area attribute description is designed. By introducing prompt engineering propt and image text generation decoder decoder, fine-grained features of the detection target are extracted, and the image text generation decoder is used to obtain the target attribute text description of the corresponding prediction box.

Benefits of technology

The problem of object detection sub-category affecting model performance reduction and regional attribute judgment is solved, the average accuracy mean (mAP) of object detection tasks is improved, and the content of detection target area description is provided to effectively solve the problem of sub-category and event judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118609149B_ABST
    Figure CN118609149B_ABST
Patent Text Reader

Abstract

The present invention discloses a target detection method and system based on prompt engineering and regional text description, the method comprising: constructing a labeled data set, the labeled data set comprising a plurality of labeled images; constructing a target detection model with regional attribute description, the target detection model with regional attribute description is a new model formed by adding an adaptive regional feature extraction module and a text generation decoding module on the basis of the RTDetr model structure; based on the labeled data set, the target detection model with regional attribute description is trained to obtain a trained target detection model with regional attribute description; obtaining an image to be detected, using the trained target detection model with regional attribute description to perform target detection on the image to be detected, and obtaining a target prediction frame, a prediction category and a regional attribute description. The present invention solves the problem that target detection subclassification affects the performance reduction of the model and the category attribute judgment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to a target detection method and system based on prompt engineering and regional text description. Background Art

[0002] The smart city vehicle-mounted edge AI detection system realizes target detection through the target detection model. Usually, the edge device will use the CNN-based YOLO series model or the Transformer-based end-to-end target detection DETR series model to realize the detection of city-related objects with the help of sub-classification data, and use its corresponding category or add some logic to judge the event. It is difficult to solve the two types of problems of smart city category attribute distinction by relying only on a large amount of data and a single target detection model: first, the sub-classification conflict problem, which is manifested as a category conflict with extremely similar features (such as sedans and large sedans SUVs), resulting in reduced model accuracy and even missed detection; second, the category attribute event problem, the detection model can only give the category without additional attribute prompts, such as the detection model cannot distinguish whether the car category is parked on the sidewalk or lane, and still cannot distinguish the city management event (sidewalk and lane cars are different events). Of course, the existing methods also use multiple identical target detection models, each target detection model is responsible for a certain number of category detections, and relies on scene logic combination to realize urban management sub-classification and event judgment. This can improve the performance of the detection model to a certain extent, but it cannot fundamentally solve the problem of urban management, but instead increases the resource consumption of edge devices. At the same time, existing methods use detailed classification for target detection, which causes model feature conflicts and insufficient feature expression, loss of model performance, and lack of text-related descriptions, making it impossible to judge urban management events for category-attached attributes.

[0003] Therefore, the smart city vehicle-mounted edge AI detection system relies on a large amount of data and subdivided categories, uses a single or multiple identical target detection models, and relies on scene logic combination to achieve smart city object category and event judgment. However, the above existing target detection methods are based on subdivision, which will lead to feature conflicts and thus reduce model performance, and cannot provide category attribute assignment for urban management event judgment. Summary of the invention

[0004] The technical problem to be solved by the present invention is that the existing object detection method based on fine classification will lead to feature conflicts and thus reduce the model performance, and it cannot provide the judgment of the target area attribute description problem. The purpose of the present invention is to provide an object detection method and system based on prompt engineering and regional text description. The present invention focuses on designing a new architecture of an object detection model: an object detection model with regional attribute description. Specifically, a prompt and an image text generation decoder are introduced into the existing object detection model. With the help of the prompt, the prediction boxes obtained by the existing detection model are embedded and encoded as accurate feature extraction query vectors Q. The features extracted by the existing detection model are used as key vectors K and numerical vectors V to extract the fine-grained features of the detection target area, and then the fine-grained features are used by the image text generation decoder to obtain the corresponding prediction box target attribute text description. The present invention solves the problems of the reduction of the model performance affected by the fine classification of object detection and the regional attribute judgment problem.

[0005] The present invention is realized through the following technical solutions:

[0006] In the first aspect, the present invention provides an object detection method based on prompt engineering and regional text description, and the method includes:

[0007] Construct a labeled data set, which includes multiple labeled image object detection labels and labeled image target area text descriptions; the labeled image object detection labels are used for detecting category labels and position labels with scene objects, and the category labels are labeled with superclasses (i.e., large classes); the labeled image target area text descriptions are short text descriptions of the corresponding detection target areas, including regional category attribute text descriptions and regional event attribute text descriptions, and they are integrated to construct text descriptions; the regional category attribute text descriptions are content descriptions of the subclasses of the corresponding detection target area superclass; the regional event attribute text descriptions are content descriptions of the events in the corresponding detection target area;

[0008] Construct an object detection model with regional attribute description. The object detection model with regional attribute description is a new model with a text description branch structure constructed by adding an adaptive regional feature extraction module and a text generation decoding module on the basis of not changing the structure of the RTDetr detection model;

[0009] Based on the labeled dataset, use the labeled image object detection annotation to train the RTDetr detection model for the detection task as the first-stage model training; use the labeled image target region text description to train the object detection model with region attribute description for the region short text description task as the second-stage model training, where the RTDetr detection model in the second stage uses the weights obtained from the first-stage training and freezes the RTDetr model structure to obtain the trained object detection model with region attribute description;

[0010] Obtain the image to be detected, and use the trained object detection model with region attribute description to perform object detection on the image to be detected to obtain the target prediction box, prediction category, and region attribute description.

[0011] Furthermore, the superclass is to merge the fine-grained categories with high feature similarity into a large class.

[0012] Furthermore, the object detection model with region attribute description is based on the RTDetr model as a benchmark, and the added adaptive region feature extraction module APF and text generation decoding module TGD are integrated with the RTDetr model to construct a new model for region description object detection based on prompt engineering and image description caption;

[0013] The adaptive region feature extraction module APF, based on the features extracted by the backbone network module in the RTDetr model and the target prediction box output by the prediction module (Head), converts the region coordinate information of the target prediction box into a vector representation, and further extracts the features of the RTDetr model according to the vector representation to obtain the region-related feature representation;

[0014] The text generation decoding module TGD, based on the region-related feature representation, uses the Transformer structure to decode the region-related feature representation to obtain the region attribute description.

[0015] Furthermore, converting the region coordinate information of the target prediction box into a vector representation and further extracting the features of the RTDetr model according to the vector representation to obtain the region-related feature representation includes:

[0016] Based on prompt engineering, convert the region coordinate information of the target prediction box into a vector representation X b ;

[0017] Take the vector representation X b as the query vector Q of the Transformer structure, and take the features of the RTDetr model as the key vector K and value vector V of the Transformer structure; use the Transformer structure to first process the vector representation X bPerform self-attention mechanism encoding, then perform cross-attention mechanism encoding on the features of the RTDetr model, and then use the FFN structure to repeat n times to extract region-related feature expressions.

[0018] Furthermore, the formula for the extracted region-related feature expression is:

[0019] F p = f n f ca (f sa (F t , X b ))

[0020] where f sa represents the self-attention mechanism structure, f ca represents the cross-attention mechanism structure, f n is the FFN structure, F t is the feature of the RTDetr model, and X b is the vector expression obtained by converting the regional coordinate information of the target prediction box.

[0021] Furthermore, the region attribute description is based on the text attribute data of the prediction box.

[0022] Furthermore, based on the labeled dataset, the RTDetr detection model is trained in the first stage using the labeled image target detection annotations; the target detection model with region attribute description is trained in the second stage using the labeled image target region text descriptions. In the second stage, the RTDetrr detection model uses the weights obtained in the first stage training and freezes the RTDetr detection model structure to obtain the trained target detection model with region attribute description, including:

[0023] The first training stage: Based on the superclass data, train the RTDetr model to obtain the prediction classes and bounding box prediction tasks of the RTDetr model, and save the weights W r of the RTDetr model; the superclass data does not contain the corresponding event attribute data descriptions;

[0024] The second training stage: Assign the weights W r saved in the first stage training to the corresponding weights of the target detection model with region attribute description, and freeze its corresponding RTDetr model structure; train the target detection model with region attribute description based on the labeled dataset and only train the adaptive region feature extraction module APF and the text generation and decoding module TGD, and update the weights of the adaptive region feature extraction module APF and the text generation and decoding module TGD to maintain the target detection ability of the original RTDetr model and obtain region attribute descriptions;

[0025] After the training is completed, save the best weights of the object detection model with regional attribute descriptions to obtain a trained object detection model with regional attribute descriptions.

[0026] In a second aspect, the present invention further provides an object detection system based on prompt engineering and regional text descriptions. This system uses the above-mentioned object detection method based on prompt engineering and regional text descriptions. The system includes:

[0027] A dataset construction unit for constructing an annotated dataset. The annotated dataset includes multiple annotated image object detection annotations and annotated image object region text descriptions. The annotated image object detection annotations are used to perform detection category annotations and position annotations on scene objects. The category annotations are performed using superclasses. The annotated image object region text descriptions are short text descriptions of the regions corresponding to the detection targets, including region category attribute text descriptions and region event attribute text descriptions. The region category attribute text descriptions are content descriptions of the subclasses of the superclasses corresponding to the detection target regions. The region event attribute text descriptions are content descriptions of the events corresponding to the detection target regions.

[0028] A model construction unit for constructing an object detection model with regional attribute descriptions. The object detection model with regional attribute descriptions is a new model with a text description branch structure constructed by adding an adaptive region feature extraction module APF and a text generation decoding module TGD on the basis of not changing the structure of the RTDetr detection model.

[0029] A model training unit for, based on the annotated dataset, using the annotated image object detection annotations to train the RTDetr detection model for detection tasks as the first-stage model training; using the annotated image object region text descriptions to train the object detection model with regional attribute descriptions for short text description tasks of regions as the second-stage model training. In the second stage, the RTDetrr detection model uses the weights obtained from the first-stage training and freezes the structure of the RTDetr detection model to obtain a trained object detection model with regional attribute descriptions.

[0030] An object detection unit for obtaining an image to be detected and using the trained object detection model with regional attribute descriptions to perform object detection on the image to be detected to obtain object prediction boxes, prediction categories, and regional attribute descriptions.

[0031] Furthermore, the object detection model with regional attribute description is based on the RTDetr model. The added adaptive regional feature extraction module APF and text generation and decoding module TGD are integrated with the RTDetr model to construct a new model for regional description object detection based on prompt engineering and image description caption.

[0032] The adaptive regional feature extraction module APF is used to convert the regional coordinate information of the target prediction box into a vector representation according to the features extracted by the backbone network module in the RTDetr model and the target prediction box output by the prediction module (Head), and further extract the features of the RTDetr model according to the vector representation to obtain region-related feature expressions.

[0033] The text generation and decoding module TGD is used to decode the region-related feature expressions using the Transformer structure according to the region-related feature expressions to obtain regional attribute descriptions.

[0034] In a third aspect, the present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above object detection method based on prompt engineering and regional text description.

[0035] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0036] 1. The object detection method and system based on prompt engineering and regional text description of the present invention focuses on designing a new architecture of an object detection model: an object detection model with regional attribute description. The present invention introduces prompt engineering and an image text generation decoder into the existing object detection model. With the help of prompt engineering, the prediction box obtained by the existing detection model is embedded and encoded as an accurate feature extraction query vector Q, and the features extracted by the existing detection model are used as the key vector K and the numerical vector V to extract the fine-grained feature Fc of the detection target. Then, the fine-grained feature Fc is used with the image text generation decoder to obtain the target attribute text description corresponding to the prediction box. The present invention solves the problems of the reduction of model performance affected by the fine classification of object detection and the judgment of category attributes.

[0037] 2. For the object detection method and system based on prompt engineering and regional text description of the present invention, the training of the object detection model PG-RTDetr with regional attribute description is divided into two stages. The first training stage is to train the RTDetr model based on superclass data. The second training stage is to use the weight W saved in the first stage of training. rAssign the weights to the corresponding weights of the PG-RTDetr model and freeze its corresponding RTDetr model structure; then train the PG-RTDetr model based on the labeled dataset and only train the Adaptive Region Feature Extraction Module (APF) and the Text Generation and Decoding Module (TGD). The above training strategy not only enhances the object detection task ability of the original RTDetr model, but also adds an additional text description of the detected target area, effectively improving the mean Average Precision (mAP) of the object detection task and providing the description content of the detected target area, effectively solving the problems of fine classification and event judgment. Brief Description of the Drawings

[0038] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not limit the embodiments of the present invention. In the drawings:

[0039] Figure 1 It is a flowchart of the object detection method based on prompt engineering and regional text description of the present invention;

[0040] Figure 2 It is a schematic diagram of the Transformer structure of the present invention;

[0041] Figure 3 It is a schematic diagram of the object detection model structure with regional attribute description of the present invention;

[0042] Figure 4 It is a schematic diagram of the training mode of the object detection model structure with regional attribute description of the present invention;

[0043] Figure 5 It is a schematic diagram of the inference mode of the object detection model structure with regional attribute description of the present invention;

[0044] Figure 6 It is a schematic diagram of the target area detection in Embodiment 1 of the present invention;

[0045] Figure 7 It is a block diagram of the object detection system structure based on prompt engineering and regional text description of the present invention. Detailed Embodiments

[0046] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. The illustrative embodiments and descriptions thereof of the present invention are only used to explain the present invention and do not limit the present invention.

[0047] At present, the in-vehicle edge AI detection system for smart cities relies on a large amount of data and fine-grained categories, uses a single or multiple identical object detection models, and relies on scenario logic combination to achieve the judgment of object categories and events in smart cities. However, the existing object detection methods based on fine-grained classification will lead to feature conflicts, which in turn will reduce the model performance, and they cannot provide category attribute assignment for the judgment of urban management events.

[0048] Based on the above problems, in order to solve the problems that fine-grained classification of object detection will lead to feature conflicts, which in turn affect the reduction of model performance and the judgment of target area attribute description, and the existing basic methods do not provide an object detection model for regional text description, let alone an end-to-end detection model that reduces resources. Therefore, the present invention designs and constructs an object detection model with regional attribute description.

[0049] So how to build such a model? We face two problems. First, how to extract regional features in the object detection model? To solve this problem, the present invention proposes the structure of the object region fine-grained feature module (APF). Second, how to realize the text description attribute of regional features? To solve this problem, the present invention proposes the structure of the object feature attribute description module (TGD), so as to build a new model architecture with regional attributes.

[0050] The specific design idea is as follows: The present invention does not consider introducing a text model (BERT) with a larger number of parameters, nor does it consider using a multi-modal structure model (CLIP). The present invention considers realizing the description of the target area attributes by means of image caption content. However, the image caption model only describes the content of the whole image and cannot accurately realize the description of the target area attributes (corresponding to the target area). Based on this, the present invention draws on the prompt engineering of the features of the Segment Anything Model (SAM), which includes point, box, and text prompts to segment the target. Similar to our detection box, we directly draw on the embedding method of the box to complete the extraction of the target area features, which are used as the encoded feature expression of the image caption, and then use the text decoding of the image caption to realize the description of the target area attributes. At the same time, both the object region fine-grained feature module (APF) and the object feature attribute description module (TGD) need to use the Transformer structure to realize regional feature extraction and text description respectively. We naturally introduce the Transformer structure detection model to realize the new architecture model of regional attribute description. Finally, the present invention uses the RTDetr object detection model as a baseline, introduces the object region fine-grained feature module (APF) and the object feature attribute description module (TGD), and constructs an object detection model with regional attribute description (PG-RTDetr model), effectively solving the problems of model performance and regional attribute judgment in urban management events.

[0051] Example 1

[0052] As Figure 1 shown, the object detection method of the present invention based on prompt engineering and regional text description depends on the following hardware components: (1) a high-definition camera, parameters: a high-definition camera with 2 million pixels (1920*1080), the distance between the detection area and the camera is less than 10 meters and greater than 1 meter, and the waterproof level is IPX6.

[0053] (2) A computing platform, parameters: Nvidia NX, TX edge computing devices, the memory and video memory are not less than 4G, and the main frequency of the processor is not less than 2.3GHz.

[0054] The method of the present invention includes:

[0055] Step 1, construct a labeled dataset, which includes multiple labeled image object detection annotations and labeled image object area text descriptions; the labeled image object detection annotations are used to label the detection categories and positions of scene objects, and the category annotations are labeled with superclasses (i.e., large classes); the labeled image object area text descriptions are short text descriptions of the areas corresponding to the detected objects, including area category attribute text descriptions and area event attribute text descriptions, and integrate them to construct text descriptions; the area category attribute text descriptions are content descriptions of the subclasses of the superclasses corresponding to the detected object areas; the area event attribute text descriptions are content descriptions of the events corresponding to the detected object areas; for example, the superclass description is a car, and the area event attribute text description is a sidewalk, and the constructed text description is a sidewalk car;

[0056] Step 2, construct an object detection model with area attribute descriptions. The object detection model with area attribute descriptions is constructed by adding an adaptive area feature extraction module and a text generation decoding module on the basis of not changing the structure of the RTDetr detection model to construct a new model with a text description branch structure;

[0057] Step 3, based on the labeled dataset, use the labeled image object detection annotations to train the RTDetr detection model for detection tasks as the first-stage model training; use the labeled image object area text descriptions to train the object detection model with area attribute descriptions for short text description tasks of areas as the second-stage model training, where the RTDetrr detection model in the second stage uses the weights obtained from the first-stage training and freezes the structure of the RTDetr detection model to obtain a trained object detection model with area attribute descriptions;

[0058] Step 4, obtain the image to be detected, and use the trained object detection model with area attribute descriptions to perform object detection on the image to be detected to obtain the target prediction box, prediction category and area attribute description.

[0059] The target detection model PG-RTDetr with regional attribute description constructed by the present invention solves the problems of reduced model performance caused by fine-grained classification and the judgment of urban management events. This model can implement fine-grained classification detection tasks and urban management event judgment tasks in the mobile scenario of in-vehicle edge devices and the arm edge computing platform. The present invention proposes a method of integrating a prompt method for extracting precise regional features and a method for generating feature texts into the existing target detection model RTDetr to construct a detection model PG-RTDetr with regional (category) attribute description. The present invention improves the structure of the existing target detection model and proposes a new architecture for target detection with regional description. The improved part of this structure mainly includes two parts. The first part is the adaptive regional feature extraction module APF, which is composed of the prompt conversion of the prediction box and multiple Transformer structures; the second part is the text generation decoding module TGD, which is also composed of multiple Transformers. Next, the present invention will introduce technical details such as the basic structure of the Transformer, the overall architecture of the PG-RTDetr model, the APF feature extraction structure, the TGD text generation structure, and the loss structure.

[0060] (1) Basic structure of the Transformer

[0061] The present invention uses the prompt engineering prompt of the prediction box as the query vector Q (query) and further encodes the encoder to extract precise regional features, and the text generation encoder decoder also uses the description caption as the query and uses the Transformer structure to realize the generation of text description content (regional attribute description).

[0062] The source of the existing Transformer structure is "Attention Is All You Need" proposed by Ashish Vaswani et al. This model proposes an encoder-decoder structure as Figure 2 shown, and the underlying structure uses the Transformer and the neural network structure (FFN structure).

[0063] The Transformer structure is composed of three input representation vectors (query vector Q, key vector K, and value vector V). Among them, the query vector Q, key vector K, and value vector V are implemented using conventional linear convolutions. Finally, the following formula can be used to implement the transformer structure, and its calculation formula is as follows:

[0064]

[0065] Among them, d k represents the feature expression dimension.

[0066] The FFN structure is composed of two linear convolutions and a ReLU activation function, and its calculation formula is as follows:

[0067] FFN(x) = max(0, xW + b)W2 + b2(2)

[0068] Among them, x is the input feature, W 1 is the one-dimensional convolution kernel weight, b 1 is the one-dimensional convolution corresponding bias, b 2 is the one-dimensional convolution kernel weight, w 2 is the one-dimensional convolution corresponding bias.

[0069] (2) The PG-RTDetr architecture of the object detection model with regional attribute description

[0070] The object detection model PG-RTDetr with regional (category) attribute description designed by the present invention (such as Figure 3 ), the PG-RTDetr model takes the RTDetr model as a benchmark, and integrates the added adaptive region feature extraction module APF, text generation decoding module TGD with the RTDetr model to construct a new model PG-RTDetr for region description object detection based on prompt engineering prompt and image description caption.

[0071] The PG-RTDetr model of the present invention maintains the original object detection structure RTDetr (such as Figure 3 the yellow part); adds an adaptive feature extraction module (such as Figure 3 the blue part); uses prompt engineering prompt to process the RTDetr features, further precisely extracts the region features Fp, and also adds a text generation decoding module (such as Figure 3 the green part), uses a stacked decoding structure to process Fp, and realizes text generation description Hc. In this way, the model proposed by the present invention can not only improve the detection ability of the original RTDetr model, but also provide attribute descriptions for regions (categories), and solve the problems of fine classification and events in urban management.

[0072] In the training stage, such as Figure 4As shown, the training strategy adopted by the present invention for the model. Experiments show that our training strategy can effectively improve the mean average precision (mAP) of the detection model. The present invention first trains the original RTDetr model with superclass data (merging sub-categories with relatively high feature similarity into a large category, for example, sedans and large SUVs are merged into cars), obtains the category prediction and prediction box prediction tasks of the detection model, and uses the loss of the original RTDetr model to achieve training. Then, the entire RTDetr detection model is frozen (such as Figure 3 the yellow part) to train the Adaptive Region Feature Extraction Module APF and the Text Generation Decoding Module TGD, and perform cross-entropy loss on the obtained Hc description output and the text attribute description of the labeled text gt (such as lane cars). In the inference stage (such as Figure 5 ), the PG-RTDetr model maintains the original RTDetr prediction task, realizes the prediction tasks of the target category and the region prediction box, and the Adaptive Region Feature Extraction Module APF and the Text Generation Decoding Module TGD use the features of RTDetr for further processing to complete the prediction of the region attribute description of the prediction box.

[0073] Specifically, the Adaptive Region Feature Extraction Module APF mainly converts the region coordinate information of the target prediction box into a vector representation based on the features extracted by the backbone network module (backbone) in the RTDetr model and the target prediction box output by the prediction module (Head), and further extracts the Ft features of the RTDetr model according to the vector representation to obtain the region-related feature expression Fp. This module mainly includes two parts. The first is to convert the prompt of the prediction box into a query vector Q (query), mainly using the prompt method of the prediction box of the SAM model; the second is to construct an n-layer feature extraction and encoding module based on Transformer.

[0074] First, the prompt encoding of the prediction box:

[0075] Based on the prompt engineering prompt, using the prompt method of the prediction box of the SAM model, convert the region prediction box coordinates [x 1 , y 1 , x 2 , y 2 into a d-dimensional vector representation X b ; it can be expressed by the following formula:

[0076] X b = f([x 1 , y 1 , x 2 , y 2 ) (3)

[0077] Among them, the f() function uses the calculation method proposed by Matthew Tancik et al. by referring to the SAM model.

[0078] Second, Prompt adaptive feature encoding:

[0079] The present invention takes the vector expression X b as the query vector Q of the Transformer structure, and takes the features of the RTDetr model as the key vector K and the value vector V of the Transformer structure; the vector expression X b is first encoded by the self-attention mechanism of the Transformer structure, and then the features of the RTDetr model are encoded by the cross-attention mechanism, and then the FFN structure is used to repeat n times to extract the region-related feature expression. The formula for the extracted region-related feature expression is:

[0080] F p = f n f ca (f sa (F t , X b ))(4)

[0081] Among them, f sa represents the self-attention mechanism structure, f ca represents the cross-attention mechanism structure, f n is the FFN structure, F t is the feature of the RTDetr model, and X b is the vector expression obtained by converting the regional coordinate information of the target prediction box.

[0082] Specifically, the text generation decoding module TGD mainly decodes the region-related feature expression obtained by the adaptive region feature extraction module APF by using the Transformer structure to obtain the region attribute description.

[0083] The text generation decoding module TGD also decodes in the way of formula (4), and uses a linear convolution to obtain the vocabulary corresponding vocabulary description. In the training stage, the input of the text generation decoding module TGD is F p and the text label token, and the cross-entropy is used to calculate the loss of the text. In the prediction stage, the input of the text generation decoding module TGD is F p and the start vocabulary cls as the text label token, and the model predicts word by word and superimposes it as the input of the new text label token to obtain the final region description text.

[0084] Such asFigure 6 As shown below, the target area detection is illustrated as follows:

[0085] Assumptions: ① Feature-similar subcategories have been merged into supercategories (e.g., sedans and large SUVs are merged into cars); ② There are existing text attribute data for regions or categories; ③ The RTDetr model has been trained using images and corresponding labels; ④ Only the APF and TGD modules need to be trained, which are the added structures in this invention.

[0086] Specific process:

[0087] ① Freeze the weights of the RTDetr model (e.g., Figure 6 the green box). Using the trained RTDetr model, the supercategory car and the corresponding position prediction box (4 coordinate points) can be predicted, and the classification is called the predicted category pred_class and the predicted box pred_box;

[0088] ② Feed the coordinates of the predicted box pred_box into the Adaptive Region Feature Extraction Module APF to convert it into the embedding of the prompt engineering prompt, and use the embedding as the query vector (query). Use the above Transformer structure to further extract more region-related features from the Ft feature, called Fp;

[0089] ③ Obtain the Fp feature through the Text Generation Decoding Module TGD, and decode the Fp feature into a simple description content, that is, the region or category text description;

[0090] ④ jointly determine the event or subcategory by the description of the lane sedan and the RTDetr predicted category car.

[0091] The above description is "lane sedan", and the predicted category is "car", so it can be known that Figure 6 this car is "a sedan parked in the lane". This invention uses the category attribute description to provide a method to solve the problem that the subcategory (region) causes feature conflicts and reduces the model performance (e.g., sedans, SUVs); in addition, it uses the category attribute to provide a solution to the problem of using category attributes to assist in judging smart city events.

[0092] As a further implementation, based on the labeled dataset, the RTDetr detection model is trained in the first stage using the labeled image target detection annotation; the target detection model with region attribute description is trained in the second stage using the labeled image target region text description, where the RTDetrr detection model in the second stage uses the weights obtained from the first stage training and freezes the RTDetr detection model structure to obtain a trained target detection model with region attribute description, including:

[0093] The first training stage: Based on the superclass data, train the RTDetr model to obtain the prediction classes and bounding box prediction tasks of the RTDetr model, and save the weights W of the RTDetr model. r The superclass data does not contain the corresponding event attribute data description.

[0094] The second training stage: Assign the weights W saved in the first stage of training r to the corresponding weights of the object detection model with regional attribute descriptions, and freeze its corresponding RTDetr model structure; Based on the labeled dataset, train the object detection model with regional attribute descriptions and only train the adaptive region feature extraction module APF and the text generation decoding module TGD, and update the weights of the adaptive region feature extraction module APF and the text generation decoding module TGD to achieve maintaining the object detection ability of the original RTDetr model and obtaining regional attribute descriptions.

[0095] After training is completed, save the best weights of the object detection model with regional attribute descriptions to obtain the trained object detection model with regional attribute descriptions.

[0096] For the above technical solutions, the two-stage training method is inseparable. By experimentally training the new model of the present invention (the object detection model PG-RTDetr architecture with regional attribute descriptions) at one time, the end-to-end detection task and the regional attribute description task training are realized. Here, the predicted bounding box of the prompt is changed to the real gt bounding box. Affected by the loss of the text description task, it is not as prominent as the two-stage training method of the present invention (first training the detection task and then training the regional attribute text description task). The present invention believes that this situation is affected by the characteristics of the branch task and the RTDetr detection task. In addition, using the predicted bounding box box trained in the first stage of RTDert in the prompt can provide more information than the real bounding box box, helping the branch task to better provide feature information.

[0097] The present invention conducts object detection in the smart city scenario, and the specific implementation is as follows:

[0098] 1. Construct regional (category) attribute description data. According to the urban scenario and industry nature description data, describe the superclass vehicle data: one is the feature class description (such as described as large sedans, SUVs, and compact cars), and the other is the event attribute data description (such as described as lane vehicles and sidewalk vehicles).

[0099] 2. Build the code for the Adaptive Region Feature Extraction Module (APF) and the Text Generation and Decoding Module (TGD). At the same time, build the code for the text loss module. Select the RTDetr object detection model as the baseline, and integrate the APF module, the TGD module with the base model to construct a new model PG-RTDetr for object detection based on prompt engineering and regional text description;

[0100] 3. Use the data of the superclass (the large class formed by merging similar small classes) without including the corresponding text description to train an original RTDetr object detection model, and save its weights W r ;

[0101] 4. Assign the W saved in step 3 r to the corresponding weights of the PG-RTDetr model, and freeze the corresponding RTDetr model structure (such as Figure 3 the yellow part);

[0102] 5. Use the dataset constructed in step 1 (including attribute descriptions) to train the PG-RTDetr model obtained in step 4, update the weights of the APF module and the TGD module, and achieve maintaining the original object detection ability and obtaining regional attribute descriptions;

[0103] 6. After training in step 5, save the best weights W best (including the weights of the detection model, APF, and TGD);

[0104] 7. Deploy the PG-RTDetr model to the arm edge device and load the weights W in step 6 best ;

[0105] 8. Perform smart city object detection and attribute description, combine the urban management standards with the model prediction categories and descriptions, and output the content for category task or event judgment.

[0106] In the present invention, on the basis of the existing RTDetr object detection model baseline, the prediction box is used as a prompt engineering prompt to obtain an accurate feature expression of the prediction box area (which is also the corresponding category). Then, the accurate regional features obtained are decoded using a structure similar to the caption to achieve regional attribute description. The new model PG-RTDetr of the present invention solves the fine-grained classification conflict problem and the industry event judgment problem with short attributes. The new model PG-RTDetr of the present invention uses the prompt engineering prompt to accurately extract features and integrate the regional description module, realizes detection and text description content, can replace multiple models to achieve the same effect, reduces the resources and inference time of the edge arm device to a certain extent, and also provides an idea for integrating regional attributes into the detection model.

[0107] Example 2

[0108] As shown Figure 7 in the figure, the difference between this embodiment and Embodiment 1 is that this embodiment provides an object detection system based on prompt engineering and regional text description, and this system uses the object detection method based on prompt engineering and regional text description in Embodiment 1; the system includes:

[0109] A dataset construction unit for constructing an annotated dataset, where the annotated dataset includes multiple annotated image object detection annotations and annotated image object region text descriptions; the annotated image object detection annotations are used to perform detection category annotations and position annotations on scene objects, and the category annotations are performed with superclasses; the annotated image object region text descriptions are short text descriptions of the regions corresponding to the detected objects, including region category attribute text descriptions and region event attribute text descriptions; the region category attribute text descriptions are content descriptions of the subclasses of the superclass corresponding to the detected object region; the region event attribute text descriptions are content descriptions of the events corresponding to the detected object region;

[0110] A model construction unit for constructing an object detection model with region attribute descriptions, where the object detection model with region attribute descriptions is a new model with a text description branch structure constructed by adding an adaptive region feature extraction module APF and a text generation decoding module TGD on the basis of not changing the structure of the RTDetr detection model;

[0111] A model training unit for, based on the annotated dataset, using the annotated image object detection annotations to train the RTDetr detection model for the detection task as the first-stage model training; using the annotated image object region text descriptions to train the object detection model with region attribute descriptions for the short region text description task as the second-stage model training, where in the second stage, the RTDetrr detection model uses the weights obtained from the first-stage training and freezes the structure of the RTDetr detection model to obtain a trained object detection model with region attribute descriptions;

[0112] An object detection unit for obtaining an image to be detected and using the trained object detection model with region attribute descriptions to perform object detection on the image to be detected to obtain an object prediction box, a predicted category, and a region attribute description.

[0113] As a further implementation, the object detection model with region attribute descriptions is based on the RTDetr model, and the added adaptive region feature extraction module APF, text generation decoding module TGD and the RTDetr model are integrated to construct a new model PG-RTDetr for region description object detection based on prompt engineering prompt and image description caption;

[0114] The Adaptive Region Feature Extraction Module APF is used to convert the regional coordinate information of the target prediction box into a vector representation based on the features extracted by the backbone network module in the RTDetr model and the target prediction box output by the Head module, and further extract the features of the RTDetr model according to the vector representation to obtain region-related feature expressions;

[0115] The Text Generation Decoding Module TGD is used to decode the region-related feature expressions using the Transformer structure based on the region-related feature expressions to obtain region attribute descriptions.

[0116] Among them, the execution process of each unit can be carried out according to the steps of the object detection method based on prompt engineering and region text description in Embodiment 1, and will not be elaborated one by one in this embodiment.

[0117] The above object detection method and system based on prompt engineering and region text description of the present invention focuses on designing a new architecture of an object detection model: an object detection model with region attribute descriptions. The present invention introduces prompt engineering prompt and image text generation decoder decoder into the existing object detection model. With the help of prompt engineering prompt, the prediction box obtained by the existing detection model is encoded by embeding as the precise feature extraction query vector Q, the features extracted by the existing detection model are used as the key vector K and the numerical vector V, the fine-grained features Fc of the detection target are extracted, and then the fine-grained features Fc are used by the image text generation decoder decoder to obtain the target attribute text description corresponding to the prediction box. The present invention solves the problems of the reduction of the model performance affected by the fine classification of object detection and the judgment of category attributes.

[0118] The present invention only adds a small amount of prompt encoder structure and text decoder structure to the original object detection model to achieve the extraction of local fine-grained features of the prediction box and short attribute descriptions. At the same time, the present invention integrates fine-grained categories with very high feature similarity into a superclass at the data level without changing the structure of the object detection model, so as to improve the performance of the detection model. Finally, the present invention selects the RTDetr model as the benchmark of the object detection model, maintains the original detection structure of the RTDetr model, and adds a prompt engineering prompt method to accurately extract the fine-grained feature module of the target (i.e., the adaptive region feature extraction module, called Adapter Prompter extract Feature, APF) and the target feature attribute description module (i.e., the text generation decoding module called Text Generation Decoder, TGD), and integrates the APF and TGD modules into the RTDetr model to propose a new object detection model with category (region) attribute descriptions (PG-RTDetr). The purpose of the present invention is to add a small number of parameters to integrate the detection model, obtain category attributes, and implement category attribute-assisted urban detection tasks in the edge vehicle-mounted AI system of a smart city.

[0119] Meanwhile, the present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above object detection method based on prompt engineering and regional text description.

[0120] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0121] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0122] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one or more processes and / or blocks Figure 1 in the process Figure 1 or processes and / or boxes

[0123] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in the process Figure 1 or processes and / or boxes

[0124] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. The object detection method based on prompt engineering and regional text description is characterized by: The method includes: Constructing a labeled data set, the labeled data set includes a plurality of labeled image target detection annotations and labeled image target region text descriptions; the labeled image target detection annotations are detection category annotations and position annotations based on scene targets, and the category annotations are annotated based on superclasses; the labeled image target region text descriptions are region short text descriptions corresponding to the detection target, including region category attribute text descriptions and region event attribute text descriptions; the region category attribute text descriptions are subdivided subclass content descriptions of the corresponding detection target region superclass; the region event attribute text descriptions are content descriptions of the corresponding detection target region events; Constructing a target detection model with regional attribute description, wherein the target detection model with regional attribute description is a new model with a text description branch structure by adding an adaptive regional feature extraction module and a text generation decoding module without changing the RTDetr detection model structure; Based on the labeled data set, the labeled image target detection annotation is used to perform detection task training on the RTDetr detection model as the first stage model training; the labeled image target region text description is used to perform regional short text description task training on the target detection model with regional attribute description as the second stage model training, wherein the second stage RTDetr detection model uses the weights obtained by the first stage training, and freezes the RTDetr detection model structure to obtain a trained target detection model with regional attribute description; An image to be detected is obtained, and a trained target detection model with regional attribute description is used to perform target detection on the image to be detected, so as to obtain a target prediction box, a prediction category and a regional attribute description.

2. The target detection method based on prompt engineering and regional text description according to claim 1, characterized in that: The superclass is to merge subcategories with high feature similarity into one large category.

3. The object detection method based on prompt engineering and regional text description according to claim 1, characterized in that: The target detection model with regional attribute description is based on the RTDetr detection model, and the added adaptive regional feature extraction module, the text generation and decoding module and the RTDetr model are integrated to form a new model for regional description target detection based on prompt engineering and image description; The adaptive regional feature extraction module converts the regional coordinate information of the target prediction box into a vector expression based on the features extracted by the backbone network module in the RTDetr model and the target prediction box output by the prediction module, and further extracts the features of the RTDetr model according to the vector expression to obtain a regional related feature expression; The text generation and decoding module, based on the region-related feature expression, uses a Transformer structure to decode the region-related feature expression to obtain a region attribute description.

4. The target detection method based on prompt engineering and regional text description according to claim 3 is characterized in that: Converting the region coordinate information of the target prediction box into a vector expression, further extracting the features of the RTDetr model according to the vector expression, and obtaining the region-related feature expression, including: Based on the prompt engineering, the regional coordinate information of the target prediction box is converted into a vector expression; The vector expression is used as the query vector of the Transformer structure, and the features of the RTDetr model are used as the key vector and numerical vector of the Transformer structure; the Transformer structure is used to first encode the vector expression through a self-attention mechanism, and then the features of the RTDetr model are encoded through a cross-attention mechanism, and then the FFN structure is used to repeat n times to realize the extraction of region-related feature expressions.

5. The object detection method based on prompt engineering and regional text description according to claim 4 is characterized in that: The expression formula of the extracted regional related features is: F p =f n f ca (f sa (F t ,X b )) Among them, f sa represents the self-attention mechanism structure, f ca represents the cross attention mechanism structure, f n is the FFN structure, F t is the characteristic of the RTDetr model, X b The region coordinate information of the target prediction box is converted into a vector expression.

6. The object detection method based on prompt engineering and regional text description according to claim 3 is characterized in that: The region attribute description is based on text attribute data of the prediction box.

7. The object detection method based on prompt engineering and regional text description according to claim 1 is characterized in that: Based on the labeled data set, the first stage model training is performed on the RTDetr detection model using the labeled image target detection annotations; the second stage model training is performed on the target detection model with regional attribute description using the labeled image target region text description, wherein the second stage RTDetr detection model uses the weights obtained by the first stage training, and freezes the RTDetr detection model structure to obtain a trained target detection model with regional attribute description, including: The first training stage: Based on the superclass data, the RTDetr model is trained to obtain the prediction category and prediction box task of the RTDetr model, and the weight W of the RTDetr model is saved. r ; The superclass data does not contain the corresponding event attribute data description; Second training stage: The weight W saved in the first stage training r Assigning corresponding weights to the target detection model with regional attribute description, and freezing its corresponding RTDetr model structure; training the target detection model with regional attribute description based on the labeled data set and only training the adaptive regional feature extraction module and the text generation decoding module, updating the weights of the adaptive regional feature extraction module and the text generation decoding module, so as to maintain the target detection capability of the original RTDetr model and obtain the regional attribute description; After the training is completed, the optimal weight of the target detection model with regional attribute description is saved to obtain a trained target detection model with regional attribute description.

8. Object detection system based on prompt engineering and regional text description, characterized in that: The system includes: A data set construction unit is used to construct a labeled data set, wherein the labeled data set includes a plurality of labeled image target detection annotations and labeled image target area text descriptions; the labeled image target detection annotations are detection category annotations and position annotations based on scene targets, and the category annotations are annotated based on superclasses; the labeled image target area text descriptions are short text descriptions of the area corresponding to the detection target, including area category attribute text descriptions and area event attribute text descriptions; the area category attribute text descriptions are subdivided subclass content descriptions of the corresponding detection target area superclass; the area event attribute text descriptions are content descriptions of the corresponding detection target area events; A model building unit, used to build a target detection model with regional attribute description, wherein the target detection model with regional attribute description is a new model with a text description branch structure by adding an adaptive regional feature extraction module and a text generation decoding module on the basis of not changing the RTDetr detection model structure; A model training unit is used to perform detection task training on the RTDetr detection model based on the labeled data set using the labeled image target detection annotations as the first stage model training; and perform regional short text description task training on the target detection model with regional attribute description using the labeled image target region text description as the second stage model training, wherein the second stage RTDetrr detection model uses the weights obtained by the first stage training, and freezes the RTDetr detection model structure to obtain a trained target detection model with regional attribute description; The target detection unit is used to obtain an image to be detected, use a trained target detection model with regional attribute description to perform target detection on the image to be detected, and obtain a target prediction box, a prediction category and a regional attribute description.

9. The object detection system based on prompt engineering and regional text description according to claim 8, characterized in that: The target detection model with regional attribute description is based on the RTDetr model, and the added adaptive regional feature extraction module, the text generation and decoding module and the RTDetr model are integrated to form a new model for regional description target detection based on prompt engineering and image description; The adaptive regional feature extraction module is used to convert the regional coordinate information of the target prediction box into a vector expression according to the features extracted by the backbone network module in the RTDetr model and the target prediction box output by the prediction module, and further extract the features of the RTDetr model according to the vector expression to obtain the regional related feature expression; The text generation and decoding module is used to decode the region-related feature expression using a Transformer structure according to the region-related feature expression to obtain a region attribute description.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the target detection method based on prompt engineering and regional text description is implemented as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Open vocabulary target detection method, system and device and medium

    CN118673465A