A detection method, device, and medium based on progressive candidate box highlighting

By combining the progressive candidate box highlighting method of Conformer network and transformer network, the shortcomings of convolutional neural network and transformer detector in semantic dependence and local features are solved, and efficient object detection and positioning are achieved.

CN114821036BActive Publication Date: 2025-07-22UNIV OF CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210391460.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-04-06
Filing Date
2022-04-14
Publication Date
2025-07-22
Estimated Expiration
2042-04-14

AI Technical Summary

Technical Problem

Existing convolutional neural network detectors are difficult to capture semantic dependencies between long-distance objects, while transformer-based detectors will deteriorate the details of local features, resulting in poor detection results.

Method used

By combining the Conformer network and transformer-based network, the progressive candidate box highlighting method is used to extract candidate box embedding vectors and feature vectors using the Conformer network, and local features are fused through regional feature aggregation and cross-attention methods to realize the detection of multiple iterations of target candidate boxes.

Benefits of technology

It significantly improves the accuracy of detection, enhances the representation ability of target features, takes into account local feature details and long-distance semantic dependence, and improves the overall performance of the detector.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821036B_ABST
    Figure CN114821036B_ABST
Patent Text Reader

Abstract

The present invention discloses a detection method based on progressive candidate box highlighting. An image is input into a detector, and the detector identifies the targets in the image to obtain the positions of the targets in the image through the following steps: extracting the candidate box embedding vector and the feature vector of the image through a Conformer network; obtaining the target candidate boxes according to the candidate box embedding vector; extracting the feature vectors within the target candidate boxes through regional feature aggregation to obtain local feature vectors; fusing the candidate box embedding vector and the local feature vectors to obtain a new candidate box embedding vector; replacing the original candidate box embedding vector with the new candidate box embedding vector, and repeating the above steps in sequence. After repeating multiple times, the positions of the targets in the image are obtained. The method disclosed by the present invention enhances the target features and significantly improves the accuracy of detection and positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a detection method based on progressive candidate box highlighting, belonging to the technical field of image recognition. Background Art

[0002] Computer vision is a technology that uses a computer to recognize and analyze visual targets such as images and videos, thereby assisting or replacing the human visual system to work and reducing the workload of humans in obtaining and processing this visual information.

[0003] Currently, the methods for visual target detection have developed two branches, one is the convolutional neural network and the other is the network based on transformer.

[0004] The detector based on the convolutional neural network can make full use of the features of the local area, but it will encounter difficulties in capturing the semantic dependencies between distant objects; the target detector based on transformer is good at capturing the semantic dependencies between distant objects, but it will deteriorate the details of the local features.

[0005] Therefore, it is necessary to study a detection method to solve the above problems. Summary of the Invention

[0006] In order to overcome the above problems, the inventors of the present invention have conducted in-depth research and proposed a detection method based on progressive candidate box highlighting. By reasonably using and fusing the features of the two frameworks of the convolutional neural network and the network based on transformer at the same time, the detector can not only fully absorb the details of the local features, but also be good at capturing the semantic dependencies between distant objects, achieving good detection results.

[0007] Specifically, the present invention discloses a detection method based on progressive candidate box highlighting. An image is input into the detector, and the detector identifies the target in the image to obtain the position of the target in the image.

[0008] Further, the detector identifies the position of the target in the image through the following steps:

[0009] S1. Extract the candidate box embedding vector and feature vector of the image through the Conformer network;

[0010] S2. Obtain the target candidate box according to the candidate box embedding vector;

[0011] S3. Extract the feature vector within the target candidate box through regional feature aggregation to obtain the local feature vector;

[0012] S4. Fuse the candidate box embedding vector and the local feature vector to obtain a new candidate box embedding vector;

[0013] S5. Replace the original candidate box embedding vector with a new candidate box embedding vector, and repeat steps S2 - S4 in sequence. After repeating multiple times, obtain the position of the target in the image, and the position is marked by the target candidate box obtained in the last time.

[0014] Further, the target candidate box includes a classification score and a bounding box.

[0015] The classification score is used to represent the probability that there is a target within the candidate box, and the bounding box is used to represent the boundary range of the candidate box.

[0016] Preferably, in S2, the initial target bounding box is corrected by a perception layer to obtain a target bounding box, and the initial target bounding box is generated by a Conformer network.

[0017] Preferably, the parameters of the perception layer are obtained by training using the candidate box embedding vector through the method of minimizing the matching loss, including the following steps:

[0018] S21. Input the candidate box embedding vector into a linear layer to obtain a predicted classification score.

[0019] S22. Obtain a bounding box offset by inputting the candidate box embedding vector into the perception layer, and correct the initial target bounding box through the bounding box offset to obtain a predicted bounding box.

[0020] S23. Combine the predicted bounding box and the predicted classification score into a predicted set R, match the predicted set R with the ground - truth target set, and update and obtain the parameters of the perception layer by minimizing the matching loss.

[0021] Preferably, in S23, the Hungarian method is used to minimize the matching loss, and the loss function for matching is the weighted sum of a classification loss function, a target localization L1 loss function, and a GioU loss function.

[0022] The classification loss function characterizes the loss between the predicted classification score and the ground - truth target classification score; the target localization L1 loss function and the GioU loss function characterize the loss between the predicted bounding box and the ground - truth target bounding box.

[0023] Preferably, in S3, the RoIAlign method is used to extract feature vectors to obtain local feature vectors.

[0024] Preferably, in S4, the fusion includes the following sub - steps:

[0025] S41. Enhance the candidate box embedding vector to I through a linear layer to form an embedding vector set.

[0026] S42. Embed the candidate box embedding vector set and the local feature vector through the cross-attention method to obtain an enhanced feature vector;

[0027] S43. Input the enhanced feature vector into a convolutional layer to obtain a new candidate box embedding vector.

[0028] Preferably, in S42, the cross-attention method can be expressed as:

[0029]

[0030] where is the enhanced feature vector, f i is the local feature vector, is the embedding vector set, θ q , θ k , θ v are the linear transformation layer parameters, C represents the dimension of the feature vector, h′ is a hyperparameter, and the superscript T represents transpose.

[0031] The present invention also provides an electronic device, including:

[0032] At least one processor; and

[0033] A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any one of the above.

[0034] The present invention also provides a computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in any one of the above.

[0035] The beneficial effects of the present invention include:

[0036] (1) Enhance the target features in target candidate box prediction and in a learnable iterative optimization framework;

[0037] (2) Through the cross-attention method, fuse the sparse transformer embedding vectors and the dense convolutional neural network local features, effectively enhance the representation ability of the embedding vectors, and significantly improve the detection accuracy. Description of the Drawings

[0038] Figure 1 Schematic diagram showing the method for the detector to identify the target position in the detection method based on progressive candidate box highlighting according to a preferred embodiment of the present invention. Detailed Embodiments

[0039] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Through these descriptions, the features and advantages of the present invention will become more clearly defined.

[0040] As used herein, the term "exemplary" means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or better than other embodiments. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0041] The present invention provides a detection method based on progressive candidate box highlighting. An image is input into a detector, and the detector identifies the targets in the image to obtain the positions of the targets in the image. The detection method of the present invention can simultaneously take into account and efficiently focus on the texture details of the targets and the long-distance semantic dependencies, effectively solving the disadvantages that the detector based on convolutional neural network cannot capture long-distance semantic dependencies and the detector based on transformer loses local details.

[0042] Specifically, the detector in the present invention identifies the positions of the targets in the image through the following steps:

[0043] S1. Extract the candidate box embedding vector and feature vector of the image through the Conformer network;

[0044] S2. Obtain the target candidate boxes according to the candidate box embedding vector;

[0045] S3. Extract the feature vectors within the target candidate boxes through regional feature aggregation to obtain local feature vectors;

[0046] S4. Fuse the candidate box embedding vector and the local feature vector to obtain a new candidate box embedding vector;

[0047] S5. Replace the original candidate box embedding vector with the new candidate box embedding vector, and repeat steps S2 to S4 in sequence. After multiple repetitions, obtain the positions of the targets in the image, and the positions are indicated by the target candidate boxes obtained in the last time.

[0048] In S1, the Conformer network is a network model proposed in 2020. In the Conformer network, there are a transformer branch and a convolutional neural network branch. For the specific structure of the Conformer network, please refer to the paper Gulati, Anmol, et al. 'Conformer: Convolution-augmented Transformer for Speech Recognition.' arXiv preprint arXiv:2005.08100 (2020). This will not be elaborated in this invention.

[0049] In this invention, the Conformer network can be the original Conformer network or any of its variants, such as the Conformer-S network.

[0050] Further, in the transformer branch, each target region in the image can be represented as a candidate box embedding vector Then the candidate box embedding vectors of each image are represented as where i represents different candidate box embedding vectors, N represents the total number of candidate box embedding vectors in each image, represents the set of real numbers, and C is the dimension of the vector.

[0051] Further, in this invention, the target candidate box includes a classification score and a bounding box.

[0052] The classification score is used to represent the probability that there is a target within the candidate box, and the bounding box is used to represent the boundary range of the candidate box.

[0053] In the Conformer network, the input image can be recognized to output the target bounding box. In S2, the target bounding box generated by the Conformer network is called the initial bounding box, and the initial bounding box is corrected to enhance the accuracy of its positioning.

[0054] Further, in this invention, the initial bounding box is represented as b i represents the bounding box, b 0i ={x 0i ,y 0i ,w 0i ,h 0i}, {x,y} are the center coordinates of the target candidate box, and {w,h} are the width and height.

[0055] Specifically, the initial target bounding box generated by the Conformer network is corrected through the perception layer to obtain the target bounding box.

[0056] Further, the parameters of the perception layer are obtained by training using the candidate box embedding vector through the minimum matching loss method, including the following steps:

[0057] S21. Input the candidate box embedding vector e i into the linear layer to obtain the predicted classification score S i ;

[0058] S22. Obtain the bounding box offset by inputting the candidate box embedding vector into the perception layer, and correct the initial target bounding box through the bounding box offset to obtain the predicted bounding box;

[0059] S23. Combine the predicted bounding box and the predicted classification score into a prediction set R, match the prediction set R with the ground truth target set, and update and obtain the parameters of the perception layer by minimizing the matching loss.

[0060] The linear layer and the perception layer are both common layer structures in neural networks, and their specific structural compositions are not elaborated in the present invention.

[0061] In S21, input each candidate box embedding vector in the image into the linear layer to obtain the corresponding predicted classification score S i ;

[0062] In S22, the perception layer has 3 layers, the activation function of the perception layer is set as the ReLU function, and the candidate box embedding vector e i obtains the bounding box offset δb i ={δx i , δy i , δw i , δh i}, and correct the initial target bounding box through the bounding box offset to obtain the predicted bounding box where

[0063] Further, combine the predicted bounding box with the predicted classification score S i to form a prediction set Match the prediction set with the ground truth target set.

[0064] The ground truth target set is expressed as where g k is the ground truth bounding box, and y k is the one-hot label of the kth target.

[0065] In a preferred embodiment, in S23, use the bipartite graph matching strategy to match the prediction set R with the ground truth target set G.

[0066] More preferably, the Hungarian method in the bipartite graph is used for matching, and the minimum matching loss is achieved through the Hungarian method, where the loss function of the matching is the weighted sum of the classification loss function, the object localization L1 loss function, and the GioU loss function;

[0067] The classification loss function represents the loss between the predicted classification score and the true target classification score; the object localization L1 loss function and the GioU loss function represent the loss between the predicted bounding box and the true target bounding box.

[0068] Preferably, the weight of the classification loss function is 1 to 3, more preferably 2; the weight of the object localization L1 loss function is 4 to 6, more preferably 5; the weight of the GioU loss function is 1 to 2, more preferably 1. The above weights are obtained by the inventor based on experience and can significantly improve the matching accuracy.

[0069] Preferably, the classification loss function is the Focal Loss function, the object localization L1 loss function is the L1 Loss function, and the GioU loss function is the Giou Loss function.

[0070] Furthermore, only the classification loss function is calculated for the predicted bounding box of the negative sample.

[0071] According to the present invention, under the supervision of the above loss function, the network parameters are updated by the stochastic gradient descent method during the training iteration.

[0072] In S2, the candidate box embedding vector will learn to predict the candidate region under the supervision of the candidate box prediction loss function, enhancing the object features.

[0073] In a preferred embodiment, in S3, the feature vector is extracted by the RoIAlign method to obtain the local feature vector.

[0074] In the present invention, the local feature vector is represented as f i , and the set of local feature vectors is represented as N and C have the same parameter meanings as N and C in the candidate box embedding vector, and S = 14, representing the length and width of the feature vector.

[0075] Preferably, in S4, the fusion includes the following sub-steps:

[0076] S41. The candidate box embedding vector is enhanced to I through a linear layer to form an embedding vector set;

[0077] S42. The candidate box embedding vector set and the local feature vector are enhanced through the cross-attention method to obtain the enhanced feature vector;

[0078] S43. Input the enhanced feature vector into the convolutional layer to obtain a new candidate box embedding vector.

[0079] In S41, each of the I enhanced candidate box embedding vectors is the same as the original candidate box embedding vector, and the embedding vector set can be expressed as:

[0080] In S42, the cross-attention method can be expressed as:

[0081]

[0082] where is the enhanced feature vector, f i is the local feature vector, is the embedding vector set, θ q 、θ k 、θ v are the parameters of the linear transformation layer, C represents the dimension of the feature vector, h′ is a hyperparameter, preferably set to 8, and the superscript T represents the transpose.

[0083] In the present invention, through the cross-attention method, the fusion of the output of the convolutional layer and the embedding vector is achieved, thereby enhancing the representation ability of the embedding vector and significantly improving the detection accuracy.

[0084] Through S4, the candidate box embedding vector of the transformer and the local features of the dense convolutional neural network are fused, strengthening the discriminability of the candidate box and the accuracy of localization.

[0085] In S5, preferably, the multiple repetitions are 3 to 8 repetitions, more preferably 6 repetitions.

[0086] Since the new candidate box embedding vector obtained after fusion contains local details and long-distance semantic dependencies, effectively highlighting the candidate box in the context. Through multiple repetitions, the alternating iteration of candidate box prediction and target feature enhancement is achieved, so that the target feature is gradually highlighted and the target position is gradually improved.

[0087] The various embodiments of the methods described above in the present invention can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0088] The program code for implementing the methods of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0089] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0090] To provide interaction with a user, the methods and apparatuses described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0091] The methods and apparatuses described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0092] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.

[0093] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this invention can be achieved, and no limitations are imposed herein.

[0094] Embodiment

[0095] Embodiment 1

[0096] Set up a simulation experiment to evaluate the detector on the MSCOCO and CrowdHuman datasets, and observe the effect of the detector in identifying the location of the target in the image.

[0097] Among them, the MSCOCO dataset is a large and rich object detection, segmentation, and caption dataset. This dataset aims at scene understanding and is mainly intercepted from complex daily scenes. The location of the target in the image is calibrated through precise segmentation. The images include 80 classes. The training set contains 118k images, the validation set contains 5k images, and the test set contains 20k images; CrowdHuman is specifically provided for pedestrian detection in crowded scenes. The semantic dependence between features is particularly crucial for improving detection performance. The training set has 15k images, the test set has 5k images, and the validation set has 4370 images.

[0098] The detector identifies the location of the target in the image through the following steps:

[0099] S1. Extract the candidate box embedding vector and feature vector of the image through the Conformer network;

[0100] S2. Obtain the target candidate box according to the candidate box embedding vector;

[0101] S3. Extract the feature vector within the target candidate box through regional feature aggregation to obtain the local feature vector;

[0102] S4. Fuse the candidate box embedding vector and the local feature vector to obtain a new candidate box embedding vector;

[0103] S5. Replace the original candidate box embedding vector with the new candidate box embedding vector, and repeat steps S2 - S4 in sequence. After multiple repetitions, obtain the location of the target in the image.

[0104] In S2, the initial target bounding box is corrected through the perception layer to obtain the target bounding box. The initial target bounding box is generated by the Conformer network. The perception layer has 3 layers, and the activation function of the perception layer is set to the ReLU function. The parameters of the perception layer are trained using the candidate box embedding vector through the minimum matching loss method, including the following steps:

[0105] S21. Input the candidate box embedding vector into the linear layer to obtain the predicted classification score;

[0106] S22. Obtain the bounding box offset by inputting the candidate box embedding vector into the perception layer, and correct the initial target bounding box through the bounding box offset to obtain the predicted bounding box;

[0107] S23. The predicted bounding boxes and predicted classification scores are combined into a prediction set R. The prediction set R is matched with the ground-truth target set, and the parameters of the perception layer are updated and obtained by minimizing the matching loss.

[0108] Further, the Hungarian method is used to minimize the matching loss, where the loss function of the matching is the weighted sum of the classification loss function, the target localization L1 loss function, and the GioU loss function.

[0109] Among them, the weight of the classification loss function is 2; the weight of the L1 loss function is 5; the weight of the GioU loss function is 1. In S3, the feature vectors are extracted by the RoIAlign method to obtain local feature vectors.

[0110] In S4, the fusion includes the following sub-steps:

[0111] S41. The candidate box embedding vectors are enhanced to I through a linear layer to form an embedding vector set.

[0112] S42. The candidate box embedding vector set and the local feature vectors are enhanced by the cross-attention method to obtain enhanced feature vectors.

[0113] S43. The enhanced feature vectors are input into a convolutional layer to obtain new candidate box embedding vectors.

[0114] In S42, the cross-attention method is as follows:

[0115]

[0116] Among them, h′ = 8.

[0117] In S5, it is repeated 6 times in total. The AP (Average Precision) index is used to evaluate the performance of the recognition results of the detector. For the specific method of AP, please refer to the literature Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John M. Winn, and Andrew Zisserman. The pascal visual object classes (VOC) challenge. Int. J. Comput. Vis., pages 303–338, 2010.

[0118] Comparative Example 1

[0119] Set up a simulation experiment. On the MSCOCO and CrowdHuman datasets, use Sparse R-CNN to identify the target positions in the images. The specific method of Sparse R-CNN can refer to the literature "Peize Sun, Rufeng Zhang, YiJiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, and Ping Luo. Sparse R-CNN: end-to-end object detection with learnable proposals. In IEEE CVPR, pages 14454–14463, 2021."

[0120] Comparative Example 2

[0121] Set up a simulation experiment. On the MSCOCO and CrowdHuman datasets, use ViDT to identify the target positions in the images. The specific method of ViDT can refer to the literature "Hwanjun Song, Deqing Sun, Sanghyuk Chun, Varun Jampani, Dongyoon Han, Byeongho Heo, Wonjae Kim, and Ming-Hsuan Yang. Vidt: An efficient and effective fully transformer-based object detector. arXiv preprint arXiv:2110.03921, 2021".

[0122] Comparative Example 3

[0123] Set up a simulation experiment. On the MSCOCO and CrowdHuman datasets, use EfficientDet to identify the target positions in the images. The specific method of EfficientDet can refer to the literature "Mingxing Tan, Ruoming Pang, and Quoc V. Le. Efficientdet: Scalable and efficient object detection. In IEEE CVPR, pages 10778–10787, 2020.".

[0124] Comparative Example 4

[0125] Set up a simulation experiment. On the MSCOCO and CrowdHuman datasets, use Deformerable DETR to identify the target positions in images. The specific method of Deformerable DETR can refer to the literature "Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable DETR: deformable transformers for end-to-end object detection. In ICLR, 2021."

[0126] Comparative Example 5

[0127] Set up a simulation experiment. On the MSCOCO and CrowdHuman datasets, use Dyhead to identify the target positions in images. The specific method of Dyhead can refer to the literature "Xiyang Dai, Yinpeng Chen, Bin Xiao, Dongdong Chen, Mengchen Liu, Lu Yuan, and Lei Zhang. Dynamic head: Unifying object detection heads with attentions. In IEEE CVPR, pages 7373–7382, 2021."

[0128] Experimental Example

[0129] Compare the test performances of Example 1 and Comparative Examples 1 - 5 in the dataset. The results are shown in Tables 1 - 3. Table 1 shows the test performance results on the MSCOCO validation set, Table 2 shows the test performance results on the MSCOCO test set, and Table 3 shows the test performance results on the CrowdHuman dataset.

[0130] Table 1 Test Performance Results on the MSCOCO Validation Set

[0131] Method AP Sparse R-CNN (Comparative Example 1) 47.9 ViDT (Comparative Example 2) 49.2 Example 1 50.1

[0132] Table 2 Test Performance Results on the MSCOCO Test Set

[0133] Method AP EfficientDet (Comparative Example 3) 52.2 Deformerable DETR (Comparative Example 4) 52.3 Dyhead (Comparative Example 5) 52.3 Example 1 52.5

[0134] Table 3 Test Performance Results on the CrowdHuman Dataset

[0135] Method AP Deformable DETR (Comparative Example 4) 86.7 Sparse R-CNN (Comparative Example 1) 89.2 Example 1 89.4

[0136] As can be seen from Table 1, compared with the current advanced detector ViDT, the performance of Example 1 has increased by 0.9%; compared with Sparse R-CNN, there is a significant improvement, and the performance is 1.3% higher. This not only verifies the rationality of the fusion of CNN local features and transformer representations, but also verifies the effectiveness of the iterative optimization strategy for candidate box prediction and target feature enhancement.

[0137] In Table 2, on the test branch of MSCOCO, when compared with the current advanced state-of-the-art detectors, Example 1 achieved 52.5% AP, which is comparable to the detection performance reported by the current advanced detectors.

[0138] As can be seen from Table 3, Example 1 achieved 89.4% AP, which is 2.7% higher than the transformer-based method (Comparative Example 4). This shows that the method in Example 1 has higher potential in complex scenarios. In crowded scenarios, it shows that the occluded or relatively small targets are emphasized through the long-distance semantic dependence function.

[0139] The present invention has been described in combination with preferred embodiments above. However, these embodiments are merely exemplary and only serve an illustrative purpose. On this basis, various substitutions and improvements can be made to the present invention, and these all fall within the protection scope of the present invention.

Claims

1. A detection method based on progressive candidate box highlighting, characterized in that Input an image into a detector, identify the target in the image through the detector, and obtain the position of the target in the image; The detector identifies the position of the target in the image through the following steps: S1. Extract the candidate box embedding vector and feature vector of the image through the Conformer network; S2. Obtain the target candidate box according to the candidate box embedding vector; S3. Extract the feature vector within the target candidate box through regional feature aggregation to obtain the local feature vector; S4. Fuse the candidate box embedding vector and the local feature vector to obtain a new candidate box embedding vector; S5. Replace the original candidate box embedding vector with the new candidate box embedding vector, and sequentially repeat steps S2 to S4. After multiple repetitions, obtain the position of the target in the image, and the position is marked by the target candidate box obtained in the last time; In S4, the fusion includes the following sub-steps: S41. Enhance the candidate box embedding vector to I through a linear layer to form an embedding vector set; S42. Enhance the candidate box embedding vector set and the local feature vector through the cross-attention method to obtain the enhanced feature vector; S43. Input the enhanced feature vector into the convolutional layer to obtain a new candidate box embedding vector.

2. The detection method based on progressive candidate box highlighting according to claim 1, wherein In S2, the initial target bounding box is corrected through a perception layer to obtain the target bounding box, and the initial target bounding box is generated by the Conformer network.

3. The detection method based on progressive candidate box highlighting according to claim 2, wherein The parameters of the perception layer are obtained by training using the candidate box embedding vector through the minimum matching loss method, including the following steps: S21. Input the candidate box embedding vector into the linear layer to obtain the predicted classification score; S22. Obtain the bounding box offset by inputting the candidate box embedding vector into the perception layer, and correct the initial target bounding box through the bounding box offset to obtain the predicted bounding box; S23. Combine the predicted bounding box and the predicted classification score into a prediction set R, match the prediction set R with the ground truth target set, and update and obtain the parameters of the perception layer by minimizing the matching loss.

4. The detection method based on progressive candidate box highlighting according to claim 3, wherein In S23, the minimum matching loss is realized through the Hungarian method, and the loss function of the matching is the weighted sum of the classification loss function, the target localization L1 loss function, and the GioU loss function; The classification loss function represents the loss between the predicted classification score and the ground truth target classification score; the target localization L1 loss function and the GioU loss function represent the loss between the predicted bounding box and the ground truth target bounding box.

5. The detection method based on progressive candidate box highlighting according to claim 1, wherein In S3, the RoIAlign method is used to extract the feature vector to obtain the local feature vector.

6. The detection method based on progressive candidate box highlighting according to claim 1, wherein In S42, the cross-attention method can be expressed as: Among them, is the enhanced feature vector, f i is the local feature vector, is the set of embedding vectors, θ q , θ k , θ v are the parameters of the linear transformation layer, C represents the dimension of the feature vector, h′ is the hyperparameter, and the superscript T represents the transpose.

7. An electronic device, comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-6.

8. A computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Pedestrian cross-mirror re-identification method and system based on background and attitude normalization

    CN114120363A

  • Contextual grounding of natural language phrases in images

    US20210081728A1