High-Resolution Image Target Detection Acceleration Method and System Based on Image Block Screening
By constructing an image block tree and filtering key image blocks using reinforcement learning training strategy networks, the problem of slow detection of high-resolution image objects in the prior art is solved, and faster object detection is achieved.
Patent Information
- Application Number
- CN202210216971.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-07
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-03-07
AI Technical Summary
Existing high-resolution image object detection technology slows down when processing a large number of image blocks, resulting in slowing down the target detection speed.
By constructing an image block tree, filtering the image block nodes on the tree based on the content and connections of the image blocks, and using reinforcement learning training strategy networks to filter out key image blocks for object detection.
While ensuring the accuracy of object detection, it streamlines the number of image blocks, significantly saves the running time of the object detection algorithm and improves detection speed.
Smart Images

Figure CN114937160B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of image processing, specifically an acceleration method and system for high-resolution image target detection based on image block screening. Background Art
[0002] Existing target detection technologies for high-resolution images generally first evenly divide a high-resolution image into several low-resolution image blocks, perform target detection on the low-resolution image blocks, and then map the detection boxes back to the high-resolution image. Although such technologies avoid directly processing high-resolution images, the existence of a large number of image blocks will slow down the speed of target detection. Summary of the Invention
[0003] In view of the above deficiencies in the prior art, the present invention proposes an acceleration method and system for high-resolution image target detection based on image block screening. By constructing a tree with image blocks as nodes and the spatial position relationships between image blocks as edges to model the connections between image blocks, and combining the content of image blocks and the connections between image blocks to screen the image block nodes on the tree, performing target detection on the screened image blocks, and then mapping the detection boxes back to the high-resolution image, the present invention can streamline the number of image blocks while ensuring the accuracy of target detection, achieving the purpose of saving the running time of the target detection algorithm.
[0004] The present invention is specifically implemented through the following technical solutions:
[0005] The present invention relates to an acceleration method for high-resolution image target detection based on image block screening. After extracting image blocks from a high-resolution image and constructing an image block tree in the offline stage, using the method of reinforcement learning to train a policy network with ① the image block tree as the state, ② selecting an image block in the tree as the action, and ③ the target detection accuracy as the reward; in the online stage, using the trained policy network, obtaining the serial numbers of the screened image blocks according to the input image block tree, performing target detection on the corresponding image blocks and mapping the detection results back to the original image to obtain the final detection result.
[0006] The described image block tree is constructed by downsampling a high-resolution image, using a target detection algorithm to extract several image blocks, and then according to the spatial position relationships between the image blocks. That is, when image block A contains image blocks B and C, then image block A is the parent node, and image blocks B and C form two child nodes of image block A.
[0007] The target detection algorithm mentioned above refers to an algorithm that outputs the category information and location information of the targets in a given image, and is implemented by using the method described in Ren S, He K, Girshick R, et al. Faster r-cnn: Towards real-time object detection with region proposal networks[J]. Advances in neural information processing systems, 2015, 28.
[0008] The image patch in the selection tree mentioned above means: number the nodes in the tree in the pre-order traversal order. Each time the policy network outputs a serial number, the image patch corresponding to the serial number is retained. All the nodes on the subtree with this image patch as the root node and the path connecting this node to the root node of the entire image patch tree cannot be selected again. Then, input the current image patch into the decoder of the policy network, and the policy network decodes the serial number of the next image patch. Repeat this process until there are no selectable image patches on the image patch tree.
[0009] The policy network mentioned above includes: an encoder and a decoder, where: the encoder extracts the content features of the image patches and the relationship between the image patches from the input image patch tree, and the decoder filters out the serial numbers of the image patches according to the information extracted by the encoder.
[0010] The accuracy of the target detection mentioned above is a common objective quality evaluation index for target detection: mean average precision (mAP).
[0011] The training mentioned above, that is, the process of optimizing the parameters of the policy network by using the method of reinforcement learning based on the state, action, reward and the preset network structure, specifically includes:
[0012] Step 1) Initialization: Randomly initialize the parameters θ of the policy network; initialize the trajectory set T as an empty set; initialize the sampling times N, the learning rate α, and the iteration times L.
[0013] Step 2) Collect training data: Use the image patch tree S i as the state, use the image patches in the selection tree as the action, use the target detection accuracy as the reward R i , use the policy network with the current parameters fixed as θ as the policy, execute the Markov decision process, and obtain the filtered image patch set τ i ={v0, v1, v2,...}, where v j represents a single image patch, and add τ i to the trajectory set T.
[0014] Step 3) Update network parameters: When the size of the trajectory set T, |T| = N, update θ according to the policy gradient algorithm as Calculate the average reward And clear the trajectory set T, then jump to Step 4. When the size of the trajectory set T, |T| < N, jump to Step 2.
[0015] Step 4) Determine whether the policy network converges: When calculating the average reward If it does not increase in L consecutive iterations, terminate the training; otherwise, jump to Step 2.
[0016] The mapping of the detection result back to the original image means: According to the position of the image patch in the original image, map the object detection result in the image patch coordinate system to the original image coordinate system to obtain the corresponding object position.
[0017] The present invention relates to an image object detection system for implementing the above method, including: an image patch tree construction module, a policy network training module, and an image patch screening module, where: the image patch tree construction module is connected to the policy network training module and transmits the image patch tree, and the policy network module is connected to the image patch screening module and transmits the trained policy network.
[0018] Technical effects
[0019] The present invention uses a fully connected layer, a tree-shaped long short-term memory network (Tree Long Short-Term Memory,
[0020] Tree-LSTM) and a chain-shaped long short-term memory (Chain Long Short-Term Memory, Chain-LSTM) network to encode the image patch tree. Among them, the tree-shaped long short-term memory network uses the inter-layer relationship of the tree to encode from the leaves to the root, while the chain-shaped long short-term memory network uses the intra-layer relationship of the tree to encode in the order of the tree layer traversal. The cascading of the two long short-term memory networks realizes the encoding of the entire tree. Screening is performed on a more complex tree structure, and the connections between tree nodes are fully considered during the screening process. The present invention models the image patch screening as a Markov decision process with the image patch tree as the state, selecting image patches in the tree as actions, and the object detection accuracy as the reward, and uses a policy network to model the policy. The screened image patches are obtained by executing this process. Description of the drawings
[0021] Figure 1 is a flowchart of the present invention;
[0022] Figure 2 is a schematic diagram of the image patch tree;
[0023] Figure 3 It is the structural diagram of the policy network;
[0024] Figure 4 It is the schematic diagram of image patch screening;
[0025] Figure 5 It is the schematic diagram of target detection. Specific implementation manners
[0026] Such as Figure 1 shown, this embodiment relates to a high-resolution image target detection acceleration method based on image patch screening, including:
[0027] Step 1) Construct an image patch tree, specifically including:
[0028] 1-1) Downsample the input high-resolution image by a factor of 4, use the Faster R-CNN algorithm to extract rough detection boxes, and then use the Mean Shift clustering algorithm to cluster the detection boxes. The bandwidth of the clustering is set to 0.025, and the minimum bounding rectangle of the detection boxes of each category is used as the image patch.
[0029] 1-2) Estimate the number of targets o i in each image patch and the average area n i of the targets in the image patch according to the number and area of the detection boxes in each image patch v i . Using the width w i , length h i , image patch area a i , average area o i of the targets in the image patch, number n i of the targets in the image patch, each image patch v i is represented as (w i , h i , a i , o i , n i ).
[0030] 1-3) Use the Mean Shift clustering algorithm to cluster the candidate image patches. The bandwidth of the clustering is set to 0.05. Consider each category obtained by clustering separately. Use the image patches of each category as child nodes, and the minimum bounding rectangle of the image patches of each category as the parent node. Connect the child nodes and the parent nodes to generate several image patch subtrees. Finally, connect the node representing the whole image with the root nodes of each subtree to generate the final image patch tree S = (V, E), where V = {v i}, is the set of image patches, E = {e i}, is the set of edges. Number the nodes in the tree in the preorder traversal order.
[0031] Step 2) Train a policy network including an encoder and a decoder, specifically including:
[0032] 2-1) Initialization: Randomly initialize the parameters θ of the policy network; initialize the set V of selected image patches as an empty set; initialize the trajectory set T as an empty set; initialize the sampling times N, the learning rate α, and the number of iterations L. In this embodiment, N is set to 64, α is set to 0.001, and L is set to 5.
[0033] The encoder includes: a fully connected layer for extracting image patch features, a tree long short-term memory network (Tree Long Short-Term Memory, Tree-LSTM) for encoding the relationship between image patches, and a chain long short-term memory (Chain Long Short-Term Memory, Chain-LSTM) network; the decoder includes: a long short-term memory network for sequential decoding and an attention head for selecting image patches.
[0034] The attention head is implemented in the manner described in Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly. Pointer networks. arXiv preprint arXiv:1506.03134, 2015.
[0035] 2-2) Collect training data: For the i-th image in the training set, generate the image patch tree S according to the method described in step 1) i , the policy network takes the image patch tree S i as input and outputs an image patch number j.
[0036] 2-3) Obtain the image patch v corresponding to the current image patch number j according to the node number on the tree j , add it to the set V. Modify the image patch tree to add masks to all nodes in the subtree with v j as the root node and the connection path from v j to the root node of the entire image patch tree S i . When all nodes on the tree have masks, jump to 2-4; otherwise, input the image patch v j into the decoder of the policy network, decode to obtain the next image patch number j+1, and use j+1 as the current image patch number, then jump to 2-3.
[0037] 2-4) Perform object detection on all image patches in the set V using, but not limited to, the Faster R-CNN algorithm, calculate the accuracy as the reward R i , let τ i = {v0, v1, v2,..}, where vj Represents a single image patch and adds τ i to the trajectory set T, and then clears the set V.
[0038] 2-5) Update network parameters: When the size of the trajectory set T, |T| = N, update θ according to the policy gradient algorithm as Calculate the average reward and clear the trajectory set T, then jump to 2-6. When the size of the trajectory set T, |T| < N, jump to 2-2.
[0039] 2-6) Determine whether the policy network converges: When calculating the average reward does not increase in L consecutive iterations, terminate the training; otherwise, jump to 2-2.
[0040] Step 3) Screen image patches, specifically including:
[0041] 3-1) Initialization: Load the trained policy network parameter θ in step 2), initialize the selected image patch set V as an empty set; for the test pictures in the offline stage, initialize the image patch tree S according to the method described in step 1).
[0042] 3-2) Execute an action: The policy network takes the image patch tree S as input and outputs an image patch serial number j.
[0043] 3-3) Obtain the image patch v corresponding to the current image patch serial number j according to the number of the node on the tree j , add it to the set V. Modify the image patch tree to add masks to all nodes on the subtree with v j as the root node and the connection path from v j to the root node of the entire image patch tree S i . When all nodes on the tree have masks, jump to 3-4; otherwise, input the image patch v j into the decoder of the policy network, decode to obtain the next image patch serial number j + 1, and use j + 1 as the current image patch serial number, then jump to 3-3.
[0044] 3-4) Object detection: Perform object detection on all image patches in the set V using but not limited to the Faster R-CNN algorithm to obtain the category and location information of the object.
[0045] 3-5) Coordinate mapping: Map the location information of the object in the image patch to the original image according to the location of the image patch in the original image to obtain the final detection result.
[0046] After specific actual experiments, in the environment of Ubuntu 16.04 operating system with a single GTX 1080Ti graphics card, a complete set of systems was built based on the open-source framework Pytorch, trained and tested on the billion-pixel object detection dataset PANDA, and the evaluation index mAP was calculated when IOU = 0.5. The experimental data obtained are shown in Table 1.
[0047] Table 1 Experimental Results
[0048] Down Sample+Sliding Window This embodiment mAP 0.715 0.722 Number of target detection runs 15019 8979 Speed (figures / second) 0.07 0.11
[0049] Compared with the existing technology of downsampling high-resolution images and then performing detection by sliding window block (Down Sample + Sliding Window), under similar object detection accuracy (mAP), the object detection speed of this method is 1.5 times that of the original.
[0050] In summary, this method uses the trained policy network to screen image patches, screen the nodes in the image patch tree, and then perform object detection on the screened image patches, significantly reducing the number of object detection runs, and thus the average detection time per high-resolution image is less.
[0051] The above specific implementation can be locally adjusted in different ways by those skilled in the art without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation, and all implementation solutions within its scope are subject to the present invention.
Claims
1. A high-resolution image target detection acceleration method based on image block screening, characterized in that, After extracting image patches from a high-resolution image and constructing an image patch tree in the offline phase, a policy network is trained using reinforcement learning with ① the image patch tree as the state, ② selecting an image patch in the tree as the action, and ③ the target detection accuracy as the reward; in the online phase, the trained policy network is used to obtain the selected image patch numbers according to the input image patch tree, and the final detection result is obtained by performing target detection on the corresponding image patches and mapping the detection results back to the original image. The described policy network includes: an encoder and a decoder, where: the encoder extracts the content features of the image patches and the relationship between the image patches from the input image patch tree, and the decoder filters out the image patch numbers according to the information extracted by the encoder. The encoder includes: a fully connected layer for extracting image patch features, a tree-shaped long short-term memory network for encoding the relationship between image patches, and a chain-shaped long short-term memory network; the decoder includes: a long short-term memory network for sequential decoding and an attention head for selecting image patches.
2. The high-resolution image target detection acceleration method based on image block screening according to claim 1, characterized in that, The described image patch tree is constructed by downsampling a high-resolution image, extracting several image patches using a target detection algorithm, and then constructing according to the spatial position relationship between the image patches.
3. The high-resolution image target detection acceleration method based on image block screening according to claim 1, characterized in that, Selecting an image patch in the tree means: numbering the nodes in the tree in preorder traversal order, the policy network outputs a number each time, the image patch corresponding to the number is retained, and all nodes on the subtree with this image patch as the root node and the path connecting this node to the root node of the entire image patch tree cannot be selected again. Then, the current image patch is input into the decoder of the policy network, and the policy network decodes the number of the next image patch, and so on until there are no selectable image patches on the image patch tree.
4. The high-resolution image target detection acceleration method based on image block screening according to claim 1, characterized in that, The accuracy of the described target detection is a general objective quality evaluation index for target detection: the average accuracy of classes.
5. The high-resolution image target detection acceleration method based on image block screening according to claim 1, characterized in that, The described training is a process of optimizing the parameters of the policy network using reinforcement learning based on the state, action, reward, and a preset network structure, specifically including: Step 1) Initialization: Randomly initialize the parameters of the policy network ; Initialize the trajectory set as an empty set; Initialize the number of sampling times , the learning rate , and the number of iterations ; Step 2) Collect training data: Using the image patch tree as the state, selecting an image patch in the tree as the action, and using the object detection accuracy as the reward , with the current parameters fixed as the policy network as the policy, perform a Markov decision process to obtain the filtered set of image patches , where represents a single image patch, and add to the trajectory set ; Step 3) Update network parameters: When the size of the trajectory set is , update according to the policy gradient algorithm as , calculate the average reward , and clear the trajectory set , then jump to Step 4; When the size of the trajectory set is , then jump to Step 2; Step 4) Determine whether the policy network has converged: When calculating the average reward During consecutive iterations without increase, terminate the training; otherwise, jump to Step 2.
6. The high-resolution image target detection acceleration method based on image block screening according to claim 1, characterized in that, Mapping the detection result back to the original image means: according to the position of the image patch in the original image, mapping the target detection result in the image patch coordinate system to the original image coordinate system to obtain the corresponding target position.
7. An image target detection system for implementing the high-resolution image target detection acceleration method based on image block screening according to any one of claims 1 to 6, characterized in that, Including: An image patch tree construction module, a policy network training module, and an image patch screening module, where: the image patch tree construction module is connected to the policy network training module and transmits the image patch tree, and the policy network module is connected to the image patch screening module and transmits its trained policy network.