Feature Map Generation Method, Training Method and Device for Object Detection Model
By partitioning the sample image candidate areas and randomly determining anchor points to generate feature maps, the problem of insufficient samples is solved, and the performance improvement and generalization ability of the object detection model are achieved.
Patent Information
- Application Number
- CN202310431499.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-04-20
AI Technical Summary
In the object detection task, insufficient sample size or poor quality leads to insufficient model performance and generalization capabilities, and the existing data enhancement methods are complex in the process, making it difficult to efficiently improve model performance.
By partitioning the candidate areas of the sample image and randomly determining the anchor points in the sub-region, a second feature map is generated based on the position information of the anchor points and the mapped positions, and used to predict the object categories in the candidate areas, simplifying the data enhancement process.
It enriches feature expression, simplifies the model training process, improves the performance and generalization capabilities of the object detection model, and does not require pre-data enhancement of the sample image itself.
Smart Images

Figure CN116612295B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, in particular to computer vision, deep learning, image processing and other technical fields, and specifically relates to a method for generating a feature map, a method for training an object detection model, and an apparatus therefor. Background Art
[0002] In vision tasks, such as object detection tasks, the more the number of image samples, the better the effect of the trained object detection model and the stronger the generalization ability of the model. However, in practice, the number of samples is often insufficient or the sample quality is not good enough, which requires data augmentation of the samples to improve the performance of the model. Summary of the Invention
[0003] The present disclosure provides a method for generating a feature map, a method for training an object detection model, and an apparatus therefor.
[0004] According to one aspect of the present disclosure, there is provided a method for generating a feature map, the method including: partitioning candidate regions of a sample image; randomly determining anchor points in sub-regions obtained by partitioning the candidate regions; determining feature values of the anchor points according to position information of the anchor points in the sub-regions and mapping positions of the candidate regions in a first feature map, where the first feature map is a feature map of the sample image; and generating a second feature map of the candidate regions according to the feature values of the anchor points, where the second feature map is used to predict the category of an object in the candidate regions.
[0005] According to another aspect of the present disclosure, there is provided a method for training an object detection model, the method including: obtaining a sample image and annotation information of the sample image; inputting the sample image into the object detection model to obtain a prediction result of the sample image; and in the process of obtaining the prediction result of the sample image, the object detection model is used to obtain a second feature map of candidate regions in the sample image according to the method for generating a feature map described in any one of the above embodiments, where the second feature map is used to obtain the prediction result of the sample image; and adjusting parameters of the object detection model according to the prediction result of the sample image and the annotation information of the sample image to obtain a trained object detection model.
[0006] According to another aspect of the present disclosure, there is provided an object detection method, including: obtaining a to-be-detected image; and inputting the to-be-detected image into the trained object detection model to obtain a detection result of the to-be-detected image, where the trained object detection model is trained by using the training method described in any one of the above embodiments.
[0007] According to another aspect of the present disclosure, there is provided a feature map generation device, including: a partitioning unit for partitioning candidate regions of a sample image; a first determination unit for randomly determining anchor points in sub-regions obtained by partitioning the candidate regions; a second determination unit for determining feature values of the anchor points according to the position information of the anchor points in the sub-regions and the mapping positions of the candidate regions in a first feature map, where the first feature map is the feature map of the sample image; a generation unit for generating a second feature map of the candidate regions according to the feature values of the anchor points, and the second feature map is used to predict the category of the object in the candidate regions.
[0008] According to another aspect of the present disclosure, there is provided a training device for an object detection model, including: a sample acquisition unit for acquiring a sample image and annotation information of the sample image; a result prediction unit for inputting the sample image into the object detection model to obtain a prediction result of the sample image; and in the process of obtaining the prediction result of the sample image, the object detection model is used to obtain a second feature map of candidate regions in the sample image according to the feature map generation method described in any one of the above embodiments, and the second feature map is used to obtain the prediction result of the sample image; an adjustment unit for adjusting parameters of the object detection model according to the prediction result of the sample image and the annotation information of the sample image to obtain a trained object detection model.
[0009] According to another aspect of the present disclosure, there is provided an object detection device, including: an image acquisition unit for acquiring a to-be-detected image; a detection unit for inputting the to-be-detected image into the trained object detection model to obtain a detection result of the to-be-detected image, where the trained object detection model is trained by using the training method described in any one of the above embodiments.
[0010] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any method in the embodiments of the present disclosure.
[0011] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.
[0012] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements any method in the embodiments of the present disclosure.
[0013] The feature map generation method, the training method and device of the object detection model provided by the embodiments of the present disclosure partition the candidate regions of the sample image; randomly determine anchor points in the sub-regions obtained by partitioning the candidate regions; determine the feature values of the anchor points according to the position information of the anchor points in the sub-regions and the mapping positions of the candidate regions in the first feature map, where the first feature map is the feature map of the sample image; generate the second feature map of the candidate regions according to the feature values of the anchor points, and the second feature map is used to predict the category of the object in the candidate regions. Since the positions of the anchor points are randomly determined, in the process of training the object detection model, each parameter adjustment process can be randomly assigned to different anchor points, and thus more feature information in the first feature map can be used for training, enriching the feature expression, thereby realizing data augmentation at the feature level, simplifying the processing process of the sample pictures and the training process of the model, and further improving the performance and generalization ability of the trained object detection model.
[0014] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0016] Figure 1 is a schematic flowchart of a feature map generation method according to an embodiment of the present disclosure;
[0017] Figure 2 is a schematic diagram of a feature map generation method in the related art;
[0018] Figure 3 is a schematic diagram of a feature map generation method according to an embodiment of the present disclosure;
[0019] Figure 4 is a schematic diagram of a feature map generation method according to another embodiment of the present disclosure;
[0020] Figure 5 is a schematic diagram of a feature map generation method according to another embodiment of the present disclosure;
[0021] Figure 6 is a schematic flowchart of a training method of an object detection model according to an embodiment of the present disclosure;
[0022] Figure 7 is a schematic flowchart of an object detection method according to an embodiment of the present disclosure;
[0023] Figure 8It is a structural diagram of a feature map generation device provided according to an embodiment of the present disclosure;
[0024] Figure 9 It is a structural diagram of a training device for an object detection model provided according to an embodiment of the present disclosure;
[0025] Figure 10 It is a structural diagram of an object detection device provided according to an embodiment of the present disclosure;
[0026] Figure 11 It is a block diagram of an electronic device for implementing the feature map generation method according to an embodiment of the present disclosure. Detailed implementation manners
[0027] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0028] Currently, deep learning technology has been actively researched in the field of artificial intelligence and has demonstrated excellent performance in fields such as image classification, object detection, and semantic segmentation. In visual tasks, generally, the more the number of image samples, the better the performance of the trained model and the stronger the generalization ability of the model. However, in practice, it is often encountered that the number of samples is insufficient or the quality of the samples is not good enough. This requires data augmentation of the samples to increase the data volume and, to a certain extent, improve the quality of the samples, thereby improving the model performance. Especially in object detection, data augmentation is a crucial link and plays a key role in the detection performance of the model. How to reasonably and efficiently implement data augmentation to improve the performance of the object detection model has also become a research hotspot.
[0029] In related technologies, data augmentation in object detection usually uses the following schemes: 1. Perform perturbations at the RGB (Red Green Blue) pixel level on the entire image, such as adjusting brightness, contrast, etc.; 2. Perform scale transformation on the entire image, such as randomly adjusting the image size, rotating the image, etc.; 3. Perform fusion on the target region of the image, such as using the MixUp (fusion) method, etc.
[0030] However, the above methods are all data augmentations performed on image samples, and then the image samples after data augmentation are input into the object detection model for training, and the process is relatively complex.
[0031] To solve at least one of the above problems, embodiments of the present disclosure provide a feature map generation method, a training method and apparatus for an object detection model. The method partitions candidate regions of a sample image, randomly determines anchor points in sub-regions obtained by partitioning the candidate regions, determines the feature values of the anchor points according to the position information of the anchor points in the sub-regions and the mapping positions of the candidate regions in a first feature map, where the first feature map is the feature map of the sample image, and generates a second feature map of the candidate regions according to the feature values of the anchor points. The second feature map is used to predict the category of an object in the candidate region. Since the positions of the anchor points are randomly determined, during the training process of the object detection model, different anchor points can be randomly selected in each parameter adjustment process. Thus, more feature information in the first feature map can be utilized for training, enriching the feature representation, thereby achieving data augmentation at the feature level, simplifying the processing process of sample images and the training process of the model, and further improving the performance and generalization ability of the trained object detection model.
[0032] The following elaborates on the embodiments of the present disclosure in detail with reference to the accompanying drawings. Figure 1 It is a schematic flowchart of a feature map generation method provided according to an embodiment of the present disclosure. Please refer to Figure 1 Embodiments of the present disclosure provide a feature map generation method 100. The method 100 includes the following steps S101 to S104.
[0033] Step S101: Partition the candidate regions of the sample image.
[0034] Step S102: Randomly determine anchor points in the sub-regions obtained by partitioning the candidate regions.
[0035] Step S103: Determine the feature values of the anchor points according to the position information of the anchor points in the sub-regions and the mapping positions of the candidate regions in a first feature map, where the first feature map is the feature map of the sample image.
[0036] Step S104: Generate a second feature map of the candidate regions according to the feature values of the anchor points. The second feature map is used to predict the category of an object in the candidate region.
[0037] Among them, the method 100 can be used in the training process of an object detection model. The object detection model can be used to identify multiple objects in a picture and can also locate different objects. The object detection model can be a one-stage object detection model or a two-stage object detection model, etc. The method 100 is used in the training process of the object detection model.
[0038] Taking a two-stage object detection model as an example, the model can be roughly divided into two stages. After the sample image is input into the object detection model, in the first stage, the object regions can be extracted, that is, the respective candidate regions (Proposals) of the sample image are obtained. In addition, in this stage, the first feature map of the sample image can also be obtained. In the second stage, the objects can be classified and recognized.
[0039] A candidate region refers to a relatively small region that may contain an object to be recognized or classified, which can be represented as a selection box in the sample image and frame the object of interest that may be contained in the sample image. The first feature map is the feature map corresponding to the sample image, which records the feature information in the sample image. Both the candidate region and the first feature map are the intermediate layer outputs of the object detection model. It can be understood that the sizes of the respective candidate regions generated from the sample image are not fixed, and method 100 can generate respective second feature maps of a specific size from the respective candidate regions of different sizes, that is, the sizes of the respective second feature maps generated from the respective candidate regions are fixed. It can be understood that the specific size can be preset, such as 2 × 2, or, 7 × 7 and so on.
[0040] In step S101, after obtaining the candidate regions, the candidate regions can be partitioned to obtain multiple sub-regions. It can be understood that the number of sub-regions is the same as the size of the second feature map to be generated. For example, to generate a second feature map of 2 × 2, the candidate regions need to be divided into 2 × 2 sub-regions, and the sizes of each sub-region can be different or the same.
[0041] In step S102, it is necessary to randomly determine anchor points in each sub-region. It can be understood that the position of the anchor point 222 is randomly generated by the system and its position is not fixed. After the anchor point is generated, the position information of the anchor point in the sub-region can be obtained, that is, the coordinates of the anchor point.
[0042] It can be understood that the mapping position of the candidate region in the first feature map can be obtained by mapping the candidate region to the first feature map. During the mapping process, if the sample image is an 800 × 800 picture, and there is an initial candidate region of 665 × 665 (with a dog framed inside) in the upper middle of the picture. After the picture extracts features through the object detection model, the scaling step size of the first feature map is 32. Therefore, the side lengths of the scaled sample image (i.e., the first feature map) and the scaled initial candidate region (i.e., the candidate region) are both 1 / 32 of the input. That is, the first feature map becomes 25 × 25. The candidate region will become 20.78 ×20.78 (retaining the decimal part), the candidate region still remains in the center of the first feature map. After mapping the candidate region to the first feature map, the positional relationship and ratio between the two are the same as those between the initial candidate region and the sample image. Additionally, the step of mapping the candidate region to the first feature map can be performed before step S101.
[0043] In step S103, based on the position information of the anchor point in the sub-region and the mapped position of the candidate region in the first feature map, the feature value of the anchor point is determined.
[0044] In step S104, the feature values of each sub-region can be determined through the feature values of the anchor points in the sub-region, thereby generating the second feature map.
[0045] The second feature map can be used to predict the category of the object in the candidate region, and can also be used to predict the confidence level of this category and the position information of the object. For example, within the candidate region is a dog, and the confidence level (probability) that the object is a dog, as well as the position coordinates of the dog in the sample image.
[0046] It can be understood that method 100 can be implemented in the ROI Align (Region of Interest Align) network of the object detection model, and the generated second feature map can also be transformed into a discrete point feature set, that is, the discrete point feature set corresponding to the candidate region. Subsequently, the discrete point feature set can be input into multiple fully connected layers of the object detection model for classification or regression, etc., to obtain the classification score (confidence level) and position information (localization coordinates) of the final detected object, which can be specifically set according to the actual situation.
[0047] Figure 2 is a schematic diagram of a feature map generation method in the related art, please refer to Figure 2 , Figure 2 shows the mapped position of the candidate region in the first feature map in the related art. The dashed box 210 represents the first feature map, and the first feature map includes multiple grids 211. The solid box 220 represents the candidate region, and the candidate region is partitioned into 2 × 2 sub-regions 221. The black dots 222 represent the anchor points in the sub-regions, and the white dots 212 represent the feature values at the corner points of the grids. In the related art, the ROI Align scheme is adopted. When generating the second feature map of the candidate region 220, the midpoint of the sub-region 221 is selected as the anchor point 222, that is, the point located at the geometric center of the sub-region 221, and then the feature values 212 of the four corner points of the grid corresponding to the anchor point 222 are used to calculate the feature value of the anchor point 222. During multiple parameter adjustment processes in model training, each time the feature value of the anchor point is calculated using the same four corner points, and less feature information in the first feature map is utilized.
[0048] Figure 3 It is a schematic block diagram of a feature map generation method provided according to an embodiment of the present disclosure. Please refer to Figure 3 , Figure 3 which shows the mapping position of the candidate region in the first feature map. The dashed box 310 represents the first feature map, and the first feature map includes multiple grids 311. The solid box 320 represents the candidate region, and the candidate region is partitioned into 2 × 2 sub-regions 321. The black dots 322 represent the anchor points in the sub-regions, and the white dots 312 represent the feature values at the corner points of the grids.
[0049] After partitioning the candidate region 320 into 2 × 2 sub-regions 321, taking the upper left sub-region 321 as an example, a random anchor point 322 is generated within the sub-region 321. Then, according to the position information of the anchor point 322 in the sub-region and the mapping position of the candidate region in the first feature map, the grid 311 corresponding to the anchor point 322 can be found. Then, the feature values of the four vertices of the grid 311 are used to calculate the feature value of the anchor point 322, and the feature value of the anchor point 322 is used as the feature value of the upper left sub-region 321. In this way, the feature values of the four sub-regions in the candidate region 320 can be determined, that is, the second feature map of the candidate region is determined.
[0050] It can be understood that during the training process of the object detection model, a sample image needs to be input, then candidate regions are generated through the sample image, and then the second feature map of the candidate region is calculated by method 100. Then, the category of the object in the candidate region is predicted using the feature map. Then, the loss function of the object detection model is calculated. When the loss function does not meet the conditions, the parameters of the model need to be adjusted, and then the steps of generating candidate regions, calculating the second feature map of the candidate region, predicting the category, and calculating the loss function are repeated. That is, the above training process needs to be continuously repeated to make the time function meet the conditions.
[0051] As Figure 3 shown, the upper left sub-region 321 overlaps with 6 grids 311 in the first feature map (that is, the features of the sub-region are related to these 6 grids). If the method of Figure 2 is used to select the anchor point, the feature value of the anchor point is always related to the same one of the 6 grids 211, while in the method 100 provided in this embodiment, during multiple adjustments of the parameters, each time the anchor point can be randomly selected into any one of these 6 grids. That is, during multiple iterations of the model, the feature values of the corner points of the 6 grids can all be used to calculate the feature value of the anchor point, thereby using more features in the first feature map and achieving data augmentation at the feature level.
[0052] It can be understood that in the related art, since the midpoint of the sub-region is used as the anchor point, that is, in multiple repeated processes of model training, the same anchor point is always selected for each sub-region, and the position information of the anchor point remains unchanged. Therefore, the grid corresponding to the anchor point in the first feature map is fixed, that is, the features in the first feature map used in the calculation process of the eigenvalue of the anchor point do not change either. Figure 2 (the eigenvalues of the four corner points 212 of the grid in
[0053] ) result in relatively single features of the first feature map used.
[0054] Figure 4 is a schematic diagram of a feature map generation method according to another embodiment of the present disclosure. Please refer to Figure 4 , in some embodiments, randomly determining an anchor point in the sub-region obtained by dividing the candidate region in step S102 may include: dividing the sub-region into multiple units; randomly determining an anchor point in each of the multiple units.
[0055] For example, Figure 4 , the candidate region 420 is partitioned into 2 × 2 sub-regions 421. Then, each sub-region 421 can be divided into 4 units 423. For example, Figure 4 for the upper left sub-region 421 in × , it can be evenly divided into 2
[0056] 2 units 423 by a thick dashed line, and an anchor point 422 can be randomly determined in each unit 423. Figure 3 ) means that an anchor point is randomly selected for the entire region. Or, the sub-region can be divided into multiple units ( Figure 4 ) means that multiple anchor points can be randomly selected for the entire region. Since the grids 411 where each unit intersects with the first feature map 410 are different, by dividing the units, it is also possible to make the anchor points more likely to be randomly and relatively evenly located in the grids related to the sub-region, so that more features in the first feature map can be utilized, further enriching the expression of features.
[0057] Please continue to refer to Figure 4, in some embodiments, in step S104, generating a second feature map of the candidate region according to the eigenvalue of the anchor point may include the following steps: taking the eigenvalue of the anchor point as the eigenvalue of the unit where the anchor point is located, and averaging the eigenvalues of multiple units in the sub-region to obtain the eigenvalue of the sub-region; generating a second feature map of the candidate region according to the eigenvalue of the sub-region in the candidate region.
[0058] Taking Figure 4 the upper left sub-region 421 in
[0059] as an example, after randomly determining the anchor point 422 in each unit 423, the grid 411 in the first feature map 410 corresponding to the anchor point 423 can be determined according to the position information of the anchor point and the mapping position of the candidate region in the first feature map, and then the eigenvalue of the anchor point 422 is calculated by using the eigenvalues of the four corner points 412 of the grid 411. Figure 4 Then, the calculated eigenvalue of the anchor point 422 is taken as the eigenvalue of the unit where the anchor point is located. In this way, the eigenvalues of multiple units in each sub-region can be determined. For example,
[0060] the eigenvalues of the four units of the upper left sub-region 421 respectively. Then, the eigenvalue of the sub-region 421 can be calculated by averaging the eigenvalues of these four units. × After determining the eigenvalues of the four sub-regions of the candidate region 420, a 2
[0061] × 2 second feature map can be obtained.
[0062] Of course, in other embodiments, the maximum value of the eigenvalues of multiple units in the sub-region can also be taken as the eigenvalue of the sub-region.
[0063] By comprehensively determining the eigenvalue of the sub-region based on the eigenvalues of multiple units in the sub-region, and then generating the second feature map of the candidate region, the second feature map can better represent the features of the candidate region, thereby improving the accuracy of subsequent prediction results.
[0063] Continuing to refer to Figure 3 , in some embodiments, in step S103, determining the eigenvalue of the anchor point according to the position information of the anchor point in the sub-region and the mapping position of the candidate region in the first feature map includes the following step 1 and step 2.
[0064] Step 1, determining the target grid corresponding to the anchor point from multiple grids included in the first feature map according to the position information of the anchor point in the sub-region and the mapping position of the candidate region in the first feature map.
[0065] Step 2, determining the eigenvalue of the anchor point by using the bilinear interpolation algorithm according to the feature information of the target grid.
[0066] For example Figure 3As shown, the first feature map 310 may include an 8×8 grid, and the feature information of each grid includes the feature values of the four corner points of the grid.
[0067] After randomly selecting the anchor point 322, since the mapping position of the candidate region 320 in the first feature map 310 is known and the position of the anchor point 322 in the sub-region 322 is also known, the positional relationship of the anchor point 322 relative to the first feature map 310 can be obtained, and then the target grid 311 corresponding to the anchor point 322, that is, the grid where the anchor point 322 is located, can be determined.
[0068] Then, the feature information of the target grid 311, that is, the feature values at its four corner points 312, can be used to calculate the feature value of the anchor point 322 by using the bilinear interpolation method.
[0069] The bilinear interpolation algorithm is a method for calculating the feature value of the anchor point by using 4 points in the first feature map. The calculation result is relatively accurate and the calculation speed is also relatively fast, which is beneficial to improving the calculation speed of the model and the accuracy of the prediction result.
[0070] Similarly, as in Figure 4 each anchor point 422 in the unit can find its corresponding target grid 411 in the first feature map 410, and then the feature value of the anchor point 422 is calculated by using the feature values of the four corner points 412 of the target grid.
[0071] Of course, in addition to the bilinear interpolation algorithm, other interpolation algorithms such as the bicubic interpolation algorithm and the nearest neighbor method can also be used to calculate the feature value of the anchor point.
[0072] Figure 5 is a schematic diagram of a feature map generation method according to another embodiment of the present disclosure. Please refer to Figure 5 , on the basis of the above embodiment, before partitioning the candidate region of the sample image in step S101, the method further includes: randomly expanding the size of the candidate region.
[0073] It can be understood that the candidate region can be a rectangular box. After obtaining the candidate region, the size of the rectangular box can be expanded in a random manner, that is, the expansion coefficient can be any number. Then, the expanded candidate region is used to generate the second feature map.
[0074] After generating the candidate region in the one-stage of the object detection model, the candidate region can be randomly expanded. Figure 5 In , the solid line box 520 represents the expanded candidate region, and the thick dashed line box 530 represents the candidate region before expansion. After the candidate region is randomly expanded, the expanded candidate region can be partitioned, and the anchor point is randomly determined in the sub-region, and then the feature value of the anchor point is calculated, and then the second feature map of the expanded candidate region is generated.
[0075] It can be understood that the range of the expanded candidate region in the first feature map 510 is relatively large. When calculating the feature value of the anchor point 522, it will utilize some grids outside the original candidate region 530, so that the semantic information of the context can be considered. That is, when making predictions, the features of the background region outside the candidate region can be comprehensively utilized, which is beneficial to improving the accuracy of result prediction.
[0076] For example, if a candidate region in the image to be detected frames a shadow, it may be difficult to accurately predict what object the shadow belongs to just from the shadow itself. However, by expanding the size of the candidate region, the content of the background region outside the candidate region in the image to be detected can be utilized to accurately predict that the shadow is, for example, the shadow of a person or a building.
[0077] In some embodiments, randomly expanding the size of the candidate region may include the following sub-steps 1 to 3.
[0078] Sub-step 1, randomly generate an expansion coefficient, and the expansion coefficient is greater than 1.
[0079] Sub-step 2, according to the expansion coefficient, the length of the candidate region, and the width of the candidate region, determine the target length and target width of the expanded candidate region;
[0080] Sub-step 3, while keeping the center position of the candidate region unchanged, expand the candidate region according to the target length and target width.
[0081] The expansion coefficient can be a random number randomly generated by the system. In order to expand the candidate region, the expansion coefficient should take any value greater than 1.
[0082] The expansion coefficient can be the expansion coefficient of the length and width of the candidate region, that is, the target length of the expanded candidate region = the length of the original candidate region × the expansion coefficient, and the target width of the expanded candidate region = the width of the original candidate region × the expansion coefficient.
[0083] Then, the center position of the candidate region can be kept unchanged, and the candidate region can be expanded outwards, that is, the length and width of the candidate region become the target length and target width respectively, so as to realize the expansion of the candidate region, and it is relatively simple and easy to implement.
[0084] Similarly, it can be understood that during the training process of the target detection model, due to multiple parameter adjustments, in each parameter adjustment process, the expansion coefficients randomly obtained by the candidate regions are all different, so that more context semantic information can be utilized to realize data enhancement at the feature level and enrich the feature expression.
[0085] Continue to refer to Figure 5, in a specific embodiment, the following two methods can be used to achieve data augmentation at the feature level.
[0086] (1) Random perturbation and interpolation of regional features
[0087] Taking the two-stage object detector as an example of the object detection model (such as Faster RCNN (Faster Region-Convolutional-Neural-Network, Fast Region Convolutional Neural Network)). After obtaining the candidate regions (proposals) in the first stage, it is necessary to aggregate the features corresponding to the candidate regions, that is, divide the candidate region features into multiple sub-regions, select one or more anchor points in each sub-region in a uniform distribution manner, and obtain the feature values of the corresponding anchor points by bilinear interpolation. In this embodiment, during the training stage of the object detection model, the anchor point positions in the sub-regions are randomly perturbed, which is equivalent to applying different weights to the features at each corner of the grid according to the position information, enriching the feature expression.
[0088] (2) Randomly expand the boundaries of candidate regions
[0089] To further improve the feature richness, while randomly perturbing the anchor point positions, the boundaries of the candidate regions are further randomly expanded outward. When expanded within an appropriate range, this operation can better utilize the context semantic information of the candidate regions and enrich the feature expression.
[0090] The object detection data augmentation method based on regional feature interpolation provided in this embodiment can be used in a two-stage object detector to perform data augmentation at the feature level and improve the performance of the detection model. The feature enhancement method of this embodiment can innovatively perform data augmentation at the feature level, make more full use of the context semantic information corresponding to the target region on the feature map, and can more efficiently improve the performance of object detection.
[0091] Figure 6 is a schematic flowchart of a training method for an object detection model according to an embodiment of the present disclosure; please refer to Figure 6 , the present disclosure embodiment also provides a training method 600 for an object detection model. The method 600 includes the following steps S601 to S603.
[0092] Step S601, obtain a sample image and the annotation information of the sample image.
[0093] Step S602, input the sample image into the object detection model to obtain the prediction result of the sample image; and during the process of obtaining the prediction result of the sample image, the object detection model is used to obtain the second feature map of the candidate regions in the sample image according to the feature map generation method described in any one of the above embodiments, and the second feature map is used to obtain the prediction result of the sample image.
[0094] Step S603: Adjust the parameters of the object detection model according to the prediction result of the sample image and the annotation information of the sample image to obtain a trained object detection model.
[0095] The object detection model can be used to identify multiple objects in a picture and can also locate different objects.
[0096] The sample image can be a picture used for training the object detection model. The annotation information of the sample image can include the true category of the object in each candidate region of the sample image and the true position information of the object, etc.
[0097] Input the sample image into the object detection model to obtain a prediction result. It can be understood that the object detection model can first use the sample image to obtain the candidate regions of the sample image, and then use the feature map generation method described in any of the above embodiments to obtain the second feature map corresponding to the candidate regions. Then use the second feature map to obtain the prediction result.
[0098] The prediction result can include the predicted category of each object included in the sample image, the confidence level (classification score) of the predicted category, and the predicted position information (localization coordinates, that is, the position coordinates of the object in the sample image) of the object.
[0099] Then, the loss function can be calculated using the annotation information and the prediction result. In the case where the loss function does not meet the convergence condition, the parameters of the object detection model can be adjusted. Then repeat the above step S602 until the loss function meets the convergence condition to obtain a trained object detection model. The trained object detection model can be used to detect the image to be tested to obtain the category of the object in the image to be tested, the confidence level of the category, and the position information of the object.
[0100] It can be understood that since the feature map generation method of any of the above embodiments is adopted in the process of outputting the prediction result, when adjusting the parameters of the model, anchor points at different positions within the sub-region can be randomly selected. Since the grids in the first feature map corresponding to each anchor point may be different, more grids in the first feature map, that is, the features in the first feature map, can be utilized during the model training process, thereby enriching the feature expression at the feature level, realizing data augmentation at the feature level, and further improving the performance of the model. And it is not necessary to perform data augmentation on the sample image itself in the way of the related art in advance, the process is simpler, and the model training process is simplified.
[0101] In some embodiments, the object detection model includes a first intermediate layer, a second intermediate layer, and a third intermediate layer; inputting the sample image into the object detection model in step S602 to obtain the prediction result of the sample image may include: inputting the sample image into the first intermediate layer to obtain the candidate regions of the sample image and the first feature map of the sample image; inputting the candidate regions and the first feature map into the second intermediate layer to obtain the second feature map of the candidate regions; inputting the second feature map into the third intermediate layer to obtain the prediction result of the sample image.
[0102] The object detection model may include a first intermediate layer, a second intermediate layer, and a third intermediate layer, and at least one of the intermediate layers may be composed of one or more neural networks.
[0103] It can be understood that the first intermediate layer may include a backbone network and a Region Proposal Network (RPN). When the sample image is input into the backbone network, the first feature map of the sample image can be obtained. Then, inputting the first feature map into the region candidate network can obtain each candidate region in the sample image, and the sizes of the candidate regions are usually inconsistent.
[0104] The second intermediate layer may include an ROI Align network. The second intermediate layer can be used to perform the feature map generation method described in any of the above embodiments, that is, partitioning the candidate regions of the sample image; randomly determining anchor points in the sub-regions obtained by partitioning the candidate regions; determining the feature values of the anchor points according to the position information of the anchor points in the sub-regions and the mapping position of the candidate regions in the first feature map; generating the second feature map of the candidate regions according to the feature values of the anchor points.
[0105] In some embodiments, when the candidate regions and the first feature map are input into the second intermediate layer, the second intermediate layer can map the candidate regions to the first feature map, find the features of the corresponding regions from the first feature map using the coordinates of the candidate regions, and then use the ROI Align network to sample points (select anchor points) from the features of various candidate regions with different shapes to generate the second feature maps of each candidate region, and the sizes of these second feature maps are the same.
[0106] The third intermediate layer may include a plurality of fully connected layers, which can use the second feature map to implement classification and bounding box regression to obtain the prediction result. For example, the third intermediate layer can convert the second feature map into a discrete point feature set. Since the sizes of the second feature maps are the same, the lengths of the discrete point feature sets corresponding to the obtained candidate regions are also the same. Then, inputting the discrete point feature set into the plurality of fully connected layers can obtain the prediction result.
[0107] Then, the loss function can be calculated using the annotation information and the prediction results. In the case where the loss function does not meet the convergence condition, the parameters of the object detection model can be adjusted. For example, the parameters of the backbone network, the region proposal network, and multiple fully connected layers can be adjusted until the loss function meets the convergence condition to obtain a trained object detection model.
[0108] In this embodiment, the second intermediate layer can be used to implement data augmentation at the feature level, thereby improving the performance of the object detection model and simplifying the training process.
[0109] Figure 7 is a schematic flowchart of an object detection method provided by an embodiment of the present disclosure; please refer to Figure 7 In this embodiment, an object detection method 700 is provided, including the following steps S701 and S702.
[0110] Step S701, obtain an image to be detected.
[0111] Step S702, input the image to be detected into the trained object detection model to obtain the detection result of the image to be detected, where the trained object detection model is trained by using the object detection model training method described in any of the above embodiments.
[0112] The image to be detected can be a picture that needs to be subjected to object detection. The trained object detection model can be trained by using the object detection model training method provided in any of the above embodiments. The detection result can include the category of the object included in the image to be detected, the confidence level of this category, and the position information of the object, etc., and the accuracy is high.
[0113] It can be understood that both the random selection of anchor points and the random expansion of the size of the candidate region in the above feature map generation method are used in the process of model training. After the model training is completed, that is, during the object detection process of the image to be detected, the position of the anchor point can still adopt the traditional ROI Align scheme, that is, the anchor point is taken as the midpoint of the sub-region (or the unit in the sub-region), and the candidate region does not need to be randomly expanded either, so as to accurately predict the image to be detected.
[0114] Figure 8 is a structural diagram of a feature map generation device provided by an embodiment of the present disclosure; please refer to Figure 8 In an embodiment of the present disclosure, a feature map generation device 800 is provided, and the device 800 includes the following units.
[0115] A partitioning unit 801, configured to partition the candidate region of the sample image.
[0116] A first determination unit 802, configured to randomly determine an anchor point in the sub-region obtained by partitioning the candidate region.
[0117] A second determination unit 803, configured to determine a feature value of an anchor point according to the position information of the anchor point in the sub-region and the mapping position of the candidate region in the first feature map, where the first feature map is the feature map of the sample image.
[0118] A generation unit 804, configured to generate a second feature map of the candidate region according to the feature value of the anchor point, where the second feature map is used to predict the category of the object in the candidate region.
[0119] In some embodiments, the first determination unit 802 is further configured to: divide the sub-region into multiple cells; randomly determine an anchor point in each of the multiple cells.
[0120] In some embodiments, the generation unit 804 is further configured to: use the feature value of the anchor point as the feature value of the cell where the anchor point is located, and average the feature values of the multiple cells in the sub-region to obtain the feature value of the sub-region; generate a second feature map of the candidate region according to the feature value of the sub-region in the candidate region.
[0121] In some embodiments, the second determination unit 803 is further configured to: determine a target grid corresponding to the anchor point from multiple grids included in the first feature map according to the position information of the anchor point in the sub-region and the mapping position of the candidate region in the first feature map; determine the feature value of the anchor point by using a bilinear interpolation algorithm according to the feature information of the target grid.
[0122] In some embodiments, before the partitioning unit 801 is configured to partition the candidate region of the sample image, the apparatus further includes: an expansion unit, configured to randomly expand the size of the candidate region.
[0123] In some embodiments, the expansion unit is further configured to: randomly generate an expansion coefficient, where the expansion coefficient is greater than 1; determine a target length and a target width of the expanded candidate region according to the expansion coefficient, the length of the candidate region, and the width of the candidate region; expand the candidate region according to the target length and the target width while keeping the center position of the candidate region unchanged.
[0124] The feature map generation device provided in this embodiment partitions the candidate regions of the sample image; randomly determines anchor points in the sub-regions obtained by partitioning the candidate regions; determines the feature values of the anchor points according to the position information of the anchor points in the sub-regions and the mapping positions of the candidate regions in the first feature map, where the first feature map is the feature map of the sample image; generates a second feature map of the candidate regions according to the feature values of the anchor points, and the second feature map is used to predict the category of the object in the candidate region. Since the positions of the anchor points are randomly determined, during the training process of the target detection model, each parameter adjustment process can be randomly assigned to different anchor points, so that more feature information in the first feature map can be used for training, enriching the feature expression, thereby realizing data augmentation at the feature level, simplifying the processing process of the sample pictures and the training process of the model, and further improving the performance and generalization ability of the trained target detection model.
[0125] Figure 9 is a structural diagram of a training device for a target detection model provided according to an embodiment of the present disclosure; please refer to Figure 9 As shown in, the present disclosure provides a training device 900 for a target detection model, and the device 900 includes the following units.
[0126] A sample acquisition unit 901, configured to acquire a sample image and the annotation information of the sample image;
[0127] A result prediction unit 902, configured to input the sample image into the target detection model to obtain the prediction result of the sample image; and in the process of obtaining the prediction result of the sample image, the target detection model is used to obtain a second feature map of the candidate regions in the sample image according to the feature map generation method described in any one of the above embodiments, and the second feature map is used to obtain the prediction result of the sample image.
[0128] An adjustment unit 903, configured to adjust the parameters of the target detection model according to the prediction result of the sample image and the annotation information of the sample image to obtain a trained target detection model.
[0129] In some embodiments, the target detection model includes a first intermediate layer, a second intermediate layer, and a third intermediate layer; the result prediction unit 902 is further configured to: input the sample image into the first intermediate layer to obtain the candidate regions of the sample image and the first feature map of the sample image; input the candidate regions and the first feature map into the second intermediate layer to obtain a second feature map of the candidate regions; input the second feature map into the third intermediate layer to obtain the prediction result of the sample image.
[0130] Figure 10 is a structural diagram of a target detection device provided according to an embodiment of the present disclosure; please refer to Figure 10 As shown in, the present disclosure provides a target detection device 1000, including the following units.
[0131] An image acquisition unit 1001 for acquiring an image to be measured.
[0132] A detection unit 1002 for inputting the image to be measured into a trained object detection model to obtain a detection result of the image to be measured, where the trained object detection model is trained by using the object detection model training method described in any of the above embodiments.
[0133] For the specific functions and examples of the modules and sub-modules of the device in the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated here.
[0134] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0135] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0136] The embodiment of the present disclosure provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any of the above embodiments.
[0137] The embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method described in any of the above embodiments.
[0138] The embodiment of the present disclosure provides a computer program product, including a computer program, and the computer program implements the method described in any of the above embodiments when executed by a processor.
[0139] Figure 11 is a block diagram of an electronic device for implementing the feature map generation method in the embodiments of the present disclosure. Please refer to Figure 11 , the electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0140] As Figure 11As shown, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the device 1100 can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0141] Multiple components in the device 1100 are connected to the I / O interface 1105, including: an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a magnetic disk, an optical disk, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows the device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0142] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1101 executes the various methods and processes described above, such as a feature map generation method, a training method of an object detection model, and an object detection method. For example, in some embodiments, the feature map generation method, the training method of the object detection model, and the object detection method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the feature map generation method, the training method of the object detection model, and the object detection method described above can be executed. Alternatively, in other embodiments, the computing unit 1101 can be configured to execute the feature map generation method, the training method of the object detection model, and the object detection method in any other appropriate manner (e.g., by means of firmware).
[0143] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0144] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0145] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0146] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, speech input, or tactile input).
[0147] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0148] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client - server relationship is generated by computer programs running on the respective computers and having a client - server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server combined with a blockchain.
[0149] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0150] The above - described specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for generating a feature map, the method comprising: Partitioning candidate regions of a sample image; Randomly determining anchor points in sub-regions obtained by partitioning the candidate regions; Determining the feature values of the anchor points according to the position information of the anchor points in the sub-regions and the mapping positions of the candidate regions in a first feature map, wherein the first feature map is the feature map of the sample image; Taking the feature values of the anchor points as the feature values of the units where the anchor points are located, and averaging the feature values of multiple units in the sub-regions to obtain the feature values of the sub-regions; Generating a second feature map of the candidate regions according to the feature values of the sub-regions in the candidate regions, the second feature map being used to predict the category of an object in the candidate regions.
2. The method according to claim 1, wherein Randomly determining anchor points in sub-regions obtained by partitioning the candidate regions, including: Dividing the sub-regions into multiple units; Randomly determining anchor points in each of the multiple units.
3. The method according to claim 1 or 2, wherein Determining the feature values of the anchor points according to the position information of the anchor points in the sub-regions and the mapping positions of the candidate regions in the first feature map, including: Determining a target grid corresponding to the anchor point from multiple grids included in the first feature map according to the position information of the anchor point in the sub-region and the mapping position of the candidate region in the first feature map; Determining the feature values of the anchor points by using a bilinear interpolation algorithm according to the feature information of the target grid.
4. The method according to any one of claims 1-3, before partitioning candidate regions of a sample image, the method further comprising: Randomly expanding the size of the candidate regions.
5. The method according to claim 4, wherein, Randomly expanding the size of the candidate regions, including: Randomly generating an expansion coefficient, the expansion coefficient being greater than 1; Determining the target length and target width of the expanded candidate regions according to the expansion coefficient, the length of the candidate regions, and the width of the candidate regions; Expanding the candidate regions according to the target length and the target width while keeping the central position of the candidate regions unchanged.
6. A method for training an object detection model, the method comprising: Obtaining a sample image and the annotation information of the sample image; Inputting the sample image into the object detection model to obtain the prediction result of the sample image; And in the process of obtaining the prediction result of the sample image, the object detection model is used to obtain a second feature map of candidate regions in the sample image according to the method according to any one of claims 1-5, the second feature map being used to obtain the prediction result of the sample image; Adjusting the parameters of the object detection model according to the prediction result of the sample image and the annotation information of the sample image to obtain a trained object detection model.
7. The method according to claim 6, wherein, The object detection model includes a first intermediate layer, a second intermediate layer, and a third intermediate layer; Inputting the sample image into the object detection model to obtain the prediction result of the sample image, including: Inputting the sample image into the first intermediate layer to obtain the candidate regions of the sample image and the first feature map of the sample image; Input the candidate region and the first feature map into the second intermediate layer to obtain the second feature map of the candidate region; Input the second feature map into the third intermediate layer to obtain the prediction result of the sample image.
8. An object detection method, comprising: Obtain an image to be detected; Input the image to be detected into a trained object detection model to obtain the detection result of the image to be detected, where the trained object detection model is trained by using the method according to any one of claims 6 or 7.
9. A feature map generation device, the device comprising: A partitioning unit, configured to partition candidate regions of a sample image; A first determination unit, configured to randomly determine anchor points in sub-regions obtained by partitioning the candidate regions; A second determination unit, configured to determine the feature value of an anchor point according to the position information of the anchor point in the sub-region and the mapping position of the candidate region in the first feature map, where the first feature map is the feature map of the sample image; A generation unit, configured to use the feature value of the anchor point as the feature value of the unit where the anchor point is located, and average the feature values of multiple units in the sub-region to obtain the feature value of the sub-region; generate the second feature map of the candidate region according to the feature value of the sub-region in the candidate region, where the second feature map is used to predict the category of an object in the candidate region.
10. The apparatus according to claim 9, wherein, The first determination unit is further configured to: Divide the sub-region into multiple units; Randomly determine anchor points in each of the multiple units.
11. The device according to claim 9 or 10, wherein, The second determination unit is further configured to: Determine a target grid corresponding to the anchor point from multiple grids included in the first feature map according to the position information of the anchor point in the sub-region and the mapping position of the candidate region in the first feature map; Determine the feature value of the anchor point by using a bilinear interpolation algorithm according to the feature information of the target grid.
12. The device according to any one of claims 9-11, before the partitioning unit is configured to partition candidate regions of a sample image, the device further comprises: An expansion unit, configured to randomly expand the size of the candidate region.
13. The apparatus according to claim 12, wherein, The expansion unit is further configured to: Randomly generate an expansion coefficient, where the expansion coefficient is greater than 1; Determine the target length and target width of the expanded candidate region according to the expansion coefficient, the length of the candidate region, and the width of the candidate region; Expand the candidate region according to the target length and the target width while keeping the center position of the candidate region unchanged.
14. A training device for an object detection model, the device comprising: A sample acquisition unit, configured to acquire a sample image and the annotation information of the sample image; A result prediction unit, configured to input the sample image into an object detection model to obtain the prediction result of the sample image; and in the process of obtaining the prediction result of the sample image, the object detection model is configured to obtain the second feature map of candidate regions in the sample image according to the method according to any one of claims 1-5, and the second feature map is used to obtain the prediction result of the sample image; An adjustment unit, configured to adjust parameters of the object detection model according to a prediction result of the sample image and annotation information of the sample image, so as to obtain a trained object detection model.
15. The apparatus according to claim 14, wherein The object detection model includes a first intermediate layer, a second intermediate layer, and a third intermediate layer; The result prediction unit is further configured to: Input the sample image into the first intermediate layer to obtain candidate regions of the sample image and a first feature map of the sample image; Input the candidate regions and the first feature map into the second intermediate layer to obtain a second feature map of the candidate regions; Input the second feature map into the third intermediate layer to obtain a prediction result of the sample image.
16. An object detection device, comprising: An image acquisition unit, configured to acquire an image to be detected; A detection unit, configured to input the image to be detected into the trained object detection model to obtain a detection result of the image to be detected, where the trained object detection model is trained by using the method according to claim 6 or 7.
17. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.
19. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Deep learning remote sensing image target detection method and device based on candidate box selection
CN110956157A
Damaged corn particle detection and classification method based on deep learning
CN111310756A