Training methods for object detection models and object detection methods
By constructing a second object detection model based on a shallow feature extraction submodule, the problems of slow detection speed and easy missed detection of small objects are solved, and fast and accurate detection of small objects is achieved.
Patent Information
- Application Number
- CN202111648973.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-26
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2040-11-26
AI Technical Summary
Small objects are prone to being missed during detection in images or videos, and the detection speed is relatively slow.
By training a first object detection model and constructing a second object detection model based on a shallow feature extraction submodule, the model parameters are adjusted using the feature map information of the first and second regions of interest, thereby achieving fast and accurate detection of small-sized objects.
A precise small-sized object detection model can be trained in a shorter time, improving detection efficiency and accuracy.
Smart Images

Figure CN114549819B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular to training methods for object detection models and object detection methods. Background Technology
[0002] Object detection is an important branch of computer vision. It is defined as the ability of a computer to determine whether a given video or image contains a target object, and if so, to determine the specific location of the target object in the video or image.
[0003] Small objects refer to objects in a video or image whose size is relatively small compared to the overall size of the video or image.
[0004] Because small objects are relatively small in size compared to the size of videos or images, they are easy to miss during detection and the detection speed is slow. Summary of the Invention
[0005] The technical problem to be solved by this application is to provide a training method for an object detection model and an object detection method, so that the trained second object detection model can quickly and accurately detect small objects.
[0006] To address the aforementioned technical problems, this application provides a training method for an object detection model. The method includes: acquiring an image training set, the image training set including training images with annotation information; processing the training images using a first object detection model to obtain first region of interest (ROI) feature map information corresponding to the training images, the first object detection model including a first feature extraction network, the first feature extraction network including at least one feature extraction module; processing the training images using a second object detection model to obtain second ROI feature map information corresponding to the training images, and obtaining a second object detection result based on the second ROI feature map information, the second object detection model including a second feature extraction network, the second feature extraction network being constructed based on a shallow feature extraction submodule in the feature extraction module; and adjusting the model parameters in the second object detection model based on the first ROI feature map information, the second ROI feature map information, the annotation information of the training images, and the second object detection result.
[0007] On the other hand, this application provides an object detection method, the method comprising: acquiring an image to be detected; detecting the image to be detected using a second object detection model trained to convergence, and obtaining a third object detection result corresponding to the image to be detected, the third object detection result including the type and location of the object to be detected in the image to be detected; wherein the second object detection model is trained using any of the object detection model training methods described above.
[0008] On the other hand, this application provides a training apparatus for an object detection model, the apparatus comprising: an image training set acquisition module for acquiring an image training set, the image training set including training images with annotation information; a first image processing module for processing the training images using a first object detection model to obtain first region of interest feature map information corresponding to the training images, the first object detection model including a first feature extraction network, the first feature extraction network including at least one feature extraction module; a second image processing module for processing the training images using a second object detection model to obtain second region of interest feature map information corresponding to the training images, and obtaining a second object detection result based on the second region of interest feature map information, the second object detection model including a second feature extraction network, the second feature extraction network being constructed based on a shallow feature extraction submodule in the feature extraction module; and a model parameter adjustment module for adjusting the model parameters in the second object detection model based on the first region of interest feature map information, the second region of interest feature map information, the annotation information of the training images, and the second object detection result.
[0009] On the other hand, this application provides an object detection apparatus, the apparatus comprising: a target image acquisition module for acquiring a target image; and a target image detection module for detecting the target image using a second object detection model trained to convergence, thereby obtaining a third object detection result corresponding to the target image, the third object detection result including the type and location of the target object in the target image; wherein the second object detection model is trained using any of the object detection model training methods described above.
[0010] On the other hand, this application provides an electronic device including a central processing unit configured to perform the method described above.
[0011] On the other hand, this application provides a computer storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded by a processor and executed as described above.
[0012] Implementing the embodiments of this application has the following beneficial effects:
[0013] In this embodiment, a first object detection model is pre-trained using an image training set. The first object detection model includes a first feature extraction network, which includes at least one feature extraction module. A second object detection model is constructed based on a shallow feature extraction sub-module within the at least one feature extraction module. The model parameters in the second object detection model are adjusted based on the first region of interest feature map information, the second region of interest feature map information, the annotation information of the training image, and the second object detection result. This allows for the rapid and accurate detection of small objects using a second object detection model. Furthermore, the trained second object detection model can be used to quickly and accurately determine the object type and location of small target objects in videos or images. Since the second object detection model is directly constructed from the shallow feature extraction sub-module of the trained first object detection model, a relatively accurate second object detection model can be trained in a shorter time, improving training efficiency. Attached Figure Description
[0014] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a schematic diagram of the hardware environment provided in the embodiments of this application;
[0016] Figure 2 This is a flowchart of a training method for an object detection model provided in an embodiment of this application;
[0017] Figure 3 This is a schematic diagram illustrating the construction of the second object detection model in a training method for an object detection model provided in an embodiment of this application;
[0018] Figure 4 This is a flowchart of an object detection method provided in an embodiment of this application;
[0019] Figure 5 This is a schematic diagram of the structure of a training device for an object detection model provided in an embodiment of this application;
[0020] Figure 6 This is a schematic diagram of the structure of an object detection device provided in an embodiment of this application;
[0021] Figure 7 This is a schematic diagram of the structure of a training device for an object detection model and an object detection device provided in an embodiment of this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0024] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0025] To facilitate understanding, the following brief explanations are provided for some of the terms:
[0026] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0027] Computer vision (CV) is the science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, and simultaneous localization and mapping (SLAM).
[0028] This invention provides a method for training an object detection model and an object detection method. Optionally, in this invention, the above-mentioned method for training an object detection model and the object detection method can be applied to, for example... Figure 1 The hardware environment comprised of server 102 and terminal 104 is shown. Figure 1 As shown, server 102 is connected to terminal 104 via a network, which includes, but is not limited to, wide area networks (WANs), metropolitan area networks (MANs), or local area networks (LANs). Terminal 104 is not limited to PCs, mobile phones, tablets, etc. The training method of the object detection model in this embodiment of the invention can be executed by server 102 or by terminal 104. The object detection method in this embodiment of the invention can be executed jointly by server 102 and terminal 104. Specifically, terminal 104 can acquire images, and when image processing is required, terminal 104 can send the images to server 102.
[0029] The terminal 104 can install various applications needed by users, such as instant messaging applications, media applications, and browser applications. The terminal 104 can send images to the server 102 for processing based on the installed applications.
[0030] The following combination Figure 2 A training method for an object detection model provided in an embodiment of the present invention will be described. For example... Figure 2 As shown, the method includes:
[0031] Step S201: Obtain an image training set, which includes training images with annotation information.
[0032] In this embodiment of the invention, the image training set may include at least one training image, and each training image may have annotation information characterizing the object type and object location to which the training image belongs. The source of the training images may include, but is not limited to, video frame images, photographs, web page images, live stream images, application images, or images from online resources. In this embodiment of the invention, the annotation information may be used to characterize the type and location of the objects included in the training images.
[0033] It is understood that, in this application, when the source of the training images is used in a specific product or technology, user permission or consent is required, and the collection, acquisition, use and processing of the training images must comply with the relevant laws, regulations and national standards of the country in which they are used.
[0034] Step S203: The training image is processed by a first object detection model to obtain a first region of interest feature map information corresponding to the training image. The first object detection model includes a first feature extraction network, and the first feature extraction network includes at least one feature extraction module.
[0035] In this embodiment of the invention, the first region of interest feature map information refers to the information obtained by projecting the first target candidate region corresponding to the training image onto the corresponding first feature map information. The first feature map information may refer to the features of the image generated by convolving the training image and the convolution kernel in the first object detection model. The first target candidate region may refer to the first target candidate region obtained by classifying the anchors in the first feature map information through the normalization function in the first object detection model to obtain positive classification information and negative classification information, determining the positive classification information as the candidate region, calculating the bounding box regression offset of the anchor, and adjusting the candidate region according to the bounding box regression offset.
[0036] The first object detection model is a model that has been trained to convergence using the image training set in advance, and the first object detection model includes knowledge of object detection and recognition in the image training set.
[0037] The first feature extraction network is used to extract the first feature map information in the training image. The first feature extraction network may include at least one feature extraction module. For example, the first feature extraction network may include four feature extraction modules, each of which is used to extract feature map information of the training image in different dimensions.
[0038] Step S205: The training image is processed by the second object detection model to obtain the second region of interest feature map information corresponding to the training image, and the second object detection result is obtained based on the second region of interest feature map information. The second object detection model includes a second feature extraction network, which is constructed based on the shallow feature extraction submodule in the feature extraction module.
[0039] In this embodiment of the invention, the second region of interest feature map information refers to the information obtained by projecting the second target candidate region corresponding to the training image onto the corresponding second feature map information. The second feature map information may refer to the features of the image generated after convolving the training image and the convolution kernel in the second object detection model. The second target candidate region may be obtained by classifying the anchors in the second feature map information through the normalization function in the second object detection model to obtain positive classification information and negative classification information, determining the positive classification information as the candidate region, calculating the bounding box regression offset of the anchor, and adjusting the candidate region according to the bounding box regression offset to obtain the second target candidate region.
[0040] Each of the feature extraction modules may include at least one feature extraction sub-module. When constructing the second object detection model, at least one relatively shallow feature sub-module in the feature extraction module may be used as a shallow feature extraction sub-module, and the second feature extraction network of the second object detection model may be constructed using the shallow feature extraction sub-module.
[0041] Specifically, the model structure of the second feature extraction network in the second object detection model is formed by sequentially connecting the shallow feature extraction sub-modules in each feature extraction module, and the second object detection model directly loads the weight parameters corresponding to the shallow feature extraction sub-modules as the parameters of the second feature extraction network.
[0042] Step S207: Based on the first region of interest feature map information, the second region of interest feature map information, the annotation information of the training image, and the second object detection result, adjust the model parameters in the second object detection model.
[0043] In this embodiment of the invention, the adjustment of the model parameters can be stopped when the second object detection model meets preset conditions, so as to obtain a second object detection model that can be used for small-sized object type and location detection. The preset conditions may specifically be that the loss function converges, the number of iterations reaches a preset number, and the total error value (for example, the sum of the error between the first region of interest feature map information and the second region of interest feature map information and the error between the annotation information of the training image and the second object detection result) is less than a preset value.
[0044] In this embodiment of the invention, a first object detection model is pre-trained using an image training set. The first object detection model includes a first feature extraction network, which includes at least one feature extraction module. A second object detection model is constructed based on a shallow feature extraction sub-module within the at least one feature extraction module. The model parameters in the second object detection model are adjusted based on the first region of interest feature map information, the second region of interest feature map information, the annotation information of the training images, and the second object detection results. With sufficient data, the first object detection model exhibits high detection accuracy and strong generalization performance for complex scenes or complex targets. By transferring the knowledge learned from the first object detection model to the structurally simple second object detection model, a second object detection model capable of quickly and accurately detecting small objects can be obtained. Furthermore, the trained second object detection model can be used to quickly and accurately determine the type and location of small objects in videos or images. Since the second object detection model is directly constructed based on the shallow feature extraction sub-module of the trained first object detection model, a second object detection model that meets the requirements for small object detection can be trained in a shorter time, improving training efficiency.
[0045] In some embodiments, the first object detection model further includes a first region generation network. Correspondingly, processing the training image using the first object detection model to obtain the first region of interest feature map information corresponding to the training image may include:
[0046] Freeze the model parameters in the first object detection model;
[0047] The training image is input into the first object detection model after the parameters are frozen, and the first feature extraction network in the first object detection model is used to extract the first feature map information of the training image.
[0048] The first feature map information is input into the first region generation network, and the first region generation network is used to process the first feature map information to obtain a first target candidate region corresponding to the training image.
[0049] The first target candidate region is projected onto the first feature map information to obtain the first region of interest feature map information.
[0050] In this embodiment of the invention, the first region generation network is used to generate a first target candidate region. Specifically, it obtains positive classification information and negative classification information by classifying the anchors in the first feature map information through a normalization function, determines the positive classification information as a candidate region, calculates the bounding box regression offset of the anchor, and adjusts the candidate region according to the bounding box regression offset to obtain the first target candidate region.
[0051] In this embodiment of the invention, freezing the model parameters in the first object detection model ensures that the model parameters remain unchanged during knowledge transfer, preventing the knowledge transferred by the first object detection model from being gradually forgotten by the second object detection model, thereby guaranteeing the accuracy of the trained second object detection model. Furthermore, this application directly uses the first region of interest feature map information as a supervision signal to train the second object detection model, instead of using a region feature map processed into a fixed size as a supervision signal, thereby reducing noise interference during training and improving the accuracy of the trained second object detection model.
[0052] In some embodiments, the second object detection model further includes a second region generation network, a second interest pooling layer, and a second classifier. Correspondingly, the process of processing the training image using the second object detection model to obtain second region of interest feature map information corresponding to the training image, and the second object detection result obtained based on the second region of interest feature map information, may include:
[0053] The training image is input into the second object detection model, and the second feature extraction network in the second object detection model is used to extract the second feature map information of the training image;
[0054] The second feature map information is input into the second region generation network, and the second region generation network is used to process the second feature map information to obtain a second target candidate region corresponding to the training image.
[0055] The second target candidate region is projected onto the second feature map information to obtain the second region of interest feature map information;
[0056] The feature map information of the second region of interest is input into the second interest pooling layer, and the output information of the second interest pooling layer is input into the second classifier to obtain the second object detection result.
[0057] In this embodiment of the invention, the second region generation network is used to generate a second target candidate region. Specifically, it obtains positive classification information and negative classification information by classifying the anchors in the second feature map information through a normalization function, determines the positive classification information as a candidate region, calculates the bounding box regression offset of the anchor, and adjusts the candidate region according to the bounding box regression offset to obtain the second target candidate region.
[0058] The second interest pooling layer is used to calculate a fixed-size regional feature map based on the second feature map information.
[0059] The second classifier may include a fully connected layer and a normalization layer. The second classifier combines the fixed-size region feature map information output by the second interest pooling layer through its fully connected layer and normalization layer to calculate the corresponding object classification result of the region feature map. At the same time, it can fine-tune the second target candidate region according to the object classification result and determine the fine-tuned second target candidate region as the object detection region.
[0060] In this embodiment of the invention, the second object detection model is trained by directly using the feature map information of the first region of interest as a supervision signal, instead of using a region feature map processed into a fixed size as a supervision signal. This reduces the interference of noise during the training process and improves the accuracy of the trained second object detection model.
[0061] In some embodiments, adjusting the model parameters in the second object detection model based on the first region of interest feature map information, the second region of interest feature map information, the annotation information of the training image, and the second object detection result may include:
[0062] Based on the feature map information of the first region of interest and the feature map information of the second region of interest, the first loss function of the second object detection model is determined;
[0063] Based on the annotation information of the training images and the second object detection results, a second loss function of the second object detection model is determined;
[0064] The target loss function is obtained by weighted summation of the first loss function and the second loss function.
[0065] The model parameters in the second object detection model are adjusted according to the target loss function.
[0066] Specifically, the target loss function It consists of the first loss function and the second loss function, as shown in the formula below:
[0067]
[0068] in, Denotes the first loss function. Let λ represent the second loss function, and λ1 represent the weights of the first loss function.
[0069] like Figure 3 As shown, the first loss function represents the distance between the first region of interest feature map information output by the first object detection model and the second region of interest feature map information output by the second object detection model. The calculation method is as follows:
[0070]
[0071] Where N represents the number of region of interest feature maps output by the first object detection model and the second object detection model, m represents the dimension of the region of interest feature map, μ represents the first region of interest feature map output by the first object detection model, ν represents the second region of interest feature map output by the second object detection model, and r is the dimension transformation function.
[0072] The second loss function represents the difference between the output information of the second object detection model and the true labeled information. This function consists of two parts:
[0073]
[0074] in, This represents the classification loss function between the predicted bounding box output by the second object detection model and the ground truth bounding box. λ represents the regression loss function between the predicted bounding box output by the second object detection model and the ground truth bounding box, and λ2 is the weight of the regression loss function.
[0075] λ1 and λ2 can be determined empirically. For example, in this embodiment of the invention, λ1 can be set to 0.9, which can ensure that the trained second detection model has good detection capabilities.
[0076] In some embodiments, such as Figure 3As shown, each feature extraction module includes at least one convolutional submodule. Accordingly, before the step of processing the training image using the second object detection model to obtain second region of interest feature map information corresponding to the training image, and obtaining the second object detection result based on the second region of interest feature map information, the method may further include:
[0077] Extract the first preset number of convolutional sub-modules from each of the feature extraction modules;
[0078] The extracted convolutional sub-modules are connected in the order of their corresponding feature extraction modules in the first feature extraction network to obtain the second feature extraction network;
[0079] After the second feature extraction network, a pre-set second region generation network, a second interest pooling layer, and a second classifier are added in sequence to obtain the second object detection model.
[0080] In this embodiment of the invention, the preset number of layers can be the first 2 layers or the first 3 layers. For example, the current preset number of layers is the first 2 layers, that is, the first 2 convolutional sub-modules of each feature extraction module are extracted.
[0081] The convolutional submodule of the pre-set layer in the feature extraction module extracts shallow features from the feature extraction module, which is beneficial for the rapid fitting of the second object detection model.
[0082] The model structure of the second feature extraction network is formed by sequentially connecting each of the convolutional sub-modules, and the second object detection model directly loads the weight parameters corresponding to the convolutional modules as the parameters of the second feature extraction network.
[0083] Optionally, the feature extraction module can be a residual structure, and the convolutional submodule can be a bottleneck block. The purpose of using a bottleneck block is to reduce the number of channels through a low-cost 1X1 convolution, so that the subsequent 3X3 convolution has fewer parameters. Finally, a 1X1 convolution is used to broaden the network.
[0084] As shown in Table 1, the feature extraction module conv2x includes 3 bottleneck blocks, and the first 2 bottleneck blocks are extracted from the 3 bottleneck blocks; the feature extraction module conv3x includes 4 bottleneck blocks, and the first 2 bottleneck blocks are extracted from the 4 bottleneck blocks; the feature extraction module conv4x includes 23 bottleneck blocks, and the first 2 bottleneck blocks are extracted from the 23 bottleneck blocks; the feature extraction module conv5x includes 3 bottleneck blocks, and the first 2 bottleneck blocks are extracted from the 3 bottleneck blocks. All the extracted bottleneck blocks are assembled according to the order of the bottleneck blocks in the first feature extraction network to obtain the second feature extraction network.
[0085] Table 1
[0086]
[0087]
[0088] Optionally, other network forms can be used for the first feature extraction network, such as the ResNet series, ResNeXt series, ResNeSt series, and Efficient series networks.
[0089] In this embodiment of the invention, by extracting the convolutional sub-modules of the pre-preset layer in the feature extraction module, and then sequentially assembling each convolutional sub-module into the second feature extraction network of the second object detection model, and directly loading the weight parameters corresponding to the convolutional sub-modules as the parameters of the second feature extraction network, it is possible to ensure that the trained second detection model has good small-sized object detection capability and facilitates the rapid fitting of the second object detection model.
[0090] In some embodiments, prior to the step of processing the training image using a first object detection model to obtain feature map information of a first region of interest, the method may further include:
[0091] Obtain the first initial object detection model;
[0092] The training images in the training image set are input into the first initial object detection model to obtain the predicted first object detection result corresponding to the training images;
[0093] Based on the annotation information of the training images and the first object detection result, the third loss function of the first initial object detection model is determined;
[0094] The first initial object detection model is trained to convergence according to the third loss function to obtain the first object detection model.
[0095] Optionally, ResNet101-FPN can be used as the first initial object detection model, which has initialization parameters.
[0096] The image training set includes training images with annotation information, which is used to characterize the type and location of objects included in the training images.
[0097] The first object detection result includes the predicted object type and location, and the third loss function can be determined by the mean square error between the annotation information of the training image (annotated object type and location) and the first object detection result (predicted object type and location).
[0098] During training, when the third loss function converges, the adjustment of the model parameters of the first initial object detection model is stopped, and the final first initial object detection model is used as the first object detection model.
[0099] This invention also provides an object detection method, such as... Figure 4 As shown, the method includes:
[0100] Step S401: Acquire the image to be detected;
[0101] Step S403: Use the second object detection model trained to convergence to detect the image to be detected, and obtain a third object detection result corresponding to the image to be detected. The third object detection result includes the type and location of the object to be detected in the image to be detected.
[0102] The second object detection model is trained using any of the object detection model training methods described above.
[0103] In one embodiment, the second object detection model is obtained by knowledge transfer from the first object detection model, which is obtained by the above-described object detection model training method.
[0104] The second feature extraction network of the second object detection model can be obtained by sequentially connecting shallow feature extraction sub-modules. The shallow feature extraction sub-module is a feature sub-module of the first feature extraction network in the first object detection model, and the second object detection model loads the weight parameters corresponding to the shallow feature extraction sub-module as the parameters of the second feature extraction network.
[0105] In one specific embodiment, the first object detection model and the second object detection model trained using the first object detection model are evaluated on a V100 GPU using an internal dataset. Specifically, the internal test set includes 7228 video images, and the evaluation metrics include mAP, AP@0.5, and inference time per image.
[0106] The test results are shown in Table 2.
[0107]
[0108] In this embodiment of the invention, the trained second object detection model can achieve a detection accuracy similar to that of the first object detection model, and the detection speed of the trained second object detection model is 2.5 times that of the first object detection model, thus meeting the requirements of real-time detection.
[0109] This invention also provides a training apparatus for an object detection model; please refer to [link to relevant documentation]. Figure 5 The device may include:
[0110] The image training set acquisition module 610 is used to acquire an image training set, which includes training images with annotation information.
[0111] The first image processing module 620 is used to process the training image through a first object detection model to obtain first region of interest feature map information corresponding to the training image. The first object detection model includes a first feature extraction network, and the first feature extraction network includes at least one feature extraction module.
[0112] The second image processing module 630 is used to process the training image through a second object detection model to obtain second region of interest feature map information corresponding to the training image, and to obtain a second object detection result based on the second region of interest feature map information. The second object detection model includes a second feature extraction network, which is constructed based on the shallow feature extraction submodule in the feature extraction module.
[0113] The model parameter adjustment module 640 is used to adjust the model parameters in the second object detection model based on the first region of interest feature map information, the second region of interest feature map information, the annotation information of the training image, and the second object detection result.
[0114] In some embodiments, the first object detection model further includes a first region generation network, and correspondingly, the first image processing module may include:
[0115] The parameter freezing submodule is used to freeze the model parameters in the first object detection model.
[0116] The first feature map information acquisition submodule is used to input the training image into the first object detection model after freezing parameters, and use the first feature extraction network in the first object detection model to extract the first feature map information of the training image.
[0117] The first target candidate region determination submodule is used to input the first feature map information into the first region generation network, and use the first region generation network to process the first feature map information to obtain the first target candidate region corresponding to the training image.
[0118] The first region of interest feature map information determination submodule is used to project the first target candidate region onto the first feature map information to obtain the first region of interest feature map information.
[0119] In some embodiments, the second object detection model further includes a second region generation network, a second interest pooling layer, and a second classifier; correspondingly, the second image processing module may include:
[0120] The second feature map information acquisition submodule is used to input the training image into the second object detection model and use the second feature extraction network in the second object detection model to extract the second feature map information of the training image.
[0121] The second target candidate region determination submodule is used to input the second feature map information into the second region generation network, and use the second region generation network to process the second feature map information to obtain the second target candidate region corresponding to the training image.
[0122] The second region of interest feature map information determination submodule is used to project the second target candidate region onto the second feature map information to obtain the second region of interest feature map information;
[0123] The second object detection result determination submodule is used to input the feature map information of the second region of interest into the second interest pooling layer, and input the output information of the second interest pooling layer into the second classifier to obtain the second object detection result.
[0124] In some embodiments, the model parameter adjustment module may include:
[0125] The first loss function determination submodule is used to determine the first loss function of the second object detection model based on the first region of interest feature map information and the second region of interest feature map information;
[0126] The second loss function determination submodule is used to determine the second loss function of the second object detection model based on the annotation information of the training image and the second object detection result;
[0127] The target loss function determination submodule is used to perform a weighted summation of the first loss function and the second loss function to obtain the target loss function;
[0128] The model parameter adjustment submodule is used to adjust the model parameters in the second object detection model according to the target loss function.
[0129] In some embodiments, the feature extraction module includes at least one convolutional submodule, and correspondingly, the apparatus may further include:
[0130] A convolutional submodule extraction module is used to extract the convolutional submodules of the first preset number of layers from each of the feature extraction modules;
[0131] The second feature extraction network generation module is used to connect the extracted convolutional sub-modules according to the order of the corresponding feature extraction modules in the first feature extraction network to obtain the second feature extraction network.
[0132] The second object detection model generation module is used to sequentially add a pre-set second region generation network, a second interest pooling layer, and a second classifier after the second feature extraction network to obtain the second object detection model.
[0133] In some embodiments, the apparatus may further include:
[0134] The first initial object detection model acquisition module is used to acquire the first initial object detection model;
[0135] The first object detection result determination module is used to input the training image in the training image set into the first initial object detection model to obtain the predicted first object detection result corresponding to the training image;
[0136] The third loss function determination module is used to determine the third loss function of the first initial object detection model based on the annotation information of the training image and the first object detection result.
[0137] The first object detection model generation module is used to train the first initial object detection model to convergence according to the third loss function, so as to obtain the first object detection model.
[0138] This invention also provides an object detection device; please refer to [link to relevant documentation]. Figure 6 The device may include:
[0139] The image acquisition module 710 is used to acquire the image to be detected.
[0140] The image detection module 720 is used to detect the image to be detected using a second object detection model trained to convergence, and to obtain a third object detection result corresponding to the image to be detected. The third object detection result includes the type and location of the object to be detected in the image to be detected.
[0141] The second object detection model is trained using any of the object detection model training methods described above.
[0142] The apparatus provided in the above embodiments can execute the methods provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in the above embodiments can be found in the methods provided in any embodiment of this application.
[0143] This embodiment also provides a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded by a processor and executed as any of the methods described in this embodiment.
[0144] This embodiment also provides a device, the structural diagram of which can be found in the following figure. Figure 7 The device 800 can vary considerably depending on its configuration or performance, and may include one or more central processing units (CPUs) 822 (e.g., one or more processors) and memory 832, and one or more storage media 830 (e.g., one or more mass storage devices) storing application programs 842 or data 844. The memory 832 and storage media 830 can be temporary or persistent storage. The program stored in the storage media 830 may include one or more modules (not shown), each module including a series of instruction operations on the device. Furthermore, the CPU 822 may be configured to communicate with the storage media 830 and execute the series of instruction operations in the storage media 830 on the device 800. The device 800 may also include one or more power supplies 826, one or more wired or wireless network interfaces 850, one or more input / output interfaces 858, and / or one or more operating systems 841, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc. Any of the methods described above in this embodiment can be based on Figure 7 The equipment shown is used for implementation.
[0145] This specification provides the operational steps of the methods described in the embodiments or flowcharts, but more or fewer operational steps may be included based on conventional or non-inventive labor. The steps and order listed in the embodiments are merely one possible execution order among many steps and do not represent the only execution order. In actual system or interrupt product execution, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment).
[0146] The structure shown in this embodiment is only a partial structure related to the solution of this application and does not constitute a limitation on the device to which the solution of this application is applied. Specific devices may include more or fewer components than shown, or combinations of certain components, or arrangements of different components. It should be understood that the methods, apparatuses, etc., disclosed in this embodiment can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or unit modules through some interfaces.
[0147] Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0148] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this specification can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0149] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A training method for an object detection model, characterized in that, The method includes: Obtain an image training set, wherein the image training set includes training images with annotation information; The training image is processed by a first object detection model to obtain a first region of interest feature map information corresponding to the training image. The first object detection model includes a first feature extraction network, and the first feature extraction network includes at least one feature extraction module. The training image is processed by a second object detection model to obtain a second region of interest feature map information corresponding to the training image, and a second object detection result is obtained based on the second region of interest feature map information. The second object detection model includes a second feature extraction network, which is constructed based on the shallow feature extraction submodule in the feature extraction module. Based on the feature map information of the first region of interest, the feature map information of the second region of interest, the annotation information of the training image, and the second object detection result, the model parameters in the second object detection model are adjusted.
2. The training method according to claim 1, characterized in that, The first object detection model further includes a first region generation network. Correspondingly, processing the training image using the first object detection model to obtain the first region of interest feature map information corresponding to the training image includes: Freeze the model parameters in the first object detection model; The training image is input into the first object detection model after the parameters are frozen, and the first feature extraction network in the first object detection model is used to extract the first feature map information of the training image. The first feature map information is input into the first region generation network, and the first region generation network is used to process the first feature map information to obtain a first target candidate region corresponding to the training image. The first target candidate region is projected onto the first feature map information to obtain the first region of interest feature map information.
3. The training method according to claim 1, characterized in that, The second object detection model further includes a second region generation network. Accordingly, processing the training image using the second object detection model to obtain the second region of interest feature map information corresponding to the training image includes: The training image is input into the second object detection model, and the second feature extraction network in the second object detection model is used to extract the second feature map information of the training image; The second feature map information is input into the second region generation network, and the second region generation network is used to process the second feature map information to obtain a second target candidate region corresponding to the training image. The second target candidate region is projected onto the second feature map information to obtain the second region of interest feature map information.
4. The training method according to claim 1 or 3, characterized in that, The second object detection model further includes a second interest pooling layer and a second classifier. Obtaining the second object detection result based on the feature map information of the second region of interest includes: The feature map information of the second region of interest is input into the second interest pooling layer, and the output information of the second interest pooling layer is input into the second classifier to obtain the second object detection result.
5. The training method according to claim 1, characterized in that, The step of adjusting the model parameters in the second object detection model based on the first region of interest feature map information, the second region of interest feature map information, the annotation information of the training image, and the second object detection result includes: Based on the feature map information of the first region of interest and the feature map information of the second region of interest, the first loss function of the second object detection model is determined; Based on the annotation information of the training images and the second object detection results, a second loss function of the second object detection model is determined; The target loss function is obtained by weighted summation of the first loss function and the second loss function. The model parameters in the second object detection model are adjusted according to the target loss function.
6. The training method according to claim 1, characterized in that, The feature extraction module includes at least one convolutional submodule. Correspondingly, before processing the training image using the second object detection model to obtain the second region of interest feature map information corresponding to the training image, the method further includes: Extract the first preset number of convolutional sub-modules from each of the feature extraction modules; The extracted convolutional sub-modules are connected in the order of their corresponding feature extraction modules in the first feature extraction network to obtain the second feature extraction network; After the second feature extraction network, a pre-set second region generation network, a second interest pooling layer, and a second classifier are added in sequence to obtain the second object detection model.
7. The training method according to claim 1, characterized in that, Before the step of processing the training image using a first object detection model to obtain the feature map information of the first region of interest, the method further includes: Obtain the first initial object detection model; The training images in the training image set are input into the first initial object detection model to obtain the first object detection result corresponding to the training images; Based on the annotation information of the training images and the first object detection result, the third loss function of the first initial object detection model is determined; The first initial object detection model is trained to convergence according to the third loss function to obtain the first object detection model.
8. An object detection method, characterized in that, The method includes: Acquire the image to be detected; The image to be detected is detected using a second object detection model trained to convergence, and a third object detection result corresponding to the image to be detected is obtained. The third object detection result includes the type and location of the object to be detected in the image to be detected. The second object detection model is trained using the training method of the object detection model according to any one of claims 1-3 and 5-7.
9. The object detection method according to claim 8, characterized in that, The second object detection model is obtained by knowledge transfer from the first object detection model, which is obtained by the training method of the object detection model according to any one of claims 1, 2 or 7.
10. The object detection method according to claim 9, characterized in that, The second feature extraction network of the second object detection model is obtained by sequentially connecting shallow feature extraction sub-modules. The shallow feature extraction sub-module is the feature sub-module of the first feature extraction network in the first object detection model, and the second object detection model loads the weight parameters corresponding to the shallow feature extraction sub-module as the parameters of the second feature extraction network.
11. A training device for an object detection model, characterized in that, The device includes: An image training set acquisition module is used to acquire an image training set, which includes training images with annotation information. A first image processing module is used to process the training image through a first object detection model to obtain a first region of interest feature map information corresponding to the training image. The first object detection model includes a first feature extraction network, and the first feature extraction network includes at least one feature extraction module. The second image processing module is used to process the training image through the second object detection model to obtain the second region of interest feature map information corresponding to the training image, and to obtain the second object detection result based on the second region of interest feature map information. The second object detection model includes a second feature extraction network, which is constructed based on the shallow feature extraction submodule in the feature extraction module. The model parameter adjustment module is used to adjust the model parameters in the second object detection model based on the first region of interest feature map information, the second region of interest feature map information, the annotation information of the training image, and the second object detection result.
12. An object detection device, characterized in that, The device includes: The image acquisition module is used to acquire the image to be detected. The image detection module is used to detect the image to be detected using a second object detection model trained to convergence, and to obtain a third object detection result corresponding to the image to be detected. The third object detection result includes the type and location of the object to be detected in the image to be detected. The second object detection model is trained using the training method of the object detection model according to any one of claims 1-3 and 5-7.
13. An electronic device, characterized in that, include: A central processing unit configured to perform a training method for an object detection model as claimed in any one of claims 1 to 7 or to perform an object detection method as claimed in any one of claims 8 to 10.
14. A computer storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded by a processor and executed by the training method of the object detection model as described in any one of claims 1 to 7 or by executing the object detection method as described in any one of claims 8 to 10.