Training method, object detection method and device for multi-scale object detection model

Through joint training of multi-scale object detection model, the appropriate branches are automatically selected, which solves the detection problem of large range of object size changes in the prior art, improves detection accuracy and reduces the calculation amount.

CN112862002BActive Publication Date: 2025-06-27刘丽
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110291177.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-18
Publication Date
2025-06-27
Estimated Expiration
2041-03-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the problem of large range of object size variations in object detection in target detection, and the method of manually setting thresholds is not accurate enough, which increases the amount of calculation.

Method used

Using the training method of multi-scale object detection model, through joint training of multi-scale identification network and branch selection network, the appropriate branches are automatically selected to improve detection accuracy without adding additional computational amount.

Benefits of technology

Improve the accuracy of object detection, avoid the shortage of manually setting thresholds, and reduce the amount of extra calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112862002B_ABST
    Figure CN112862002B_ABST
Patent Text Reader

Abstract

The present application relates to a training method, an object detection method and a device for a multi-scale object detection model. The training method includes: obtaining first sample data, where each sample image carries object annotation information; respectively inputting the sample images into a multi-scale recognition network and a branch selection network of the multi-scale object detection model to obtain branch selection results and recognition results of each branch; obtaining a model output result according to the branch selection results and the recognition results of each branch; calculating a first loss function according to the model output result and the object annotation information, and performing backpropagation based on the first loss function to adjust the parameters in the multi-scale recognition network and the branch selection network until the first loss function meets a preset value, thereby obtaining the multi-scale object detection model. The object detection method is a method for performing object detection according to the multi-scale object detection model. Using this method improves the accuracy and does not require additional computational effort.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a training method for a multi-scale object detection model, an object detection method, and a device. Background Art

[0002] Object detection, that is, finding all objects of interest in an image, includes two subtasks: object localization and object classification, and at the same time determines the category and location of the object. It is a popular direction in computer vision and digital image processing, and is widely used in many fields such as robot navigation, intelligent video surveillance, industrial inspection, and aerospace. Reducing the consumption of human capital through computer vision has important practical significance.

[0003] In real life, objects of various sizes exist, and even the same object may have large scale variations. The deeper features in a convolutional neural network have a large receptive field and rich semantic information. Deep features are robust to changes such as object pose changes, occlusions, and local deformations, but due to the reduction in resolution, geometric detail information is lost. On the contrary, shallow features have a small receptive field and rich geometric detail information, but the problem is that the resolution is high and the semantic information is scarce. In a convolutional neural network, the semantic information of an object can appear in different layers (related to the size of the object). For small objects, shallow features contain some of their detail information. As the number of layers deepens, the geometric detail information in the extracted features may completely disappear (the receptive field is too large), and it becomes very difficult to detect small objects through deep features. For large objects, their semantic information will appear in deeper features. This will affect the performance of the convolutional neural network to a certain extent.

[0004] In traditional technologies, in order to solve the problem of a large range of detection target size variations, three parallel networks can be used to separately identify targets of different scales.

[0005] However, in the current method of using three parallel networks for identification, the method of manually setting network selection, for example, first obtaining the target size and then comparing it with a threshold to determine which network's output result to use among the three networks. Such a method has insufficient accuracy due to manually setting the threshold, and requires calculating the target size, increasing additional computational complexity. Summary of the Invention

[0006] Based on this, in view of the above technical problems, it is necessary to provide a training method for a multi-scale object detection model, an object detection method, and a device that can improve accuracy without increasing additional computational complexity.

[0007] A training method for a multi-scale object detection model, the training method for the multi-scale object detection model includes:

[0008] Obtain first sample data, where each sample image in the first sample data carries target annotation information;

[0009] Input the sample images into the multi-scale recognition network and the branch selection network of the multi-scale object detection model respectively, to obtain the branch selection result corresponding to the branch selection network and the recognition result of each branch corresponding to the multi-scale recognition network. Each branch network structure of the multi-scale recognition network is the same and parameter-sharing, but the coefficients of the dilated convolution are different;

[0010] Obtain the model output result according to the branch selection result and the recognition result of each branch;

[0011] Calculate a first loss function according to the model output result and the target annotation information, and backpropagate based on the first loss function to adjust the parameters in the multi-scale recognition network and the branch selection network until the first loss function meets a preset value, to obtain the multi-scale object detection model.

[0012] In one embodiment, before obtaining the model output result according to the branch selection result and the recognition result of each branch, it further includes:

[0013] Determine a target branch according to the branch selection result, and use the recognition result of the target branch as the pseudo-label of the corresponding sample image;

[0014] Train the non-target branches according to the pseudo-label of the sample image and the recognition results of the non-target branches.

[0015] In one embodiment, training the non-target branches according to the pseudo-label of the sample image and the recognition results of the non-target branches includes:

[0016] Change the branch selection result of the branch selection network to select a non-target branch;

[0017] Calculate a second loss function according to the pseudo-label of the sample image and the recognition results of the non-target branches, and backpropagate based on the second loss function to adjust the parameters in the multi-scale recognition network until the second loss function meets a preset value, to obtain the multi-scale object detection model.

[0018] In one embodiment, after obtaining the first sample data, where each sample image in the first sample data carries target annotation information, it further includes:

[0019] Extract the image features of the sample images, and generate candidate regions based on the image features.

[0020] In one embodiment, the method further includes:

[0021] Obtain test data, where the test pictures in the test data carry test labels;

[0022] Input the test pictures into the trained multi-scale object detection model, so as to obtain a test selection result through the branch selection network, and obtain a test recognition result for each branch through the multi-scale recognition network;

[0023] Obtain an initial prediction result according to the test selection result and the test recognition result;

[0024] Filter the prediction result to obtain a target prediction result;

[0025] Compare the test label with the target prediction result to obtain a test result.

[0026] An object detection method based on a multi-scale object detection model, the object detection method based on the multi-scale object detection model includes:

[0027] Obtain a picture to be detected;

[0028] Input the picture to be detected into the multi-scale object detection model trained by the training method of the above multi-scale object detection model to obtain an object detection result.

[0029] A training device for a multi-scale object detection model, the training device for the multi-scale object detection model includes:

[0030] A sample acquisition module, configured to acquire first sample data, where each sample picture in the first sample data carries target annotation information;

[0031] A data processing module, configured to input the sample pictures into the multi-scale recognition network and the branch selection network of the multi-scale object detection model respectively, to obtain a branch selection result corresponding to the branch selection network and an identification result for each branch corresponding to the multi-scale recognition network, and each branch network structure of the multi-scale recognition network is the same and parameter-sharing, but the coefficients of the dilated convolution are different;

[0032] A model output result acquisition module, configured to obtain a model output result according to the branch selection result and the identification result of each branch;

[0033] A training module, configured to calculate a first loss function according to the model output result and the target annotation information, and backpropagate based on the first loss function to adjust the parameters in the multi-scale recognition network and the branch selection network until the first loss function meets a preset value, to obtain a multi-scale object detection model.

[0034] An object detection device based on a multi-scale object detection model, the object detection device based on the multi-scale object detection model includes:

[0035] A to-be-detected picture acquisition module, configured to acquire a to-be-detected picture;

[0036] A detection module, configured to input the to-be-detected picture into a multi-scale object detection model obtained by training a training device of the above multi-scale object detection model to obtain an object detection result.

[0037] A computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the above embodiments are implemented.

[0038] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any one of the above embodiments are implemented.

[0039] In the above multi-scale object detection model training method, object detection method and device, during training, the multi-scale recognition network and the branch selection network are jointly trained, that is, the loss function is calculated and backpropagated according to the recognition result of each branch output by the multi-scale recognition network and the branch selection result output by the branch selection network, so as to obtain a multi-scale object detection model. That is, the branch selection result is obtained according to the model and does not need to be manually set, which improves the accuracy. In addition, it is not necessary to calculate the object size, so there is no need to increase additional computational complexity. Description of the Drawings

[0040] Figure 1 It is a schematic flowchart of a training method of a multi-scale object detection model in an embodiment;

[0041] Figure 2 It is a schematic diagram of the training process of a multi-scale object detection model in an embodiment;

[0042] Figure 3 It is a schematic diagram of a multi-scale recognition network in an embodiment;

[0043] Figure 4 It is a schematic flowchart of an object detection method based on a multi-scale object detection model in an embodiment;

[0044] Figure 5 It is a structural block diagram of a training device of a multi-scale object detection model in an embodiment;

[0045] Figure 6It is a structural block diagram of an object detection device based on a multi-scale object detection model in an embodiment;

[0046] Figure 7 It is an internal structure diagram of a computer device in an embodiment. Specific embodiments

[0047] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0048] In one embodiment, as Figure 1 shown, a training method for a multi-scale object detection model is provided. In this embodiment, the method is illustrated by applying it to a terminal. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The so-called training stage means that the training samples contain specific label information. Based on these label information, supervised learning can be formed to generate a loss function, and then the parameters in the model can be adjusted. In this embodiment, the method includes the following steps:

[0049] S102: Obtain first sample data, where each sample picture in the first sample data carries target annotation information.

[0050] Specifically, the first sample data can be pre-annotated data, which includes target annotation information, and the target annotation information refers to the position of the target in each picture.

[0051] S104: Input the sample pictures into the multi-scale recognition network and the branch selection network of the multi-scale object detection model respectively to obtain the branch selection results corresponding to the branch selection network and the recognition results of each branch of the corresponding multi-scale recognition network. Each branch network structure of the multi-scale recognition network is the same and the parameters are shared, but the coefficients of the dilated convolution are different.

[0052] Specifically, the multi-scale object detection model at least includes a multi-scale recognition network and a branch selection network. The multi-scale recognition network is used to extract multi-scale features of the sample pictures, and the branch selection network is used to select one of the branches in the multi-scale recognition network according to the picture features. Specifically, it can be combined with Figure 2 shown, Figure 2 It is a schematic diagram of the training process of a multi-scale object detection model in an embodiment.

[0053] Among them, the multi-scale recognition network may include multiple branches, and the network structures of each branch are the same and the parameters are shared, but the coefficients of the dilated convolution are different. Here, the parameter sharing means that the convolution parameters are the same, and the different dilation coefficients are used to obtain different receptive fields, that is, the image features of different scales. In addition, it should be noted that the branches of the multi-scale recognition network may include at least two. In this embodiment, three branches are taken as an example for illustration. Specifically, please refer to Figure 3 , Figure 3 in which the convolution parameters of the three branches are the same, but the dilation coefficients d are different. In this way, when the sample image is input into the multi-scale recognition network, three different-scale features corresponding to the sample image can be obtained. In other embodiments, when the sample image is input into the multi-scale recognition network, multiple different-scale features corresponding to the number of branches can be obtained.

[0054] The branch selection network can extract the features of the sample image and select the corresponding branch according to the features of the sample image. For example, it can use the gumbel layer network, put the extracted features as input into the gumbel softmax, and the output is an n-dimensional one-hot vector. Through this vector, a certain branch of the multi-scale recognition network is determined to be used. Here, n is the number of branches. Taking the above three branches as an example, the obtained vector may be (1, 0, 0), which indicates that the first branch is selected.

[0055] S106: Obtain the model output result according to the branch selection result and the recognition result of each branch.

[0056] Specifically, the model output result is obtained according to the branch selection result and the recognition result of each branch, that is, by multiplying the recognition result of each branch by the vector in the above branch selection result, so as to filter out the model output results of the unselected branches and only retain the results of the selected branch.

[0057] S108: Calculate the first loss function according to the model output result and the target annotation information, and backpropagate based on the first loss function to adjust the parameters in the multi-scale recognition network and the branch selection network until the first loss function meets the preset value, and obtain the multi-scale object detection model.

[0058] Specifically, the terminal calculates a first loss function from the obtained model output result and the target annotation information, where the first loss function includes the losses of the multi-scale recognition network and the branch selection network. That is, the first loss function can be regarded as the loss function of the multi-scale recognition network multiplied by the loss function of the branch selection network. In this way, during backpropagation, the multi-scale recognition network is backpropagated through the loss function of the multi-scale recognition network to modify the parameters of the corresponding branches, and the branch selection network is backpropagated through the loss function of the branch selection network to modify its parameters. Among them, the modification of the parameters of the branches in the multi-scale recognition network refers to modifying the weights of the convolutions at each position.

[0059] In this way, the terminal performs backpropagation to modify the network parameters until the first loss function meets the preset value, and a multi-scale object detection model is obtained. For example, when the first loss function is minimized, or when the first loss function is less than a certain value, the multi-scale object detection model is obtained.

[0060] In the above training method of the multi-scale object detection model, during training, the multi-scale recognition network and the branch selection network are jointly trained. That is, the loss function is calculated and backpropagated based on the recognition result of each branch output by the multi-scale recognition network and the branch selection result output by the branch selection network, so as to obtain the multi-scale object detection model. That is, the branch selection result is obtained according to the model and does not need to be set manually, which improves the accuracy. In addition, the target size does not need to be calculated, so no additional computational effort is required.

[0061] In one embodiment, before obtaining the model output result according to the branch selection result and the recognition result of each branch, it further includes: determining the target branch according to the branch selection result, and using the recognition result of the target branch as the pseudo-label of the corresponding sample picture; training the non-target branches according to the pseudo-label of the sample picture and the recognition results of the non-target branches.

[0062] Specifically, in this embodiment, in order to increase the sample size and improve the robustness of each branch, self-training can be performed. That is, after the branch selection network selects a branch, the recognition result of this branch is obtained as the pseudo-label of the corresponding sample. For example, the non-maximum suppression operation can be performed on the model recognition result of this branch to obtain a unique result as the pseudo-label, and then this pseudo-label is used as the true value of other branches to train other branches.

[0063] Optionally, the non-target branch is trained according to the pseudo-labels of the sample images and the recognition results of the non-target branch, including: changing the branch selection result of the branch selection network to select the non-target branch; calculating a second loss function according to the pseudo-labels of the sample images and the recognition results of the non-target branch, and backpropagating based on the second loss function to adjust the parameters in the multi-scale recognition network until the second loss function meets a preset value, obtaining a multi-scale object detection model.

[0064] Take Figure 2 as an example. Suppose the output result of the branch selection network is to select the first branch. First, it is trained to obtain a loss function. Then, change the branch selection result of the branch selection network to select the non-target branches, that is, select the second branch and the third branch. In this way, a second loss function is calculated according to the pseudo-labels of the sample images and the recognition results of the non-target branches, so that the parameters in the multi-scale recognition network can be adjusted by backpropagating based on the second loss function until the second loss function meets the preset value, obtaining a multi-scale object detection model. The training of the non-target branches can be carried out separately, for example, in parallel or sequentially. Training each branch separately in this way can enhance the robustness of each branch, and using the pseudo-labels output by the target branch as the ground truth increases the number of samples.

[0065] In one embodiment, after obtaining the first sample data, it further includes: extracting the image features of the sample image and generating candidate regions based on the image features. The candidate regions refer to the partial images where there is a high probability of the presence of the target.

[0066] Specifically, the first sample data can be an image with only one target, or an image including two or more targets. If it is an image with only one target, the terminal can directly extract the foreground for object detection, etc. If it is an image including two or more targets, the image features of the sample image are extracted, and candidate regions are generated based on the image features.

[0067] For the selection of candidate regions, the features of the image can be extracted first through a convolutional network, and then the candidate regions are extracted, for example, extracting RPN according to the method of faster-rcnn.

[0068] Continue to refer to Figure 2As shown in the figure, the multi-scale object detection model may further include a candidate region extraction module, which includes a feature extraction unit and a candidate region extraction unit. Among them, by using a common convolutional network, features are extracted from the input image network. For example, the input image is processed by a convolutional neural network to generate a deep feature map, and then various algorithms are used to complete region generation and loss calculation. This part of the convolutional neural network is the "skeleton" of the entire detection algorithm and is also called the Backbone. Then, the RPN is extracted according to the method of faster-rcnn. In other embodiments, the candidate region extraction module may be before the multi-scale object detection model, that is, the multi-scale object detection model only processes the candidate regions. Thus, the first sample data may be a picture with only one object, that is to say, the first sample data is the picture processed by the candidate region extraction module.

[0069] In the above embodiments, the candidate region extraction module extracts features and candidate regions from the picture, laying a foundation for the subsequent processing of the multi-scale object detection model.

[0070] In one of the embodiments, the training method of the multi-scale object detection model further includes a testing process, which mainly includes: obtaining test data, where the test pictures in the test data carry test labels; inputting the test pictures into the trained multi-scale object detection model to obtain test selection results through the branch selection network and test recognition results for each branch through the multi-scale recognition network; obtaining an initial prediction result according to the test selection result and the test recognition result; filtering the prediction result to obtain the target prediction result; and comparing the test label with the target prediction result to obtain the test result.

[0071] Specifically, the so-called testing stage means that the input data in the model does not have label information, and the network is only responsible for extracting features from the pictures.

[0072] First, the test pictures are input into the trained multi-scale object detection model to obtain test selection results through the branch selection network and test recognition results for each branch through the multi-scale recognition network; an initial prediction result is obtained according to the test selection result and the test recognition result.

[0073] Optionally, before inputting the test pictures into the trained multi-scale object detection model, the test pictures are first input into the candidate region extraction module to extract the features and candidate regions of the test pictures, and then the candidate regions are input into the multi-scale object detection model to obtain the above initial prediction result.

[0074] Finally, the prediction results are filtered to obtain the target prediction results. For example, the final target prediction results are obtained through non-maximum suppression filtering, and then the test labels and the target prediction results are compared to obtain the test results, i.e., the test is successful or the test fails.

[0075] In one embodiment, multiple test data can also be used for testing, so as to calculate the test success rate to determine whether the model meets the requirements.

[0076] In the above embodiment, in the model testing stage, when inputting, an image is input, then the prediction is made using the branch corresponding to the result of gumbel softmax, and finally the final prediction result is obtained through non-maximum suppression filtering for testing.

[0077] In one embodiment, as Figure 4 shown, a target detection method based on a multi-scale target detection model is provided. In this embodiment, an example is given where this method is applied to a terminal. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0078] S402: Obtain the image to be detected.

[0079] Specifically, the image to be detected can be sent to the terminal by other terminals, or selected by the user himself / herself, or obtained from a database, and no specific limitation is made here.

[0080] S404: Input the image to be detected into the multi-scale target detection model trained by the training method of the multi-scale target detection model in any of the above embodiments to obtain the target detection result.

[0081] Specifically, for the training method of the multi-scale target detection model, reference can be made to the above text and will not be elaborated here.

[0082] When inputting the image to be detected into the multi-scale target detection model, the features of the image to be detected can be extracted by the multi-scale target detection model, and then the candidate regions can be extracted according to the method of faster-rcnn. In other embodiments, if the multi-scale target detection model does not include a candidate region extraction module, the image to be detected is first input into the candidate region extraction module to extract the features of the image to be detected, and then the candidate regions are extracted according to the method of faster-rcnn.

[0083] Then, the multi-scale detection model processes each candidate region through a multi-scale recognition network and a branch selection network respectively. Thus, according to the results of the multi-scale recognition network and the branch selection results, the detection result corresponding to each candidate region can be obtained. Then, the detection results are filtered, such as non-maximum suppression processing, to obtain the object detection results. Optionally, the branch selection results can be obtained by processing through the branch selection network first, and then the branch processing results can be obtained by processing according to the branch selection results and the multi-scale recognition network. In this way, when in use, only one branch of the multi-scale recognition network in the multi-scale object detection model can be used to achieve object detection.

[0084] In the above object detection method based on a multi-scale object detection model, during training, the multi-scale recognition network and the branch selection network are jointly trained, that is, the loss function is calculated and backpropagated according to the recognition results of each branch output by the multi-scale recognition network and the branch selection results output by the branch selection network, so as to obtain the multi-scale object detection model. That is, the branch selection results are obtained according to the model, without the need for manual setting, which improves the accuracy. In addition, the object size does not need to be calculated, so no additional computational amount needs to be added. In this way, when in use, the corresponding branch can be selected by the model itself, improving the accuracy.

[0085] It should be understood that although Figure 1 and Figure 4 the steps in the flowcharts of Figure 1 and Figure 4 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover,

[0086] In one embodiment, as Figure 5 shown, a training device for a multi-scale object detection model is provided, including: a sample acquisition module 502, a data processing module 504, a model output result acquisition module 506, and a training module 508, where:

[0087] The sample acquisition module 502 is used to acquire first sample data, and each sample picture in the first sample data carries object annotation information;

[0088] A data processing module 504, configured to input the sample pictures into a multi-scale recognition network and a branch selection network of a multi-scale object detection model respectively, to obtain a branch selection result corresponding to the branch selection network and recognition results of each branch of the multi-scale recognition network. Each branch network structure of the multi-scale recognition network is the same and parameter-shared, but the coefficients of the dilated convolution are different;

[0089] A model output result acquisition module 506, configured to obtain a model output result according to the branch selection result and the recognition results of each branch;

[0090] A training module 508, configured to calculate a first loss function according to the model output result and the target annotation information, and perform backpropagation based on the first loss function to adjust the parameters in the multi-scale recognition network and the branch selection network until the first loss function meets a preset value, so as to obtain a multi-scale object detection model.

[0091] In one embodiment, the above training device of the multi-scale object detection model may further include:

[0092] A pseudo-label generation module, configured to determine a target branch according to the branch selection result, and use the recognition result of the target branch as the pseudo-label of the corresponding sample picture;

[0093] A branch training module 508, configured to train non-target branches according to the pseudo-labels of the sample pictures and the recognition results of non-target branches.

[0094] In one embodiment, the above branch training module 508 includes:

[0095] A change unit, configured to change the branch selection result of the branch selection network to select a non-target branch;

[0096] A training unit, configured to calculate a second loss function according to the pseudo-labels of the sample pictures and the recognition results of non-target branches, and perform backpropagation based on the second loss function to adjust the parameters in the multi-scale recognition network until the second loss function meets a preset value, so as to obtain a multi-scale object detection model.

[0097] In one embodiment, the above training device of the multi-scale object detection model may further include:

[0098] An extraction module, configured to extract image features of the sample pictures and generate candidate regions based on the image features.

[0099] In one embodiment, the above training device of the multi-scale object detection model may further include:

[0100] A test data acquisition module, configured to acquire test data, and the test pictures in the test data carry test labels;

[0101] A test module for inputting a test image into a trained multi-scale object detection model to obtain a test selection result through a branch selection network and obtain a test recognition result for each branch through a multi-scale recognition network;

[0102] An initial prediction result generation module for obtaining an initial prediction result according to the test selection result and the test recognition result;

[0103] A filtering module for filtering the prediction result to obtain a target prediction result;

[0104] A test result acquisition module for obtaining a test result by comparing a test label with the target prediction result.

[0105] In one embodiment, as Figure 6 shown, an object detection device based on a multi-scale object detection model is provided, including: a to-be-detected image acquisition module 602 and a detection module 604, where:

[0106] The to-be-detected image acquisition module 602 is used to acquire a to-be-detected image;

[0107] The detection module 604 is used to input the to-be-detected image into a multi-scale object detection model trained by the training device of the multi-scale object detection model in any of the above embodiments to obtain an object detection result.

[0108] For the specific definitions of the training device of the multi-scale object detection model and the object detection device based on the multi-scale object detection model, reference can be made to the definitions of the training method of the multi-scale object detection model and the object detection method based on the multi-scale object detection model in the above text, which will not be elaborated here. Each module in the above training device of the multi-scale object detection model and the object detection device based on the multi-scale object detection model can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above modules.

[0109] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 7As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it realizes a training method for a multi-scale object detection model. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0110] Those skilled in the art can understand that Figure 7 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0111] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented: obtaining first sample data, where each sample picture in the first sample data carries target annotation information; respectively inputting the sample pictures into a multi-scale recognition network and a branch selection network of a multi-scale object detection model to obtain a branch selection result corresponding to the branch selection network and an identification result of each branch of the corresponding multi-scale recognition network. Each branch network structure of the multi-scale recognition network is the same and parameter-sharing, but the coefficients of the dilated convolution are different; obtaining a model output result according to the branch selection result and the identification result of each branch; calculating a first loss function according to the model output result and the target annotation information, and backpropagating based on the first loss function to adjust the parameters in the multi-scale recognition network and the branch selection network until the first loss function meets a preset value to obtain a multi-scale object detection model.

[0112] In one embodiment, before obtaining the model output result according to the branch selection result and the identification result of each branch, which is implemented when the processor executes the computer program, it further includes: determining a target branch according to the branch selection result, and using the identification result of the target branch as the pseudo-label of the corresponding sample picture; training the non-target branches according to the pseudo-label of the sample picture and the identification results of the non-target branches.

[0113] In one embodiment, when the processor executes a computer program, training the non-target branch according to the pseudo-label of the sample picture and the recognition result of the non-target branch includes: changing the branch selection result of the branch selection network to select the non-target branch; calculating a second loss function according to the pseudo-label of the sample picture and the recognition result of the non-target branch, and backpropagating based on the second loss function to adjust the parameters in the multi-scale recognition network until the second loss function meets a preset value, thereby obtaining a multi-scale object detection model.

[0114] In one embodiment, after the processor executes a computer program to obtain first sample data, where each sample picture in the first sample data carries target annotation information, it further includes: extracting the image features of the sample picture and generating candidate regions based on the image features.

[0115] In one embodiment, the processor further implements the following steps when executing a computer program: obtaining test data, where the test pictures in the test data carry test labels; inputting the test pictures into the trained multi-scale object detection model to obtain a test selection result through the branch selection network and obtain the test recognition result of each branch through the multi-scale recognition network; obtaining an initial prediction result according to the test selection result and the test recognition result; filtering the prediction result to obtain a target prediction result; and comparing the test label with the target prediction result to obtain a test result.

[0116] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented: obtaining a picture to be detected; inputting the picture to be detected into the multi-scale object detection model obtained by training with the training method of the multi-scale object detection model in any of the above embodiments to obtain an object detection result.

[0117] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining first sample data, where each sample picture in the first sample data carries target annotation information; respectively inputting the sample pictures into a multi-scale recognition network and a branch selection network of a multi-scale object detection model to obtain a branch selection result corresponding to the branch selection network and an identification result of each branch of the corresponding multi-scale recognition network. Each branch network structure of the multi-scale recognition network is the same and parameter-sharing, but the coefficients of the dilated convolution are different; obtaining a model output result according to the branch selection result and the identification result of each branch; calculating a first loss function according to the model output result and the target annotation information, and backpropagating based on the first loss function to adjust the parameters in the multi-scale recognition network and the branch selection network until the first loss function meets a preset value, thereby obtaining the multi-scale object detection model.

[0118] In one embodiment, before obtaining the model output result according to the branch selection result and the identification result of each branch when the computer program is executed by the processor, the following steps are further included: determining a target branch according to the branch selection result, and using the identification result of the target branch as the pseudo-label of the corresponding sample picture; training the non-target branches according to the pseudo-label of the sample picture and the identification results of the non-target branches.

[0119] In one embodiment, training the non-target branches according to the pseudo-label of the sample picture and the identification results of the non-target branches when the computer program is executed by the processor includes: changing the branch selection result of the branch selection network to select a non-target branch; calculating a second loss function according to the pseudo-label of the sample picture and the identification results of the non-target branches, and backpropagating based on the second loss function to adjust the parameters in the multi-scale recognition network until the second loss function meets a preset value, thereby obtaining the multi-scale object detection model.

[0120] In one embodiment, after obtaining the first sample data, where each sample picture in the first sample data carries target annotation information, when the computer program is executed by the processor, the following steps are further included: extracting the image features of the sample pictures and generating candidate regions based on the image features.

[0121] In one embodiment, the following steps are further implemented when the computer program is executed by the processor: obtaining test data, where the test pictures in the test data carry test labels; inputting the test pictures into the trained multi-scale object detection model to obtain a test selection result through the branch selection network and a test identification result of each branch through the multi-scale recognition network; obtaining an initial prediction result according to the test selection result and the test identification result; filtering the prediction result to obtain a target prediction result; comparing the test label with the target prediction result to obtain a test result.

[0122] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining a picture to be detected; inputting the picture to be detected into a multi-scale object detection model obtained by training with the training method of the multi-scale object detection model in any of the above embodiments to obtain an object detection result.

[0123] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0124] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0125] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A training method for a multi-scale object detection model, characterized in that, The training method of the multi-scale object detection model includes: Obtain first sample data, where each sample picture in the first sample data carries object annotation information; Input the sample pictures into the multi-scale recognition network and the branch selection network of the multi-scale object detection model respectively, to obtain the branch selection result corresponding to the branch selection network and the recognition result of each branch of the multi-scale recognition network. Each branch network structure of the multi-scale recognition network is the same and the parameters are shared, but the coefficients of the dilated convolution are different; Obtain the model output result according to the branch selection result and the recognition result of each branch; Calculate the first loss function according to the model output result and the object annotation information, and backpropagate based on the first loss function to adjust the parameters in the multi-scale recognition network and the branch selection network until the first loss function meets the preset value, to obtain the multi-scale object detection model; Before obtaining the model output result according to the branch selection result and the recognition result of each branch, it further includes: Determine the target branch according to the branch selection result, and use the recognition result of the target branch as the pseudo-label of the corresponding sample picture; Train the non-target branches according to the pseudo-label of the sample picture and the recognition results of the non-target branches.

2. The training method of the multi-scale object detection model according to claim 1, wherein, The training of the non-target branches according to the pseudo-label of the sample picture and the recognition results of the non-target branches includes: Change the branch selection result of the branch selection network to select the non-target branch; Calculate the second loss function according to the pseudo-label of the sample picture and the recognition results of the non-target branches, and backpropagate based on the second loss function to adjust the parameters in the multi-scale recognition network until the second loss function meets the preset value, to obtain the multi-scale object detection model.

3. The training method of the multi-scale object detection model according to claim 1 or 2, characterized in that After obtaining the first sample data, where each sample picture in the first sample data carries object annotation information, it further includes: Extract the image features of the sample pictures, and generate candidate regions based on the image features.

4. The training method of the multi-scale object detection model according to claim 1 or 2, characterized in that, The method further includes: Obtain test data, where the test pictures in the test data carry test labels; Input the test pictures into the trained multi-scale object detection model, to obtain the test selection result through the branch selection network and the test recognition result of each branch through the multi-scale recognition network; Obtain the initial prediction result according to the test selection result and the test recognition result; Filter the prediction result to obtain the target prediction result; Compare the test label and the target prediction result to obtain the test result.

5. A target detection method based on a multi-scale target detection model, characterized in that, The object detection method based on the multi-scale object detection model includes: Obtain the picture to be detected; Input the picture to be detected into the multi-scale object detection model trained by the training method of the multi-scale object detection model according to any one of claims 1 to 4 to obtain the object detection result.

6. A training device for a multi-scale object detection model, characterized in that, The training device of the multi-scale object detection model includes: A sample acquisition module, configured to obtain first sample data, where each sample picture in the first sample data carries object annotation information; A data processing module, configured to input the sample pictures into a multi-scale recognition network and a branch selection network of a multi-scale object detection model respectively, to obtain a branch selection result corresponding to the branch selection network and recognition results corresponding to each branch of the multi-scale recognition network, where each branch network structure of the multi-scale recognition network is the same and the parameters are shared, but the coefficients of the dilated convolution are different; A model output result acquisition module, configured to obtain a model output result according to the branch selection result and the recognition results of each branch; A training module, configured to calculate a first loss function according to the model output result and the target annotation information, and perform backpropagation based on the first loss function to adjust the parameters in the multi-scale recognition network and the branch selection network until the first loss function meets a preset value, so as to obtain a multi-scale object detection model; A pseudo-label generation module, configured to determine a target branch according to the branch selection result, and use the recognition result of the target branch as the pseudo-label of the corresponding sample picture; A branch training module, configured to train the non-target branches according to the pseudo-labels of the sample pictures and the recognition results of the non-target branches.

7. The device according to claim 6, characterized in that, The branch training module includes: A change unit, configured to change the branch selection result of the branch selection network to select a non-target branch; A training unit, configured to calculate a second loss function according to the pseudo-labels of the sample pictures and the recognition results of the non-target branches, and perform backpropagation based on the second loss function to adjust the parameters in the multi-scale recognition network until the second loss function meets a preset value, so as to obtain a multi-scale object detection model.

8. An object detection device based on a multi-scale object detection model, characterized in that, The object detection device based on the multi-scale object detection model includes: A to-be-detected picture acquisition module, configured to acquire a to-be-detected picture; A detection module, configured to input the to-be-detected picture into the multi-scale object detection model trained by the training device of the multi-scale object detection model according to claim 6 or 7 to obtain an object detection result.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 or 5 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 4 or 5 are implemented.

Citation Information

Patent Citations

  • Target detection network construction method and training method, and target detection method

    CN109784194A