Target detection method and device, computer device and storage medium
Patent Information
- Application Number
- CN202310217325.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-08
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2043-03-08
AI Technical Summary
[0003]但现有的图像自适应动态网络主要应用在图像分类功能方面,在很多下游功能如视频目标检测还未进行深入研究;同时,也未形成一套快速且精确的规则进行图像自适应推断以完成对图像的目标检测,因此,图像目标检测的效率和精确度还需提高
[0066] The aforementioned target detection method, apparatus, computer equipment, and storage medium introduce multiple branch networks arranged in order of precision into the target detection model, and configure a corresponding threshold for each branch network. When performing target detection on an image to be detected, the image to be detected is first input into the first branch network of the target detection model (i.e., the branch network with the lowest precision in the target detection model), and the first branch network outputs the target detection result and the confidence level of the target detection result. Then, the score of the image to be detected, determined based on the confidence level of the target detection result, is compared with the threshold corresponding to the first branch network. If the score of the image to be detected is greater than the threshold corresponding to the first branch network, the target detection result output by the first branch network is taken as the final detection result of the image to be detected. If the score of the image to be detected is less than the threshold corresponding to the first branch network, the image to be detected is input into the next branch network of the first branch network in order of precision, and the above process is repeated until there is a score determined based on the confidence level of the target detection result output by a branch network that is greater than the threshold corresponding to that branch network, and the target detection result output by that branch network is taken as the final detection result of the image to be detected. Compared to related target detection methods, this application, on the one hand, enables dynamic inference of the target detection process of the image to be detected by configuring a corresponding threshold for each branch network. That is, it can dynamically select which branch network to output the final detection result of the image to be detected based on the relationship between the score of the image to be detected and the threshold corresponding to the branch network. On the other hand, when performing target detection on the image, not all branch networks are used, but target detection is performed on the image starting from the branch network with the lowest precision according to the precision order of the branch networks, until there is a score determined by the confidence of the target detection result output by a branch network that is greater than the threshold corresponding to that branch network. This improves the efficiency of target detection while ensuring the accuracy of target detection.
Smart Images

Figure CN116206157B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more particularly to the field of deep learning technology, specifically to a target detection method, apparatus, computer device, and storage medium. Background Technology
[0002] Existing image adaptive dynamic networks improve upon the traditional paradigm of static inference in static networks by enabling adaptive image inference based on actual conditions, thus greatly improving inference efficiency.
[0003] However, existing image adaptive dynamic networks are mainly used for image classification, and their application in many downstream functions such as video object detection has not been studied in depth. At the same time, a set of fast and accurate rules for image adaptive inference to complete image object detection has not been formed. Therefore, the efficiency and accuracy of image object detection need to be improved. Summary of the Invention
[0004] Therefore, it is necessary to provide a target detection method, apparatus, computer equipment, and storage medium to address the aforementioned technical problems, which can perform adaptive inference of images quickly and accurately, ultimately achieving efficient and accurate target detection of images.
[0005] Firstly, this application provides a target detection method, which includes:
[0006] The first branch network of the target detection model is taken as the target network; wherein the target detection model includes at least two branch networks arranged in order of accuracy, and the first branch network has the lowest accuracy;
[0007] The image to be detected is input into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network;
[0008] The score of the image to be detected is determined based on the confidence level of the target detection results;
[0009] If the score of the image to be detected is less than the threshold corresponding to the target network, then the next branch network of the target network is used as the new target network, and the process is repeated to input the image to be detected into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network, until the score of the image to be detected is greater than the threshold corresponding to the target network.
[0010] The target detection results output by the target network are used as the final detection results for the image to be detected.
[0011] In one embodiment, each branch network includes an initial feature extraction layer, a feature extraction subnetwork, and a detection subnetwork;
[0012] The initial feature extraction layers of each branch network are connected sequentially according to the arrangement order of each branch network;
[0013] The number of downsampling layers included in the feature extraction subnetwork of each branch network decreases sequentially according to the order of the branch networks.
[0014] In one embodiment, the image to be detected is input into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network, including:
[0015] The image to be detected or its previous basic feature map is input into the initial feature extraction layer of the target network to obtain the current basic feature map of the image to be detected; wherein, the previous basic feature map is output by the initial feature extraction layer in the previous branch network of the target network.
[0016] The current base feature map is input into the feature extraction subnetwork of the target network to obtain a depth feature map; wherein, the depth feature maps output by different target networks have the same scale;
[0017] The deep feature map is input into the detection subnetwork of the target network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0018] In one embodiment, each detection subnetwork includes a region recommendation network and a detection head network;
[0019] The deep feature map is input into the detection subnetwork of the target network to obtain the target detection result and the confidence level of the target detection result output by the target network, including:
[0020] The deep feature map is input into the region recommendation network to obtain at least one candidate box;
[0021] Input at least one candidate bounding box and a depth feature map into the detection head network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0022] In one embodiment, the target detection result includes at least one target detection box;
[0023] Based on the confidence level of the target detection results, a score is determined for the image to be detected, including:
[0024] The average confidence score of each target detection box is used as the score of the image to be detected.
[0025] In one embodiment, the training process of the object detection model is as follows:
[0026] Obtain the training set of sample images;
[0027] The training sample images in the sample image training set are input into each branch network in the target detection model to obtain the first sample detection result output by the region recommendation network in each branch network and the second sample detection result output by the detection head network in each branch network.
[0028] Based on the training sample labels of the training sample images, and the first sample detection results and second sample detection results output by each branch network, the branch loss of each branch network is determined.
[0029] The total loss of the target detection model is determined based on the path loss of each branch network;
[0030] The target detection model is trained using the total loss.
[0031] In one embodiment, the splitting loss of each splitting network is determined based on the training sample labels of the training sample images and the first and second sample detection results output by each splitting network, including:
[0032] For each branch network, the first loss is determined based on the training sample labels of the training sample images and the first sample detection result output by the region recommendation network in the branch network;
[0033] The second loss is determined based on the training sample labels of the training sample images and the second sample detection results output by the detection head network in the split network;
[0034] The routing loss of the routing network is determined based on the first loss and the second loss.
[0035] In one embodiment, the total loss of the target detection model is determined based on the branch loss of each branch network, including:
[0036] The total loss of the object detection model is determined based on the branch loss and weight of each branch network, as well as the number of samples in the training set of sample images.
[0037] In one embodiment, after training the object detection model, the method further includes:
[0038] Obtain a sample image validation set;
[0039] The validation sample images in the sample image validation set are input into the target detection model to obtain the confidence of the third sample detection result output by each branch network in the target detection model;
[0040] Based on the confidence level of the third sample detection results output by each branch network, determine the score of the verification sample image input to each branch network;
[0041] The threshold for each branch network is determined based on the scores of the validation sample images input to each branch network.
[0042] Secondly, this application also provides a target detection device, which includes:
[0043] The first determining module is used to select the first branch network of the target detection model as the target network; wherein the target detection model includes at least two branch networks arranged in order of accuracy, and the first branch network has the lowest accuracy;
[0044] The detection module is used to input the image to be detected into the target network and obtain the target detection result and the confidence level of the target detection result output by the target network;
[0045] The scoring determination module is used to determine the score of the image to be detected based on the confidence level of the target detection results;
[0046] The second determining module is used to, if the score of the image to be detected is less than the threshold corresponding to the target network, take the next branch network of the target network as the new target network, return to execute the input of the image to be detected into the target network, and obtain the target detection result and the confidence of the target detection result output by the target network, until the score of the image to be detected is greater than the threshold corresponding to the target network.
[0047] The output module is used to take the target detection results output by the target network as the final detection result of the image to be detected.
[0048] Thirdly, this application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0049] The first branch network of the target detection model is taken as the target network; wherein the target detection model includes at least two branch networks arranged in order of accuracy, and the first branch network has the lowest accuracy;
[0050] The image to be detected is input into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0051] The score of the image to be detected is determined based on the confidence level of the target detection results;
[0052] If the score of the image to be detected is less than the threshold corresponding to the target network, then the next branch network of the target network is used as the new target network, and the process is repeated to input the image to be detected into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network, until the score of the image to be detected is greater than the threshold corresponding to the target network.
[0053] The target detection results output by the target network are used as the final detection results for the image to be detected.
[0054] Fourthly, this application also provides a computer-readable storage medium on which a computer program is stored, and when executed by a processor, the computer program performs the following steps:
[0055] The first branch network of the target detection model is taken as the target network; wherein the target detection model includes at least two branch networks arranged in order of accuracy, and the first branch network has the lowest accuracy;
[0056] The image to be detected is input into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network;
[0057] The score of the image to be detected is determined based on the confidence level of the target detection results;
[0058] If the score of the image to be detected is less than the threshold corresponding to the target network, then the next branch network of the target network is used as the new target network, and the process is repeated to input the image to be detected into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network, until the score of the image to be detected is greater than the threshold corresponding to the target network.
[0059] The target detection results output by the target network are used as the final detection results for the image to be detected.
[0060] Fifthly, this application also provides a computer program product comprising a computer program that, when executed by a processor, performs the following steps:
[0061] The first branch network of the target detection model is taken as the target network; wherein the target detection model includes at least two branch networks arranged in order of accuracy, and the first branch network has the lowest accuracy;
[0062] The image to be detected is input into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network;
[0063] The score of the image to be detected is determined based on the confidence level of the target detection results;
[0064] If the score of the image to be detected is less than the threshold corresponding to the target network, then the next branch network of the target network is used as the new target network, and the process is repeated to input the image to be detected into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network, until the score of the image to be detected is greater than the threshold corresponding to the target network.
[0065] The target detection results output by the target network are used as the final detection results for the image to be detected.
[0066] The aforementioned target detection method, apparatus, computer equipment, and storage medium introduce multiple branch networks arranged in order of precision into the target detection model, and configure a corresponding threshold for each branch network. When performing target detection on an image to be detected, the image to be detected is first input into the first branch network of the target detection model (i.e., the branch network with the lowest precision in the target detection model), and the first branch network outputs the target detection result and the confidence level of the target detection result. Then, the score of the image to be detected, determined based on the confidence level of the target detection result, is compared with the threshold corresponding to the first branch network. If the score of the image to be detected is greater than the threshold corresponding to the first branch network, the target detection result output by the first branch network is taken as the final detection result of the image to be detected. If the score of the image to be detected is less than the threshold corresponding to the first branch network, the image to be detected is input into the next branch network of the first branch network in order of precision, and the above process is repeated until there is a score determined based on the confidence level of the target detection result output by a branch network that is greater than the threshold corresponding to that branch network, and the target detection result output by that branch network is taken as the final detection result of the image to be detected. Compared to related target detection methods, this application, on the one hand, enables dynamic inference of the target detection process of the image to be detected by configuring a corresponding threshold for each branch network. That is, it can dynamically select which branch network to output the final detection result of the image to be detected based on the relationship between the score of the image to be detected and the threshold corresponding to the branch network. On the other hand, when performing target detection on the image, not all branch networks are used, but target detection is performed on the image starting from the branch network with the lowest precision according to the precision order of the branch networks, until there is a score determined by the confidence of the target detection result output by a branch network that is greater than the threshold corresponding to that branch network. This improves the efficiency of target detection while ensuring the accuracy of target detection. Attached Figure Description
[0067] Figure 1A This is a flowchart illustrating a target detection method in one embodiment;
[0068] Figure 1B This is a schematic diagram of the target detection model in one embodiment;
[0069] Figure 2A This is a schematic diagram of the target detection model in another embodiment;
[0070] Figure 2B This is a schematic diagram illustrating the process of obtaining target detection results and the confidence level of target detection results from a target network in one embodiment;
[0071] Figure 3This is a schematic diagram of the structure of a detection subnetwork of a target detection model in one embodiment;
[0072] Figure 4 This is a flowchart illustrating the target detection method in another embodiment;
[0073] Figure 5 This is a flowchart illustrating a training method for an object detection model in one embodiment.
[0074] Figure 6 This is a flowchart illustrating a verification method for a target detection model in one embodiment;
[0075] Figure 7 This is a structural block diagram of a target detection device in one embodiment;
[0076] Figure 8 This is a structural block diagram of the target detection device in another embodiment;
[0077] Figure 9 This is a structural block diagram of the target detection device in another embodiment;
[0078] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0079] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0080] The target detection method provided in this application can be applied to the detection of targets in images. This embodiment illustrates the method by applying it to a server. It is understood that the method can also be applied to a terminal, or to a system including a terminal and a server, and implemented through the interaction between the terminal and the server.
[0081] In one embodiment, Figure 1A This is a schematic flowchart of a target detection method provided according to an embodiment of this application; Figure 1B This is a schematic diagram of the structure of a target detection model provided according to an embodiment of this application. Taking the application of this method to a server as an example, the method includes the following steps:
[0082] S101, the first branch network of the object detection model is used as the target network.
[0083] In this embodiment, the target detection model includes at least two branch networks arranged in order of accuracy, with the first branch network having the lowest accuracy. Optionally, each branch network has the same function and can be used to detect targets in the image.
[0084] It should be noted that the number of branch networks in the object detection model can be selected as needed, and this application does not impose any restrictions on this. For the sake of facilitating the subsequent description, this application selects a number of 4 branch networks in the object detection model as an example for illustration.
[0085] like Figure 1B As shown, the target detection model 1 includes four branch networks, which are arranged in ascending order of accuracy as branch network 1, branch network 2, branch network 3 and branch network 4; that is, branch network 1 has the lowest accuracy and is the first branch network. Each branch network is used to detect the input image to be detected.
[0086] S102, input the image to be detected into the target network, and obtain the target detection result and the confidence level of the target detection result output by the target network.
[0087] Specifically, the image to be detected is input into the target network, for example... Figure 1B In the branch network 1, the target network performs target detection on the image to be detected based on its own network parameters, and outputs the target detection result and the confidence level of the target detection result.
[0088] The target detection result may include the location information of the target in the image to be detected as predicted by the target network, specifically including at least one target detection box; furthermore, the target detection result may also include the category information of the target in the image to be detected as predicted by the target network.
[0089] Optionally, the confidence level of the target detection result refers to the credibility of the target detection result output by the target network based on the image to be detected. Furthermore, the confidence level can take any value between 0 and 1; a higher confidence level indicates a more realistic target detection result output by the target network based on the image to be detected. It should be noted that the target detection result can include at least one target detection box, and correspondingly, the confidence level of the target detection result can include the confidence level of each target detection box. Furthermore, if the target detection result can also include the category information of the target in the image to be detected predicted by the target network, then the confidence level of the target detection result can also include the confidence level of each category information, etc.
[0090] S103, Determine the score of the image to be detected based on the confidence level of the target detection results.
[0091] The so-called score of the image to be detected can be used to characterize the quality of the target detection result of the image to be detected.
[0092] In one implementation, the confidence level of the target detection result output by the target network can be processed based on a pre-defined logic for evaluating the quality of image detection results to obtain a score for the image to be detected. Optionally, if the target detection result includes at least one target detection box, the average confidence level of each target detection box can be used as the score for the image to be detected. Furthermore, if the target detection result also includes the category information of the target in the image to be detected predicted by the target network, the score for the image to be detected can be determined by combining the confidence levels of each category information and the confidence levels of each target detection box.
[0093] Because different branch networks have different accuracies, for the same image to be detected, the confidence levels of the target detection results output by different branch networks based on the image to be detected are different. Therefore, different branch networks give different scores to the same image to be detected.
[0094] S104, determine whether the score of the image to be detected is greater than the threshold corresponding to the target network; if yes, proceed to S105; if no, proceed to S106.
[0095] S105, take the target detection result output by the target network as the final detection result of the image to be detected.
[0096] S106, take the next branch network of the target network as the new target network, and return to execute S102.
[0097] Specifically, the score of the image to be detected is compared with the threshold of the target network; if the score of the image to be detected is greater than the threshold of the target network, the target detection result output by the target network is taken as the final detection result of the image to be detected.
[0098] If the score of the image to be detected is less than the threshold corresponding to the target network, then according to the accuracy order, the next branch network of the target network is taken as the new target network, and the process returns to S102 until the score of the image to be detected is greater than the threshold corresponding to the target network, and then the target network is taken as the target network for the final output target detection result.
[0099] The following are Figure 1BTaking an example, the process of obtaining the final detection result of the image to be detected is described. Specifically, the image to be detected is input into the split network 1 to obtain the target detection result and the confidence level of the target detection result output by the split network 1; the average confidence level of the target detection result output by the split network 1 is used as the score of the image to be detected; it is determined whether the score of the image to be detected is greater than the threshold corresponding to the split network 1. If it is, the target detection result output by the split network 1 is used as the final detection result of the image to be detected; if not, the split network 2 is used as a new target network, the image to be detected is input into the split network 2 to obtain the target detection result and the confidence level of the target detection result output by the split network 2; the average confidence level of the target detection result output by the split network 2 is used as the score of the image to be detected; it is determined whether the score of the image to be detected is greater than the threshold corresponding to the split network 2. If it is, the target detection result output by the split network 2 is used as the final detection result of the image to be detected; if not, the split network 3 is used as a new target network, and the above process is repeated until the final detection result of the image to be detected is obtained.
[0100] In the aforementioned object detection method, multiple branch networks arranged in order of accuracy are introduced into the object detection model, and a corresponding threshold is configured for each branch network. When performing object detection on the image to be detected, the image to be detected is first input into the first branch network of the object detection model (i.e., the branch network with the lowest accuracy in the object detection model). The first branch network outputs the object detection result and its confidence score. Then, the score of the image to be detected, determined based on the confidence score of the object detection result, is compared with the threshold corresponding to the first branch network. If the score of the image to be detected is greater than the threshold corresponding to the first branch network, the object detection result output by the first branch network is taken as the final detection result of the image to be detected. If the score of the image to be detected is less than the threshold corresponding to the first branch network, the image to be detected is input into the next branch network of the first branch network in order of accuracy, and the above process is repeated until there is a score determined based on the confidence score of the object detection result output by a branch network that is greater than the threshold corresponding to that branch network. The object detection result output by that branch network is then taken as the final detection result of the image to be detected. Compared to related target detection methods, this application, on the one hand, enables dynamic inference of the target detection process of the image to be detected by configuring a corresponding threshold for each branch network. That is, it can dynamically select which branch network to output the final detection result of the image to be detected based on the relationship between the score of the image to be detected and the threshold corresponding to the branch network. On the other hand, when performing target detection on the image, not all branch networks are used, but target detection is performed on the image starting from the branch network with the lowest precision according to the precision order of the branch networks, until there is a score determined by the confidence of the target detection result output by a branch network that is greater than the threshold corresponding to that branch network. This improves the efficiency of target detection while ensuring the accuracy of target detection.
[0101] To make the extracted image features more accurate, the split-path network is further refined based on the above embodiments. Optionally, such as... Figure 2A As shown, each branch network includes an initial feature extraction layer, a feature extraction subnetwork, and a detection subnetwork.
[0102] The initial feature extraction layer can be a convolutional layer. In this embodiment, the convolutional depth of the initial feature extraction layer varies for different branch networks. Figure 2A The initial feature extraction layers 1, 2, 3, and 4 in the model are different; furthermore, for example... Figure 2A As shown, in this embodiment, the initial feature extraction layers of each branch network are connected sequentially according to the arrangement order of each branch network, so that each branch network in the target detection model 1 is connected together.
[0103] In this embodiment, the initial feature extraction layer 1 corresponds to a 7*7 convolution kernel with a stride of 2. Its input is the image to be detected, and its output is a coarse basic feature map with a scale of input scale / 2. The features in this part are relatively simple features such as texture, color, and contour. The initial feature extraction layers 2, 3, and 4 are stacks of different bottleneck modules. The number of bottleneck modules stacked in each initial feature extraction layer is 3, 4, and 6, respectively. The idea of the bottleneck modules is to reduce the dimension first and then increase the dimension to reduce the computation. The input of the initial feature extraction layers 2, 3, and 4 is the previous basic feature map of the image to be detected, and the output scales are input scale / 4, input scale / 8, and input scale / 16, respectively, which can be used to generate progressively finer feature maps. The features in the feature maps generated by the initial feature extraction layers 2, 3, and 4 are mainly high-level features such as semantic information, association information, and main structure.
[0104] Optionally, the initial feature extraction layer in each branch network is connected to the feature extraction subnetworks within that branch network. For example, Figure 2A As shown, the initial feature extraction layer 1 in the branch network 1 is connected to the feature extraction sub-network 1 in the branch network 1.
[0105] Furthermore, each feature extraction subnetwork in the branch network can include a downsampling layer, and all downsampling layers can be the same; optionally, the number of downsampling layers included in different branch networks is different; specifically, in this embodiment, the number of downsampling layers included in the feature extraction subnetwork of each branch network decreases sequentially according to the arrangement order of the branch networks. For example, Figure 2A As shown, the feature extraction subnetwork in the branch network can include downsampling layers. Since the scales of the basic feature maps output by the initial feature extraction layers 1, 2, 3, and 4 are input scale / 2, input scale / 4, input scale / 8, and input scale / 16, respectively, the number of downsampling layers in the feature extraction subnetwork decreases sequentially according to the order of the branch networks. The number of downsampling layers in feature extraction subnetworks 1, 2, 3, and 4 are 3, 2, 1, and 0, respectively. This ensures that the depth feature maps output by different target networks have the same scale. On the one hand, the depth feature maps output by the feature extraction subnetworks at the same scale ensure that the feature map size of each target is at the same scale, which is convenient for detection. On the other hand, although branch networks 2, 3, and 4 are used to extract increasingly fine depth feature maps from the image to be detected, their receptive fields become smaller and smaller. The corresponding depth feature maps may lack partial global information. Therefore, feature extraction at different scales is performed in the branch network to provide more semantic information for subsequent detection.
[0106] It should be noted that in this embodiment, feature extraction subnetworks with different numbers of downsampling layers are introduced for different branch networks, which can ensure that the depth feature maps output by the feature extraction subnetworks of different branch networks have the same scale.
[0107] Furthermore, each feature extraction subnetwork in the branch network is also connected to the detection subnetwork in that branch network. For example, Figure 2A As shown, feature extraction subnetwork 1 in branch network 1 is connected to detection subnetwork 1 in branch network 1, feature extraction subnetwork 2 in branch network 2 is connected to detection subnetwork 2 in branch network 2, feature extraction subnetwork 3 in branch network 3 is connected to detection subnetwork 3 in branch network 3, and feature extraction subnetwork 4 in branch network 4 is connected to detection subnetwork 4 in branch network 4.
[0108] Based on this, such as Figure 2B As shown, the image to be detected is input into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network. Specifically, this may include the following steps:
[0109] S201, input the image to be detected or the previous basic feature map of the image to be detected into the initial feature extraction layer of the target network to obtain the current basic feature map of the image to be detected.
[0110] Optionally, if the target network is the first branch network, the input to the initial feature extraction layer of the target network is the image to be detected; in this case, the initial feature extraction layer of the target network extracts features from the image to be detected to obtain the current basic feature map of the image to be detected.
[0111] If the target network is not the first branch network, the initial feature extraction layer of the target network is the previous basic feature map of the image to be detected. In this case, the initial feature extraction layer extracts features from the previous basic feature map of the image to be detected to obtain the current basic feature map of the image to be detected. The previous basic feature map is output by the initial feature extraction layer in the previous branch network of the target network.
[0112] For example, such as Figure 2A As shown, if the target network is a branch network 1, the image to be detected is input to the initial feature extraction layer of branch network 1, and the initial feature extraction layer of branch network 1 performs feature extraction on the image to be detected to obtain the current basic feature map of the image to be detected; if the target network is a branch network 2, the previous basic feature map output by the initial feature extraction layer in branch network 1 is input to the initial feature extraction layer of branch network 2, and the initial feature extraction layer of branch network 2 performs feature extraction on the previous basic feature map to obtain the current basic feature map of the image to be detected.
[0113] S202, input the current basic feature map into the feature extraction subnetwork of the target network to obtain the deep feature map.
[0114] Optionally, in this embodiment, the depth feature maps output by different target networks have the same scale.
[0115] Specifically, the current basic feature map is input into the downsampling layer of the feature extraction subnetwork of the target network. The downsampling layer downsamples the current basic feature map to perform deeper feature extraction, resulting in a deep feature map.
[0116] S203, input the deep feature map into the detection subnetwork of the target network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0117] Specifically, the deep feature map is input into the detection subnetwork of the target network, and the detection subnetwork performs target detection on the deep feature map to obtain the target detection result and the confidence level of the target detection result.
[0118] It is understandable that the introduction of an initial feature extraction layer and a feature extraction sub-network in this embodiment to perform progressive feature extraction on the image to be detected can make the extracted features richer and more accurate.
[0119] In one possible implementation, the detection sub-network is refined based on the above embodiments. Combined with... Figure 3 As shown, each detection sub-network includes a region recommendation network 30 and a detection head network 40; wherein, the region recommendation network 30 includes an anchor box generation module 31, a classification regression module 32 and a non-maximum suppression module 33; the detection head network 40 includes a pooling module 41 and a classification regression module 42.
[0120] Furthermore, step S203 can be further refined as follows: inputting the deep feature map into the region recommendation network to obtain at least one candidate box; inputting the at least one candidate box and the deep feature map into the detection head network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0121] Specifically, the process of inputting the depth feature map into the region recommendation network 30 to obtain at least one candidate box can be as follows: the depth feature map is convolved using a 3*3 convolution kernel, and the result is input into the region recommendation network 30. Using each feature point of the depth feature map as an anchor, the anchor box generation module 31 generates nine different anchor boxes at each feature point. The classification and regression module 32 included in the region recommendation network 30 has two branches: the first branch classifies the anchor boxes using a softmax activation function; the second branch performs offset regression on the bounding boxes generated by the anchor box generation module to obtain accurate candidate boxes; then, candidate boxes that are too small or out of bounds are deleted, and duplicate candidate boxes are deleted using a non-maximum suppression module 33, finally selecting candidate boxes that meet the above conditions from the anchor boxes.
[0122] Furthermore, the detection head network 40 has two inputs: a depth feature map and candidate boxes generated by the region recommendation network 30. Specifically, the process of inputting at least one candidate box and the depth feature map into the detection head network to obtain the target detection result and its confidence level from the target network can be as follows: the candidate boxes generated by the region recommendation network 30 provide the target's location information, and the depth feature map provides the features corresponding to that location information. The pooling module 41 in the detection head network 40 collects the candidate boxes generated by the region recommendation network 30. Due to the different scales of the targets, the candidate boxes corresponding to the targets are divided into a 7x7 grid, with the maximum value of each grid cell used as the output, thus achieving a fixed-size output for each candidate box. The classification and regression module 42 calculates the category to which each candidate box belongs, outputs the probability value of each category, regresses the candidate box, and calculates the positional offset, ultimately obtaining the target detection result and its confidence level from the target network.
[0123] It should be noted that in this embodiment, the backbone network ResNet used in the prior art for object detection is changed to multi-head ResNet, i.e., MH-ResNet. MH-ResNet includes an initial feature extraction layer and a feature extraction sub-network in each branch network. At the same time, Faster R-CNN is introduced as the main framework of MH-ResNet. By combining Faster R-CNN and MH-ResNet, object detection based on dynamic neural networks is realized, that is, dynamic networks and dynamic inference are applied to object detection.
[0124] Additionally, in one embodiment, this application also provides an optional example of an object detection method. (Combined with...) Figure 4 As shown, the specific process includes:
[0125] S401 uses the first branch network of the object detection model as the target network.
[0126] S402, input the image to be detected or the previous basic feature map of the image to be detected into the initial feature extraction layer of the target network to obtain the current basic feature map of the image to be detected.
[0127] S403: Input the current basic feature map into the feature extraction subnetwork of the target network to obtain the deep feature map.
[0128] S404, input the deep feature map into the region recommendation network of the target network to obtain at least one candidate box.
[0129] S405, input at least one candidate box and a depth feature map into the detection head network of the target network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0130] S406, the average confidence score of each object detection box is used as the score of the image to be detected.
[0131] S407, determine whether the score of the image to be detected is greater than the threshold corresponding to the target network; if yes, proceed to S408; if no, proceed to S409.
[0132] S408 uses the target detection results output by the target network as the final detection result for the image to be detected.
[0133] S409, take the next branch network of the target network as the new target network, and return to execute S402.
[0134] The specific processes of S401-S409 described above can be found in the description of the above method embodiments. Their implementation principles and technical effects are similar, and will not be repeated here.
[0135] It should be noted that the object detection model used in the above object detection method needs to be obtained through training. Therefore, in one embodiment, this application also provides a flowchart of an object detection model training method, combined with... Figure 5 As shown, the specific process includes:
[0136] S501, Obtain the training set of sample images.
[0137] In this embodiment, the sample image training set includes multiple training sample images, and each training sample image has a training sample label.
[0138] S502, input the training sample images from the sample image training set into each branch network of the target detection model to obtain the first sample detection result output by the region recommendation network in each branch network and the second sample detection result output by the detection head network in each branch network.
[0139] Specifically, the training sample images in the sample image training set are input into each branch network in the target detection model. The region recommendation network in each branch network outputs the first sample detection result, and the first sample detection result is input into the detection head network in each branch network. The detection head network outputs the second sample detection result.
[0140] The first sample detection result may include at least one candidate box and category information predicted by the region recommendation network; the second sample detection result may include at least one target detection box and category information predicted by the detection head network.
[0141] S503, based on the training sample labels of the training sample images and the first and second sample detection results output by each branch network, determine the branch loss of each branch network.
[0142] This application designs a corresponding loss function when training the object detection model. During the training process of the object detection model, the main losses are the classification loss and regression loss of the region recommendation network and the classification loss and regression loss of the detection head network, while other losses are negligible.
[0143] For each branch network, the branching loss of that branch network can be determined by the following steps.
[0144] The first step is to determine the first loss based on the training sample labels of the training sample images and the first sample detection results output by the region recommendation network in the split network.
[0145] The training sample labels for each training sample image may include ground truth bounding boxes, as well as type or category labels.
[0146] In this embodiment, the first loss includes the classification loss and regression loss of the regional recommendation network.
[0147] Specifically, the classification loss of a region recommendation network can be labeled as L. rpn_cls It is used to determine whether the anchor box becomes the foreground of the candidate box, and is expressed as formula (1).
[0148] L rpn_cls (p i ,p i * )=-log[p i * p i +(1-p i * ),(1-pi)] (1)
[0149] Where, p iThe probability that the anchor box is predicted as a candidate box, and the true value p. i * The type of the label in the anchor box is determined by formula (2).
[0150]
[0151] Specifically, when the tag type is a negative tag, the real value is 0; when the tag type is a positive tag, the real value is 1.
[0152] Specifically, the regression loss of the region recommendation network can be labeled as L. rpn_loc It is used to fine-tune the position of the anchor frame, and it is expressed as formula (3).
[0153]
[0154] Among them, t i ={t x ,t y ,t w ,t h} is a vector representing the four parametric coordinates of the predicted candidate box, where t x =(xx) a ) / w a t y =(yy) a ) / h a t w =log(w / w) a ),t h =log(h / h) a ), t i * For p corresponding to the active anchoring box i * The coordinate vector, i.e., the coordinates of the actual bounding box, where t x * =(xx) a ) / w a t y * =(yy) a ) / h a t w * =log(w / w) a ),t h * =log(h / h) a ), where R is the smooth L1 function, which is expressed as formula (4).
[0155]
[0156] The second step is to determine the second loss based on the training sample labels of the training sample images and the second sample detection results output by the detection head network in the split network.
[0157] In this embodiment, the second loss includes the classification loss and regression loss of the detection head network.
[0158] L roi_cls and L roi_loc These are the classification loss and regression loss of the detection head network, respectively. Their calculation methods are the same as those for calculating the classification loss of the region recommendation network, and will not be repeated here.
[0159] The third step is to determine the branching loss of the branching network based on the first loss and the second loss.
[0160] After obtaining the first loss and the second loss, the first loss and the second loss are summed to obtain the branch loss of the branch network. The branch loss of the branch network can be expressed as the following formula (5).
[0161] L(f k ) = L rpn_loc +L rpn_cls +L roi_loc +L roi_cls (5)
[0162] S504: Determine the total loss of the target detection model based on the branch loss of each branch network.
[0163] One possible approach is to use the sum of the branch losses of each branch network as the total loss of the entire object detection model.
[0164] Optionally, to ensure that the determined total loss is more accurate, another possible approach is to determine the total loss of the target detection model using the following formula (6), based on the branch loss and weight of each branch network and the number of samples in the sample image training set.
[0165]
[0166] Where ω is the weight of each classifier, and D is the number of samples in the training image set.
[0167] S505 uses total loss to train the object detection model.
[0168] Specifically, after obtaining the total loss of the entire object detection network, the model parameters in the object detection model can be updated based on the total loss using the batch stochastic gradient descent method. Furthermore, the above process of training the object detection model is repeated until a training threshold is reached, or the accuracy of the object detection model reaches the set requirements.
[0169] Understandably, the above-mentioned object detection model training method trains the object detection model by designing a loss function to calculate the total loss of the model detection model and subdividing the total loss into the losses in each subnetwork. This can result in a high-precision object detection model, thereby improving the accuracy of subsequent object detection.
[0170] Optionally, after training the object detection model, it is also necessary to validate the model to ensure its effectiveness and accuracy. Therefore, in one embodiment, this application also provides a method for validating an object detection model, which involves obtaining the threshold values corresponding to each branch network in the object detection model during the validation process. Combined with... Figure 6 As shown, the specific process includes:
[0171] S601, Obtain the sample image verification set.
[0172] In this embodiment, the sample image verification set includes multiple verification sample images.
[0173] S602, input the validation sample images from the sample image validation set into the target detection model to obtain the confidence level of the third sample detection result output by each branch network in the target detection model.
[0174] Specifically, for each verification sample image, the verification sample image is input into each branch network to obtain the third sample detection result output by each branch network based on the verification sample image, as well as the confidence level of the third sample detection result.
[0175] Optionally, for each branch network, the confidence level of the third sample detection result output by the branch network can be the confidence level of the target detection box predicted by the detection subnetwork in the branch network.
[0176] S603, based on the confidence level of the third sample detection results output by each branch network, determine the score of the verification sample image input to each branch network.
[0177] For each branch network, the average confidence score of each object detection box predicted by the detection subnetwork in that network is used as the score of that branch network for the input verification sample image.
[0178] S604, determine the threshold corresponding to each branch network based on the score of the verification sample image input to each branch network.
[0179] This application also designs an adaptive threshold finding method for object detection to validate the object detection model. The idea is to first perform forward inference on the validation set to obtain at least one object detection box and its confidence score for each validation sample image in each branch network, and then use the summation average of the confidence scores of the branch networks as the score of the validation sample image input to that branch network.
[0180] The scores of each validation sample image in each branch network are used to form a matrix M. Here, M is an m*n matrix, where m is the branch network number, n is the number of validation sample images, and each value in the matrix represents the score of each validation sample image corresponding to all branch networks.
[0181] After obtaining matrix M, it is necessary to sort the scores of matrix M on each branch network, and determine whether each verification sample image can pass through each branch network according to the sorting. The lowest score of the verification sample image that can pass through each branch network is used as the threshold corresponding to each branch network.
[0182] Understandably, the above-mentioned object detection model validation method involves forming a matrix from the scores on each sub-network and sorting the matrix. Then, it determines whether each validation sample image can pass through each sub-network according to the order of the scores on each sub-network in the matrix. Finally, it determines the threshold corresponding to each sub-network, which provides a foundation for subsequent object detection and also ensures the effectiveness and accuracy of the object detection model.
[0183] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0184] Based on the same inventive concept, this application also provides a target detection apparatus for implementing the target detection method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more target detection apparatus embodiments provided below can be found in the limitations of the target detection method described above, and will not be repeated here.
[0185] In one embodiment, such as Figure 7 As shown, a target detection device 7 is provided, comprising: a first determination module 70, a detection module 71, a scoring determination module 72, a second determination module 73, and an output module 74, wherein:
[0186] The first determining module 70 is used to take the first branch network of the target detection model as the target network.
[0187] The target detection model includes at least two branch networks arranged in order of accuracy, with the first branch network having the lowest accuracy.
[0188] The detection module 71 is used to input the image to be detected into the target network and obtain the target detection result and the confidence level of the target detection result output by the target network.
[0189] The scoring determination module 72 is used to determine the score of the image to be detected based on the confidence level of the target detection result.
[0190] The second determining module 73 is used to, if the score of the image to be detected is less than the threshold corresponding to the target network, take the next branch network of the target network as the new target network, return to execute the input of the image to be detected into the target network, and obtain the target detection result and the confidence level of the target detection result output by the target network, until the score of the image to be detected is greater than the threshold corresponding to the target network.
[0191] Output module 74 is used to take the target detection results output by the target network as the final detection result of the image to be detected.
[0192] In one embodiment, each branch network includes an initial feature extraction layer, a feature extraction subnetwork, and a detection subnetwork; the initial feature extraction layers of each branch network are connected sequentially according to the arrangement order of the branch networks; the number of downsampling layers included in the feature extraction subnetwork of each branch network decreases sequentially according to the arrangement order of the branch networks.
[0193] In one embodiment, such as Figure 8 As shown, above Figure 7 The detection module 71 in the middle may specifically include:
[0194] The first detection unit 711 is used to input the image to be detected or the previous basic feature map of the image to be detected into the initial feature extraction layer of the target network to obtain the current basic feature map of the image to be detected.
[0195] The previous basic feature map is output from the initial feature extraction layer in the previous branch network of the target network.
[0196] The second detection unit 712 is used to input the current basic feature map into the feature extraction subnetwork of the target network to obtain a deep feature map.
[0197] In this case, the depth feature maps output by different target networks have the same scale.
[0198] The third detection unit 713 is used to input the depth feature map into the detection subnetwork of the target network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0199] In one embodiment, each detection subnetwork includes a region recommendation network and a detection head network; such as Figure 9 As shown, above Figure 8 The third detection unit 713 may specifically include:
[0200] The first detection subunit 7131 is used to input the depth feature map into the region recommendation network to obtain at least one candidate box.
[0201] The second detection subunit 7132 is used to input at least one candidate box and a depth feature map into the detection head network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0202] In one embodiment, the target detection result includes at least one target detection box; the scoring determination module 72 is specifically used to: take the average confidence level of each target detection box as the score of the image to be detected.
[0203] In one embodiment, the target detection device 7 further includes:
[0204] The sample acquisition module is used to acquire a training set of sample images;
[0205] The detection result determination module is used to input the training sample images in the sample image training set into each branch network in the target detection model to obtain the first sample detection result output by the region recommendation network in each branch network and the second sample detection result output by the detection head network in each branch network.
[0206] The first loss determination module is used to determine the branch loss of each branch network based on the training sample labels of the training sample images and the first sample detection results and second sample detection results output by each branch network.
[0207] The second loss determination module is used to determine the total loss of the target detection model based on the branch loss of each branch network;
[0208] The training module is used to train the object detection model using the total loss.
[0209] In one embodiment, the first loss determination module is specifically used for:
[0210] For each branch network, a first loss is determined based on the training sample labels of the training sample images and the first sample detection result output by the region recommendation network in the branch network; a second loss is determined based on the training sample labels of the training sample images and the second sample detection result output by the detection head network in the branch network; and the branch loss of the branch network is determined based on the first loss and the second loss.
[0211] In one embodiment, the second loss determination module is specifically used for:
[0212] The total loss of the object detection model is determined based on the branch loss and weight of each branch network, as well as the number of samples in the training set of sample images.
[0213] In one embodiment, the target detection device 7 further includes a verification model, the verification module being specifically used for:
[0214] Obtain a sample image validation set; input the validation sample images from the sample image validation set into the target detection model to obtain the confidence level of the third sample detection result output by each branch network in the target detection model; determine the score of the validation sample image input to each branch network based on the confidence level of the third sample detection result output by each branch network; determine the threshold corresponding to each branch network based on the score of the validation sample image input to each branch network.
[0215] Each module in the aforementioned target detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0216] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows. Figure 10 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores target detection data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a target detection method.
[0217] Those skilled in the art will understand that Figure 10The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specifically, the computer device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0218] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0219] The first branch network of the target detection model is taken as the target network; wherein the target detection model includes at least two branch networks arranged in order of accuracy, and the first branch network has the lowest accuracy;
[0220] The image to be detected is input into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0221] The score of the image to be detected is determined based on the confidence level of the target detection results;
[0222] If the score of the image to be detected is less than the threshold corresponding to the target network, then the next branch network of the target network is used as the new target network, and the process is repeated to input the image to be detected into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network, until the score of the image to be detected is greater than the threshold corresponding to the target network.
[0223] The target detection results output by the target network are used as the final detection results for the image to be detected.
[0224] In one embodiment, each branch network includes an initial feature extraction layer, a feature extraction subnetwork, and a detection subnetwork; the initial feature extraction layers of each branch network are connected sequentially according to the arrangement order of the branch networks; the number of downsampling layers included in the feature extraction subnetwork of each branch network decreases sequentially according to the arrangement order of the branch networks.
[0225] In one embodiment, when the processor executes the logic in the computer program that inputs the image to be detected into the target network and obtains the target detection result and the confidence level of the target detection result output by the target network, the following steps are specifically implemented:
[0226] The image to be detected, or its previous basic feature map, is input into the initial feature extraction layer of the target network to obtain the current basic feature map of the image to be detected. The previous basic feature map is output by the initial feature extraction layer in the previous branch network of the target network. The current basic feature map is input into the feature extraction subnetwork of the target network to obtain the depth feature map. The depth feature maps output by different target networks have the same scale. The depth feature map is input into the detection subnetwork of the target network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0227] In one embodiment, each detection subnetwork includes a region recommendation network and a detection head network. When the processor executes the logic in the computer program that inputs the deep feature map into the detection subnetwork of the target network and obtains the target detection result and the confidence level of the target detection result output by the target network, the following steps are specifically implemented:
[0228] The deep feature map is input into the region recommendation network to obtain at least one candidate box; the at least one candidate box and the deep feature map are input into the detection head network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0229] In one embodiment, the target detection result includes at least one target detection box. When the processor executes the logic in the computer program that determines the score of the image to be detected based on the confidence level of the target detection result, the following steps are specifically implemented:
[0230] The average confidence score of each target detection box is used as the score of the image to be detected.
[0231] In one embodiment, when the processor executes the logic of the training process of the object detection model in the computer program, it specifically implements the following steps:
[0232] Obtain a training set of sample images; input the training sample images from the training set into each branch network of the object detection model to obtain the first sample detection result output by the region recommendation network in each branch network and the second sample detection result output by the detection head network in each branch network; determine the branch loss of each branch network based on the training sample labels of the training sample images and the first and second sample detection results output by each branch network; determine the total loss of the object detection model based on the branch loss of each branch network; train the object detection model using the total loss.
[0233] In one embodiment, when the processor executes the logic in the computer program to determine the branching loss of each branching network based on the training sample labels of the sample images and the first and second sample detection results output by each branching network, the following steps are specifically implemented:
[0234] For each branch network, a first loss is determined based on the training sample labels of the sample image and the first sample detection result output by the region recommendation network in that branch network; a second loss is determined based on the training sample labels of the sample image and the second sample detection result output by the detection head network in that branch network; and the branch loss of that branch network is determined based on the first loss and the second loss.
[0235] In one embodiment, when the processor executes the logic in the computer program that determines the total loss of the target detection model based on the branch loss of each branch network, it specifically implements the following steps:
[0236] The total loss of the object detection model is determined based on the branch loss and weight of each branch network, as well as the number of samples in the training set of sample images.
[0237] In one embodiment, when the processor executes the logic in the computer program after training the object detection model, it specifically implements the following steps:
[0238] Obtain a sample image validation set; input the validation sample images from the sample image validation set into the target detection model to obtain the confidence level of the third sample detection result output by each branch network in the target detection model; determine the score of the validation sample image input to each branch network based on the confidence level of the third sample detection result output by each branch network; determine the threshold corresponding to each branch network based on the score of the validation sample image input to each branch network.
[0239] The principles and specific processes of the computer equipment provided above in implementing the various embodiments can be found in the descriptions of the target detection method embodiments in the foregoing embodiments, and will not be repeated here.
[0240] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0241] The first branch network of the target detection model is taken as the target network; wherein the target detection model includes at least two branch networks arranged in order of accuracy, and the first branch network has the lowest accuracy;
[0242] The image to be detected is input into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0243] The score of the image to be detected is determined based on the confidence level of the target detection results;
[0244] If the score of the image to be detected is less than the threshold corresponding to the target network, then the next branch network of the target network is used as the new target network, and the process is repeated to input the image to be detected into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network, until the score of the image to be detected is greater than the threshold corresponding to the target network.
[0245] The target detection results output by the target network are used as the final detection results for the image to be detected.
[0246] In one embodiment, each branch network includes an initial feature extraction layer, a feature extraction subnetwork, and a detection subnetwork; the initial feature extraction layers of each branch network are connected sequentially according to the arrangement order of the branch networks; the number of downsampling layers included in the feature extraction subnetwork of each branch network decreases sequentially according to the arrangement order of the branch networks.
[0247] In one embodiment, when the logic in the computer program that inputs the image to be detected into the target network and obtains the target detection result and the confidence level of the target detection result output by the target network is executed by the processor, the following steps are specifically implemented:
[0248] The image to be detected, or its previous basic feature map, is input into the initial feature extraction layer of the target network to obtain the current basic feature map of the image to be detected. The previous basic feature map is output by the initial feature extraction layer in the previous branch network of the target network. The current basic feature map is input into the feature extraction subnetwork of the target network to obtain the depth feature map. The depth feature maps output by different target networks have the same scale. The depth feature map is input into the detection subnetwork of the target network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0249] In one embodiment, each detection subnetwork includes a region recommendation network and a detection head network. When the logic in the computer program that inputs the deep feature map into the detection subnetwork of the target network to obtain the target detection result and the confidence level of the target detection result output by the target network is executed by the processor, the following steps are specifically implemented:
[0250] The deep feature map is input into the region recommendation network to obtain at least one candidate box; the at least one candidate box and the deep feature map are input into the detection head network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0251] In one embodiment, the target detection result includes at least one target detection box. When the logic in the computer program that determines the score of the image to be detected based on the confidence level of the target detection result is executed by the processor, the following steps are specifically implemented:
[0252] The average confidence score of each target detection box is used as the score of the image to be detected.
[0253] In one embodiment, when the logic of the training process of the object detection model in the computer program is executed by the processor, the following steps are specifically implemented:
[0254] Obtain a training set of sample images; input the training sample images from the training set into each branch network of the object detection model to obtain the first sample detection result output by the region recommendation network in each branch network and the second sample detection result output by the detection head network in each branch network; determine the branch loss of each branch network based on the training sample labels of the training sample images and the first and second sample detection results output by each branch network; determine the total loss of the object detection model based on the branch loss of each branch network; train the object detection model using the total loss.
[0255] In one embodiment, when the logic in the computer program that determines the branch loss of each branch network based on the training sample labels of the sample images and the first and second sample detection results output by each branch network is executed by the processor, the following steps are specifically implemented:
[0256] For each branch network, a first loss is determined based on the training sample labels of the sample image and the first sample detection result output by the region recommendation network in that branch network; a second loss is determined based on the training sample labels of the sample image and the second sample detection result output by the detection head network in that branch network; and the branch loss of that branch network is determined based on the first loss and the second loss.
[0257] In one embodiment, when the logic in the computer program that determines the total loss of the target detection model based on the branch loss of each branch network is executed by the processor, the following steps are specifically implemented:
[0258] The total loss of the object detection model is determined based on the branch loss and weight of each branch network, as well as the number of samples in the training set of sample images.
[0259] In one embodiment, when the logic of training the object detection model in the computer program is executed by the processor, the following steps are specifically implemented:
[0260] Obtain a sample image validation set; input the validation sample images from the sample image validation set into the target detection model to obtain the confidence level of the third sample detection result output by each branch network in the target detection model; determine the score of the validation sample image input to each branch network based on the confidence level of the third sample detection result output by each branch network; determine the threshold corresponding to each branch network based on the score of the validation sample image input to each branch network.
[0261] The principles and specific processes of the computer-readable storage medium provided above in implementing the various embodiments can be found in the descriptions of the target detection method embodiments in the foregoing embodiments, and will not be repeated here.
[0262] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:
[0263] The first branch network of the target detection model is taken as the target network; wherein the target detection model includes at least two branch networks arranged in order of accuracy, and the first branch network has the lowest accuracy;
[0264] The image to be detected is input into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0265] The score of the image to be detected is determined based on the confidence level of the target detection results;
[0266] If the score of the image to be detected is less than the threshold corresponding to the target network, then the next branch network of the target network is used as the new target network, and the process is repeated to input the image to be detected into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network, until the score of the image to be detected is greater than the threshold corresponding to the target network.
[0267] The target detection results output by the target network are used as the final detection results for the image to be detected.
[0268] In one embodiment, each branch network includes an initial feature extraction layer, a feature extraction subnetwork, and a detection subnetwork; the initial feature extraction layers of each branch network are connected sequentially according to the arrangement order of the branch networks; the number of downsampling layers included in the feature extraction subnetwork of each branch network decreases sequentially according to the arrangement order of the branch networks.
[0269] When the logic in a computer program that inputs the image to be detected into the target network and obtains the target detection result and the confidence level of the target detection result from the target network is executed by the processor, the following steps are specifically implemented:
[0270] The image to be detected, or its previous basic feature map, is input into the initial feature extraction layer of the target network to obtain the current basic feature map of the image to be detected. The previous basic feature map is output by the initial feature extraction layer in the previous branch network of the target network. The current basic feature map is input into the feature extraction subnetwork of the target network to obtain the depth feature map. The depth feature maps output by different target networks have the same scale. The depth feature map is input into the detection subnetwork of the target network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0271] In one embodiment, each detection subnetwork includes a region recommendation network and a detection head network. When the logic in the computer program that inputs the deep feature map into the detection subnetwork of the target network to obtain the target detection result and the confidence level of the target detection result output by the target network is executed by the processor, the following steps are specifically implemented:
[0272] The deep feature map is input into the region recommendation network to obtain at least one candidate box; the at least one candidate box and the deep feature map are input into the detection head network to obtain the target detection result and the confidence level of the target detection result output by the target network.
[0273] In one embodiment, the target detection result includes at least one target detection box. When the logic in the computer program that determines the score of the image to be detected based on the confidence level of the target detection result is executed by the processor, the following steps are specifically implemented:
[0274] The average confidence score of each target detection box is used as the score of the image to be detected.
[0275] In one embodiment, when the logic of the training process of the object detection model in the computer program is executed by the processor, the following steps are specifically implemented:
[0276] Obtain a training set of sample images; input the training sample images from the training set into each branch network of the object detection model to obtain the first sample detection result output by the region recommendation network in each branch network and the second sample detection result output by the detection head network in each branch network; determine the branch loss of each branch network based on the training sample labels of the training sample images and the first and second sample detection results output by each branch network; determine the total loss of the object detection model based on the branch loss of each branch network; train the object detection model using the total loss.
[0277] In one embodiment, when the logic in the computer program that determines the branch loss of each branch network based on the training sample labels of the sample images and the first and second sample detection results output by each branch network is executed by the processor, the following steps are specifically implemented:
[0278] For each branch network, a first loss is determined based on the training sample labels of the sample image and the first sample detection result output by the region recommendation network in that branch network; a second loss is determined based on the training sample labels of the sample image and the second sample detection result output by the detection head network in that branch network; and the branch loss of that branch network is determined based on the first loss and the second loss.
[0279] In one embodiment, when the logic in the computer program that determines the total loss of the target detection model based on the branch loss of each branch network is executed by the processor, the following steps are specifically implemented:
[0280] The total loss of the object detection model is determined based on the branch loss and weight of each branch network, as well as the number of samples in the training set of sample images.
[0281] In one embodiment, when the logic of training the object detection model in the computer program is executed by the processor, the following steps are specifically implemented:
[0282] Obtain a sample image validation set; input the validation sample images from the sample image validation set into the target detection model to obtain the confidence level of the third sample detection result output by each branch network in the target detection model; determine the score of the validation sample image input to each branch network based on the confidence level of the third sample detection result output by each branch network; determine the threshold corresponding to each branch network based on the score of the validation sample image input to each branch network.
[0283] The principles and specific processes of implementing the computer program products provided above can be found in the descriptions of the target detection method embodiments in the foregoing embodiments, and will not be repeated here.
[0284] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0285] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0286] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A target detection method, characterized in that, include: The first branch network of the target detection model is taken as the target network; wherein, the target detection model includes at least two branch networks arranged in order of accuracy, and the first branch network has the lowest accuracy; each branch network includes an initial feature extraction layer, a feature extraction sub-network, and a detection sub-network; the initial feature extraction layers of each branch network are connected sequentially according to the arrangement order of the branch networks; the number of downsampling layers included in the feature extraction sub-network of each branch network decreases sequentially according to the arrangement order of the branch networks, and the number of downsampling layers are 3, 2, 1, and 0 respectively; each branch network in the target detection model corresponds to a different threshold; The image to be detected or its previous basic feature map is input into the initial feature extraction layer of the target network to obtain the current basic feature map of the image to be detected; wherein the previous basic feature map is output by the initial feature extraction layer in the previous branch network of the target network. The current basic feature map is input into the feature extraction subnetwork of the target network to obtain a depth feature map; wherein, the depth feature maps output by different target networks have the same scale; The deep feature map is input into the detection subnetwork of the target network to obtain the target detection result output by the target network and the confidence level of the target detection result; the target detection result includes at least one target detection box; The average confidence score of each target detection box is used as the score of the image to be detected. If the score of the image to be detected is less than the threshold corresponding to the target network, then the next branch network of the target network is used as the new target network, and the process is repeated to input the image to be detected into the target network to obtain the target detection result and the confidence level of the target detection result output by the target network, until the score of the image to be detected is greater than the threshold corresponding to the target network. The target detection result output by the target network is used as the final detection result for the image to be detected; The threshold values corresponding to each branch network are determined in the following way: Obtain a sample image validation set; input the validation sample images in the sample image validation set into the target detection model to obtain the confidence of the third sample detection result output by each branch network in the target detection model; Based on the confidence level of the third sample detection results output by each branch network, the score of the verification sample image input to each branch network is determined; The scores of each verification sample image in each branch network are used to form a matrix M; The scores of the matrix M on each branch network are sorted, and each verification sample image is judged according to the sorting to see if it can pass through each branch network. The lowest score of the verification sample image that can pass through each branch network is taken as the threshold corresponding to each branch network.
2. The method according to claim 1, characterized in that, Each detection subnetwork includes a region recommendation network and a detection head network; The step of inputting the depth feature map into the detection subnetwork of the target network to obtain the target detection result output by the target network and the confidence level of the target detection result includes: The deep feature map is input into the region recommendation network to obtain at least one candidate box; The at least one candidate box and the depth feature map are input into the detection head network to obtain the target detection result output by the target network and the confidence level of the target detection result.
3. The method according to claim 2, characterized in that, The training process of the target detection model is as follows: Obtain the training set of sample images; The training sample images in the sample image training set are input into each branch network in the target detection model to obtain the first sample detection result output by the region recommendation network in each branch network and the second sample detection result output by the detection head network in each branch network. Based on the training sample labels of the training sample images, and the first sample detection results and second sample detection results output by each branch network, the branch loss of each branch network is determined. The total loss of the target detection model is determined based on the branch loss of each branch network. The target detection model is trained using the total loss.
4. The method according to claim 3, characterized in that, The step of determining the splitting loss of each splitting network based on the training sample labels of the training sample images and the first sample detection result and the second sample detection result output by each splitting network includes: For each branch network, a first loss is determined based on the training sample labels of the training sample images and the first sample detection result output by the region recommendation network in that branch network; The second loss is determined based on the training sample labels of the training sample images and the second sample detection results output by the detection head network in the split network; Based on the first loss and the second loss, the routing loss of the routing network is determined.
5. The method according to claim 4, characterized in that, The step of determining the total loss of the target detection model based on the branch loss of each branch network includes: The total loss of the target detection model is determined based on the branch loss and weight of each branch network and the number of samples in the training set of the sample images.
6. A target detection device, characterized in that, The target detection device is used to implement the target detection method according to any one of claims 1-5, the device comprising: The first determining module is used to take the first branch network of the target detection model as the target network; wherein, the target detection model includes at least two branch networks arranged in order of accuracy, and the first branch network has the lowest accuracy; The detection module is used to input the image to be detected into the target network and obtain the target detection result output by the target network and the confidence level of the target detection result; The scoring determination module is used to determine the score of the image to be detected based on the confidence level of the target detection result; The second determining module is used to, if the score of the image to be detected is less than the threshold corresponding to the target network, take the next branch network of the target network as the new target network, return to execute the input of the image to be detected to the target network, obtain the target detection result output by the target network and the confidence of the target detection result, until the score of the image to be detected is greater than the threshold corresponding to the target network; The output module is used to take the target detection result output by the target network as the final detection result of the image to be detected.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Safety helmet wearing detection method, device and equipment and storage medium
CN113177513A