Method and device for training target detection network
By using the fusion technology of different sample distribution characteristics and candidate area characteristics of multiple image sample sets, the problem of low accuracy of the target detection network under weak supervision technology is solved, and higher detection accuracy is achieved.
Patent Information
- Application Number
- CN202311553750.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-20
AI Technical Summary
During the training process, the target detection network based on weak supervision technology trained, lacks accurate detection accuracy and cannot accurately predict the area and category of the target object.
By acquiring multiple image sample sets, each sample set has different sample distribution characteristics. The overall feature extraction network, the distribution feature extraction network and the candidate area feature network are used to fuse the sample distribution features and candidate area features, optimize the candidate area features, and input it to the detection head network for prediction.
By reducing the interference of sample distribution characteristics on candidate region characteristics, the accuracy of the target detection network in region and category prediction is improved, thereby improving the accuracy of the entire detection network.
Smart Images

Figure CN120020903A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet technologies, and particularly to a method and apparatus for training an object detection network, an object detection method and apparatus, a computing device, a computer-readable storage medium, and a computer program product. Background Art
[0002] At present, the technology of neural networks trained based on weak supervision technology has been booming. For an object detection network that detects image data, because weak supervision technology is adopted, it is no longer necessary to calibrate the region where the target object is located for each image sample before training, which saves a lot of labor. However, during the training process, since the accurate region where the target object is located is not pre-calibrated, the trained object detection network often has certain errors and cannot accurately predict the region and category where the target object is located. How to improve the accuracy of the object detection network is an urgent problem to be solved. Summary of the Invention
[0003] In view of this, the present application provides a method and apparatus for training an object detection network, an object detection method and apparatus, a computing device, a computer-readable storage medium, and a computer program product, in the hope of alleviating or overcoming some or all of the above-mentioned defects and other possible defects.
[0004] According to a first aspect of the present application, a method for training an object detection network is provided, including: obtaining a plurality of image sample sets, each image sample in each image sample set includes at least one target object, and each target object has a corresponding class label. Each image sample set among the plurality of image sample sets has at least one sample distribution feature different from other image sample sets; for each image sample in each image sample set among the plurality of image sample sets, perform the following steps: input the image sample into the overall feature extraction network in the object detection network to obtain the overall feature of the image sample; input the overall feature of the image sample into the distribution feature extraction network in the object detection network to obtain at least one sample distribution feature of the image sample set corresponding to the image sample; input the overall feature of the image sample into the candidate region feature network in the object detection network to determine the features of at least one candidate region where the at least one target object is located in the image sample, and the at least one candidate region corresponds to the at least one target object one by one; fuse the features of each candidate region of the image sample with the at least one sample distribution feature of the image sample set corresponding to the image sample to obtain the optimized features of each candidate region of the image sample; input the optimized features of each candidate region of the image sample into the detection head network of the object detection network to obtain the predicted classes and the predicted regions where each target object is located in the image sample; adjust the parameters of the object detection network according to the class labels, the predicted classes, and the predicted regions of each target object in each image sample in each image sample set among the plurality of image sample sets.
[0005] In some embodiments according to the present application, the distribution feature extraction network includes a pooling layer, a preset number of fully connected layers, an activation function layer, and a distribution feature layer. The distribution feature layer is a matrix composed of a plurality of one-dimensional vectors, and the step of inputting the overall feature of the image sample into the distribution feature extraction network in the object detection network to obtain at least one sample distribution feature of the image sample set corresponding to the image sample includes: inputting the overall feature of the image sample into the pooling layer to obtain the overall feature after pooling; inputting the overall feature after pooling into the preset number of fully connected layers to obtain the distribution feature coefficients of the image sample; inputting the distribution feature coefficients of the image sample into the activation function layer to obtain the optimized distribution feature coefficients; inputting the optimized distribution feature coefficients into the distribution feature layer to obtain at least one sample distribution feature of the image sample set corresponding to the image sample.
[0006] In some embodiments according to the present application, the pooling layer is one of a max pooling layer and an average pooling layer.
[0007] In some embodiments according to the present application, the activation function layer is one of a hyperbolic tangent (Tanh) function layer and a rectified linear unit (ReLU) function layer.
[0008] In some embodiments according to the present application, the step of inputting the optimized distribution feature coefficients into the distribution feature layer to obtain at least one sample distribution feature of the image sample set corresponding to the image sample includes: multiplying the optimized distribution feature coefficients by each one-dimensional vector of the distribution feature layer to obtain a two-dimensional matrix; and determining the two-dimensional matrix as at least one sample distribution feature of the image sample set corresponding to the image sample.
[0009] In some embodiments according to the present application, the candidate region feature network includes a candidate region determination network and a candidate region global feature extraction network, and the step of inputting the global feature of the image sample into the candidate region feature network in the object detection network to determine the features of at least one candidate region where at least one target object is located in the image sample includes: inputting the image sample into the candidate region determination network to obtain at least one candidate region of the image sample; and inputting at least one candidate region of the image sample and the global feature of the image sample into the candidate region global feature extraction network to obtain the features of each candidate region of the image sample.
[0010] In some embodiments according to the present application, at least one sample distribution feature of the image sample set corresponding to the image sample is respectively represented by at least one sample distribution vector, and the step of fusing the features of each candidate region of the image sample with the at least one sample distribution feature of the image sample set corresponding to the image sample to obtain the optimized features of each candidate region of the image sample includes: for the features of each candidate region of the image sample, performing the following operations: performing an addition operation on the features of the candidate region and at least one sample distribution vector of the image sample set corresponding to the image sample to obtain a first fusion result; performing a subtraction operation on the features of the candidate region and at least one sample distribution vector of the image sample set corresponding to the image sample to obtain a second fusion result; performing a multiplication operation on the features of the candidate region and at least one sample distribution vector of the image sample set corresponding to the image sample to obtain a third fusion result; performing a division operation on the features of the candidate region and at least one sample distribution vector of the image sample set corresponding to the image sample to obtain a fourth fusion result; and determining the optimized features of the candidate region of the image sample according to at least one of the first fusion result, the second fusion result, the third fusion result, and the fourth fusion result.
[0011] According to a second aspect of the present application, there is provided a method for detecting a target object, including: obtaining an image to be detected, where the image to be detected includes at least one target object to be detected; inputting the image to be detected into the target detection network according to claim 1 to obtain the category and location area of the at least one target object in the image to be detected.
[0012] According to a third aspect of the present application, there is provided an apparatus for training a target detection network, including: a processing module configured to obtain a plurality of image sample sets, each image sample in each image sample set includes at least one target object, and each target object has a corresponding category label, and each image sample set in the plurality of image sample sets has at least one sample distribution feature different from other image sample sets; for each image sample in each image sample set in the plurality of image sample sets, perform the following steps: input the image sample into the overall feature extraction network in the target detection network to obtain the overall feature of the image sample; input the overall feature of the image sample into the distribution feature extraction network in the target detection network to obtain at least one sample distribution feature of the image sample set corresponding to the image sample; input the overall feature of the image sample into the candidate region feature network in the target detection network to determine the features of at least one candidate region where the at least one target object in the image sample is located, and the at least one candidate region corresponds to the at least one target object one by one; fuse the features of each candidate region of the image sample with the at least one sample distribution feature of the image sample set corresponding to the image sample to obtain the optimized features of each candidate region of the image sample; input the optimized features of each candidate region of the image sample into the detection head network of the target detection network to obtain the predicted category and the predicted location area of each target object in the image sample; an adjustment module configured to adjust the parameters of the target detection network according to the category label of each target object in each image sample in each image sample set in the plurality of image sample sets and the predicted category and the predicted location area of the target object.
[0013] According to a fourth aspect of the present application, there is provided a target object detection apparatus, including: an acquisition module configured to acquire an image to be detected, where the image to be detected includes at least one target object to be detected; a detection module configured to input the image to be detected into the target detection network according to claim 1 to obtain the category and location area of the at least one target object in the image to be detected.
[0014] According to a fifth aspect of the present application, there is provided a computing device including a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, it causes the processor to execute the steps of a method for training an object detection network according to some embodiments of the present application.
[0015] According to a sixth aspect of the present application, there is provided a computer-readable storage medium having computer-readable instructions stored thereon, and when the computer-readable instructions are executed, they implement a method for training an object detection network according to some embodiments of the present application.
[0016] According to a seventh aspect of the present application, there is provided a computer program product including computer instructions, and when the computer instructions are executed by a processor, they implement a method for training an object detection network according to some embodiments of the present application.
[0017] In the method and apparatus for training an object detection network according to some embodiments of the present application, multiple image sample sets are used to train the object detection network. Since each image sample set has different sample distribution characteristics, different sample distribution characteristics can be provided for the training process. The distribution feature extraction network in the object detection network can extract the sample distribution characteristics from the overall features of the image samples and fuse the features of the candidate regions with the sample distribution characteristics, thereby reducing the influence brought by the sample distribution characteristics in the features of each candidate region. In this way, interference factors can be excluded at each candidate region, and more accurate optimized features can be obtained, which is beneficial to improving the accuracy of the region for detecting the target object. The improvement of the accuracy of the region will also lead to the improvement of the accuracy of the category. Therefore, the trained object detection network has higher accuracy.
[0018] According to the embodiments described below, these and other advantages of the present application will become clear, and these and other advantages of the present application are illustrated with reference to the embodiments described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Embodiments of the present application will now be described in more detail and with reference to the drawings, wherein:
[0020] Figure 1 is a schematic diagram of a common method for training an object detection network;
[0021] Figure 2 is a schematic diagram of the prediction result of an object detection network trained based on weak supervision technology;
[0022] Figure 3 is a schematic diagram of an example scenario of a method for training an object detection network according to some embodiments of the present application;
[0023] Figure 4A flowchart of a method for training an object detection network according to some embodiments of the present application;
[0024] Figure 5 A schematic diagram of a distribution feature extraction network according to some embodiments of the present application;
[0025] Figure 6 A flowchart of extracting distribution features according to some embodiments of the present application;
[0026] Figure 7 A schematic diagram of a training object detection network according to some embodiments of the present application;
[0027] Figure 8 A flowchart of a method for detecting an object target according to some embodiments of the present application;
[0028] Figure 9 An exemplary structural block diagram of an apparatus for training an object detection network according to some embodiments of the present application;
[0029] Figure 10 An exemplary structural block diagram of an object target detection apparatus according to some embodiments of the present application;
[0030] Figure 11 An example system is shown, which includes an example computing device representing one or more systems and / or devices that can implement the various methods described herein. Detailed Description of Specific Embodiments
[0031] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar parts, and thus their repeated description will be omitted.
[0032] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0033] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0034] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor do they have to be executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0035] It should be understood that although terms such as first, second, and third may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Thus, the first component discussed below may be referred to as the second component without departing from the teachings of the concept of the present application. As used herein, the term "and / or" and similar terms include any, multiple, and all combinations of the associated listed items.
[0036] Those skilled in the art can understand that the drawings are only schematic diagrams of exemplary embodiments, and the modules or processes in the drawings are not necessarily essential for implementing the present application, so they cannot be used to limit the protection scope of the present application.
[0037] Before introducing the embodiments of the present application in detail, for the sake of clarity, some related concepts are first explained.
[0038] The present application relates to artificial intelligence and machine learning. The related technologies are briefly introduced below:
[0039] Artificial Intelligence (AI): So-called AI is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0040] AI technology is a comprehensive discipline that covers a wide range of fields, including both hardware and software technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, processing technologies for large application programs, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0041] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. In this application, it mainly involves the training of neural networks for processing image data. More specifically, it involves object detection networks trained based on weak supervision techniques.
[0042] Object detection network trained based on weak supervision techniques: The detection task of this object detection network is to train an object detection network with only class labels, which is different from the fully supervised object detection that requires instance-level (instance level, which requires annotating the center coordinates, height, and width of the maximum bounding rectangle of the target object in the image data) labels. Annotating instance-level information requires a large amount of manpower, material resources, and financial resources.
[0043] Convolutional Neural Networks (CNN): It is a type of feed-forward neural network (FNN) that contains convolutional calculations and has a deep structure, and is one of the representative algorithms of deep learning (DeepLearning). Convolutional neural networks have the ability of representation learning and can perform shift-invariant classification on the input images according to their hierarchical structures.
[0044] Convolutional Layer: Each convolutional layer in a convolutional neural network consists of several convolutional units, and the parameters of each convolutional unit are optimized through the backpropagation algorithm. The purpose of the convolution operation is to extract different features of the input. The first convolutional layer may only be able to extract some low-level features such as edges, lines, and corners. More layers of the network can iteratively extract more complex features from the low-level features.
[0045] Pooling Layer: The network layer that performs pooling processing on the feature map using a filter. Pooling processing specifically aggregates and statistically analyzes a pixel point with the data points around it, reduces the size of the feature map, and then takes the mean or maximum value of its adjacent regions to further reduce the number of parameters. The feature map here can be the overall feature of the image sample mentioned in this application.
[0046] Fully Connected Layer: Each node in the fully connected layer is connected to all nodes in the previous layer, and is used to synthesize the features extracted previously.
[0047] PASCAL VOC 2007: Abbreviated as VOC07, this image sample set was initially a project initiated by the European Conference on Computer Vision, mainly used for object detection, image classification, and semantic segmentation tasks. The PASCAL VOC 2007 image sample set has a total of 9963 images, including 5011 images in the training set and validation set, and 4952 images in the test set, covering 20 categories.
[0048] PASCAL VOC 2012: Abbreviated as VOC12, this image sample set was created by the Visual Object Classes (VOC) team to promote the development of computer vision research. This image sample set contains a series of image samples in real scenes, and the objects in them are labeled with different categories, such as people, cars, cats, etc.
[0049] Microsoft COCO: Abbreviated as COCO, this image sample set is an image sample set released by Microsoft in 2014 that can be used for image recognition.
[0050] mean Average Precision (mAP): For multiple target objects, after predicting the category of each target object, the consistency between the predicted category and the true category of each target object is determined. The average value of the consistencies of each target object is the mean average precision. The larger the mAP value, the more accurate the object detection network. In this application, this indicator is used to evaluate the performance of the object detection network.
[0051] Correct Localization (CorLoC): For multiple target objects, after predicting the regions of each target object, the consistency between the predicted region and the true region of each target object is determined, and the average value of the consistencies of each target object is the correct localization. The larger the CorLoC value, the more accurate the target detection network. In this application, this metric is used to evaluate the performance of the target detection network.
[0052] First, a common method for training a target detection network is introduced. Figure 1 It is a schematic diagram of a common method for training a target detection network. In Figure 1 the illustrated embodiment, the target detection network 102 is composed of 4 sub-networks, namely the overall feature extraction network 1021, the candidate region determination network 1022, the candidate region overall feature extraction network 1023, and the detection head network 1024 of the target detection network. Generally, an image sample set 101 is used to train the target detection network.
[0053] In one training process, an image sample is obtained from the image sample set 101, and the image sample is input into the overall feature extraction network 1021 and the candidate region determination network 1022 respectively, and the overall feature of the image sample and each candidate region are calculated respectively. Then, the overall feature of the image sample and each candidate region are input into the candidate region overall feature extraction network 1023, and the region overall feature extraction network 1023 determines the features of each candidate region based on the overall feature of the image sample. Finally, the features of each candidate region are provided to the detection head network 1024 of the target detection network to predict the category and location area of the target object of the image sample.
[0054] Figure 2 It is a schematic diagram of the prediction result of a target detection network trained based on weak supervision technology. For easy understanding, the specific prediction effect of the target detection network trained based on weak supervision technology can be referred to Figure 2 .
[0055] The image sample 210 includes the first target object "bird" 220 and the second target object "car" 230. The area where the bird is located is the box 221, and the area where the car is located is the box 231. During the training process, the label of the image sample 210 only includes the category labels "bird" 220 and "car" 230, and does not include the label for the area. After inputting the image sample 210 into the target detection network 260 trained based on weak supervision technology, in the case of accurate prediction, the obtained results are the categories "bird" 240 and "car" 250, and the area 221 where the "bird" 220 is located and the area 231 where the "car" 230 is located.
[0056] However, in such asFigure 1 In the target detection network shown, the image samples in the image sample set often have certain sample distribution characteristics. For example, they are all images with a daytime background. This causes factors such as daytime to be considered when identifying the categories of target objects during the training process, and daytime is also considered when determining the regions where the target objects are located. That is, the sample distribution characteristics interfere with the prediction accuracy.
[0057] In view of this, the present application proposes a target detection network and its training method that can exclude the interference of sample distribution characteristics.
[0058] Figure 3 It is a schematic diagram of an example scenario 300 of a method for training a target detection network according to some embodiments of the present application. The scenario 300 may include a client 301, a network 302, and a service unit 303. The service unit 303 is communicatively coupled to the client 301 through the network 302. The network 302 can be, for example, a wide area network (WAN), a local area network (LAN), a wireless network, a public telephone network, an intranet, and any other type of network well-known to those skilled in the art.
[0059] In this embodiment, the method for training the target detection network is executed by the service unit 303. The method executed by the service unit 303 includes the following steps.
[0060] The service unit 303 performs the following steps for each image sample in each image sample set among multiple image sample sets.
[0061] First, the service unit 303 obtains an image sample from the client 301, where the image sample includes at least one target object, and each target object has a corresponding category label. In this embodiment, the service unit 303 can pre-store the target detection network to be trained, and apply the content of this step after obtaining the image sample set including the image sample from the client 301.
[0062] Second, the service unit 303 inputs the image sample into the overall feature extraction network in the target detection network to obtain the overall feature of the image sample. These computing tasks can be implemented by the processor and memory of the service unit 303.
[0063] Third, the service unit 303 inputs the overall feature of the image sample into the distribution feature extraction network in the target detection network to obtain the sample distribution feature of the image sample set corresponding to the image sample.
[0064] Then, the service unit 303 inputs the overall feature of the image sample into the candidate region feature network in the target detection network to determine the features of each candidate region of the image sample, where each candidate region indicates the region where each target object is located obtained through the calculation of the candidate region feature network.
[0065] Next, the service unit 303 fuses the features of each candidate region of the image sample with at least one sample distribution feature of the image sample set corresponding to the image sample to obtain the optimized features of each candidate region of the image sample.
[0066] Next, the service unit 303 inputs the optimized features of each candidate region of the image sample into the detection head network of the object detection network to predict the categories and regions where each target object is located in the image sample. The final prediction result can also be transmitted to the client 301 through the network 302 for display, so that the trainer can view the training effect in real time.
[0067] Finally, the service unit 303 adjusts the parameters of the object detection network according to the category labels of each target object in the image sample and the categories and regions predicted by the object detection network for each target object. Here, the service unit 303 completes one training of the object detection network. In order to achieve a better effect, the above steps can be repeatedly executed according to specific convergence conditions.
[0068] In Figure 3 In the scenario 300 shown, the method for training the object detection network according to some embodiments of the present application is implemented on the service unit 303, but this is only illustrative and not restrictive. The method for determining the image classification network according to some embodiments of the present application can also be implemented on other entities with sufficient computing resources and computing capabilities, such as on the client 301 with sufficient computing resources and computing capabilities. Of course, it can also be partially implemented on the service unit 303 and partially implemented on the client 301, which is not restrictive.
[0069] As understood by those of ordinary skill in the art, an instance of the service unit 303 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The servers can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this here.
[0070] The client 301 can be any type of mobile computing device, including mobile computers (e.g., personal digital assistants (PDAs), laptop computers, notebook computers, tablet computers, netbooks, etc.), mobile phones (e.g., cellular phones, smartphones, etc.), wearable computing devices (e.g., smartwatches, head-mounted devices, including smart glasses, etc.), or other types of mobile devices. In some embodiments, the client 301 can also be a fixed computing device, such as a desktop computer, a gaming console, a smart TV, etc.
[0071] As Figure 3 shown, the client 301 can include a display screen and a terminal application that can interact with the end user via the display screen. The terminal application can be a native application, a web application, or a mini program (Lite App, such as a mobile mini program, a WeChat mini program) as a lightweight application. In the case where the terminal application is a native application that needs to be installed, the terminal application can be installed in the client 301. In the case where the terminal application is a web application, the terminal application can be accessed through a browser. In the case where the terminal application is a mini program, the terminal application can be directly opened on the client 301 by searching for relevant information of the terminal application (such as the name of the terminal application), scanning the graphic code of the terminal application (such as a barcode, a QR code, etc.), etc., without installing the terminal application.
[0072] Figure 4 is a flowchart of a method for training an object detection network according to some embodiments of the present application. The method 400 shown can be implemented on Figure 3 the service unit 303 shown. In some embodiments, when Figure 3 the client 301 shown has sufficient computing resources and computing capabilities, the method for training an object detection network according to some embodiments of the present application can be directly executed on the client 301. In other embodiments, the method for training an object detection network according to some embodiments of the present application can also be executed in combination by the service unit 303 and the client 301. As Figure 4 shown, for each image sample in each image sample set of a plurality of image sample sets, the method for training an object detection network according to some embodiments of the present application performs steps S401 - S407.
[0073] In step S401, multiple image sample sets are obtained. Each image sample in each image sample set includes at least one target object, and each target object has a corresponding category label. Each image sample set among the multiple image sample sets has at least one sample distribution feature different from that of other image sample sets. The image samples can be any type of image data. For example, they can be taken photos or a frame of video data. The target objects in the image samples can be any type of target objects, such as animals, plants, objects, etc.
[0074] The sample distribution feature is a feature common to all image samples in an image sample set. For example, if the sample distribution feature of the first image sample set is that the background is daytime, it means that each image sample in the first image sample set is an image sample with daytime as the background. However, the sample distribution feature does not only refer to the background feature of the image sample but can also be any other feature. For another example, if the sample distribution feature of the second image sample set is that the grayscale of the image samples is less than a specific value, it means that the grayscale of each image sample in the second image sample set is less than this specific value. In the case where an image sample set includes multiple sample distribution features, all image samples in this image sample set have these multiple sample distribution features.
[0075] In the prior art, usually, the training of the target detection network is completed using a certain image sample set, resulting in the sample distribution feature of this image sample set affecting the prediction behavior of each sub-network of the target detection network, causing interference. Even if multiple image sample sets are used for training, since the existing target detection network cannot perform different processing according to the sample distribution features of different image sample sets, during the training process, the sample distribution features of each image sample set still subtly have a confounding negative impact on the prediction results of the target detection network.
[0076] In the embodiments of the present application, in order to enable the target detection network to learn the ability to extract sample distribution features, multiple image sample sets are provided. Each image sample set among the multiple image sample sets has at least one sample distribution feature different from that of other image sample sets, so as to provide different sample distribution features. For example, the image sample set A for training can include sample distribution features A1 - A10, the image sample set B can include sample distribution features A1 - A5 and B6 - B10, and the image sample set C can include sample distribution features C1 - C10. That is, there can be the same sample distribution features among the image sample sets, or there can be no same sample distribution features, but there must be different sample distribution features. This helps the target detection network master the ability to extract sample distribution features.
[0077] For each image sample in each image sample set among the multiple image sample sets, steps S402 - S407 are executed.
[0078] In step S402, the image sample is input into the overall feature extraction network in the object detection network to obtain the overall feature of the image sample. In this step, the overall feature extraction network can be, for example, a convolutional neural network composed of several convolutional layers, which can accurately extract the features of the image data. During the convolution calculation process, the convolution kernel slides while performing convolution operations on the image sample, thereby obtaining the overall feature of the image sample. This overall feature includes both the sample distribution feature, the feature related to the category of the target object, and the feature related to the region where the target object is located.
[0079] In step S403, the overall feature of the image sample is input into the distribution feature extraction network in the object detection network to obtain at least one sample distribution feature of the image sample set corresponding to the image sample. In this step, the distribution feature extraction network is used to extract at least one sample distribution feature of the image sample set corresponding to the image sample, realizing the purification of the sample distribution feature. Only after determining the sample distribution feature is it possible to reduce the negative impact of the sample distribution feature on the prediction behavior.
[0080] In step S404, the overall feature of the image sample is input into the candidate region feature network in the object detection network to determine the features of at least one candidate region where at least one target object is located in the image sample, and at least one candidate region corresponds to at least one target object one by one. Since the object detection network in this embodiment is a neural network trained based on weak supervision technology, the regions where the respective target objects are located are not calibrated in the image sample, so the features at the regions where the respective target objects are located cannot be obtained. Therefore, it is necessary to predict the regions where the respective target objects are located in the image sample, that is, the candidate regions, and then determine the features of the respective candidate regions according to the overall feature of the image sample and the respective candidate regions. The candidate region can be, for example, Figure 2 the frame 2111 as shown, which can be represented by the coordinates at the four vertices of the frame 2111. After obtaining the respective candidate regions, when calculating the features at the respective candidate regions (for example, there are 3 candidate regions in total), the feature of one candidate region is calculated each time according to the overall feature of the image sample. For example, the features of these 3 candidate regions are determined through 3 calculations.
[0081] In step S405, the features of each candidate region of the image sample are fused with at least one sample distribution feature of the image sample set corresponding to the image sample, so as to obtain the optimized features of each candidate region of the image sample. Since the candidate region features are obtained based on the overall features, the candidate region features are also affected by the sample distribution features. Because at least one sample distribution feature of the image sample set where the image sample is located has been obtained in the previous step, in this step, the at least one sample distribution feature and the features of the candidate regions can be fused to reduce the influence caused by the sample distribution features. The optimized features of the candidate regions obtained through such processing can more accurately reflect the category and location area of the target object.
[0082] In step S406, the optimized features of each candidate region of the image sample are input into the detection head network of the object detection network, so as to obtain the predicted category and the predicted region where each target object is located in the image sample. The detection head network of the object detection network can be a network including several fully connected layers, or other networks. After the optimized features of the candidate regions are input into the detection head network of the object detection network, the detection head network of the object detection network can analyze the optimized features and calculate to predict the category and location area of each target object. In this step, for example, the optimized features of one candidate region can be input into the detection head network of the object detection network each time, and then the detection head network of the object detection network determines the category of the target object and the area where the target object is located according to the optimized features of the candidate region and outputs. As an example, the detection head network of the object detection network can be the detection head network of a single shot multibox detector (SSD), etc.
[0083] In step S407, according to the category label of each target object in each image sample in multiple image sample sets and the predicted category and the predicted region where the target object is located, the parameters of the object detection network are adjusted. For the calculation of the loss in terms of category, for example, the cross-entropy loss function, the focal loss function, etc. can be used to achieve it. After calculating the loss in terms of category, various parameters in the object detection network are optimized through the backpropagation gradient algorithm or other algorithms. For the calculation of the loss in terms of region, for example, loss functions such as the L1 norm loss, the L2 norm loss, etc. can be used to achieve it. After calculating the loss in terms of region, various parameters in the object detection network are optimized through the backpropagation gradient algorithm or other algorithms.
[0084] In the method of training an object detection network according to some embodiments of the present application, multiple image sample sets are used to train the object detection network. Since each image sample set has different sample distribution characteristics, different sample distribution characteristics can be provided for the training process. The distribution feature extraction network in the object detection network can extract the sample distribution characteristics from the overall features of the image samples and fuse the features of the candidate regions with the sample distribution characteristics, so as to reduce the influence brought by the sample distribution characteristics in the features of each candidate region. In this way, interference factors can be excluded at each candidate region, and more accurate optimized features can be obtained, which is beneficial to improving the accuracy of the region for detecting the target object. The improvement of the accuracy of the region will also lead to the improvement of the accuracy of the category. Therefore, the trained object detection network has higher accuracy.
[0085] For the distribution feature extraction network, the present application gives some specific embodiments. Figure 5 It is a schematic diagram of a distribution feature extraction network according to some embodiments of the present application. As Figure 5 shown, the distribution feature extraction network 501 includes four sub-networks, namely a pooling layer 5011, a preset number of fully connected layers 5012, an activation function layer 5013, and a distribution feature layer 5014. The distribution feature layer 5014 is a matrix composed of multiple one-dimensional vectors. These four sub-networks are connected together in the described order to jointly achieve the purpose of extracting sample distribution characteristics.
[0086] It should be noted that the parameter form of the distribution feature layer 5014 is different from that of the other three sub-networks. The neurons of the pooling layer 5011, the preset number of fully connected layers 5012, and the activation function layer 5013 are all scalars. For example, each neuron of a fully connected layer can be a real number, and these neurons are arranged together to form a one-dimensional vector, that is, the fully connected layer. And a neuron of the distribution feature layer 5014 is a one-dimensional vector. If the distribution feature layer 5014 includes 10 neurons and each neuron is a one-dimensional vector with a length of 10, then the distribution feature layer 5014 is actually a 10x10 two-dimensional matrix.
[0087] Figure 6 It is a flowchart of extracting distribution features according to some embodiments of the present application. Similar to Figure 5 the distribution feature extraction network 501 shown, in Figure 6 the embodiment, the distribution feature extraction network includes a pooling layer, a preset number of fully connected layers, an activation function layer, and a distribution feature layer. The distribution feature layer is a matrix composed of multiple one-dimensional vectors. In Figure 6 , the sample distribution characteristics of the image samples are extracted through steps S601 - S604.
[0088] In step S601, the overall features of the image sample are input into the pooling layer to obtain the pooled overall features. Through the pooling process, it can reduce the number of output feature quantities of the convolutional layer, thereby reducing the number of parameters in the subsequent network layers and improving the overfitting phenomenon. In some embodiments, the pooling layer is one of the max pooling layer and the average pooling layer. The max pooling layer performs max pooling, which means that in the field of view of the filter, the largest value in the overall features is used as the value of each feature in this field of view. The average pooling layer performs average pooling, which means that in the field of view of the filter, the average value of all values in the overall features is used as the value of each feature in this field of view.
[0089] In step S602, the pooled overall features are input into a preset number of fully connected layers to obtain the distribution feature coefficients of the image sample. Since the fully connected layer has the ability to comprehensively predict all features, by using a preset number of fully connected layers, the distribution feature coefficients of the image sample can be predicted. The distribution feature coefficients can be used to represent which sample distribution features in the set composed of each sample distribution feature are hit by the sample distribution features of the corresponding image sample set. For example, in set A composed of sample distribution features, including a total of 10 sample distribution features A1 - A10, then the distribution feature coefficients represent which of these 10 sample distribution features appear in the image sample set corresponding to this image sample, or to what extent they appear in the image sample set corresponding to this image sample.
[0090] In some embodiments, the preset number of fully connected layers includes two fully connected layers.
[0091] In step S603, the distribution feature coefficients of the image sample are input into the activation function layer to obtain the optimized distribution feature coefficients. The activation function layer is the network layer that runs the activation function. The activation function is a function that runs on the neurons of the artificial neural network and is responsible for mapping the input of the neuron to the output end. In this step, the activation function is introduced to introduce non - linear factors, so that the object detection network can approximate any non - linear function arbitrarily, and thus the neural network can be applied to many non - linear models.
[0092] In some embodiments, the activation function layer is one of the hyperbolic tangent (Tanh) function layer and the rectified linear unit (ReLU) function layer.
[0093] In step S604, the optimized distribution feature coefficients are input into the distribution feature layer to obtain at least one sample distribution feature of the image sample set corresponding to the image sample. As described above, the distribution feature layer is a matrix composed of multiple one-dimensional vectors. The reason for determining the coefficient of each neuron in the distribution feature layer as a vector is that each neuron in the distribution feature layer can be understood as representing a sample distribution feature, and representing the sample distribution feature with a vector can express rich dimensions, so it is relatively accurate.
[0094] In some embodiments, inputting the optimized distribution feature coefficients into the distribution feature layer to obtain at least one sample distribution feature of the image sample set corresponding to the image sample includes: first multiplying the optimized distribution feature coefficients by each one-dimensional vector of the distribution feature layer to obtain a two-dimensional matrix; then determining the two-dimensional matrix as at least one sample distribution feature of the image sample set corresponding to the image sample. In this embodiment, each one-dimensional vector represents a sample distribution feature that may exist in the image sample set. Multiplying the optimized distribution feature coefficients by each one-dimensional vector of the distribution feature layer can achieve the fusion of the optimized distribution feature coefficients and the one-dimensional vectors, and the obtained two-dimensional matrix accurately expresses at least one sample distribution feature in the image sample set in a mathematical way.
[0095] As an example, if the optimized distribution feature coefficients are (1, 0, 0, 1, 1, 0), the first one-dimensional vector of the distribution feature layer is A1, the second one-dimensional vector is A2, the third one-dimensional vector is A3, the fourth one-dimensional vector is A4, the fifth one-dimensional vector is A5, and the sixth one-dimensional vector is A6, then the obtained two-dimensional matrix is (A1, 0, 0, A4, A5, 0). It can be found from this that there are 3 sample distribution features in the image sample set corresponding to this image sample, which are A1, A4, and A5 respectively.
[0096] In some embodiments, the candidate region feature network includes a candidate region determination network and a candidate region global feature extraction network, and inputting the global feature of the image sample into the candidate region feature network in the object detection network to determine the features of at least one candidate region where at least one target object is located in the image sample includes: first inputting the image sample into the candidate region determination network to obtain at least one candidate region of the image sample; then inputting at least one candidate region of the image sample and the global feature of the image sample into the candidate region global feature extraction network to obtain the features of each candidate region of the image sample.
[0097] In this embodiment, the candidate region feature network is divided into two sub-networks, namely, the candidate region determination network and the candidate region global feature extraction network. The candidate region determination network can be, for example, a Region Proposal Network (RPN), or a Region Convolutional Neural Network (R-CNN) combined with a selective search algorithm can be used to perform offline processing on the image sample to obtain candidate regions. It should be noted that since the purpose of the candidate region determination network is only to determine candidate regions, rather than determining the category of the target object and the actual region where it is located (i.e., the region predicted by the detection head network of the target detection network), this network can be pre-run before training to pre-obtain the candidate regions of each image sample in each image sample set, which can improve the training efficiency during training.
[0098] In some embodiments, before obtaining the image sample, the image sample is input into the candidate region determination network to pre-obtain at least one candidate region of the image sample. That is to say, the process of determining candidate regions can be carried out in advance, which can improve the training efficiency.
[0099] In some embodiments, at least one sample distribution feature of the image sample set corresponding to the image sample is respectively represented by at least one sample distribution vector, and the feature of each candidate region of the image sample is fused with at least one sample distribution feature of the image sample set corresponding to the image sample to obtain the optimized features of each candidate region of the image sample, including the following steps.
[0100] For the features of each candidate region of an image sample, perform the following operations: Add the features of the candidate region to at least one sample distribution vector of the image sample set corresponding to the image sample to obtain a first fusion result; Subtract the features of the candidate region from at least one sample distribution vector of the image sample set corresponding to the image sample to obtain a second fusion result; Multiply the features of the candidate region by at least one sample distribution vector of the image sample set corresponding to the image sample to obtain a third fusion result; Divide the features of the candidate region by at least one sample distribution vector of the image sample set corresponding to the image sample to obtain a fourth fusion result; Determine the optimized features of the candidate region of the image sample based on at least one of the first fusion result, the second fusion result, the third fusion result, and the fourth fusion result. During the training of the object detection network, no matter which one of the addition operation, subtraction operation, multiplication operation, and division operation is performed on the features of the candidate region and at least one sample distribution feature of the image sample set corresponding to the image sample, the purpose of reducing the influence of the sample distribution feature can be achieved. This is because the various parameters in the object detection network need to be adjusted according to the prediction loss at the same time, and the adjustment content is related to the type of operation selected in this embodiment.
[0101] Figure 7 is a schematic diagram of training an object detection network according to some embodiments of the present application. As Figure 7 shown, four image sample sets 7011, 7012, 7013, and 7014 are used to train the object detection network 702. The four image sample sets 7011 - 7014 respectively have sample distribution features 1 - 4. In one training, an image sample is obtained from these four image sample sets. Assuming that the image sample comes from the image sample set 7012, then the image sample has the sample distribution feature 2 of the image sample set 7012. The image sample is input into the overall feature extraction network 7021 to obtain the overall feature of the image sample. The overall feature of the image sample is input into the distribution feature extraction network 7024 to obtain the sample distribution feature 2 of the image sample set 7012 corresponding to the image sample.
[0102] The image sample is input into the candidate region determination network 7022 to obtain multiple candidate regions of the image sample. The overall feature of the image sample and the multiple candidate regions are input into the candidate region overall feature extraction network 7023, and then the features of each candidate region are obtained. The sample distribution feature 2 and the features of each candidate region are fused at 7025, that is, the influence brought by the sample distribution feature 2 is reduced, and then the optimized features of each candidate region are obtained. The optimized features of each candidate region are input into the detection head network 7026 of the object detection network, and then the detection head network 7026 of the object detection network outputs the categories and locations of each target object in the image sample.
[0103] According to the class label of the image sample and the classes and regions of each target object in the image sample predicted by the detection head network 7026 of the object detection network, use a loss function to determine the prediction loss and adjust each parameter of the object detection network 702.
[0104] Table 1 is a comparison table of prediction effects, which reflects Figure 1 the prediction effect of the object detection network in the embodiment shown and Figure 7 the comparison between the prediction effects of the object detection networks in the embodiments shown.
[0105]
[0106] Table 1: Comparison table of prediction effects.
[0107] In Table 1, when using the VOC07 image sample set to train the object detection network 102 and the object detection network 702, the mAP of the obtained object detection network 102 is 62.6 and the CorLoc is 78.7, which are respectively less than the mAP 63.0 and CorLoc 80.6 of the object detection network 702. When using the VOC07 and VOC12 image sample sets to train the object detection network 102 and the object detection network 702, the mAP of the obtained object detection network 102 is 63.5 and the CorLoc is 79.2, which are respectively less than the mAP 64.1 and CorLoc 82.2 of the object detection network 702. When using the VOC07 and COCO image sample sets to train the object detection network 102 and the object detection network 702, the mAP of the obtained object detection network 102 is 61.4 and the CorLoc is 78.2, which are respectively less than the mAP 63.0 and CorLoc 80.5 of the object detection network 702.
[0108] It can be seen from the content of Table 1 that whether using one image sample set to train the object detection networks 102 and 702 or using two image sample sets to train the object detection networks 102 and 702, the accuracy of the object detection network 702 is higher than that of the object detection network 102. It shows that the object detection network according to the embodiments disclosed in this application has good accuracy.
[0109] Figure 8 is a flowchart of a method for detecting target objects according to some embodiments of this application. In Figure 8 the embodiment shown, it includes steps S801 - S802.
[0110] In step S801, obtain an image to be detected, where the image to be detected includes at least one target object to be detected;
[0111] The image to be detected is used to identify the category and location area of the target object therein. For example, it can be a captured photo or a video frame in a video.
[0112] In step S802, the image to be detected is input into the target detection network according to the embodiments of the present application to obtain the category and location area of the at least one target object in the image to be detected. Since the target detection networks disclosed in the present application can all relatively well determine the category and location area of each target object in the image to be detected, the target detection network in the above-described embodiments can be used to achieve this purpose.
[0113] Figure 9 It is an exemplary structural block diagram of a device 900 for training a target detection network according to some embodiments of the present application. The device 900 for training the target detection network includes: a processing module 901 and an adjustment module 902. The processing module 901 is configured to obtain a plurality of image sample sets, each image sample in each image sample set includes at least one target object, and each target object has a corresponding category label. Each image sample set in the plurality of image sample sets has at least one sample distribution feature different from other image sample sets; for each image sample in each image sample set in the plurality of image sample sets, perform the following steps: input the image sample into the overall feature extraction network in the target detection network to obtain the overall feature of the image sample; input the overall feature of the image sample into the distribution feature extraction network in the target detection network to obtain at least one sample distribution feature of the image sample set corresponding to the image sample; input the overall feature of the image sample into the candidate region feature network in the target detection network to determine the features of at least one candidate region where the at least one target object in the image sample is located, and the at least one candidate region corresponds to the at least one target object one by one; fuse the features of each candidate region of the image sample with the at least one sample distribution feature of the image sample set corresponding to the image sample to obtain the optimized features of each candidate region of the image sample; input the optimized features of each candidate region of the image sample into the detection head network of the target detection network to obtain the predicted category and the predicted region where each target object in the image sample is located; the adjustment module 902 is configured to adjust the parameters of the target detection network according to the category label of each target object in each image sample in each image sample set in the plurality of image sample sets and the predicted category and the predicted region where the target object is located.
[0114] It should be noted that the above various modules can be implemented by software or hardware or a combination of both. Multiple different modules can be implemented in the same software or hardware structure, or one module can be implemented by multiple different software or hardware structures.
[0115] In an apparatus for training an object detection network according to some embodiments of the present application, a plurality of image sample sets are used to train the object detection network. Since each image sample set has different sample distribution characteristics, different sample distribution characteristics can be provided for the training process. The distribution feature extraction network in the object detection network can extract the sample distribution characteristics from the overall features of the image samples and fuse the features of the candidate regions with the sample distribution characteristics, thereby reducing the influence brought by the sample distribution characteristics in the features of each candidate region. In this way, interference factors can be excluded at each candidate region, and more accurate optimized features can be obtained, which is beneficial to improving the accuracy of the region for detecting the target object. The improvement of the accuracy of the region will also lead to the improvement of the accuracy of the category. Therefore, the trained object detection network has higher accuracy.
[0116] Figure 10 FIG. 5 is an exemplary structural block diagram of an object detection apparatus 1000 according to some embodiments of the present application. The object detection apparatus 1000 includes: an acquisition module 1001 and a detection module 1002. The acquisition module 1001 is configured to acquire an image to be detected, and the image to be detected includes at least one object to be detected; the detection module 1002 is configured to input the image to be detected into the object detection network according to the embodiments of the present application to obtain the category and location of the at least one object in the image to be detected.
[0117] It should be noted that the above various modules can be implemented in software or hardware or a combination of both. Multiple different modules can be implemented in the same software or hardware structure, or one module can be implemented by multiple different software or hardware structures.
[0118] Figure 11 FIG. 11 illustrates an example system 1100 that includes an example computing device 1110 representing one or more systems and / or devices that can implement the various methods described herein. The computing device 1110 can be, for example, a server of a service provider, a device associated with the server, a system on a chip, and / or any other suitable computing device or computing system. The apparatus 900 for training an object detection network described above with reference to Figure 9 can take the form of the computing device 1110. Alternatively, the apparatus 900 for training an object detection network can be implemented as a computer program in the form of an application 1116.
[0119] The example computing device 1110 shown in the figure includes a processing system 1111, one or more computer-readable media 1112, and one or more I / O interfaces 1113 that are communicatively coupled to each other. Although not shown, the computing device 1110 may also include a system bus or other data and command transfer systems that couple the various components to each other. The system bus may include any one or combination of different bus architectures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus that utilizes any of a variety of bus architectures. Also contemplated are various other examples, such as control and data lines.
[0120] The processing system 1111 represents the functionality to perform one or more operations using hardware. Thus, the processing system 1111 is shown as including hardware elements 1114 that can be configured as a processor, functional blocks, etc. This can include being implemented in hardware as an application-specific integrated circuit or other logic devices formed using one or more semiconductors. The hardware elements 1114 are not limited by the materials from which they are formed or the processing mechanisms employed therein. For example, a processor can be composed of (multiple) semiconductors and / or transistors (e.g., an electronic integrated circuit (IC)). In such a context, the instructions executable by the processor can be electronically executable instructions.
[0121] The computer-readable media 1112 is shown as including a memory 1115. The memory 1115 represents the memory / storage capacity associated with one or more computer-readable media. The memory / 1115 can include volatile media (such as random access memory (RAM)) and / or non-volatile media (such as read-only memory (ROM), flash memory, optical discs, magnetic disks, etc.). The memory 1115 can include fixed media (e.g., RAM, ROM, fixed hard disk drives, etc.) as well as removable media (e.g., flash memory, removable hard disk drives, optical discs, etc.). The computer-readable media 1112 can be configured in various other ways as further described below.
[0122] One or more I / O interfaces 1113 represent the functionality that allows a user to input commands and information into the computing device 1110 using various input devices and optionally also allows information to be presented to the user and / or other components or devices using various output devices. Examples of input devices include a keyboard, a cursor control device (e.g., a mouse), a microphone (e.g., for voice input), a scanner, a touch function (e.g., a capacitive or other sensor configured to detect physical touch), a camera (e.g., that can detect motion not involving touch as a gesture using visible or non-visible wavelengths such as infrared frequencies), etc. Examples of output devices include a display device, a speaker, a printer, a network card, a haptic response device, etc. Thus, the computing device 1110 can be configured in various ways as further described below to support user interaction.
[0123] The computing device 1110 further includes an application 1116. The application 1116 can be, for example, a software instance of the apparatus 900 for training an object detection network, and implements the techniques described herein in combination with other elements in the computing device 1110.
[0124] Various techniques may be described herein in the general context of software, hardware elements, or program modules. Generally, these modules include routines, programs, objects, elements, components, data structures, etc. that perform particular tasks or implement particular abstract data types. As used herein, the terms “module,” “function,” and “component” generally refer to software, firmware, hardware, or a combination thereof. The features of the techniques described herein are platform-independent, meaning that these techniques may be implemented on various computing platforms having a variety of processors.
[0125] In some embodiments of the present application, the term “module” or “unit” refers to a computer program or a part of a computer program having a predetermined function, and works together with other related parts to achieve a predetermined goal, and may be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) may be used to implement one or more modules or units. In addition, each module or unit may be a part of an overall module or unit that includes the function of that module or unit.
[0126] The implementation of the described modules and techniques may be stored on or transmitted across some form of computer-readable medium. The computer-readable medium may include various media accessible by the computing device 1110. By way of example and not limitation, the computer-readable medium may include “computer-readable storage medium” and “computer-readable signal medium”.
[0127] Contrary to mere signal transmission, carrier, or signal itself, a “computer-readable storage medium” refers to a medium and / or device that can persistently store information, and / or a tangible storage device. Thus, a computer-readable storage medium refers to a non-signal-bearing medium. Computer-readable storage media include hardware such as volatile and non-volatile, removable and non-removable media and / or storage devices implemented in a method or technology suitable for storing information such as computer-readable instructions, data structures, program modules, logic elements / circuits, or other data. Examples of computer-readable storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical storage devices, hard disks, cassette tapes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or other storage devices, tangible media, or articles of manufacture suitable for storing the desired information and accessible by a computer.
[0128] "Computer-readable signal medium" refers to a signal-bearing medium configured to send instructions to computing device 1110, such as via a network. A signal medium typically can embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, data signal, or other transmission mechanism. The signal medium also includes any information delivery medium. The term "modulated data signal" refers to a signal in which one or more of the characteristics are set or changed in order to encode information in the signal. By way of example, and not limitation, communication media include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media.
[0129] As described above, hardware elements 1114 and computer-readable media 1112 represent instructions, modules, programmable device logic, and / or fixed device logic implemented in hardware, which in some embodiments can be used to implement at least some aspects of the techniques described herein. Hardware elements can include integrated circuits or systems-on-a-chip, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), and components of other hardware devices implemented in silicon or other hardware. In this context, the hardware elements can serve as processing devices that execute program tasks defined by the instructions, modules, and / or logic embodied by the hardware elements, and as hardware devices that store instructions for execution, e.g., the computer-readable storage media described previously.
[0130] The foregoing combinations can also be used to implement the various techniques and modules herein. Thus, software, hardware, or program modules and other program modules can be implemented as one or more instructions and / or logic on some form of computer-readable storage medium and / or embodied by one or more hardware elements 1114. Computing device 1110 can be configured to implement specific instructions and / or functions corresponding to the software and / or hardware modules. Thus, for example, by using the computer-readable storage medium of the processing system and / or hardware elements 1114, a module can be implemented at least in part in hardware as a module executable by computing device 1110 as software. The instructions and / or functions can be executable / operable by one or more articles of manufacture (e.g., one or more computing devices 1110 and / or processing system 1111) to implement the techniques, modules, and examples described herein.
[0131] In various embodiments, computing device 1110 may be configured in a variety of different ways. For example, computing device 1110 may be implemented as a computer-like device including a personal computer, a desktop computer, a multi-screen computer, a laptop computer, a netbook, etc. Computing device 1110 may also be implemented as a mobile device-like device including mobile devices such as mobile phones, portable music players, portable gaming devices, tablet computers, multi-screen computers, etc. Computing device 1110 may also be implemented as a television-like device, which includes a device having or connected to a generally larger screen in a casual viewing environment. These devices include televisions, set-top boxes, gaming consoles, etc.
[0132] The techniques described herein may be supported by these various configurations of computing device 1110 and are not limited to the specific examples of the techniques described herein. The functionality may also be implemented in whole or in part using a distributed system, such as on “cloud” 1120 via platform 1122 as described below.
[0133] Cloud 1120 includes and / or represents platform 1122 for resources 1124. Platform 1122 abstracts the underlying functionality of the hardware (e.g., servers) and software resources of cloud 1120. Resources 1124 may include applications and / or data that may be used when performing computer processing on servers remote from computing device 1110. Resources 1124 may also include services provided via the Internet and / or via a subscriber network such as a cellular or Wi-Fi network.
[0134] Platform 1122 may abstract resources and functionality to connect computing device 1110 with other computing devices. Platform 1122 may also be used to abstract the hierarchical nature of resources to provide a corresponding level of hierarchy for the demands encountered for resources 1124 implemented via platform 1122. Thus, in an interconnected device embodiment, the implementation of the functionality described herein may be distributed throughout system 1100. For example, the functionality may be implemented in part on computing device 1110 and via platform 1122 that abstracts the functionality of cloud 1120.
[0135] It should be understood that, for clarity, embodiments of the present application have been described with reference to different functional units. However, it will be apparent that, without departing from the present application, the functionality of each functional unit may be implemented in a single unit, implemented in multiple units, or implemented as part of other functional units. For example, functionality described as being performed by a single unit may be performed by multiple different units. Thus, reference to a particular functional unit is only considered a reference to an appropriate unit for providing the described functionality, rather than indicating a strict logical or physical structure or organization. Thus, the present application may be implemented in a single unit or may be physically and functionally distributed between different units and circuits.
[0136] The present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of a computing device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computing device executes the method for training a target detection network provided in the above-mentioned various optional implementations.
[0137] While the present application has been described in conjunction with some embodiments, it is not intended to be limited to the specific forms set forth herein. Instead, the scope of the present application is limited only by the appended claims. Additionally, although individual features may be included in different claims, these may possibly be advantageously combined, and the inclusion in different claims does not imply that a combination of features is not feasible and / or advantageous. The order of features in the claims does not imply any specific order in which the features must work. Furthermore, in the claims, the word "comprising" does not exclude other elements, and the term "a" or "an" does not exclude a plurality. The reference signs in the claims are provided merely as a clarifying example and should not be construed as limiting the scope of the claims in any way.
[0138] It is understandable that in the specific implementation of this application, image sample sets and other related data are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
Claims
1. A method for training a target detection network, characterized in that: The method comprises: Acquire multiple image sample sets, each image sample in each image sample set includes at least one target object, each target object has a corresponding category label, and each image sample set in the multiple image sample sets has at least one sample distribution feature different from other image sample sets; For each image sample of each image sample set in the multiple image sample sets, the following steps are performed: Inputting the image sample into the overall feature extraction network in the target detection network to obtain the overall features of the image sample; Inputting the overall feature of the image sample into a distribution feature extraction network in the target detection network to obtain at least one sample distribution feature of an image sample set corresponding to the image sample; Inputting the overall features of the image sample into a candidate region feature network in the target detection network to determine features of at least one candidate region where the at least one target object in the image sample is located, the at least one candidate region corresponding to the at least one target object in one-to-one correspondence; fusing the feature of each candidate region of the image sample with the at least one sample distribution feature of the image sample set corresponding to the image sample to obtain optimized features of each candidate region of the image sample; Inputting the optimized features of each candidate region of the image sample into the detection head network of the target detection network to obtain the predicted category and predicted region of each target object in the image sample; According to the category label of each target object in each image sample in each image sample set of the multiple image sample sets and the predicted category and predicted area of the target object, the parameters of the target detection network are adjusted.
2. The method according to claim 1, characterized in that The distribution feature extraction network includes a pooling layer, a preset number of fully connected layers, an activation function layer and a distribution feature layer, wherein the distribution feature layer is a matrix composed of multiple one-dimensional vectors. And the step of inputting the overall feature of the image sample into the distribution feature extraction network in the target detection network to obtain at least one sample distribution feature of the image sample set corresponding to the image sample comprises: Inputting the overall features of the image sample into the pooling layer to obtain the overall features processed by pooling; Inputting the pooled overall features into the preset number of fully connected layers to obtain distribution feature coefficients of the image samples; Inputting the distribution characteristic coefficient of the image sample into the activation function layer to obtain an optimized distribution characteristic coefficient; The optimized distribution feature coefficient is input into the distribution feature layer to obtain at least one sample distribution feature of the image sample set corresponding to the image sample.
3. The method according to claim 2, characterized in that The pooling layer is one of a maximum pooling layer and an average pooling layer.
4. The method according to claim 2, characterized in that: The activation function layer is one of a hyperbolic tangent (Tanh) function layer and a rectified linear unit (ReLU) function layer.
5. The method according to claim 2, characterized in that: The step of inputting the optimized distribution feature coefficient into the distribution feature layer to obtain at least one sample distribution feature of the image sample set corresponding to the image sample comprises: Multiplying the optimized distribution feature coefficients by each one-dimensional vector of the distribution feature layer to obtain a two-dimensional matrix; The two-dimensional matrix is determined as at least one sample distribution feature of an image sample set corresponding to the image sample.
6. The method according to claim 1, characterized in that The candidate region feature network includes a candidate region determination network and a candidate region overall feature extraction network. And the step of inputting the overall features of the image sample into a candidate region feature network in the target detection network to determine features of at least one candidate region where the at least one target object in the image sample is located comprises: Inputting the image sample into the candidate region determination network to obtain at least one candidate region of the image sample; At least one candidate region of the image sample and the overall feature of the image sample are input into the candidate region overall feature extraction network to obtain the features of each candidate region of the image sample.
7. The method according to claim 1, characterized in that At least one sample distribution feature of the image sample set corresponding to the image sample is represented by at least one sample distribution vector. Furthermore, the step of fusing the feature of each candidate region of the image sample with the at least one sample distribution feature of the image sample set corresponding to the image sample to obtain optimized features of each candidate region of the image sample includes: For each candidate region feature of the image sample, perform the following operations: Performing an addition operation on the feature of the candidate region and at least one sample distribution vector of the image sample set corresponding to the image sample to obtain a first fusion result; Subtracting the feature of the candidate region from at least one sample distribution vector of the image sample set corresponding to the image sample to obtain a second fusion result; Performing a multiplication operation on the feature of the candidate region and at least one sample distribution vector of the image sample set corresponding to the image sample to obtain a third fusion result; Performing a division operation on the feature of the candidate region and at least one sample distribution vector of the image sample set corresponding to the image sample to obtain a fourth fusion result; Determine an optimization feature of the candidate region of the image sample according to at least one of the first fusion result, the second fusion result, the third fusion result, and the fourth fusion result.
8. A target object detection method, characterized in that: The method comprises: Acquire an image to be detected, wherein the image to be detected includes at least one target object to be detected; The image to be detected is input into the target detection network according to claim 1 to obtain the category and the region where the at least one target object in the image to be detected is located.
9. A device for training a target detection network, characterized in that: The device comprises: A processing module is configured to obtain a plurality of image sample sets, each image sample in each image sample set includes at least one target object, each target object has a corresponding category label, and each image sample set in the plurality of image sample sets has at least one sample distribution feature different from other image sample sets; For each image sample of each image sample set in the multiple image sample sets, the following steps are performed: inputting the image sample into an overall feature extraction network in the target detection network to obtain an overall feature of the image sample; Inputting the overall feature of the image sample into a distribution feature extraction network in the target detection network to obtain at least one sample distribution feature of an image sample set corresponding to the image sample; Inputting the overall features of the image sample into a candidate region feature network in the target detection network to determine features of at least one candidate region where the at least one target object in the image sample is located, the at least one candidate region corresponding to the at least one target object in one-to-one correspondence; fusing the feature of each candidate region of the image sample with the at least one sample distribution feature of the image sample set corresponding to the image sample to obtain optimized features of each candidate region of the image sample; Inputting the optimized features of each candidate region of the image sample into the detection head network of the target detection network to obtain the predicted category and predicted region of each target object in the image sample; The adjustment module is configured to adjust the parameters of the target detection network according to the category label of each target object in each image sample in each image sample set of the multiple image sample sets and the predicted category and predicted area of the target object.
10. A target object detection device, characterized in that: The device comprises: An acquisition module is configured to acquire an image to be detected, wherein the image to be detected includes at least one target object to be detected; A detection module is configured to input the image to be detected into the target detection network according to claim 1 to obtain the category and area of the at least one target object in the image to be detected.
11. A computing device comprising: a memory configured to store computer-executable instructions; A processor, which is configured to perform the method according to any one of claims 1 to 8 when the computer executable instructions are executed by the processor.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed, the method according to any one of claims 1 to 8 is performed.
13. A computer program product, characterized in that The computer program product comprises computer executable instructions which, when executed, implement the method according to any one of claims 1 to 8.