Neural Network Search Method, Device and Electronic Device for Target Detection Model
By constructing search methods for search space and neural network structures, the backbone, neck and head networks of the object detection model are automatically designed, solving the problem of detection model optimization under computing resource limitations, and achieving efficient and accurate object detection.
Patent Information
- Application Number
- CN202210542828.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-05-18
AI Technical Summary
Under the limitation of computing resources, artificially designed detection models cannot effectively explore endless detection models, and the effects of backbone detection models are difficult to achieve optimal results in different learning tasks.
By building a search space, the combination of backbone network, neck network and head network is automatically searched, and multiple candidate detection models are generated using neural network structure search method, and the target detection model that meets the set conditions is selected through training and adjustment.
It reduces the generation cost of object detection model, saves training time, and enhances the consistency and accuracy of the model.
Smart Images

Figure CN114819100B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the fields of intelligent search, deep learning, image processing, and computer vision technology, and particularly relates to a neural network search method, device, electronic device, and storage medium for a target detection model. Background Art
[0002] Deep learning is a branch of machine learning, aiming to establish a neural network that simulates the human brain for analysis and learning, and processes data by imitating the working mechanism of the human brain. For example, images are usually applied in the fields of video recognition, image recognition, or sound recognition.
[0003] However, under the limitation of computing resources, manually designed detection models can only rely on continuous experiments and experience in the past, and it is impossible to experiment with endless detection models. The target detection model depends on the backbone detection model to learn image features. Therefore, the effect of the backbone detection model has a great impact on the target detection model. However, due to different learning tasks, it is difficult to say that the optimal backbone detection model can also bring the optimal effect in the detection task. Summary of the Invention
[0004] The present disclosure provides a neural network search method, device, electronic device, and storage medium for a target detection model.
[0005] According to one aspect of the present disclosure, there is provided a neural network search method for a target detection model, including:
[0006] Obtaining a search space, where the search space includes a plurality of network modules, and the plurality of network modules are used to form a backbone network, a detection neck network, and a detection head network;
[0007] Performing multiple neural network structure searches on the network modules in the search space to obtain a plurality of candidate detection models, where the candidate detection models include a candidate backbone network, a candidate detection neck network, and a candidate detection head network;
[0008] For any candidate detection model, training the candidate detection model based on the first training sample image to adjust the model of the candidate detection model to obtain a trained candidate detection model;
[0009] Selecting the trained candidate detection model that meets the set conditions as the target detection model.
[0010] According to another aspect of the present disclosure, there is provided a neural network search device for a target detection model, including:
[0011] An obtaining module, configured to obtain a search space, where the search space includes a plurality of network modules, and the plurality of network modules are used to form a backbone network, a detection neck network, and a detection head network;
[0012] A search module, configured to perform multiple neural network structure searches on network modules in a search space to obtain multiple candidate detection models, where the candidate detection models include candidate backbone networks, candidate detection neck networks, and candidate detection head networks;
[0013] A training module, configured to train any candidate detection model based on first training sample images to adjust the candidate detection model and obtain a trained candidate detection model;
[0014] A determination module, configured to select a trained candidate detection model that meets set conditions as a target detection model.
[0015] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the neural network search method for the target detection model in the above-mentioned embodiment of one aspect.
[0019] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, on which a computer program is stored, and the computer instructions are used to cause a computer to execute the neural network search method for the target detection model in the above-mentioned embodiment of one aspect.
[0020] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the neural network search method for the target detection model in the above-mentioned embodiment of one aspect.
[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0023] Figure 1 is a schematic flowchart of a neural network search method for a target detection model provided by an embodiment of the present disclosure;
[0024] Figure 2Structural schematic diagram of the search space of a neural network search method for an object detection model provided by an embodiment of the present disclosure;
[0025] Figure 3 Flow schematic diagram of another neural network search method for an object detection model provided by an embodiment of the present disclosure;
[0026] Figure 4 Flow schematic diagram of another neural network search method for an object detection model provided by an embodiment of the present disclosure;
[0027] Figure 5 Flow schematic diagram of another neural network search method for an object detection model provided by an embodiment of the present disclosure;
[0028] Figure 6 Flow schematic diagram of another neural network search method for an object detection model provided by an embodiment of the present disclosure;
[0029] Figure 7 Schematic diagram of sampling samples of a neural network search method for an object detection model provided by an embodiment of the present disclosure;
[0030] Figure 8 Structural schematic diagram of a neural network search device for an object detection model provided by an embodiment of the present disclosure;
[0031] Figure 9 Block diagram of an electronic device for a neural network search method of an object detection model according to an embodiment of the present disclosure. Detailed implementation manners
[0032] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0033] The following describes a neural network search method, device, and electronic device for an object detection model according to an embodiment of the present disclosure with reference to the accompanying drawings.
[0034] Artificial intelligence is a discipline that studies the use of computers to simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and it has technical fields at both the hardware and software levels. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, as well as deep learning, big data processing technology, and knowledge graph technology.
[0035] Deep learning is a new research direction in the field of machine learning. Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability to analyze and learn like humans, and be able to recognize data such as text, images, and sounds. Deep learning is a complex machine learning algorithm, and the effects achieved in speech and image recognition far exceed those of previous related technologies.
[0036] Image processing technology is a technology that uses a computer to process image information. It mainly includes image digitization, image enhancement and restoration, image data encoding, image segmentation, and image recognition, etc. The most basic method for processing images is the point processing method. Since the object processed by this method is pixels, it gets this name. The point processing method is simple and effective, and is mainly used for image brightness adjustment, image contrast adjustment, and image brightness inversion processing, etc.
[0037] Computer vision is a science that studies how to make machines "see". Further speaking, it refers to using cameras and computers to replace the human eye to perform machine vision such as target recognition, tracking, and measurement on targets, and further perform graphic processing to make the computer-processed images more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain 'information' from images or multi-dimensional data. The information referred to here is the information defined by Shannon, which can be used to help make a "decision". Because perception can be regarded as extracting information from sensory signals, computer vision can also be regarded as the science of studying how to make artificial systems "perceive" from images or multi-dimensional data.
[0038] Intelligent search is a new generation of search engine that combines artificial intelligence technology. In addition to providing traditional functions such as fast retrieval and relevance ranking, it can also provide functions such as user role registration, automatic recognition of user interests, semantic understanding of content, intelligent information filtering, and pushing.
[0039] Figure 1 Schematic flow diagram of a neural network search method for an object detection model provided by an embodiment of the present disclosure.
[0040] S101. Obtain a search space, where the search space includes multiple network modules, and the multiple network modules are used to form a backbone network, a detection neck network, and a detection head network.
[0041] The search space includes multiple layers of neural networks. For each layer of the network, there is a corresponding set of network modules of the search neural network. The set of network modules includes basic network modules (blocks) composed of multiple neurons. The design of the search space is crucial in Neural Architecture Search (NAS).
[0042] Based on the fact that a neural network structure includes the number of network modules (blocks) and the number of channels, the embodiments of the present disclosure propose a search space that performs different combinations of blocks and channels for a detection model to construct different detection models, that is, perform arbitrary combinations in terms of model depth and model width to obtain detection models with different depths and widths. The search space proposed in the solution of the present disclosure can achieve automatic search to design a detection model, making it more practical.
[0043] The search space may include multiple blocks. It should be noted that the blocks can form a detection super network. In the embodiments of the present disclosure, the block is the basic network unit that forms the detection backbone network, the detection neck network, and the detection head network. Each block is a module composed of multiple convolution functions and can achieve specific functions through different combinations.
[0044] For a detection model, the functions of the backbone network, the detection neck network, and the detection head network are different.
[0045] In the embodiments of the present disclosure, the backbone network is the basic feature extractor for the object detection task. For example, the backbone network can be used to extract image data, sound data, etc. from the input data. The structures and functions of the backbone network required in different scenarios may be different. For different requirements for accuracy and efficiency, people can choose a deeper and more connection-dense backbone network, such as ResNet, ResNeXt, or AmoebaNet, etc., or a lightweight backbone network, such as MobileNet, ShuffleNet, SqueezeNet, Xception, or MobileNetV2, etc. The detection neck network and the detection head network are used for data transmission and processing of the data extracted by the backbone network.
[0046] It should be noted that the backbone network can control the number of blocks and the number of channels output to the detection neck, and the output of the neck further controls the number of channels of the detection head.
[0047] S102. Perform multiple neural network architecture searches on the network modules in the search space to obtain multiple candidate detection models. Among them, the candidate detection model includes a candidate backbone network, a candidate detection neck network, and a candidate detection head network.
[0048] For a detection network, the functions contributed by the backbone network, the detection neck network, and the detection head network are different. Different designs of the backbone network, the detection neck network, and the detection head network are pieced together to form a detection network, and the proportion of the three parts should be appropriately modified according to the specific situation to enable the detection network to achieve optimal performance.
[0049] However, in the search space proposed in the embodiments of the present disclosure, the entire detection model is directly searched instead of decoupling the backbone network, the detection neck network, and the detection head network separately. The retrieved model obtained by this search has higher consistency and a simpler process.
[0050] In the embodiments of the present disclosure, multiple candidate detection models can be constructed in the search space by using the NAS technology after setting the maximum number of blocks and the maximum number of channels in advance. It should be noted that the number of blocks and the number of channels of each candidate detection model are randomly generated and can be different. The number of blocks of each generated candidate detection model is less than the set maximum number of blocks, and the number of channels of each candidate detection model is less than the set maximum number of channels. NAS can obtain a candidate detection model in the search space through an activation function. For example, this activation function can be various and is not limited here.
[0051] In the embodiments of the present disclosure, in one multiple neural network architecture search, when calculating the forward output of the network each time, each block can select a path, and one path represents a candidate detection model, that is, a sub-network corresponding to a detection super-network. For example, as Figure 2 shown, block1 is selected, while block2 is skipped, and blocki is selected. At this time, block1 and blocki form a sub-network. It should be noted that whether each block is selected or skipped is random.
[0052] S103. For any candidate detection model, train the candidate detection model based on the first training sample image to adjust the model of the candidate detection model and obtain the trained candidate detection model.
[0053] After obtaining multiple candidate detection models, the first training sample image can be input into any candidate detection model, and the model of the candidate detection model can be adjusted according to the result output by the candidate detection model to preliminarily optimize the candidate detection model.
[0054] Optionally, a loss function may be determined based on the candidate detection model, and then a loss value may be determined based on the output result and the ground truth result of the first training sample image. Then, the candidate detection model may be adjusted based on the loss value until the loss value is less than the loss threshold, at which point the training ends and the trained candidate detection model is output. It should be noted that the loss threshold may be set in advance and can be set according to actual training needs, and no specific limitation is imposed here. It should be noted that there may be multiple post-candidate detection models, and the loss functions of each post-candidate detection model may be the same or different, which can be determined according to the actual situation.
[0055] Optionally, the training time or the number of training times may also be set in advance. When the training reaches the training time or the number of training times, the training ends and the trained candidate detection model is output. It should be noted that the training time or the number of training times may be set in advance and can be set according to actual training needs, and no specific limitation is imposed here.
[0056] S104. Select the trained candidate detection model that meets the set conditions as the target detection model.
[0057] In the embodiments of the present disclosure, the set conditions may be various.
[0058] Optionally, after obtaining the trained candidate detection model, the accuracy of the trained candidate detection model may be analyzed. When the accuracy of the trained candidate detection model reaches the set accuracy, it is considered that the trained candidate detection model meets the set conditions and is identified as the target detection model. For example, after the post-candidate detection model completes training, it enters the model screening stage, that is, the accuracy of any combination of sub-networks in the post-candidate detection model is confirmed.
[0059] Optionally, it may also be determined according to the loss value of the trained candidate detection model. When the loss value of the trained candidate detection model is less than the loss threshold, the trained candidate detection model may be identified as the target detection model.
[0060] In an embodiment of the present disclosure, first, a search space is obtained. The search space includes a plurality of network modules, and the plurality of network modules are used to form a backbone network, a detection neck network, and a detection head network. Then, the network modules in the search space are subjected to multiple neural network architecture searches to obtain a plurality of candidate detection models. Among them, the candidate detection models include a candidate backbone network, a candidate detection neck network, and a candidate detection head network. Then, for any candidate detection model, the candidate detection model is trained based on the first training sample images to adjust the candidate detection model to obtain a trained candidate detection model. Finally, the trained candidate detection model that meets the set conditions is selected as the target detection model. Thus, the target detection model is directly searched and obtained. Compared with the prior art, there is no need to separately load the backbone network, which can reduce the cost of generating the target detection model, save training time, and enhance the consistency of the generated target detection model.
[0061] In the above embodiment, when selecting the trained candidate detection model that meets the set conditions as the target detection model, it can also be achieved through Figure 3 For further explanation, the method includes:
[0062] S301, select the trained candidate detection model to be verified from the multiple trained candidate detection models according to the evolutionary algorithm.
[0063] In an embodiment of the present disclosure, after obtaining multiple trained candidate detection models, the evolutionary algorithm can be used. That is, a part of the sub-networks are randomly selected for verification to obtain the K sub-networks with the highest accuracy. Other sub-networks are derived from these K sub-networks, and the above steps are repeated until all the multiple trained candidate detection models are traversed. If the K sub-networks remain unchanged, the trained candidate detection model to be verified is output. It should be noted that K is a positive integer and can be adjusted according to actual needs, and no specific limitation is made here.
[0064] S302, perform accuracy verification on the trained candidate detection model to be verified based on the verification sample images, and obtain the model accuracy of the trained candidate detection model to be verified.
[0065] In an embodiment of the present disclosure, the accuracy verification may include the following steps: after the candidate detection model is trained, enter the search stage, that is, confirm the accuracy of any combination of sub-networks in the candidate detection model, and then summarize the accuracy of the sub-networks in the candidate detection model to determine the model accuracy of the trained candidate detection model.
[0066] S303, select the trained candidate detection model to be verified whose model performance parameters and model accuracy meet the conditions as the target detection model.
[0067] It should be noted that the model performance parameters can include various types. For example, they can include model computing power (floating point operations per second, FLOPS), error rate, recall, etc.
[0068] In the embodiments of the present disclosure, the post-training candidate detection models to be verified can be determined based on the model computing power and the model accuracy. That is, when the model computing power is greater than the preset computing power and the model accuracy is greater than the preset accuracy, it can be considered that the post-training candidate detection model meets the conditions and can be used as the target detection model. It should be noted that the preset computing power and the preset accuracy can be set in advance and can be set according to actual needs.
[0069] It should be noted that the finally output target detection model can be one, or multiple, or the output can be 0, that is, there is no candidate detection model that meets the conditions, which specifically depends on the actual situation.
[0070] In the embodiments of the present disclosure, first, the post-training candidate detection models to be verified are selected from multiple post-training candidate detection models according to the evolutionary algorithm, and then the accuracy of the post-training candidate detection models to be verified is verified based on the verification sample images to obtain the model accuracy of the post-training candidate detection models to be verified. Finally, the post-training candidate detection models whose model performance parameters and model accuracy meet the conditions are selected as the target detection models. Thus, further screening the post-training candidate detection models based on the preset accuracy and model performance parameters to determine the target detection model can improve the accuracy and computing power of the target detection model.
[0071] Figure 4 The following is a schematic flowchart of another neural network search method for the target detection model provided by the embodiments of the present disclosure. As shown in the figure, the method includes:
[0072] S401, for the i-th neural network structure search, determine the network modules that have been searched in the search space in the previous i - 1 times, where i is a positive integer greater than 1 and not greater than the set total number of searches.
[0073] In the embodiments of the present disclosure, before searching for candidate detection models in the search space, it is necessary to first set the maximum number of searches. In this way, it can be prevented that the randomly obtained candidate detection models exceed the set standards of the models and cannot be used or trained.
[0074] It should be noted that the maximum number of searches can be changed according to actual design requirements and is not limited here.
[0075] Optionally, in the embodiments of the present disclosure, during the next multiple neural network architecture search, the blocks that have been used in the previous search will not be selected again, that is, the blocks of each candidate detection model are different. For example, the search space contains N network modules such as block1, block2, block3... block i... block N. The candidate detection model obtained by the first NAS search includes block1 and block2. The second search needs to determine the next candidate detection model among the N-2 network modules such as block3... block i... block N.
[0076] S402. Based on the searched network modules, determine the remaining network modules in the search space and perform a search among the remaining network modules.
[0077] In implementation, in a relatively large search space, the probability of sampling the same detection model is almost 0. For example, a certain detection model includes 10 network modules, that is, the search quantity is 10. Each module has the possibility of ending or forming a detection model with the other nine network modules, that is, there are 10 choices in total. Then the permutation and combination situation of the detection model is 10 to the 10th power. Therefore, it can be considered that in a huge detection model pool, the probability of repeatedly selecting the same detection model is 0. The sampling of the search space can use the method of sampling without replacement to ensure the fairness of the search point sampling.
[0078] In the embodiments of the present disclosure, first, for the i-th neural network architecture search, determine the network modules that have been searched in the search space in the previous i-1 times, where i is a positive integer greater than 1 and not greater than the set total number of searches. Then, based on the searched network modules, determine the remaining network modules in the search space and perform a search among the remaining network modules. Thus, by using the method of sampling without replacement, the fairness of the search point sampling can be ensured, providing a basis for generating more accurate candidate detection models.
[0079] In the above embodiments, the process of performing a neural network architecture search in the search space to obtain a candidate detection model can also be achieved through Figure 5 For further explanation, the method includes:
[0080] S501. Perform a search among the currently searchable network modules in the search space to obtain the first network module for constructing the candidate detection model.
[0081] The process of obtaining the first network module is random. After selecting a set number of first network modules, the search stops.
[0082] It should be noted that each first network module will not be selected repeatedly. Therefore, during each search, the network modules obtained from the previous search need to be removed from the search space, and the remaining network modules serve as the currently searchable network modules. For example, the search space contains N network modules such as block1, block2, block3... block i... block N. When searching for the first time, the searchable network modules are the entire search space. If the candidate detection model obtained from the first search includes block1 and block2, the searchable network modules for the second search are N-2 network modules such as block3... block i... block N.
[0083] S502. Replace the remaining second network modules in the searchable network modules with a set convolutional layer.
[0084] In the embodiments of the present disclosure, the second network module is the network module skipped during the search process. As Figure 2 shown, if block1 is selected, block2 is skipped, and blocki is selected, then block1 and blocki are the first network modules, and block2 is the second network module.
[0085] It should be noted that searching for the modules in the search space in a direct skipping manner will cause instability of the detection model. This instability is reflected in the inconsistency between the accuracy of the detection model and the actual retraining accuracy. This will greatly affect the accuracy of the detection model.
[0086] Therefore, in the embodiments of the present disclosure, when choosing to skip the second network module, the module that is directly skipped can be replaced by using a 1×1 convolutional layer, which can solve the above problems. Thus, the stability of the detection model is increased.
[0087] S503. Generate a candidate detection model based on the set convolutional layer corresponding to the first network module and the second network module.
[0088] In the embodiments of the present disclosure, first search in the currently searchable network modules in the search space to obtain the first network modules used to construct the candidate detection model, then replace the remaining second network modules in the searchable network modules with a set convolutional layer, and finally generate a candidate detection model based on the set convolutional layer corresponding to the first network module and the second network module. Thus, by setting the convolutional layer, the stability and accuracy of the candidate detection model can be increased, and the effect of model training can be improved.
[0089] In the embodiments of the present disclosure, the network module includes multiple channels, and the weights of the maximum number of channels are shared among the multiple channels. It can also be Figure 6 Further explained that the method includes:
[0090] S601, randomly select channels for the first network module during the neural network architecture search process.
[0091] In the embodiments of the present disclosure, the channels of the first network module are randomly selected and not fixed. For example, if the minimum number of channels of the first network module is 64 and the maximum is 128, then the number of channels of the first network module can be randomly selected from between 64 and 128.
[0092] S602, select the weights with the maximum number of channels shared for training during training.
[0093] It should be noted that the number of channels of the convolution operation in the selected block can have different selections, but there is no concept of different paths.
[0094] Therefore, when the selected number of channels is less than the maximum number of channels, the weights corresponding to the randomly selected channels are sliced from the weights with the maximum number of channels, and the weights corresponding to the randomly selected channels are updated with gradients. For example, if the minimum number of channels of the first network module is 64 and the maximum is 128, when 64 channels are selected, the weights with the maximum number of channels are still used, that is, the weights with 128 channels are used to train the detection model. After the training of the detection model is completed, the gradients of the selected 64 channels are updated, while the remaining 64 unselected channels remain unchanged, thereby ensuring the accuracy of training.
[0095] In the embodiments of the present disclosure, it is also necessary to retrain the target detection model based on the second training sample image, thereby further optimizing the target detection model and making the target detection model more accurate. It is also necessary to delete the set convolution layer in the target detection model and retrain the target detection model with the set convolution layer deleted based on the second training sample image. Thus, a complete target detection model can be generated.
[0096] As Figure 7 shown, C3 is the maximum number of channels, C2 is included in C3, and C1 is included in C2. Then when C1 or C2 is selected as the number of channels for training, the weight value is the weight value of C3.
[0097] After determining the target detection model, it is also possible to retrain the target detection model based on the second training sample image to obtain the final target detection model. It should be noted that the second training sample image and the first training sample image can be the same or different, and no limitation is made here.
[0098] Optionally, the loss function can be determined based on the object detection model. Then, the second training sample image is input into the object detection model, and the loss value is determined based on the output result and the ground truth result of the second training sample image. Then, the candidate detection model is adjusted based on the loss value until the loss value is less than the loss threshold, at which point the training ends and the trained candidate detection model is output. It should be noted that the loss threshold and loss function for training the object detection model may be the same as or different from those for training the candidate detection model, and specific settings are required, which are not limited here.
[0099] Optionally, the training time or the number of training times can also be set in advance. When the training reaches the training time or the number of training times, the training ends and the trained object detection model is output. It should be noted that the training time or the number of training times for training the object detection model may be the same as or different from those for training the candidate detection model, and specific settings are required according to actual needs, which are not limited here.
[0100] In the embodiments of the present disclosure, during the neural network architecture search process, channels are randomly selected for the first network module, and then the weights with the largest number of shared channels are selected for training during training. Thus, by sharing the weights with the largest number of channels among multiple channels, the training effect of the detection model can be increased and the accuracy of training the detection model can be improved.
[0101] Corresponding to the neural network search methods for the object detection model provided in the above several embodiments, an embodiment of the present disclosure also provides a neural network search device for an object detection model. Since the neural network search device for the object detection model provided in the embodiments of the present disclosure corresponds to the neural network search methods for the object detection model provided in the above several embodiments, the implementation manners of the above neural network search methods for the object detection model are also applicable to the neural network search device for the object detection model provided in the embodiments of the present disclosure and will not be described in detail in the following embodiments.
[0102] Figure 8 FIG. is a schematic structural diagram of a neural network search device for an object detection model provided in an embodiment of the present disclosure.
[0103] As Figure 8 shown, the neural network search device 800 for the object detection model may include: an acquisition module 801, a search module 802, a training module 803, and a determination module 804.
[0104] Among them, the acquisition module 801 is used to acquire a search space, and the search space includes a plurality of network modules, and the plurality of network modules are used to form a backbone network, a detection neck network, and a detection head network.
[0105] A search module 802 is configured to perform multiple neural network architecture searches on network modules in a search space to obtain multiple candidate detection models. Among them, the candidate detection models include candidate backbone networks, candidate detection neck networks, and candidate detection head networks.
[0106] A training module 803 is configured to, for any candidate detection model, train the candidate detection model based on the first training sample images to adjust the candidate detection model and obtain a trained candidate detection model.
[0107] A determination module 804 is configured to select a trained candidate detection model that meets the set conditions as the target detection model.
[0108] In an embodiment of the present disclosure, the determination module 804 is further configured to: select a to-be-verified trained candidate detection model from multiple trained candidate detection models according to an evolutionary algorithm; perform accuracy verification on the to-be-verified trained candidate detection model based on verification sample images to obtain the model accuracy of the to-be-verified trained candidate detection model; select a to-be-verified trained candidate detection model whose model performance parameters and detection model accuracy meet the conditions as the target detection model.
[0109] In an embodiment of the present disclosure, the determination module 804 is further configured to: retrain the target detection model based on the second training sample images to obtain the final target detection model.
[0110] In an embodiment of the present disclosure, the neural network search device 800 for the target detection model is further configured to: for the i-th neural network architecture search, determine the network modules that have been searched in the search space in the previous i - 1 times, where i is a positive integer greater than 1 and not greater than the set total number of searches; based on the searched network modules, determine the remaining network modules in the search space and perform a search among the remaining network modules.
[0111] In an embodiment of the present disclosure, the search module 802 is further configured to: perform a search among the currently searchable network modules in the search space to obtain a first network module for constructing a candidate detection model; replace the remaining second network modules in the searchable network modules with set convolutional layers; generate a candidate detection model based on the first network module and the set convolutional layers corresponding to the second network modules.
[0112] In an embodiment of the present disclosure, the determination module 804 is further configured to: delete the set convolutional layers in the target detection model; retrain the target detection model with the set convolutional layers deleted based on the second training sample images.
[0113] In an embodiment of the present disclosure, the determining module 804 is further configured to: randomly select channels for the first network module during the neural network architecture search process; and select the weights with the maximum number of shared channels for training during training.
[0114] In an embodiment of the present disclosure, the training module 803 is further configured to: in response to the number of selected channels being less than the maximum number of channels, split out the weights corresponding to the randomly selected channels from the weights with the maximum number of channels; and perform gradient update on the weights corresponding to the randomly selected channels.
[0115] Thus, through the neural network search device of the object detection model, the object detection model can be directly searched and obtained. Compared with the prior art, there is no need to separately load the backbone network, which can reduce the cost of generating the object detection model, save training time, and enhance the consistency of the neural network model.
[0116] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0117] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0118] Figure 9 FIG. shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0119] As Figure 9 shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 902 or the computer program loaded from the storage unit 906 to the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.
[0120] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906 such as a keyboard, mouse, etc.; output unit 907, such as various types of displays, speakers, etc.; storage unit 908, such as a disk, optical disc, etc.; and communication unit 909, such as a network card, modem, wireless communication transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0121] Computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning detection model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 901 executes the various methods and processes described above, such as the neural network search method for the object detection model. For example, in some embodiments, the neural network search method for the object detection model can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as storage unit 906. In some of these embodiments, part or all of the computer program can be loaded and / or installed onto device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by computing unit 901, one or more steps of the neural network search method for the object detection model described above can be executed. Alternatively, in other embodiments, computing unit 901 can be configured to execute the neural network search method for the object detection model in any other suitable manner (e.g., by means of firmware).
[0122] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0123] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0124] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0125] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0126] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain networks.
[0127] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating blockchain.
[0128] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitation is imposed herein. The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A neural network search method for a target detection model, comprising: Obtaining a search space, where the search space includes a plurality of network modules, and the plurality of network modules are used to form a backbone network, a detection neck network, and a detection head network; Searching in the currently searchable network modules in the search space to obtain a first network module for constructing a candidate detection model; Replacing the remaining second network modules in the searchable network modules with a set convolutional layer; Generating a plurality of the candidate detection models based on the first network module and the set convolutional layer corresponding to the second network module, where the candidate detection models include a candidate backbone network, a candidate detection neck network, and a candidate detection head network, and the number of network modules and the number of channels of each candidate detection model are randomly generated; For any candidate detection model, training the candidate detection model based on a first training sample image to adjust the model of the candidate detection model to obtain a trained candidate detection model; Selecting the trained candidate detection model that meets the set conditions as the target detection model.
2. The method according to claim 1, wherein, The step of selecting the trained candidate detection model that meets the set conditions as the target detection model includes: Selecting a to-be-verified trained candidate detection model from the plurality of trained candidate detection models according to an evolutionary algorithm; Performing accuracy verification on the to-be-verified trained candidate detection model based on a verification sample image to obtain the model accuracy of the to-be-verified trained candidate detection model; Selecting the to-be-verified trained candidate detection model whose model performance parameters and the detection model accuracy meet the conditions as the target detection model.
3. The method according to claim 1, wherein, After the target detection model is determined, it further includes: Retraining the target detection model based on a second training sample image to obtain a final target detection model.
4. The method according to claim 1, wherein The method further includes: For the i th neural network architecture search, determine the network modules that have been searched in the search space in the previous i -1 times, where the i is a positive integer greater than 1 and not greater than the set total number of searches; Determining the remaining network modules in the search space based on the searched network modules, and searching in the remaining network modules.
5. The method according to claim 3, wherein The step of retraining the target detection model based on the second training sample image includes: Deleting the set convolutional layer in the target detection model; Retraining the target detection model with the set convolutional layer deleted based on the second training sample image.
6. The method according to claim 1, wherein The network module includes a plurality of channels, and the weights of the maximum number of channels are shared by the plurality of channels. The method further includes: Randomly selecting channels for the first network module during the neural network structure search process; During training, the selected channels share the weights of the maximum number of channels for training.
7. The method according to claim 6, wherein, Adjusting the model of the candidate detection model includes: In response to the number of selected channels being less than the maximum number of channels, splitting out the weights corresponding to the randomly selected channels from the weights of the maximum number of channels; Performing gradient update on the weights corresponding to the randomly selected channels.
8. A neural network search device for a target detection model, comprising: An acquisition module for acquiring a search space, where the search space includes a plurality of network modules, and the plurality of network modules are used to form a backbone network, a detection neck network, and a detection head network; A search module, configured to perform multiple neural network structure searches on the network modules in the search space to obtain multiple candidate detection models, where the candidate detection models include candidate backbone networks, candidate detection neck networks, and candidate detection head networks, and the number of network modules and the number of channels of each candidate detection model are randomly generated; A training module, configured to, for any candidate detection model, train the candidate detection model based on the first training sample images to adjust the candidate detection model and obtain a trained candidate detection model; A determination module, configured to select a trained candidate detection model that meets the set conditions as the target detection model; The search module is specifically configured to: Search in the currently searchable network modules in the search space to obtain a first network module for constructing the candidate detection model; Replace the remaining second network modules in the searchable network modules with set convolutional layers; Generate multiple candidate detection models based on the first network module and the set convolutional layers corresponding to the second network modules.
9. The device according to claim 8, wherein The determination module is further configured to: Select a trained candidate detection model to be verified from the multiple trained candidate detection models according to an evolutionary algorithm; Perform accuracy verification on the trained candidate detection model to be verified based on verification sample images to obtain the model accuracy of the trained candidate detection model to be verified; Select a trained candidate detection model to be verified whose model performance parameters and detection model accuracy meet the conditions as the target detection model.
10. The device according to claim 8, wherein, The determination module is further configured to: Retrain the target detection model based on the second training sample images to obtain a final target detection model.
11. The device according to claim 8, wherein, The apparatus is further configured to: For the i -th neural network architecture search, determine the network modules that have been searched in the search space in the previous i - 1 times, where the i is a positive integer greater than 1 and not greater than the set total number of searches; Determine the remaining network modules in the search space based on the searched network modules, and perform a search in the remaining network modules.
12. The apparatus according to claim 10, wherein, The determination module is further configured to: Delete the set convolutional layers in the target detection model; Retrain the target detection model with the set convolutional layers deleted based on the second training sample images.
13. The device according to claim 8, wherein, The network module includes multiple channels, and the weights of the maximum number of channels are shared by the multiple channels. The determination module is further configured to: Randomly select channels for the first network module during the neural network structure search; During training, the selected channels share the weights of the maximum number of channels for training.
14. The device according to claim 13, wherein, The training module is further configured to: In response to the number of selected channels being less than the maximum number of channels, split out the weights corresponding to the randomly selected channels from the weights of the maximum number of channels; Perform gradient update on the weights corresponding to the randomly selected channels.
15. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the neural network search method of the target detection model according to any one of claims 1-7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the neural network search method of the object detection model according to any one of claims 1-7.
17. A computer program product, comprising a computer program which, when executed by a processor, implements the neural network search method of the object detection model according to any one of claims 1-7.
Citation Information
Patent Citations
Infrared video multi-target tracking method
CN113808164A