Target detection method and system based on neural network architecture search

Through neural network architecture search technology, an object detection neural network is automatically built, which solves the problems of robot object detection accuracy and speed, and achieves efficient and accurate target recognition and capture.

CN120279249APending Publication Date: 2025-07-08FOSHAN UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510346404.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing robot object detection algorithms rely on manual design and are difficult to adapt to diverse objects. The detection accuracy is limited and the speed is slow, which affects real-time detection performance.

Method used

The object detection method based on neural network architecture search is adopted, combined with the NAS-CSP module and the SEC-Blcok module, an object detection neural network model is automatically constructed, and the detection accuracy and efficiency are improved through feature extraction, multi-scale fusion and semantic information processing.

Benefits of technology

The automated construction of the target detection network is realized, the detection accuracy and efficiency are improved, the manual design workload is reduced, time and cost are saved, and the stability of robot grabbing is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279249A_ABST
    Figure CN120279249A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method and system based on neural network architecture search, and the method comprises the steps: obtaining an RGB image of a target object, carrying out the image data preprocessing, and obtaining a preprocessed RGB image of the target object; a neural network architecture search method is introduced, and a target detection neural network model is constructed in combination with an NAS-CSP module and an SEC-Blcok module; and performing target detection on the preprocessed target object RGB image based on the target detection neural network model to obtain a target object detection result. According to the invention, through a neural network architecture search technology, the neural network architecture can be automatically adjusted according to the characteristics of the specific object, so that the detection efficiency and precision of the target object are improved. The target detection method and system based on neural network architecture search can be widely applied to the technical field of robot target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot target detection, and in particular to an object detection method and system based on neural network architecture search. Background Art

[0002] In real-world applications, robots often face challenges from object diversity and environmental complexity when performing target detection. Faced with objects of different sizes in a chaotic background, relying solely on traditional object detection algorithms is difficult to meet the requirements, and it is also necessary to consider the size, physical properties of the objects and the environmental factors in which they are located. The design of traditional object detection CNN models requires relying on manual experience and a large amount of time investment, and their ability to be transferred to other tasks is limited, that is, they rely on manual design, which requires a large amount of time to build a neural network suitable for object detection. This dependence means that it may be difficult for robots to accurately identify targets when dealing with diverse objects. In contrast, neural network architecture search (NAS) can automatically design models for specific tasks. In recent years, researchers have combined deep learning with NAS technology and applied it to the object detection task of robots, significantly improving the overall grasping performance of robots through automatically constructed neural networks. However, existing algorithms are usually designed for general tasks and lack optimization for specific objects, so the detection accuracy is limited, which may lead to failures when the robot grasps, and due to the complex structure and numerous parameters, traditional object detection algorithms may reduce the detection speed, affecting the real-time detection performance of the robot, and thus reducing the probability of successful grasping. Summary of the Invention

[0003] In order to solve the above technical problems, the purpose of the present invention is to provide an object detection method and system based on neural network architecture search, which can automatically adjust the neural network architecture according to the characteristics of specific objects through neural network architecture search technology, thereby improving the detection efficiency and accuracy of target objects.

[0004] The first technical solution adopted by the present invention is: an object detection method based on neural network architecture search, comprising the following steps:

[0005] Obtain the RGB image of the target object and perform image data preprocessing to obtain the preprocessed RGB image of the target object;

[0006] Introduce the neural network architecture search method, combine the NAS-CSP module and the SEC-Block module to construct an object detection neural network model;

[0007] Perform object detection on the preprocessed RGB image of the target object based on the object detection neural network model to obtain the object detection result of the target object.

[0008] Furthermore, the target detection neural network model specifically includes a feature extraction module, a multi-scale fusion module, an SEC-Block module, and a head module. The output end of the feature extraction module is connected to the input end of the multi-scale fusion module. The output end of the multi-scale fusion module is connected to the first input end of the SEC-Block module. The output end of the SEC-Block module is connected to the input end of the head module. The output end of the head module is connected to the second input end of the SEC-Block module, where:

[0009] The feature extraction module includes a CBM module and several NAS-CSP modules;

[0010] The multi-scale fusion module includes a first CBL module, a first max pooling layer, a second max pooling layer, a third max pooling layer, and a second CBL module. Among them, the output end of the first CBL module is respectively connected to the input ends of the first max pooling layer, the second max pooling layer, and the third max pooling layer. The output ends of the first max pooling layer, the second max pooling layer, and the third max pooling layer are all connected to the input end of the second CBL module;

[0011] The SEC-Block module includes a first CBL layer, a second CBL layer, and an SE attention mechanism layer. Among them, the output end of the SE attention mechanism layer is respectively connected to the input ends of the first CBL layer and the second CBL layer;

[0012] The head module includes a third CBL layer and a convolutional layer.

[0013] Furthermore, the NAS-CSP module specifically includes a first CBM layer, a second CBM layer, a neural network architecture search layer, a third CBM layer, a fourth CBM layer, and a fifth CBM layer. Among them, the first output end of the first CBM layer is connected to the input end of the second CBM layer. The output end of the second CBM layer is connected to the input end of the neural network architecture search layer. The output end of the neural network architecture search layer is connected to the input end of the third CBM layer. The second output end of the first CBM layer is connected to the input end of the fourth CBM layer. The output ends of the third CBM layer and the fourth CBM layer are connected to the input end of the fifth CBM layer.

[0014] Further, the neural network architecture search layer includes a first normal unit, a second normal unit, a first attenuation unit, a third normal unit, a fourth normal unit, and a second attenuation unit. The first normal unit, the second normal unit, the first attenuation unit, the third normal unit, the fourth normal unit, and the second attenuation unit are connected in sequence. Among them, the first normal unit, the second normal unit, the third normal unit, and the fourth normal unit constitute a first sub-search space, and the first attenuation unit, the third normal unit, the fourth normal unit, and the second attenuation unit constitute a second sub-search space.

[0015] Further, the step of performing object detection on the preprocessed RGB image of the target object by the object detection neural network model to obtain the object detection result of the target object specifically includes:

[0016] Input the preprocessed RGB image of the target object into the object detection neural network model;

[0017] Based on the feature extraction module of the object detection neural network model, perform feature extraction processing on the preprocessed RGB image of the target object to obtain the feature image of the target object;

[0018] Based on the multi-scale fusion module of the object detection neural network model, perform feature fusion processing on the feature image of the target object to obtain the preliminarily fused feature image of the target object;

[0019] Based on the SEC-Blcok module of the object detection neural network model, perform semantic information feature fusion processing on the preliminarily fused feature image of the target object to obtain the fused feature image of the target object;

[0020] Based on the head module of the object detection neural network model, perform prediction classification on the fused feature image of the target object to obtain the object detection result of the target object.

[0021] Further, the step of performing feature extraction processing on the preprocessed RGB image of the target object by the feature extraction module of the object detection neural network model to obtain the feature image of the target object specifically includes:

[0022] Input the preprocessed RGB image of the target object into the feature extraction module of the object detection neural network model;

[0023] Based on the CBM module of the feature extraction module, perform local feature extraction processing on the preprocessed RGB image of the target object to obtain the preliminary feature image of the target object;

[0024] The NAS-CSP module based on the feature extraction module performs global feature extraction and splicing processing on the preliminary target object feature image to obtain the target object feature image.

[0025] Furthermore, the step of the NAS-CSP module based on the feature extraction module performing global feature extraction and splicing processing on the preliminary target object feature image to obtain the target object feature image specifically includes:

[0026] Input the preliminary target object feature image into the NAS-CSP module of the feature extraction module;

[0027] Based on the first CBM layer, the second CBM layer, the neural network architecture search layer, and the third CBM layer of the NAS-CSP module, perform feature extraction processing on the preliminary target object feature image in sequence to obtain the first target object feature image;

[0028] Based on the fourth CBM layer of the NAS-CSP module, perform feature extraction processing on the preliminary target object feature image in sequence to obtain the second target object feature image;

[0029] Based on the fifth CBM layer of the NAS-CSP module, perform splicing processing on the first target object feature image and the second target object feature image to obtain the target object feature image.

[0030] Furthermore, the step of the multi-scale fusion module based on the target detection neural network model performing feature fusion processing on the target object feature image to obtain the preliminarily fused target object feature image specifically includes:

[0031] Input the target object feature image into the multi-scale fusion module of the target detection neural network model;

[0032] Based on the first CBL module of the multi-scale fusion module, perform feature extraction processing on the target object feature image to obtain the refined target object feature image;

[0033] Based on the first max pooling layer, the second max pooling layer, and the third max pooling layer of the multi-scale fusion module, perform pooling processing on the refined target object feature image to obtain the pooled refined target object feature image;

[0034] Based on the second CBL module of the multi-scale fusion module, perform feature fusion processing on the pooled refined target object feature image to obtain the preliminarily fused target object feature image.

[0035] Further, for the SEC-Block module based on the object detection neural network model, the step of performing semantic information feature fusion processing on the preliminarily fused object feature image to obtain the fused object feature image specifically includes:

[0036] Input the preliminarily fused object feature image into the SEC-Block module of the object detection neural network model;

[0037] Based on the SE attention mechanism layer of the SEC-Block module, perform compression activation processing on the preliminarily fused object feature image to obtain the compressed and activated object feature image;

[0038] Based on the first CBL layer and the second CBL layer of the SEC-Block module, perform feature extraction and fusion processing on the compressed and activated object feature image to obtain the fused object feature image.

[0039] The second technical solution adopted by the present invention is: an object detection system based on neural network architecture search, including:

[0040] The first module 201 is used to obtain the RGB image of the object and perform image data preprocessing to obtain the preprocessed RGB image of the object;

[0041] The second module 202 is used to introduce the neural network architecture search method, combine the NAS-CSP module and the SEC-Block module to construct an object detection neural network model;

[0042] The third module 203 is used to perform object detection on the preprocessed RGB image of the object based on the object detection neural network model to obtain the object detection result.

[0043] The beneficial effects of the method and system of the present invention are as follows: By obtaining the RGB image of the object and performing image data preprocessing, and then introducing the neural network architecture search method, combining the NAS-CSP module and the SEC-Block module to construct an object detection neural network model, the present invention adopts the neural network architecture search technology to realize the automatic construction of the object detection network, thereby improving the performance and efficiency of the network, saving a large amount of time and cost, and being able to automatically adjust the neural network architecture according to the characteristics of specific objects to improve the accuracy of object detection. The SEC-Block module is constructed to strengthen the useful features of the object and suppress the unimportant features. Finally, object detection is performed on the preprocessed RGB image of the object based on the object detection neural network model to improve the detection accuracy of the object. Description of the Drawings

[0044] Figure 1It is a flowchart of the steps of an object detection method based on neural network architecture search according to the present invention;

[0045] Figure 2 It is a block diagram of the structure of an object detection system based on neural network architecture search according to the present invention;

[0046] Figure 3 It is a schematic structural diagram of an object detection neural network model provided by a specific embodiment of the present invention;

[0047] Figure 4 It is a schematic structural diagram of the CBM module provided by a specific embodiment of the present invention;

[0048] Figure 5 It is a schematic structural diagram of the CBL module provided by a specific embodiment of the present invention;

[0049] Figure 6 It is a schematic structural diagram of the NAS-CSP module provided by a specific embodiment of the present invention;

[0050] Figure 7 It is a schematic structural diagram of the neural network architecture search layer provided by a specific embodiment of the present invention;

[0051] Figure 8 It is a schematic diagram of the search space Cell structure provided by a specific embodiment of the present invention. Specific Embodiments

[0052] The following further elaborates on the present invention in detail in conjunction with the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0053] First of all, it should be noted that the application of artificial intelligence technology in modern industrial production has significantly improved the adaptability of tasks and manufacturing processes, and has become a core part of intelligent manufacturing. Among the many research fields of artificial intelligence, robotics has attracted much attention due to its important impact on social progress. With the wide application of robots in various fields, the demand for their precise operation is also increasing, which makes the grasping detection ability of robots particularly important. The vision system is one of the main means for robots to perceive the environment. It can provide accurate position information of the target object for the robot, help predict the best grasping area, and evaluate the suitability of the grasping position and size. All of this depends on the stable and accurate information feedback of the vision system. At present, the research focus of robots is concentrated on the development of vision technology, which plays a crucial role in many key fields such as industrial manufacturing, autonomous driving vehicles, and remote medical services. Neural architecture search (NAS) is an automated technology for generating efficient convolutional neural network (CNN) architectures for detection tasks. The core process of NAS includes three main parts: constructing the search space, selecting the search strategy, and formulating the performance evaluation plan. The search space covers all possible network structures, and there are usually two construction methods: one is to define the search space based on existing neural network structures to search the entire network; the other is to construct the search space based on cells to find the optimal cell structure. The search strategy involves how to find the best network structure in this space, and common methods include using techniques such as reinforcement learning or gradient descent. The performance evaluation strategy is used to predict and measure the generalization ability of the sampled network structures during the NAS process.

[0054] Based on this, referring to Figure 1 , the present invention provides a target detection method based on neural architecture search, and this method includes the following steps:

[0055] S100. Obtain the RGB image of the target object and perform image data preprocessing to obtain the preprocessed RGB image of the target object;

[0056] Specifically, use the Intel RealSense D435i camera to capture the RGB image of the target object by using its active infrared binocular stereo ranging function. The collected images are used to make the training data set, and the scale of the data set is enlarged through data augmentation techniques such as image flipping, cropping, and translation.

[0057] S200. Introduce the neural architecture search method, combine the NAS-CSP module and the SEC-Blcok module to construct a target detection neural network model;

[0058] Specifically, as Figure 3As shown in the figure, the target detection neural network model specifically includes a feature extraction module, a multi-scale fusion module, an SEC-Block module, and a head module. The output end of the feature extraction module is connected to the input end of the multi-scale fusion module. The output end of the multi-scale fusion module is connected to the first input end of the SEC-Block module. The output end of the SEC-Block module is connected to the input end of the head module. The output end of the head module is connected to the second input end of the SEC-Block module. Among them, the feature extraction module includes a CBM module and several NAS-CSP modules; the multi-scale fusion module includes a first CBL module, a first max pooling layer, a second max pooling layer, a third max pooling layer, and a second CBL module. Among them, the output end of the first CBL module is respectively connected to the input ends of the first max pooling layer, the second max pooling layer, and the third max pooling layer. The output ends of the first max pooling layer, the second max pooling layer, and the third max pooling layer are all connected to the input end of the second CBL module; the SEC-Block module includes a first CBL layer, a second CBL layer, and an SE attention mechanism layer. Among them, the output end of the SE attention mechanism layer is respectively connected to the input ends of the first CBL layer and the second CBL layer; the head module includes a third CBL layer and a convolutional layer.

[0059] Among them, it should be noted that the SEC-Block module integrates the SE attention mechanism and the CBL module, aiming to strengthen the extraction of the important information part of the input image data, effectively combine the semantic information above, and promote feature fusion.

[0060] In this embodiment, as Figure 4 shown, the CBM (Conv + Batch Normalization + Mish) module is a module composed of convolution, normalization, and the Mish activation function. The convolution part uses the convolution kernel to perform dot product calculations with the input data to generate feature maps, and these feature maps contain the important information of the input image; the normalization part normalizes the output of the convolutional layer, so that the data of each batch follows the same distribution. Such a setting helps to accelerate the convergence speed of the network, improve the stability of the network, and reduce the risk of overfitting; the Mish activation function is to introduce non-linearity, so that our network can learn complex patterns and improve the robustness of the model.

[0061] As Figure 5As shown, the function of CBL (Conv + Batch Normalization + Leak relu) is the same as that of CBM, except that the activation function of CBL uses Leak relu to make slightly different non-linear changes to the input data.

[0062] As Figure 6 shown, the NAS-CSP module specifically includes a first CBM layer, a second CBM layer, a neural network architecture search layer, a third CBM layer, a fourth CBM layer and a fifth CBM layer. Among them, the first output end of the first CBM layer is connected to the input end of the second CBM layer, the output end of the second CBM layer is connected to the input end of the neural network architecture search layer, the output end of the neural network architecture search layer is connected to the input end of the third CBM layer, the second output end of the first CBM layer is connected to the input end of the fourth CBM layer, and the output ends of the third CBM layer and the fourth CBM layer are connected to the input end of the fifth CBM layer.

[0063] Among them, it should be noted that the NAS-CSP (NAS-Cross Stage Partial) module divides the input feature layer into two parts for processing, and then splices the processed features. Such a design can reduce redundant calculations in the network, effectively capture details and context information in the input data, and improve the accuracy of object detection.

[0064] As Figure 7 shown, the neural network architecture search layer includes a first normal unit, a second normal unit, a first attenuation unit, a third normal unit, a fourth normal unit and a second attenuation unit. The first normal unit, the second normal unit, the first attenuation unit, the third normal unit, the fourth normal unit and the second attenuation unit are connected in sequence. Among them, the first normal unit, the second normal unit, the third normal unit and the fourth normal unit form a first sub-search space, and the first attenuation unit, the third normal unit, the fourth normal unit and the second attenuation unit form a second sub-search space.

[0065] S300. Perform object detection on the preprocessed RGB image of the target object based on the object detection neural network model to obtain the object detection result of the target object.

[0066] Specifically, first of all, it should be noted that the embodiment of the present invention receives a three-channel RGB image with a size of 416×416 pixels as input, and then uses a feature extraction module to obtain the global and local features of the image. In the Backbone stage, first, a CBM (Conv+Batch Normalization+Mish) module is used for preliminary feature extraction, and then n NAS-CSP modules integrating NAS technology are used to further refine the feature extraction process; after that, the extracted features are deeply fused through a multi-scale fusion module. In the Neck stage, first, it is processed by three CBL (Conv+BatchNormalization+Leakly Relu) modules, then through three max-pooling layers of different sizes, stacked with the original input image, and output the fused features after being processed by another three CBL modules; finally, the predicted bounding boxes are output through the head module. In the Head stage, it will first pass through a CBL module and a convolutional layer, and in the transition part between the Neck and the Head, it will first perform an upsampling operation through the SEC-Block.

[0067] S310. Input the preprocessed RGB image of the target object into the target detection neural network model;

[0068] S320. Based on the feature extraction module of the target detection neural network model, perform feature extraction processing on the preprocessed RGB image of the target object to obtain the feature image of the target object;

[0069] Specifically, input the preprocessed RGB image of the target object into the feature extraction module of the target detection neural network model; based on the CBM module of the feature extraction module, perform local feature extraction processing on the preprocessed RGB image of the target object to obtain a preliminary feature image of the target object; based on the NAS-CSP module of the feature extraction module, perform global feature extraction and splicing processing on the preliminary feature image of the target object to obtain the feature image of the target object.

[0070] It should also be noted that the preliminary feature image of the target object is input into the NAS-CSP module of the feature extraction module; based on the first CBM layer, the second CBM layer, the neural network architecture search layer, and the third CBM layer of the NAS-CSP module, perform feature extraction processing on the preliminary feature image of the target object in sequence to obtain the first feature image of the target object; based on the fourth CBM layer of the NAS-CSP module, perform feature extraction processing on the preliminary feature image of the target object in sequence to obtain the second feature image of the target object; based on the fifth CBM layer of the NAS-CSP module, perform splicing processing on the first feature image of the target object and the second feature image of the target object to obtain the feature image of the target object.

[0071] In this embodiment, the diversity of objects and complex background environments pose challenges to object recognition. Traditional convolution operations often ignore the multi-scale features of objects and are difficult to effectively filter out background interference. To accurately recognize diverse objects in complex environments, the embodiment of the present invention integrates NAS technology in the feature extraction module to accurately recognize specific targets through automated neural network architecture design, which enhances the feature extraction ability of the model and improves the accuracy of recognition. In the feature extraction module, the CSP module is a cross-stage partial connection module, which includes two parts: feature extraction and residual connection. Among them, the feature extraction stage can reduce the number of channels while extracting features, effectively reducing the computational amount; the residual connection stage helps to prevent the vanishing gradient. In specific implementation, the embodiment of the present invention integrates NAS technology in the feature extraction stage of the CSP module to form a new module NAS-CSP, which can automatically construct a neural network suitable for the task for specific objects and further enhance the feature extraction ability while keeping the number of parameters roughly unchanged.

[0072] Among them, the network structure of NAS is as Figure 7 shown. It includes a Normal Cell, a Reduce Cell, and input and output. Its processing flow is: the input image first passes through two Normal Cells, then through a Reduce Cell, then through two more Normal Cells and a Reduce Cell, and finally outputs the result. Usually, the initially defined search space by artificial is relatively wide and requires high computing resources. By adopting an adaptive search space reduction strategy, NAS can automatically evaluate the performance of cells in the initial space and gradually eliminate poorly performing cells, finally forming a small and efficient search space, thereby improving the search efficiency and performance of NAS. The cell structure in the search space defined by the embodiment of the present invention is as Figure 8 shown, including basic operations such as 3×3 convolution, 5×5 convolution, 3×3 depth convolution, normalization, Mish activation function, and max pooling. Among them, Figure 8 Cells (A), (B), (C), and (D) in belong to sub-search space A, and the image size remains unchanged after being processed by these cells; Figure 8 Cells (E), (F), (G), and (H) in belong to sub-search space B, and the image size is halved after being processed by these cells.

[0073] S330. The multi-scale fusion module based on the object detection neural network model performs feature fusion processing on the target object feature image to obtain a preliminarily fused target object feature image;

[0074] Specifically, the target object feature image is input into the multi-scale fusion module of the target detection neural network model; based on the first CBL module of the multi-scale fusion module, feature extraction processing is performed on the target object feature image to obtain a refined target object feature image; based on the first max pooling layer, the second max pooling layer, and the third max pooling layer of the multi-scale fusion module, pooling processing is performed on the refined target object feature image to obtain a pooled refined target object feature image; based on the second CBL module of the multi-scale fusion module, feature fusion processing is performed on the pooled refined target object feature image to obtain a preliminarily fused target object feature image.

[0075] In this embodiment, the multi-scale fusion technology plays a crucial role in the fields of deep learning image processing and computer vision. It enhances the model's efficiency by integrating features at different scales. Although the feature extraction modules can capture detailed features, they often fail to fully utilize these features, thus unable to strengthen the mutual relationship between different features. Spatial Pyramid Pooling (SPP) allows the network to process input images of different sizes without cropping or resizing the images. SPP performs pooling on the feature map by applying pooling windows of different sizes and merges these feature vectors to generate a fixed-length output vector. The main advantage of the SPP layer is its ability to maintain spatial information and generate feature descriptors of a unified length for input images of different sizes, which enables the model to adapt to different-sized image inputs while maintaining its performance. To further enhance the output of the feature extraction module, we adopted the SPP method to fuse features. Specifically, first, the features are processed through three CBL modules, then these features are respectively sent into max pooling layers of three different sizes, namely 5×5, 9×9, and 13×13, for pooling. After that, these pooled features are stacked with the original features, and finally, the merged features are output.

[0076] S340. Based on the SEC-Block module of the target detection neural network model, semantic information feature fusion processing is performed on the preliminarily fused target object feature image to obtain a fused target object feature image.

[0077] Specifically, the preliminarily fused target object feature image is input into the SEC-Block module of the target detection neural network model; based on the SE attention mechanism layer of the SEC-Block module, compression activation processing is performed on the preliminarily fused target object feature image to obtain a compressed and activated target object feature image; based on the first CBL layer and the second CBL layer of the SEC-Block module, feature extraction and fusion processing are performed on the compressed and activated target object feature image to obtain a fused target object feature image.

[0078] In this embodiment, as Figure 3As shown in the figure, the structure of the SEC-Block consists of several key steps: First, it is processed through an SE attention module, which focuses on enhancing the representation ability of features; then, the processed features are sent into two CBL modules for further feature extraction; finally, these features are merged. The SE attention mechanism plays a role in strengthening useful features and suppressing unimportant features in this process, and it achieves this by learning the mutual relationships between features in the channel dimension.

[0079] S350. The head module based on the object detection neural network model predicts and classifies the fused object feature image to obtain the object detection result.

[0080] Specifically, in the object detection model, the role of the Head part is crucial. It is located at the end of the network and is responsible for the final prediction tasks, including the regression of bounding boxes and the classification of categories. When the input image passes through the feature extraction and multi-scale fusion module, its resolution is reduced from 416×416 to 56×56. This reduction in size helps the network focus on extracting higher-level features, but it may also lead to the loss of image details. To improve the image's resolvability and retain its spatial features, the embodiment of the present invention designs an SEC-Block module for image upsampling processing. In this way, at the end of the network, after passing through the CBL module and the convolutional layer, the Head part can generate prediction boxes of three different sizes: 13×13, 26×26, and 52×52.

[0081] In summary, the embodiment of the present invention is applicable to detecting multi-scale target objects in unstructured scenarios, such as the 3C part assembly scenario, where the parts are of different sizes and are randomly distributed. Traditional detection algorithms are difficult to accurately identify each object. Therefore, the robot vision module is required to accurately identify the target in a complex environment and transmit the world coordinates to the robotic arm. After being converted into rotation and depth matrices through hand-eye calibration, the robotic arm is guided above the target to complete the grasping. It can be seen from this that object detection plays an important guiding role in the subsequent grasping task. The solution of the embodiment of the present invention incorporates NAS technology, making the object detection network more diverse, specifically improving the recognition accuracy, and thus enhancing the stability of robot grasping.

[0082] The embodiment of the present invention has the following improvement points compared with the prior art:

[0083] 1) It adopts the neural network architecture search technology to realize the automatic construction of the object detection network, which not only reduces the workload of manual design, but also improves the performance and efficiency of the network, thus saving a large amount of time and cost.

[0084] 2) It can automatically adjust the neural network architecture according to the characteristics of specific objects to improve the accuracy of object detection.

[0085] 3) The redundant structure is removed by the automatically constructed neural network architecture, reducing the number of parameters required for network training, making the network more lightweight, accelerating the training speed, and improving the efficiency of real-time detection.

[0086] Referring to Figure 2 , an object detection system based on neural network architecture search, comprising:

[0087] The first module 201 is used to obtain the RGB image of the target object and perform image data preprocessing to obtain the preprocessed RGB image of the target object;

[0088] The second module 202 is used to introduce the neural network architecture search method, combine the NAS-CSP module and the SEC-Block module to construct an object detection neural network model;

[0089] The third module 203 is used to perform object detection on the preprocessed RGB image of the target object based on the object detection neural network model to obtain the object detection result.

[0090] The content in the above method embodiments is applicable to the present system embodiment. The functions specifically implemented by the present system embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0091] The above is a specific description of the preferred embodiment of the present invention. However, the present invention is not limited to the described embodiment. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A target detection method based on neural network architecture search, characterized in that, It includes the following steps: Obtain the RGB image of the target object and perform preprocessing on the image data to obtain the preprocessed RGB image of the target object; Introduce the neural network architecture search method, combine the NAS-CSP module and the SEC-Blcok module to construct a target detection neural network model; Perform target detection on the preprocessed RGB image of the target object based on the target detection neural network model to obtain the target object detection result.

2. The object detection method based on neural network architecture search according to claim 1, wherein The target detection neural network model specifically includes a feature extraction module, a multi-scale fusion module, a SEC-Blcok module, and a head module. The output end of the feature extraction module is connected to the input end of the multi-scale fusion module. The output end of the multi-scale fusion module is connected to the first input end of the SEC-Blcok module. The output end of the SEC-Blcok module is connected to the input end of the head module. The output end of the head module is connected to the second input end of the SEC-Blcok module, where: The feature extraction module includes a CBM module and several NAS-CSP modules; The multi-scale fusion module includes a first CBL module, a first max pooling layer, a second max pooling layer, a third max pooling layer, and a second CBL module. Among them, the output end of the first CBL module is respectively connected to the input ends of the first max pooling layer, the second max pooling layer, and the third max pooling layer. The output ends of the first max pooling layer, the second max pooling layer, and the third max pooling layer are all connected to the input end of the second CBL module; The SEC-Blcok module includes a first CBL layer, a second CBL layer, and an SE attention mechanism layer. Among them, the output end of the SE attention mechanism layer is respectively connected to the input ends of the first CBL layer and the second CBL layer; The head module includes a third CBL layer and a convolutional layer.

3. The object detection method based on neural network architecture search according to claim 2, wherein The NAS-CSP module specifically includes a first CBM layer, a second CBM layer, a neural network architecture search layer, a third CBM layer, a fourth CBM layer, and a fifth CBM layer. Among them, the first output end of the first CBM layer is connected to the input end of the second CBM layer. The output end of the second CBM layer is connected to the input end of the neural network architecture search layer. The output end of the neural network architecture search layer is connected to the input end of the third CBM layer. The second output end of the first CBM layer is connected to the input end of the fourth CBM layer. The output ends of the third CBM layer and the fourth CBM layer are connected to the input end of the fifth CBM layer.

4. The object detection method based on neural network architecture search according to claim 3, wherein, The neural network architecture search layer includes a first normal unit, a second normal unit, a first attenuation unit, a third normal unit, a fourth normal unit, and a second attenuation unit. The first normal unit, the second normal unit, the first attenuation unit, the third normal unit, the fourth normal unit, and the second attenuation unit are connected in sequence. Among them, the first normal unit, the second normal unit, the third normal unit, and the fourth normal unit constitute a first sub-search space, and the first attenuation unit, the third normal unit, the fourth normal unit, and the second attenuation unit constitute a second sub-search space.

5. The object detection method based on neural network architecture search according to claim 4, wherein The step of performing object detection on the preprocessed RGB image of the target object based on the object detection neural network model to obtain the object detection result of the target object specifically includes: Input the preprocessed RGB image of the target object into the object detection neural network model; Based on the feature extraction module of the object detection neural network model, perform feature extraction processing on the preprocessed RGB image of the target object to obtain the feature image of the target object; Based on the multi-scale fusion module of the object detection neural network model, perform feature fusion processing on the feature image of the target object to obtain the preliminarily fused feature image of the target object; Based on the SEC-Block module of the object detection neural network model, perform semantic information feature fusion processing on the preliminarily fused feature image of the target object to obtain the fused feature image of the target object; Based on the head module of the object detection neural network model, perform prediction classification on the fused feature image of the target object to obtain the object detection result of the target object.

6. The object detection method based on neural network architecture search according to claim 5, wherein The step of performing feature extraction processing on the preprocessed RGB image of the target object based on the feature extraction module of the object detection neural network model to obtain the feature image of the target object specifically includes: Input the preprocessed RGB image of the target object into the feature extraction module of the object detection neural network model; Based on the CBM module of the feature extraction module, perform local feature extraction processing on the preprocessed RGB image of the target object to obtain the preliminary feature image of the target object; Based on the NAS-CSP module of the feature extraction module, perform global feature extraction and splicing processing on the preliminary feature image of the target object to obtain the feature image of the target object.

7. The object detection method based on neural network architecture search according to claim 6, wherein, The step of performing global feature extraction and splicing processing on the preliminary feature image of the target object based on the NAS-CSP module of the feature extraction module to obtain the feature image of the target object specifically includes: Input the preliminary feature image of the target object into the NAS-CSP module of the feature extraction module; Based on the first CBM layer, the second CBM layer, the neural network architecture search layer, and the third CBM layer of the NAS-CSP module, perform feature extraction processing on the preliminary feature image of the target object in sequence to obtain the first feature image of the target object; Based on the fourth CBM layer of the NAS-CSP module, perform feature extraction processing on the preliminary feature image of the target object in sequence to obtain the second feature image of the target object. The fifth CBM layer based on the NAS-CSP module splices the first target object feature image and the second target object feature image to obtain the target object feature image.

8. The object detection method based on neural network architecture search according to claim 7, characterized in that, The step of using the multi-scale fusion module based on the object detection neural network model to perform feature fusion processing on the target object feature image to obtain the preliminarily fused target object feature image specifically includes: Input the target object feature image into the multi-scale fusion module of the object detection neural network model; Based on the first CBL module of the multi-scale fusion module, perform feature extraction processing on the target object feature image to obtain the refined target object feature image; Based on the first max pooling layer, the second max pooling layer, and the third max pooling layer of the multi-scale fusion module, perform pooling processing on the refined target object feature image to obtain the pooled refined target object feature image; Based on the second CBL module of the multi-scale fusion module, perform feature fusion processing on the pooled refined target object feature image to obtain the preliminarily fused target object feature image.

9. The object detection method based on neural network architecture search according to claim 8, characterized in that, The step of using the SEC-Block module based on the object detection neural network model to perform semantic information feature fusion processing on the preliminarily fused target object feature image to obtain the fused target object feature image specifically includes: Input the preliminarily fused target object feature image into the SEC-Block module of the object detection neural network model; Based on the SE attention mechanism layer of the SEC-Block module, perform compression activation processing on the preliminarily fused target object feature image to obtain the compressed and activated target object feature image; Based on the first CBL layer and the second CBL layer of the SEC-Block module, perform feature extraction and fusion processing on the compressed and activated target object feature image to obtain the fused target object feature image.

10. An object detection system based on neural network architecture search, characterized in that, It includes the following modules: The first module is used to obtain the RGB image of the target object and perform image data preprocessing to obtain the preprocessed RGB image of the target object; The second module is used to introduce the neural network architecture search method, combine the NAS-CSP module and the SEC-Block module to build an object detection neural network model; The third module is used to perform object detection on the preprocessed RGB image of the target object based on the object detection neural network model to obtain the object detection result.