An Improved YOLOX Object Detection Model Construction Method and Its Application

Through the improved YOLOX object detection model, combined with the SE attention mechanism and ASFF adaptive spatial feature fusion structure, the problem of difficult to identify and locate fruits in complex environments is solved, and high-precision and rapid fruit recognition and positioning are achieved, and the picking success rate is improved.

CN115019302BActive Publication Date: 2025-06-10JIANGSU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210660801.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-13
Publication Date
2025-06-10
Estimated Expiration
2042-06-13

AI Technical Summary

Technical Problem

Existing fruit picking robots are difficult to efficiently identify and locate fruits in complex crop growth environments, especially due to the occlusion of leaves and branches, resulting in a low picking success rate.

Method used

The improved YOLOX object detection model is adopted, and by building a backbone feature extraction network, strengthening feature extraction network and predicting feature network, combining the SE attention mechanism and ASFF adaptive spatial feature fusion structure, fast and high-precision fruit recognition and positioning are achieved.

Benefits of technology

It improves the recognition and positioning accuracy of fruit picking robots in complex environments, enhances the picking success rate, and improves the generalization ability of the model through data enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115019302B_ABST
    Figure CN115019302B_ABST
Patent Text Reader

Abstract

The present invention discloses an improved YOLOX target detection model construction method and its application. This application targets the basic YOLOX target detection model and improves the network, which is mainly divided into three modules: the backbone feature extraction module, the enhanced feature extraction module, and the prediction feature module. An SE attention mechanism module is added on the basis of the backbone feature extraction network. By means of self-learning, the importance degree of each feature channel is obtained, and the mutual dependence relationship between the convolutional feature channels of the modeling network is clearly defined to improve the representation quality generated by the network, so as to screen out the attention for channels and effectively improve the network performance. The enhanced feature extraction network adopts an ASFF adaptive spatial feature fusion structure; based on the constructed improved YOLOX target detection model, the depth information obtained by the RealSense camera is combined to output three-dimensional coordinates to achieve target positioning. In the problem of fruit picking in the natural environment, the present invention has higher recognition accuracy and faster recognition speed compared with traditional methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine vision, and particularly relates to a method for constructing an improved YOLOX target detection model and its application. Background Art

[0002] With the increasing consumption demand for fruits, the planting area and output of fruits have also increased accordingly. The labor force used in the picking operation accounts for about 30% of the total labor force used in the entire production process. If manual picking is used, not only is the labor intensity high, but also the timely harvesting of fruits cannot be guaranteed. To solve the practical problems in agricultural picking, the research and application of fruit picking robots have become an urgent need. Although great progress has been made in the development of picking robots at present, in fact, it is very difficult to commercialize fruit picking robots even abroad. When the current fruit picking robots are put into actual work, there are still great defects.

[0003] The operating environment of picking robots is complex. In addition to being restricted by topographical conditions, the growth environment of crops is directly affected by natural conditions such as seasons and weather, with many uncertain factors, resulting in great picking difficulties. At the same time, most of the operation objects are blocked by leaves and branches, increasing the difficulty of visual recognition and positioning of the robot, reducing the picking success rate, which puts higher requirements on the recognition, positioning and obstacle avoidance functions of the robot. With the rapid development of intelligentization, the research on target recognition and positioning methods in computer vision is beneficial to accelerating the development of the subsequent agricultural system and providing certain theoretical reference for the research of other subsequent agricultural picking robots. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention proposes a method for constructing an improved YOLOX target detection model and a method for fruit recognition and positioning using the improved YOLOX target detection model, and uses the improved YOLOX target detection model to achieve fast and high-precision detection of fruits.

[0005] The technical solution adopted by the present invention is as follows:

[0006] A method for constructing an improved YOLOX target detection model, comprising the following steps:

[0007] Step 1, construct the structure of the improved YOLOX target detection model, including a backbone feature extraction network, an enhanced feature extraction network, and a prediction feature network;

[0008] The main backbone feature extraction network is mainly composed of a Focus network structure, a CBL network structure, a CSPnet network structure, an SPP network structure, and an SE attention mechanism structure. The three SE attention mechanism modules are respectively denoted as the SE-1 network structure, the SE-2 network structure, and the SE-3 network structure. It is sequentially connected by a Focus network structure, a first CBL network structure, a first CBL + CSPnet network structure, a second CBL + CSPnet network structure, an SE-1 network structure, a third CBL + CSPnet network structure, an SE-2 network structure, a second CBL network structure, an SPP network structure, a CSPnet network structure, and an SE-3 network structure. The input is an image input to the Focus network structure, and the SE-1 network structure, the SE-2 network structure, and the SE-3 network structure respectively output a first effective feature layer, a second effective feature layer, and a third effective feature layer.

[0009] The enhanced feature extraction network adopts a PAN + FPN feature pyramid and an ASFF adaptive spatial feature fusion structure to perform feature fusion on feature layers of different shapes. The enhanced feature extraction network is composed of 2 Conv2ds, 4 Concat + CSPnet network structures, 2 upsamplings, 2 downsamplings, and 3 ASFF network structures.

[0010] The prediction feature network receives the output of the enhanced feature extraction network and obtains feature information through the prediction feature network.

[0011] Step 2: Train the improved YOLOX target detection model structure built in Step 1.

[0012] Furthermore, the CBL network structure is composed of a Conv2d, a BN layer, and a SiLU activation function.

[0013] Furthermore, the CSPnet network structure is composed of 3 CBL network structures, a Residual residual block, and a Concat connection layer. The input of the CSPnet network structure is processed in two paths. One path is sequentially processed through the CBL network structure and the Residual residual block, and the other path directly passes through the CBL network structure. The outputs of the two paths are spliced through the Concat connection layer and then input into the CBL network structure to obtain the final output.

[0014] Furthermore, the CBL + CSPnet network structure is a sequentially connected CBL network structure and CSPnet network structure, and the output of the CBL network structure serves as the input of the CSPnet network structure.

[0015] Furthermore, the SE attention mechanism module includes two parts: compression and excitation. The feature map is compressed into a 1×1×C vector through global pooling; weights are generated for each feature channel through learning, and the obtained channel weights are multiplied by the two-dimensional matrix of the corresponding channel of the original feature map, that is, the original feature map is weighted in the channel direction to obtain the output, where C represents the number of channels of the feature map.

[0016] Furthermore, a learnable coefficient is added on the basis of the PAN+FPN feature pyramid fusion to achieve an adaptive fusion effect. The formula is as follows:

[0017]

[0018] Among them, represents the (i, j)-th feature vector of the output feature map Y between channels l ; represents the feature vector at the position (i, j) on the feature map adjusted from level 1 to level l; represents the feature vector at the position (i, j) on the feature map adjusted from level 2 to level l; represents the feature vector at the position (i, j) on the feature map adjusted from level 3 to level l; refers to the spatial importance weights of the feature maps of three different levels to level l adaptively learned by the network. The superscript l represents the feature map of the l-th level.

[0019] Furthermore, the prediction feature network adopts three decoupled heads, and each decoupled head represents a branch, which are respectively represented as:

[0020] The first branch Cls(H, W, C): predicts the category of the target box;

[0021] The second branch Reg(H, W, 4): predicts the coordinate information of the target box;

[0022] The second branch Obj(H, W, 1): judges whether the target box contains an object;

[0023] After the feature information obtained by the three decoupled heads is concatenated and fused through Concat, the prediction information is obtained.

[0024] Furthermore, a depth camera is used to obtain the color map of the fruit to be recognized in the natural scene, and the color map is labeled to form a data set;

[0025] The data set is divided into a training set, a validation set, and a test set according to (training set + validation set): test set = 9:1, and training set: validation set = 9:1, which are used for the training, validation, and testing of the improved YOLOX object detection model structure.

[0026] A fruit recognition and positioning method based on improved YOLOX combined with RealSense, comprising the following steps:

[0027] Step 1, call the RealSense camera to collect the color image and depth image of the fruit to be recognized;

[0028] Step 2, input the color image obtained by the RealSense camera into the trained improved YOLOX target detection model to identify the target fruit and generate a two-dimensional target box.

[0029] Step 3, calculate the pixel coordinates (X, Y) of the center point of the two-dimensional target box by identifying the coordinates of the upper left point and the lower right point of the two-dimensional target box; extract the depth value Z corresponding to the pixel coordinates of the center point of the two-dimensional target box in the aligned depth image, and obtain the three-dimensional coordinates (X, Y, Z) of the center point of the target fruit in real time.

[0030] Step 4, use camera calibration to obtain the camera internal parameter to convert the three-dimensional coordinates (X, Y, Z) of the center point of the target fruit into the camera coordinate system representation.

[0031] Further, the three-dimensional coordinates of the center point of the target fruit in the camera coordinate system are expressed as (X 1 , Y 1 , Z), where,

[0032] X 1 =(X - C x ) / f x * Z

[0033] Y 1 =(Y - C y ) / f y * Z

[0034] Where f x represents the length of the focal length in the x-axis direction described in pixels, and f y represents the length of the focal length in the y-axis direction described in pixels, and (C x , C y ) respectively represent the offsets of the center point of the camera sensor chip in the x and y directions.

[0035] Advantages of the present invention:

[0036] Based on obtaining a large number of original apple images, the present invention adopts data augmentation to improve the generalization ability of the detection model. While ensuring the detection accuracy, the selected target detection network YOLOX has the advantages of fast response, high precision, simple structure and convenient deployment, and adds SE and ASFF modules to the original YOLOX structure, achieving good results in detection speed and accuracy.

[0037] Meanwhile, the present invention adopts the method of obtaining the three-dimensional coordinates of two-dimensional pixel points and selects a depth camera with a short range, which can achieve the best depth positioning detection at any distance between 0.2 and 1.5 meters. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a schematic diagram of the construction method of the improved YOLOX target detection model of the present invention;

[0039] Figure 2 It is a schematic diagram of the structure of the improved YOLOX target detection model of the present invention;

[0040] Figure 3 It is a schematic diagram of the CBL and CSPnet network structures;

[0041] Figure 4 It is a schematic diagram of the network structure of the SE attention mechanism module;

[0042] Figure 5 It is a schematic diagram of the prediction decoupling head structure;

[0043] Figure 6 It is an effect diagram of the improved YOLOX target detection model of the present invention for identifying apples;

[0044] Figure 7 It is a flow chart of the present invention for combining the RealSense camera to output the three-dimensional coordinates of apples. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0046] Combined with the attached Figures 1-5 , a method for constructing an improved YOLOX target detection model of the present application includes the following steps:

[0047] Step 1, construct the structure of the improved YOLOX target detection model. Combined with the attached Figure 2 , the structure of the improved YOLOX target detection model is specifically as follows:

[0048] 1. The backbone feature extraction network is CSPDarknet, which is mainly composed of a Focus network structure, a CBL network structure, a CSPnet network structure, an SPP network structure, and an SE attention mechanism structure. The three SE attention mechanism modules are respectively denoted as the SE-1 network structure, the SE-2 network structure, and the SE-3 network structure. Specifically, it is sequentially connected by a Focus network structure, a first CBL network structure, a first CBL + CSPnet network structure, a second CBL + CSPnet network structure, an SE-1 network structure, a third CBL + CSPnet network structure, an SE-2 network structure, a second CBL network structure, an SPP network structure, a CSPnet network structure, and an SE-3 network structure. The input of the backbone feature extraction network is an image, and the backbone feature extraction network has three outputs, namely, the first effective feature layer output by the SE-1 network structure, the second effective feature layer output by the SE-2 network structure, and the third effective feature layer output by the SE-3 network structure.

[0049] Furthermore, as Figure 3 shown, the CBL network structure is composed of Conv2d, a BN layer, and a SiLU activation function.

[0050] Furthermore, as Figure 3 shown, the CSPnet network structure is composed of three CBL network structures, a Residual residual block, and a Concat connection layer. The input of the CSPnet network structure is processed in two paths. One path is sequentially processed by the CBL network structure and the Residual residual block, and the other path directly passes through the CBL network structure. The outputs of the two paths are spliced by the Concat connection layer and then input into the CBL network structure to obtain the final output.

[0051] Furthermore, the CBL + CSPnet network structure in the above text is the CBL network structure and the CSPnet network structure connected in sequence, and the output of the CBL network structure is used as the input of the CSPnet network structure.

[0052] Furthermore, the Focus network follows the structure of YOLOV5. Before the picture enters the backbone feature extraction network, a value is extracted every other pixel point, and the width and height information is concentrated into the channel space, and the input channels are expanded four times, that is, from the original RGB three-channel mode to twelve channels. Without losing information, it is more conducive to subsequent calculations.

[0053] Furthermore, the SiLU activation function is adopted. SiLU is an improved version of Sigmoid and ReLU. SiLU has the characteristics of having no upper bound but a lower bound, being smooth, and non-monotonic, which helps to prevent the gradient from gradually approaching 0 and causing saturation during slow training. At the same time, its smoothness plays an important role in optimization and generalization, especially in deeper networks. The function expression is as follows:

[0054] SiLU(x) = x · Sigmoid(x)

[0055] The improved YOLOX target detection model structure designed in this application adds the SE attention mechanism module on the basis of the original YOLOX backbone feature extraction network, which is added after the second CBL + CSPnet network structure, the third CBL + CSPnet network structure, and the CSPnet network structure respectively. The SE attention mechanism module belongs to the channel attention module, which obtains the importance of each feature channel through self-learning, and clearly models the mutual dependence between the convolutional feature channels of the network to improve the representation quality generated by the network, so as to screen out the attention for channels and effectively improve the network performance.

[0056] Combined with the attached Figure 4 , the SE attention mechanism mainly includes two parts: Squeeze (compression) and Excitation (excitation). W and H respectively represent the width and height of the feature map, and C represents the number of channels of the feature map. In the Squeeze operation, the feature map is compressed into a 1×1×C vector through global pooling. In the Excitation operation, weights are generated for each feature channel through learning. In the Scale operation, the channel weights obtained in Excitation are multiplied by the two-dimensional matrix of the corresponding channels of the original feature map, that is, weighted in the channel direction of the original feature map, and the output can be directly input to the subsequent layers of the network.

[0057] Furthermore, after the picture undergoes multiple convolutions, three effective feature layers are finally extracted through the backbone feature extraction network. The shapes of the three feature layers are: C1 = (80, 80, 256), C2 = (40, 40, 512), and C3 = (20, 20, 1024).

[0058] 2. The enhanced feature extraction network adopts the PAN + FPN feature pyramid and the ASFF adaptive spatial feature fusion structure to fuse the feature layers with different shapes, so as to achieve the purpose of enhancing feature extraction.

[0059] The enhanced feature extraction network consists of 2 Conv2d, 4 Concat + CSPnet network structures, 2 upsamplings, 2 downsamplings, and 3 ASFF network structures. The three effective feature layers output by the backbone feature extraction network are used as the input of the enhanced feature extraction network. Among them, the third effective feature layer is input into the first Conv2d to obtain the feature layer U1. After U1 is upsampled, it is input into the second Concat + CSPnet together with the second effective feature layer. After splicing and feature extraction, the feature layer U2=(40, 40, 512) is obtained; the feature layer of U2=(40, 40, 512) is input into the second Conv2d to obtain the feature layer U3. After U3 is upsampled, it is input into the fourth Concat + CSPnet together with the first effective feature layer. After splicing and feature extraction, the obtained feature layer is F1=(80, 80, 256);

[0060] The feature layer of F1=(80, 80, 256) is downsampled and input into the third Concat + CSPnet together with U3. After splicing and feature extraction, the obtained feature layer is F2=(40, 40, 512);

[0061] The feature layer of feature F2=(40, 40, 512) is downsampled and input into the first Concat + CSPnet together with U1. After splicing and feature extraction, the feature layer F3=(20, 20, 1024) is obtained.

[0062] Furthermore, F1 is input into ASFF-1, F2 is input into ASFF-2, and F3 is input into ASFF-3 for feature filtering respectively, without changing the size and number of channels of the feature layer. Finally, the shapes of the three feature layers obtained through the enhanced feature extraction network are: F1=(80, 80, 256), F2=(40, 40, 512), F3=(20, 20, 1024).

[0063] Furthermore, the ASFF adaptive spatial feature fusion structure is adopted to solve the problem of inconsistency between different feature scales. Especially for one-stage detectors, this inconsistency will interfere with the gradient calculation during training and reduce the effectiveness of the feature pyramid. The ASFF adaptive spatial feature fusion structure can enable the network to directly learn how to perform spatial filtering on features at other feature levels so as to retain only useful information for combination. Specifically, based on the original PAN + FPN feature pyramid fusion, learnable coefficients are added to achieve the adaptive fusion effect. The formula is as follows:

[0064]

[0065] Among them, represents the output feature map Y between channels lThe (i, j)-th eigenvector represents the eigenvector at position (i, j) on the feature map adjusted from level 1 to level l represents the eigenvector at position (i, j) on the feature map adjusted from level 2 to level l represents the eigenvector at position (i, j) on the feature map adjusted from level 3 to level l refers to the spatial importance weights of the feature maps of three different levels to level l adaptively learned by the network. The superscript l represents the feature map of the l-th level

[0066] 3. As Figure 5 shown, the prediction feature network adopts three decoupled heads to separately implement classification and regression, reducing the number of parameters while improving the convergence speed of the network. Specifically, it includes a 1×1 convolutional layer to reduce the channel size, and then adds two parallel branches, each branch having two 3×3 convolutional layers for classification and regression tasks respectively. The three branches are as follows

[0067] The first branch Cls(H, W, C): predicts the category of the target box. In the present invention, there is only one category of apple

[0068] The second branch Reg(H, W, 4): predicts the coordinate information of the target box, that is, the center point coordinates and width and height of the target box

[0069] The second branch Obj(H, W, 1): determines whether the target box contains an object

[0070] The feature information obtained through the three decoupled heads of the prediction feature network is respectively: H1=(80, 80, 6), H2=(40, 40, 6), H3=(20, 20, 6). After final Concat splicing and fusion, 8400*6 prediction information is obtained. Here, 8400 is the number of prediction boxes, and 6 is the information of each prediction box: (Reg, Obj, Cls).

[0071] Combined with Figure 5 , score screening and non-maximum suppression screening are performed on the above prediction information. Loop through all the pictures, screen out the box with the largest score belonging to the same category within a certain area, and find the boxes with scores greater than the threshold in the picture. In object detection, non-maximum suppression can eliminate redundant detection boxes and find the best object detection position. Sort the categories from largest to smallest according to the scores, take out the box with the largest score each time, calculate the overlap degree with all other prediction boxes, and eliminate those with too large overlap degree. The results after score screening and non-maximum suppression can draw prediction boxes on the pictures to achieve apple target recognition

[0072] Step 2: After completing the construction of the improved YOLOX target detection model structure, it is necessary to train, verify and test the built improved YOLOX target detection model structure.

[0073] Step 2.1, dataset preparation.

[0074] First, a depth camera is used to obtain 800 color images (RGB images) of the fruits to be identified in natural scenes, and Labelimg is used to annotate the single category "apple" to form an apple dataset; in this embodiment, the annotation format adopts the PascalVOC format.

[0075] The images in the data set in step 1 are preprocessed; the preprocessing specifically includes selecting translation, flipping, rotation, contrast, random scaling, brightness and saturation to enhance the data of the apple data set, wherein the brightness and saturation are expanded by 1.5 times, and the rest are expanded by 1 times, so as to increase the data set while improving the stability and robustness of the network. Then the apple data set is resized without distortion to a uniform size of 640*640. At the same time, the Mosaic data enhancement method is used to stitch four pictures, which increases the batch_size in disguise, so that a single GPU can achieve a better training effect. The enhanced data set is divided into a training set, a validation set and a test set, and the division ratio is as follows: (training set + validation set): test set = 9:1, wherein the training set: validation set = 9:1.

[0076] Step 2.2, model training.

[0077] The improved YOLOX object detection model was trained using the training set and validation set of the dataset. The training epochs were set to 300. The backbone feature extraction network was frozen in the first 100 epochs and the initial learning rate was set to 0.01. The backbone feature extraction network was unfrozen in the last 200 epochs and the initial learning rate was set to 0.0001. Object detection was verified on the test set of the constructed apple dataset to complete the object recognition of apples.

[0078] Combination Figure 7 This application proposes a fruit recognition and positioning method based on improved YOLOX combined with RealSense. In this embodiment, the fruit to be picked is apple, but not limited to apple. Similar fruits such as pears and oranges are also suitable for recognition and positioning by this method; the method includes the following steps:

[0079] Step 1: Call the RealSense camera to collect the color image and depth image of the fruit to be recognized. The depth image, also known as the distance image, is the image frame obtained from the depth data stream of the camera. Each pixel point in it represents the distance from the scene captured by the camera to the camera plane. The color image and depth image obtained by the camera are registered and aligned, and the pixel points have a one-to-one correspondence.

[0080] Step 2: Input the color image (RGB image) obtained by the RealSense camera into the trained improved YOLOX object detection model to identify the target fruit and generate a two-dimensional target box, as Figure 6 shown.

[0081] Step 3: Calculate the pixel coordinates (X, Y) of the center point of the two-dimensional target box by identifying the coordinates of the upper left point and the lower right point of the two-dimensional target box; extract the depth value Z corresponding to the pixel coordinates of the center point of the two-dimensional target box in the aligned depth image, and the three-dimensional coordinates (X, Y, Z) of the center point of the target apple can be obtained in real time.

[0082] Step 4: The apple center point coordinates (X, Y) in the three-dimensional coordinates (X, Y, Z) of the target apple center point in Step 3 are pixel coordinates, and the pixel position cannot be used for the robotic arm to grasp. It is necessary to convert the coordinates for use in the camera coordinate system to achieve the true apple positioning. Use the camera calibration to obtain the camera internal parameter for conversion. The camera internal parameter matrix is as follows:

[0083]

[0084] Among them, f x represents the length of the focal length in the x-axis direction described in pixels, f y represents the length of the focal length in the y-axis direction described in pixels, (C x , C y ) represent the offsets of the center point of the camera sensor chip in the x and y directions respectively.

[0085] The expression for converting pixel coordinates to camera coordinates is as follows:

[0086] X 1 =(X - C x ) / f x * Z

[0087] Y 1 =(Y - C y ) / f y * Z

[0088] Finally, the three-dimensional coordinates of the apple center point are obtained as (X 1 , Y 1 , Z), realizing the target positioning of the apple.

[0089] The above embodiments are only used to illustrate the design concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed by the present invention are within the protection scope of the present invention.

Claims

1. An improved method for constructing a YOLOX object detection model, characterized in that, it includes the following steps: Step 1: Build an improved YOLOX object detection model structure, including a backbone feature extraction network, an enhanced feature extraction network, and a prediction feature network; The backbone feature extraction network is mainly composed of a Focus network structure, a CBL network structure, a CSPnet network structure, an SPP network structure, and an SE attention mechanism structure. The 3 SE attention mechanism modules are respectively denoted as the SE-1 network structure, the SE-2 network structure, and the SE-3 network structure; It is successively connected by a Focus network structure, a first CBL network structure, a first CBL + CSPnet network structure, a second CBL + CSPnet network structure, an SE-1 network structure, a third CBL + CSPnet network structure, an SE-2 network structure, a second CBL network structure, an SPP network structure, a CSPnet network structure, and an SE-3 network structure; The input is an image input to the Focus network structure, and the SE-1 network structure, the SE-2 network structure, and the SE-3 network structure respectively output the first effective feature layer, the second effective feature layer, and the third effective feature layer; The enhanced feature extraction network adopts a PAN + FPN feature pyramid and an ASFF adaptive spatial feature fusion structure to perform feature fusion on feature layers of different shapes. The enhanced feature extraction network is composed of 2 Conv2d, 4 Concat + CSPnet network structures, 2 upsamplings, 2 downsamplings, and 3 ASFF network structures; The prediction feature network receives the output of the enhanced feature extraction network and obtains feature information through the prediction feature network; The prediction feature network adopts three decoupled heads, and each decoupled head represents a branch, which are respectively expressed as: The first branch Cls(H, W, C): predicts the category of the target box; The second branch Reg(H, W, 4): predicts the coordinate information of the target box; The second branch Obj(H, W, 1): determines whether the target box contains an object; After the feature information obtained by the three decoupled heads is concatenated and fused through Concat, the prediction information is obtained; Step 2: Train the improved YOLOX object detection model structure built in Step 1.

2. An improved method for constructing a YOLOX object detection model according to claim 1, characterized in that, the CBL network structure is composed of a Conv2d, a BN layer, and a SiLU activation function.

3. An improved method for constructing a YOLOX object detection model according to claim 2, characterized in that, the CSPnet network structure is composed of 3 CBL network structures, a Residual residual block, and a Concat connection layer. The input of the CSPnet network structure is processed in two paths. One path is successively processed by the CBL network structure and the Residual residual block, and the other path directly passes through the CBL network structure. And the outputs of the two paths are concatenated through the Concat connection layer and then input into the CBL network structure to obtain the final output.

4. An improved YOLOX object detection model construction method according to claim 3, characterized in that, the CBL + CSPnet network structure is a CBL network structure and a CSPnet network structure connected in sequence, and the output of the CBL network structure is used as the input of the CSPnet network structure.

5. An improved YOLOX object detection model construction method according to claim 1, characterized in that, the SE attention mechanism module includes two parts: compression and excitation. The feature map is compressed into a vector of 1×1×C through global pooling; weights are generated for each feature channel through learning, and the obtained channel weights are multiplied by the two-dimensional matrix of the corresponding channel of the original feature map, that is, the original feature map is weighted in the channel direction to obtain the output, where C represents the number of channels of the feature map.

6. An improved YOLOX object detection model construction method according to claim 1, characterized in that, a learnable coefficient is added on the basis of the PAN + FPN feature pyramid fusion to achieve an adaptive fusion effect, and the formula is as follows: Among them, represents the (i, j)-th eigenvector of the output feature map Y between channels l ; represents the eigenvector at the position (i, j) on the feature map adjusted from level 1 to level l, represents the eigenvector at the position (i, j) on the feature map adjusted from level 2 to level l, represents the eigenvector at the position (i, j) on the feature map adjusted from level 3 to level l, refers to the spatial importance weights of the feature maps of three different levels to level l adaptively learned by the network, and the superscript l represents the feature map of the l-th level.

7. An improved YOLOX object detection model construction method according to claim 1, characterized in that, a color image of the fruit to be recognized in the natural scene is obtained by using a depth camera, and the color image is labeled to form a data set; the data set is divided into a training set, a validation set and a test set according to (training set + validation set): test set = 9:1, and training set: validation set = 9:1, which are used for the training, validation and testing of the improved YOLOX object detection model structure.

8. A fruit recognition and positioning method based on improved YOLOX combined with RealSense, characterized in that, it includes the following steps: Step 1, call the RealSense camera to collect the color image and depth image of the fruit to be recognized; Step 2, input the color image obtained by the RealSense camera into the improved YOLOX object detection model trained by using an improved YOLOX object detection model construction method as described in claim 1 to recognize the target fruit and generate a two-dimensional target box, Step 3, calculate the pixel coordinates (X, Y) of the center point of the two-dimensional target box by identifying the coordinates of the upper left point and the lower right point of the two-dimensional target box; extract the depth value Z corresponding to the pixel coordinates of the center point of the two-dimensional target box in the aligned depth image, and obtain the three-dimensional coordinates (X, Y, Z) of the center point of the target fruit in real time; Step 4, use camera calibration to obtain the camera internal parameter to convert the three-dimensional coordinates (X, Y, Z) of the center point of the target fruit into the camera coordinate system representation.

9. A fruit recognition and positioning method based on improved YOLOX combined with RealSense according to claim 8, characterized in that, The three-dimensional coordinate representation of the center point of the target fruit in the camera coordinate system is (X 1 , Y 1 , Z), where X 1 = (X - C x ) / f x * Z Y 1 = (Y - C y ) / f y * Z Among them, f x represents the length of the focal length in the x-axis direction described using pixels, and f y represents the length of the focal length in the y-axis direction described using pixels. (C x , C y ) respectively represent the offsets of the center point of the camera's photosensitive chip in the x and y directions.

Citation Information

Patent Citations

  • Unmanned aerial vehicle aerial image target detection method based on improved YOLO V5

    CN113807464A

  • Weld defect detection method based on improved YOLOX

    CN114240821A