Method and device for detecting appearance quality of fresh corn ears

Through the principle of multi-threading and space-division multiplexing combined with ShuffleNetV2 and YOLOv8 networks, the 360° appearance quality detection of fresh corn ears in high-speed motion is achieved, solving the problems of slow detection speed and low accuracy in the existing technology, and real-time and efficient detection effects are achieved.

CN120107956AActive Publication Date: 2025-06-06ACADEMY OF PLANNING & DESIGNING OF THE MINIST OF AGRI
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510175056.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-06
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The prior art is difficult to achieve 360° appearance quality detection of fresh corn ears in high-speed motion states, and traditional machine vision technology has a slow detection speed and low accuracy, which cannot meet the real-time detection needs.

Method used

Using multi-threaded processing strategy and space division multiplexing principle, combined with ShuffleNetV2 and YOLOv8 networks, a lightweight convolution layer and SPPF module are designed to realize 360° rotation and real-time detection of the seeds.

Benefits of technology

Real-time non-destructive testing of the 360° appearance quality of the fruit ear during processing and transportation is realized, which improves the detection speed and accuracy, and is suitable for real-time inspection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107956A_ABST
    Figure CN120107956A_ABST
Patent Text Reader

Abstract

The invention discloses a fresh corn ear appearance quality detection method and device, and belongs to the field of fresh corn appearance detection. The appearance quality detection method comprises the steps that SS-check is introduced into YOLOv8, the SS-check adopts a maximum pooling convolution layer to reduce the dimension of an original fresh corn ear image feature map, a basic unit of ShuffleNetV2 is adopted to carry out feature information extraction on the feature map after dimension reduction, a parameter-free attention mechanism layer is adopted to enhance feature extraction of the space dimension and the channel dimension, and the feature information of the original fresh corn ear image is extracted. A downsampling unit of ShuffleNetV2 is adopted to perform downsampling operation on the extracted feature information, a lightweight convolutional layer is adopted to map a series of sub-feature maps, and then the extracted sub-feature maps are spliced in a channel dimension; according to the method, the model parameter quantity is reduced, the detection speed is fastest, it is ensured that the extracted corn ear feature information is not lost, the detection precision is high, and the method is suitable for real-time detection tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of fresh corn appearance detection, and in particular to a method and a device for detecting the appearance quality of fresh corn ears. Background Art

[0002] my country is a major producer and consumer of fresh corn, and the price of corn ears is determined by taste and appearance. In the actual production process, appearance inspection is mainly based on manual visual inspection. The heavy workload and long-term fatigue work often lead to missed inspections and inadequate inspections, which leads to many problems in the subsequent stages of corn ears. The appearance defects of corn ears can be found all over the surface, and there is a lack of special inspection equipment in the market.

[0003] Fresh corn ears are harvested at the milky stage, with a moisture content of more than 60%, thin skin and high juice content, and are extremely easy to be damaged during the production process. At present, there are related visual inspection technologies for mature ears and seed testing, but the material properties are quite different from fresh corn ears. For example, the patent application number is 2012105716413, and its name is an invention patent application for a method for splicing ordered images of corn ears. It mainly focuses on the splicing of ordered images of mature ears and the calculation of seed testing indicators; the patent application number is 2015104017970, and its name is an invention patent for a method and system for testing the seed of corn ears. A method for rotating ears driven by double rollers is designed, but the rotating state of the ear processing during the entire operation cycle, especially the high-speed operation state, is not suitable for fresh corn; the patent application number is 2010102881528, and its name is an invention patent application for a computer vision detection and grading method and device for the quality of fresh corn ears. It discloses a method for obtaining fresh corn ear images for grading using machine vision, but the ear processing is in a static state and cannot meet the actual requirements of rotating one circle.

[0004] To achieve full detection of the ear surface, it is necessary to combine detection technology with equipment technology to ensure that all surface information is obtained while minimizing hard contact. Existing visual inspection technology lacks material protection, which will cause the material to be in a state of friction or hard contact all the time, and is not suitable for long-term appearance inspection of fresh corn ear processing.

[0005] In addition, traditional machine vision technology is widely used in corn cob and seed detection, which is mainly achieved through a combination of image processing technology, artificial feature extraction, and classification algorithms. Traditional machine learning technology achieves classification goals by extracting effective features such as texture, color, and size of corn cobs and seeds. Although it can achieve automatic recognition and detection, the required manual feature extraction is laborious and time-consuming, and as the cob variety changes, the model robustness deteriorates, the detection accuracy decreases, and the detection speed is slow.

[0006] In the existing technology, NASNet-mobile network can be used to apply transfer learning methods to classify and identify corn varieties, but as the varieties change, the model parameters and structure need to be adjusted, which has certain limitations and cannot be deployed on mobile devices. The improved GAN network can also be used for data enhancement, and then a corn ear detection model is proposed in combination with transfer learning. Although the detection accuracy is higher than other models, it has a large number of parameters, is not easy to train, has a slow detection speed, and cannot be deployed on edge devices. Summary of the invention

[0007] In order to solve the deficiencies in the prior art, the present invention provides a method and device for detecting the appearance quality of fresh corn ears, which are suitable for real-time nondestructive detection of ears in high-speed motion during factory processing production lines, and can realize 360° appearance quality detection of ears during processing and transportation, replacing manual inspection, improving production efficiency and product commerciality, and adapting to the synchronous detection of large size and small surface defect characteristics of fresh corn ears, and a lightweight design of the YOLOv8 backbone network is performed, with the fastest detection speed, which is suitable for real-time detection tasks.

[0008] The present invention adopts the following technical solution:

[0009] The first aspect of the present invention discloses a method for detecting the appearance quality of fresh corn ears, comprising the following steps:

[0010] A set of fresh corn ear images is obtained through a camera; the obtained images are processed: SS-neck is introduced in YOLOv8, and SS-neck uses a maximum pooling convolution layer to reduce the dimension of the original fresh corn ear image feature map to obtain a feature map after dimensionality reduction, and the basic unit of ShuffleNetV2 is used to extract feature information from the feature map after dimensionality reduction, and a parameterless attention mechanism layer is used to enhance the feature extraction of spatial dimensions and channel dimensions, and the downsampling unit of ShuffleNetV2 is used to downsample the extracted feature information; the feature map is repeatedly passed through the parameterless attention mechanism layer, the downsampling unit of ShuffleNetV2, and the basic unit of ShuffleNetV2 to obtain a feature map after channel shuffling, and a series of sub-feature maps are mapped out by a lightweight convolution layer, and then the extracted sub-feature maps are spliced ​​in the channel dimension to obtain a feature map with increased channel dimension and reduced spatial dimension, and the SPPF module is used to obtain the final feature map; the neck network fuses the feature information extracted by the SS-neck feature to make the feature information fully fused; the detection head network makes a regression decision on the feature information fused by the neck network, and finally outputs the detected defect results.

[0011] According to the described method for detecting the appearance quality of fresh corn ears, the camera starts a multi-threaded tracking shooting mode, and the drag chain drives the tray to rotate clockwise at the same time. The camera contains m threads, and the value of m is not less than four. Thread one is used to obtain pictures, and the remaining threads are used to process pictures. The tray moves to the left boundary of the camera's field of view, and the camera thread one obtains the Pn1 picture at time Tn1. The camera takes pictures at a fixed time, and the time interval is T. After capturing the picture, thread one passes it to thread two, and thread two processes the picture obtained by thread one within the waiting time. At the same time, thread one times the time, and thread one continues to capture the next picture at time Tn2. Thread one obtains n pictures Pn1-Pnn after n-1 intervals T, and there are n trays in the Pnn pictures corresponding to ear one to ear n, respectively, which are passed to thread two to thread m, respectively; ear one moves to the right boundary of the camera's field of view, and the acquisition of the ear one picture ends; after the statistics are completed, thread two is released and used for the processing of the next round of image ear n+1.

[0012] According to the described method for detecting the appearance quality of fresh corn ears, the ears are placed in a tray and move with the tray, at which time the ears are in a relatively static state relative to the tray; the tray continues to move, the ears contact the inclined surface of the guide rail, and after passing the inclined surface of the guide rail, they float on the guide rail; the inner wall on the left side of the tray pushes the ears, and the ears rotate clockwise under the action of the guide rail.

[0013] According to the described method for detecting the appearance quality of fresh corn ears, thread one obtains the Pn2 picture at time Tn2, at which time the image contains ear one and ear two, ear one is a second angle picture, ear two is a first angle picture, and ear one is passed to thread two, while ear two is passed to thread three; thread one obtains the Pn3 picture at time Tn3, at which time the image contains ear one, ear two and ear three, ear one is a third angle picture, ear two is a second angle picture, and ear three is a first angle picture, and ear one is passed to thread two, ear two is passed to thread three, while ear three is passed to thread four.

[0014] According to the method for detecting the appearance quality of fresh corn ears, the ears are photographed n times in total, each time at an angle of Thread 2 processes n images separately and counts the defects of each image.

[0015] According to the described method for detecting the appearance quality of fresh corn ears, the maximum pooling convolution layer includes: a convolution structure, a batch normalization structure, a ReLU activation function and a maximum pooling structure; the maximum pooling convolution layer first extracts features through a 3×3 convolution structure, then performs normalization processing through a batch normalization structure, then activates through a ReLU activation function, and finally obtains a feature map after dimensionality reduction through a 3×3 maximum pooling operation.

[0016] According to the described method for detecting the appearance quality of fresh corn ears, the information of several feature channels of the feature map after dimensionality reduction is subjected to channel separation operation, and after separation, they enter the identity mapping path and the convolution re-extraction path respectively; the feature information of the convolution re-extraction path is subjected to convolution re-extraction, and the feature information of the identity mapping path is directly subjected to identity mapping, and the two feature information of the identity mapping and the convolution re-extraction are spliced ​​through the Concat function, and then the channels are shuffled to obtain the feature map after feature fusion.

[0017] According to the described method for detecting the appearance quality of fresh corn ears, the convolution re-extraction path first reduces the dimension of the channel feature information through a 1×1 convolution, then undergoes a 3×3 deep convolution, and finally undergoes a 1×1 convolution to increase the dimension, while the number of output feature channels remains unchanged.

[0018] According to the described method for detecting the appearance quality of fresh corn ears, the enhanced channel feature information flows into the first convolution branch and the second convolution branch respectively, the feature information is downsampled, then spliced ​​through the Concat function, and then the channels are shuffled. After downsampling, the number of output channels of the feature map is doubled and the size is halved; the first convolution branch includes: 3×3 depth convolution and 1×1 ordinary convolution in sequence; the second convolution branch includes: 1×1 ordinary convolution, 3×3 depth convolution and 1×1 ordinary convolution in sequence.

[0019] According to the described method for detecting the appearance quality of fresh corn ears, the characteristic value after channel shuffling is obtained. Figure X ,feature Figure X The size is S×S×C 1 , a series of sub-features are mapped out through lightweight convolutional layers Figure X ', sub-feature Figure X The size of ' is S'×S'×C1. When the scaling factor is 2, 4 sub-feature maps will appear. Connect the 4 mapped feature maps along the channel dimension to obtain a feature map with a 4-fold increase in channel dimension and a 2-fold decrease in spatial dimension. Figure X "; then input the feature map into the C 2 After a non-strided convolutional layer with 10 filters, the final output feature map size is S / 2×S / 2×C2.

[0020] Another aspect of the present application also provides a device for detecting the appearance quality of fresh corn ears, including: a complete machine support, a drag chain, a tray, a guide rail mechanism and a camera; a drag chain is arranged on the top of the complete machine support, and trays are fixed at equal intervals above the drag chain; a camera support is arranged at the middle position of the top of the complete machine support, and the camera is fixed on the cross bar of the camera support; the vertical direction of the camera is perpendicular to the drag chain, and the camera is located on the central plane of the guide rail mechanism.

[0021] According to the device for detecting the appearance quality of fresh corn ears, the drag chain includes: a first drag chain, a second drag chain, a front gear, a rear gear, a front connecting shaft and a rear connecting shaft; the front gear includes: a first drag chain front gear and a second drag chain front gear, and the first drag chain front gear and the second drag chain front gear are connected through the front connecting shaft; the rear gear includes: a first drag chain rear gear and a second drag chain rear gear, and the first drag chain rear gear and the second drag chain rear gear are connected through the rear connecting shaft; the protruding parts of the front connecting shaft and the rear connecting shaft are connected to the whole machine bracket; the protruding part of the front connecting shaft can be externally connected to a power device.

[0022] According to the device for detecting the appearance quality of fresh corn ears, the guide rail mechanism comprises: a guide rail, a guide rail bracket and a support; the guide rail is in the shape of a long bar, and inclined surfaces are arranged at both ends of the guide rail. The upper and lower sides of the guide rail are planes, and the upper plane has equally spaced grooves. The upper plane has no less than two mounting holes at equal intervals along the axial direction, and the lower plane is connected to the guide rail bracket by a connecting piece; the guide rail bracket is located on the support, and the support is connected to the whole machine bracket.

[0023] According to the device for inspecting the appearance quality of fresh corn ears, the tray is in the shape of a long semi-cylindrical strip, with an operating groove in the middle for cooperating with a guide rail to make the ears roll, arc-shaped on both sides for limiting the position of the ears, and a boss at the bottom for connecting to a drag chain; the groove surface of the operating groove is higher than the bottom of the tray.

[0024] According to the device for detecting the appearance quality of fresh corn ears, the guide rail mechanism is provided with no less than one guide rail, the guide rail is located below the tray, the upper planes of all the guide rails are at the same height, the side surfaces of the guide rails are parallel to the plane where the drag chain is located, and the side surfaces of the guide rails are perpendicular to the tray; the upper plane of the guide rails is higher than the bottom of the tray and lower than the operating groove surface in the middle of the tray.

[0025] Compared with the prior art, the beneficial effects of the present invention include at least the following points:

[0026] (1) The present application aims at high-speed movement of fruit ears during processing, and utilizes the principle of space division multiplexing and multi-threaded processing strategy. The present application designs the structure of the drag chain, pallet and guide rail mechanism. The drag chain, pallet and guide rail mechanism cooperate with each other to realize 360° rotation of the fruit ears during processing and transportation, and the pallet space is reasonably allocated; the camera can process pictures of the fruit ears at various angles in a short time through the multi-threaded processing strategy, thereby increasing the timeliness of visual inspection;

[0027] (2) This application performs a lightweight design on the backbone network, introduces ShuffleNetV2 downsampling unit, ShuffleNetV2 basic unit, maximum pooling convolution and lightweight convolution layer, which not only reduces the number of model parameters, but also has the fastest detection speed, ensuring that the extracted corn ear feature information is not lost, with high detection accuracy, and is suitable for real-time detection tasks;

[0028] (3) This application introduces the SimAM module, which directly calculates the three-dimensional attention weights in the feature map and unifies the feature weights, thereby enhancing the model's ability to focus on defective ear features;

[0029] (4) This application introduces Wise-IoU, which improves the convergence speed and reduces the impact of low-quality samples on the generalization ability of the model;

[0030] (5) This application designs a long strip tray for fresh corn cob-like materials, which is hollow in the middle and curved on both sides, and cooperates with the guide rail designed at the bottom to achieve moderate floating and uniform rotation of the material during the forward movement of the material in the visual inspection area;

[0031] (6) In the effective area of ​​visual inspection, the camera shooting frame rate matches the pallet movement speed to achieve material tracking, fixed-angle multi-image acquisition and quality inspection. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a schematic diagram of the detection method of the present application;

[0033] Figure 2 This is the YOLOv8 network architecture diagram in this application;

[0034] Figure 3 This is the maximum pooling layer structure diagram of this application;

[0035] Figure 4 This is the structure diagram of the lightweight convolutional layer of this application;

[0036] Figure 5 This is a comparison chart of the convergence of the Loss curve;

[0037] Figure 6 It is a three-dimensional structural diagram of the detection device of the present application;

[0038] Figure 7 It is a schematic diagram of the pallet of this application;

[0039] Figure 8 It is a schematic diagram of the guide rail of this application;

[0040] Fig. 9 It is a side view of the single tray and guide rail mechanism of the present application after assembly;

[0041] Fig.10It is a top view of the single tray and guide rail of the present application after assembly;

[0042] Fig.11 This is a diagram of the interaction between the fruit ears, trays and guide rails of the present application.

[0043] In the figure: 1. whole machine bracket; 2. drag chain; 21. first drag chain; 22. second drag chain; 23. front gear; 24. rear gear; 25. front connecting shaft; 26. rear connecting shaft; 3. tray; 31. operating slot; 32. countersunk positioning hole; 33. boss; 4. guide rail mechanism; 41. guide rail; 42. guide rail bracket; 43. support; 44. mounting hole; 5. camera; 51. camera bracket. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical scheme and advantages of the present invention clearer, the technical scheme of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. The embodiments described in this application are only embodiments of a part of the present invention, rather than all embodiments. Based on the spirit of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the protection scope of the present invention.

[0045] In the description of the present invention, it should be noted that the terms "front", "rear", "inside", "outside", "right", "left", "both ends", "one end", "the other end" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance, and the term "clockwise" is limited to the purpose of description and cannot be understood as indicating that the device can only rotate clockwise.

[0046] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "provided with", "connected", etc. should be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0047] Embodiment 1:

[0048] like Figure 1-Figure 2 As shown, the present invention provides a method for detecting the appearance quality of fresh corn ears, comprising the following steps:

[0049] S1. Acquire a set of fresh corn ear images. The image set is mainly acquired by camera 5. In other embodiments, the image set can also be acquired by a shooting device such as a mobile phone. Camera 5 starts a multi-threaded tracking shooting mode. Camera 5 includes m threads, and the value of m is not less than four. Thread 1 is responsible for acquiring images, and the remaining threads are responsible for processing images. At the same time, the drag chain 2 drives the tray 3 to rotate clockwise at a rotation speed of v.

[0050] S1.1. The fresh corn cob is placed in the tray 3 and moves with the tray 3. At this time, the cob is relatively stationary relative to the tray 3. The radius of the cob is r, and the circumference is S=2πr.

[0051] S1.2, the tray 3 continues to move, the fruit ear contacts the inclined surface of the guide rail 41, passes through the inclined surface of the guide rail 41, and floats on the guide rail 41. The length of the guide rail 41 is L.

[0052] S1.3. The left inner wall of the tray 3 pushes the fruit ear, and the fruit ear rotates clockwise under the action of the guide rail 41.

[0053] S1.4, tray 3 moves to the left edge of camera 5's field of view, and camera 5 thread 1 obtains Pn1 image at time Tn1, which only contains the first angle image of ear 1, and the image occupies the entire ear surface n=m-1.

[0054] S1.5, camera 5 takes pictures at a fixed time interval of T. Thread 1 captures the picture and passes it to thread 2, which processes the picture obtained by thread 1 during the waiting time. At the same time, thread 1 starts timing, and thread 1 continues to capture the next picture at time Tn2.

[0055] S1.6. Thread one obtains the Pn2 image at time Tn2. At this moment, the image contains ear one and ear two. Ear one is the second angle image, and ear two is the first angle image. Ear one is passed to thread two, and ear two is passed to thread three.

[0056] S1.7, thread 1 obtains the Pn3 image at time Tn3. At this moment, there are ear 1, ear 2 and ear 3 in the image. Ear 1 is the third angle image, ear 2 is the second angle image, and ear 3 is the first angle image. Ear 1 is passed to thread 2, ear 2 is passed to thread 3, and ear 3 is passed to thread 4. Similarly, thread 1 obtains n images Pn1-Pnn after n-1 intervals T. There are n trays 3 in the Pnn image, corresponding to ear 1-ear n, which are passed to thread 2-thread m respectively.

[0057] S1.8, the ear of fruit moves to the right edge of the field of view of camera 5, and the acquisition of the image of the ear of fruit is completed. The ear of fruit is photographed n times in total, and the angle of each time is Thread 2 processes n images respectively and counts the defects of each image. After the statistics are completed, thread 2 is released and used to process the next round of image ear n+1.

[0058] like Figure 1 As shown, in this embodiment, the value of m is seven, and the value of n is six, that is, the camera 5 includes seven threads, thread one is responsible for acquiring pictures, and threads two to seven are responsible for processing pictures; camera thread one acquires Pn1 picture at time Tn1, which only includes the first angle picture of ear one, and the picture accounts for 1 / 6 of the entire ear surface. Thread one acquires 6 pictures Pn1-Pn6 after 6 intervals T, and there are 6 trays in the Pn6 picture corresponding to ear one to ear six, which are respectively transmitted to thread two to thread seven. Ear one is photographed 6 times in total, each time at an angle of 60°. Thread two processes 6 pictures respectively and counts the defects of each image. After the statistics are completed, thread two is released and is used to process the next round of image ear seven.

[0059] In another embodiment, the camera 5 comprises thirteen threads, thread one is responsible for acquiring pictures, and threads two to thirteen are responsible for processing pictures. The ear of fruit is photographed twelve times in total, each time at an angle of 30°.

[0060] In another embodiment, the camera 5 includes four threads, thread one is responsible for acquiring pictures, and threads two to four are responsible for processing pictures. The ear of fruit is photographed twelve times in total, each time at an angle of 120°.

[0061] Data constraints:

[0062] (1) L>S, the optimal ratio is L=1.5S=3πr;

[0063] (2)vT=S / N=2πr / N.

[0064] Embodiment 2:

[0065] Based on the feature that fresh corn ears are composed of a large number of independent and closely arranged kernels, the present application improves the backbone feature extraction network, adopts a strategy combining feature reuse of ShuffleNetV2, lightweight convolutional layers to ensure fine-grained information, and maximum pooling convolutional layers to reduce the amount of calculation, to ensure that the model does not lose local and fine-grained information of the ear, while achieving model lightweighting; in view of the fact that the phenotypic characteristics of missing and dropped kernels in inferior ears of fresh corn account for a small proportion of the entire ear, resulting in false detection due to failure of capture by the backbone feature extraction network, a parameter-free attention module is introduced into the backbone feature extraction network module to improve the model's feature extraction capability for dropped and missing kernels; Wise-IoU is introduced as the bounding box regression loss function to make up for the deficiency in the Complete-IoU loss function that the length and width of the prediction box cannot be changed at the same time, resulting in large size differences and deformed ears affecting the training convergence speed and the decline in model performance, thereby further ensuring the detection performance of the present application.

[0066] The yellow fresh corn cob products of the production line are collected by the method of Example 1 to obtain high-quality and low-quality cob images, and a high-quality and low-quality fresh corn cob dataset is established as the data for this research, wherein the high-quality cobs include first-class and second-class cobs, and the low-quality cobs include deformed, fallen, missing, and mechanically damaged cobs.

[0067] Figure 2 As shown in the figure, the obtained image is processed: SS-neck is introduced in YOLOv8, SS-neck uses the maximum pooling convolution layer to reduce the dimension of the original fresh corn ear image feature map to obtain the reduced feature map, the basic unit of ShuffleNetV2 is used to extract feature information from the reduced feature map, the non-parameter attention mechanism layer is used to enhance the feature extraction of spatial dimension and channel dimension, and the downsampling unit of ShuffleNetV2 is used to downsample the extracted feature information; the feature map is repeatedly passed through the non-parameter attention mechanism layer, the downsampling unit of ShuffleNetV2 and the basic unit of ShuffleNetV2 to obtain the feature map after channel shuffling, the lightweight convolution layer is used to map out a series of sub-feature maps, and then the extracted sub-feature maps are spliced ​​in the channel dimension to obtain a feature map with increased channel dimension and reduced spatial dimension, and the SPPF module is used to obtain the final feature map.

[0068] S2.1, such as Figure 3As shown, the input layer receives the image acquired by the camera 5, changes the size of the image to a predetermined size, such as 640×640, and inputs the resized image into the maximum pooling convolution layer, which includes a convolution structure, a batch normalization structure, a ReLU activation function, and a maximum pooling structure; the maximum pooling convolution layer first extracts features through a 3×3 convolution structure, then performs normalization processing through a batch normalization structure, and then activates through a ReLU activation function, and finally obtains a feature map after dimensionality reduction through a 3×3 maximum pooling operation, the number of channels of which is C;

[0069] Among them, the convolution structure extracts the input corn ear features for the first time, and the output feature map enters the maximum pooling structure for downsampling operation to further extract the feature information. Compared with the convolution layer, the maximum pooling structure can not only complete the downsampling operation, but also reduce the amount of calculation.

[0070] S2.2, input the feature map after dimensionality reduction into the ShuffleNetV2 basic unit to obtain the feature map after feature fusion;

[0071] S2.2.1. Perform channel separation operation on the information of C feature channels, and enter the identity mapping path and convolution re-extraction path respectively after separation; let the number of feature channels in the identity mapping path be C 3 , the number of feature channels in the convolution re-extraction path is C 4 indivual;

[0072] Among them, C 3 +C 4 =C;

[0073] S2.2.2.1, C 2 The channel feature information is first reduced in dimension by a 1×1 convolution, then passed through a 3×3 depth convolution, and finally passed through a 1×1 convolution for dimension increase, while the number of output feature channels remains unchanged;

[0074] S2.2.2.2, C 1 The channel feature information is directly mapped identically;

[0075] S2.2.3, the two feature information extracted by the identity mapping and the convolution are concatenated through the Concat function to ensure that the extracted feature information is not lost, and then the channels are shuffled to ensure that the obtained feature information is fully fused to obtain the feature map after feature fusion;

[0076] S2.3. The feature map after feature fusion is input into the parameter-free attention mechanism layer to enhance the feature extraction of spatial dimension and channel dimension, thereby enhancing the feature information representation capability, and directly calculating the three-dimensional attention weight. The parameter-free attention mechanism layer directly calculates the feature weight of the feature map, assigns a unique weight to each neuron, and each neuron has an independent energy function. Minimizing the energy function eventually obtains the minimum energy function of a single neuron:

[0077]

[0078] In the formula and is the mean and variance of all neurons in the channel except t. From the above formula, we can see that the lower the energy, the more obvious the distinction between neuron t and surrounding neurons, the more significant the attention to features, and the enhanced ability to represent feature information.

[0079] It can improve the model's attention to characteristic information such as missing kernels, mechanical damage, and kernel loss in low-quality fruit clusters;

[0080] S2.4, the enhanced feature information is input into the ShuffleNetV2 downsampling unit, and the channel feature information flows into the first convolution branch and the second convolution branch respectively, and the feature information is downsampled, and then concatenated through the Concat function, and then the channels are shuffled. After downsampling, the number of output channels of the feature map is doubled and the size is halved to ensure that the extracted feature information is not lost;

[0081] The first convolution branch includes: 3×3 depth convolution and 1×1 normal convolution in sequence; the second convolution branch includes: 1×1 normal convolution, 3×3 depth convolution and 1×1 normal convolution in sequence;

[0082] S2.5, such as Figure 4 As shown, S2.3, S2.4, and S2.2 are repeated in sequence to obtain the features after channel shuffling. Figure X , assuming its size is S×S×C 1 , a series of sub-features are mapped out through the lightweight convolution layer Figure X ', the size is S'×S'×C1. When the scaling factor is 2, 4 sub-feature maps will appear. Connect the 4 mapped feature maps along the channel dimension to obtain features with the channel dimension increased by 4 times and the spatial dimension reduced by 2 times. Figure X ”, X” size is S / 2×S / 2×4C 1 This feature map is then input into the 2 After a non-strided convolutional layer with 10 filters, the final output feature map size is S / 2×S / 2×C2.

[0083] The lightweight convolution layer includes: a space-to-depth layer and a non-stride convolution layer; the space-to-depth layer downsamples the feature information of fresh corn ears extracted by ShuffleNetV2, while retaining all feature information in the channel dimension, thereby avoiding feature loss; the non-stride convolution layer can realize feature extraction without reducing the size of the feature map, which can further reduce the loss of fine-grained information in the corn ear image, thereby ensuring the accuracy of the network model.

[0084] S2.6. Input the feature information of the feature map obtained above into the SPPF module, perform 1×1 convolution operations in sequence to halve the channels, then perform pooling operations with kernel sizes of 5, 9, and 13 in sequence, then splice the pooled feature information, and finally perform 1×1 convolution operations to adjust the channels to obtain the spliced ​​feature map.

[0085] S2.7, the neck network fuses the feature information extracted from the trunk features, so that the feature information is fully integrated;

[0086] S2.8. The detection head network performs regression decision on the feature information fused by the neck network and finally outputs the detected defect results.

[0087] The SS-neck structure of the present application uses the ShuffleNetV2 structure to perform feature extraction to avoid the loss of feature information when the channel of the backbone network is halved; the convolution structure using the maximum pooling convolution layer and the lightweight convolution layer for downsampling not only reduces the number of model parameters, but also ensures that the extracted corn ear feature information is not lost, thereby ensuring high-precision detection of the model.

[0088] like Figure 2 As shown, the dotted box part represents the backbone network of the present application, and the original model backbone network is replaced by SS-neck, Input represents the input layer, Conv_Maxpool represents the maximum pooling convolution layer, ShuffleNetV2_B represents the ShuffleNetV2 basic unit, SimAM represents the parameterless attention mechanism layer, ShuffleNetV2_U represents the ShuffleNetV2 downsampling unit; SPDConv represents the lightweight convolution layer, SPPF represents the SPPF module in the YOLOv8 network, concat represents the concat function, Upsample represents the Upsample function, Detect represents the Detect layer in the YOLOv8 network, conv represents the convolution layer, and C2f represents the C2f module.

[0089] The activation function of the maximum pooling convolution layer, ShuffleNetV2 basic unit, and ShuffleNetV2 downsampling unit uses the ReLU function.

[0090] In the low-quality samples of fresh corn ears, the size of deformed corn ears is much different from that of other ears. If the length and width of the prediction box cannot be changed at the same time, the training results will be affected. This application cites the loss function Wise-IoU with a dynamic gradient gain allocation strategy and length-width separation calculation. Wise-IoU can dynamically adjust the loss function according to the quality of the sample through a dynamic non-monotonic focused gradient gain allocation strategy, reduce the competitive advantage of high-quality samples, and reduce the adverse gradient effects caused by low-quality samples. At the same time, it improves the detection accuracy and robustness of the model. According to the distance metric, Wise-IoU with two layers of attention is obtained, as shown in formula (2):

[0091]

[0092]

[0093] In the formula, R WIoU W IoU The distance metric function is defined as shown in formula (4); L of the normal quality anchor box is significantly enlarged IoU ; (Wg, Hg) represent the length and width of the optimal bounding box respectively; * is the separation operation, in order to prevent R WIoU Generate gradients that hinder convergence, separate Wg and Hg from the calculation, avoid the situation where the length and width of the prediction box cannot change at the same time, and reduce interference with model training.

[0094] like Figure 5 As shown in the figure, the red and black curves represent the use of WIoU and CIoU respectively, and the blue, green and purple curves represent the use of DIoU, GIoU and SIoU respectively. It can be seen that after the introduction of Wise-IoU, the model converges in the area of ​​about 50 epochs of training, and the convergence effect is good, while the other four loss functions tend to converge after nearly 180 epochs, and the loss value during training is much lower than that of the other four loss functions.

[0095] Table 1. Detection performance comparison of YOLO detection model

[0096]

[0097] The results in Table 1 show that the precision, recall, and mAP results of this application are higher than those of the YOLOv5 and YOLOv8 models. Compared with YOLOv6, the precision and recall are slightly lower, and the reduction range is within 1 percentage point. Compared with several other detection models, the improved model size and the amount of calculation parameters are smaller, and the model parameters are reduced by 0.52, 1.02, and 2.25M respectively; the model size is reduced by 1.3, 2.3, and 4.7MB respectively. Taking into account the detection accuracy of the model, the calculation amount of the model, and the model size, the improved model is more suitable for the task of corn ear quality detection.

[0098] Embodiment three:

[0099] like Figure 6-11 As shown, the present invention also provides a fresh corn ear appearance quality detection device, comprising: a whole machine support 1, a drag chain 2, a tray 3, a guide rail mechanism 4 and a camera 5. The drag chain 2 is arranged on the top of the whole machine support 1, and the tray 3 is fixed above the drag chain 2. A camera support 51 is arranged at the middle position of the top of the whole machine support 1, and the camera 5 is fixed on the crossbar of the camera support 51; the vertical direction of the camera 5 is perpendicular to the drag chain 2, and the camera 5 is on the central plane of the guide rail mechanism 4.

[0100] The whole machine bracket provides support for the device and is made of stainless steel.

[0101] like Figure 6 As shown, the drag chain 2 includes: a first drag chain 21, a second drag chain 22, a front gear 23, a rear gear 24, a front connecting shaft 25 and a rear connecting shaft 26; the front gear 23 includes: a first drag chain front gear and a second drag chain front gear, and the first drag chain front gear and the second drag chain front gear are connected through the front connecting shaft 25; the rear gear 24 includes: a first drag chain rear gear and a second drag chain rear gear, and the first drag chain rear gear and the second drag chain rear gear are connected through the rear connecting shaft 26. The protruding parts of the front connecting shaft 25 and the rear connecting shaft 26 are connected to the whole machine bracket 1; the top of the drag chain 2 is evenly spaced with trays 3 installed, the drag chain 2 is located on the whole machine bracket 1 to hold up the tray 3, and the protruding part of the front connecting shaft 25 protrudes outward by 10cm to be externally connected to the power device.

[0102] like Figure 7 As shown, the tray 3 is a long semi-cylindrical shape, with an operating groove 31 in the middle for cooperating with the guide rail 41 to make the fruit ears roll, arc-shaped on both sides for limiting the position of the fruit ears, countersunk positioning holes 32 at the front and rear ends of the tray 3, and a boss 33 at the bottom for connecting with the drag chain 2. The groove surface of the operating groove 31 is higher than the bottom of the tray 3.

[0103] like Figure 8 or Fig. 9As shown, the guide rail mechanism 4 includes: a guide rail 41, a guide rail bracket 42 and a support 43. The guide rail 41 can be made of silicone material. The guide rail 41 is in the shape of a long bar. Both ends of the guide rail 41 are provided with inclined surfaces. The upper and lower sides of the guide rail 41 are planes. The upper plane has equally spaced grooves. The upper plane has no less than two mounting holes 44 at equal intervals along the axial direction. The lower plane is connected to the guide rail bracket 42 by connecting members such as bolts or pins. In this embodiment, three mounting holes 44 are opened to fix the guide rail 41 on the guide rail bracket 42. In another embodiment, three mounting holes 44 can be opened to fix the guide rail 41 on the guide rail bracket 42. The equally spaced grooves are used to increase the friction force and ensure the uniform rotation of the fruit ears. The guide rail bracket 42 is located on the support 43, and the support 43 is connected to the whole machine bracket 1.

[0104] The tray 3 cooperates with the guide rail mechanism 4 to realize the rotation of the fruit cluster.

[0105] The guide rail mechanism 4 is provided with at least one guide rail 41. In the present embodiment, two guide rails 41 are provided, one at each end and placed in parallel. The guide rails 41 are located below the tray 3. The upper plane of the guide rails 41 is at the same height. The two side surfaces of the guide rails 41 are parallel to the plane where the drag chain 2 is located. The two side surfaces of the guide rails 41 are perpendicular to the tray 3. The spacing between the two guide rails 41 can be adjusted according to the length of the fruit cluster, generally at a spacing of 20 cm. In another embodiment, one guide rail 41 can be provided, which is placed at the center below the tray 3 and perpendicular to the tray 3.

[0106] like Fig.11 As shown, the upper plane of the guide rail 41 is higher than the bottom of the tray 3 and lower than the surface of the operating groove 31 in the middle of the tray 3. The reason why it is higher than the bottom of the tray 3 is to lift the fruit ears, and the reason why it is lower than the surface of the operating groove 31 in the middle of the tray 3 is to scrape the surface against the tray 3, so that the fruit ears can rotate as the tray 3 moves on the guide rail 41.

[0107] The present invention discloses a method and device for detecting the appearance quality of fresh corn ears. The appearance quality detection method comprises: using a towline 2 to drive a tray 3 within a limited length, using the tray 3, a guide rail 41 and the ear to interact synchronously to achieve forward movement and uniform flipping, and a camera 5 to obtain a certain number of fresh corn images within a suitable field of view; using the space division multiplexing principle and a multi-threaded processing strategy, images of each ear at different angles are obtained and processed within a shooting cycle to obtain quality data, thereby realizing rapid detection of the appearance quality of fresh corn.

[0108] Compared with the prior art, the beneficial effects of the present invention include at least the following points:

[0109] (1) The present application aims at high-speed movement of fruit ears during processing, and utilizes the principle of space division multiplexing and multi-threaded processing strategy. The present application designs the structure of the drag chain, pallet and guide rail mechanism. The drag chain, pallet and guide rail mechanism cooperate with each other to realize 360° rotation of the fruit ears during processing and transportation, and the pallet space is reasonably allocated; the camera can process pictures of the fruit ears at various angles in a short time through the multi-threaded processing strategy, thereby increasing the timeliness of visual inspection;

[0110] (2) This application performs a lightweight design on the backbone network, introduces ShuffleNetV2 downsampling unit, ShuffleNetV2 basic unit, maximum pooling convolution and lightweight convolution layer, which not only reduces the number of model parameters, but also has the fastest detection speed, ensuring that the extracted corn ear feature information is not lost, with high detection accuracy, and is suitable for real-time detection tasks;

[0111] (3) This application introduces the SimAM module, which directly calculates the three-dimensional attention weights in the feature map and unifies the feature weights, thereby enhancing the model's ability to focus on defective ear features;

[0112] (4) This application introduces Wise-IoU, which improves the convergence speed and reduces the impact of low-quality samples on the generalization ability of the model;

[0113] (5) This application designs a long strip tray for fresh corn cob-like materials, which is hollow in the middle and curved on both sides, and cooperates with the guide rail designed at the bottom to achieve moderate floating and uniform rotation of the material during the forward movement of the material in the visual inspection area;

[0114] (6) In the effective area of ​​visual inspection, the camera shooting frame rate matches the pallet movement speed to achieve material tracking, fixed-angle multi-image acquisition and quality inspection.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for detecting the appearance quality of fresh corn ears, characterized in that: The following steps are involved: Acquiring a set of fresh corn ear images through a camera (5); The obtained image is processed: SS-neck is introduced in YOLOv8. SS-neck uses the maximum pooling convolution layer to reduce the dimension of the original fresh corn ear image feature map to obtain the reduced feature map. The basic unit of ShuffleNetV2 is used to extract feature information from the reduced feature map. The parameter-free attention mechanism layer is used to enhance the feature extraction of spatial dimensions and channel dimensions. The downsampling unit of ShuffleNetV2 is used to downsample the extracted feature information. The feature map is repeatedly passed through the parameter-free attention mechanism layer, the downsampling unit of ShuffleNetV2 and the basic unit of ShuffleNetV2 to obtain the channel shuffled feature map. A lightweight convolution layer is used to map out a series of sub-feature maps. The extracted sub-feature maps are then spliced ​​in the channel dimension to obtain a feature map with increased channel dimension and reduced spatial dimension. The SPPF module is used to obtain the final feature map. The neck network fuses the feature information extracted from the SS-neck feature, so that the feature information is fully integrated; The detection head network performs regression decision on the feature information fused by the neck network and finally outputs the detected defect results.

2. A method for detecting the appearance quality of fresh corn ears according to claim 1, characterized in that: The camera (5) starts a multi-threaded tracking shooting mode, and the drag chain (2) drives the tray (3) to rotate clockwise. The camera (5) includes m threads, and the value of m is not less than four. Thread 1 is used to obtain pictures, and the remaining threads are used to process pictures. The tray (3) moves to the left edge of the camera (5) field of view, and thread 1 of the camera (5) obtains Pn1 picture at time Tn1. The camera (5) takes pictures at a fixed time interval of T; thread 1 captures the picture and passes it to thread 2, which processes the picture obtained by thread 1 during the waiting time; at the same time, thread 1 starts timing, and thread 1 continues to capture the next picture at time Tn2; Thread 1 obtains n pictures Pn1-Pnn after n-1 intervals T. There are n trays (3) in the Pnn picture, corresponding to ear 1-ear n, and they are passed to thread 2-thread m respectively; The ear of fruit moves to the right edge of the camera (5) field of view, and the acquisition of the image of the ear of fruit is completed; after the statistics are completed, thread 2 is released and used for the next round of processing of the image of ear of fruit n+1.

3. A method for detecting the appearance quality of fresh corn ears according to claim 2, characterized in that: The fruit ear is placed in the tray (3) and moves with the tray (3), and at this time the fruit ear is in a relatively static state relative to the tray (3); the tray (3) continues to move, the fruit ear contacts the inclined surface of the guide rail (41), and after passing the inclined surface of the guide rail (41), floats on the guide rail (41); the left inner wall of the tray (3) pushes the fruit ear, and the fruit ear rotates clockwise under the action of the guide rail (41).

4. A method for detecting the appearance quality of fresh corn ears according to claim 2, characterized in that: Thread 1 obtains the Pn2 image at time Tn2. At this moment, the image contains ear 1 and ear 2. Ear 1 is the second angle image, and ear 2 is the first angle image. Ear 1 is passed to thread 2, and ear 2 is passed to thread 3. Thread one obtains the Pn3 picture at time Tn3. At this moment, there are ear one, ear two and ear three in the image. Ear one is the third angle picture, ear two is the second angle picture, and ear three is the first angle picture. Ear one is passed to thread two, ear two is passed to thread three, and ear three is passed to thread four.

5. The method for detecting the appearance quality of fresh corn ears according to claim 2, characterized in that: The ear was photographed n times, each time at an angle of Thread 2 processes n images separately and counts the defects of each image.

6. A method for detecting the appearance quality of fresh corn ears according to claim 1, characterized in that: The maximum pooling convolution layer includes: convolution structure, batch normalization structure, ReLU activation function and maximum pooling structure; the maximum pooling convolution layer first extracts features through a 3×3 convolution structure, then normalizes through a batch normalization structure, then activates through a ReLU activation function, and finally obtains a feature map after dimensionality reduction through a 3×3 maximum pooling operation.

7. A method for detecting the appearance quality of fresh corn ears according to claim 6, characterized in that: The information of several feature channels of the feature map after dimensionality reduction is subjected to channel separation operation, and after separation, they enter the identity mapping path and the convolution re-extraction path respectively; the feature information of the convolution re-extraction path is re-extracted by convolution, and the feature information of the identity mapping path is directly subjected to identity mapping. The two feature information of the identity mapping and the convolution re-extraction are spliced ​​through the Concat function, and then the channels are shuffled to obtain the feature map after feature fusion.

8. A method for detecting the appearance quality of fresh corn ears according to claim 7, characterized in that: The convolution re-extraction path first reduces the dimension of the channel feature information through a 1×1 convolution, then passes it through a 3×3 depth convolution, and finally increases the dimension through a 1×1 convolution, while the number of output feature channels remains unchanged.

9. A method for detecting the appearance quality of fresh corn ears according to claim 7, characterized in that: The enhanced channel feature information flows into the first convolution branch and the second convolution branch respectively, and the feature information is downsampled, then concatenated through the Concat function, and then the channels are shuffled. After downsampling, the number of output channels of the feature map is doubled and the size is halved; The first convolution branch includes: 3×3 depth convolution and 1×1 normal convolution in sequence; the second convolution branch includes: 1×1 normal convolution, 3×3 depth convolution and 1×1 normal convolution in sequence.

10. A method for detecting the appearance quality of fresh corn ears according to claim 9, characterized in that: The feature map X after channel shuffling is obtained, and the size of the feature map X is S×S×C1. A series of sub-feature maps X' are mapped out through the lightweight convolution layer. The size of the sub-feature map X' is S'×S'×C1. When the scaling factor is 2, 4 sub-feature maps will appear. The 4 mapped feature maps are connected along the channel dimension to obtain a feature map X" with the channel dimension increased by 4 times and the spatial dimension reduced by 2 times; then the feature map is input into the non-strided convolution layer with C2 filters. Finally, the output feature map size is S / 2×S / 2×C2.

11. A device for detecting the appearance quality of fresh corn ears, comprising: The whole machine support (1), the drag chain (2), the tray (3), the guide rail mechanism (4) and the camera (5); the characteristics are: A drag chain (2) is arranged on the top of the whole machine support (1), trays (3) are fixed at equal intervals above the drag chain (2), a camera support (51) is arranged at the middle position of the top of the whole machine support (1), and the camera (5) is fixed on the crossbar of the camera (5) support (51); the vertical direction of the camera (5) is perpendicular to the drag chain (2), and the camera (5) is located on the central plane of the guide rail mechanism (4).

12. A device for detecting the appearance quality of fresh corn ears according to claim 11, characterized in that: The drag chain (2) comprises: a first drag chain (21), a second drag chain (22), a front gear (23), a rear gear (24), a front connecting shaft (25) and a rear connecting shaft (26); the front gear (23) comprises: a first drag chain front gear and a second drag chain front gear, and the first drag chain front gear and the second drag chain front gear are connected via the front connecting shaft (25); the rear gear (24) comprises: a first drag chain rear gear and a second drag chain rear gear, and the first drag chain rear gear and the second drag chain rear gear are connected via the rear connecting shaft (26); The protruding parts of the front connecting shaft (25) and the rear connecting shaft (26) are connected to the whole machine bracket (1); and the protruding part of the front connecting shaft (25) can be externally connected to a power device.

13. The device for detecting the appearance quality of fresh corn ears according to claim 11, characterized in that: The guide rail mechanism (4) comprises: a guide rail (41), a guide rail bracket (42) and a support (43); the guide rail (41) is in the shape of a long bar, and inclined surfaces are arranged at both ends of the guide rail (41); the upper and lower sides of the guide rail (41) are planes, the upper plane has grooves at equal intervals, and the upper plane has no less than two mounting holes (44) at equal intervals along the axial direction, and the lower plane is connected to the guide rail bracket (42) through a connecting piece; the guide rail bracket (42) is located on the support (43), and the support (43) is connected to the whole machine bracket (1).

14. A device for detecting the appearance quality of fresh corn ears according to claim 13, characterized in that: The tray (3) is in the shape of a long semi-cylindrical strip, with an operating groove (31) in the middle for cooperating with a guide rail (41) to cause the fruit ears to roll, arc-shaped sides for limiting the position of the fruit ears, and a boss (33) at the bottom for connecting with the drag chain (2); the groove surface of the operating groove (31) is higher than the bottom of the tray (3).

15. The device for detecting the appearance quality of fresh corn ears according to claim 13, characterized in that: The guide rail mechanism (4) is provided with at least one guide rail (41), the guide rail (41) is located below the tray (3), the upper planes of all the guide rails (41) are at the same height, the two side surfaces of the guide rails (41) are parallel to the plane where the drag chain (2) is located, and the two side surfaces of the guide rails (41) are perpendicular to the tray (3); The upper plane of the guide rail (41) is higher than the bottom of the tray (3) and lower than the surface of the operating groove (31) in the middle of the tray (3).

Citation Information

Patent Citations

  • Plant disease area detection method

    CN111968087A

  • Lightweight target detection method, apparatus and device, and storage medium

    CN115187820A

  • High-quality low-carbon drying method for corn ears

    CN115902131A

  • Lightweight pomegranate identification method based on improved YOLOv8s

    CN116958961A

  • Target detection method based on improved YOLOv8

    CN118485822A