A yolo v7-based ginseng bruise area detection and grade evaluation method
By using a YOLOv7-based method for detecting the odorous region of Panax notoginseng, and leveraging the Laplacian variance algorithm and collaborative attention mechanism, a high-quality training dataset is generated. By optimizing feature fusion and bounding box regression loss, the problem of missed detection and false detection in small regions of Panax notoginseng odor is solved, achieving efficient and accurate screening and grade evaluation.
Patent Information
- Application Number
- CN202310654052.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-05
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-06-05
AI Technical Summary
Existing technologies are insufficient for effectively detecting and screening small areas of blemishes in Panax notoginseng, resulting in high rates of missed and false detections. Furthermore, manual screening is inefficient and cannot guarantee the quality and safety of Panax notoginseng products.
A YOLOv7-based method for detecting the odorous region of *Panax notoginseng* is adopted. The image dataset is processed by the Laplacian variance algorithm, the feature fusion module is optimized, a collaborative attention mechanism is introduced, a high-quality training dataset is generated, prior boxes with large size spans are generated, and angular cost is introduced into the regression loss function of the predicted box to improve detection accuracy and efficiency.
It enables efficient detection of small areas of odor in Panax notoginseng, reduces the rate of missed and false detections, provides a basis for the grading of Panax notoginseng, and improves screening efficiency and accuracy.
Smart Images

Figure CN116797925B_ABST
Abstract
Description
TECHNICAL FIELD
[0002] The application discloses a YOLOv7-based panax notoginseng bruise area detection and grade evaluation method, relates to detection of bruise areas in panax notoginseng images and evaluation of panax notoginseng grades according to detection results, and belongs to the field of target detection. BACKGROUND
[0004] Panax notoginseng is a perennial herbaceous plant of the Araliaceae family, and generally, its rhizome is used as medicine. The root system of panax notoginseng is cylindrical or conical, with a diameter of 1-4 cm and a length of about 1-6 cm, and the surface is grayish yellow or grayish brown. The pharmacological effects of panax notoginseng are manifested in the blood system, central nervous system, cerebrovascular system, and immune regulation system, and it is one of the main components of Xuesaitong, Compound Danshen Dropping Pills, and Yunnan Baiyao. The huge medicinal value of panax notoginseng puts forward strict requirements for quality control, and it is very important to sort and screen panax notoginseng rhizomes to make consumers feel at ease. The medicinal part of panax notoginseng is the rhizome, which is easily crushed and broken during picking or transportation, leading to bruising or odor of panax notoginseng rhizomes. If these bruised or odorous panax notoginseng are not screened out, the therapeutic effect of panax notoginseng cannot be guaranteed, and it is harmful to human health.
[0005] Currently, the screening of bruised panax notoginseng is mainly carried out by manual methods, which requires a large amount of human resources and time cost, has low screening efficiency, and the bruised area of panax notoginseng has large differences. Large bruised areas can be quickly detected, while small bruised areas and normal panax notoginseng areas are similar, and it is difficult to distinguish them using manual methods, resulting in serious missed selection and misselection and low detection accuracy.
[0006] In order to assist panax notoginseng sales and pharmaceutical enterprises in quickly and accurately screening bruised panax notoginseng and ensuring product quality, some known methods use deep learning-based target detection technology for panax notoginseng bruise area detection and similar scenarios. For example, He Jun et al. (Patent (Panax notoginseng quality evaluation method, server and evaluation system based on image recognition) 202211398376.3, 2022) improved the feature fusion module of YOLOv5 (You Only Look Once v5) and introduced an attention mechanism to fully fuse panax notoginseng image features, which can effectively improve the accuracy of panax notoginseng bruise area detection, but does not consider the missed detection and misdiagnosis problems caused by insufficient positioning of small bruised areas of panax notoginseng.
[0007] In a similar scenario as the detection of the pungent area of Panax notoginseng, some known methods enrich the small target feature information by incorporating low-level features, but this leads to an increase in computational complexity and a decrease in detection speed, and does not consider the problem of insufficient small target feature extraction caused by large target size variation. For example, Li Pan (< Degree Thesis (Research on Apple Surface Defect Detection Based on Deep Learning) >, 2022) aimed at the problem of low detection accuracy caused by small apple surface defect area and unobvious feature information, replaced the traditional convolution with the dilated convolution to expand the receptive field of the convolution kernel, which can well capture the context information in the image, but did not consider the problem of small defect omission caused by color and light in the image. Zhang Rongshuang (< Degree Thesis (Research and Implementation of Crop Disease Detection Based on Deep Learning) >, 2022) aimed at the problem of small disease area in crop leaves that is difficult to accurately detect, used a feature fusion method based on attention mechanism, focused on the position information and disease information of the small disease area, thereby effectively improving the focusing ability on crop disease information, but the extraction and fusion of multi-scale features were insufficient. Fan Cheng et al. (< Degree Thesis (Mango Skin Defect Detection Based on Convolutional Neural Network) >, 2021) aimed at the problem of omission in mango skin defect detection caused by occlusion and small defect area, increased the receptive field of each layer of the network, and incorporated more shallow features in the decoder module to improve the detection accuracy of small defects, thereby reducing the omission rate of small defects in mangoes, but the incorporation of shallow features would increase the computational complexity, thereby reducing the detection speed.
[0008] To overcome the shortcomings of the above known methods, the present application discloses a Panax notoginseng pungent area detection and grade evaluation method based on YOLOv7, which uses the Laplacian variance algorithm to process the Panax notoginseng image dataset with large quality difference into a high-quality training dataset; optimizes the selection process of the Panax notoginseng image real box clustering center, generates a Panax notoginseng image prior box with large size span, solves the problem of insufficient small pungent area feature extraction in Panax notoginseng images; improve the feature fusion module of YOLOv7, fully fuse the feature information missed by the parallel pooling operation, thereby obtaining more rich Panax notoginseng deep features; introduce a collaborative attention mechanism to suppress the interference of background noise on small Panax notoginseng pungent area feature extraction, and introduce an angle cost in the Panax notoginseng image prediction box regression loss function, improve the positioning accuracy by fixing the regression direction, and solve the problem of small pungent area omission and misdiagnosis. SUMMARY
[0010] I. Objectives of the Invention
[0011] In order to solve the problems of insufficient feature extraction of small bruise area of ginseng and missed detection and false detection caused by complex background noise, the present application proposes a ginseng bruise area detection and grade evaluation method based on YOLOv7, realizes the detection of bruise area in ginseng image, solves the problem of missed detection and false detection of small bruise area of ginseng, and proposes a ginseng grade evaluation method according to the detection result, which provides a basis for efficient screening of ginseng.
[0012] Secondly, the steps of the present application
[0013] The execution process of the present application is divided into four steps.
[0014] (1) Ginseng image data preprocessing: ginseng image data set is constructed by collecting ginseng stem image, and Laplacian variance algorithm is used to filter the fuzzy and poor quality images, then the filtered ginseng image data set is labeled and divided, and finally the images in the data set are scaled, normalized and data enhanced.
[0015] (2) Ginseng bruise area detection model construction: a ginseng bruise area detection model based on YOLOv7 is constructed, first, a ginseng feature extraction network is used to extract and process multi-scale features of the input ginseng image, then an improved feature fusion module is used to obtain more rich ginseng deep features, and a collaborative attention mechanism is introduced in the feature re-extraction module to suppress the interference of background noise, and finally the ginseng detection network is processed to output the prediction result.
[0016] (3) Ginseng bruise area detection model training: the ginseng image training set and validation set constructed in step (1) are input into the model in step (2), the loss function is calculated and back propagation is carried out, and the weight of the model is iteratively updated.
[0017] (4) Ginseng bruise area detection and grade evaluation: the ginseng image test set is input into the model trained in step (3), the ginseng bruise area is detected, and the ginseng is graded according to the detection result.
[0018] The specific steps are as follows:
[0019] 1: Ginseng image data preprocessing
[0020] 1.1: Ginseng image data screening
[0021] The collected ginseng image data set is denoted as D, D=(d1,d2,…,d k ), d i (1≤i≤k) is any image in D. Using Laplacian variance algorithm, d i that meets the requirements is selected from D. The calculation process is as follows: select a 3x3 Laplacian operator matrix A to di After performing convolution, the variance is calculated. Set the threshold to Y (Y>0), if Then determine d i If a high-quality Panax notoginseng image is obtained, it should be stored in D; otherwise, the result should be determined as d. i The image is of low quality and is removed from D.
[0022] 1.2: Annotation of the Panax notoginseng image dataset
[0023] The images retained in step 1.1 are used as the new Panax notoginseng image dataset D', where D' = {d'1, d'2, ..., d'} n}, where d' i (1≤i≤n) represents the unlabeled raw images in the image dataset D'. For d' i Perform annotation to obtain the corresponding real labeled data g. i (1≤i≤b), g i ={g cls ,g x ,g y ,g w ,g h}. Where g cls The images of Panax notoginseng are numbered according to their odor, injury, and normal characteristics, with numbers 1, 2, and 3 representing the odor, injury, and normal categories, respectively; the image d' is... i The top-left pixel is set as the origin, the x-axis points to the right, and the y-axis points downwards. g is then calculated. x =x' / w, g y =y' / h, g w =w' / w and g h =h' / h, where x' and y' represent the x and y coordinates of the target center point, respectively, w' and h' represent the width and height of the true bounding box of the 37 image, respectively, and w and h represent the total width and total height of the image, respectively.
[0024] 1.3: Partitioning of the Panax notoginseng image dataset
[0025] Divide the Panax notoginseng image dataset D' from step 1.2 into a training set D'. train , Validation set D' val and test set D' test Three parts, of which D' train 80% of the data was used for training the odor detection model for Panax notoginseng, D' val 5% was used for the validation of the odor detection model in the Panax notoginseng septicemia area, D' test It accounts for 15% and is used for testing the odor detection model in the area where Panax notoginseng is damaged.
[0026] 1.4: Image scaling, normalization, and data augmentation of Panax notoginseng
[0027] Firstly, the three-seven image in D' is adjusted to 640x640 pixels conforming to the model input, and the adjustment process is: the three-seven image is scaled by the same ratio according to the width and height, and the blank area appearing in the scaling process is filled with a gray bar. Then, each pixel value in D' is mapped to a value in the range of 0-1 by dividing each pixel value by 255. Finally, in order to make the model learn more features of the three-seven image and improve the generalization ability of the model, the three-seven image in D' is processed using the Mosaic data enhancement algorithm. train and D' val each pixel value in the three-seven image is mapped to a value in the range of 0-1. Finally, in order to make the model learn more features of the three-seven image and improve the generalization ability of the model, the three-seven image in D' train is processed using the Mosaic data enhancement algorithm.
[0028] 2: Three-seven stench area detection model construction
[0029] The three-seven stench area detection model built is based on YOLOv7, and its overall architecture includes a three-seven feature extraction network, a three-seven feature fusion network, and a three-seven detection network. Among them, the three-seven feature extraction network is used to extract three-seven shallow features. The three-seven feature fusion network is composed of a feature fusion module and a feature re-extraction module. By improving the feature fusion module, the feature information missed by the parallel pooling operation is fully fused, so that more rich three-seven deep features are obtained. In view of the problem of missing detection and mis-detection of small stench areas caused by complex background noise, a collaborative attention mechanism is introduced in the feature re-extraction module to suppress the interference of background noise, and the three-seven deep features and the three-seven shallow features are more carefully re-extracted and fused, thereby reducing the missing detection rate and mis-detection rate of the three-seven small stench area. The three-seven detection network is responsible for processing the three-seven features output by the three-seven feature fusion network to obtain a feature map containing result prediction boxes.
[0030] 2.1: Three-seven feature extraction network construction
[0031] The three-seven feature extraction network is stacked by multiple CBS modules, ELAN modules and MP modules, which are used to extract three-seven shallow features. Among them, the CBS module is composed of a convolution layer, a batch normalization layer and a SiLU activation function; the ELAN module is stacked by a CBS module and a residual connection; the MP module uses a maximum pooling layer with a size of 2 and a step of 2 and a 3x3 convolution layer with a step of 2 to down-sample the feature map processed by the ELAN module, then uses two 1x1 convolution layers for dimension reduction processing respectively, and finally uses a Concat module to add the two feature maps processed by dimension reduction to obtain a three-seven shallow feature map. The three-seven shallow feature extraction process is shown in Table 1, wherein, C i ((x i ×y i ×c i ), (1≤i≤11)) is the three-seven shallow feature map extracted by the corresponding CBS convolution layer, ELAN module and MP module, xi x y i denotes the feature map size, c i is the number of feature map channels, C0 is the feature map of the input ginseng image with a size fixed to 640x640 by step 1.4, and the number of channels is 3.
[0032] Table 1. Ginseng shallow feature extraction process
[0033] Feature extraction module Input feature map Output feature map CBS module 1 [C0(x0×y0×c0)] [C1(x1×y1×c1)] CBS module 2 [C1] [C2(x2 x y2 x c2)] CBS module 3 [C2] [C3(x3×y3×c3)] CBS module 4 [C3] [C4(x4×y4×c4)] ELAN module 1 [C4] [C5(x5×y5×c5)] MP module 1 [C5] [C6(x6×y6×c6)] ELAN module 2 [C6] [C7(x7×y7×c7)] MP module 2 [C7] [C8(x8 x y8 x c8) <!-- 3 -->]]> ELAN module 3 [C8] [C9(x9×y9×c9)] MP module 3 [C9] C 10 (x 10 ×y 10 ×c 10 )]]> ELAN module 4 [C 10 ]]> C 11 (x 11 ×y 11 ×c 11 )]]>
[0034] 2.2: Ginseng feature fusion network construction
[0035] 2.2.1: SPPFCSPC module construction
[0036] The spatial pyramid pooling cross-stage-partial-connections (SPPCSPC) module of the YOLOv7 model is used to fuse and enhance the ginseng deep features obtained from the ginseng shallow features extracted in step 2.1. The spatial pyramid pooling-fast cross-stage-partial-connections (SPPFCSPC) module proposed in the present application improves the original parallel max pooling with sizes of 5, 9 and 13 to three series max pooling with a size of 5 when processing ginseng shallow features. Not only can it achieve the same fusion effect as the SPPCSPC module, but also can reduce the computational complexity while keeping the receptive field unchanged, thereby improving the detection efficiency of the model.
[0037] The SPPFCSPC module includes 7 CBS modules, 3 max pooling modules with a size of 5, and two Concat modules. The ginseng feature fusion process is as follows: first, use three 1x1 CBS modules to reduce the dimension of the ginseng shallow feature map C 11 output in step 2.1 to obtain the feature map C 12 ; then input C 12 into three series max pooling layers with a convolution kernel size of 5 and a step of 1 to obtain the ginseng feature map C 13 ; then use the Concat module to splice C 12 and C 13 , and use two 3x3 CBS modules to convolve the spliced ginseng feature map to obtain the feature map C 14 . Finally, use the Concat module to splice C 14 and the original input feature map C 11 , and then pass through a 1x1 CBS module to obtain the ginseng deep feature map C15 .
[0038] 2.2.2: PANet module construction
[0039] The path aggregation network (PANet) of the pseudo-ginseng odor area detection model reextracts and deeply fuses the pseudo-ginseng shallow features extracted in step 2.1 and the pseudo-ginseng deep features fused in step 2.2.1. This process is shown in Table 2.
[0040] Table 2. Pseudo-ginseng feature reextraction and fusion process
[0041]
[0042]
[0043] The feature reextraction module ELAN-H of YOLOv7 realizes the extraction and fusion of pseudo-ginseng features through dense residual stacking. The present application introduces a coordinate attention mechanism (CA) in the second layer of residual stacking of ELAN-H, and the obtained feature reextraction module ELAN-CA not only can enhance the extraction and fusion of small odor area features in pseudo-ginseng deep features, but also can suppress the interference of background noise on small odor area detection, and reduce the missed detection rate and false detection rate of pseudo-ginseng small odor areas.
[0044] The feature reextraction module ELAN-CA of the pseudo-ginseng odor area detection model is constructed, and the feature map C i (i = 19, 24, 26, 28) to be processed is denoted as Wherein, C, H and W represent the channel number, height and width of the pseudo-ginseng feature map respectively. The feature reextraction process is specifically as follows:
[0045] Firstly, for the input pseudo-ginseng feature map C CA decomposes the global max pooling into a pair of one-dimensional feature encoding operations, and uses convolution kernels with sizes of (H, 1) and (1, W) to encode each channel along the horizontal and vertical coordinate directions respectively, so the output of the cth (1 ≤ c ≤ C) channel with a height of H and the output of the cth channel with a width of W are respectively shown in formula (2-1) and formula (2-2).
[0046]
[0047] Wherein, x c (h, j) represents the pixel value at the coordinate (c, h, j) in the pseudo-ginseng feature map. and The position information of the captured three-seven deep feature in the wound-smell area feature can be fully utilized to suppress background noise interference and accurately locate the small wound-smell area in the three-seven image, and the relationship between the channels can also be effectively captured.
[0048] Then, the two three-seven feature maps containing rich wound-smell area information are cascaded and F1 convolution transformed, as shown in equation (2-3).
[0049] f = δ (F1 ([z h , z w ])) (2-3)
[0050] Where δ represents the ReLU function; the generated is an intermediate feature map that encodes the spatial information of the three-seven feature map in the horizontal and vertical directions, and r (r > 0) represents the down-sampling ratio used to control the size of the module.
[0051] Next, f is divided into two separate three-seven feature vectors and Two convolution transformations F H and F W are used to transform the three-seven feature vectors f h and f w to the same number of channels as the input , resulting in the results shown in equations (2-4) and (2-5).
[0052] g h = σ (F h (f h )) (2-4)
[0053] g w = σ (F w (f w )) (2-5)
[0054] Where σ represents the sigmoid function.
[0055] Finally, g h and g w are multiplied as the attention weight and the input three-seven feature map C i (i = 19, 24, 26, 28) to obtain the output of the three-seven feature map C i (i = 20, 25, 27, 29) at a certain pixel, as shown in equation (2-6).
[0056]
[0057] 2.3: Three-seven detection network construction
[0058] The Sanqi detection network uses the YOLO_Head module to process the Sanqi features extracted and fused by the Sanqi feature fusion network, and finally outputs a feature map containing Sanqi smell, injury and normal area prediction box information. The network layers of the YOLO_Head module and the input and output feature maps are shown in Table 3.
[0059] Table 3. Yolo_Head network layers and feature maps
[0060] Yolo_Head network layer Input feature map Output feature map CBS module 5 + RepConv module [C 25 ]]> C 30 (x 30 ×y 30 ×c 30 )]]> CBS module 6 + RepConv module C 27 ]]> C 31 (x 31 ×y 31 ×c 31 )]]> CBS module 7 + RepConv module [C 29 ]]> C 32 (x 32 ×y 32 ×c 32 )]]>
[0061] The YOLO_Head module takes C 25 , C 27 and C 29 in step 2.2 as input, and goes through the corresponding CBS convolutional layers and RepConv to obtain outputs C 30 , C 31 and C 32 , where RepConv uses the idea of structured model reparameterization to separate the network structure of the Sanqi injury and smell area detection model in the training and prediction stages. In the training stage, the network structure of RepConv consists of two Conv layers and three BN operations, thereby achieving higher detection accuracy; while in the detection stage, the network structure of RepConv consists of one Conv and one BN operation, which is equivalent to the network structure in the training stage and improves the detection efficiency of the Sanqi injury and smell area detection model under the premise of unchanged detection accuracy. 30 , C 31 and C 32 are three feature maps of different scales of the output Sanqi image, and their sizes are shown in Table 4.
[0062] Table 4. Size of Yolo_Head output feature map
[0063] Output of Yolo_Head Size of feature map [C 30 ]]> x 30 x y 30 x [(fes + classes) x n anchonr ]]]> [C 31 ]]> x 31 x y 31 x [(fes + classes) x n anchonr ]]]> [C 32 ]]> x 32 x y 32 x [(fes + classes) x n anchonr ]]]>
[0064] In Table 4, fes={p o , t x , t y , t w , t h} represents a set of predicted values corresponding to the 5 data values labeled in the Sanqi image dataset D', p o represents the confidence that the target is contained in the prediction box, t x and t y represent the center point coordinates of the Sanqi image prediction box, t w and t h represent the width and height of the Sanqi image prediction box, respectively, and classes={score1, score2, score3} represents a set of predicted values of the three target objects labeled in D', i.e. smell, injury and normal, nanchonr = 3 represents the number of initial prior boxes on each three-seven feature point. C 30 The length x 30 = 80, the width y 30 = 80, represents the feature map obtained by downsampling the input 640x640 three-seven image by 8 times, a total of 80x80 = 6400 feature points, each feature point has a minimum receptive field relative to the original three-seven image, and is used to detect smaller regions in the three-seven image; C 31 The length x 31 = 40, the width y 31 = 40, represents the feature map obtained by downsampling the input 640x640 three-seven image by 16 times, a total of 1600 feature points, each feature point has a moderate receptive field relative to the original three-seven image, and is used to detect medium-sized regions in the three-seven image; C 32 The length x 32 = 80, the width y 32 = 80, represents the feature map obtained by downsampling the input 640x640 three-seven image by 32 times, a total of 400 feature points, each feature point has a maximum receptive field relative to the original three-seven image, and is used to detect larger regions in the three-seven image.
[0065] 3: Three-seven stench region detection model training
[0066] 3.1: Three-seven image real box feature map calculation
[0067] 3.1.1: Prior box generation
[0068] The prior box is a pre-set width and height frame area, and the prediction box predicted by the network model is represented by the prior box and the corresponding offset. The present application uses the K-Means++ algorithm to optimize the selection process of the three-seven image real box clustering center, thereby generating three-seven image prior boxes with large size span, which not only avoids the local optimal problem, but also effectively solves the problem of insufficient feature extraction for small stench regions in the three-seven image. The steps of generating a prior box using the K-Means++ algorithm are as follows: first, randomly select a three-seven image real box of a certain size from the three-seven image training set D' train , denoted as c j (1≤j≤n); then calculate the intersection over union (IoU) using c j and the remaining three-seven image real boxes x train (1≤m≤n-1) in D' m , and the Euclidean distance between x m and the existing clustering center is denoted as D(x m , c j ); finally, calculate the remaining size real box xm The probability P (P ≥ 0) of being selected as a cluster center. Repeat the above step of generating prior boxes until the maximum P of 9 different sizes of prior boxes is obtained.
[0069]
[0070] 3.1.2: Real box feature map calculation
[0071] First, according to the real box information g i of the ginseng image, the grid where the center of the real box is located in the feature map C i (i = 30, 31, 32) is determined, and the grid and the two nearest neighbor grids are selected as the screening area of the ginseng image positive sample. Then, the IoU of the ginseng image real box and the prior box in the screening area is calculated, and the ginseng image prior box anchor gt with the maximum IoU is selected to perform regression calculation with the real box, so as to obtain the ginseng image prediction box. The ginseng image real box and the prior box anchor gt are extracted together as a feature vector v, v = {(p0, a x , a y , a w , a h , p1, p2, p3) x n a}, p0 is the confidence of the grid containing the ginseng target to be detected, n a is the number of initial prior boxes, a x , a y , a w and a h respectively represent the offset of the ginseng image real box and the corresponding prior box anchor gt , and the calculation formula is shown in (3-2) ~ (3-5).
[0072] a x = g x -c x (3-2)
[0073] a y = g y -c y (3-3)
[0074] a w = log(g w / p w ) (3-4)
[0075] a h = log(g h / p h ) (3-5)
[0076] Where cx and c y are the width and height of the selected three ginger image anchor i (i = 30, 31, 32) are the x and y coordinates of the top-left corner grid, p w and p h are the width and height of the selected three ginger image anchor gt . p1, p2 and p3 represent the confidence scores of the three ginger stink, injury and normal targets, respectively. Finally, the feature vectors obtained by each grid are spliced to obtain the feature map and
[0077] 3.2: Three ginger image optimal prediction box selection
[0078] Based on the three ginger feature map C i (i = 30, 31, 32) obtained in step 2.3, all prediction boxes of the three ginger stink, injury and normal targets are predicted, and then the non-maximum suppression method (NMS) is used to select the three ginger image prediction box B best with the highest confidence score. The steps of NMS to select the optimal three ginger image prediction box are as follows: select the prediction box B c with the maximum confidence score in the current category prediction box, calculate the IoU of B c and the remaining three ginger image prediction boxes by formula (3-6), and if the IoU is greater than the set threshold μ (0 ≤ μ ≤ 1), it is discarded. Finally, all the remaining n three ginger image prediction boxes B = {B1, B2, …, B n} are retained.
[0079]
[0080] where s in is the intersection of the three ginger image prediction box and the true box, and s union is the union of the three ginger image prediction box and the true box.
[0081] The n three ginger image prediction boxes obtained are expressed in the form of feature vectors using the corresponding anchor and offset, and are spliced to form feature maps C b1 , C b2 and C b3 .
[0082] 3.3: Loss function calculation
[0083] Because there are a large number of small bruise areas in the pseudo-ginseng image, and the pseudo-ginseng image prediction box regression loss has insufficient positioning ability for small bruise areas, positioning deviation is prone to occur. Therefore, the pseudo-ginseng image prediction box regression loss is reconstructed, and an angle cost is introduced therein to enhance the positioning accuracy of the model for small bruise areas by fixing the regression direction, effectively solving the problem of missed detection and false detection of small bruise areas, and enhancing the robustness of the model. The loss function of the pseudo-ginseng image detection model includes a pseudo-ginseng image classification loss, a pseudo-ginseng image confidence loss, and a pseudo-ginseng image prediction box regression loss.
[0084] (1) Pseudo-ginseng image classification loss used to calculate the error between the predicted class and the true class, and the calculation formula is shown in (3-7).
[0085]
[0086] where m represents the number of all prediction boxes in the pseudo-ginseng feature map, q represents the number of classes, y i represents the true class label in the pseudo-ginseng image, represents the class score predicted by the pseudo-ginseng bruise area detection model.
[0087] (2) Pseudo-ginseng image confidence loss used to calculate the error between the confidence score of the pseudo-ginseng contained in the prediction box and the true score, and the calculation formula is shown in (3-8).
[0088]
[0089] where x i is the pseudo-ginseng image label confidence, is the confidence score predicted by the pseudo-ginseng bruise area detection model.
[0090] (3) Pseudo-ginseng image prediction box regression loss used to calculate the error between the pseudo-ginseng image prediction box and the true box, and the pseudo-ginseng bruise area detection model uses a Complete IoU (Complete Intersection over Union, CIoU) loss function to calculate the regression loss of the network model, and the calculation formula is shown in (3-9) to (3-11).
[0091]
[0092] where, ρ 2 (b,b gt ) is the Euclidean distance between the center points of the pseudo-ginseng image prediction box and the true box, l is the diagonal distance of the smallest closed region containing the pseudo-ginseng image prediction box and the true box, w and h are the width and height of the pseudo-ginseng image prediction box, respectively, w gt and hgt α and v represent the width and height of the ground truth bounding box in the 3 / 7 image, respectively. α is the balance factor, and v is the aspect ratio penalty term.
[0093] As shown in formula (3-9), when the aspect ratio of the predicted bounding box of the Panax notoginseng image is the same as that of the true bounding box, the aspect ratio penalty term v is 0, the balance factor fails, and the expression of the CIoU loss function becomes unstable, resulting in a decrease in the model's detection accuracy. Therefore, this invention reconstructs the regression loss function of the predicted bounding box of the Panax notoginseng image, introduces angle cost to redescribe the distance, reduces the degrees of freedom of the loss function, enhances the model's localization accuracy for small damaged and smelly regions by fixing the regression direction, and improves the model's training and detection efficiency. The calculation of the reconstructed regression loss function of the predicted bounding box of the Panax notoginseng image is divided into angle calculation, distance calculation, and shape calculation. The angle calculation formulas are shown in (3-12) to (3-15).
[0094]
[0095] Where Λ is the final result of the angle calculation, x is the sine of the angle α between the center point of the ground truth bounding box and the center point of the predicted bounding box in the 3 / 7 image, d is the distance between them, and B h It is the relative height difference between two points.
[0096] The distance calculation formulas are shown in (3-16) to (3-17).
[0097]
[0098] Where, ρ x and ρ y It is the square of the ratio of the relative distance between the center points of the ground truth bounding box and the predicted bounding box in the 37 image on the x-axis and y-axis, and the ratio of the width and height of their smallest bounding rectangle, where e is Euler's constant.
[0099] The shape calculation formulas are shown in (3-18) to (3-19).
[0100]
[0101] Wherein, θ (θ>0) is the attention coefficient in the shape calculation formula, which is defined as a value between 2 and 6 in the Panax notoginseng odor detection model.
[0102] The formula for calculating the regression loss of the reconstructed Panax notoginseng image prediction box is shown in (3-20).
[0103]
[0104] (4) The overall loss of the Panax notoginseng odor detection model is the sum of the Panax notoginseng image classification loss, the Panax notoginseng image confidence loss and the regression loss of the reconstructed Panax notoginseng image prediction box, and the calculation formula is shown in (3-21).
[0105]
[0106] 3.4: Iterative update of the weight of the Panax notoginseng bruise area detection model
[0107] According to the overall loss function of the Panax notoginseng bruise area detection model calculated in step 3.3 Calculate the gradient g, which is calculated by taking the partial derivative of the loss function with respect to each variable. The Stochastic Gradient Descent (SGD) method is used to optimize the weight parameters in the network model, and the calculation formula is shown in (3-22).
[0108] W t+1 = W t - εg t (3-22)
[0109] where W t and W t+1 are the parameters after SGD optimization of the weight parameters, ε is the learning rate, and g t represents the average gradient of a batch of samples.
[0110] After multiple iterations of calculation, the network weight is updated, and the weight parameter W of the Panax notoginseng bruise area detection model is finally obtained. When training the network model, the iteration period epoch (i.e., the number of times the Panax notoginseng bruise area detection model traverses D' train and D' val ) needs to be set. After the training is completed, the weights generated in each iteration period are tested by the D' test test data set, and the weight parameter W best with the highest prediction accuracy is selected as the final weight of the Panax notoginseng bruise area detection model.
[0111] 4: Panax notoginseng bruise area detection and grade evaluation
[0112] 4.1: Panax notoginseng bruise area detection
[0113] Adjust the Panax notoginseng image to be detected to 640x640 pixels, and perform normalization processing on the adjusted Panax notoginseng image according to step 1.4 to make it meet the input standard of the Panax notoginseng bruise area detection model.
[0114] The weight parameter W bestThe preprocessed Panax notoginseng images are loaded into the YOLOv7 target detection network, and then input into the Panax notoginseng foul area detection model based on YOLOv7 to obtain the prediction information of the model on the foul, damaged and normal areas in the Panax notoginseng image, including the prediction of the category of the target area in the Panax notoginseng image, and the prediction box and confidence score. The prediction box of the Panax notoginseng image obtained is denoted as L, L = {l1, l2,..., ln} (n≥1), wherein l n = {t i ,t x ,t y ,t w ,t h}(1≤i≤n) represents any prediction box coordinates.
[0115] 4.2: Panax notoginseng grade evaluation
[0116] For any prediction box in step 4.1, l i is converted into coordinates in the real Panax notoginseng image, the left upper corner coordinates are (x' ij ,y' ij ), and the right lower corner coordinates are wherein W'(W'>0) and H'(H'>0) represent the width and height of the Panax notoginseng image respectively, i∈(1,2,3), 1, 2 and 3 represent that the detected target category is foul Panax notoginseng, damaged Panax notoginseng and normal Panax notoginseng respectively, and j(j≥0) represents the number of prediction boxes in the Panax notoginseng image.
[0117] When the prediction box coordinates of the Panax notoginseng image are output as (x' 3j ,y' 3j ) and (x" 3j ,y" 3j ), it indicates that the Panax notoginseng is normal Panax notoginseng; when the prediction box coordinates are output as (x' 1j ,y' 1j ) and (x" 1j ,y" 1j ), it indicates that the Panax notoginseng is foul Panax notoginseng and has lost medicinal value, which is discarded; when the prediction box coordinates are output as (x' 1j ,y' 1j ), (x" 1j ,y" 1j ) and (x' 2j ,y' 2j ), (x" 2j ,y" 2j ), it indicates that the Panax notoginseng is not only damaged but also foul and deteriorated, which has lost medicinal value and is discarded; when the prediction box coordinates are output as (x' 2j ,y' 2j ) and (x" 2j ,y"2j ) indicates that the ginseng is damaged ginseng, only due to the fracture extrusion caused damage, has not yet rotten metamorphism, still has medicinal value. Using η (0≤η≤1) to represent the damage degree of ginseng, the formula of η is shown in (4-1).
[0118]
[0119] wherein, (x” 2j -x' 2j ) and (y” 2j -y' 2j ) represent the length and width of the ginseng image wound area prediction box respectively, s hurt =(x” 2j -x' 2j )(y” 2j -y' 2j ) represents the area of a single prediction box of ginseng image wound area, represents the area of all prediction boxes in ginseng image, N represents the number of ginseng image wound area prediction boxes, and Q represents the number of ginseng image normal area prediction boxes.
[0120] Further, the ginseng grade evaluation can be carried out based on the detection result of the model, for example, Table 5 gives a ginseng grade evaluation and recycling value standard.
[0121] Table 5. Ginseng grade evaluation and recycling value standard
[0122] Value of η Level Recycling value η ≤ 0.1 A Extremely high 0.1 < η ≤ 0.3 B High 0.3 < η < 0.5 C Normal 0.5 < η < 0.7 D Low η > 0.7 E Extremely low III. DETAILED DESCRIPTION
[0124] The specific embodiments of the present application are described below in conjunction with the accompanying drawings, so that those skilled in the art can better understand the present application. Further detailed description, it should be particularly noted that in the following description, when the detailed description of the known function and design may dilute the main content of the present application, these descriptions will be ignored here.
[0125] 1: Ginseng image data preprocessing
[0126] According to step 1.1, the ginseng image data set D collected by Yunnan Baiyao Group Co., Ltd. is used as the initial data set, and the operator matrix and the threshold value Y=120 are set according to experience. For d i , its S di is calculated according to the Laplacian variance algorithm, if , d i is determined as high-quality ginseng image and retained in D, otherwise d i is determined as low-quality ginseng image and removed from D. The ginseng image data screening process is shown in Table 6.
[0127] Table 6. Screening process for Panax notoginseng image data
[0128]
[0129] Following step 1.2, the Panax notoginseng image dataset D'={d'1,d'2,…,d'... after step 1.1 is processed... n The images in the table are manually annotated sequentially. The target objects for annotation are the odorous area, damaged area and normal area of Panax notoginseng. The true bounding box of each target object is represented by 5 data values of the annotation. The annotation process of some Panax notoginseng image information is shown in Table 7.
[0130] Table 7. Annotation process of Panax notoginseng image information
[0131] Serial number Labeled object g x ]]> g y ]]> g w ]]> g h ]]> class_id [d'1] Smell 0.782473 0.309245 0.328292 0.410156 1 [d'2] Wound 0.488851 0.557116 0.490566 0.616105 2 [d'3] Normal 0.65407 0.638143 0.459302 0.636465 3 …… …… …… …… …… …… …… d' n-1 ]]> Wound 0.888851 0.441569 0.211665 0.473565 2 d' n ]]> Smell 0.687563 0.510356 0.139407 0.324275 1
[0132] Following step 1.3, the training set, validation set, and test set of the 37 image dataset D' are divided. train D' val and D' test The division ratio is 80:5:15, and the data distribution of different categories of samples in D' is 1:1.
[0133] Following step 1.4, the image patches in the Panax notoginseng image dataset D' are sequentially scaled and normalized. Then, D'... train and D' val Mosaic data augmentation was performed on images of Panax notoginseng.
[0134] 2: Construction of a detection model for the odor-causing area of Panax notoginseng
[0135] Following step 2.1, a Panax notoginseng feature extraction network is used to extract shallow features from the Panax notoginseng image. The feature extraction process can be divided into five stages. The size of the output feature map in each stage is reduced by a factor of {2, 4, 8, 16, 32} compared to the input Panax notoginseng image C0 (640×640×3). The input and output feature map sizes for each stage are shown in Table 8.
[0136] Table 8. Feature extraction process of Panax notoginseng feature extraction network
[0137]
[0138] Following step 2.2, the shallow features of Panax notoginseng obtained in step 2.1 are re-extracted and deeply fused using a Panax notoginseng feature fusion network. The feature processing consists of five stages, where stage 1 corresponds to the operation in step 2.2.1, and stages 2 to 5 correspond to the operations in step 2.2.2. The input and output feature map dimensions for each stage are shown in Table 9.
[0139] According to step 2.2.1, the feature fusion module SPPFCSPC of the Sanqi feature fusion network is used to fully fuse the feature information missed by the parallel pooling operation, so as to obtain more rich Sanqi deep features. First, 3 CBS modules with a size of 1x1 are used to reduce the dimension of the Sanqi shallow feature map C 11 (20x20x1024) output in step 2.1 to obtain a feature map C 12 (20x20x512); C 12 (20x20x512) is input into 3 maximum pooling layers in series with a convolution kernel size of 5 and a step size of 1 to obtain a Sanqi feature map C 13 (20x20x256); C 12 (20x20x512) and C 13 (20x20x256) are spliced using a Concat module, and then 2 CBS modules with a size of 3x3 are used to perform convolution operation on the spliced Sanqi feature map to obtain a Sanqi feature map C 14 (20x20x1024). Finally, C 14 (20x20x1024) and the original input Sanqi feature map C 11 (20x20x1024) are spliced using a Concat module, and then dimension reduction is performed on the spliced Sanqi feature map using a CBS module with a size of 1x1 to obtain a Sanqi deep feature map C 15 (20x20x512). Experimental verification shows that after using the SPPFCSPC module proposed in the application, the mAP is improved from 80.1% to 81.3, and the recall rate is improved from 68.4% to 70.5%.
[0140] According to step 2.2.2, the feature re-extraction module ELAN-CA of the Sanqi feature fusion network is used to re-extract the Sanqi deep features and perform deep fusion with the Sanqi shallow features obtained in step 2.1. Stage 2 is responsible for up-sampling the Sanqi deep features in stage 1. First, a CBS module with a size of 1x1 is used to reduce the dimension of the Sanqi deep feature map C 15 (20x20x512) output in stage 1 to obtain a feature map C 16 (20x20x256); C 16 (20x20x256) is input into an Upsample module to perform up-sampling to obtain a feature map C 17 (40x40x256), and C 10 (20x20x1024) is also up-sampled to obtain a feature map C 18 (40x40x256), and a Concat module is used to splice the feature map C 17 (40x40x256) and the feature map C 18(40x40x256) is spliced by channel to obtain the ginseng feature map C 19 (40x40x512). Finally, the input ginseng feature map C 19 (40x40x512) is subjected to deep feature extraction and fusion using the ELAN-CA module to obtain the ginseng feature map C 20 (40x40x256). Stage 3 performs up-sampling processing on the feature map output by stage 2 to obtain the ginseng feature map C 25 (80x80x128).
[0141] Stages 4 and 5 are responsible for down-sampling processing of the ginseng feature map. First, the ginseng feature map C 25 (80x80x128) is input into the ginseng detection network, which is responsible for detecting small regions in the ginseng image; at the same time, C 25 (80x80x128) is input into the MP module of stage 4, which is subjected to down-sampling to obtain the feature map C 26 (40x40x512); then the EALN-CA module is used to perform deep feature extraction and fusion on C 26 (40x40x512) to obtain the feature map C 27 (40x40x256), the ginseng feature map C 27 (40x40x256) is input into the ginseng detection network, which is responsible for detecting medium regions in the ginseng image; at the same time, C 27 (40x40x256) is input into the MP module of stage 5, which is subjected to down-sampling to obtain the feature map C 28 (20x20x1024); finally, the EALN-CA module is used to perform deep feature extraction and fusion on C 28 (20x20x1024) to obtain the feature map C 29 (20x20x512), the ginseng feature map C 29 (20x20x512) is input into the ginseng detection network, which is responsible for detecting large regions in the ginseng image. Experimental verification shows that after using the improved ELAN-CA module of the application, the mAP is improved from 80.1% to 88.1%, and the recall rate is improved from 68.4% to 83%.
[0142] Table 9. Feature processing process of ginseng feature fusion network
[0143]
[0144]
[0145] According to step 2.3, the three-seven detection network is used to process the three-seven features extracted and fused by the three-seven feature fusion network, and finally outputs a feature map containing three-seven smell, wound and normal area prediction box information, and the process is shown in Table 10.
[0146] Table 10. Three-seven detection network detection process
[0147] Module Input size Output size CBS module 9 + RepConv module [C 25 ]]> C 30 (80 x 80 x 24) CBS module 10 + RepConv module [C 27 ]]> C 31 (40 x 40 x 24) CBS module 11 + RepConv module [C 29 ]]> C 32 (20 x 20 x 24)
[0148] C 30 (80x80x24), C 31 (40x40x24) and C 32 (20x20x24) are the prediction results output by the three-seven detection network. The three-seven feature map C 30 (80x80x24) has a length of 80 and a width of 80, indicating that the three-seven image of 640x640 is obtained by downsampling 8 times, and there are 80x80=6400 feature points, each feature point has a minimum receptive field relative to the original three-seven image, and is used to detect smaller areas in the three-seven image. The three-seven feature map C 31 (40x40x24) has a length of 40 and a width of 40, indicating that the three-seven image of 640x640 is obtained by downsampling 16 times, and there are 40x40=1600 feature points, each feature point has a moderate receptive field relative to the original three-seven image, and is used to detect medium-sized areas in the three-seven image. The three-seven feature map C 32 (20x20x24) has a length of 20 and a width of 20, indicating that the three-seven image of 640x640 is obtained by downsampling 32 times, and there are 20x20=400 feature points, each feature point has a maximum receptive field relative to the original three-seven image, and is used to detect larger areas in the three-seven image, where 24 represents the dimension of the three-seven feature map, which can be represented as [(5+3)x3]; 5 represents the center point coordinates, width and height of the prediction box, and target confidence; 3 in the small parentheses represents the number of prediction target categories, which are three categories of smell, wound and normal; 3 outside the small parentheses represents 3 prior boxes assigned to each feature point.
[0149] 3: Three-seven wound and smell area detection model training
[0150] According to step 3, the three-seven image training set D' train and the test set D' val are input into the three-seven wound and smell area detection model constructed in step 2 for training. During the training process, first, the three-seven image real box feature map is calculated according to step 3.1, then the feature map containing the optimal prediction box of the three-seven image is obtained according to step 3.2, and the offset amount of the real box and the predicted box of the paeony image the offset amount of the real box and the predicted box of the paeony image test test the weight parameters generated in each iteration period, and select the weight parameter W with the highest prediction accuracy best as the parameter of the paeony odor area detection model.
[0151] 4: Paeony odor area detection and grade evaluation
[0152] According to step 4.1, first adjust the paeony image to be detected to 640x640 pixels, and normalize the adjusted paeony image according to step 1.4 to make it meet the input standard of the paeony odor area detection model. Then input the processed paeony image into the paeony odor area detection model trained in step 3 to obtain the detection result containing the predicted box of the paeony image.
[0153] According to step 4.2, determine whether there is an odor area in the paeony image from the detection result obtained in step 4.1. If there is, it means that the paeony is a smelly paeony and cannot be used for medicine; if there is no odor area, it means that the paeony is a normal paeony or a damaged paeony, and the damaged paeony can be returned to medicine. Therefore, calculate the damage degree of the paeony according to formula (4-1) to determine the grade of the paeony. Table 11 is the calculation process of the damage degree of the paeony.
[0154] Table 11. Calculation process of damage degree of paeony
[0155]
[0156] For image 1, there are a total of 3 damages, and the damage degree is The damage level is B level, and the recycling value is high; for image 2, there are a total of 2 damages, and the damage degree is The damage level is C level, and the recycling value is general; for image 3, there are a total of 3 damages, and the damage degree is The damage level is D level, and the recycling value is low. For image n-1, there are a total of 4 damages, and the damage degree is The damage level is E level, and the recycling value is very low; for image n, there are a total of 2 damages, and the damage degree is The damage level is A level, and the recycling value is very high
[0157] Four, compared with the prior art, the application has the advantages and positive effects
[0158] (1) The application provides a YOLOv7-based ginseng bruise area detection and grade evaluation method, which overcomes the problems of the prior art, such as not considering the interference of background noise on small target feature extraction and insufficient small target feature extraction caused by large target size variation, and provides a ginseng grade evaluation method according to the detection results, providing a new grade evaluation scheme for other fields similar to ginseng bruise area detection.
[0159] (2) The application provides an improved feature fusion module for feature fusion and enhancement of a ginseng bruise area detection model, and optimizes the selection process of the ginseng image real box clustering center, solving the problem of insufficient feature extraction of small bruise areas in ginseng images.
[0160] (3) The application aims at the problem of missed detection and false detection of small bruise areas of ginseng, introduces a cooperative attention mechanism and reconstructs the ginseng image prediction box regression loss function, effectively suppresses the interference of background noise on ginseng image feature extraction, improves the positioning accuracy of small bruise areas, and greatly improves the detection accuracy of the ginseng bruise area detection model. BRIEF DESCRIPTION OF DRAWINGS
[0162] The drawings described herein are used to provide further understanding of the present application, and constitute a part of the present application, the illustrative embodiments of the present application and the description thereof are used to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:
[0163] Figure 1 The flowchart of the present application
[0164] Figure 2 The SPPFCSPC pooling processing structure
[0165] Figure 3 The ginseng image bruise area detection example.
Claims
1. A method for detecting and grading the odor-causing areas of Panax notoginseng based on YOLOv7, characterized in that, Includes the following steps: S1: Preprocessing of Panax notoginseng image data: Constructing a Panax notoginseng image dataset by collecting images of Panax notoginseng rhizomes, and using the Laplacian variance algorithm to filter out blurry and poor-quality images. The filtered Panax notoginseng image dataset is labeled and divided. Finally, the images in the dataset are scaled, normalized and data augmented. S2: Construction of the Panax notoginseng blemish and odor region detection model. A Panax notoginseng blemish and odor region detection model based on YOLOv7 is constructed. First, the Panax notoginseng feature extraction network is used to extract and process multi-scale features from the input Panax notoginseng image. Then, a richer deep Panax notoginseng feature is obtained through an improved feature fusion module. In the feature re-extraction module, a collaborative attention mechanism is introduced to suppress the interference of background noise. Finally, the prediction result is output after processing by the Panax notoginseng detection network. Specifically, it includes the following three sub-steps. S2.1: Construction of Panax notoginseng feature extraction network The Panax notoginseng feature extraction network consists of multiple stacked CBS, ELAN, and MP modules, used to extract shallow features of Panax notoginseng. The CBS module comprises convolutional layers, batch normalization layers, and the SiLU activation function; the ELAN module consists of stacked CBS modules and residual connections; and the MP module uses max pooling layers of size 2 and stride 2, and stride 2, respectively. The convolutional layer downsamples the feature map processed by the ELAN module, and then uses two... The convolutional layers perform dimensionality reduction processing separately, and finally, the Concat module is used to add the two dimensionality-reduced feature maps to obtain the 3 / 7 shallow feature map. The three-seven shallow feature map is extracted through the corresponding CBS convolutional layer, ELAN module and MP module. Representation of feature map Size, For feature map The number of channels, It is a feature map of an input Panax notoginseng image with a fixed size of 640×640 and 3 channels; S2.2: Construction of a Panax notoginseng Feature Fusion Network S2.2.1: SPPFCSPC Module Setup The spatial pyramid pooling module of the YOLOv7 model is used to fuse and enhance the shallow features of Panax notoginseng extracted from step S2.1 to obtain deep features. When processing the shallow features of Panax notoginseng, the original parallel max pooling with sizes of 5, 9 and 13 is improved to three cascaded max pooling with a size of 5. This not only achieves the same fusion effect as the SPPCSPC module, but also reduces the amount of computation while keeping the receptive field unchanged, thereby improving the detection efficiency of the model. The SPPFCSPC module includes 7 CBS modules, 3 max-pooling modules of size 5, and 2 Concat modules. Its 3 / 7 feature fusion process is as follows: First, using 3... The CBS module outputs the shallow feature map of Panax notoginseng from step 2.
1. Dimensionality reduction is performed to obtain the feature map. Then, The input is processed into three cascaded max-pooling layers with a kernel size of 5 and a stride of 1 to obtain the 3 / 7 feature map. Next, and Use the Concat module to concatenate, then use 2 The CBS module performs convolution on the concatenated 3 / 7 feature maps to obtain the feature map. Finally, and the feature map of the original input Use the Concat module to concatenate, then go through 1... The CBS module obtains the deep feature map of Panax notoginseng. ; S2.2.2: PANet Module Setup The path aggregation network of the Panax notoginseng blemish and odor region detection model will re-extract and deeply fuse the shallow features of Panax notoginseng extracted in step S2.1 and the deep features of Panax notoginseng fused in step S2.2.
1. The feature re-extraction module ELAN-H of YOLOv7 realizes the extraction and fusion of Panax notoginseng features through dense residual stacking. In the second layer of residual stacking of ELAN-H, a collaborative attention mechanism is introduced. The resulting feature re-extraction module ELAN-CA can not only enhance the extraction and fusion of blemish and odor region features in the deep features of Panax notoginseng, but also suppress the interference of background noise on the detection of blemish and odor region, and reduce the false negative rate and false positive rate of blemish and odor region of Panax notoginseng. The ELAN-CA feature re-extraction module, which constructs the odor region detection model for Panax notoginseng, extracts the feature maps to be processed. Recorded as ,in, , H and W The number of channels, height, and width of the 37 feature map are represented respectively. The specific feature re-extraction process is as follows: First, for the input 3 / 7 feature map CA decomposes global max pooling into a one-to-one one-dimensional feature encoding operation, using a size of and The convolutional kernels encode each channel along both the horizontal and vertical coordinate directions, therefore the height is... The The output of each channel has a width of The The outputs of each channel are shown in equations (2-1) and (2-2), respectively: (2-1) (2-2) Representing the coordinates in the characteristic diagram of Panax notoginseng Pixel value at that location, and It can fully utilize the location information of the damaged and odorous areas in the captured deep features of Panax notoginseng, suppress background noise interference, accurately locate the damaged and odorous small areas in the Panax notoginseng image, and effectively capture the relationship between channels; Then, the two Panax notoginseng feature maps containing rich information on the odorous regions are concatenated and summed. The convolution transformation is shown in equation (2-3): (2-3) in, Represents the ReLU function; the generated It is an intermediate feature map that encodes the spatial information of the 3 / 7 feature map in the horizontal and vertical directions. Indicates the downsampling ratio and Used to control the size of the module; Next, along the spatial dimension It is split into two separate 3 / 7 feature vectors. and Then use two convolution transformations and The eigenvectors of 37 and Transform to and input With the same number of channels, we obtain the results shown in equations (2-4) and (2-5): (2-4) (2-5) in, Represents the sigmoid function; Finally, and The 3 / 7 feature map serves as the weights for attention and the input. Multiplying them together yields the characteristic diagram of the three-seven combination. The output at a certain pixel is calculated using the formula shown in (2-6): (2-6) S2.3: Construction of Panax notoginseng detection network The Panax notoginseng detection network uses the YOLO_Head module to process the Panax notoginseng features extracted and fused by the Panax notoginseng feature fusion network, and finally outputs a feature map containing prediction box information of Panax notoginseng odor, injury and normal areas. The YOLO_Head module uses step S2.
2. As input, the data passes through the corresponding CBS convolutional layer and RepConv to obtain the output. Among them, RepConv uses the idea of reparameterization of structured models to separate the network structure of the Panax notoginseng odor detection model for training and prediction. In the training phase, the network structure of RepConv consists of two Conv layers and three BN operations, thereby achieving higher detection accuracy. In the detection phase, the network structure of RepConv consists of one Conv and one BN operation, which is equivalent to the network structure during training and improves the detection efficiency of the Panax notoginseng odor detection model without changing the detection accuracy. These are three feature maps at different scales of the output 37 image, and their sizes can be represented as follows: in, Represents the image dataset of 37. The set of predicted values corresponding to the 5 data values marked in the figure. This indicates that the prediction box contains the confidence level of the target. and This indicates the coordinates of the center point of the prediction box in the Panax notoginseng image. and These represent the width and height of the prediction bounding box for the Panax notoginseng image, respectively. express The set of predicted values for the three categories of target objects labeled as odorous, injured, and normal. This represents the number of initial prior boxes at each 3 / 7 feature point. length ,Width This indicates that the feature map obtained by downsampling a 640×640 Panax notoginseng image by 8 times has a total of 80×80=6400 feature points. Each feature point has the smallest receptive field relative to the original Panax notoginseng image, which is used to detect smaller areas in the Panax notoginseng image. length ,Width This indicates that the feature map obtained by downsampling a 640×640 Panax notoginseng image by 16 times has a total of 1600 feature points. Each feature point has a moderate receptive field relative to the original Panax notoginseng image, which is used to detect medium-sized regions in the Panax notoginseng image. length ,Width This indicates that the feature map is obtained by downsampling a 640×640 Panax notoginseng image by 32 times. It contains 400 feature points, each with the largest receptive field relative to the original Panax notoginseng image, which is used to detect larger areas in the Panax notoginseng image. S3: Training the Panax notoginseng blemish detection model. Input the Panax notoginseng image training set and validation set constructed in step S1 into the model in step S2, calculate the loss function and perform backpropagation, and iteratively update the model weights. S4: Detection and grading of scorched and odorous areas of Panax notoginseng. Input the Panax notoginseng image test set into the model trained in step S3, detect the scorched and odorous areas of Panax notoginseng, and grade the Panax notoginseng according to the detection results.
2. The method for detecting and grading the odor-causing areas of Panax notoginseng based on YOLOv7 according to claim 1, characterized in that, Step S1 further includes the following specific steps: S1.1: Screening of Panax notoginseng image data The collected Panax notoginseng image dataset is denoted as... , , for For any image in the dataset, use the Laplacian variance algorithm to obtain the variance from... Select those that meet the requirements The calculation process is as follows: Select The Laplacian operator matrix A pairs After performing convolution, the variance is calculated. Set the threshold to and ,like Then determine For high-quality Panax notoginseng images and preserve them If it is correct, otherwise it will be judged. For low-quality Panax notoginseng images, and from Remove from; S1.2: Annotation of Panax notoginseng image dataset The images retained in step S1.1 are used as the new Sanqi image dataset. , ,in For image datasets The original, unlabeled image in the middle, for Perform annotation to obtain the corresponding actual annotation data. , ,in The images of Panax notoginseng are numbered according to their odor, injury, and normal condition. Numbers 1, 2, and 3 represent the odor, injury, and normal condition categories, respectively. The images are then... The top-left pixel is set as the origin, and the pixels to the right are... x The axis, downwards is y Axis, calculated ,in, and Representing the target center point respectively x and y coordinates and These represent the width and height of the true bounding box of the 37 image, respectively. and These represent the total width and total height of the image, respectively. S1.3: Partitioning of the Panax notoginseng image dataset The Panax notoginseng image dataset from step S1.2 Divided into training set Validation set and test set Three parts, among which 80% of the data was used for training the model to detect areas of odor caused by Panax notoginseng. 5% was used for the validation of the odor detection model in areas affected by Panax notoginseng. 15% was used for testing the odor detection model in areas affected by Panax notoginseng; S1.4: Image scaling, normalization, and data augmentation for Panax notoginseng First, The Panax notoginseng image was adjusted to conform to the model input of 640×640 pixels. The adjustment process was as follows: the Panax notoginseng image was scaled proportionally by width and height, and blank areas appearing during the scaling process were filled with gray bars. Then, each pixel value was divided by 255. and Each pixel value in the Panax notoginseng image is mapped to a value in the range of 0 to 1. Finally, to enable the model to learn more Panax notoginseng image features and improve the model's generalization ability, [the following steps are taken]. The Panax notoginseng images were processed using the Mosaic data augmentation algorithm.
3. The method for detecting and grading the odor-causing areas of Panax notoginseng based on YOLOv7 according to claim 1, characterized in that, Step S3 further includes the following specific steps: S3.1: Calculation of the true bounding box feature map of Panax notoginseng image S3.1.1: Prior Box Generation A prior bounding box is a pre-defined region with a pre-set width and height. The predicted bounding boxes generated by the network model are represented by the prior bounding boxes and their corresponding offsets. The K-Means++ algorithm is used to optimize the selection process of the cluster centers of ground truth bounding boxes in Panax notoginseng images, thereby generating prior bounding boxes with a larger size range. This not only avoids the local optima problem but also effectively solves the problem of insufficient feature extraction of small damaged and smelly regions in Panax notoginseng images. The steps of generating prior bounding boxes using the K-Means++ algorithm are as follows: First, from the Panax notoginseng image training set... A real bounding box of a 37-inch image of random size is selected as the cluster center, denoted as . Then use and The true frame of the 37 image in other sizes Calculate the Intersectionover Union (IoU) ratio and then... The Euclidean distance from the existing cluster centers is denoted as Finally, use formula (3-1) to calculate the remaining true frames. The probability of being selected as a cluster center and Repeat the above steps to generate the prior box until the result is obtained. The nine largest different sizes of prior boxes; (3-1) S3.1.2: Calculation of the feature map of the ground truth bounding box First, based on the true frame information of the Panax notoginseng image. Determine the center of the ground truth bounding box in the feature map. The grid in which the image is located is selected, and the grid and its two nearest neighboring grids are chosen as the filtering region for positive samples of Panax notoginseng images. Then, the IoU between the ground truth bounding box of the Panax notoginseng image and the prior bounding box in the filtering region is calculated, and the prior bounding box of the Panax notoginseng image with the largest IoU is selected. Regression calculations are performed with the ground truth bounding boxes to obtain the predicted bounding boxes for the Panax notoginseng image. The predicted bounding boxes for the Panax notoginseng image are then compared with the predicted bounding boxes. They are all extracted as feature vectors v. , The confidence level is the presence of the target *Panax notoginseng* in the grid. The initial number of prior bounding boxes. These represent the ground truth bounding box and the corresponding prior bounding box in the 37 image, respectively. The offset is calculated using formulas (3-2) to (3-5): (3-2) (3-3) (3-4) (3-5) in, and These are the characteristic diagrams of Panax notoginseng. The x and y coordinates with the top-left corner of the grid as the origin. and The selected prior bounding boxes for the Panax notoginseng image are respectively Width and height, The confidence scores for the three categories of targets—smelly, damaged, and normal—are represented respectively. Finally, the feature vectors obtained from each grid are concatenated to obtain the feature map. ; S3.2: Selection of the optimal prediction box for Panax notoginseng image Based on step S2.3, the Panax notoginseng feature map is obtained. All predicted bounding boxes for the three target categories of odor, injury, and normal were analyzed, and then non-maximum suppression (NMS) was used to select the predicted bounding boxes with the highest confidence scores. The steps for NMS to select the optimal Panax notoginseng image prediction box are as follows: Select the prediction box with the highest confidence score in the current category prediction box. Calculated using formula (3-6) If the IoU with the other predicted bounding boxes in the 37-image is greater than a set threshold μ and μ≥0, it is discarded, and all remaining bounding boxes are ultimately retained. Three-seven image prediction box ; (3-6) in, This represents the intersection of the predicted bounding box and the ground truth bounding box in the 37 image. It is the union of the predicted bounding boxes and the ground truth bounding boxes in the 37 image; The result Each predicted bounding box in the 37-point image is represented as a feature vector using its corresponding prior bounding box and offset, and these vectors are then concatenated to form a feature map. ; S3.3: Loss Function Calculation Because there are a large number of small, odorous regions in Panax notoginseng images, and the Panax notoginseng image prediction bounding box regression loss is insufficient for locating these small regions, it is prone to localization errors. Therefore, the Panax notoginseng image prediction bounding box regression loss is reconstructed, and an angle cost is introduced into it. By fixing the regression direction, the model's localization accuracy for small, odorous regions is enhanced, effectively solving the problem of missed and false detections of small, odorous regions. At the same time, the robustness of the model is enhanced. The loss function of the Panax notoginseng image detection model includes Panax notoginseng image classification loss, Panax notoginseng image confidence loss, and Panax notoginseng image prediction bounding box regression loss. (1) Image classification loss of 37 The formula used to calculate the error between the predicted category and the true category is shown in (3-7): (3-7) in This represents the number of all predicted boxes in the 3 / 7 feature map. Indicates the number of categories. This represents the true category label in the Panax notoginseng image. This represents the category score predicted by the Panax notoginseng odor detection model. (2) Confidence loss of Panax notoginseng image The error between the confidence score and the true score for predicting that the box contains 37 is calculated using the formula shown in (3-8): (3-8) in Confidence of Panax notoginseng image labels The confidence score predicted by the Panax notoginseng odor detection model in the affected area; (3) Regression loss of Panax notoginseng image prediction box To calculate the error between the predicted bounding box and the ground truth bounding box in the Panax notoginseng image, the Panax notoginseng blemish region detection model uses the Complete IoU (Complete Intersection over Union, CIoU) loss function to calculate the regression loss of the network model, as shown in formulas (3-9) to (3-11): (3-9) (3-10) (3-11) in, Let be the Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box in the 37 image. The diagonal distance of the smallest closure region that simultaneously contains both the predicted bounding box and the ground truth bounding box of the 37 image. and These are the width and height of the prediction bounding box for the Panax notoginseng image, respectively. and These are the width and height of the true frame of the Panax notoginseng image, respectively. As a balance factor, This is a penalty for aspect ratio; As shown in formula (3-9), when the aspect ratios of the predicted bounding box and the true bounding box of the Panax notoginseng image are the same, the aspect ratio penalty term... When the value is 0, the balance factor fails, leading to instability in the expression of the CIoU loss function and a decrease in the model's detection accuracy. Therefore, the regression loss function of the 37-image prediction box is reconstructed. By introducing angle cost to redescribe the distance, the degrees of freedom of the loss function are reduced. By fixing the regression direction, the model's localization accuracy for small damaged areas is enhanced, while improving the model's training and detection efficiency. The calculation of the reconstructed regression loss function of the 37-image prediction box is divided into angle calculation, distance calculation, and shape calculation. The angle calculation formulas are shown in (3-12) to (3-15): (3-12) (3-13) (3-14) (3-15) in, It is the final result of the angle calculation. It is the angle between the center point of the ground truth bounding box and the center point of the predicted bounding box in the 37 image. The sine value, It is the distance between them. It is the relative height difference between two points; The distance calculation formulas are shown in (3-16) to (3-17): (3-16) (3-17) in, and It is the square of the ratio of the relative distance between the center points of the ground truth bounding box and the predicted bounding box in the 37 image on the x-axis and y-axis, and the ratio of the width and height of their smallest bounding rectangle, where e is Euler's constant; The shape calculation formulas are shown in (3-18) to (3-19): (3-18) (3-19) in, The attention coefficient in the shape calculation formula and ; The formula for calculating the regression loss of the reconstructed Panax notoginseng image prediction box is shown in (3-20): (3-20) (4) The overall loss of the Panax notoginseng odor region detection model is the sum of the Panax notoginseng image classification loss, the Panax notoginseng image confidence loss, and the regression loss of the reconstructed Panax notoginseng image prediction box. The calculation formula is shown in (3-21): (3-21) S3.4: Iterative update of weights in the Panax notoginseng odor detection model The overall loss function of the Panax notoginseng odor detection model calculated in step S3.3 is used. Calculate gradient , The weight parameters in the network model are calculated by taking the partial derivatives of the loss function with respect to each variable. The method of stochastic gradient descent (SGD) is used to optimize the weight parameters, and the calculation formula is shown in (3-22). (3-22) in, and These are the weight parameters after SGD optimization. For learning rate, This represents the average gradient of a batch of samples. After multiple iterations of calculation and updating of network weights, the final weight parameters W of the Panax notoginseng odor detection model are obtained. When training the network model, an iteration period of epoch needs to be set, i.e., the Panax notoginseng odor detection model traverses and processes... and The number of times, after the training, through The test dataset is used to test the weights generated in each iteration cycle, and the weight parameters with the highest prediction accuracy are selected. As the final weight of the Panax notoginseng odor detection model.
4. The method for detecting and grading the odor-causing areas of Panax notoginseng based on YOLOv7 according to claim 1, characterized in that, Step S4 further includes the following specific steps: S4.1: Detection of the odor-causing area of Panax notoginseng The image of Panax notoginseng to be detected is adjusted to 640×640 pixels, and the adjusted image is normalized according to step S1.4 to make it conform to the input standard of the Panax notoginseng blemish area detection model. The weight parameters of the Panax notoginseng odor detection model trained in step S3.4 are... The data is loaded into a YOLOv7 object detection network, and then the preprocessed Panax notoginseng image is input into a YOLOv7-based Panax notoginseng bruise and odor region detection model. The model's predictions of odorous, bruised, and normal regions in the Panax notoginseng image are obtained, including the predicted category of the target region, the predicted bounding box, and the confidence score. The predicted bounding box of the obtained Panax notoginseng image is denoted as L. ,in Represents the coordinates of any prediction box; S4.2: Panax notoginseng grading For any predicted bounding box in step S4.1, Converted to coordinates in a real Panax notoginseng image, the coordinates of the top left corner are: , , The coordinates of the lower right corner are , ,in, and These represent the width and height of the original image of Panax notoginseng, respectively. 1, 2, and 3 represent the detected target categories as rotten Panax notoginseng, damaged Panax notoginseng, and normal Panax notoginseng, respectively, and "j" represents the number of prediction boxes in the Panax notoginseng image; When the predicted bounding box coordinates of the Panax notoginseng image are output as and When the output of the prediction box coordinates is , it indicates that the Panax notoginseng is normal; when the output of the prediction box coordinates is and" When the output of the prediction box coordinates is 0, it indicates that the Panax notoginseng is rotten and has lost its medicinal value, and should be discarded; When the output of the prediction box coordinates is..., it indicates that the Panax notoginseng is not only damaged but also smelly and spoiled, and has lost its medicinal value; therefore, the Panax notoginseng should be discarded. When the damage is indicated, it means that the Panax notoginseng is damaged due to breakage and compression, but has not yet become smelly or spoiled, and still has medicinal value. The degree of damage is represented by "η". The formula for calculating "η" is shown in (4-1). (4-1) in, Let represent the length and width of the predicted bounding box for the wound region in the Panax notoginseng image, respectively. This represents the area of a single predicted bounding box in the damaged region of the Panax notoginseng image. denoted by , N represents the area of all predicted boxes in the Panax notoginseng image, N represents the number of predicted boxes in the injured area of the Panax notoginseng image, and Q represents the number of predicted boxes in the normal area of the Panax notoginseng image.
Citation Information
Patent Citations
Pseudo-ginseng quality evaluation method based on image recognition, server and evaluation system
CN115620281A