Watermelon stripe phenotype detection method

By combining the improved YOLOv8s neural network and the RealSense D455 depth camera, the problems of insufficient recognition accuracy and environmental adaptability in watermelon peel stripe analysis were solved, and high-precision stripe segmentation and multi-dimensional texture feature extraction were achieved. It can adapt to different lighting and environments and support watermelon variety identification and quality evaluation.

CN120635890APending Publication Date: 2025-09-12INSTITUTE OF VEGETABLES & FLOWERS CHINESE ACADEMY OF AGRICULTURAL SCIENCES
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510791459.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies for watermelon peel stripe analysis have problems such as insufficient stripe recognition accuracy and consistency, poor environmental adaptability, difficulty in extracting three-dimensional structure and surface detail information, and the inability of two-dimensional image acquisition equipment to accurately capture the curved surface structure and depth changes of the peel.

Method used

An improved YOLOv8s neural network was used in combination with the RealSenseD455 depth camera and illumination correction technology. The network structure was improved through the DSConv, DPAG and FEM modules. The depth camera was used to collect watermelon stripe images and perform illumination correction. Combined with LBP encoding and grayscale conversion processing, the stripe segmentation image and LBP histogram were generated.

Benefits of technology

It achieves adaptive recognition of complex peel structures and diverse stripes, improves stripe segmentation accuracy and multi-dimensional texture feature extraction, enhances the stability and accuracy of stripe analysis, adapts to different lighting and environmental conditions, and supports automated watermelon variety identification and quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635890A_ABST
    Figure CN120635890A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of peel stripe analysis, and provides a watermelon stripe phenotype detection method, which comprises the steps of YOLOv8s neural network improvement, watermelon stripe image acquisition, illumination correction and illumination compensation, stripe detection and detection result post-processing. According to the method, the YOLOv8s neural network is improved, so that the reasoning speed of the model, the quality of feature extraction and the fineness of feature extraction are improved; lBP coding is carried out on an image detection result, so that multi-level and multi-dimensional texture features of peel stripes are extracted, and the description capability of stripe forms and textures is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fruit peel stripe analysis, in particular to a method for detecting watermelon stripe phenotypes. Background Art

[0002] Existing technologies for analyzing watermelon peel stripes rely on traditional computer vision algorithms and image processing techniques to detect and analyze stripes. These methods typically use basic operations such as edge detection, grayscale threshold segmentation, and texture feature extraction to identify and analyze peel stripes.

[0003] However, these methods have obvious limitations in application, especially when faced with multiple varieties and diverse watermelon peel textures, it is difficult to ensure the accuracy and consistency of stripe recognition. The following analyzes the deficiencies of existing technologies from multiple aspects such as stripe segmentation, texture feature extraction, environmental adaptability and image acquisition technology. Insufficient stripe segmentation accuracy: Existing stripe analysis systems usually use grayscale threshold segmentation or edge detection to identify peel stripes. This method can provide basic recognition effects for stripe patterns with simple structures and obvious color contrast, but when faced with complex peel structures or stripes with unclear contrast, segmentation errors often occur. In addition, existing methods generally use fixed segmentation thresholds, which are prone to segmentation errors, affecting the accuracy of stripe feature analysis. Poor adaptability to environmental changes: Fruit peel stripe image acquisition is typically performed in natural or semi-natural environments, where lighting conditions, acquisition angles, and distances may vary significantly. Existing technologies are not very adaptable to environmental changes, especially under conditions of uneven or complex lighting. Stripe recognition is prone to distortion or noise. For example, under sunlight, strong light can easily create shadows in the stripe area, interfering with texture feature extraction, resulting in blurred stripe boundaries or loss of stripe information. Under insufficient or uneven lighting conditions, the smooth surface of the watermelon peel will produce reflections, significantly reducing the effectiveness of texture feature extraction and stripe recognition, affecting the stability and accuracy of the analysis. In addition, traditional image preprocessing methods have low adaptability to environmental changes and cannot dynamically adjust thresholds and parameters, resulting in unstable performance of existing systems under different lighting and environmental conditions. Limitations of image acquisition technology: Existing fruit peel stripe analysis is mostly based on two-dimensional image acquisition equipment, using ordinary cameras to capture stripes. This method cannot accurately capture the three-dimensional structure and surface details of the peel, and the depth, width, and morphological information of the stripes cannot be fully extracted. For example, watermelon peel has a certain curved surface structure, and the stripes observed in the two-dimensional image may be distorted due to the perspective effect, causing the actual width and direction of the stripes to deviate. In addition, the two-dimensional image is difficult to reflect the concave and convex characteristics of the stripes, and cannot identify the changes in the depth of the stripes, resulting in incomplete extracted stripe information and a lack of spatial structural information in the analysis results. Summary of the Invention

[0004] In order to overcome the deficiencies of the prior art, the object of the present invention is to provide a method for detecting watermelon stripe phenotypes to improve the accuracy and consistency of stripe identification.

[0005] To achieve the above object, the present invention provides the following solutions:

[0006] A method for detecting watermelon stripe phenotypes, comprising:

[0007] The convolution modules in the backbone network and the neck network of the YOLOv8s neural network are replaced by the DSConv module, a DPAG module is set at the connection between the backbone network and the neck network, and the C2f module in the backbone network of the YOLOv8s neural network is replaced by the FEM module to obtain an improved YOLOv8s model.

[0008] The left imager, right imager, and infrared projector of the RealSense D455 depth camera are used to collect watermelon stripe images for target detection.

[0009] Performing illumination correction and illumination compensation on the watermelon stripe image to obtain an illumination optimized image;

[0010] Inputting the illumination optimized image into the pre-trained improved YOLOv8s model to perform stripe detection to obtain an image detection result;

[0011] According to the image detection result, the watermelon stripe image is subjected to stripe segmentation, grayscale conversion, LBP coding calculation, statistics and normalization processing to obtain a stripe segmentation image, a stripe grayscale map and an LBP histogram.

[0012] Preferably, the shooting resolution and bit depth of the RealSense D455 depth camera are 848×480 and 24 respectively.

[0013] Preferably, the illumination compensation includes: logarithmic transformation and exponential transformation.

[0014] Preferably, the expression for illumination correction is:

[0015]

[0016] in, is the i-th target component x i The mapping output of β i is the i-th correction coefficient; n is the number of components; β0 is the basic component.

[0017] Preferably, stripe segmentation, grayscale conversion, LBP coding calculation, statistics and normalization processing are performed on the watermelon stripe image according to the image detection result to obtain a stripe segmentation image, a stripe grayscale image and an LBP histogram, including:

[0018] Convert the image detection result into a plurality of 8-neighborhood masks;

[0019] Convert the 8-neighborhood mask into several 8-bit binary codes using a preset coding formula;

[0020] Convert the 8-bit binary code into decimal data to obtain a pixel code;

[0021] Normalization is performed on all the pixel codes of the image detection result to obtain the LBP histogram.

[0022] Preferably, the DPAG module integrates a branch channel attention submodule and a channel attention submodule.

[0023] Preferably, the expression of the preset encoding formula is:

[0024]

[0025] Wherein, pixel_LBP(i) is the i-th bit in the 8-bit binary code; pixel(i) is the i-th neighborhood pixel around the center pixel.

[0026] Preferably, the expression of the DPAG module includes:

[0027] W=Conv3(X)+RepConv5(X),

[0028] A=Softmax(C(reshape(W)×⊙,GAP(X)))、

[0029] Λ=X×Conv(A)+X×RepConv(A),

[0030] Ω=Avgpool(Λ)+Maxpool(Λ)、

[0031] Δ = Sigmoid(Conv(Ω));

[0032] Among them, X is the input feature map; Conv3(·) represents a 3×3 convolution operation; RepConv5(·) represents a 5×5 RepVGG convolution operation; W is the intermediate feature map; C(·) represents a channel splicing operation; reshape(·) represents a tensor reshape operation; ⊙ represents element-wise product; GAP(·) represents global average pooling; Softmax(·) is the Softmax activation function; A is the attention weight; Conv(·) represents a standard convolution operation; RepConv(·) represents a RepVGG-style convolution operation; Λ is a weighted feature; Avgpool(·) represents average pooling; Maxpool(·) represents maximum pooling; Ω is the pooling result; Sigmoid(·) is the Sigmoid activation function; Δ is the final output feature.

[0033] The present invention discloses the following technical effects:

[0034] The present invention provides a method for detecting watermelon stripe phenotypes. By improving the YOLOv8s neural network, the present invention solves the problem that existing stripe analysis systems usually use grayscale threshold segmentation or edge detection to identify peel stripes, resulting in segmentation errors when facing complex peel structures or stripes with unclear contrast, and achieves self-adaptation to the complexity and diversity of stripes. By performing LBP encoding on image detection results, the present invention solves the problem that existing texture analysis methods are relatively single and difficult to extract multi-dimensional texture information of stripes, and achieves the extraction of multi-level and multi-dimensional texture features of peel stripes. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0036] Figure 1 A schematic diagram of the watermelon stripe phenotype detection process according to an embodiment of the present invention;

[0037] Figure 2 This is a flow chart for detecting watermelon stripe phenotypes according to an embodiment of the present invention. DETAILED DESCRIPTION

[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0039] The present invention aims to provide a method for detecting watermelon stripe phenotypes to improve the accuracy and consistency of stripe identification.

[0040] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0041] Figure 1 Schematic diagram of the watermelon stripe phenotype detection process provided by the embodiment of the present invention, as shown in FIG. Figure 1 As shown, the present invention provides a method for detecting watermelon stripe phenotype, comprising:

[0042] Step 100: Use the DSConv module to replace the convolution modules in the backbone network and the neck network of the YOLOv8s neural network, set the DPAG module at the connection between the backbone network and the neck network, and use the FEM module to replace the C2f module in the backbone network of the YOLOv8s neural network to obtain an improved YOLOv8s model;

[0043] Step 200: using the left imager, the right imager, and the infrared projector of the RealSense D455 depth camera to collect a watermelon stripe image of the target watermelon for detection;

[0044] Step 300: performing illumination correction and illumination compensation on the watermelon stripe image to obtain an illumination optimized image;

[0045] Step 400: Input the illumination optimized image into the pre-trained improved YOLOv8s model to perform stripe detection to obtain an image detection result;

[0046] Step 500: According to the image detection result, the watermelon stripe image is subjected to stripe segmentation, grayscale conversion, LBP coding calculation, statistics and normalization processing to obtain a stripe segmentation image, a stripe grayscale map and an LBP histogram.

[0047] Specifically, the shooting resolution and bit depth of the RealSense D455 depth camera are 848×480 and 24 respectively.

[0048] Optionally, the illumination compensation includes: logarithmic transformation and exponential transformation.

[0049] Preferably, the expression for illumination correction is:

[0050]

[0051] in, is the i-th target component x i The mapping output of β i is the i-th correction coefficient; n is the number of components; β0 is the basic component.

[0052] Furthermore, according to the image detection result, stripe segmentation, grayscale conversion, LBP coding calculation, statistics and normalization processing are performed on the watermelon stripe image to obtain a stripe segmentation image, a stripe grayscale image and an LBP histogram, including:

[0053] Convert the image detection result into a plurality of 8-neighborhood masks;

[0054] Convert the 8-neighborhood mask into several 8-bit binary codes using a preset coding formula;

[0055] Convert the 8-bit binary code into decimal data to obtain a pixel code;

[0056] Normalization is performed on all the pixel codes of the image detection result to obtain the LBP histogram.

[0057] Specifically, the DPAG module integrates the branch channel attention submodule and the channel attention submodule.

[0058] Preferably, the expression of the preset encoding formula is:

[0059]

[0060] Wherein, pixel_LBP(i) is the i-th bit in the 8-bit binary code; pixel(i) is the i-th neighborhood pixel around the center pixel.

[0061] Specifically, the expression of the DPAG module includes:

[0062] W=Conv3(X)+RepConv5(X),

[0063] A=Softmax(C(reshape(W)×⊙,GAP(X)))、

[0064] Λ=X×Conv(A)+X×RepConv(A),

[0065] Ω=Avgpool(Λ)+Maxpool(Λ)、

[0066] Δ = Sigmoid(Conv(Ω));

[0067] Among them, X is the input feature map; Conv3(·) represents a 3×3 convolution operation; RepConv5(·) represents a 5×5 RepVGG convolution operation; W is the intermediate feature map; C(·) represents a channel splicing operation; reshape(·) represents a tensor reshape operation; ⊙ represents element-wise product; GAP(·) represents global average pooling; Softmax(·) is the Softmax activation function; A is the attention weight; Conv(·) represents a standard convolution operation; RepConv(·) represents a RepVGG-style convolution operation; Λ is a weighted feature; Avgpool(·) represents average pooling; Maxpool(·) represents maximum pooling; Ω is the pooling result; Sigmoid(·) is the Sigmoid activation function; Δ is the final output feature.

[0068] Specifically, the RealSense D455 depth camera used includes a left imager, a right imager, and an infrared projector. The infrared projector projects an invisible static infrared pattern to improve the depth accuracy of low-texture scenes. The left and right imagers capture the scene and send the imager data to the depth imaging (vision) processor, which calculates the depth distance of each pixel in the image, associates the points on the left image with the right image, and calculates the disparity by the movement distance between the points on the left image and the right image, thereby inferring the depth information of each pixel. A depth frame is generated by processing the depth pixel values, and subsequent depth frames will create a depth video stream; and the shooting parameters are calibrated; a D455 camera is used in the watermelon greenhouse to shoot watermelon stripe pictures with a resolution of 848×480 and a bit depth of 24, and the pictures are sorted and saved as a watermelon stripe dataset.

[0069] Furthermore, the illumination correction and illumination compensation technology of the stripe image is described. Generally, color correction involves operations such as color balance, color temperature adjustment, and curve adjustment. The color correction method selected in this embodiment is a color correction method based on polynomial fitting; polynomials can be used to approximate the mapping relationship from the source space to the target space. A certain output component can be represented by the following polynomial:

[0070]

[0071] In the RGB device color space, x iThis includes the original value of the red channel, the original value of the green channel, the original value of the blue channel, the square of the red channel value, the square of the green channel value, the square of the blue channel value, the product of the red and green channel values, the product of the red and blue channel values, the product of the green and blue channel values, and so on. The source space is sampled and the corresponding value of the sampled point in the target space is obtained. Color correction is achieved by fitting the coefficients of the polynomial using mathematical methods, usually the least squares method.

[0072] Furthermore, nonlinear transformation is a grayscale transformation method, which can be described as follows: suppose the grayscale of the monochrome image f(x, y) is r, and it is transformed into another image g(x, y) with a grayscale of s through a nonlinear transformation function s=T(r). The histograms of the two are p(r) and p(s), respectively. The new histogram p(s) should be determined based on the human visual perception model, and computer vision requires that the visual sensitivity response curve of its visual sensor matches the histogram p(x, y) of g(x, y) in order to achieve the best visual effect. Some data show that subjective brightness is a logarithmic function of the incident light intensity (illuminance) of the eye. The logarithmic transformation form is as follows:

[0073] g(x,y)=a+ln(f(x,y)+1) / blnc

[0074] The parameters a, b, and c are introduced to facilitate the adjustment of the position and shape of the curve. In the formula, f(x, y)+1, that is, each pixel is added by one, in order to ensure that the logarithm of numbers greater than zero is taken. When selecting parameters, experiments have shown that when a is 0, b is 1 / (255*ln1.2), and c is 255, a better illumination compensation effect can be obtained. The image processed by logarithmic transformation is softer, the image level is clearer, which conforms to the visual characteristics of human beings, and has a better illumination compensation effect for images that are too dark or polarized. The image processed by logarithmic transformation is softer, the image level is clearer, which conforms to the visual characteristics of human beings, and has a better illumination compensation effect for images that are too dark or polarized. However, for images that are too bright, the edges are not clear after logarithmic transformation, and the edge information cannot be extracted. Exponential transformation can solve this problem.

[0075] Specifically, the form of the exponential transformation is as follows:

[0076] X_train_log=np.log(X_train+1)X_test_log=np.log(X_test+1);

[0077] Where X_train is the feature matrix of the training dataset; X_test is the feature matrix of the test dataset; np.log(·) is the natural logarithm function in the NumPy library. Adding 1 is to prevent the presence of 0 in the original data, so 1 is added first and then the logarithm is taken.

[0078] After extensive experiments and comparisons, we found that using a = 0 and c = 1 / 255, converting the original grayscale image f(x, y) to the interval [0, 1], and b = 255, yields an image g(x, y) with a grayscale range of [0, 255]. This has the opposite effect of the logarithmic transform, significantly expanding high-grayscale areas and compressing low-grayscale areas. The exponential transform is particularly effective for compensating for overly bright faces.

[0079] Furthermore, the watermelon stripes dataset was amplified by data augmentation, and then the model was trained through a neural network, and the attention mechanism and connection layer were added, and the local binary pattern (LBP) algorithm was added.

[0080] Specifically, a deep learning network structure and optimization method for watermelon stripe segmentation: In this embodiment, a lightweight YOLOv8s was selected. YOLOv8s is a lightweight parameter structure derived from the YOLOv8 algorithm. It includes a backbone network, a neck network, and a prediction output head. The backbone network uses convolution operations to extract features of various scales from RGB (red, green, and blue) color images. At the same time, the role of the neck network is to merge the features extracted by the backbone network. Feature pyramid structures (feature pyramid networks, FPN) are generally used to aggregate low-level features into higher-level representations. The head layer is responsible for predicting the target category and uses three groups of detection detectors of different sizes to select and detect image content.

[0081] Furthermore, this embodiment proposes an improved object detection model for quickly and accurately segmenting watermelon stripes in a natural environment. Depthwise separable convolution (DSConv) is used to replace the common convolutions in the backbone and neck parts of the original network, reducing the model size and improving the inference speed. The specific changes are shown in the figure, which is significantly different from YOLOv8s. In addition, the network also introduces a dual-path attention gate (DPAG) to overcome the weaknesses of lightweight neural networks in feature extraction. In addition, the model also integrates a feature enhancement module (FEM) to facilitate the network to extract more refined target features.

[0082] Preferably, the attention module has had a significant impact on the field of deep learning, enabling advanced techniques for channel attention such as SE-Net, channel and spatial dual attention such as CBAM, and non-local techniques that emphasize global information in feature maps. To improve edge detection performance, this embodiment adds an attention module to the connection layer. This preserves a large amount of detail when fusing low-level features into higher-level features. Adding a dual-path attention mechanism to the neck connection layer (Concat) can balance detection speed and feature extraction capabilities.

[0083] Specifically, the DPAG module fuses the attention gate AG and the CBAM attention module; it innovatively introduces an additional path on the channel layer to facilitate information extraction. DPAG integrates two continuous attention mechanisms, namely the branched channel attention module (BCAM) and the channel attention module (CAM); the former enables dual-path channel attention, while the latter learns image position information. BCAM and CAM interact closely to extract channel and spatial features, where BCAM enhances channel correlation and feature accuracy through channel relationship gates and position relationship gates, and CAM locates entities by grasping spatial information. Through DPAG's feature absorption and optimization process, pixels receive individual weights, which determine their necessity based on the weight value. This in turn improves feature utilization and the efficiency of recognition functions.

[0084] Preferably, the element summation operation is represented as “+” (element summation), the element multiplication operation is represented as “×” (element generation), and the channel summation is represented as “⊕” (concatenation), which is represented as C. The operation of the channel attention BCAM is as follows: the transmitted feature map is processed by standard convolution Conv3 and Conv5 respectively, and then the two are merged to obtain a shallow convolution layer, which is represented as After global average pooling (GAP), the reshaped row and column information are multiplied and added to obtain a feature map, which is then passed through the Softmax layer to obtain a set of learned weights, denoted as δ. The learned weights are multiplied and added with the standard convolution Conv and RepConv to obtain denoted as Λ. The spatial attention CAM is based on the output of the channel attention and uses average pooling (Avgpool) and maximum pooling (Maxpool). The intermediate amount is obtained by connecting Avgpool and Maxpool. After 1×1 convolution and S-shaped layer, the final spatial attention is Δ. The expression of the above process is as follows:

[0085] W=Conv3(X)+RepConv5(X)

[0086] A=Softmax(C(reshape(W)×⊙,GAP(X)))

[0087] Λ=X×Conv(A)+X×RepConv(A)

[0088] Ω=Avgpool(Λ)+Maxpool(Λ)

[0089] Δ=Sigmoid(Conv(Ω))

[0090] Specifically, the LBP encoding and histogram generation method of the stripe area: LBP has grayscale monotonic transformation invariance and rotation invariance, and binarizes the pixel values ​​of the neighborhood according to the central pixel value of the mask. The calculation form of LBP encoding is:

[0091]

[0092] in, P represents the number of pixels in a circular neighborhood with a radius of R, where R is the radius of the neighborhood. Measured in pixels, larger P values ​​and smaller R values ​​provide strong texture descriptiveness, but are computationally complex and highly susceptible to noise. Conversely, smaller values ​​are less susceptible to noise, with details such as edge information being weakened. (xc, yc) represent the coordinates of the center pixel C in the neighborhood, and the pixel values ​​are gc and gn, representing the values ​​of P equally spaced neighboring pixels with a radius of R from the center pixel. d represents the dth pixel in the neighborhood of the center gc, and a weight of 2d is assigned based on the value of S(gn-gc). The basic LBP model uses a mask with a 3×3 neighborhood size, i.e., an 8-neighborhood mask, where (8, 1) represents the standard LBP definition for an 8-neighborhood.

[0093] Furthermore, LBP can extract the texture features of the image and save them as a series of numbers. Usually, a pixel is taken as the center and its eight neighbors are observed to see whether they are larger or smaller than the center pixel. If center is represented as the center pixel and pixel(i) is represented as the eight neighbors of the center pixel, the calculation is:

[0094]

[0095] Finally, these eight numbers are sequentially concatenated to form the LBP eigenvalue for that pixel. Ultimately, the LBP eigenvalues ​​for each pixel are arranged according to the original pixel position, forming an image, commonly known as an LBP atlas. Each pixel value in the atlas is the corresponding LBP eigenvalue, an 8-bit binary number, typically converted to a decimal number ranging from 0 to 255 as the pixel value in the atlas. However, in practice, for example, the number of distinct values ​​in the LBP eigenvalues ​​is used in a classifier. The simplest process involves extracting LBP eigenvalues ​​from an image to obtain a series of numbers ranging from 0 to 255. The number of each 0-255 number is then calculated, and after normalization, the result can be fed into the classifier. This result is more intuitively represented using a histogram than an atlas.

[0096] The beneficial effects of the present invention are as follows:

[0097] (1) High stripe segmentation accuracy: By using a deep learning algorithm to accurately segment watermelon peel stripes, the present invention can adapt to different lighting conditions and stripe complexity, significantly improving the accuracy and consistency of stripe segmentation, and overcoming the limitations of traditional edge detection and grayscale threshold methods in distinguishing stripe morphology.

[0098] (2) Multidimensional texture feature extraction: After segmentation, the present invention combines the local binary pattern (LBP) algorithm to perform a detailed analysis of the local texture and grayscale distribution characteristics of the stripes, generating histograms and grayscale images, providing rich texture information. This multidimensional feature extraction greatly enhances the ability to describe stripe morphology and texture, facilitating more accurate watermelon variety identification and quality evaluation.

[0099] (3) Standardized Data Storage and Management: This paper systematically stores information such as segmented fringe images, LBP histograms, and grayscale images, creating a structured fringe database that facilitates subsequent analysis, retrieval, and model upgrades. This data management approach improves analysis efficiency and reproducibility, providing reliable support for standardized management of research data.

[0100] (4) Improved adaptability and robustness: The present invention adopts an adaptive deep learning model that can be adjusted according to different stripe structures, color characteristics, and lighting changes, ensuring good adaptability under diverse conditions and reducing the impact of environmental changes on detection results.

[0101] (5) Improve the level of automation: By combining deep learning and machine vision automation solutions, the present invention greatly reduces manual participation and can efficiently and automatically complete the detection and analysis process of stripe features. It is suitable for the construction of a large-scale watermelon phenotypic database and provides technical support for subsequent smart agricultural applications.

[0102] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0103] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. A method for detecting watermelon stripe phenotype, characterized in that: include: The convolution modules in the backbone network and the neck network of the YOLOv8s neural network are replaced by the DSConv module, a DPAG module is set at the connection between the backbone network and the neck network, and the C2f module in the backbone network of the YOLOv8s neural network is replaced by the FEM module to obtain an improved YOLOv8s model. The left imager, right imager, and infrared projector of the RealSense D455 depth camera are used to collect watermelon stripe images for target detection. Performing illumination correction and illumination compensation on the watermelon stripe image to obtain an illumination optimized image; Inputting the illumination optimized image into the pre-trained improved YOLOv8s model for stripe detection to obtain an image detection result; According to the image detection result, the watermelon stripe image is subjected to stripe segmentation, grayscale conversion, LBP coding calculation, statistics and normalization processing to obtain a stripe segmentation image, a stripe grayscale map and an LBP histogram.

2. A watermelon stripe phenotype detection method according to claim 1, characterized in that: The RealSense D455 depth camera has a capture resolution of 848×480 and a bit depth of 24.

3. The method for detecting watermelon stripe phenotype according to claim 1, wherein: The illumination compensation includes logarithmic transformation and exponential transformation.

4. A watermelon stripe phenotype detection method according to claim 1, characterized in that: The expression of the illumination correction is: in, is the i-th target component x i The mapping output of β i is the i-th correction coefficient; n is the number of components; β0 is the basic component.

5. The method for detecting watermelon stripe phenotype according to claim 1, wherein: According to the image detection result, the watermelon stripe image is subjected to stripe segmentation, grayscale conversion, LBP coding calculation, statistics and normalization processing to obtain a stripe segmentation image, a stripe grayscale image and an LBP histogram, including: Convert the image detection result into a plurality of 8-neighborhood masks; Convert the 8-neighborhood mask into several 8-bit binary codes using a preset coding formula; Convert the 8-bit binary code into decimal data to obtain a pixel code; Normalization is performed on all the pixel codes of the image detection result to obtain the LBP histogram.

6. The method for detecting watermelon stripe phenotype according to claim 1, wherein: The DPAG module integrates the branch channel attention submodule and the channel attention submodule.

7. The method for detecting watermelon stripe phenotype according to claim 5, wherein: The expression of the preset encoding formula is: Wherein, pixel_LBP(i) is the i-th bit in the 8-bit binary code; pixel(i) is the i-th neighborhood pixel around the center pixel.

8. The method for detecting watermelon stripe phenotype according to claim 6, wherein: The expression of the DPAG module includes: W=Conv3(X)+RepConv5(X), A=Softmax(C(reshape(W)×⊙,GAP(X)))、 Λ=X×Conv(A)+X×RepConv(A), Ω=Avgpool(Λ)+Maxpool(Λ)、 Δ = Sigmoid(Conv(Ω)); Among them, X is the input feature map; Conv3(·) represents a 3×3 convolution operation; RepConv5(·) represents a 5×5 RepVGG convolution operation; W is the intermediate feature map; C(·) represents a channel splicing operation; reshape(·) represents a tensor reshape operation; ⊙ represents element-wise product; GAP(·) represents global average pooling; Softmax(·) is the Softmax activation function; A is the attention weight; Conv(·) represents a standard convolution operation; RepConv(·) represents a RepVGG-style convolution operation; Λ is a weighted feature; Avgpool(·) represents average pooling; Maxpool(·) represents maximum pooling; Ω is the pooling result; Sigmoid(·) is the Sigmoid activation function; Δ is the final output feature.

Citation Information

Patent Citations

  • Face detecting and tracking method and device

    CN103116756A

  • Illumination correction method for image

    CN110264411A

  • Apple identification and positioning method suitable for orchard complex open environment

    CN116385536A

  • Sequence image dodging method and device and readable medium

    CN117197328A

  • Watermelon and watermelon stem identification method based on improved YOLOv8n-obb

    CN119131780A