Watermelon fruit contour cross section parameter determination method based on deep learning
Through a deep learning-based method, the watermelon cross-sectional image is collected using the left imager, the right imager and the infrared projector, converted into a three-dimensional point cloud image and performed light correction and compensation, and U-Net is used to improve model segmentation and Delaunay triangulation to calculate the fruit cross-sectional area, solving the problems of low efficiency, labor dependence and high cost in traditional methods, and achieving automated and high-precision measurement of watermelon fruit parameters.
Patent Information
- Application Number
- CN202510707807.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Traditional watermelon fruit parameter measurement methods are inefficient, labor-dependent, poor versatility and high cost, making it difficult to meet the high-throughput measurement needs of modern agriculture.
Using a deep learning-based method, the watermelon cross-section image is collected using the left imager, the right imager and the infrared projector, and converted into a three-dimensional point cloud image. After illumination correction and compensation, the pre-trained U-Net improved model is used for segmentation and smoothing, and the fruit cross-sectional area is calculated in combination with Delaunay triangulation.
Automatic measurement of watermelon fruit parameters is realized, the measurement accuracy and efficiency are improved, different fruit forms are adapted to, and manual intervention and equipment costs are reduced.
Smart Images

Figure CN120580281A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fruit parameter measurement, and in particular to a method for measuring watermelon fruit contour cross-sectional parameters based on deep learning. Background Art
[0002] The measurement of watermelon fruit parameters has traditionally relied mainly on manual operations or simple image processing techniques. These methods have developed to a certain application basis: manual measurement method: manually measuring the long diameter, short diameter and surface area of the fruit using tools such as tape measures and calipers; two-dimensional image processing technology: capturing fruit images with an ordinary camera, and calculating fruit parameters through image edge detection and geometric shape fitting; three-dimensional modeling technology: in recent years, three-dimensional laser scanning and structured light technology have begun to be used for fruit phenotyping, which can accurately restore the three-dimensional structure of the fruit and improve measurement accuracy.
[0003] However, manual measurement and two-dimensional image processing technologies are inefficient in large-scale fruit sample processing and cannot meet the needs of modern agriculture for high-throughput measurement. Manual measurement is easily affected by operational errors. Two-dimensional image processing technology cannot accurately restore the three-dimensional characteristics of fruit, especially when the fruit contour is irregular or there is occlusion, the error is more significant, and the measurement results are easily affected by light, background and fruit morphology, and lack stability. Existing technologies generally rely on human intervention, including manual adjustment of measurement position or parameters, and lack automated data processing and storage functions. Many systems can only adapt to specific types of fruit and have poor versatility. High-precision equipment such as three-dimensional laser scanning and structured light are expensive and complex to operate, which limits their application in actual production. Summary of the Invention
[0004] In order to overcome the shortcomings of the existing technology, the purpose of the present invention is to provide a method for measuring the cross-sectional parameters of watermelon fruit contour based on deep learning, so as to solve the problems of low efficiency and accuracy, reliance on manual labor, poor versatility and high cost of traditional methods.
[0005] To achieve the above object, the present invention provides the following solutions:
[0006] A method for measuring watermelon fruit contour cross-sectional parameters based on deep learning, comprising:
[0007] The left imager, the right imager, and the infrared projector are used to collect the cross-sectional image of the watermelon for target detection;
[0008] Mapping the watermelon cross-section image into a three-dimensional point cloud image;
[0009] Performing illumination correction and illumination compensation on the three-dimensional point cloud image to obtain a point cloud enhanced image;
[0010] The watermelon cross-section is segmented using the pre-trained U-Net improved model embedded with the DCA module to obtain a watermelon plane point cloud model.
[0011] Performing a downsampling operation on the watermelon plane point cloud model under a preset voxel grid side length standard to obtain a sparse point cloud model;
[0012] Eliminate outliers from the sparse point cloud model under a preset adjacent point value and outlier threshold to obtain a noise-removed point cloud model;
[0013] Smoothing the noise-removed cloud model using a moving least squares method to obtain a smoothed point cloud model;
[0014] Performing Delaunay triangulation on the smooth point cloud model under preset maximum side length, maximum angle, and minimum angle of a triangle to obtain a triangular face;
[0015] The areas of the triangular facets are calculated and accumulated using Heron's formula to obtain the cross-sectional area of the watermelon fruit.
[0016] Preferably, the watermelon cross-section image has a resolution of 848×480 and a bit depth of 24.
[0017] Preferably, the illumination compensation includes: logarithmic transformation and exponential transformation.
[0018] Preferably, the side length standard of the voxel grid is 1.00 mm.
[0019] Preferably, the adjacent point value is 50; the outlier threshold is 1.0.
[0020] Preferably, the maximum side length of the triangle is 2.00 mm; the maximum angle is 120°; and the minimum angle is 10°.
[0021] Preferably, the mapping relationship of the watermelon cross-section image includes:
[0022] xw=d / 1000、
[0023] yw=d(u-cx) / 1000fx and
[0024] zw=d(v-cy) / 1000fy;
[0025] Among them, xw, yw, zw are the horizontal coordinate, vertical coordinate and vertical coordinate of the three-dimensional point cloud image respectively; d is the depth value; u and v are the U-axis coordinate and V-axis coordinate of the pixel in the watermelon cross-section image respectively; cx and cy are the horizontal coordinate and vertical coordinate of the principal point of the depth camera respectively; fx and fy are the focal lengths of the depth camera in the X-axis and Y-axis directions respectively.
[0026] Preferably, the expression for illumination correction is:
[0027]
[0028] in, is the i-th target component x i The mapping output of β i is the i-th correction coefficient; n is the number of components; β0 is the basic component.
[0029] The present invention discloses the following technical effects:
[0030] The present invention provides a method for measuring the cross-sectional parameters of watermelon fruit contours based on deep learning. Automatic segmentation is performed through an improved U-Net model, which solves the defects of traditional manual measurement relying on manual operation and two-dimensional image processing requiring a large amount of manual proofreading and adjustment, thereby realizing the automated measurement of fruit parameters. By converting the watermelon cross-sectional image into a three-dimensional point cloud image, the problem of inaccurate measurement caused by viewing angle error is solved, thereby achieving improved measurement accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0032] Figure 1 A schematic diagram of the watermelon fruit contour segmentation process based on deep learning provided by an embodiment of the present invention;
[0033] Figure 2 A flow chart of watermelon fruit contour segmentation based on deep learning provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0035] The purpose of the present invention is to provide a method for measuring the cross-sectional parameters of watermelon fruit contour based on deep learning, so as to solve the problems of low efficiency and accuracy, reliance on manual labor, poor versatility and high cost of traditional methods.
[0036] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0037] Figure 1 A schematic diagram of a watermelon fruit contour segmentation process based on deep learning provided by an embodiment of the present invention. Figure 2 The flow chart of watermelon fruit contour segmentation based on deep learning provided by the embodiment of the present invention is as follows: Figure 1 and Figure 2 As shown, the present invention provides a method for measuring watermelon fruit contour cross-sectional parameters based on deep learning, comprising:
[0038] Step 100: using a left imager, a right imager, and an infrared projector to capture a cross-sectional image of the watermelon of the target detection watermelon;
[0039] Step 200: mapping the watermelon cross-section image into a three-dimensional point cloud image;
[0040] Step 300: performing illumination correction and illumination compensation on the three-dimensional point cloud image to obtain a point cloud enhanced image;
[0041] Step 400: Segment the watermelon cross section using a pre-trained U-Net improved model embedded with a DCA module to obtain a watermelon plane point cloud model;
[0042] Step 500: performing a downsampling operation on the watermelon plane point cloud model under a preset voxel grid side length standard to obtain a sparse point cloud model;
[0043] Step 600: removing outliers from the sparse point cloud model under a preset adjacent point value and outlier threshold to obtain a noise-removed point cloud model;
[0044] Step 700: Smoothing the noise-removed cloud model using a moving least squares method to obtain a smoothed point cloud model;
[0045] Step 800: performing Delaunay triangulation on the smooth point cloud model under preset maximum triangle side length, maximum angle, and minimum angle to obtain triangular facets;
[0046] Step 900: Calculate and accumulate the areas of the triangular facets using Heron's formula to obtain the cross-sectional area of the watermelon fruit.
[0047] Preferably, the watermelon cross-section image has a resolution of 848×480 and a bit depth of 24.
[0048] Specifically, the illumination compensation includes: logarithmic transformation and exponential transformation.
[0049] Preferably, the side length standard of the voxel grid is 1.00 mm.
[0050] Optionally, the adjacent point value is 50; the outlier threshold is 1.0.
[0051] Preferably, the maximum side length of the triangle is 2.00 mm; the maximum angle is 120°; and the minimum angle is 10°.
[0052] Specifically, the mapping relationship of the watermelon cross-section image includes:
[0053] xw=d / 1000、
[0054] yw=d(u-cx) / 1000fx and
[0055] zw=d(v-cy) / 1000fy;
[0056] Among them, xw, yw, zw are the horizontal coordinate, vertical coordinate and vertical coordinate of the three-dimensional point cloud image respectively; d is the depth value; u and v are the U-axis coordinate and V-axis coordinate of the pixel in the watermelon cross-section image respectively; cx and cy are the horizontal coordinate and vertical coordinate of the principal point of the depth camera respectively; fx and fy are the focal lengths of the depth camera in the X-axis and Y-axis directions respectively.
[0057] Furthermore, the expression of the illumination correction is:
[0058]
[0059] in, is the i-th target component x i The mapping output of β i is the i-th correction coefficient; n is the number of components; β0 is the basic component.
[0060] Specifically, the RealSense D455 depth camera includes a left imager, a right imager, and an infrared projector. The infrared projector projects an invisible static infrared pattern to improve the depth accuracy of low-texture scenes. The left and right imagers capture the scene and send the imager data to the depth imaging (vision) processor, which associates the points on the left image with the right image by calculating the depth distance of each pixel in the image, and calculates the disparity by the movement distance between the points on the left image and the right image, thereby inferring the depth information of each pixel. A depth frame is generated by processing the depth pixel value, and subsequent depth frames will create a depth video stream; to generate a three-dimensional point cloud from the depth map captured by the camera, it is necessary to construct a mapping relationship between the pixel point and the position in space. This relationship requires the internal and external parameters of the depth camera to determine. The process of solving these two parameters is camera calibration. According to the RealSense camera imaging principle, the infrared image and the depth map have the same viewing angle and resolution. In this embodiment, infrared images shot at multiple angles and distances are used instead of depth maps as the calibration object, and calibration is performed using the calibration box provided by MATLAB software. In a watermelon greenhouse, a D455 camera was used to capture cross-sectional images of cut watermelons with a resolution of 848×480 and a bit depth of 24, and the images were organized and saved as a watermelon cross-sectional dataset.
[0061] Further preparations:
[0062] 1) Install Anaconda (Anaconda3-2021.04), CUDA (cuda_11.7.1_516.94), and Pyc harm;
[0063] 2) Create a virtual environment: condacreate -nvisual_intelpython=3.10;
[0064] 3) Activate the virtual environment: condaactivatevisual_intel;
[0065] 4) Install dependent packages: torch (1.13.1+cu117-cp310), torchaudio (0.13.1+cu117-cp310), torchvision (0.14.1+cu117-cp310), opencv-python (4.10.0.84), pyrealsense2 (2.55.16486), numpy (2.1.0);
[0066] Specifically, generate a three-dimensional point cloud: the depth map is converted into a three-dimensional point cloud by mapping the pixel coordinate system to the world coordinate system. The conversion formula of the coordinates (xw, yw, zw) is as follows:
[0067] xw=d / 1000
[0068] yw=d(u-cx) / 1000fx
[0069] zw=d(v-cy) / 1000fy
[0070] Given the camera's intrinsic and extrinsic parameters, combined with the above equation, we transform all non-zero pixel values in the depth map to generate a 3D point cloud. This point cloud is saved in a .pcd file in the PointXYZ format. If color coding is added to the corresponding point locations, the point cloud data contains both 3D coordinate information and RGB color information, forming a colored point cloud in the PointXYZRGB format.
[0071] Preferably, the illumination correction and illumination compensation technology of the fringe image:
[0072] Generally, color correction involves operations such as color balance, color temperature adjustment, and curve adjustment. The color correction method used in this embodiment is a color correction method based on polynomial fitting; polynomials can be used to approximate the mapping relationship from the source space to the target space. A certain output component can be represented by the following polynomial:
[0073]
[0074] In the RGB device color space, x i This includes the original value of the red channel, the original value of the green channel, the original value of the blue channel, the square of the red channel value, the square of the green channel value, the square of the blue channel value, the product of the red and green channel values, the product of the red and blue channel values, the product of the green and blue channel values, and so on. The source space is sampled and the corresponding value of the sampled point in the target space is obtained. Color correction is achieved by fitting the coefficients of the polynomial using mathematical methods, usually the least squares method.
[0075] Furthermore, nonlinear transformation is a grayscale transformation method, which can be described as follows: suppose the grayscale of the monochrome image f(x,y) is r, and it is transformed into another image g(x,y) with a grayscale of s through a nonlinear transformation function s=T(r). The histograms of the two are p(r) and p(s), respectively. The new histogram p(s) should be determined based on the human visual perception model, and computer vision requires that the visual sensitivity response curve of its visual sensor matches the histogram p(x,y) of g(x,y) in order to achieve the best visual effect. Some data show that subjective brightness is a logarithmic function of the incident light intensity (illuminance) of the eye. The logarithmic transformation form is as follows:
[0076] g(x,y)=a+ln(f(x,y)+1) / blnc
[0077] The parameters a, b, and c are introduced to facilitate adjustment of the curve's position and shape. The formula f(x,y)+1, which adds one to each pixel, ensures that the logarithm of numbers greater than zero is taken. When selecting parameters, experiments have shown that using 0 for a, 1 / (255*ln1.2) for b, and 255 for c achieves optimal lighting compensation. Logarithmic transformation produces a softer image with clearer layers, which aligns with human visual characteristics and provides excellent lighting compensation for overly dark or polarized images. However, for overly bright images, the logarithmic transformation blurs edges and makes it impossible to extract edge information. Exponential transformation can address this issue.
[0078] Specifically, the form of the exponential transformation is as follows:
[0079] X_train_log=np.log(X_train+1)X_test_log=np.log(X_test+1)
[0080] Where X_train is the feature matrix of the training dataset; X_test is the feature matrix of the test dataset; and np.log(·) is the natural logarithm function in the NumPy library. The addition of 1 prevents the presence of zeros in the original data, so 1 is added before taking the logarithm. After extensive experimentation and comparison, we found that using a = 0, c = 1 / 255, and converting the original grayscale image f(x,y) to the interval [0,1], with b = 255, results in an image g(x,y) with a grayscale range [0,255]. This has the opposite effect of the logarithm, significantly expanding high-intensity areas and compressing low-intensity areas. The exponential transform is particularly effective for compensating for overly bright faces.
[0081] Furthermore, the watermelon cross-section image dataset is amplified by data augmentation (corresponding to the aforementioned Gaussian blur, rotation, brightness adjustment, etc.), and then the improved U-Net neural network training model is used to accurately obtain the long and short diameters. All point cloud data of the watermelon cross-section edge after U-Net segmentation are obtained, and the Euclidean distance between all two points is calculated to obtain the longest and shortest, that is, the long and short diameters of the watermelon. In this embodiment, the dual cross attention module (DCA) is selected to improve the U-Net structure.
[0082] Specifically, DCA in skip connections directly connects shallow features (rich in spatial information) in the encoder to deep features (rich in semantic information) in the decoder. The DCA module replaces the traditional simple concatenation operation, enhancing the interaction between shallow and deep features. DCA in bridge layers, located in the bottleneck between the encoder and decoder, captures global context and multi-scale information, enhancing the network's understanding of details and overall structure.
[0083] Furthermore, the encoder output is: X, which usually represents a low-level feature map with high spatial resolution, the number of channels is C1, and the size is H × W. The decoder input is: Y, which represents a high-level feature map with strong semantic information, the number of channels is C2, and the size is H × W.
[0084] Specifically, the DCA module processes the input features through the following steps:
[0085] 1) Feature transformation:
[0086] Perform a linear transformation on X and Y to generate a query vector (Query), a key vector (Key), and a value vector (Value). This is accomplished using learnable weight matrices, each of which independently learns a different feature subspace. For example, for X, the transformation weights WQ for query vector Q are generated, while the transformation weights WK and WV for key vector K and value vector V are generated, respectively.
[0087] 2) Cross-attention calculation:
[0088] The encoder feature X and the decoder feature Y are taken as input to each other for cross information extraction.
[0089] 3) Attention weight: Calculate the similarity between the query of X and the key of Y, and normalize the weight through the softmax function to control the contribution of the value vector.
[0090] 4) Weighted operation: Use attention weights to weight the value vector and extract the most relevant features.
[0091] 5) Feature fusion: The interaction result is added to the original feature through residual connection to retain the original feature information and enhance the new feature expression.
[0092] Furthermore, the feature mapping module transforms the input features X and Y into a high-dimensional feature space through linear transformation, enabling different feature subspaces to interact through the attention mechanism. This provides the basis for subsequent cross-attention calculations, mapping spatial and semantic information into a unified representation.
[0093] Specifically, the attention mechanism module inputs: encoder features as queries, and decoder features as keys and values (or vice versa). It assigns attention weights based on the similarity between the query and the key, capturing long-range dependencies between encoder and decoder features. It extracts highly relevant information from the decoder features and injects it into the encoder features. The fusion module inputs: features calculated by cross-attention and original features. Residual connections ensure information integrity. The fused features contain complementary information from the encoder and decoder features, providing richer semantic guidance for the decoder.
[0094] Furthermore, the cross-attention mechanism achieves feature interaction and information fusion through the following process: The query vector represents the feature requirements to be searched. The key and query together measure the relevance of two features. The value provides specific feature information, which is ultimately selectively fused based on the attention weights.
[0095] Specifically, while traditional skip connections directly concatenate features, the DCA module uses an attention mechanism to align features at different levels, making the interaction between shallow and deep features more precise and reducing the interference of irrelevant information. The DCA module uses residual connections to preserve the original information, while the shared weight matrix design reduces the number of parameters and improves computational efficiency.
[0096] Preferably, DCA addresses the semantic gap between encoder and decoder features by capturing the channel and spatial dependencies between multi-scale encoder features in sequence:
[0097] First, cross-channel attention (CCA) extracts global channel dependencies by cross-attention of cross-channel tokens leveraging multi-scale encoder features.
[0098] Then, the Spatial Cross Attention (SCA) module performs a cross-attention operation to capture the spatial dependencies across spatial tokens.
[0099] Finally, these fine-grained encoder features are upsampled and connected to the corresponding decoder parts, forming a skip-connection scheme.
[0100] Furthermore, DCA can be divided into two main phases and three steps:
[0101] The first stage consists of a multi-scale patchbedding module to obtain encoder tokens.
[0102] In the second stage, DCA is implemented using channel-wise cross attention (CCA) and spatial cross attention (SCA) modules on these encoder tokens to capture long-range dependencies.
[0103] Finally, these tokens are serialized and upsampled using layer normalization and GeLU, and connected to their decoder counterparts.
[0104] Specifically, Patch extraction in the multi-scale encoder stage:
[0105] First, patches are extracted from N multi-scale encoder stages.
[0106] Given s encoder stages of different scales, the output feature map is Its width and height are
[0107] Define the block size Where s = 1, 2,…, N.
[0108] Use size and step size of P s The average pooling is used to extract the patch, and a 1×1 one-dimensional separable convolution (DConv1D) is used to map the flattened two-dimensional patch. The formula is as follows:
[0109] T s =DConv1D(Reshape(AvgPool2D(E s )))
[0110] in, represents the patch sequence extracted by the s-th encoder stage, and P represents each T s The number of patches in is constant and is uniform across all scales, so the interactions between these tokens can be directly exploited.
[0111] Specifically, CCA is used to process the Ti of each token. The channel-wise attention module (CCA) is a key component of the dual-cross attention module (DCA) in the improved U-Net architecture. Its design aims to enhance the network's feature representation capabilities by capturing the inter-channel dependencies of input features. CCCA receives multiple input features and normalizes them through layer normalization to eliminate scale differences. Then, independent projection operations are performed to unify the number of channels in the features, generating projected features of shape P × Cc. The projected features are concatenated into a unified representation, which is then further combined to generate the query (Q), key (K), and value (V) to construct the inter-channel attention mechanism. The inter-channel correlation matrix is calculated by transposing Q and K (KT), and the weight matrix is normalized through softmax. This is used to weight the value and generate weighted features that incorporate inter-channel dependencies. Finally, through projection and scaling, the weighted features are restored to the same shape as the input, facilitating subsequent network processing. The CCA module improves the network's ability to capture global and local features by explicitly modeling the interaction between channels, thereby significantly enhancing the performance of U-Net in segmentation tasks. First, layer normalization (LN) is performed on each Ti, and then T is normalized along the channel dimension. s Splice and get T c , by T c Query and Value are further generated. Depthwise separable convolution is introduced into self-attention to capture local information and reduce computational complexity.
[0112] Furthermore, the area of the watermelon plane point cloud model generated by segmentation is calculated. In order to ensure the effect of Delaunay triangulation and the accuracy of watermelon phenotyping, the following steps are performed:
[0113] 1) Perform downsampling on the original point cloud to reduce the number of points while still preserving the shape of the watermelon fruit. The side length of the voxel grid is set to 1.00 mm.
[0114] 2) Perform statistical filtering on the leaf point cloud, set the parameters (the value of the adjacent points in the statistical filtering query is 50, and the outlier threshold is set to 1.0) to remove outliers;
[0115] 3) Use the moving least squares method to smooth the data points and fit the blade point cloud into a better three-dimensional surface effect;
[0116] 4) Delaunay triangulation, setting parameters (maximum triangle side length is 2.00 mm, maximum angle is 120°, minimum angle is 10°) to obtain a triangular facet of the plane.
[0117] Using Heron's formula, we can calculate the area of a triangle. That is, by traversing all the triangles in the cross section of the watermelon fruit and summing them up, we can get the area S. The formula is as follows:
[0118] sq=pq(pq-a)(pq-b)(pq-c)
[0119] pq=(a+b+c) / 2
[0120] S=∑sq
[0121] Where sq is the area of the qth triangle, and a, b, and c are the side lengths of the triangle.
[0122] The beneficial effects of the present invention are as follows:
[0123] (1) This invention uses a deep learning model to automatically segment the contours of watermelon fruits, replacing traditional methods that rely on manual annotation or simple thresholds to achieve automated measurement. Furthermore, the RealSense D455 camera uses point cloud technology to collect three-dimensional fruit data in real time, eliminating manual measurement and tedious data preprocessing. The entire measurement process is highly automated, significantly reducing the time required to obtain fruit parameters.
[0124] (2) This invention accurately models fruit cross-sections using 3D point cloud data, effectively avoiding measurement inaccuracies caused by viewing angle errors. Furthermore, the fruit segmentation technology based on deep learning can handle complex backgrounds and different fruit morphologies, ensuring that the extracted contours are authentic and reliable, further improving the accuracy of parameter calculations.
[0125] (3) The present invention rigorously calibrates the RealSense D455 depth camera to ensure the continuity and consistency of the measurement data. In addition, the present invention uses point cloud data compensation and optimization algorithms to reduce errors caused by fluctuations in the acquisition environment or equipment, ensuring that the measurement results of different samples and different batches remain highly consistent.
[0126] (4) The present invention automatically saves the measured long diameter, short diameter, and area results, reducing the workload of manual recording and data management. The associated storage of image and point cloud results also facilitates subsequent analysis and optimizes the operational process for researchers.
[0127] (5) Through the generalization capabilities of deep learning models and the flexible adaptation of point cloud technology, the present invention can adapt to the measurement needs of watermelon fruits of different sizes and shapes. At the same time, the system has high scalability and can be extended to measure parameters of other fruits or crops.
[0128] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0129] The present invention is described in detail using specific examples. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will appreciate that the specific implementation and scope of application may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.
Claims
1. A method for measuring watermelon fruit contour cross-sectional parameters based on deep learning, characterized in that: include: The left imager, the right imager, and the infrared projector are used to collect the cross-sectional image of the watermelon for target detection; Mapping the watermelon cross-section image into a three-dimensional point cloud image; Performing illumination correction and illumination compensation on the three-dimensional point cloud image to obtain a point cloud enhanced image; The watermelon cross-section is segmented using the pre-trained U-Net improved model embedded with the DCA module to obtain a watermelon plane point cloud model. Performing a downsampling operation on the watermelon plane point cloud model under a preset voxel grid side length standard to obtain a sparse point cloud model; Eliminate outliers from the sparse point cloud model under a preset adjacent point value and outlier threshold to obtain a noise-removed point cloud model; Smoothing the noise-removed cloud model using a moving least squares method to obtain a smoothed point cloud model; Performing Delaunay triangulation on the smooth point cloud model under preset maximum side length, maximum angle, and minimum angle of a triangle to obtain a triangular face; The areas of the triangular facets are calculated and accumulated using Heron's formula to obtain the cross-sectional area of the watermelon fruit.
2. A method for measuring watermelon fruit contour cross-section parameters based on deep learning according to claim 1, characterized in that: The watermelon cross-section image has a resolution of 848×480 and a bit depth of 24.
3. A method for measuring watermelon fruit contour cross-section parameters based on deep learning according to claim 1, characterized in that: The illumination compensation includes logarithmic transformation and exponential transformation.
4. A method for measuring watermelon fruit contour cross-section parameters based on deep learning according to claim 1, characterized in that: The standard side length of the voxel grid is 1.00 mm.
5. A method for measuring watermelon fruit contour cross-section parameters based on deep learning according to claim 1, characterized in that: The adjacent point value is 50; the outlier threshold is 1.
0.
6. The method for measuring watermelon fruit contour cross-section parameters based on deep learning according to claim 1, wherein: The maximum side length of the triangle is 2.00 mm; the maximum angle is 120°; and the minimum angle is 10°.
7. The method for measuring watermelon fruit contour cross-sectional parameters based on deep learning according to claim 1, wherein: The mapping relationship of the watermelon cross-section image includes: xw=d / 1000、 yw=d(u-cx) / 1000fx and zw=d(v-cy) / 1000fy; Among them, xw, yw, zw are the horizontal coordinate, vertical coordinate and vertical coordinate of the three-dimensional point cloud image respectively; d is the depth value; u and v are the U-axis coordinate and V-axis coordinate of the pixel in the watermelon cross-section image respectively; cx and cy are the horizontal coordinate and vertical coordinate of the principal point of the depth camera respectively; fx and fy are the focal lengths of the depth camera in the X-axis and Y-axis directions respectively.
8. The method for measuring watermelon fruit contour cross-section parameters based on deep learning according to claim 1, wherein: The expression of the illumination correction is: in, is the i-th target component x i The mapping output of β i is the i-th correction coefficient; n is the number of components; β0 is the basic component.
Citation Information
Patent Citations
Rice quality detection method and device based on three-dimensional image processing and electronic equipment
CN115496756A
Watermelon section phenotype analysis method, system and device based on vision
CN117392162A
Method for estimating weight of cut watermelon in image recognition based on ellipsoid missing model
CN118261962A
Three-dimensional lidar point cloud semantic segmentation method and apparatus based on deep learning
WO2024130776A1