A multi-modal grading method of seedless watermelon fusing appearance texture and single fruit weight

By using multi-view image reconstruction and distortion correction methods, combined with three-dimensional surface geometry models and weight data, the problems of perspective distortion and incomplete information in seedless watermelon grading were solved, achieving high-precision fusion of appearance texture and weight features, and improving the accuracy and intelligence of grading.

CN122336489APending Publication Date: 2026-07-03东台市综合检验检测中心(东台市农产品质量检测中心) +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
东台市综合检验检测中心(东台市农产品质量检测中心)
Filing Date
2026-04-07
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing technologies for grading seedless watermelons suffer from problems such as perspective distortion, incomplete information collection, inaccurate feature extraction, and insufficient integration of weight and appearance information, resulting in insufficient grading accuracy and intelligence, especially in high-value-added fruits where high precision requirements are difficult to meet.

Method used

A three-dimensional surface point cloud was reconstructed using multi-view image sequences, a parameterized three-dimensional surface geometric model was fitted, and the appearance texture information was mapped to a two-dimensional texture map through a distortion correction weight map. Combined with single fruit weight data, a cross-modal attention mechanism was used to perform feature fusion to extract the appearance texture and weight features of watermelon.

Benefits of technology

It achieves high reproducibility and consistency in the grading results of seedless watermelons, improves the accuracy of judging key commercial indicators, enhances the ability to identify internal defects, and enables accurate inference from external characteristics to internal quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122336489A_ABST
    Figure CN122336489A_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal grading method for seedless watermelons that integrates appearance texture and single-fruit weight. The method includes: simultaneously capturing images of the watermelon placed on a weighing device using a circular camera array and a top-view camera to obtain multi-view image sequences and single-fruit weight data; reconstructing a three-dimensional surface point cloud of the watermelon based on the image sequence and fitting it as a hyperellipsoidal model; mapping the three-dimensional surface texture to a two-dimensional texture map and simultaneously generating a distortion-corrected weight map; using this weight map to weight the two-dimensional texture map and extracting appearance features with real physical meaning, such as defect area and texture statistics; inputting these distortion-corrected appearance features and single-fruit weight data into a deep fusion network based on a cross-modal attention mechanism; and determining the final quality grade of the watermelon by mining deep correlations between modalities. This invention improves the objectivity of grading standards and the accuracy of predicting internal quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and agricultural automation, and in particular to a multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight. Background Technology

[0002] As modern agriculture develops towards intensification and intelligence, the use of machine vision technology for automated and non-destructive quality grading of agricultural products has become a research hotspot and industry trend. In the grading of fruits and vegetables such as watermelons, existing technologies typically employ two-dimensional image processing methods, using single or multiple cameras to capture images of the fruit from different angles, and then analyzing its color, shape, and surface texture. Furthermore, some solutions incorporate a weighing module on a conveyor belt, integrating weight information as an independent indicator into the grading system. While these technologies improve grading efficiency and objectivity to some extent, their inherent limitations become increasingly apparent when dealing with near-spherical curved surfaces and complex surface textures like watermelons, making it difficult to meet the high-precision, high-standard market demands. The shortcomings of existing technologies are even more pronounced in grading high-value-added fruits like seedless watermelons, as consumers have extremely high requirements for their appearance quality, such as the roundness of the fruit, the clarity of the skin stripes, and color contrast, and generally have zero tolerance for internal physiological defects such as hollowness.

[0003] Specifically, existing technologies suffer from the following shortcomings: First, the two-dimensional imaging principle inevitably leads to severe perspective distortion and nonlinear stretching when photographing three-dimensional spheres or ellipsoids, especially at the fruit's edge. This causes significant distortion in key quality indicators for seedless watermelons (such as stripe width, spacing, clarity, and morphology), resulting in a substantial decrease in the accuracy and reliability of feature extraction. Second, traditional single-shot or limited-angle shooting methods suffer from incomplete information acquisition, easily missing local defects such as scars and blemishes on the fruit's back. While complex mechanical rotation and image stitching techniques can be used to obtain full-surface information for these defects, this not only increases the complexity and cost of the equipment but may also introduce new errors during the stitching process. Third, existing methods struggle to quantify and preserve the features extracted from distorted images. For example, the true physical area of ​​defects cannot be accurately calculated; estimations based on pixel count are necessary, resulting in a lack of uniformity and stability in the grading standards regarding physical dimensions. Furthermore, when integrating weight and appearance information, existing technologies often employ simple threshold judgments or linear weighted models, failing to effectively uncover the deep coupling relationship between appearance quality and single fruit weight (e.g., fruit maturity, flesh density, etc.). Therefore, they cannot effectively utilize weight and volume data to assist in screening for secondary fruits that may have "hollow" problems, thus limiting the accuracy and intelligence level of grading decisions. Summary of the Invention

[0004] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0005] In view of the aforementioned existing problems, this invention is proposed. Therefore, this invention provides a multimodal grading method for seedless watermelons that integrates appearance texture and single-fruit weight, to solve the problems mentioned in the background art.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight, comprising: Simultaneously acquire multi-view image sequences and single-fruit weight data of the watermelon under test; Based on the multi-view image sequence, the three-dimensional surface point cloud of the watermelon is reconstructed, and a parameterized three-dimensional surface geometric model is obtained by fitting. The surface texture information of the three-dimensional surface geometry model is mapped to a two-dimensional texture map, and a distortion correction weight map is generated simultaneously. Each pixel value of the distortion correction weight map represents the real physical area scale on the three-dimensional surface geometry model represented by the corresponding pixel on the two-dimensional texture map. Based on the two-dimensional texture map, the appearance texture features of the watermelon are extracted by weighting the distortion correction weight map. By combining the appearance and texture features with the weight data of a single fruit, the quality grade of the watermelon is determined.

[0007] As a preferred embodiment of the multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight as described in this invention, the method includes: acquiring a multi-view image sequence of the watermelon to be tested, including: Multiple side-view cameras and one top-view camera, deployed in a circular array, are used to capture a single, synchronous image of the watermelon to be tested, which is placed statically on the weighing equipment.

[0008] As a preferred embodiment of the multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight as described in this invention, wherein: the parameterized three-dimensional surface geometry model is a hyperellipsoidal model, and the fitting process of the model includes: By using a nonlinear optimization algorithm, the algebraic distance from all points in the 3D surface point cloud to the surface of the hyperellipsoidal model is minimized, and the geometric and pose parameters of the model are solved.

[0009] As a preferred embodiment of the multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight as described in this invention, the generation of the distortion correction weight map includes: For any pixel on the two-dimensional texture map, calculate the corresponding parameter coordinates of the pixel on the three-dimensional surface geometry model through the mapping relationship, and solve the determinant of the metric tensor of the three-dimensional surface geometry model under the parameter coordinates based on the first fundamental form in differential geometry. The value of the determinant is used as the basis for determining the weight value of the pixel.

[0010] As a preferred embodiment of the multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight as described in this invention, wherein: mapping the surface texture information of the three-dimensional surface geometric model to a two-dimensional texture map includes: The color value of a pixel is determined by traversing each pixel of the two-dimensional texture map, determining the corresponding three-dimensional spatial point of the pixel on the three-dimensional surface geometry model through inverse mapping, finding the neighboring points of the three-dimensional spatial point in the three-dimensional surface point cloud, and determining the color value of the pixel by performing inverse distance weighted averaging on the color information of the neighboring points.

[0011] As a preferred embodiment of the multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight as described in this invention, the extraction of the appearance texture features includes: The defect area is segmented on the two-dimensional texture map, and the physical area of ​​the defect is determined by summing the weight values ​​of all pixels in the defect area on the distortion correction weight map.

[0012] As a preferred embodiment of the multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight as described in this invention, the extraction of the appearance texture features further includes: Construct the gray-level co-occurrence matrix of the two-dimensional texture map; When counting the frequency of pixel pairs with specific gray values, the contribution of each pair of pixels is weighted using the combination of their corresponding weight values ​​on the distortion correction weight map to obtain a weighted gray co-occurrence matrix, and texture features are calculated based on this matrix.

[0013] As a preferred embodiment of the multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight as described in this invention, the extraction of the appearance texture features further includes: Frequency domain analysis was performed on the two-dimensional texture map to extract the periodicity and sharpness features of the watermelon stripes; When calculating spectral energy, the contribution of each pixel is adjusted according to its corresponding weight value in the distortion correction weight map.

[0014] As a preferred embodiment of the multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight according to the present invention, the integration of the appearance texture features and the single fruit weight data includes: The appearance texture features and the single fruit weight data are respectively input into independent feature encoders to obtain the depth features of the first mode and the depth features of the second mode; The depth features of the first modality and the depth features of the second modality are input into an attention module to dynamically calculate the fusion weight of the depth features of the first modality and the depth features of the second modality.

[0015] As a preferred embodiment of the multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight as described in this invention, the attention module is a cross-modal attention module. The cross-modal attention module calculates the attention score of the first modality on the second modality by using the depth feature representation of the first modality as a query and the depth feature representation of the second modality as a key and value, generates enhanced features that reflect the mutual influence between modalities, and uses the original features and the enhanced features together to determine the quality grade of the watermelon.

[0016] Compared with existing technologies, the beneficial effects of this solution are: 1. This invention uses 3D reconstruction and parametric model fitting to losslessly unfold the irregular curved surface texture of seedless watermelons into a 2D texture map, generating a distortion correction weight map. This weight map quantifies the area scale change from the 2D mapping space to the 3D physical space, ensuring that subsequently extracted features such as defect area and texture statistics have clear physical units. This overcomes the instability of feature values ​​caused by shooting angle, distance, and perspective distortion in traditional 2D image analysis, ensuring high reproducibility and consistency of grading results.

[0017] 2. Traditional two-dimensional image processing methods cannot accurately assess the true texture morphology on curved surfaces. This invention, by analyzing distortion-free two-dimensional texture maps, can calculate the true periodicity, clarity, and physical area of ​​blemishes, avoiding feature distortion caused by pixel stretching in edge regions. For seedless watermelons, where appearance quality is a key selling point, this significantly improves the accuracy of judging key commercial indicators such as "clear stripes" and "smooth skin."

[0018] 3. By designing a deep fusion network based on a cross-modal attention mechanism, the interpretation weight of weight information can be dynamically adjusted according to appearance characteristics (such as maturity characteristics). This effectively learns the comprehensive performance of complex quality traits such as "thin skin and crisp flesh" and "sandy and hollow flesh" in terms of appearance and weight, improves the ability to identify internal defects such as "hollow melons", and achieves accurate inference from external characteristics to internal quality. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart illustrating the overall process of a multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight, as described in one embodiment of the present invention. Detailed Implementation

[0020] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0021] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0022] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0023] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.

[0024] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0025] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. Example 1

[0026] Reference Figure 1 This is the first embodiment of the present invention, which provides a multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight, including: S1. Simultaneously acquire multi-view image sequences and single fruit weight data of the watermelon to be tested.

[0027] It should be noted that, in order to achieve full coverage information collection on the surface of seedless watermelons and avoid the efficiency bottleneck and motion blur caused by traditional mechanical rotation, the present invention adopts a specific static multi-camera acquisition scheme.

[0028] Specifically, N high-resolution industrial cameras (in this embodiment, N is 8-12) are evenly deployed along a horizontal circular track to form a side-view camera array. The radius and height of this camera array are set to ensure that the field of view of all cameras completely covers the lateral spherical surface of the seedless watermelon placed at the center of the platform, and that there is sufficient overlap between adjacent cameras. Simultaneously, a top-view camera is positioned vertically downwards directly above the center of the camera array to specifically capture the top region of the watermelon (including the stem and navel). The addition of this top-view camera aims to address the image quality degradation and information loss issues caused by the excessive grazing angle of the side-view cameras in the polar regions of the watermelon, ensuring the integrity of the 3D point cloud. Finally, a high-precision electronic weighing device is positioned below the center of the camera array, its support plate also serving as a platform for placing the watermelon under test. The watermelon is placed on this platform, and while remaining stationary, weight measurement and image acquisition are performed simultaneously.

[0029] It should be noted that the core of this acquisition scheme lies in "trading space for time," replacing the time consumption of mechanical rotation with the spatial redundancy of multiple cameras. Furthermore, the use of a single-shot mode improves detection efficiency and is suitable for industrialized assembly line operations. Simultaneously, static acquisition avoids any image artifacts introduced by motion, providing fundamental data support for obtaining high-quality texture information and 3D geometric reconstruction.

[0030] Furthermore, before data acquisition, the entire static multi-camera acquisition scheme must undergo precise geometric calibration to determine the intrinsic and extrinsic parameters of each camera and their relative poses in a unified world coordinate system. This calibration process follows standard computer vision procedures: First, the intrinsic parameters of each camera (including N side-view cameras and 1 top-view camera) are calibrated independently, usually using the Zhang Zhengyou calibration method. This process solves for the intrinsic parameter matrix K of each camera: in, and These are the normalized focal lengths of the camera along the U-axis and V-axis, respectively. is the principal point coordinate, i.e., the intersection of the camera's optical axis and the imaging plane; s is the torsion parameter describing the non-perpendicularity of the two axes, which can usually be approximated as 0. In addition, the camera's distortion parameters (such as radial and tangential distortion coefficients) need to be calibrated to correct image distortion caused by lens optical characteristics.

[0031] Secondly, by capturing images of a high-precision calibration board (such as a checkerboard or dot array) in multiple different poses within the acquisition space, the extrinsic parameters of all N+1 cameras are simultaneously solved. These extrinsic parameters describe the transformation relationship from the camera coordinate system to the world coordinate system (usually with the center of the weighing platform as the origin), and are composed of the rotation matrix R and the translation vector t. Therefore, for a three-dimensional point in the world coordinate system... Its projected pixel on the i-th camera imaging plane The following projection relationship must be satisfied: in, This represents the depth of the point in the camera coordinate system. It is obtained through precise calibration. The set is the necessary input data for subsequent execution of 3D surface point cloud reconstruction algorithms (such as SfM based motion reconstruction or spatial carving algorithms).

[0032] Furthermore, once the seedless watermelon to be tested is placed on the weighing platform and stabilized, a central controller sends a synchronous hardware trigger signal. This trigger signal is simultaneously sent to all N+1 cameras and the electronic weighing device. At the instant all cameras receive the trigger signal, they simultaneously expose and capture a frame, forming a multi-view image sequence containing N+1 images. This ensures that all images capture the watermelon at a static state at the same moment, eliminating temporal inconsistencies. Then, the electronic weighing device locks in and reads a stable weight reading at the same trigger moment, obtaining the single-fruit weight data W (in grams). Finally, by assigning a unique identifier (ID) to each watermelon collected and binding this ID to the collected image sequence I and weight data W, a complete multimodal data sample can be obtained. .

[0033] S2. Based on multi-view image sequences, reconstruct the three-dimensional surface point cloud of watermelon and fit it to obtain a parameterized three-dimensional surface geometric model.

[0034] Furthermore, using the obtained precise calibration... The set is used to generate a dense 3D point cloud describing the geometry of the watermelon surface, denoted as . Based on this, this embodiment adopts a hybrid reconstruction strategy that combines contour information and multi-view stereo (MVS) to fully utilize the characteristics of the smooth surface of seedless watermelons and improve the robustness and accuracy of the reconstruction.

[0035] Furthermore, for each image in the multi-view image sequence I... A pre-trained deep learning semantic segmentation network (such as U-Net or DeepLabv3+) is applied for foreground segmentation to extract the binarized mask of the watermelon region. This semantic segmentation network is trained on a large number of annotated watermelon images and can resist interference from lighting changes and complex backgrounds to achieve pixel-level segmentation.

[0036] It should be noted that semantic segmentation is introduced because the surface texture of seedless watermelons is sometimes quite simple, lacking sufficient corner points and other features, which poses a challenge to MVS based purely on feature matching. However, its contour information is very stable and clear. By back-projecting the binarized masks of all cameras into 3D space, multiple view frustums are formed. The intersection of these view frustums constitutes the visual shell of the watermelon. This visual shell can effectively eliminate mismatched points in the background and limit the search range of MVS, thereby significantly improving computational efficiency and reconstruction quality.

[0037] Furthermore, within the spatial range defined by the visual shell, a PlaneSweeping-based MVS algorithm is applied to generate a dense point cloud. Specifically, one camera is first used as a reference camera, and a series of planes with different depths are assumed within its field of view. For each assumed plane, the images of other views (source views) are projected onto that plane through homography transformation, and photometric consistency matching metrics (e.g., Normalized Cross-Correlation (NCC) or Zero-Mean Normalized Cross-Correlation (ZNCC)) are performed with the reference view. Then, for each pixel p in the reference view, its optimal depth is determined by the plane that achieves the highest photometric consistency score among all assumed depth planes. By traversing all pixels, a depth map is obtained. This process is applied to the case where each camera is used as a reference camera, and all generated depth maps are fused and then back-projected back to the world coordinate system, ultimately yielding a dense 3D surface point cloud with color information. .

[0038] Furthermore, due to the original point cloud obtained from the reconstruction Although it describes the true geometry of a watermelon, its data structure is discrete and unordered, and may contain noise and a few outliers. Therefore, in order to obtain a smooth, continuous surface representation that is easy to perform mathematical analysis, this invention further fits the point cloud into a parameterized three-dimensional surface geometry model.

[0039] Specifically, considering that the overall shape of seedless watermelons is mostly near-spherical or ellipsoidal, this embodiment preferably uses a hyperellipsoid model as its parametric geometric model. This is because the hyperellipsoid model can flexibly describe various shapes from spheres and ellipsoids to approximately cuboids, and it is controlled by only a few parameters, exhibiting strong compactness and noise resistance. Therefore, based on this, the implicit equation of a hyperellipsoid centered at the origin with its principal axes aligned with the coordinate axes can be defined as: in, , , It refers to the length of the semi-axis in the three main axis directions, which controls the size of the object; and These are shape parameters. The squareness / sharpness in the vertical direction was controlled. The squareness / sharpness in the horizontal direction was controlled. When At that time, the model degenerates into a standard ellipsoid.

[0040] In addition, to fit a seedless watermelon in any pose, it is also necessary to introduce pose parameters, including the center position. And rotational attitude (which can be represented by quaternions q or Euler angles). Therefore, the total set of parameters to be optimized is .

[0041] Furthermore, the fitting process is essentially solving an optimization problem, the goal of which is to find a set of optimal parameters. This makes all points in the point cloud ( The sum of the distances to the surface of the hyperellipsoid model is minimized.

[0042] Specifically, this invention uses minimizing the algebraic distance as the optimization objective function for points after pose transformation. Its algebraic distance to the surface of the hyperellipsoid is Therefore, the objective function Defined as: in, This indicates that the points are based on the pose parameters. Perform rigid body transformation. .

[0043] Furthermore, for the aforementioned objective function, a nonlinear optimization algorithm, such as the Levenberg-Marquardt (LM) algorithm, can be used to solve the fitting process. Specifically, the nonlinear optimization algorithm starts with reasonable initial parameters (e.g., the initial orientation and size of the point cloud estimated through principal component analysis (PCA)) and iteratively updates the parameters. until the objective function It converges to the minimum value.

[0044] S3. Map the surface texture information of the three-dimensional surface geometry model to a two-dimensional texture map, and simultaneously generate a distortion correction weight map. Each pixel value in the distortion correction weight map represents the actual physical area scale on the three-dimensional surface geometry model represented by the corresponding pixel on the two-dimensional texture map.

[0045] Furthermore, a rectangular two-dimensional image is created, which records the color and texture information of the entire seedless watermelon surface completely and without overlap. Simultaneously, this invention employs an inverse mapping strategy. This strategy starts from the two-dimensional target (texture image) and inversely calculates the information of that target on the three-dimensional source (watermelon surface) to ensure that each pixel of the texture image is effectively filled, avoiding the holes or overlaps that may occur with forward mapping.

[0046] Specifically, a global parametric coordinate system is established for the obtained hyperellipsoidal model. This process is similar to the latitude and longitude system in geography, and we can use extended spherical coordinates. To parameterize any point on the surface of the hyperellipsoid ,in, Similar to latitude, Similar to longitude. Based on this, the parametric equations for points on the surface of a hyperellipsoid can be expressed as: It should be noted that the parametric equations for the points on the surface of the hyperellipsoid established a two-dimensional parameter domain. Points on the surface of a three-dimensional hyperellipsoid A one-to-one correspondence between them.

[0047] Furthermore, create a blank 2D texture map with a size of For example, it can be set , The horizontal axis u of this texture map corresponds to longitude. The vertical axis v corresponds to latitude. By traversing every pixel on the texture map The range of the horizontal axis u is: The range of the vertical axis v is For each pixel, use coordinates. This indicates that the following reverse mapping process will be executed: 301.1. Pixel coordinates Linear transformation to two-dimensional parameter domain : 301.2. Convert the parameter coordinates Substituting the parametric equations of the hyperellipsoid above The pose transformation (rotation and translation) obtained in step S2 is applied to obtain the 3D surface point corresponding to the pixel in the world coordinate system. The three-dimensional surface point is the point on the ideal smooth model that we fitted.

[0048] 301.3 Since a point on a 3D surface is only a geometric location, its color needs to be determined from the original point cloud reconstructed in step S2, which contains true color information. Therefore, we can obtain it through [the following]. Perform a K-Nearest Neighbors (k-NN) search to find the k nearest point cloud points to a point on the 3D surface. .

[0049] 301.4. Furthermore, to obtain a stable and noise-resistant color value, we also need to perform inverse distance weighting (IDW) on the colors of these k neighboring points. Let... The color is Its Euclidean distance to points on the three-dimensional surface is Then the pixel final color value The calculation is as follows: in, It is a very small positive number, used to prevent when The denominator is zero. In this embodiment, k typically takes the value of 5 to 8.

[0050] It should be noted that, through the aforementioned inverse mapping strategy, the present invention cleverly combines the geometric regularity of the parametric model with the texture fidelity of the original point cloud. By using an ideal smooth model fitted from 3D surface points to guide the sampling position, the continuity of texture mapping is ensured. Simultaneously, sampling colors from the real point cloud rather than multiple original images avoids complex issues such as multi-view visibility and inconsistent lighting, making the texture reconstruction process more robust and efficient.

[0051] Furthermore, when unfolding a three-dimensional surface onto a two-dimensional plane, area distortion inevitably occurs. For example, the two poles of a watermelon (points on a sphere) are stretched into a line in the unfolded image (rectangle), and the area of ​​their surrounding region is greatly magnified. If this distortion is ignored and pixel statistics are directly performed on the unfolded image, the results will be completely unrelated to measurements in the real physical world. Therefore, it is necessary to quantify and correct this distortion.

[0052] Specifically, for each pixel on the texture map, calculating the true physical area of ​​the tiny facet represented by that pixel on the surface of a 3D hyperellipsoid requires the use of the first fundamental form of differential geometry. It should be explained that the first fundamental form describes the differential of the arc length on the surface, and its coefficients constitute the metric tensor g, quantifying the local stretching and shearing from the parameter space to 3D space. Therefore, for a given surface from the parameter domain... The defined surface, i.e., the parametric equation The components of its metric tensor are: in, in, and Is it a curved surface? The tangent vector at a point can be obtained by taking the partial derivatives of the parametric equations of the hyperellipsoid. E represents... The stretching factor in the direction. G represents... The stretching factor in direction. F represents the shear distortion or angular distortion factor. If F=0, it indicates that the two tangent vectors... and In three-dimensional space, they are perpendicular to each other, meaning the mapping is orthogonal or conformal. That is, a 90-degree grid angle is still 90 degrees on the surface. This is an ideal but not usually global case (for example, standard spherical coordinate mapping is orthogonal at the equator, but not elsewhere). If F≠0, it means that the two tangent vectors are no longer perpendicular in three-dimensional space, and the original right angle is distorted into an acute or obtuse angle.

[0053] Furthermore, in the parameter domain an infinitesimal rectangular region Its corresponding area on a three-dimensional curved surface We obtain it from the following formula: in, It is the area element, which describes the actual area of ​​a unit area in three-dimensional space under a parametric coordinate system.

[0054] Furthermore, a single-channel floating-point image of the same size as the texture map is created as a distortion correction weight map, denoted as... Then, iterate through each pixel of the distortion correction weight map: 302.1. Similar to the first step of the aforementioned reverse mapping process, first, the pixel coordinates... Transform to parametric coordinates .

[0055] 302.2 Calculate the hyperellipsoid in Tangent vector of a point and Then calculate the components of the metric tensor. .

[0056] 302.3 Calculate the determinant of the metric tensor .

[0057] 302.4, assign weight value to this pixel Set to the square root of the area element, i.e. This value is the basis for determining the pixel weight value, representing the area per unit parameter on the two-dimensional texture map. To the true physical area of ​​a three-dimensional curved surface The scaling factor.

[0058] It should be noted that, after the above processing, if you want to calculate the true area of ​​a blemish, you only need to segment the pixel mask of the blemish on the two-dimensional texture map, and then sum up the corresponding weight values ​​of all pixels in the distortion correction weight map to obtain the accurate physical area of ​​the blemish (unit: square millimeters).

[0059] S4. Based on the two-dimensional texture map, the appearance texture features of the watermelon are extracted by weighting the distortion correction weight map.

[0060] Furthermore, the defective regions are segmented at the pixel level on the two-dimensional texture map. For this segmentation, this embodiment employs a deep learning-based semantic segmentation network, such as an optimized U-Net++ model.

[0061] Specifically, the input to this model is a two-dimensional texture map generated in step S3, with a size of [missing information]. (e.g., 1024×512×3). The output of this model is a binary defect mask of the same size as the input. ,in Represents pixels This area is considered a defect. This indicates a normal watermelon surface. Furthermore, the U-Net++ model is trained under supervision using a dataset containing a large number of seedless watermelon texture images and their corresponding finely annotated blemish masks. During model training, a combination of the Dice loss function and the cross-entropy loss function is used as the loss function to address potential issues such as small objects and class imbalance in blemish areas.

[0062] Furthermore, in obtaining the defect mask Traditional methods estimate defect size by simply counting pixels with a value of 1 in the mask. However, this method completely ignores mapping distortion. Therefore, the present invention utilizes a distortion correction weight map for calculation. As mentioned above, In essence, it is the area scaling factor from parameter space to three-dimensional Euclidean space, and a microelement in parameter space is... Its corresponding physical area is Meanwhile, in a discrete pixel grid, and Therefore, the total physical area of ​​the defects. This can be obtained by summing (i.e., integrating) the physical areas of all pixels within the mask region: in, Is The defect mask value (0 or 1) at the location. Is The distortion correction weight value at the location. and These are the parameter domains. and The distances of the pixels are denoted by 0, and their product represents the "area" of a single pixel in the parameter space.

[0063] It should be noted that the contribution of each segmented defective pixel is determined through its corresponding... The value is corrected. That is, a flawed pixel at either end of the watermelon (the top and bottom edges of the 2D texture map) has its... The value will be very small, and therefore its contribution to the total area will also be small; conversely, a flawed pixel in the watermelon equator (the middle of the 2D texture map) will have a much larger value. A larger value indicates a larger contribution. Therefore, the final result is... It is a stable value with a definite physical unit (e.g., square millimeter), completely unaffected by the watermelon's placement during harvesting, thus achieving true measurement preservation.

[0064] Furthermore, since the standard Gray-Level Co-occurrence Matrix (GLCM) describes texture by statistically analyzing the frequency of occurrence of pixel pairs with specific gray values, it implicitly assumes that the physical region represented by each pixel in the image is equal. Applying the standard GLCM directly to a two-dimensional texture map will produce a serious bias because pixels in extreme regions are excessively "magnified," and their texture patterns will naturally have a disproportionately large impact on the final statistical results. Therefore, this invention uses a weighted approach to GLCM to correct this bias.

[0065] Specifically, the two-dimensional texture image is converted into a grayscale image and its grayscale levels are quantized to... Each level (in this embodiment, Take 64). Then the weighted GLCM is denoted as... , defined as: at a distance of , direction is Under the given conditions, the weighted sum of the occurrences of pixel pairs with gray levels i and j. The calculation method is as follows: in, It is a quantized grayscale image. It is a pixel Along direction The neighboring pixels after moving a distance d. It is a Kronecker function, which is 1 when a=b and 0 otherwise. This is a weighted term for the contribution of pixel pairs. Here, the geometric mean of the weight values ​​of the two corresponding pixels is used as the contribution weight of the pixel pair.

[0066] It should be noted that the geometric mean was chosen as the weighting method because it can smoothly reflect the importance of the "micro-region" jointly defined by two spatial points. For a pixel pair, if one of them is located in a region with severe distortion (small weight), even if the other is located in a region with less distortion (large weight), the overall reliability or representativeness should be reduced accordingly. Therefore, the geometric mean can precisely reflect this mutually restrictive relationship.

[0067] Furthermore, the weighted gray-level co-occurrence matrix is ​​normalized to obtain a weighted probability matrix. Unlike standard GLCM, this normalization factor is the sum of all weight elements in the matrix: Furthermore, based on this normalized weighted probability matrix By calculating a series of standard Haralick texture features, a feature set capable of describing the macroscopic statistical properties of the texture of a real watermelon surface can be constructed. These texture features include, but are not limited to: Weighted contrast ratio: used to measure the sharpness and darkness of stripes. Weighted correlation: used to reflect the linear direction and continuity of the stripes.

[0068] Weighted energy: Used to measure the uniformity of texture; the weighted energy is higher in smooth areas.

[0069] Weighted homogeneity measures local variations in texture; the more similar the stripes, the higher the value.

[0070] Furthermore, by performing weighted frequency domain analysis on the texture map, the core quality characteristics representing the stripes of the watermelon were extracted.

[0071] Specifically, directly performing a Fourier Transform (FFT) on a two-dimensional texture image also presents problems. This is because the FFT assumes that the input signal is uniformly sampled in space, while a two-dimensional texture image is in the parameter domain. The signal is uniformly sampled in the middle. This can easily lead to oversampling of the signal in the pole region, producing artifacts in the frequency domain and interfering with the determination of the true fringe frequency. Therefore, to solve this problem, the present invention uses a method that first performs frequency domain analysis before conducting frequency domain analysis. The spatial signal is weighted to simulate the non-uniform sampling effect on a real curved surface. The process is as follows: S401.1 Generating a weighted grayscale image : It should be noted that the physical meaning of this step is to multiply the gray value of each pixel by the actual physical area it represents, so that the "energy" of the signal is proportional to its physical scale. The signal in regions with small distortion (equator) is preserved, while the signal in regions with large distortion (polar regions) is suppressed.

[0072] S401.2, Weighted image Perform a two-dimensional fast Fourier transform (2D-FFT) to obtain its spectrum. .

[0073] S401.3, Analyzing the power spectrum By extracting periodic and clarity features, the power spectrum reflects the energy distribution of different frequency components.

[0074] Specifically, for periodic features, since watermelon stripes are usually distributed along a specific direction (mainly the meridian direction, corresponding to the horizontal axis of the spectrum), they will form bright spots (peaks) symmetrical about the origin in the power spectrum. The periodic features are determined by the dominant frequency, and the processing steps are as follows: In power spectrum In the middle, the DC component (i.e., the origin) is shielded. (A small neighborhood nearby).

[0075] Search for the frequency point with the highest energy value in the remaining region. .

[0076] Calculate the Euclidean distance from the peak point to the origin. The Euclidean distance is the dominant spatial frequency.

[0077] Periodic characteristics Defined as the reciprocal of the dominant frequency: This directly corresponds to the average physical spacing of the stripes on the watermelon surface.

[0078] Specifically, regarding sharpness features, sharp, well-defined stripes will generate richer high-frequency components in the spectrum. Based on this, a high-frequency energy ratio is defined as a sharpness feature, and its processing steps are as follows: Define a low-frequency cutoff radius This radius can be set based on statistical analysis of a large number of samples; for example, it can be set to... This is to ensure that the dominant frequency itself is not classified as a high-frequency region.

[0079] Calculate the high-frequency region (i.e.) The total energy of the region .

[0080] Calculate total energy (Remove DC component)

[0081] Sharpness features Defined as the ratio of high-frequency energy to total energy: The larger this ratio, the richer the high-frequency components and the clearer the stripes.

[0082] It should be noted that, after the above processing, a comprehensive appearance texture feature vector can be obtained. ,in, This represents the feature set describing the macroscopic statistical characteristics of the surface texture of a real watermelon. Each component in this vector has undergone distortion correction, accurately reflecting the true physical state of the seedless watermelon surface.

[0083] S5. Integrate appearance and texture characteristics with single fruit weight data to determine the quality grade of watermelon.

[0084] It should be noted that this step aims to fuse the appearance texture feature vector with the one-dimensional single-fruit weight data W to output an accurate quality grade. However, traditional methods such as feature concatenation or linear weighting cannot capture the complex, non-linear intrinsic relationship between these two heterogeneous modal data. For example, a watermelon that is lighter in weight but larger in volume, combined with its appearance features of "blurred stripes and low contrast," may have the defect of being "hollow" inside; while a watermelon of the same weight, if its appearance has "clear stripes and zero defect area," is very likely to be a high-quality fruit. Therefore, in order to effectively explore this coupling relationship, the present invention designs a deep fusion network model based on a cross-modal attention mechanism.

[0085] Furthermore, before inputting the data into the fusion network model, the data from both modalities needs to be processed to make them suitable for the input format of the deep learning model.

[0086] Specifically, for the first modality (appearance texture), the appearance texture feature vector from step S4 is input. ,in This refers to the dimension of the feature vector. The appearance texture feature vector is input into a separate Multi-Layer Perceptron (MLP) as its feature encoder. This encoder is... It consists of several fully connected layers, designed to map the original features to a higher-dimensional deep feature space. In this embodiment, a three-layer MLP is used, with the following structure: Each layer is followed by a ReLU activation function and a batch normalization layer to accelerate convergence and prevent overfitting. After passing through this feature encoder, the deep features of the first modality are obtained. ,in, It is a unified depth feature dimension (in this embodiment, ).

[0087] Specifically, for the second mode (single fruit weight), the single fruit weight data W obtained in step S1 is a scalar. First, this single fruit weight data is standardized using Z-score. ,in, This represents the mean of the single fruit weight data. The standard deviation of the single fruit weight data is used to ensure it follows a standard normal distribution with a mean of 0 and a variance of 1, thus eliminating the influence of dimensions. Similarly, the standardized single fruit weight data... The input is fed into a separate MLP encoder. Since the input is only one-dimensional, the encoder's structure can be relatively simple, and its structure is as follows: Similarly, a ReLU activation and batch normalization layer is appended after each layer. By obtaining the output of this encoder, the depth features of the second modality can be obtained. .

[0088] It should be noted that setting up independent encoders allows models to first learn effective feature representations within their respective modalities before engaging in cross-modal interactions. This is much more flexible than concatenating the raw data from the outset and can better handle differences in the statistical characteristics and information density of data from different modalities.

[0089] Furthermore, a feature fusion module based on cross-modal attention is constructed. This module aims to allow features from one modality to "focus on" and "extract" relevant information from another modality, thereby generating more discriminative enhanced features.

[0090] Specifically, the input to this module is two deep feature vectors. and Meanwhile, in this embodiment, richer and more dimensional appearance texture features are used. As a query (Q), the relatively simple but crucial single-fruit weight feature will be used. The model uses Q as the key (K) and value (V) to dynamically adjust the representation of texture features by calculating the similarity between Q and K to determine how much "attention" weight should be assigned to V. The calculation process is as follows: S501.1, through three independent and learnable linear transformation matrices Project the input depth features onto Space, in which, yes Dimensions (in general) ): S501.2 Calculate the dot product similarity between query Q and key K, and scale it to prevent the gradient from being too small. The attention score is included in this calculation. It can be represented as: S501.3. Apply the Softmax function to the scores to obtain the normalized attention weights. : S501.4 Using attention weights We perform a weighted summation on the value V to obtain the enhanced feature that reflects the modulation of texture information by weight information. : Furthermore, the original depth features , Enhanced features generated by the feature fusion module The features are then concatenated to form the final fused feature vector. : It should be noted that the present invention employs a concatenation rather than addition approach, and preserves the original features through residual connections, ensuring lossless information transfer. This allows the final classifier to utilize not only the original, independent modal information, but also the enhanced information with deep semantics generated by inter-modal interactions, thereby making more robust decisions.

[0091] Furthermore, for the model's classification head, its input is the fused feature vector. This classification head consists of one or more fully connected layers, used to map fused features to the final class space. The structure of each fully connected layer is as follows: The MLP's output layer does not use ReLU; instead, it is directly connected to a Softmax activation function, outputting the predicted probability for each quality level. ,in, It is a predefined number of quality grades (e.g., 3 grades: excellent, first-class, and second-class). Each element Both represent the probability that the watermelon to be tested belongs to the i-th level, and Finally, the category with the highest probability is selected as the final quality level. .

[0092] Furthermore, training this deep fusion network model requires a labeled multimodal dataset. Each sample contains a sample of a seedless watermelon to be tested. The data pairs, along with their corresponding true quality grade labels, are presented. These labels are determined by agricultural experts through fruit breakage inspection (measuring sugar content, checking for hollow centers, and examining flesh texture, etc.) and then one-hot encoded. Simultaneously, the standard cross-entropy loss function is used to measure the model's predicted probability. With real labels The difference between them: Preferably, during training, the AdamW optimizer is used. This optimizer adds weight decay decoupling to the Adam optimizer, resulting in better generalization performance. In this embodiment, the initial learning rate is set to... Furthermore, a learning rate decay strategy (such as Cosine Annealing) is employed. Through Mini-batch Stochastic Gradient Descent (Mini-batch SGD), all learnable parameters of the entire network model (including the two encoders, the transformation matrix of the feature fusion module, and the classification head) are iteratively updated on the training set, with the goal of minimizing the cross-entropy loss function until the model's performance on the validation set converges.

[0093] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0094] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0095] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0096] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0097] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0098] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A multi-modal grading method for seedless watermelon that fuses appearance texture and single fruit weight, characterized in that, include: Simultaneously acquire multi-view image sequences and single-fruit weight data of the watermelon under test; Based on the multi-view image sequence, the three-dimensional surface point cloud of the watermelon is reconstructed, and a parameterized three-dimensional surface geometric model is obtained by fitting. The surface texture information of the three-dimensional surface geometry model is mapped to a two-dimensional texture map, and a distortion correction weight map is generated simultaneously. Each pixel value of the distortion correction weight map represents the real physical area scale on the three-dimensional surface geometry model represented by the corresponding pixel on the two-dimensional texture map. Based on the two-dimensional texture map, the appearance texture features of the watermelon are extracted by weighting the distortion correction weight map. By combining the appearance and texture features with the weight data of a single fruit, the quality grade of the watermelon is determined.

2. The multi-modal grading method of seedless watermelons fusing appearance texture with individual fruit weight as claimed in claim 1, wherein, Obtain a multi-view image sequence of the watermelon to be tested, including: Multiple side-view cameras and one top-view camera, deployed in a circular array, are used to capture a single, synchronous image of the watermelon to be tested, which is placed statically on the weighing equipment.

3. The multi-modal seedless watermelon grading method of fusing appearance texture with individual fruit weight of claim 1, wherein, The parameterized three-dimensional surface geometry model is a hyperellipsoid model, and the fitting process of the model includes: By using a nonlinear optimization algorithm, the algebraic distance from all points in the 3D surface point cloud to the surface of the hyperellipsoidal model is minimized, and the geometric and pose parameters of the model are solved.

4. The multi-modal seedless watermelon grading method of fusing appearance texture with individual fruit weight of claim 1, wherein, Generating the distortion correction weight map includes: For any pixel on the two-dimensional texture map, calculate the corresponding parameter coordinates of the pixel on the three-dimensional surface geometry model through the mapping relationship, and solve the determinant of the metric tensor of the three-dimensional surface geometry model under the parameter coordinates based on the first fundamental form in differential geometry. The value of the determinant is used as the basis for determining the weight value of the pixel.

5. The multi-modal seedless watermelon grading method of fusing appearance texture with individual fruit weight of claim 1, wherein, Mapping the surface texture information of the three-dimensional surface geometry model to a two-dimensional texture map includes: The color value of a pixel is determined by traversing each pixel of the two-dimensional texture map, determining the corresponding three-dimensional spatial point of the pixel on the three-dimensional surface geometry model through inverse mapping, finding the neighboring points of the three-dimensional spatial point in the three-dimensional surface point cloud, and determining the color value of the pixel by performing inverse distance weighted averaging on the color information of the neighboring points.

6. The multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight as described in claim 1 or 5, characterized in that, Extracting the appearance texture features includes: The defect area is segmented on the two-dimensional texture map, and the physical area of ​​the defect is determined by summing the weight values ​​of all pixels in the defect area on the distortion correction weight map.

7. The multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight as described in claim 1, characterized in that, Extracting the appearance texture features also includes: Construct the gray-level co-occurrence matrix of the two-dimensional texture map; When counting the frequency of pixel pairs with specific gray values, the contribution of each pair of pixels is weighted using the combination of their corresponding weight values ​​on the distortion correction weight map to obtain a weighted gray co-occurrence matrix, and texture features are calculated based on this matrix.

8. The multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight as described in claim 1, characterized in that, Extracting the appearance texture features also includes: Frequency domain analysis was performed on the two-dimensional texture map to extract the periodicity and sharpness features of the watermelon stripes; When calculating spectral energy, the contribution of each pixel is adjusted according to its corresponding weight value in the distortion correction weight map.

9. The multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight as described in claim 1, characterized in that, The fusion of the appearance texture features and the single fruit weight data includes: The appearance texture features and the single fruit weight data are respectively input into independent feature encoders to obtain the depth features of the first mode and the depth features of the second mode; The depth features of the first modality and the depth features of the second modality are input into an attention module to dynamically calculate the fusion weight of the depth features of the first modality and the depth features of the second modality.

10. The multimodal grading method for seedless watermelons that integrates appearance texture and single fruit weight as described in claim 9, characterized in that, The attention module is a cross-modal attention module. The cross-modal attention module calculates the attention score of the first modality to the second modality by using the deep feature representation of the first modality as the query and the deep feature representation of the second modality as the key and value, generates enhanced features that reflect the mutual influence between modalities, and uses the original features and the enhanced features together to determine the quality grade of the watermelon.