A computer vision-based full-automatic vegetable sorting method and system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU SHENGXIAO AGRI DEV CO LTD
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-04
AI Technical Summary
[0003]然而,在实际生产过程中,现有基于可见光图像的视觉分拣技术存在以下不足:首先,蔬菜作为鲜活农产品,其品质不仅取决于表皮是否完好、颜色是否均匀等外部形态指标,更与内部组织是否发生病变、水分含量是否正常、成熟度是否一致等生化属性密切相关
1、本发明利用高光谱成像技术获取物体的多维空谱特征。通过对预设波段范围的精细解析,系统不仅能够识别蔬菜表面的视觉特征及外部缺陷,更能穿透表皮层探测内部生化信息,填补现有技术无法进行无损内部品质检测的空白,实现全方位的品质评估。
Smart Images

Figure CN122499989A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer vision, specifically relating to a fully automated vegetable sorting method and system based on computer vision. Background Technology
[0002] With the continuous improvement of agricultural automation and intelligence, computer vision-based vegetable sorting technology has become a research hotspot in the field of post-harvest processing of agricultural products. Traditional vegetable sorting equipment typically uses visible light image sensors to acquire surface images of vegetables, analyzes their external geometric features such as color, shape, and size to determine their quality grade, and drives actuators such as air jets or robotic arms to complete the sorting. This method has been applied to some extent in the fruit and vegetable primary processing industry, and can partially replace manual sorting, improving production efficiency.
[0003] However, in actual production processes, existing visual sorting technologies based on visible light images have the following shortcomings: First, as fresh agricultural products, the quality of vegetables depends not only on external morphological indicators such as whether the skin is intact and the color is uniform, but also on biochemical attributes such as whether the internal tissues are diseased, whether the moisture content is normal, and whether the maturity is consistent. Visible light cannot penetrate the vegetable skin, making it difficult to detect hidden quality defects such as black heart, browning, and internal rot, resulting in an inability to accurately evaluate the overall quality of vegetables based solely on appearance information. Second, existing technologies typically treat the external morphological characteristics of vegetables separately from their internal quality information, lacking analytical methods to deeply integrate the two. This makes it difficult to capture the intrinsic correlation between spatial structural anomalies and spectral response shifts, thus limiting the discrimination ability and robustness of the sorting model. Furthermore, due to the randomness of the posture and position of vegetables on the conveyor belt, coupled with the response delay of the actuator, existing systems often struggle to achieve precise synchronous sorting actions, easily leading to missorting or missed sorting.
[0004] Therefore, how to simultaneously acquire the external morphology and internal biochemical information of vegetables without damage, and effectively integrate these two types of cross-modal features to improve the accuracy of quality judgment, while achieving precise coordination with the execution mechanism, is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide a fully automated vegetable sorting method and system based on computer vision, which can effectively solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A fully automated vegetable sorting method based on computer vision includes the following specific steps: Step S1: Obtain the original hyperspectral data cube of the vegetables to be sorted within a preset wavelength range using a hyperspectral imaging device, and perform reflectance correction processing on the original hyperspectral data cube to obtain a standard reflectance data cube. Step S2: Extract the region of interest of the vegetables to be sorted from the standard reflectance data cube, and construct a spatial feature matrix describing the external morphology and a spectral feature vector describing the internal biochemical properties based on the region of interest; Step S3: Input the spatial feature matrix and the spectral feature vector into a pre-constructed deep fusion recognition model. The model includes a multi-dimensional spatial-spectral fusion convolutional layer, which is used to simultaneously capture the cross-modal coupling features between the spatial feature matrix and the spectral feature vector, and output the quality grade determination result of the vegetables to be sorted based on the cross-modal coupling features. Step S4: Based on the quality grade determination result and the real-time position coordinates of the vegetables to be sorted on the conveyor belt, calculate the sorting trigger delay including multiple compensation corrections, and drive the end effector to sort the vegetables to be sorted into the corresponding collection channel when the delay arrives.
[0007] A fully automated vegetable sorting system based on computer vision includes: Conveyor belts are used to transport vegetables to be sorted. A hyperspectral imaging device is installed above the conveyor belt to acquire the original hyperspectral data cube of the vegetables to be sorted within a preset wavelength range. An industrial control computer, connected to the hyperspectral imaging device, is configured to: The original hyperspectral data cube is subjected to reflectance correction processing to obtain a standard reflectance data cube; The region of interest (ROI) of the vegetables to be sorted is extracted from the standard reflectance data cube, and a spatial feature matrix describing the external morphology and a spectral feature vector describing the internal biochemical properties are constructed based on the ROI. The spatial feature matrix and the spectral feature vector are input into a pre-built deep fusion recognition model. The model contains a multi-dimensional spatial-spectral fusion convolutional layer, which is used to simultaneously capture cross-modal coupled features and output the quality grade determination result of the vegetables to be sorted. Based on the quality grade determination result and the real-time position coordinates of the vegetables to be sorted on the conveyor belt, the sorting trigger delay, which includes multiple compensation and correction amounts, is calculated. An end effector, connected to the industrial control computer, is used to sort the vegetables to be sorted into the corresponding collection channels according to the instructions of the industrial control computer when the sorting trigger delay arrives.
[0008] In summary, this application includes at least one of the following beneficial technical effects: 1. This invention utilizes hyperspectral imaging technology to acquire multidimensional spatial spectral features of objects. Through fine analysis of preset wavelength ranges, the system can not only identify visual features and external defects on the surface of vegetables, but also penetrate the epidermis to detect internal biochemical information, filling the gap in existing technologies that cannot perform non-destructive internal quality testing, and achieving comprehensive quality assessment.
[0009] 2. This invention achieves deep coupling between spatial morphological information and the spectral properties of substances. Spectral features can effectively eliminate interference from light fluctuations in the production environment. For vegetables with similar surface features but differences in internal quality, hyperspectral data can reflect subtle changes in their internal substances, significantly improving identification accuracy and demonstrating strong robustness.
[0010] 3. This invention employs a non-contact scanning and sorting method, eliminating the need for physical contact with vegetables throughout the process and minimizing secondary mechanical damage. Simultaneously, through an optimized algorithm architecture and precise control of sorting trigger delays, the system achieves highly automated assembly line operations, reaching a preset high-speed standard, significantly improving sorting efficiency and reducing labor costs.
[0011] 4. The method of this invention supports the fusion and analysis of multi-source features and possesses transfer learning capabilities, enabling dynamic adjustment of the recognition model and control parameters according to different vegetable varieties. By establishing a comprehensive evaluation system covering external morphology and internal biochemical indicators, the system can adapt to the sorting needs of various agricultural products. The collaborative work of industrial-grade control and real-time sensors ensures the stability and accuracy of system operation. Attached Figure Description
[0012] Figure 1 This is a schematic diagram of the overall technical solution for a fully automated vegetable sorting method based on computer vision. Figure 2 A schematic diagram illustrating the core principles of a deep fusion recognition model; Figure 3 This is a flowchart illustrating the logic of hyperspectral image preprocessing and spatial spectral feature extraction. Figure 4 A schematic diagram illustrating the multi-level interaction and data flow between industrial control computers and actuators; Figure 5 This is a schematic diagram illustrating the principle of sorting trigger delay calculation and actuator control. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this invention clearer, the following description is provided in conjunction with the appendix. Figure 1 To be continued Figure 5 The present invention will be further described in detail below with reference to specific embodiments.
[0014] The first aspect is the fully automated vegetable sorting method based on computer vision disclosed in this application, which is carried out according to the following implementation.
[0015] In this embodiment, a fully automated vegetable sorting method based on computer vision is provided, where the entire sorting process is implemented in a highly integrated industrial automation environment. This environment includes a continuously operating conveyor system, a hyperspectral imaging device, an industrial control computer, a sensor array, and an end effector.
[0016] The first step, S1, involves acquiring a cube of raw hyperspectral data covering the 400 nm to 1000 nm wavelength range using a linear scan camera in a closed, light-proof environment. Then, using pre-calibrated whiteboard reference data W and dark field reference data B, the raw image I acquired in real time is normalized to obtain standard reflectance data R. This eliminates interference from uneven light sources and equipment noise, providing a physically meaningful and highly comparable spectral information basis for subsequent feature extraction. The specific implementation follows the sub-steps described below.
[0017] Step S101: Set up and initialize the hyperspectral imaging acquisition environment. The hyperspectral imaging device is installed inside a light-shielding cover above the conveyor belt to isolate it from interference from external ambient light. The core imaging component of the device is a linear scan camera with an effective wavelength coverage range of 400 nm to 1000 nm and a spectral resolution of 2.8 nm.
[0018] This configuration ensures sufficiently precise detection capabilities across the visible to near-infrared bands, allowing for the continuous acquisition of spectral information across 176 bands in a single scan. Simultaneously, the system's light source must be ensured to be in a stable operating state before the formal sorting operation commences.
[0019] In this embodiment, the light source array is composed of a combination of halogen tungsten lamps and LEDs. Its incident angle is at a 45-degree angle to the vertical direction of the scanning camera. This geometric layout can effectively suppress the interference of diffuse reflection and shadows on subsequent spectral imaging.
[0020] In step S102, raw hyperspectral image data is acquired using a linear array scanning method. When the vegetables to be sorted enter the scanning area along the conveyor belt, a trigger signal is generated by a photoelectric sensor to activate the scanning camera. The linear array scanning camera then performs line-by-line scanning at a preset constant line frequency.
[0021] Each scanned row contains both the positional information of the horizontal pixels in the spatial dimension and the spectral intensity information of each pixel in that row across various wavelengths. As the conveyor belt continues to move longitudinally, multiple consecutive scanned rows are stitched together in real time, ultimately generating a three-dimensional hyperspectral data cube for each vegetable. The three dimensions of this cube are the horizontal coordinate x, the vertical coordinate y, and the one-dimensional spectral axis λ in two-dimensional spatial coordinates.
[0022] Step S103: Execute the system pre-calibration process to obtain the reference parameters required for reflectance correction. After acquiring the original hyperspectral image data, precise reflectance correction processing must be performed. This is because the spectral radiation intensity of the light source varies in different wavelength bands, and the camera itself has dark current noise and lens edge light intensity attenuation effects, causing the originally acquired digital brightness value (DN) to not directly and accurately reflect the physical reflectance properties of the vegetable surface. Therefore, the system performs a calibration process before starting the sorting operation. The calibration process specifically consists of the following two steps.
[0023] Step 1: Place a standard polytetrafluoroethylene whiteboard with a reflectivity of 99% within the scanning field of view, collect its reflectivity data across the entire wavelength range, and record this data as the reference whiteboard data W.
[0024] Step 2: Completely block the camera lens to reduce it to zero light intake, collect the background noise data of the system at this time, i.e., dark current data, and record this data as the reference dark field data B.
[0025] Step S104: Perform normalized reflectance correction calculation on the real-time acquired raw hyperspectral image data. During the real-time sorting process, for each newly acquired frame of raw hyperspectral image data, denoted as raw image data I, the industrial control computer will call a preset normalization algorithm to perform reflectance conversion. The specific physical conversion logic follows the calculation formula: In the above calculation formula, the physical meaning of each letter symbol is as follows: R: represents the standard reflectance data obtained after correction, with a value range of 0 to 1, dimensionless; I: represents the digital brightness value of the original hyperspectral image acquired in real time; B: represents the system reference dark field data, i.e., dark current noise value, which is pre-acquired and stored in step S103; W: represents the reference reflectance data of the standard white board in the full band, which is pre-acquired and stored in step S103.
[0026] Through this calculation, the interference caused by the uneven radiation intensity of the light source in different wavelength bands is effectively eliminated, thereby ensuring the stability of the subsequently extracted spectral features and their comparability between different sorting batches.
[0027] Step S105: Perform polarization filtering compensation processing on vegetables with high reflectivity. For vegetable varieties with high reflectivity due to the presence of a natural wax layer or other factors on the surface, such as bell peppers or eggplants, an additional polarization filtering step is introduced in the correction preprocessing.
[0028] The specific operation involves selectively filtering out bright spots caused by specular reflection by adjusting the angle of the polarizer in the optical path. This process effectively avoids saturation distortion of hyperspectral data in local bright areas on the vegetable surface, ensuring the integrity and validity of the reflectance data R across the entire analysis area.
[0029] In summary, step S1 obtains high-quality reflectance data characterizing the appearance and internal biochemical information of vegetables. This data cube not only eliminates errors caused by equipment noise and ambient light variations but also compensates for special optical surfaces, thus providing a data foundation for the subsequent steps S2 (accurate extraction of spatial feature matrices and spectral feature vectors), S3 (quality assessment using a deep fusion recognition model), and S4 (achieving high-precision sorting).
[0030] Furthermore, for step S2, the region of interest of the vegetable is extracted from the corrected hyperspectral data cube, and a spatial feature matrix describing its external morphology and a spectral feature vector reflecting its internal biochemical properties are constructed respectively.
[0031] After obtaining the standard reflectance data cube in step S1, the system immediately enters the feature engineering stage. The primary goal of this stage is to separate the vegetable target from the conveyor belt background, and on this basis, to transform the original high-dimensional hyperspectral data into two sets of compact and discriminative numerical descriptors for subsequent analysis by the deep fusion recognition model. The specific implementation process of step S2 is subdivided into the following sub-steps.
[0032] Step S201: Select a feature band image to enhance the contrast between the target and the background. From the corrected 176 band data, the system selects a single-band image with a wavelength of 800 nanometers as the input source for background removal. This band was chosen because at this wavelength, there is a significant difference between the reflectivity of most vegetable tissues and the reflectivity of the rubber material used in the conveyor belt, resulting in peak contrast between the target and the background. Utilizing this characteristic, subsequent segmentation algorithms can be performed on a clearer image, thereby improving the accuracy of contour extraction.
[0033] Step S202: An initial binarized mask is generated using an adaptive thresholding algorithm. For the 800 nm band image selected in step S201, the system calculates a dynamic adaptive threshold using the maximum inter-class variance method. This algorithm searches for a segmentation threshold that maximizes the variance between the foreground pixel class and the background pixel class by traversing the image's grayscale histogram. The image is then binarized using this threshold, assigning a value of 0 to pixels in the conveyor belt region and a value of 1 to pixels in the vegetable target region, thus obtaining a preliminary binarized mask image.
[0034] Step S203 involves applying morphological processing operators to refine the binarized mask. Since the vegetable surface may have localized shadows or highlight spots, the initial mask generated in step S202 often contains tiny internal voids and sporadic noise points from the conveyor belt edge. Therefore, the system continuously performs the following two morphological operations.
[0035] First, a closing operation is performed. Specifically, the mask is first expanded to fill the small holes inside the vegetable outline caused by texture or spots. Then, the expanded result is eroded to restore the original boundary size of the target.
[0036] Next, an opening operation is performed. Specifically, the mask is first eroded to remove scattered, isolated noise points in the conveyor belt area, and then the eroded result is expanded to smooth the edge contour of the vegetable target.
[0037] After a combination of closing and opening operations, the system finally obtains a connected, complete, and smooth vegetable target outline. The region enclosed by this outline is the region of interest for all subsequent analyses in this method.
[0038] Step S204: Based on the region of interest (ROI), spatial geometric parameters and texture parameters are extracted to construct a spatial feature matrix. After obtaining an accurate ROI mask, the system initiates feature reconstruction tasks in both the spatial and spectral dimensions in parallel. In the spatial dimension, the system performs the following calculation process for the set of vegetable pixels identified by the mask.
[0039] First, calculate the minimum bounding rectangle of the region of interest, and extract the length and width parameters from this rectangle as the basic geometric quantities describing the size specifications of the vegetables.
[0040] Secondly, the centroid coordinates of the region of interest are calculated, which will also be used to determine the center of gravity of the mechanical forces during sorting.
[0041] Next, calculate the eccentricity value of the region of interest. Eccentricity reflects the degree to which the shape of the vegetable deviates from the standard circle and can be used to assess its appearance regularity.
[0042] Meanwhile, the system uses a gray-level co-occurrence matrix algorithm to quantify the texture features of the vegetable surface. Specifically extracted parameters include contrast, which describes the depth of the grooves in the epidermal texture; energy, which describes the uniformity of the gray-level distribution of the texture; and correlation, which describes the extension pattern of the texture in a specific direction. These three texture parameters can effectively reflect whether there are wrinkles, scratches, or abnormal roughness on the epidermis.
[0043] The eight geometric and texture parameters calculated above—length, width, centroid x-coordinate, centroid y-coordinate, eccentricity, contrast, energy, and correlation—are normalized and arranged in a predetermined order to form a 16-dimensional spatial feature matrix. This matrix will serve as a standardized input vector describing the external physical morphology and surface integrity of the vegetable.
[0044] Step S205: Extract the average spectral curve based on the region of interest and perform derivative transformation to enhance signal features. In the spectral dimension, the system traverses all pixels in the region of interest and calculates the arithmetic mean of the reflectance of all pixels in each band along the direction of the spectral axis λ, thereby obtaining an average spectral curve representing the overall spectral response of the vegetable.
[0045] To suppress potential baseline drift in the spectral curve and enhance the weak absorption characteristics caused by minute differences in material composition, the system performs a second-order derivative mathematical transformation on the average spectral curve. This second-order derivative processing effectively amplifies the variation in the positions of inflection points and shoulder peaks in the original spectral curve, making previously difficult-to-discern spectral details more prominent.
[0046] Step S206: Select physiological and biochemical sensitive bands and calculate the slope of change to construct the original high-dimensional spectral feature vector. Based on the plant physiology and biochemistry principles of vegetable quality detection, the system selects three types of sensitive band regions with clear indicative significance from the spectral curves processed by the second derivative.
[0047] The first category is the freshness indicator band, which selects the spectral response value near 760 nm. This band corresponds to the absorption peak of the stretching vibration of the OH bond of water molecules, and its reflectance is closely related to the water content of vegetable tissue.
[0048] The second category is the maturity indicator band, which selects the spectral response value near 680 nm. This band corresponds to the strong absorption characteristics of chlorophyll, and its reflectance changes can reflect the ripening stage or color change of vegetables.
[0049] The third category is the sugar indicator band, which selects the spectral response value near 920 nm. This band corresponds to the absorption characteristics of CH bonds in carbohydrates, and its reflectance value is somewhat correlated with the content of soluble solids inside vegetables.
[0050] In addition, the system also calculates the slope of the reflectance change of adjacent bands on both sides of the above three bands, and uses the slope value as a supplementary feature to describe the local trend of the spectral curve. These selected band reflectance values and their corresponding slope values are sequentially concatenated to form a high-dimensional spectral feature vector.
[0051] Step S207 involves using principal component analysis to reduce the dimensionality of the high-dimensional spectral feature vector. The spectral feature vector constructed in step S206 has a high dimensionality; directly inputting it into the subsequent deep fusion recognition model would significantly increase the computational load and potentially introduce redundant information. Therefore, the system employs principal component analysis to perform a linear transformation to reduce the dimensionality of this vector.
[0052] The execution process of principal component analysis is as follows: First, calculate the covariance matrix of the spectral eigenvectors of all samples in the training set; then solve for the eigenvalues and corresponding eigenvectors of the covariance matrix; then select the eigenvectors corresponding to the first 5 largest eigenvalues in descending order of eigenvalues to construct a dimension-reduced projection matrix; finally, multiply the original high-dimensional spectral eigenvectors with the projection matrix to obtain the compressed 5-dimensional principal component eigenvectors.
[0053] Through experimental verification, retaining the first 5 principal components can compress the dimension of the original spectral feature vector to one-tenth of its original size while ensuring that the cumulative variance contribution rate reaches more than 99.5%, thereby significantly reducing the computational burden of the model with minimal information loss.
[0054] In summary, step S2 transforms the massive amount of raw hyperspectral data output from step S1 into two sets of highly condensed and physically meaningful numerical features: a 16-dimensional spatial feature matrix and a 5-dimensional spectral feature vector. These two sets of feature data provide precise quantitative descriptions of each vegetable from the dimensions of external morphology and internal biochemical properties, respectively, providing input for the multi-source feature fusion analysis and quality grade determination of the deep fusion recognition model in step S3.
[0055] Furthermore, for step S3, the spatial feature matrix and spectral feature vector generated in step S2 are input into the pre-constructed and trained deep fusion recognition model to output the quality grade determination result of vegetables.
[0056] Step S3 utilizes the multi-source information fusion capability of deep learning networks to deeply couple and analyze the spatial features describing the external morphology of vegetables with the spectral features characterizing their internal biochemical properties, thereby achieving accurate grading of the overall quality of vegetables. The specific implementation process of step S3 is subdivided into the following sub-steps.
[0057] Step S301: Construct the network architecture of the deep fusion recognition model. The deep fusion recognition model used in this embodiment is a multi-branch parallel convolutional neural network architecture, which specifically includes three processing branches with different functions.
[0058] The first branch is the spatial feature extraction branch. This branch takes the 16-dimensional spatial feature matrix output from step S2 as input and uses a residual network-based structure for feature learning. This branch consists of alternately stacked convolutional layers and residual connection modules, where the convolutional kernel size includes both 1×1 and 3×3 specifications.
[0059] 1×1 convolutional kernels are used for cross-channel information integration and dimensionality adjustment, while 3×3 convolutional kernels are used to capture local correlation patterns between different geometric and texture parameters in the spatial feature matrix. The residual connection module establishes skip connections between adjacent layers, directly transferring shallow features to deeper layers, thereby alleviating the gradient decay problem during deep network training.
[0060] This branch can automatically learn subtle visual anomaly patterns on the surface of vegetables through layer-by-layer abstraction, such as dotted spots caused by mold and latent subcutaneous tissue damage caused by squeezing.
[0061] The second branch is the spectral feature processing branch. This branch takes the 5-dimensional principal component spectral feature vector output from step S2 as input and employs a 1-dimensional convolutional neural network structure. The convolution kernel of this branch performs a 1-dimensional sliding convolution along the dimension of the spectral feature vector, calculated as follows: In the above formula: This represents the output activation value of the k-th convolutional kernel at position j; This represents the i-th weight parameter in the k-th convolutional kernel; This represents the element value at the corresponding position in the input spectral feature vector; This represents the bias term of the k-th convolutional kernel; The width of the convolution kernel; For nonlinear activation functions, this embodiment uses a modified linear unit.
[0062] By alternating the execution of multi-layer 1D convolution and pooling operations, this branch can effectively capture the key evolution laws reflecting the vibrational modes of matter molecules in the spectral feature vector, and then analyze the abstract features closely related to the internal biochemical indicators such as vegetable freshness, sugar content, and moisture distribution.
[0063] The third branch is a multidimensional spatial-spectral fusion convolutional layer. This layer is located at the intersection of high-level features extracted by the two branches mentioned above, and uses a three-dimensional convolutional kernel to simultaneously perform convolution operations on the reconstructed spatial-spectral joint feature tensor. The three sliding directions of the three-dimensional convolutional kernel correspond to the spatial feature dimension, the spectral feature dimension, and the cross-modal interaction dimension, respectively. This synchronous sliding operation mechanism enables the model to capture the coupling relationship between spatial morphological anomalies and spectral response shifts.
[0064] For example, a certain area on the surface of a vegetable may show a slight depression in the spatial branch, while the same area may show an abnormal attenuation of the water absorption peak near 760 nm in the spectral branch. The multidimensional convolutional layer can correlate these two types of cross-modal cues, thereby accurately determining that there is tissue collapse inside the vegetable due to water loss or disease, rather than simply surface mechanical indentation.
[0065] After the multidimensional spatial spectrum fusion convolutional layer, the model connects several fully connected layers to further abstract the fused features, and finally generates the quality level classification result through the output layer.
[0066] Step S302: Define the model output layer structure and quality level determination rules. The model output layer uses the normalized exponential function Softmax as the activation function, and its calculation formula is: In the above formula: This represents the probability value that the model predicts the input sample belongs to the kth quality level, and its value ranges from 0 to 1. This represents the linear combination output value of the k-th neuron in the output layer, which is the inactive value passed from the fully connected layer to the output layer. This represents the linear combination output value of the i-th neuron in the output layer; This represents the total number of quality grade categories; in this embodiment, K is set to 4.
[0067] The four quality grades are as follows: Grade 1 is the premium grade, which corresponds to vegetables with no defects in appearance and internal biochemical indicators in the best range; Grade 2 is the qualified grade, which corresponds to vegetables with minor appearance defects but whose internal quality still meets the market standards; Grade 3 is the substandard grade, which corresponds to vegetables with obvious defects in appearance or internal indicators and are only suitable for processing; Grade 4 is the discard grade, which corresponds to vegetables that have rotted, become diseased, or suffered severe mechanical damage.
[0068] Industrial control computers execute sorting decisions based on the probability distribution output by the model. When a certain level corresponds to a probability value... When the probability value is greater than 0.85, the system directly confirms this level as the final judgment result and generates the corresponding sorting instruction. If the probability values of all levels are less than 0.6, it indicates that the model's judgment confidence for the current sample is insufficient. The system marks the vegetable as suspected and guides it to the manual re-inspection channel at the end for secondary confirmation via conveyor belt.
[0069] Step S303: Train the deep fusion recognition model. Before being deployed to the sorting production line, the model needs to be fully trained in an offline environment to obtain stable and reliable quality judgment capabilities. The training process is carried out according to the following procedure.
[0070] Construction of the training dataset. A total of 50,000 vegetable samples were collected, covering different varieties, origins, harvest batches, and quality conditions. The sample varieties included common vegetables such as tomatoes, cucumbers, and potatoes. For each sample, its spatial feature matrix and spectral feature vector were first obtained according to the process described in steps S1 and S2, forming a complete input feature record.
[0071] Subsequently, the true internal quality indicators of the sample were obtained through destructive physicochemical testing methods as supervisory labels. Specifically, a digital saccharimeter was used to measure the soluble solids content of the sample tissue, a fruit firmness tester was used to measure the compressive strength of the sample tissue, an oven-drying and weighing method was used to determine the moisture content of the sample tissue, and trained quality inspectors conducted a comprehensive appearance score of the sample according to industry grading standards.
[0072] Based on the measurement results of the above multiple indicators, each sample is labeled as one of the aforementioned four quality levels. The final training dataset contains 50,000 pairs of input feature-quality level label samples, which are randomly divided into 40,000 training samples and 10,000 validation samples in a 4:1 ratio.
[0073] Loss function design. The goal of model training is to minimize the difference between the model's predicted probability distribution and the true label distribution. In this embodiment, the cross-entropy loss function is chosen as the optimization objective, and its calculation formula is: In the above formula: This represents the average loss value for the current training batch; This indicates the number of samples included in the current training batch; This is a sign function, which takes a value of 1 when the true quality level of the nth sample is the kth level, and a value of 0 otherwise. This represents the predicted probability value of the model for the nth sample belonging to the kth quality level; This is the operation for the natural logarithm.
[0074] Training Process and Optimization Algorithm. The training process employs a mini-batch stochastic gradient descent algorithm for iterative updates of the model weights. In each iteration, 64 samples are randomly selected from the training set to form a mini-batch, and the loss function for that batch is calculated. Regarding the gradient of the weights in each layer, the weight values are adjusted according to the following update rules: In the above formula: This represents the model weight matrix at the t-th iteration; This represents the model weight matrix after the (t+1)th iteration update; This represents the learning rate; in this embodiment, the initial learning rate is set to 0.01. This represents the gradient vector of the loss function under the current weights.
[0075] Learning rate A dynamic decay strategy is employed during training: every 10 complete training epochs, if the loss value on the validation set does not show a significant decrease, the learning rate is decayed to 0.5 times the current value. This helps the model converge more precisely to the vicinity of the global optimum in the later stages of optimization. The training termination condition is set when the loss value on the validation set does not improve for 20 consecutive epochs, or when the total number of training epochs reaches 200.
[0076] Through the training process described above, the deep fusion recognition model can automatically learn complex and nonlinear coupling discrimination rules between spatial and spectral features from a large number of labeled samples. After training, the model can output accurate quality grade probability distributions for unseen test samples, thus providing reliable decision-making basis in real-time sorting scenarios.
[0077] Step S304: Using the trained model, inference and judgment are performed on real-time vegetable samples. In the real-time sorting operation, for each piece of vegetable flowing through the scanning area, after completing the feature extraction in step S2, the system immediately loads the generated spatial feature matrix and spectral feature vector into the trained deep fusion recognition model. The model calculates through forward propagation and outputs the probability distribution values of the vegetable corresponding to the four quality grades within milliseconds. The industrial control computer reads this probability distribution and determines the final quality grade according to the judgment rules described in step S302.
[0078] In summary, step S3 achieves deep coupling of spatial morphological information and material spectral properties at the feature level by employing a deep fusion recognition model with a multi-branch parallel architecture. During model training, the methods for collecting and labeling training data, the specific form of the loss function, and the configuration parameters of the optimization algorithm are clearly defined to ensure sufficient learning of network parameters and effective improvement of generalization ability. The quality level judgment result output in this step will be directly passed to step S4 to guide the calculation of sorting trigger delay and the generation of specific action instructions for the execution mechanism.
[0079] Furthermore, the final step is S4, which calculates the sorting trigger delay based on the quality grade judgment result output in step S3 and the real-time position coordinates of the vegetables on the conveyor belt, and drives the end effector to sort the vegetables to the corresponding collection channel. The specific implementation process is divided into the following sub-steps.
[0080] Step S401: Establish clock synchronization and signal links between the industrial control computer and the field sensors. The industrial control computer establishes real-time data communication links with photoelectric sensors, rotary encoders, and electromagnetic pneumatic valve drive modules distributed along the conveyor belt via high-speed industrial Ethernet protocol. The Ethernet communication cycle is set to 2 milliseconds to ensure that the transmission delay between control commands and sensor feedback signals is compressed to a negligible range.
[0081] When the vegetables to be sorted move to the entrance of the scanning area along the conveyor belt, their front end first cuts off the infrared beam of a through-beam photoelectric sensor. Upon detecting the beam being blocked, the sensor's receiver immediately sends a level-change signal to the hardware interrupt pin of the industrial control computer. Upon receiving this interrupt signal, the industrial control computer's hardware timer synchronously starts, recording the start time of the vegetable entering the sorting control process.
[0082] Meanwhile, the rotary encoder is rigidly connected to the shaft of the conveyor belt drive roller via a coupling, continuously outputting orthogonal encoded signals with a resolution of 4096 pulses per revolution. The industrial control computer decodes and counts these encoded signals to obtain the linear displacement and instantaneous linear velocity values of the conveyor belt surface relative to the scanning reference point in real time.
[0083] Step S402: Obtain the pre-calibrated geometric parameters required for the sorting trigger delay calculation. During the system installation and commissioning phase, the key geometric dimensions on the sorting production line have been precisely calibrated in advance, and the calibration results have been stored as constant parameters in the configuration file of the industrial control computer. The specific calibrated parameters include the following two items.
[0084] The first item is the scanning reference distance, denoted as the fixed physical distance D. This distance is defined as the straight-line distance along the conveyor belt movement direction between the plane containing the scan line of the linear array camera in the hyperspectral imaging device and the plane containing the central axis of the end effector's pneumatic valve array. In this embodiment, the calibrated value of D is 1200 mm.
[0085] The second item is a mapping table of target collection trough numbers corresponding to each quality grade. This mapping table associates the four quality grades defined in step S302 with the index numbers of one or more specific air spray valves in the air spray valve array, thereby determining the location of the collection trough where the vegetables should fall when blown off the conveyor belt. The physical installation positions of the air spray valves corresponding to different collection troughs are slightly different, so the system will fine-tune the fixed physical distance D according to the specific position of the target air spray valve when actually calculating the delay.
[0086] Step S403: Calculate the basic sorting trigger delay based on the real-time linear velocity. The industrial control computer reads the current instantaneous linear velocity of the conveyor belt from the rotary encoder at a frequency of 500Hz, and records it as the real-time velocity value v. The basic sorting trigger delay is denoted as t0, and its calculation follows the physical relationship of uniform linear motion of an object. The calculation formula is: In the above formula: t0 represents the basic sorting trigger delay in milliseconds; D represents the pre-calibrated scanning reference distance in step S402 in millimeters; v represents the instantaneous linear velocity of the conveyor belt fed back by the rotary encoder in real time in millimeters per second.
[0087] It should be noted that the base delay t0 calculated in this step only holds true under ideal conditions, namely, assuming no response lag in the solenoid valve, no airflow transmission time, and the center of mass of the vegetable is exactly located on the center line of the conveyor belt. In actual working conditions, none of the above ideal conditions are met, therefore t0 must be corrected and compensated.
[0088] Step S404: Apply multi-source error compensation correction to the basic sorting trigger delay. To ensure that the air jet can accurately act on the geometric center of the vegetables, the industrial control computer superimposes three compensation correction components on the basic delay t0, resulting in the final execution delay denoted as t. The calculation formula is as follows: The physical meanings and determination methods of each symbol in the above formula are explained below.
[0089] Δt1 is the solenoid valve response hysteresis compensation value. This component is used to compensate for the electromechanical delay time between the issuance of the drive signal from the industrial control computer and the completion of the opening action of the air injection valve core. The value of Δt1 is pre-determined through offline testing. Specifically, a high-speed camera system is used to record the time difference between the rising edge of the drive signal and the appearance of airflow at the air injection nozzle. This test is repeated 100 times, and the average value is stored in the system as a fixed compensation parameter. In this embodiment, the measured value of Δt1 is 7 milliseconds.
[0090] Δt2 is the compensation amount for airflow travel time. This component is used to compensate for the time required for high-pressure air to travel from the outlet of the air jet valve to the surface of the vegetables. The value of Δt2 is related to the vertical distance from the outlet of the air jet valve to the surface of the conveyor belt and the initial injection velocity of the airflow, and is also obtained through offline calibration. In this embodiment, the calibrated value of Δt2 is 3 milliseconds.
[0091] Δt3 represents the compensation amount for the vegetable's center of mass offset. In step S204, the coordinates of the center of mass of the vegetable's region of interest have been calculated. When the lateral coordinates of this center of mass deviate from the conveyor belt centerline, the vegetable's trajectory after being impacted by the air jet will shift laterally, causing it to fall into the collection trough at a position deviating from the expected location. Therefore, the system fine-tunes the triggering time proportionally based on the magnitude of the lateral offset Δx, making the air jet's point of impact closer to the vegetable's actual center of mass, thereby suppressing trajectory scattering caused by the eccentric torque. The formula for calculating Δt3 is: In the above formula: k is the proportionality coefficient, with the unit being milliseconds per millimeter. Its value is determined through on-site calibration tests for vegetables of different weights and sizes. In this embodiment, k is taken as 0.2 milliseconds per millimeter for medium-sized vegetables; Δx is the deviation value between the transverse coordinate of the vegetable's centroid and the center line of the conveyor belt, with the unit being millimeters. Its positive or negative sign indicates the direction of deviation.
[0092] In step S405, a pulse trigger signal is sent to the solenoid valve drive module when the delay arrives. When the accumulated time of the hardware timer started in step S401 reaches the final execution delay t, the industrial control computer immediately sends one pulse signal to the solenoid valve drive module corresponding to the target quality level via industrial Ethernet. The width of this pulse signal can be dynamically programmed according to the weight of the vegetables. In this embodiment, the adjustable range of the pulse width is 50 milliseconds to 150 milliseconds.
[0093] When the solenoid valve drive module receives the rising edge of the pulse signal, it immediately initiates the conduction of the solenoid coil of the air injection valve. The electromagnetic force generated by the coil overcomes the spring force of the return spring, pushing the valve core to move and connecting the high-pressure air path with the air injection port. The high-pressure air released by the air injection valve is preset to 0.6 MPa.
[0094] Step S406: Dynamically adjust the air jet parameters according to the weight characteristics of the vegetables to optimize the sorting effect and reduce damage. In order to prevent sorting failure or vegetable damage caused by excessive or insufficient air jet impact force, the system estimates the approximate weight grade of the vegetables based on the size parameters extracted in step S204, and makes differentiated adjustments to the air jet parameters accordingly.
[0095] For heavier vegetable varieties, such as large potatoes or radishes, the system automatically extends the pulse width to the range of 100 to 150 milliseconds, prolonging the action time of the high-pressure air to provide sufficient impulse to remove the vegetables from the conveyor belt.
[0096] Meanwhile, if the estimated weight of a single vegetable exceeds the rated blowing capacity of a single air jet valve, the system can simultaneously trigger two or three adjacent air jet valves through logic control to form a planar thrust to ensure the smooth movement of large vegetables.
[0097] For lighter vegetable varieties, such as leafy greens or small berries, the system reduces the air jet pressure from 0.6 MPa to the range of 0.3 MPa to 0.4 MPa through a proportional pressure regulating valve while maintaining a pulse width of 50 milliseconds. This reduces the impact force of the airflow and avoids damage to the leaves or abrasions to the fruit skin.
[0098] Step S407: A buffer device is installed in the collection trough to absorb the kinetic energy of the falling vegetables. A 20 mm thick polymer foam buffer pad is laid on the bottom and side walls of the collection trough corresponding to each quality grade. The buffer pad is made of closed-cell polyethylene foam, which has a moderate compression resilience and good impact absorption performance. When the vegetables leave the conveyor belt under the action of air jets and fall into the collection trough along a parabolic trajectory, the buffer pad absorbs the falling kinetic energy of the vegetables through elastic deformation, significantly reducing the risk of mechanical damage caused by the collision between the vegetables and the rigid surface of the trough.
[0099] Step S408 ensures high-speed sorting through a multi-threaded parallel processing mechanism. To achieve an ultra-high processing speed of over 30 pieces per second, the industrial control computer's software architecture adopts a multi-threaded parallel processing design. Hyperspectral data acquisition and reflectance correction are handled by one independent thread; feature extraction in step S2 and model inference in step S3 are completed by another independent thread with the assistance of a graphics processor; and delay calculation and execution control in step S4 are handled by a third high-priority real-time thread. The three threads exchange vegetable identification information, quality grade data, and centroid coordinate data through a shared memory area. Thread synchronization uses a mutex lock mechanism to avoid data read / write conflicts. This parallel architecture compresses the entire processing cycle of a single vegetable from entering the scanning area to completing the sorting action to within 15 milliseconds.
[0100] In summary, step S4, by closely integrating hyperspectral visual perception results with real-time motion control technology, constructs a complete closed-loop execution link from quality assessment to physical sorting. After precise sorting in step S4, vegetables of different quality grades are guided to their corresponding collection channels, thus completing the entire fully automated sorting process.
[0101] On the other hand, the fully automated vegetable sorting system based on computer vision disclosed in this application includes: Conveyor belts are used to transport vegetables to be sorted. A hyperspectral imaging device, installed above the conveyor belt, is used to acquire raw hyperspectral data cubes of vegetables to be sorted within a preset wavelength range. An industrial control computer, connected to a hyperspectral imaging device, is configured as follows: The original hyperspectral data cube is subjected to reflectance correction to obtain a standard reflectance data cube; The region of interest (ROI) of the vegetables to be sorted is extracted from the standard reflectance data cube, and a spatial feature matrix describing the external morphology and a spectral feature vector describing the internal biochemical properties are constructed based on the ROI. The spatial feature matrix and spectral feature vector are input into a pre-built deep fusion recognition model. The model contains a multi-dimensional spatial-spectral fusion convolutional layer, which is used to simultaneously capture cross-modal coupled features and output the quality grade judgment result of the vegetables to be sorted. Based on the quality grade determination results and the real-time position coordinates of the vegetables to be sorted on the conveyor belt, the sorting trigger delay, which includes multiple compensation and correction amounts, is calculated. The end effector, connected to the industrial control computer, is used to sort the vegetables to be sorted into the corresponding collection channels according to the instructions of the industrial control computer when the sorting trigger delay arrives.
[0102] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, the embodiments should be regarded as exemplary and non-limiting in all respects.
[0103] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A fully automated vegetable sorting method based on computer vision, characterized in that, Includes the following steps: Step S1: Obtain the original hyperspectral data cube of the vegetables to be sorted within a preset wavelength range using a hyperspectral imaging device, and perform reflectance correction processing on the original hyperspectral data cube to obtain a standard reflectance data cube. Step S2: Extract the region of interest of the vegetables to be sorted from the standard reflectance data cube, and construct a spatial feature matrix describing the external morphology and a spectral feature vector describing the internal biochemical properties based on the region of interest; Step S3: Input the spatial feature matrix and the spectral feature vector into a pre-constructed deep fusion recognition model. The model includes a multi-dimensional spatial-spectral fusion convolutional layer, which is used to simultaneously capture the cross-modal coupling features between the spatial feature matrix and the spectral feature vector, and output the quality grade determination result of the vegetables to be sorted based on the cross-modal coupling features. Step S4: Based on the quality grade determination result and the real-time position coordinates of the vegetables to be sorted on the conveyor belt, calculate the sorting trigger delay including multiple compensation corrections, and drive the end effector to sort the vegetables to be sorted into the corresponding collection channel when the delay arrives.
2. The method according to claim 1, characterized in that, The reflectivity correction process in step S1 specifically includes: Using pre-calibrated standard whiteboard reference data W and dark field reference data B, the real-time acquired raw image data I is normalized to obtain standard reflectance data R; where W is the standard whiteboard reference data, B is the dark field reference data, I is the raw image data, and R is the standard reflectance data, calculated using the following formula: ; Furthermore, for vegetables with highly reflective surfaces, polarization filtering compensation is performed before or after the normalization calculation to filter out bright spots caused by specular reflection by adjusting the angle of the polarizer in the optical path.
3. The method according to claim 1, characterized in that, The construction of the spatial feature matrix in step S2 specifically includes: Calculate the minimum bounding rectangle of the region of interest to extract the length and width parameters; Calculate the centroid coordinates and eccentricity of the region of interest; The contrast, energy, and correlation texture parameters of the region of interest are extracted using the gray-level co-occurrence matrix algorithm. The spatial feature matrix is formed by arranging the above parameters, including length, width, centroid x-coordinate, centroid y-coordinate, eccentricity, contrast, energy, and correlation, after numerical normalization.
4. The method according to claim 1, characterized in that, The construction of the spectral feature vector in step S2 specifically includes: Calculate the arithmetic mean of the reflectance of all pixels in the region of interest under each band to obtain the average spectral curve; A second-order derivative mathematical transformation is performed on the average spectral curve to enhance signal characteristics; From the spectral curve after second derivative transformation, spectral response values at wavelengths of 760 nm, 680 nm, and 920 nm are selected in a directional manner, and the slope of the change in reflectance of adjacent bands on both sides of each wavelength is calculated. The selected reflectance values and slope values are combined to form the original high-dimensional spectral feature vector. Principal component analysis (PCA) is used to reduce the dimensionality of the original high-dimensional spectral feature vector, retaining the first few principal components whose cumulative variance contribution rate reaches a preset threshold, thus obtaining the dimensionality-reduced spectral feature vector.
5. The method according to claim 1, characterized in that, The deep fusion recognition model is a multi-branch parallel convolutional neural network architecture, specifically including: A spatial feature extraction branch, whose input is the spatial feature matrix, uses a residual network-based structure for feature learning to extract the external morphological abnormal features of vegetables; A spectral feature processing branch, whose input is the spectral feature vector, adopts a one-dimensional convolutional neural network structure to capture the spectral evolution pattern reflecting the internal biochemical indicators of vegetables. The multidimensional spatial-spectral fusion convolutional layer takes as input the high-level features output by the spatial feature extraction branch and the spectral feature processing branch. It uses a three-dimensional convolution kernel to perform convolution operations synchronously on the reconstructed spatial-spectral joint feature tensor to capture the coupling relationship between spatial morphological anomalies and spectral response shifts, and finally outputs the classification result of quality level.
6. The method according to claim 1, characterized in that, The calculation of the sorting trigger delay in step S4 specifically includes: Obtain the pre-calibrated scanning reference distance D and the instantaneous linear velocity v of the conveyor belt fed back in real time by the rotary encoder, and calculate the basic sorting trigger delay. ; Obtain and superimpose the solenoid valve response hysteresis compensation amount Airflow transmission time compensation And vegetable centroid offset compensation amount The final execution delay is obtained. ; The centroid offset compensation amount Based on the deviation value between the lateral component of the centroid position coordinates in the spatial feature matrix and the centerline of the conveyor belt Calculated proportionally, the formula is: , where k is the proportionality coefficient.
7. The method according to claim 6, characterized in that, The step S4 of driving the end effector also includes: The weight grade of the vegetables to be sorted is estimated based on the size parameters in the spatial feature matrix; Based on the weight level, the execution parameters of the air spray are dynamically adjusted; for vegetables with an estimated weight greater than a preset threshold, the air spray pulse width is extended or multiple adjacent air spray valves are triggered to perform combined spraying; for vegetables with an estimated weight less than a preset threshold, the air spray pressure is reduced.
8. The method according to claim 1, characterized in that, In step S3, the quality level judgment result output by the deep fusion recognition model is a probability distribution, and the method further includes: When the probability value corresponding to a certain quality level is greater than the first confidence threshold, the level is directly confirmed as the final judgment result. When the probability values of all quality grades are lower than the second confidence threshold, the vegetables to be sorted are marked as suspected and guided to the manual re-inspection channel.
9. The method according to claim 1, characterized in that, In step S4, the sorting process is executed through a multi-threaded parallel processing mechanism: the first thread is responsible for the acquisition of hyperspectral data and reflectance correction, the second thread is responsible for feature extraction and model inference, and the third thread is responsible for delay calculation and execution control. The threads exchange data through shared memory.
10. A fully automated vegetable sorting system based on computer vision, used to execute the method according to any one of claims 1 to 9, characterized in that, include: Conveyor belts are used to transport vegetables to be sorted. A hyperspectral imaging device is installed above the conveyor belt to acquire the original hyperspectral data cube of the vegetables to be sorted within a preset wavelength range. An industrial control computer, connected to the hyperspectral imaging device, is configured to: The original hyperspectral data cube is subjected to reflectance correction processing to obtain a standard reflectance data cube; The region of interest (ROI) of the vegetables to be sorted is extracted from the standard reflectance data cube, and a spatial feature matrix describing the external morphology and a spectral feature vector describing the internal biochemical properties are constructed based on the ROI. The spatial feature matrix and the spectral feature vector are input into a pre-built deep fusion recognition model. The model contains a multi-dimensional spatial-spectral fusion convolutional layer, which is used to simultaneously capture cross-modal coupled features and output the quality grade determination result of the vegetables to be sorted. Based on the quality grade determination result and the real-time position coordinates of the vegetables to be sorted on the conveyor belt, the sorting trigger delay, which includes multiple compensation and correction amounts, is calculated. An end effector, connected to the industrial control computer, is used to sort the vegetables to be sorted into the corresponding collection channels according to the instructions of the industrial control computer when the sorting trigger delay arrives.