A Machine Learning-Based Laser Marking Method and System for Fruits and Vegetables

CN121360895BActive Publication Date: 2026-09-18ZHEJIANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511486502.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-09-18
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

[0005]上述现有技术方案在应用于水果激光标记时,暴露出诸多固有缺陷:

Benefits of technology

1、本发明旨在通过建立预测模型,取代传统依赖大量重复实验的“经验试凑法”,从而大幅度缩短针对水果的激光标记工艺开发周期,显著降低材料、设备和人力的消耗。由此训练出的模型,将能够基于一组包含“设备参数+工艺参数”的输入,来预测不同设备在不同工艺参数下的色差效果,极大增强了模型的通用性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121360895B_ABST
    Figure CN121360895B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of post-harvest logistics technology for fruits and vegetables, and discloses a laser marking method and system for fruits and vegetables based on machine learning. The method includes the following steps: (1) identification of the skin data of fruits and vegetables; (2) selection of the marking area; (3) laser marking energy output based on machine learning; and (4) data correction and model usage. This invention aims to significantly shorten the development cycle of laser marking technology for fruits and vegetables and significantly reduce the consumption of materials, equipment, and manpower by establishing a predictive model to replace the traditional "trial and error" method that relies on a large number of repeated experiments. The model trained in this way can predict the color difference effect of different equipment under different process parameters based on a set of inputs including "equipment parameters + process parameters", which greatly enhances the versatility and industrial deployment value of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of postharvest logistics technology for fruits and vegetables, and in particular, it is a method and system for laser marking of fruits and vegetables based on machine learning. Background Technology

[0002] Laser marking technology, as an advanced labeling method to replace traditional self-adhesive stickers, is increasingly being used on the surfaces of horticultural products such as fruits and vegetables. This technology uses a low-energy laser beam (such as a CO2 laser) to etch permanent marks composed of letters, numbers, or graphics onto the surface of fruits and vegetables, meeting the growing market demands for product traceability, brand identification, and anti-counterfeiting.

[0003] The core of this technology lies in the interaction between laser and fruit peel tissue. Laser energy induces physical ablation, tissue carbonization, and chemical reactions in epidermal cells, resulting in visually recognizable color changes on the product surface, i.e., "color difference." However, the final color difference effect is influenced by a multitude of factors, including not only equipment process parameters such as scanning speed, laser power, pulse frequency, and pulse width, but also the biological characteristics of the product itself, such as fruit variety, surface curvature, peel water content, and tissue structure. The color difference value reflects the color difference before and after laser marking, and it is a reflection of contrast. Its magnitude can reflect the clarity of the marking pattern and the degree of damage to the fruit and vegetable surface caused by the laser. The complex nonlinear relationships between these factors make accurately controlling and predicting the color difference value after laser marking to ensure the quality and consistency of product marking a major technical challenge in this field.

[0004] like Figure 1 As shown, currently, when developing laser marking processes for specific fruits (such as peaches), the industry generally uses an "experimental trial-and-error method" or an "experimental verification method" to determine the optimal process parameters. The specific technical solution involves process engineers setting a preliminary set of laser process parameter combinations for mainstream peach varieties based on personal experience or equipment recommendations. After marking a small area of ​​selected peaches with these parameters, the effect is tested to check if it meets the requirements. If it does not meet the requirements, the parameters need to be reselected and the test repeated.

[0005] The aforementioned existing technical solutions reveal several inherent defects when applied to laser marking of fruits: 1. Long R&D cycle and high cost: This method is essentially a process of repeated trial and error, requiring numerous experiments to determine suitable parameters. Each parameter adjustment means consuming new fruit samples, equipment operating time, and labor costs. This problem is particularly prominent when developing labeling processes for new varieties or fruits at different ripeness levels, resulting in low efficiency and high costs.

[0006] 2. Difficulty in achieving precise and optimal control: As biological materials, fruits vary in size, curvature, and peel thickness even among individuals (within the same batch), introducing uncertainty into experimental results. Furthermore, due to the complex coupling effects between laser parameters, it is difficult to find the globally optimal parameter combination through a limited number of experiments; often, only a "usable" rather than "optimal" process solution can be obtained.

[0007] 3. Unpredictable results and poor process consistency: Before actual marking, it is impossible to scientifically predict the final color difference that a specific combination of parameters will produce. The determination of the process heavily relies on the personal experience of engineers, which leads to fluctuations in marking effects between different operators or different batches of products, making it difficult to guarantee quality consistency in large-scale production.

[0008] A search revealed that the application of machine learning to predict the relationship between "process parameters and material properties" has been explored in other fields. For example, Chinese invention patent CN118430701B discloses a method for predicting the phase transformation temperature of NiTi alloys using selective laser melting (SLM) based on machine learning. This prior art collects parameters such as alloy composition and preparation process as input features to train various machine learning models (such as GRNN, KNN, RF, etc.) to predict the phase transformation temperature of NiTi alloys. The GRNN model ultimately selected as the optimal choice has a coefficient of determination R² of 0.97. Summary of the Invention

[0009] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for laser marking of fruits and vegetables based on machine learning.

[0010] The technical solution adopted by this invention to solve the problem is: A machine learning-based laser marking method for fruits and vegetables, the method comprising the following steps: (1) Fruit and vegetable skin data recognition: The camera dynamically captures the image in the working area of ​​the scanning laser and transmits the captured image information to the computer. The computer calculates the dynamic information of the image information to identify the skin color and curvature of the fruit and vegetable to be marked. The computer also displays the captured image information on the computer screen. (2) Selection of the marking area: The calculator determines the best marking area for fruits and vegetables based on the data identified in step (1) and displays it on the screen. Then, the location information of the marking area is transmitted to the fruit and vegetable placement platform so that the fruit and vegetable placement platform moves the fruits and vegetables to be laser-marked on it to the direct under the laser, making it easier to perform the marking operation. (3) Laser marking energy output based on machine learning: Data of the skin of standard samples of fruits and vegetables after marking in actual production is collected into the computer. The parameters used by the laser and scanning galvanometer are input into the computer. Then, the laser and scanning galvanometer are operated to laser mark the skin of fruits and vegetables in the marking area in step (2). Multiple fruits and vegetables are laser marked with different output times. At the same time, the laser output frequency and output time of each fruit and vegetable during laser marking are recorded. The camera collects the color values ​​before and after marking. The collected data and standard samples are used to calculate the color difference based on the following formula: ; in, , , The color value before laser processing. , , The color value is 24 hours after laser treatment, ΔE is the total color difference, where L represents the brightness of the fruit surface, a represents the position of the color on the red-green axis, reflecting the distribution of red and green components, b represents the position of the color on the yellow-green axis, reflecting the distribution of yellow and green components, the subscript initial indicates the value before treatment, and final indicates the value measured 24 hours after treatment; By importing the color difference calculation data of multiple fruits and vegetables and the color difference calculation data of standard samples, along with the parameters of the laser, the parameters of the scanning galvanometer, and the marking time, into various machine learning models, the original input features are the three core adjustable process parameters of the laser and the scanning galvanometer: scanning speed mm / s, repetition frequency KHz, and pulse width μs. Machine learning is performed in this way, and the various models are compared to select the best initial model. (4) Data correction and model use: The initial model after machine learning in step (3) is corrected to reduce the impact of outliers on the model and make it closer to the real value. The corrected model is directly imported into the laser marking machine used in production. The model can then use the data of scanning fruits and vegetables by the camera that is matched with the laser used in production, combined with the target color difference ΔE set by the operator, to solve the best output parameters of the laser and scanning galvanometer through optimization algorithms (such as grid search, genetic algorithm, etc.) and guide the laser and scanning galvanometer to perform laser marking on fruits and vegetables.

[0011] Furthermore, the color recognition method in step (1) is as follows: S1 uses a hyperspectral camera to capture and scan the surface information of fruits and vegetables on the fruit and vegetable placement platform. The hyperspectral camera can identify extremely subtle color differences on the surface of fruits and vegetables, and can also identify defects on the surface of fruits and vegetables. S2 uses a trained deep learning segmentation algorithm, U-Net network model or Mask R-CNN instance segmentation model, to analyze the captured images.

[0012] Furthermore, the method for identifying the marking area in step (2) is as follows: A vision-guided robot is used for operation, specifically: S1 collects information about the surface of fruits and vegetables through a camera. The computer determines the position and orientation of the best marking area in the image. Based on this position and orientation information, it guides a multi-axis robotic arm with flexible grippers to move over and accurately grab the fruit or vegetable. The S2 robot's robotic arm moves the fruits and vegetables directly under the laser, while rotating its wrist joint and adjusting the posture of the fruits and vegetables in conjunction with the image information captured by the camera. This ensures that the surface of the fruits and vegetables facing the laser lens is an undamaged marking area, avoiding any impact on the laser marking process. After S3 is aligned in both posture and position, the robotic arm sends a signal to trigger the marking machine to complete the operation.

[0013] Furthermore, in step (3), to capture the interaction between the parameters of several markers, new interaction features are constructed based on the three original input features, including: scanning speed. Pulse frequency, scan speed Pulse width, pulse frequency Pulse width; The dataset above was examined to remove samples with missing values. Outlier detection and correction were performed using the following methods: 1) For the data from three parallel replicate experiments under each set of process parameters, the statistically recognized Grubbs' Test was used to identify and remove outliers when processing the data from each set of parallel replicate experiments under each set of process parameters. First, the arithmetic mean of each set of measurements, i.e., ΔE, was calculated. The standard deviation s is given, where the number of parallel experiments n ≥ 3; subsequently, suspicious data points with the largest deviation from the group mean are identified, and their G-statistics are calculated, defined as... , where x i Data points that are considered suspicious; 2) Compare the calculated G value with the Grubbs critical value G corresponding to the sample size n at a preset confidence level of 95%. crit For comparison, when n=3, G crit The threshold is 1.155. If the calculated value of G is greater than this threshold, the data point is determined to be an outlier and removed from the data set; otherwise, all data is retained. 3) Finally, based on the remaining valid data after removing any identified outliers, the arithmetic mean is recalculated and this result is used as the final representative measurement for this set of process parameters. The data in the above correction method is standardized using the StandardScaler method. This process is applied to all input features, including both original and interactive features, to eliminate the dimensional differences between different features, enabling the model to treat each feature more fairly. The standardized formula is as follows: ; Where z is the standardized new value, and x is the original data value. σ is the mean of all the original data for this feature, and σ is the standard deviation of all the original data for this feature.

[0014] Furthermore, the method for comparing and selecting multiple machine learning models in step (3) includes the following steps: S1. Perform stratified sampling on the corrected dataset and divide it into training and test sets in a ratio of 85%:15%. S2. On the training set, train multiple machine learning regression models in parallel, including multinomial regression, random forest regressor, gradient boosting regressor, support vector regression (SVR), XGBoost regressor, or other models. S3. For each model, the GridSearchCV method with 5-fold cross-validation is used to systematically find the optimal combination of hyperparameters with the coefficient of determination R² as the evaluation index. The final performance evaluation of all optimized models is performed using an independent test set, and the model with the highest R² score on the test set is selected as the final color difference prediction model. Among them, R 2 The formula is as follows: ; Among them, y i Let i be the true value of the i-th sample. Let be the predicted value for the i-th sample. is the average of the true values, and n is the number of samples.

[0015] Furthermore, the data correction method in step (4) is as follows: traverse each parameter setting and check its corresponding multiple parallel measurement values; if the G value of a certain value is found to be greater than 1.155, it is determined to be an outlier and replaced with the average value of the other normal values ​​in the same group.

[0016] Furthermore, the specific steps are as follows: Step 1: Constructing a laser-marked color difference experimental dataset 1) Determine the input features and output target: The input features are three core adjustable process parameters of the laser and scanning galvanometer: scanning speed (mm / s), repetition frequency (KHz), and pulse width (μs). The output target is the total color difference ΔE generated on the sample surface 24 hours after laser treatment; the total color difference ΔE is calculated according to the following formula, where , , The color value before laser processing. , , Color values ​​24 hours after laser treatment: ; 2) Perform the experiment and collect data: Orthogonal experimental design method is used to systematically combine different input feature parameters; The scanning method uses a line spacing of 0.05mm, a spot diameter of 0.05mm at the focal plane, one scan, a square area with a side length of 1.8cm, and a serpentine scanning path. The self-comparison method was adopted. Before marking, the initial color value of the area to be processed was measured with a colorimeter. The final color value of the same area was measured again 24 hours after marking was completed. To ensure data reliability, each parameter group was tested in three parallel replicates. The original dataset is formed by combining the recorded input features and their corresponding output targets, which are the average values ​​of the three repeated experiments. Step 2: Feature Engineering and Data Preprocessing 1) Feature Engineering: To capture the interactive effects between various parameters, new interactive features are constructed based on the three original input features, including: scanning speed. Pulse frequency, scan speed Pulse width, pulse frequency The pulse width is used to obtain the feature list of the expanded dataset. 2) Data Cleaning: The expanded dataset is inspected to remove samples with missing values, and outlier detection and correction are performed. The correction method is as follows: For each set of process parameters, the mean and standard deviation of three parallel repeated experiments are calculated; each parameter setting is iterated through, and its corresponding three parallel measurements are checked; if the G value of a certain value is found to be greater than 1.155, it is identified as an outlier and replaced with the average of the other normal values ​​in the same group; finally, a cleaned and fused robust measurement result is obtained for each parameter setting. 3) Data Standardization: A standardization method is used to standardize all input features of the expanded dataset, including the original features (scan speed, scan frequency, and pulse width) and the interactive feature (scan speed). Pulse frequency, scan speed Pulse width, pulse frequency Pulse width is processed to eliminate the dimensional differences between different features, enabling the model to treat each feature more fairly. For each feature to be standardized, the standardization method formula is: , where z is the new value of the feature to be standardized after standardization, and x is the original data value of the feature. σ is the mean of all the original data for this feature, and σ is the standard deviation of all the original data for this feature. Step 3: Training, evaluating, and selecting the best machine learning model 1) Data set partitioning: The data after feature engineering and data preprocessing is first binned, and then stratified sampling is performed to divide it into training set and test set in a ratio of 85%:15% to ensure that the distribution of color difference values ​​in the training set and test set is consistent with the distribution of color difference values ​​in the entire dataset. 2) Multi-model training: Train various machine learning regression models on the training set, including multinomial regression, random forest regressor, gradient boosting regressor, support vector regression, XGBoost regressor, or other models; 3) Hyperparameter optimization: For each model, the GridSearchCV method with 5-fold cross-validation is used to systematically find the optimal combination of hyperparameters with the coefficient of determination R² as the evaluation index. The formula for the R² coefficient of determination is as follows:

[0017] Among them, y i Let i be the true value of the i-th sample. Let be the predicted value for the i-th sample. y is the average of the true values ​​of y, and n is the sample size; 4) Model Evaluation and Selection: The final performance of all optimized models was evaluated using an independent test set; the model with the highest R² score on the test set was selected as the final color difference prediction model. The results showed that the gradient boosting regressor performed optimally. Step 4: Application of color difference prediction based on the optimal model The trained optimal model, i.e., the gradient boosting regressor, is then deployed and labeled.

[0018] Furthermore, the laser marking system is a 5W 355nm ultraviolet laser marking machine.

[0019] A machine learning-based laser marking system for fruits and vegetables, utilizing the method described above, includes a camera, a laser, a fruit and vegetable placement platform, a computer, a robot, and a scanning galvanometer. The scanning galvanometer, camera, laser, and robot are all connected to the computer via electrical signals. The scanning galvanometer is connected to the laser, and the laser is connected to the camera. The side of the fruit and vegetable placement platform is connected to the robot. The laser, scanning galvanometer, and camera are spaced apart above the fruit and vegetable placement platform. The robot can place fruits and vegetables on the platform, the camera can capture image information of the fruit and vegetable surface on the platform, the laser can perform laser marking operations on the surface of the fruits and vegetables on the platform through the scanning galvanometer, and the computer can collect image information from the camera via electrical signals and control the working state of the laser, scanning galvanometer, and robot.

[0020] The advantages and positive effects of this invention are as follows: 1. This invention aims to significantly shorten the development cycle of laser marking technology for fruits and substantially reduce the consumption of materials, equipment, and manpower by establishing a predictive model to replace the traditional "trial and error" method that relies on a large number of repeated experiments. The model trained in this way can predict the color difference effect of different equipment under different process parameters based on a set of inputs including "equipment parameters + process parameters", which greatly enhances the versatility of the model.

[0021] 2. This invention introduces machine learning algorithms to deeply analyze and learn the complex nonlinear coupling relationship between laser process parameters (scanning speed, pulse frequency, pulse width, etc.) and the final color difference value. This aims to accurately predict the color difference effect produced by a specific parameter combination before actual processing, thereby achieving precise control over marking quality. By collecting a comprehensive large dataset containing various fruits and laser parameters, a more powerful general prediction model can be trained. This model can understand the different response characteristics of different fruits to laser action, thus enabling the prediction of color difference based on the input of "characteristics of fruit A" + "process parameters," truly solving the core problem in the background technology where process development is difficult due to differences in the biological characteristics of products.

[0022] 3. The purpose of this invention is to provide a data-driven, repeatable scientific model to replace the current debugging process, which heavily relies on the personal experience of process engineers. This not only improves the scientific rigor and accuracy of process parameter settings but also ensures a high degree of consistency and stability in marking effects across different batches and operators in large-scale production.

[0023] 4. This invention transforms the traditional "trial and error" process, which requires extensive physical experiments, into rapid computer simulation prediction through a constructed predictive model. Operators only need to input parameters to obtain prediction results, eliminating the need for actual peach samples and laser equipment operation time. This significantly shortens the development and optimization cycle of process parameters and substantially reduces R&D costs.

[0024] 5. Comparative analysis with existing technologies: The application areas and technical challenges are completely different: the existing technology is applied to the field of metal additive manufacturing, dealing with industrial alloy materials that have relatively stable and uniform physicochemical properties. This invention, however, is applied to the laser processing of fruits and vegetables, dealing with "living" biological materials. Fruits and vegetables, as heterogeneous organisms, exhibit significant individual differences and uncertainties in their water content, curvature, pigment distribution, and tissue structure, making their response mechanism to laser energy far more complex than that of metallic materials. Therefore, successfully applying machine learning to this highly variable subject of fruits and vegetables faces even greater technical challenges.

[0025] The technical solutions differ in their targeted approach and innovation: To overcome the challenge of high variability in biological materials, this invention proposes a series of targeted technical designs. For example, this invention constructs a dataset through a systematic orthogonal experimental design and employs a local outlier correction method based on parallel repeated experiments to ensure data quality, which is crucial for handling the inherent differences in biological samples. In contrast, the data preprocessing in patent CN118430701B mainly addresses the imputation of missing values ​​commonly found in literature data, without addressing the specific challenge of handling the inherent differences in repeated experiments of biological samples.

[0026] Superior Technical Performance and Predictive Accuracy: Despite facing more complex technical challenges, the prediction model constructed in this invention still demonstrates superior performance. The optimal model in patent CN118430701B has an R² value of 0.97. In contrast, this invention, through in-depth data processing, construction of interactive features, and rigorous model optimization, ultimately obtains a gradient boosting regression model with a higher coefficient of determination (R²) on independent test sets. 2 The R value is as high as 0.979. This higher R value... 2 The values ​​indicate that the methodology proposed in this invention can more accurately capture the mapping relationship between laser parameters and the complex visual effect of color difference in fruit and vegetable peels.

[0027] In summary, while there are precedents for using machine learning to predict process outcomes in fields such as metal processing, this invention is not a simple technology transfer. This invention is the first to successfully apply this technological concept to highly unstable fruit and vegetable biomaterials, and proposes a more robust data processing and modeling method tailored to their characteristics. This solves the long-standing problem of "trial and error" in the industry and achieves predictive accuracy superior to existing technologies in related fields. Therefore, this invention possesses outstanding substantive features and significant progress. Attached Figure Description

[0028] Figure 1 Images showing the effect of traditional techniques marking patterns with different parameters; Figure 2 This is a flowchart of the experimental method in Embodiment 1 of the present invention; Figure 3 These are laser scanning images of peaches with different parameters in Example 1 of this invention. Figure 4 This is a graph showing the R^2 performance of cross-validation for some different models in this invention; Figure 5 This is a graph showing the R^2 performance of cross-validation for another part of the different models in this invention; Figure 6 This is a graph showing the excellent performance of the gradient boosting regressor in this invention on the evaluation metrics; Figure 7 This is a training and testing result diagram of the gradient boosting regressor in this invention; Figure 8 This is a schematic diagram of the structural connection of the fruit and vegetable laser marking system of the present invention. Detailed Implementation

[0029] The present invention will be further described below with reference to the embodiments. The following embodiments are descriptive and not limiting, and should not be used to limit the scope of protection of the present invention.

[0030] The various experimental operations involved in the specific embodiments are all conventional techniques in the field. For parts not specifically annotated in this document, those skilled in the art can refer to various commonly used reference books, scientific and technological documents or related instructions and manuals prior to the filing date of this invention to carry out the operations.

[0031] A machine learning-based laser marking method for fruits and vegetables, the method comprising the following steps: (1) Fruit and vegetable skin data recognition: The camera dynamically captures the image in the laser working area and transmits the captured and scanned image information to the computer. The computer calculates the dynamic information of the image information to identify the skin color and curvature of the fruit and vegetable to be marked. The computer also displays the captured image information on the computer screen. (2) Selection of the marking area: The computer determines the best marking area for fruits and vegetables based on the data identified in step (1), displays it on the screen, and then transmits the location information of the marking area to the fruit and vegetable placement platform, so that the fruit and vegetable placement platform moves the fruits and vegetables to be laser-marked on it to the direct under the laser, so as to facilitate the marking operation. (3) Laser marking energy output based on machine learning: Data of the skin of standard samples of fruits and vegetables after marking in actual production is collected into the computer. The parameters used by the laser and scanning galvanometer are input into the computer. Then, the laser and scanning galvanometer are operated to laser mark the skin of fruits and vegetables in the marking area in step (2). Multiple fruits and vegetables are laser marked with different output times. At the same time, the laser output frequency and output time of each fruit and vegetable during laser marking are recorded. The camera collects the color values ​​before and after marking. The collected data and standard samples are used to calculate the color difference based on the following formula: ; in, , , The color value before laser processing. , , The color value is 24 hours after laser treatment, ΔE is the total color difference, where L represents the brightness of the fruit surface, a represents the position of the color on the red-green axis, reflecting the distribution of red and green components, b represents the position of the color on the yellow-green axis, reflecting the distribution of yellow and green components, the subscript initial indicates the value before treatment, and final indicates the value measured 24 hours after treatment; By importing the color difference calculation data of multiple fruits and vegetables and the color difference calculation data of standard samples, along with the parameters of the laser, the parameters of the scanning galvanometer, and the marking time, into various machine learning models, the original input features are the three core adjustable process parameters of the laser and the scanning galvanometer: scanning speed (mm / s), repetition frequency (KHz), and pulse width (μs). Machine learning is then performed, and the various models are compared to select the best initial model. (4) Data correction and model use: The initial model after machine learning in step (3) is corrected to reduce the impact of outliers on the model and make it closer to the real value. The corrected model is directly imported into the laser marking machine used in production. The model can then use the data of scanning fruits and vegetables by the camera that is matched with the laser used in production, combined with the target color difference ΔE set by the operator, to solve the best output parameters of the laser and scanning galvanometer through optimization algorithms (such as grid search, genetic algorithm, etc.) and guide the laser and scanning galvanometer to perform laser marking on fruits and vegetables.

[0032] A machine learning-based laser marking system for fruits and vegetables, utilizing the methods described above, such as... Figure 8 As shown, the system includes a camera 1, a laser 2, a fruit and vegetable placement platform 3, a computer 4, a robot 5, and a scanning galvanometer 6. The scanning galvanometer, camera, laser, and robot are all connected to the computer via electrical signals. The scanning galvanometer is connected to the laser, and the laser is connected to the camera. The side of the fruit and vegetable placement platform is connected to the robot. The laser, scanning galvanometer, and camera are all spaced apart above the fruit and vegetable placement platform. The robot can place fruits and vegetables on the fruit and vegetable placement platform. The camera can collect image information of the surface of the fruits and vegetables on the platform. The laser can perform laser marking operations on the surface of the fruits and vegetables on the platform through the scanning galvanometer. The computer can collect image information from the camera via electrical signals and control the working status of the laser, scanning galvanometer, and robot.

[0033] Preferably, the color recognition method in step (1) is as follows: S1 uses a hyperspectral camera to capture information about the surface of fruits and vegetables on the platform. The hyperspectral camera can identify extremely subtle color differences on the surface of fruits and vegetables, and can also identify defects on the surface of fruits and vegetables. S2 uses a trained deep learning segmentation algorithm, U-Net network model or Mask R-CNN instance segmentation model, to analyze the captured images. These algorithms are existing technologies specifically designed for pixel-level image recognition. They can automatically and accurately identify and segment the contours of fruits and vegetables in the image, thereby completely separating the fruit and vegetable skin data from irrelevant data such as conveyor belts and backgrounds.

[0034] By integrating the above technologies, the accuracy of color recognition and target segmentation can be greatly improved, and it is less affected by changes in ambient lighting, resulting in better overall recognition performance.

[0035] Preferably, the method for identifying the marking area in step (2) is as follows: using a vision-guided robot for operation, specifically: S1 uses a camera to collect information about the surface of fruits and vegetables, and the computer determines the location and orientation of the optimal marking area in the image. Based on this location and orientation information, it guides a multi-axis robotic arm with flexible grippers to move over and precisely grasp the fruit or vegetable. The S2 robot's robotic arm moves the fruits and vegetables directly under the laser, while rotating its wrist joint and adjusting the posture of the fruits and vegetables in conjunction with the image information captured by the camera. This ensures that the surface of the fruits and vegetables facing the laser lens is an undamaged marking area, avoiding any impact on the laser marking process.

[0036] After S3 is aligned in both posture and position, the robotic arm sends a signal to trigger the marking machine to complete the operation.

[0037] This method is highly automated and has precise positioning, making it especially suitable for processing irregularly shaped fruits and vegetables.

[0038] The method described above, in which the computer determines the optimal marking area, can be based on the color continuity judgment in traditional techniques.

[0039] Preferably, in step (3), to capture the interaction between the parameters of several markers, new interaction features are constructed based on three original input features, including: "scanning speed". "Pulse frequency", "scan speed" "Pulse width" and "Pulse frequency" "Pulse width"; The dataset above was examined to remove samples with missing values. Outlier detection and correction were performed using the following methods: 1) For the data from three parallel replicate experiments under each set of process parameters, to ensure the reliability and accuracy of the experimental data, the statistically recognized Grubbs' Test was used to identify and remove outliers when processing the parallel replicate experimental data under each set of process parameters. This method first calculates the arithmetic mean of the measured values ​​(i.e., ΔE) for each set (where the number of parallel experiments n≥ 3). The standard deviation (s) and the mean of the set were then used to determine the suspicious data points. The G-statistic, defined as (s), was then calculated. , where x i (Data points considered suspicious). The calculated G value will be compared with the Grubbs' critical value (G0) corresponding to the sample size n at a pre-set confidence level (95%). crit ) for comparison, when n=3, G critThe threshold is 1.155. If the calculated value of G is greater than this threshold, the data point is considered an outlier and removed from the data set; otherwise, all data is retained. Finally, based on the remaining valid data after removing any identified outliers, the arithmetic mean is recalculated, and this result is used as the final representative measurement value for this set of process parameters. The StandardScaler standardization method is used to process all input features (including original features and interaction features) to eliminate the differences in scale between different features, so that the model can treat each feature more fairly. The standardized formula is as follows: ; Where z is the standardized new value, and x is the original data value. σ is the mean of all the original data for this feature, and σ is the standard deviation of all the original data for this feature.

[0040] Preferably, the method for comparing and selecting multiple machine learning models in step (3) includes the following steps: S1. The corrected dataset is stratified and divided into training and test sets in a ratio of 85%:15%. The stratification strategy ensures that the distribution of color difference values ​​in the training and test sets is consistent with that in the original dataset, making the model evaluation more reliable. S2. On the training set, train multiple machine learning regression models in parallel, including multinomial regression, random forest regressor, gradient boosting regressor, support vector regression (SVR), and XGBoost regressor. S3. For each model, use the GridSearchCV method with 5-fold cross-validation to determine the coefficient of determination R. 2 To evaluate the metrics, we systematically searched for the optimal combination of hyperparameters, used an independent test set to perform a final performance evaluation on all optimized models, and selected the model with the highest R² score on the test set as the final color difference prediction model. Among them, R 2 The formula is as follows:

[0041] Among them, y i Let i be the true value of the i-th sample. Let be the predicted value for the i-th sample. is the average of the true values, and n is the number of samples.

[0042] Preferably, the data correction method in step (4) is as follows: traverse each parameter setting and check its corresponding multiple parallel measurements; if the G value of a certain value is found to be greater than 1.155, it is determined to be an outlier and replaced with the average of the other normal values ​​in the same group. Finally, a cleaned and fused robust measurement result is obtained for each parameter setting.

[0043] The computer-based marking region recognition method and image recognition method described above can both employ existing technologies such as the bounding box method or other methods capable of directional region recognition.

[0044] Example 1 To illustrate the technical solution of this invention in detail, this embodiment uses a 355nm ultraviolet laser marking machine with a power of 5W as an example to conduct a marking experiment on a peach sample. The overall technical process of this embodiment is as follows: Figure 2 As shown, the method of this invention is completed by constructing an experimental dataset, performing feature engineering and data preprocessing, training, evaluating and selecting multiple models, and finally applying the optimal model to color difference prediction.

[0045] Step 1: Constructing a laser-marked color difference experimental dataset 1. Determine the input features and output target: The input features are three core adjustable process parameters of the laser and scanning galvanometer: scanning speed (mm / s), repetition frequency (kHz), and pulse width (μs). Parameters of the laser marking machine, such as laser power (W) and laser wavelength (nm), are also collected.

[0046] The output target is the total color difference (ΔE) generated on the sample surface 24 hours after laser treatment. The total color difference (ΔE) is calculated according to the following formula, where... , , The color value before laser processing. , , Color values ​​24 hours after laser treatment: ; 2. Perform the experiment and collect data: Orthogonal experimental design method is used to systematically combine different input feature parameters.

[0047] The scanning parameters are 0.05mm line spacing, 0.05mm focal spot diameter, one scan, a square area with a side length of 1.8cm, and a serpentine scanning path.

[0048] The self-comparison method was adopted. Before marking, the initial color value of the area to be processed was measured with a colorimeter. The final color value of the same area was measured again 24 hours after marking was completed.

[0049] To ensure data reliability, each parameter group was tested in three parallel replicates.

[0050] The recorded input features and their corresponding output targets (the average of three repeated experiments) are compiled to form the original dataset. The original dataset is shown in Table 1 below: Table 1 Summary of Original Datasets

[0051]

[0052]

[0053]

[0054] Scan results as follows Figure 3 As shown, Figure 3 These are peach samples used for data acquisition in an embodiment of the present invention. (Compared to those in the background art...) Figure 1 In stark contrast to the inconsistent marking effects and difficulty in quality control caused by the "trial and error method" shown, this invention employs a systematic orthogonal experimental design method to perform planned and standardized laser scanning on "Zhonghua Crisp Peach" samples. Figure 3 The study clearly demonstrates the independent scanning regions delineated on multiple samples for testing with different parameter combinations. This rigorous and repeatable experimental method forms the basis for building a high-quality dataset, ensuring that the acquired color difference data (ΔE) can truly and accurately reflect the intrinsic relationship between laser process parameters (scanning speed, repetition frequency, and pulse width) and the marking effect. This provides reliable data support for the subsequent training of a high-precision prediction model, fundamentally overcoming the shortcomings of existing technologies, such as long R&D cycles and unpredictable results due to a lack of scientific experimental design.

[0055] Step 2: Feature Engineering and Data Preprocessing 1. Feature Engineering: To capture the interactive effects between various parameters, new interactive features are constructed based on the three original input features, including: "scanning speed". "Pulse frequency", "scan speed" "Pulse width" and "Pulse frequency" The header of the feature list of the expanded dataset, obtained by changing the "pulse width", is as shown in Table 2 below, which is different from the header of Table 1. Table 2 Feature Table of Extended Dataset

[0056]

[0057]

[0058]

[0059] 2. Data Cleaning: The expanded dataset is examined to remove samples with missing values, and outlier detection and correction are performed. The correction method is as follows: For each set of process parameters, the mean and standard deviation of three parallel replicate experiments are calculated. Each parameter setting is iterated through, and its corresponding three parallel measurements are checked. If a value with a G-value greater than 1.155 is found, it is identified as an outlier and replaced with the average of the remaining normal values ​​in the same group. Finally, a cleaned and fused robust measurement result is obtained for each parameter setting.

[0060] 3. Data Standardization: A standardization method (e.g., StandardScaler) is used to standardize all input features of the expanded dataset, including original features (i.e., scan speed, scan frequency, and pulse width) and interactive features ("scan speed"). "Pulse frequency", "scan speed" "Pulse width" and "Pulse frequency" The pulse width is processed to eliminate the dimensional differences between different features, so that the model can treat each feature more fairly.

[0061] For each feature to be standardized, the standardization method formula is: , where z is the new value of the feature to be standardized after standardization, and x is the original data value of the feature. σ is the mean of all the original data for this feature, and σ is the standard deviation of all the original data for this feature.

[0062] Step 3: Training, evaluating, and selecting the best machine learning model 1. Dataset Splitting: After feature engineering and data preprocessing, the data is first binned, then stratified sampling is performed, dividing it into training and test sets at an 85%:15% ratio. This ensures that the distribution of color difference values ​​in the training and test sets is consistent with the distribution of color difference values ​​in the entire dataset (for example, if the proportion of color difference values ​​1-5 in the entire dataset is 30%, then the proportion of color difference values ​​1-5 in the training and test sets is also 30%). This stratification strategy ensures that the distribution of color difference values ​​in the training and test sets is consistent with the original dataset, making model evaluation more reliable.

[0063] 2. Multi-model training: Multiple machine learning regression models are trained in parallel on the training set, including multinomial regression, random forest regressor, gradient boosting regressor, support vector regression (SVR), and XGBoost regressor.

[0064] 3. Hyperparameter optimization: For each model, the GridSearchCV method with 5-fold cross-validation is used to systematically find the optimal combination of hyperparameters, with the coefficient of determination (R²) as the evaluation index.

[0065] The formula for R² (coefficient of determination) is as follows:

[0066] Among them, y i Let i be the true value of the i-th sample. Let be the predicted value for the i-th sample. y is the average of the true values, and n is the number of samples.

[0067] 4. Model Evaluation and Selection: All optimized models were evaluated using an independent test set. The model with the highest R² score on the test set was selected as the final color difference prediction model. For example... Figures 4 to 7 As shown in Table 3, the experimental results indicate that the Gradient Boosting Regressor performs optimally, and its cross-validation R^2 value is high, demonstrating good data generalization ability.

[0068] Table 3. Cross-validation scores of the gradient boosting regressor under different hyperparameters.

[0069] Based on the original input features, we added interactive features through feature engineering to improve model performance, and standardized all data. Subsequently, we split the dataset into training and test sets and compared the prediction performance of several mainstream machine learning regression models (such as multinomial regression, random forest, XGBoost, etc.). After 5-fold cross-validation and rigorous hyperparameter optimization, the Gradient Boosting Regressor model was confirmed as the best performing model, with its coefficient of determination R0 on the test set being [value missing]. 2 With a value as high as 0.979, it demonstrates extremely high prediction accuracy and stability.

[0070] Step 4: Application of color difference prediction based on the optimal model 1. Deploy the trained optimal model (gradient boosting regressor).

[0071] 2. When operators are conducting actual production or research and development, they do not need to conduct physical experiments. They only need to input a new combination of laser process parameters (scanning speed, repetition frequency, pulse width) into the system. Some of the predictors are shown in Table 4.

[0072] Table 4 Partial Prediction Sub-Data Table

[0073] 3. The model will instantly output an accurate prediction of the color difference ΔE after 24 hours under this parameter combination, thereby guiding the rapid setting and optimization of process parameters.

[0074] When specific marking effects are required (such as controlling the color difference ΔE within a target range), operators do not need to conduct tedious physical experiments. They only need to input a new set of process parameters into the prediction system, and the model can instantly and cost-free output the predicted color difference value. This prediction result can quickly verify the feasibility of the new parameter combination, thereby greatly shortening the R&D cycle and improving the efficiency and accuracy of process development.

[0075] Although embodiments of the invention have been disclosed for illustrative purposes, those skilled in the art will understand that various substitutions, variations, and modifications are possible without departing from the spirit and scope of the invention and the appended claims. Therefore, the scope of the invention is not limited to the contents disclosed in the embodiments.

Claims

1. A machine learning-based laser marking method for fruits and vegetables, characterized in that: The method includes the following steps: (1) Fruit and vegetable skin data recognition: The image in the working area of ​​the laser is dynamically captured by the camera and the captured image information is transmitted to the computer. The computer calculates the dynamic information of the image information to identify the skin color and curvature of the fruit and vegetable to be marked. The computer also displays the captured image information on the computer screen. (2) Selection of the marking area: The computer determines the best marking area for fruits and vegetables based on the data identified in step (1), displays it on the screen, and then transmits the location information of the marking area to the fruit and vegetable placement platform, so that the fruit and vegetable placement platform moves the fruits and vegetables placed on it to be laser-marked directly below the laser. (3) Laser marking energy output based on machine learning: Data of the skin of standard samples of fruits and vegetables after marking in actual production is collected into the computer. The parameters used by the laser and scanning galvanometer are input into the computer. Then, the laser and scanning galvanometer are operated to laser mark the skin of fruits and vegetables in the marking area in step (2). Multiple fruits and vegetables are marked with lasers of different parameters. At the same time, the parameters of the laser used and the marking time are recorded when each fruit and vegetable is laser marked. The camera collects the color values ​​before and after marking. The collected data and standard samples are used to calculate the color difference based on the following formula: ; in, , , The color value before laser processing. , , The color value is 24 hours after laser treatment, ΔE is the total color difference, where L represents the brightness of the fruit surface, a represents the position of the color on the red-green axis, reflecting the distribution of red and green components, b represents the position of the color on the yellow-green axis, reflecting the distribution of yellow and green components, the subscript initial indicates the value before treatment, and final indicates the value measured 24 hours after treatment; By importing the color difference calculation data of multiple fruits and vegetables and the color difference calculation data of standard samples, along with the parameters of the laser used, the parameters of the scanning galvanometer, and the marking time, into various machine learning models, the original input features are the three core adjustable process parameters of the laser and the scanning galvanometer: scanning speed mm / s, repetition frequency KHz, and pulse width μs. Machine learning is performed in this way, and the various models are compared to select the best initial model. (4) Data correction and model use: Correct the initial model after machine learning in step (3); import the corrected model directly into the laser marking machine used in production. The model can then use the data from the camera that scans the fruits and vegetables with the laser used in production, combined with the target color difference ΔE set by the operator, to solve the best output parameters of the laser and scanning galvanometer through the optimization algorithm, and guide the laser and scanning galvanometer to perform laser marking on the fruits and vegetables.

2. The method according to claim 1, characterized in that: The color recognition method in step (1) is as follows: S1 uses a hyperspectral camera to capture and scan the surface information of fruits and vegetables on the fruit and vegetable placement platform. The hyperspectral camera can identify extremely subtle color differences on the surface of fruits and vegetables, and can also identify defects on the surface of fruits and vegetables. S2 uses a trained deep learning segmentation algorithm, U-Net network model or Mask R-CNN instance segmentation model, to analyze the captured images.

3. The method according to claim 1, characterized in that: The method for identifying the marking area in step (2) is as follows: a vision-guided robot is used for operation, specifically: S1 collects information about the surface of fruits and vegetables through a camera. The computer determines the position and orientation of the best marking area in the image. Based on this position and orientation information, it guides a multi-axis robotic arm with flexible grippers to move over and accurately grab the fruit or vegetable. The S2 robot's robotic arm moves the fruits and vegetables directly under the laser, while rotating its wrist joint and adjusting the posture of the fruits and vegetables in conjunction with the image information captured by the camera, so that the surface of the fruits and vegetables facing the laser is a non-damaging marking area. After S3 is aligned in both posture and position, the robotic arm sends a signal to trigger the laser and scanning galvanometer to complete the task.

4. The method according to claim 1, characterized in that: In step (3), to capture the interaction between various system parameters, new interaction features are constructed based on three original input features, including: scanning speed. Pulse frequency, scan speed Pulse width, pulse frequency Pulse width; The dataset above was examined to remove samples with missing values. Outlier detection and correction were performed using the following methods: 1) For the data from three parallel replicate experiments under each set of process parameters, the statistically recognized Grubbs' Test was used to identify and remove outliers when processing the data from each set of parallel replicate experiments under each set of process parameters. First, the arithmetic mean of each set of measurements, i.e., ΔE, was calculated. The standard deviation s is given, where the number of parallel experiments n ≥ 3; subsequently, suspicious data points with the largest deviation from the group mean are identified, and their G-statistics are calculated, defined as... , where x i Data points that are considered suspicious; 2) Compare the calculated G value with the Grubbs critical value G corresponding to the sample size n at a preset confidence level of 95%. crit For comparison, when n=3, G crit The threshold is 1.

155. If the calculated value of G is greater than this threshold, the data point is determined to be an outlier and removed from the data set; otherwise, all data is retained. 3) Finally, based on the remaining valid data after removing any identified outliers, the arithmetic mean is recalculated and this result is used as the final representative measurement for this set of process parameters. The data in the above correction method is standardized using the StandardScaler method to process all input features of original features and interactive features, including original features and interactive features, to eliminate the dimensional differences between different features. The standardized formula is as follows: ; Where z is the standardized new value, and x is the original data value. σ is the mean of all the original data for this feature, and σ is the standard deviation of all the original data for this feature.

5. The method according to claim 1, characterized in that: The method for comparing and selecting multiple machine learning models in step (3) includes the following steps: S1. Perform stratified sampling on the corrected dataset and divide it into training and test sets in a ratio of 85%:15%. S2. On the training set, train multiple machine learning regression models in parallel, including multinomial regression, random forest regressor, gradient boosting regressor, support vector regression (SVR), XGBoost regressor, or other models. S3. For each model, use the GridSearchCV method with 5-fold cross-validation to determine the coefficient of determination R. 2 To evaluate the metrics, we systematically search for the optimal combination of hyperparameters, and use independent test sets to perform a final performance evaluation on all optimized models. We then select R on the test set. 2 The model with the highest score will be used as the final color difference prediction model; Among them, R 2 The formula is as follows: ; Among them, y i Let i be the true value of the i-th sample. Let be the predicted value for the i-th sample. is the average of the true values, and n is the number of samples.

6. The method according to any one of claims 1 to 5, characterized in that: The data correction method in step (4) is as follows: traverse each parameter setting and check its corresponding multiple parallel measurement values; if the G value of a certain value is found to be greater than 1.155, it is determined to be an outlier and replaced with the average value of the other normal values ​​in the same group.

7. The method according to any one of claims 1 to 5, characterized in that: The specific steps are as follows: Step 1: Constructing a laser-marked color difference experimental dataset 1) Determine the input features and output target: The input features are the three core adjustable process parameters of the system: scanning speed (mm / s), repetition frequency (KHz), and pulse width (μs). The output target is the total color difference ΔE generated on the sample surface 24 hours after laser treatment; the total color difference ΔE is calculated according to the following formula, where , , The color value before laser processing. , , Color values ​​24 hours after laser treatment: ; 2) Perform the experiment and collect data: Orthogonal experimental design method is used to systematically combine different input feature parameters; The scanning method uses a line spacing of 0.05mm, a spot diameter of 0.05mm at the focal plane, one scan, a square area with a side length of 1.8cm, and a serpentine scanning path. The self-comparison method was adopted. Before marking, the initial color value of the area to be processed was measured with a colorimeter. The final color value of the same area was measured again 24 hours after marking was completed. To ensure data reliability, each parameter group was tested in three parallel replicates. The original dataset is formed by combining the recorded input features and their corresponding output targets, which are the average values ​​of the three repeated experiments. Step 2: Feature Engineering and Data Preprocessing 1) Feature Engineering: To capture the interactive effects between various parameters, new interactive features are constructed based on the three original input features, including: scanning speed. Pulse frequency, scan speed Pulse width, pulse frequency The pulse width is used to obtain the feature list of the expanded dataset. 2) Data Cleaning: The expanded dataset is inspected to remove samples with missing values, and outlier detection and correction are performed. The correction method is as follows: For each set of process parameters, the mean and standard deviation of three parallel repeated experiments are calculated; each parameter setting is iterated through, and its corresponding three parallel measurements are checked; if the G value of a certain value is found to be greater than 1.155, it is identified as an outlier and replaced with the average of the other normal values ​​in the same group; finally, a cleaned and fused robust measurement result is obtained for each parameter setting. 3) Data Standardization: A standardization method is used to standardize all input features of the expanded dataset, including the original features (scan speed, repetition frequency, and pulse width) and the interactive features (scan speed). Pulse frequency, scan speed Pulse width, pulse frequency Pulse width is processed to eliminate the dimensional differences between different features, enabling the model to treat each feature more fairly. For each feature to be standardized, the standardization method formula is: , where z is the new value of the feature to be standardized after standardization, and x is the original data value of the feature. σ is the mean of all the original data for this feature, and σ is the standard deviation of all the original data for this feature. Step 3: Training, evaluating, and selecting the best machine learning model 1) Data set partitioning: The data after feature engineering and data preprocessing is first binned, and then stratified sampling is performed to divide it into training set and test set in a ratio of 85%:15% to ensure that the distribution of color difference values ​​in the training set and test set is consistent with the distribution of color difference values ​​in the entire dataset. 2) Multi-model training: Train various machine learning regression models on the training set, including multinomial regression, random forest regressor, gradient boosting regressor, support vector regression, XGBoost regressor, or other models; 3) Hyperparameter optimization: For each model, the GridSearchCV method with 5-fold cross-validation is used to determine the coefficient of determination R. 2 To evaluate the metrics, we systematically search for the optimal combination of hyperparameters. Among them, R 2 The formula for the coefficient of determination is as follows: Among them, y i Let i be the true value of the i-th sample. Let be the predicted value for the i-th sample. y is the average of the true values ​​of y, and n is the sample size; 4) Model Evaluation and Selection: Perform a final performance evaluation of all optimized models using independent test sets; select the appropriate R model for the test set. 2 The model with the highest score was selected as the final color difference prediction model, and the results show that the gradient boosting regressor performed optimally. Step 4: Application of color difference prediction based on the optimal model The trained optimal model, i.e., the gradient boosting regressor, is then deployed and labeled.

8. The method according to claim 7, characterized in that: The system is a 5W 355nm ultraviolet laser marking machine.

9. A machine learning-based laser marking system for fruits and vegetables using the method described in any one of claims 1 to 8, characterized in that: The system includes a camera, a laser, a fruit and vegetable placement platform, a computer, a robot, and a scanning galvanometer. The scanning galvanometer, camera, laser, and robot are all connected to the computer via electrical signals. The scanning galvanometer is connected to the laser, and the laser is connected to the camera. The side of the fruit and vegetable placement platform is connected to the robot. The laser, scanning galvanometer, and camera are spaced apart above the fruit and vegetable placement platform. The robot can place fruits and vegetables on the platform, the camera can capture image information of the surface of the fruits and vegetables on the platform, the laser can perform laser marking on the surface of the fruits and vegetables on the platform through the scanning galvanometer, and the computer can collect image information from the camera via electrical signals and control the working status of the laser, scanning galvanometer, and robot.

Citation Information

Patent Citations

  • Prediction method of phase transition temperature in selective laser melting NiTi alloy based on machine learning

    CN118430701B

  • Laser marking control method and system adaptive to marking workpiece surface attributes

    CN115837520A

  • Laser automatic cutting and UV ink cartridge based marking system

    KR102836427B1