Sweet potato near-infrared multi-index modeling method based on spectral characteristics and multi-task learning

CN122814524APending Publication Date: 2026-09-25CROP RES INST GUANGDONG ACAD OF AGRI SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611244645.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-17
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,不同品质组分在近红外波段内的吸收响应存在重叠,样品粒度、含水状态和装样状态产生的散射变化还会对较弱的组分光谱响应形成干扰,使多个品质指标同步建模时容易发生光谱信息混杂

Benefits of technology

[0022]与现有技术相比,本发明具有如下实质性特点和显著进步:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122814524A_ABST
    Figure CN122814524A_ABST
Patent Text Reader

Abstract

The application discloses a sweet potato near-infrared multi-index modeling method based on spectral characteristics and multi-task learning, and belongs to the technical field of near-infrared spectrum analysis. The method collects the near-infrared reflection spectrum of a sweet potato calibration sample and performs standardization processing to generate a reference value of amylopectin and a difference mode label of amylose-amylopectin; a near-equal sample pair is constructed according to the reference value of total starch, a common mode spectrum base of total starch is extracted and common mode projection is removed, and a difference mode spectrum base orthogonal to the common mode spectrum base is obtained according to the remaining spectrum difference; total starch, difference mode and the remaining quality index prediction tasks are established respectively, a multi-index prediction model is formed through joint training, and the prediction values of amylose and amylopectin are restored from the prediction values of total starch and difference mode. The application can weaken the masking of the strong response of total starch to the difference of starch composition, and maintain the quality relationship between each starch component.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of near-infrared spectroscopy analysis technology, specifically involving a near-infrared multi-index modeling method for sweet potatoes based on spectral features and multi-task learning. Background Technology

[0002] The contents of total starch, amylose, amylopectin, soluble sugar, reducing sugar, crude protein, and crude fiber in sweet potato tubers are important indicators for evaluating the eating quality, processing suitability, and breeding material value of sweet potatoes. These indicators are typically determined separately using chemical detection methods such as enzymatic hydrolysis, colorimetry, titration, or gravimetric analysis. However, this approach suffers from numerous sample pretreatment steps, long detection cycles, high reagent consumption, and difficulty in rapidly and simultaneously detecting large numbers of samples. Near-infrared spectroscopy, on the other hand, can reflect the internal component information of a sample through its absorption and scattering responses to different wavelengths of near-infrared light. Combined with chemical reference values ​​of calibrated samples, it can establish predictive relationships for quality indicators, making it suitable for rapid, reagent-free, or low-reagent detection of sweet potato quality. However, the absorption responses of different quality components in the near-infrared band overlap, and scattering variations caused by sample particle size, moisture content, and loading conditions can interfere with the spectral responses of weaker components, easily leading to spectral information confounding when modeling multiple quality indicators simultaneously.

[0003] Total starch content primarily reflects the overall mass of starch in sweet potato samples, while amylose and amylopectin reflect the internal compositional distribution of total starch. Among samples with approximately the same total starch content but different amylose to amylopectin ratios, the strong spectral response of total starch shows relatively small changes, and the spectral differences corresponding to relative changes in amylose and amylopectin are usually weak. If total starch, amylose, and amylopectin are treated as parallel prediction targets and directly share the same spectral characterization, the strong spectral response corresponding to total starch can easily mask the weak differences in response between amylose and amylopectin. An increase in amylose is usually accompanied by a corresponding decrease in amylopectin; the same compositional change contributes in opposite directions to the two prediction results. Establishing separate prediction relationships can easily lead to inter-task interference, resulting in a relatively stable prediction result for total starch, while the prediction results for amylose and amylopectin show larger deviations, and it is difficult to maintain a clear component mass relationship among the three. Summary of the Invention

[0004] The present invention aims to provide a method for near-infrared multi-index modeling of sweet potatoes based on spectral features and multi-task learning, which solves the technical problems mentioned in the background art.

[0005] To achieve the above-mentioned objectives, the present invention provides the following technical solution.

[0006] A near-infrared multi-index modeling method for sweet potatoes based on spectral features and multi-task learning includes:

[0007] Near-infrared reflectance spectra of sweet potato calibration samples were collected. These spectra were then subjected to absorbance conversion and standardization to obtain standard spectra for each sweet potato calibration sample. Chemical reference detection methods were used to obtain reference values ​​for total starch, amylose, soluble sugar, reducing sugar, crude protein, and crude fiber corresponding to the standard spectra, ensuring that all reference values ​​used the same quality standard and unit of measurement.

[0008] Subtracting the amylose reference value from the total starch reference value yields the amylopectin reference value; subtracting the amylopectin reference value from the amylose reference value yields the amylose-amylopectin differential label.

[0009] For each sweet potato calibration sample, the sample with the closest total starch reference value from the remaining sweet potato calibration samples is selected as the paired sample. The corresponding two sweet potato calibration samples are combined into a near-equal sample pair, and the linear-branched differential label of the two sweet potato calibration samples in the near-equal sample pair is subtracted to obtain the differential label difference. When there are more than two candidate samples with the same absolute value of the difference in total starch reference value, the candidate sample with the largest absolute value of the linear-branched differential label difference with the current sweet potato calibration sample is selected as the paired sample.

[0010] A first partial least squares regression relationship is established based on the standard spectra and total starch reference values ​​of each sweet potato calibration sample. The common mode spectral basis is obtained based on the spectral loading direction of the first partial least squares regression relationship. The standard spectra of each sweet potato calibration sample are projected onto the common mode subspace spanned by the common mode spectral basis, and the corresponding projection components are removed from the standard spectra to obtain the remaining spectra.

[0011] The residual spectra of the two sweet potato correction samples in each near-equal sample pair are subtracted to obtain the residual spectral difference. A second partial least squares regression relationship is established based on the residual spectral difference and the corresponding differential mode label difference to obtain the initial straight-chain-branched differential mode spectral basis. The initial straight-chain-branched differential mode spectral basis is projected onto the orthogonal complement space of the common mode subspace, and the projection result is normalized and orthogonalized to obtain the straight-chain-branched differential mode spectral basis orthogonal to the common mode spectral basis.

[0012] The standard spectra of each sweet potato calibration sample were projected onto the common mode spectral base and the linear-branched differential mode spectral base to obtain the common mode score and differential mode score. The common mode score was input into the total starch prediction task, and the differential mode score was input into the linear-branched differential mode prediction task. A shared spectral characterization was formed based on the standard spectra. The shared spectral characterization was input into the soluble sugar prediction task, the reducing sugar prediction task, the crude protein prediction task, and the crude fiber prediction task, respectively.

[0013] The prediction errors for the total starch prediction task, the amylose-branched differential prediction task, the soluble sugar prediction task, the reducing sugar prediction task, the crude protein prediction task, and the crude fiber prediction task were calculated separately. Each prediction error was divided by the variance of the corresponding reference value or label to obtain the normalized task error. The amylose-branched differential prediction values ​​of two sweet potato correction samples in a pair of nearly equal sample sizes were subtracted to obtain the differential prediction difference. The differential preservation error was obtained based on the error between the differential prediction difference and the corresponding differential label difference. The arithmetic mean of each normalized task error and the differential preservation error was used as the joint training loss.

[0014] The multi-task learning relationship has a separable parameter structure, wherein the training error of the total starch prediction task is used to determine the parameters of the total starch prediction task, the training error and the difference mode retention error of the linear-branched differential mode prediction task are used to determine the parameters of the linear-branched differential mode prediction task, and the training errors of the soluble sugar prediction task, the reducing sugar prediction task, the crude protein prediction task, and the crude fiber prediction task are used together to determine the shared parameters and the parameters of the corresponding prediction tasks.

[0015] According to sweet potato variety or harvest batch, the sweet potato calibration samples are divided into training group and verification group. In each training and verification process, the sweet potato calibration samples in the training group are used to construct near-equal sample pairs, obtain common mode spectral basis and linear-branched differential mode spectral basis, and solve the prediction task parameters. Then, the spectral basis and prediction task parameters obtained from the training group are used to calculate the prediction results of the verification group, so as to reduce the evaluation bias caused by the simultaneous entry of sample information of the same variety or the same harvest batch into the training and verification processes.

[0016] The total starch prediction value is obtained using the trained total starch prediction task, and the amylose-branch differential prediction value is obtained using the trained amylose-branch differential prediction task. Half of the sum of the total starch prediction value and the amylose-branch differential prediction value is determined as the amylose prediction value, and half of the difference between the total starch prediction value and the amylose-branch differential prediction value is determined as the amylopectin prediction value, so that the sum of the amylose prediction value and the amylopectin prediction value equals the total starch prediction value.

[0017] Furthermore, before performing component recovery of the amylose and amylopectin predicted values, the original output of the total starch prediction task is non-negatively processed, and the original output of the amylose-amylose differential prediction task is limited to between the negative total starch predicted value and the total starch predicted value, so that both the amylose and amylopectin predicted values ​​are not less than zero.

[0018] Furthermore, the standardization process of the near-infrared reflectance spectrum includes: taking the negative logarithm of the near-infrared reflectance at each wavelength position to obtain the absorbance spectrum; calculating the average and standard deviation of the absorbance values ​​at all wavelength positions of the same sweet potato calibration sample; subtracting the average value from each absorbance value and dividing by the standard deviation to obtain the standard spectrum.

[0019] Furthermore, the loading directions of the first partial least squares regression relation are normalized and orthogonalized to obtain the common-mode spectral basis; the loading directions of the second partial least squares regression relation are projected onto the orthogonal complement space of the common-mode subspace, and the projection results are normalized and orthogonalized to obtain the straight-chain-branched differential-mode spectral basis.

[0020] Furthermore, when the L2 norm of the spectral weight direction is zero or the sum of squares of the latent variable scores is zero during partial least squares regression, the extraction of latent variables is stopped; before performing orthogonal decomposition on the loading matrix, the zero vector column and the column linearly related to the retained columns are deleted.

[0021] Furthermore, the sweet potato near-infrared multi-index prediction model saves the wavelength position and arrangement order of the near-infrared spectrum, standardized processing parameters, common mode spectral basis, linear-branched differential mode spectral basis, total starch prediction task parameters, linear-branched differential mode prediction task parameters, shared parameters, shared task direction, and the average and standard deviation of the reference values ​​for each shared prediction task.

[0022] Compared with the prior art, the present invention has the following substantial features and significant progress:

[0023] This invention collects near-infrared reflectance spectra of sweet potato samples and performs standardized processing. It then establishes a correspondence between the spectral data and reference values ​​for total starch, amylose, soluble sugars, reducing sugars, crude protein, and crude fiber. Based on a single spectral acquisition, it simultaneously constructs predictive relationships for multiple quality indicators, reducing the sample pretreatment, reagent consumption, and detection time required for destructive chemical testing of each quality indicator. By setting shared spectral characterization for the prediction tasks of soluble sugars, reducing sugars, crude protein, and crude fiber, multiple quality indicators can utilize their common response information in the near-infrared spectrum, improving the utilization of spectral information. Furthermore, by dividing the training and validation groups according to sweet potato variety or harvest batch, it reduces evaluation bias caused by information from samples of the same source entering both training and validation phases, thus improving the applicability of the prediction model to sweet potato samples of different varieties and harvest batches.

[0024] This invention constructs nearly equal sample pairs based on the principle of having the closest total starch reference value. First, it extracts the common-mode spectral basis related to changes in total starch content and removes the corresponding projected components from the standard spectrum. Then, based on the remaining spectral difference of the nearly equal sample pairs, it extracts the amylose-branched differential-mode spectral basis orthogonal to the common-mode spectral basis, thereby reducing the masking effect of the strong spectral response of total starch on the weak differential responses of amylose and amylopectin. Furthermore, it sets the total starch prediction task and the amylose-branched differential-mode prediction task as separate prediction tasks, and uses the differential-mode retention error constraint to constrain the differential-mode prediction relationship between nearly equal sample pairs, which can reduce the interference of other quality tasks on the starch component prediction direction. By recovering the predicted values ​​of amylose and amylopectin from the total starch prediction value and the amylose-branched differential-mode prediction value, the sum of the predicted values ​​of amylose and amylopectin remains the total starch prediction value, ensuring that the prediction results of each starch component conform to a non-negative quality relationship, thus improving the consistency and interpretability of the sweet potato starch composition prediction results. Attached Figure Description

[0025] Figure 1 Flowchart of a method for modeling multiple indicators of near-infrared radiation for sweet potatoes;

[0026] Figure 2 A flowchart illustrating the multi-task learning relationship structure of sweet potato near-infrared spectroscopy. Detailed Implementation

[0027] The following description is provided in conjunction with the accompanying drawings. It should be understood that the following embodiments are for illustrative purposes only and are not intended to limit the invention; those skilled in the art can make equivalent substitutions or combinations for the specific implementations.

[0028] This embodiment uses freeze-dried and pulverized sweet potato powder as the spectral detection object, ensuring that the calibration sample and the test sample share the same physical state, sample preparation conditions, spectral acquisition conditions, and data processing rules. In this embodiment, near-infrared reflectance is expressed as a symbol... The reference value for reducing sugars is indicated by the symbol. The crude protein reference value is indicated by the symbol. The response regression coefficients in the partial least squares regression of total starch are expressed as follows: express.

[0029] In one embodiment, a near-infrared multi-index prediction model for sweet potatoes based on spectral features and multi-task learning is established to obtain prediction results for total starch, amylose, amylopectin, soluble sugar, reducing sugar, crude protein, and crude fiber based on the near-infrared reflectance spectrum of the same sweet potato sample; such as Figure 1As shown, firstly, sweet potato calibration samples were prepared and numbered. Near-infrared reflectance spectra of each sweet potato calibration sample were collected. The near-infrared reflectance spectra were standardized, and chemical reference values ​​for total starch, amylose, soluble sugar, reducing sugar, crude protein, and crude fiber were determined. Based on the total starch reference value and amylose reference value, amylopectin reference value and amylose-amylose differential mode labels were generated, and near-equal sample pairs were constructed based on the total starch reference value. Then, the common-mode spectral basis of total starch was extracted based on the standard spectrum and the total starch reference value. The common-mode projection component was removed from the standard spectrum, and the amylose-amylose differential mode spectral basis was extracted based on the remaining spectral difference of the near-equal sample pairs. The common-mode score, differential mode score, and shared spectral characterization were calculated respectively. Total starch prediction tasks, amylose-amylose differential mode prediction tasks, soluble sugar prediction tasks, reducing sugar prediction tasks, crude protein prediction tasks, and crude fiber prediction tasks were established and jointly trained. Finally, the predicted values ​​of amylose and amylopectin were recovered based on the predicted values ​​of total starch and differential mode, and the sweet potato near-infrared multi-index prediction model was output.

[0030] The sweet potato calibration samples referred to in this embodiment are sweet potato samples that simultaneously possess near-infrared reflectance spectra and corresponding chemical reference values, used to establish a near-infrared multi-index prediction model for sweet potatoes. The sweet potato calibration samples cover different sweet potato varieties, different harvest batches, and different quality levels, ensuring that each quality reference value reflects actual changes within the sample set. The number of sweet potato calibration samples is denoted as [missing information]. The number of wavelength positions contained in each near-infrared spectrum is denoted as ;

[0031] Sweet potato tubers of different varieties and harvested batches were selected, and each tuber was assigned a unique sample number. Soil and impurities attached to the surface of the tubers were removed. Based on the longitudinal length of the tuber, flesh samples were taken from one-quarter of the way from the top, one-quarter of the way from the middle, and one-quarter of the way from the bottom. The flesh samples from the three locations of the same tuber were mixed to serve as the representative sample for that tuber, so as to reduce the impact of the uneven spatial distribution of internal quality components of the tuber on the modeling results.

[0032] The mixed sweet potato flesh was cut into slices with a thickness of 2 mm and pre-frozen at -40 degrees Celsius for 12 hours. The pre-frozen slices were then placed in a freeze dryer with a cold trap temperature not exceeding -50 degrees Celsius and an absolute pressure in the drying chamber not exceeding 100 Pa for 48 hours. After drying, the samples were crushed and passed through a sieve with a pore size of 0.25 mm to ensure that the particle size of each sweet potato calibration sample was consistent. The sieved sweet potato powder was thoroughly mixed and sealed for storage. Before near-infrared spectroscopy acquisition, the sealed samples were placed in an environment of 25 degrees Celsius for equilibration for 2 hours.

[0033] Five .00 grams of sweet potato powder were weighed from each sweet potato calibration sample and placed into a circular spectral sampling cup with an inner diameter of 30 mm, ensuring that the sample layer thickness was not less than 10 mm. After loading the sample, the sample surface was smoothed, and no additional compaction was applied to the sample to avoid changes in optical path and scattering state caused by different degrees of compaction. All sweet potato calibration samples in the same prediction model used the same sampling location, slice thickness, pre-freezing conditions, drying conditions, particle size, loading mass, and loading thickness. Samples prepared using different drying methods or different physical states were not mixed to establish the same prediction model.

[0034] The diffuse reflectance spectra of each sweet potato calibration sample in the range of 900 nm to 1700 nm were acquired using a near-infrared spectrometer with a wavelength sampling interval of 2 nm. Each spectrum was obtained by taking the arithmetic mean of 32 consecutive scans. The ambient temperature for spectral acquisition was controlled within 25 degrees Celsius. The near-infrared spectrometer was preheated for 30 minutes after startup, and white reference calibration was performed before the first sample spectral acquisition. Thereafter, white reference calibration was performed again after spectral acquisition of every ten sweet potato calibration samples.

[0035] Each sweet potato calibration sample was independently reloaded three times; after each reloading, the spectral sampling cup was rotated and the near-infrared reflectance spectrum was reacquired. The reflectance of the three spectra at the same wavelength position was calculated arithmetically to obtain the original near-infrared reflectance spectrum of the sweet potato calibration sample; the first... The first sweet potato calibration sample was in the... The average near-infrared reflectance at each wavelength position is denoted as:

[0036]

[0037] in, , When the near-infrared reflectance at any wavelength position is less than or equal to zero, the corresponding spectral record is determined as an invalid record, and white reference correction and sample spectrum acquisition are performed again. Logarithmic transformation is not performed on reflectance less than or equal to zero.

[0038] Converting near-infrared reflectance to absorbance:

[0039]

[0040] In the formula, Indicates the first The first sweet potato calibration sample was in the... The absorbance values ​​at each wavelength position are given; both near-infrared reflectance and absorbance are dimensionless. Calculate the absorbance value at the [number]th wavelength position. The average absorbance of the sweet potato calibration sample across all wavelength positions:

[0041]

[0042] Calculate the standard deviation of the absorbance values ​​at all wavelengths of the sweet potato calibration sample:

[0043]

[0044] Perform a standard normal transformation on the absorbance values ​​at each wavelength position:

[0045]

[0046] In the formula, Represents the standardized spectral value; [The text abruptly ends here, likely due to an incomplete sentence or a format All standardized spectral values ​​of the sweet potato calibration samples are arranged in wavelength order to form a standard spectral row vector:

[0047]

[0048] Arrange the standard spectral row vectors of all sweet potato calibration samples sequentially to form a standard spectral matrix:

[0049]

[0050] If a certain sweet potato calibration sample A value of zero indicates that the sample did not form an effective spectral change at all wavelength positions. The instrument light source, reference calibration, sample loading, and data reading process should be checked, and the near-infrared reflectance spectrum of the sample should be reacquired. If the spectrum obtained by reloading the same sweet potato calibration sample three times shows obvious discontinuities, signal saturation, missing data, or abnormal jumps, the sample should be reloaded and reacquired. The corresponding invalid spectrum should not be used for modeling.

[0051] Approximately 2.00 grams of sweet potato powder were weighed from the same sweet potato calibration sample after spectral acquisition. The mass of the sample before drying was recorded, and the powder was dried at 105 degrees Celsius to a constant mass. The constant mass refers to the mass difference between two consecutive weighings taken two hours apart, which does not exceed 0.2% of the previous weighing result. The first... Dry matter mass fraction of each sweet potato calibration sample:

[0052]

[0053] In the formula, This indicates the mass of the sample before drying. This indicates the mass of the sample after it has been dried to a constant mass. Indicates the first Dry matter mass fraction of each sweet potato calibration sample;

[0054] In this embodiment, the total starch, amylose, amylopectin, soluble sugar, reducing sugar, crude protein, and crude fiber contents are all converted to the mass of the corresponding components per 100 grams of dried sweet potato, with the unit being grams per 100 grams of dry basis. Each quality indicator is measured twice in parallel. When the difference between the two parallel measurements is less than 5% of the arithmetic mean of the two measurements, the arithmetic mean of the two measurements is used as the corresponding reference value. When the ratio exceeds 5%, parallel measurements are performed again.

[0055] Weigh 0.1000 grams of sweet potato powder and add 0.2 ml of 80% ethanol to fully wet the sample. Add 3 ml of phosphate buffer solution with a concentration of 50 mmol / L and a hydrogen ion concentration index of 7.0, and thermostable α-amylase, ensuring that the α-amylase activity added to each sample is not less than 300 units. One unit of α-amylase activity refers to the amount of enzyme required to release one micromole of reducing sugar equivalent per minute using soluble starch as a substrate under the temperature and pH conditions specified in the enzyme preparation test certificate. The same batch of α-amylase should be used for the same batch of sweet potato calibration samples, and the amount of enzyme added should be converted according to the activity recorded in the enzyme preparation test certificate.

[0056] The reaction tubes were placed in a boiling water bath for six minutes, and then shaken thoroughly at the second and fourth minutes after the start of the reaction to ensure that the starch in the sample was gelatinized and liquefied. After the reaction tubes were cooled to 50 degrees Celsius, four milliliters of sodium acetate buffer solution with a concentration of 100 mmol / L and a hydrogen ion concentration index of 4.5, along with glucosyl amylase, were added to ensure that the glucosyl amylase activity added to each sample was not less than 30 units. One unit of glucosyl amylase activity refers to the amount of enzyme required to release one micromole of glucose per minute using soluble starch as a substrate under the temperature and pH conditions specified in the enzyme preparation test certificate. The same batch of glucosyl amylase was used for the same batch of sweet potato calibration samples.

[0057] The reaction tube was placed at 50 degrees Celsius for 30 minutes. After the reaction, the reaction solution was brought to a final volume of 100 ml, thoroughly mixed, and centrifuged. The supernatant was used for glucose oxidase-peroxidase colorimetric reaction. A standard curve was established using a glucose standard solution to correlate glucose mass with absorbance. The absorbance of the supernatant was measured at 510 nm. The mass of glucose released in the entire reaction solution was calculated. The mass of glucose released was calculated according to the following formula. Reference values ​​for total starch in one corrected sweet potato sample:

[0058]

[0059] In the formula, This represents the glucose mass calculated from the standard curve and converted to the total reaction solution based on the liquid sampling ratio and fixed volume, in grams; The value represents the mass of the sweet potato sample used in the test, in grams; 0.9 represents the mass coefficient of glucose converted into anhydrous glucose units that make up starch. This indicates the reference value for total starch, expressed in grams per 100 grams of dry weight.

[0060] Weigh 2.00 g of dried sweet potato powder, add 40 ml of water, stir continuously for 30 minutes, and filter through a sieve with a pore size of 0.15 mm; centrifuge the filtrate at 3,000 times the acceleration of gravity for 10 minutes and collect the precipitate; wash the precipitate twice with water, then wash it once with 80% ethanol, and dry it at 40 degrees Celsius to obtain the starch to be tested; weigh 0.1000 g of the starch to be tested, add 1 ml of anhydrous ethanol and 9 ml of sodium hydroxide solution with a concentration of 1 mol / L, treat it in a water bath at 85 degrees Celsius for 10 minutes, cool it and make up to 100 ml to obtain the starch dispersion;

[0061] Five milliliters of starch dispersion were transferred to one milliliter of acetic acid solution with a concentration of 1 mol / L and two milliliters of iodine-potassium iodide solution, and then the volume was adjusted to 100 milliliters. The iodine-potassium iodide solution was prepared by dissolving 0.2 grams of iodine and 2 grams of potassium iodide and adjusting the volume to 100 milliliters. After 20 minutes of color development, the absorbance was measured at 620 nm. A standard curve was established using a mixed standard of amylose and amylopectin with a known mass ratio of amylose to obtain the mass ratio of amylose in the starch to be tested. The amylose mass percentage was converted to an amylose reference value with the same mass basis as the total starch reference value:

[0062]

[0063] In the formula, Use decimal form in calculations. This represents the reference value for amylose, expressed in grams per 100 grams of dry basis. When the amylose detection method used already directly yields the mass of amylose per 100 grams of sweet potato dry matter, the corresponding detection result is directly used as the reference value. The proportional conversion will no longer be performed.

[0064] Weigh 0.500 g of dried sweet potato powder, add 25 mL of 80% ethanol, extract in an 80°C water bath for 30 minutes, centrifuge and collect the supernatant; repeat the same ethanol extraction on the centrifuged precipitate once, combine the supernatants from the two extractions, and dilute to 100 mL to obtain a soluble sugar extract; transfer 1 mL of the soluble sugar extract, evaporate the ethanol at 60°C, add 1 mL of water and 5 mL of anthrone-sulfuric acid colorimetric solution, react in a boiling water bath for 10 minutes, cool and measure the absorbance at 620 nm; the anthrone-sulfuric acid colorimetric solution is prepared by adding 0.2 g of anthrone to concentrated sulfuric acid and diluting to 100 mL, the colorimetric solution should be stored away from light and used within the specified shelf life after preparation;

[0065] A standard curve was established using a glucose standard solution. Based on the standard curve, the volume of the extract, and the proportion of extract taken for color development, the glucose equivalent mass of soluble sugars in the entire extract was calculated. Calculate the reference value for soluble sugars using the following formula:

[0066]

[0067] In the formula, This indicates the glucose equivalent mass of soluble sugars converted to the total extract, in grams. The reference value for soluble sugars is expressed in grams per 100 grams of dry weight.

[0068] Use the same ethanol extract as for the soluble sugar determination; transfer one milliliter of the ethanol extract, evaporate the ethanol at 60 degrees Celsius, add one milliliter of water to redissolve the extract, then add one milliliter of dinitrosalicylic acid colorimetric solution, react in a boiling water bath for five minutes, cool, and bring the volume to 10 milliliters. Measure the absorbance at 540 nanometers; each 100 milliliters of the dinitrosalicylic acid colorimetric solution contains 1.0 g of dinitrosalicylic acid, 1.0 g of sodium hydroxide, and 30.0 g of potassium sodium tartrate, dissolved in water and brought to a volume of 100 milliliters, and stored protected from light.

[0069] Using the same aqueous phase colorimetric conditions as the test solution, a standard curve was established with glucose standard solution. Based on the standard curve, the volume of the extract, and the proportion of the colorimetric sample, the glucose equivalent mass of reducing sugar in the total extract was calculated. Calculate the reference value for reducing sugars using the following formula:

[0070]

[0071] In the formula, This represents the glucose equivalent mass of reducing sugars converted to all extracts, in grams. The reference values ​​for reducing sugars are expressed in grams per 100 grams of dry weight; the reference values ​​for soluble sugars are also provided. This indicates the glucose equivalent of sugars that can be extracted by ethanol; reference value for reducing sugars. This represents the glucose equivalent of sugars in the ethanol extract that can undergo a reduction reaction with dinitrosalicylic acid; the two are used as different prediction task labels.

[0072] Weigh 0.500 grams of dried sweet potato powder, add 5 grams of a mixed catalyst composed of potassium sulfate and copper sulfate in a 10:1 mass ratio, and 12 ml of concentrated sulfuric acid. Heat and digest until the digestion solution is clear. After cooling, add excess sodium hydroxide solution for distillation. Absorb the ammonia released during distillation with boric acid solution, and titrate with a standard hydrochloric acid solution of a specified concentration. Calculate the total nitrogen mass in the sample based on the amount of hydrochloric acid standard solution consumed. Use 6.25 as the nitrogen-protein conversion factor and calculate the crude protein reference value according to the following formula:

[0073]

[0074] In the formula, This indicates the total nitrogen mass in the sample, expressed in grams. This indicates the crude protein reference value, in grams per 100 grams dry basis.

[0075] Weigh 1.000 grams of dried sweet potato powder and uniformly degrease all samples using petroleum ether with a boiling range of 30°C to 60°C for six hours using Soxhlet. After degreasing, allow the petroleum ether to evaporate completely, and then perform acid-base digestion. Do not set any criteria for determining whether degreasing has been performed based on the crude fat content. Add 200 ml of 1.25% sulfuric acid solution to the degreased sample, maintain a gentle boil for 30 minutes, filter, and wash with hot water until the filtrate is nearly neutral. Then add 200 ml of 1.25% sodium hydroxide solution to the obtained residue, maintain a gentle boil for 30 minutes, filter, and wash with hot water until the filtrate is nearly neutral.

[0076] The residue after digestion was dried at 105°C to a constant mass, and the mass of the dried residue was recorded. The residue was then ashed at 550°C to a constant mass, and the ash content after ashing was recorded. The reference value for crude fiber was calculated using the following formula:

[0077]

[0078] In the formula, This indicates the mass of the dried residue after cooking. This indicates the ash content of the residue after ashing. This indicates a reference value for crude fiber, in grams per 100 grams dry basis.

[0079] Before generating the amylopectin reference value, the component mass relationship of the total starch reference value and the amylose reference value was verified to ensure that it met the following requirements:

[0080]

[0081] as well as:

[0082]

[0083] When the amylose reference value is less than zero or greater than the total starch reference value, the total starch and amylose of the corresponding sweet potato calibration sample should be re-measured in parallel. If the above component mass relationship is still not met after re-measurement, the sweet potato calibration sample should not be used for model training. Subtract the amylose reference value from the total starch reference value to obtain the amylopectin reference value.

[0084]

[0085] In the formula, This represents the reference value for amylopectin, in grams per 100 grams dry basis; after verification of the component mass relationship, we have:

[0086]

[0087] Subtracting the amylopectin reference value from the amylose reference value yields the amylose-amylopectin differential label:

[0088]

[0089] Will Substituting into the above equation, we get:

[0090]

[0091] In the formula, The amylose-branched starch differential label is expressed in grams per 100 grams of dry basis. With the total starch reference value remaining constant, an increase in the amylose-branched starch differential label indicates an increase in the amylose reference value and a corresponding decrease in the amylopectin reference value; a decrease in the amylose-branched starch differential label indicates a decrease in the amylose reference value and a corresponding increase in the amylopectin reference value. A single amylose-branched starch differential label represents the inverse component changes between amylose and amylopectin, thus avoiding setting amylose and amylopectin as two competing independent prediction tasks.

[0092] All sweet potato calibration samples are grouped according to variety or harvest batch, ensuring that calibration samples of the same variety or harvest batch are in the same data group. There are at least three data groups, and a cross-validation method is used, sequentially selecting one data group as the validation group and the remaining data groups as training groups. In each cross-validation round, only the standard spectra and reference values ​​from the training group are used for near-equal sample pairing, common-mode spectral basis extraction, linear-branched differential-mode spectral basis extraction, and the solution of each prediction task parameter. The validation group does not participate in near-equal sample pairing, spectral basis extraction, or model parameter solution in the corresponding cross-validation round; only the spectral processing parameters, common-mode spectral basis, linear-branched differential-mode spectral basis, and prediction task parameters obtained from the training group are applied to the validation group.

[0093] In each cross-validation round, the training mean of the standard spectra of the training group is calculated according to the wavelength column, and this training mean is subtracted from the standard spectra of both the training and validation groups. The training mean remains unchanged within the current cross-validation round. When finally building the prediction model using all sweet potato calibration samples, the wavelength column mean corresponding to all sweet potato calibration samples is recalculated, and this wavelength column mean is saved as the data processing parameter of the final prediction model. For ease of description, the standard spectral matrix after subtracting the corresponding training mean will still be referred to as... ;

[0094] Within the training group of each cross-validation round, for the first... One sweet potato calibration sample was selected, and the sweet potato calibration sample with the closest total starch reference value was selected from the remaining training samples as the paired sample; the paired sample number was determined according to the following formula:

[0095]

[0096] When there are two or more candidate samples with the same absolute difference in total starch reference values, the values ​​of the candidate samples and the first sample are calculated respectively. The absolute value of the linear-branched differential label difference between each sweet potato calibration sample is used, and the candidate sample with the largest absolute value of the difference is selected as the paired sample.

[0097]

[0098] The second selection relationship is only performed when two or more candidate samples have the same absolute value of the difference between their minimum total starch reference values; the second selection relationship is performed only when two or more candidate samples have the same absolute value of the difference between their minimum total starch reference values. The first sweet potato calibration sample and the first The sweet potato calibration samples were grouped into nearly equal sample pairs, and the corresponding differential label difference was calculated:

[0099]

[0100] All valid near-equal sample pairs are grouped into a valid sample pair set, and the number of valid sample pairs is denoted as . This embodiment uses directional, sample-by-sample pairing records, with each training sample forming one pairing record. When some pairing records are deleted due to invalid spectra, missing reference values, or zero-difference mode label differences, Determined based on the actual number of paired records retained.

[0101] Using the standard spectral matrix of the training group as input and the column vector composed of the total starch reference values ​​of the training group as output, a first partial least squares regression relationship is established; the total starch reference values ​​of the training group are then used to construct a total starch reference vector.

[0102]

[0103] In the formula, This represents the number of sweet potato correction samples in the current training group; let the initial spectral residual matrix be:

[0104]

[0105] Let the initial total starch residual vector be:

[0106]

[0107] In the formula, This represents the average total starch reference value for the training groups. Represents a column vector with all elements equal to one; for the th... 1 latent variable, calculate the spectral weight direction:

[0108]

[0109] when When the extraction of latent variables stops, the number of latent variables successfully extracted is taken as the number of effective latent variables for the current first partial least squares regression relationship; when When, calculate:

[0110]

[0111] Calculate the latent variable score:

[0112]

[0113] when When, stop extracting latent variables; when At that time, calculate the direction of the spectral load:

[0114]

[0115] Calculate the total starch response regression coefficient:

[0116]

[0117] Update the spectral residual matrix:

[0118]

[0119] Update the total starch residual vector:

[0120]

[0121] The successfully extracted spectral load directions are arranged in the extraction order to form the first load matrix:

[0122]

[0123] In the formula, This represents the number of effective latent variables in the first partial least squares regression relationship; before orthogonal decomposition, the zero vector column and columns linearly correlated with the retained columns are deleted from the first loading matrix, and then an economical orthogonal triangular decomposition is used:

[0124]

[0125] In the formula, The columns are mutually orthogonal and have a modulus of one. It was determined to be a common-mode spectral basis; This represents the corresponding upper triangular matrix.

[0126] Within a range not exceeding the rank of the standard spectral matrix of the training group, candidate values ​​for the number of first latent variables are sequentially set; for each candidate value for the number of first latent variables, a first partial least squares regression relationship is established using only the training group, and then the total starch prediction result of the validation group is calculated using the established first partial least squares regression relationship; the root mean square error of the total starch prediction of the validation group is calculated according to the following formula:

[0127]

[0128] In the formula, Indicates the number of samples in the validation group. This represents the predicted total starch value; the root mean square error of total starch validation across all cross-validation rounds is summarized, and the candidate value with the smallest summarized root mean square error is determined as the final number of first latent variables. When multiple candidate values ​​correspond to the same minimum aggregate root mean square error (RMSE), select the candidate value with the smaller value; Number of first latent variables Once determined, the number of second latent variables and the shared representation dimension will not be changed in subsequent selection processes.

[0129] For the training group For each sweet potato calibration sample, the standard spectrum is projected onto the common-mode spectral basis to obtain the common-mode score:

[0130]

[0131] The projection components of the standard spectrum in the common-mode subspace are:

[0132]

[0133] Remove the corresponding projected components from the standard spectrum to obtain the remaining spectrum:

[0134]

[0135] Or it can be expressed as:

[0136]

[0137] In the formula, Indicates the remaining spectrum, Indicates the order is The identity matrix;

[0138] For each valid pair of nearly equal samples, calculate the residual spectral difference between the two calibrated sweet potato samples:

[0139]

[0140] In the formula, , and They represent the first The master sample number and paired sample number are recorded in each valid pairing record; all remaining spectral differences are arranged sequentially to form a remaining spectral difference matrix:

[0141]

[0142] Arrange all the differential modulus label differences sequentially to form the differential modulus label difference vector:

[0143]

[0144] Using the residual spectral difference matrix as input and the difference mode label difference vector as output, a second partial least squares regression relationship is established using the same single-response iteration method as the first partial least squares regression relationship, resulting in the second loading matrix:

[0145]

[0146] In the formula, This represents the number of effective latent variables in the second partial least squares regression relationship; the second loading matrix is ​​projected onto the orthogonal complement space of the common mode subspace:

[0147]

[0148] delete The zero vector column and the columns linearly related to the retained columns are then subjected to an economical orthogonal triangular decomposition:

[0149]

[0150] In the formula, The columns are mutually orthogonal and have a modulus of one. It is determined to be a linear-branched differential mode spectral basis; from the projection relation, we can obtain:

[0151]

[0152] The first The standard spectra of each sweet potato calibration sample are projected onto the straight-chain-branched differential mode spectral basis to obtain the differential mode score:

[0153]

[0154] Two differential mode prediction relationships are established under the same training group, validation group, and second latent variable candidate range. The first differential mode prediction relationship takes the standard spectral difference of nearly equal sample pairs without removing common mode projection components as input and the corresponding differential mode label difference as output. The second differential mode prediction relationship takes the remaining spectral difference after removing common mode projection components as input and the corresponding differential mode label difference as output. The root mean square error of validation for the two differential mode prediction relationships is calculated respectively. When the second differential mode prediction relationship forms a non-zero straight-chain-branched differential mode spectral basis and its root mean square error of validation is not higher than that of the first differential mode prediction relationship, the corresponding candidate value of the number of second latent variables is retained.

[0155] When all candidate differential modes become zero vectors after common mode elimination, or when the second differential mode prediction relationship cannot form an effective differential mode prediction, the number of candidate values ​​for the first latent variable is reduced and the verification of the number of first latent variables is re-executed; the final number of first latent variables retained is still the candidate value that satisfies the differential mode validity condition and has the smallest root mean square error of total starch verification.

[0156] Number of final first latent variables After determining and fixing the number of second latent variables, candidate values ​​for the number of second latent variables are set within a range not exceeding the rank of the remaining spectral difference matrix. For each candidate value for the number of second latent variables, nearly equal sample pairs are constructed using only the training group, and common-mode spectral bases and linear-branch differential-mode spectral bases are extracted. Then, the differential-mode score of the validation group is calculated using the spectral bases obtained from the training group. The root mean square error of the prediction of the linear-branch differential-mode label of the validation group is calculated, and the root mean square error of the validation of all cross-validation rounds is summarized. The candidate value with the smallest summarized root mean square error is determined as the final number of second latent variables. When multiple candidate values ​​correspond to the same minimum aggregate root mean square error (RMSE), select the candidate value with the smaller value; number of second latent variables. Once determined, it will not be changed during the selection of shared representation dimensions;

[0157] Multi-task learning relationships include total starch prediction tasks, linear-branched differential prediction tasks, soluble sugar prediction tasks, reducing sugar prediction tasks, crude protein prediction tasks, and crude fiber prediction tasks; such as Figure 2As shown, standard spectra are input into the common-mode spectral basis projection path, the differential-mode spectral basis projection path, and the shared transformation path, respectively. Specifically, the common-mode score is obtained by projecting the standard spectra onto the common-mode spectral basis, and this score is input into the total starch prediction task to output the total starch prediction value. The differential-mode score is obtained by projecting the standard spectra onto the differential-mode spectral basis, and this score is input into the amylose-branched differential-mode prediction task to output the differential-mode prediction value. The shared spectral characterization is obtained by transforming the standard spectra into shared spectral characteristics, which are input into the soluble sugar prediction task, reducing sugar prediction task, crude protein prediction task, and crude fiber prediction task, respectively, and output the soluble sugar prediction value, reducing sugar prediction value, crude protein prediction value, and crude fiber prediction value, respectively. The total starch prediction value and the differential-mode prediction value are jointly input into the component recovery process to obtain the amylose prediction value and amylopectin prediction value. Figure 2 In this context, no connection is established between the straight-chain-branched differential mode prediction task and the shared spectral characterization, indicating that the straight-chain-branched differential mode prediction task does not receive the shared spectral characterization, thereby reducing the influence of other quality prediction tasks on the differential mode prediction direction.

[0158] The total starch prediction task only accepts the common-mode score, and its raw linear output is:

[0159]

[0160] The straight-chain-branch differential mode prediction task only accepts the differential mode score, and its raw linear output is:

[0161]

[0162] The prediction tasks for soluble sugars, reducing sugars, crude protein, and crude fiber employ shared linear spectral characterization; a reference value matrix is ​​constructed from the four reference values.

[0163]

[0164] Calculate the mean and standard deviation of the four reference values ​​in the training groups, and standardize each column:

[0165]

[0166] Calculate the multi-output least squares coefficients based on the standard spectral matrix and dimensionless reference matrix of the training groups:

[0167]

[0168] Calculate the multi-output regression fitting matrix:

[0169]

[0170] Perform eigenvalue decomposition on the following expression:

[0171]

[0172] Select the first ones in descending order of eigenvalues. The eigenvectors form a shared task direction matrix. The shared parameter matrix is:

[0173]

[0174] No. The shared spectral characterization of the sweet potato calibration samples is as follows:

[0175]

[0176] The four dimensionless prediction results are as follows:

[0177]

[0178] Restore to the original content scale:

[0179]

[0180] in:

[0181]

[0182] Before calculating the normalized task error, it was verified that the variances of the reference values, linear-branch differential label variances, and differential label variances in the training groups were all greater than zero; the total starch normalized task error was:

[0183]

[0184] The normalized task error for straight-chain versus branched differential modulus is:

[0185]

[0186] The normalization task error for soluble sugars was calculated using the same method. Normalization error of reducing sugar task Crude protein normalization task error and the error of coarse fiber normalization task ; Regarding the first For each pair of nearly equal-sized valid samples, calculate the difference in prediction mode:

[0187]

[0188] The differential mode retention error is:

[0189]

[0190] The joint training loss is uniformly defined as:

[0191]

[0192] In the training group, each parameter group is obtained by minimizing the joint training loss; in the validation group, the validation joint loss is obtained according to the same calculation relationship and is used to select the shared representation dimension.

[0193] The parameters for the total starch prediction task are denoted as:

[0194]

[0195] The parameters for the straight-chain-branch differential prediction task are denoted as:

[0196]

[0197] Let the four shared prediction task parameters be denoted as... ;because Only parameter groups , and Only parameter groups The four shared task errors only include the parameter set. Therefore, minimizing the joint training loss can be decomposed into three independent parameter subproblems; dividing by a constant does not change the position of the minimum value of the joint training loss. Each parameter set is obtained by the closed-form solution of the corresponding sub-objective, and together they constitute the joint training result.

[0198] The common mode score and constant of each sweet potato calibration sample are used to construct the total starch design matrix:

[0199]

[0200] The parameter vector for the total starch prediction task is:

[0201]

[0202] In the formula, include and The differential mode design matrix is ​​constructed by combining the differential mode scores and constants of each sweet potato calibration sample.

[0203]

[0204] Establish with OK, The difference matrix of the sample pairs of the column ;No. line in Take one of the columns, in the first column. Take negative 1 for each column and zero for the rest; construct the augmented design matrix:

[0205]

[0206] Construct augmented supervision vectors:

[0207]

[0208] The parameter vector for the differential prediction task is:

[0209]

[0210] In the formula, include and ;

[0211] Number of first latent variables Number of second latent variables Once all dimensions are determined and kept fixed, only the shared representation dimension is considered. A selection is made; the candidate values ​​for the shared representation dimension are one, two, three, and four; for each candidate value, the shared parameter matrix and the shared task direction matrix are obtained using only the training group, and the joint validation loss of the validation group is calculated using all the parameters obtained from the training group; since and At this stage, the total starch, differential mode, and differential mode remain unchanged across different shared representation dimensions. The difference in the joint validation loss is generated by the four shared prediction tasks. The joint validation loss from all cross-validation rounds is summed, and the candidate value with the smallest summed joint validation loss is selected as the final shared representation dimension. When multiple candidate values ​​correspond to the same minimum aggregate validation joint loss, select the candidate value with the smaller value.

[0212] Sure , and Then, the mean calculation of wavelength columns, near-equal sample pairing, common mode spectral basis extraction, differential mode spectral basis extraction and solution of each task parameter were re-executed using all sweet potato calibration samples to obtain the final sweet potato near-infrared multi-index prediction model.

[0213] The original linear output of the total starch prediction task This serves as an intermediate result within the task; physical range processing is used as the output processing layer for the total starch prediction task.

[0214]

[0215] The output obtained after the output processing layer The official output of the total starch prediction task; the raw linear output of the linear-branched differential mode prediction task. This serves as an intermediate result within the task; physical range processing is used as the output processing layer for the differential prediction task.

[0216]

[0217] The output obtained after the output processing layer As the formal output of the linear-branched differential modulus prediction task; calculate the predicted value of amylose:

[0218]

[0219] Calculate the predicted value of amylopectin:

[0220]

[0221] Therefore, the following is satisfied:

[0222]

[0223] The performance of the evaluation model was assessed using validation samples independent of the sweet potato calibration samples. These validation samples were processed using the same sampling, drying, pulverizing, loading, spectral acquisition, and chemical reference detection methods as the sweet potato calibration samples. For any quality index... The root mean square error of the prediction is:

[0224]

[0225] The mean absolute error is:

[0226]

[0227] The coefficient of determination is:

[0228]

[0229] Calculate the closure error of the starch component:

[0230]

[0231] According to the component recovery relationship in this embodiment, theoretically there is... Based on the specific purpose of sweet potato breeding screening, quality grading, or processing raw material evaluation, the allowable prediction error for each quality indicator is determined before model establishment. When the prediction error of each quality indicator in the independently verified sample meets the requirements of the corresponding detection purpose, the established prediction relationship is determined as an usable sweet potato near-infrared multi-index prediction model. When a certain quality indicator does not meet the requirements of the corresponding detection purpose, sweet potato correction samples covering the content range not fully covered by that quality indicator are added, and the model establishment process is repeated.

[0232] In another embodiment, the sweet potato sample to be tested is processed using the same sampling, freeze-drying, crushing, sieving, loading, and spectral acquisition methods as the aforementioned sweet potato calibration sample; the sweet potato sample to be tested is then placed in the... The near-infrared reflectance at each wavelength position is denoted as And convert it to absorbance:

[0233]

[0234] The standard spectrum to be measured is obtained by transforming the variables according to the standard normal distribution method described above. Then, the mean value of the training spectrum at the corresponding wavelength position saved in the final prediction model is subtracted to form the row vector of the standard spectrum to be measured. ; Calculate the common modulus score:

[0235]

[0236] Calculate the differential score:

[0237]

[0238] Calculate the raw output of total starch:

[0239]

[0240] Calculate the original output of the differential mode:

[0241]

[0242] After physical range processing, the predicted values ​​for total starch and differential model are obtained:

[0243]

[0244]

[0245] Calculate the predicted value of amylose:

[0246]

[0247] Calculate the predicted value of amylopectin:

[0248]

[0249] Computation of shared spectral characterization:

[0250]

[0251] Calculate the four dimensionless prediction results:

[0252]

[0253] Restore to the original content scale:

[0254]

[0255] Finally, predicted values ​​for total starch, amylose, amylopectin, soluble sugar, reducing sugar, crude protein, and crude fiber were obtained.

[0256] In one calculation example, the component mass relationships and prediction relationships in this invention are used to verify the calculation of amylopectin reference values, amylose-branched differential labeling, component recovery relationships, and evaluation indicators; the unit of each content data is grams per 100 grams dry basis.

[0257] Table 1 shows the reference and predicted values ​​of starch components;

[0258]

[0259] All data in Table 1 satisfy:

[0260]

[0261] as well as:

[0262]

[0263] Therefore, all twelve samples meet the requirements. ;

[0264] Table 2 shows the reference and predicted values ​​for the remaining four quality indicators;

[0265]

[0266] Table 3 shows the evaluation indicators for the prediction results;

[0267]

[0268] Table 4 shows the modulus preservation results for nearly equal total starch sample pairs;

[0269]

[0270] This invention can also establish an independent near-infrared multi-index prediction model for whole sweet potato tubers. Since the spectra of whole sweet potato tubers and freeze-dried powder have different water content, effective optical path, surface curvature, and scattering conditions, the prediction model for whole tubers should re-establish a calibration sample set, chemical reference values, common-mode spectral basis, and straight-chain-branched differential-mode spectral basis using whole tuber samples, and should not directly use the powder sample prediction model to process the whole tuber spectrum. Without obtaining calibration data and independent verification data for whole tubers, the prediction accuracy and technical effectiveness of this invention will not be demonstrated using whole tuber detection methods. The aforementioned freeze-dried powder sample preparation and spectral acquisition methods constitute a specific implementation method of this invention.

[0271] The embodiments of this example have been described above. However, this example is not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of this example, and all of them are within the protection scope of this example.

Claims

1. A method for modeling multiple indicators of sweet potato near-infrared radiation based on spectral features and multi-task learning, characterized in that, include: Near-infrared reflectance spectra of sweet potato calibration samples were collected and standardized to obtain standard spectra. Reference values ​​for total starch, amylose, soluble sugar, reducing sugar, crude protein, and crude fiber were obtained for each sweet potato calibration sample. Based on the total starch reference value and the amylose reference value, an amylose-branched differential label is generated; for each sweet potato calibration sample, the sample with the closest total starch reference value is selected from the remaining sweet potato calibration samples to form a near-equal sample pair, and the corresponding differential label difference is obtained; The common-mode spectral base is obtained based on the standard spectrum and the total starch reference value. The remaining spectrum is obtained by removing the projection of the standard spectrum into the common-mode subspace. The residual spectral difference is obtained by subtracting the residual spectra of nearly equal sample pairs. Based on the residual spectral difference and the differential mode label difference, the straight-chain-branched differential mode spectral basis orthogonal to the common mode spectral basis is obtained, and the common mode score and differential mode score are calculated. The common mode score is input into the total starch prediction task, the differential mode score is input into the linear-branched differential mode prediction task, and the shared spectral characterization of the standard spectrum is input into the soluble sugar, reducing sugar, crude protein and crude fiber prediction tasks. The multi-task learning relationship is obtained through joint training. Based on the outputs of the first two prediction tasks, the component recovery relationship between amylose and amylopectin was established, resulting in a near-infrared multi-index prediction model for sweet potatoes.

2. The method for near-infrared multi-index modeling of sweet potato based on spectral features and multi-task learning according to claim 1, characterized in that, The acquisition of the standard spectrum and reference values ​​includes: Each sweet potato calibration sample was processed using the same sample preparation method, and the near-infrared reflectance of each sweet potato calibration sample at each wavelength was collected. The absorbance spectrum is obtained by taking the negative logarithm of the near-infrared reflectance at each wavelength position to the base 10. The standard spectrum is obtained by subtracting the average of all absorbance values ​​of the corresponding sweet potato calibration sample from each absorbance value, and then dividing by the standard deviation of all absorbance values ​​of the corresponding sweet potato calibration sample. Measure each reference value, ensuring that all reference values ​​use the same quality standard and the same content unit, and uniformly express them on a dry basis or a wet basis.

3. The method for near-infrared multi-index modeling of sweet potato based on spectral features and multi-task learning according to claim 1, characterized in that, The generation of the straight-chain-branched differential tag includes: Subtract the amylose reference value from the total starch reference value to obtain the amylopectin reference value; Subtracting the amylopectin reference value from the amylose reference value yields the amylose-amylopectin differential label.

4. The method for near-infrared multi-index modeling of sweet potato based on spectral features and multi-task learning according to claim 3, characterized in that, The acquisition of the sum and difference modulus label difference of the nearly equal sample pairs includes: For each sweet potato calibration sample, the sample with the smallest absolute value of the difference between the total starch reference value and that of the sweet potato calibration sample is selected from the remaining sweet potato calibration samples as the paired sample. When there are two or more candidate samples with the same absolute value of the difference in total starch reference value, the sample with the largest absolute value of the difference in the linear-branched differential label with the sweet potato correction sample is selected as the paired sample. The sweet potato calibration sample and the paired sample were combined to form a near-equal sample pair. The linear-branch differential label of the sweet potato calibration sample was subtracted from the linear-branch differential label of the paired sample to obtain the differential label difference.

5. The method for near-infrared multi-index modeling of sweet potato based on spectral features and multi-task learning according to claim 4, characterized in that, The acquisition of the common-mode spectral base and the straight-chain-branched differential-mode spectral base includes: A first partial least squares regression relationship was established based on the standard spectrum and total starch reference value of each sweet potato calibration sample. The loading direction of the first partial least squares regression relationship was normalized and orthogonalized to obtain the common mode spectral basis. The standard spectrum is projected onto the common-mode subspace spanned by the common-mode spectral basis, and the corresponding projected components are removed from the standard spectrum to obtain the remaining spectrum. The residual spectral difference is obtained by subtracting the residual spectra of two sweet potato calibration samples from a nearly equal sample pair. A second partial least squares regression relationship is established based on the remaining spectral difference and the differential mode label difference to obtain the initial straight-chain-branched differential mode spectral basis. The initial straight-chain-branched differential mode spectral basis is projected onto the orthogonal complement space of the common mode subspace, and the projection result is normalized and orthogonalized to obtain the straight-chain-branched differential mode spectral basis orthogonal to the common mode spectral basis. The standard spectrum is projected onto the common-mode spectral basis and the straight-chain-branched differential-mode spectral basis, respectively, to obtain the common-mode score and the differential-mode score.

6. The method for near-infrared multi-index modeling of sweet potato based on spectral features and multi-task learning according to claim 5, characterized in that, The number of latent variables in the first and second partial least squares regression relationships is determined in the following way: Set multiple candidate values ​​for the number of latent variables within a range not exceeding the rank of the corresponding input matrix; The sweet potato calibration samples were divided into training and validation groups according to the sweet potato variety or harvest batch. In each cross-validation, only the sweet potato correction samples in the training group are used to establish the corresponding partial least squares regression relationship, and the established partial least squares regression relationship is used to calculate the prediction results of the validation group. Based on the prediction results of the validation group, the root mean square error of validation corresponding to the candidate values ​​of the number of latent variables is calculated, and the candidate value of the number of latent variables with the smallest root mean square error of validation is determined as the number of latent variables in the corresponding partial least squares regression relationship.

7. The method for near-infrared multi-index modeling of sweet potato based on spectral features and multi-task learning according to claim 5, characterized in that, The establishment of the multi-task learning relationship includes: By setting shared parameters for common transformation of the standard spectrum, a shared spectral characterization can be obtained; The common mode score was input into the total starch prediction task, the differential mode score was input into the linear-branched differential mode prediction task, and the shared spectral characterization was input into the soluble sugar prediction task, reducing sugar prediction task, crude protein prediction task, and crude fiber prediction task, respectively. The straight-chain-branched differential mode prediction task does not receive shared spectral representations, and the common-mode spectral basis and the straight-chain-branched differential mode spectral basis remain unchanged during joint training; The parameters for the total starch prediction task are updated based on the training error of the total starch prediction task. The parameters for the linear-branched differential prediction task are updated based on the training error of the linear-branched differential prediction task. The shared parameters and the parameters for the corresponding prediction tasks are updated based on the training errors of the soluble sugar prediction task, reducing sugar prediction task, crude protein prediction task, and crude fiber prediction task.

8. The method for near-infrared multi-index modeling of sweet potato based on spectral features and multi-task learning according to claim 7, characterized in that, The joint training includes: Calculate the average of the squared differences between the predicted values ​​and the corresponding reference values ​​for total starch, soluble sugar, reducing sugar, crude protein, and crude fiber, and divide each average by the variance of the corresponding reference value to obtain the normalized task error for the corresponding prediction task. Calculate the average of the squared differences between the straight-chain and branch differential modulus predicted values ​​and the straight-chain and branch differential modulus labels, and divide this average by the variance of the straight-chain and branch differential modulus labels to obtain the normalized task error of the straight-chain and branch differential modulus prediction task. The linear-branch differential prediction values ​​of two sweet potato correction samples in a nearly equal sample pair are subtracted to obtain the differential prediction difference. The average of the squared differences between the differential prediction difference and the corresponding differential label difference is calculated, and the average value is divided by the variance of the differential label difference to obtain the differential preservation error. The arithmetic mean of the normalized task error and the differential mode preservation error is used as the joint training loss; Training and validation groups are divided according to sweet potato variety or harvest batch. In each training and validation process, only the parameters of the common mode spectral basis, straight-chain-branched differential mode spectral basis, and multi-task learning relationship are obtained using the training group. The common mode score and differential mode score of the validation group are calculated using the spectral basis, and the parameter state of the multi-task learning relationship is selected according to the joint training loss of the validation group.

9. The method for near-infrared multi-index modeling of sweet potato based on spectral features and multi-task learning according to claim 8, characterized in that, The component recovery relationship includes: The total starch prediction value output by the total starch prediction task and the amylose-branch differential prediction value output by the amylose-branch differential prediction task are combined to determine half of the amylose prediction value. The amylopectin prediction value is determined by half the difference between the total starch prediction value and the amylose-branched differential prediction value. The component recovery relationship and the trained multi-task learning relationship together constitute a sweet potato near-infrared multi-index prediction model, so that the sum of the amylose prediction value and the amylopectin prediction value equals the total starch prediction value.