Rapid nondestructive testing method for internal defects of torreya grandis seeds based on near infrared

By establishing the conversion relationship between the Chinese torreya seed spectrum and the seed kernel spectrum, eliminating the spectral interference between the shell and the black coat, and extracting the characteristic wavelength, the problem of low detection accuracy of internal defects in Chinese torreya seeds is solved, and rapid non-destructive detection and high-precision prediction of internal defects in Chinese torreya seeds is achieved.

CN120213853APending Publication Date: 2025-06-27ZHEJIANG FORESTRY UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510264482.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and non-destructively detect internal defects of torreya seeds, such as empty seeds, black seeds and mold, and the spectral information of the shell and black clothing reduces the detection accuracy.

Method used

By establishing the conversion relationship between the complete Torreya seed spectrum and the Torreya seed kernel spectrum, the spectral information interference between the shell and the black coat is eliminated, and a variety of methods are combined to extract the characteristic wavelengths to establish a prediction model for the internal defects of Torreya seeds.

Benefits of technology

The rapid non-destructive detection of internal defects of Torreya seeds is achieved, the detection accuracy is improved, and the accuracy and stability of the prediction model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FDA0005300816650000022
    Figure FDA0005300816650000022
  • Figure FDA0005300816650000031
    Figure FDA0005300816650000031
Patent Text Reader

Abstract

The invention discloses a method for detecting internal defects of torreya grandis seeds based on near infrared. By establishing the conversion relation between the torreya grandis seed spectrum and the torreya grandis seed kernel spectrum, the torreya grandis seed spectrum is approximately converted into the kernel spectrum, the interference of the spectrum such as the shell is eliminated, and the nondestructive testing precision of the internal defects of the torreya grandis seeds is improved. In addition, the characteristic wavelength of the torreya grandis seeds is extracted by adopting a two-dimensional synchronous spectrum and mutual information measurement method, so that the precision and the stability of the torreya grandis seed internal defect prediction model are further improved. According to the method, multiple internal defects of the torreya grandis seeds can be rapidly and nondestructively detected at the same time, the detection precision is high, and the problems existing in an existing method are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of quality inspection of torreya grandis seeds, and relates to a rapid non-destructive detection method for internal defects of torreya grandis seeds based on near-infrared. Background Art

[0002] Torreya grandis, also known as Chinese torreya, is a tree species native to China under the order Taxales and the family Taxaceae. It is a rare economic tree species in the world. Its fruit, torreya grandis seed, belongs to the category of nuts and has rich nutritional value and unique flavor. However, during the growth and processing of torreya grandis seeds, some common defects may occur, including empty seeds, black seeds, and mildew. Due to growth and development obstacles, or premature picking, the kernel of torreya grandis seeds fails to grow completely, resulting in empty and shriveled kernels, which are called empty torreya grandis seeds. The kernel of torreya grandis contains tannins and can only be eaten after after-ripening treatment. Usually, after harvesting the green fruits of torreya grandis, the stacking method is used to release heat through its own respiration and carry out low-temperature after-ripening. However, there may be a phenomenon of false seed coat rot caused by poor ventilation and too high temperature. At this time, the essential oil in the false seed coat will penetrate into the kernel, making it produce the smell of torreya grandis, which is called black torreya grandis seeds. At the same time, due to changes in environmental factors, torreya grandis seeds may be infected by molds and mildew. The above-mentioned several defects of torreya grandis seeds are not easy or impossible to distinguish from the appearance. There is an urgent need for a rapid non-destructive detection method that can simultaneously detect the internal defects of the above torreya grandis seeds.

[0003] Near-infrared spectroscopy is a rapid, non-destructive, and green optical detection method. The combination frequencies and overtones of the vibrations of various groups of most organic compounds can find corresponding signals in the near-infrared spectrum, so as to obtain information on molecular chemical bonds or functional groups, and reflect the structure and composition of substances based on this.

[0004] The morphological structure of torreya grandis seeds mainly consists of the outer seed coat (shell), the inner seed coat (black skin), and the kernel. When using near-infrared spectroscopy to obtain the spectral information of torreya grandis seeds, the spectrum of torreya grandis seeds contains all the spectral information of the shell, the black skin, and the kernel. Since the main differences in the internal defects of torreya grandis seeds are reflected in the kernel, the spectral information of the shell and the black skin belongs to interference information, which will reduce the detection accuracy of the internal defects of torreya grandis seeds. And separately collecting the spectral information of the kernel of torreya grandis seeds requires shelling the torreya grandis seeds, which is a destructive operation and cannot achieve non-destructive detection. Therefore, it is necessary to adopt a certain method to reduce the interference of the spectral information of the shell and the black skin of torreya grandis seeds, so as to improve the detection accuracy of the internal defects of torreya grandis seeds. In addition, there are many wavelengths in the near-infrared spectral data, and there is redundant and noise information. It is necessary to extract effective characteristic wavelengths from them for establishing a prediction model for the internal defects of torreya grandis seeds to improve the accuracy and stability of the model. Summary of the Invention

[0005] To solve the problems existing in the background technology, the present invention proposes to establish a conversion relationship between the spectra of whole torreya grandis seeds and the spectra of torreya grandis seed kernels, and approximately convert the collected spectra of whole torreya grandis seeds into the spectra of torreya grandis seed kernels through the conversion relationship, so as to eliminate the interference of the spectral information of the torreya grandis seed shell and black skin, and improve the accuracy of rapid non-destructive detection of internal defects of torreya grandis seeds. In addition, the present invention jointly extracts the characteristic wavelengths of internal defects of torreya grandis seeds by multiple methods to further improve the accuracy and stability of the prediction model for internal defects of torreya grandis seeds.

[0006] The technical solution adopted by the present invention to solve its technical problems is: a method for rapid non-destructive detection of internal defects of torreya grandis seeds based on near-infrared, characterized in that the steps of the method are as follows: Step 1: Establishment of a prediction model for internal defects of torreya grandis seeds; Step 1.1, Collect n samples of whole torreya grandis seeds with uniform size and normal appearance; Step 1.2, For the n samples of torreya grandis seeds, collect the near-infrared spectra of each whole torreya grandis seed sample in a semi-transmissive manner in turn under room temperature conditions (25°C, 40% relative humidity); Step 1.3, After the spectral collection is completed, break the shells of the n samples of torreya grandis seeds in turn, and screen out the torreya grandis seeds with empty seeds, black seeds and mildew defects and normal torreya grandis seeds respectively, where there are a1 samples of torreya grandis seeds with empty seed defects, a2 samples of torreya grandis seeds with black seed defects, a3 samples of torreya grandis seeds with mildew defects, and b samples of normal torreya grandis seeds, where a1 + a2 + a3 + b ≤ n; and assign the class values Y of the normal torreya grandis seeds and the samples of torreya grandis seeds with empty seeds, black seeds and mildew seeds to 0, 1, 2 and 3 respectively; Step 1.4, Among the n samples of torreya grandis seeds, select a certain number of torreya grandis seeds with empty seeds, black seeds and mildew defects and normal torreya grandis seeds respectively, with a total of m, and then select the near-infrared spectra corresponding to these m samples, and record the spectral matrix as A m ; Step 1.5, Separate the m shelled samples of torreya grandis seeds in Step 1.4 to obtain the kernels of each torreya grandis seed sample, and then crush them into powders to obtain the kernel powders of m samples of torreya grandis seeds; Step 1.6, For the kernel powders of m samples of torreya grandis seeds, collect the near-infrared spectra of the kernel powders of each sample in turn to obtain the spectral data matrix of the kernel powders, which is recorded as B m ; Step 1.7, In addition to carrying the chemical information of the torreya grandis samples themselves, the near-infrared original spectral data is also doped with other noises and irrelevant information, and is often affected by systematic offsets and scale changes; to improve the comparability and interpretability of the spectral data, for the spectral matrix A m and B mPerform baseline correction and normalization to eliminate the systematic offset caused by baseline drift and the scale change caused by spectral intensity differences, and achieve unified data scales among samples; A m and B m The corrected spectral data matrices are denoted as A' m and B' m ; Step 1.8, since the near-infrared spectra of whole torreya grandis seeds contain information such as the shell and kernel, and the differences in internal defects are mainly reflected in the kernel, the spectral information of the shell belongs to interfering information, which will reduce the detection accuracy of internal defects in torreya grandis seeds; therefore, consider approximately converting the spectra of whole torreya grandis seeds into the spectra of the kernel by constructing a transfer matrix to improve the detection accuracy of internal defects in torreya grandis seeds; for the spectral data matrix A' of m whole torreya grandis seed samples m and the corresponding kernel spectral data matrix B' m , establish the relationship between the two, that is, B' m = A' m ·T, where T is the transfer matrix to be solved; Step 1.9, for B' m = A' m ·T, solve for T by matrix inversion, and get Step 1.10, among n torreya grandis seed samples, after removing the m samples selected in Step 1.4, select a certain number of empty seeds, black seeds, moldy seeds and normal torreya grandis seeds, totaling c, and then select the near-infrared spectra corresponding to these c complete samples, and correct the near-infrared spectra according to Step 1.7. The corrected spectral data matrix is denoted as D c ; Step 1.11, use the transfer matrix T to approximately convert the spectral data matrix D of whole torreya grandis seeds c into the spectral data matrix of the kernel, that is, S c = D c ·T, where S c is the spectral data matrix after conversion by matrix T; Step 1.12, for the data of the spectral matrix S c , calculate the Pearson correlation coefficient between all wavelengths to obtain a correlation coefficient matrix. The correlation coefficient calculation formula is: where c is the number of samples, S kp is the spectral value of the p-th wavelength of the k-th sample, is the spectral mean of the p-th wavelength of the spectra of c samples, and α(p, q) is the correlation coefficient between the p-th wavelength and the q-th wavelength; Step 1.13. For the c sample spectra, based on the correlation coefficient matrix obtained in Step 1.12, perform data transformation to obtain the two-dimensional synchronous spectra of the c samples. The mathematical transformation formula is: Φ(p, q) = Δλ(p, q) × α(p, q), where Δλ(p, q) represents the difference in spectral data between the p-th wavelength and the q-th wavelength, and α(p, q) is the Pearson correlation coefficient between the p-th wavelength and the q-th wavelength calculated in Step 1.12; the two-dimensional synchronous spectrogram converts the correlation information of each pair of wavelengths into spectral features on a two-dimensional plane, which can more effectively reveal the co-variation and correlation between different wavelengths, separate overlapping absorption peaks, identify the co-variation in spectral signals, further improve the spectral resolution, and reveal the differences between samples; Step 1.14. Calculate the average of the two-dimensional synchronous spectra of all normal torreya grandis seeds and defective seed samples among the c samples respectively, then subtract the two-dimensional synchronous average spectrum of the defective seed samples from the two-dimensional synchronous average spectrum of the normal torreya grandis seeds to obtain the difference between the two-dimensional synchronous average spectra of the two. Then select 20 data points with the largest difference values in the two-dimensional synchronous spectra, and select the wavelengths corresponding to the 20 data points as the characteristic wavelengths, denoted as the characteristic wavelength variable set R1; Step 1.15. In order not to lose useful spectral information and ensure that the established prediction model has higher accuracy and robustness, further extract characteristic wavelengths from the sample spectral data; for the spectral matrix S c of the data, use the mutual information method to measure the correlation between the spectral wavelengths and the torreya grandis seed category value Y. The calculation formula of the mutual information method is: where v and Y are the spectral wavelength and the torreya grandis seed category value respectively, p(x, y) is the joint probability, and p(x) and p(y) are their respective marginal probabilities; calculate the mutual information MI(v i between each wavelength v i and the torreya grandis seed category value Y, and then sort the wavelengths in descending order according to the mutual information MI value. Denote the 1st - 10th wavelengths as the characteristic wavelength variable set R2, and the 11th - 50th wavelengths as the candidate wavelength variable set R3; Step 1.16. For the 1st wavelength in R2, calculate the mutual information value between it and each candidate wavelength in R3, and select the wavelength corresponding to the minimum mutual information MI value in R3 as the characteristic wavelength; for the 2nd wavelength in R2, calculate the mutual information value between it and the remaining candidate wavelengths in R3, and select the wavelength corresponding to the minimum mutual information MI value in R3 as the characteristic wavelength; for the 3rd - 10th wavelengths in R2, calculate the mutual information values between them and the remaining candidate wavelengths in R3 in sequence and respectively, and then select the wavelengths corresponding to the minimum mutual information MI values in R3 as the characteristic wavelengths respectively; Step 1.17: Combine the characteristic wavelength variable set R2 with the 10 characteristic wavelength variables selected in Step 1.16, and denote it as the characteristic wavelength variable set R4. Step 1.18: Merge the characteristic wavelength variable set R1 and the characteristic wavelength variable set R4 together, and denote it as R1∪R4. If there are duplicate characteristic wavelengths in the two characteristic wavelength variable sets, denote it as R repetitive ; then eliminate the duplicate characteristic wavelengths to form a new characteristic wavelength variable set R=(R1∪R4)-R repetitive ; Step 1.19: For the spectral matrix S c data, extract the spectral values of the characteristic wavelength variable set R; then use the support vector machine method to associate the spectral values of the characteristic wavelength variable set R with the torreya grandis seed category value Y, and establish a classification prediction model for the internal defects of torreya grandis seeds. The model form is Y = f(R), where Y is the category value of torreya grandis seeds, the category value of normal torreya grandis seeds is 0, and the category values of empty seeds, black seeds, and mildewed seeds of torreya grandis are 1, 2, and 3 respectively; denote the model as Model Y. Step 2: Prediction of internal defects of torreya grandis seeds; Step 2.1: For g torreya grandis seed samples with unknown defects, collect the near-infrared spectra of each sample according to Step 1.2, and denote the spectral matrix as S g ; Step 2.2: Preprocess the spectral matrix data S g according to Step 1.7; then, through the transfer matrix T obtained in Step 1.9, approximately convert the spectral matrix S of the whole torreya grandis seed g into the spectral matrix of the torreya grandis kernel, that is, S′ g =S g ·T, where S′ g is the spectral matrix after being converted by the matrix T; Step 2.3: For the spectral matrix data S′ g , extract the spectral values of its characteristic wavelength variable set R, and then substitute them into the model Model Y established in Step 1.19 to obtain the category prediction values of g torreya grandis seed samples with unknown defects, thereby realizing the rapid non-destructive detection of the internal defects of the whole torreya grandis seeds.

[0007] The beneficial effects of the present invention are as follows: A rapid non-destructive detection method for internal defects of torreya grandis seeds based on near-infrared spectroscopy of the present invention can rapidly and non-destructively detect the internal defects and types of torreya grandis seeds; in addition, by converting the relationship, the spectrum of the whole torreya grandis seed is approximately converted into the spectrum of the torreya grandis kernel, eliminating the interference of the spectral information of the torreya grandis seed shell and black skin, and improving the rapid non-destructive detection accuracy of the internal defects of torreya grandis seeds; and by jointly extracting the characteristic wavelengths of the internal defects of torreya grandis seeds by multiple methods, the accuracy and stability of the internal defect prediction model of torreya grandis seeds are further improved. Detailed implementation manners

[0008] The present invention will be further described below through specific embodiments.

[0009] Embodiment: A rapid non-destructive detection method for internal defects of torreya grandis seeds based on near-infrared, and the specific steps are as follows: Step 1: Establishment of a prediction model for internal defects of torreya grandis seeds; Step 1.1, Collect n samples of complete torreya grandis seeds with uniform size and normal appearance; Step 1.2, For the n samples of torreya grandis seeds, sequentially collect the near-infrared spectra of each complete torreya grandis seed sample in a semi-transmission manner under room temperature conditions (25 °C, 40% relative humidity); Step 1.3, After the spectral collection is completed, sequentially crack the n samples of torreya grandis seeds, and respectively screen out torreya grandis seeds with empty-seed, black-seed and mildew defects and normal torreya grandis seeds, where there are a1 samples of torreya grandis seeds with empty-seed defects, a2 samples of torreya grandis seeds with black-seed defects, a3 samples of torreya grandis seeds with mildew defects, and b samples of normal torreya grandis seeds, where a1 + a2 + a3 + b ≤ n; and assign the class values Y of the normal torreya grandis seeds and the samples of torreya grandis seeds with empty-seed, black-seed and mildew to 0, 1, 2 and 3 respectively; Step 1.4, Among the n samples of torreya grandis seeds, respectively select a certain number of torreya grandis seeds with empty-seed, black-seed and mildew defects and normal torreya grandis seeds, with a total of m, and then select the near-infrared spectra corresponding to these m samples, and the spectral matrix is denoted as A m ; Step 1.5, Perform a separation operation on the m cracked samples of torreya grandis seeds in Step 1.4 to obtain the kernels of each torreya grandis seed sample, and then crush them into powders to obtain the kernel powders of the m samples of torreya grandis seeds; Step 1.6, For the kernel powders of the m samples of torreya grandis seeds, sequentially collect the near-infrared spectra of the kernel powders of each sample, and the spectral data matrix of the kernel powders is denoted as B m ; Step 1.7, In addition to carrying the chemical information of the torreya grandis samples themselves, the near-infrared original spectral data is also doped with other noises and irrelevant information, and is often affected by systematic offsets and scale changes; to improve the comparability and interpretability of the spectral data, perform baseline correction and normalization processing on the spectral matrix A m and B m to eliminate the systematic offset caused by baseline drift and the scale change caused by spectral intensity differences, and achieve the unification of data scales among samples; the corrected spectral data matrices of A m and B m are respectively denoted as A' m and B' m ; Step 1.8. Since the near-infrared spectrum of a whole torreya grandis seed contains information such as the shell and the kernel, and the differences in internal defects are mainly reflected in the kernel, the spectral information of the shell belongs to interfering information, which will reduce the detection accuracy of internal defects in torreya grandis seeds. Therefore, it is considered to approximately convert the spectrum of the whole torreya grandis seed into the spectrum of the kernel by constructing a transfer matrix to improve the detection accuracy of internal defects in torreya grandis seeds. For the spectral data matrix A′ of m whole torreya grandis seed samples m and the corresponding spectral data matrix B′ of the kernels m , establish the relationship between the two, that is, B′ m =A′ m ·T, where T is the transfer matrix to be solved; Step 1.9. For B′ m =A′ m ·T, solve for T by matrix inversion to obtain Step 1.10. Among n torreya grandis seed samples, after removing the m samples selected in Step 1.4, select a certain number of empty seeds, black seeds, moldy seeds, and normal torreya grandis seeds, totaling c. Then select the near-infrared spectra corresponding to these c complete samples and correct the near-infrared spectra according to Step 1.7. The corrected spectral data matrix is denoted as D c ; Step 1.11. Use the transfer matrix T to approximately convert the spectral data matrix D c of the whole torreya grandis seed into the spectral data matrix of the kernel, that is, S c =D c ·T, where S c is the spectral data matrix after conversion by matrix T; Step 1.12. For the data of the spectral matrix S c , calculate the Pearson correlation coefficient between all wavelengths to obtain a correlation coefficient matrix. The calculation formula for the correlation coefficient is: where c is the number of samples, S kp is the spectral value of the p-th wavelength of the k-th sample, is the spectral mean of the p-th wavelength of the spectra of c samples, and α(p, q) is the correlation coefficient between the p-th wavelength and the q-th wavelength; Step 1.13: For the c sample spectra, based on the correlation coefficient matrix obtained in Step 1.12, perform data transformation to obtain the two-dimensional synchronous spectra of the c samples. The mathematical transformation formula is: Φ(p, q) = Δλ(p, q) × α(p, q), where Δλ(p, q) represents the difference in spectral data between the p-th wavelength and the q-th wavelength, and α(p, q) is the Pearson correlation coefficient between the p-th wavelength and the q-th wavelength calculated in Step 1.12. The two-dimensional synchronous spectrogram converts the correlation information of each pair of wavelengths into spectral features on a two-dimensional plane, which can more effectively reveal the co-variation and correlation between different wavelengths, separate overlapping absorption peaks, identify the co-variation in spectral signals, further improve the spectral resolution, and reveal the differences between samples. Step 1.14: Calculate the average of the two-dimensional synchronous spectra of all normal torreya grandis seeds and defective seed samples among the c samples respectively. Then subtract the two-dimensional synchronous average spectrum of the defective seed samples from the two-dimensional synchronous average spectrum of the normal torreya grandis seeds to obtain the difference between the two-dimensional synchronous average spectra of the two. Then select the 20 data points with the largest difference in the two-dimensional synchronous spectra, and select the wavelengths corresponding to the 20 data points as characteristic wavelengths, denoted as the characteristic wavelength variable set R1. Step 1.15: In order not to lose useful spectral information and ensure that the established prediction model has higher accuracy and robustness, further extract characteristic wavelengths from the sample spectral data. For the spectral matrix S c of the data, use the mutual information method to measure the correlation between the spectral wavelengths and the torreya grandis seed class value Y. The calculation formula of the mutual information method is: where v and Y are the spectral wavelength and the torreya grandis seed class value respectively, p(x, y) is the joint probability, and p(x) and p(y) are their respective marginal probabilities. Calculate the mutual information MI(v i , Y) between each wavelength v i and the torreya grandis seed class value Y. Then sort the wavelengths in descending order according to the mutual information MI value. Denote the 1st - 10th wavelengths as the characteristic wavelength variable set R2, and the 11th - 50th wavelengths as the candidate wavelength variable set R3. Step 1.16: For the 1st wavelength in R2, calculate its mutual information value with each candidate wavelength in R3, and select the wavelength corresponding to R3 when the mutual information MI value is the smallest as the characteristic wavelength. For the 2nd wavelength in R2, calculate its mutual information value with the remaining candidate wavelengths in R3, and select the wavelength corresponding to R3 when the mutual information MI value is the smallest as the characteristic wavelength. For the 3rd - 10th wavelengths in R2, calculate their mutual information values with the remaining candidate wavelengths in R3 in sequence and respectively, and then select the wavelengths corresponding to R3 when the mutual information MI values are the smallest as the characteristic wavelengths. Step 1.17: Combine the characteristic wavelength variable set R2 and the 10 characteristic wavelength variables selected in Step 1.16, and denote it as the characteristic wavelength variable set R4. Step 1.18: Merge the characteristic wavelength variable set R1 and the characteristic wavelength variable set R4 together, denoted as R1∪R4. If there are duplicate characteristic wavelengths in the two characteristic wavelength variable sets, denote it as R repetitive ; Then eliminate the duplicate characteristic wavelengths to form a new characteristic wavelength variable set R = (R1∪R4) - R repetitive ; Step 1.19: For the spectral matrix S c data, extract the spectral values of the characteristic wavelength variable set R; Then use the support vector machine method to associate the spectral values of the characteristic wavelength variable set R with the torreya grandis seed category values Y, and establish a classification prediction model for the internal defects of torreya grandis seeds. The model form is Y = f(R), where Y is the category value of torreya grandis seeds, the category value of normal torreya grandis seeds is 0, and the category values of empty seeds, black seeds, and mildewed seeds of torreya grandis are 1, 2, and 3 respectively; Denote this model as Model Y. Step 2: Prediction of internal defects of torreya grandis seeds; Step 2.1: For g torreya grandis seed samples with unknown defects, collect the near-infrared spectra of each sample according to Step 1.2, and denote the spectral matrix as S g ; Step 2.2: Preprocess the spectral matrix data S g according to Step 1.7; Then, through the transfer matrix T obtained in Step 1.9, approximately convert the spectral matrix S of the whole torreya grandis seed g into the spectral matrix of the torreya grandis kernel, that is, S′ g = S g ·T, where S′ g is the spectral matrix after being converted by the matrix T; Step 2.3: For the spectral matrix data S′ g , extract the spectral values of its characteristic wavelength variable set R, and then substitute them into the model Model Y established in Step 1.19 to obtain the category prediction values of g torreya grandis seed samples with unknown defects, so as to realize the rapid non-destructive detection of the internal defects of the whole torreya grandis seed.

[0010] The above specific implementation manners are used to explain and illustrate the present invention, rather than limiting the present invention. Within the spirit and protection scope of the claims of the present invention, any modification and change made to the present invention fall within the protection scope of the present invention.

Claims

1. A near-infrared rapid nondestructive detection method for internal defects of Torreya grandis seeds, characterized in that The steps of the method are as follows: Step 1: Establishment of a prediction model for internal defects of Torreya grandis seeds; Step 1.1, collect n complete Torreya grandis seed samples with uniform size and normal appearance; Step 1.2, for n Torreya grandis seed samples, collect the near infrared spectrum of each complete Torreya grandis seed sample in turn in a semi-transmission mode at room temperature (25°C, 40% relative humidity); Step 1.3, after the spectrum collection is completed, the n Torreya grandis seed samples are shelled in turn, and Torreya grandis seeds with empty seeds, black seeds, and moldy defects and normal Torreya grandis seeds are screened out respectively, among which there are a1 Torreya grandis empty seed defect samples, a2 Torreya grandis black seed defect samples, a3 Torreya grandis moldy seed defect samples, and b normal Torreya grandis seed samples, where a1+a2+a3+b≤n; and the category values ​​Y of normal Torreya grandis seeds and Torreya grandis empty seed, black seed, and moldy seed samples are assigned to 0, 1, 2, and 3 respectively; Step 1.4, among the n Torreya grandis seed samples, select a certain number of empty seeds, black seeds, moldy Torreya grandis seeds and normal Torreya grandis seeds, totaling m, and then select the near infrared spectra corresponding to these m samples, and the spectral matrix is ​​recorded as A m ; Step 1.5, performing a separation operation on the m shelled Torreya grandis seed samples in step 1.4 to obtain the seed kernel of each Torreya grandis seed sample, and then crushing them into powder, thereby obtaining the seed kernel powder of the m Torreya grandis seed samples; Step 1.6: for the seed kernel powder of m Torreya grandis seeds, collect the near infrared spectrum of each sample seed kernel powder in turn, and obtain the spectral data matrix of the seed kernel powder, which is recorded as B m ; Step 1.7, in addition to the chemical information of the Torreya grandis sample itself, the raw near-infrared spectral data is also mixed with other noise and irrelevant information, and is often affected by systematic offset and scale changes; In order to improve the comparability and interpretability of spectral data, the spectral matrix A m and B m Baseline correction and normalization are performed to eliminate the systematic offset caused by baseline drift and the scale change caused by spectral intensity difference, so as to achieve the uniformity of data scale among samples; A m and B m The corrected spectral data matrix is ​​denoted as A′ m and B′ m ; Step 1.8, since the near-infrared spectrum of the complete Torreya grandis seed contains information about the shell and kernel, and the difference in internal defects is mainly reflected in the kernel, the spectral information of the shell is interference information, which will reduce the detection accuracy of the internal defects of Torreya grandis seeds; therefore, consider converting the spectrum of the complete Torreya grandis seed into the spectrum of the kernel by constructing a transfer matrix to improve the detection accuracy of the internal defects of Torreya grandis seeds; for the spectral data matrix A′ of m complete Torreya grandis seed samples m And the corresponding kernel spectrum data matrix B′ m , establish the relationship between the two, namely B′ m =A′ m T, where T is the transfer matrix to be solved; Step 1.9, for B′ m =A′ m T, solve T by matrix inversion, and get Step 1.10, among the n Torreya grandis seed samples, after removing the m samples selected in step 1.4, select a certain number of empty seeds, black seeds, moldy seeds and normal Torreya grandis seeds, totaling c, and then select the near-infrared spectra corresponding to these c complete samples, and calibrate the near-infrared spectra according to step 1.

7. The corrected spectral data matrix is ​​recorded as D c ; Step 1.11, use the transfer matrix T to transform the spectral data matrix D of the complete Torreya grandis seeds c Approximately converted into the seed kernel spectral data matrix, namely S c =D c T, where S c is the spectrum data matrix after transformation by matrix T; Step 1.12, for the spectral matrix S c The Pearson correlation coefficients between all wavelengths are calculated to obtain a correlation coefficient matrix. The correlation coefficient calculation formula is: Where c is the sample size, S kp is the spectral value of the pth wavelength of the kth sample, is the spectral mean of the pth wavelength of c sample spectra, α(p,q) is the correlation coefficient between the pth wavelength and the qth wavelength; Step 1.13, for c sample spectra, based on the correlation coefficient matrix obtained in step 1.12, obtain the two-dimensional synchronous spectra of c samples through data transformation, and the mathematical transformation formula is: Φ(p, q) = Δλ(p, q) × α(p, q), where Δλ(p, q) represents the difference between the spectral data of the p-th wavelength and the q-th wavelength, and α(p, q) is the Pearson correlation coefficient between the p-th wavelength and the q-th wavelength calculated in step 1.12; the two-dimensional synchronous spectrum converts the correlation information of each pair of wavelengths into spectral features on a two-dimensional plane, which can more effectively reveal the coordinated changes and correlations between different wavelengths, separate overlapping absorption peaks, identify coordinated changes in spectral signals, further improve the resolution of the spectrum and reveal the differences between samples; Step 1.14, average the two-dimensional synchronous spectra of all normal Torreya grandis seeds and defective seed samples in c samples respectively, then subtract the two-dimensional synchronous average spectrum of the defective seed sample from the two-dimensional synchronous average spectrum of the normal Torreya grandis seeds to obtain the difference between the two-dimensional synchronous average spectra, then select the 20 data points with the largest difference in the two-dimensional synchronous spectrum, select the wavelengths corresponding to the 20 data points as characteristic wavelengths, and record them as characteristic wavelength variable set R1; Step 1.15: In order not to lose useful spectral information and ensure that the established prediction model has higher accuracy and robustness, the characteristic wavelength is further extracted from the sample spectral data; for the spectral matrix S of c samples c The data was used to measure the correlation between the spectral wavelength and the Torreya grandis seed category value Y using the mutual information method. The calculation formula of the mutual information method is: Where v and Y are the spectral wavelength and Torreya grandis seed category value, respectively, p(x, y) is the joint probability, and p(x) and p(y) are their respective marginal probabilities; calculate each wavelength v i The mutual information MI(v i , Y), then sort the wavelengths from large to small according to the mutual information MI value, record the 1st to 10th wavelengths as the characteristic wavelength variable set R2, and record the 11th to 50th wavelengths as the candidate wavelength variable set R3; Step 1.16, for the first wavelength in R2, calculate the mutual information value between it and each of the candidate wavelengths in R3, and select the wavelength in R3 corresponding to the minimum mutual information M value as the characteristic wavelength; for the second wavelength in R2, calculate the mutual information value between it and each of the remaining candidate wavelengths in R3, and select the wavelength in R3 corresponding to the minimum mutual information MI value as the characteristic wavelength; for the third to tenth wavelengths in R2, calculate the mutual information value between it and each of the remaining candidate wavelengths in R3 in turn, and then select the wavelength in R3 corresponding to the minimum mutual information MI value as the characteristic wavelength; Step 1.17, combining the characteristic wavelength variable set R2 and the 10 characteristic wavelength variables selected in step 1.16, and recording them as characteristic wavelength variable set R4; Step 1.18, merge the characteristic wavelength variable set R1 and the characteristic wavelength variable set R4 together, denoted as R1∪R4. If there are repeated characteristic wavelengths in the two characteristic wavelength variable sets, then denoted as R repetitive ; Then the repeated characteristic wavelengths are removed to form a new characteristic wavelength variable set R = (R1∪R4)-R repedtitive ; Step 1.19, for the spectral matrix S c Data, extract the spectral value of the characteristic wavelength variable set R; then use the support vector machine method to associate the spectral value of the characteristic wavelength variable set R with the category value Y of Torreya grandis seeds, and establish a classification prediction model for Torreya grandis seeds internal defects. The model form is Y=f(R), where Y is the category value of Torreya grandis seeds, the category value of normal Torreya grandis seeds is 0, and the category values ​​of empty Torreya grandis seeds, black seeds, and moldy seeds are 1, 2, and 3 respectively; the model is recorded as ModelY; Step 2: Prediction of internal defects of Torreya grandis seeds; Step 2.1: For g Torreya grandis seeds samples with unknown defects, collect the near infrared spectrum of each sample according to step 1.2, and the spectrum matrix is ​​recorded as S g ; Step 2.2: Follow step 1.7 to process the spectral matrix data S g Preprocessing is performed; then, the spectral matrix S of the complete Torreya grandis seeds is converted into g Approximately converted into the spectrum matrix of Torreya grandis kernel, namely S′ g =S g T, where S′ g is the spectrum matrix after transformation by matrix T; Step 2.3, for the spectral matrix data S′ g ,extract The spectral value of its characteristic wavelength variable set R is then substituted into the model ModelY established in step 1.19 to obtain the category prediction values ​​of g unknown defect Torreya grandis seed samples, thereby realizing rapid non-destructive detection of internal defects of complete Torreya grandis seeds.