Yeast extract amino acid content prediction method based on hyperspectral imaging and data enhancement

By spectral purification and multi-strategy data enhancement of hyperspectral data of yeast extract, the problems of noise and insufficient sample size in the detection of amino acid content in yeast extract were solved, and more stable and efficient prediction results were achieved.

CN121746828APending Publication Date: 2026-03-27GUILIN UNIV OF ELECTRONIC TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511968531.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies for detecting amino acid content in yeast extracts suffer from several drawbacks: spectral data is susceptible to noise and baseline drift, characteristic signal identification is low, and the model's generalization ability is insufficient when the sample size is limited, resulting in poor stability of prediction results.

Method used

By performing spectral purification processing on hyperspectral data, including bias elimination, region screening and feature enhancement, followed by multi-strategy collaborative data enhancement and principal component analysis feature extraction, a core principal component feature matrix of the spectrum is constructed, which is then combined with a specially designed amino acid prediction model.

Benefits of technology

It significantly improved the stability and adaptability of predicting amino acid content in yeast extract, reduced noise and redundant interference, and improved the prediction accuracy and efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746828A_ABST
    Figure CN121746828A_ABST
Patent Text Reader

Abstract

A yeast extract amino acid content prediction method based on hyperspectral imaging and data enhancement comprises the following steps: step 1, performing spectrum purification treatment on original hyperspectral data of target yeast to obtain a purified spectrum; step 2, performing multi-strategy collaborative data enhancement processing on the purified spectrum to obtain an enhanced spectrum, and performing principal component analysis feature extraction processing on the enhanced spectrum to obtain a spectrum core principal component feature matrix; and step 3, predicting the amino acid content of the target yeast based on the spectrum core principal component characteristic matrix. According to the scheme, the problems of noise interference and unobvious features of the original spectrum are solved through spectrum purification, the defect of insufficient model generalization ability under small samples is overcome through multi-strategy collaborative data enhancement and principal component analysis, and finally the prediction result is more reliable and applicable. The requirements on rapid and stable detection of the amino acid content of the yeast extract in industrial production can be better met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of spectral detection technology, and more specifically, to a method for predicting the amino acid content of yeast extract based on hyperspectral imaging and data enhancement. Background Technology

[0002] Yeast extract is an important raw material in the food industry and bio-fermentation fields. Its amino acid content directly affects the nutritional value and application effect of the product. Rapid and accurate content detection is crucial for production quality control.

[0003] In existing technologies, the detection of amino acid content in yeast extracts mostly relies on a scheme combining near-infrared spectroscopy with conventional modeling. This scheme involves acquiring near-infrared spectral data from samples, performing simple preprocessing, and then inputting the data into a regression model to directly output the predicted amino acid content.

[0004] However, this scheme has obvious technical defects: on the one hand, the raw spectral data is susceptible to noise and baseline drift, and lacks targeted purification processing, resulting in low identification of feature signals; on the other hand, no effective data augmentation strategy is designed for the characteristics of spectral data, and the model has insufficient generalization ability when the sample size is limited, making it difficult to adapt to the detection needs of yeast extracts of different batches and different processes, ultimately resulting in poor stability of prediction results. Summary of the Invention

[0005] To address the aforementioned technical problems, this application provides a method for predicting the amino acid content of yeast extract based on hyperspectral imaging and data augmentation, thereby at least alleviating the aforementioned technical problems.

[0006] The technical solutions provided in this application are as follows: A method for predicting the amino acid content of yeast extract based on hyperspectral imaging and data augmentation includes the following steps: Step 1: Perform spectral purification on the raw hyperspectral data of the target yeast to obtain the purified spectrum; Step 2: Perform multi-strategy collaborative data augmentation on the purified spectrum to obtain the enhanced spectrum, and perform principal component analysis feature extraction on the enhanced spectrum to obtain the core principal component feature matrix of the spectrum. The technical solution provided in this application has the following technical advantages: To address the issue of low spectral signal identification in existing technologies, this application employs a technical design that "purifies the raw hyperspectral data of the target yeast to obtain purified spectra." Through a systematic purification process involving bias elimination, region screening, and feature enhancement, interference from ambient light, system noise, and non-sample regions is eliminated at the source. Simultaneously, local spectral features are highlighted and baseline drift is eliminated, allowing the purified spectral data to more accurately reflect the true spectral response of amino acids. This provides a solid foundation of high-quality data for subsequent modeling, fundamentally solving the problem of difficult feature extraction caused by cluttered raw spectral signals, and providing reliable data support for subsequent data enhancement and modeling.

[0007] To address the technical bottleneck of insufficient model generalization ability when sample size is limited, the dual processing design of this application—"multi-strategy collaborative data augmentation processing to obtain enhanced spectra, and principal component analysis feature extraction processing to obtain the core principal component feature matrix of the spectrum"—plays a crucial role. Multi-strategy collaborative data augmentation achieves spectral quality grading through multi-dimensional quality assessment, identifies and locates high-difficulty samples based on difficulty, and then employs targeted augmentation strategies such as interpolation generation and neighborhood perturbation. This ensures the diversity of augmented samples while verifying their physical rationality to ensure they conform to spectral characteristics, effectively expanding the effective sample space. This targeted augmentation method, compared to traditional indiscriminate augmentation, more accurately compensates for sample distribution defects and significantly improves the model's adaptability to yeast extracts from different batches and processes. Meanwhile, principal component analysis feature extraction, while retaining key information, significantly reduces data dimensionality, minimizes redundant information interference in modeling, improves model computational efficiency, and makes the correlation between features and amino acid content more direct.

[0008] This application's technical design, "Predicting the Amino Acid Content of Target Yeast Based on the Spectral Core Principal Component Feature Matrix," relies on high-quality data that has been purified, enhanced, and feature-extracted. Combined with a specially designed amino acid prediction model, it achieves precise correlation between features and amino acid content through multi-layer processing. Compared to traditional direct modeling methods, this design avoids noise and redundant interference in the original data, allowing the model to focus on core features. This results in better stability of the prediction results and effectively reduces prediction bias caused by data quality issues or insufficient model adaptability. Attached Figure Description

[0009] Figure 1 This is a flowchart illustrating a method for predicting the amino acid content of yeast extract based on hyperspectral imaging and data augmentation, as described in this application. Figure 2a A schematic diagram comparing the original spectrum of standard yeast extract with the enhanced spectrum generated by multiple strategies; Figure 2b A schematic diagram illustrating the prediction effect of total amino acid content of standard yeast extract based on the original spectrum; Figure 2c A schematic diagram illustrating the prediction effect of total amino acid content of standard yeast extract based on expanded spectrum; Figure 3 A schematic diagram showing the comparison of the root mean square error in the prediction of the content of 16 amino acids in standard yeast extract; Figure 4a A schematic diagram comparing the original spectrum of the low-hydrolyzed yeast extract with the enhanced spectrum generated by the multi-strategy process. Figure 4b A schematic diagram illustrating the prediction effect of total amino acid content of low-hydrolyzed yeast extract based on the original spectrum; Figure 4c A schematic diagram illustrating the prediction effect of total amino acid content in low-hydrolyzed yeast extract based on expanded spectra; Figure 5 A schematic diagram showing the comparison of the root mean square error in the prediction of the content of 16 amino acids in low-hydrolyzed yeast extract. Figure 6a A schematic diagram comparing the original spectrum and the enhanced spectrum generated by the multi-strategy production of highly hydrolyzed yeast extract; Figure 6b A schematic diagram illustrating the prediction effect of total amino acid content of highly hydrolyzed yeast extract based on the original spectrum; Figure 6c A schematic diagram illustrating the prediction effect of total amino acid content in highly hydrolyzed yeast extract based on expanded spectra; Figure 7 A schematic diagram showing the comparison of the root mean square error in predicting the content of 16 amino acids in highly hydrolyzed yeast extract. Detailed Implementation

[0010] like Figure 1 The image shows a method for predicting the amino acid content of yeast extract based on hyperspectral imaging and data augmentation, comprising the following steps: Step 1, performing spectral purification processing on the original hyperspectral data of the target yeast to obtain the purified spectrum; Step 2, performing multi-strategy collaborative data augmentation processing on the purified spectrum to obtain the enhanced spectrum, and performing principal component analysis feature extraction processing on the enhanced spectrum to obtain the spectral core principal component feature matrix; Step 3, predicting the amino acid content of the target yeast based on the spectral core principal component feature matrix.

[0011] Optionally, step 1 specifically includes the following steps: Step 11: Obtain the dark current spectrum and the white board reference spectrum, perform deviation elimination processing on the original hyperspectral data to obtain the reference correction spectrum; Step 12: Perform region screening processing on the reference correction spectrum to avoid non-sample regions and extract the spectrum of the target region of interest; Step 13: Perform feature enhancement processing on the spectrum of the target region to highlight the local spectral features and eliminate baseline interference to obtain the purified spectrum.

[0012] Optionally, step 11 specifically includes: using the dark current spectrum as the background noise reference and the whiteboard reference spectrum as the standard reflection reference, by calculating the difference between the original hyperspectral data and the dark current spectrum, and then comparing it with the difference between the whiteboard reference spectrum and the dark current spectrum, the deviation between ambient light and system noise is eliminated to obtain the reference correction spectrum.

[0013] Preferably, the specific implementation process of step 11 is as follows: The reference spectrum acquisition layer acquires three types of basic spectral data according to the set acquisition sequence and parameters. The acquisition sequence is set as "dark current spectrum → white board reference spectrum → raw hyperspectral data", and the acquisition interval of the three types of spectra is controlled within a preset time range (e.g., 5-10 minutes) to ensure that the light intensity, temperature and other conditions of the acquisition environment remain consistent; the acquisition parameters of dark current spectrum B are the shutter closed state of the hyperspectral imaging system. At this time, the sensor only records the dark current signal generated by its own electronic thermal motion. Dark current spectrum B is a one-dimensional vector. (N is the total number of spectral wavelengths, Represents the i-th wavelength point The corresponding dark current signal value); the acquisition parameters for the white board reference spectrum W are as follows: a standard diffuse reflection white board (with known reflectivity and stable within the measurement band, e.g., reflectivity ≥ 99%) is placed at the sample acquisition position, and the sensor records the reflectance spectrum signal of the white board. The white board reference spectrum W is a one-dimensional vector. ,in Represents the i-th wavelength point Corresponding whiteboard reflectance signal values; raw hyperspectral data The spectral data of the target yeast extract sample are compared with the dark current spectrum (B) and the white board reference spectrum. Consistent, original hyperspectral data A one-dimensional vector ,in Representing the wavelength points The corresponding original spectral signal value of the sample.

[0014] Preferably, in the specific technical implementation of step 11, the noise dynamic modeling layer receives the dark current spectrum output by the reference spectrum acquisition layer. Whiteboard reference spectrum First, calculate the noise fluctuation coefficient at each wavelength point. This coefficient reflects the intensity difference between system noise and ambient light noise at different wavelengths. The calculation formula is:

[0015] in This is the function for calculating standard deviation. to For the first wavelength points Repeatedly collect dark current signal values ​​( The value range is 3-10, for example The fluctuation characteristics of dark current noise are captured by repeated sampling. For the first The whiteboard reference spectral signal values ​​at each wavelength point are used to normalize the noise fluctuation coefficient, making... It is comparable across different wavelengths; subsequently, based on the noise fluctuation coefficient... Constructing a wavelength-adaptive noise weight matrix The matrix is ​​an N×N diagonal matrix, with the elements on the diagonal being... ,in The weighting adjustment parameter (within the range of 2-8, for example) (), used to control the decay rate of noise weights, noise fluctuation coefficient The larger the value, the higher the corresponding weight. The smaller the value, the stronger the noise interference at that wavelength, requiring a higher suppression weight in subsequent corrections; the noise dynamic modeling layer outputs a wavelength adaptive noise weight matrix. The product of the dark current spectrum B (Matrix multiplication operation) yields the weighted dark current spectrum. ,in The spectrum enhances the dark current signal at low-noise wavelengths and weakens the interference of dark current signals at high-noise wavelengths through weighted allocation.

[0016] Preferably, in one scenario, when step 11 is specifically implemented, the adaptive correction layer receives the raw hyperspectral data. Weighted dark current spectrum Whiteboard reference spectrum First, calculate the effective signal difference at each wavelength point, i.e., the effective signal difference of the sample. ,in Raw hyperspectral data The signal value at the i-th wavelength point, Weighted dark current spectrum In the The signal values ​​at each wavelength point are used to obtain the effective signal difference vector of the sample. Simultaneously calculate the difference in effective signals on the whiteboard. ,in Whiteboard reference spectrum In the Signal values ​​at each wavelength point Dark current spectrum In the The original signal values ​​at each wavelength point are used to obtain the effective signal difference vector of the whiteboard. To avoid differences in the effective signal of the whiteboard The calculation anomaly caused by a value of zero, for Perform minimum value constraint processing and set The minimum effective signal threshold (ranging from 0.01% to 0.1% of the maximum value of the original whiteboard reference spectral signal, for example, 0.05%), when season Subsequently, the adaptive correction layer incorporates a wavelength adaptive noise weight matrix. Perform a weighted ratio correction operation to obtain the reference corrected spectrum. , its first Correction signal value at each wavelength point The calculation formula is:

[0017] The formula uses a weighted fusion of two correction results: the first is a correction result based on weighted dark current spectroscopy, and the second is a traditional ratio correction result. The weights are determined by a wavelength-adaptive noise weight matrix. Control; finally, the reference correction spectrum. Numerical normalization is performed to map the correction signal values ​​at each wavelength point to the [0,1] interval, resulting in the final reference correction spectrum. ,in For the first wavelength points The corresponding final reference-corrected spectral value.

[0018] In this technology, the timing control of the reference spectrum acquisition layer ensures the consistency of the acquisition environment for the three types of spectral data, providing a reliable data foundation for subsequent noise modeling and correction. The fluctuation coefficient calculation and weight matrix construction of the noise dynamic modeling layer enable noise suppression to adapt to the characteristics of different wavelength points, avoiding the problems of signal distortion at low-noise wavelength points or insufficient correction at high-noise wavelength points caused by the traditional method of using a uniform correction strategy for the entire spectrum. The weighted fusion strategy of the adaptive correction layer, while retaining the advantages of the traditional ratio correction method, introduces dynamic adjustment of noise weights, further improving the signal-to-noise ratio and feature fidelity of the reference correction spectrum. The final output reference correction spectrum effectively eliminates the deviation between ambient light and system noise, providing high-quality spectral data support for subsequent target area spectral extraction and feature enhancement.

[0019] Optionally, step 12 specifically includes: displaying the sample image in the RGB imaging mode corresponding to the reference correction spectrum, excluding non-sample areas including at least the sample disk edge and shadows by threshold segmentation, locating and extracting the spectrum of the effective sample area, and obtaining the spectrum of the target area.

[0020] First, the spectral-RGB mapping layer receives the reference correction spectrum output from step 11. The reference correction spectrum An M×N dimensional matrix ( The total number of imaging pixels, (Total number of wavelength points), elements in the matrix Representing the The pixel in the first wavelength points The baseline calibrated spectral value; the spectral-RGB mapping layer first filters characteristic wavelength points that match the RGB three primary color bands, and selects the wavelength corresponding to the red band. (Values ​​range from 620-750nm, e.g., 650nm), wavelength of the green band (Values ​​range from 495-570nm, e.g., 532nm), wavelength of the blue band (Values ​​range from 450-495 nm, e.g., 470 nm), from the reference calibration spectrum. Extract the spectral data from these three wavelength points to form an M×3 dimensional characteristic wavelength spectral matrix. Elements in the matrix , , ,in , , They are respectively , , The corresponding wavelength point index; subsequently, the characteristic wavelength spectral matrix... Normalization is performed to map the three-channel spectral values ​​of each pixel to an 8-bit grayscale range [0, 255], resulting in an RGB image matrix. Its element calculation formula is:

[0021] in Corresponding to the R, G, and B channels respectively. For rounding operations, For minimum value operation, For maximum value calculation, this RGB image matrix The grayscale difference between the sample and the background in the reference calibration spectrum was fully preserved.

[0022] The multi-feature fusion threshold generation layer receives an RGB image matrix. Corrected spectrum with reference First, calculate the RGB image matrix. The gray-level co-occurrence matrix texture features are analyzed, selecting three texture parameters: contrast, energy, and entropy. Contrast reflects the degree of difference in gray levels in the image, energy reflects the uniformity of gray-level distribution, and entropy reflects the richness of image information. The calculation of the three texture parameters is based on pixel neighborhood (neighborhood size is set to 3×3 or 5×5, for example, 3×3); simultaneously, the reference correction spectrum is calculated. Spectral mean of each pixel The calculation formula is:

[0023] Where mean The mean operation is performed to obtain the pixel spectral mean vector. Subsequently, a multi-feature fusion vector is constructed. ,in RGB image matrix No. The average grayscale value of each pixel is calculated using the following formula:

[0024] These are the contrast, energy, and entropy texture parameters of the m-th pixel, respectively; based on a multi-feature fusion vector. Construct an M×5 dimensional feature matrix, perform principal component analysis on the feature matrix, extract the first principal component PC1 (cumulative variance contribution rate not less than 85%), and obtain the principal component vector. The Otsu algorithm (maximum inter-class variance method) is used to perform threshold segmentation on the principal component vector PC1, and the optimal segmentation threshold is calculated. The threshold This maximizes the inter-class variance between the sample region and the non-sample region, and includes principal component vectors PC1 that are greater than or equal to... The pixels marked as candidate sample pixels are smaller than The pixels are marked as candidate background pixels to obtain the initial binary mask. Its elements (candidate sample pixels) or (Candidate background pixels).

[0025] The region morphology optimization layer receives the initial binary mask. First, an erosion operation is performed to remove small noise regions. The erosion operation uses a circular structuring element (with a radius set to 1-3 pixels, for example, 2 pixels). The resulting mask... By removing isolated pixels at the edges of pixel clusters, fine shadows and noise points at the edges of the sample disk are eliminated; the etched mask... Perform the dilation operation using a circular structuring element of the same size as the erosion operation; the resulting mask... It can fill in tiny holes and boundary defects inside the sample area; then, it performs connected component analysis, calculates the area and roundness of each connected region in the mask, and identifies connected regions with an area less than a set threshold (set to 0.5%-1% of the total number of pixels, for example 0.8%) as pseudo sample regions, and connected regions with a roundness less than a set threshold (roundness calculation formula is as follows) as non-sample regions (such as irregular areas on the edge of the sample disk):

[0026] The roundness threshold is set to 0.6-0.8 (e.g., 0.7). After removing these spurious sample areas and non-sample areas, an optimized binary mask is obtained. Finally, the optimized binary mask is... Boundary smoothing is performed using a Gaussian filter (kernel size set to 3×3, standard deviation set to 0.5-1.5, e.g., 1.0) to obtain the final region mask. The mask Pixels marked as 1 constitute a complete and continuous effective sample area, while pixels marked as 0 constitute non-sample areas (including sample disk edges, shadows, etc.).

[0027] The spectral extraction layer receives the final region mask. Corrected spectrum with reference First, filter the final region mask. The set of pixel indices marked as 1 ( (Total number of pixels in the effective region of the sample); based on the index set From the reference correction spectrum Extract the spectral data of the corresponding pixels to obtain the K×N dimensional spectral matrix of the target region. Its element calculation formula is:

[0028] in , Each row of this matrix represents the complete spectrum of a pixel in the effective region of a sample, and each column represents the spectral value of a wavelength point across all effective pixels; for the target region spectral matrix Perform row mean operation to obtain the average spectrum of the target region in 1×N dimensions. Its element calculation formula is:

[0029] The average spectrum eliminates spectral fluctuations of pixels within the effective area of ​​the sample, while retaining the core spectral features related to amino acid content, which is the target region spectrum finally output in step 12.

[0030] This application's technology differs from traditional single grayscale threshold segmentation methods. Through a designed multi-feature fusion threshold generation mechanism, it combines spectral features with image texture features, avoiding segmentation errors caused by traditional methods relying solely on grayscale information. This is particularly suitable for scenarios where yeast extract samples have low contrast with the background. The combined operations of erosion-dilation-connected region analysis-boundary smoothing in the region morphology optimization layer accurately remove non-sample areas such as sample plate edges and shadows, while simultaneously filling in sample area defects, ensuring the integrity and accuracy of the effective sample area. The feature wavelength screening and normalization processing in the spectral-RGB mapping layer achieves a high-fidelity conversion of spectral data to RGB images, providing a reliable image foundation for subsequent feature extraction. The pixel screening and mean calculation in the spectral extraction layer further purifies the sample spectrum and reduces interference from non-effective region spectra. In this technology, the construction of multi-feature fusion vectors allows the segmentation threshold to adapt to the spectral and image characteristics of different samples. The introduction of principal component analysis reduces feature dimensionality and enhances the difference between the sample and the background. The final output target region spectrum provides high-quality data support for subsequent feature enhancement processing.

[0031] Optionally, step 13 specifically includes: calculating the spectral value difference of the target region at adjacent wavelength points, enhancing the local features of the spectral absorption peaks and reflection peaks through differential operations, and eliminating baseline drift interference to obtain the purified spectrum after feature enhancement.

[0032] Preferably, the specific implementation process of step 13 is as follows: the spectral local feature sensing layer receives the target region spectrum output in step 12. The spectrum of the target region A 1×N dimensional vector ( (Total number of wavelength points), elements in the vector Representing the wavelength points The corresponding spectral reflectance or absorptivity values; the spectral local feature sensing layer first calculates the spectrum of the target region. The local fluctuation intensity is determined by iterating through all wavelength points using a sliding window (window size set to 5-11 wavelength points, e.g., 7 wavelength points), and the standard deviation of the spectral values ​​within each window is calculated. ( ), to obtain the local wave intensity vector Based on local wave intensity vector Divide the region into a characteristic region and a flat region, and set a fluctuation threshold. ( To adjust the coefficient, the value range is 1.2-2.0 (e.g., 1.5). The wavelength point is determined as a characteristic region point. The wavelength points are identified as points in a flat region, forming a region identifier vector. ,in Representative feature region points, Points representing flat regions; simultaneously calculating the spectrum of the target region. global trend slope ( (The actual wavelength value at the nth wavelength point), the slope of this global trend. Used to characterize the overall direction and intensity of baseline drift changes.

[0033] Preferably, in the specific technical implementation of step 13, the target region spectrum output by the adaptive weighted micro-layer receiving spectral local feature sensing layer. Local wave intensity vector The weighted differential formula designed with respect to the region identifier vector Flag is as follows:

[0034] in, For the first The weighted differential results at each wavelength point The adaptive weighting coefficients are calculated as follows: , The weights for the feature regions (with values ​​ranging from 1.5 to 2.5, for example, 2.0) The basic weights for the smooth regions (within the range of 0.5-1.0, for example, 0.8) Local wave intensity vector The maximum value; The wavelength interval between adjacent wavelength points is denoted as . If the hyperspectral data is sampled at equal intervals, then... For fixed value At this point, the denominator can be simplified to The weighted differential vector is calculated using this formula. The differential results corresponding to the feature region points in the vector are amplified, while the differential results corresponding to the smooth region points are adaptively adjusted according to the fluctuation intensity. This not only strengthens the edge difference between the absorption peak and the reflection peak, but also avoids the excessive amplification of noise in the smooth region.

[0035] Preferably, in one scenario, when step 13 is specifically implemented, the baseline drift cancellation layer receives the weighted differential vector D output by the adaptive weighted micro-layer, and the global trend slope output by the spectral local feature perception layer. First, a baseline trend model is constructed based on the set of points in the flat areas; then all... For points in the flat region where (n)=0, extract their spectral and wavelength values ​​to construct a dataset. The baseline fitting function is obtained by using linear fitting or low-order polynomial fitting (the polynomial degree is set to 1-3, for example, 2). (m is the number of fitting iterations, (These are the fitting coefficients); calculate the differential curve of the baseline fitting function. The baseline differential vector is obtained. This vector characterizes the differential interference caused by baseline drift; the weighted differential vector The net differential vector is obtained by canceling out the baseline differential vector B'. ,in The offsetting factor is determined by the intensity of the local fluctuations. , Local wave intensity vector The maximum value of this design reduces the baseline cancellation intensity in feature regions and increases the baseline cancellation intensity in smooth regions, ensuring that the baseline elimination process does not damage feature information.

[0036] Preferably, in the specific implementation of step 13, the feature fusion enhancement layer receives the net differential vector output by the baseline drift cancellation layer. and the region identifier vector output by the spectral local feature perception layer. First, consider the net differential vector. To complete the boundary, add an element at the beginning and end that is equal to the adjacent differential value, resulting in the completed net differential vector. To make its dimensions match the spectrum of the target region Consistency; Constructing a feature enhancement matrix Elements in the matrix The calculation method is as follows ,in This is the feature region enhancement coefficient (with a value range of 0.8-1.5, for example, 1.2). The enhancement coefficient for smooth regions (range 0.3-0.7, e.g., 0.5); for the feature enhancement matrix. Smoothing constraint processing is performed using a three-point Savitzky-Golay filter. Convolution operations are performed, and the filter coefficients are obtained through least-squares fitting to ensure that feature sharpness is preserved while noise interference is reduced; the final output is a 1×N dimensional vector. This is the spectrum after purification. The local features of the amino acid-related absorption and reflection peaks in this spectrum are significantly enhanced, and the overall shift caused by baseline drift is completely eliminated.

[0037] This application differs from traditional fixed-coefficient differentiation methods by employing a local spectral feature sensing layer to intelligently segment spectral regions, enabling differentiation operations to accurately adapt to the spectral characteristics of different regions. The adaptive weighted differentiation layer dynamically adjusts its weight coefficients based on local fluctuation intensity, addressing the technical challenge of traditional differentiation methods amplifying noise while enhancing features. The baseline drift cancellation layer, based on a baseline model constructed from smooth regions, accurately captures baseline change trends, and the dynamic adjustment mechanism of the cancellation coefficients prevents over-correction of feature regions. The feature fusion enhancement layer achieves a balance between feature enhancement and noise suppression through a combination of dual-coefficient enhancement and smoothing constraints. In this implementation, the introduction of local fluctuation intensity vectors allows all processing parameters to be adaptively generated based on the spectral quality itself, requiring no manual intervention. This approach is suitable for high-quality target region spectra and can also be used to specifically optimize low-quality spectra. The purified output spectrum provides high signal-to-noise ratio and high feature recognition as foundational data for subsequent data enhancement and feature extraction.

[0038] Optionally, step 2 specifically includes the following steps: Step 21: Perform multidimensional quality assessment on the purified spectrum to obtain spectral quality rating results; Step 22: Based on the spectral quality rating results, perform prediction difficulty identification on the purified spectrum to obtain high-difficulty spectral labeling results; Step 23: For spectra with different quality levels and difficulty labels, perform interpolation generation, neighborhood perturbation, distribution completion, and density filling processing respectively to obtain multi-type enhanced spectrum sets; Step 24: Perform physical rationality verification on the multi-type enhanced spectrum sets and select compliant enhanced spectra that meet the spectral characteristics as enhanced spectra.

[0039] Optionally, step 21 specifically includes: calculating the numerical stability, derivative continuity, and signal-to-noise ratio of the purified spectrum; calculating the comprehensive quality score by weighting the three indicators according to preset weights; dividing the spectrum into high-quality, medium-quality, and low-quality levels according to the score range to obtain the spectral quality rating result.

[0040] Preferably, the specific implementation process of step 21 is as follows: the multidimensional evaluation layer of spectral quality receives the purified spectrum output from step 13. The spectrum after purification treatment Given an M×N dimensional matrix (M being the number of spectral samples and N being the total number of wavelength points in each spectrum), the elements in the matrix... This represents the m-th spectral sample at the n-th wavelength. The corresponding spectral values; the multidimensional evaluation layer of spectral quality first calculates the numerical stability index. For each spectral sample m, a sliding window (the window size is set to 3-9 wavelength points, for example, 5 wavelength points) is used to traverse all its wavelength points, and the coefficient of variation of the spectral values ​​within each window is calculated. ( Index of the center wavelength point of the window. For window size, For standard deviation calculation, (For mean calculation), each spectral sample is obtained. coefficient of variation sequence Take the coefficient of variation sequence The mean value is used as an index of the numerical stability of the spectral sample. Numerical stability index The smaller the value, the smoother the fluctuation of the spectral sample values, and the better the quality.

[0041] Preferably, in the specific technical implementation of step 21, the multidimensional evaluation layer of spectral quality continues to calculate the derivative continuity index for each spectral sample. First, calculate its first derivative sequence. ( ), The wavelength interval between adjacent wavelength points; the second derivative sequence is calculated based on the first derivative sequence D1(m). ( This second derivative sequence reflects the rate of change of the spectral slope; a derivative continuity index is defined. A sequence of second derivatives The absolute value of the score is calculated as follows: ,in A sequence of second derivatives Standard deviation, derivative continuity index The larger the value, the smoother the slope of the spectrum, the more regular the contours of the absorption and reflection peaks, and the higher the feature recognition. At the same time, in order to adapt to the characteristic band characteristics of the amino acid spectrum of yeast extract, when calculating the derivative sequence, the derivative of the characteristic wavelength range of amino acids (e.g., a specific sub-range in the range of 400-800nm) is given a weight of 1.2-1.5 times (e.g., 1.3 times), so that the continuity of this range is reasonably highlighted in the index.

[0042] Preferably, in one scenario, when step 21 is specifically implemented, the multi-dimensional evaluation layer of spectral quality calculates the signal-to-noise ratio index, which is characterized by the ratio of signal power to noise power. First, a smooth region without obvious absorption peaks or reflection peaks is selected in each spectral sample m (using the region identification vector in step 13). Sure, (Wavelength points form a flat region), from which the spectral value sequence of this region is extracted. ; Calculate the spectral value sequence of the flat region Standard deviation as noise power Then extract the absorption and reflection peak regions associated with amino acid characteristics in the spectral sample m. The wavelength points constitute a characteristic region, and the peak-to-peak value of the spectral values ​​in this region is calculated. As a signal power; signal-to-noise ratio indicator (Dividing by 2 is because the peak-to-peak value corresponds to the double amplitude of the signal), signal-to-noise ratio (SNR) The larger the value, the higher the proportion of effective feature signals in the spectrum relative to noise, and the stronger the reliability of feature extraction.

[0043] Preferably, in the specific implementation of step 21, the adaptive weight allocation layer receives the numerical stability index output by the spectral quality multidimensional evaluation layer. Derivative continuity index With signal-to-noise ratio First, the three indicators are normalized to eliminate dimensional differences; numerical stability index Using reverse normalization, the calculation formula is as follows:

[0044] in This represents the maximum value of the numerical stability index for all spectral samples. The minimum value is chosen so that a larger normalized value indicates better stability; the derivative continuity index... With signal-to-noise ratio Using forward normalization, the calculation formulas are as follows:

[0045]

[0046] Based on the scenario requirement of predicting the amino acid content of yeast extract, a dynamic weight allocation rule is designed, and the weight vector is... The method for determining this is as follows: calculate the mean signal-to-noise ratio of all spectral samples. When a certain spectral sample hour, (e.g., 0.25) (e.g., 0.35) (e.g., 0.4), highlighting the weight of the signal-to-noise ratio; when hour, (e.g., 0.35) (e.g., 0.45) (For example, 0.2), to increase the weight of numerical stability and derivative continuity; the comprehensive quality score of each spectral sample is obtained by weighted summation, and the calculation formula is:

[0047] Overall quality score The value range is [0,1].

[0048] Preferably, in the specific technical implementation of step 21, the quality level classification layer receives the comprehensive quality score output by the adaptive weight allocation layer. Quality levels are determined based on statistical distribution characteristics; a comprehensive quality score is calculated for all spectral samples. mean with standard deviation Set a high-quality level threshold (e.g., 0.7 times) Low quality threshold (e.g., 0.7 times) ); The overall quality score The spectral samples are classified into high-quality grades. The quality is classified as medium. The samples are classified as low quality. To avoid imbalances in the classification due to extreme values, boundary constraints are applied to the thresholds to ensure that the proportions of high-quality and low-quality spectral samples are controlled between 20%-30% (e.g., 25%) and 15%-25% (e.g., 20%), respectively. If the proportions exceed these ranges, the values ​​of T1 and T2 are adjusted proportionally. The final output is a spectral quality rating result containing the quality level identifier for each spectral sample. .

[0049] This application's technology differs from traditional fixed-weight spectral quality assessment methods. Through a targeted index calculation method designed in a multi-dimensional spectral quality assessment layer, it directly correlates numerical stability, derivative continuity, and signal-to-noise ratio with the spectral characteristics of amino acids in yeast extracts. In particular, the differentiated processing of characteristic bands enhances the assessment's adaptability to various scenarios. The adaptive weight allocation layer, based on a dynamic weight adjustment rule using the signal-to-noise ratio, addresses the issue of varying importance of indicators in spectral samples of different quality, ensuring that the overall quality score better reflects the actual contribution of the spectrum to subsequent prediction tasks. The quality grade classification layer, combining statistical distribution and proportion-constrained threshold setting, avoids the limitations of a single threshold classification, ensuring the rationality of the grade distribution. In this technical implementation, the three assessment indicators comprehensively characterize spectral quality from three dimensions: numerical fluctuation, morphological regularity, and signal-to-noise ratio. Weight allocation and grade classification are adaptively generated based on the data's inherent characteristics. The output spectral quality rating results can accurately guide subsequent multi-strategy data augmentation, enabling efficient utilization of high-quality spectra and targeted optimization of low-quality spectra, providing a fundamental guarantee for improving the accuracy of amino acid content prediction.

[0050] Preferably, the specific implementation process of step 22 is as follows: The prediction difficulty recognition layer receives the spectral quality rating result Grade and the corresponding purified spectral set S_refine output in step 21. The spectral set S_refine is an M×N dimensional matrix (M is the total number of spectral samples, and N is the number of wavelength points in each spectrum). The matrix element S_refine(m,n) represents the spectral value of the m-th spectral sample at the n-th wavelength point λ_n. At the same time, it receives the true value vector of amino acid content Y_true corresponding to each spectral sample (Y_true is an M×1 dimensional vector, and Y_true(m) is the true value of amino acid content of the m-th spectral sample). The prediction difficulty recognition layer divides the spectral set into a high-quality spectral subset S_high, a medium-quality spectral subset S_mid, and a low-quality spectral subset S_low according to the spectral quality rating result Grade. These correspond to spectral samples with Grade(m) of high quality, medium quality, and low quality, respectively. The number of samples in each subset is M_high, M_mid, and M_low, respectively, and M_high + M_mid + M_low = M.

[0051] Preferably, in the specific technical implementation of step 22, the adaptive ridge regression model layer loads a preset model structure, which includes an input feature adaptation module, a regularization parameter dynamic adjustment module, and a regression calculation module. The input feature adaptation module receives the spectral subsets of each quality level output by the prediction difficulty recognition layer. For each spectral sample m0 to be evaluated (m0∈{1,2,...,M}), it extracts the remaining M-1 spectral samples to form a training spectral matrix X_train (with dimensions (M-1)×N), and the corresponding true values ​​of amino acid content form a training label vector Y_train (with dimensions (M-1)×1). The regularization parameter dynamic adjustment module calculates the dynamic regularization parameter λ based on the quality level distribution corresponding to the training spectral matrix X_train, and the calculation method is λ=λ0× λ0 is the basic regularization parameter (ranging from 0.01 to 0.1, e.g., 0.05), P_high is the proportion of high-quality spectral samples in the training spectral matrix X_train, P_mid is the proportion of medium-quality samples, P_low is the proportion of low-quality samples, and α, β, and γ are quality weight coefficients (α ranges from 0.2 to 0.3, β ranges from 0.3 to 0.4, and γ ranges from 0.4 to 0.5, e.g., α=0.25, β=0.35, γ=0.4). The higher the proportion of low-quality samples, the larger the regularization parameter λ, in order to enhance the model's ability to suppress noise.

[0052] Preferably, in a scenario, when step 22 is specifically implemented, the regression calculation module of the adaptive ridge regression model layer constructs a regression equation and solves for the regression coefficient vector W based on the training spectral matrix X_train, the training label vector Y_train, and the dynamic regularization parameter λ. The objective function of the regression equation is:

[0053] Where ||·||_2 is the L2 norm, the first term is the sum of squares of the model prediction errors, and the second term is the L2 regularization term, used to limit the magnitude of the regression coefficient vector W to avoid model overfitting; the regression coefficient vector is obtained by solving the least squares method. , where I_N is an N×N dimensional identity matrix, and X_train^T is the transpose of the training spectral matrix X_train; the regression calculation module inputs the spectral vector X_test (dimension 1×N, i.e. S_refine(m0,:)) of the spectral sample m0 to be evaluated into the regression model obtained by solving, and calculates the predicted value of amino acid content Y_pred(m0)=X_test×W for the sample.

[0054] Preferably, in the specific implementation of step 22, the error evaluation identification layer receives the predicted value vector Y_pred (dimension M×1) of all spectral samples to be evaluated output by the adaptive ridge regression model layer. First, it calculates the absolute error E_abs(m) = |Y_pred(m) - Y_true(m)| for each sample, and then performs hierarchical normalization on the absolute error. The hierarchical normalization is performed according to the spectral quality level. For high-quality spectral samples, the normalization error E_norm_high(m) = E_ Where E_abs_high is the set of absolute errors for all high-quality spectral samples. This is the maximum value of the set; for medium-quality spectral samples, E_norm_mid(m) = E_ E_abs_mid is the set of absolute errors for medium-quality spectral samples; for low-quality spectral samples, E_norm_low(m) = E_ E_abs_low represents the set of absolute errors for low-quality spectral samples. The error assessment label layer sets stratified thresholds based on stratified normalized errors. The error threshold T_high for high-quality spectral samples ranges from 0.2 to 0.3 (e.g., 0.25), the error threshold T_mid for medium-quality spectral samples ranges from 0.3 to 0.4 (e.g., 0.35), and the error threshold T_low for low-quality spectral samples ranges from 0.4 to 0.5 (e.g., 0.45). The thresholds are set based on the prediction error distribution characteristics of spectral samples of different qualities, and higher normalized error thresholds are allowed for low-quality spectral samples.

[0055] Preferably, in the specific technical implementation of step 22, the error evaluation labeling layer compares the normalized error of each spectral sample m with the error threshold of the corresponding quality level. If sample m is a high-quality spectral sample and E_norm_high(m)>T_high, or a medium-quality spectral sample and E_norm_mid(m)>T_mid, or a low-quality spectral sample and E_norm_low(m)>T_low, then the sample is marked as a high-difficulty spectrum, and a high-difficulty spectrum labeling vector Flag_hard (with dimensions M×1) is generated. Flag_hard(m)=1 indicates that the m-th sample is a high-difficulty spectrum. hard(m)=0 indicates a non-high-difficulty spectrum. Simultaneously, the error assessment labeling layer calculates the proportion of high-difficulty spectra in each quality level. If the proportion of high-difficulty spectra in a certain quality level exceeds a set ratio (ranging from 30% to 40%, for example, 35%), the error threshold for that level is dynamically corrected using the formula T'=T×(1+0.1×(R-R0)), where R is the current proportion of high-difficulty spectra and R0 is the set ratio, ensuring the threshold adapts to the actual error distribution. Finally, the output includes a high-difficulty spectrum label for each spectral sample, which is a combination of the Flag_hard vector and the corresponding normalized error matrix (with dimensions M×1).

[0056] This application differs from traditional ridge regression models with fixed regularization parameters. By employing a dynamic adjustment module for regularization parameters in the adaptive ridge regression model layer, it incorporates the spectral quality level distribution in the training dataset into the parameter calculation. This ensures that the regularization strength matches the data quality, enhancing noise suppression when low-quality samples constitute a higher proportion and preserving more feature information when high-quality samples constitute a higher proportion. The error assessment and labeling layer employs a hierarchical normalization error and hierarchical threshold strategy, avoiding the bias of a single threshold for evaluating spectral samples of different quality levels. This makes the identification of high-difficulty spectra more closely aligned with the predictive characteristics of spectra at each quality level. The quality level classification preprocessing in the prediction difficulty identification layer provides a targeted foundation for subsequent model parameter adjustments and error assessment, ensuring that the entire identification process revolves around the quality differences in the amino acid spectra of yeast extracts. In this implementation, leave-one-out cross-validation ensures that the prediction difficulty assessment for each sample is based on an independent training set, avoiding mutual interference between samples. The design of dynamic regularization parameters and hierarchical thresholds addresses the insufficient adaptability of traditional methods to spectral quality differences. The output high-difficulty spectra labeling results can accurately locate samples requiring focused enhancement, laying the foundation for improving the overall accuracy of amino acid content prediction.

[0057] Optionally, step 23 specifically includes: selecting spectral pairs from the purified spectra of high-quality grades, calculating the Pearson correlation coefficient of the spectral pairs, and only retaining the spectral pairs with a correlation coefficient greater than the set value; adaptively allocating interpolation weights according to the difference in the comprehensive quality scores of the spectral pairs, and generating new spectra through linear interpolation to obtain interpolated enhanced spectra.

[0058] Preferably, the specific implementation process of step 23 is as follows: The multi-strategy collaborative data enhancement module receives the spectral quality rating results output in step 21 and the corresponding set of purified spectra, and filters out the spectral samples with high quality grades to form a high-quality spectral subset S_high. This subset is an M_high×N-dimensional matrix (M_high is the number of high-quality spectral samples, and N is the number of wavelength points of each spectrum), and the matrix element S_high(i,n) represents the spectral value of the i-th high-quality spectral sample at the n-th wavelength point λ_n. At the same time, the comprehensive quality score Q(i) corresponding to each high-quality spectral sample is obtained (Q(i) is the score calculated by weighted calculation of numerical stability, derivative continuity, and signal-to-noise ratio in step 21, and the value range is 0-100).

[0059] Preferably, in the specific technical implementation of step 23, the spectral pairing and screening layer performs full pairwise pairing on the high-quality spectral subset S_high to form a spectral pair set ∈{1,2,...,M_high}, j∈{1,2,...,M_high}, i < j}, and this set contains spectral pairs (C is the combination number); the spectral pairing and screening layer calculates the Pearson correlation coefficient r_ij of the two spectra in each spectral pair, and the calculation method is:

[0060] where is the spectral mean of the i-th high-quality spectral sample, is the spectral mean of the j-th high-quality spectral sample; the value range of the Pearson correlation coefficient r_ij is [-1,1], which is used to characterize the similarity of the change trends of the two spectra in the wavelength dimension. The closer the value is to 1, the higher the similarity; the spectral pairing and screening layer sets the correlation coefficient threshold r0 (the value range is 0.8-0.95, for example, 0.9), and retains the spectral pairs with r_ij > r0 as the effective spectral pair set Pair_eff, and eliminates the spectral pairs with insufficient similarity to ensure that the newly generated spectra by subsequent interpolation have a consistent spectral feature trend.

[0061] Preferably, in a scenario, when step 23 is specifically implemented, the adaptive weight allocation layer receives the effective spectral pair set Pair_eff output by the spectral pairing and screening layer, and for each effective spectral pair The difference in overall quality scores for the spectral pair, ΔQ_ij = |Q(i) - Q(j)|, is calculated, and an adaptive weight allocation function is constructed based on ΔQ_ij. The designed weight allocation function is a non-linear mapping relationship. When ΔQ_ij is small, the weight difference between the two spectra is small; when ΔQ_ij is large, the spectrum with the higher overall quality score is given a larger weight. The specific calculation method is as follows:

[0062]

[0063] Where w_i is the interpolation weight of the i-th spectrum, w_j is the interpolation weight of the j-th spectrum, Q_max is the maximum comprehensive quality score in the high-quality spectrum subset S_high, and k is the weight adjustment coefficient (ranging from 0.1 to 0.3, for example, 0.2). The reason for this weight allocation method is that the spectrum with a higher comprehensive quality score contains more reliable amino acid content correlation features. Giving it a higher weight can make the generated new spectrum closer to the feature distribution of the high-quality spectrum. At the same time, the introduction of ΔQ_ij avoids the distortion of the new spectrum features due to excessive quality differences.

[0064] Preferably, in the specific implementation of step 23, the linear interpolation generation layer receives the interpolation weights of each effective spectral pair output by the adaptive weight allocation layer. For each effective spectral pair Perform linear interpolation operation point by point to generate a new spectrum S_gen(k,n) (where k is the index of the new spectrum and n is the index of the wavelength point). The calculation method is as follows:

[0065] Where S_gen(k,n) is the spectral value of the k-th new spectrum at the n-th wavelength; to control the balance between the amount of new data and computational efficiency, the linear interpolation generation layer sets the number of interpolations t (ranging from 1 to 3, for example, 2) for each effective spectral pair, and generates t different new spectra by adjusting the weight step size. The weight step size Δw ranges from 0.1 to 0.2 (for example, 0.15), that is, the weight of subsequent new spectra is... , This ensures that the generated new spectra are uniformly distributed in the feature space of the two original spectra; all generated new spectra form an interpolated enhanced spectrum subset S_gen, which is an M_gen×N dimensional matrix (M_gen is the total number of new spectra generated by interpolation, M_gen=t×C_eff, C_eff is the number of spectral pairs in the effective spectral pair set Pair_eff).

[0066] Preferably, in the specific technical implementation of step 23, the linear interpolation generation layer also assigns a corresponding virtual comprehensive quality score Q_gen(k) to each interpolated enhanced spectrum S_gen(k), calculated as Q_gen(k) = w_i × Q(i) + w_j × Q(j), which is used for physical rationality verification and quality traceability in subsequent step 24. At the same time, the linear interpolation generation layer associates and stores the interpolated enhanced spectrum subset S_gen with the corresponding virtual comprehensive quality score and source spectrum pair identification information to form a complete interpolated enhanced spectrum set. This set, together with the spectrum set generated by neighborhood perturbation, distribution completion, and density filling processing, constitutes a multi-type enhanced spectrum set.

[0067] This application's technology differs from traditional fixed-weight or random-pairing interpolation methods. It employs a Pearson correlation coefficient screening layer for spectral pairing to ensure high similarity among interpolated spectral pairs, avoiding the generation of anomalous spectra with conflicting features. The adaptive weight allocation layer's nonlinear weight function incorporates spectral quality differences into the weight calculation, solving the problem of traditional fixed-weight interpolation's inability to distinguish spectral quality priorities, thus making the new spectra more likely to retain the effective features of high-quality spectra. The linear interpolation generation layer's multi-quantity, step-size generation strategy enriches data diversity while ensuring spectral authenticity, and the allocation of virtual comprehensive quality scores provides a basis for subsequent data management and quality assessment. In this technical implementation, the collaborative operation of each structural layer revolves around the preservation and reasonable expansion of high-quality spectral features. The generated interpolated enhanced spectra maintain consistency with the original high-quality spectra while achieving data enhancement diversity through weight adjustment and multi-quantity generation, providing a high-quality data foundation for improving the generalization ability and prediction stability of amino acid content prediction models.

[0068] Optionally, step 23 specifically includes: for the high-difficulty spectrum, finding a set number of neighboring spectra in the feature space, calculating the mean of the neighboring spectra; applying a local perturbation of a preset amplitude in the direction of the neighboring mean to generate a new spectrum that is similar to the high-difficulty spectrum but has subtle differences, thus obtaining the neighborhood perturbation enhanced spectrum.

[0069] Preferably, the specific implementation process of step 23 is as follows: The multi-strategy collaborative data augmentation module receives the high-difficulty spectral identification result and the corresponding purified spectral set output in step 22, and selects the samples identified as high-difficulty spectra to form a high-difficulty spectral subset S_hard. This subset is an M_hard×N dimensional matrix (M_hard is the number of high-difficulty spectral samples, and N is the number of wavelength points for each spectrum). The matrix element S_hard(p,n) represents the spectral value of the p-th high-difficulty spectral sample at the n-th wavelength point λ_n. At the same time, the comprehensive quality score Q(p) and difficulty score D(p) corresponding to each high-difficulty spectral sample are obtained (Q(p) is the weighted quality score in step 21, and D(p) is the normalized difficulty score in step 22).

[0070] Preferably, in the specific technical implementation of step 23, the feature space mapping layer performs principal component analysis to reduce the dimensionality of the high-difficulty spectral subset S_hard and the fully purified spectral set S_refine, constructing a suitable feature space. First, S_refine is standardized to S_norm, the covariance matrix of S_norm is calculated and eigenvalue decomposition is performed, and the top K principal components with a cumulative variance contribution rate of not less than 90% are selected (K ranges from 3 to 10, for example, 5), resulting in a feature vector matrix W (dimension N×K). The high-difficulty spectral subset S_hard is mapped to this feature space to obtain the high-difficulty feature matrix F_hard = S_hard × W (dimension M_hard × K). At the same time, the fully purified spectral set S_refine is mapped to the global feature matrix F_global = S_refine × W (dimension M×K). Feature space mapping can reduce the spectral dimension and highlight the core features, making the nearest neighbor search more targeted.

[0071] Preferably, in a scenario, when step 23 is specifically implemented, the dynamic nearest neighbor selection layer searches for nearest neighbors in the global feature matrix F_global for each high-difficulty feature vector F_hard(p,:) (the feature vector of the p-th high-difficulty spectrum); firstly, it calculates the Euclidean distance d(p,q) between F_hard(p,:) and each feature vector F_global(q,:) in F_global:

[0072] Where F_hard(p,k) is the k-th principal component value of the p-th high-difficulty feature vector, and F_global(q,k) is the k-th principal component value of the q-th global feature vector; based on Euclidean distance sorting, the top L candidate nearest neighbors with the smallest distance are initially selected (L ranges from 5 to 15, for example, 8); combined with spectral quality and difficulty characteristics, further screening is carried out, and samples among the candidate nearest neighbors whose comprehensive quality score Q(q) is lower than a set threshold (for example, Q(q) < 0.5) or whose difficulty score D(q) differs from D(p) by more than a set range (for example, |D(q) - D(p)| > 0.3) are removed. The final number of effective nearest neighbors is L_eff (L_eff ranges from 3 to 8, for example, 5), denoted as the nearest neighbor index set Idx_neighbor(p); the dynamic nearest neighbor screening layer calculates the original spectral mean S_neighbor_mean(p,:) corresponding to the effective nearest neighbors, and the calculation method is as follows:

[0073] Where S_refine(q,n) is the spectral value of the qth effective nearest neighbor at the nth wavelength point. This mean spectrum reflects the core neighborhood features of the high-difficulty spectrum in the feature space.

[0074] Preferably, in the specific implementation of step 23, the gradient perturbation generation layer receives the high-difficulty spectrum S_hard(p,:) and the nearest-neighbor mean spectrum S_neighbor_mean(p,:) output by the dynamic nearest-neighbor screening layer. First, it calculates the difference vector ΔS(p,:) = S_neighbor_mean(p,:) - S_hard(p,:), which represents the characteristic offset direction between the high-difficulty spectrum and the neighborhood mean. Based on the quality score Q(p) and difficulty score D(p) of the high-difficulty spectrum, a dynamic perturbation coefficient γ(p) is designed. The calculation method is γ(p) = γ0 × (a × D(p) + b × (1 - Q(p))), where γ0 is the basic perturbation coefficient (ranging from 0.02 to 0.1, for example, 0.05), and a and b are adjustment weights (a ranging from 0.6 to 0.8, and b ranging from 0.2 to 0.4, for example, a = 0.7, b = 0.3). The higher the difficulty score and the lower the quality score, the larger the perturbation coefficient γ(p), so that the difficult and poor quality spectrum receives a more appropriate perturbation. Then, a gradient-based local perturbation is applied to generate the enhanced spectrum S_perturb(p,:).

[0075] Where ω(n) is the wavelength adaptive weight, determined by the signal strength at that wavelength point, ω(n) = S_si S_signal(n) is the average signal intensity of the spectrum at the nth wavelength after all purification processes (S_si The wavelength adaptive weighting ensures that the perturbation is more gentle in the characteristic bands with strong signals and more moderate in the bands with weak signals, avoiding damage to the core features; the enhanced spectra corresponding to all high-difficulty spectra form a neighborhood perturbation enhanced spectrum subset S_perturb, which together with the spectrum sets generated by other enhancement strategies constitutes a multi-type enhanced spectrum set.

[0076] Preferably, in the specific technical implementation of step 23, the gradient perturbation generation layer also performs constraint checks on the generated enhanced spectrum S_perturb(p,:) to ensure that the spectral values ​​of each wavelength point are within [0.8, 1.2] times the corresponding wavelength point values ​​of the original high-difficulty spectrum subset S_hard. If they exceed this range, they are truncated to the boundary value. At the same time, the Pearson correlation coefficient between the enhanced spectrum and the original high-difficulty spectrum is calculated, and only enhanced spectra with a correlation coefficient greater than 0.95 are retained to ensure that the perturbation does not cause distortion of the core features. The final output neighborhood perturbation enhanced spectrum subset S_perturb is an M_perturb×N-dimensional matrix (M_perturb is the number of enhanced spectra that pass the constraint check).

[0077] This application differs from traditional methods that fix the number of nearest neighbors and the perturbation amplitude. It uses principal component analysis in the feature space mapping layer to reduce dimensionality, focusing the nearest neighbor search on core features rather than the original noise. A dynamic nearest neighbor selection layer combines quality and difficulty characteristics to select effective nearest neighbors, avoiding interference from low-quality or excessively different samples in the mean calculation. The gradient perturbation generation layer uses dynamic perturbation coefficients and wavelength-adaptive weights to adapt the perturbation intensity to the spectral characteristics, ensuring both the diversity of the enhanced spectrum and preventing the destruction of core features. In this implementation, each structural layer is designed around the "difficulty" and "quality" characteristics of highly complex spectra. The generated enhanced spectra can accurately fill sample gaps in the model's boundary regions, ensuring the effectiveness and reliability of the enhanced spectra and providing strong support for improving the model's generalization ability.

[0078] Optionally, step 23 specifically includes: dividing the amino acid content range into multiple intervals according to a set interval, and statistically analyzing the number density of the purified spectra in each interval; for intervals with a density lower than the set proportion of the global mean, selecting the spectra in that interval to perform interpolation to generate new spectra, filling the spectral distribution gaps in the intervals, and obtaining distribution-completed enhanced spectra.

[0079] Preferably, the specific implementation process of step 23 is as follows: The multi-strategy collaborative data enhancement module receives the purified spectral set S_refine and the corresponding amino acid content true value vector Y_glu output from step 13. The purified spectral set S_refine is an M×N dimensional matrix (M is the total number of samples, and N is the number of wavelength points for each spectrum). The matrix element S_refine(m,n) represents the spectral value of the m-th sample at the n-th wavelength point λ_n. The amino acid content true value vector Y_glu is an M×1 dimensional vector. Y_glu(m) represents the true value of the amino acid content of the m-th sample (the unit is consistent with chemical detection, such as %).

[0080] Preferably, in the specific technical implementation of step 23, the content range division layer performs dynamic range division on the true value vector Y_glu of amino acid content, first calculating the minimum value Y_glu. Maximum value Y_ The range ΔY = Y_max - Y_min; the number of intervals K is determined based on the range and sample distribution characteristics, and K is calculated as follows: (rounding is used for rounding operations), the value range is 5-15 (e.g., when M=100, K=10), ensuring that each interval contains enough samples and reflects the distribution differences; the value range is divided using an equal interval method. Divide into K consecutive intervals, the range of the kth interval is Where Y_low(k) = Y_min + (k-1) × ΔY / K, Y_high(k) = Y_min + k × ΔY / K (k = 1, 2, ..., K); at the same time, in order to avoid missing boundary samples, the lower limit of the first interval is extended to Y_min - ε (ε is the minimum value, which takes the range of 0.5%-1% of ΔY, for example 0.8%), and the upper limit of the Kth interval is extended to Y_max + ε, to ensure that all samples can fall into the corresponding interval.

[0081] Preferably, in one scenario, when step 23 is specifically implemented, the sample density evaluation layer evaluates each interval To calculate the sample density ρ(k), first count the number of samples n(k) within each interval, i.e., n(k) = count{m|Y_glu(m)∈ ,m=1,2,...,M}; calculate the global average sample density ρ_avg=M / K, representing the number of samples that each interval should contain under an ideal uniform distribution; define the interval sample density ρ(k)=n(k) / ρ_avg, where a density value greater than 1 indicates that the interval sample is sufficient, and a density value less than 1 indicates that the interval sample is sparse; set a density threshold ρ_th (within the range of 0.3-0.6, for example 0.5), mark the intervals where ρ(k)<ρ_th as sparse intervals, and record the sparse interval index set K_sparse={k|ρ(k)<ρ_th,k=1,2,...,K}; for each sparse interval k∈K_sparse, extract the samples in the interval to form a sparse interval sample subset S_sparse(k) (with a dimension of n(k)×N), and the corresponding amino acid content forms Y_sparse(k) (with a dimension of n(k)×1).

[0082] Preferably, in the specific implementation of step 23, the interval interpolation generation layer performs interpolation to generate a new spectrum for each sparse interval k∈K_sparse based on its sample subset S_sparse(k); firstly, the samples in the sparse interval are sorted in ascending order of amino acid content to obtain the sorted spectral matrix S_sparse. (Rows are arranged in ascending order of content) and the sorted content vector Y_so The number of interpolation generators t(k) is determined based on the number of samples n(k) in the sparse interval, where t(k) = r. η is the generation coefficient (ranging from 0.8 to 1.2, for example, 1.0), ensuring that the number of generated samples is sufficient to fill gaps without causing redundancy; a combination of linear interpolation and spectral feature constraints is used to generate new spectra. For two adjacent samples i and i+1 after sorting (i=1,2,...,n(k)-1), the interpolation weight α∈[0.1,0.9] is calculated (t(k) weight values ​​are selected according to a uniform distribution), generating a new spectrum S_new(k,i,α) and the corresponding new content Y_new(k,i,α):

[0083]

[0084] Where S_sorted(k,i,:) is the spectral vector of the i-th sample after sorting within the k-th sparse interval, and Y_sorted(k,i) is the corresponding amino acid content; to ensure the physical rationality of the new spectrum, constraints are applied to the spectral values ​​at each wavelength point during the interpolation process, making them fall within the range of S_sorted(k,i,n) and S_sorted Within the interval (n=1,2,...,N).

[0085] Preferably, in the specific technical implementation of step 23, the interval interpolation generation layer summarizes the new spectra generated by all sparse intervals to form a distribution completion enhanced spectrum subset S_dist (with a dimension of M_dist×N, where M_dist is the total number of new samples generated by all sparse intervals, M_dist=∑t(k), k∈K_sparse); at the same time, a corresponding virtual amino acid content Y_dist (with a dimension of M_dist×1) is assigned to each new spectrum, which is Y_new(k,i,α) generated by the above interpolation; finally, the distribution completion enhanced spectrum subset S_dist is associated and stored with the corresponding virtual content Y_dist, and merged with the spectrum sets generated by other enhancement strategies to form a multi-type enhanced spectrum set.

[0086] This application differs from traditional fixed interval division and uniform interpolation methods. It dynamically calculates the number of intervals in the content value range division layer, ensuring the interval division matches the total sample size and avoiding inaccurate distribution characterization caused by too many or too few intervals. The sample density evaluation layer, based on global average density and dynamic thresholds, accurately identifies truly sparse intervals, ensuring enhanced targeting. The interval interpolation generation layer combines content ranking and spectral feature constraints in its interpolation method, guaranteeing both the continuity of the new spectrum in the content dimension and maintaining the correlation between spectral features and content, avoiding the generation of unrealistic or abnormal spectra. In this implementation, each structural layer revolves around the distribution characteristics of amino acid content. The generated distribution completion enhanced spectrum accurately fills the gaps in sparse intervals, making the training set sample coverage more comprehensive. This effectively solves the problem of poor model prediction performance for amino acids with extreme content values ​​or in sparse intervals. Its technical logic is connected with the physical rationality verification in step 24, ensuring that the generated spectrum both meets distribution requirements and possesses spectral authenticity, providing high-quality data support with a uniform distribution for improving the model's generalization ability.

[0087] Optionally, step 23 specifically includes: calculating the local density of each purified spectrum in the feature space; performing conservative interpolation on the spectrum with a local density lower than a set quantile and its nearest neighbor spectrum to generate a new spectrum to increase the density of the feature space, thereby obtaining a density-filled enhanced spectrum.

[0088] Preferably, the specific implementation process of step 23 is as follows: the multi-strategy collaborative data enhancement module receives the purified spectral set output in step 13. Compared with the spectral quality rating results output in step 21, the purified spectral set An M×N dimensional matrix ( This represents the total number of samples. (Number of wavelength points for each spectrum), matrix elements Representing the The sample at the th wavelength points The spectral values; simultaneously receiving the true amino acid content values ​​corresponding to each sample. With overall quality score .

[0089] Preferably, in the specific technical implementation of step 23, the feature space construction layer constructs the purified spectral set. Compared with the true value of amino acid content Joint feature construction is performed; firstly, for Standardized processing is performed to obtain ,calculate Principal component characteristics, selecting the top components with a cumulative variance contribution rate of not less than 85%. Principal components ( The value range is 3-8, for example 4), thus obtaining the spectral feature matrix. ( (The feature vector matrix has dimensions N×K); the true values ​​of amino acid content. Standardized to The standardized formula is:

[0090] Constructing the joint feature matrix (dimension is) This matrix contains both core spectral features and content information, making the feature space more aligned with the prediction task requirements; the feature space construction layer also outputs a joint feature matrix. Mapping relationship with the corresponding original spectral index.

[0091] Preferably, in a scenario, when step 23 is specifically implemented, the local density calculation layer targets the joint feature matrix. Each feature vector in Calculate its local density First, set the number of nearest neighbors. (Values ​​range from 3 to 10, for example, 5), calculate and All other feature vectors Mahalanobis distance The calculation formula is:

[0092] in Joint characteristic matrix The covariance matrix (with dimensions (K+1)×(K+1)). The Mahalanobis distance is the inverse of the covariance matrix. It can eliminate the influence of correlation between features, making distance calculation more reasonable. Based on the Mahalanobis distance ranking, the features with the smallest distance are selected. Each sample is used as the nearest neighbor sample set. ; Calculate local density The weighted average of the overall quality scores of samples in the nearest neighbor sample set is calculated using the following formula:

[0093] in The design uses the comprehensive quality score of the nearest neighbor sample m' to make the local density reflect both the sample distribution density in the feature space and the sample quality, thus avoiding misclassification of low-quality sample clusters as high-density regions.

[0094] Preferably, in the specific implementation of step 23, the local density calculation layer calculates the local density of all samples. Perform statistical analysis to determine the density quantile threshold. Set quantile ratio (The value range is 20%-40%, for example, 30%) for The t-quantile, i.e. , local density The samples are labeled as sparse region samples, forming a sparse sample subset. (dimension is) , (Number of samples in the sparse region), and the corresponding joint feature vectors. (dimension is) ).

[0095] Preferably, in the specific technical implementation of step 23, the conservative interpolation generation layer is designed for sparse sample subsets. Each sample in (The spectrum of the p-th sparse region sample), in the joint feature matrix Find its nearest neighbor sample (Minimum distance and overall quality score) (sample); construct conservative interpolation constraints, interpolation weights The value range is [0.4, 0.6] (e.g., 0.5), ensuring that the new spectrum is closer to the middle region of the feature space and avoiding excessive deviation from the features of the original sample; generating a new spectrum. With corresponding new content The calculation formulas are as follows:

[0096]

[0097] in for The index of the nearest neighbor sample. The original content of sparse samples. The content of the nearest neighbor sample; to further ensure the rationality of the interpolation, the new spectrum... Apply spectral morphology constraints and calculate its correlation with... , The Pearson correlation coefficient was used to retain only the new spectra with correlation coefficients greater than 0.92, ensuring the consistency of the characteristics of the generated spectra.

[0098] Preferably, in the specific implementation of step 23, the conservative interpolation generation layer summarizes all the new spectra that pass the constraints to form a density-filled enhanced spectral subset. (dimension is) , (The total number of new samples generated for density filling); and simultaneously assign a virtual comprehensive quality score to each new spectrum, calculated using the following formula:

[0099] This virtual composite quality score is used for physical plausibility verification in subsequent step 24; density-filled enhancement spectral subsets are then used. It is stored in association with the corresponding virtual quality score and source sample identifier, and combined with the spectral sets generated by other enhancement strategies to form a multi-type enhanced spectral set.

[0100] This application differs from traditional density enhancement methods based solely on spectral features. By incorporating amino acid content information through joint feature design in the feature space construction layer, the feature space better aligns with the essential needs of the prediction task. The local density calculation layer employs a density evaluation method weighted by Mahalanobis distance and quality score, avoiding the shortcomings of traditional Euclidean distance which fails to consider feature correlation and sample quality, thus enabling more accurate identification of sparse regions. The conservative interpolation generation layer's weight constraints and morphological verification ensure that the new spectrum fills density gaps without deviating from the original sample's feature distribution range. In this implementation, each structural layer revolves around density optimization in the feature space. The generated density-filled enhanced spectrum effectively improves the uniformity of the feature space, helping the model learn the characteristic patterns of sparse regions. Its technical logic, along with the quality evaluation in step 21 and the rationality verification in step 24, forms a closed loop, ensuring the targeted and reliable nature of the enhancement process and providing comprehensive feature space coverage support for improving the generalization ability of the amino acid content prediction model.

[0101] Optionally, step 24 specifically includes: verifying whether the value of the enhanced spectrum is within a set proportion of the original hyperspectral data value range, whether the Pearson correlation coefficient with the reference spectrum is greater than a set threshold, whether the change in the number of spectral zero crossings is within a set proportion of the expected value of the number of spectral zero crossings in the original spectrum, and whether the relative change amplitude of the spectral values ​​at each wavelength does not exceed a set proportion; and retaining only the enhanced spectra that have passed all verifications.

[0102] Preferably, the specific implementation process of step 24 is as follows: The multi-strategy collaborative data enhancement module receives the multi-type enhanced spectral sets output in step 23. (dimension is) , To increase the total number of spectra, (Number of wavelength points), while simultaneously receiving reference spectral information for generating each enhanced spectrum. (dimension is) , This is the master reference spectral index for the m-th enhanced spectrum. To assist in the reference spectral indexing, the reference spectra are derived from the purified spectral set. ); Simultaneously load global statistical information of the original hyperspectral data, including the minimum value matrix for each wavelength point. (Dimension is 1×N, ), maximum value matrix (Dimension is 1×N, ), and the expected value of the zero-crossing number of the original spectrum. ( (A function for calculating zero crossing times).

[0103] Preferably, in the specific technical implementation of step 24, the numerical range constraint layer enhances the spectral set. Each enhanced spectrum in Perform wavelength adaptive range verification; set global range coefficient. (Values ​​range from 0.8 to 0.9, for example, 0.85) and (Values ​​range from 1.1 to 1.2, for example, 1.15), for each wavelength point Calculate the dynamic constraint range at this wavelength point. Simultaneously, the constraint range is adjusted based on the numerical distribution of the reference spectrum, if the reference spectrum... In to If the range is low (low value range), then the lower limit will be adjusted to... (Stricter constraints); if in to If the range is high (high value range), then the upper limit will be adjusted to... ;check If all wavelengths are within the adjusted dynamic constraint range, the layer is verified; otherwise, it is marked as an invalid spectrum.

[0104] Preferably, in one scenario, when step 24 is specifically implemented, the spectral correlation verification layer calculates the Pearson correlation coefficient between the enhanced spectrum constrained by the numerical range and the reference spectrum; for the interpolated enhanced spectrum, the average correlation coefficient with the two reference spectra is calculated using the following formula:

[0105] in for and The Pearson correlation coefficient, To and The Pearson correlation coefficient; for the enhanced spectrum generated by neighborhood perturbation, the correlation coefficient with the master reference spectrum (the hard sample itself) is calculated. and the correlation coefficient with the mean spectrum of the neighborhood ,Pick Set a correlation threshold (Values ​​range from 0.92 to 0.96, e.g., 0.94), the interpolated spectrum must meet the following requirements. The neighborhood perturbation spectrum must satisfy If the verification passes, the next level of verification is performed; otherwise, the spectrum is marked as invalid.

[0106] Preferably, in the specific implementation of step 24, the number of zero-crossings of the focused spectrum of the physical property consistency layer ( Verification, calculation of enhanced spectrum Zero crossing times This number reflects the stability of the spectral oscillation structure; a threshold for the change ratio is set. (Values ​​range from 0.2 to 0.3, for example, 0.25), calculate the allowable range of variation:

[0107] Simultaneously, combining the zero-crossing number of the reference spectrum Fine-tune the range of variation: if (Low oscillation spectrum), then the lower limit is adjusted to ;like (High oscillation spectrum), then the upper limit is adjusted to ;check If the spectrum is within the fine-tuned range, proceed to the final verification layer; otherwise, mark it as an invalid spectrum.

[0108] Preferably, in the specific technical implementation of step 24, the local variation limiting layer performs relative change verification on each wavelength point of the enhanced spectrum; and sets a local variation threshold. (Values ​​range from 0.08 to 0.12, e.g., 0.10). For each wavelength point n, calculate the relative change between the enhanced spectrum and the main reference spectrum using the following formula:

[0109] in To avoid the denominator being zero; for wavelengths within the characteristic wavelength range associated with amino acids (e.g., 450-750 nm), the threshold is adjusted to... (Stricter constraints) because the spectral changes in this range are highly correlated with amino acid content, and excessive perturbation must be avoided; calibrate all wavelength points. If the wavelength does not exceed the corresponding threshold, and the number of wavelengths that consecutively exceed the threshold does not exceed 5% of the total number of wavelengths (to avoid localized anomalies), then the spectrum is considered compliant. Finally, all enhanced spectra that pass the four-layer verification are collected to form the enhanced spectrum. The output is then used in subsequent feature extraction steps.

[0110] This application's technology differs from traditional single-dimensional verification methods. It employs a four-layer progressive architecture to achieve multi-dimensional, adaptive verification logic: a numerical range constraint layer dynamically adjusts the range based on wavelength characteristics and reference spectral distribution, avoiding the loss of valid samples due to one-size-fits-all constraints; a spectral correlation verification layer designs differentiated correlation coefficient calculation methods for different enhancement strategies, adapting to different generation logics such as interpolation and perturbation; a physical characteristic consistency layer combines the global expected value with the individual characteristics of the reference spectrum to ensure the rationality of the oscillation structure; and a local change restriction layer strengthens constraints on feature intervals, ensuring that key spectral information is not destroyed. In this technology, abnormally enhanced spectra that do not conform to physical laws and spectral characteristics are filtered comprehensively, ensuring the reliability and effectiveness of the enhanced spectra. This provides a high-quality data foundation for subsequent principal component analysis and amino acid content prediction, making data enhancement both rich in sample diversity and free from introducing interference information.

[0111] Optionally, step 2, "Principal Component Analysis Feature Extraction Processing," specifically includes the following steps: Step 25: Perform centering processing on the enhanced spectrum, subtracting the mean of the corresponding wavelength point from the spectral value of each wavelength point to obtain centered spectral data; Step 26: Calculate the covariance between each wavelength point based on the centered spectral data, and construct a spectral covariance matrix; Step 27: Perform eigenvalue decomposition on the spectral covariance matrix to obtain the eigenvalues ​​and eigenvectors corresponding to each principal component, forming a spectral eigenvalue-vector set containing the eigenvalue sequence and the corresponding eigenvector matrix; Step 28: Calculate the variance contribution rate and cumulative variance contribution rate of each principal component based on the eigenvalues ​​in the spectral eigenvalue-vector set, select principal components with a cumulative variance contribution rate not lower than a set threshold, and construct the spectral core principal component feature matrix by combining the eigenvectors of the corresponding principal components in the spectral eigenvalue-vector set.

[0112] Preferably, the specific implementation process of step 25 is as follows: the principal component analysis feature extraction module receives the enhanced spectrum output in step 24. The enhanced spectrum for 3D matrix ( The total number of enhanced spectra that have passed physical rationality verification. (Number of wavelength points for each spectrum), matrix elements Representing the The enhanced spectral sample at the ... wavelength points The corresponding spectral values; simultaneously receiving the spectral quality rating result output from step 21, the enhanced spectral source identifier (original purified spectrum / generated enhanced spectrum) output from step 23, and the characteristic band index set determined based on the amino acid spectral characteristics of yeast extract. ( and The wavelength range corresponding to the characteristic absorption peaks of amino acids, for example , This index set is used to mark wavelength points that are strongly correlated with amino acid content.

[0113] Preferably, in the specific technical implementation of step 25, the spectral grouping statistical layer enhances the spectrum. Classified by source type into the original purified spectral group (sample size) ) and the generation of enhanced spectra (sample size) , ); Calculate the local statistical characteristics of the two sets of spectra at each wavelength: local mean of the original purified spectra group. Local standard deviation Generate local mean of enhanced spectral group Local standard deviation Simultaneously calculate the overall mean of the global spectrum at each wavelength point. These statistical characteristics provide a dual basis for mean calculation and centering correction.

[0114] Preferably, in a scenario, when step 25 is specifically implemented, the wavelength adaptive mean calculation layer calculates a dynamically centered mean for each wavelength point n. The design means are calculated as follows:

[0115] in For wavelength adaptive weights, the value selection rule is as follows: If (Characteristic band), then ( The characteristic band weights, ranging from 0.6 to 0.8 (e.g., 0.7), are assigned higher weights to the original purified spectral group to preserve the true characteristic distribution; if (Non-characteristic band), then ( This is the weight for non-characteristic bands, ranging from 0.4 to 0.6 (e.g., 0.5), balancing the mean contribution of the two sets of spectra; if the local standard deviation of the enhanced spectral group is generated... Local standard deviation from the original purified spectral set If the ratio exceeds a set threshold (e.g., 1.2), then... Adjusted to To avoid excessive fluctuations in the generated spectrum affecting the reliability of the mean; dynamically center the mean. It is a 1×N dimensional vector, where each element corresponds to the adaptive fusion mean of a wavelength point.

[0116] Preferably, in the specific implementation of step 25, the dynamic centering correction layer receives the dynamic centering mean output by the wavelength adaptive mean calculation layer. For the enhanced spectrum Perform a differential centering operation on each sample to obtain centralized spectral data. The calculation method is as follows:

[0117] in These are dynamic correction coefficients used to adapt to different sample types and wavelength characteristics: If If it belongs to the original purified spectrum (source identified as "original"), then To maintain the centered purity of the original sample; if it belongs to the generated enhanced spectrum (source identified as "generated"), then hour (e.g., 0.95) hour (For example, 1.05), which both suppresses excessive shift in the characteristic bands of the generated samples and moderately preserves the diversity of their non-characteristic bands; at the same time, if a certain wavelength point global mean Mean of the original purified spectral set The absolute value of the difference exceeds (If the global offset is large), then all samples at that wavelength point will be included. Unified adjustment to Strengthen centralization to offset the offset.

[0118] This application differs from traditional centralized methods with fixed global means. Through source grouping and dual-dimensional statistics in the spectral grouping statistical layer, it fully considers the distribution differences between the original and generated spectra. The weight allocation mechanism of the wavelength-adaptive mean calculation layer ensures that the mean calculation closely matches the actual sample distribution while also adapting to the supplementary characteristics of the generated samples. The design of the differential coefficient in the dynamic centralization correction layer further optimizes the centralization effect for different types of samples and different wavelength bands. In this technical implementation, each structural layer is precisely designed around the spectral characteristics of amino acids. The centralized spectral data not only meets the requirement of zero mean for principal component analysis but also avoids the feature overload problem that may occur with traditional centralization through targeted adjustments. This provides high signal-to-noise ratio and high feature recognition data support for the subsequent covariance matrix construction in step 26 and eigenvalue decomposition in step 27. Its technical logic forms a coherent closed loop with the subsequent principal component extraction process, ensuring that principal component analysis can accurately capture the core features related to amino acid content.

[0119] Preferably, the specific implementation process of step 26 is as follows: The principal component analysis feature extraction module receives the centered spectral data output in step 25. This centralized spectral data for 3D matrix ( To enhance the total number of spectra, (Number of wavelength points for each spectrum), matrix elements Representing the The enhanced spectral sample at the ... wavelength points The centered spectral value; simultaneously receiving the characteristic band index set determined based on the amino acid spectral characteristics of yeast extract. (For example , ) and non-characteristic band index set And the signal-to-noise ratio at each wavelength point output in step 21. ( For the first The signal-to-noise ratio at each wavelength point, with a value range of [value missing]. ).

[0120] Preferably, in the specific technical implementation of step 26, the band importance weighting layer applies to the centered spectral data. Perform wavelength point weighting processing, and design wavelength weights. The calculation method is as follows:

[0121] in The characteristic band weighting coefficient (with a value range of 1.2-1.5, for example 1.3). These are the non-characteristic band weighting coefficients (with values ​​ranging from 0.7 to 1.0, for example, 0.8). This represents the maximum signal-to-noise ratio (SNR) across all wavelengths. This formula assigns higher weights to characteristic bands and dynamically adjusts the contribution of each wavelength point based on the SNR, resulting in a more reasonable weight allocation for high SNR bands in covariance calculations. This allows for the centralization of spectral data. With weight vector By performing column-by-column weighting, we obtain weighted centered spectral data. ( (An N×N dimensional diagonal matrix constructed for the weight vector).

[0122] Preferably, in one scenario, when step 26 is specifically implemented, the local covariance enhancement layer applies weighted centered spectral data. Perform local window covariance calculation to capture local correlation features of adjacent wavelength points; set the local window size. (The value range is 3-7 wavelength points, for example 5), for each wavelength point Select the set of wavelength points within the window. (The boundary wavelength points are filled in using a mirror-fill method to complete the window); calculate the local covariance matrix within the window. :

[0123] in For the first A sample in the window The weighted centered spectral vector within (dimension 1) Local covariance matrix for A 3D matrix, reflecting the local correlation strength of wavelength points within the window; for all local covariance matrices... Normalization is performed to obtain the standardized local covariance matrix. Its elements satisfy the condition that the sum of rows is 1, ensuring the comparability of local correlation strength.

[0124] Preferably, in the specific implementation of step 26, the noise suppression covariance construction layer combines global and local covariance information to construct the final spectral covariance matrix. First, calculate the weighted centered spectral data. global covariance matrix :

[0125] in For the first The and the first The global covariance at each wavelength point reflects the overall correlation between the two across all samples; subsequently, based on the local covariance matrix... Constructing a local correlation weight matrix , For the first The and the first Local correlation weights for each wavelength point, if and Belonging to the same window ,but ( for In the window (position index in the middle), otherwise Finally, the spectral covariance matrix is ​​constructed by integrating the global covariance and the local correlation weights. :

[0126] in This is the fusion coefficient (with a value ranging from 0.6 to 0.8, for example, 0.7). The minimum regularization constant (within the range of values) ,For example ), For the elements of the identity matrix ( (The value is 1 if the condition is met, and 0 otherwise). This formula strengthens the covariance value of strongly correlated wavelength points through local correlation weights, and avoids matrix singularity through regularization terms, ultimately yielding an N×N dimensional spectral covariance matrix. Its elements Accurately reflects the first The and the first The weighted correlation strength at each wavelength point.

[0127] This application differs from traditional unweighted global covariance calculation methods. Through a differentiated weighting design in the band importance weighting layer, it highlights the contributions of amino acid characteristic bands and high signal-to-noise ratio bands, solving the problem of core feature dilution caused by the traditional method's indiscriminate treatment of all bands. The local covariance enhancement layer introduces local correlation information, enabling the covariance matrix to better capture local spectral structural correlations and adapt to the local absorption peak structures of amino acid characteristic bands. The fusion strategy and regularization processing of the noise suppression covariance construction layer effectively suppress the interference of isolated noise bands, improving the numerical stability of the covariance matrix. In this technical implementation, each structural layer is designed around the core requirement of amino acid prediction. The constructed spectral covariance matrix retains global correlation features while strengthening the importance of local structures and core bands, providing high-quality data support for the eigenvalue decomposition and high-value principal component extraction in step 27. Its technical logic forms a closed loop with subsequent principal component screening, ensuring that the extracted principal components can accurately correlate with changes in amino acid content.

[0128] Preferably, the specific implementation process of step 27 is as follows: The principal component analysis feature extraction module receives the spectral covariance matrix output in step 26. The spectral covariance matrix An N×N dimensional symmetric matrix ( (Number of wavelength points), matrix elements Representing the The and the first Weighted correlation strength at each wavelength point; simultaneously receiving the amino acid feature band index set. Signal-to-noise ratio weights at each wavelength (Calculated by the band importance weighting layer in step 26), and the reliability score of each wavelength point associated in the spectral quality rating result output in step 21. ( For the first The reliability quantization value of spectral data at each wavelength point, with a value range of [0,1]).

[0129] Preferably, in the specific technical implementation of step 27, the covariance matrix preprocessing layer processes the spectral covariance matrix. Numerical stability optimization and feature-guided enhancement are performed; firstly, the matrix is ​​calculated. condition number ( The largest eigenvalue of the matrix. (where the minimum non-zero eigenvalue is), if the condition number Exceeding the set threshold (e.g.) Then, regularization correction is performed on matrix C:

[0130] in Regularization coefficient (range of values ​​is ) ,For example ), An N×N diagonal matrix is ​​constructed for the reliability score vector. Through reliability score weighted regularization, matrix singularity is improved while retaining the correlation information of high-reliability wavelength points. Subsequently, based on the feature band index set... Constructing the feature guidance matrix ,in:

[0131] The regularized matrix is ​​then fused with the guiding matrix to obtain the optimized covariance matrix:

[0132] This fusion operation strengthens the correlation weights of the feature bands, making subsequent decomposition more inclined to extract amino acid-related features.

[0133] Preferably, in a scenario, when step 27 is specifically implemented, the adaptive eigenvalue decomposition layer optimizes the covariance matrix. The eigenvalue decomposition is performed with the objective of finding the eigenvector matrix that satisfies the following equation. diagonal matrix with eigenvalues :

[0134] in For an N×N dimensional matrix, the column vectors are... Representing the The eigenvectors corresponding to each principal component It is a diagonal matrix. For the first Each eigenvalue represents the variance contribution of the corresponding principal component. To adapt to the high-dimensionality of spectral data, an iterative eigenvalue decomposition algorithm is adopted, first extracting the first eigenvalues... The largest eigenvalues ​​and their corresponding eigenvectors ( The range of values ​​is ,For example hour Then, the remaining feature values ​​are rapidly decomposed, ensuring that core features are not lost while improving computational efficiency; a numerical verification mechanism is introduced during the decomposition process for each feature vector. Calculate its correlation score with the characteristic band:

[0135] The correlation score is used for subsequent feature vector selection.

[0136] Preferably, in the specific implementation of step 27, the feature vector filtering layer processes the feature values ​​obtained from the decomposition. With corresponding feature vectors Perform targeted screening; first, by feature value Sort in descending order to obtain the sorted sequence of feature values. ( ) and the corresponding eigenvector matrix Based on the weighted guidance of amino acid characteristic bands, the effective contribution of each feature vector is calculated:

[0137] The effective contribution rate comprehensively considers the variance contribution of eigenvalues, their correlation with characteristic bands, and the signal-to-noise ratio weight of characteristic bands; a contribution rate threshold is set. (The value range is 10%-20% of the maximum effective contribution, for example, 15%), retained. The eigenvectors are calculated; simultaneously, the orthogonality deviation between the eigenvectors is calculated, if the cosine of the angle between two eigenvectors satisfies:

[0138] Then, the feature vector with the lowest effective contribution is removed to ensure that the filtered feature vectors have good orthogonality. Finally, the spectral feature value-vector set is composed of the retained feature value sequence and the corresponding feature vector matrix. This set contains the feature value sequence. ( (Number of principal components retained after screening) and eigenvector matrix .

[0139] This application differs from traditional unguided eigenvalue decomposition methods. Through regularization correction and feature-guided fusion in the covariance matrix preprocessing layer, it addresses the potential singularity problem in high-dimensional spectral covariance matrices while strengthening the correlation information of amino acid-related bands. The iterative algorithm and correlation verification in the adaptive eigenvalue decomposition layer better adapt the decomposition process to the characteristics of spectral data, preventing core features from being overwhelmed. The multi-dimensional screening criteria in the feature vector selection layer ensure that the final extracted principal components have both high variance contribution and are highly correlated with changes in amino acid content. In this technical implementation, each structural layer is designed around the core requirement of amino acid prediction. The resulting spectral eigenvalue-vector set accurately captures key information in the spectral data, providing a high-quality foundation for subsequent principal component selection and the construction of the spectral core principal component feature matrix. Its technical logic forms a closed loop with the cumulative variance contribution rate screening in step 28, ensuring that the extracted principal components possess both statistical significance and practical predictive value.

[0140] Preferably, the specific implementation process of step 28 is as follows: The principal component analysis feature extraction module receives the spectral eigenvalue-vector set output in step 27, which contains an eigenvalue diagonal matrix. (Dimensions are M×M, (The number of principal components retained after screening in step 27) and the eigenvector matrix (Dimensions are N×M,) (Number of wavelength points for each spectrum) Representing the The eigenvalues ​​of the principal components Representing the The wavelength point at the ... The loading coefficients on each principal component; simultaneously receiving the signal-to-noise ratio weights at each wavelength point output from step 26. Amino acid characteristic band index set and the enhanced spectrum output from step 24. (dimension is) , (Total number of spectra after enhancement).

[0141] Preferably, in the specific technical implementation of step 28, the variance contribution quantization layer performs normalization processing and contribution calculation on the eigenvalue sequence, first calculating the sum of all retained principal component eigenvalues. The variance contribution rate P(k) of the kth principal component is calculated as follows: in Characterizing the first The proportion of each principal component that explains the total variance of the spectral data, with values ​​ranging from [0,1]; the cumulative variance contribution rate is obtained by sequentially summing the variance contribution rates. , Before characterization The cumulative proportion of total variance explained by each principal component, with values ​​ranging from [0,1]; simultaneously, the feature correlation contribution of the principal components is calculated in conjunction with the feature band weights. This indicator reflects the first The loading coefficients of each principal component are concentrated in the characteristic bands of amino acids, with values ​​ranging from [0,1]. The larger the value, the closer the main component is to the characteristics of the amino acid.

[0142] Preferably, in a scenario, when step 28 is specifically implemented, the dynamic threshold filtering layer combines the cumulative variance contribution rate and the feature association contribution rate to design a two-dimensional filtering rule. First, a basic cumulative variance contribution rate threshold is set. (Values ​​range from 0.85 to 0.95, for example, 0.90), initially filtering out those that meet the requirements. Minimum principal component subset ( If there is a feature correlation contribution within this subset; Below the set association threshold For principal components (within the range of 0.3-0.5, e.g., 0.4), then continue selecting features with higher correlation contributions from that principal component. Principal components are added to subsets until the cumulative variance contribution rate of the added subsets is not less than [amount missing]. (e.g., 0.945), thus obtaining the candidate principal component index set. Meanwhile, if the number of candidate principal components exceeds (For example If there are more than 5, then proceed as follows: Sort the products in descending order, keeping the first few products. To avoid model redundancy due to an excessive number of principal components, the final set of principal component indexes is determined. .

[0143] Preferably, in the specific implementation of step 28, the core feature matrix construction layer extracts the corresponding feature vectors based on the filtered principal component index set and constructs the principal component loading matrix. The matrix has dimensions of ( (This refers to the number of principal components ultimately retained after screening). Still representing the The wavelength point at the ... Loading coefficients on the principal components after screening; enhanced spectra With principal component loading matrix Perform matrix multiplication to obtain the characteristic matrix of the spectral core principal components. The matrix has dimensions of ,in This represents the score of the m-th enhanced spectral sample on the principal component after the k-th selection. This score comprehensively reflects the sample's performance in the spectral feature dimensions represented by the corresponding principal component. To ensure the numerical stability of the feature matrix, [the following is omitted as the text is incomplete and requires further context]. Perform row standardization, dividing the principal component score of each sample by the standard deviation of all principal component scores for that sample (standard deviation plus standard deviation). (To avoid a denominator of zero), we obtain the standardized spectral core principal component feature matrix. The output is then used in the subsequent amino acid content prediction step.

[0144] This application differs from traditional principal component screening methods that rely solely on cumulative variance contribution rates. It introduces feature correlation contribution through a variance contribution quantification layer, enabling the screening process to not only focus on information coverage but also strengthen the feature orientation related to amino acid content. The dual-dimensional rules and quantity constraints of the dynamic threshold screening layer avoid the loss or redundancy of core features caused by a single threshold, while ensuring the conciseness and relevance of the principal component set. The standardized processing of the core feature matrix construction layer further enhances the consistency and reliability of the feature data. In this technical implementation, each structural layer, based on the scenario requirements of amino acid prediction from yeast extract, uses multi-dimensional screening and precise calculation to ensure that the constructed spectral core principal component feature matrix retains the core information of the spectral data while highlighting the feature dimensions strongly related to the prediction task. Its technical logic is coherently adapted to the principal component extraction in step 27 and the input requirements of the subsequent prediction model, providing crucial support for improving the accuracy and generalization ability of amino acid content prediction.

[0145] Optionally, step 3, "predicting the amino acid content of the target yeast based on the spectral core principal component feature matrix," specifically includes: inputting the spectral core principal component feature matrix into the amino acid prediction model to predict the amino acid content value in the target yeast.

[0146] Optionally, the amino acid prediction model includes an input layer, a latent variable extraction layer, a regression coefficient calculation layer, and a prediction output layer. The spectral core principal component feature matrix is ​​input into the amino acid prediction model to predict the amino acid content in the target yeast. This includes: Step 31, performing adaptation and standardization processing on the spectral core principal component feature matrix through the input layer to obtain standardized features for yeast-derived amino acid prediction; Step 32, performing co-decomposition on the standardized features for yeast-derived amino acid prediction and the baseline data of amino acid content in yeast extract through the latent variable extraction layer to obtain amino acid-related latent variables; Step 33, performing regression calculation on the amino acid-related latent variables through the regression coefficient calculation layer to obtain the regression coefficient matrix for amino acid prediction; and Step 34, performing matrix operations on the regression coefficient matrix and the standardized features for yeast-derived amino acid prediction through the prediction output layer to obtain the amino acid content value of the target yeast.

[0147] Preferably, the specific implementation process of step 31 is as follows: The input layer receives the spectral core principal component feature matrix output in step 28. The matrix has dimensions of ( To enhance the total number of spectra, (Number of principal components ultimately retained after screening) Matrix elements Representing the The enhanced spectral sample at the ... The standardized scores of each principal component after screening are received; simultaneously, the effective contribution of each principal component output from step 27 is received. This index comprehensively reflects the correlation between the variance contribution of the principal components and the amino acid characteristics, and its value ranges from [0,1].

[0148] Preferably, in the specific technical implementation of step 31, the input layer first processes the feature matrix of the spectral core principal components. Perform principal component weight adaptation and design the principal component weight vector. The calculation method is as follows: in Characterizing the first The importance weights of each principal component in the prediction model range from [0,1]. Principal components with higher effective contributions receive higher weights, making the model more focused on high-value features. The spectral core principal component feature matrix is ​​used to... With principal component weight vector By performing column-by-column weighting, the weighted principal component feature matrix is ​​obtained. ( Constructed for weight vectors (diagonal matrix); for the weighted principal component characteristic matrix Perform row standardization: subtract the mean of all principal component scores for that sample from the principal component score for each sample, and then divide by the standard deviation of all principal component scores for that sample (standard deviation plus...). (To avoid a denominator of zero), we obtain the normalized features for predicting yeast-derived amino acids. The dimension of this matrix remains the same. ,element Representing the The standardized eigenvalues ​​of a sample on the k-th weighted principal component.

[0149] Preferably, the specific implementation process of step 32 is as follows: the latent variable extraction layer receives the yeast-derived amino acid prediction normalization features output by the input layer. Simultaneously, it receives benchmark data on the amino acid content of yeast extract. (dimension is) , The number of baseline data samples, Representing the The true amino acid content of each benchmark sample), and the comprehensive quality score in the spectral quality rating result corresponding to the benchmark sample output in step 21. (The value range is [0,1]).

[0150] Preferably, in the specific technical implementation of step 32, the latent variable extraction layer first constructs a benchmark feature-content joint matrix. ,in The normalized feature matrix for predicting yeast-derived amino acids corresponding to the baseline sample (dimension 1). Based on comprehensive quality score Constructing the benchmark sample weight matrix (dimension is) The higher the quality score of the benchmark sample, the greater its weight in the collaborative decomposition, thus improving the reliability of the decomposition results. A collaborative decomposition objective function is designed, and the amino acid association latent variable matrix satisfying the following relationship is solved. (dimension is) , The number of latent variables, with a range of values. For example, 3) with the load matrix (Dimensions are) ), content correlation vector (Dimension is 1×L): in For the square of the Frobenius norm, for Norm square, The content-related weighting coefficient (ranging from 1.5 to 2.5, e.g., 2.0) is used to balance the importance of feature fitting and content fitting; the objective function is solved using an alternating iterative optimization algorithm, first fixing... Solve and Then fix and Solve The number of iterations is set to 30-50 (e.g., 40) until the objective function value converges (the difference between two adjacent iterations is less than 1). The final amino acid correlation latent variable matrix was obtained. middle, Representing the The benchmark sample at the ... The score on a latent variable contains both principal component feature information and amino acid content correlation information.

[0151] Preferably, in one scenario, when step 32 is specifically implemented, the latent variable extraction layer performs validity verification on the amino acid association latent variable matrix T, and calculates the Pearson correlation coefficient between each latent variable and the amino acid content. Latent variables with correlation coefficients greater than a set threshold (ranging from 0.6 to 0.8, for example, 0.7) are retained to form the final amino acid association latent variable matrix. (dimension is) , (Number of effective latent variables), the corresponding loading matrix is ​​updated as follows: (dimension is) This ensures that all extracted latent variables are strongly correlated with amino acid content.

[0152] Preferably, the specific implementation process of step 33 is as follows: the regression coefficient calculation layer receives the amino acid association latent variable matrix output by the latent variable extraction layer. Load matrix and the normalized predictive features of yeast-derived amino acids from the benchmark sample. Compared with amino acid content benchmark data .

[0153] Preferably, in the specific technical implementation of step 33, the regression coefficient calculation layer first constructs a latent variable regression model based on the amino acid-related latent variable, and designs the regression coefficient matrix. (dimension is) The solution method is as follows: in Regularization parameter (value range is) ,For example ), used to suppress model overfitting, for 3D identity matrix The square of the baseline sample weight matrix further strengthens the influence of high-quality samples on the regression coefficients. This formula uses latent variables as an intermediary bridge to link principal component features with amino acid content, allowing the regression coefficients to simultaneously consider the correlation between feature loadings and latent variable content. The regression coefficient matrix... Perform feature band-guided correction based on the feature band index set output in step 26. With principal component loading matrix Calculate the characteristic band loading percentage of each principal component. Construct correction coefficients ( To correct the intensity coefficient (with a value ranging from 0.3 to 0.5, for example, 0.4), the first value of the regression coefficient matrix B is... Each element multiplied by The final amino acid prediction regression coefficient matrix is ​​obtained. This correction moderately amplifies the regression coefficients corresponding to principal components closely associated with amino acid characteristic bands, thereby improving the predictive accuracy.

[0154] Preferably, the specific implementation process of step 34 is as follows: the prediction output layer receives the amino acid prediction regression coefficient matrix output by the regression coefficient calculation layer. The normalized features of yeast-derived amino acid predictions corresponding to the target yeast sample output from the input layer. (dimension is) , The number of samples to be predicted. Represents the m-th sample to be predicted in the... Standardized eigenvalues ​​on each weighted principal component).

[0155] Preferably, in the specific technical implementation of step 34, the prediction output layer first normalizes the yeast-derived amino acid prediction features of the sample to be predicted. With amino acid prediction regression coefficient matrix Perform matrix multiplication to obtain the initial amino acid prediction values. (dimension is) Based on the spectral quality score of the sample to be predicted output in step 21. (Values ​​range [0,1]) Construct prediction correction coefficients (Values ​​range from [0.9, 1.1]), the correction coefficient for samples with higher quality scores is closer to 1.1, and the correction coefficient for samples with lower quality scores is closer to 0.9. This correction balances the prediction bias of samples of different quality. The initial amino acid prediction values ​​are... With prediction correction coefficient Element-by-element multiplication yields the corrected amino acid prediction values. Finally, the corrected amino acid prediction values ​​were analyzed. Applying range constraints will result in values ​​exceeding the amino acid content benchmark data. Range of values ( The minimum value of the baseline data. The predicted value (based on the maximum value of the baseline data) is truncated to the boundary of this range to obtain the final target yeast amino acid content value. This ensures that the predicted results are within a reasonable range of the actual content.

[0156] This application's technology differs from traditional direct regression prediction methods. Through principal component weight adaptation in the input layer, the model focuses on high-value principal component features, avoiding interference from ineffective features. The collaborative decomposition mechanism in the latent variable extraction layer integrates principal component features and amino acid content information, resulting in latent variables with both feature representation and content correlation attributes. Regularization design and feature band-guided correction in the regression coefficient calculation layer enhance the stability and relevance of the regression coefficients. Quality adaptation correction and range constraints in the prediction output layer further optimize the rationality of the prediction results. In this technical implementation, each structural layer, based on the scenario requirements of predicting amino acids from yeast extracts, uses multi-dimensional constraints and adaptation mechanisms to ensure that the model can fully utilize the rich features of the enhanced spectrum while accurately capturing the correlation patterns with amino acid content.

[0157] As shown in Figure 2(a), the horizontal axis represents wavelength (unit: nm) and the vertical axis represents reflectance. The black waveform in the figure represents the original spectrum of the standard yeast extract, while the red, blue, and green waveforms represent the new spectra generated by quality-driven interpolation, hard sample neighborhood perturbation, and feature space density optimization, respectively. The waveform trends of all generated spectra are highly consistent with the original spectra, and no abnormal fluctuations exceeding physical laws have appeared. It can be seen that this application ensures the authenticity of the enhanced spectrum and avoids the introduction of invalid abnormal samples through a four-layer physical rationality verification mechanism. As shown in Figures 2(b) and 2(c), the horizontal axis represents the true value of amino acid content (unit: %), and the vertical axis represents the predicted value. The blue circles in the figures represent the prediction results of a single sample. In Figure 2(b), the prediction results based on the original spectral model have a large degree of dispersion, with some samples deviating far from the ideal prediction line (diagonal). In contrast, the prediction results based on the extended spectral model in Figure 2(c) are more concentrated near the ideal prediction line, and the overall distribution is more uniform. Combined with the data in the table, it can be seen that the determination coefficient of the test set of the standard yeast extract increased from 0.8877 to 0.9127, and the root mean square error decreased from 0.8850 to 0.7804. It can be seen that this application effectively improves the prediction accuracy and stability of the model through the synergistic enhancement of multiple strategies such as quality-driven interpolation enhancement and hard sample neighborhood enhancement, especially improving the prediction effect of boundary samples.

[0158] like Figure 3As shown, the horizontal axis represents the types of 16 amino acids (from left to right: aspartic acid, threonine, serine, glutamic acid, glycine, alanine, valine, methionine, isoleucine, leucine, tyrosine, phenylalanine, lysine, histidine, arginine, and proline), and the vertical axis represents the root mean square error (RMSE). The blue bars in the figure represent the prediction error of each amino acid. The lower the bar height, the higher the prediction accuracy. It can be seen that the prediction errors of all 16 amino acids are at a low level, and there are no amino acids with significantly abnormally high errors. Therefore, the scheme combining principal component analysis feature extraction and PLSR modeling in this application can effectively capture the spectral correlation features of different amino acids, achieve high-precision synchronous prediction of multiple amino acids, and meet the detection needs of complex multi-component systems in yeast extracts.

[0159] As shown in Figure 4(a), the horizontal axis represents wavelength (unit: nm), and the vertical axis represents reflectance. The absorption and reflection peaks of the black original spectrum of the low-hydrolyzed yeast extract and the red, blue, and green generated spectra in the characteristic band (400-800 nm) are highly consistent, and no characteristic distortion is observed in the generated spectrum, demonstrating the rationality of the conservative interpolation and neighborhood perturbation strategy of this application. As shown in Figures 4(b) and 4(c), the coefficient of determination for the test set modeled based on the original spectrum of the low-hydrolyzed yeast extract is 0.9492, and the root mean square error is 0.5619. However, the coefficient of determination for the test set modeled based on the extended spectrum increases to 0.9583, and the root mean square error decreases to 0.5088. The clustering of the predicted result points is significantly improved, especially the predicted values ​​of samples in the low content range are closer to the true values. It can be seen that the distribution balance enhancement strategy of this application effectively fills the sample gap in the sparse content range, making the model more adaptable to the entire content range.

[0160] like Figure 5 As shown, the horizontal axis represents the types of 16 amino acids (from left to right: aspartic acid, threonine, serine, glutamic acid, glycine, alanine, valine, methionine, isoleucine, leucine, tyrosine, phenylalanine, lysine, histidine, arginine, and proline), and the vertical axis represents the root mean square error (RMSE). The height of the prediction error bars for each amino acid in the low-hydrolyzed yeast extract is lower than that of the original modeling scheme, and the error reduction is more significant for key amino acids such as glutamic acid and lysine. This demonstrates that this application, through feature space density optimization and spectral feature enhancement, highlights the signals of the amino acid feature-related bands, effectively suppresses noise interference, and improves the prediction accuracy of key amino acids.

[0161] As shown in Figure 6(a), the fluctuation trends of the original black spectrum and the generated red, blue, and green spectra of the highly hydrolyzed yeast extract in the characteristic bands are consistent with those of the original spectrum, and the local details are richer. This demonstrates the advantages of the adaptive weighted differential and baseline drift cancellation techniques of this application, which enable the enhanced spectrum to retain core features while supplementing detailed information. As shown in Figures 6(b) and 6(c), the coefficient of determination of the test set based on the augmented spectrum model of the highly hydrolyzed yeast extract increased from 0.9650 to 0.9735, and the root mean square error decreased from 0.5077 to 0.4416. The predicted result points are almost close to the ideal prediction line, with extremely small dispersion. It can be seen that the synergistic effect of the multi-strategy collaborative data augmentation and principal component analysis feature extraction of this application can fully explore the effective information in the spectrum, significantly improve the generalization ability and prediction accuracy of the model, and achieve high-precision prediction even for samples with more complex components such as high degree of hydrolysis.

[0162] like Figure 7 As shown, the horizontal axis represents the types of 16 amino acids (from left to right: aspartic acid, threonine, serine, glutamic acid, glycine, alanine, valine, methionine, isoleucine, leucine, tyrosine, phenylalanine, lysine, histidine, arginine, and proline), and the vertical axis represents the root mean square error (RMSE). The prediction errors of the 16 amino acids in the highly hydrolyzed yeast extract all reached a low level, and the error differences between each amino acid were small. This shows that the technical solution of this application has good stability and versatility, and can be adapted to different types of yeast extracts such as standard, low-hydrolyzed, and high-hydrolyzed types. At the same time, it meets the high-precision prediction requirements of 16 amino acids and solves the problem of insufficient prediction accuracy of traditional methods for complex systems and multi-component systems.

Claims

1. A method for predicting the amino acid content of yeast extract based on hyperspectral imaging and data augmentation, characterized in that, The method includes the following steps: Step 1: Perform spectral purification on the raw hyperspectral data of the target yeast to obtain the purified spectrum; Step 2: Perform multi-strategy collaborative data augmentation on the purified spectrum to obtain the enhanced spectrum, and perform principal component analysis feature extraction on the enhanced spectrum to obtain the core principal component feature matrix of the spectrum. Step 3: Based on the spectral core principal component feature matrix, predict the amino acid content of the target yeast.

2. The method for predicting amino acid content in yeast extract based on hyperspectral imaging and data augmentation according to claim 1, characterized in that, Step 1 specifically includes the following steps: Step 11: Obtain the dark current spectrum and the white board reference spectrum, perform bias elimination processing on the original hyperspectral data, and obtain the reference correction spectrum; Step 12: Perform region screening on the reference calibration spectrum to avoid non-sample areas and extract the spectrum of the target region of interest; Step 13: Perform feature enhancement processing on the spectrum of the target region to highlight local spectral features and eliminate baseline interference, resulting in a purified spectrum.

3. The method for predicting amino acid content in yeast extract based on hyperspectral imaging and data augmentation according to claim 2, characterized in that, Step 11 specifically includes: using the dark current spectrum as the background noise reference and the whiteboard reference spectrum as the standard reflection reference, by calculating the difference between the original hyperspectral data and the dark current spectrum, and then comparing it with the difference between the whiteboard reference spectrum and the dark current spectrum, the deviation between ambient light and system noise is eliminated, and the reference correction spectrum is obtained.

4. The method for predicting amino acid content in yeast extract based on hyperspectral imaging and data augmentation according to claim 2, characterized in that, Step 12 specifically includes: displaying the sample image in the RGB imaging mode corresponding to the reference correction spectrum, excluding non-sample areas including at least the sample disk edge and shadows by threshold segmentation, locating and extracting the spectrum of the effective sample area, and obtaining the spectrum of the target area.

5. The method for predicting amino acid content in yeast extract based on hyperspectral imaging and data augmentation according to claim 2, characterized in that, Step 13 specifically includes: calculating the spectral value difference of the target region at adjacent wavelength points, enhancing the local features of the spectral absorption peaks and reflection peaks through differential operations, and eliminating baseline drift interference to obtain the purified spectrum after feature enhancement.

6. The method for predicting amino acid content in yeast extract based on hyperspectral imaging and data augmentation according to claim 1, characterized in that, Step 2 specifically includes the following steps: Step 21: Perform a multidimensional quality assessment on the purified spectrum to obtain the spectral quality rating results; Step 22: Based on the spectral quality rating results, perform prediction difficulty identification on the purified spectrum to obtain high-difficulty spectral labeling results; Step 23: For spectra with different quality levels and difficulty labels, perform interpolation generation, neighborhood perturbation, distribution completion and density filling processing respectively to obtain multi-type enhanced spectrum sets; Step 24: Perform physical rationality verification on the multi-type enhanced spectrum set and select the compliant enhanced spectrum that meets the spectral characteristics as the enhanced spectrum.

7. The method for predicting amino acid content in yeast extract based on hyperspectral imaging and data augmentation according to claim 6, characterized in that, Step 21 specifically includes: calculating the numerical stability, derivative continuity, and signal-to-noise ratio of the purified spectrum; calculating the comprehensive quality score by weighting the three indicators according to preset weights; dividing the spectrum into high-quality, medium-quality, and low-quality levels according to the score range to obtain the spectral quality rating result.

8. The method for predicting amino acid content in yeast extract based on hyperspectral imaging and data augmentation according to claim 6, characterized in that, Step 22 specifically includes: calling the trained ridge regression model, inputting the purified spectra and corresponding amino acid content data other than the spectrum to be evaluated, predicting the amino acid content of the spectrum to be evaluated through the model, calculating and normalizing the error between the predicted value and the true value, marking the spectra with errors exceeding the set threshold as high-difficulty spectra, and obtaining the high-difficulty spectrum identification results.

9. The method for predicting amino acid content in yeast extract based on hyperspectral imaging and data augmentation according to claim 6, characterized in that, Step 23 specifically includes: selecting spectral pairs from the purified spectra of high quality grade, calculating the Pearson correlation coefficient of the spectral pairs, and retaining only spectral pairs with a correlation coefficient greater than a set value; adaptively allocating interpolation weights according to the difference in the comprehensive quality scores of the spectral pairs, generating new spectra through linear interpolation, and obtaining interpolated enhanced spectra.

10. The method for predicting amino acid content in yeast extract based on hyperspectral imaging and data augmentation according to claim 6, characterized in that, Step 23 specifically includes: for the high-difficulty spectrum, finding a set number of neighboring spectra in the feature space and calculating the mean of the neighboring spectra; applying a local perturbation of a preset amplitude in the direction of the neighboring mean to generate a new spectrum that is similar to the high-difficulty spectrum but has subtle differences, thus obtaining the neighborhood perturbation enhanced spectrum.