Spectral feature extraction and quantitative analysis method and system based on multi-modal fusion
Patent Information
- Application Number
- CN202610990665.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-25
AI Technical Summary
然而,单一模态光谱技术存在固有局限性:分子振动类光谱(如红外、拉曼、近红外光谱)虽能反映分子结构和官能团信息,但对低浓度目标成分的检测灵敏度不足,易受复杂基质干扰;电子跃迁类光谱(如紫外-可见吸收、荧光、磷光光谱)具有高灵敏度优势,但特征峰重叠严重,对分子结构的表征能力有限,单一模态光谱提取的特征信息难以全面、准确地表征复杂样本的化学特性,导致定量分析精度和稳定性受限
[0016]本发明的一种基于多模态融合的光谱特征提取与定量分析方法及系统,采用所述多模态光谱数据采集模块、所述单模态光谱特征提取模块、所述多模态特征融合输出模块进行如下步骤:获取目标样本异质模态光谱数据,经预处理后得到标准化多模态光谱数据集;其中预处理过程包括基线校正、噪声抑制、数据标准化;针对多模态光谱数据,分别提取包含特征峰核心参数、谱图形态表征参数、统计分布特征参数的多维度特征,形成各单模态特征集;构建多模态特征关联矩阵,经动态权重分配得到增强特征集,混合降维后获取最终融合特征集,根据最终融合特征集进行分析处理,并输出分析结果;通过上述方式,实现提高特征表征能力,能够全面反映光谱数据,提升复杂样本目标成分定量分析的精度和稳定性,满足各领域高精度检测的实际应用需求。
Smart Images

Figure CN122821165A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of spectral analysis technology, and in particular to a method and system for spectral feature extraction and quantitative analysis based on multimodal fusion. Background Technology
[0002] Spectroscopic analysis technology, with its advantages of speed, non-destructive nature, and environmental friendliness, has become a core technical means in fields such as food quality testing, pharmaceutical quality control, environmental monitoring, and agricultural product evaluation. Quantitative analysis, as the core application of spectroscopic technology, directly depends on the quality of spectral data and the effective utilization of characteristic information. However, single-modal spectroscopy has inherent limitations: while molecular vibrational spectroscopy (such as infrared, Raman, and near-infrared spectroscopy) can reflect molecular structure and functional group information, it lacks sensitivity for detecting low-concentration target components and is easily affected by complex matrices; electronic transition spectroscopy (such as ultraviolet-visible absorption, fluorescence, and phosphorescence spectroscopy) has high sensitivity, but its characteristic peaks overlap significantly, limiting its ability to characterize molecular structures. The characteristic information extracted by single-modal spectroscopy is insufficient to comprehensively and accurately characterize the chemical properties of complex samples, thus limiting the accuracy and stability of quantitative analysis.
[0003] To address the limitations of single-mode spectroscopy, multi-mode spectral fusion analysis methods have gradually emerged in existing technologies. However, they still suffer from the problem of failing to fully reflect the potential information in spectral data. Single-modal feature extraction has a limited dimension. Existing methods mostly focus only on the basic parameters of the feature peaks (such as peak position and peak intensity), without fully exploring the spectral morphological features (such as slope changes and inflection point distribution) and statistical distribution features (such as data dispersion and distribution symmetry). This results in insufficient feature representation capabilities and an inability to fully reflect the potential information of spectral data.
[0004] Therefore, it is essential to propose a method and system for spectral feature extraction and quantitative analysis that can improve feature representation capabilities and comprehensively reflect the potential information of spectral data. Summary of the Invention
[0005] The purpose of this invention is to provide a method and system for spectral feature extraction and quantitative analysis based on multimodal fusion, which aims to improve feature characterization capabilities and comprehensively reflect the effect of spectral data.
[0006] To achieve the above objectives, this invention employs a method for spectral feature extraction and quantitative analysis based on multimodal fusion, comprising the following steps: Heterogeneous modal spectral data of the target sample are acquired and preprocessed to obtain a standardized multimodal spectral dataset; the preprocessing process includes baseline correction, noise suppression, and data standardization. For multimodal spectral data, multidimensional features including core parameters of characteristic peaks, spectral morphology characterization parameters, and statistical distribution characteristic parameters are extracted to form feature sets for each single mode. A multimodal feature correlation matrix is constructed, and an enhanced feature set is obtained through dynamic weight allocation. After hybrid dimensionality reduction, the final fused feature set is obtained. The final fused feature set is then analyzed and processed, and the analysis results are output.
[0007] In the step of obtaining heterogeneous modal spectral data of the target sample and obtaining a standardized multimodal spectral dataset after preprocessing: Multiple heterogeneous modal spectral data were selected as target samples from molecular vibrational spectra and electronic transition spectra, respectively, and the spectral data of the target samples were collected. Baseline correction is performed on the spectral data to eliminate baseline drift interference. Standardize the spectral data to unify the scale and distribution range of different modal spectral data, eliminate scale differences between data, and obtain a standardized multimodal spectral dataset.
[0008] After the step of baseline correction processing of the spectral data to eliminate baseline drift interference in the spectral data: Noise suppression processing is performed on the baseline-corrected spectral data to filter out random noise and interference signals.
[0009] In the step of extracting multi-dimensional features, including core parameters of characteristic peaks, spectral morphology parameters, and statistical distribution parameters, from multimodal spectral data to form feature sets for each single mode: Identify characteristic peaks in multimodal spectral datasets and select effective characteristic peaks based on the local signal-to-noise ratio of the spectral data and the standard spectrum information of the target components; Extract core parameters from effective characteristic peaks; these core parameters include peak position, peak intensity, peak area, peak width, full width at half maximum (FWHM), and peak shape asymmetry factor. Analyze the overall morphological characteristics of spectral data, extract the mean slope of the spectrum segments, inflection point density, number of local extreme points, and rate of change of waveform curvature to characterize the morphological change law of the spectrum.
[0010] Among these steps, after analyzing the overall morphological characteristics of the spectral data and extracting the mean slope of the spectrum segments, inflection point density, number of local extrema, and rate of change of waveform curvature to characterize the morphological change law of the spectrum: Perform statistical analysis on the spectral data, calculate the mean, variance, standard deviation, skewness, kurtosis, and interquartile range, and explore the distribution characteristics of the data.
[0011] In the step of identifying characteristic peaks in a multimodal spectral dataset and selecting effective characteristic peaks based on the local signal-to-noise ratio of the spectral data and the standard spectrum information of the target components: For each standardized modal spectral data, the data is traversed point by point in the order of wavelength / wavenumber of data acquisition, and multiple local data windows are divided. Calculate the noise standard deviation for the spectral data within each local window, and use four times this standard deviation as the peak height threshold for the current window; Set a peak width threshold range and, in conjunction with the standard spectrum of the target component, clarify the position range and approximate peak width of the known characteristic peaks as a screening reference.
[0012] In the step of identifying characteristic peaks in a multimodal spectral dataset and selecting effective characteristic peaks based on the local signal-to-noise ratio of the spectral data and the standard spectrum information of the target components: During the traversal, when the intensity value of a certain data point exceeds the peak height threshold of the window, and the peak shape and peak width formed by its adjacent data points fall within the set range, it is initially determined to be a candidate feature peak. Compare the positions of candidate characteristic peaks with the positions of characteristic peaks in the standard spectrum of the target component, and eliminate false peaks whose position deviation exceeds a set threshold. Calculate the signal-to-noise ratio (SNR) of the candidate feature peaks, retain peaks with an SNR ≥ 3, and finally obtain the effective feature peaks.
[0013] The steps include: constructing a multimodal feature correlation matrix, obtaining an enhanced feature set through dynamic weight allocation, obtaining a final fused feature set after hybrid dimensionality reduction, performing analysis and processing based on the final fused feature set, and outputting the analysis results. Semantic matching is performed on the feature parameters of each single-modal feature set to unify the feature dimensions and construct a multimodal feature association matrix; Dynamic weights are assigned to each feature parameter in the multimodal feature correlation matrix, and the feature parameters are weighted and fused based on the assigned weights to obtain an enhanced multimodal feature set. The enhanced multimodal feature set is initially dimensionality reduced to retain the main feature information. The feature dimension is then optimized through nonlinear dimensionality reduction to remove redundant information and obtain a final fused feature set with reduced dimensions.
[0014] The process involves several steps, including initial dimensionality reduction of the enhanced multimodal feature set to retain key feature information, optimization of feature dimensions through nonlinear dimensionality reduction to remove redundant information, and obtaining a final fused feature set with reduced dimensions: A machine learning quantitative regression model is constructed based on the final fused feature set. The model is trained using the training set, the model parameters are optimized using the validation set, and the model performance is evaluated using the test set. The samples to be analyzed are then subjected to quantitative analysis, and the quantitative analysis results of the target components are output.
[0015] This invention also provides a spectral feature extraction and quantitative analysis system based on multimodal fusion, comprising a multimodal spectral data acquisition module, a single-modal spectral feature extraction module, and a multimodal feature fusion output module; wherein: The multimodal spectral data acquisition module is used to acquire heterogeneous modal spectral data of the target sample, and obtain a standardized multimodal spectral dataset after preprocessing; wherein the preprocessing process includes baseline correction, noise suppression, and data standardization; The single-mode spectral feature extraction module is used to extract multi-dimensional features, including core parameters of characteristic peaks, spectral morphology representation parameters, and statistical distribution characteristic parameters, for multi-mode spectral data, to form a single-mode feature set; The multimodal feature fusion output module is used to construct a multimodal feature correlation matrix, obtain an enhanced feature set through dynamic weight allocation, obtain the final fused feature set after hybrid dimensionality reduction, perform analysis and processing based on the final fused feature set, and output the analysis results.
[0016] This invention discloses a method and system for spectral feature extraction and quantitative analysis based on multimodal fusion. The method comprises a multimodal spectral data acquisition module, a single-modal spectral feature extraction module, and a multimodal feature fusion output module, performing the following steps: acquiring heterogeneous modal spectral data of the target sample, and obtaining a standardized multimodal spectral dataset after preprocessing; the preprocessing includes baseline correction, noise suppression, and data standardization; for the multimodal spectral data, extracting multidimensional features including core parameters of characteristic peaks, spectral morphology parameters, and statistical distribution parameters to form single-modal feature sets; constructing a multimodal feature correlation matrix, obtaining an enhanced feature set through dynamic weight allocation, and obtaining the final fused feature set after dimensionality reduction; performing analysis based on the final fused feature set and outputting the analysis results; through the above methods, improving feature representation capabilities, comprehensively reflecting spectral data, enhancing the accuracy and stability of quantitative analysis of target components in complex samples, and meeting the practical application needs of high-precision detection in various fields. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of the steps of the spectral feature extraction and quantitative analysis method based on multimodal fusion of the present invention.
[0019] Figure 2This is a flowchart of steps S100 of the present invention.
[0020] Figure 3 This is a flowchart of steps S200 of the present invention.
[0021] Figure 4 This is a flowchart of steps S300 of the present invention.
[0022] Figure 5 This is a schematic diagram of the structural principle of the spectral feature extraction and quantitative analysis system based on multimodal fusion of the present invention.
[0023] Figure 6 This is a schematic diagram of the electronic device of the present invention.
[0024] 401 - Multimodal spectral data acquisition module, 402 - Single-modal spectral feature extraction module, 403 - Multimodal feature fusion output module. Detailed Implementation
[0025] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0026] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0027] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0028] Please see Figures 1-4 This invention provides a method for spectral feature extraction and quantitative analysis based on multimodal fusion, comprising the following steps: S100: Obtain heterogeneous modal spectral data of the target sample, and obtain a standardized multimodal spectral dataset after preprocessing; the preprocessing process includes baseline correction, noise suppression, and data standardization.
[0029] In this embodiment, heterogeneous modal spectral data of the target sample are acquired, and a standardized multimodal spectral dataset is obtained after preprocessing. The preprocessing includes baseline correction, noise suppression, and data standardization. The specific process is as follows: S101: Select multiple heterogeneous modal spectral data from molecular vibrational spectra and electronic transition spectra respectively as target samples, and collect the spectral data of the target samples; S102: Perform baseline correction processing on the spectral data to eliminate baseline drift interference in the spectral data; S103: Perform noise suppression processing on the baseline-corrected spectral data to filter out random noise and interference signals in the data; S104: Standardize the spectral data to unify the scale and distribution range of different modal spectral data, eliminate scale differences between data, and obtain a standardized multimodal spectral dataset.
[0030] In the above process, multiple heterogeneous modal spectral data are selected from molecular vibrational spectroscopy and electronic transition spectroscopy as target samples, and the spectral data of the target samples are acquired. For molecular vibrational spectroscopy, at least two of infrared spectroscopy, Raman spectroscopy, and near-infrared spectroscopy are selected, and for electronic transition spectroscopy, at least two of ultraviolet-visible absorption spectroscopy, fluorescence spectroscopy, and phosphorescence spectroscopy are selected to ensure that the selected modal spectral data have complementary information. A multimodal spectral synchronous acquisition system is adopted, and the acquisition sequence of each spectral detection device is controlled by a synchronous trigger control module to ensure that the acquisition timestamp deviation of different modal spectral data is ≤10ms, so as to avoid the decrease in data correlation caused by time difference. The acquisition environment is controlled at temperature 20±2℃ and humidity 50±5%RH to avoid the influence of temperature and humidity fluctuations on spectral detection equipment and sample characteristics. For target samples in different states such as solids and liquids, appropriate sample cells or detection accessories are used to ensure that the sample detection conditions are consistent and the original multimodal spectral data are acquired.
[0031] An improved adaptive iterative reweighted penalized least squares method (m-airPLS) is used for baseline correction. This method dynamically fits the spectral baseline by constructing a penalty function and an iterative reweighting strategy; the number of iterations is set to 50-100 and the penalty factor is 10. 4 ~10 6 The parameters are adaptively adjusted according to the baseline drift of the spectral data; baseline correction is performed on the original spectral data of each mode, and baseline tilt or drift caused by instrument noise, sample scattering and other factors is eliminated by point-by-point fitting, so that the baseline of the spectral data tends to be stable and the characteristic peak signal is highlighted.
[0032] A threshold denoising method based on wavelet packet transform is adopted. Appropriate wavelet basis functions are selected according to the modal type of the spectral data (e.g., db6 wavelet for vibrational spectra, sym5 wavelet for electronic transition spectra). The wavelet decomposition level is set to 4-6 levels. Wavelet packet decomposition is performed on the baseline-corrected spectral data to obtain wavelet coefficients at different frequency scales. An improved heuristic threshold function is used to threshold the wavelet coefficients, retaining large coefficients reflecting spectral characteristics and suppressing small coefficients corresponding to noise. The spectral data is then reconstructed through inverse wavelet packet transform, achieving the filtering of random noise and interference signals and improving the discriminability of spectral features.
[0033] Based on the characteristics of different modal spectral data, an appropriate standardization method is selected: Standard Normal Variable Transform (SNV) is used for modes with large peak intensity differences (such as Raman spectroscopy and fluorescence spectroscopy), and Z-score standardization is used for modes with significant differences in data distribution range. SNV processing converts the original data into standardized data with a mean of 0 and a standard deviation of 1 by calculating the mean and standard deviation of each data point relative to the full spectrum data, thereby eliminating baseline shifts and intensity fluctuations caused by uneven sample granularity and concentration distribution. Z-score standardization achieves numerical scale uniformity for different modal spectral data by calculating the difference between each data point and the mean of the dataset and dividing it by the standard deviation of the dataset. All modal spectral data after baseline correction and noise suppression are standardized separately, and finally integrated to form a standardized multimodal spectral dataset with uniform structure and consistent scale.
[0034] S200: For multimodal spectral data, extract multi-dimensional features including core parameters of characteristic peaks, spectral morphology characterization parameters, and statistical distribution characteristic parameters to form feature sets for each single mode.
[0035] In this embodiment, for multimodal spectral data, multidimensional features including characteristic peak core parameters, spectral morphology characterization parameters, and statistical distribution characteristic parameters are extracted to form feature sets for each single mode. The specific process is as follows: S201: Identify the characteristic peaks in the multimodal spectral dataset and select the effective characteristic peaks based on the local signal-to-noise ratio of the spectral data and the standard spectrum information of the target components; For each standardized modal spectral data, the data is traversed point by point in the order of wavelength / wavenumber of data acquisition, and multiple local data windows are divided. Calculate the noise standard deviation for the spectral data within each local window, and use four times this standard deviation as the peak height threshold for the current window; Set a peak width threshold range and combine it with the standard spectrum of the target component to clarify the position range and approximate peak width of the known characteristic peaks as a screening reference. During the traversal, when the intensity value of a certain data point exceeds the peak height threshold of the window, and the peak shape and peak width formed by its adjacent data points fall within the set range, it is initially determined to be a candidate feature peak. Compare the positions of candidate characteristic peaks with the positions of characteristic peaks in the standard spectrum of the target component, and eliminate false peaks whose position deviation exceeds a set threshold. Calculate the signal-to-noise ratio of the candidate feature peaks, retain the peaks with a signal-to-noise ratio ≥ 3, and finally obtain the effective feature peaks; S202: Extract core parameters from effective characteristic peaks; the core parameters include peak position, peak intensity, peak area, peak width, full width at half maximum (FWHM), and peak shape asymmetry factor. S203: Analyze the overall morphological characteristics of spectral data, extract the average slope of the spectrum segments, inflection point density, number of local extreme points, and waveform curvature change rate to characterize the morphological change law of the spectrum. S204: Perform statistical analysis on spectral data, calculate the mean, variance, standard deviation, skewness, kurtosis, and interquartile range, and explore the distribution characteristics of the data.
[0036] In the above process, for each standardized modal spectral data, the data is traversed point by point according to the wavelength / wavenumber order of data acquisition, and multiple local data windows are divided (the vibrational spectral window size is 5~10cm). -1 The electronic transition spectral window size is 3~8nm, ensuring that each window covers the local spectral structure without missing any feature information. The noise standard deviation is calculated for the spectral data within each local window, and four times this standard deviation is used as the peak height threshold for the current window. This dynamically adapts to the noise level in different regions, avoiding missed or false detections caused by fixed thresholds. Combined with standard spectra of the target component (such as standard spectra in national standard spectral libraries or professional spectral databases), the location range of known characteristic peaks is clearly defined (allowing ±2cm). -1 (or ±1 nm deviation) and approximate peak width range (vibrational spectrum 1-20 cm⁻¹) -1 (electronic transition spectrum 1~15nm) was used as a screening reference. During the traversal, when the intensity value of a certain data point exceeds the peak height threshold of its window, and its adjacent data points form a complete peak structure of "rising first and then falling", and the peak width falls within the set range, it is initially determined to be a candidate feature peak. The position of the candidate feature peak is compared with the position of the feature peak in the standard spectrum of the target component, and false peaks with position deviations exceeding the set threshold (such as false signals caused by instrument interference or matrix background) are eliminated. The signal-to-noise ratio (SNR) of the candidate feature peak (the ratio of the peak intensity to the standard deviation of the local window noise) is calculated, and peaks with an SNR ≥ 3 are retained. Finally, the effective feature peaks after noise filtering and position verification are obtained.
[0037] Extracting core parameters: Peak position is determined by the wavelength or wavenumber corresponding to the peak of the effective feature. If there are multiple consecutive data points of equal intensity at the peak, the middle value of the interval is taken as the peak position. Peak intensity: Directly read the spectral intensity value of the effective characteristic peak (such as absorbance, Raman shift intensity, fluorescence intensity, etc., consistent with the spectral data type); Peak area is calculated by integrating the spectral data within the peak boundary using the intersection of the baselines on both sides of the effective characteristic peak (the baselines are obtained by fitting the lowest values on both sides of the peak). This yields the peak area value, which reflects the relevant information about the component content corresponding to the characteristic peak. Peak width is defined as the wavelength / wavenumber span between the baseline intersection points on both sides of the effective characteristic peak, i.e., the difference in the horizontal coordinates between the two intersection points. For half-width at half-height, first determine the intensity value corresponding to half the peak height of the effective feature peak, then find the two intersection points of this intensity value with the left and right sides of the peak curve. The horizontal coordinate span between the two intersection points is the half-width at half-height. The peak shape asymmetry factor is calculated by taking the effective feature peak as the center, measuring the width on the left (the horizontal coordinate distance from the peak peak to the left intersection point) and the width on the right (the horizontal coordinate distance from the peak peak to the right intersection point) at 1 / 2 of the peak height, and calculating the ratio of the right width to the left width as the peak shape asymmetry factor (ratio = 1 indicates a symmetrical peak, ratio > 1 indicates a right-skewed peak, and ratio < 1 indicates a left-skewed peak).
[0038] Analyze the overall morphological characteristics of the spectral data, extract the mean slope of the spectrum segments, inflection point density, number of local extrema, and rate of change of waveform curvature to characterize the morphological variation patterns of the spectrum: The average slope of the spectral segments is used to uniformly divide the single-mode normalized spectral data into several continuous segments (every 50 cm⁻¹) according to the wavelength / wavenumber range. -1(or 30nm as a segment), the slope of adjacent data points in each segment is calculated sequentially (the difference between the intensity of the next data point and the intensity of the previous data point divided by the difference in the horizontal coordinate), and then the arithmetic mean of all slopes in the segment is calculated. Each segment corresponds to a slope mean, forming a segmented slope mean sequence, which reflects the overall trend of the spectrum in different intervals. Inflection point density is determined by first identifying inflection points in the spectral data (defined as the critical point where the slope changes from positive to negative or from negative to positive, determined by judging the direction of slope change of three adjacent data points), counting the total number of inflection points in the entire spectral data, and then dividing by the total length of the spectrum (wavelength / wavenumber span) to obtain the inflection point density per unit length, which reflects the complexity of the spectral morphology. The number of local extrema is determined by identifying local maxima (the point with the highest intensity among adjacent data points) and local minima (the point with the lowest intensity among adjacent data points). The total number of both is counted as the number of local extrema, reflecting the frequency of fluctuations in the spectrum. The waveform curvature change rate is obtained by fitting a local curve using the three-point method for each data point (using the current point and one adjacent point before and after as the fitting object), obtaining the curvature value of that point, then calculating the curvature difference between adjacent data points, summing the absolute values of all curvature differences and dividing by the total number of data points, thus obtaining the waveform curvature change rate, which reflects the intensity of fluctuations in the spectral shape.
[0039] Calculate the mean, variance, standard deviation, skewness, kurtosis, and interquartile range of the data to uncover its distribution characteristics. The average intensity of the intensity data is obtained by collecting the intensity values of all valid data points in the single-mode spectral data to form an intensity data sequence, summing all intensity values and dividing by the total number of data points. Variance is calculated by first determining the difference between each intensity value and the mean, then squaring each difference, and finally averaging all the squared values to reflect the dispersion of the data. Standard deviation, which is the arithmetic square root of variance, has the same unit as intensity data and more intuitively reflects the range of data fluctuation. Skewness, based on variance and mean, is calculated by taking the cube of the difference between each intensity value and the mean, averaging all cubed values, and then dividing by the cube of the standard deviation. A positive skewness value indicates that the data distribution is right-skewed, a negative value indicates left-skewed, and a value close to 0 indicates that the distribution is symmetrical. Kurtosis is calculated by summing the fourth power of the differences between each intensity value and the mean, averaging the sums, dividing by the fourth power of the standard deviation, and finally subtracting 3 (the kurtosis of a normal distribution is 3). A kurtosis greater than 0 indicates a steeper data distribution, while a kurtosis less than 0 indicates a flatter distribution. The interquartile range (IQR) is calculated by first sorting the intensity data sequence in ascending order, then dividing the sorted sequence into four equal parts. The data at the 25th percentile is the first quartile (Q1), and the data at the 75th percentile is the third quartile (Q3). The IQR is the difference between Q3 and Q1, reflecting the dispersion of the middle 50% of the data, and is not affected by extreme values.
[0040] S300: Construct a multimodal feature correlation matrix, obtain an enhanced feature set through dynamic weight allocation, obtain the final fused feature set after hybrid dimensionality reduction, perform analysis and processing based on the final fused feature set, and output the analysis results.
[0041] In this embodiment, a multimodal feature correlation matrix is constructed, an enhanced feature set is obtained through dynamic weight allocation, and a final fused feature set is obtained after hybrid dimensionality reduction. The final fused feature set is then analyzed and processed, and the analysis results are output. The specific process is as follows: S301: Perform semantic matching on the feature parameters of each single-modal feature set, unify the feature dimensions, and construct a multimodal feature association matrix; S302: Dynamically assign weights to each feature parameter in the multimodal feature correlation matrix, and perform weighted fusion of the feature parameters based on the assigned weights to obtain an enhanced multimodal feature set; S303: Perform preliminary dimensionality reduction on the enhanced multimodal feature set, retain the main feature information, optimize the feature dimension through nonlinear dimensionality reduction, remove redundant information, and obtain the final fused feature set with reduced dimensions; S304: Construct a machine learning quantitative regression model based on the final fused feature set, train the model using the training set, optimize the model parameters using the validation set, evaluate the model performance using the test set, perform quantitative analysis on the samples to be analyzed, and output the quantitative analysis results of the target components.
[0042] In the above process, based on the correlation rules of spectral physical properties (combining Lambert-Beer law and molecular spectral transition mechanism), the physical meaning and correlation of different modal characteristic parameters are clarified. For example, the peak position parameters corresponding to amide bond vibration in infrared spectrum are semantically correlated with the peak position parameters corresponding to electronic transition of aromatic amino acids in ultraviolet-visible spectrum. The feature parameters of each single-modal feature set are classified and organized, and the naming rules and data formats are unified according to the classification logic of "feature peak core parameters - spectral morphology characterization parameters - statistical distribution feature parameters" to achieve semantic matching of feature parameters; Count the number of feature parameters in each single modality feature set, and unify the feature dimensions of all modalities by padding with zeros or feature expansion (for example, expand the 16 feature parameters of a certain modality to 27 dimensions consistent with other modalities) to ensure that the feature sets of each modality have the same dimensions. The single-modal feature sets after unifying the dimensions are concatenated column by column to form a multi-row, multi-column multimodal feature association matrix. The rows of the matrix correspond to the number of samples, and the columns correspond to the feature parameters of all modalities.
[0043] Construct a hybrid deep attention network model combining multi-head attention mechanism (Multi-HeadAttention) and convolutional neural network (CNN), with the input layer dimension of the model having the same number of columns as the multimodal feature association matrix; The model is trained using the actual content of the target component in the target sample (determined by standard detection methods such as high performance liquid chromatography and Kjeldahl nitrogen determination) as a supervision signal: the multimodal feature correlation matrix is standardized and input into the model, the root mean square error (RMSE) is used as the loss function, the model parameters are optimized by the adaptive momentum estimation algorithm (AdamW), and iterative training is performed until the loss function converges and the accuracy of the validation set is stable. After training, the model outputs the dynamic weight coefficients of each feature parameter (weight range is [0,1]). The higher the weight coefficient, the greater the influence of the feature parameter on the quantitative analysis results. Based on the output weight coefficients, each feature parameter in the multimodal feature correlation matrix is weighted (feature parameter value × corresponding weight coefficient), and then all weighted feature parameters are concatenated column by column to obtain an enhanced multimodal feature set that highlights effective features and suppresses redundant information.
[0044] The enhanced multimodal feature set is initially reduced in dimensionality to retain the main feature information. The feature dimension is then optimized through nonlinear dimensionality reduction to remove redundant information and obtain a final fused feature set with reduced dimensionality. Principal component analysis (PCA) is used for preliminary dimensionality reduction: the covariance matrix of the enhanced multimodal feature set is calculated, the eigenvalues and eigenvectors are solved, the eigenvalues are sorted from largest to smallest, and the principal components with a cumulative contribution rate ≥90% are selected to obtain the feature set after preliminary dimensionality reduction, thus achieving linear dimensionality compression. The Local Linear Embedding (LLE) algorithm is used for nonlinear dimensionality reduction optimization: the feature set after initial dimensionality reduction is used as input to construct the k nearest neighbor graph of the sample (k takes the value of 5~10), the nonlinear relationship of the sample is represented by the local linear reconstruction, and the local structural information of the sample is preserved by the low-dimensional space mapping, and redundant information and noise are further removed. After PCA+LLE hybrid dimensionality reduction, a final fused feature set with reduced dimensions and high information density is obtained, which not only retains the complementary information of each modality feature, but also improves the effectiveness and computational efficiency of the features.
[0045] A machine learning quantitative regression model is constructed based on the final fused feature set. The model is trained using a training set, its parameters are optimized using a validation set, and its performance is evaluated using a test set. The samples to be analyzed are then subjected to quantitative analysis, and the quantitative analysis results of the target components are output. The dataset is divided into training, validation and test sets in a 7:1:2 ratio, containing the final fused feature set and the true content of the corresponding target components, to ensure the consistency of the dataset distribution. For model construction and training, Lightweight Gradient Boosting Machine (LightGBM) was selected as the quantitative regression model. The final fused feature set and corresponding true values of the input training set were used to train the model, and ten-fold cross-validation was employed to avoid overfitting. Model parameter optimization uses the RMSE of the validation set as the evaluation metric. The model hyperparameters are optimized using the grid search method, including the learning rate (0.05~0.2), the number of decision trees (100~200), and the maximum tree depth (3~8), until the model achieves optimal performance on the validation set. Model performance evaluation involves inputting the final fused feature set of the test set into the optimized model and outputting predicted values using the coefficient of determination (R²). 2 RMSE and mean absolute percentage error (MAPE) are used as evaluation metrics, requiring the model to satisfy R... 2 ≥0.96, RMSE≤3%, MAPE≤2%; Quantitative analysis of the sample to be analyzed: Data collection, preprocessing, feature extraction and fusion are performed on the target sample to be analyzed according to steps S100-S303 to obtain the final fused feature set of the sample to be analyzed. The fused feature set is then input into the trained and optimized quantitative regression model, and the model outputs the quantitative analysis results of the target components in the sample to be analyzed.
[0046] Corresponding to the aforementioned embodiments of the spectral feature extraction and quantitative analysis method based on multimodal fusion, this application also provides embodiments of a spectral feature extraction and quantitative analysis system based on multimodal fusion.
[0047] Figure 5 This is a block diagram of a spectral feature extraction and quantitative analysis system based on multimodal fusion, according to an exemplary embodiment. (Refer to...) Figure 5 The system may include: a multimodal spectral data acquisition module 401, a single-modal spectral feature extraction module 402, and a multimodal feature fusion output module 403; wherein: The multimodal spectral data acquisition module 401 is used to acquire heterogeneous modal spectral data of the target sample, and obtain a standardized multimodal spectral dataset after preprocessing; wherein the preprocessing process includes baseline correction, noise suppression, and data standardization; The single-mode spectral feature extraction module 402 is used to extract multi-dimensional features, including core parameters of characteristic peaks, spectral morphology characterization parameters, and statistical distribution characteristic parameters, for multi-mode spectral data, to form a single-mode feature set; The multimodal feature fusion output module 403 is used to construct a multimodal feature correlation matrix, obtain an enhanced feature set through dynamic weight allocation, obtain the final fused feature set after hybrid dimensionality reduction, perform analysis and processing based on the final fused feature set, and output the analysis results.
[0048] In this embodiment, the multimodal spectral data acquisition module 401 acquires heterogeneous modal spectral data of the target sample, and obtains a standardized multimodal spectral dataset after preprocessing. The preprocessing process includes baseline correction, noise suppression, and data standardization. The single-modal spectral feature extraction module 402 extracts multi-dimensional features, including core parameters of characteristic peaks, spectral morphology parameters, and statistical distribution parameters, from the multimodal spectral data to form single-modal feature sets. The multimodal feature fusion output module 403 constructs a multimodal feature correlation matrix, obtains an enhanced feature set through dynamic weight allocation, and obtains the final fused feature set after dimensionality reduction. The final fused feature set is then analyzed and processed, and the analysis results are output. Through the above methods, the feature representation capability is improved, which can comprehensively reflect the spectral data, improve the accuracy and stability of quantitative analysis of target components in complex samples, and meet the practical application needs of high-precision detection in various fields.
[0049] Regarding the system in the above embodiments, the specific ways in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0050] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0051] Accordingly, this application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; and, when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the above-described method for spectral feature extraction and quantitative analysis based on multimodal fusion. Figure 6The diagram shown is a hardware structure diagram of any data processing device, in which a spectral feature extraction and quantitative analysis system based on multimodal fusion based on an embodiment of the present invention is located. Except for... Figure 6 In addition to the processor, memory, and network interface shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.
[0052] Accordingly, this application also provides a computer-readable storage medium storing computer instructions thereon, which, when executed by a processor, implement the multimodal fusion-based spectral feature extraction and quantitative analysis method described above. The computer-readable storage medium can be an internal storage unit of any data processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units of any data processing device and external storage devices. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.
[0053] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0054] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A method for spectral feature extraction and quantitative analysis based on multimodal fusion, characterized in that, Includes the following steps: Obtain heterogeneous modal spectral data of the target sample, and obtain a standardized multimodal spectral dataset after preprocessing; The preprocessing process includes baseline correction, noise suppression, and data standardization. For multimodal spectral data, multidimensional features including core parameters of characteristic peaks, spectral morphology characterization parameters, and statistical distribution characteristic parameters are extracted to form feature sets for each single mode. A multimodal feature correlation matrix is constructed, and an enhanced feature set is obtained through dynamic weight allocation. After hybrid dimensionality reduction, the final fused feature set is obtained. The final fused feature set is then analyzed and processed, and the analysis results are output.
2. The method for spectral feature extraction and quantitative analysis based on multimodal fusion as described in claim 1, characterized in that, In the step of obtaining heterogeneous modal spectral data of the target sample and preprocessing it to obtain a standardized multimodal spectral dataset: Multiple heterogeneous modal spectral data were selected as target samples from molecular vibrational spectra and electronic transition spectra, respectively, and the spectral data of the target samples were collected. Baseline correction is performed on the spectral data to eliminate baseline drift interference. Standardize the spectral data to unify the scale and distribution range of different modal spectral data, eliminate scale differences between data, and obtain a standardized multimodal spectral dataset.
3. The method for spectral feature extraction and quantitative analysis based on multimodal fusion as described in claim 2, characterized in that, After performing baseline correction on the spectral data to eliminate baseline drift interference: Noise suppression processing is performed on the baseline-corrected spectral data to filter out random noise and interference signals.
4. The method for spectral feature extraction and quantitative analysis based on multimodal fusion as described in claim 1, characterized in that, In the step of extracting multi-dimensional features, including core parameters of characteristic peaks, spectral morphology parameters, and statistical distribution parameters, from multimodal spectral data to form feature sets for each single mode: Identify characteristic peaks in multimodal spectral datasets and select effective characteristic peaks based on the local signal-to-noise ratio of the spectral data and the standard spectrum information of the target components; Extract core parameters from effective characteristic peaks; these core parameters include peak position, peak intensity, peak area, peak width, full width at half maximum (FWHM), and peak shape asymmetry factor. Analyze the overall morphological characteristics of spectral data, extract the mean slope of the spectrum segments, inflection point density, number of local extreme points, and rate of change of waveform curvature to characterize the morphological change law of the spectrum.
5. The method for spectral feature extraction and quantitative analysis based on multimodal fusion as described in claim 4, characterized in that, After analyzing the overall morphological characteristics of the spectral data and extracting the mean slope of the spectrum segments, inflection point density, number of local extrema, and rate of change of waveform curvature to characterize the morphological variation law of the spectrum: Perform statistical analysis on the spectral data, calculate the mean, variance, standard deviation, skewness, kurtosis, and interquartile range, and explore the distribution characteristics of the data.
6. The method for spectral feature extraction and quantitative analysis based on multimodal fusion as described in claim 4, characterized in that, In the step of identifying characteristic peaks in a multimodal spectral dataset and selecting effective characteristic peaks based on the local signal-to-noise ratio of the spectral data and the standard spectrum information of the target component: For each standardized modal spectral data, the data is traversed point by point in the order of wavelength / wavenumber of data acquisition, and multiple local data windows are divided. Calculate the noise standard deviation for the spectral data within each local window, and use four times this standard deviation as the peak height threshold for the current window; Set a peak width threshold range and, in conjunction with the standard spectrum of the target component, clarify the position range and approximate peak width of the known characteristic peaks as a screening reference.
7. The method for spectral feature extraction and quantitative analysis based on multimodal fusion as described in claim 6, characterized in that, In the step of identifying characteristic peaks in a multimodal spectral dataset and selecting effective characteristic peaks based on the local signal-to-noise ratio of the spectral data and the standard spectrum information of the target component: During the traversal, when the intensity value of a certain data point exceeds the peak height threshold of the window, and the peak shape and peak width formed by its adjacent data points fall within the set range, it is initially determined to be a candidate feature peak. Compare the positions of candidate characteristic peaks with the positions of characteristic peaks in the standard spectrum of the target component, and eliminate false peaks whose position deviation exceeds a set threshold. Calculate the signal-to-noise ratio (SNR) of the candidate feature peaks, retain peaks with an SNR ≥ 3, and finally obtain the effective feature peaks.
8. The method for spectral feature extraction and quantitative analysis based on multimodal fusion as described in claim 1, characterized in that, In the steps of constructing a multimodal feature correlation matrix, obtaining an enhanced feature set through dynamic weight allocation, obtaining a final fused feature set after hybrid dimensionality reduction, performing analysis and processing based on the final fused feature set, and outputting the analysis results: Semantic matching is performed on the feature parameters of each single-modal feature set to unify the feature dimensions and construct a multimodal feature association matrix; Dynamic weights are assigned to each feature parameter in the multimodal feature correlation matrix, and the feature parameters are weighted and fused based on the assigned weights to obtain an enhanced multimodal feature set. The enhanced multimodal feature set is initially dimensionality reduced to retain the main feature information. The feature dimension is then optimized through nonlinear dimensionality reduction to remove redundant information and obtain a final fused feature set with reduced dimensions.
9. The method for spectral feature extraction and quantitative analysis based on multimodal fusion as described in claim 8, characterized in that, After performing initial dimensionality reduction on the enhanced multimodal feature set, retaining the main feature information, optimizing the feature dimension through nonlinear dimensionality reduction, eliminating redundant information, and obtaining the final fused feature set with reduced dimensions: A machine learning quantitative regression model is constructed based on the final fused feature set. The model is trained using the training set, the model parameters are optimized using the validation set, and the model performance is evaluated using the test set. The samples to be analyzed are then subjected to quantitative analysis, and the quantitative analysis results of the target components are output.
10. A spectral feature extraction and quantitative analysis system based on multimodal fusion, employing the spectral feature extraction and quantitative analysis method based on multimodal fusion as described in claim 1, characterized in that, It includes a multimodal spectral data acquisition module, a single-modal spectral feature extraction module, and a multimodal feature fusion output module; wherein: The multimodal spectral data acquisition module is used to acquire heterogeneous modal spectral data of the target sample, and obtain a standardized multimodal spectral dataset after preprocessing; wherein the preprocessing process includes baseline correction, noise suppression, and data standardization; The single-mode spectral feature extraction module is used to extract multi-dimensional features, including core parameters of characteristic peaks, spectral morphology representation parameters, and statistical distribution characteristic parameters, for multi-mode spectral data, to form a single-mode feature set; The multimodal feature fusion output module is used to construct a multimodal feature correlation matrix, obtain an enhanced feature set through dynamic weight allocation, obtain the final fused feature set after hybrid dimensionality reduction, perform analysis and processing based on the final fused feature set, and output the analysis results.