Method and apparatus for identifying cashmere and wool composition based on infrared spectroscopy
By preparing standard samples covering the entire range, collecting infrared spectral data, normalizing and smoothing them, and constructing a Mahalanobis distance-weighted regression model, the problems of spectral overlap and insufficient model stability in cashmere and wool composition identification were solved, and high-precision cashmere content prediction was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies for identifying cashmere and wool components suffer from severe spectral overlap, insufficient feature extraction, and inadequate model prediction stability, especially in identifying low-content components and in complex sample backgrounds where accuracy is limited.
By preparing standard samples covering the entire range, collecting infrared spectral data, performing normalization and smoothing, calculating the second derivative, constructing a Mahalanobis distance-weighted regression model, and using Mahalanobis distance to reflect the similarity between spectral features and class centers, an accurate method for predicting cashmere content is established.
It achieves accurate identification of cashmere and wool components, improves identification accuracy and robustness, and can stably predict cashmere content in complex sample backgrounds.
Smart Images

Figure CN121409902B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cashmere and wool content detection technology, specifically to a method and apparatus for identifying cashmere and wool component content based on infrared spectroscopy technology. Background Technology
[0002] As high-end textile raw materials, the accurate identification of the composition of cashmere and wool is crucial for product quality control, market regulation, and consumer protection. Traditional composition detection methods, such as microscopic observation and chemical dissolution, are cumbersome, time-consuming, highly dependent on operator experience, and may damage the sample, making them unsuitable for the modern textile industry's demands for rapid, non-destructive, and high-precision detection. While infrared spectroscopy has been applied to fiber composition analysis, it still faces challenges in the quantitative identification of mixed cashmere and wool samples, including severe spectral overlap, insufficient feature extraction, and inadequate model prediction stability, particularly in identifying low-content components and complex sample backgrounds. Therefore, developing a method that can effectively extract the spectral characteristics of cashmere and wool and improve the accuracy and robustness of composition identification has become an urgent technical problem to be solved in this field.
[0003] In the prior art, CN117092059A discloses a method for identifying the composition and content of cashmere and wool based on infrared spectroscopy, which includes: first measuring the cashmere and wool textiles with known composition and content using a spectrometer; normalizing the measured spectral data and then removing abnormal samples; converting the processed data into a two-dimensional infrared correlation spectrum; dividing the two-dimensional infrared correlation spectrum into a training set and a test set using the KS method; establishing a model using the training set data; and verifying the model's effectiveness using the validation set data.
[0004] The main problems with the above scheme are: two-dimensional infrared spectroscopy cannot completely solve the spectral overlap problem of cashmere and wool, two protein fibers with extremely similar chemical structures, resulting in insufficient accuracy of the extracted data; the infrared spectra of cashmere and wool are not a simple linear superposition, and their spectral characteristics change in complex ways with the change of mixing ratio. The above scheme fails to take this into account, resulting in insufficient adaptability to actual samples; when using the training set to build the model, it is assumed that all samples contribute equally to the model. In reality, the spectral characteristics of samples with different content ratios are quite different, and their reliability is also different. If modeling is performed directly, it will lead to inaccurate modeling results.
[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this invention is to provide a method and apparatus for identifying the content of cashmere and wool components based on infrared spectroscopy technology, so as to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution:
[0008] A method for identifying the composition content of cashmere and wool based on infrared spectroscopy technology, comprising the following steps:
[0009] Step 1: Prepare several standard samples with known cashmere and wool content, and collect infrared spectral data of each standard sample;
[0010] Step 2: Normalize the infrared spectral data of each standard sample and extract the amide I and amide II bands. Smooth the normalized infrared spectral data and calculate the second derivative. Concatenate the second derivatives of the two spectral bands to generate a feature vector.
[0011] Step 3: Construct a data matrix based on the feature vectors of all standard samples, perform principal component analysis on the data matrix and generate a score matrix, extract each row vector from the score matrix as the principal component score vector of different standard samples, and use the principal component score vectors of pure wool samples and pure cashmere samples as the wool center vector and cashmere center vector, respectively.
[0012] Step 4: Calculate the Mahalanobis distance from the principal component score vector of each standard sample to the wool center vector and the cashmere center vector, respectively, construct a Mahalanobis distance weighted regression model, calculate the principal component score vector of the sample to be tested, and substitute it into the Mahalanobis distance weighted regression model to predict the cashmere content.
[0013] Furthermore, the cashmere content in the standard samples ranges from 0% to 100%, with a gradient set at 10% intervals, and one standard sample prepared for each gradient.
[0014] The specific principle for acquiring infrared spectral data of standard samples is as follows: a Fourier transform infrared spectrometer equipped with an ATR crystal is used, with a scanning wavenumber range of 4000~600. The resolution is set to 4. Several scanning wavenumber samples are generated within the scanning wavenumber range according to the resolution. For each scanning wavenumber sample, the ATR crystal is first scanned without any sample, and the background interferogram is acquired. Then, each standard sample is pressed tightly onto the ATR crystal, and its single-spectral interferogram is acquired. Fourier transform is performed on the background interferogram and the single-spectral interferogram of each standard sample to convert them into background single-beam spectra and sample single-beam spectra, respectively. The sample single-beam spectrum is divided by the background single-beam spectrum to obtain the transmittance spectrum of each standard sample. The negative logarithm of the transmittance spectrum is taken to obtain the absorbance spectrum corresponding to each scanning wavenumber sample. The absorbance spectra corresponding to all scanning wavenumber samples under the same standard sample are summarized to construct a spectral vector, which serves as the infrared spectral data of the standard sample.
[0015] Furthermore, the principle for generating feature vectors is as follows:
[0016] For each standard sample, its spectral vector is: ,in, Represents the spectral vector of the standard sample. Indicates the first in the standard sample Absorbance spectrum of each scan wavenumber sample value Indicates the index of the scan wavenumber sample value, and , This represents the total number of wavenumber samples taken for a single standard sample.
[0017] The formula for normalizing the spectral vector is:
[0018]
[0019]
[0020] in, Represents the normalized spectral vector. express The modulus length;
[0021] Extracting spectral frequency bands from the normalized spectral vector is specifically as follows:
[0022] Amide I band: wavenumber range 1700~1600 ;
[0023] Amide II band: wavenumber range 1580~1480 ;
[0024] The principle of smoothing is as follows: For each absorbance spectrum in the extracted spectral band, an equal-length window is divided with that window as the center. The average absorbance spectrum within the equal-length window is calculated and replaced with that absorbance spectrum. The second derivative of the replaced absorbance spectrum is then calculated, and the second derivatives of all equal-length windows are concatenated to form a feature vector. ,in, Represents the eigenvector. This indicates the first frequency band of the extracted spectrum. The second derivative of the absorbance spectrum corresponding to each scan wavenumber sample value This represents the index of the scan wavenumber sample value in the extracted spectral frequency band, and , This represents the total number of scan wavenumber samples in the extracted spectral band.
[0025] Furthermore, the principle for generating the score matrix is as follows:
[0026] The constructed data matrix is as follows:
[0027]
[0028] in, Represents a data matrix. Indicates the first The feature vector of a standard sample Indicates the index of the standard sample, and , Indicates the quantity of standard samples. Represents the transpose matrix;
[0029] Calculate the average of each column in the data matrix and generate a centered data matrix. The specific formula is as follows:
[0030]
[0031]
[0032] in, Indicates the first The average of the second derivatives of the absorbance spectrum corresponding to each scan wavenumber sample value. Indicates the first The standard sample at the first The second derivative of the absorbance spectrum corresponding to each scan wavenumber sample value Represents a centralized data matrix. express A column vector of all 1s Let represent the vector of the second derivative of the absorbance spectrum, and Represents the th element in the centered matrix The row vectors corresponding to each standard sample;
[0033] The covariance matrix is calculated based on a centralized data matrix, using the following formula:
[0034]
[0035] in, Represents the covariance matrix. express The transpose of the matrix;
[0036] The eigenvalue decomposition of the covariance matrix is based on the following formula:
[0037]
[0038] in, Indicates the first The eigenvectors of the principal components Indicates the first The eigenvalues of the principal components Indicates the index of the principal component, and ,
[0039] The variance contribution rate of each principal component is calculated, and then the cumulative variance contribution rate of the principal components is calculated using the following formula:
[0040]
[0041]
[0042] in, Indicates the first The variance contribution rate of each principal component Indicates the preceding The cumulative variance contribution rate of each principal component Indicates the index of the principal component, and ;
[0043] from Start, gradually increase The value of is calculated, and the corresponding value is determined. until the smallest one is found. , making Before choosing The eigenvectors constitute the loading matrix: ;
[0044] The formula for calculating the score matrix is:
[0045]
[0046] in, This represents the score matrix.
[0047] Furthermore, extract each row vector from the score matrix. , serving as the principal component score vector for each standard sample, where, Represents the score matrix of the first... There are n row vectors, and the center vector of the wool class is n. The center vector of the cashmere class is .
[0048] Furthermore, the principle of constructing the Mahalanobis distance regression model is as follows:
[0049] The formula used to calculate the Mahalanobis distance from the standard sample to the center vector of the wool is as follows:
[0050]
[0051] in, Indicates the first Mahalanobis distance from the standard sample to the center vector of the wool class Let represent the covariance matrix of the principal component scores, and For the front A diagonal matrix of principal component eigenvalues;
[0052] The formula used to calculate the Mahalanobis distance from the standard sample to the center vector of the cashmere sample is:
[0053]
[0054] in, Indicates the first Mahalanobis distance from a standard sample to the center vector of the cashmere category;
[0055] The formula used to calculate the weights of standard samples in the regression fitting is:
[0056]
[0057] in, Indicates the first The weight of each standard sample, Indicates the attenuation coefficient. express and The minimum value between;
[0058] A dependent variable vector is constructed based on the known cashmere content of all standard samples, specifically: ,in, Represents the dependent variable vector. Indicates the first The cashmere content of each standard sample was determined; based on the scores of all standard samples on the first k principal components, a [structure / construction] was built. The design matrix is as follows:
[0059]
[0060] in, Represents the design matrix. Indicates the first The sample in the first Scores on each principal component;
[0061] The weight matrix is constructed based on the weights of the standard samples, specifically as follows:
[0062]
[0063] in, Represents the weight matrix;
[0064] Find a regression coefficient vector This makes the weighted sum of squared residuals of all standard samples... Minimum; fitted based on least squares method The Mahalanobis distance regression model is as follows:
[0065]
[0066] in, Represents the regression coefficient vector. Represents the design matrix of the first... Row vectors.
[0067] Furthermore, the principle for predicting cashmere content is as follows:
[0068] For the sample to be tested, feature vectors are generated using the same method as for the standard sample. ,in, This indicates the frequency of the sample under test within the extracted spectral band. The second derivative of the absorbance spectrum corresponding to each scan wavenumber sample value, After centering, the principal component score vector of the sample to be tested is calculated. The specific formula is as follows: Then construct the prediction vector. ,Will Substitute the values into the Mahalanobis distance regression model to generate predicted cashmere content values. .
[0069] The present invention also provides a cashmere and wool composition content identification device based on infrared spectroscopy technology. The device is used to implement the above-mentioned cashmere and wool composition content identification method based on infrared spectroscopy technology, specifically including:
[0070] The data acquisition module is used to prepare several standard samples with known cashmere and wool content and to acquire the infrared spectral data of each standard sample.
[0071] The feature extraction module is used to normalize the infrared spectral data of each standard sample and extract the two spectral bands, amide I and amide II. The normalized infrared spectral data is smoothed and the second derivative is calculated. The second derivatives of the two spectral bands are concatenated to generate a feature vector.
[0072] The principal component analysis module is used to construct a data matrix based on the feature vectors of all standard samples, perform principal component analysis on the data matrix and generate a score matrix, extract each row vector from the score matrix as the principal component score vector of different standard samples, and use the principal component score vectors of pure wool samples and pure cashmere samples as the wool center vector and cashmere center vector, respectively.
[0073] The component prediction module is used to calculate the Mahalanobis distance from the principal component score vector of each standard sample to the wool center vector and the cashmere center vector, respectively, to construct a Mahalanobis distance weighted regression model, calculate the principal component score vector of the sample to be tested, and substitute it into the Mahalanobis distance weighted regression model to predict the cashmere content.
[0074] Compared with the prior art, the beneficial effects of the present invention are:
[0075] This invention establishes a calibration set covering the entire range by collecting samples with known cashmere content, ensuring that the spectral variation patterns across the entire concentration range can be learned. The second derivative of the normalized and smoothed spectra is calculated. The eigenvectors of amide I and II bands are extracted and spliced to effectively distinguish overlapping bands, highlighting subtle differences between wool and cashmere. Furthermore, characteristic bands of protein fibers are selectively chosen to eliminate interference from irrelevant bands. By calculating the principal component scores of a full range of gradient samples from 0% to 100%, the invention describes how the principal component scores change with cashmere content, thus achieving accurate cashmere content detection given the known spectral data of the sample to be tested.
[0076] This invention also calculates the Mahalanobis distance from each sample to the centers of the two classes; a weighted regression model is constructed based on the exponential weights of the minimum distance. The Mahalanobis distance takes into account the data covariance structure. When constructing the regression model, samples with typical spectral characteristics and clear classifications contribute more to the model parameters, while the influence of samples with ambiguous characteristics, potential noise, or outliers is weakened. This significantly enhances the model's ability to combat potential noise and uncertainty in the training data, improving the stability and accuracy of predictions. Attached Figure Description
[0077] Figure 1 This is a schematic diagram of the method flow of an embodiment of the present invention;
[0078] Figure 2 This is a schematic diagram of the device module in an embodiment of the present invention. Detailed Implementation
[0079] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0080] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0081] Example:
[0082] Please see Figure 1 The present invention provides a technical solution:
[0083] A method for identifying the composition content of cashmere and wool based on infrared spectroscopy technology, comprising the following steps:
[0084] Step 1: Prepare several standard samples with known cashmere and wool content, and collect infrared spectral data of each standard sample;
[0085] In this embodiment, the cashmere content in the standard samples ranges from 0% to 100%, and a gradient is set at every 10% interval, with one standard sample prepared for each gradient.
[0086] The cashmere content gradients are 0%, 10%, 20%, ..., 90%, 100%, with one standard sample prepared for each gradient, for a total of 11 standard samples, ensuring that the sum of the cashmere and wool content ratios in each standard sample is 1.
[0087] The specific principle for acquiring infrared spectral data of standard samples is as follows: a Fourier transform infrared spectrometer equipped with an ATR crystal is used, with a scanning wavenumber range of 4000~600. The resolution is set to 4. Several scanning wavenumber samples are generated within the scanning wavenumber range according to the resolution. For each scanning wavenumber sample, the ATR crystal is first scanned without any sample, and the background interferogram is acquired. Then, each standard sample is pressed tightly onto the ATR crystal, and its single-spectral interferogram is acquired. Fourier transform is performed on the background interferogram and the single-spectral interferogram of each standard sample to convert them into background single-beam spectra and sample single-beam spectra, respectively. The sample single-beam spectrum is divided by the background single-beam spectrum to obtain the transmittance spectrum of each standard sample. The negative logarithm of the transmittance spectrum is taken to obtain the absorbance spectrum corresponding to each scanning wavenumber sample. The absorbance spectra corresponding to all scanning wavenumber samples under the same standard sample are summarized to construct a spectral vector, which serves as the infrared spectral data of the standard sample.
[0088] Before acquiring the spectral data of the standard sample, a background scan is performed on the clean ATR crystal without any sample placed on it to record the spectral signal of the environment, i.e., the instrument itself, thus eliminating interference from these factors. Then, the prepared standard sample is tightly pressed onto the ATR crystal. The formula for calculating the absorbance spectrum is: The infrared light intensity signal after passing through the standard sample is obtained based on the single-beam spectrum of the sample, and the initial infrared light intensity signal is obtained based on the background single-beam spectrum. For each standard sample, the absorbance spectrum obtained at each scanning wavenumber sampling value is obtained and combined into a spectral vector, which is the infrared spectral data of the standard sample.
[0089] Step 2: Normalize the infrared spectral data of each standard sample and extract the amide I and amide II bands. Smooth the normalized infrared spectral data and calculate the second derivative. Concatenate the second derivatives of the two spectral bands to generate a feature vector.
[0090] In this embodiment, the principle for generating feature vectors is as follows:
[0091] For each standard sample, its spectral vector is: ,in, Represents the spectral vector of the standard sample. Indicates the first in the standard sample Absorbance spectrum of each scan wavenumber sample value Indicates the index of the scan wavenumber sample value, and , This represents the total number of wavenumber samples taken for a single standard sample.
[0092] The formula for normalizing the spectral vector is:
[0093]
[0094]
[0095] in, Represents the normalized spectral vector. express The modulus length;
[0096] The purpose of normalizing spectral vectors is to eliminate spectral intensity variations caused by factors such as sample thickness, density, and measurement conditions. By dividing each spectral vector by its modulus, all spectra are unified to the same scale, making spectral data from different samples comparable.
[0097] Extracting spectral frequency bands from the normalized spectral vector is specifically as follows:
[0098] Amide I band: wavenumber range 1700~1600 ;
[0099] Amide II band: wavenumber range 1580~1480 ;
[0100] In the complete infrared spectrum, the spectra of cashmere and wool are very similar overall, with a large number of absorption peaks overlapping. Directly using full-spectrum modeling would introduce a lot of noise and irrelevant information. The chemical basis of cashmere and wool is keratin, but they differ in the proportion of secondary structures. The two characteristic spectral bands, amide I and amide II, contain the C=O stretching vibration, NH bending vibration, and CN stretching vibration of the protein backbone, which are sensitive to changes in the secondary structure of proteins. Selecting these two characteristic spectral bands can effectively distinguish the overlapping bands of cashmere and wool.
[0101] The principle of smoothing is as follows: For each absorbance spectrum in the extracted spectral band, an equal-length window is divided with that window as the center. The average absorbance spectrum within the equal-length window is calculated and replaced with that absorbance spectrum. The second derivative of the replaced absorbance spectrum is then calculated, and the second derivatives of all equal-length windows are concatenated to form a feature vector. ,in, Represents the eigenvector. This indicates the first frequency band of the extracted spectrum. The second derivative of the absorbance spectrum corresponding to each scan wavenumber sample value This represents the index of the scan wavenumber sample value in the extracted spectral frequency band, and , This represents the total number of scan wavenumber samples in the extracted spectral band.
[0102] The infrared spectra of cashmere and wool are very similar, with significant overlap in absorption peaks. The second derivative can be used to highlight subtle differences within these overlapping peaks. Infrared spectra are often affected by baseline drift, impacting the accuracy of quantitative analysis. The second derivative is insensitive to linear or slowly changing baselines, effectively eliminating such systematic errors; the eigenvector formed by splicing these peaks... Each element represents the second derivative value at a specific wavenumber, reflecting the concavity and convexity of the spectral curve at that location. By extracting the second derivative, overlapping absorption peaks are effectively separated, making the differences between cashmere and wool in the amide I and II bands more obvious. By splicing the two characteristic bands, a comprehensive feature vector is formed, which centrally expresses the spectral information most relevant to the component content. The original spectral data has high dimensionality and contains a lot of information unrelated to component identification. By truncating key bands and extracting the second derivative, the data dimensionality is significantly reduced, while baseline drift and noise interference are suppressed.
[0103] Step 3: Construct a data matrix based on the feature vectors of all standard samples, perform principal component analysis on the data matrix and generate a score matrix, extract each row vector from the score matrix as the principal component score vector of different standard samples, and use the principal component score vectors of pure wool samples and pure cashmere samples as the wool center vector and cashmere center vector, respectively.
[0104] In this embodiment, the principle for generating the score matrix is as follows:
[0105] The constructed data matrix is as follows:
[0106]
[0107] in, Represents a data matrix. Indicates the first The feature vector of a standard sample Indicates the index of the standard sample, and , Indicates the quantity of standard samples. Represents the transpose matrix;
[0108] Each standard sample generates a feature vector. The feature vectors of all L standard samples are stacked row-wise to form a... The matrix, Each row in the data matrix represents all the spectral features of a standard sample, and each column represents the second derivative value of the absorbance spectrum of all spectral samples at a specific wavenumber. The data matrix reflects the spectral features of all standard samples in the intercepted spectral region and reflects the spectral differences between standard samples, which are related to the proportion of cashmere and wool content.
[0109] Calculate the average of each column in the data matrix and generate a centered data matrix. The specific formula is as follows:
[0110]
[0111]
[0112] in, Indicates the first The average of the second derivatives of the absorbance spectrum corresponding to each scan wavenumber sample value. Indicates the first The standard sample at the first The second derivative of the absorbance spectrum corresponding to each scan wavenumber sample value Represents a centralized data matrix. express A column vector of all 1s Let represent the vector of the second derivative of the absorbance spectrum, and Represents the th element in the centered matrix The row vectors corresponding to each standard sample;
[0113] When calculating the centered data matrix, the average value of each column is first calculated, which is the second derivative of the average spectral response intensity of all standard samples at each specific wavenumber. This finds the center point of all samples at a given wavenumber. Then, the mean of its column is subtracted from each element in the data matrix, eliminating common baselines across wavenumbers and focusing on the relative differences between cashmere and wool at specific wavenumbers. After centering, the subsequent principal component scores clearly represent the positive and negative deviations of the samples in each principal component direction. This reflects the deviation or difference of the spectral characteristics of each sample relative to the average level of the entire sample set. Each element in If the value is positive, it means that the first... The standard sample at the first The spectral response of a scan wavenumber sample is higher than the average level of all standard samples. If the value is negative, it indicates that it is lower than the average level. If it is 0, it indicates that it is consistent with the average level.
[0114] The covariance matrix is calculated based on a centralized data matrix, using the following formula:
[0115]
[0116] in, Represents the covariance matrix. express The transpose of the matrix;
[0117] The mean of each column is already 0, and the covariance matrix... It can be calculated directly using matrix multiplication. It is The matrix, It is The matrix, when multiplied by the two, yields C, which is a... The matrix, covariance matrix The diagonal elements represent the expected covariance of the scan wavenumber sample value, which measures the degree of fluctuation in the spectral data corresponding to that scan wavenumber sample value. The larger the value, the greater the difference between different standard samples at this wavenumber point. The off-diagonal elements measure the linear correlation between the spectral data corresponding to two scan wavenumber sample values. A positive value indicates that when the second derivative of the absorbance spectrum of one scan wavenumber sample value is high, the other tends to be high as well. A negative value indicates that when the second derivative of the absorbance spectrum of one scan wavenumber sample value is high, the other tends to be low. A value of 0 indicates that the changes of the two scan wavenumber sample values are independent of each other.
[0118] The eigenvalue decomposition of the covariance matrix is based on the following formula:
[0119]
[0120] in, Indicates the first The eigenvectors of the principal components Indicates the first The eigenvalues of the principal components Indicates the index of the principal component, and ,
[0121] The variance contribution rate of each principal component is calculated, and then the cumulative variance contribution rate of the principal components is calculated using the following formula:
[0122]
[0123]
[0124] in, Indicates the first The variance contribution rate of each principal component Indicates the preceding The cumulative variance contribution rate of each principal component Indicates the index of the principal component, and ;
[0125] Feature vector The orientation of the principal components and eigenvalues are defined. This reflects the variance of the data projected along the direction of that principal component. This represents the sum of the variances of the projections of the data onto all principal components, where the data is on the [missing information]. The projection variance over each principal component is: , No. The contribution rate of each principal component is... The proportion of all principal components, i.e. .
[0126] from Start, gradually increase The value of is calculated, and the corresponding value is determined. until the smallest one is found. , making Before choosing The eigenvectors constitute the loading matrix: ;
[0127] The purpose of calculating the cumulative variance contribution rate is to determine how many principal components need to be retained in principal component analysis in order to retain most of the information in the original data while minimizing data dimensionality; the cumulative variance contribution rate represents the percentage of principal components retained in the original data. The sum of the variance contribution rates of each principal component, by setting a threshold for the cumulative variance contribution rate, determines the minimum number of principal components needed to represent the original data. Cashmere and wool have very similar infrared spectra with minimal differences; therefore, it is necessary to retain as many principal components as possible to capture these subtle differences. This threshold is set to... While reducing dimensionality, it retains useful spectral information to the maximum extent.
[0128] The loading matrix is constructed based on the eigenvectors of the first k principal components. Each column represents a principal component, and each row represents the second derivative of the absorbance spectrum corresponding to a specific scan wavenumber sampling point. The loading matrix is used to define the projection direction, where each eigenvector defines a new direction. Projecting the centered data matrix onto these directions yields a low-dimensional principal component score vector. The loading matrix reflects which wavenumber points contribute the most to the composition of the principal components.
[0129] The formula for calculating the score matrix is:
[0130]
[0131] in, This represents the score matrix.
[0132] The score matrix reflects the position and structure of the original high-dimensional data in the reduced principal component space. Each row of the score matrix corresponds to a standard sample, which is the principal component score vector of that standard sample. This matrix reflects the distribution of different standard samples along a few principal component directions, thus preserving most of the information in the original spectral data while eliminating multicollinearity. The score matrix is a... A matrix can be represented as ,in, Indicates the first The standard sample at the first The larger the absolute value of the projection score along the principal component direction, the stronger the spectral characteristics of the standard sample in the 1st principal component direction. The more significant the performance in each principal component direction, the higher the row vector of the score matrix. Each row vector corresponds to a principal component score vector of a standard sample, and from arrive The cashmere content of the corresponding standard samples gradually increased.
[0133] Extract each row vector from the score matrix , serving as the principal component score vector for each standard sample, where, Represents the score matrix of the first... There are n row vectors, and the center vector of the wool class is n. The center vector of the cashmere class is .
[0134] Step 4: Calculate the Mahalanobis distance from the principal component score vector of each standard sample to the wool center vector and the cashmere center vector, respectively, construct a Mahalanobis distance weighted regression model, calculate the principal component score vector of the sample to be tested, and substitute it into the Mahalanobis distance weighted regression model to predict the cashmere content.
[0135] In this embodiment, the principle of constructing the Mahalanobis distance regression model is as follows:
[0136] The formula used to calculate the Mahalanobis distance from the standard sample to the center vector of the wool is as follows:
[0137]
[0138] in, Indicates the first Mahalanobis distance from the standard sample to the center vector of the wool class Let represent the covariance matrix of the principal component scores, and For the front A diagonal matrix of principal component eigenvalues;
[0139] The formula used to calculate the Mahalanobis distance from the standard sample to the center vector of the cashmere sample is:
[0140]
[0141] in, Indicates the first Mahalanobis distance from a standard sample to the center vector of the cashmere category;
[0142] Mahalanobis distance reflects the similarity between the spectral characteristics of a standard sample and those of "pure wool" or "pure cashmere." This similarity is derived after considering the variance and correlation of all principal component directions. A small Mahalanobis distance indicates that the sample's spectral characteristics are very similar to the class center, and it is very likely a pure wool or pure cashmere sample. A large Mahalanobis distance indicates that the sample's spectral characteristics differ significantly from the class center, and the sample may be located on the boundary between the two classes. Mahalanobis distance considers the distribution structure of the data, not just the absolute difference in Euclidean distance. In principal component space, the importance of each principal component direction is determined by its variance, i.e., the eigenvalue. The larger the eigenvalue, the greater the sample difference along that principal component direction. Using the diagonal matrix formed by the eigenvalues as the covariance matrix is equivalent to standardizing each principal component direction, making the distance metric comparable across different directions. The covariance matrix is... .
[0143] The formula used to calculate the weights of standard samples in the regression fitting is:
[0144]
[0145] in, Indicates the first The weight of each standard sample, Indicates the attenuation coefficient. express and The minimum value between;
[0146] Pick The aim is to determine which class center the standard sample is closer to, thus determining whether the standard sample is more like a cashmere or wool sample. An exponential function is used to map distance to weights: the smaller the distance, the larger the weight, indicating the sample is very close to a particular class center; the larger the distance, the smaller the weight, indicating the sample's spectral data is closer to the center of either class, and the more blurred the sample is. The attenuation coefficient is first determined. If the distance between samples varies greatly, then reduce the adjustment. To slow down the decay, if you want to screen samples more rigorously, increase the [value]. .
[0147] A dependent variable vector is constructed based on the known cashmere content of all standard samples, specifically: ,in, Represents the dependent variable vector. Indicates the first The cashmere content of each standard sample; based on the scores of all standard samples on the first k principal components, construct... The design matrix is as follows:
[0148]
[0149] in, Represents the design matrix. Indicates the first The sample in the first Scores on each principal component;
[0150] The 1 in matrix B is reserved for the constant term of the regression model, corresponding to... Each row in the design matrix reflects the coordinates of the standard sample in the reduced principal component space.
[0151] The weight matrix is constructed based on the weights of the standard samples, specifically as follows:
[0152]
[0153] in, Represents the weight matrix;
[0154] Find a regression coefficient vector This makes the weighted sum of squared residuals of all standard samples... Minimum; fitted based on least squares method The Mahalanobis distance regression model is as follows:
[0155]
[0156] in, Represents the regression coefficient vector. Represents the design matrix of the first... Row vectors.
[0157] For a physical mixture like wool and cashmere, its infrared spectrum can ideally be considered as a linear weighted sum of the wool and cashmere spectra. There is also an approximately linear relationship between the principal component scores and content obtained from the spectral data. Given the principal component score vectors of several standard samples and the design matrix they construct, as well as the cashmere content of these standard samples, the regression coefficient vector is fitted using the least squares method. Under the premise of knowing the principal component score vectors, the corresponding cashmere content is determined by the Mahalanobis distance regression model.
[0158] The principle for predicting cashmere content is as follows:
[0159] For the sample to be tested, feature vectors are generated using the same method as for the standard sample. ,in, This indicates the frequency of the sample under test within the extracted spectral band. The second derivative of the absorbance spectrum corresponding to each scan wavenumber sample value, After centering, the principal component score vector of the sample to be tested is calculated. The specific formula is as follows: Then construct the prediction vector. ,Will Substitute the values into the Mahalanobis distance regression model to generate predicted cashmere content values. .
[0160] Please see Figure 2 The present invention also provides a cashmere and wool composition content identification device based on infrared spectroscopy technology. The device is used to implement the above-mentioned cashmere and wool composition content identification method based on infrared spectroscopy technology, specifically including:
[0161] The data acquisition module is used to prepare several standard samples with known cashmere and wool content and to acquire the infrared spectral data of each standard sample.
[0162] The feature extraction module is used to normalize the infrared spectral data of each standard sample and extract the two spectral bands, amide I and amide II. The normalized infrared spectral data is smoothed and the second derivative is calculated. The second derivatives of the two spectral bands are concatenated to generate a feature vector.
[0163] The principal component analysis module is used to construct a data matrix based on the feature vectors of all standard samples, perform principal component analysis on the data matrix and generate a score matrix, extract each row vector from the score matrix as the principal component score vector of different standard samples, and use the principal component score vectors of pure wool samples and pure cashmere samples as the wool center vector and cashmere center vector, respectively.
[0164] The component prediction module is used to calculate the Mahalanobis distance from the principal component score vector of each standard sample to the wool center vector and the cashmere center vector, respectively, to construct a Mahalanobis distance weighted regression model, calculate the principal component score vector of the sample to be tested, and substitute it into the Mahalanobis distance weighted regression model to predict the cashmere content.
[0165] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0166] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.
[0167] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0168] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for identifying the content of cashmere and wool components based on infrared spectroscopy, characterized in that, The specific steps include: Step 1: Prepare several standard samples with known cashmere and wool content, and collect infrared spectral data of each standard sample; Step 2: Normalize the infrared spectral data of each standard sample and extract the amide I and amide II bands. Smooth the normalized infrared spectral data and calculate the second derivative. Concatenate the second derivatives of the two spectral bands to generate a feature vector. Step 3: Construct a data matrix based on the feature vectors of all standard samples, perform principal component analysis on the data matrix and generate a score matrix, extract each row vector from the score matrix as the principal component score vector of different standard samples, and use the principal component score vectors of pure wool samples and pure cashmere samples as the wool center vector and cashmere center vector, respectively. Step 4: Calculate the Mahalanobis distance from the principal component score vector of each standard sample to the wool center vector and the cashmere center vector, respectively, construct a Mahalanobis distance weighted regression model, calculate the principal component score vector of the sample to be tested, and substitute it into the Mahalanobis distance weighted regression model to predict the cashmere content.
2. The method for identifying the cashmere and wool composition based on infrared spectroscopy technology according to claim 1, characterized in that: In step 1, the cashmere content of the standard samples ranges from 0% to 100%, and a gradient is set at 10% intervals, with one standard sample prepared for each gradient. The specific principle for acquiring infrared spectral data of standard samples is as follows: a Fourier transform infrared spectrometer equipped with an ATR crystal is used, with a scanning wavenumber range of 4000~600. The resolution is set to 4. Several scanning wavenumber samples are generated within the scanning wavenumber range according to the resolution. For each scanning wavenumber sample, the ATR crystal is first scanned without any sample, and the background interferogram is acquired. Then, each standard sample is pressed tightly onto the ATR crystal, and its single-spectral interferogram is acquired. Fourier transform is performed on the background interferogram and the single-spectral interferogram of each standard sample to convert them into background single-beam spectra and sample single-beam spectra, respectively. The sample single-beam spectrum is divided by the background single-beam spectrum to obtain the transmittance spectrum of each standard sample. The negative logarithm of the transmittance spectrum is taken to obtain the absorbance spectrum corresponding to each scanning wavenumber sample. The absorbance spectra corresponding to all scanning wavenumber samples under the same standard sample are summarized to construct a spectral vector, which serves as the infrared spectral data of the standard sample.
3. The method for identifying cashmere and wool composition based on infrared spectroscopy technology according to claim 1, characterized in that: The principle behind generating the feature vector in step 2 is as follows: For each standard sample, its spectral vector is: ,in, Represents the spectral vector of the standard sample. Indicates the first in the standard sample Absorbance spectrum of each scan wavenumber sample value Indicates the index of the scan wavenumber sample value, and , This represents the total number of wavenumber samples taken for a single standard sample. The formula for normalizing the spectral vector is: in, Represents the normalized spectral vector. express The modulus length; Extracting spectral frequency bands from the normalized spectral vector is specifically as follows: Amide I band: wavenumber range 1700~1600 ; Amide II band: wavenumber range 1580~1480 ; The principle of smoothing is as follows: For each absorbance spectrum in the extracted spectral band, an equal-length window is divided with that window as the center. The average absorbance spectrum within the equal-length window is calculated and replaced with that absorbance spectrum. The second derivative of the replaced absorbance spectrum is then calculated, and the second derivatives of all equal-length windows are concatenated to form a feature vector. ,in, Represents the eigenvector. This indicates the first frequency band of the extracted spectrum. The second derivative of the absorbance spectrum corresponding to each scan wavenumber sample value This represents the index of the scan wavenumber sample value in the extracted spectral frequency band, and , This represents the total number of scan wavenumber samples in the extracted spectral band.
4. The method for identifying cashmere and wool composition based on infrared spectroscopy technology according to claim 3, characterized in that: The principle behind generating the score matrix in step 3 is as follows: The constructed data matrix is as follows: in, Represents a data matrix. Indicates the first The feature vector of a standard sample Indicates the index of the standard sample, and , Indicates the quantity of standard samples. Represents the transpose matrix; Calculate the average of each column in the data matrix and generate a centered data matrix. The specific formula is as follows: in, Indicates the first The average of the second derivatives of the absorbance spectrum corresponding to each scan wavenumber sample value. Indicates the first The standard sample at the first The second derivative of the absorbance spectrum corresponding to each scan wavenumber sample value Represents a centralized data matrix. express A column vector of all 1s Let represent the vector of the second derivative of the absorbance spectrum, and Represents the th element in the centered matrix The row vectors corresponding to each standard sample; The covariance matrix is calculated based on the centralized data matrix, and the specific formula is as follows: in, Represents the covariance matrix. express The transpose of the matrix; The eigenvalue decomposition of the covariance matrix is based on the following formula: in, Indicates the first The eigenvectors of the principal components Indicates the first The eigenvalues of the principal components Indicates the index of the principal component, and , The variance contribution rate of each principal component is calculated, and then the cumulative variance contribution rate of the principal components is calculated using the following formula: in, Indicates the first The variance contribution rate of each principal component Indicates the preceding The cumulative variance contribution rate of each principal component Indicates the index of the principal component, and ; from Start, gradually increase The value of is calculated, and the corresponding value is determined. until the smallest one is found. , making Before choosing The eigenvectors constitute the loading matrix: ; The formula for calculating the score matrix is: in, This represents the score matrix.
5. The method for identifying cashmere and wool composition based on infrared spectroscopy technology according to claim 4, characterized in that: Extract each row vector from the score matrix , serving as the principal component score vector for each standard sample, where, Represents the score matrix of the first... There are n row vectors, and the center vector of the wool class is n. The center vector of the cashmere class is .
6. The method for identifying the cashmere and wool composition based on infrared spectroscopy technology according to claim 5, characterized in that: The principle of constructing the Mahalanobis distance regression model in step 4 is as follows: The formula used to calculate the Mahalanobis distance from the standard sample to the center vector of the wool is as follows: in, Indicates the first Mahalanobis distance from the standard sample to the center vector of the wool class Let represent the covariance matrix of the principal component scores, and For the front A diagonal matrix of principal component eigenvalues; The formula used to calculate the Mahalanobis distance from the standard sample to the center vector of the cashmere sample is: in, Indicates the first Mahalanobis distance from a standard sample to the center vector of the cashmere category; The formula used to calculate the weights of standard samples in the regression fitting is: in, Indicates the first The weight of each standard sample, Indicates the attenuation coefficient. express and The minimum value between; A dependent variable vector is constructed based on the known cashmere content of all standard samples, specifically: ,in, Represents the dependent variable vector. Indicates the first The cashmere content of each standard sample was determined; based on the scores of all standard samples on the first k principal components, a [structure / construction] was built. The design matrix is as follows: in, Represents the design matrix. Indicates the first The sample in the first Scores on each principal component; The weight matrix is constructed based on the weights of the standard samples, specifically as follows: in, Represents the weight matrix; Find a regression coefficient vector This makes the weighted sum of squared residuals of all standard samples... Minimum; fitted based on least squares method The Mahalanobis distance regression model is as follows: in, Represents the regression coefficient vector. Represents the design matrix of the first... Row vectors.
7. The method for identifying cashmere and wool composition based on infrared spectroscopy technology according to claim 6, characterized in that: The principle for predicting cashmere content is as follows: For the sample to be tested, feature vectors are generated using the same method as for the standard sample. ,in, This indicates the frequency of the sample under test within the extracted spectral band. The second derivative of the absorbance spectrum corresponding to each scan wavenumber sample value, After centering, the principal component score vector of the sample to be tested is calculated. The specific formula is as follows: Then construct the prediction vector. ,Will Substitute the values into the Mahalanobis distance regression model to generate predicted cashmere content values. .
8. A cashmere and wool composition identification device based on infrared spectroscopy technology, characterized in that: The device is used to implement the cashmere and wool composition content identification method based on infrared spectroscopy technology as described in any one of claims 1-7, specifically including: The data acquisition module is used to prepare several standard samples with known cashmere and wool content and to acquire the infrared spectral data of each standard sample. The feature extraction module is used to normalize the infrared spectral data of each standard sample and extract the two spectral bands, amide I and amide II. The normalized infrared spectral data is smoothed and the second derivative is calculated. The second derivatives of the two spectral bands are concatenated to generate a feature vector. The principal component analysis module is used to construct a data matrix based on the feature vectors of all standard samples, perform principal component analysis on the data matrix and generate a score matrix, extract each row vector from the score matrix as the principal component score vector of different standard samples, and use the principal component score vectors of pure wool samples and pure cashmere samples as the wool center vector and cashmere center vector, respectively. The component prediction module is used to calculate the Mahalanobis distance from the principal component score vector of each standard sample to the wool center vector and the cashmere center vector, respectively, to construct a Mahalanobis distance weighted regression model, calculate the principal component score vector of the sample to be tested, and substitute it into the Mahalanobis distance weighted regression model to predict the cashmere content.
Citation Information
Patent Citations
Cashmere and wool component content identification method based on infrared spectrum technology
CN117092059A
Foundation soil slurry quality analysis method, system and equipment and storage medium
CN119691679A
Tool performance evaluation method and device
CN121119441A