Near-infrared quantitative model construction method and device, and storage medium

By optimizing the preprocessing and modeling algorithms for near-infrared spectral data in the global search space using the particle swarm optimization algorithm, the problems of noise interference and reliance on human experience are solved, and efficient and robust near-infrared quantitative model construction is achieved.

CN116434870BActive Publication Date: 2026-03-31CHINA TOBACCO HUNAN IND CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In the current near-infrared spectral data modeling and analysis process, the noise interference is severe due to the influence of measurement conditions and sample state, resulting in poor preprocessing and modeling effects. Furthermore, the optimization of algorithm parameters and sequence is complex, relies on human experience, and is inefficient.

Method used

The particle swarm optimization algorithm is used to optimize the parameters and order of the preprocessing algorithm, variable selection algorithm, and modeling algorithm for near-infrared spectral data in the global optimization search space. The optimal algorithm combination is quickly determined through global and local searches in the particle swarm search space.

Benefits of technology

It enables the efficient construction of high-quality near-infrared quantitative models without the need for human experience, reduces the requirements for analysts, significantly improves the predictive power and robustness of the models, and has broad applicability and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116434870B_ABST
    Figure CN116434870B_ABST
Patent Text Reader

Abstract

The application discloses a near-infrared quantitative model construction method and device and a storage medium, wherein the method comprises the following steps: acquiring near-infrared spectrum data, and constructing a near-infrared quantitative model training sample set; acquiring a plurality of near-infrared spectrum data preprocessing algorithms, variable selection algorithms and modeling algorithms and the value range of parameters thereof, and constructing a particle swarm search space composed of the parameters and the near-infrared spectrum data preprocessing algorithm sequence; based on the near-infrared quantitative model training sample set, the parameters of each algorithm and the preprocessing algorithm sequence are optimized in the particle swarm search space by using a particle swarm optimization algorithm; and the global optimal solution obtained by selection is used as the parameters of each algorithm and the preprocessing algorithm sequence, near-infrared spectrum data is preprocessed, variable selection and modeling are performed, and a near-infrared quantitative model is obtained. The modeling requirement is reduced, and a high-quality quantitative model can be simply and efficiently obtained on the basis of the data modeling experience of researchers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of chemometrics, and in particular to a method, apparatus, and storage medium for constructing a near-infrared quantitative model based on particle swarm optimization. Background Technology

[0002] Near-infrared spectroscopy (NIRS) offers advantages such as fast analysis speed and non-destructive testing, leading to its widespread application in agriculture, food, petrochemicals, tobacco, and many other industries. However, external factors such as measurement conditions and sample state inevitably introduce signals and noise unrelated to the measured target into the measured signal, affecting the modeling effectiveness of NIRS data. Therefore, various methods such as data augmentation, derivative analysis, standard normal transformation, and multivariate scattering correction are needed to process NIRS data and eliminate the noise. Because NIRS data modeling and analysis involves a lengthy process, numerous algorithms and parameters, and complex interrelationships between related algorithm parameters and the order of preprocessing algorithms, the preprocessing, variable selection, and modeling of NIRS data have become pressing issues in the field of NIRS research.

[0003] Currently, the methods for determining near-infrared spectral preprocessing algorithms and their parameters are still mainly based on manual experience, supplemented by local optimization methods such as grid search. The former determines the appropriate spectral preprocessing analysis scheme based on the researcher's past experience, which is heavily dependent on the researcher's preprocessing experience for a certain type of near-infrared spectral analysis; the latter, taking grid search algorithms as an example, often requires experimental analysis of a large number of grid nodes, resulting in low overall performance.

[0004] Based on this, this invention studies an intelligent construction and optimization scheme for near-infrared quantitative models based on the particle swarm optimization method. This scheme, based on the global optimization search concept of the particle swarm optimization method, performs a rapid search in the global space constituted by a series of preprocessing algorithms, variable selection algorithms, modeling algorithm parameters, and the order of the preprocessing algorithms in the near-infrared quantitative model construction. This ultimately determines the relevant algorithm parameters that match the characteristics of the near-infrared data, and optimizes the order of the preprocessing algorithms to improve the predictive ability and robustness of the near-infrared model. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a method, apparatus, and storage medium for constructing near-infrared quantitative models, thereby solving the problems of low overall performance and low efficiency in the construction of existing near-infrared quantitative models.

[0006] Firstly, a method for constructing a near-infrared quantitative model is provided, including:

[0007] Acquire near-infrared spectral data and construct a training sample set for a near-infrared quantitative model;

[0008] A particle swarm search space is constructed, and the value ranges of various near-infrared spectral data preprocessing algorithms, variable selection algorithms, and modeling algorithms and their parameters are obtained. A particle swarm search space is constructed consisting of these parameters and various infrared spectral data preprocessing algorithms in sequence.

[0009] Based on the training sample set of the near-infrared quantitative model, the parameters of various near-infrared spectral data preprocessing algorithms, variable selection algorithms, and modeling algorithms, as well as the order of various infrared spectral data preprocessing algorithms, are optimized in the particle swarm optimization algorithm in the particle swarm search space; where each particle is a vector composed of the parameters of each algorithm and the order of various infrared spectral data preprocessing algorithms.

[0010] Based on the particle swarm iteration output, the obtained global optimal solution is selected as the parameter of each algorithm and the order of various infrared spectral data preprocessing algorithms. The near-infrared spectral data is preprocessed, variables are selected and modeled to obtain a near-infrared quantitative model.

[0011] Furthermore, the particle swarm search space construction process is as follows:

[0012] Construct a set of near-infrared spectral data preprocessing algorithms, variable selection algorithms, and modeling algorithms, A = {A1, A2, ..., A...} m},in, For m-2 near-infrared spectral data preprocessing algorithms, A m-1 For the variable selection algorithm, A m For modeling algorithms;

[0013] For algorithm A i The parameter of (i∈{1,2,…,m}) is denoted as P. i ={P i1 ,P i2 ,…,P iki}, where parameter P ij (i∈{1,2,…,m},j∈{1,2,…,k i The range of values ​​for}) is denoted as Then, a search space S of n = n' + m - 2 dimensions is constructed sequentially from the parameters of all algorithms and the near-infrared spectral data preprocessing algorithms. n ,in: This represents the total number of parameters for the near-infrared spectral data preprocessing algorithm, variable selection algorithm, and modeling algorithm; m-2 represents the order of the m-2 preprocessing algorithms.

[0014] Search Space S n For each position in the vector, an n-dimensional data vector P = (p1, p2, ..., p...) is used. n' ,p n'+1 ,p n'+2 ,…,p n'+m-2) represents; where, in the first n' dimensions, the value range of each dimension is the value range of the corresponding algorithm parameter; in the last m-2 dimensional sub-vector, the value range of each dimension is {1,2,…,m-2}, and the value of each component in the last m-2 dimensional sub-vector is distinct from the values ​​of other components, used to represent the execution sequence number of the corresponding near-infrared spectral data preprocessing algorithm in a series of preprocessing processes;

[0015] Assume the search space S n Let O be the region enclosed by the outer envelope, and let Φ be the non-search region within space O. Then, we have the particle swarm search space S. n for:

[0016] S n =O-Φ.

[0017] Furthermore, the training sample set based on the near-infrared quantitative model utilizes a particle swarm optimization algorithm to optimize the parameters and order of various near-infrared spectral data preprocessing algorithms, variable selection algorithms, and modeling algorithms within the particle swarm search space. Specifically, this includes:

[0018] A1: Set the particle velocity iteration strategy, and use the linear decreasing weight method to set the particle flight velocity iteration strategy;

[0019] A2: Set a particle position iteration strategy, calculate the particle position after iteration based on the particle flight velocity, and verify the particle position according to the parameter value range of various near-infrared spectral data preprocessing algorithms, variable selection algorithms and modeling algorithms, as well as the sequential value range of various infrared spectral data preprocessing algorithms, and determine the final particle position after iteration.

[0020] A3: Particle swarm initialization, initializing the number of particles, the initial position and initial velocity of each particle, and the particle iteration exit mechanism; where the particle position is represented by a vector composed of the parameters of various near-infrared spectral data preprocessing algorithms, variable selection algorithms, and modeling algorithms, as well as the sequence of various infrared spectral data preprocessing algorithms;

[0021] A4: Particle swarm position verification and adjustment. Determine whether the positions of all particles are feasible parameters for near-infrared preprocessing algorithms, variable selection algorithms, and modeling algorithms, and whether the order of infrared spectral data preprocessing algorithms is feasible. If so, calculate the distance between the particles and the optimal solution. Otherwise, adjust and optimize the positions of the particles, and then calculate the distance between their positions and the optimal solution.

[0022] A5: Particle iteration. If the particle iteration exit mechanism is satisfied, the historical global optimal position is output as the optimal solution of the particle swarm optimization algorithm; otherwise, the next flight position of the particle is calculated according to the particle flight speed iteration strategy and the particle position iteration strategy, and the process returns to step A4.

[0023] Furthermore, in step A1, the iterative formula for particle flight velocity is as follows:

[0024]

[0025] In the formula, Let c1 and c2 represent the flight velocity of the i-th particle in the t-th iteration, and let rand() represent a random number between (0,1). gbest represents the optimal position in the search history of the i-th particle. (t) This represents the optimal position in the entire particle search history. ω represents the position of the i-th particle in the t-th iteration; (t) Let represent the particle inertial weight at the t-th iteration, and its calculation formula is as follows:

[0026]

[0027] In the formula, G k ω represents the maximum number of iterations. ini Let ω be the initial inertia weight. end The inertia weight is the value at the maximum number of iterations.

[0028] Furthermore, in step A2, the position of the particle in the next iteration is predicted based on the particle's flight velocity. x' i =[x' i1 ,x' i2 ,…,x' in Let x' be an n-dimensional vector. i Not in the particle swarm search space S n Inside, it represents x' i There are k algorithm parameters whose values ​​are not within the range of their respective algorithm parameters, denoted as P'={P' τ1 ,P' τ2 …,P' τk}, from x' i Extract the corresponding k-dimensional data and search the n-dimensional particle swarm space S. n Extract the corresponding k-dimensional sub-region And find the k-dimensional subregion using Euclidean distance. Surface and x' i The closest point Replace x' with the value in y' i The values ​​of the corresponding dimensions are used to obtain a vector. As the next search location for the particle;

[0029] The final determined particle position update strategy is shown in the following equation:

[0030]

[0031] Furthermore, in step A3, the particle iteration exit mechanism includes reaching the maximum number of iterations and model R. 2 There are two scenarios: the indicator value reaches the set value.

[0032] Furthermore, in which model R 2 The particle iteration exit mechanism when the indicator value reaches the set value specifically includes:

[0033] During particle swarm optimization, after each iteration, the algorithm parameters represented by the particles and various infrared spectral data preprocessing algorithms are sequentially substituted into the corresponding near-infrared spectral data preprocessing algorithm, variable selection algorithm, and modeling algorithm. Modeling is then performed based on the near-infrared quantitative model training sample set, and the R-squared value of the model is utilized. 2 The index value is used to calculate the distance Dist between the particle and the optimal solution. If the distance Dist is less than a preset value, the iteration is terminated. The formula for calculating Dist is as follows:

[0034] Dist = 1 - R 2

[0035] Among them, R 2 The formula for calculating the index value is shown below:

[0036]

[0037] Among them, y i This represents the true value of the i-th training sample. This represents the predicted value of the i-th training sample. q represents the average of the true values ​​of all training samples; q represents the number of experimental samples.

[0038] Secondly, a near-infrared quantitative model construction device is provided, comprising:

[0039] The data acquisition module is used to acquire near-infrared spectral data and construct a training sample set for the near-infrared quantitative model.

[0040] The particle swarm search space construction module is used to obtain the value ranges of various near-infrared spectral data preprocessing algorithms, variable selection algorithms, and modeling algorithms and their parameters, and construct a particle swarm search space composed of these parameters and the sequence of various infrared spectral data preprocessing algorithms.

[0041] The parameter optimization module is used to optimize the parameters of various near-infrared spectral data preprocessing algorithms, variable selection algorithms, and modeling algorithms, as well as the order of various infrared spectral data preprocessing algorithms, based on the training sample set of the near-infrared quantitative model. The particle swarm optimization algorithm is used to optimize the parameters of various near-infrared spectral data preprocessing algorithms, variable selection algorithms, and modeling algorithms in the particle swarm search space. Each particle is a vector composed of the parameters of each algorithm and the order of various infrared spectral data preprocessing algorithms.

[0042] The model building module is used to select the globally optimal solution as the parameters of each algorithm and the order of various infrared spectral data preprocessing algorithms based on the particle swarm iteration output results, to preprocess, select variables and model near-infrared spectral data, and obtain a near-infrared quantitative model.

[0043] Thirdly, an electronic device is provided, comprising:

[0044] A memory that stores computer programs;

[0045] When the processor loads and executes the computer program stored in the memory, it implements the near-infrared quantitative model construction method as described above.

[0046] Fourthly, a computer-readable storage medium is provided that stores a computer program, which, when executed by a processor, implements the near-infrared quantitative model construction method as described above.

[0047] Beneficial effects

[0048] This invention proposes a near-infrared quantitative model construction method, device, and storage medium, which has the following advantages compared with the prior art:

[0049] (1) Unaffected by sample characteristics and their corresponding near-infrared spectral data characteristics, the data obtained from the analysis of samples from different fields, types and batches do not require the reference of researchers' modeling and analysis experience. High-quality near-infrared quantitative models can be obtained simply and efficiently, greatly reducing the requirements for analysts and meeting the requirements of near-infrared technology for data modeling and analysis.

[0050] (2) By using a fast global search in the early stage and a strengthened local search in the later stage, a high-dimensional space search scheme is constructed to optimize the parameters of the preprocessing algorithm, variable selection algorithm and modeling algorithm and to optimize the order of the preprocessing algorithm. Its performance and model results are significantly better than traditional analysis and other parameter optimization methods.

[0051] (3) The proposed solution can further optimize the algorithm parameters and order involved in the modeling process by constructing a more comprehensive and broader near-infrared spectral optimization space, and has strong scalability and portability. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 This is a flowchart of a near-infrared quantitative model construction method provided by an embodiment of the present invention;

[0054] Figure 2 These are the raw near-infrared spectral data provided in the embodiments of the present invention;

[0055] Figure 3 This is a graph showing the experimental results of the number of particle swarm iterations provided in an embodiment of the present invention;

[0056] Figure 4 The R-squared model of the near-infrared spectral data provided in this embodiment of the invention, after particle swarm optimization, is compared with the modeling model using traditional methods. 2 Value, Q 2 Value comparison analysis results. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be described in detail below. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other implementation methods obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0058] Example 1

[0059] like Figure 1 As shown, this embodiment provides a method for constructing a near-infrared quantitative model, including:

[0060] S1: Obtain near-infrared spectral data and construct a training sample set for the near-infrared quantitative model.

[0061] S2: Construct a particle swarm search space, obtain the value ranges of various near-infrared spectral data preprocessing algorithms, variable selection algorithms, and modeling algorithms and their parameters, and construct a particle swarm search space composed of these parameters and various infrared spectral data preprocessing algorithms in sequence.

[0062] Specifically, the particle swarm search space construction process is as follows:

[0063] Construct a set of near-infrared spectral data preprocessing algorithms, variable selection algorithms, and modeling algorithms, A = {A1, A2, ..., A...} m},in, For m-2 near-infrared spectral data preprocessing algorithms, A m-1 For the variable selection algorithm, A m For modeling algorithms;

[0064] For algorithm A i The parameters of (i∈{1,2,…,m}) are denoted as Each parameter has a certain range of values, among which parameter P ij (i∈{1,2,…,m},j∈{1,2,…,k i The range of values ​​for}) is denoted as Then, a search space S of n = n' + m - 2 dimensions is constructed sequentially from the parameters of all algorithms and the near-infrared spectral data preprocessing algorithms. n ,in: This represents the total number of parameters for the near-infrared spectral data preprocessing algorithm, variable selection algorithm, and modeling algorithm; m-2 represents the order of the m-2 preprocessing algorithms.

[0065] Search Space S n For each position in the vector, an n-dimensional data vector P = (p1, p2, ..., p...) is used. n' ,p n'+1 ,p n'+2 ,…,p n'+m-2 The expression ) represents the search space S. In the first n' dimensions, the value range of each dimension corresponds to the value range of the algorithm parameters. In the last m-2 dimensional sub-vector, the value range of each dimension is {1,2,…,m-2}, and the value of each component in this last m-2 dimensional sub-vector is distinct from the values ​​of other components. This sub-vector represents the execution sequence number of the corresponding near-infrared spectral data preprocessing algorithm in a series of preprocessing steps. n The position in the middle is represented as P = (p1, p2, ..., p n' ,p n'+1 ,p n'+2 ,…,,p n'+m-2 ), where {p n'+1 ,p n'+2 ,…,,p n'+m-2 The range of values ​​for} is {1,2,…,m-2}, and p n'+1 ≠p n'+2 ≠…≠p n'+m-2 For preprocessing algorithm A i For (i∈{1,2,…,m-2}), its execution order number among the m-2 preprocessing algorithms is p. n'+i ;

[0066] Since the parameter ranges of different algorithms can be either a continuous region of real numbers (e.g., the parameter range of the SNV algorithm is a preset continuous interval of real numbers Γ = (a, b)) or discrete integers (e.g., the parameter range of the derivative data processing method is Γ = {1, 2}, representing the calculation of the first or second derivative), the search space S... n The interior is discontinuous;

[0067] Assume the search space S n Let O be the region enclosed by the outer envelope, and let Φ be the non-search region within space O. Then, we have the particle swarm search space S. n for:

[0068] S n =O-Φ.

[0069] S3: Based on the training sample set of the near-infrared quantitative model, the particle swarm optimization algorithm is used to optimize the parameters of various near-infrared spectral data preprocessing algorithms, variable selection algorithms, and modeling algorithms, as well as the order of various infrared spectral data preprocessing algorithms, in the particle swarm search space; where each particle is a vector composed of the parameters of each algorithm and the order of various infrared spectral data preprocessing algorithms.

[0070] Specifically, the above process includes:

[0071] A1: Set the particle velocity iteration strategy, and use the linear decreasing weight method to set the particle flight velocity iteration strategy.

[0072] More specifically, a linear decreasing weighting mechanism is adopted as shown in the following formula:

[0073]

[0074] Where, ω (t) Let G represent the particle inertial weights at the t-th iteration. k ω represents the maximum number of iterations. ini Let ω be the initial inertia weight. end This represents the inertia weight at the maximum number of iterations. Initially, particles are given a larger inertia to improve their global search capability. Later in the iterations, smaller inertia is set to allow for better local search capability, thus improving both the algorithm's global search capability and convergence speed.

[0075] Therefore, the iterative formula for particle flight velocity can be expressed as follows:

[0076]

[0077] In the formula, Let c1 and c2 represent the flight velocity of the i-th particle in the t-th iteration, and let rand() represent a random number between (0,1). gbest represents the optimal position in the search history of the i-th particle. (t) This represents the optimal position in the entire particle search history. This represents the position of the i-th particle in the t-th iteration.

[0078] A2: Set a particle position iteration strategy, calculate the particle position after iteration based on the particle flight velocity, and verify the particle position according to the parameter value range of various near-infrared spectral data preprocessing algorithms, variable selection algorithms and modeling algorithms, as well as the sequential value range of various infrared spectral data preprocessing algorithms, and determine the final particle position after iteration.

[0079] More specifically, regarding the discontinuous parameter values ​​in the near-infrared spectral data preprocessing and modeling algorithms, the particle search space S... n There may be internal discontinuous regions, causing particles to potentially fall outside region O or inside region Φ during the search iteration process, potentially leading to infeasible solutions. After each particle iteration, the particle position is verified, and fine-tuning and optimization are performed based on the verification results to prevent particle i from falling into the search space S. n outside.

[0080] Suppose that the position of the i-th particle after t iterations is Predict the particle's position in the next iteration based on its flight speed. x' i =[x' i1 ,x' i2 ,…,x' in Let ] be an n-dimensional vector, representing each parameter P. ij (i∈{1,2,…,m},j∈{1,2,…,k i}) and the specific values ​​of the preprocessing algorithm order; if x' i Not in the particle swarm search space S n Inside, it represents x' i If k algorithm parameters have values ​​that are not within the range of their respective parameter ranges, they are denoted as k. From x' i Extract the corresponding k-dimensional data and search the n-dimensional particle swarm space S. n Extract the corresponding k-dimensional sub-region And find the k-dimensional subregion using Euclidean distance. Surface and x i 'Nearest point' Replace x with the value in y' iThe values ​​of the corresponding dimensions are used to obtain a vector. This serves as the next search location for the particle; the Euclidean distance is calculated using the following formula:

[0081]

[0082] The final determined particle position update strategy is shown in the following equation:

[0083]

[0084] A3: Particle swarm initialization, initializing the number of particles, the initial position and initial velocity of each particle, and the particle iteration exit mechanism; wherein, the position of the particle is represented by a vector composed of the parameters of various near-infrared spectral data preprocessing algorithms, variable selection algorithms and modeling algorithms, and the sequence of various infrared spectral data preprocessing algorithms.

[0085] More specifically, the particle iteration exit mechanism includes reaching the maximum number of iterations and model R. 2 There are two scenarios: the indicator value reaches the set value.

[0086] Among them, model R 2 The particle iteration exit mechanism when the indicator value reaches the set value specifically includes:

[0087] During particle swarm optimization, after each iteration, the algorithm parameters represented by the particles and various infrared spectral data preprocessing algorithms are sequentially substituted into the corresponding near-infrared spectral data preprocessing algorithm, variable selection algorithm, and modeling algorithm. Modeling is then performed based on the near-infrared quantitative model training sample set, and the R-squared value of the model is utilized. 2 The index value calculates the distance Dist between the particle and the optimal solution (i.e., the fitness function). If the distance Dist is less than a preset value, the iteration is terminated. The formula for calculating Dist is as follows:

[0088] Dist = 1 - R 2

[0089] Among them, R 2 The formula for calculating the index value is shown below:

[0090]

[0091] Among them, y i This represents the true value of the i-th training sample. This represents the predicted value of the i-th training sample. q represents the average of the true values ​​of all training samples; q represents the number of experimental samples.

[0092] Combining the above two equations, the distance between the particle's current position and the optimal solution during the particle iterative search process is determined by R0, which is derived from the model trained based on the particle's current position. 2 The index value, and the R of the optimal near-infrared model 2 Indicator value (R) 2 The difference between (=1) is represented by the following formula.

[0093]

[0094] A4: Particle swarm position verification and adjustment. Determine whether the positions of all particles are feasible parameters for near-infrared preprocessing algorithms, variable selection algorithms, and modeling algorithms, and whether the order of infrared spectral data preprocessing algorithms is feasible. If so, calculate the distance between the particles and the optimal solution. Otherwise, adjust and optimize the positions of the particles, and then calculate the distance between their positions and the optimal solution.

[0095] A5: Particle iteration. If the particle iteration exit mechanism is satisfied, the historical global optimal position is output as the optimal solution of the particle swarm optimization algorithm; otherwise, the next flight position of the particle is calculated according to the particle flight speed iteration strategy and the particle position iteration strategy, and the process returns to step A4.

[0096] S4: Based on the particle swarm iteration output, the obtained global optimal solution is selected as the parameter of each algorithm and the order of various infrared spectral data preprocessing algorithms. The near-infrared spectral data is preprocessed, variables are selected and modeled to obtain a near-infrared quantitative model.

[0097] Example 2

[0098] This embodiment provides a near-infrared quantitative model construction device, including:

[0099] The data acquisition module is used to acquire near-infrared spectral data and construct a training sample set for the near-infrared quantitative model.

[0100] The particle swarm search space construction module is used to obtain the value ranges of various near-infrared spectral data preprocessing algorithms, variable selection algorithms, and modeling algorithms and their parameters, and construct a particle swarm search space composed of these parameters and the sequence of various infrared spectral data preprocessing algorithms.

[0101] The parameter optimization module is used to optimize the parameters of various near-infrared spectral data preprocessing algorithms, variable selection algorithms, and modeling algorithms, as well as the order of various infrared spectral data preprocessing algorithms, based on the training sample set of the near-infrared quantitative model. The particle swarm optimization algorithm is used to optimize the parameters of various near-infrared spectral data preprocessing algorithms, variable selection algorithms, and modeling algorithms in the particle swarm search space. Each particle is a vector composed of the parameters of each algorithm and the order of various infrared spectral data preprocessing algorithms.

[0102] The model building module is used to select the globally optimal solution as the parameters of each algorithm and the order of various infrared spectral data preprocessing algorithms based on the particle swarm iteration output results, to preprocess, select variables and model near-infrared spectral data, and obtain a near-infrared quantitative model.

[0103] Example 3

[0104] This embodiment provides an electronic device, including:

[0105] A memory that stores computer programs;

[0106] When the processor loads and executes the computer program stored in the memory, it implements the near-infrared quantitative model construction method as described in Example 1.

[0107] Example 4

[0108] This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the near-infrared quantitative model construction method as described in Embodiment 1.

[0109] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0110] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0111] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0112] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0113] It is understood that the same or similar parts in the above embodiments can be referred to each other, and the contents not described in detail in some embodiments can be referred to the same or similar contents in other embodiments.

[0114] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0115] To enhance understanding of the technical solution of the present invention, a specific example is provided below to further illustrate the technical solution of the present invention.

[0116] Example data: using Figure 2 The near-infrared spectral data of the three batches, totaling 408 fragrance and flavor samples, as shown, along with their corresponding aroma identification results and physical indicators, were used for modeling and analysis.

[0117] The instance data is shown in the table below:

[0118] Implementation data Sample size Sample classification (near-infrared data acquisition) Scent and physical properties (dependent variable) Data 1 89 Undiluted sample Nine aroma indicators, including light fragrance, wine aroma, and sour aroma. Data 2 268 Undiluted sample Two physical indicators: refractive index and relative density Data 3 51 Sample diluted with distilled water Two physical indicators: refractive index and relative density

[0119] (1) Selection of preprocessing algorithm and variable selection algorithm, and construction of parameter optimization search scenario.

[0120] In this example, the PLS algorithm is used as the modeling algorithm. Due to the special nature of the latent variable parameter calculation in the PLS algorithm, the latent variable parameter optimization of the modeling algorithm adopts external validation set selection. Therefore, in this example, it is not necessary to optimize the modeling algorithm parameters during the particle swarm optimization process. Finally, the particle swarm optimization algorithm is determined to be used for preprocessing, variable selection (where non-informative variables are removed as part of the variable selection algorithm), and algorithm parameter optimization, as shown in the table below:

[0121]

[0122]

[0123] As shown in the table above, the preprocessing and variable selection of experimental spectral data involves a total of 9 preprocessing algorithms and 1 variable selection algorithm. Except for the Standard Normal Variable Transform (SNV) algorithm, which requires no parameter settings, 15 algorithm parameters need to be set and optimized together. These, along with the order of the 9 preprocessing algorithms, together construct a 24-dimensional algorithm parameter space S. 24 It should be noted that for the sample normalization preprocessing algorithm, the centering preprocessing algorithm, and the scaling preprocessing algorithm, the numerical labels of each specific processing method are set in advance, thus using each specific processing method as a parameter. Taking the sample normalization preprocessing algorithm as an example, its parameter is the normalization method, and its value range is [0,1,2,3,4,5]. When the value is 0, the normalization method uses area normalization; when the value is 1, the normalization method uses maximum value normalization; when the value is 2, the normalization method uses unit vector normalization; when the value is 3, the normalization method uses range normalization; when the value is 4, the normalization method uses average value normalization; and when the value is 5, the normalization method uses peak value normalization. Other preprocessing algorithms can be found in the table above and will not be described in detail here.

[0124] (2) Particle Iterative Search Strategy Design

[0125] As mentioned earlier, in the above search space, the velocity update formula for the particle search process is designed as follows:

[0126]

[0127]

[0128] That is, setting the learning factors in the particle velocity update algorithm to c1 = c2 = 2, and the linearly decreasing weight inertia function ω (t) In the middle, set the initial inertia weight ω ini =0.7, the inertia weight ω at the maximum number of iterations. end =0.3, to improve the global optimization ability of particles in the early search process and the local optimization ability in the later search process.

[0129] (3) Particle initialization design

[0130] The initial particle swarm size is 200, the particle exit threshold is 0.2 (i.e., model R2 ≥ 0.8), and the maximum number of iterations is 20.

[0131] For each particle, its first 15 dimensions are randomly determined within the range of corresponding parameter values, and its last 9 dimensions are determined by randomly shuffling the order of any components in the vector (1,2,3,4,5,6,7,8,9). The final 15 + 9 = 24 values ​​are used as the initial position of the particle. After performing the above processing on all 200 particles, the initial positions of each particle are obtained.

[0132] (4) Evaluation of the model

[0133] After completing the preprocessing and variable selection analysis of the near-infrared spectral data, two-thirds of the data were randomly selected as the model training set, and the remaining one-third was used as the model validation set. A partial least squares regression model was constructed, and the R-squared of the model was calculated. 2 Value. After model training is complete, the Q-value is calculated through external validation of the model. 2 The value is used as the final evaluation metric for the model.

[0134] like Figure 3 As shown, in the particle swarm optimization process, the fastest iteration to find the exit condition (R0) is 5 iterations. 2 For data preprocessing and variable selection parameters ≥0.8, a maximum of 9 iterations are required, and the construction of 13 models averages 6.7 iterations. This is significantly lower than the number of iterations required by traditional analysis methods, indicating that the particle swarm optimization scheme proposed in this invention has performance far superior to traditional modeling and analysis methods.

[0135] like Figure 4 As shown, overall, the external validation Q of the model constructed using particle swarm optimization parameters is comparable to that of the model constructed using traditional analysis. 2 The average values ​​are 0.87 and 0.66, respectively. Looking at the 13 models, the model Q constructed after particle swarm optimization of the parameters... 2 The results were almost all superior to those obtained by models constructed using traditional analysis methods. This fully demonstrates the effectiveness of the particle swarm optimization method in quantitative modeling and analysis of near-infrared data.

[0136] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A near infrared quantitative model construction method, characterized in that, The method comprises the following steps: obtaining near-infrared spectrum data, and constructing a near-infrared quantitative model training sample set; constructing a particle swarm search space, obtaining a value range of a plurality of near-infrared spectrum data preprocessing algorithms, variable selection algorithms and modeling algorithms and parameters, and constructing a particle swarm search space composed of the parameters and a plurality of near-infrared spectrum data preprocessing algorithm sequences; based on the near-infrared quantitative model training sample set, the particle swarm optimization algorithm is used to optimize the parameters of the plurality of near-infrared spectrum data preprocessing algorithms, variable selection algorithms and modeling algorithms and the plurality of near-infrared spectrum data preprocessing algorithm sequences in the particle swarm search space; wherein each particle is a vector composed of the parameters of the algorithms and the plurality of near-infrared spectrum data preprocessing algorithm sequences; specifically comprising: A1: setting a particle velocity iteration strategy, and setting a particle flight velocity iteration strategy based on a linear decreasing weight method; A2: setting a particle position iteration strategy, calculating the particle position after iteration based on the particle flight velocity, verifying the particle position after iteration according to the value range of the parameters of the plurality of near-infrared spectrum data preprocessing algorithms, variable selection algorithms and modeling algorithms and the value range of the plurality of near-infrared spectrum data preprocessing algorithm sequences, and determining the final particle position after iteration; A3: particle swarm initialization, initializing the number of particles, the initial position and initial velocity of each particle, and the particle iteration exit mechanism; wherein the position of the particle is represented by a vector composed of the parameters of the plurality of near-infrared spectrum data preprocessing algorithms, variable selection algorithms and modeling algorithms and the plurality of near-infrared spectrum data preprocessing algorithm sequences; A4: particle swarm position verification and adjustment, judging whether the position of all particles is a feasible near-infrared preprocessing algorithm, variable selection algorithm, modeling algorithm parameter and feasible near-infrared spectrum data preprocessing algorithm sequence; if yes, the distance between the particle and the optimal solution is calculated; otherwise, the position of the particle is adjusted and optimized, and then the distance between the position of the particle and the optimal solution is calculated; A5: particle iteration, if the particle iteration exit mechanism is met, the historical global optimal position is taken as the optimal solution of the particle swarm optimization algorithm and output; otherwise, the next flight position of the particle is calculated according to the particle flight velocity iteration strategy and the particle position iteration strategy, and the step A4 is returned to; according to the particle swarm iteration output result, the global optimal solution obtained is selected as the parameters of the algorithms and the plurality of near-infrared spectrum data preprocessing algorithm sequences, near-infrared spectrum data is preprocessed, variable selection and modeling are performed, and a near-infrared quantitative model is obtained.

2. The near infrared quantitative model construction method according to claim 1, characterized by, The particle swarm search space construction process is as follows: A set of near infrared spectroscopy data preprocessing algorithms, variable selection algorithms and modeling algorithms are constructed wherein, is a near infrared spectroscopy data preprocessing algorithm, a variable selection algorithm, a modeling algorithm; For the algorithm the parameter is denoted as where the parameter the value range of the parameter is denoted as ; then a search space of dimension is constructed by all the parameters of the algorithm and the sequence of the near infrared spectral data preprocessing algorithm, where: , represents the total number of parameters of the near infrared spectral data preprocessing algorithm, variable selection algorithm and modeling algorithm, represents the sequence of the preprocessing algorithm;​ Search space Each position in the search space is represented by a dimensional data vector ; wherein, in the former dimensional vector, the value range of each dimension is the value range of the corresponding algorithm parameter; in the latter dimensional sub-vector, the value range of each dimension is , and the value of each component in the latter dimensional sub-vector is different from that of other components, and is used to represent the execution order number of the corresponding near-infrared spectrum data preprocessing algorithm in a series of preprocessing processes. Assume the search space The area surrounded by the outer envelope surface of The non-search area in the space is recorded as The particle swarm search space is: 。 3. The near infrared quantitative model construction method according to claim 1, characterized by, in the step A1, the particle flight velocity iteration formula is as follows: ; wherein represents the velocity of the i-th particle at the t-th iteration, c1 and c2 each represent a learning factor, and rand() represents a random number between (0, 1), represents the optimal position in the search history of the i-th particle, represents the optimal position in the search history of all particles, represents the position of the i-th particle at the t-th iteration; represents the inertia weight of the particle at the t-th iteration, and the calculation formula is as follows: ; In the formula, is the maximum number of iterations, is the initial inertia weight, is the inertia weight when the maximum number of iterations is reached.

4. The near infrared quantitative model construction method according to claim 1, characterized by, In step A2, the position of the particle in the next iteration is predicted based on the particle's flight velocity. , For one dimensional vector, if Not in the particle swarm search space Inside, it means There is If the value of an algorithm parameter is not within the range of the corresponding algorithm parameter, it is denoted as... ,from Extract the corresponding Dimensional data, from 3D particle swarm search space Extract the corresponding Vic region And find it through Euclidean distance Vic region Surface and The closest point ,Will Value replacement The values ​​of the corresponding dimensions are used to obtain a vector. As the next search location for the particle; the final determined particle position update strategy is as follows: 。 5. The near infrared quantitative model construction method according to claim 1, characterized by, The particle iteration exit mechanism in step A3 includes reaching a maximum number of iterations and a model The index value reaches a set value.

6. The near infrared quantitative model construction method according to claim 5, characterized in that, wherein, Model The particle iteration exit mechanism when the index value reaches the set value specifically includes: In the particle swarm optimization process, after each iteration is completed, the algorithm parameters represented by the particles and various infrared spectral data preprocessing algorithms are sequentially substituted into the corresponding near-infrared spectral data preprocessing algorithms, variable selection algorithms and modeling algorithms, modeling is performed based on the near-infrared quantitative model training sample set, and the index value is calculated to calculate the distance between the current solution and the optimal solution If the distance is less than a preset value, the iteration is exited. The calculation formula is as follows: ; wherein, The calculation formula of the index value is as follows: ; wherein, represents the true value of the i-th training sample, represents the predicted value of the i-th training sample, represents the average value of all training sample true values; q represents the number of experimental samples.

7. A near-infrared quantitative model construction device, characterized in that, the method and device for constructing a near-infrared quantitative model according to any one of claims 1 to 6 comprise: a data acquisition module for acquiring near-infrared spectrum data and constructing a near-infrared quantitative model training sample set; A particle swarm search space construction module is configured to acquire a value range of a plurality of near-infrared spectral data preprocessing algorithms, variable selection algorithms, modeling algorithms and parameters thereof, and construct a particle swarm search space composed of the parameters and a sequence of the plurality of near-infrared spectral data preprocessing algorithms; A parameter optimization module is configured to train a near-infrared quantitative model sample set, and optimize the parameters of the plurality of near-infrared spectral data preprocessing algorithms, variable selection algorithms and modeling algorithms and the sequence of the plurality of near-infrared spectral data preprocessing algorithms in the particle swarm search space by using a particle swarm optimization algorithm, wherein each particle is a vector composed of the parameters of the algorithms and the sequence of the plurality of near-infrared spectral data preprocessing algorithms; A model construction module is configured to select a global optimal solution obtained according to a particle swarm iteration output result as the parameters of the algorithms and the sequence of the plurality of near-infrared spectral data preprocessing algorithms, perform preprocessing, variable selection and modeling on near-infrared spectral data, and obtain a near-infrared quantitative model.

8. An electronic device, comprising: It comprises: a memory storing a computer program; a processor, when loading and executing the computer program stored on the memory, implements the near-infrared quantitative model construction method according to any one of claims 1 to 6.

9. A computer readable storage medium storing a computer program, characterized in that, The computer program, when executed by the processor, implements the near-infrared quantitative model construction method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Characteristic wavelength selecting method for near infrared spectrum in ant colony optimization algorithm

    CN103344600A

  • Near infrared spectrum wavelength selecting method based on particle swarm optimization

    CN103913432A