A waveband selection method

CN122528111BActive Publication Date: 2026-09-15NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610991690.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-15
Estimated Expiration
2046-07-06

AI Technical Summary

Technical Problem

[0005]本申请提供了一种波段选择方法,以至少解决现有技术中传统遗传算法在光谱波段选择中缺乏对种群进化状态的自适应动态调控机制,导致算法极易陷入局部最优而早熟收敛,最终造成波段筛选精度低下的技术问题

Benefits of technology

[0005] This application provides a band selection method to at least solve the technical problem that traditional genetic algorithms in the prior art lack an adaptive dynamic control mechanism for the population evolution state in spectral band selection, which makes the algorithm prone to getting trapped in local optima and premature convergence, ultimately resulting in low band selection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122528111B_ABST
    Figure CN122528111B_ABST
Patent Text Reader

Abstract

The application discloses a wave band selection method, belonging to the technical field of spectral analysis. It includes: based on the spectral physical continuity prior and the regression model verification set performance index, constructing the task specialization genetic algorithm, which represents the wave band subset with a binary mask vector, takes the performance index as the fitness function, and adopts the structured search operator based on the continuity prior and the traditional genetic operator to form a hybrid search mechanism; based on the population evolution state representation and the fitness improvement feedback, constructing the reinforcement learning meta-control framework for synergistically regulating the algorithm hyperparameters and the structured operator; based on the recurrent neural network, constructing the timing-enhanced reinforcement learning agent; based on the spectral physical interference mechanism and the feature variation law, generating the simulation data for the iterative training of the agent; and taking the trained agent as the meta-controller to drive the algorithm to run and complete the wave band selection. The application solves the problems of the traditional genetic algorithm, such as lack of adaptive regulation, easy premature convergence and low screening precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of spectral analysis technology, and in particular to a band selection method. Background Technology

[0002] Visible-near-infrared spectroscopy is widely used in various fields due to its advantages such as rapid detection and non-destructive analysis. However, the high-dimensional data acquired by modern spectrometers exhibits strong collinearity and a large amount of redundancy, necessitating band selection to remove useless information and build robust models. Among existing band selection methods, genetic algorithms have become a commonly used wrapper method for hyperspectral band selection because they do not rely on data distribution assumptions and are adaptable to discrete search spaces.

[0003] However, traditional genetic algorithms have significant drawbacks in the selection of visible-near-infrared spectral bands: on the one hand, they lack an adaptive dynamic control mechanism for the population's evolutionary state, and hyperparameter adjustment relies heavily on human experience. Existing adaptive variants also rely on rigid manual rules and cannot achieve real-time and accurate feedback, making the algorithm prone to premature convergence in local optima. On the other hand, random crossover and mutation operations can easily disrupt the continuous distribution characteristics of chemical bond absorption peaks, resulting in low band selection accuracy and difficulty in meeting the requirements of high-precision, high-stability, and automated quantitative spectral analysis.

[0004] In summary, existing technologies suffer from the problem that traditional genetic algorithms lack an adaptive dynamic control mechanism for the population's evolutionary state in spectral band selection. This makes the algorithm prone to getting trapped in local optima and premature convergence, ultimately resulting in low band selection accuracy. Summary of the Invention

[0005] This application provides a band selection method to at least solve the technical problem that traditional genetic algorithms in the prior art lack an adaptive dynamic control mechanism for the population evolution state in spectral band selection, which makes the algorithm prone to getting trapped in local optima and premature convergence, ultimately resulting in low band selection accuracy.

[0006] Firstly, this application provides a band selection method, the method comprising: Based on the spectral physical continuity prior and the performance index of the regression model validation set, a task-specific genetic algorithm is constructed. The task-specific genetic algorithm represents the band subset with a binary mask vector, uses the performance index as the fitness function, and adopts a hybrid search mechanism consisting of a structured search operator based on the continuity prior and a traditional genetic operator. Based on population evolution state representation and fitness improvement feedback, a reinforcement learning meta-control framework is constructed. The reinforcement learning meta-control framework is used to coordinate the hyperparameters and structured operators of the task-specific genetic algorithm. Based on recurrent neural networks, a temporally enhanced reinforcement learning agent is constructed within the aforementioned reinforcement learning meta-control framework. Based on the spectral physical interference mechanism and the variation law of spectral features, simulated training data is generated to iteratively train the reinforcement learning agent. The trained reinforcement learning agent is used as the meta-controller of the task-specialized genetic algorithm, outputting control actions to drive the task-specialized genetic algorithm to run, and applied to real spectral data to complete band selection.

[0007] The above technical solution introduces a structured search operator based on spectral continuity priors to preserve the continuous characteristics of absorption peaks. It combines a temporally enhanced reinforcement learning agent to dynamically adjust the hyperparameters and operators of the genetic algorithm according to the real-time evolutionary state of the population. Furthermore, it uses physically heuristic simulation data to pre-train the agent. This achieves the technical effects of reducing the destruction of effective bands by random search, alleviating the premature convergence phenomenon of the algorithm, and improving the generalization ability of the model on real spectral data. Ultimately, it obtains more accurate, stable, and human-intervention-free spectral band selection results.

[0008] Secondly, this application provides a band selection system, the system comprising: The genetic algorithm construction module constructs a task-specific genetic algorithm based on the spectral physical continuity prior and the performance index of the regression model validation set. The task-specific genetic algorithm represents the band subset with a binary mask vector, uses the performance index as the fitness function, and adopts a hybrid search mechanism consisting of a structured search operator based on the continuity prior and a traditional genetic operator. The meta-control framework construction module constructs a reinforcement learning meta-control framework based on population evolution state representation and fitness improvement feedback. The reinforcement learning meta-control framework is used to coordinate the regulation of the hyperparameters and structured operators of the task-specific genetic algorithm. The agent construction module, based on a recurrent neural network, constructs a temporally enhanced reinforcement learning agent under the reinforcement learning meta-control framework; The data generation and training module generates simulated training data based on the spectral physical interference mechanism and the variation law of spectral features, which is used to iteratively train the reinforcement learning agent. The band selection module uses the trained reinforcement learning agent as the meta-controller of the task-specialized genetic algorithm, outputting control actions to drive the task-specialized genetic algorithm to run, and applying it to real spectral data to complete band selection.

[0009] Thirdly, this application provides an electronic device comprising one or more processors and one or more memories, wherein at least one piece of program code is stored in the one or more memories, the program code being loaded and executed by the one or more processors to implement the operations performed by the band selection method.

[0010] Fourthly, this application also provides a computer-readable storage medium storing at least one piece of program code, which is loaded and executed by a processor to implement the operations performed by the band selection method.

[0011] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described band selection methods. Attached Figure Description

[0012] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0013] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 A flowchart illustrating a band selection method provided in this application embodiment. Figure 1 ; Figure 2 A flowchart illustrating a band selection method provided in this application embodiment. Figure 2 ; Figure 3 A flowchart illustrating a band selection method provided in this application embodiment. Figure 3 ; Figure 4 A flowchart illustrating a band selection method provided in this application embodiment. Figure 4 ; Figure 5 A flowchart illustrating a band selection method provided in this application embodiment. Figure 5 ; Figure 6 A flowchart illustrating a band selection method provided in this application embodiment. Figure 6 ; Figure 7 A flowchart illustrating a band selection method provided in this application embodiment. Figure 7 ; Figure 8 A flowchart illustrating a band selection method provided in this application embodiment. Figure 8 ; Figure 9 A flowchart illustrating a band selection method provided in this application embodiment. Figure 9 ; Figure 10 A flowchart illustrating a band selection method provided in this application embodiment. Figure 10 ; Figure 11 This is a schematic diagram of the structure of a band selection system provided in an embodiment of this application; Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Figure 13 This is a schematic diagram illustrating the operation of the three structured operators provided in the embodiments of this application; Figure 14 A schematic diagram comparing the convergence accuracy of PPO-GA and fixed-strategy GA provided in the embodiments of this application; Figure 15 A comparison of the R² distribution of the RL-GA, fixed-policy GA, SPA and CARS algorithms provided in the embodiments of this application on four datasets; Figure 16 The action probability evolution curves of the PPO agent provided in the embodiments of this application on four datasets. Detailed Implementation

[0015] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0016] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0017] Among related technologies, visible-near-infrared (Vis-NIR) spectroscopy, with its core advantages of rapid detection, non-destructive analysis, and simultaneous quantification of multiple components, has been widely applied in various fields such as agriculture, chemical industry, biomedicine, environmental monitoring, and mineral detection. Modern spectrometers can collect spectral data from hundreds to thousands of adjacent continuous bands. This type of high-dimensional data has strong multicollinearity and a large amount of redundant information, which can easily lead to the "Hughes phenomenon," causing problems such as overfitting and decreased prediction accuracy during modeling. At the same time, spectral signals are often affected by noise such as background baseline shift and scattering interference. Directly using full-spectrum data for modeling increases computational overhead and reduces the model's generalization ability. Therefore, selecting the most informative band subset from the full spectrum is a key step in building robust chemometric models.

[0018] Currently, spectral band selection methods are mainly classified into three categories: filtering, embedded, and wrapper methods. Filtering methods, such as the Successive Projections Algorithm (SPA), are computationally efficient but struggle to capture the complex nonlinear relationships between features and target variables. Embedded methods, such as the Lasso algorithm, may disrupt the physical continuity of spectral bands, causing the selection results to lose their physical meaning. Among wrapper methods, the Genetic Algorithm (GA) is widely used in hyperspectral band selection because it does not rely on data distribution assumptions and is adaptable to discrete search spaces.

[0019] However, traditional genetic algorithms have significant drawbacks: First, their performance is highly dependent on empirical adjustments of hyperparameters such as mutation rate and selection pressure. Improper hyperparameter settings can easily lead to premature convergence, making it difficult to reach the global optimum. Second, existing adaptive GA variants rely on manually designed empirical rules and lack a principled feedback mechanism for the real-time state of population evolution, making it impossible to achieve precise adaptive adjustment of hyperparameters. Third, they do not fully utilize the physical continuity of spectral bands, and random searches can easily disrupt the continuous distribution of chemical bond absorption peaks, affecting screening accuracy. Fourth, the limited number of real spectral samples results in insufficient training data for the algorithm, leading to poor generalization ability and Sim-to-Real transfer performance.

[0020] Deep Reinforcement Learning (DRL) has unique advantages in sequence decision-making and delayed feedback processing, providing new ideas for solving dynamic optimization problems. However, existing methods for end-to-end band selection using RL suffer from exponentially expanding action spaces when dealing with high-dimensional spectral data, leading to low convergence efficiency. Therefore, there is an urgent need for a method that can integrate the advantages of Reinforcement Learning (RL) and genetic algorithms, overcome the inherent shortcomings of genetic algorithms, and achieve accurate, stable, and automated selection of Vis-NIR spectral bands.

[0021] In summary, existing technologies suffer from the problem that traditional genetic algorithms lack an adaptive dynamic control mechanism for the population's evolutionary state in spectral band selection. This makes the algorithm prone to getting trapped in local optima and premature convergence, ultimately resulting in low band selection accuracy.

[0022] To address the aforementioned technical challenges, this application proposes an adaptive band selection mechanism that deeply integrates reinforcement learning and genetic algorithms. This mechanism introduces a reinforcement learning agent as a meta-controller, dynamically adjusting the hyperparameters and structured operators of the genetic algorithm based on the real-time evolutionary state of the population to overcome premature convergence. Simultaneously, it incorporates prior knowledge of spectral physical continuity to design structured search operators, thereby improving screening accuracy. Furthermore, it utilizes recurrent neural networks to enhance temporal decision-making capabilities and employs physically heuristic data synthesis to improve generalization performance, thus achieving precise, stable, and automated spectral band selection.

[0023] The application scenarios of the technical solutions provided in the embodiments of this application are described below.

[0024] The technical solutions provided in this application are mainly applied in the field of visible-near-infrared spectroscopy analysis technology, and are particularly suitable for quantitative analysis scenarios that require dimensionality reduction of high-dimensional spectral data and screening of characteristic bands. Specifically, they can be widely applied in the following fields: 1. Agricultural and food inspection For example, it can be used for rapid, non-destructive testing of the internal quality of agricultural products such as the dry matter content of fruits like mangoes and the protein content of grains like wheat. Faced with complex lighting and background interference in the field, this application can screen out robust bands resistant to interference, improving the prediction accuracy of portable instruments.

[0025] 2. Mineral composition analysis For example, rapid component determination is performed during the mining and beneficiation of iron ore, such as the iron ore in Anshan. Given the significant influence of scattering effects and complex noise in ore spectra, this application utilizes physical heuristics to generate data, enhancing generalization capabilities and accurately locating absorption peak bands.

[0026] 3. Biopharmaceuticals For example, it can be used for online monitoring of the content of active ingredients during the production of tablets. The band selection method provided in this application can automatically eliminate redundant bands and screen out characteristic bands that are highly correlated with the active ingredients, meeting the stringent requirements of the pharmaceutical industry for model stability and accuracy.

[0027] 4. Energy and chemical production For example, it enables rapid analysis of physicochemical properties of diesel fuel, such as cetane number. In cases of multicollinearity in chemical spectral data, this application effectively avoids premature convergence of the algorithm and extracts the most informative subset of bands.

[0028] 5. Environmental monitoring For example, it can be used to quantitatively detect the concentration of organic pollutants in environmental water bodies. In real-world scenarios with low concentrations and strong interference, the structured search operator in this application can preserve the physical continuity of the spectrum, avoid disrupting weak characteristic absorption peaks, and improve the detection reliability in low signal-to-noise ratio scenarios.

[0029] In the aforementioned application scenarios, traditional full-spectrum modeling often faces problems such as high data dimensionality, large computational overhead, susceptibility to overfitting, and poor model generalization ability. This application achieves highly accurate, stable, and automated band selection through deep collaboration between reinforcement learning and genetic algorithms. This not only improves the performance of quantitative analysis models but also lowers the modeling threshold, adapting to the intelligent detection needs of multiple fields and multi-source spectra.

[0030] The technical solutions provided in the embodiments of this application are described below.

[0031] This application provides a band selection method. Figure 1 A flowchart of a band selection method provided in an embodiment of this application is shown. This method can be applied to spectral analysis equipment or electronic devices with data processing capabilities. The following will be combined with... Figure 1 Each step is explained in detail.

[0032] Step S101: Based on the prior knowledge of spectral physical continuity and the performance index of the regression model validation set, construct a task-specific genetic algorithm.

[0033] The task-specific genetic algorithm uses a binary mask vector to represent a band subset, the performance index as the fitness function, and a hybrid search mechanism consisting of a structured search operator based on continuous priors and a traditional genetic operator.

[0034] Spectroscopic physical continuity priors can be used to indicate the physical property that chemical bond absorption peaks in a spectrum exhibit a continuous distribution. In traditional genetic algorithms, random crossover and mutation operations can easily disrupt this continuity, causing the selected bands to lose their physical meaning.

[0035] Binary mask vectors can be used to encode a subset of bands. For example, a "1" in the mask vector can indicate that the corresponding band is selected, and a "0" can indicate that the corresponding band is removed.

[0036] Performance metrics for the validation set of a regression model can include, but are not limited to, the coefficient of determination (R²) and the root mean square error (RMSE). It can be understood that using this performance metric as a fitness function links band selection to the final model prediction performance, making the fitness assessment more aligned with the needs of actual regression tasks.

[0037] Structured search operators based on continuity priors refer to operators that perform overall preservation, elimination, or boundary fine-tuning on continuous band regions. The aforementioned traditional genetic operators can include traditional crossover and mutation operators. By combining structured search operators with traditional genetic operators to form a hybrid search mechanism, the complete structure of absorption peaks can be effectively preserved during the search process, avoiding the disruption of characteristic band continuity caused by random operations.

[0038] Specifically, in spectral analysis, the absorption peaks of chemical bonds often exhibit a continuous distribution. However, the random crossover and mutation operations of traditional genetic algorithms easily disrupt this continuity, causing the selected band subsets to lose their physical meaning and limiting accuracy. Therefore, this step specializes the standard genetic algorithm based on the prior knowledge of spectral physical continuity. On the one hand, a binary mask vector is used, such as 0 indicating elimination and 1 indicating selection, to intuitively encode the band subsets. The performance index of a regression model, such as partial least squares regression, on the validation set is used as the fitness function, with the performance index, such as the coefficient of determination R², directly linking band selection with model prediction performance. On the other hand, a purely random search method is abandoned, and a structured search operator based on a continuity prior is introduced. These operators can preserve entire continuous bands or fine-tune boundaries, forming a hybrid search mechanism in conjunction with traditional crossover and mutation operators. This effectively preserves the complete structure of absorption peaks during the search process and avoids feature destruction caused by random operations.

[0039] Step S102: Based on the population evolution state representation and fitness improvement feedback, construct a reinforcement learning meta-control framework.

[0040] The reinforcement learning meta-control framework is used to coordinate the hyperparameters and structured operators of the task-specific genetic algorithm.

[0041] Population evolutionary state representation can be used to indicate information such as the current evolutionary progress and diversity distribution of a population. For example, population evolutionary state representation can include population fitness variance, the gap between the best individual and the average fitness, etc.

[0042] Fitness improvement feedback can be used to indicate the extent to which the current evolutionary strategy contributes to fitness improvement. It can be understood that by sensing the improvement in fitness, the effectiveness of current regulatory actions can be assessed.

[0043] The reinforcement learning meta-control framework can formalize the dynamic adjustment problem of hyperparameters and structured operators in task-specific genetic algorithms into a Markov decision process. This framework guides the agent to make decisions by perceiving the evolutionary state representation of the population and combining it with fitness improvement feedback, thereby achieving closed-loop synergistic control of hyperparameters of task-specific genetic algorithms, such as mutation rate and selection pressure, with structured operators.

[0044] Specifically, in traditional genetic algorithms, hyperparameters such as mutation rate and selection pressure are usually fixed or adjusted by manual rules during operation, failing to provide dynamic feedback based on the real-time evolutionary state of the population. This makes the algorithm prone to getting trapped in local optima and premature convergence. To address this issue, this step formalizes the dynamic adjustment problem of task-specialized genetic algorithm hyperparameters and structured operators as a Markov decision process, constructing a reinforcement learning meta-control framework. This framework, by sensing the evolutionary state representation of the population, captures information such as current evolutionary progress and diversity, and combines this with feedback signals from fitness improvement to guide the agent's decision-making. This achieves closed-loop synergistic control of task-specialized genetic algorithm hyperparameters and structured operators, enabling genetics to adaptively balance global exploration and local development at different evolutionary stages.

[0045] Step S103: Based on recurrent neural networks, construct a temporally enhanced reinforcement learning agent under the reinforcement learning meta-control framework.

[0046] Recurrent Neural Networks (RNNs) can include network structures with memory functions, such as Long Short-Term Memory Networks (LSTMs) or Gated Recurrent Units (GRUs).

[0047] In some embodiments of this application, the evolution of the task-specialized genetic algorithm is a dynamic process with time-dependent characteristics, where the population state and decisions of the previous stage directly affect the evolutionary trajectory of the next stage. If decisions are made solely based on the isolated state at the current moment, it is difficult to capture the long-term evolutionary trend.

[0048] By introducing recurrent neural networks, reinforcement learning agents can extract features from continuous population state sequences and capture temporal dependencies in the evolutionary process of task-specific genetic algorithms, thereby achieving temporal enhancement. This allows reinforcement learning agents to make regulatory decisions not only by considering the current instantaneous population state but also by combining historical evolutionary trajectories, resulting in more forward-looking and global regulatory strategies.

[0049] Specifically, the evolution of task-specialized genetic algorithms is a dynamic process with temporal dependencies; the population state and decisions of previous stages directly influence the evolutionary trajectory of subsequent stages. If decisions are made solely based on the isolated state at the current moment, the reinforcement learning agent struggles to capture long-term evolutionary trends. Therefore, this step introduces a recurrent neural network (RNN) to construct the reinforcement learning agent within the reinforcement learning meta-control framework. RNNs possess the ability to memorize historical states, enabling feature extraction from continuous population state sequences and capturing the temporal dependencies in the evolutionary process of task-specialized genetic algorithms, thus achieving temporal reinforcement. This allows the reinforcement learning agent to consider not only the current instantaneous population state but also historical evolutionary trajectories when making regulatory decisions, resulting in more forward-looking and global regulatory strategies.

[0050] Step S104: Based on the spectral physical interference mechanism and the variation law of spectral features, generate simulated training data for iterative training of the reinforcement learning agent.

[0051] Spectral physical interference mechanisms can include common interference factors encountered during real-world spectral acquisition. Examples include baseline drift, background fluctuations, and scattering effects.

[0052] The variation pattern of spectral characteristics can refer to the random variation pattern of characteristics such as the position, width, and height of absorption peaks in a spectrum.

[0053] Due to the high cost and limited sample size of acquiring real visible-near-infrared spectral data, training reinforcement learning agents directly on real data is prone to overfitting. Therefore, a data generator is constructed based on physical heuristics to generate large-scale, high-quality simulated training data by simulating various disturbances and feature variations present in real spectra.

[0054] In a simulated environment, the trained reinforcement learning agent can engage in extensive trial-and-error interactions with the simulated task-specific genetic algorithm evolution process, updating its policy based on state representations and feedback rewards to complete iterative training.

[0055] Specifically, training reinforcement learning agents typically requires massive amounts of interactive data. However, acquiring real visible-near-infrared spectral data is costly and has limited sample size. Training directly on real data easily leads to overfitting and poor generalization ability. Therefore, this step constructs a physically heuristic data generator based on spectral physical interference mechanisms, such as baseline drift, background fluctuations, and scattering effects, and the variation patterns of spectral features, such as random changes in the position, width, and height of absorption peaks. By simulating various interferences and feature variations present in real spectra, large-scale, high-quality simulated training data is generated. Iterative training of the reinforcement learning agent on this simulated data enables it to learn robust strategies to cope with different interferences and feature distributions, effectively improving the generalization performance of the reinforcement learning agent from simulated environments to real-world data scenarios.

[0056] Step S105: Use the trained reinforcement learning agent as the meta-controller of the task-specialized genetic algorithm, output the control action to drive the task-specialized genetic algorithm to run, and apply it to real spectral data to complete band selection.

[0057] After offline training, the trained reinforcement learning agent possesses the ability to dynamically adjust algorithm parameters based on its evolutionary state. In the inference application phase, the trained reinforcement learning agent is used as a meta-controller and combined with a task-specific genetic algorithm.

[0058] Faced with real spectral data to be selected, during the evolution of the task-specialized genetic algorithm, the trained reinforcement learning agent obtains the current state vector of the population in real time at each control time step, outputs regulatory actions based on the policy network it has learned, and dynamically adjusts the hyperparameters and structured operators of the task-specialized genetic algorithm.

[0059] Specifically, after offline training, the trained reinforcement learning agent possesses the ability to dynamically adjust algorithm parameters based on its evolutionary state. In the inference application phase, the trained reinforcement learning agent is combined with a task-specific genetic algorithm as a meta-controller. Faced with real visible-near-infrared spectral data to be selected, during the evolution of the trained reinforcement learning agent, the task-specific genetic algorithm acquires the current population's state vector in real time at each control time step. Based on the learned policy network, it outputs control actions, dynamically adjusting the hyperparameters and structured operators of the trained reinforcement learning agent. Driven by the trained reinforcement learning agent, the task-specific genetic algorithm adaptively avoids premature convergence traps and searches for the optimal band subset. When the preset termination condition is met, the task-specific genetic algorithm outputs the final optimal band subset, thus achieving highly accurate and automated band selection from real spectral data.

[0060] This embodiment introduces a structured search operator based on spectral continuity priors to preserve the continuous characteristics of absorption peaks. It combines a temporally enhanced reinforcement learning agent to dynamically adjust the hyperparameters and operators of the genetic algorithm according to the real-time evolutionary state of the population. Furthermore, it utilizes physically heuristic simulation data to pre-train the reinforcement learning agent. This achieves the technical effects of reducing the destruction of effective bands by random search, alleviating premature convergence of the algorithm, and improving the model's generalization ability on real spectral data. Ultimately, it obtains more accurate, stable, and human-intervention-free spectral band selection results.

[0061] It should be noted that the above steps S101-S105 are a simplified description of the embodiments provided in this application.

[0062] The band selection method provided in this application will be described in more detail below with some examples. See [link to relevant documentation]. Figure 2 Based on the prior knowledge of spectral physical continuity and the performance index of the regression model validation set, a task-specific genetic algorithm is constructed, which includes the following steps.

[0063] Step S201: Determine the fitness function and encoding method based on the performance metrics of the regression model validation set.

[0064] The encoding method is used to represent a subset of bands using a binary mask vector.

[0065] Specifically, the aforementioned performance metrics for the validation set of the regression model are used to measure the prediction accuracy of the regression model trained on the validation set using a selected band subset, such as the coefficient of determination (R²) and root mean square error (RMSE). A higher fitness function value indicates better prediction performance of the regression model corresponding to that band subset. The aforementioned binary mask vector is used to discretize and encode the spectral band subset. Assuming the full spectrum contains N bands, it can be encoded using a binary vector of length N, where "1" indicates that the corresponding band is selected and added to the band subset, and "0" indicates that the corresponding band is removed.

[0066] Step S202: Determine the structured search operator based on the prior knowledge of spectral physical continuity.

[0067] The structured search operator is used to preserve the entire band and perform boundary variations on continuous bands.

[0068] Specifically, the spectrophysical continuity prior indicates that the absorption peaks of chemical bonds in a spectrum often exhibit a continuous distribution. Random single-point mutations in traditional genetic algorithms easily disrupt this continuous absorption peak structure. Therefore, it is necessary to design structured search operators based on the spectrophysical continuity prior. (See [reference needed]). Figure 13 , Figure 13 The diagram illustrates the operation of three structured operators provided in embodiments of this application. For example... Figure 13As shown, during the search process, the intervals of consecutive "1" in the binary mask vector are treated as a whole gene block for genetic operation to avoid random cutting within the interval. At the same time, mutation operation is only performed on the edge bands of the feature band intervals of consecutive "1", i.e., boundary mutation.

[0069] Step S203: Based on structured search operators and traditional genetic operators, a hybrid search mechanism is formed.

[0070] Specifically, the aforementioned hybrid search mechanism refers to selectively using structured search operators or traditional genetic operators, such as single-point crossover and basic position mutation, to manipulate individuals in the population during each generation of the task-specialized genetic algorithm, according to certain rules or probabilities. This retains the advantages of traditional genetic algorithms, such as strong global search capabilities and ease of generating new gene patterns, while incorporating the ability of structured search operators to preserve physically continuous characteristics.

[0071] Step S204: Construct a task-specific genetic method based on fitness function, encoding method and hybrid search mechanism.

[0072] Specifically, the fitness function and binary mask encoding method based on the performance index of the regression model validation set, as well as the hybrid search mechanism, are integrated to construct a genetic algorithm specializing in spectral band selection tasks. This task-specialized genetic algorithm is no longer a general optimization tool, but a dedicated algorithm that deeply integrates the continuity of spectral physics and the requirements of regression modeling, ensuring that the search direction conforms to the laws of spectral physics.

[0073] This embodiment achieves a direct correlation between band selection and model prediction performance, as well as effective protection of the continuous physical structure of absorption peaks, by using techniques such as determining the fitness function and binary encoding based on regression indices, and determining structured operators for whole-segment retention and boundary variation based on prior knowledge to form a hybrid search mechanism. This makes the search target of the task-specific genetic algorithm highly consistent with the needs of the actual regression task, avoids the destruction of the continuity of feature bands by random operations, and improves the physical interpretability and search accuracy of the screening results.

[0074] In some embodiments, see Figure 3 and Figure 13 The determination of the structured search operator based on the prior of spectral physical continuity includes: Step S301: Based on the prior knowledge of spectral physical continuity, determine the requirements for locating the absorption peak center, reducing redundant noise, and expanding the broad peak characteristics.

[0075] Specifically, in spectral analysis, different spectral features require different optimization strategies: to accurately capture the strongest response position of the absorption peak, there is a need to locate the center of the absorption peak; to eliminate invalid background interference and redundant bands, there is a need to reduce redundant noise; and to cover the overall features of a wider absorption band, there is a need to expand the broad peak features. Based on the prior of the continuity of spectral physics, these three needs correspond to different search operation logics.

[0076] Step S302: Determine the translation operator based on the positioning requirements of the absorption peak center.

[0077] The translation operator is used to shift the currently selected subset of bands one position to the left or right.

[0078] Specifically, when the AI ​​inference or search process finds that the currently selected subset of bands covers a certain absorption peak but may be deviating from the peak extreme point, the translation operator can shift the entire interval of consecutive "1"s in the binary mask vector to the left or right by one band position, thereby fine-tuning the positioning of the absorption peak and making it more accurately aligned with the center position of the strongest physical absorption.

[0079] Step S303: Determine the pruning operator based on the simplification requirements of redundant noise.

[0080] The pruning operator is used to randomly remove the currently selected band with a preset probability.

[0081] Specifically, the selected band subset may contain some redundant bands that contribute little to the model or even introduce noise. The pruning operator randomly sets the mask bits corresponding to certain selected bands from "1" to "0" with a certain probability, thereby simplifying and removing bands, which helps improve the model's noise resistance and simplify its structure.

[0082] Step S304: Based on the expansion requirements of the broad peak feature, determine the growth operator.

[0083] The growth operator is used to activate the neighboring bands of the selected band subset.

[0084] Specifically, some chemical components have broad absorption peaks. If the currently selected subset of bands cannot completely cover the width of the entire absorption peak, feature information will be lost. The growth operator expands the feature interval by activating the unselected bands adjacent to the boundary of the selected subset of bands (i.e., the neighborhood bits with a mask of "0") to "1", thus ensuring the complete capture of broad peak features.

[0085] Step S305: Determine the structured search operator based on the translation operator, pruning operator, and growth operator.

[0086] Specifically, the translation, pruning, and growth operators designed for different physical requirements are combined to form a structured search operator set. During the evolution of the task-specialized genetic algorithm, these operators are selected and invoked based on the actual situation and probability of the population, thereby achieving refined operations on spectral features for different physical requirements.

[0087] This embodiment clarifies the requirements for absorption peak location, redundancy reduction, and broad peak expansion based on the prior knowledge of spectral physical continuity, and designs corresponding translation operators, pruning operators, and growth operators to achieve refined search operations for different physical requirements of spectral features. The translation operator locates the center of the absorption peak, the pruning operator removes redundant noise, and the growth operator captures the complete features of the broad peak, avoiding the destruction of physical structure by traditional random operators and improving the physical interpretability and search accuracy of band selection.

[0088] In some embodiments, see Figure 4 The reinforcement learning meta-control framework, based on population evolution state representation and fitness improvement feedback, is constructed to coordinate the hyperparameters and structured operators of the task-specific genetic algorithm. This framework includes: Step S401: Determine the state space for reinforcement learning based on the population evolution state representation.

[0089] The state space is used to perceive the progress of evolution and population diversity.

[0090] Specifically, to enable reinforcement learning agents to accurately perceive the current operational state of task-specific genetic algorithms, a state space needs to be constructed. This state space consists of population evolutionary state representations, including but not limited to indicators such as population fitness variance, the difference between the optimal individual and the average fitness, and the current generation number. These indicators comprehensively reflect the speed of progress and the level of genetic diversity in the population, providing environmental information for the reinforcement learning agent's decision-making.

[0091] Step S402: Based on the requirements of hyperparameters and structured operators of the collaborative regulation task-specific genetic algorithm, determine the action space of reinforcement learning.

[0092] The action space is used to characterize the control operations on the hyperparameters and structured operators.

[0093] Specifically, the output of the reinforcement learning agent needs to directly influence the task-specific genetic algorithm. Therefore, the action space design must encompass the control of hyperparameters (such as mutation rate and crossover rate) and structured operators (such as the calling probabilities of translation, pruning, and growth operators). By outputting specific action vectors, the reinforcement learning agent can directly intervene in and coordinately control the search behavior of the task-specific genetic algorithm.

[0094] Step S403: Based on the state space and action space, construct the basic interaction mechanism of the reinforcement learning meta-control framework.

[0095] Specifically, based on Markov decision processes, the state space is used as the observation input of the environment, and the action space is used as the decision output of the reinforcement learning agent. A basic interaction loop of "state observation - action output - environment transfer" is constructed, which enables the reinforcement learning agent to continuously interact with the evolutionary process of the task-specialized genetic algorithm.

[0096] Step S404: Determine the reward function for reinforcement learning based on the fitness improvement feedback.

[0097] The reward function employs a reward backfilling mechanism to back-allocate accumulated rewards when the fitness improvement of the optimal individual exceeds a preset fitness improvement threshold.

[0098] Specifically, in genetic algorithms, fitness improvement is often delayed; a good parameter adjustment at a particular step may only show its effect after several generations. If rewards are designed solely based on immediate feedback, reinforcement learning agents struggle to learn long-term effective strategies. Therefore, this reward function employs a reward backfilling mechanism. When the fitness improvement of the optimal individual exceeds a preset threshold, not only is a reward given for the current step, but the accumulated rewards are also backfilled to historical decisions, thus addressing the problem of delayed rewards.

[0099] Step S405: Construct a reinforcement learning meta-control framework based on the reward function and basic interaction mechanism.

[0100] Specifically, the reward function is combined with the basic interaction mechanism to form a complete reinforcement learning training loop. The reinforcement learning agent interacts with the environment through the basic interaction mechanism and updates the policy network based on the signals fed back by the reward function, gradually learning the regulation policy that can maximize long-term cumulative rewards, thus constructing a reinforcement learning meta-control framework.

[0101] This embodiment achieves comprehensive perception and effective handling of delayed feedback of the evolutionary state of the task-specialized genetic algorithm by constructing a state space for perceiving the evolutionary state, an action space for representing the regulatory operation, a basic interaction mechanism, and a reward function with a reward backfilling mechanism. It provides a principled feedback mechanism for reinforcement learning agents, solves the credit allocation problem caused by the delay in fitness improvement, and enables reinforcement learning agents to learn more accurate and long-term effective adaptive regulation strategies.

[0102] In some embodiments, see Figure 5 The step of determining the state space for reinforcement learning based on the population evolutionary state representation includes: Step S501: Determine the scale-invariant state vector based on the population evolution state representation.

[0103] The state vector integrates three major categories of indicators: evolutionary progress characteristics, fitness distribution patterns, diversity, and quality indicators.

[0104] Specifically, to eliminate numerical scale differences caused by different datasets or evolutionary stages and improve the generalization ability of reinforcement learning agents, the state space is represented by a scale-invariant state vector. This vector integrates three major categories of indicators: evolutionary progress features (such as the ratio of the current generation to the maximum generation), fitness distribution patterns (such as fitness skewness and kurtosis), and diversity and quality indicators (such as the population gene diversity index and the current best fitness), thereby comprehensively and scale-independently characterizing the population state.

[0105] Step S502: The fitness-related indices in the state vector are normalized using interquartile range to obtain the normalized state vector.

[0106] Specifically, fitness values ​​vary greatly across different tasks, and direct input can lead to difficulties in training neural networks. This step employs interquartile range normalization, which uses the interquartile range of the fitness values ​​instead of the range for scaling. This effectively resists the interference of extreme and abnormal fitness values, making the normalized state vector more robust and stable.

[0107] Step S503: Determine the state space for reinforcement learning based on the normalized state vector.

[0108] Specifically, the state vector after interquartile range normalization is used as the final representation of the reinforcement learning state space and input to the reinforcement learning agent, which ensures that the state features remain consistent in optimization problems of different scales and improves the training stability and convergence speed of the policy network.

[0109] This embodiment constructs a scale-invariant state vector by fusing three major categories of indicators and employs interquartile range normalization to achieve an unbiased and robust representation of the population's evolutionary state. This eliminates the numerical scale differences between different datasets and evolutionary stages, effectively resists the interference of extreme outliers, and improves the generalization ability and training stability of the reinforcement learning agent on different spectral band selection tasks.

[0110] In some embodiments, see Figure 6 The determination of the action space for reinforcement learning based on the requirements of synergistic regulation of the hyperparameters and structured operators of the task-specialized genetic algorithm includes: Step S601: Determine the synergistic regulation requirements based on the coupling characteristics of hyperparameters and structured operators in the task-specific genetic algorithm.

[0111] Specifically, the hyperparameters (such as mutation rate) and structured operators (such as growth operator probabilities) of task-specific genetic algorithms are not independent but strongly coupled. For example, when the mutation rate is high, the strength of structured operators should be reduced to avoid destroying superior individuals. Therefore, it is necessary to determine the requirements for coordinated regulation based on this coupling characteristic, that is, the design of the action space should be able to adjust these two types of parameters simultaneously and in a coordinated manner.

[0112] Step S602: Based on the need for coordinated regulation, determine the multi-head independent control architecture.

[0113] Specifically, to achieve coordinated and flexible control, the action space adopts a multi-head independent control architecture. This architecture, based on a shared underlying feature extraction network, separates multiple independent output heads, each responsible for controlling a specific type of parameter or operator. This ensures information sharing among parameters while achieving decoupled control at the output end. Complex hyperparameter combinations are discretized into seven fixed parameter combinations with clear semantics, i.e., action presets. Each action is coupled with the basic evolutionary parameters of the genetic algorithm and structured operators. The reinforcement learning agent selects one action at each control stage to drive the genetic algorithm's evolution.

[0114] Step S603: Based on the multi-head independent control architecture, determine the multi-head independent action space.

[0115] The multi-head independent action space includes a mutation rate head, a diversity rate head, a selection pressure head, a structured operator head, and a structured strength head.

[0116] Specifically, the multi-head independent action space is subdivided into five heads: the mutation rate head controls the probability of traditional mutations; the diversity rate head controls the intensity of introducing new genes; the selection pressure head controls the strength of survival of the fittest; the structured operator head determines the probability of calling translation, pruning, and growth operators; and the structured intensity head controls the magnitude of structured operations. This fine division enables reinforcement learning agents to collaboratively intervene in the evolutionary process from multiple dimensions. The mutation rate head offers seven fine-tuning values ​​from 0.001 to 0.0015, adapting to different needs of local fine-tuning and global exploration; the diversity rate head offers seven control options to adjust the proportion of random perturbations introduced into the population, maintaining genetic diversity; the selection pressure head offers seven discrete levels, ranging from 2 to 8 for tournament scales, gradient-adjusting selection pressure to balance exploration and convergence; the structured operator head includes four modes: no operation, pruning, growth, and translation, aligning with spectral physical characteristics; and the structured intensity head offers six intensity levels to precisely control the magnitude of structured operator action, avoiding over-adjustment or under-adjustment.

[0117] Step S604: Use the multi-head independent action space as the action space for reinforcement learning.

[0118] Specifically, the aforementioned multi-head independent action space is used as the final output dimension of the reinforcement learning agent. The reinforcement learning agent outputs a multi-dimensional action vector at each decision step, corresponding to the control values ​​of the five heads, thereby achieving comprehensive and fine-grained coordinated control of the task-specialized genetic algorithm.

[0119] This embodiment achieves decoupling and collaborative control of task-specific genetic algorithm hyperparameters and structured operators by determining collaborative control requirements based on coupling characteristics and designing a multi-head independent control architecture and a multi-head independent action space. Its beneficial effects are: it ensures information sharing and coordination among control parameters, achieves fine-grained independent adjustment, avoids the dimensionality curse caused by a single action space, and improves the flexibility and optimization efficiency of reinforcement learning agent control.

[0120] In some embodiments, see Figure 7 The reward function for reinforcement learning is determined based on fitness improvement feedback. This reward function employs a reward backfilling mechanism to retroactively allocate accumulated rewards when the fitness improvement of the optimal individual exceeds a preset fitness improvement threshold. This includes: Step S701: Determine the reward feedback mechanism based on the fitness improvement feedback.

[0121] Specifically, due to the delayed evolutionary effect of task-specific genetic algorithms—meaning that current excellent decisions may only lead to significant improvements in fitness in future generations—a reward backfilling mechanism is needed to accurately evaluate the contribution of historical decisions. This mechanism retroactively allocates rewards for positive outcomes to previous key decision-making steps.

[0122] Step S702: Generate cumulative rewards by obtaining the cumulative improvement of the fitness of the current best individual relative to the previous improvement.

[0123] Specifically, when a new optimal fitness is detected, the difference between it and the previously recorded optimal fitness is calculated, which is the cumulative improvement. The larger this improvement is, the more effective the recent series of regulatory decisions have been, and a cumulative reward of a corresponding scale is generated accordingly.

[0124] Step S703: Determine historical decisions by recording the decision time steps since the last fitness improvement.

[0125] Specifically, all decision time steps between the last optimal fitness improvement and the current improvement are recorded. These steps constitute the historical decision set, which together contribute to the current fitness improvement.

[0126] Step S704: When the fitness improvement is detected to exceed the preset fitness improvement threshold, the accumulated reward is retroactively allocated to the historical decision.

[0127] Specifically, a breakthrough is considered valid only when the increase in fitness exceeds a preset threshold, rather than a random fluctuation. At this point, the generated cumulative reward is back-allocated to all historical decisions since the last improvement with a certain decay weight, ensuring that early decisions that contributed to the breakthrough receive the positive feedback they deserve.

[0128] In one specific implementation, when the model performance of the optimal solution of GA is effectively improved (i.e. R²>10 -4 When the cumulative reward is redistributed back to all GA evolution microsteps since the last performance boost, the cumulative reward will be redistributed back to the last microstep. i The reward value for each microstep is defined as:

[0129] in, α This is the reward scaling factor, with a default value of 1.0, used to adjust the magnitude of the reward signal; ΔR² represents the cumulative improvement relative to the most recent performance boost. best, current ΔR² represents the change in R² of the current best individual (i.e., the difference between the current best R² value and the previous best R² value, reflecting the performance improvement of the current best individual); test, last - improve This represents the change in R² since the last improvement on the test set (i.e., the difference between the current R² value on the test set and the previous R² value on the test set, reflecting the performance improvement of the test set), and K is the total number of microsteps in the GA evolution since the last performance improvement. This formula is used to calculate the cumulative improvement as the basis for reward allocation.

[0130] Step S705: Determine the penalty for inhibiting decision-making stagnation.

[0131] Specifically, to prevent reinforcement learning agents from adopting conservative strategies that lead to long-term evolutionary stagnation, a penalty term β(i) is added to the reward function to suppress the agent's decision-making stagnation behavior. When the fitness does not improve or the population diversity is too low for several consecutive decision cycles, a negative penalty signal is given, forcing the reinforcement learning agent to try new regulatory strategies.

[0132] The expression for the penalty term β(i) is: .

[0133] Step S706: Construct a reward function for reinforcement learning based on the cumulative reward and penalty items allocated by the reward backfilling mechanism.

[0134] Specifically, the cumulative reward from backfill allocation is combined with a penalty term to suppress stagnation to construct the final reward function. This reward function effectively incentivizes exploratory behavior that leads to long-term gains while avoiding evolutionary stagnation, guiding the reinforcement learning agent to continuously drive the task-specialized genetic algorithm toward the global optimum.

[0135] This embodiment designs a reward backfilling mechanism to retroactively allocate accumulated rewards to historical decisions and introduces a penalty term to suppress decision stagnation, thereby achieving reasonable credit allocation for delayed feedback and active suppression of conservative strategies. Its beneficial effects are: it effectively solves the credit allocation problem caused by delayed rewards in the evolution of task-specialized genetic algorithms, incentivizes reinforcement learning agents to make regulatory decisions with long-term benefits, avoids evolutionary stagnation, and enhances the ability of reinforcement learning agents to break through local optima.

[0136] In some embodiments, see Figure 8 The construction of a temporally enhanced reinforcement learning agent based on a recurrent neural network within the reinforcement learning meta-control framework includes: Step S801: Construct a historical state sequence based on the state space of the reinforcement learning meta-control framework.

[0137] Specifically, in order to capture the temporal dependencies of evolution, the reinforcement learning agent not only receives the current state vector but also constructs a sequence of historical states. The current state is combined with the state vectors of the previous N consecutive control time steps in chronological order to form a state sequence reflecting the recent evolutionary trajectory, which serves as the input to the recurrent neural network.

[0138] Step S802: Determine that the recurrent neural network is a two-layer gated recurrent unit.

[0139] Specifically, compared to traditional RNNs, Gated Recurrent Units (GRUs) effectively alleviate the gradient vanishing problem in long sequence training through a gating mechanism, and their computational efficiency is higher than that of LSTMs. A two-layer GRU structure is adopted, with 64 hidden layers. The first layer extracts short-term local temporal features, and the second layer, based on this, extracts more macroscopic long-term evolutionary trends, thereby enhancing the network's ability to fit complex temporal patterns. The Actor network and the Critic network share this two-layer GRU feature extraction layer.

[0140] Step S803: Construct a temporal feature extraction module based on a two-layer gated loop unit.

[0141] The temporal feature extraction module extracts features from the historical state sequence by updating the hidden state memory historical evolution information, thereby obtaining a temporally enhanced state representation.

[0142] Specifically, the historical state sequence is input into a temporal feature extraction module composed of a two-layer GRU. The GRU controls the flow of information through update and reset gates, updating the hidden state at each time step, compressing and storing the historical evolution information in the hidden state. The output of the hidden state in the final step is the temporal enhanced state representation that integrates the information of the entire historical sequence.

[0143] Step S804: Based on the temporally enhanced state representation, construct a temporally enhanced reinforcement learning agent under the reinforcement learning meta-control framework.

[0144] Specifically, the original single-step state vector is replaced by a temporally enhanced state representation as input to the reinforcement learning policy network. Based on this representation, which contains the historical evolutionary trajectory, the reinforcement learning policy network outputs the regulatory action at the current moment, thereby constructing a temporally enhanced reinforcement learning agent with temporal awareness and memory capabilities.

[0145] The training of reinforcement learning agents follows an iterative process of "task generation - trajectory acquisition - model optimization": New spectral simulation tasks are generated from a physically heuristic simulation environment, and a task-specific genetic algorithm is initiated for band search; the reinforcement learning agent selects the optimal action at each control time step, driving the task-specific genetic algorithm to complete multi-step evolution until a preset termination condition is reached, forming a complete decision sequence; decision trajectories, such as state, action, and reward, are collected, and the agent network is iteratively optimized using a PPO (Proximal Policy Optimization) pruning objective function. Specifically, the training parameters can be set as follows: the RL-GA framework in the band selection method is implemented based on the PyTorch deep learning framework, with a total of 3000 training rounds, 4 optimization iterations per round, and a learning rate of 1×10⁻⁶. - ³, clipping factor 0.2, entropy regularization factor 0.03.

[0146] This embodiment achieves deep capture of the temporal dependencies in the evolutionary process of a task-specific genetic algorithm by constructing a historical state sequence and using a two-layer gated recurrent unit to extract temporal features to obtain temporally enhanced state representations. Its beneficial effects are: enabling reinforcement learning agents to remember historical evolutionary trajectories and understand the dynamic changes in population states, thereby making more forward-looking and global regulatory decisions and avoiding short-sighted behavior caused by relying solely on the current state.

[0147] In some embodiments, see Figure 9 The generation of simulated training data based on spectral physical interference mechanisms and spectral feature variation patterns includes: Step S901: Based on the spectral physical interference mechanism, generate a random walk sequence and perform one-dimensional Gaussian smoothing to determine the baseline spectrum of the simulated baseline drift and background fluctuation.

[0148] Specifically, real spectra are often affected by instrument hardware and environmental factors, resulting in baseline drift and background fluctuations. This step generates a random walk sequence based on a physical interference mechanism to simulate low-frequency drift trends, and processes it with a one-dimensional Gaussian smoothing filter to make it closer to the real smooth baseline drift characteristics, thereby determining the simulated baseline spectrum.

[0149] Step S902: Based on the variation law of spectral characteristics, multiple Gaussian absorption peaks with random center positions, widths and peak heights are superimposed on the baseline spectrum to determine the synthetic spectrum simulating the absorption characteristics of chemical components.

[0150] Specifically, the absorption peaks of chemical components exhibit variability in morphology and position. Based on the characteristic variation law, this step randomly superimposes multiple Gaussian absorption peaks onto the baseline spectrum. The center wavelength, half width at half maximum (FWHM), and peak height of each absorption peak are randomly generated according to a certain distribution to simulate the absorption characteristics of different components at different concentrations and states, thus forming a synthetic spectrum.

[0151] Step S903: Based on the spectral physical interference mechanism, record the mean of the synthesized spectrum as the original scattering intensity, and perform standard normalization on the synthesized spectrum to determine the normalized simulated spectrum.

[0152] Specifically, scattering effects are a common physical interference in spectral acquisition, usually related to particle size. This step records the mean of the synthesized spectrum without scattering interference as the original scattering intensity. Subsequently, the synthesized spectrum is normalized using Standard Normal Variate (SNV), a commonly used spectral data preprocessing method to eliminate the scattering effects caused by particle size and optical path differences, to simulate the spectral shape after eliminating scattering effects and determine the normalized simulated spectrum.

[0153] Step S904: Construct stoichiometric relationships based on real spectral information bands, introduce multiplicative bias and Gaussian measurement noise, and generate regression target values ​​corresponding to the normalized simulated spectra.

[0154] Specifically, in order to give the simulated data a target label for supervised learning, a stoichiometric relationship model is constructed based on the physicochemical properties of the real spectra. Multiplicative bias and measurement noise following a Gaussian distribution are introduced to calculate the concentration or property target value corresponding to the normalized simulated spectrum, ensuring the logical consistency between the input spectrum and the output target.

[0155] Step S905: Based on the normalized simulated spectrum and the regression target value, construct the simulated training data.

[0156] Specifically, the generated normalized simulated spectrum is used as the input feature, and the corresponding regression target value is used as the label. The two are paired to form a complete simulated training data sample, which is used for large-scale interactive training of reinforcement learning agents in the simulated environment.

[0157] This embodiment achieves highly realistic physical simulation of spectral data by simulating baseline drift and background fluctuations, superimposing random Gaussian absorption peaks, performing SNV normalization, and introducing noise to generate regression target values ​​by constructing stoichiometric relationships. It can generate high-quality training data covering a variety of interferences and feature variations on a large scale, effectively solving the overfitting problem of reinforcement learning agents caused by the scarcity of real spectral samples, and improving the generalization ability of reinforcement learning agents from simulated environments to real data scenarios.

[0158] In some embodiments, see Figure 10 The step of using the trained reinforcement learning agent as the meta-controller of the task-specialized genetic algorithm, outputting regulatory actions to drive the task-specialized genetic algorithm to run, and applying it to real spectral data to complete band selection, includes: Step S1001: Obtain the real visible-near infrared spectral data to be selected.

[0159] Samples of agricultural products, ores, or pharmaceuticals to be analyzed are illuminated using a visible-near-infrared spectrometer to collect their reflected or transmitted spectral signals, thereby obtaining absorbance or reflectance data of diffuse reflectance spectra covering the visible and near-infrared regions. For example, an iron ore dataset from Anshan, containing 125 ore samples with spectral bands covering the Vis-NIR range, can be obtained, using total iron content (TFe) as the regression target; a mango dataset, using dry matter content (DMC, an internal quality indicator for agricultural products) as the target; a pill dataset, using active ingredient content as the target; and a diesel dataset, using cetane number (CN, a diesel combustion performance indicator) as the target. A unified preprocessing workflow is applied to these datasets: outlier samples are removed using the Z-score method, followed by standard normal variable (SNV) transformation to eliminate baseline drift and light scattering interference. The actual visible-near-infrared spectral data is then input into an edge computing device or server deployed with a task-specific genetic algorithm and a trained reinforcement learning agent.

[0160] Step S1002: Input real visible-near-infrared spectral data into the task-specific genetic algorithm, and use the trained reinforcement learning agent as the meta-controller of the task-specific genetic algorithm.

[0161] Specifically, the population of the task-specialized genetic algorithm is initialized; in this embodiment, the population size is set to 20. Simultaneously, the trained reinforcement learning agent model is loaded. The trained reinforcement learning agent is set as a meta-controller, putting it in a ready state to output control commands based on the real-time running status of the task-specialized genetic algorithm.

[0162] Step S1003: During the evolution of the task-specialized genetic algorithm, the trained reinforcement learning agent obtains the current population state vector at each control time step.

[0163] Specifically, during the generational evolution of the task-specialized genetic algorithm, the number of iterations was uniformly set to 200 generations. The trained reinforcement learning agent did not passively wait; instead, every preset control time step, such as every generation or every 5 generations, it automatically extracted the evolutionary state representation of the current population, such as fitness variance and optimal individual fitness, to construct the current state vector as the basis for decision-making. All experiments were repeated under 10 different random seeds to ensure the statistical reliability of the results.

[0164] Step S1004: Based on the state vector, output the control action to control the hyperparameters and structured operators of the task-specific genetic algorithm.

[0165] Specifically, after receiving the state vector, the trained reinforcement learning agent extracts temporal features through its internal two-layer GRU network, maps them through the policy network, and outputs a multi-dimensional action vector. This action vector directly acts on the task-specific genetic algorithm, dynamically adjusting the mutation rate, selection pressure, and the probability and intensity of calling structured operators such as translation, pruning, and growth in the next stage, thus achieving closed-loop control.

[0166] Step S1005: Driven by the control action, the task-specialized genetic algorithm iterates and evolves until the preset termination condition is reached, outputting the optimal band subset and completing the precise selection of bands.

[0167] Specifically, under the continuous and dynamic regulation of the trained reinforcement learning agent, the task-specialized genetic algorithm can adaptively balance exploration and development, effectively avoiding premature convergence. When the preset termination conditions such as reaching the maximum number of generations or the fitness no longer improving over a long period are met, the task-specialized genetic algorithm stops evolving, extracts the individual with the highest fitness in the population, and decodes its binary mask vector into the corresponding band combination, which is the final optimal band subset selected, thus completing the accurate band selection of real spectral data.

[0168] This embodiment uses a trained reinforcement learning agent as a meta-controller to acquire the state in real time and output regulatory actions to drive iterative evolution in the task-specific genetic algorithm evolution of real spectral data. This achieves deep collaboration between the trained reinforcement learning agent and the task-specific genetic algorithm in real reasoning scenarios. Its beneficial effects are: it can dynamically adjust the search strategy according to the real-time evolution state of real data, adaptively avoid premature convergence traps, and thus efficiently and accurately search for the optimal band subset with the most physical meaning and predictive performance, completing automated band selection without human intervention.

[0169] In one specific implementation, this embodiment uses four types of Vis-NIR spectral datasets—iron ore, mango, pharmaceutical tablets, and diesel fuel—from Anshan as the verification objects. The specific implementation steps can be summarized as follows: 1. Dataset preprocessing A unified preprocessing workflow was adopted for the four types of datasets: outlier samples were removed by Z-score method, and then standard normal variable (SNV) transformation was performed to eliminate baseline drift and light scattering interference. Among them, the iron ore dataset from Anshan contains 125 ore samples with spectral bands covering the Vis-NIR range, and the total iron content (TFe) is used as the regression target; the mango dataset uses dry matter content (DMC) as the target, the pill dataset uses active ingredient content as the target, and the diesel dataset uses cetane number (CN) as the target.

[0170] 2. Experimental parameter settings The RL-GA framework in the band selection method is implemented based on the PyTorch deep learning framework. The RL agent adopts a two-layer GRU structure with a hidden layer dimension of 64. The Actor network and Critic network share the feature extraction layer. The number of GA iterations is uniformly set to 200 generations, and the population size is set to 20. The training parameters of the RL agent are: a total of 3000 training rounds, 4 optimization iterations per round, and a learning rate of 1×10. - ³, pruning factor 0.2, entropy regularization factor 0.03; all experiments were repeated under 10 different random seeds to ensure the statistical reliability of the results.

[0171] 3. Algorithm Execution Process (1) Start the physical heuristic synthetic spectrum generator to simulate interferences such as baseline drift, absorption peak changes, and scattering effects, and generate large-scale simulated training data for training RL agents; (2) Training PPO agent: Input the simulated spectral task into GA, the agent selects actions according to the population state, drives GA evolution, collects decision trajectories and optimizes the agent network until the agent's decision performance converges; (3) Selection of bands in real dataset: The trained agent is combined with GA, and the preprocessed real spectral data is input. The agent dynamically adjusts the hyperparameters and structured operators of GA to complete the selection of band subsets. (4) Model validation: Based on the selected band subset, the PLSR model is constructed and compared with the traditional SPA, CARS and fixed parameter GA algorithms using the test set R² as the evaluation index.

[0172] 4. Implementation Results Experimental results show that the band selection method proposed in this invention outperforms the control method on the test sets of four types of datasets: the R² of the iron ore dataset in Anshan reaches 0.8448, which is significantly higher than that of SPA (0.5645) and CARS (0.5190); the R² of the mango dataset reaches 0.8916; the R² of the pill dataset reaches 0.9415; and the R² of the diesel dataset reaches 0.6036. At the same time, the standard deviation of RL-GA is significantly lower than that of the fixed parameter GA, and the stability is improved by more than 90%, which verifies the superiority and practicality of the proposed method.

[0173] To more intuitively demonstrate the convergence performance and stability advantages of the method proposed in this application, please refer to... Figure 14 , Figure 14 This diagram illustrates a comparison of the convergence accuracy of RL-GA and Random GA provided in an embodiment of this application. Figure 14 As shown, RL-GA (Reinforcement Learning-based Genetic Algorithm) and Random GA (Random Genetic Algorithm) are used. The colored regions represent the standard deviation range of 10 random seeds. It can be clearly seen that the band selection method converges faster and has higher final accuracy.

[0174] Further, see Figure 15 , Figure 15 The diagram shows a comparison of the R² distributions of the RL-GA, Random GA, SPA (Successive Projections Algorithm), and CARS (Competitive Adaptive Reweighted Sampling) algorithms provided in this application across four datasets. Figure 15 As shown, the box plots for each dataset demonstrate the distribution of results with effective random seeds. The median and overall distribution of the band selection method in this application are significantly higher than those of other comparative methods, and the box plots are more compact, confirming its high accuracy and high stability.

[0175] In addition, see Figure 16, Figure 16 The diagram illustrates the action probability evolution curves of the PPO agent provided in this application across four datasets. For example... Figure 16 As shown, each column corresponds to one of five control heads: mutation rate, diversity rate, selection pressure, structured operator type, and structure strength, while each row corresponds to a different dataset. This evolutionary curve demonstrates that the meta-controller can make adaptive and targeted regulatory decisions based on the population state, rather than using rigid, fixed parameters.

[0176] In another specific implementation, the method of the present invention was applied to a wheat spectral dataset with 26 bands, using protein content as the regression target and an environmental water organic pollutant spectral dataset. The band selection steps were repeated, and the results showed that the PLSR model constructed from the band subset selected by the present method achieved R² values ​​of 0.8723 and 0.7856, respectively, both of which were superior to traditional methods. Furthermore, the method demonstrated excellent stability and generalization ability in spectral tasks with different complexities and noise levels, proving that the present method can be adapted to Vis-NIR spectral detection scenarios in multiple fields.

[0177] Figure 11 This is a schematic diagram of a band selection system provided in an embodiment of this application. See also... Figure 11 The band selection system 11 includes: The genetic algorithm construction module 1101 constructs a task-specific genetic algorithm based on the spectral physical continuity prior and the performance index of the regression model validation set. The task-specific genetic algorithm represents the band subset with a binary mask vector, uses the performance index as the fitness function, and adopts a hybrid search mechanism consisting of a structured search operator based on the continuity prior and a traditional genetic operator. The meta-control framework construction module 1102 constructs a reinforcement learning meta-control framework based on population evolution state representation and fitness improvement feedback. The reinforcement learning meta-control framework is used to coordinate the regulation of the hyperparameters and structured operators of the task-specific genetic algorithm. The agent construction module 1103, based on a recurrent neural network, constructs a temporally enhanced reinforcement learning agent under the reinforcement learning meta-control framework; The data generation and training module 1104 generates simulated training data based on the spectral physical interference mechanism and the variation law of spectral features, which is used to iteratively train the reinforcement learning agent. The band selection module 1105 uses the trained reinforcement learning agent as the meta-controller of the task-specialized genetic algorithm, outputs control actions to drive the task-specialized genetic algorithm to run, and applies it to real spectral data to complete band selection.

[0178] It should be noted that the band selection system provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the band selection system and the band selection method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the band selection method embodiments described above, which will not be repeated here.

[0179] This application also provides an electronic device. Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0180] Typically, electronic device 12 includes one or more processors 1201 and one or more memories 1202.

[0181] Processor 1201 may include one or more processing cores, such as a quad-core processor, a hexa-core processor, etc. Processor 1201 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1201 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1201 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1201 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0182] The memory 1202 may include one or more computer-readable storage media, which may be non-transitory. The memory 1202 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices.

[0183] In some embodiments, the non-transitory computer-readable storage medium in memory 1202 is used to store at least one computer program, which is executed by processor 1201 to implement the band selection method provided in the method embodiments of this application.

[0184] Those skilled in the art will understand that Figure 12 The structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0185] In addition, the device provided in the embodiments of this application may specifically be a chip, component or module. The chip may include a connected processor and a memory. The memory is used to store instructions. When the processor calls and executes the instructions, the chip can execute a band selection method provided in the above embodiments.

[0186] This embodiment also provides a computer-readable storage medium storing computer program code. When the computer program code is run on a computer, the computer executes the above-described related method steps to implement the depot selection method provided in the above embodiment.

[0187] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement a segment selection method provided in the above embodiment.

[0188] In this embodiment, the device, computer-readable storage medium, computer program product, or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.

[0189] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0190] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0191] The above description is only a specific implementation of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application.

Claims

1. A band selection method, characterized in that, The method includes: Based on the spectral physical continuity prior and the performance index of the regression model validation set, a task-specific genetic algorithm is constructed. The task-specific genetic algorithm represents the band subset with a binary mask vector, uses the performance index as the fitness function, and adopts a hybrid search mechanism consisting of a structured search operator based on the continuity prior and a traditional genetic operator. Based on population evolution state representation and fitness improvement feedback, a reinforcement learning meta-control framework is constructed. The reinforcement learning meta-control framework is used to coordinate the hyperparameters and structured operators of the task-specific genetic algorithm. Based on recurrent neural networks, a temporally enhanced reinforcement learning agent is constructed within the aforementioned reinforcement learning meta-control framework. Based on the spectral physical interference mechanism and the variation law of spectral features, simulated training data is generated to iteratively train the reinforcement learning agent. The trained reinforcement learning agent is used as the meta-controller of the task-specialized genetic algorithm, and outputs control actions to drive the task-specialized genetic algorithm to run, which is applied to real spectral data to complete band selection. Based on the spectral physical continuity prior and the performance index of the regression model validation set, a task-specific genetic algorithm is constructed. This algorithm represents band subsets using binary mask vectors, uses the performance index as the fitness function, and employs a hybrid search mechanism combining a structured search operator based on continuity priors and traditional genetic operators. Based on the performance metrics of the regression model validation set, the fitness function and encoding method are determined. The encoding method is used to represent a band subset using a binary mask vector. Based on the prior knowledge of spectral physical continuity, a structured search operator is determined. This structured search operator is used to perform whole-band preservation and boundary variation on continuous bands. A hybrid search mechanism is formed based on the structured search operator and the traditional genetic operator; Based on the fitness function, the encoding method, and the hybrid search mechanism, the task-specific genetic algorithm is constructed. The determination of the structured search operator based on the prior of spectral physical continuity includes: Based on the prior knowledge of the continuity of spectral physics, the requirements for locating the absorption peak center, reducing redundant noise, and expanding the broad peak characteristics are determined. Based on the positioning requirements of the absorption peak center, a translation operator is determined. The translation operator is used to shift the currently selected band subset as a whole one position to the left or right. Based on the simplification requirement of the aforementioned redundant noise, a pruning operator is determined, which is used to randomly remove the currently selected band with a preset probability. Based on the expansion requirements of the aforementioned broad peak feature, a growth operator is determined, which is used to activate the neighboring bands of the selected band subset. Based on the translation operator, the pruning operator, and the growth operator, the structured search operator is determined.

2. The band selection method according to claim 1, characterized in that, The reinforcement learning meta-control framework is constructed based on population evolution state representation and fitness improvement feedback. This framework is used to coordinate the hyperparameters and structured operators of the task-specific genetic algorithm, including: Based on the population evolution state representation, a state space for reinforcement learning is determined, which is used to perceive evolutionary progress and population diversity. Based on the requirement of coordinating the hyperparameters and structured operators of the task-specific genetic algorithm, the action space of reinforcement learning is determined, which is used to represent the control operations on the hyperparameters and structured operators. Based on the state space and the action space, the basic interaction mechanism of the reinforcement learning meta-control framework is constructed. Based on the fitness improvement feedback, a reward function for reinforcement learning is determined. The reward function employs a reward backfilling mechanism to back-allocate accumulated rewards when the fitness improvement of the optimal individual exceeds a preset fitness improvement threshold. Based on the reward function and the basic interaction mechanism, the reinforcement learning meta-control framework is constructed.

3. The band selection method according to claim 2, characterized in that, The process of determining the state space for reinforcement learning based on the population evolutionary state representation includes: Based on the population evolutionary state representation, a scale-invariant state vector is determined, which integrates three major categories of indicators: evolutionary progress characteristics, fitness distribution patterns, and diversity and quality indicators. The fitness-related indices in the state vector are normalized using interquartile range (IIR) to obtain the normalized state vector. The state space of the reinforcement learning is determined based on the normalized state vector.

4. The band selection method according to claim 2, characterized in that, The requirement for co-regulating the hyperparameters and structured operators of the task-specialized genetic algorithm determines the action space for reinforcement learning, including: Based on the coupling characteristics of hyperparameters and structured operators in task-specific genetic algorithms, the requirements for coordinated regulation are determined. Based on the aforementioned collaborative regulation requirements, a multi-head independent control architecture is determined; Based on the aforementioned multi-head independent control architecture, a multi-head independent action space is determined, which includes a mutation rate head, a diversity rate head, a selection pressure head, a structured operator head, and a structured strength head. The multi-head independent action space is used as the action space for reinforcement learning.

5. The band selection method according to claim 2, characterized in that, The reward function for reinforcement learning is determined based on fitness improvement feedback. This reward function employs a reward backfilling mechanism to retroactively allocate accumulated rewards when the fitness improvement of the optimal individual exceeds a preset fitness improvement threshold. This includes: Based on the feedback from the improvement in fitness, a reward feedback mechanism was determined; Cumulative rewards are generated by obtaining the cumulative improvement in the fitness of the current best individual relative to the previous improvement. Historical decisions are determined by recording the decision time steps since the last fitness boost; When the fitness improvement is detected to exceed a preset fitness improvement threshold, the accumulated reward is back-allocated to the historical decision. Identify the penalties for suppressing decision-making stagnation; The reward function for reinforcement learning is constructed based on the cumulative reward allocated by the reward feedback mechanism and the penalty term.

6. The band selection method according to claim 1, characterized in that, The method of constructing a temporally enhanced reinforcement learning agent based on a recurrent neural network within the reinforcement learning meta-control framework includes: Based on the state space of the reinforcement learning meta-control framework, a historical state sequence is constructed. The recurrent neural network is determined to be a two-layer gated recurrent unit; Based on the aforementioned two-layer gated recurrent unit, a temporal feature extraction module is constructed. The temporal feature extraction module extracts features from the historical state sequence by updating the hidden state memory historical evolution information, thereby obtaining a temporally enhanced state representation. Based on the aforementioned temporally enhanced state representation, the temporally enhanced reinforcement learning agent is constructed within the reinforcement learning meta-control framework.

7. The band selection method according to claim 1, characterized in that, The simulation training data generated based on the spectral physical interference mechanism and the variation law of spectral features includes: Based on the spectral physical interference mechanism, a random walk sequence is generated and smoothed by one-dimensional Gaussian to determine the baseline spectrum of simulated baseline drift and background fluctuation. Based on the variation law of spectral characteristics, multiple Gaussian absorption peaks with random center positions, widths and peak heights are superimposed on the baseline spectrum to determine the synthetic spectrum simulating the absorption characteristics of chemical components; Based on the spectral physical interference mechanism, the mean value of the synthesized spectrum is recorded as the original scattering intensity, and the synthesized spectrum is normalized by standard normal variables to determine the normalized simulated spectrum. Based on the real spectral information bands, a stoichiometric relationship is constructed, and a multiplicative bias and Gaussian measurement noise are introduced to generate a regression target value corresponding to the normalized simulated spectrum. The simulation training data is constructed based on the normalized simulated spectrum and the regression target value.

8. The band selection method according to claim 1, characterized in that, The step of using the trained reinforcement learning agent as the meta-controller of the task-specialized genetic algorithm, outputting regulatory actions to drive the task-specialized genetic algorithm to run, and applying it to real spectral data to complete band selection includes: Acquire the true visible-near-infrared spectral data to be selected; The real visible-near-infrared spectral data is input into the task-specific genetic algorithm, and the trained reinforcement learning agent is used as the meta-controller of the task-specific genetic algorithm. During the evolution of the task-specialized genetic algorithm, the trained reinforcement learning agent obtains the current population state vector at each control time step. Based on the state vector, output control actions are used to control the hyperparameters and structured operators of the task-specific genetic algorithm; Driven by the aforementioned control action, the task-specialized genetic algorithm iterates and evolves until a preset termination condition is met, outputting the optimal band subset and completing precise band selection.

Citation Information

Patent Citations

  • High spectral image waveband selection method based on binary coded ant colony algorithm with improved differential evolution

    CN107437098A

  • Renewable resource recovery data management system based on Internet of Things

    CN121599657A