Data processing method for high-throughput screening of bimetallic catalysts based on machine learning
By constructing a set of variation amplitudes and reconstructing the feature numerical sequence, the problem of sudden jumps in prediction results during machine learning screening was solved, the stability and reliability of the screening path were achieved, and the accuracy of high-throughput screening of bimetallic catalysts was improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FOSHAN UNIVERSITY
- Filing Date
- 2026-05-05
- Publication Date
- 2026-07-31
AI Technical Summary
In high-throughput screening based on machine learning, candidate systems are prone to entering the vicinity of the training data distribution boundary when dynamically advancing, leading to sudden jumps in prediction results and affecting the stability and reliability of the screening process.
By constructing a set of change amplitudes, abnormal jump segments are identified and reconstructed. The feature value sequence is adjusted to maintain a continuous connection. The sequence is optimized by continuous constraint adjustment and balanced distribution conditions to ensure the continuity and consistency of the prediction results.
This effectively avoids path deviations caused by local fluctuations, improves the overall consistency and reliability of large-scale candidate system evaluation results, and ensures the stable progress of the screening process and the accuracy of the results.
Smart Images

Figure CN122494043A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computational materials science and intelligent catalyst design technology, specifically to a data processing method for high-throughput screening of bimetallic catalysts based on machine learning. Background Technology
[0002] Data processing for high-throughput screening of bimetallic catalysts based on machine learning refers to the systematic organization, feature representation, and mapping modeling of structural information and reaction performance-related data for a large number of candidate bimetallic catalyst systems. Machine learning is then used to achieve rapid performance prediction and screening. Specifically, this process first uses density functional theory calculations as the data source to quantify the adsorption characteristics of key reaction intermediates in specific support environments for different metal combinations, and selects core indicators characterizing catalytic behavior as learning objectives. Subsequently, multidimensional features are constructed from the original structural parameters, transforming information such as metal type, interatomic spacing, and electronic structure characteristics into calculable numerical expressions, and forming efficient descriptive vectors through correlation analysis and dimensionality reduction. Based on this, the nonlinear mapping relationships obtained from training are used to rapidly predict candidate systems that have not yet been calculated, thus completing the performance evaluation and selection of a large number of bimetallic combinations in a very short time. Finally, the screening results are verified through fine calculations, achieving a closed-loop data processing system from a small amount of accurate data to massive, rapid predictions. This method essentially transforms the high-cost process of traditional step-by-step calculations into a data-driven batch inference process, representing a highly efficient processing mechanism for reconstructing material screening paths using data.
[0003] The existing technology has the following shortcomings: In existing technologies, during high-throughput screening based on machine learning, candidate systems are continuously predicted by mapping feature representations to the feature space. When some candidate systems gradually enter the vicinity of the training data distribution boundary or even extend into the uncovered area during dynamic advancement, the model's response in this area is prone to discontinuous changes, causing sudden jumps in the prediction results. Such jump results often appear as abnormally good values in numerical terms, and are easily included directly in the preferred set and participate in the subsequent screening process. This leads to the screening path continuously advancing in a direction that deviates from the true pattern, further causing the prediction results to gradually accumulate biases during continuous screening. Ultimately, this can easily lead to overall distortion of the evaluation results of large-scale candidate systems, seriously affecting the stability and reliability of the screening process.
[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide a data processing method for high-throughput screening of bimetallic catalysts based on machine learning, so as to solve the problems in the background art mentioned above.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a data processing method for high-throughput screening of bimetallic catalysts based on machine learning, comprising the following steps: The feature values and prediction results corresponding to each candidate result in the continuous screening process are obtained. The feature values are uniformly organized, and the adjacent differences are calculated point by point according to the screening sequence. A set of change amplitudes is constructed based on the size of the adjacent differences to characterize the degree of change of the prediction results in the screening sequence. Based on the set of change ranges, continuous intervals with change ranges exceeding a preset threshold are identified, and the feature value distribution of the corresponding screening sequence position is aggregated to form a set of change segments used to characterize the range of abnormal jumps. For the set of changing segments, the feature values at the corresponding screening sequence positions are rearranged. The feature values located in the set of changing segments are discretized and expanded along the screening sequence direction. During the expansion process, the trend direction of the prediction results is kept consistent, and the adjusted feature value sequence is obtained. Based on the adjusted feature value sequence, the candidate results that subsequently enter the screening sequence are continuously constrained and adjusted so that the prediction results corresponding to the subsequent candidate results maintain a continuous connection with the adjusted feature value sequence in the screening sequence. Based on the prediction results after continuous constraint adjustment, the screening sequence is reconstructed as a whole so that the prediction results at each position in the screening sequence meet the preset balanced distribution conditions in terms of the magnitude of change, and form a corresponding constraint relationship with the set of change segments.
[0007] Preferably, focusing on the ordered representation of changes in prediction results during continuous screening, a data structure describing the degree of change in prediction results is formed by serializing and organizing candidate results and constructing a set of change magnitudes. The specific steps are as follows: Each candidate result in the continuous screening process is numbered and the corresponding feature values and prediction results are extracted. They are arranged in the screening order to form a feature value sequence and a prediction result sequence. At the same time, a position identifier is established to achieve a one-to-one correspondence between the two. By combining the prediction result sequence, the difference between adjacent positions is calculated to form a difference sequence and the corresponding candidate result number and feature value position are recorded to obtain difference data with positional attributes and feature correlation. The difference sequence is summarized and organized and classified according to the numerical value to form a set of change ranges. At the same time, the candidate result number and feature value position corresponding to the difference are retained to construct a mapping relationship. By mapping the difference categories in the set of change magnitudes to the positions of the selected sequences and then to the feature value sequences, the overall expression of the degree of change in the prediction results is completed, forming a continuous and consistent data structure for subsequent processing.
[0008] Preferably, each difference in the difference sequence has a unique correspondence with the position of the screening sequence, the set of change magnitudes is divided into intervals according to the range of difference distribution, and the differences in the same interval are mapped to continuous positions in the corresponding feature value sequence, thereby limiting the distribution range of the degree of change of the prediction result in the screening sequence.
[0009] Preferably, focusing on the segmented representation of abnormal changes in the set of change amplitudes, a set of change segments with positional correlation and feature mapping relationships is constructed by filtering and continuously aggregating the change amplitudes. The specific steps are as follows: Read the differences in the set of change ranges and compare them with the preset threshold. Mark the positions of the filtering sequences corresponding to the differences that exceed the preset threshold. At the same time, record the corresponding candidate result numbers and feature value distributions to form an abnormal position set. Sort the set of abnormal positions and determine the numbering relationship between adjacent positions. Merge the continuously increasing abnormal positions into the same continuous interval, record the start and end positions of the continuous interval, and form a set of continuous intervals. Extract the feature values within the corresponding position range of the continuous interval set, arrange them in a predetermined order and establish position binding relationships to form a feature value distribution set corresponding to the continuous interval; By summarizing the continuous interval set and the characteristic value distribution set, interval description information is established and arranged in the order of the screening sequence to form a set of changing segments used to characterize the range of abnormal jumps.
[0010] Preferably, each continuous interval in the set of change intervals is arranged in relation to the position of the screening sequence, and the difference sequence corresponding to each continuous interval maintains a positional mapping relationship with the feature value distribution set. The boundary position of the abnormal jump range is limited by the interval description information, forming a set of change intervals with a correspondence between continuous interval identifiers and feature distributions.
[0011] Preferably, the feature values at corresponding positions within the set of changing segments are reconstructed. This process involves rearranging the positions and discretizing the data to form a continuous distribution structure, while maintaining consistency with the trend of the predicted results. The resulting adjusted feature value sequence is obtained through the following steps: Extract the feature values and prediction results of each segment corresponding to the selected sequence position in the set of changing segments, arrange them in the original order to form a local feature value sequence and establish position identification relationship; The positions of each feature value in the local feature value sequence are rearranged, the position numbers are reassigned according to the direction of the screening sequence, and the correspondence between the feature values and the prediction results is maintained to form a new position distribution sequence. The rearranged feature value sequence is expanded, arranged sequentially according to the position number, and an interval distribution relationship is established. At the same time, the arrangement direction is adjusted according to the prediction results to maintain a consistent trend. The results of each segment expansion are integrated and sequentially concatenated with the feature values of the unchanging segments to form a continuously arranged sequence of adjusted feature values.
[0012] Preferably, the position numbers are arranged progressively according to the screening sequence direction, and the feature values are expanded in the order of the position numbers to form an interval distribution structure. The arrangement direction of the prediction results is consistent with the original change order, and the feature values at each position in the adjusted feature value sequence maintain a one-to-one correspondence with the prediction results.
[0013] Preferably, based on the extended process of the adjusted feature value sequence, a continuous constraint adjustment mechanism is introduced to guide the candidate results entering the screening sequence, forming a sequence structure with continuous prediction results. The specific steps are as follows: Extract the feature values and corresponding prediction results from the end positions of the adjusted feature value sequence and establish position identifiers. Then, determine the end positions as candidate results and connect them to the reference positions. The feature values of the candidate results are entered and compared sequentially with the feature values of the end positions to complete the insertion of the candidate results in the filtering sequence and assign them consecutive numbers. Adjust the arrangement direction of candidate results corresponding to prediction results, and combine the relationship between the changes in prediction results at the end position to form a continuous change sequence; The candidate results are integrated with the adjusted feature value sequence and then sequentially spliced to form a continuous screening sequence structure.
[0014] Preferably, in the process of adjusting the arrangement direction of the candidate results corresponding to the prediction results, the arrangement order is determined according to the change relationship between the prediction result at the end position and the prediction result at the previous position, and a continuous progressive relationship is maintained between the candidate result prediction results and the prediction result at the end position, while maintaining the correspondence between the feature values of the candidate results and the prediction results.
[0015] Preferably, the overall distribution optimization of the prediction results after continuous constraint adjustment is carried out by constructing a balanced arrangement structure through variation amplitude coordination and segment constraints, forming a stable screening sequence corresponding to the set of variation segments. The specific steps are as follows: Read the prediction results after continuous constraint adjustment and record the position number and corresponding feature value according to the screening sequence. Perform comparison of prediction results of adjacent positions to generate a change amplitude sequence and establish position correlation. The change amplitude sequence is divided into multiple position intervals, and the change amplitude distribution within each interval is organized. The change amplitude sequence is then redistributed while maintaining positional correlation to form a balanced distribution structure. The order of the predicted results in the screening sequence is rearranged and the corresponding change magnitude distribution results are re-arranged to form a continuous change relationship and maintain the consistency between the predicted results and the feature values. Map the set of changed segments to the position interval of the rearranged screening sequence and establish constraint relationships to form a screening sequence that satisfies the condition of balanced distribution and has a segment constraint structure.
[0016] The technical effects and advantages provided by the present invention in the above technical solution are as follows: This invention constructs a set of change amplitudes and identifies and reconstructs abnormal jump segments, transforming the changes in the prediction results in the screening sequence from local bursts to a continuous transition state. This avoids individual candidate results being misjudged as preferred objects due to numerical anomalies, ensures that the screening path proceeds along the established change pattern, maintains the sequential correlation between prediction results during sequence expansion, reduces path deviations caused by local fluctuations, keeps the overall screening process stable, and reduces the risk of accumulated deviations.
[0017] This invention rearranges and discretizes the feature numerical sequence, and combines continuous constraint adjustment and overall reconstruction processing to enable subsequent candidate results to form a continuous connection with the preceding data when they are included in the screening sequence. At the same time, it coordinates the prediction results globally through balanced distribution conditions to ensure that the variation amplitude of the screening sequence at different positions is consistently expressed, thereby improving the overall consistency and reliability of the evaluation results of large-scale candidate systems and ensuring that the screening results have a stable distribution structure. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0019] Figure 1 This is a flowchart of the data processing method for high-throughput screening of bimetallic catalysts based on machine learning, as described in this invention. Detailed Implementation
[0020] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.
[0021] This invention provides, for example Figure 1 The data processing method for high-throughput screening of bimetallic catalysts based on machine learning, as shown, includes the following steps: The feature values and prediction results corresponding to each candidate result in the continuous screening process are obtained. The feature values are uniformly organized, and the adjacent differences are calculated point by point according to the screening sequence. A set of change amplitudes is constructed based on the size of the adjacent differences to characterize the degree of change of the prediction results in the screening sequence. In the continuous screening process, to achieve a detailed characterization and sequential expression of the predicted changes in candidate results, the specific implementation steps for unifying and organizing the feature values of candidate results and prediction results and constructing a set of change magnitudes are as follows: Each candidate result generated during the continuous screening process is numbered according to the screening order, and the corresponding feature value and prediction result are extracted for each number position. Each set of feature values consists of multiple fixed-dimensional values, and the arrangement order and numerical meaning of each dimension remain consistent across different candidate results. After extraction, all candidate results are arranged in order of number to form a sequential feature value sequence, and the prediction results corresponding to each candidate result are arranged in a completely consistent number order to form a prediction result sequence.
[0022] During the arrangement process, a specific position identifier for the feature value in the sequence is established for each candidate result, so that the feature value of any candidate result can be directly mapped to a specific position in the prediction result sequence through the position identifier. On this basis, the feature value sequence is uniformly organized and processed. Specifically, the feature values of each dimension of all candidate results are arranged one by one in a predetermined order, and the values of the same dimension between different candidate results are aligned to ensure that the feature expression structure of each candidate result in the sequence is consistent, thereby forming a set of feature values with a unified structure and a prediction result sequence that strictly corresponds to it.
[0023] Around the already formed prediction result sequence, the difference calculation is performed on the prediction results of adjacent positions one by one in the screening order. In specific implementation, starting from the beginning of the prediction result sequence, the prediction result of the current number position is selected and compared with the prediction result of the immediately preceding number position one by one. The numerical difference between the two is recorded as the difference value of that position, and a correspondence is established between the difference value and the current number position. Then, the process continues to advance along the screening order, and the above difference calculation process is repeated for each pair of adjacent number positions until the entire prediction result sequence is covered, thus forming a set of difference sequences arranged in the screening order.
[0024] During the difference calculation process, the candidate result number corresponding to each difference and the feature value of the position of that number are recorded simultaneously. This ensures that the difference not only has numerical attributes but also clear positional attributes and feature relationships, thus providing a traceable data foundation for subsequent change analysis based on the difference.
[0025] The resulting difference sequences are centrally organized to construct a set of variation ranges. In the specific implementation process, all differences are first uniformly summarized according to their numbering order in the screening sequence, and a unique identifier is assigned to each difference, so that each difference can form a one-to-one correspondence with the corresponding candidate result number and feature value. Then, all differences are classified and grouped according to their specific numerical values, and differences with similar values are grouped into the same category. Within each category, they are rearranged according to the numbering order, thus forming a set of variation ranges composed of multiple difference categories.
[0026] In constructing the set of variation amplitudes, all the difference data contained in each difference category are retained, and the candidate result number and feature value position corresponding to each difference are recorded. This allows the set of variation amplitudes to express both the numerical change and the distribution of the corresponding candidate results. In this way, the set of variation amplitudes not only reflects the numerical differences between the prediction results, but also reflects the distribution of the differences in the screening sequence and the correspondence with the feature values.
[0027] The pre-constructed set of variation amplitudes is used to express the overall degree of change of the prediction results in the screening sequence. In the specific implementation process, for each difference category in the set of variation amplitudes, the distribution position of the differences it contains in the screening sequence is marked, and these position marks are mapped back to the corresponding feature value sequence, so that each variation amplitude category can form a corresponding distribution area in the feature value set.
[0028] Furthermore, by integrating the distribution range of different difference categories in the screening sequence, the changes in the prediction results throughout the entire screening sequence are uniformly expressed in a set form, so that the degree of change in the prediction results at any position can be reflected by its corresponding difference category. In this process, the set of change magnitudes and the set of feature values maintain a one-to-one correspondence, thereby realizing a complete processing chain from feature value acquisition, prediction result sorting, adjacent difference calculation to the construction of the set of change magnitudes and the expression of the degree of change. This allows the prediction changes in the entire screening sequence to be presented in a structured manner and provides a continuous and consistent data foundation for subsequent processing based on the change segments.
[0029] Based on the set of change ranges, continuous intervals with change ranges exceeding a preset threshold are identified, and the feature value distribution of the corresponding screening sequence position is aggregated to form a set of change segments used to characterize the range of abnormal jumps. In continuous filtering paths, to effectively characterize potential anomalous jumps in the prediction results, it is necessary to further extract anomalous intervals with continuous characteristics based on the set of change amplitudes, and then aggregate them by combining the feature value distribution of the filtering sequence positions, thereby forming a set of change segments that can reflect the range of anomalous jumps. The specific implementation steps are as follows: In the established set of variation ranges, each difference is read one by one according to its number in the filtering sequence, and a corresponding variation range judgment condition is set for each difference. The variation range judgment condition is defined by a pre-set numerical limit, which is used to distinguish between normal variation range and abnormal variation range. In the actual processing, each difference in the variation range set is compared with the preset threshold one by one. When the value of a certain difference exceeds the preset threshold, the filtering sequence position corresponding to the difference is marked as an abnormal position. At the same time, the candidate result number corresponding to the abnormal position and the specific distribution of its feature value in the feature value set are recorded.
[0030] By performing the above-mentioned comparison process on all differences in the set of variation amplitudes, an initial set of abnormal locations is formed, consisting of multiple abnormal locations. Each location in this set can be directly associated with the difference in the set of variation amplitudes and the corresponding location in the set of feature values, thus providing basic data for the identification of subsequent continuous intervals.
[0031] Based on the initial set of abnormal locations, the abnormal locations in the screening sequence are processed for continuous identification. In the specific implementation process, all abnormal locations are sorted according to the numbering order of the screening sequence, and it is determined whether the number difference between adjacent abnormal locations is a fixed interval. When the numbers of adjacent abnormal locations increase continuously, these consecutively numbered locations are merged into the same continuous interval, and a unified interval identifier is assigned to the continuous interval. During the merging process, the start position number and end position number of each continuous interval are recorded, and all abnormal locations contained in the interval are uniformly collected to form multiple continuous interval sets.
[0032] After completing the continuous identification of all abnormal locations, each continuous interval corresponds to a set of continuous differences in the set of change amplitudes, and also corresponds to a continuous distribution area in the set of feature values. This makes the continuous interval not only reflect the distribution range of abnormal changes in the screening sequence, but also retain the mapping relationship between it and the feature values.
[0033] After identifying continuous intervals, the distribution of corresponding positions in the feature value set is aggregated based on the screening sequence position corresponding to each continuous interval. In the specific implementation process, for each continuous interval, the feature values of all candidate results within the range from the start position to the end position of the interval are extracted and centrally organized according to the predetermined arrangement order in the feature value set, so that these feature values form a continuous distribution segment in the set.
[0034] During the processing, each position in the continuous interval is bound to its corresponding feature value, so that each feature value can identify the range of the continuous interval to which it belongs. At the same time, the feature values in the same continuous interval are aggregated as a whole to form a feature distribution set with clear boundaries. This feature distribution set can reflect the feature value distribution status of each candidate result in the abnormal change interval, thereby realizing the mapping and association from the set of change amplitudes to the set of feature values.
[0035] After completing the aggregation and processing of continuous intervals and characteristic value distributions, all continuous intervals are uniformly organized to form a set of change segments used to characterize the range of abnormal jumps. In the specific implementation process, complete segment description information is established for each continuous interval. This segment description information includes the start position number, end position number, difference sequence in the corresponding change amplitude set, and the corresponding characteristic value distribution set of the interval. Subsequently, all interval description information is uniformly arranged according to the screening sequence order to form a set of change segments, so that each segment in the set can completely express the position, change amplitude, and characteristic value distribution of an abnormal jump range in the screening sequence.
[0036] This set of change segments allows for segmented representation of anomalous change behaviors throughout the screening sequence, transforming the originally discrete anomalous changes into segment structures with continuous boundaries, thus providing a clear data foundation for subsequent segment-based processing.
[0037] For the set of changing segments, the feature values at the corresponding screening sequence positions are rearranged. The feature values located in the set of changing segments are discretized and expanded along the screening sequence direction. During the expansion process, the trend direction of the prediction results is kept consistent, and the adjusted feature value sequence is obtained. In the continuous screening path, when further processing the already determined set of change segments, the feature values at the corresponding screening sequence positions are rearranged, and discretization is performed along the screening sequence direction. This transforms the originally concentrated change segments into a continuous transitional expression within the sequence, thereby obtaining the adjusted feature value sequence. This process unfolds segment by segment around the set of change segments, combining the positional relationship of the screening sequence with the direction of change in the prediction results to progressively reconstruct the feature values. The specific implementation steps are as follows: For each change segment in the set of change segments, its start and end position numbers in the filtering sequence are read one by one, and feature value data corresponding to the corresponding position is extracted from the feature value set according to the range of the interval. At the same time, the prediction result data corresponding to these positions are extracted simultaneously. During the extraction process, all feature values within the change segment are recorded in the original arrangement order in the filtering sequence, and clear position identification information is attached to each feature value so that each feature value can indicate its specific position in the filtering sequence.
[0038] Furthermore, all feature values extracted within the same change range are arranged in their original order to form a continuous sequence of local feature values. At the same time, a one-to-one correspondence is established between this local feature value sequence and the corresponding prediction result sequence, so that each feature value can correspond to its original prediction result, providing a complete and continuous data foundation for subsequent position rearrangement processing.
[0039] For the already formed local feature value sequence, the feature values are rearranged along the direction of the screening sequence. In the specific implementation process, according to the span of the change segment in the screening sequence, each feature value in the local feature value sequence is redistributed to a new screening sequence position. During the redistribution, the originally closely arranged feature values are inserted one by one into the new position interval in sequence, so that the position interval between adjacent feature values in the screening sequence changes, thereby making the feature values form a dispersed distribution in the new arrangement.
[0040] In this process, each feature value is assigned a new position number, which is arranged in ascending order according to the direction of the screening sequence, so that the new positions are distributed in the sequence in a continuous and progressive relationship. At the same time, the relationship between each feature value and its corresponding prediction result is kept unchanged during the position rearrangement process, so that the feature value can still point to the original prediction result under the new position number, thereby completing the position rearrangement while maintaining the data correspondence.
[0041] For the feature value sequence after position rearrangement, discretization expansion is performed along the direction of the screening sequence. In the specific implementation process, based on the position of the rearranged feature value, each feature value is expanded sequentially according to the new position number, so that each feature value occupies an independent and clear position in the screening sequence, while forming a fixed interval between adjacent feature values, so that the feature values originally concentrated in the range of change are distributed point by point in the new sequence structure.
[0042] During the unfolding process, the prediction results corresponding to each feature value are synchronously associated, so that the unfolded feature value sequence can fully reflect the change process of the prediction results within the change range. Furthermore, combined with the original change order of the prediction results within the change range, the arrangement order of the unfolded feature values is adjusted so that the arrangement direction of the feature values in the selection sequence is consistent with the change direction of the prediction results. That is, when the prediction results show an arrangement trend from low to high within the change range, the feature values are arranged in the same direction; when the prediction results show an arrangement trend from high to low within the change range, the feature values are arranged in the opposite direction. This ensures that the arrangement direction of the discretized unfolded feature value sequence is consistent with the change trend of the prediction results.
[0043] After completing the discretization and expansion of all the change segments, the expansion results corresponding to each change segment are integrated in a unified manner according to the screening sequence. During the integration process, the expanded feature values within the change segment are arranged sequentially according to their new position numbers and concatenated with the feature values not within the range of the change segment, so that the feature values in the entire screening sequence form a continuous arrangement relationship.
[0044] During the splicing process, the connection between the boundary positions of the changed segments and the positions of the non-changed segments is sequentially aligned to ensure that the arrangement of feature values in the sequence remains continuous and unbroken, while ensuring that the feature value at each position corresponds to a unique prediction result. Through the above integration process, a complete adjusted feature value sequence is formed. This sequence presents a continuous distribution in the screening sequence, transforming the concentrated changes in the original changed segments into a gradual transition form in the new sequence structure. This achieves the overall reconstruction of the set of changed segments and provides a stable and consistent data foundation for subsequent continuous adjustment processing.
[0045] Based on the adjusted feature value sequence, the candidate results that subsequently enter the screening sequence are continuously constrained and adjusted so that the prediction results corresponding to the subsequent candidate results maintain a continuous connection with the adjusted feature value sequence in the screening sequence. As the continuous screening process progresses, to prevent sudden changes in the predicted expression of newly entered candidate results, it is necessary to continuously constrain and adjust the candidate results entering the screening sequence using the already obtained adjusted feature value sequence as a reference. This ensures that the predicted results in the screening sequence can form a sequential connection with the aforementioned adjusted feature value sequence. This process revolves around the sequence extension process, guiding the newly entered candidate results point by point to maintain the continuity of change and the consistency of distribution throughout the expansion process of the entire screening sequence. The specific implementation steps are as follows: Around the already formed adjusted feature value sequence, the position of the sequence at the end of the screening sequence is marked, and the feature value corresponding to the end position and its corresponding prediction result are extracted. The end position is used as the reference starting point when subsequent candidate results enter the screening sequence.
[0046] In the specific implementation process, a complete record is established for the last position in the adjusted feature value sequence, including the number of the position in the screening sequence, the corresponding feature value arrangement status, and the corresponding prediction result value, so that the last position not only has numerical attributes, but also has clear sequence position attributes; at the same time, the last position is used as the starting reference position when the subsequent candidate results are entered into the screening sequence, so that the newly entered candidate results can be arranged around the position, thereby forming a continuous extension relationship at the sequence level.
[0047] When a new candidate result enters the filtering sequence, for each new candidate result, its corresponding feature value is extracted, and this feature value is compared one by one with the feature value at the end of the adjusted feature value sequence in a predetermined order, so that the feature value of the new candidate result can correspond to the adjusted feature value sequence in terms of arrangement order. In the specific processing, according to the relative relationship between the feature value of the new candidate result and the end feature value, the new candidate result is inserted into the next number position immediately after the end position in the filtering sequence, so that its position in the sequence is immediately adjacent to the adjusted feature value sequence.
[0048] During the insertion process, new candidate results are assigned new sequence numbers, and these numbers are ensured to remain continuously progressive in the overall screening sequence, so that the newly entered candidate results are closely connected with the adjusted feature value sequence in terms of position.
[0049] For new candidate results that have completed location access, continuous constraint adjustment processing is applied to their corresponding prediction results. In the specific implementation process, the prediction results of the new candidate results are arranged sequentially with reference to the prediction results at the end of the adjusted feature value sequence, so that the new prediction results are consistent with the previous prediction results in the direction of numerical change. Specifically, when the prediction results at the end of the adjusted feature value sequence show an increasing relationship at its previous position, the prediction results of the new candidate results are arranged in the same direction. When the prediction results at the end show a decreasing relationship at its previous position, the prediction results of the new candidate results are arranged in the corresponding direction, so that the prediction results form a continuous change relationship in the screening sequence.
[0050] In this process, by constraining the arrangement direction of the new candidate result prediction results, the change state is kept consistent with the change trend in the adjusted feature value sequence, thus avoiding sudden changes that are inconsistent with the change direction of the preceding sequence.
[0051] After completing the feature value input of new candidate results and the continuous constraint adjustment of prediction results, the entire screening sequence is uniformly organized so that the adjusted feature value sequence and the newly entered candidate results together form a complete continuous sequence structure. During the organization process, the feature values and prediction results corresponding to the new candidate results are arranged in order according to their position numbers in the screening sequence, and sequentially spliced with the original adjusted feature value sequence so that the entire screening sequence maintains a continuous arrangement relationship in structure.
[0052] Meanwhile, by maintaining the correspondence between the feature values at each position and the prediction results, the numerical changes in the entire screening sequence are continuously connected between the preceding and following positions, thereby achieving continuous constraint adjustment on subsequent candidate results. This ensures that the prediction results in the screening sequence maintain a consistent change and continuity with the adjusted feature value sequence, and provides a continuous data foundation for the stable advancement of the subsequent screening process.
[0053] Based on the prediction results after continuous constraint adjustment, the screening sequence is reconstructed as a whole so that the prediction results at each position in the screening sequence meet the preset balanced distribution conditions in terms of the magnitude of change, and form a corresponding constraint relationship with the set of change segments. After continuous constraint adjustment of the continuous screening path, to further optimize the overall screening sequence, a unified reconstruction process needs to be carried out around the prediction results after continuous constraint adjustment. This involves coordinating the variation amplitude of the prediction results at each position in the screening sequence to form an arrangement state that conforms to the preset balanced distribution conditions across the entire sequence, while establishing a stable correspondence with the set of change segments. This results in a continuous and controllable change structure at the overall level. The specific implementation steps are as follows: For the prediction results after continuous constraint adjustment, the prediction results corresponding to each position are read one by one in the order of the position numbers in the screening sequence, and each prediction result is bound and recorded with its position number and corresponding feature value, so that each position forms a complete data unit. After the reading is completed, the prediction results of adjacent positions in the screening sequence are compared one by one, the numerical difference between each pair of adjacent positions is recorded, and the starting position number and ending position number of each numerical difference are marked, so that the numerical difference can reflect the changes in the specific position interval.
[0054] In this process, all numerical differences are arranged in the order of the screening sequence to form a complete change range sequence. At the same time, the position information corresponding to each difference and the associated feature value identifier are retained in the change range sequence, so that the change range sequence can fully express the change distribution of the prediction result in the screening sequence and maintain the same expression method as the existing screening data.
[0055] Based on the established sequence of variation amplitudes, the variation amplitudes are adjusted item by item according to the preset balanced distribution conditions. In the specific implementation process, the variation amplitude sequence is first divided into intervals, and the entire sequence is segmented according to a fixed number of position intervals. Each interval contains a set of consecutively numbered positions. Then, the variation amplitudes within each interval are statistically analyzed and sorted so that the variation amplitudes within each interval are arranged in numerical order, and the distribution position of these variation amplitudes within the interval is recorded. On this basis, the intervals with relatively concentrated variation amplitude distribution are redistributed, and some variation amplitudes within the interval are moved to adjacent intervals, so that the originally concentrated variation amplitudes are dispersed in multiple intervals. At the same time, the intervals with relatively sparse variation amplitude distribution are supplemented so that continuous variation relationships can be formed within these intervals.
[0056] During the adjustment process, each change magnitude is bound to its corresponding screening sequence position, so that the redistribution of the change magnitude can directly affect the reconstruction process of subsequent prediction results, thereby ensuring that the overall distribution of the change magnitude in the screening sequence meets the preset equilibrium distribution condition.
[0057] After adjusting the distribution of the variation amplitude, the prediction results in the screening sequence are reconstructed as a whole. In the specific implementation process, based on the adjusted variation amplitude sequence, each variation amplitude is mapped to the prediction results between adjacent positions in the screening sequence, and the prediction results are rearranged point by point according to the order of the variation amplitude, so that the prediction results between adjacent positions form a new numerical relationship according to the adjusted variation amplitude. When performing reconstruction, each position in the screening sequence is processed one by one, so that the prediction result of each position can form a corresponding variation relationship with the previous position, and form a continuous variation state in the entire sequence.
[0058] During the reconstruction process, the binding relationship between each prediction result and its corresponding feature value remains unchanged, so that the reconstructed screening sequence still maintains a complete correspondence in the data structure. This allows the prediction results to be rearranged in the screening sequence while ensuring data consistency, so that the prediction results at each position show a balanced distribution in terms of variation.
[0059] After the overall reconstruction of the screening sequence is completed, the reconstructed screening sequence and the set of changing segments are subjected to corresponding constraint processing. In the specific implementation process, the start position number and end position number of each changing segment in the set of changing segments are read one by one, and the range of the segment is mapped to the corresponding position interval in the reconstructed screening sequence. During the mapping process, each segment in the set of changing segments corresponds to a continuous position in the screening sequence, and the distribution of the change amplitude within that position range is kept consistent with the preset equilibrium distribution condition. Furthermore, within the interval corresponding to the changing segment, the change relationship of the prediction results is continuously maintained, so that these intervals maintain a stable change state in the subsequent screening process, thereby making the set of changing segments a constraint region in the overall screening sequence. Through the above-mentioned corresponding constraint processing, the distribution of the prediction results in the screening sequence not only satisfies the equilibrium distribution condition, but also maintains a consistent correspondence with the set of changing segments, thereby achieving a stable expression of the overall structure of the screening sequence.
[0060] Example: Predictive fluctuation control data based on a continuous screening process; In the high-throughput screening of bimetallic catalysts, 30 consecutive candidate systems were selected, and the prediction results corresponding to their characteristic values were recorded. The results are then presented in the screening order as follows: Table 1: Original prediction results data; 1 -0.27 16 -0.36 2 -0.30 17 -0.35 3 -0.32 18 -0.34 4 -0.34 19 -0.33 5 -0.35 20 -0.37 6 -0.36 21 -0.36 7 -0.33 22 -0.35 8 -0.63 23 -0.60 9 -0.36 24 -0.38 10 -0.34 25 -0.36 11 -0.33 26 -0.34 12 -0.35 27 -0.35 13 -0.36 28 -0.33 14 -0.59 29 -0.34 15 -0.37 30 -0.36 (2) Statistical data on the magnitude of change The key fluctuation range is obtained by calculating the difference between adjacent results: Table 2: Key Fluctuation Range Data; Normal fluctuation range: 0.01~0.05eV; Abnormal jump range: above 0.20.
[0061] (3) Distribution of abnormal sections Three sets of abnormal segments were identified: Section A: 7-9; Section B: 13-15; Section C: 22-24; Outliers are typically characterized as follows: 0.63 eV; 0.59eV; 0.60 eV.
[0062] (4) Abnormal impact analysis data Table 3: Data table for analysis of abnormal impacts; Extreme value offset range -0.63~-0.27 Deviation from optimal interval proportion 38% Number of misjudged optimal choices 9 Outliers were mistakenly identified as having “superior adsorption capacity”, interfering with the screening process.
[0063] The processed prediction results; (5) The processed prediction results Table 4: Corrected data obtained after dispersing and adjusting the abnormal sections: 8 -0.36 14 -0.35 23 -0.37 The overall changes in the corresponding sections are as follows: A 0.30 0.05 B 0.23 0.04 C 0.25 0.05 Table 5: Comparison data of optimized screening results: Table 6: Stability Improvement Data: Average fluctuation range 0.11 0.03 Maximum jump amplitude 0.30 0.06 Consistency of continuous change Low high The processed prediction results are mainly concentrated in the range of -0.33 to -0.37 eV, which corresponds to the range of relatively ideal adsorption performance and can reflect the stable binding ability of the candidate system in the catalytic process. At the same time, this distribution range is consistent with the adsorption energy range corresponding to the known preferred bimetallic combination, indicating that the results after the above processing are more reasonable in numerical terms, which is conducive to improving the credibility and practical reference value of the screening results.
[0064] This invention constructs a set of change amplitudes and identifies and reconstructs abnormal jump segments, transforming the changes in the prediction results in the screening sequence from local bursts to a continuous transition state. This avoids individual candidate results being misjudged as preferred objects due to numerical anomalies, ensures that the screening path proceeds along the established change pattern, maintains the sequential correlation between prediction results during sequence expansion, reduces path deviations caused by local fluctuations, keeps the overall screening process stable, and reduces the risk of accumulated deviations.
[0065] This invention rearranges and discretizes the feature numerical sequence, and combines continuous constraint adjustment and overall reconstruction processing to enable subsequent candidate results to form a continuous connection with the preceding data when they are included in the screening sequence. At the same time, it coordinates the prediction results globally through balanced distribution conditions to ensure that the variation amplitude of the screening sequence at different positions is consistently expressed, thereby improving the overall consistency and reliability of the evaluation results of large-scale candidate systems and ensuring that the screening results have a stable distribution structure.
[0066] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A data processing method for high-throughput screening of bimetallic catalysts based on machine learning, characterized in that, Includes the following steps: The feature values and prediction results corresponding to each candidate result in the continuous screening process are obtained. The feature values are uniformly organized, and the adjacent differences are calculated point by point according to the screening sequence. A set of change ranges is constructed based on the size of the adjacent differences. Based on the set of change ranges, continuous intervals with change ranges exceeding a preset threshold are identified, and then aggregated by combining the feature value distribution of the corresponding screening sequence position to form a set of change segments. For the set of changing segments, the feature values at the corresponding screening sequence positions are rearranged. The feature values located in the set of changing segments are discretized and expanded along the screening sequence direction. During the expansion process, the trend direction of the prediction results is kept consistent, and the adjusted feature value sequence is obtained. Based on the adjusted feature value sequence, the candidate results that subsequently enter the screening sequence are continuously constrained and adjusted so that the prediction results corresponding to the subsequent candidate results maintain a continuous connection with the adjusted feature value sequence in the screening sequence. Based on the prediction results after continuous constraint adjustment, the screening sequence is reconstructed as a whole so that the prediction results at each position in the screening sequence meet the preset balanced distribution conditions in terms of the magnitude of change, and form a corresponding constraint relationship with the set of change segments.
2. The data processing method for high-throughput screening of bimetallic catalysts based on machine learning according to claim 1, characterized in that, To address the ordered representation of changes in prediction results during continuous screening, the candidate results are sequentially organized and a set of change magnitudes is constructed. The specific steps are as follows: Each candidate result in the continuous screening process is numbered and the corresponding feature values and prediction results are extracted. They are arranged in the screening order to form a feature value sequence and a prediction result sequence. At the same time, a position identifier is established to achieve a one-to-one correspondence between the two. By combining the prediction result sequence, the difference between adjacent positions is calculated to form a difference sequence and the corresponding candidate result number and feature value position are recorded to obtain difference data with positional attributes and feature correlation. The difference sequence is summarized and organized and classified according to the numerical value to form a set of change ranges. At the same time, the candidate result number and feature value position corresponding to the difference are retained to construct a mapping relationship. Based on the difference category in the set of change magnitudes, the position of the selection sequence is mapped to the feature value sequence to complete the overall expression of the degree of change of the prediction result, forming a continuous and consistent data structure for subsequent processing.
3. The data processing method for high-throughput screening of bimetallic catalysts based on machine learning according to claim 2, characterized in that, Each difference in the difference sequence has a unique correspondence with the position in the screening sequence. The set of change magnitudes is divided into intervals according to the range of difference distribution, and the differences in the same interval are mapped to continuous positions in the corresponding feature value sequence, thereby limiting the distribution range of the degree of change of the prediction result in the screening sequence.
4. The data processing method for high-throughput screening of bimetallic catalysts based on machine learning according to claim 2, characterized in that, To address the segmented representation of abnormal changes within a set of change magnitudes, a set of change segments is constructed by filtering and continuously aggregating the change magnitudes. The specific steps are as follows: Read the differences in the set of change ranges and compare them with the preset threshold. Mark the positions of the filtering sequences corresponding to the differences that exceed the preset threshold. At the same time, record the corresponding candidate result numbers and feature value distributions to form an abnormal position set. Sort the set of abnormal positions and determine the numbering relationship between adjacent positions. Merge the continuously increasing abnormal positions into the same continuous interval, record the start and end positions of the continuous interval, and form a set of continuous intervals. Extract the feature values within the corresponding position range of the continuous interval set, arrange them in a predetermined order and establish position binding relationships to form a feature value distribution set corresponding to the continuous interval; Summarize the continuous interval set and the characteristic value distribution set, establish interval description information, and arrange them in the order of the screening sequence to form a set of change segments.
5. The data processing method for high-throughput screening of bimetallic catalysts based on machine learning according to claim 4, characterized in that, The continuous intervals in the set of change intervals are arranged in relation to each other according to the position order of the screening sequence. The difference sequence corresponding to each continuous interval maintains a positional mapping relationship with the feature value distribution set. The boundary position of the abnormal jump range is limited by the interval description information to form the set of change intervals.
6. The data processing method for high-throughput screening of bimetallic catalysts based on machine learning according to claim 4, characterized in that, The feature values at corresponding positions within the set of changing segments are reconstructed. A continuous distribution structure is formed through position rearrangement and discretization, while maintaining consistency with the predicted trend, resulting in an adjusted feature value sequence. The specific steps are as follows: Extract the feature values and prediction results of each segment corresponding to the selected sequence position in the set of changing segments, arrange them in the original order to form a local feature value sequence and establish position identification relationship; The positions of each feature value in the local feature value sequence are rearranged, the position numbers are reassigned according to the direction of the screening sequence, and the correspondence between the feature values and the prediction results is maintained to form a new position distribution sequence. The rearranged feature value sequence is expanded, arranged sequentially according to the position number, and an interval distribution relationship is established. At the same time, the arrangement direction is adjusted according to the prediction results to maintain a consistent trend. The results of each segment expansion are integrated and sequentially concatenated with the feature values of the unchanging segments to form a continuously arranged sequence of adjusted feature values.
7. The data processing method for high-throughput screening of bimetallic catalysts based on machine learning according to claim 6, characterized in that, The position numbers are arranged progressively according to the screening sequence, and the feature values are expanded in the order of the position numbers to form an interval distribution structure. The arrangement direction of the prediction results is consistent with the original change order. After adjustment, the feature values at each position in the feature value sequence maintain a one-to-one correspondence with the prediction results.
8. The data processing method for high-throughput screening of bimetallic catalysts based on machine learning according to claim 6, characterized in that, Based on the extended process of the adjusted feature value sequence, a continuous constraint adjustment mechanism is introduced to guide the candidate results entering the screening sequence, forming a sequence structure of continuously connected prediction results. The specific steps are as follows: Extract the feature values and corresponding prediction results from the end positions of the adjusted feature value sequence and establish position identifiers. Then, determine the end positions as candidate results and connect them to the reference positions. The feature values of the candidate results are entered and compared sequentially with the feature values of the end positions to complete the insertion of the candidate results in the filtering sequence and assign them consecutive numbers. Adjust the arrangement direction of candidate results corresponding to prediction results, and combine the relationship between the changes in prediction results at the end position to form a continuous change sequence; The candidate results are integrated with the adjusted feature value sequence and then sequentially spliced to form a continuous screening sequence structure.
9. The data processing method for high-throughput screening of bimetallic catalysts based on machine learning according to claim 8, characterized in that, In the process of adjusting the arrangement direction of the candidate results and the prediction results, the arrangement order is determined according to the change relationship between the prediction results at the end position and the prediction results at the previous position, and a continuous progressive relationship is maintained between the candidate results and the prediction results at the end position, while maintaining the correspondence between the feature values of the candidate results and the prediction results.
10. The data processing method for high-throughput screening of bimetallic catalysts based on machine learning according to claim 8, characterized in that, To optimize the overall distribution of prediction results after continuous constraint adjustment, a balanced arrangement structure is constructed by coordinating the magnitude of change and segment constraints, forming a stable screening sequence corresponding to the set of change segments. The specific steps are as follows: Read the prediction results after continuous constraint adjustment and record the position number and corresponding feature value according to the screening sequence. Perform comparison of prediction results of adjacent positions to generate a change amplitude sequence and establish position correlation. The change amplitude sequence is divided into multiple position intervals, and the change amplitude distribution within each interval is organized. The change amplitude sequence is then redistributed while maintaining positional correlation to form a balanced distribution structure. The order of the predicted results in the screening sequence is rearranged and the corresponding change magnitude distribution results are re-arranged to form a continuous change relationship and maintain the consistency between the predicted results and the feature values. Map the set of changed segments to the position interval of the rearranged screening sequence and establish constraint relationships to form a screening sequence that satisfies the condition of balanced distribution and has a segment constraint structure.