Virtual sample generation method and system based on non-stationary neural network gaussian process

By using a non-stationary neural network Gaussian process generation method, the problems of sample distortion and insufficient diversity in traditional virtual sample generation are solved, and high-quality virtual sample generation is achieved under small sample conditions, thereby improving the model's generalization ability and applicability.

CN121579961BActive Publication Date: 2026-05-05CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2026-01-29
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Most existing virtual sample generation methods are based on the stationarity assumption, which makes it difficult to handle high-dimensional, multi-source heterogeneous, and strongly nonlinear datasets. This results in distorted and insufficiently diverse virtual samples, and poor model generalization ability under small sample conditions.

Method used

A method based on non-stationary neural networks and Gaussian processes is adopted. Virtual samples are generated through non-stationarity determination, latent variable modeling, probability sampling, and high-dimensional mapping. The latent variable model is constructed by combining neural networks and Gaussian processes, and the similarity measure is adaptively adjusted. Consistency checks are performed to screen high-quality virtual samples.

Benefits of technology

The generated virtual samples maintain statistical consistency and structural rationality under small sample conditions, improve the generalization ability of the model, adapt to various data characteristics, and enhance the diversity and realism of virtual samples, thus solving the problems of sample distortion and insufficient diversity in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121579961B_ABST
    Figure CN121579961B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for generating virtual samples based on a non-stationary neural network and a Gaussian process. The method includes: acquiring raw data and determining its non-stationarity; determining whether statistical properties change with the input position; if non-stationary, constructing a latent variable model that integrates a neural network and a Gaussian process, and training it to obtain a latent variable space representing the non-stationary distribution; sampling based on the probability distribution of the latent variable space; and generating a virtual sample = (Z, Z) using the model through a non-stationary high-dimensional mapping. ‑1 X, where Z is a non-stationary kernel function, Z is a low-dimensional latent variable, and X is an input variable; the consistency between high-dimensional virtual samples and original data is tested, and a set of virtual samples that meet statistical consistency is selected; this invention can generate high-quality and highly diverse virtual samples in small sample and non-stationary scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual sample generation technology, and in particular to a virtual sample generation method and system based on a non-stationary neural network Gaussian process. Background Technology

[0002] When applying machine learning (ML) to the real world, a major bottleneck is the scarcity of data and the high demands on the model. To address this issue, the industry commonly uses virtual sample generation techniques. These techniques aim to extract the inherent distribution patterns and feature associations from a small number of existing real samples, and then synthesize new virtual samples with the same statistical properties as the real samples, thereby expanding the size of the training dataset. Numerous studies have shown that virtual sample generation methods can effectively alleviate overfitting problems under small sample conditions and significantly improve the performance of downstream prediction models.

[0003] However, existing traditional virtual sample generation methods generally suffer from a fundamental theoretical limitation: most of them are built on the strong assumption that the data distribution satisfies "stationarity". Specifically, this assumption holds that the covariance between any two sample points depends only on their absolute distance in the feature space, and is unrelated to the inherent physicochemical properties or other semantic information of the sample points themselves.

[0004] While the stationarity assumption simplifies the mathematical complexity of model building, it often proves inadequate when dealing with the high-dimensional, multi-source, heterogeneous, and strongly nonlinear datasets prevalent in modern applications. For example, in materials property prediction tasks, fundamental properties such as atomic number and electronegativity are intrinsic factors determining the macroscopic properties of materials. Ignoring this information and relying solely on the geometric distance of sample points to measure correlation inevitably leads to misjudgments of the data's intrinsic structure. Numerous academic studies have clearly pointed out that models relying on the stationarity assumption are highly susceptible to two serious problems: first, a significant decline in predictive performance; and second, inaccurate quantification of model prediction uncertainty, a fatal flaw for risk assessment and safety-critical applications. Furthermore, the stationarity assumption is naturally more suitable for describing continuous, smoothly changing data processes, but it falls short in addressing the discontinuities, abrupt changes, and local heterogeneity commonly found in real-world data. Therefore, virtual samples generated by traditional methods often exhibit pattern distortion and a lack of diversity, failing to substantially improve the generalization ability of the target model and potentially even impairing model performance by introducing noisy samples.

[0005] Therefore, there is an urgent need for a virtual sample generation method and system based on Gaussian processes of non-stationary neural networks, which can generate high-quality and highly diverse virtual samples in both small sample and non-stationary scenarios. Summary of the Invention

[0006] To address the aforementioned shortcomings of the existing technology, the purpose of this invention is to provide a virtual sample generation method and system based on a non-stationary neural network Gaussian process, aiming to solve the technical problems of virtual sample distortion, insufficient diversity, and poor model generalization caused by the traditional stationary assumption in small sample learning.

[0007] To achieve the above objectives, in a first aspect, the present invention provides a method for generating virtual samples based on a non-stationary Gaussian process of a neural network, the steps of which include:

[0008] S1. Obtain the original sample data and perform a non-stationarity determination on the original sample data to determine whether the statistical characteristics of the sample change with the input position.

[0009] S2. When the original sample data is determined to be non-stationary, a Gaussian Process Latent Variable Model (GPLVM) is constructed to describe the characteristics of the changes in the distribution location of the samples. The Gaussian Process Latent Variable Model is trained on the original sample data by fusing neural networks and Gaussian processes to obtain a latent variable space that characterizes the non-stationary distribution characteristics of the samples.

[0010] S3. Sampling is performed based on the probability distribution of the latent variable space to obtain virtual samples of latent variables. ;

[0011] S4. Using the Gaussian process latent variable model, perform a non-stationary high-dimensional mapping on the latent variable virtual samples to generate high-dimensional virtual samples. ,in Z is a non-stationary kernel function, Z is a low-dimensional hidden variable, and X is the input variable;

[0012] S5. Perform a consistency test on the high-dimensional virtual sample and the original sample data, and select a set of virtual samples that meet the statistical consistency requirements based on the consistency test results.

[0013] As a further improvement to the above scheme, when the original sample data are labeled samples, the high-dimensional virtual samples are processed based on the Gaussian process latent variable model or Gaussian process regression model. Generate corresponding virtual tags .

[0014] As a further improvement to the above scheme, in step S1, the non-stationarity determination includes any or more of the following methods:

[0015] The sample space is divided into multiple overlapping regions and the statistical differences between the regions are calculated.

[0016] Comparative analysis of the distribution of kernel function hyperparameters for samples at different locations;

[0017] Alternatively, the variation of the sample distribution with the input location can be evaluated using an empirical distribution function.

[0018] As a further improvement to the above scheme, the Gaussian process latent variable model is a non-stationary Gaussian process latent variable model, which constructs a non-stationary kernel function by embedding a neural network into a stationary kernel function, so as to achieve adaptive modeling of the characteristics of samples in different input regions.

[0019] As a further improvement to the above scheme, the non-stationary kernel function The construction method is as follows:

[0020] Using neural networks The original input coordinates x are nonlinearly mapped to the latent space φ(x), and the distance in the latent space is calculated using a stationary kernel function K. ;

[0021] These represent the coordinates or feature vectors of the i-th and j-th original input samples, respectively.

[0022] For neural network mapping functions, it is used to distort the nonlinear structure of the original space to the latent space, so that the distance in the latent space can better reflect the similarity between samples;

[0023] , The original inputs are respectively The representation in the latent space after mapping via a neural network;

[0024] K is a pre-selected stationary kernel function, whose input is the latent space distance and output is the kernel function value;

[0025] The non-stationary kernel function constructed in the above manner can adaptively adjust the similarity metric in different regions of the original input space, thereby achieving modeling of the non-stationary distribution characteristics of the samples.

[0026] As a further improvement to the above scheme, the stationary kernel function includes one or more of the following: radial basis function kernel, polynomial kernel, exponential kernel, periodic kernel, and linear kernel.

[0027] As a further improvement to the above scheme, when constructing the latent variable space, the model parameters are initialized, and the scaling conjugate gradient class optimization algorithm is used to optimize the training of the model, so as to improve the model convergence efficiency and stability.

[0028] As a further improvement to the above scheme, the consistency check includes any one or more of the following methods: visual comparison of virtual samples and original samples in latent variable space or feature space;

[0029] Two-sample consistency test based on statistical hypothesis testing;

[0030] Or a global consistency assessment based on a measure of distributional differences.

[0031] As a further improvement to the above scheme, the probability distribution of the latent variable space is expressed as: ,sampling , ;

[0032] in, The i-th latent variable is a latent feature representation used to capture latent patterns or structures in the data that are difficult to observe directly.

[0033] Representing latent variables The probability distribution, which we assume here to be a multivariate normal distribution. ;

[0034] It is a latent variable The mean vector of the hidden variable, which follows a normal distribution, reflects the central tendency of the values ​​of the hidden variable.

[0035] It is a latent variable The covariance matrix, which follows a normal distribution, describes the correlation and dispersion among the dimensions of the latent variables;

[0036] It is a noise term sampled from the standard normal distribution N(0,I), where I is the identity matrix. This noise term is introduced to implement the reparameterization technique so that gradient propagation can be performed in the subsequent generation process to optimize the model parameters.

[0037] These are new latent variables obtained after reparameterization, and their mean is used to calculate them. The product of the noise term and the covariance matrix is ​​added to generate the latent variable, which is then used for subsequent operations such as generating high-dimensional virtual samples based on this latent variable.

[0038] As a further improvement to the above scheme, in step S4, the probability space of the non-stationary high-dimensional mapping is represented as:

[0039] ;

[0040] Therefore, based on the above formula, the high-dimensional virtual sample representation is obtained as follows:

[0041] ,

[0042] in, This represents the generated high-dimensional virtual sample. It is a non-stationary kernel function. Let Z be the new latent variable, and Z be the original set of low-dimensional latent variables. The kernel matrix represents the relationship between the new latent variables and the original latent variables; It is the inverse of the original latent variable kernel matrix. This is the kernel matrix between the new hidden variables.

[0043] Secondly, the present invention also provides a virtual sample generation system based on a non-stationary neural network Gaussian process, comprising:

[0044] The nonstationarity determination module is used to perform nonstationarity analysis on the original sample data;

[0045] The latent variable modeling module is used to construct and train a Gaussian process latent variable model to form a latent variable space;

[0046] The latent variable sampling module is used to perform probability sampling in the latent variable space;

[0047] The high-dimensional mapping module is used to map latent variable virtual samples to high-dimensional virtual samples;

[0048] The consistency check module is used to evaluate the consistency of the generated high-dimensional virtual samples and select the set of virtual samples that meet the statistical consistency requirements.

[0049] Because the present invention adopts the above technical solutions, the beneficial effects of this application are as follows:

[0050] 1. This invention provides a virtual sample generation method based on a non-stationary neural network Gaussian process. By organically combining techniques such as non-stationarity determination, Gaussian process latent variable model construction, latent variable sampling, and non-stationary high-dimensional mapping, it can effectively overcome the problems of sample distortion and insufficient diversity caused by the stationarity assumption in traditional virtual sample generation methods. Specifically, firstly, by determining non-stationarity in step S1, scenarios where the traditional stationarity assumption is inapplicable can be identified in a timely manner when the data distribution changes with the input position. This avoids the covariance estimation bias caused by directly applying a stationary model to data that does not meet the stationarity requirement. This determination step provides the necessary prerequisite for subsequent non-stationary modeling, ensuring that the method only activates the corresponding mechanism when non-stationarity processing is truly necessary, thus improving the applicability and reliability of the model.

[0051] Secondly, the Gaussian process latent variable model constructed in step S2 integrates neural networks and Gaussian processes. The neural network nonlinearly distorts the input coordinates, enabling data that was originally non-stationary in the original space to obtain a more reasonable similarity measure in the latent space. Combined with the probabilistic modeling of the latent variable space using Gaussian processes, it can more accurately capture the location-related changes in the sample distribution. Since the model is no longer limited to the assumption of stationarity based solely on the distance between samples, it can more flexibly express the structure of high-dimensional, heterogeneous, and locally heterogeneous data, reducing misjudgments of distribution caused by incorrect assumptions.

[0052] Furthermore, step S3 is based on the probability distribution of the latent variable space. Reparameterized sampling is performed by introducing standard normal noise. and with Generating new latent variables can maintain the randomness of sampling while making the sampling process differentiable, which is beneficial for the stable propagation of gradients and parameter optimization during the model training phase, thereby improving the quality of latent variable generation and the stability of subsequent mapping.

[0053] In addition, step S4 utilizes a non-stationary kernel function. Virtual samples with latent variables Perform high-dimensional mapping based on the mean formula of the posterior distribution of a Gaussian process. Generate high-dimensional virtual samples. Because... The distance metric, embedded with a distorted neural network, adaptively adjusts similarity weights for different input regions. This allows the generated virtual samples to inherit the statistical characteristics of the original data while exhibiting reasonable structure and variation in non-stationary regions, thereby improving sample diversity and realism and mitigating the risk of overfitting under small sample conditions. Step S5 sets up a consistency check and screening mechanism to compare the statistical characteristics of the generated high-dimensional virtual samples with the original sample data, filtering out abnormal samples that deviate from the original distribution. This ensures that the final virtual sample set remains consistent with the original data in key statistical measures such as mean and covariance. This step effectively suppresses distorted samples caused by model errors or sampling fluctuations from entering the training set, further improving the generalization ability of downstream models.

[0054] This invention, through a technical chain from "non-stationarity determination," "non-stationary latent variable modeling," "differentiable sampling," "non-stationary high-dimensional mapping" to "consistency testing," can generate high-quality virtual samples that are statistically consistent, structurally reasonable, and sufficiently diverse in scenarios with small samples and non-stationary data. This improves upon the shortcomings of traditional stationary assumption methods in terms of sample distortion, insufficient diversity, and poor model generalization ability, thus meeting the needs of practical applications for data augmentation and model performance improvement.

[0055] 2. This invention provides a virtual sample generation method based on a non-stationary neural network Gaussian process. By embedding the neural network into a stationary kernel function, a non-stationary kernel function is constructed that can adaptively adjust hyperparameters according to the spatial location of the input data. The stationary kernel function can be selected from radial basis function kernels, polynomial kernels, exponential kernels, periodic kernels, linear kernels, and any combination thereof. This construction method allows the kernel function to no longer rely solely on the geometric distance between samples, but also incorporates prior information such as the physicochemical properties of sample points to measure similarity, thus better reflecting the essential structure of the data. Since data in different regions can exhibit different smoothness or heterogeneity, non-stationary kernel functions can apply differentiated modeling to local regions, effectively capturing complex, nonlinear, and non-stationary data patterns, reducing the distortion in similarity measurement caused by a single kernel function assumption, and thereby generating higher-quality virtual samples. Furthermore, since stationary kernel functions are diverse and can be freely combined, the non-stationary kernel function of this invention can select or mix multiple basis kernels according to the characteristics of the task to accommodate different data characteristics, such as periodicity, trend, and abrupt changes. This flexibility expands the applicability of the method, enabling it to handle both continuously smooth data and non-stationary data containing abrupt changes, discontinuities, or multimodal distributions, thereby improving the robustness of virtual sample generation in cross-domain applications. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0057] Figure 1 This is a flowchart illustrating the virtual sample generation method based on a non-stationary neural network Gaussian process disclosed in Embodiment 1 of the present invention.

[0058] Figure 2 This is a visualization diagram of the data nonstationarity test disclosed in this invention, wherein... Figure 2 (a) is a scatter plot of length scale and signal variance. Figure 2 (b) is a violin plot of length scale and signal variance;

[0059] Figure 3 This is a visualization diagram of the Gaussian kernel function (RBF kernel function) and the non-stationary neural network kernel function disclosed in this invention, wherein... Figure 3 (a) is the RBF kernel function. Figure 3 (b) is the kernel function of a non-stationary neural network; Figure 3In (a), the left plot shows the characteristics of the RBF kernel function, and the right plot shows the sample functions generated by the RBF kernel function. Figure 3 In (b), the left side plot shows the characteristics of the non-stationary neural network kernel function, and the right side plot shows the sample function generated by the non-stationary neural network kernel function.

[0060] Figure 4 This is a two-dimensional visualization diagram of the latent variables disclosed in this invention, wherein... Figure 4 (a) represents the latent variables of the original data. Figure 4 (b) is a two-dimensional visualization diagram of latent variables using the method provided by the present invention;

[0061] Figure 5 This is a radar chart showing the maximum and minimum eigenvalues ​​of the original sample and the high-dimensional virtual sample in each dimension, as disclosed in this invention. Figure 5 (a) shows the maximum feature values ​​of the original and virtual samples. Figure 5 (b) shows the minimum eigenvalue;

[0062] Figure 6 This is a schematic diagram showing the comparison of prediction performance of different learning models after introducing virtual samples generated by different virtual sample generation algorithms, as disclosed in this invention.

[0063] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0064] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0065] It should be noted that the technical solutions of the various embodiments of the present invention can be combined with each other, but only if they are based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0066] Example 1

[0067] This invention provides a virtual sample generation method based on a non-stationary neural network and a Gaussian process, aiming to solve the problems of sample distortion and insufficient diversity caused by the assumption of stationary data distribution in existing technologies. This method, by fusing neural networks and Gaussian processes, can effectively capture the high-dimensional non-stationary characteristics of data.

[0068] See Figure 1 The specific steps are as follows:

[0069] S1. Obtaining Original Sample Data and Determining Non-Stationarity: First, obtain the original sample data to be processed. The stationarity of the original data is determined by examining whether its statistical characteristics drift with the input location. For example, check whether the distribution shape of the original data drifts with absolute coordinates; if a drift occurs, it is determined to be non-stationary.

[0070] Preferably, nonstationarity can be determined by methods including but not limited to dividing the coordinate axes into overlapping windows and calculating the statistics within each window, or by plotting an empirical distribution function.

[0071] This ability to accurately identify non-stationary features in the data provides a basis for subsequent targeted modeling using non-stationary kernel functions, avoiding model distortion caused by blindly assuming data stationarity in traditional methods.

[0072] S2. Construct and train a non-stationary Gaussian process latent variable model (GPLVM): After determining that the data is non-stationary, construct a GPLVM model that integrates neural networks and Gaussian processes to obtain a latent variable space that represents the non-stationary distribution characteristics of the samples.

[0073] Specifically, model construction and optimization: First, the original data, non-stationary kernel function, and initial parameters are input into the GPLVM model. Principal component analysis (PCA) is used to initialize the latent variable space with directions having high variance, and the scaled conjugate gradient (SCG) algorithm is used to optimize the model loss function.

[0074] Non-stationary kernel function design: The original input coordinates x are nonlinearly mapped to the latent space φ(x), and the distance in the latent space is calculated using a stationary kernel function K, and then expressed by the formula. Obtain the non-stationary kernel function;

[0075] in, These represent the coordinates or feature vectors of the i-th and j-th original input samples, respectively.

[0076] For neural network mapping functions, it is used to distort the nonlinear structure of the original space to the latent space, so that the distance in the latent space can better reflect the similarity between samples;

[0077] , The original inputs are respectively The representation in the latent space after mapping via a neural network;

[0078] K is a pre-selected stationary kernel function, whose input is the latent space distance and output is the kernel function value;

[0079] The non-stationary kernel function constructed in the above manner can adaptively adjust the similarity metric in different regions of the original input space, thereby achieving modeling of the non-stationary distribution characteristics of the samples.

[0080] By utilizing the nonlinear mapping capability of neural networks to distort the original coordinates, the kernel function can adaptively adjust the covariance variation pattern according to the input position, thereby achieving differentiated modeling of data characteristics in different regions and deeply integrating the physicochemical properties of the samples.

[0081] S3. Sampling based on the probability distribution of the latent variable space: Extract the probability distribution of low-dimensional latent variables from the trained GPLVM model, and perform sampling using reparameterization techniques to obtain virtual samples of latent variables. .

[0082] Specifically, in the constructed latent variable model of a non-stationary Gaussian process, the projection of each original sample into the latent space is not an isolated point, but rather represents a probability distribution. This is specifically expressed as:

[0083] ;

[0084] To maintain the optimizability of model parameters while generating virtual samples, this embodiment employs a reparameterized sampling mechanism. The specific calculation formula is as follows:

[0085] ,

[0086] in ; The i-th latent variable is a latent feature representation used to capture latent patterns or structures in the data that are difficult to observe directly. Representing latent variables The probability distribution, which we assume here to be a multivariate normal distribution. ; It is a latent variable The mean vector of the hidden variable, which follows a normal distribution, reflects the central tendency of the values ​​of the hidden variable. It is a latent variable The covariance matrix, which follows a normal distribution, describes the correlation and dispersion among the dimensions of the latent variables; It is a noise term sampled from the standard normal distribution N(0,I), where I is the identity matrix. This noise term is introduced to implement the reparameterization technique so that gradient propagation can be performed in the subsequent generation process to optimize the model parameters. These are new latent variables obtained after reparameterization, and their mean is used to calculate them. The product of the noise term and the covariance matrix is ​​added to generate the latent variable, which is then used for subsequent operations such as generating high-dimensional virtual samples based on this latent variable. Defining the latent variable as a probability distribution rather than a fixed value can effectively quantify the uncertainty in small sample data. This representation allows the model to delve into the deep structure hidden behind physicochemical characteristics, ensuring that the generated virtual samples logically conform to the inherent evolutionary laws of materials science. A reparameterization technique is introduced to transfer randomness to an external noise term. In this way, the sampling process affects the model parameters. and Differentiability. This feature ensures that gradients can propagate smoothly during subsequent generative model training, allowing for precise fine-tuning of the latent space structure using optimization algorithms.

[0087] Sampling in a low-dimensional latent variable space not only maintains the clarity and interpretability of the sample patterns, but also allows the generated virtual samples to exhibit a non-linear arrangement in the latent space, more effectively covering the potential distribution patterns in the original data.

[0088] S4. Generate high-dimensional virtual samples: Utilize the forward propagation of a non-stationary Gaussian process model to generate high-dimensional virtual samples with latent variables. Map back to the original high-dimensional space to generate high-dimensional virtual samples. The specific mapping formula is as follows: This mapping process fully considers the non-stationary characteristics, making the generated samples more closely resemble the real physical scene in terms of representation. It can effectively expand the range of feature values, explore meaningful samples outside the original feature space, and enhance the generalization ability of the model.

[0089] S5. Consistency Check and Screening: The generated high-dimensional virtual samples are checked for consistency with the original samples to screen out a set of high-quality virtual samples that meet statistical consistency requirements. Through multi-dimensional consistency filtering, it is ensured that the virtual samples maintain data diversity while not deviating from the physical essence of the original data, fundamentally guaranteeing the effectiveness of data augmentation.

[0090] As a preferred embodiment, based on the above-described virtual sample generation method, a scheme for synchronously generating virtual labels is provided for labeled original data. When the original sample data is a labeled sample (X,y), this application utilizes the non-stationary mapping relationship established during training to generate high-dimensional virtual samples. Matching high-quality virtual tags .

[0091] The specific implementation steps are as follows:

[0092] Constructing a non-stationary label mapping model:

[0093] During step S2, while training the Gaussian Process Latent Variable Model (GPLVM), a Gaussian Process Regression (GPR) model is simultaneously fitted using the input features X and corresponding labels y of the original samples. This model incorporates a non-stationary kernel function constructed from a neural network φ. To capture the non-stationary nature of the label as it changes with the input position;

[0094] Virtual tags Generation: Using the fitted GPR model, the high-dimensional virtual samples generated in step S4 are transformed... The virtual sample is used as input for forward inference. By calculating the cross-covariance between the virtual sample and the original sample set, and considering the non-stationary distribution of the data, the response value or category attribute corresponding to the virtual sample is predicted, i.e., the virtual label. .

[0095] This setup addresses the label prediction distortion problem inherent in traditional augmentation methods when dealing with non-stationary data. By adaptively adjusting the covariance patterns of different regions using a non-stationary kernel function, it ensures that the generated label y* accurately reflects the data's evolution within a specific physicochemical context. The generated feature-label pairs ( , It exhibits a high degree of consistency and logical self-consistency. Experimental data shows that... (See...) Figure 6 After introducing such labeled virtual samples, learning models such as Lasso and SVR improved their performance in MAE, MSE, and R... 2 Significant improvements were observed in all predictive performance metrics.

[0096] In a preferred embodiment, determining the non-stationarity of the original sample data is a prerequisite for constructing the subsequent model. The core of this determination lies in identifying whether the statistical characteristics of the samples significantly shift with changes in the input spatial location. These statistical characteristics include the mean, variance, or distribution pattern, etc. In this embodiment, the non-stationarity determination specifically includes one or more of the following methods:

[0097] Method 1: Statistical Difference Analysis Based on Overlapping Region Division: The input space containing the original samples is divided into multiple local regions with overlapping parts, and statistical characteristics, including but not limited to mean, standard deviation, or higher-order moments, are calculated for the sample points within each region. By comparing the fluctuations of statistics between different overlapping regions, it is determined whether there is a trend shift. Overlapping window technology can smooth out noise interference and more sensitively capture subtle changes in the data in the local space. This method can intuitively reveal the local characteristics of the data distribution, providing direct prior evidence for the adaptive adjustment of the non-stationary kernel function in different spaces in S2.

[0098] Method 2: Comparison of kernel function hyperparameter distributions based on samples at different locations: Using a preliminary Gaussian process fitting, the distribution of kernel function hyperparameters corresponding to samples at different spatial locations is extracted. If the hyperparameters exhibit a non-uniform distribution or significant differences in space, the data is determined to be non-stationary. Deeply exploring the non-stationary nature of the data from the perspective of model parameters can quantify the complexity of data changes. By comparing hyperparameter distributions, the regions with the strongest non-stationarity in the data can be accurately located, thereby enabling the generated virtual samples to have higher representational accuracy in these key regions.

[0099] Method 3: Evaluating the distribution variation with location using empirical distribution functions: Construct empirical distribution functions for the sample under different input subsets and use statistical tests (such as the KS test) to compare the consistency of the empirical distributions among different subsets. If the test results show that the distribution drifts with the input coordinates, it is determined to be non-stationary. This method does not rely on a specific parameterized model and has stronger universality. By evaluating the change in the total distribution, the distribution law of the physicochemical properties of the original sample can be obtained more comprehensively, ensuring that the generated virtual sample is not only reasonable in mean but also logically consistent with the real physical scene in terms of overall probability structure.

[0100] To further illustrate the specific implementation of the nonstationarity determination in step S1, this embodiment uses a dataset of thermal expansion coefficient (CTE) values ​​of 200 compounds collected from the Automated Materials Discovery Process (AFLOW) database as an example.

[0101] Spatial drift test of statistical properties: In this example, 62 physicochemical features of each compound were first extracted using the Magpie tool as input locations / coordinates, and the two key statistical indicators, “signal variance” and “length scale”, were selected to evaluate how they change with the spatial distribution of the input.

[0102] Judgment logic: If the data is stationary, then in different input subspaces, the length scale (representing the distance of correlation) and signal variance fitted by the model should tend to be consistent or fluctuate within a small range.

[0103] like Figure 2 Figure (a) is a scatter plot of length scale and signal variance, with the horizontal axis representing signal variance and the vertical axis representing length scale. By estimating the hyperparameters of samples from different regions, a significant linear relationship was found between length scale and signal variance, with no obvious clustering trend. This means that the statistical characteristics (variance and correlation length) are shifting regularly with changes in the input variable (physicochemical properties of the compound).

[0104] Analysis of the dispersion of the empirical distribution: Further analysis using a violin plot, such as... Figure 2 As shown in (b), the horizontal axis represents Parameter (parameter category, here corresponding to the two parameters "Signal Variance" and "Length Scale"), and the vertical axis represents Value (parameter value), to visualize and analyze the distribution of the above hyperparameters.

[0105] Judgment logic: The "width" and "length" of the violin plot reflect the range of fluctuation of the statistic in space.

[0106] The results show that the hyperparameters are distributed extremely widely, not concentrated around a single constant value. This wide distribution further reveals the potential non-stationarity of the dataset, that is, changes in the data magnitude induce dramatic fluctuations in spatial statistical properties.

[0107] Through the above-described determination process based on thermal expansion coefficient data, this embodiment can scientifically verify that the dataset has non-stationary properties, thereby avoiding the translation invariance limitation of traditional methods that blindly assume "covariance is only related to distance".

[0108] By combining one or more of the above-mentioned determination methods, this embodiment can accurately define the non-stationary boundary of the original data from multiple dimensions such as statistical characteristics, model parameters, and probability distribution. This solves the problem of virtual sample distortion caused by blindly applying the stationarity assumption and ignoring positional correlation in small-sample learning, and lays a solid data foundation for generating high-quality and highly diverse virtual samples in the latent variable space.

[0109] In a preferred embodiment, in step S2, a non-stationary Gaussian process latent variable model (GPLVM) is constructed to capture the complex distribution of the samples. The kernel function design of GPLVM is the core of achieving non-stationary modeling, and the construction logic of the non-stationary kernel function is as follows:

[0110] This embodiment employs a strategy of embedding a neural network into a traditional stationary kernel function to construct a non-stationary kernel function that can adaptively change with the input position. The mathematical expression of this kernel function is:

[0111] (x i ,x j )=K(|| ||);

[0112] Where, x i and x j These represent the coordinates or high-dimensional feature vectors of the i-th and j-th original input samples, respectively. As a nonlinear mapper, its function is to map the coordinates x in the original input space to the latent space. By training a neural network, the complex nonlinear structure in the original space is "distorted" or "reconstructed." Latent space representation. and This is the feature vector after the original input has been transformed by the neural network. In the latent space, the sample distribution that was originally non-stationary and non-uniform in the original space is remapped. Stationary kernel function K: At the latent space scale, this embodiment pre-selects a stationary kernel function, such as the radial basis function (RBF), etc., and calculates the stationary kernel function K in the latent space. and The distance between them ultimately outputs a non-stationary kernel function. value.

[0113] Due to neural networks The mapping capability makes The similarity metric can be adaptively adjusted across different regions of the original input space. Even if two sample pairs are physically equal in the original space, their latent space distances after neural network mapping will differ depending on the characteristics of their respective regions, thus reflecting different levels of correlation. This construction method breaks away from the traditional Gaussian process's notion that "covariance depends only on the input displacement x". i -x j The model overcomes the stationarity constraint of the kernel function. When dealing with data exhibiting significant locational correlations, such as the coefficient of thermal expansion of materials, the model can identify the differences between regions of dramatic fluctuations and regions of relative stability, making the latent variable space more closely resemble the actual physicochemical distribution. Through this non-stationary mapping, the latent variable model can learn deeper nonlinear relationships within the data. Sampling based on the latent variable space obtained through training with this type of kernel function generates high-dimensional virtual samples that are not only numerically reasonable but also effectively fill the "distribution gaps" between the original samples, significantly improving the virtual samples' ability to cover complex and unknown spaces.

[0114] In a preferred embodiment, the stationary kernel function K used in the above-mentioned construction process of the non-stationary kernel function is further limited and optimized. Specifically, the stationary kernel function K is calculated based on the latent space distance, and it is constructed from one or more of the following kernel types, including radial basis function (RBF), polynomial kernel, exponential kernel, periodic kernel, and linear kernel, according to specific data feature requirements.

[0115] The specific application scenarios and implementation logic are as follows:

[0116] In processing data with continuously and smoothly varying physical and chemical properties of materials, the radial basis function (RBF) kernel and exponential kernel are preferred in this embodiment as the basic stationary kernel. This is achieved through neural networks. Map the original space x to the latent space. Then, the similarity of sample points in the latent space, i.e., K, is calculated using the RBF kernel function. RBF =exp(-γ|| || 2 The RBF kernel provides infinite-order differentiability, which, combined with the nonlinear distortion capability of neural networks, allows the model to capture extremely complex and smooth non-stationary local features. This ensures that the generated virtual samples have good continuity in the feature space, avoiding abrupt distortions during the data generation process.

[0117] For sample data exhibiting a clear global shift trend or a specific order of correlation, this embodiment employs a combination of polynomial and linear kernels. Linear kernels are used to capture linear evolution patterns in the latent space, while polynomial kernels are used to model the interactions between features.

[0118] The application of periodic kernels is aimed at non-stationary scenarios with cyclical characteristics or repetitive patterns, such as parameters with periodic fluctuations in the material preparation process. In this embodiment, periodic kernels are selected for modeling.

[0119] The synergistic effect of combined kernel functions allows this embodiment to combine multiple kernel functions through addition or multiplication in complex, low-sample scenarios.

[0120] By combining kernel functions, this invention can simultaneously process mixed non-stationary data with multiple statistical properties, such as data exhibiting both periodic fluctuations and linear growth trends. This flexible construction method greatly broadens the applicability of the virtual sample generation system, enabling the generation of virtual sample sets to be more comprehensive. The data, which is closer to the output of real physical experiments in terms of structural complexity and feature diversity, provides richer and higher-quality training materials for subsequent machine learning models.

[0121] To further illustrate the inventive concept of this invention, experiments were conducted on virtual sample generation and data analysis. After generating 600 virtual samples using the method provided by this invention, 39 outliers were identified and removed, ultimately resulting in a usable dataset containing 561 valid virtual samples, providing a foundation for subsequent analysis and modeling.

[0122] Kernel function visualization and feature comparison:

[0123] like Figure 3 As shown, two typical kernel functions (Radial Basis Function (RBF) and Neural Network Kernel Function (NN)) are analyzed visually from two dimensions: "characteristics of the kernel function itself" and "form of the generated sample function".

[0124] Kernel function characteristics: see Figure 3In the left-hand plot of (b), for the NN kernel function, when the input x and x′ are close, the kernel function value approaches 1, reflecting a strong correlation; as the distance between x and x′ increases, the kernel function value slowly decreases. For comparison, see [link to relevant documentation]. Figure 3 In the left-hand plot of (a), the RBF kernel function rapidly decreases to 0 with increasing distance. The difference in the decay trends of the two types of kernel functions intuitively reflects the characteristic that "NN kernel functions have a more complex correlation structure".

[0125] Sample generation function: see Figure 3 The right-hand plot in (b) shows function samples generated from a Gaussian process using the NN kernel function. In the initial stage (when the input value is small), the fluctuations are large; as the input value increases, the fluctuations gradually converge, and the overall change tends to be smoother, exhibiting "complex and irregular" dynamic characteristics. See also Figure 3 The right side of (a) shows that, in contrast, the samples generated by the RBF kernel function tend to have smoother variations.

[0126] The RBF kernel function, due to its rapidly decaying correlation and smooth output characteristics, is more suitable for processing smooth data; while the neural network kernel function (NN), with its long-range dependency and complex decaying correlation structure, is better suited for complex and nonlinear data pattern analysis and modeling tasks. Through the above data analysis experiments, not only was the effectiveness of the method provided in this invention in generating virtual samples verified, but the differences in data pattern adaptability among different kernel functions were also clarified through kernel function visualization and comparison, providing a practical basis for subsequent selection of kernel functions for targeted analysis.

[0127] In a preferred embodiment, during the construction of the Gaussian Process Latent Variable Model (GPLVM) in step S2, an optimized parameter initialization strategy and an efficient numerical optimization algorithm ensure the training efficiency and robustness of the model when dealing with non-stationary high-dimensional data. The specific implementation steps are as follows:

[0128] Model parameter initialization: Before starting model training, the latent variable space and related parameters are first initialized.

[0129] Latent variable initialization: Principal component analysis (PCA) is used to reduce the dimensionality of the original sample data X and project it onto the directions of the first q principal components with the largest variance, which are used as the initial values ​​of the latent variable Z.

[0130] Hyperparameter initialization: Assign initial values ​​to hyperparameters such as neural network weights, biases, noise terms of Gaussian processes, and signal variance in non-stationary kernel functions.

[0131] Initialization using PCA provides a starting search point with good physical context and geometric meaning for non-stationary GPLVM models. Compared to random initialization, this strategy allows the latent space to reflect the main distribution characteristics of the original data in the early stages of training, significantly reducing the risk of the model getting trapped in local optima and shortening the time required to establish non-stationary mapping relationships.

[0132] After the optimized training model based on scaled conjugate gradient is constructed, the latent variable Z and the model hyperparameters are jointly iteratively optimized using the scaled conjugate gradient (SCG) optimization algorithm.

[0133] The SCG algorithm adaptively adjusts the step size in each iteration by calculating the gradient of the loss function with respect to the latent variables and kernel function parameters, utilizing second-order derivative information. In the high-dimensional parameter space of the non-stationary neural network kernel function, the algorithm searches along conjugate directions until the model's likelihood function converges. The SCG algorithm avoids the time-consuming linear search process of traditional conjugate gradient methods, achieving convergence with fewer iterations when dealing with non-stationary models containing a large number of neural network parameters. The energy landscape of non-stationary kernel functions is relatively complex due to the embedding of neural networks. The SCG algorithm effectively addresses the non-positive definiteness of the Hessian matrix through a scaling mechanism, ensuring numerical stability of the training process under highly nonlinear environments.

[0134] The synergistic application of the aforementioned initialization and optimization schemes enables the method provided by this invention to quickly and stably learn the latent variable space representing the non-stationary characteristics of the samples. This not only ensures that the latent variable virtual samples Z* sampled in the subsequent step S3 have extremely high statistical representativeness, but also provides a solid model foundation for the final generation of physically and logically consistent, diverse, high-dimensional virtual samples. Especially in scenarios with few samples, this efficient training strategy can extract non-stationary features from limited data to the maximum extent, significantly enhancing the practicality of the generation system.

[0135] In a preferred embodiment, in step S5, a consistency check is performed between the generated high-dimensional virtual sample and the original sample data to ensure that the generated sample possesses both diversity and conforms to the inherent physical logic and statistical characteristics of the original data. The consistency check specifically includes one or more of the following methods:

[0136] Method 1: Visual comparison of latent variable space or feature space: Using dimensionality reduction techniques (such as PCA, t-SNE, or the GPLVM latent space mapping constructed in this invention), high-dimensional virtual samples and original samples are projected into two-dimensional or three-dimensional space for intuitive comparison.

[0137] Specifically, observe whether the coverage area of ​​virtual sample points in space exceeds the envelope of the original samples, and whether their distribution density is logically correlated with the clustering trend of the original samples. Visualization provides the most intuitive basis for judging "reasonableness." By comparing the distribution in the latent variable space, the neural network mapping function can be intuitively verified. Whether the topological structure of non-stationary characteristics has been successfully captured, ensuring that the generated samples are not simply noise superpositions, but an effective exploration of the original feature space.

[0138] Method 2: Two-sample consistency test based on statistical hypothesis testing: This method uses rigorous statistical methods to test the consistency between the virtual sample set and the original sample set, such as Hotelling's T. 2 The test, or two-sample t-test, is used to assess whether there are significant differences between two groups of data in terms of mean vector and covariance structure.

[0139] Taking the first six dimensions of the data as an example, the results of the two-sample t-test are shown in Table 1 below. At a significance level of 0.05, the null hypothesis is not rejected. A two-sample t-test was performed on all 62 dimensions, and the results show that at a significance level of 0.05, the null hypothesis is not rejected, indicating that all 62 dimensions exhibit comparable characteristics in both the original and virtual datasets. (In Hotelling'T...) 2 In the test, the calculated F-statistic was 0.0235, which fell outside the rejection region. At a significance level of 0.05, the mean vectors of the original data and the dummy data were considered to be equal, indicating that they came from the same population.

[0140] ;

[0141] Specifically, a confidence level is set, and if the test statistic falls within the acceptance region, the virtual sample is considered statistically homologous to the original data. Quantitative hypothesis testing provides rigorous mathematical support for the quality of the virtual sample. This approach effectively identifies and eliminates outliers that deviate from physical realities, ensuring the high reliability of the data augmentation set used for subsequent model training and preventing overfitting caused by the introduction of "spurious data."

[0142] Method 3: Global Consistency Assessment Based on Distribution Difference Metrics. This method assesses the global similarity between the generated virtual sample distribution P(X*) and the original data distribution P(X) using a distance metric between probability distributions. Specifically, it calculates the MMD value of the two sample sets in the feature space. The smaller the MMD value, the closer the generated virtual samples are to the real data in terms of global statistical characteristics. Global assessment can quantify the generation effect of non-stationary Gaussian process models from a macroscopic perspective. In scenarios with few samples, this assessment method ensures that the generated virtual sample sets maintain high diversity while their overall probability structure does not shift, thereby significantly enhancing the model's generalization ability under unknown conditions while improving the model's prediction accuracy.

[0143] Through the above-mentioned multi-dimensional consistency checks and screenings, the virtual sample set selected by this invention can deeply integrate the physicochemical properties of the samples, making the enhanced dataset closely resemble the real physical scenario in terms of information content and representation accuracy. It is particularly suitable for scientific fields with extremely high requirements for data rigor, such as battery material design.

[0144] In a preferred embodiment, in step S4, the sampled latent variable virtual sample Z* is mapped to the high-dimensional original space through forward propagation of a non-stationary Gaussian process model, thereby generating a high-dimensional virtual sample with physical consistency. The probabilistic mechanism and computational implementation of this mapping process are as follows:

[0145] The probability space construction of non-stationary high-dimensional mappings involves not performing simple point-to-point mappings when generating high-dimensional virtual samples, but rather inferring based on the conditional probability distribution of a non-stationary Gaussian process. The probability space of a non-stationary high-dimensional mapping is represented as:

[0146] ;

[0147] Based on the conditional distribution characteristics of the Gaussian process described above, this embodiment obtains high-dimensional virtual samples by calculating the mean function. The deterministic representation of is given by the following formula:

[0148] ;

[0149] in, This represents the generated high-dimensional virtual sample. It is a non-stationary kernel function. Let Z be the new latent variable, and Z be the original set of low-dimensional latent variables. The kernel matrix represents the relationship between the new latent variables and the original latent variables; It is the inverse of the original latent variable kernel matrix. This is the kernel matrix between the new hidden variables.

[0150] This mapping mechanism utilizes a non-stationary kernel function. Instead of traditional stationary kernel functions, it can adaptively adjust mapping weights according to different locations in the latent space. This ensures the generated... Not only are the numerical values ​​within a reasonable range, but the nonlinear proportional relationships between features also strictly adhere to the physicochemical constraints of the original data, effectively solving the sample distortion problem caused by traditional linear or stationary mappings. The mapping formula mathematically guarantees that the generated virtual samples are a logical extension of the original data distribution. In low-sample environments, this generation method based on the posterior mean of a Gaussian process can maximize the use of known structural information to "interpolate" or "extrapolate" new samples with physical meaning, providing high-quality data support for subsequently improving the generalization ability of machine learning models.

[0151] The following section uses raw CTE (Coefficient of Thermal Expansion) data to verify the performance of the virtual sample generation method proposed in this invention. (See [link to relevant documentation]). Figure 4 Figure (a) shows the two-dimensional visualization of the latent variables in the original CTE data. The horizontal axis represents latent variable z1, and the vertical axis represents latent variable z2. The scatter distribution reflects the overall shape and dispersion of the original samples in the latent variable space.

[0152] Then, using the virtual sample generation method based on the Non-Stationary Neural Gaussian Process (NNGP) proposed in this invention, corresponding virtual samples are generated, and the virtual samples and the original samples are plotted together on the same two-dimensional coordinate system of latent variables to obtain... Figure 4 (b)

[0153] from Figure 4 As can be seen in (b), the virtual samples exhibit a non-linear arrangement within the "neighborhood" of the original samples: they neither simply extend along a straight line / plane nor are they overly concentrated in a certain local area; this non-linear distribution allows the virtual samples to effectively cover the outlier areas in the original data; compared with traditional linear methods, such as linear interpolation, simple expansion after PCA dimensionality reduction, etc., this invention, by leveraging NNGP's ability to model non-stationary data distributions, achieves more flexible and inclusive sample expansion in the latent variable space;

[0154] Therefore, in terms of core capabilities such as covering the details of the original distribution and expanding the representation of abnormal scenarios, the method of this invention demonstrates superior virtual sample generation performance compared to linear methods. The NNGP-based virtual sample generation not only reasonably reproduces the original data distribution in the latent variable space, but also effectively fills in "boundary / abnormal" regions with its nonlinear characteristics, providing higher-quality virtual sample support for subsequent tasks such as data augmentation and model robustness training.

[0155] During the experiment, the maximum and minimum values ​​of features were also analyzed for the original CTE samples and the generated virtual samples.

[0156] For the analysis of eigenvalues, see [link to relevant documentation]. Figure 5 Figure (a) is a polar coordinate plot showing the maximum feature values ​​of the original and virtual CTE samples. Each "ray" in the figure represents a dimension, and the length of the ray corresponds to the maximum feature value in that dimension. It can be intuitively seen from the figure that in some directions, the "ray" length of the virtual sample is shorter, meaning the maximum feature value of the virtual sample in that dimension is smaller than that of the original sample; while in many other directions, the "ray" length of the virtual sample is not shorter than that of the original sample, demonstrating that the maximum feature value of the virtual sample effectively covers most dimensions and even expands in some dimensions.

[0157] For the analysis of eigenvalue minima, see [link to relevant documentation]. Figure 5 Figure (b), also a polar coordinate plot, shows the distribution of eigenvalues ​​for the original and virtual CTE samples. Each "ray" represents a dimension, and the ray length reflects the eigenvalue minimum in the corresponding dimension. It can be observed from the figure that some directions have longer "rays," meaning that the eigenvalue minimum of the virtual samples in these dimensions is greater than that of the original samples. This may be related to the removal of virtual samples containing minimum values ​​during the outlier screening step. Through a comprehensive analysis of these two figures, the differences in eigenvalues ​​between the virtual and original samples, as well as the effectiveness of the method of this invention in expanding the range of eigenvalues, can be clearly understood.

[0158] The method proposed in this invention successfully expands the range of eigenvalues, thereby enabling the effective exploration of samples outside the original feature space. These samples outside the original feature space often offer meaningful insights, indicating alternative pathways for discovering new materials.

[0159] To verify the advantages of the method (NNGP-VSG) provided by this invention in improving the predictive performance of machine learning models, the following comparative experiment was designed:

[0160] Comparison method:

[0161] Virtual samples are not used; the raw data contains only real samples.

[0162] GMM-VSG (AIC / BIC) generates virtual samples based on Gaussian Mixture Model (GMM) and optimizes the model using the AIC / BIC criteria respectively.

[0163] VSG-GP is a traditional virtual sample generation method based on variational autoencoder (VSG) and Gaussian process (GP).

[0164] The NNGP-VSG provided by this invention combines non-stationary neural networks and Gaussian processes to dynamically model the nonlinear characteristics of data.

[0165] Evaluation models: Support Vector Regression (SVR), Lasso Regression, Random Forest (RF), and XGBoost.

[0166] Evaluation metrics: Mean Squared Error (MSE), Mean Absolute Error (MAE), Coefficient of Determination (R²) 2 ).

[0167] Figure 6 The first image, from top to bottom, compares four different machine learning models using MSE as the evaluation metric; the second image compares four different machine learning models using MAE as the evaluation metric; and the third image compares four different machine learning models using R... 2 To evaluate the results, four different machine learning models were compared; the four machine learning models, from left to right, are SVR, Lasso, Random Forest, and XGBoost.

[0168] Of all the comparison methods, NNGP-VSG achieved the lowest values ​​for both MSE and MAE; while in R... 2 The results show that the virtual samples generated by this method are significantly better than those of other methods, indicating that the virtual samples generated by this method can significantly improve the model's prediction accuracy and generalization ability.

[0169] Compared to GMM-VSG, NNGP-VSG captures the complex nonlinear relationships in data through nonstationary kernel functions. Especially in models that are sensitive to feature nonlinearity, such as Lasso and SVR, MSE and MAE are reduced by about 30% and 25%, respectively, highlighting the advantages of nonlinear modeling.

[0170] Compared to VSG-GP, NNGP-VSG performs better in RF and XGBoost models. For example, the MAE of XGBoost decreased from 0.2 to 0.18, proving that non-stationary neural networks are more adaptable to handling high-dimensional and dynamic data distributions.

[0171] Experimental results show that, among all models, NNGP-VSG has the shortest MSE and MAE histograms, and R0 is the lowest. 2 The bar chart shows the highest overall performance, outperforming other methods. This invention significantly improves the quality of virtual samples by combining non-stationary kernel functions and non-parametric Gaussian processes, especially excelling in nonlinear data scenarios. NNGP-VSG not only reduces model error but also enhances the ability to characterize feature interactions, providing an efficient solution for modeling complex datasets such as battery materials.

[0172] Example 2

[0173] This invention also provides a virtual sample generation system based on a non-stationary neural network Gaussian process, comprising:

[0174] The non-stationarity determination module acquires raw sample data and determines whether the data exhibits non-stationarity by analyzing how the statistical characteristics of the samples change with the input location. This module incorporates various determination logics, including dividing the sample space into overlapping regions and calculating local statistical differences, or using an empirical distribution function to assess the shift in the distribution. As the system's entry point, this module accurately identifies the non-stationar nature of the data. This provides a crucial decision-making basis for subsequent modules to select non-stationary kernel functions, ensuring the system's ability to handle complex distributed data and avoiding modeling biases caused by the single stationary assumption in traditional systems.

[0175] The latent variable modeling module constructs a GPLVM model that integrates neural networks and Gaussian processes. Through training on original samples, it extracts a low-dimensional latent variable space representing the non-stationary distribution characteristics of the data. This module constructs a non-stationary kernel matrix by embedding a stationary kernel function into a neural network and integrates PCA initialization and scaling conjugate gradient optimization algorithms for model training. By deconstructing high-dimensional features into latent variables with probabilistic meaning, this module effectively compresses redundant information and captures hidden physicochemical correlations. The introduction of the non-stationary kernel function enables the model to adaptively adjust the mapping scale for different regions, significantly improving the system's accuracy in representing heterogeneous data.

[0176] The latent variable sampling module is used to extract virtual latent variables from the latent variable space formed during training, based on the learned probability distribution p(Z). This module employs a reparameterized sampling mechanism, generating virtual latent variable samples by superimposing random noise weighted by the covariance matrix onto the mean vector. This module ensures the randomness and diversity of generated samples by sampling in the probability space. The application of reparameterization techniques makes the sampling process differentiable, ensuring that the system can continuously correct the latent space structure during iterative optimization.

[0177] The high-dimensional mapping module is responsible for mapping low-dimensional latent variable virtual samples. By using forward propagation through a non-stationary Gaussian process, the original feature space is restored and mapped to a higher dimension, generating a higher-dimensional virtual sample. Using non-stationary kernel matrices The conditional posterior mean is calculated to achieve a nonlinear mapping from the latent space to the physical feature space. The mapping process fully considers the non-stationary fluctuations of the data, ensuring that the generated virtual samples are logically consistent across the feature dimensions. In scenarios with few samples, this module can effectively fill data gaps, providing the system with a high-fidelity augmented dataset.

[0178] The consistency verification module is used to perform multi-dimensional quality assessment and screening of the generated high-dimensional virtual samples, producing the final virtual sample set. This module integrates a series of evaluation tools, including latent space visualization analysis and Hotelling's T... 2 Statistical hypothesis testing and global evaluation based on distributional variance measures such as MMD are performed. Through a rigorous screening mechanism, this module eliminates potential outliers or illogical "spurious samples," ensuring that the final output virtual sample set is statistically consistent with the original data. This fundamentally guarantees the effectiveness of data augmentation and provides reliable data support for subsequently improving the generalization ability of the materials prediction model.

[0179] The system provided by this invention can solve the technical problems of data distortion and lack of diversity in small sample learning through the coordinated cooperation between various modules. It shows great application value, especially in the field of battery material research and development with complex and non-stationary distribution characteristics.

[0180] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. All equivalent structural transformations made using the contents of the present invention's specification and drawings under the inventive concept of the present invention, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.

Claims

1. A virtual sample generation method based on a non-stationary neural network Gaussian process, applied to material property prediction, characterized in that, The steps include: S1. Obtain the original physicochemical property dataset for battery material design, and determine the non-stationarity of the original sample data to determine whether the statistical characteristics of the sample change with the input position. S2. Given that the original sample data is determined to be non-stationary, a Gaussian process latent variable model is constructed to describe the characteristics of changes in the sample distribution location. This Gaussian process latent variable model is a non-stationary Gaussian process latent variable model, which constructs a non-stationary kernel function by embedding a neural network into a stationary kernel function to achieve adaptive modeling of sample characteristics in different input regions. The non-stationary kernel function... The construction method is as follows: Using neural networks The original input coordinates x are nonlinearly mapped to the latent space φ(x), and the distance in the latent space is calculated using a stationary kernel function K. ; in, These represent the coordinates or feature vectors of the i-th and j-th original input samples, respectively. For neural network mapping functions; , The original inputs are respectively The representation in the latent space after mapping by the neural network; K is a stationary kernel function; It is a non-stationary kernel function; The Gaussian process latent variable model is trained on the original sample data by fusing neural networks and Gaussian processes to obtain a latent variable space that characterizes the non-stationary distribution of the samples. S3. Sampling is performed based on the probability distribution of the latent variable space to obtain virtual samples of latent variables. ; S4. Using the Gaussian process latent variable model, perform non-stationary high-dimensional mapping on the latent variable virtual samples to generate physically consistent high-dimensional material virtual samples. ,in Z is a non-stationary kernel function, Z is a low-dimensional hidden variable, and X is the input variable; S5. Perform a consistency check on the high-dimensional virtual samples and the original sample data, and based on the consistency check results, select a set of virtual samples that meet the physical nature and statistical consistency of the materials to expand the training dataset of battery materials.

2. The virtual sample generation method based on a non-stationary neural network Gaussian process according to claim 1, characterized in that, When the original sample data are labeled samples, the high-dimensional virtual samples are processed based on the Gaussian process latent variable model or Gaussian process regression model. Generate corresponding virtual tags .

3. The virtual sample generation method based on a non-stationary neural network Gaussian process according to claim 1 or 2, characterized in that, In step S1, the non-stationarity determination includes any or more of the following methods: The sample space is divided into multiple overlapping regions and the statistical differences between the regions are calculated. Comparative analysis of the distribution of kernel function hyperparameters for samples at different locations; Alternatively, the variation of the sample distribution with the input location can be evaluated using an empirical distribution function.

4. The virtual sample generation method based on a non-stationary neural network Gaussian process according to claim 1 or 2, characterized in that, When constructing the latent variable space, the model parameters are initialized, and the scaling conjugate gradient class optimization algorithm is used to optimize and train the model.

5. The virtual sample generation method based on a non-stationary neural network Gaussian process according to claim 1 or 2, characterized in that, The consistency check includes one or more of the following methods: Visual comparison of virtual samples and original samples in latent variable space or feature space; Two-sample consistency test based on statistical hypothesis testing; Or a global consistency assessment based on a measure of distributional differences.

6. The virtual sample generation method based on a non-stationary neural network Gaussian process according to claim 1 or 2, characterized in that, The probability distribution of the latent variable space is expressed as: ; sampling And calculate: ; in, This represents the i-th hidden variable; Representing latent variables The probability distribution is assumed to be a multivariate normal distribution. ; It is a latent variable The mean vector of the distribution it follows; It is a latent variable The covariance matrix of the distribution that follows a normal distribution; It is a noise term sampled from the standard normal distribution N(0,I), where I is the identity matrix; These are new latent variables obtained after reparameterization.

7. The virtual sample generation method based on a non-stationary neural network Gaussian process according to claim 1 or 2, characterized in that, In step S4, the probability space of the non-stationary high-dimensional mapping is represented as: ; Therefore, based on the above formula, the high-dimensional virtual sample representation is obtained as follows: , in, This represents the generated high-dimensional virtual sample. It is a non-stationary kernel function. Let Z be the new latent variable, and Z be the original set of low-dimensional latent variables. The kernel matrix represents the relationship between the new latent variables and the original latent variables; It is the inverse of the original latent variable kernel matrix. This is the kernel matrix between the new hidden variables.

8. A virtual sample generation system based on a non-stationary neural network Gaussian process, employing the virtual sample generation method based on a non-stationary neural network Gaussian process as described in any one of claims 1-7, characterized in that, include: The nonstationarity determination module is used to perform nonstationarity analysis on the original physicochemical property dataset of battery materials; The latent variable modeling module is used to construct and train a Gaussian process latent variable model to form a latent variable space; The latent variable sampling module is used to perform probability sampling in the latent variable space; The high-dimensional mapping module is used to map latent variable virtual samples to physically consistent high-dimensional material virtual samples. The consistency verification module is used to evaluate the consistency of the generated high-dimensional virtual samples and select a set of virtual samples that meet the requirements of material physical properties and statistical consistency.

Citation Information

Patent Citations

  • A method and system of data modelling

    CA2689789A1

  • Antenna electromagnetic optimization method and system based on non-stationary Gaussian process model

    CN111625923A