Method, apparatus, and computer-readable storage medium for determining metabolic status

By performing normal distribution preprocessing and singular value decomposition and interpolation on metabolomic data, the problem of multiple types of metabolomic data is solved, and the accuracy of the metabolic state of organisms and data processing efficiency is improved.

CN119153122BActive Publication Date: 2025-07-29JIANGXI UNIVERSITY OF TRADITIONAL CHINESE MEDICINE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411659385.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-20
Publication Date
2025-07-29
Estimated Expiration
2044-11-20

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with multiple deletion types in metabolomics data, resulting in low accuracy in determining the metabolic status of the organism.

Method used

The missing data matrix of metabolomics is preprocessed by normal distribution, and the data is interpolated by singular value decomposition method. The corresponding interpolation method is selected according to the missing type, and the data is interpolated by combining the preprocessed data matrix and the first approximate matrix until the iteration condition is reached.

Benefits of technology

It improves the accuracy of metabolomics and organisms' metabolic state, simplifies the difficulty of data processing, increases the application scenarios of data interpolation, and reduces the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119153122B_ABST
    Figure CN119153122B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, and computer-readable storage medium for determining a metabolic state. The method includes: preprocessing a missing data matrix of metabolomics using a normal distribution to obtain a preprocessed data matrix; performing singular value decomposition on the preprocessed data matrix to obtain a diagonal matrix; updating the missing data matrix according to the diagonal matrix to obtain a first approximation matrix; imputing the missing data in the first approximation matrix according to the missing data matrix, the preprocessed data matrix, and the first approximation matrix to obtain a second approximation matrix; repeating from taking the second approximation matrix as the preprocessed data matrix for singular value decomposition until the iteration condition is reached. The finally obtained second approximation matrix is the final imputation result, and the metabolic state of an organism is determined according to the imputed metabolomics data. By imputing missing data by selecting a corresponding imputation method according to the missing type, the present application increases the application scenarios of the metabolic state determination method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and more particularly, to a method and apparatus for determining a metabolic state, and a computer-readable storage medium. Background Art

[0002] Due to technical limitations, errors in sample collection and processing, and insufficient sensitivity and coverage of metabolite detection techniques, the missingness of metabolomics data has become a widespread problem. The missingness of metabolomics data directly leads to a decrease in the accuracy of analyzing the metabolic state of an organism. Currently, the common solution is to impute the missing metabolomics data and then determine the metabolic state of the organism.

[0003] Traditional methods for determining metabolic states mainly include mean imputation, median imputation, mode imputation, regression imputation, K-nearest neighbor imputation (KNN), and multiple imputation, etc. These methods have their own advantages and disadvantages and are suitable for different data types and missing patterns. For example, mean or median imputation is simple and fast, but may not be accurate enough for non-randomly missing or highly skewed data; regression imputation takes into account the relationships between variables but assumes that the data missingness is random; KNN imputation is based on the principle of similarity and is suitable for more complex data structures, but has a large computational cost; multiple imputation reflects the uncertainty of imputation by generating multiple complete datasets and is a relatively advanced method.

[0004] Despite the existence of various data imputation techniques, most existing methods are mainly designed for a single type of missing data, such as only dealing with randomly missing, completely randomly missing, or specific types of non-randomly missing data. This limitation makes it difficult for them to face complex and changing practical application scenarios, thereby affecting the accuracy of metabolic data and resulting in a relatively low accuracy in determining the metabolic state of an organism. Summary of the Invention

[0005] In view of this, the purpose of the embodiments of this application is to provide a method and apparatus for determining a metabolic state, and a computer-readable storage medium, which can improve the accuracy of metabolomics and the metabolic state of an organism, and at the same time increase the application scenarios of the method for determining the metabolic state.

[0006] In a first aspect, an embodiment of the present application provides a method for determining a metabolic state, including: preprocessing a missing data matrix of metabolomics using a normal distribution to obtain a preprocessed data matrix; wherein, the metabolomics is configured to reveal the metabolic state of an organism under different physiological or pathological states; performing singular value decomposition on the preprocessed data matrix to obtain a diagonal matrix; updating the missing data matrix according to the diagonal matrix to obtain a first approximation matrix; imputing the data in the first approximation matrix according to the missing data matrix, the preprocessed data matrix, and the first approximation matrix to obtain a second approximation matrix; repeating from taking the second approximation matrix as the preprocessed data matrix and performing singular value decomposition until the iteration condition is reached, and the finally obtained second approximation matrix is the final imputation result; determining the metabolic state of the organism according to the imputed metabolomics data.

[0007] In the above implementation process, when imputing data, the data in the first approximation matrix is imputed jointly according to the missing data matrix, the preprocessed data matrix, and the first approximation matrix. The preprocessed data matrix can be used to determine the missing type, so that the corresponding imputation method can be selected according to the missing type to impute the missing data, which can improve the accuracy of metabolomics and the metabolic state of the organism, and at the same time increase the application scenarios of this method for determining the metabolic state. In addition, since the singular value decomposition method has certain advantages in matrix dimensionality reduction and data reconstruction, in the method of data imputation, using the singular value decomposition method for imputation can reduce the imputation difficulty and improve the imputation accuracy. Moreover, since the normal distribution only depends on two parameters, the mean and the variance, by preprocessing the missing data matrix of metabolomics using the normal distribution, the processing difficulty of the missing data matrix can be simplified and the processing efficiency can be improved.

[0008] In one embodiment, the diagonal matrix includes a first diagonal matrix and a second diagonal matrix; the performing singular value decomposition on the preprocessed data matrix to obtain a diagonal matrix includes: performing singular value decomposition on the preprocessed data matrix to obtain a first diagonal matrix containing singular values; processing the singular values in the first diagonal matrix based on a set threshold to obtain a second diagonal matrix; the updating the missing data matrix according to the diagonal matrix to obtain a first approximation matrix includes: updating the missing data matrix according to the second diagonal matrix to obtain a first approximation matrix.

[0009] In the above implementation process, by performing singular value decomposition on the preprocessed data matrix, high-dimensional data can be reduced to low-dimensional data, effectively reducing the dimension of the data and lowering the computational complexity. In addition, by retaining the features corresponding to the singular values, the important information in the preprocessed data matrix can be retained, improving the accuracy of data imputation.

[0010] In one embodiment, the method for obtaining the first diagonal matrix includes: classifying the missing data by using the PX-MDC model to obtain the missing type; in the case that the missing type of the missing data is a non-random missing type, randomly weighting the preprocessed data of the non-random missing type to obtain a randomly weighted preprocessed data matrix; performing singular value decomposition on the randomly weighted preprocessed data matrix to obtain a first diagonal matrix including singular values.

[0011] In the above implementation process, in the case that the missing type is non-random missing, by weighting the data before imputation to assign higher weights to the observed data, the reliable data can be utilized to the greatest extent, while reducing the weights of the non-random missing type data, minimizing the negative impact of the missing data on matrix decomposition, and improving the flexibility and robustness of the model.

[0012] In one embodiment, processing the singular values in the first diagonal matrix based on a set threshold to obtain a second diagonal matrix includes: subtracting the set threshold from all the singular values in the first diagonal matrix to obtain an intermediate matrix; retaining a first preset ratio of the data in the intermediate matrix to obtain the second diagonal matrix; wherein the first preset ratio is the percentage of the sum of the retained singular values in the total singular values.

[0013] In the above implementation process, after determining the first diagonal matrix, by setting the first preset ratio and retaining a first preset ratio of the data in the intermediate matrix, the numerical value of the first preset ratio can be adjusted to retain as many singular values as possible and optimize the imputation effect.

[0014] In one embodiment, preprocessing the missing data matrix of metabolomics by using a normal distribution to obtain a preprocessed data matrix includes: replacing the missing values with the minimum value of the metabolites without missing in the original data; calculating the mean and variance of the metabolites after replacing the missing values; generating normal distribution data based on the mean and the variance; and performing corresponding preprocessing on the normal distribution data according to the missing type of the missing data to obtain a preprocessed data matrix.

[0015] In the above implementation process, by performing normal distribution processing on the missing data matrix, the data distribution characteristics in the missing data matrix become more intuitive and easy to interpret, simplifying the processing difficulty of the missing data matrix and improving the convenience of processing the missing data matrix.

[0016] In one embodiment, the corresponding preprocessing is performed on the missing data according to the missing type of the missing data by using the normal distribution data to obtain a preprocessed data matrix, including: in the case where the missing data is of the random missing type, filling the missing data with randomly selected values from the normal distribution data; in the case where the missing data is of the non-random missing type, selecting data smaller than the minimum value of the metabolite from the normal distribution data to fill the missing data; and obtaining the preprocessed data matrix from the missing data matrix after filling.

[0017] In the above implementation process, when preprocessing the missing data matrix, corresponding filling methods are used to fill the missing data according to the missing type of the missing data, which can improve the filling accuracy.

[0018] In one embodiment, interpolating the data in the first approximation matrix according to the missing data matrix, the preprocessed data matrix, and the first approximation matrix to obtain a second approximation matrix includes: determining the missing position of the missing data according to the missing data matrix; determining the missing type of the missing data according to the preprocessed data matrix; determining the missing position and the missing type of the missing data in the first approximation matrix through the missing position and the missing type of the missing data; and interpolating the corresponding missing position in the first approximation matrix by using a corresponding interpolation method according to the missing type of the missing data in the first approximation matrix to determine the second approximation matrix.

[0019] In the above implementation process, when determining the second approximation matrix, the missing position of the missing data is determined according to the missing data matrix, the missing type of the missing data is determined according to the preprocessed data matrix, and then the corresponding interpolation method is determined to perform interpolation at the corresponding processing position of the first approximation matrix, so that the interpolated data can be accurately interpolated to the corresponding position, improving the accuracy of the second approximation matrix.

[0020] In one embodiment, interpolating corresponding missing positions in the first approximate matrix according to the missing type of the missing data in the first approximate matrix by using a corresponding interpolation method to determine the second approximate matrix includes: in the case where the missing data is of the missing at random (MAR) type, filling the missing data with a randomly selected value from the normal distribution data; in the case where the missing data is of the non-MAR type and the value at the corresponding position of the missing data in the first approximate matrix is higher than the minimum value of the metabolites without missing in the original data, filling the missing data with the first weight value at the corresponding position of the preprocessed data matrix and the second weight value at the corresponding position in the first approximate matrix; in the case where the missing data is of the non-MAR type and the value of the missing data at the corresponding position in the first approximate matrix is not higher than the minimum value of the metabolites without missing in the original data, filling the missing data with the value at the corresponding position in the first approximate matrix.

[0021] In the above implementation process, for the missing data of different missing types, corresponding interpolation methods are used for missing data interpolation, which can improve the accuracy of missing data interpolation, increase the missing types that the data interpolation method can target, and increase the application scenarios.

[0022] In a second aspect, an embodiment of the present application further provides a metabolic state determination device, including: a preprocessing module configured to preprocess a missing data matrix of metabolomics by using a normal distribution to obtain a preprocessed data matrix; wherein, the metabolomics is configured to reveal the metabolic state of an organism under different physiological or pathological states; a decomposition module configured to perform singular value decomposition on the preprocessed data matrix to obtain a diagonal matrix; an updating module configured to update the missing data matrix according to the diagonal matrix to obtain a first approximate matrix; an interpolation module configured to interpolate the data in the first approximate matrix according to the missing data matrix, the preprocessed data matrix, and the first approximate matrix to obtain a second approximate matrix; an iteration module configured to repeat starting from performing singular value decomposition on the second approximate matrix as the preprocessed data matrix until an iteration condition is reached, and the finally obtained second approximate matrix is the final interpolation result; a determination module configured to determine the metabolic state of the organism according to the interpolated metabolomics data.

[0023] In a third aspect, an embodiment of the present application further provides an electronic device, including: a processor and a memory, where the memory stores machine-readable instructions executable by the processor, and when the electronic device runs, the machine-readable instructions are executed by the processor to perform the steps of the method in the first aspect or any possible implementation manner of the first aspect.

[0024] Fourthly, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the metabolic state determination method in the above first aspect or any possible implementation manner of the first aspect.

[0025] To make the above objects, features, and advantages of the present application more obvious and understandable, specific embodiments are hereinafter given and described in detail in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 A block diagram of the electronic device provided by the embodiment of the present application;

[0028] Figure 2 A flowchart of the metabolic state determination method provided by the embodiment of the present application;

[0029] Figure 3 NRMSE of different imputation methods on the ST000419 dataset in the experiment provided by the embodiment of the present application;

[0030] Figure 4 NRMSE of different imputation methods on the ST000118 dataset in the experiment provided by the embodiment of the present application;

[0031] Figure 5 SOR of different imputation methods on the ST000419 dataset in the experiment provided by the embodiment of the present application;

[0032] Figure 6 SOR of different imputation methods on the ST000118 dataset in the experiment provided by the embodiment of the present application;

[0033] Figure 7 A functional module diagram of the metabolic state determination device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0035] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, terms such as "first" and "second" are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0036] The missing of metabolomics data stems from various technical and biological factors, such as the missing of low-concentration metabolites due to insufficient instrument detection sensitivity, randomly introduced missing, and data loss caused by changes in the measurement environment, etc. Based on these factors, the missing values can be classified into missing types including: completely random missing, random missing, and non-random missing, etc. Completely random missing means that the probability of missing is independent of any observed or unobserved values, so it is considered completely random. Random missing occurs when the probability of missing depends on the observed data, for example, caused by technical problems such as inaccurate peak detection, signal overlap deconvolution, or ion suppression. Non-random missing is the data missing due to metabolite concentration below the detection limit, usually showing uneven distribution and concentrating on specific metabolites.

[0037] To solve the problem of missing data, technicians have proposed a series of data imputation methods. For example, for data with completely random missing, common imputation methods include random forest imputation, K-nearest neighbor imputation, singular value decomposition imputation, and multiple imputation chained equations, etc.

[0038] Among them, random forest imputation uses multiple decision trees to capture complex relationships in the data, is suitable for high-dimensional data and is robust to outliers. K-nearest neighbor imputation fills in the missing samples by finding the k nearest neighbors most similar to the missing samples and based on the values of these neighbors. The method is simple and intuitive but has a high computational cost in high-dimensional data. Singular value decomposition imputation captures the main structure of the data through matrix decomposition technology, can effectively reduce the dimension and reduce noise, but depends on the linear assumption of the data. Multiple imputation chained equations reflect the uncertainty of missing values by generating multiple imputation datasets, providing more reliable statistical inferences, but with a high computational complexity.

[0039] The processing of non-random missing data is more complex. Commonly used methods include imputation methods such as semi-minimum value imputation, left-censored missing data quantile regression imputation, and No-Skip K-nearest neighbor imputation. Among them, semi-minimum value imputation is a method of inferring missing values based on assumptions about the data, which can estimate the statistical characteristics of the existing data. However, this method may lead to serious biases when the assumptions do not hold, especially when the data distribution is uneven, and the accuracy of the imputation results will be affected. The left-censored missing data quantile regression imputation method can effectively handle the imputation of left-censored data by constructing a quantile regression model, especially suitable for the case of asymmetric data distribution. The advantage of this method is its strong robustness to extreme values, but it has a high dependence on model setting, requires ensuring that the selected quantiles are suitable for the data characteristics, and the computational complexity is relatively high. No-Skip K-nearest neighbor imputation retains all available data, avoiding information loss in traditional KNN methods. Its main advantage is that it can use more information for imputation, enhancing the accuracy of imputation.

[0040] However, the inventors of the present application have found through long-term research that the missing situation of metabolomics data in reality is more complex, and there are usually multiple types of mixed missing, such as completely random missing, random missing, and non-random missing. Therefore, these traditional single strategies are difficult to effectively cope with, which in turn affects the accuracy of metabolic data and leads to a relatively low accuracy in determining the metabolic state of organisms.

[0041] In view of this, the present application proposes a method for determining a metabolic state. When imputing data, the data in the first approximation matrix is imputed based on the missing data matrix, the preprocessed data matrix, and the first approximation matrix. The preprocessed data matrix can be used to determine the missing type, so that the corresponding imputation method can be selected according to the missing type to impute the missing data, which can improve the accuracy of metabolomics and the metabolic state of organisms, and at the same time increase the application scenarios of this method for determining the metabolic state. In addition, since the singular value decomposition method has certain advantages in matrix dimensionality reduction and data reconstruction, in the method of data imputation, the singular value decomposition method is used for imputation, which can reduce the imputation difficulty and improve the imputation accuracy. Moreover, since the normal distribution depends only on two parameters, the mean and variance, by preprocessing the missing data matrix of metabolomics using the normal distribution, the processing difficulty of the missing data matrix can be simplified and the processing efficiency can be improved.

[0042] To facilitate the understanding of this embodiment, the electronic device for executing the method for determining a metabolic state disclosed in the embodiments of the present application will be introduced in detail first.

[0043] As Figure 1As shown, it is a block diagram of an electronic device. The electronic device 100 may include a memory 111 and a processor 113. Those of ordinary skill in the art can understand that Figure 1 The structure shown is only illustrative and does not limit the structure of the electronic device 100. For example, the electronic device 100 may further include more or fewer components than those shown Figure 1 in the figure, or have a different configuration from that shown Figure 1 in the figure.

[0044] The above-mentioned memory 111 and processor 113 are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components may be electrically connected to each other through one or more communication buses or signal lines. The above-mentioned processor 113 is used to execute the executable module stored in the memory.

[0045] Among them, the memory 111 may be, but is not limited to, a random access memory (Random Access Memory, abbreviated as RAM), a read-only memory (Read Only Memory, abbreviated as ROM), a programmable read-only memory (Programmable Read-Only Memory, abbreviated as PROM), an erasable programmable read-only memory (Erasable Programmable Read-Only Memory, abbreviated as EPROM), an electrically erasable programmable read-only memory (Electric Erasable Programmable Read-Only Memory, abbreviated as EEPROM), etc. Among them, the memory 111 is used to store a program. After receiving an execution instruction, the processor 113 executes the program. The method executed by the electronic device 100 defined by the process disclosed in any embodiment of the present application can be applied to the processor 113 or implemented by the processor 113.

[0046] The above-mentioned processor 113 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 113 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0047] The electronic device 100 in this embodiment can be used to execute each step in the various methods provided in the embodiments of the present application. The implementation process of the metabolic state determination method will be described in detail through several embodiments below.

[0048] Please refer to Figure 2 , which is a flowchart of the metabolic state determination method provided in the embodiments of the present application. The following will elaborate in detail on Figure 2 the specific process shown.

[0049] Step 201, preprocess the missing data matrix of metabolomics using a normal distribution to obtain a preprocessed data matrix.

[0050] Among them, metabolomics is a new discipline that simultaneously performs qualitative and quantitative analysis on all low-molecular-weight metabolites in a specific physiological period of a certain organism or cell. By analyzing and interpreting the metabolite profiles in biological samples, the metabolic state of the organism under different physiological or pathological states can be revealed.

[0051] The above-mentioned normal distribution is a probability distribution used to describe the probability distribution of a random variable taking values within a certain range.

[0052] It can be understood that in real data, metabolites may not follow a normal distribution. To improve this situation, the data can be logarithmically transformed to make it close to a normal distribution. After completing data imputation, the exponential transformation is then applied to restore the data to the original scale, which can ensure that the imputation process does not change the essential distribution characteristics of the data.

[0053] Step 202, perform singular value decomposition on the preprocessed data matrix to obtain a diagonal matrix.

[0054] Among them, the diagonal matrix may include a first diagonal matrix and a second diagonal matrix. The first diagonal matrix is the matrix obtained by performing singular value decomposition on the preprocessed data matrix, and the second diagonal matrix is the matrix obtained after processing the first diagonal matrix.

[0055] It can be understood that the diagonal matrix after singular value decomposition may contain singular values.

[0056] Step 203: Update the missing data matrix according to the diagonal matrix to obtain a first approximation matrix.

[0057] The first approximation matrix here can be further deduced inversely according to the singular value decomposition formula.

[0058] It should be understood that after the preprocessed data matrix is subjected to singular value decomposition, a diagonal matrix is obtained. By further processing the diagonal matrix, a processed diagonal matrix (i.e., the second diagonal matrix) may be obtained. Substituting the processed diagonal matrix into the above singular value decomposition formula, a new preprocessed data matrix (i.e., the first approximation matrix) can be obtained.

[0059] Step 204: Impute the data in the first approximation matrix according to the missing data matrix, the preprocessed data matrix, and the first approximation matrix to obtain a second approximation matrix.

[0060] Among them, the missing data matrix is the original metabolomics data with missing data, and this missing data matrix can be used to determine the missing positions of the missing data. The preprocessed data matrix refers to the matrix obtained after simple data processing of the missing data matrix, and this preprocessed data matrix can be used to determine the missing types of the missing data in the missing data matrix. The first approximation matrix is the missing data matrix obtained again after processing the diagonal matrix after singular value decomposition.

[0061] It can be understood that when imputing the data in the first approximation matrix, the corresponding imputation method can be determined according to the missing type of the missing data for imputation.

[0062] The missing of the metabolomics data here includes various types such as completely random missing, random missing, and non-random missing.

[0063] Among them, the data imputation methods corresponding to different missing types may be different.

[0064] Step 205: Repeat from taking the second approximation matrix as the preprocessed data matrix for singular value decomposition until the iteration condition is reached. The finally obtained second approximation matrix is the final imputation result.

[0065] It should be understood that after the second approximate matrix is determined, the second approximate matrix is used as the preprocessed data matrix to continue performing singular value decomposition to obtain a new diagonal matrix. Then, a new first approximate matrix is obtained based on the new diagonal matrix, and the new first approximate matrix is interpolated to obtain a new second diagonal matrix. Then, the new second diagonal matrix is used as the preprocessed data matrix to continue performing singular value decomposition... Repeat this process until the iteration condition is met. The finally obtained second approximate matrix is the final interpolation result.

[0066] The iteration condition here can be to reach the convergence condition or reach the maximum number of iterations, etc. The iteration condition can be selected according to the actual situation.

[0067] Step 206: Determine the metabolic state of the organism based on the interpolated metabolomics data. In the above implementation process, when interpolating the data, the data in the first approximate matrix is interpolated jointly according to the missing data matrix, the preprocessed data matrix, and the first approximate matrix. The preprocessed data matrix can be used to determine the missing type, so that the corresponding interpolation method can be selected according to the missing type to interpolate the missing data, which can improve the accuracy of metabolomics and the metabolic state of the organism, and at the same time increase the application scenarios of this metabolic state determination method. In addition, since the singular value decomposition method has certain advantages in matrix dimensionality reduction and data reconstruction, in the method of data interpolation, the singular value decomposition method is used for interpolation, which can reduce the interpolation difficulty and improve the interpolation accuracy. Moreover, since the normal distribution only depends on two parameters, the mean and the variance, by preprocessing the missing data matrix of metabolomics using the normal distribution, the processing difficulty of the missing data matrix can be simplified and the processing efficiency can be improved.

[0068] In a possible implementation manner, step 202 includes: performing singular value decomposition on the preprocessed data matrix to obtain a first diagonal matrix containing singular values; processing the singular values in the first diagonal matrix based on a set threshold to obtain a second diagonal matrix.

[0069] In one embodiment, the formula for singular value decomposition can be:

[0070] ;

[0071] where, is the preprocessed data matrix, and are orthogonal matrices, is the first diagonal matrix.

[0072] The set threshold here can be a pre-set value. For example, the set threshold can be 0, 1, 5, etc. The set threshold can be adjusted according to the actual situation.

[0073] Understandably, after performing singular value decomposition, a new diagonal matrix (i.e., the second diagonal matrix) can be re-determined based on the singular values in the matrix obtained after singular value decomposition and a set threshold.

[0074] In one embodiment, step 203 includes: updating the missing data matrix according to the second diagonal matrix to obtain a first approximation matrix.

[0075] Specifically, the first approximation matrix can be obtained through the following formula:

[0076] ;

[0077] where is the first approximation matrix, and are orthogonal matrices, is the second diagonal matrix.

[0078] In the above implementation process, by performing singular value decomposition on the preprocessed data matrix, high-dimensional data can be reduced to low-dimensional data, effectively reducing the dimension of the data and the computational complexity. In addition, by retaining the features corresponding to the singular values, important information in the preprocessed data matrix can be retained, improving the accuracy of data imputation.

[0079] In a possible implementation manner, the method for obtaining the first diagonal matrix includes: classifying the missing data using the PX-MDC model to obtain the missing type; in the case where the missing type of the missing data is a non-random missing type, randomly weighting the preprocessed data of the non-random missing type to obtain a randomly weighted preprocessed data matrix; performing singular value decomposition on the randomly weighted preprocessed data matrix to obtain a first diagonal matrix containing singular values.

[0080] The PX-MDC model here is a missing data classification model based on the particle swarm algorithm and XGBoost.

[0081] The above non-random missing type means that the probability of missing is independent of any observed or unobserved values, and the non-random missing type is completely random.

[0082] Among them, random weighting means selecting random numbers within a set range for weighting. Using random weighting introduces a certain degree of randomness to the singular value decomposition imputation, which helps to reduce the impact of the systematic bias of the non-random missing type and makes the imputation model more flexible and robust. Among them, the set range can be set according to actual needs. For example, the set range can be [0.25, 0.5], and the set range can also be [0.5, 0.75], etc. The set range can be adjusted according to actual needs.

[0083] In one embodiment, random weighting can be achieved through the following formula:

[0084] ;

[0085] wherein, is the preprocessed data matrix, is the matrix after being processed by the normal distribution, is a random array. Here, the has the same size and quantity as those of the non-random missing type.

[0086] In the above case, when the value of is non-random missing, the value of is a random number; when the value of is not non-random missing, the value of

[0087] is 1. In the above implementation process, when the missing type is non-random missing, by weighting the data before imputation to assign higher weights to the observed data, reliable data can be utilized to the greatest extent, while reducing the weights of the non-random missing type data, minimizing the negative impact of the missing data on matrix decomposition, and improving the flexibility and robustness of the model.

[0088] In a possible implementation manner, processing the singular values in the first diagonal matrix based on a set threshold to obtain a second diagonal matrix includes: subtracting the set threshold from all the singular values in the first diagonal matrix to obtain an intermediate matrix; retaining data of a first preset ratio in the intermediate matrix to obtain the second diagonal matrix.

[0089] wherein, the first preset ratio is the percentage of the sum of the retained singular values in the total singular values.

[0090] It can be understood that after determining the first diagonal matrix, the singular values can be processed through the set threshold, and then by retaining data of the first preset ratio of the sum of the singular values in the total singular values, the second diagonal matrix is obtained, achieving the purpose of optimizing the first diagonal matrix.

[0091] The above intermediate matrix can be obtained in the following manner: namely, subtracting the set threshold from the singular values, and setting the singular values less than the set threshold to 0 to obtain the intermediate matrix. Among them, this intermediate matrix is also a diagonal matrix.

[0092] Here, the first preset ratio can be a preset ratio value. For example, the first preset ratio can be 50%, 80%, 98%, etc. The first preset ratio can be adjusted according to the actual situation.

[0093] In one embodiment, the set threshold is 98%. Since the data contains a large amount of information, to ensure the rationality of the imputation results, it is recommended to retain as many singular values as possible by setting the first preset ratio to 98%. The first diagonal matrix is optimized by retaining 98% of the singular values to obtain the second diagonal matrix.

[0094] In the above implementation process, after the first diagonal matrix is determined, by setting the first preset ratio and retaining the data of the first preset ratio in the intermediate matrix, the value of the first preset ratio can be adjusted to retain as many singular values as possible and optimize the imputation effect.

[0095] In a possible implementation manner, step 201 includes: replacing the missing values with the minimum value of the metabolites without missing in the original data; calculating the mean and variance of the metabolites after replacing the missing values; generating normal distribution data based on the mean and variance; and performing corresponding preprocessing on the normal distribution data according to the missing type of the missing data to obtain a preprocessing data matrix.

[0096] The minimum value of the metabolite here refers to the minimum value in the current metabolomics data.

[0097] It can be understood that for the missing data in the missing data matrix, the minimum value of the metabolites without missing in the original data can be used for filling first, and the minimum value of the metabolites without missing in the original data is filled at the position of the missing data in the missing data matrix.

[0098] The calculation method of the above mean can be as follows:

[0099] ;

[0100] Among them, is the mean, is the number of data in the missing data matrix of the metabolites after replacing the missing values, is the th data.

[0101] The calculation method of the variance can be as follows:

[0102] ;

[0103] Among them, is the variance, is the number of data in the missing data matrix of the metabolites after replacing the missing values, is the th data, is the mean.

[0104] The above normal distribution data can be generated by setting functions. For example, functions such as np.random.normal, randn, NORM.INV, etc. The setting function can be selected according to the actual situation.

[0105] It should be understood that for missing data of different missing types, due to different missing reasons, the corresponding processing methods may also be different. Therefore, for missing data of different missing types, adopting corresponding preprocessing methods can improve the accuracy of missing data processing.

[0106] In the above implementation process, by performing normal distribution processing on the missing data matrix, the data distribution characteristics in the missing data matrix become more intuitive and easy to interpret, simplifying the processing difficulty of the missing data matrix and improving the convenience of processing the missing data matrix.

[0107] In a possible implementation manner, corresponding preprocessing is performed on the basis of the normal distribution data according to the missing type of the missing data to obtain a preprocessing data matrix, including: in the case where the missing data is of the random missing type, filling the missing data with randomly selected values in the normal distribution data; in the case where the missing data is of the non-random missing type, selecting data smaller than the minimum value of the metabolite from the normal distribution data to fill the missing data; obtaining the preprocessing data matrix through the filled missing data matrix.

[0108] It can be understood that a missing data matrix may include multiple missing data, and the multiple missing data may correspond to one or more missing types. For missing data of different missing types, corresponding filling methods can be adopted to fill the missing data in the missing data matrix.

[0109] In the above implementation process, when preprocessing the missing data matrix, adopting corresponding filling methods to fill the missing data according to the missing type of the missing data can improve the filling accuracy.

[0110] In a possible implementation manner, step 204 includes: determining the missing position of the missing data according to the missing data matrix; determining the missing type of the missing data according to the preprocessing data matrix; determining the missing position and missing type of the missing data in the first approximation matrix through the missing position and missing type of the missing data; performing interpolation on the corresponding missing position in the first approximation matrix according to the missing type of the missing data in the first approximation matrix to determine the second approximation matrix.

[0111] Understandably, when interpolating the data in the first approximation matrix, the missing positions of the missing data in the missing data matrix can be determined first according to the missing data matrix, and then the corresponding positions in the first approximation matrix can be mapped to these missing positions. Then, according to the preprocessed data matrix, the missing types of the missing data at the missing positions are determined, and corresponding interpolation methods are determined according to these missing types. The interpolated data obtained by these interpolation methods are interpolated into the corresponding positions in the first approximation matrix to complete the data interpolation at the corresponding positions in the first approximation matrix. Repeat the above steps until all the missing positions in the missing data matrix have undergone corresponding data interpolation at the corresponding positions in the first approximation matrix, completing the interpolation of the first approximation matrix and obtaining the second approximation matrix.

[0112] In the above implementation process, when determining the second approximation matrix, the missing positions of the missing data are determined according to the missing data matrix, and the missing types of the missing data are determined according to the preprocessed data matrix. Then, corresponding interpolation methods are determined to perform interpolation at the corresponding positions in the first approximation matrix, so that the interpolated data can be accurately interpolated to the corresponding positions, improving the accuracy of the second approximation matrix.

[0113] In a possible implementation manner, corresponding interpolation methods are used to interpolate the corresponding missing positions in the first approximation matrix according to the missing types of the missing data in the first approximation matrix to determine the second approximation matrix, including: in the case where the missing data is of the random missing type, the missing data is filled with randomly selected values from the normal distribution data.

[0114] In the case where the missing data is of the non-random missing type and the value at the corresponding position of the missing data in the first approximation matrix is higher than the minimum value of the metabolites without missing in the original data, the missing data is filled with the first weight value at the corresponding position in the preprocessed data matrix and the second weight value at the corresponding position in the first approximation matrix.

[0115] In the case where the missing data is of the non-random missing type and the value of the missing data at the corresponding position in the first approximation matrix is not higher than the minimum value of the metabolites without missing in the original data, the missing data is filled with the value at the corresponding position in the first approximation matrix.

[0116] In an embodiment, when the missing data is of the non-random missing type, the missing data can be interpolated in the following manner:

[0117] ;

[0118] wherein, is the second approximation matrix, is the first approximation matrix, It means that the position of the missing value is not the position where the value at the corresponding position in the first approximate matrix is greater than the minimum value of the metabolite without missing in the original data. It means that the position of the missing value is the position where the value at the corresponding position in the first approximate matrix is greater than the minimum value of the metabolite without missing in the original data. is the first weight value. is the second weight value.

[0119] The data imputation method involved in the embodiments of the present application can improve the defects of traditional data imputation methods in imputation of non-random missing types.

[0120] In order to verify that the data imputation method in the embodiments of the present application improves the defects of traditional data imputation methods in imputation of non-random missing types, the experiment designed experimental data with the proportion of missing data of non-random missing types ranging from about 50% to 100% and a step size of about 5%. For each experiment with the proportion of non-random missing types, the missing rate ranges from 5% to 40% with a step size of about 2.5%, and six different concentrations were randomly designed for the experiment below.

[0121] In order to better evaluate the application of the data imputation method in the embodiments of the present application in metabolomics missing data, real metabolomics datasets are used in the embodiments of the present application, and all datasets can be accessed through the Metabolomics Workbench (https: / / www.metabolomicsworkbench.org / ).

[0122] Specifically, the datasets in this time come from two different experiments. One includes two datasets, namely the asthma / obesity-related study (mice) and the protein family data (bacteria). Among them, the asthma / obesity-related study (mice) dataset is used to study how three different diets affect the arginine metabolic pathway in the lungs of mice. This experiment contains 80 mouse lung tissue samples, and 612 metabolites are measured by GC-MS. The ID of this dataset is ST000419. The protein family data (bacteria) is used to examine the differences in intracellular metabolite concentrations between wild-type and knockout strains. This experiment includes 5 groups, with 6 replicate samples in each group, and 1 sample is missing (a total of 29 samples), and 249 metabolites are measured by GC-MS. The ID of this dataset is ST000118.

[0123] In order to quantify the imputation error, the normalized root mean square error is often used to evaluate the difference between the imputed value and the true value. The calculation formula is as follows:

[0124] ;

[0125] Among them, represents the value of the true data, while It is the data value after interpolation. NRMSE can measure the standardized level of the error between the interpolated value and the true value, and is sensitive to outliers.

[0126] Considering that the missing values of the non-random missing type are not randomly distributed, directly using NRMSE for evaluation may lead to result bias. The traditional processing method is as follows: First, calculate the NRMSE of each missing variable, rank it among different interpolation methods, sum the ranks of all missing variables for each method, and compare them based on SOR. SOR can be expressed by the following formula:

[0127] ;

[0128] where is the number of missing variables, refers to the rank of different interpolation methods in the th missing variable.

[0129] As Figures 3 - 6 shown, in this study, random weighting is to select random numbers in [0.5, 0.75] for weighting. By comparing the data interpolation method in the embodiment of the present application ( Figures 3 - 6 CWSVD_13 shown in Figures 3 - 6 with the traditional data interpolation methods [including: SVD ( Figures 3 - 6 soft_svd shown in Figures 3 - 6 rf shown in

[0130] The above comparison not only proves the efficiency and reliability of the data interpolation method in the embodiment of the present application when dealing with complex non-random missing data patterns, but also highlights its advantages in accurately interpolating missing data, which is crucial for data integrity and analysis accuracy.

[0131] In the above implementation process, for missing data of different missing types, corresponding imputation methods are used to impute the missing data, which can improve the accuracy of missing data imputation, increase the missing types that the data imputation method can target, and increase the application scenarios.

[0132] Based on the same application concept, an embodiment of the present application also provides a metabolic state determination device corresponding to the metabolic state determination method. Since the principle of solving problems by the device in the embodiment of the present application is similar to that of the foregoing embodiment of the metabolic state determination method, the implementation of the device in this embodiment can refer to the description in the embodiment of the above method, and the repeated parts will not be elaborated.

[0133] Please refer to Figure 7 , which is a schematic diagram of the functional modules of the metabolic state determination device provided by the embodiment of the present application. Each module in the metabolic state determination device in this embodiment is used to execute each step in the above method embodiment. The metabolic state determination device includes a preprocessing module 301, a decomposition module 302, an update module 303, an imputation module 304, an iteration module 305, and a determination module 306; among them,

[0134] The preprocessing module 301 is used to preprocess the missing data matrix of metabolomics by using a normal distribution to obtain a preprocessed data matrix; wherein, the metabolomics is configured to reveal the metabolic state of an organism under different physiological or pathological states.

[0135] The decomposition module 302 is used to perform singular value decomposition on the preprocessed data matrix to obtain a diagonal matrix.

[0136] The update module 303 is used to update the missing data matrix according to the diagonal matrix to obtain a first approximation matrix.

[0137] The imputation module 304 is used to impute the data in the first approximation matrix according to the missing data matrix, the preprocessed data matrix, and the first approximation matrix to obtain a second approximation matrix.

[0138] The iteration module 305 is used to start repeating from performing singular value decomposition on the second approximation matrix as the preprocessed data matrix until the iteration condition is reached. The finally obtained second approximation matrix is the final imputation result.

[0139] The determination module 306 is used to determine the metabolic state of an organism according to the imputed metabolomics data.

[0140] In a possible implementation manner, the decomposition module 302 is further used to: perform singular value decomposition on the preprocessed data matrix to obtain a first diagonal matrix containing singular values; process the singular values in the first diagonal matrix based on a set threshold to obtain a second diagonal matrix.

[0141] In a possible implementation manner, the updating module 303 is specifically configured to: update the missing data matrix according to the second diagonal matrix to obtain a first approximation matrix.

[0142] In a possible implementation manner, the decomposition module 302 is specifically configured to: classify the missing data by using the PX-MDC model to obtain a missing type; in the case that the missing type of the missing data is a non-random missing type, perform random weighting on the preprocessed data of the non-random missing type to obtain a preprocessed data matrix after random weighting; perform singular value decomposition on the preprocessed data matrix after random weighting to obtain a first diagonal matrix including singular values.

[0143] In a possible implementation manner, the decomposition module 302 is specifically configured to: subtract the set threshold from all the singular values in the first diagonal matrix to obtain an intermediate matrix; retain data of a first preset ratio in the intermediate matrix to obtain the second diagonal matrix; where the first preset ratio is the percentage of the sum of the retained singular values in the total singular values.

[0144] In a possible implementation manner, the preprocessing module 301 is further configured to: replace the missing values with the minimum value of the metabolites that are not missing in the original data; calculate the mean and variance of the metabolites after replacing the missing values; generate normal distribution data based on the mean and the variance; perform corresponding preprocessing on the normal distribution data according to the missing type of the missing data to obtain a preprocessed data matrix.

[0145] In a possible implementation manner, the preprocessing module 301 is specifically configured to: in the case that the missing data is of a random missing type, fill the missing data with randomly selected values in the normal distribution data; in the case that the missing data is of a non-random missing type, select data smaller than the minimum value of the metabolite from the normal distribution data to fill the missing data; obtain the preprocessed data matrix through the missing data matrix after filling.

[0146] In a possible implementation manner, the imputation module 304 is specifically configured to: determine the missing position of the missing data according to the missing data matrix; determine the missing type of the missing data according to the preprocessed data matrix; determine the missing position and the missing type of the missing data in the first approximation matrix through the missing position and the missing type of the missing data; perform imputation on the corresponding missing position in the first approximation matrix by using a corresponding imputation method according to the missing type of the missing data in the first approximation matrix to determine the second approximation matrix.

[0147] In a possible implementation, the interpolation module 304 is specifically configured to: when the missing data is of the random missing type, fill the missing data with a randomly selected value from the normal distribution data; when the missing data is of the non-random missing type and the value at the corresponding position of the missing data in the first approximation matrix is higher than the minimum value of the non-missing metabolites in the original data, fill the missing data with the first weight value at the corresponding position of the preprocessed data matrix and the second weight value at the corresponding position of the first approximation matrix; when the missing data is of the non-random missing type and the value of the missing data at the corresponding position in the first approximation matrix is not higher than the minimum value of the non-missing metabolites in the original data, fill the missing data with the value at the corresponding position in the first approximation matrix.

[0148] In addition, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the metabolic state determination method described in the above method embodiment.

[0149] The computer program product of the metabolic state determination method provided by the embodiment of the present application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the steps of the metabolic state determination method described in the above method embodiment. For details, please refer to the above method embodiment and will not be elaborated here.

[0150] In several embodiments provided by the present application, it should be understood that the disclosed device and method can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the device, method, and computer program product according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0151] In addition, in each embodiment of the present application, each functional module can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0152] If the above-mentioned functions are implemented in the form of software functional modules and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes. It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or device including the said elements.

[0153] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application. It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0154] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A method for determining a metabolic state, characterized in that Including: Preprocessing the missing data matrix of metabolomics using a normal distribution to obtain a preprocessed data matrix; wherein, the metabolomics is configured to reveal the metabolic state of an organism under different physiological or pathological states; Performing singular value decomposition on the preprocessed data matrix to obtain a diagonal matrix; Updating the missing data matrix according to the diagonal matrix to obtain a first approximation matrix; Imputing the data in the first approximation matrix according to the missing data matrix, the preprocessed data matrix, and the first approximation matrix to obtain a second approximation matrix; Repeating from taking the second approximation matrix as the preprocessed data matrix for singular value decomposition until the iteration condition is reached, and the finally obtained second approximation matrix is the final imputation result; Determining the metabolic state of the organism according to the imputed metabolomics data; The diagonal matrix includes a first diagonal matrix and a second diagonal matrix; The performing singular value decomposition on the preprocessed data matrix to obtain a diagonal matrix includes: Performing singular value decomposition on the preprocessed data matrix to obtain a first diagonal matrix containing singular values; Processing the singular values in the first diagonal matrix based on a set threshold to obtain a second diagonal matrix; The updating the missing data matrix according to the diagonal matrix to obtain a first approximation matrix includes: Updating the missing data matrix according to the second diagonal matrix to obtain a first approximation matrix; Wherein, the obtaining method of the first diagonal matrix includes: Classifying the missing data using the PX-MDC model to obtain a missing type; In the case that the missing type of the missing data is a non-random missing type, randomly weighting the preprocessed data of the non-random missing type to obtain a randomly weighted preprocessed data matrix; Performing singular value decomposition on the randomly weighted preprocessed data matrix to obtain a first diagonal matrix containing singular values The processing the singular values in the first diagonal matrix based on a set threshold to obtain a second diagonal matrix includes: Subtracting the set threshold from all the singular values in the first diagonal matrix respectively to obtain an intermediate matrix; Retaining data of a first preset ratio in the intermediate matrix to obtain the second diagonal matrix; Wherein, the first preset ratio is the percentage of the sum of the retained singular values in the total singular values; The imputing the data in the first approximation matrix according to the missing data matrix, the preprocessed data matrix, and the first approximation matrix to obtain a second approximation matrix includes: Determining the missing position of the missing data according to the missing data matrix; Determining the missing type of the missing data according to the preprocessed data matrix; Determining the missing position and the missing type of the missing data in the first approximation matrix through the missing position and the missing type of the missing data; Imputing the corresponding missing position in the first approximation matrix using a corresponding imputation method according to the missing type of the missing data in the first approximation matrix to determine the second approximation matrix; Imputing the corresponding missing positions in the first approximate matrix by adopting a corresponding imputation method according to the missing type of the missing data in the first approximate matrix to determine the second approximate matrix includes: In the case where the missing data is of the missing at random (MAR) type, filling the missing data with randomly selected values from the normal distribution data; In the case where the missing data is of the non-random missing type and the value at the corresponding position of the missing data in the first approximate matrix is higher than the minimum value of the metabolites without missing in the original data, filling the missing data with the first weight value at the corresponding position of the preprocessed data matrix and the second weight value at the corresponding position in the first approximate matrix; In the case where the missing data is of the non-random missing type and the value of the missing data at the corresponding position in the first approximate matrix is not higher than the minimum value of the metabolites without missing in the original data, filling the missing data with the value at the corresponding position in the first approximate matrix.

2. The method according to claim 1, wherein Preprocessing the missing data matrix of metabolomics by using the normal distribution to obtain a preprocessed data matrix, including: Replacing the missing values with the minimum value of the metabolites without missing in the original data; Calculating the mean and variance of the metabolites after replacing the missing values; Generating normal distribution data based on the mean and the variance; Performing corresponding preprocessing on the missing data according to the missing type of the missing data by using the normal distribution data to obtain a preprocessed data matrix.

3. The method according to claim 2, wherein Performing corresponding preprocessing on the missing data according to the missing type of the missing data by using the normal distribution data to obtain a preprocessed data matrix, including: In the case where the missing data is of the MAR type, filling the missing data with randomly selected values from the normal distribution data; In the case where the missing data is of the non-random missing type, selecting data smaller than the minimum value of the metabolites from the normal distribution data to fill the missing data; Obtaining the preprocessed data matrix from the missing data matrix after filling.

4. A metabolic state determination device, characterized in that Including: A preprocessing module for preprocessing the missing data matrix of metabolomics by using the normal distribution to obtain a preprocessed data matrix; wherein, the metabolomics is configured to reveal the metabolic state of an organism under different physiological or pathological states; A decomposition module for performing singular value decomposition on the preprocessed data matrix to obtain a diagonal matrix; An updating module for updating the missing data matrix according to the diagonal matrix to obtain a first approximate matrix; An imputation module for imputing the data in the first approximate matrix according to the missing data matrix, the preprocessed data matrix and the first approximate matrix to obtain a second approximate matrix; An iteration module for repeating from performing singular value decomposition on the second approximate matrix as the preprocessed data matrix until the iteration condition is reached, and the finally obtained second approximate matrix is the final imputation result; A determination module for determining the metabolic state of an organism according to the imputed metabolomics data; The decomposition module is further configured to: perform singular value decomposition on the preprocessed data matrix to obtain a first diagonal matrix containing singular values; process the singular values in the first diagonal matrix based on a set threshold to obtain a second diagonal matrix; The update module is specifically configured to: update the missing data matrix according to the second diagonal matrix to obtain a first approximation matrix; The decomposition module is specifically configured to: classify the missing data by using the PX-MDC model to obtain a missing type; in the case that the missing type of the missing data is a non-random missing type, perform random weighting on the preprocessed data of the non-random missing type to obtain a randomly weighted preprocessed data matrix; perform singular value decomposition on the randomly weighted preprocessed data matrix to obtain a first diagonal matrix containing singular values; The decomposition module is specifically further configured to: subtract the set threshold from all the singular values in the first diagonal matrix to obtain an intermediate matrix; retain data of a first preset ratio in the intermediate matrix to obtain the second diagonal matrix; wherein the first preset ratio is the percentage of the sum of the retained singular values in the total singular values; The imputation module is specifically configured to: determine the missing position of the missing data according to the missing data matrix; determine the missing type of the missing data according to the preprocessed data matrix; determine the missing position and the missing type of the missing data in the first approximation matrix through the missing position and the missing type of the missing data; perform imputation on the corresponding missing position in the first approximation matrix by using a corresponding imputation method according to the missing type of the missing data in the first approximation matrix to determine the second approximation matrix; The imputation module is specifically further configured to: in the case that the missing data is of a random missing type, fill the missing data with a randomly selected value from the normal distribution data; in the case that the missing data is of a non-random missing type and the value at the corresponding position of the missing data in the first approximation matrix is higher than the minimum value of the non-missing metabolites in the original data, fill the missing data with a first weight value at the corresponding position of the preprocessed data matrix and a second weight value at the corresponding position of the first approximation matrix; in the case that the missing data is of a non-random missing type and the value of the missing data at the corresponding position in the first approximation matrix is not higher than the minimum value of the non-missing metabolites in the original data, fill the missing data with the value at the corresponding position in the first approximation matrix.

5. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the steps of the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Feature sampling-based data restoration method and device and related equipment

    CN115543991A

  • Metabonomics data processing method, device and equipment and readable storage medium

    CN118212994A