Single-cell rna sequencing data imputation method based on robust non-negative matrix factorization

By using an iterative algorithm optimized by robust nonnegative matrix factorization and semi-quadratic theory, the problem of noise influence in single-cell RNA sequencing data interpolation was solved, achieving more accurate data interpolation and cell state inference, and improving the accuracy of single-cell RNA sequencing data analysis.

CN117238372BActive Publication Date: 2025-12-16YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311013281.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-11
Publication Date
2025-12-16
Estimated Expiration
2043-08-11

AI Technical Summary

Technical Problem

Existing single-cell RNA sequencing data interpolation methods are susceptible to noise, leading to inaccurate interpolation results and an inability to effectively reconstruct lineage trajectories and infer the differentiation and progenitor status of single cells.

Method used

A robust nonnegative matrix factorization-based approach is adopted, using the C-loss loss function to evaluate noise error and the least squares loss function to evaluate true expression. The objective function is optimized by combining semi-quadratic theory, and the model is optimized through iterative algorithms to obtain cell and gene feature matrices, thereby preventing overfitting and improving interpolation accuracy.

Benefits of technology

It improved the accuracy and rationality of data interpolation, enhanced the precision of cell type identification and pseudo-time inference, and optimized the analysis effect of single-cell RNA sequencing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117238372B_ABST
    Figure CN117238372B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of single cell RNA sequencing, and particularly relates to a single cell RNA sequencing data interpolation method based on robust non-negative matrix factorization. The single cell RNA sequencing data interpolation method based on robust non-negative matrix factorization obtains optimal parameters of a cell feature matrix W and a gene feature matrix H by using a target function of a scRNA-seq data interpolation method based on robust non-negative matrix factorization, and then predicts the interpolated cell gene expression data by using a scRNMF model. The target function includes two loss functions, namely a C-loss loss function and a least square loss function. The scRNA-seq data interpolation method based on robust non-negative matrix factorization is hereinafter referred to as scRNMF. The method provided in the application solves the target function by training, determines the scRNMF model by using the solving result, and predicts the result by using the determined scRNMF model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of single-cell RNA sequencing technology, specifically relating to a single-cell RNA sequencing data interpolation method based on robust non-negative matrix factorization. Background Technology

[0002] Single-cell RNA sequencing (scRNA-seq) is a high-throughput technology for analyzing gene expression in individual cells, providing valuable information about cellular heterogeneity and function. However, scRNA-seq data often suffers from missing data or low quality, necessitating imputation methods to fill in these missing values. The goal of imputation methods is to infer missing values ​​from existing data to more accurately describe cellular expression profiles. Matrix factorization-based single-cell RNA data imputation is a commonly used technique, representing scRNA-seq data as a product of low-rank matrices and using matrix factorization to estimate missing values. However, this matrix factorization method is susceptible to noise, leading to inaccurate imputation results.

[0003] Cell clustering is one of the most important applications for scRNA-seq data, and a series of clustering algorithms have been developed for this purpose. For example, PCA dimensionality reduction + K-means clustering is a popular single-cell clustering scheme, but it cannot solve the noise problem in scRNA-seq data.

[0004] A common task in scRNA-seq data analysis is reconstructing lineage trajectories and inferring the differentiation and progenitor states of single cells. For example, the Monocle2 package performs differential expression and time-series analysis on single-cell expression data. It classifies individual cells according to the progression of biological processes. However, Monocle2 does not perform missing imputation for data reprocessing.

[0005] Therefore, how to more accurately and reasonably interpolate scRNA-seq data is one of the urgent problems to be solved. Summary of the Invention

[0006] The purpose of this invention is to provide a single-cell RNA sequencing data interpolation method based on robust non-negative matrix factorization. This method uses the correlation entropy-induced metric loss to replace the least squares error loss for noise, resulting in more accurate and reasonable interpolated data.

[0007] To achieve the above-mentioned objectives, the technical solution of the present invention is as follows:

[0008] A robust nonnegative matrix factorization-based method for single-cell RNA sequencing data interpolation is proposed. This method utilizes the objective function of a robust nonnegative matrix factorization-based scRNA-seq data interpolation method to obtain the optimal parameters of the cell feature matrix W and the gene feature matrix H, respectively. Then, the scRNMF model is used to predict the interpolated cell gene expression data.

[0009] The objective function includes two loss functions: the C-loss function and the least squares loss function.

[0010] The robust and nonnegative matrix factorization-based scRNA-seq data interpolation method is referred to as scRNMF.

[0011] The method provided by this invention solves the objective function through training, determines the scRNMF model using the solution results, and uses the determined scRNMF model to predict the results.

[0012] This invention classifies and evaluates noise and missing gene expression values ​​in the original expression matrix. For zero values ​​in the original expression matrix, which are generated by sequencing noise, C-loss is used to evaluate the error. For non-zero values, which represent true gene expression, least squares loss is used to evaluate the error. Therefore, the loss function of this method consists of two loss functions: C-loss and least squares loss, which is both robust to noise and can fit the true gene expression well.

[0013] In the aforementioned single-cell RNA sequencing data interpolation method based on robust nonnegative matrix factorization, given a single-cell gene expression matrix X∈R G×C Obtain the cell feature matrix W∈R G×k Gene feature matrix H∈R k×C , where G and C represent the number of cells and genes, respectively, and k is the dimension of the potential features of cells and genes.

[0014] In the above-mentioned single-cell RNA sequencing data interpolation method based on robust nonnegative matrix factorization, the scRNMF model expression is as follows:

[0015]

[0016] The objective function includes, in sequence, a loss function, a regularization term, and a regularization factor. The regularization term includes two regularization terms that respectively constrain gene factor W and cytokine H.

[0017] Since there are no negative numbers in the representation matrix, non-negativity restrictions are imposed on W and H.

[0018] In this invention, to prevent the information in the original expression matrix from being decomposed and affecting the effective expression of the potential representation of cells and genes, two regularization terms and a regularization factor are set to constrain gene factor W and cytokine H to prevent overfitting of the loss function.

[0019] In the above-mentioned single-cell RNA sequencing data interpolation method based on robust nonnegative matrix factorization, the objective function of the scRNA-seq data interpolation method based on robust nonnegative matrix factorization is as shown in formula (2):

[0020]

[0021] Where K G It is a gene similarity matrix, K C Let X be a cell similarity matrix, (g, c) represents the matrix index in the g-th row and c-th column, Xgc is the element in the g-th row and c-th column of the X matrix, Wgi is the element in the g-th row and i-th column of the W matrix, and Hic is the element in the i-th row and c-th column of the H matrix. is the Frobenius norm, α and β are hyperparameters controlling the importance of the corresponding regularization terms in the objective function, and k is the dimension of the potential features of cells and genes; l c It is the C-loss function.

[0022] Furthermore, the C-loss function is defined as follows:

[0023]

[0024] Preferably, W and H are respectively composed of gene similarity matrix K G and cell similarity matrix K C Constraints, the K G and K C The definitions are as follows:

[0025]

[0026]

[0027] The above-mentioned single-cell RNA sequencing data interpolation method based on robust nonnegative matrix factorization also includes an optimization objective function formula (2) based on semi-quadratic theory.

[0028] Get v gc The calculation formula is as follows:

[0029]

[0030] Where σ is a pre-defined hyperparameter;

[0031] The objective function formula (2) is rewritten as formula (14):

[0032]

[0033] Where k, σ, α, β, and λ are pre-defined hyperparameters, W T It is the transpose of matrix W, H T It is the transpose of matrix H, e gc The definition is as follows:

[0034]

[0035] In order to reduce the amount of computation, this invention uses semi-quadratic theory to solve and derive the non-convex objective function formula (2).

[0036] First, define a convex function:

[0037] g(v)=-vlog(-v)+v (20)

[0038] Where v < 0, v is a variable that is needed when solving the problem, but has no particular biological significance.

[0039] The conjugate function of formula (20) is:

[0040]

[0041] in,

[0042] g′(v)=uv-g(v)=uv+vlog(-v)-v (22)

[0043] Taking the derivative of formula (22) and setting it to zero, we get:

[0044] v = -exp(-u) < 0 (23)

[0045] Substituting formula (23) into formula (21), we get:

[0046] g * (u)=exp(-u) (24)

[0047] Combining formulas (24), (21), and (22), we obtain:

[0048]

[0049] in,

[0050]

[0051] The maximum value based on the semi-quadratic theory formula (25) can be obtained from the following equation:

[0052]

[0053] Therefore, formula (2) can be rewritten as:

[0054]

[0055] In the above-mentioned single-cell RNA sequencing data interpolation method based on robust nonnegative matrix factorization, the objective function optimization formula (2) based on semi-quadratic theory also includes optimizing formula (14) using an iterative algorithm. The specific steps are as follows:

[0056] First, fix W and H, and use formulas (29) and (14) to solve for v. gc ;

[0057] The second step is to use v gc Using the Karush-Khun-Tucker conditions and H, solve for W;

[0058] The third step is to use v gc Using the Karush-Khun-Tucker conditions, we can solve for H, along with W.

[0059] Furthermore, the objective function optimization formula (2) based on semi-quadratic theory also includes formula (14) optimized using an iterative algorithm:

[0060] The formula for calculating the optimal parameters of W is (23), as follows:

[0061]

[0062] The formula for calculating the optimal parameters of H is (24), as follows:

[0063]

[0064] Where α, β, and λ are predefined hyperparameters; M and P are defined as follows:

[0065]

[0066]

[0067] The iterative algorithm consists of the following three steps:

[0068] First, fix W and H, and derive formula (7) based on semi-quadratic theory.

[0069]

[0070] Where σ is a hyperparameter, a constant that is set manually;

[0071] Using formula (27), the solution to formula (34) is:

[0072]

[0073] The second step is to give vv gc With H, formula (7) can be rewritten as:

[0074]

[0075] Where · represents the element-wise smart product symbol, and M and P are defined as follows:

[0076]

[0077]

[0078] Formula (36) satisfies the Karush-Khun-Tucker (KKT) conditions:

[0079]

[0080] The gradient of formula (36) is

[0081]

[0082] Substituting formula (40) into formula (39), we get:

[0083]

[0084] Rewriting formula (41), it is easy to obtain

[0085]

[0086] The third step is to give vv gc And W, solve for H. Similar to the second step, using the KKT conditions, it is easy to obtain:

[0087]

[0088] Formulas (23) and (24) are functions of W and H. The optimal W and H are determined using the corresponding functions, i.e., the optimal parameters in formulas (23) and (24). Then, formulas (23) and (24) with the optimal parameters are substituted into (1) for interpolation. The mathematical formulas listed above are mathematical derivations for solving the objective function (2) and have no biological significance.

[0089] The single-cell RNA sequencing data interpolation method based on robust nonnegative matrix factorization described above also includes an iterative algorithm for the scRNMF model. The pseudocode for this iterative algorithm is as follows:

[0090] Algorithm: Iterative Algorithm for scRNMF Model

[0091] Input: Original representation matrix X; parameters k, σ, α, β, λ;

[0092] Output: Interpolation matrix

[0093] Initialization: W, H = NNDSVD(X);

[0094] Calculate the cell similarity matrix K C ;

[0095] Calculate the gene similarity matrix K G ;

[0096] If convergence fails, loop through the following:

[0097] Update v using formula (13) gc v;

[0098] Update W using formula (23);

[0099] Update H using formula (22);

[0100] Check the convergence status;

[0101] End the loop

[0102] Return after convergence.

[0103] This invention optimizes the scRNMF model using an iterative algorithm based on half-quadratic theory, which effectively reduces the computational cost of the objective function while ensuring its convergence, thereby further guaranteeing the accuracy and rationality of the model's imputation data.

[0104] Compared with the prior art, the beneficial effects of the present invention are reflected in:

[0105] (1) This invention classifies and evaluates noise and missing values ​​of true gene expression in the original expression matrix. For zero values ​​in the original expression matrix, which are generated by sequencing technology noise, C-loss is used to evaluate the error. For non-zero values, which are true gene expression, least squares loss is used to evaluate the error. Therefore, the loss of this method consists of two loss functions, C-loss and least squares loss, which can be robust to noise and fit the true gene expression well.

[0106] (2) In order to avoid the information in the original expression matrix from being decomposed and affecting the effective expression of the potential representation of cells and genes, the present invention sets two regularization terms and regularization factors to constrain gene factor W and cell factor H to prevent overfitting of the loss function.

[0107] (3) This invention optimizes the scRNMF model by using an iterative algorithm based on half-quadratic theory, which effectively reduces the amount of computation of the objective function while ensuring the convergence of the objective function, thereby further ensuring the accuracy and rationality of the model interpolation data. Attached Figure Description

[0108] Figure 1 This is a graphical representation of the objective function of the scRNA-seq data interpolation method based on robust and nonnegative matrix factorization of the present invention.

[0109] Figure 2 This is a comparative schematic diagram illustrating the enhanced cell trajectory inference achieved by the robust and non-negative matrix factorization-based scRNA-seq data interpolation method of this invention. Detailed Implementation

[0110] The following detailed embodiments illustrate the technical solution of the present invention.

[0111] Example 1

[0112] This embodiment provides a robust and nonnegative matrix factorization-based scRNA-seq data interpolation method, namely scRNMF.

[0113] Given a single-cell gene expression matrix X∈R G×C Obtain the cell feature matrix W∈R G×k Gene feature matrix H∈R k×C , where G and C represent the number of cells and genes, respectively, and k is the dimension of the potential features of cells and genes.

[0114] The scRNMF model expression is as follows:

[0115]

[0116] The objective function of scRNMF is used to solve for W and H. After solving, the predicted result is obtained using formula (1). The graphical representation of the scRNMF objective function is as follows: Figure 1 As shown, the specific formula is as follows:

[0117]

[0118] Where K G It is a gene similarity matrix, K C Let X be a cell similarity matrix, (g, c) represents the matrix index in the g-th row and c-th column, Xgc is the element in the g-th row and c-th column of the X matrix, Wgi is the element in the g-th row and i-th column of the W matrix, and Hic is the element in the i-th row and c-th column of the H matrix. is the Frobenius norm, α and β are hyperparameters controlling the importance of the corresponding regularization terms in the objective function, and k is the dimension of the potential features of cells and genes; l c It is the C-loss function.

[0119] The C-loss function is defined as follows:

[0120]

[0121] In formula (2), the first and second terms are loss functions used to fit the original expression matrix. Zero values ​​in the original expression matrix represent missing values ​​and noise caused by sequencing technology, so C-loss is used to evaluate the error. Non-zero values ​​represent the true expression of the gene, and least squares loss is used to evaluate the error. Therefore, the loss in this method consists of both C-loss and least squares loss functions, which is robust to noise and can also fit the true gene expression well.

[0122] Because the original expression matrix may contain important information, further decomposition may fail to effectively express the potential representations of cells and genes. Therefore, a third and fourth regularization term is introduced into formula (2). Here, W and H are respectively derived from the gene similarity matrix K. G Cell similarity matrix K C Constraints. K G and K C The definition is as follows:

[0123]

[0124]

[0125] For the non-convex objective function formula (2), in order to reduce the amount of computation, semi-quadratic theory is used for the derivation. First, a convex function is defined:

[0126] g(v)=-vlog(-v)+v (49)

[0127] Where v < 0, v is a variable that is needed when solving the problem, but has no particular biological significance.

[0128] The conjugate function of formula (20) is:

[0129]

[0130] in,

[0131] g′(v)=uv-g(v)=uv+vlog(-v)-v (51)

[0132] Taking the derivative of formula (22) and setting it to zero, we get:

[0133] v = -exp(-u) < 0 (52)

[0134] Substituting formula (23) into formula (21), we get:

[0135] g * (u)=exp(-u) (53)

[0136] Combining formulas (24), (21), and (22), we obtain:

[0137]

[0138] in,

[0139]

[0140] The maximum value based on the semi-quadratic theory formula (25) can be obtained from the following equation:

[0141]

[0142] Where σ is a pre-defined hyperparameter;

[0143] Therefore, formula (2) can be rewritten as:

[0144]

[0145] Next, we use an iterative algorithm to optimize formula (14). The three steps of this iterative algorithm are as follows:

[0146] First, fix W and H, and derive formula (7) based on semi-quadratic theory.

[0147]

[0148] Where σ is a hyperparameter, a constant that is set manually;

[0149] Using formula (27), the solution to formula (34) is:

[0150]

[0151] The second step is to give v gc With H, formula (7) can be rewritten as:

[0152]

[0153] Where · represents the element-wise smart product symbol, and M and P are defined as follows:

[0154]

[0155]

[0156] Formula (36) satisfies the Karush-Khun-Tucker (KKT) conditions:

[0157]

[0158] The gradient of formula (36) is

[0159]

[0160] Substituting formula (40) into formula (39), we get:

[0161]

[0162] Rewriting formula (41), it is easy to obtain

[0163]

[0164] The third step is to give v gc And W, solve for H. Similar to the second step, using the KKT conditions, it is easy to obtain:

[0165]

[0166] Formulas (23) and (24) are functions of W and H. The optimal W and H are determined using the corresponding functions, i.e., the optimal parameters in formulas (23) and (24). Then, formulas (23) and (24) with the optimal parameters are substituted into (1) for interpolation. The mathematical formulas listed above are mathematical derivations for solving the objective function (2) and have no biological significance.

[0167] Summarizing the above steps, the pseudocode for the iterative algorithm is as follows:

[0168] Algorithm: Iterative Algorithm for scRNMF Model

[0169] Input: Original representation matrix X; parameters k, σ, α, β, λ;

[0170] Output: Interpolation matrix

[0171] Initialization: W, H = NNDSVD(X);

[0172] Calculate the cell similarity matrix K C ;

[0173] Calculate the gene similarity matrix K G ;

[0174] If convergence fails, loop through the following:

[0175] Update v using formula (13) gc ;

[0176] Update W using formula (23);

[0177] Update H using formula (22);

[0178] Check the convergence status;

[0179] End the loop

[0180] Return after convergence.

[0181] Comparative Example 1

[0182] We compared scRNMF with raw scRNA-seq data and several other popular scRNA-seq data missing value imputation methods (including AutoClass, DCA, scGCL, Magic, SAVER, scImpute, and CMF-Impute). We clustered these expression matrices using PCA+K-means, with ARI and NMI as evaluation metrics. The data used were from five publicly available scRNA-seq datasets (Buttner, Usoskin, Lake, Zeisel, and Pollen).

[0183] Other methods for handling missing data imputation impute the five scRNA-seq datasets according to their respective methods. In this approach, the objective function (2) is used to obtain the optimal parameters of the cell feature matrix W and gene feature matrix H for the five scRNA-seq datasets, and then the five scRNA-seq datasets are imputed using formula (1). The final imputation results are shown in Table 1.

[0184] Table 1 shows the clustering results on five real-world datasets.

[0185]

[0186] Clearly, based on the ARI and NMI measurements, scRNMF yields the best results. This indicates that imputation of missing values ​​using scRNMF can significantly improve the clustering accuracy of the PCA+K-means method. The scRNMF provided in this embodiment enhances cell type identification.

[0187] Comparative Example 2

[0188] scRNMF was integrated into Monocle2, and its performance on the Deng dataset in pseudo-time inference was compared. Consistency between time labels and pseudo-time order was measured using the pseudo-timing score (POS) and Kendall's rank correlation score (KOR), with results as follows: Figure 2 As shown.

[0189] Depend on Figure 2 It can be seen that scRNMF achieves optimal POS and KOR. This demonstrates its advantage over Monocle2 in improving pseudo-temporal inference, namely, scRNMF enhances the inference of cell trajectories.

[0190] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.

Claims

1. A method for interpolating single-cell RNA sequencing data based on robust nonnegative matrix factorization, characterized in that: The optimal parameters of the cell feature matrix W and gene feature matrix H were obtained using the objective function of a robust and nonnegative matrix factorization-based scRNA-seq data interpolation method. Then, the scRNMF model was used to predict the interpolated cell gene expression data. The objective function includes two loss functions: the C-loss function and the least squares loss function.

2. The single-cell RNA sequencing data interpolation method based on robust nonnegative matrix factorization as described in claim 1, characterized in that: Given a single-cell gene expression matrix X∈R G×c Obtain the cell feature matrix W∈R G×k Gene feature matrix H∈R k×C , where G and C represent the number of cells and genes, respectively, and k is the dimension of the potential features of cells and genes.

3. The single-cell RNA sequencing data interpolation method based on robust nonnegative matrix factorization as described in claim 2, characterized in that: The scRNMF model expression is as follows: The objective function includes, in sequence, a loss function, a regularization term, and a regularization factor. The regularization term includes two regularization terms that respectively constrain gene factor W and cytokine H.

4. The single-cell RNA sequencing data interpolation method based on robust nonnegative matrix factorization as described in claim 3, characterized in that: The objective function of the robust and nonnegative matrix factorization-based scRNA-seq data interpolation method is shown in Equation (2): subject to:W≥0,H≥0. (2) Where K G It is a gene similarity matrix, K C Let X be a cell similarity matrix, (g, c) represents the matrix index in the g-th row and c-th column, Xgc is the element in the g-th row and c-th column of the X matrix, Wgi is the element in the g-th row and i-th column of the W matrix, and Hic is the element in the i-th row and c-th column of the H matrix. is the Frobenius norm, α and β are hyperparameters controlling the importance of the corresponding regularization terms in the objective function, and k is the dimension of the potential features of cells and genes; l c It is the C-loss function.

5. The single-cell RNA sequencing data interpolation method based on robust nonnegative matrix factorization as described in claim 4, characterized in that: The C-loss function is defined as follows:

6. The single-cell RNA sequencing data interpolation method based on robust nonnegative matrix factorization as described in claim 4, characterized in that: The W and H are respectively derived from the gene similarity matrix K. G and cell similarity matrix K C Constraints, the K G and K C The definitions are as follows:

7. The single-cell RNA sequencing data interpolation method based on robust nonnegative matrix factorization as described in claim 4, characterized in that, It also includes formula (2) for optimizing the objective function based on semi-quadratic theory. Get v gc The calculation formula is as follows: Where σ is a pre-defined hyperparameter; The objective function formula (2) is rewritten as formula (14): subject to: W≥0, H≥0. (7) Where k, σ, α, β, and λ are pre-defined hyperparameters, W T It is the transpose of matrix W, H T It is the transpose of matrix H, e gc The definition is as follows:

8. The single-cell RNA sequencing data interpolation method based on robust nonnegative matrix factorization as described in claim 7, characterized in that: The objective function optimization formula (2) based on semi-quadratic theory also includes optimizing formula (14) using an iterative algorithm. The specific steps are as follows: First, fix W and H, and use formulas (8) and (14) to solve for v. gc ; The second step is to use v gc Using the Karush-Khun-Tucker conditions and H, solve for W; The third step is to use v gc Using the Karush-Khun-Tucker conditions, we can solve for H, along with W.

9. The single-cell RNA sequencing data interpolation method based on robust nonnegative matrix factorization as described in claim 8, characterized in that: The objective function optimization formula (2) based on semi-quadratic theory also includes the optimization formula (14) using an iterative algorithm: The formula for calculating the optimal parameters of W is (23), as follows: The formula for calculating the optimal parameters of H is (24), as follows: Among them, K c Yes; α, β, and λ are predefined hyperparameters; M and P are defined as follows:

10. The single-cell RNA sequencing data interpolation method based on robust nonnegative matrix factorization as described in claim 7, characterized in that: It also includes an iterative algorithm for the scRNMF model, the pseudocode of which is as follows: Algorithm: Iterative Algorithm for scRNMF Model Input: Original representation matrix X; parameters k, σ, α, β, λ; Output: Interpolation matrix Initialization: W, H = NNDSVD(X); Calculate the cell similarity matrix K C ; Calculate the gene similarity matrix K G ; If convergence fails, loop through the following: Update v using formula (13) gc ; Update W using formula (23); Update H using formula (22); Check the convergence status; End the loop Return after convergence

Citation Information

Patent Citations

  • Hyperspectral unmixing method and system based on entropy regular non-negative matrix factorization model

    CN111914893A

  • Single cell RNA-SEQ data processing

    CN114424287A