Classification method and device for male infertility and method and device for constructing a related prediction model

CN122422947APending Publication Date: 2026-07-17SHENZHEN HUADA GENE INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN HUADA GENE INST
Filing Date
2023-12-28
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

The existing male infertility diagnosis technology has problems with invasiveness and low detection accuracy, especially the diagnosis of non-obstructive azoospermia is difficult to accurately carry out.

Method used

By using the contribution ratio of testicularly associated single-cell types in free RNA as characterized, a male infertility prediction model was constructed, a machine learning algorithm was used for non-invasive diagnosis, and combining the deconvolution method and the testicularly single-cell transcription data set, the contribution ratio of each cell type was calculated.

Benefits of technology

It realizes non-invasive and accurate prediction of male azoospermia and its subtypes, and improves the sensitivity and specificity of diagnosis, especially the diagnostic accuracy of non-obstructive azoospermia reaches 0.88, stability and ease of operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122422947A_ABST
    Figure CN122422947A_ABST
Patent Text Reader

Abstract

本发明提供了一种男性不育症的分类方法及装置和构建相关预测模型的方法和装置。该分类方法包括:利用精浆游离RNA中睾丸相关单细胞类型的贡献比例为特征构建男性不育症预测模型;将待测样本的精浆游离RNA中睾丸相关单细胞类型的贡献比例输入男性不育症预测模型,获得待测样本的分类预测结果。该分类方法能精确预测未知样本男性无精症及其亚型,相比之前的方法在操作上更加便捷。且该方法可在临床上的睾丸穿刺之前对无精症的亚型进行预测,实现无创检测和诊断的目的。
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for classifying male infertility and method and device for constructing related prediction model Technical Field

[0001] The present invention relates to the field of male infertility, and in particular to a method and device for classifying male infertility and a method and device for constructing a related prediction model. Background Art

[0002] Approximately 10%-15% of couples worldwide experience infertility, with male factors accounting for approximately 50% of these cases. According to the World Health Organization, azoospermia, defined as the absence of sperm after centrifugation in at least two routine semen examinations, is a serious form of male infertility. Azoospermia includes obstructive azoospermia (OA) and non-obstructive azoospermia (NOA), with NOA accounting for approximately 70% of cases of azoospermia.

[0003] For many years, traditional male infertility diagnosis has relied on methods such as sperm motility analysis and sperm counting. Sperm count is significantly affected by external factors, such as smoking and alcohol consumption, which can cause significant fluctuations in results. Sperm motility analysis requires expensive equipment and is significantly correlated with the subject's recent sexual activity, resulting in low stability. Currently, the pathological diagnosis of NOA relies primarily on testicular biopsy, an invasive and inaccurate procedure involving puncturing testicular tissue with a needle, followed by HE staining and microscopic examination. Therefore, a new, non-invasive, and accurate method for diagnosing azoospermia and NOA is highly desirable.

[0004] Summary of the Invention

[0005] The main purpose of the present invention is to provide a method and device for classifying male infertility and a method and device for constructing a related prediction model to solve the problems of invasiveness or low detection accuracy of diagnostic solutions in the prior art.

[0006] In order to achieve the above-mentioned purpose, according to the first aspect of the present invention, a classification method for male infertility is provided, which classification method comprises: constructing a male infertility prediction model using the contribution ratio of testicular-related single cell types in seminal plasma free RNA as a feature; inputting the contribution ratio of testicular-related single cell types in seminal plasma free RNA of the sample to be tested into the male infertility prediction model to obtain a classification prediction result of the sample to be tested.

[0007] Furthermore, constructing a male infertility prediction model using the contribution ratio of testicular single cell types in seminal plasma free RNA as a feature includes: obtaining the contribution ratio of various testicular-related single cell types in the seminal plasma free mRNA of each sample in the training set, wherein the training set is selected from set A or set B, set A includes a fertile male sample group and an azoospermic male sample group, and set B includes an obstructive azoospermia male sample group and a non-obstructive azoospermia male sample group; using the contribution ratio as the input feature of each sample for machine learning training to obtain a male infertility prediction model.

[0008] Furthermore, obtaining the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample in the training set includes: obtaining the expression profile of the seminal plasma free mRNA of each sample in the training set; using the testicular single-cell transcriptional dataset to obtain a matrix of cell type-specific gene expression; and using a deconvolution method to calculate the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample based on the expression profile and the matrix of cell type-specific gene expression.

[0009] Furthermore, the testis-related single cell types include the following 12 types: sperm, late primary spermatocytes, round spermatids, Sertoli cells, spermatogonial stem cells, macrophages, differentiating spermatogenic cells, vascular endothelial cells, elongated spermatids, seminiferous tubule muscle cells, interstitial cells and early primary spermatocytes.

[0010] Furthermore, obtaining the expression profile of the seminal plasma free mRNA of each sample in the training set includes: obtaining the sequencing data of the seminal plasma free mRNA of each sample in the training set; performing quality control on the sequencing data to obtain quality-controlled data; aligning the quality-controlled data to the human transcriptome and quantifying it to obtain the expression profile; preferably, the quality control includes at least one of the following: 1) cutting junctions; 2) removing reads with a quality lower than a quality threshold; 3) removing reads <17 bp in length; 4) removing rRNA sequences, vt RNA and Y RNA sequences; preferably, the quality-controlled data are aligned to the human transcriptome in the following order: 1) miRNA; 2) tRNA and piRNA; 3) mRNA and lncRNA; 4) remaining RNA.

[0011] Furthermore, while calculating the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample, the classification method also includes: calculating the p-value, correlation and residual of the deconvolution of each sample; and evaluating the reliability of the deconvolution results of each sample based on the p-value, correlation and residual.

[0012] In order to achieve the above-mentioned purpose, according to the second aspect of the present invention, a classification device for male infertility is provided, which includes: a model construction module, which is configured to use the contribution ratio of testicular-related single cell types in seminal plasma free RNA as a feature to construct a male infertility prediction model; a classification module, which is configured to input the contribution ratio of testicular-related single cell types in seminal plasma free RNA of the sample to be tested into the male infertility prediction model to obtain a classification prediction result of the sample to be tested.

[0013] Furthermore, the model building module includes: an acquisition module, which is configured to obtain the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample in the training set, wherein the training set is selected from set A or set B, set A includes a fertile male sample group and an azoospermia male sample group, and set B includes an obstructive azoospermia male sample group and a non-obstructive azoospermia male sample group; a machine learning module, which is configured to use the contribution ratio as the input feature of each sample for machine learning training to obtain an azoospermia prediction model.

[0014] Furthermore, the acquisition module includes: a first acquisition unit, which is configured to obtain the expression spectrum of seminal plasma free mRNA of each sample in the training set; a second acquisition unit, which is configured to use the testicular single-cell transcription data set to obtain a matrix of cell type-specific gene expression; a calculation unit, which is configured to calculate the contribution ratio of various testicular-related single cell types in the seminal plasma free mRNA of each sample using a deconvolution method based on the expression spectrum and the matrix of cell type-specific gene expression.

[0015] Furthermore, the testis-related single cell types include the following 12 types: sperm, late primary spermatocytes, round spermatids, Sertoli cells, spermatogonial stem cells, macrophages, differentiating spermatogenic cells, vascular endothelial cells, elongated spermatids, seminiferous tubule muscle cells, interstitial cells and early primary spermatocytes.

[0016] Furthermore, the first acquisition unit includes: a sequencing data acquisition module, which is configured to obtain the sequencing data of the seminal plasma free mRNA of each sample in the training set; a quality control module, which is configured to perform quality control on the sequencing data to obtain quality-controlled data; a quantification module, which is configured to compare the quality-controlled data to the human transcriptome and perform quantification to obtain an expression profile; preferably, quality control includes at least one of the following: 1) cutting junctions; 2) removing reads with a quality lower than a quality threshold; 3) removing reads <17bp in length; 4) removing rRNA sequences, vt RNA and Y RNA sequences; preferably, the quality-controlled data are compared to the human transcriptome in the following order: 1) miRNA; 2) tRNA and piRNA; 3) mRNA and lncRNA; 4) remaining RNA.

[0017] Furthermore, the classification device also includes: a calculation submodule, which is configured to calculate the p-value, correlation and residual of each sample deconvolution; and an evaluation module, which is configured to evaluate the reliability of each sample deconvolution result based on the p-value, correlation and residual.

[0018] To achieve the above-mentioned object, according to a third aspect of the present invention, a method for constructing a male infertility prediction model is provided, the method comprising: obtaining the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample in a training set, wherein the training set is selected from set A or set B, set A includes a fertile male sample group and a male sample group with azoospermia, and set B includes a male sample group with obstructive azoospermia and a male sample group with non-obstructive azoospermia; using the contribution ratio as an input feature of each sample for machine learning training to obtain a male infertility prediction model.

[0019] Furthermore, obtaining the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample in the training set includes: obtaining the expression profile of the seminal plasma free mRNA of each sample in the training set; using the testicular single-cell transcriptional dataset to obtain a matrix of cell type-specific gene expression; and using a deconvolution method to calculate the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample based on the expression profile and the matrix of cell type-specific gene expression.

[0020] Furthermore, the testis-related single cell types include the following 12 types: sperm, late primary spermatocytes, round spermatids, Sertoli cells, spermatogonial stem cells, macrophages, differentiating spermatogenic cells, vascular endothelial cells, elongated spermatids, seminiferous tubule muscle cells, interstitial cells and early primary spermatocytes.

[0021] Furthermore, obtaining the expression profile of the seminal plasma free mRNA of each sample in the training set includes: obtaining the sequencing data of the seminal plasma free mRNA of each sample in the training set; performing quality control on the sequencing data to obtain quality-controlled data; aligning the quality-controlled data to the human transcriptome and quantifying it to obtain the expression profile; preferably, the quality control includes at least one of the following: 1) cutting junctions; 2) removing reads with a quality lower than a quality threshold; 3) removing reads <17 bp in length; 4) removing rRNA sequences, vt RNA and Y RNA sequences; preferably, the quality-controlled data are aligned to the human transcriptome in the following order: 1) miRNA; 2) tRNA and piRNA; 3) mRNA and lncRNA; 4) remaining RNA.

[0022] Furthermore, while calculating the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample, the method also includes: calculating the p-value, correlation and residual of the deconvolution of each sample; and evaluating the reliability of the deconvolution results of each sample based on the p-value, correlation and residual.

[0023] In order to achieve the above-mentioned purpose, according to the fourth aspect of the present invention, there is provided a device for constructing a male infertility prediction model, which includes: an acquisition module, configured to obtain the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample in a training set, wherein the training set is selected from set A or set B, set A includes a fertile male sample group and an azoospermia male sample group, and set B includes an obstructive azoospermia male sample group and a non-obstructive azoospermia male sample group; a machine learning module, configured to use the contribution ratio as an input feature of each sample for machine learning training to obtain an azoospermia prediction model.

[0024] According to a fifth aspect of the present invention, a non-transitory computer-readable storage medium is provided, the storage medium including a stored program, wherein, when the program is executed, the device where the storage medium is located is controlled to execute any of the above-mentioned methods for classifying male infertility or any of the above-mentioned methods for constructing a male infertility prediction model.

[0025] According to one aspect of the present invention, a processor is provided, which is used to run a program, wherein when the program is run, any of the above-mentioned male infertility classification methods or any of the above-mentioned methods for constructing a male infertility prediction model is executed.

[0026] The technical solution of the present invention uses the contribution ratio of various testicular cell types to seminal plasma free RNA to construct a non-invasive diagnostic classification model for azoospermia or its subtypes. This model can accurately predict azoospermia and its subtypes in unknown male samples. Compared with previous methods, this method is more convenient in operation and has higher prediction accuracy. The method of the present application can also predict the subtype of azoospermia before the commonly used testicular puncture in clinical practice, achieving the purpose of non-invasive detection and diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0028] FIG1 is a schematic diagram showing a method for constructing a male infertility prediction model according to Example 1 of the present application;

[0029] FIG2 shows a schematic flow chart of a method for classifying male infertility according to Example 2 of the present application;

[0030] FIG3 shows a schematic flow chart of a method for classifying male infertility according to Example 3 of the present application;

[0031] FIG4A and FIG4B show the prediction results of a healthy and azoospermia classification prediction model provided in Example 4 of the present application.

[0032] 5A and 5B show the prediction results of a classification prediction model for non-obstructive azoospermia and obstructive azoospermia provided in Example 4 of the present application;

[0033] FIG6 shows a hardware structure block diagram of a terminal for a method for classifying male infertility provided by an embodiment of the present application.

[0034] FIG7 shows a schematic structural diagram of a device for constructing a male infertility prediction model according to Example 5 of the present application;

[0035] FIG8 shows a schematic structural diagram of a male infertility classification device provided according to Example 6 of the present application. DETAILED DESCRIPTION

[0036] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below with reference to the embodiments.

[0037] Explanation of terms:

[0038] Azoospermia: refers to the absence of sperm after centrifugation in at least two routine semen examinations.

[0039] Obstructive azoospermia (OA) refers to normal testicular spermatogenesis, but sperm cannot be discharged from the body due to obstruction of the vas deferens.

[0040] Non-obstructive azoospermia (NOA): refers to functional disorders of the testicles themselves, also known as primary azoospermia.

[0041] Existing technologies for diagnosing azoospermia and NOA include using changes in plasma-free RNA, seminal plasma-free RNA, specific RNA in semen exosomes, specific proteins in testicular tissue and serum, and changes in mRNA expression levels of specific genes in testes or testicular biopsy extracts. Technically, diagnostic kits, gene chips, immunoassays, and sequencing are currently the primary tools for diagnosing azoospermia and NOA using gene and protein expression as markers in materials such as semen, seminal plasma, testicular tissue, or testicular biopsy.

[0042] The main disadvantages of the above-mentioned existing technologies include: (1) Proteomics and metabolomics are more expensive and less mature than transcriptomics. (2) Using changes in gene or protein expression levels in testicles or testicular puncture extracts to diagnose azoospermia and NOA is invasive and may carry the risk of complications such as hematoma, infection, and testicular scar formation. In other words, the existing technologies are invasive and have low sensitivity and specificity. There is currently no effective solution to effectively improve this problem.

[0043] To improve this situation, in this application, the inventors attempted to use mature transcriptomics and readily available clinical information to replace relatively less mature technologies such as proteomics or metabolomics, and to achieve non-invasive diagnosis of azoospermia and its subtypes through non-invasive means.

[0044] It should be noted that the current prior art has no way to capture all seminal plasma free mRNA, and is limited to individual RNA analysis. In addition, the transcriptional data set of the relevant complete testicular single cell in the existing research is relatively small, and therefore the analysis of the single cell contribution ratio cannot be carried out using the full transcriptome. In the present application, the full transcriptional spectrum of mRNA can be obtained using the PLAM_seq sequencing technology developed before the applicant, and the seminal plasma free RNA is composed of multiple organs (including testis) secretion, wherein, the cell-specific mRNAs of each secretory organ are included, and the contribution degree information of each cell to the seminal plasma free RNA can be more perfectly reflected. In spermatogenesis disorder disease, the applicant has found that there are differences in the free mRNA expression spectrum between the fertile group and the azoospermia group, so it is further speculated that there are differences in the contribution ratios of the various cell types that secrete seminal plasma free mRNAs, that is, the cell type contribution ratio can be used as one of the features of spermatogenesis disorder disease. On this basis, in combination with the high-quality testicular single cell data set and deconvolution tracing method published in recent years, the design of the present application is proposed.

[0045] The following is the inventive concept of this application: Seminal plasma free RNA is secreted by multiple organs (primarily the testis, epididymis, prostate, and seminal vesicles), containing more than 20,000 mRNAs, including cell-specific mRNAs from these organs. This can reflect gene expression events in seminal plasma-derived cells and also provide a relatively comprehensive picture of the contribution of each cell type to seminal plasma free RNA. By establishing a method to calculate the contribution ratio of various testicular cell types to seminal plasma free RNA and combining it with a machine learning algorithm, a model can be constructed to achieve non-invasive diagnosis of azoospermia and its subtypes.

[0046] Based on the above-mentioned inventive ideas, the present application proposes for the first time, on the basis of clinical information, an improved scheme for constructing a classification prediction model for azoospermia and its subtypes using the contribution ratio of testicular single cells of seminal plasma free RNA as a feature, and further evaluates the prediction effect of the model according to the area under the curve (Area Under Curve, AUC) of the sensitivity and specificity evaluation. The results show that the improved scheme of the present application increases the AUC for diagnosing azoospermia to 1.00, and the AUC of NOA reaches 0.88. In addition, seminal plasma free RNA has the characteristics of non-invasive, sensitive, high stability and good accuracy, so the detection results of the improved scheme of the present application are stable, reliable, and easy to operate. On the basis of the above-mentioned research results, the applicant has proposed a series of protection schemes of the present application.

[0047] Example 1

[0048] This embodiment provides a method for constructing a male infertility prediction model, as shown in FIG1 , the method comprising the following steps:

[0049] S101, obtaining the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample in the training set, wherein the training set is selected from set A or set B, set A includes a fertile male sample group and a male sample group with azoospermia, and set B includes a male sample group with obstructive azoospermia and a male sample group with non-obstructive azoospermia;

[0050] S102: Using the contribution ratio as an input feature of each sample for machine learning training to obtain a male infertility prediction model.

[0051] In the above embodiments, depending on the specific types of clinical samples collected for machine learning training, whether they are fertile male sample groups and azoospermia male sample groups, or obstructive azoospermia male sample groups and non-obstructive azoospermia male sample groups, the same feature based on the contribution ratio of cell types is used as input for machine learning training to obtain a prediction model for predicting azoospermia or a prediction model for predicting non-obstructive azoospermia.

[0052] It should be noted that in the two parallel schemes of the above embodiments, the method of tracing the origin ratio of testicular single cells is used for the first time to construct a model for classifying and predicting diseases and subtypes, that is, the contribution ratio of each cell type in each sample in different training sets is used as an input feature for machine learning training for the first time. Among them, the method for obtaining the contribution ratio of each cell type in each sample is not limited, and can be based on the deconvolution method of SVR, the non-negative matrix decomposition (NMF) method or the non-negative least squares regression (NNLS) method. In this application, the deconvolution method based on SVR is preferred.

[0053] In the above embodiments, there are multiple classification algorithms suitable for machine learning training in the present application, including but not limited to random forest, logistic regression, decision tree or K nearest neighbor classification algorithm. The random forest algorithm is preferably used in the present application. In some preferred embodiments, multiple different prediction models can be obtained by using multiple different algorithms, and the prediction effects of the multiple different prediction models are further evaluated. The prediction model with the best prediction effect is selected as the final prediction model based on the evaluation effect.

[0054] In some preferred embodiments, the step of obtaining the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample in the training set includes: obtaining the expression profile of the seminal plasma free mRNA of each sample in the training set; using the testicular single-cell transcription data set to obtain a matrix of cell type-specific gene expression; and using a deconvolution method to calculate the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample based on the expression profile and the matrix of cell type-specific gene expression.

[0055] The step of "obtaining expression profiles of cell-free mRNA in seminal plasma for each sample in the training set" includes obtaining sequencing data (including high-throughput sequencing data or other sequencing data, or qPCR data) from each sample and performing quantification based on the obtained sequencing data to obtain expression profiles. This step is further illustrated below and specifically includes: performing quality control on the off-machine sequencing data of seminal plasma cfRNA from each sample, including trimming adapters, removing low-quality reads (quality threshold of 15, here referring to reads below this quality threshold), and removing reads <17 bp in length. (For example, using bowtie) rRNA sequences, vtRNA (another RNA type that is removed, similar to rRNA), and gamma RNA sequences are removed. The remaining reads are then aligned to the human transcriptome (in the order of miRNA, tRNA, and piRNA, then mRNA and lncRNA, and finally other RNAs). mRNA and lncRNA are quantified using RSEM (a typical transcript-based quantification method), and expression levels are then normalized using FPKM.

[0056] In some preferred embodiments, obtaining the expression profile of the seminal plasma free mRNA of each sample in the training set includes: obtaining the sequencing data of the seminal plasma free mRNA of each sample in the training set; performing quality control on the sequencing data to obtain post-quality control data; comparing the post-quality control data to the human transcriptome and quantifying it to obtain the expression profile.

[0057] In some preferred embodiments, quality control includes at least one of the following: 1) cutting junctions; 2) removing reads with a quality below a quality threshold; 3) removing reads <17 bp in length; 4) removing rRNA sequences, vtRNA and YRNA sequences;

[0058] It should be noted that the sequencing data quality control requirement for the off-machine data is strict: the off-machine data must have >10 million reads. Then, using the independently developed PALM-seq technology, the corresponding sequences in 1)-4) above are removed.

[0059] In some preferred embodiments, post-quality control data are aligned to the human transcriptome in the following order: 1) miRNA; 2) tRNA and piRNA; 3) mRNA and lncRNA; 4) remaining RNA. Transcriptome expression is quantified after alignment. In this application, no sample filtering is performed at this step. For scalability, a quality control standard of >10,000 mRNA detections can be set.

[0060] The exemplary steps for the above-mentioned "Using the testis single-cell transcriptional dataset to obtain a matrix of cell type-specific gene expression" are as follows: by researching the literature (Guo, J., EJ Grow, H. Mlcochova, et al. "The adult human testis transcriptional cell atlas." Cell Research (2018). 28(12): 1141-1157), a testis single-cell transcriptional dataset (see Table 2) was collected to sort out the cell types involved, and a cell type-specific gene expression matrix was generated using Cibersortx. Combined with the cell type-specific gene expression matrix, the contribution ratio of each cell type to the free mRNA in the seminal plasma of each sample was calculated using the SVR-based deconvolution method. At the same time, the p-value, correlation, and residual of the deconvolution of each sample are calculated to evaluate the reliability of the sample deconvolution results (the evaluation results here do not determine whether the sample is included, but evaluate the reliability of the method we use to diagnose azoospermia and non-obstructive azoospermia using the cell type contribution ratio. When evaluating the deconvolution results, there will be a corresponding p-value to reflect whether the contribution score has reference value. If the p-value is very high, it will affect the credibility of the scoring results. Evaluating reliability specifically refers to the statistical significance of the deconvolution results. The smaller the p-value, the higher the correlation and the smaller the residual, which means that the tracing effect is better, and the mixing matrix and the reference matrix provided for deconvolution are more appropriate. If the p-value is small and the effect is not bad, no additional processing is required).

[0061] Therefore, in some preferred embodiments, the testis-related single cell types include the following 12 types: sperm, late primary spermatocytes, round spermatids, supporting cells, spermatogonial stem cells, macrophages, differentiating spermatogenic cells, vascular wall endothelial cells, elongated spermatids, seminiferous tubule muscle cells, interstitial cells and early primary spermatocytes.

[0062] In some preferred embodiments, while calculating the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample, the method also includes: calculating the p-value, correlation and residual of the deconvolution of each sample; and evaluating the reliability of the deconvolution results of each sample based on the p-value, correlation and residual.

[0063] Example 2

[0064] This embodiment provides a classification method for male infertility, as shown in FIG2 , which includes the following steps:

[0065] S201, constructing a male infertility prediction model using the contribution ratio of testicular-related single cell types in seminal plasma free RNA as a feature;

[0066] S202, inputting the contribution ratio of testis-related single cell types in the seminal plasma free RNA of the sample to be tested into the male infertility prediction model to obtain a classification prediction result of the sample to be tested.

[0067] Seminal plasma free RNA is secreted by multiple organs (mainly testis, epididymis, prostate, seminal vesicle), wherein containing more than 20000 mRNAs, comprising the specific mRNAs of these organ constituent cells, can reflect the gene expression events of seminal plasma derived cells, and also can more perfectly reflect the contribution degree information of each cell to seminal plasma free RNA. The classification method of male infertility in the present embodiment, by utilizing the contribution ratio of various cell types to seminal plasma free RNA in testis to build the non-invasive diagnosis classification model of azoospermia or its subtype, can accurately predict unknown sample male azoospermia and its subtype, more convenient in operation than previous method, and prediction accuracy is higher. And this method can be predicted to the subtype of azoospermia before testicular puncture commonly used clinically, realizes the purpose of non-invasive detection and diagnosis.

[0068] The step of constructing the male infertility prediction model in the above S201 can establish a classification model of normal sperm and azoospermia according to the source of the sample population; or a classification model of obstructive azoospermia and non-obstructive azoospermia. Regardless of the type of male infertility classification prediction model, it is constructed by using the contribution ratio of testicular-related single cell types in seminal plasma free RNA as a feature. The specific model construction method is as mentioned in the above embodiment 1, including: obtaining the contribution ratio of various testicular-related single cell types in the seminal plasma free mRNA of each sample in the training set, wherein the training set is selected from set A or set B, set A includes a fertile male sample group and an azoospermic male sample group, and set B includes an obstructive azoospermia male sample group and a non-obstructive azoospermia male sample group; using the contribution ratio as the input feature of each sample for machine learning training to obtain a male infertility prediction model.

[0069] Example 2 of the present application is the first to use the method of testicular single cell tracing ratio to construct a classification prediction model for classification prediction of male infertility and its subtypes, that is, the contribution ratio of each cell type in each sample in different training sets is used as an input feature for machine learning training for the first time. The method for obtaining the contribution ratio of each cell type in each sample is not limited, including but not limited to the deconvolution method based on SVR, the non-negative matrix decomposition (NMF) method or the non-negative least squares regression (NNLS) method. The deconvolution method based on SVR is preferred in this application.

[0070] In the above embodiments, there are multiple classification algorithms suitable for machine learning training in the present application, including but not limited to random forest, logistic regression, decision tree or K nearest neighbor classification algorithm. The random forest algorithm is preferably used in the present application. In some preferred embodiments, multiple different prediction models can be obtained by using multiple different algorithms, and the prediction effects of the multiple different prediction models are further evaluated. The prediction model with the best prediction effect is selected as the final prediction model based on the evaluation effect.

[0071] In some preferred embodiments, the step of obtaining the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample in the training set includes: obtaining the expression profile of the seminal plasma free mRNA of each sample in the training set; using the testicular single-cell transcription data set to obtain a matrix of cell type-specific gene expression; and using a deconvolution method to calculate the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample based on the expression profile and the matrix of cell type-specific gene expression.

[0072] The step of "obtaining expression profiles of cell-free mRNA in seminal plasma for each sample in the training set" includes obtaining sequencing data (including high-throughput sequencing data or other sequencing data, or qPCR data) from each sample and performing quantification based on the obtained sequencing data to obtain expression profiles. This step is further illustrated below and specifically includes: performing quality control on the off-machine sequencing data of seminal plasma cfRNA from each sample, including trimming adapters, removing low-quality reads (quality threshold of 15, here referring to reads below this quality threshold), and removing reads <17 bp in length. (For example, using bowtie) rRNA sequences, vtRNA (another RNA type that is removed, similar to rRNA), and gamma RNA sequences are removed. The remaining reads are then aligned to the human transcriptome (in the order of miRNA, tRNA, and piRNA, then mRNA and lncRNA, and finally other RNAs). mRNA and lncRNA are quantified using RSEM (a typical transcript-based quantification method), and expression levels are then normalized using FPKM.

[0073] In some preferred embodiments, obtaining the expression profile of the seminal plasma free mRNA of each sample in the training set includes: obtaining the sequencing data of the seminal plasma free mRNA of each sample in the training set; performing quality control on the sequencing data to obtain post-quality control data; comparing the post-quality control data to the human transcriptome and quantifying it to obtain the expression profile.

[0074] In some preferred embodiments, quality control includes at least one of the following: 1) cutting junctions; 2) removing reads with a quality below a quality threshold; 3) removing reads <17 bp in length; 4) removing rRNA sequences, vtRNA and YRNA sequences;

[0075] It should be noted that the sequencing data quality control requirement for the off-machine data is strict: the off-machine data must have >10 million reads. Then, using the independently developed PALM-seq technology, the corresponding sequences in 1)-4) above are removed.

[0076] In some preferred embodiments, post-quality control data are aligned to the human transcriptome in the following order: 1) miRNA; 2) tRNA and piRNA; 3) mRNA and lncRNA; 4) remaining RNA. Transcriptome expression is quantified after alignment. In this application, no sample filtering is performed at this step. For scalability, a quality control standard of >10,000 mRNA detections can be set.

[0077] The exemplary steps for the above-mentioned "Using the testis single-cell transcriptional dataset to obtain a matrix of cell type-specific gene expression" are as follows: by researching the literature (Guo, J., EJ Grow, H. Mlcochova, et al. "The adult human testis transcriptional cell atlas." Cell Research (2018). 28(12): 1141-1157), a testis single-cell transcriptional dataset (see Table 2) was collected to sort out the cell types involved, and a cell type-specific gene expression matrix was generated using Cibersortx. Combined with the cell type-specific gene expression matrix, the contribution ratio of each cell type to the free mRNA in the seminal plasma of each sample was calculated using the SVR-based deconvolution method. At the same time, the p-value, correlation, and residual of the deconvolution of each sample are calculated to evaluate the reliability of the sample deconvolution results (the evaluation results here do not determine whether the sample is included, but evaluate the reliability of the method used to diagnose azoospermia and non-obstructive azoospermia using the cell type contribution ratio. When evaluating the deconvolution results, there will be a corresponding p-value to reflect whether the contribution score has reference value. If the p-value is very high, it will affect the credibility of the scoring results. Evaluating reliability specifically refers to the statistical significance of the deconvolution results. The smaller the p-value, the higher the correlation and the smaller the residual, which means that the tracing effect is better, and the mixing matrix and the reference matrix provided for deconvolution are more appropriate. If the p-value is small and the effect is not bad, no additional processing is required).

[0078] Therefore, in some preferred embodiments, the testis-related single cell types include the following 12 types: sperm, late primary spermatocytes, round spermatids, supporting cells, spermatogonial stem cells, macrophages, differentiating spermatogenic cells, vascular wall endothelial cells, elongated spermatids, seminiferous tubule muscle cells, interstitial cells and early primary spermatocytes.

[0079] In some preferred embodiments, while calculating the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample, the method also includes: calculating the p-value, correlation and residual of the deconvolution of each sample; and evaluating the reliability of the deconvolution results of each sample based on the p-value, correlation and residual.

[0080] The beneficial effects of the present application will be further explained in detail below with reference to specific embodiments.

[0081] Example 3

[0082] This example provides a method for classifying male infertility. This method uses seminal plasma samples from fertile and azoospermic men to obtain expression profiles of free mRNA in seminal plasma. A testicular single-cell transcriptome dataset is collected simultaneously. CIBERSORTx is used to generate a matrix of cell-type-specific gene expression. The contribution ratio of each cell type to the free mRNA in the seminal plasma of each sample is calculated using a deconvolution method based on SVR (Support Vector Regression). Using the contribution ratio as a feature of the sample, the sample is grouped by phenotype (normal / anospermia), and a training and validation set is divided. A binary classification model is constructed based on a random forest. This model can effectively distinguish between azoospermia and normal samples in the test set (AUC = 1.00). A similar method is used to distinguish subtypes of azoospermia patients, enabling prediction of obstructive and non-obstructive azoospermia (AUC = 0.88).

[0083] As shown in the flowchart of FIG3 , the specific steps of this embodiment are as follows:

[0084] 1. Extraction and library construction of seminal plasma samples and sequencing analysis.

[0085] (1) Extraction of free mRNA from seminal plasma

[0086] All semen samples were centrifuged at 10,000 g for 30 min at 4°C to remove sperm. The supernatant was collected and added to TRIzol LS at a ratio of 1:3 and immediately mixed by vortexing. Subsequent cfRNA extraction steps can be performed using methods familiar in the art.

[0087] (2) Sequencing of seminal plasma cfRNA

[0088] Polyadenylation ligation-mediated sequencing (PALM-Seq) and other technologies were used to construct a seminal plasma cfRNA library, and next-generation sequencing was used to sequence the seminal plasma free RNA to obtain the full transcriptional profile of seminal plasma free mRNA.

[0089] (3) Quantification of seminal plasma cfRNA expression profile

[0090] Semen plasma cfRNA sequencing data were quality-controlled, including adapter trimming, removal of low-quality reads, and removal of reads <17 bp in length. rRNA sequences, vert RNA, and Y RNA sequences were removed using bowtie. Remaining reads were aligned to the human transcriptome (in the order of miRNA, tRNA, and piRNA, then mRNA and lncRNA, and finally other RNAs). mRNA and lncRNA were quantified using RSEM, and expression levels were normalized using FPKM.

[0091] 2. Generate a testicular single-cell cell type-specific expression matrix and calculate the cell contribution ratio.

[0092] We collected a testicular single-cell transcriptome dataset (Table 1) based on the literature (Jingtao Guo et al., 2018) and used Cibersortx to generate a cell-type-specific gene expression matrix. Combining this matrix with the signature matrix, we calculated the contribution of each cell type to the cell-free mRNA in seminal plasma for each sample using a SVR-based deconvolution method. We also calculated the p-value, correlation, and residual for each sample deconvolution to assess the reliability of the sample deconvolution results.

[0093] 3. Building a Machine Learning Classification Model

[0094] Input features: contribution ratios of 12 testis-related single cell types.

[0095] Included samples: samples with high reliability of deconvolution results (p-value < 0.01) were used to build an obstruction classification model. Due to the extremely small number of samples and imbalance, simulated samples similar to obstruction samples were randomly generated using SMOTE.

[0096] Sample division: 60% of the samples are randomly divided as training set and 40% of the samples are divided as validation set.

[0097] Classification model: Random Forest (upsampling via SMOTE, tree model function Random Forest Classifier parameters n_estimators = 100, max_depth = 20, random_state = 123, oob_score = True).

[0098] Classification criteria: Samples with a score higher than 0.5 were defined as case (anospermia / obstruction), and samples with a score lower than 0.5 were defined as control (healthy / non-obstruction).

[0099] Example 4

[0100] (1) Semen collection from fertile men and azoospermic men

[0101] Semen from 107 fertile men and 102 azoospermic men was obtained from a hospital ( Table 1 ). Spermatozoa were removed by centrifugation at 10,000 g for 30 min at 4°C, and the supernatant was collected.

[0102] Table 1. Clinical information of enrolled samples

[0103] Data are reported as medians

[0104] (2) cfRNA extraction and sequencing

[0105] Add Trizol LS to the seminal plasma and immediately vortex to mix. The subsequent cfRNA extraction steps are as follows:

[0106] 1) Remove seminal plasma (300 μL of original seminal plasma) preserved with 900 μL of TRIzol LS and add 250 μL of chloroform. Immediately mix by repeated inversion. Centrifuge at 12,000 rpm at 4°C for 15 minutes. Pipette 500-600 μL of the upper colorless liquid into a new 1.5 mL EP tube. Add (volume of liquid / 500 μL) glycogen (5 mg / mL) to each sample, then add an equal volume of isopropanol. Mix by inversion and store overnight at -20°C.

[0107] 2) Centrifuge at 12,000 rpm and 4°C for 15 minutes. Discard as much of the supernatant as possible.

[0108] 3) Add 400 μL of freshly prepared 75% ethanol and centrifuge at 12,000 rpm at 4°C for 5 minutes. Discard the supernatant. Repeat step 5. Centrifuge at 12,000 rpm at 4°C for 5 minutes, carefully aspirating any remaining ethanol. Allow to air dry at room temperature for 10 minutes. Finally, transfer the tubes to a 96-well cryovial and begin library preparation as soon as possible.

[0109] Database construction process:

[0110] 1) After extraction, mix 10 μL of cfRNA with 4.5 μL of DEPC water, 2 μL of 10× RNase H buffer, 2 μL of ATP, 0.5 μL of LT4PNK, 0.5 μL of poly(A) polymerase, and 0.5 μL of RNase inhibitor, and incubate at 37°C for 30 minutes. Terminate the reaction by adding 2 μL of stop-annealing buffer (40 mM EDTA, 1.5 M KCl).

[0111] 2) Then, 1 μL of depletion oligo mix was added and incubated at 95°C for 5 minutes. The temperature was then lowered to 22°C at a rate of 0.1°C / s, and the cfRNA was placed on ice.

[0112] 3) Add 4 μL of 10× T4 RNA ligase buffer, 4 μL of ATP, 7 μL of 50% PEG-8000, 1 μL of T4 RNA ligase I, 0.5 μL of RNase inhibitor, and 0.5 μL of 20 μM adapter. Incubate the sample at 25°C for 1 hour. Then, add 2 μL of rSAP and incubate at 37°C for 10 minutes to remove ATP.

[0113] 4) Add 1 μL RNase H and 2 μL DEPC water and incubate at 37°C for 30 minutes. Then, add 2 μL DNase I and 3 μL 30 mM CaCl2 and incubate at 37°C for 30 minutes to remove the probe and any remaining free DNA. Purify the modified RNA using 50 μL of RNA Clean Beads and resuspend in 12.5 μL of DEPC water.

[0114] 5) Add 1 μL of 5 μM reverse transcription (RT) primer and incubate the mixture at 65°C for 5 minutes, then place on ice for 2 minutes. Add 1 μL of dNTPs, 0.5 μL of RNase inhibitor, 4 μL of 5× HiScript II buffer, and 1 μL of HiScript II reverse transcriptase and incubate at 50°C for 15 minutes. Then, incubate the mixture at 85°C for 5 minutes to inactivate the reverse transcriptase.

[0115] 6) Add 5 μL of PCR primers and 25 μL of 2× Q5 Hot Start PCR Master Mix to each sample. Perform PCR according to the manufacturer's instructions, using an annealing temperature of 56°C, an extension time of 1 minute, and 20 amplification cycles. Purify the PCR product using 75 μL of DNA Clean Beads and circularize.

[0116] 7) Finally, sequencing was performed on DNBSEQ-T1 to generate 100-bp single-end reads.

[0117] (3) Quantification of cfRNA expression profile

[0118] Semen plasma cfRNA sequencing data were quality-controlled, including adapter trimming, removal of low-quality reads, and removal of reads <17 bp in length. rRNA sequences, vert RNA, and Y RNA sequences were removed using bowtie. Remaining reads were aligned to the human transcriptome (in the order of miRNA, tRNA, and piRNA, then mRNA and lncRNA, and finally other RNAs). Quantification was performed using RSEM, and expression levels were normalized using FPKM.

[0119] (4) Generate a testicular single-cell cell type-specific expression matrix and calculate the cell contribution ratio.

[0120] Using a human testis single-cell transcriptome dataset (Jingtao Guo et al., 2018), we generated a cell-type-specific gene expression matrix using CIBERSORTx. Combined with the signature matrix, we used a SVR-based deconvolution method to calculate the contribution of each cell type to the cell-free mRNA in seminal plasma for each sample. We also calculated the p-value, correlation, and residual for each sample deconvolution to assess the reliability of the sample deconvolution results.

[0121] (5) Build a machine learning classification model

[0122] Input features: contribution ratios of 12 testis-related single cell types

[0123] Table 2. Cell type information

[0124] Samples: Based on sequencing quality control and transcriptome expression, 104 healthy samples, 97 azoospermia samples, 26 non-obstructive azoospermia samples, and 6 obstructive azoospermia samples were selected as qualified. To construct the obstruction classification model, due to the extremely small and unbalanced sample size, simulated samples similar to the obstruction samples were randomly generated using SMOTE.

[0125] Sample division: Randomly divide 60% of the samples as training set and 40% of the samples as validation set

[0126] Classification model: Random Forest (upsampling via SMOTE, Random Forest Classifier with n_estimators = 100, max_depth = 20, random_state = 123, oob_score = True)

[0127] Classification criteria: Samples with a score higher than 0.5 were defined as case (anospermia / obstruction), and samples with a score lower than 0.5 were defined as control (healthy / non-obstruction).

[0128] (6) Evaluation of the predictive effect of testicular single cell tracing scoring on azoospermia and its subtypes.

[0129] (6.1) Predictive effect on male azoospermia

[0130] Based on the contribution ratio of various cell types in the seminal plasma tracing results, a healthy and azoospermia classification model was constructed, and the prediction effect of the validation set (see Table 3) was as follows: AUC = 1.00, the precision rate reached 0.97, and the recall rate was 100% (as shown in Figures 4A and 4B. Figure 4B is the result of the SHAP algorithm based on random forest, which calculates the influence of cells of each source of each sample on the final classification result. The recall rate result is calculated based on the specific prediction result), and the accuracy rate reached 0.98.

[0131] (6.2) Predictive effect on non-obstructive azoospermia

[0132] Based on the contribution ratios of various cell types in the seminal plasma tracing results, a classification model for non-obstructive azoospermia and obstructive azoospermia was constructed. The prediction effect of the validation set (see Table 3) was: AUC 0.87, precision 0.67, recall rate 0.89 (see Figure 5A and Figure 5B), and accuracy 0.76.

[0133] In the table below, Accuracy represents the proportion of correctly classified samples to the total number of samples. Precision, also known as the recall rate, represents the proportion of samples predicted as positive that are actually positive. Recall represents the proportion of samples predicted as correct to the total number of correct samples. The three are calculated using the following formulas:

[0134] Accuracy = (TP+TN) / (TP+FP+TN+FN); Precision = TP / (TP+FP); Recall = TP / (TP+FN).

[0135] TP: True positive; TN: True negative;

[0136] FP: False positive; FN: False negative.

[0137] Table 3. Prediction results of model validation set

[0138] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the present invention.

[0139] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus hardware devices such as detection devices. Based on this understanding, the data processing part of the technical solution of the present application can be embodied in the form of a software product, and the computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment of the present application or certain parts of the embodiments.

[0140] The present application can be used in a wide variety of general-purpose or specialized computing system environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments that include any of the above.

[0141] The method provided in this application can be executed in a terminal, a computer terminal or a similar computing device. Taking running on a terminal as an example, Figure 6 is a hardware structure block diagram of a terminal of a classification method for male infertility according to an embodiment of the present invention. As shown in Figure 6, the terminal may include one or more (only one is shown in Figure 6) processors A1 (processor A1 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory B1 for storing data. Optionally, the terminal may also include a transmission device C1 and an input and output device D1 for communication functions. It will be understood by those skilled in the art that the structure shown in Figure 6 is only for illustration and does not limit the structure of the terminal. For example, the terminal may also include more or fewer components than those shown in Figure 6, or have a configuration different from that shown in Figure 6.

[0142] The memory B1 can be used to store computer programs, for example, software programs and modules of application software, such as computer programs corresponding to methods such as read segment splicing, clustering, and consistency processing in the embodiments of the present invention. The processor A1 executes various functional applications and data processing by running the computer programs stored in the memory B1, that is, implementing the above-mentioned method. The memory B1 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory B1 may further include a memory remotely located relative to the processor A1, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0143] The transmission device C1 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the terminal's communications provider. In one embodiment, the transmission device C1 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device C1 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0144] Obviously, those skilled in the art should understand that some of the modules or steps of the present application described above can be implemented on a general-purpose computing device, they can be concentrated on a single computing device, or distributed across a network consisting of multiple computing devices. Alternatively, they can be implemented using program code executable by the computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately manufactured into individual integrated circuit modules, or multiple modules or steps can be manufactured into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.

[0145] Example 5

[0146] This embodiment provides a device for constructing a male infertility prediction model, as shown in FIG7 , the device includes: an acquisition module and a machine learning module, wherein:

[0147] an acquisition module configured to obtain contribution ratios of various testis-related single cell types in seminal plasma free mRNA of each sample in a training set, wherein the training set is selected from set A or set B, set A includes a fertile male sample group and a male sample group with azoospermia, and set B includes a male sample group with obstructive azoospermia and a male sample group with non-obstructive azoospermia;

[0148] The machine learning module is configured to use the contribution ratio as the input feature of each sample for machine learning training to obtain an azoospermia prediction model.

[0149] Optionally, the acquisition module includes: a first acquisition unit, configured to obtain the expression spectrum of the seminal plasma free mRNA of each sample in the training set; a second acquisition unit, configured to use the testicular single-cell transcription data set to obtain a matrix of cell type-specific gene expression; a calculation unit, configured to calculate the contribution ratio of various testicular-related single cell types in the seminal plasma free mRNA of each sample using a deconvolution method based on the expression spectrum and the matrix of cell type-specific gene expression.

[0150] Optionally, the testis-related single cell types include the following 12 types: sperm, late primary spermatocytes, round spermatids, supporting cells, spermatogonial stem cells, macrophages, differentiating spermatogenic cells, vascular wall endothelial cells, elongated spermatids, seminiferous tubule muscle cells, interstitial cells and early primary spermatocytes.

[0151] Optionally, the first acquisition unit includes: a sequencing data acquisition module, which is configured to obtain the sequencing data of the seminal plasma free mRNA of each sample in the training set; a quality control module, which is configured to perform quality control on the sequencing data to obtain quality-controlled data; a quantification module, which is configured to match the quality-controlled data to the human transcriptome and perform quantification to obtain an expression profile; preferably, quality control includes at least one of the following: 1) cutting junctions; 2) removing reads with a quality lower than a quality threshold; 3) removing reads <17bp in length; 4) removing rRNA sequences, vt RNA and Y RNA sequences; preferably, the quality-controlled data are matched to the human transcriptome in the following order: 1) miRNA; 2) tRNA and piRNA; 3) mRNA and lncRNA; 4) remaining RNA.

[0152] Optionally, the classification device further includes: a calculation submodule configured to calculate the p-value, correlation and residual of each sample deconvolution; and an evaluation module configured to evaluate the reliability of each sample deconvolution result based on the p-value, correlation and residual.

[0153] Example 6

[0154] The embodiment of the present invention provides a classification prediction device for male infertility, as shown in FIG8 , the classification device includes: a model building module and a classification module, wherein:

[0155] The model building module is configured to construct a male infertility prediction model using the contribution ratio of testis-related single cell types in seminal plasma free RNA as a feature;

[0156] The classification module is configured to input the contribution ratio of testis-related single cell types in the seminal plasma free RNA of the sample to be tested into the male infertility prediction model to obtain the classification prediction result of the sample to be tested.

[0157] In the above preferred embodiment, the model building module includes: an acquisition module, which is configured to obtain the contribution ratio of various testis-related single cell types in the seminal plasma free mRNA of each sample in the training set, wherein the training set is selected from set A or set B, set A includes a fertile male sample group and an azoospermia male sample group, and set B includes an obstructive azoospermia male sample group and a non-obstructive azoospermia male sample group; a machine learning module, which is configured to use the contribution ratio as the input feature of each sample for machine learning training to obtain an azoospermia prediction model.

[0158] Optionally, the acquisition module includes: a first acquisition unit, configured to obtain the expression spectrum of the seminal plasma free mRNA of each sample in the training set; a second acquisition unit, configured to use the testicular single-cell transcription data set to obtain a matrix of cell type-specific gene expression; a calculation unit, configured to calculate the contribution ratio of various testicular-related single cell types in the seminal plasma free mRNA of each sample using a deconvolution method based on the expression spectrum and the matrix of cell type-specific gene expression.

[0159] Optionally, the testis-related single cell types include the following 12 types: sperm, late primary spermatocytes, round spermatids, supporting cells, spermatogonial stem cells, macrophages, differentiating spermatogenic cells, vascular wall endothelial cells, elongated spermatids, seminiferous tubule muscle cells, interstitial cells and early primary spermatocytes.

[0160] Optionally, the first acquisition unit includes: a sequencing data acquisition module, which is configured to obtain the sequencing data of the seminal plasma free mRNA of each sample in the training set; a quality control module, which is configured to perform quality control on the sequencing data to obtain quality-controlled data; a quantification module, which is configured to match the quality-controlled data to the human transcriptome and perform quantification to obtain an expression profile; preferably, quality control includes at least one of the following: 1) cutting junctions; 2) removing reads with a quality lower than a quality threshold; 3) removing reads <17bp in length; 4) removing rRNA sequences, vt RNA and Y RNA sequences; preferably, the quality-controlled data are matched to the human transcriptome in the following order: 1) miRNA; 2) tRNA and piRNA; 3) mRNA and lncRNA; 4) remaining RNA.

[0161] Optionally, the classification device further includes: a calculation submodule configured to calculate the p-value, correlation and residual of each sample deconvolution; and an evaluation module configured to evaluate the reliability of each sample deconvolution result based on the p-value, correlation and residual.

[0162] The present application also provides a device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, the following steps are implemented: a male infertility prediction model is constructed using the contribution ratio of testicular-related single cell types in seminal plasma free RNA as a feature; the contribution ratio of testicular-related single cell types in seminal plasma free RNA of the sample to be tested is input into the male infertility prediction model to obtain a classification prediction result of the sample to be tested.

[0163] It should be noted that the devices in this article can be servers, PCs, PADs, mobile phones, etc.

[0164] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing an initialized program having the following method steps: constructing a male infertility prediction model using the contribution ratio of testicular-related single cell types in seminal plasma free RNA as a feature; inputting the contribution ratio of testicular-related single cell types in seminal plasma free RNA of the sample to be tested into the male infertility prediction model to obtain a classification prediction result of the sample to be tested.

[0165] The present application is described with reference to the flow chart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.

[0166] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0167] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0168] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0169] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0170] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0171] From the above description, it can be seen that the above embodiments of the present invention achieve the following technical effects:

[0172] (1) Based on clinical information, the present invention proposes for the first time to use the contribution score of testicular single cells in seminal plasma free RNA as a feature to construct a model for the classification and prediction of azoospermia and its subtypes. The prediction effect of the model is evaluated according to the area under the curve (AUC) of the sensitivity and specificity. The AUC for diagnosing azoospermia is increased to 1.00, and the AUC for NOA reaches 0.88, both of which are higher than the existing technical level. In addition, seminal plasma free RNA is non-invasive, sensitive, highly stable and accurate, and is currently the most economical method.

[0173] (2) Use mature transcriptomics and readily available clinical information instead of proteomics or metabolomics technologies to make the results stable, reliable, and easy to operate.

[0174] (3) In addition, the samples used in the various embodiments of this application refer to seminal plasma, but other body fluids such as plasma, serum, and urine may also have predictive effects; the various embodiments of this application detect the contribution ratio of various cell types in the testis to seminal plasma free RNA, and the expression of free RNA may also have the same effect; the prediction model algorithm used in the various embodiments of this application is one, but using other models instead may also have the same effect. In addition to the method of establishing a prediction model to achieve classification prediction, other detection methods, such as qPCR or other sequencing methods, may also have the same effect.

[0175] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A classification method for male infertility, characterized in that, The classification method includes: Constructing a male infertility prediction model using the contribution ratio of testicular - related single - cell types in seminal plasma free RNA as a feature; Inputting the contribution ratio of testicular - related single - cell types in the seminal plasma free RNA of the sample to be tested into the male infertility prediction model to obtain the classification prediction result of the sample to be tested.

2. The classification method according to claim 1, characterized in that, Constructing a male infertility prediction model using the contribution ratio of testicular single - cell types in seminal plasma free RNA as a feature includes: Obtaining the contribution ratio of various testicular - related single - cell types in the seminal plasma free mRNA of each sample in the training set, where the training set is selected from set A or set B, set A includes a fertile male sample population and a non - obstructive azoospermia male sample population, and set B includes an obstructive azoospermia male sample population and a non - obstructive azoospermia male sample population; Using the contribution ratio as the input feature of each sample for machine learning training to obtain the male infertility prediction model.

3. The classification method according to claim 2, wherein Obtaining the contribution ratio of various testicular - related single - cell types in the seminal plasma free mRNA of each sample in the training set includes: Obtaining the expression profile of the seminal plasma free mRNA of each sample in the training set; Using the testicular single - cell transcript dataset to obtain a matrix of cell - type - specific gene expression; According to the expression profile and the matrix of cell - type - specific gene expression, using a deconvolution method to calculate the contribution ratio of various testicular - related single - cell types in the seminal plasma free mRNA of each sample.

4. The classification method according to any one of claims 1-3, characterized in that The testicular - related single - cell types include the following 12 types: sperm, late primary spermatocytes, round spermatids, Sertoli cells, spermatogonial stem cells, macrophages, differentiating spermatogenic cells, vascular endothelial cells, elongated spermatids, peritubular myoid cells, interstitial cells, and early primary spermatocytes.

5. The classification method according to claim 3, wherein Obtaining the expression profile of the seminal plasma free mRNA of each sample in the training set includes: Obtaining the raw sequencing data of the seminal plasma free mRNA of each sample in the training set; Performing quality control on the raw sequencing data to obtain quality - controlled data; Aligning the quality - controlled data to the human transcriptome and performing quantification to obtain the expression profile; Preferably, the quality control includes at least one of the following: 1) trimming adapters; 2) removing reads with quality lower than the quality threshold; 3) removing reads with a length less than 17bp; 4) removing rRNA sequences, vt RNA, and Y RNA sequences; Preferably, aligning the quality - controlled data to the human transcriptome in the following order: 1) miRNA; 2) tRNA and piRNA; 3) mRNA and lncRNA; 4) remaining RNA.

6. The classification method according to claim 2, characterized in that, While calculating the contribution ratio of various testicular - related single - cell types in the seminal plasma free mRNA of each sample, the classification method further includes: Calculating the p - value, correlation, and residual of the deconvolution of each sample; And evaluating the reliability of the deconvolution result of each sample according to the p - value, correlation, and residual.

7. A classification device for male infertility, characterized in that, The classification device includes: A model construction module, configured to construct a male infertility prediction model using the contribution ratio of testicular - related single - cell types in seminal plasma free RNA as a feature; A classification module, configured to input the contribution ratio of the testicular-related single-cell types in the free RNA of seminal plasma of a sample to be tested into the male infertility prediction model, and obtain a classification prediction result of the sample to be tested.

8. The classification device according to claim 7, characterized in that The model construction module includes: An acquisition module, configured to acquire the contribution ratio of various testicular-related single-cell types in the free mRNA of seminal plasma of each sample in a training set, where the training set is selected from set A or set B, the set A includes a fertile male sample population and an azoospermia male sample population, and the set B includes an obstructive azoospermia male sample population and a non-obstructive azoospermia male sample population; A machine learning module, configured to perform machine learning training with the contribution ratio as the input feature of each sample, and obtain the azoospermia prediction model.

9. The classification device according to claim 8, characterized in that, The acquisition module includes: A first acquisition unit, configured to acquire the expression profile of the free mRNA of seminal plasma of each sample in the training set; A second acquisition unit, configured to obtain a matrix of cell type-specific gene expression by using a testicular single-cell transcript dataset; A calculation unit, configured to calculate the contribution ratio of various testicular-related single-cell types in the free mRNA of each sample by using a deconvolution method according to the expression profile and the matrix of cell type-specific gene expression.

10. The classification device according to any one of claims 7-9, characterized in that, The testicular-related single-cell types include the following 12 types: sperm, late primary spermatocytes, round spermatids, Sertoli cells, spermatogonial stem cells, macrophages, differentiating spermatogenic cells, vascular wall endothelial cells, elongated spermatids, peritubular myoid cells, interstitial cells, and early primary spermatocytes.

11. The classification device according to claim 9, characterized in that, The first acquisition unit includes: A sequencing data acquisition module, configured to acquire the sequencing off-machine data of the free mRNA of seminal plasma of each sample in the training set; A quality control module, configured to perform quality control on the sequencing off-machine data to obtain quality-controlled data; A quantification module, configured to align the quality-controlled data to the human transcriptome and perform quantification to obtain the expression profile; Preferably, the quality control includes at least one of the following: 1) trimming adapters; 2) removing reads with a quality lower than a quality threshold; 3) removing reads with a length less than 17 bp; 4) removing rRNA sequences, vt RNA, and Y RNA sequences; Preferably, the quality-controlled data is aligned to the human transcriptome in the following order: 1) miRNA; 2) tRNA and piRNA; 3) mRNA and lncRNA; 4) remaining RNA. The classification device further includes:

12. The classification device according to claim 9, characterized in that A calculation sub-module, configured to calculate the p-value, correlation, and residual of the deconvolution of each sample; An evaluation module, configured to evaluate the reliability of the deconvolution result of each sample according to the p-value, correlation, and residual. The method includes:

13. A method for constructing a prediction model for male infertility, characterized in that, Acquiring the contribution ratio of various testicular-related single-cell types in the free mRNA of seminal plasma of each sample in a training set, where the training set is selected from set A or set B, the set A includes a fertile male sample population and an azoospermia male sample population, and the set B includes an obstructive azoospermia male sample population and a non-obstructive azoospermia male sample population; ​ Using the contribution ratio as the input feature of each of the samples for machine learning training to obtain the male infertility prediction model.

14. The method according to claim 13, characterized in that, Obtaining the contribution ratio of various testicular-related single cell types in the free seminal plasma mRNA of each sample in the training set includes: Obtaining the expression profile of the free seminal plasma mRNA of each sample in the training set; Using the testicular single cell transcriptome dataset to obtain the matrix of cell type-specific gene expression; According to the expression profile and the matrix of cell type-specific gene expression, using the deconvolution method to calculate the contribution ratio of various testicular-related single cell types in the free seminal plasma mRNA of each sample.

15. The method according to claim 13 or 14, characterized in that, The testicular-related single cell types include the following 12 types: sperm, late primary spermatocytes, round spermatids, Sertoli cells, spermatogonial stem cells, macrophages, differentiating spermatogenic cells, vascular wall endothelial cells, elongated spermatids, peritubular myoid cells, interstitial cells, and early primary spermatocytes.

16. The method according to claim 14, characterized in that Obtaining the expression profile of the free seminal plasma mRNA of each sample in the training set includes: Obtaining the sequencing off-machine data of the free seminal plasma mRNA of each sample in the training set; Performing quality control on the sequencing off-machine data to obtain the quality-controlled data; Aligning the quality-controlled data to the human transcriptome and performing quantification to obtain the expression profile; Preferably, the quality control includes at least one of the following: 1) trimming adapters; 2) removing reads with quality lower than the quality threshold; 3) removing reads with a length less than 17bp; 4) removing rRNA sequences, vt RNA, and Y RNA sequences; Preferably, the quality-controlled data is aligned to the human transcriptome in the following order: 1) miRNA; 2) tRNA and piRNA; 3) mRNA and lncRNA; 4) remaining RNA.

17. The method according to claim 14, wherein While calculating the contribution ratio of various testicular-related single cell types in the free seminal plasma mRNA of each sample, the method further includes: Calculating the p-value, correlation, and residual of the deconvolution of each sample; And evaluating the reliability of the deconvolution results of each sample according to the p-value, correlation, and residual.

18. A device for constructing a prediction model for male infertility, characterized in that, The device includes: An acquisition module, configured to acquire the contribution ratio of various testicular-related single cell types in the free seminal plasma mRNA of each sample in the training set, where the training set is selected from set A or set B, the set A includes a fertile male sample population and an azoospermia male sample population, and the set B includes an obstructive azoospermia male sample population and a non-obstructive azoospermia male sample population; A machine learning module, configured to use the contribution ratio as the input feature of each of the samples for machine learning training to obtain the azoospermia prediction model.

19. A non-transitory computer-readable storage medium, characterized in that, The storage medium includes a stored program, wherein when the program runs, it controls the device where the storage medium is located to execute the classification method of male infertility according to any one of claims 1-7, or the method for constructing a male infertility prediction model according to any one of claims 13-17.

20. A processor, characterized in that, The processor is used to run a program, wherein when the program runs, it executes the classification method of male infertility described in any one of 1-7, or the method of constructing a male infertility prediction model described in any one of claims 13-17.