Multi-endpoint reproductive toxicity prediction method and system based on large language model

By collaborating with experts and large language models to acquire high-quality data and construct a multi-dimensional reproductive toxicity prediction model, the problem of insufficient resolution in reproductive toxicity prediction in existing technologies has been solved, achieving high-precision reproductive toxicity prediction, especially in terms of accuracy in terms of sex differences and intergenerational transmission effects.

CN121905338APending Publication Date: 2026-04-21NANJING UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV
Filing Date
2025-11-14
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve high-resolution predictions of reproductive toxicity, particularly in terms of considering sex differences, intergenerational transmission effects, and developmental stage specificity. Furthermore, traditional animal experiments are time-consuming and resource-intensive, and existing models rely on single large language models, resulting in insufficient accuracy and misclassification issues.

Method used

We construct a multi-terminal reproductive toxicity prediction method based on a large language model. We acquire high-quality data through a collaborative strategy between experts and the large language model, perform multi-dimensional annotation and multi-dimensional modeling, and utilize the molecular encoder and contrastive learning module of element-knowledge graph to construct binary classifiers for different phenotype-study type-generation-sex combinations, forming multiple specific prediction models. We then combine performance evaluation to select the optimal model combination.

Benefits of technology

It achieves high-precision prediction of reproductive toxicity endpoints, from overall toxicity to specific reproductive organ weight and sperm quality, solves the biological complexity of sex differences and intergenerational transmission effects, and provides multi-dimensional reproductive toxicity prediction capabilities across sexes, generations, and research types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121905338A_ABST
    Figure CN121905338A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-endpoint reproductive toxicity prediction method and system based on a large language model, and belongs to the technical field of toxicity prediction. The method comprises the following steps: extracting phenotype information of text data in a candidate data set by using the large language model; performing multi-dimensional annotation processing on the phenotype level annotation data set; inputting the reproductive toxicity training set into a pre-constructed molecular representation learning model for training, and constructing a reproductive toxicity prediction model set with phenotype specificity; performing performance evaluation on the plurality of classifiers in the reproductive toxicity prediction model set by using the test data set, and screening according to preset performance indexes to obtain an optimal prediction model combination; performing toxicity prediction on the molecular structure of the chemical substance to be detected by using the optimal prediction model combination to obtain a multi-dimensional reproductive toxicity prediction result including cross-gender, cross-generation and cross-research types; according to the method, high-resolution and multi-dimensional reproductive toxicity phenotype prediction is carried out on the environmental chemical substances.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of toxicity prediction technology, and more specifically, to a multi-terminal reproductive toxicity prediction method and system based on a large language model. Background Technology

[0002] Reproductive health issues are becoming increasingly serious globally. Numerous epidemiological surveys indicate an accelerating trend of declining sperm counts and increasing testicular cancer rates in men, as well as rising incidence of female reproductive system diseases such as polycystic ovary syndrome and endometriosis. International health organizations, including the World Health Organization (WHO), have pointed out that exposure to chemical substances is a significant driving factor behind these trends. These health crises pose new challenges to both public health and chemical regulation.

[0003] Currently, regulatory systems in major jurisdictions such as the EU, China, and the US all require reproductive toxicity assessments for chemicals. However, the sheer volume and continuous growth of chemicals used in commercial applications have exceeded the processing capacity of existing traditional toxicological testing methods. Current reproductive toxicity testing primarily relies on animal experiments, which, while having some reference value, are generally time-consuming and resource-intensive, making it difficult to meet the current demand for efficient risk screening of large-scale chemicals. Therefore, a shift from animal-based testing paradigms to computer-based virtual screening has become a development trend.

[0004] Currently, artificial intelligence (AI) and machine learning (ML) technologies have demonstrated outstanding performance in molecular property prediction, chemical reactivity, and toxicity assessment, showing particular potential in reducing reliance on animal testing. For example, the invention patent application CN109658989A, entitled "A Method for Toxicity Prediction of Drug-like Compounds Based on Deep Learning," discloses a method for predicting toxicity using molecular fingerprints as feature inputs and a deep neural network (including an autoencoder stack structure). This method also involves feature denoising, dimensionality reduction, and model structure optimization to achieve toxicity prediction of drug-like compounds. However, the prediction results of this method are mainly limited to binary classification tasks, i.e., distinguishing whether a compound has overall toxicity, and fail to perform hierarchical modeling of specific reproductive toxicity phenotypes, sex differences, and developmental stages.

[0005] For example, the invention patent application CN113380341B, entitled "A Method for Predicting Compound Toxicity Based on Gene Expression Data," discloses a method that combines gene expression data with a machine learning model to predict toxicity. This method improves the biological relevance of the prediction by integrating large-scale genomic information. However, this method mainly predicts overall toxicity and does not address the complex phenotypic characteristics specific to reproductive toxicity, nor does it consider key factors such as sex specificity and developmental stage. Therefore, it is difficult to meet the needs of high-resolution reproductive toxicity modeling, resulting in limited value of the prediction results in mechanism explanation, risk stratification, and regulatory applications.

[0006] Furthermore, existing toxicology databases generally suffer from coarse data granularity and a lack of crucial experimental background information. Most databases only provide brief toxicity conclusions, without including experimental results by exposure time, generational range, or sex. This makes it difficult to achieve high-resolution reproductive toxicity modeling, even based on existing predictive models.

[0007] In recent years, Large Language Models (LLMs) have demonstrated significant potential in processing unstructured text data, enabling the automatic extraction of biological entities and toxicological concepts. LLMs typically refer to deep learning-based natural language processing models, with a core neural network architecture boasting a massive number of parameters (e.g., a Transformer-based structure). These models, pre-trained on large corpora, learn the statistical regularities and semantic representations of natural language, thus exhibiting high performance in various natural language processing tasks such as text generation, information extraction, question answering, and translation. For example, the invention patent application CN202411359418, entitled "Method for Mining Toxicity Effect Test Information from Literature Sources," describes a method that introduces pre-defined information mining rules into a large language model and combines them with cue word engineering to process literature corpora, thereby extracting toxicity effect information including compounds, exposure duration, and effect outcomes. This method demonstrates effectiveness in the automated extraction of literature information. However, this method still has limitations: it relies primarily on a single large language model for data extraction, which can lead to insufficient accuracy, misclassification, and biases in phenotypic semantic understanding. Furthermore, it struggles to effectively integrate background information across different literature sources or under complex experimental conditions. Complete reliance on LLM for data extraction can result in error propagation and reduced data reliability. Therefore, relying solely on LLM for data extraction carries significant risks. A collaborative strategy combining expert-guided manual extraction with LLM-assisted validation is essential to ensure high accuracy and reliability, thereby effectively supporting high-resolution reproductive toxicity modeling. Summary of the Invention

[0008] To address the insufficient phenotypic resolution in existing reproductive toxicity prediction technologies, this application provides a multi-terminal reproductive toxicity prediction method and system based on a large language model. It constructs specific prediction models for different phenotypic-study-generation-sex combinations to perform high-resolution, multi-dimensional reproductive toxicity phenotypic predictions of environmental chemicals.

[0009] One aspect of this application provides a multi-terminal reproductive toxicity prediction method based on a large language model, comprising: S1, acquiring a multi-source dataset containing chemical substances and their reproductive toxicity research records; S2, performing reliability screening on the multi-source dataset to obtain candidate datasets; S3, using a large language model to extract phenotypic information from the text data in the candidate datasets to obtain a standardized phenotypic-level annotated dataset; S4, performing multi-dimensional annotation processing on the phenotypic-level annotated dataset to obtain a reproductive toxicity training set containing phenotypic resolution; the multi-dimensional annotation processing includes adding research type labels, sex labels, generation labels, phenotypic endpoint labels, and activity labels to each phenotypic record; S5, inputting the reproductive toxicity training set into a pre-constructed molecular characterization learning model for training to construct a phenotypic-specific reproductive toxicity prediction model. The test model set includes: a molecular characterization learning model (comprising an element-knowledge graph-based molecular encoder and a pre-training module based on contrastive learning); a reproductive toxicity prediction model set (comprising multiple binary classifiers trained for different phenotype-study type-generation-sex combinations); S6, performance evaluation of multiple classifiers in the reproductive toxicity prediction model set using the test dataset, and selection of the optimal prediction model combination based on preset performance indicators; the optimal prediction model combination includes multiple prediction models for different toxicity endpoints; S7, toxicity prediction using the molecular structure of the test chemical substance using the optimal prediction model combination, obtaining multi-dimensional reproductive toxicity prediction results including cross-sex, cross-generation, and cross-study types; study types include prenatal developmental toxicity studies, first-generation reproductive toxicity studies, and multi-generation reproductive toxicity studies.

[0010] Furthermore, S2, reliability screening is performed on the multi-source dataset, including: using a chemical named entity recognition model to identify text data in the multi-source dataset and screening out research data centered on chemicals; using a preset reliability scoring standard to score the research data; screening research data with scores greater than or equal to a preset threshold as preliminary screening datasets; and using a machine learning verification model to retrospectively verify the preliminary screening datasets to confirm the accuracy of the screening results and form candidate datasets.

[0011] Furthermore, the chemical named entity recognition model adopts the Chemlistem model based on a recurrent neural network.

[0012] Furthermore, in S3, phenotypic information of text data in the candidate dataset is extracted using a large language model, including: obtaining a pre-constructed phenotypic extraction baseline dataset as a performance evaluation standard; the phenotypic extraction baseline dataset includes labeled phenotypic information samples; selecting multiple candidate large language models to perform phenotypic recognition on the text data in the candidate dataset; evaluating the adaptability, structured output capability, and controllability of the phenotypic recognition results of each candidate large language model based on the phenotypic extraction baseline dataset; selecting the candidate large language model with the best performance as the target extraction model based on consistency and text parsing stability according to the results of the adaptability evaluation, structured output capability evaluation, and controllability evaluation; using the target extraction model to extract phenotypic information from the text data in the candidate dataset to obtain raw annotation data containing reproductive toxicity phenotypic endpoints; and standardizing the raw annotation data to form a standardized phenotypic-level annotation dataset.

[0013] Furthermore, the candidate large language models include at least two of GPT-4, GPT-4 mini, Gemini 2.0, LLaMa-3.1, and Gemma-3.

[0014] Furthermore, phenotypic information includes at least three of the following: chemical substance name, toxic phenotypic description, dosage information, exposure time, and experimental species.

[0015] Furthermore, in S4, multidimensional annotation processing is performed on the phenotypic-level annotated dataset, including: adding a research type label to each phenotypic record in the phenotypic-level annotated dataset based on preset research type classification rules; research type labels include prenatal developmental toxicity research labels, first-generation reproductive toxicity research labels, and multi-generation reproductive toxicity research labels; adding a sex label to each phenotypic record in the phenotypic-level annotated dataset based on preset sex identification rules; sex labels include male labels, female labels, and hermaphrodite labels; adding a generation label to each phenotypic record in the phenotypic-level annotated dataset based on preset generation division rules; generation labels include parental labels, first-generation labels, and second-generation labels; the parental label corresponds to the generation of experimental animals directly exposed to the chemical substance, the first-generation label corresponds to the first-generation offspring produced by parental reproduction, and the second-generation label corresponds to the second-generation offspring produced by the first-generation offspring reproduction. Based on a pre-defined phenotypic endpoint classification system, phenotypic endpoint labels are added to each phenotypic record in the phenotypic-level annotation dataset; each phenotypic endpoint label corresponds to a specific reproductive toxicity phenotypic category. Based on pre-defined activity determination criteria, activity labels are added to each phenotypic record in the phenotypic-level annotation dataset; these activity labels indicate whether the phenotypic record exhibits toxic activity. When different research results for the same chemical substance conflict, a conservative determination rule is used for activity labeling; the conservative determination rule is: when at least one positive phenotype exists, the corresponding chemical substance is marked as active. The phenotypic records with completed multi-dimensional annotations are combined to form a reproductive toxicity training set. Each phenotypic record in the reproductive toxicity training set contains five dimensions of label information, forming a multi-dimensional annotation structure of phenotypic-research type-generation-sex-activity.

[0016] Further, in S5, a set of phenotype-specific reproductive toxicity prediction models is constructed, including: acquiring pre-constructed molecular characterization learning models; the molecular characterization learning models include a molecular encoder based on element-knowledge graphs and a pre-training module based on contrastive learning; performing multi-view characterization transformation on the molecular structures of chemical substances in the reproductive toxicity training set to generate molecular graph structures and molecular sequence representations; wherein, the molecular graph structures are constructed based on the atoms and chemical bonds of chemical substances; the molecular sequence representations adopt the SMILES (Simplified Molecular Input Line EntrySystem) format; using the contrastive learning pre-training module, the molecular graph structures and molecular sequence representations of unlabeled compounds are pre-trained using contrastive learning to learn the general structural features of molecules and obtain pre-trained molecular embeddings; inputting the molecular graph structures and molecular sequence representations of chemical substances in the reproductive toxicity training set into the molecular encoder, and combining them with element-knowledge graphs to generate molecular characterization vectors that integrate knowledge; based on the multi-dimensional annotation information in the reproductive toxicity training set, by combining different phenotype endpoint labels, study type labels, generation labels, and sex labels, determining the number and type of binary classifiers to be constructed; each phenotype endpoint... A unique combination of endpoint-study type-generation-sex corresponds to a binary classifier. For each phenotypic endpoint-study type-generation-sex combination, a phenotypic-specific cue vector is constructed. This phenotypic-specific cue vector is fused with the molecular characterization vector of the corresponding chemical substance to obtain a phenotypic-enhanced molecular characterization vector. For each phenotypic endpoint-study type-generation-sex combination, a binary classifier is trained using the corresponding training samples and the phenotypic-enhanced molecular characterization vector. Hyperparameter optimization and model validation are performed on multiple trained binary classifiers. The optimized binary classifiers are combined to form a set of reproductive toxicity prediction models.

[0017] Furthermore, in S7, the optimal prediction model combination is used to predict the toxicity of the analyte based on its molecular structure, including: acquiring the molecular structure information of the analyte; performing multi-view representation transformation on the molecular structure of the analyte to generate its molecular graph structure and molecular sequence representation; inputting the molecular graph structure and molecular sequence representation of the analyte into a molecular encoder, and combining it with an element-knowledge graph to generate a molecular representation vector of the analyte; for each binary classifier in the optimal prediction model combination, fusing the corresponding phenotype-specific cue vector with the molecular representation vector of the analyte to obtain a phenotype-enhanced molecular representation vector of the analyte. The optimal prediction model combination uses each binary classifier to predict the phenotypic enhanced molecular characterization vector of the corresponding analyte, obtaining toxicity activity prediction results for a specific phenotypic endpoint-study type-generation-sex combination. The prediction results of all binary classifiers in the optimal prediction model combination are summarized and categorized according to study type, sex, generation, and phenotypic endpoint. Study types include prenatal developmental toxicity studies, first-generation reproductive toxicity studies, and multi-generation reproductive toxicity studies. A multi-dimensional reproductive toxicity prediction report is generated, which includes cross-sex, cross-generation, and cross-study type prediction results. The prediction results report includes the toxicity activity prediction, prediction confidence level, and toxicity risk level for each phenotypic endpoint.

[0018] Another aspect of this application provides a multi-terminal reproductive toxicity prediction system based on a large language model, comprising: a data acquisition module for acquiring a multi-source dataset containing chemical substances and their reproductive toxicity research records; a reliability screening module for performing reliability screening on the multi-source dataset to obtain candidate datasets; a phenotypic information extraction module for extracting phenotypic information from the text data in the candidate datasets using a large language model to obtain a standardized phenotypic-level annotated dataset; a multidimensional annotation module for performing multidimensional annotation on the phenotypic-level annotated dataset to obtain a reproductive toxicity training set containing phenotypic resolution; and a model training module for inputting the reproductive toxicity training set into a pre-constructed molecular characterization learning model for training to construct a phenotypic-specific reproductive toxicity model. The system includes a set of sex prediction models; a molecular characterization learning model comprising an element-knowledge graph-based molecular encoder and a pre-training module based on contrastive learning; a set of reproductive toxicity prediction models comprising multiple binary classifiers trained for different phenotype-study type-generation-sex combinations; a performance evaluation module that evaluates the performance of multiple classifiers in the reproductive toxicity prediction model set using a test dataset and selects the optimal prediction model combination based on preset performance indicators; the optimal prediction model combination comprising multiple prediction models for different toxicity endpoints; and a toxicity prediction module that uses the optimal prediction model combination to predict the toxicity of the test chemical substance based on its molecular structure, obtaining multi-dimensional reproductive toxicity prediction results that include cross-sex, cross-generation, and cross-study type characteristics.

[0019] Compared to existing technologies, the advantages of this application are: The first reproductive toxicity prediction model based on phenotypic level annotation was constructed. By using large language model-assisted extraction technology, the model achieved a leap from research-level annotation to phenotypic level annotation, refining the prediction granularity from overall toxicity to specific phenotypic endpoints (such as reproductive organ weight, sperm quality, embryonic development, etc.). Specific predictors were trained for different phenotypic-research-generation-sex combinations, enabling the model to accurately capture the biological complexities such as sex differences, intergenerational transmission effects, and developmental stage specificity observed in traditional animal experiments. This fundamentally solves the problem of insufficient phenotypic resolution in existing models and provides high-precision, multi-dimensional predictive capabilities for the risk assessment of chemical reproductive toxicity. Attached Figure Description

[0020] Figure 1 This is a flowchart of the multi-endpoint reproductive toxicity phenotypic data integration and predictive modeling method based on expert collaboration proposed in this application; Figure 2 This paper presents a comparison of the performance of expert manual extraction with the Large Language Model (GPT-4.0) benchmark test and hybrid strategy. Figure 3 This is a comparison of the accuracy and time of expert evaluation of this application with mixed and single extraction of different Large Language Models (LLMs); Figure 4 This is a schematic diagram of the AnimalRepTox model development in this application; Figure 5 This application demonstrates the predictive performance of the model on the male and female reproductive systems. Figure 6 The model in this application shows how its predictive performance varies with experimental complexity in both males and females. Figure 7 This is a comparison of the AnimalRepTox model of this application with nine QSAR models on four other additional indicators of male reproductive phenotype; Figure 8 This is a comparison of the AnimalRepTox model of this application with nine QSAR models on four other additional indicators of female reproductive phenotype; Figure 9 This is the external validation result of the AnimalRepTox model of this application for prenatal development of male reproductive phenotypes; Figure 10 This is the external validation result of the AnimalRepTox model of this application for single-generation reproduction studies of female reproductive phenotypes; Figure 11This is the external validation result of the AnimalRepTox model of this application for multigenerational reproduction studies of male reproductive phenotypes; Figure 12 This is an external validation result of the AnimalRepTox model in this application for prenatal development studies of female reproductive phenotypes; Figure 13 This is the external validation result of the AnimalRepTox model of this application for single-generation reproduction studies of female reproductive phenotypes; Figure 14 This is the external validation result of the AnimalRepTox model of this application for multigenerational reproduction studies of female reproductive phenotypes; Figure 15 This is the predicted reproductive toxicity result of the Class 1A chemical in this application under the prenatal developmental toxicity paradigm; Figure 16 This is the predicted reproductive toxicity result of the Class 1A chemical in this application under the single-generation reproductive toxicity paradigm; Figure 17 This is the predicted reproductive toxicity result of the Class 1A chemical in this application under the multigenerational reproductive toxicity paradigm. Detailed Implementation

[0021] The present application will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0022] Example 1 Step S1: Obtain multi-source datasets containing research records on chemical substances and their reproductive toxicity from scientific literature databases and regulatory databases. Step S2: Organize the multi-source datasets and have experts conduct reliability assessments to form a shortlisted candidate dataset. Step S3: Utilize a phenotypic extraction strategy combining experts and Large Language Models (LLM) to construct a standardized phenotypic-level annotated dataset. Step S4: Preprocess the annotated dataset, including multi-dimensional annotation and conflict rule handling, to form a reproductive toxicity training set with phenotypic resolution. Step S5: Input the training set into a knowledge graph-enhanced molecular contrastive learning framework for model training and optimization, constructing a set of phenotypic-specific reproductive toxicity prediction models. Step S6: Use a test dataset to evaluate model performance based on multiple classification metrics and select the optimal prediction model set. Step S7: Input a list of chemical substances to be tested into the optimal model set to generate high-resolution reproductive toxicity prediction results across sexes, generations, and multiple research paradigms, thereby identifying potential reproductive toxins.

[0023] In one embodiment of this application, step S2, the sorting process includes: first, automatically identifying and screening studies centered on chemicals using a chemical nominate entity recognition model (such as Chemlistem based on a recurrent neural network); second, having experts use ToxRTool software to evaluate the reliability of the studies using 21 standardized scoring criteria, with studies scoring 17 or higher considered reliable; and finally, verifying the results using machine learning methods, such as using the ASReview LAB platform to retrospectively verify the results of the manual screening through a simulation mode.

[0024] In a preferred embodiment of this application, step S3, to determine the most suitable language model architecture for phenotypic data extraction, firstly establishes an expert-guided manual extraction process, which serves as a performance evaluation baseline. Subsequently, comparative tests are conducted using general-purpose large language models (such as GPT-4 and GPT-4 mini (OpenAI), Gemini 2.0 (DeepMind), LLaMa-3.1 (Meta AI), and Gemma-3 (Google)). This process compares the adaptability, structured output capability, and controllability of different models in phenotypic recognition and information extraction. Based on a comprehensive consideration of result consistency and text parsing stability, the GPT-4 model is determined as the preferred language model for the collaborative phenotypic extraction process in this implementation, used in the subsequent expert-model collaborative extraction stage.

[0025] In one embodiment of this application, step S4, preprocessing, includes: adding metadata such as study type (prenatal development, first generation reproduction, multiple generations reproduction), sex (male or female), generation (P or F1), phenotypic endpoint, and activity label (active or inactive) to each phenotypic record; when different research results of the same chemical substance conflict, a conservative judgment rule is adopted, that is, as long as there is a positive phenotype, it is marked as active, thereby minimizing the risk of false negatives and conforming to the public health principles of toxicological risk assessment.

[0026] In one embodiment of this application, the model building process in step S5 includes: constructing an element-knowledge graph (ElementKG), combining molecular graphs and molecular sequence encoding, performing comparative learning pre-training on approximately 250,000 unlabeled compounds to obtain generalized molecular embeddings, and introducing phenotypic awareness cues for fine-tuning in downstream training; training a binary classifier based on a combination of phenotypic-research type-generation-sex, and using resampling, data augmentation, class weighting, and focus loss to address the data imbalance problem; and finally forming an optimal set of prediction models through hyperparameter optimization and model integration.

[0027] In a preferred embodiment of this application, the performance evaluation in step S6 is based not only on accuracy, precision, recall and F1 score, but also on a comprehensive assessment of model performance using ROC-AUC and PR-AUC indices. PR-AUC is used to highlight the model’s performance in minority class predictions to ensure the ability to identify sensitive compounds in reproductive toxicity predictions.

[0028] In a preferred embodiment of this application, the prediction results of step S7 can output a multidimensional toxicity profile that is cross-sex, cross-generational, and cross-research paradigm, rather than a single binary classification result, thereby providing support for the study of chemical mechanisms, risk prioritization, and development of non-animal testing strategies.

[0029] In a further embodiment of this application, the method is applicable to a variety of reproductive-related toxicity endpoints, including but not limited to prenatal developmental toxicity, single-generation reproductive toxicity, and multi-generation reproductive toxicity; for different toxicity endpoints, sub-models can be constructed separately and integrated into a secondary model, thereby improving the predictive applicability under different research paradigms.

[0030] In a further embodiment of this application, the method can also be applied to an external independent validation set to validate the method using research data from different sources and with different experimental designs, thereby further demonstrating the generalization ability and robustness of the method.

[0031] Example 2 like Figure 1 As shown, in one embodiment of this application, a multi-terminal reproductive toxicity prediction model based on a large language model is provided. The method includes: step S1, obtaining a multi-source dataset containing research records of chemical substances and their reproductive toxicity from scientific literature databases and regulatory databases; step S2, organizing the multi-source datasets and having experts conduct reliability assessments to form a screened candidate dataset; and step S3, constructing a standardized phenotypic-level annotated dataset by adopting a phenotypic extraction strategy that combines experts and a large language model (LLM).

[0032] To comprehensively and systematically collect and annotate sex-specific, phenotypic reproductive toxicity data, we developed an expert-large language model collaborative strategy to integrate animal experimental literature and public databases from multiple sources.

[0033] Data was collected from both scientific literature and regulatory databases, with peer-reviewed literature and standardized toxicology reports as the primary data sources. Literature mining involved searching PubMed and Web of Science using predetermined reproductive toxicity-specific keywords, resulting in 75 keywords. Of these, 41 were related to male reproductive endpoints and 34 to female reproductive endpoints. Keywords included mechanistic descriptors (such as "endocrine disruption" and "maternal exposure") and specific phenotypes (such as "anorectal distance" and "vaginal opening").

[0034] The regulatory database data comes from four authoritative databases: the Canadian Classification Results (CCR), the European Chemicals Agency REACH Communication Portal (ECHA REACH), the Japan Chemicals Collaboration Knowledge Database (J-CHECK), and the OECD Existing Chemical Screening Information Dataset (OECD SID IUCLID). All databases contain standardized experimental procedures and structured toxicity endpoint information.

[0035] To improve the relevance and consistency of the data, the literature was first filtered through multiple layers. Non-English literature, duplicate records, literature lacking full text, studies not using rodents as a model, and non-research literature (reviews, epidemiological studies) were not included.

[0036] An automated literature review was then performed using a model based on a recurrent neural network (RNN) (Chemlistem32). Chemlistem can automatically identify chemical nominate entities from the titles and abstracts of given documents. Based on the identified chemical names, documents focusing on exogenous organic compounds were retrieved, while those focusing on metals, amino acids, nucleosides, purines, and pyrimidines were deleted.

[0037] In one embodiment of this application, a dual strategy of expert manual evaluation and machine learning verification was used to ensure the scientific rationality and consistency of the collected reproductive toxicity research literature.

[0038] A review panel of five experts from national institutions, industry, and academia conducted manual quality reviews of the selected research literature. The expert panel provided professional guidance and feedback throughout the data collection and management process to ensure the scientific validity and applicability of the final literature data. In the specific implementation, the collected literature was divided into five equal parts and assigned to the five experts for evaluation. Each expert first used ToxRTool software to assess the reliability of the literature. This software provides 21 criteria (Table 1) covering five aspects: test substance identification, test organism characterization, study design description, study results documentation, and the reasonableness of the study design and results. Experts scored each criterion individually, assigning a value of 1 to those that met the criteria and 0 to those that did not. Each literature received a total score (Score_total) ranging from 0 to 21 points. Literature with a total score of 17 points or higher was considered reliable and proceeded to the subsequent phenotypic data extraction stage.

[0039] Table 1. Detailed Evaluation Criteria for ToxRTool

[0040] To further improve the consistency of literature inclusion decisions and reduce potential biases caused by expert fatigue and subjective differences, this application introduces machine learning-based tools (such as the ASReview LAB tool) to retrospectively validate the results of manual screening. In one embodiment of this application, ASReview's "Simulation mode" was used to retrospectively validate 24 manually screened literature datasets from 2000 to 2023. Each dataset contained approximately 20,000 records, and the validation was run independently five times. In a specific implementation, the simulation mode of this tool was used to construct a simulation scenario for the 24 manually screened datasets. First, a portion of relevant and irrelevant studies were manually labeled as prior data for training a classifier model. Subsequently, this classifier was applied to predict the inclusion decision of the remaining corpus. Experts systematically changed the number and content of prior labels and tested various algorithm configurations, including different feature extraction methods (TF-IDF), classifiers (Naive Bayes), query strategies (uncertainty sampling), and class balancing methods (dynamic resampling), to iteratively optimize the research inclusion boundary. After multiple rounds of repeated validation, as shown in Table 2, the classifier model achieved a recall rate of 1.0 in all 24 datasets, and was able to identify all relevant literature determined by experts within the top 15% of the selected articles.

[0041] Table 2. Validation results of reproductive toxicity-related literature selected based on the ASReview model.

[0042] The regulatory report database reports are reviewed by the same team of experts using the same evaluation criteria to ensure methodological consistency, but due to its standardized structure, no additional computational validation is required.

[0043] This collaborative filtering method involving experts and artificial intelligence ultimately yielded 4,006 high-confidence studies (2,638 from literature and 1,368 from databases), laying a solid foundation for subsequent phenotypic annotation.

[0044] We first benchmarked the performance of expert-guided manual extraction versus general large language models. In one embodiment of this application, five general large language models were selected (GPT-4 and GPT-4 mini (OpenAI), Gemini2.0 (DeepMind), LLaMa-3.1 (MetaAI), and Gemma-3 (Google)). A total of 79 peer-reviewed English articles were selected, representing diverse reproductive toxicity studies covering different chemical categories, study designs (e.g., prenatal development, first-generation reproduction, and multi-generation reproduction), and phenotypic endpoints. Table 3 shows a comparison of the evaluation of each extraction modality in the mixed extraction strategy. Manual annotation by five toxicology experts achieved the highest accuracy of 80%, but each study required 10-30 minutes. In contrast, LLMs completed extraction within 0.1-1.5 minutes but showed lower accuracy. Figure 2 In a preferred embodiment of this application, LLaMa-3.1 exhibited the worst accuracy, while GPT-4 was selected as the preferred large language model due to its overall superior performance compared to other models. Figure 3 ).

[0045] Table 3. Performance Comparison of Expert, Large Language Model, and Expert-GPT-4.0 Hybrid Extraction Strategies in Phenotypic Extraction Tasks

[0046] In the cross-validation extraction strategy assisted by manual annotation and large language models, two workflows were implemented based on the complexity of the research: For studies with simple structures (e.g., single-generation, clearly defined phenotypes, and standardized formats), LLMs use predefined prompts for initial extraction, followed by expert review, correction, and validation of the output. The model also assists experts in locating supporting evidence in the original text, improving review efficiency.

[0047] For complex studies (e.g., multi-generational designs, overlapping exposure windows, high term density), experts perform initial extraction followed by LLM-assisted validation, as illustrated in one example of reproductive toxicity data in Table 4. The model cross-checks the expert output and identifies potential omissions or inconsistencies for expert review.

[0048] Ultimately, a series of high-resolution multi-endpoint reproductive toxicity datasets were generated through a hybrid extraction strategy (Tables 5-6), which included datasets of male (Table 5) and female (Table 6) reproductive system-related chemicals and reproductive phenotypes.

[0049] This method significantly improves labeling accuracy and efficiency. The collection covers four major experimental paradigms and two widely used model organisms, reflecting significant differences in study design, exposure window, and biological context.

[0050] Table 4. Examples of reproductive toxicity data extracted from the literature.

[0051] Table 5 Summary of data on chemical substances and reproductive phenotypes related to the male reproductive system

[0052] Table 6 Summary of data on chemical substances and reproductive phenotypes related to the female reproductive system

[0053] Example 3 This application provides a multi-terminal reproductive toxicity prediction model, AnimalRepTox, based on a large language model of knowledge graph-enhanced molecular contrastive learning and functional cues (KANO). Figure 4 The method includes: Step S4, preprocessing the annotated dataset, including multidimensional annotation and conflict rule handling, to form a reproductive toxicity training set with phenotypic resolution. Step S5, inputting the training set into a knowledge graph-enhanced molecular contrastive learning framework for model training and optimization, and constructing a set of phenotypic-specific reproductive toxicity prediction models; In one embodiment of this application, to achieve phenotypic-specific toxicity prediction for structurally diverse chemicals, the application first performs systematic preprocessing on the managed reproductive toxicity dataset. The preprocessing steps include multidimensional metadata annotation for each phenotypic record, with the metadata including at least the following five experimental dimensions: study type (including prenatal developmental studies PRE, first-generation reproductive studies OGR, and multigenerational reproductive studies MGR), sex (male or female), generation (P generation or F1 generation), phenotypic endpoint, and binary activity label (active or inactive).

[0054] In another embodiment of this application, to address potential discrepancies in results for a single chemical in different studies, this application employs a conservative conflict resolution rule: when a chemical exhibits at least one activity result under the same study type, sex, generation, and phenotypic endpoint combination, the chemical is identified as the active substance of that phenotype. Through the above preprocessing steps, high-quality data records with phenotypic markers can be obtained, providing a reliable input data foundation for subsequent model construction and performance evaluation.

[0055] This application employs KANO, a collaborative pre-training strategy that integrates knowledge graph-enhanced molecular contrastive learning with functional prompting. In our study, 80% of the preprocessed dataset was used as the model training set, and the remaining 20% ​​was reserved as the test dataset. After model training and iterative optimization, as... Figure 5 As shown, this application develops an AnimalRepTox containing a set of 123 binary classification models, where each model corresponds to a specific combination of phenotype, study type, generation, and gender.

[0056] Step S6: Use the test dataset to evaluate the model performance based on multiple classification metrics, and select the optimal set of prediction models; The reproductive developmental toxicity prediction model established in this application demonstrates high predictive accuracy across different sexes and study types. The model performs best on the Prenatal Developmental Research (PRE) dataset, exhibiting higher precision and recall than both the First Generation Reproductive Research (OGR) and Multiple Generation Reproductive Research (MGR) datasets. Figure 6 As shown, the model has high applicability and stability in predicting toxicity in the early developmental stages, but its performance weakens as the complexity of the treatment increases.

[0057] To verify the effectiveness and reliability of the AnimalRepTox model in this application, a systematic performance evaluation and comparative experiment were conducted.

[0058] In terms of model evaluation, multiple evaluation metrics based on four standard classification indicators (true positive TP, false positive FP, true negative TN, and false negative FN) were used to quantitatively analyze the model's prediction results. Evaluation metrics included accuracy, precision, recall, F1 score, area under the receiver operating characteristic curve (ROC-AUC), and area under the precision-recall curve (PR-AUC).

[0059] Accuracy reflects the proportion of correctly predicted instances among all samples. Precision measures the proportion of true positives among all positive predictions, while recall quantifies the ability to identify true positives from all actual positives. The F1 score harmonizes precision and recall into a single metric. ROC-AUC evaluates performance at different classification thresholds by plotting the true positive rate (TPR, i.e., recall) against the false positive rate (FPR). In contrast, PR-AUC plots precision against recall, making it particularly suitable for evaluating model performance under imbalanced class distributions, emphasizing correct predictions of the minority (active) classes. Detailed formulas are as follows:

[0060]

[0061]

[0062]

[0063] Individual indices were calculated for each of the 123 phenotype-specific binary classifiers in the framework, and the AnimalRepTox model was systematically compared with nine traditional machine learning-based quantitative structure-activity relationship (QSAR) models under the same dataset and research conditions. Figure 7 , Figure 8 As shown, AnimalRepTox demonstrates superior performance across multiple metrics, particularly outperforming all comparable models in ROC-AUC and PR-AUC. Furthermore, it maintains high stability and consistency in accuracy, recall, and F1 score, achieving robust predictions even under conditions of data sparsity, endpoint heterogeneity, and imbalanced samples.

[0064] In one embodiment of this application, such as Figure 9As shown in Figure 14, to assess robustness under inter-laboratory variability, external validation was performed using the ToxRefDB toxicity reference database. This database contains in vivo reproductive toxicity studies conducted primarily between 1900 and 2010, which differs from the 2000-2023 dataset on which AnimalRepTox is based in terms of time and experimental procedures, thus creating strict temporal consistency testing conditions. The validation types included results predicting male and female reproductive outcomes in three main study types (prenatal development studies, single-generation reproductive studies, and multi-generation reproductive studies). The results show that the AnimalRepTox model in this application maintains high ROC-AUC and PR-AUC performance across different sex- and generation-specific endpoints, demonstrating good predictive stability and generalization ability under cross-time and cross-study type data conditions.

[0065] To verify the reliability of the model's predictions, its Applicability Domain (AD) was evaluated. The Applicability Domain is defined as the chemical structure space in which the model's predictions can be considered reliable. This evaluation is particularly crucial when chemical structures are the sole input to the model, as the structural similarity between new compounds and training compounds largely determines the confidence level of the predictions. In this embodiment, the structural coverage of the training set was depicted using kernel density estimation (KDE), and the Tanimoto similarity metric was used to test the degree of structural similarity between compounds and training compounds. Based on the similarity results, the predictions were stratified into conclusive (within the AD) and non-conclusive (outside the AD) categories to improve the interpretability and confidence level of the model's predictions. Analysis results show that AD stratification significantly improves the model's classification performance across different research types. Furthermore, although some compounds are located outside the predefined chemical space, AnimalRepTox can still correctly predict approximately 45% of them, indicating that the model has good generalization ability outside the training space, providing a reliable basis for the toxicity prediction of new structural compounds.

[0066] In one embodiment of this application, such as Figure 15 As shown in Figure 17, this application, in order to verify the species extrapolation capability of the model, conducted external validation of the AnimalRepTox model based on 21 Class 1A reproductive toxins in three paradigm studies (prenatal development, single-generation reproduction, and multi-generation reproduction). These chemicals are formally classified as substances with confirmed reproductive toxicity to humans according to CLP regulations.

[0067] Although official documentation does not disclose endpoint-level toxicity details, this embodiment systematically screened the aforementioned compounds across all endpoints predictable by the model. Validation results showed that AnimalRepTox successfully identified all 21 evaluable compounds, achieving a 100% detection rate, and observed significant predictive interference signals in multiple phenotypic clusters. These results not only confirm the broad reproductive toxicity profile of Class 1A substances but also demonstrate that the model in this application can accurately capture human-relevant toxicological signals, even though the model was trained entirely on animal-derived data.

[0068] Step S7: Input the list of chemical substances to be tested into the optimal model set to generate high-resolution reproductive toxicity prediction results that are cross-sex, cross-generational and multi-research paradigm, thereby identifying potential reproductive toxins.

[0069] In one embodiment of this application, to demonstrate the practical utility and global relevance of AnimalRepTox, we applied the model to screen a comprehensive set of industrial and commercial chemical substances previously compiled from 19 countries. These lists included 11 developed countries and territories (EU, USA, Canada, Nordic countries, South Korea, Australia, New Zealand, Japan) and 8 developing countries (China, Mexico, Philippines, Turkey, Thailand, Russia, Chile, Vietnam) (Table 5). After excluding entries with confidentiality restrictions and structural information, 270,306 structurally well-defined substances were ultimately retained for model prediction.

[0070] Table 719 National Chemical Substance Inventory Information

[0071] A compound was considered a potential reproductive toxicant if it exhibited any predicted activity across 46 male or female phenotypic endpoints in at least one of the three study types (prenatal development (PRE), one-generation reproduction (OGR), or multiple-generation reproduction (MGR) studies). After filtering by applied domain assignment, AnimalRepTox identified 70,441 substances (26.06%) as potential reproductive toxicants, while 199,865 (73.94%) did not show predicted reproductive toxicity. The application domain filtering enabled AnimalRepTox to identify a subset of potential reproductive toxicants while showing low overlap between different study types, indicating that each study paradigm can capture different biological response windows and toxicokinetics, validating the necessity of a multi-paradigm integration strategy in comprehensively assessing reproductive risks.

[0072] AnimalRepTox is an artificial intelligence platform that combines deep learning technology with a systematically organized, multi-dimensional reproductive toxicity phenotypic database to achieve high-resolution predictions of chemically induced reproductive effects. The model extracts molecular features from input chemical structures and constructs mapping relationships between structures and various phenotypic endpoints, thereby enabling toxicity predictions across sexes, generations, and multiple research paradigms.

[0073] Users can upload a SMILES string or input a chemical structure using the platform's built-in molecular editor. After submission, the model characterizes the molecule, including molecular descriptors and structural fingerprints, and uses the trained deep learning network to generate qualitative toxicity predictions. The prediction results not only show the potential toxic endpoints of the chemical substance, but also assess the prediction confidence by combining applicable domain assignment, and can provide relevant structural warnings, thereby revealing the potential mechanistic association between molecular structure and toxic phenotype.

[0074] The predicted results are presented through an interactive graphical interface, highlighting potential toxicity endpoints. Users can view detailed information such as experimental paradigms, phenotypic descriptions, and toxicity status. Results can also be exported as tables or standard file formats, facilitating integration with downstream analyses and risk assessments. This process enables users to gain mechanistic insights into specific compounds based on their understanding of the relationship between molecular structure and reproductive toxicity.

[0075] Example 4 In one embodiment of this application, in order to identify environmental chemicals associated with changes in human anorectal distance (AGD), a systematic literature search strategy was employed to retrieve relevant research literature from two major literature databases, PubMed and Web of Science (WoS), with the search scope extending to June 2025.

[0076] Literature retrieval employed a combination of controlled terms and free text keywords, constructing a composite query encompassing three thematic domains: (i) AGD variation (including terms such as "anogenital distance", "AGD", and "anogenital index"); (ii) environmental chemical exposure (including phthalates, bisphenols, perfluorinated and polyfluoroalkyl substances (PFAS), organophosphorus pesticides, etc.); and (iii) epidemiological research background (including "cohort studies", "prenatal exposure", and "population studies").

[0077] The above search yielded 516 relevant publications, including 222 from PubMed (dating back to 1998-2025) and 294 from WoS (dating back to 1994-2025). After filtering the publications based on research relevance, publication type, and recentity (limited to 2015-2025), 235 studies with epidemiological relevance were ultimately retained (79 from PubMed and 156 from WoS).

[0078] In a preferred embodiment of this application, a comprehensive manual evaluation of the 235 screened studies identified 72 environmental chemicals with statistically significant correlations to changes in population AGD. Among these, phthalates, perfluoroalkyl substances (PFAS, including PFOA), bisphenol A (BPA), polychlorinated biphenyls (PCBs), bisphenol F / S / AF, polybrominated diphenyl ethers (PBDEs), vinclozolin, and diethylstilbestrol have been repeatedly verified in multiple studies to be significantly correlated with AGD shortening, indicating that existing experimental results are consistent with predicted results.

[0079] The foregoing illustrative description of the present application and its embodiments is not restrictive and can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. The accompanying drawings are only one embodiment of the present application, and the actual structure is not limited thereto. Therefore, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the present application, such designs should fall within the scope of protection of this application. Furthermore, the word "comprising" does not exclude other elements or steps, and the word "a" preceding an element does not exclude the inclusion of "a plurality" of that element. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.

Claims

1. A multi-terminal reproductive toxicity prediction method based on a large language model, characterized in that, include: S1, Obtain a multi-source dataset containing records of studies on chemical substances and their reproductive toxicity; S2, perform reliability screening on the multi-source dataset to obtain candidate datasets; S3 utilizes a large language model to extract phenotypic information from text data in the candidate dataset, resulting in a standardized phenotypic-level annotated dataset. S4. Perform multidimensional annotation processing on the phenotypic-level annotated dataset to obtain a reproductive toxicity training set containing phenotypic resolution; the multidimensional annotation processing includes adding study type label, sex label, generation label, phenotypic endpoint label and activity label to each phenotypic record; S5. Input the reproductive toxicity training set into the pre-constructed molecular characterization learning model for training, and construct a set of phenotype-specific reproductive toxicity prediction models. The molecular characterization learning model includes a molecular encoder based on element-knowledge graph and a pre-training module based on contrastive learning. The set of reproductive toxicity prediction models includes multiple binary classifiers trained for different phenotype-research type-generation-sex combinations. S6. The performance of multiple classifiers in the reproductive toxicity prediction model set is evaluated using the test dataset, and the optimal prediction model combination is selected according to the preset performance indicators. The optimal prediction model combination includes multiple prediction models for different toxicity endpoints. S7 utilizes the optimal prediction model combination to predict the toxicity of the chemical substance under test based on its molecular structure, and obtains multi-dimensional reproductive toxicity prediction results that include cross-sex, cross-generational and cross-study types; study types include prenatal developmental toxicity studies, first-generation reproductive toxicity studies and multi-generational reproductive toxicity studies.

2. The multi-terminal reproductive toxicity prediction method based on a large language model according to claim 1, characterized in that: S2 performs reliability screening on multi-source datasets, including: A chemical named entity recognition model was used to identify text data from multi-source datasets and filter out research data centered on chemicals. The research data were scored using a pre-defined reliability scoring standard. Research data with scores greater than or equal to a preset threshold are used as the initial screening dataset; A machine learning validation model is used to retrospectively validate the initially screened dataset to confirm the accuracy of the screening results and form a candidate dataset.

3. The multi-terminal reproductive toxicity prediction method based on a large language model according to claim 2, characterized in that: The chemical named entity recognition model adopts the Chemlistem model based on recurrent neural networks.

4. The multi-terminal reproductive toxicity prediction method based on a large language model according to claim 2, characterized in that: S3 yields a standardized phenotypic-level annotated dataset, including: Obtain a pre-built phenotypic extraction baseline dataset as a performance evaluation criterion; the phenotypic extraction baseline dataset includes labeled phenotypic information samples; Multiple candidate large language models are selected to perform phenotypic recognition on the text data in the candidate dataset; Based on the baseline dataset for phenotypic extraction, the phenotypic recognition results of each candidate large language model are evaluated for adaptability, structured output capability, and controllability. Based on the results of adaptability assessment, structured output capability assessment, and controllability assessment, the candidate large language model with the best performance is selected as the target extraction model based on consistency and text parsing stability. Phenotypic information was extracted from the text data in the candidate dataset using a target extraction model to obtain raw annotation data containing reproductive toxicity phenotypic endpoints; The original annotation data is standardized to form a standardized phenotypic-level annotation dataset.

5. The multi-endpoint reproductive toxicity prediction method based on a large language model according to claim 4, characterized in that: Candidate large language models include at least two of GPT-4, GPT-4 mini, Gemini 2.0, LLaMa-3.1, and Gemma-3.

6. The multi-terminal reproductive toxicity prediction method based on a large language model according to claim 4, characterized in that: Phenotypic information includes at least three of the following: chemical substance name, toxic phenotype description, dosage information, exposure time, and experimental species.

7. The multi-terminal reproductive toxicity prediction method based on a large language model according to claim 4, characterized in that: S4 performs multidimensional annotation processing on the phenotypic level annotated dataset, including: Based on the preset research type classification rules, a research type label is added to each phenotypic record in the phenotypic-level annotated dataset; the research type labels include prenatal developmental toxicity research labels, first-generation reproductive toxicity research labels, and multi-generation reproductive toxicity research labels. Based on preset gender recognition rules, a gender label is added to each phenotypic record in the phenotypic-level annotated dataset; the gender label includes a male label and a female label. Based on the preset generational division rules, generational labels are added to each phenotypic record in the phenotypic-level annotated dataset; generational labels include parental labels, first offspring labels, and second offspring labels. Based on a pre-defined phenotypic endpoint classification system, phenotypic endpoint labels are added to each phenotypic record in the phenotypic-level annotated dataset; the phenotypic endpoint labels correspond to specific reproductive toxicity phenotypic categories. Based on preset activity determination criteria, an activity label is added to each phenotypic record in the phenotypic-level annotated dataset; the activity label is used to identify whether the corresponding phenotypic record shows toxic activity. When conflicting research findings on the same chemical substance are presented, a conservative determination rule is used for activity labeling. The conservative determination rule is: when at least one positive phenotype is present, the corresponding chemical substance is labeled as active. The phenotypic records with completed multidimensional annotations are combined to form a reproductive toxicity training set.

8. The multi-terminal reproductive toxicity prediction method based on a large language model according to claim 7, characterized in that: S5, Construct a set of phenotype-specific reproductive toxicity prediction models, including: Obtain a pre-built molecular representation learning model; the molecular representation learning model includes a molecular encoder based on element-knowledge graph and a pre-training module based on contrastive learning; Multi-perspective characterization and transformation of the molecular structures of chemicals in the reproductive toxicity training set are performed to generate molecular diagram structures and molecular sequence characterizations. The molecular diagram structure and molecular sequence characterization of unlabeled compounds are pre-trained using a contrastive learning pre-training module to obtain pre-trained molecular embeddings. The molecular graph structure and molecular sequence representation of the chemical substances in the reproductive toxicity training set are input into the molecular encoder, and combined with the element-knowledge graph, a molecular representation vector with fused knowledge is generated. Based on the multidimensional annotation information in the reproductive toxicity training set, determine the number and type of binary classifiers that need to be constructed; For each phenotypic endpoint-study type-generation-sex combination, a phenotypic-specific cue vector is constructed. The phenotypic-specific cue vector is then fused with the molecular characterization vector of the corresponding chemical substance to obtain a phenotypic-enhanced molecular characterization vector. For each phenotypic endpoint-study type-generation-sex combination, a binary classifier is trained using the corresponding training samples and phenotypic-enhanced molecular representation vectors. Hyperparameter optimization and model validation are performed on multiple trained binary classifiers; The optimized binary classifiers are combined to form a set of reproductive toxicity prediction models.

9. The multi-endpoint reproductive toxicity prediction method based on a large language model according to claim 8, characterized in that: S7, using the optimal prediction model combination to predict the toxicity of the chemical substance under test based on its molecular structure, including: To obtain molecular structure information of the chemical substance to be tested; The molecular structure of the chemical substance to be tested is transformed from multiple perspectives to generate molecular map structure and molecular sequence characterization of the chemical substance to be tested. The molecular graph structure and molecular sequence characterization of the chemical substance to be tested are input into the molecular encoder, and combined with the element-knowledge graph, a molecular characterization vector of the chemical substance to be tested is generated. For each binary classifier in the optimal prediction model combination, the corresponding phenotype-specific cue vector is fused with the molecular characterization vector of the chemical substance to be tested to obtain the phenotype-enhanced molecular characterization vector of the chemical substance to be tested. By using each binary classifier in the optimal prediction model combination to predict the phenotypic enhanced molecular characterization vector of the corresponding analyte, the toxicity activity prediction results for a specific phenotypic endpoint-study type-generation-sex combination are obtained. Summarize the prediction results of all binary classifiers in the optimal prediction model combination, and classify them according to research type, gender, generation, and phenotypic endpoint. Generate multi-dimensional reproductive toxicity prediction reports that include cross-sex, cross-generational, and cross-study types.

10. A multi-endpoint reproductive toxicity prediction system based on a large language model, characterized in that, include: The data acquisition module acquires multi-source datasets containing research records on chemical substances and their reproductive toxicity. The reliability screening module performs reliability screening on the multi-source dataset to obtain candidate datasets; The phenotypic information extraction module uses a large language model to extract phenotypic information from the text data in the candidate dataset, resulting in a standardized phenotypic-level annotated dataset. The multidimensional annotation module performs multidimensional annotation processing on the phenotypic level annotated dataset to obtain a reproductive toxicity training set containing phenotypic resolution. The model training module inputs the reproductive toxicity training set into a pre-constructed molecular characterization learning model for training, thereby constructing a set of phenotype-specific reproductive toxicity prediction models; the molecular characterization learning model includes a molecular encoder based on element-knowledge graphs and a pre-training module based on contrastive learning. The set of reproductive toxicity prediction models includes multiple binary classifiers trained for different phenotype-study type-generation-sex combinations; The performance evaluation module uses a test dataset to evaluate the performance of multiple classifiers in the reproductive toxicity prediction model set, and selects the optimal prediction model combination based on preset performance indicators; the optimal prediction model combination includes multiple prediction models for different toxicity endpoints. The toxicity prediction module uses the optimal prediction model combination to predict the toxicity of the chemical substance to be tested based on its molecular structure, and obtains multi-dimensional reproductive toxicity prediction results that include cross-sex, cross-generational and cross-study types.

Citation Information

Patent Citations

  • Drug-similar compound toxicity predicating method based on deep learning

    CN109658989A

  • A method for constructing a drug target toxicity prediction model and its application

    CN113380341B

  • Method for mining literature source toxicity effect test information

    CN119166752A