An Artificial Intelligence-Based Prognostic Assessment System and Method for Gastric Cancer
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- REHABILITATION UNIVERSITY QINGDAO CENTRAL HOSPITAL
- Filing Date
- 2026-01-26
- Publication Date
- 2026-05-26
Smart Images

Figure CN122091184A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of prognostic monitoring and assessment technology, and more specifically, to an artificial intelligence-based prognostic assessment system and method for gastric cancer. Background Technology
[0002] With the deep integration of big data and artificial intelligence technologies, building general-purpose computing systems and software development paradigms for intelligent modeling and feature discovery has become a key technological requirement in data-intensive research fields. In traditional data processing technologies for data-intensive analysis tasks, the implementation logic relies on users manually arranging the execution order of discrete tools to form a linear data flow processing pipeline that is highly dependent on the operator's experience. However, this approach has significant technical bottlenecks: the process construction and execution heavily rely on manual operation and experience judgment, resulting in low automation. Furthermore, the selection of key genes largely depends on a few algorithms, making it susceptible to subjective influence and leading to insufficient robustness and interpretability of feature discovery results.
[0003] To address the issues of fragmented processes and low efficiency in traditional technologies, existing technologies encapsulate some commonly used computational functions into reusable software libraries or algorithm modules. This refactors the specific computational steps, which were originally implemented in the form of scripts, into modular software components with standard input and output interfaces. These components are then accessible through command-line tools or simple graphical interfaces, lowering the operational threshold for non-professionals and improving the standardization of individual analysis steps to some extent.
[0004] However, in practical use, it still has some shortcomings. For example, due to the lack of a unified, intelligent, and reproducible computing system architecture at the underlying level, consulting experts have to spend a lot of energy on tedious data processing, toolchain debugging, and result integration, resulting in long response cycles for consulting services, difficulties in knowledge accumulation, insufficient large-scale delivery capabilities, and inadequate customer interaction. The integration of existing technologies is seriously insufficient, and a unified software architecture has not been built. The data dependencies and flow relationships between different modules cannot be automatically identified and managed, still requiring a lot of manual intervention from users. This results in human-caused breakpoints in the data processing flow, limited overall system throughput efficiency, and difficulty in accurately reproducing the process, which seriously affects the automation level and reliability of complex data analysis tasks. In the core screening stage, there is a lack of a configurable intelligent data fusion and feature ranking model layer, which makes the feature screening logic rigid and inflexible. Moreover, the ranking results are easily affected by noise from a single data source, making it difficult to stably and efficiently identify high-value targets from a massive number of candidate entities, thus restricting the decision-making intelligence and discovery capabilities of data analysis. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art, the present invention provides an artificial intelligence-based gastric cancer prognostic assessment system and method, which addresses the problems mentioned in the background art through the following solutions.
[0006] To achieve the above objectives, the present invention provides the following technical solution: An artificial intelligence-based prognostic assessment system for gastric cancer includes: Data acquisition module: used to preprocess the acquired first state data and corresponding category labels to generate the first entity feature; The difference queue management module is used to divide the first entity feature into samples to generate a comparison queue, and to process the comparison queue through a difference analysis algorithm to generate the second entity feature. Filtering and sorting module: Based on the second entity features, it outputs high-priority third entity features through a preset weighted decision model and combined with various topological analyses; Multi-dimensional association analysis module: used to perform at least one of the following analyses on the third entity feature: association analysis with the category label, classification performance analysis of data samples, and enrichment analysis based on a preset function set, to generate a multi-dimensional evaluation report of the third entity feature; The process coordination and optimization module supports users in making agile configurations based on templates to quickly respond to analysis needs. It schedules the execution order and data flow between the data acquisition module, the difference queue management module, the filtering and sorting module, and the multi-dimensional correlation analysis module. Based on the comparison between the external verification report and the multi-dimensional evaluation report, it optimizes the parameters of the weighted decision model in the filtering and sorting module.
[0007] Preferably, the data acquisition module, in acquiring the first state data, specifically includes: In response to the user-specified data source identifier and acquisition task, the preset data source adapter corresponding to the target biomedical database is invoked; Through the data source adapter, at least two heterogeneous raw data sets are simultaneously obtained from the database in an automated manner. The raw data sets include at least: omics data containing quantitative gene expression values, and clinical information containing clinical attributes and status indicators of samples. The acquired omics data and the clinical information are logically linked through a common sample identifier to form the first state data with a clearly defined structure.
[0008] Preferably, the difference queue management module obtains the second entity feature, specifically including: Configurable grouping rules defined based on clinical information in the first entity features; The grouping rules are executed to filter out samples that meet the rule conditions from the first entity features, forming at least two sample sets, i.e., comparison queues; The grouping rules and the generated sample set identifier are used as metadata; At least two differential analysis algorithms are invoked in parallel to perform differential gene expression analysis on the comparison queue. Generate the second entity feature, which includes a saliency metric.
[0009] Preferably, in the difference queue management module, the significance measurement index in the second entity feature includes at least the comprehensive corrected p-value and the comprehensive effect value, specifically including: The original p values Perform fusion and calculate the comprehensive chi-square value. The specific formula is as follows: in, denoted as the original p-value calculated for the gene by the i-th differential analysis algorithm, and k represents the total number of differential analysis algorithms; The inverse variance weighted average method was used to calculate the effect values. To perform fusion and calculate the combined effect value Specifically, it is expressed as: in, It is represented as a weight, and its value can be the reciprocal of the variance of each effect value; Multiple tests were performed on the original p-values to obtain the overall corrected p-values.
[0010] Preferably, the filtering and sorting module has a preset weighted decision model configured to support at least two weight configuration modes, including: In static configuration mode, static weight vectors are manually set for the various topology analysis algorithms involved in the fusion based on prior knowledge. In dynamic optimization mode, the entropy weight method is automatically invoked to calculate and allocate weights based on the dispersion of the output scores of each topology analysis algorithm. The learning optimization model adjusts the weight configuration through regression analysis to maximize the correlation between the fused ranking results and the effective target priorities verified by external experiments.
[0011] Preferably, the filtering and sorting module performs various topology analyses, including at least two selected from the following set of algorithms: Maximum clique centrality, maximum neighborhood component, degree centrality, bottleneck, eccentricity, proximity centrality, radiality, betweenness centrality, stress, and clustering coefficient; Furthermore, the gene interaction network automatically constructed based on the second entity features executes the selected algorithm in parallel and outputs a topology score matrix with a uniform format.
[0012] Preferably, the multi-dimensional association parsing module, in its association analysis with the category label, is specifically configured to execute an automated prognostic value assessment pipeline, including: Obtain expression level data corresponding to the third entity feature and clinical label data including survival time and status; Automatically calculate the optimal cutoff value for dividing the high and low expression groups; Automatically invokes the survival analysis model to generate survival curves and significance P-values; Automatically perform univariate and multivariate proportional hazards regression analysis to calculate the hazard ratio and confidence interval of the third entity feature as a risk factor; The statistical indicators and charts generated automatically.
[0013] Preferably, the multi-dimensional association parsing module, for classifying the data samples, is specifically configured as follows: Based on the category labels, the samples are clearly divided into positive and negative sets; Calculate the classification performance of the third entity feature on samples under different expression thresholds, and generate ROC curves; Automatically calculate the area under the curve, the optimal classification threshold, and the corresponding sensitivity and specificity.
[0014] Preferably, the multi-dimensional association parsing module, based on enrichment analysis of a preset function set, is specifically configured as follows: Automatically generate a gene sequence list for enrichment analysis; Connect to a local feature annotation knowledge base that can be regularly synchronized and updated; Select biological pathways that are significantly related to the gene list from the local functional annotation knowledge base; Automatically generate textual hypotheses describing the underlying biological mechanisms of the features of the third entity.
[0015] To achieve the above objectives, the present invention provides the following technical solution: an artificial intelligence-based method for gastric cancer prognostic assessment, comprising implementing the aforementioned artificial intelligence-based gastric cancer prognostic assessment system, including: S1: Preprocess the acquired first state data and corresponding category labels to generate the first entity feature; S2: Divide the first entity feature into samples to generate a comparison queue, and process the comparison queue using a difference analysis algorithm to generate the second entity feature; S3: Based on the second entity feature, a high-priority third entity feature is output through a preset weighted decision model and combined with multiple topological analyses; S4: Perform at least one of the following analyses on the third entity feature: correlation analysis with the category label, classification performance analysis of the data sample, and enrichment analysis based on a preset function set, to generate a multidimensional evaluation report of the third entity feature; S5: Schedule the execution order and data flow between the data acquisition module, the difference queue management module, the filtering and sorting module, and the multi-dimensional correlation analysis module, and optimize the parameters of the weighted decision model in the filtering and sorting module based on the comparison between the external verification report and the multi-dimensional evaluation report.
[0016] The technical effects and advantages of this invention are as follows: 1. This invention improves the reliability of early-stage data for feature screening through a differential queue management module, reduces interference from noise in a single algorithm, enhances the anti-interference capability of the screening process, and improves the responsiveness and scalable delivery capability of information technology consulting services and customer interactions. 2. This invention constructs a configurable intelligent data fusion and feature ranking model layer through a screening and ranking module, breaking the limitation of fixed feature screening logic, enhancing model flexibility, reducing the interference of noise from a single data source on the ranking results, and realizing stable, efficient, and highly reliable priority ranking of massive candidate features, significantly improving the intelligence of data analysis. 3. This invention provides multi-dimensional verification basis for the screening results through a multi-dimensional correlation analysis module, which improves the accuracy and scientific nature of the screening decision, solves the defect of limited discovery capability of existing technologies, enhances the scientific nature and reliability of the analysis conclusions, and improves the depth and value of consulting service output. Attached Figure Description
[0017] Figure 1 This is a block diagram of an artificial intelligence-based gastric cancer prognostic assessment system provided according to an embodiment of this application.
[0018] Figure 2 This is a flowchart illustrating the steps of an artificial intelligence-based prognostic assessment method for gastric cancer according to an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to and includes any or all possible combinations of one or more of the listed items.
[0021] Hereinafter, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first," "second," and "third" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0022] As attached Figure 1 The system shown is an artificial intelligence-based prognostic assessment system for gastric cancer, which includes a data acquisition module, a differential cohort management module, a screening and sorting module, a multi-dimensional correlation analysis module, and a process coordination and optimization module.
[0023] The data acquisition module is used to preprocess the acquired first state data and the corresponding category label to generate the first entity feature.
[0024] Specifically, acquiring the first state data includes: responding to a user-specified data source identifier and acquisition task, invoking a preset data source adapter corresponding to the target biomedical database; and automatically acquiring at least two heterogeneous raw data sets from the database through the data source adapter, wherein the raw data sets include at least: omics data containing gene expression quantification values, and clinical information containing sample clinical attributes and state indicators; wherein the acquired omics data and the clinical information are logically associated through a common sample identifier to constitute the first state data with a clearly defined structure.
[0025] It should be noted that the data acquisition module utilizes a built-in extensible data source adapter library. Each data source adapter encapsulates the API interface or data publishing format of a specific data source. Users only need to specify the target dataset identifier through the configuration interface or upload a task list containing multiple identifiers. After a task is triggered, the data source adapter automatically executes the following: automatically processing the login token or access key required by the database; automatically acquiring the dataset's description information, sample list, file list, and corresponding clinical data tables; automatically initiating multi-threaded / asynchronous download requests based on the file list to efficiently acquire RNA-seq quantitative files, clinical information text files, mutation data, etc.; and ensuring that the data is not corrupted during transmission by comparing the MD5 / SHA verification of the files, and automatically retrying failed download tasks.
[0026] Furthermore, based on a predefined rule base for biomedical data, which matches based on file extensions, first-row header patterns, column name keywords, etc., to automatically identify the structure and semantics of the omics data and clinical information; according to the identified file structure, the corresponding parser is invoked; in this embodiment, for the expression matrix, the parser can identify the format where rows are genes and columns are samples, and automatically process possible annotation rows; for clinical data, it can identify heterogeneous headers used by different projects and map them to unified data model fields within the system.
[0027] Furthermore, the acquired data is preprocessed to generate the first entity feature that can be directly used by downstream modules. The preprocessing includes at least: pre-setting cleaning rules for different types of data; automatically detecting gene identifiers in the expression data and calling built-in or online mapping services to unify them into system-specified standard identifiers to solve the integration failure problem caused by identifier inconsistency; automatically performing appropriate standardization processing on RNA-seq expression levels according to differential analysis requirements; automatically verifying the consistency between the sample IDs in the expression matrix and the sample IDs in the clinical information, identifying and reporting mismatched samples, and automatically performing data cropping and alignment according to preset rules.
[0028] It should be noted that the first entity feature generated after the above preprocessing is a structured data object, the core of which includes: an expression matrix, where rows correspond to genes after the unified identifier, columns correspond to sample IDs after consistency verification, and matrix values are standardized gene expression quantification values; and a sample feature association table, where each row corresponds to a sample ID, and columns include various category labels extracted from clinical information, including but not limited to metastasis status, survival time, survival status, TNM stage, etc., as well as other clinical features, and are associated with the expression matrix in columns and rows through the sample ID.
[0029] In this embodiment, a gastric cancer prognostic study is used as the designated data source, identified as "TCGA-STAD". The TCGA adapter will be invoked to automatically download RNA-seq expression data files and clinical information files. After preprocessing, the first entity feature generated contains an expression matrix containing the expression values of tens of thousands of genes in hundreds of gastric cancer samples. The sample feature association table contains key category labels such as distant metastasis status, overall survival time, and survival status for each sample, which can be directly used to drive downstream modules to perform differential analysis.
[0030] The difference queue management module is used to divide the first entity feature into samples to generate a comparison queue, and to process the comparison queue through a difference analysis algorithm to generate the second entity feature.
[0031] Specifically, obtaining the second entity feature includes: a configurable grouping rule defined based on clinical information in the first entity feature; executing the grouping rule to select samples that meet the rule conditions from the first entity feature, forming at least two sample sets with clear biological comparative significance, i.e., a comparison queue; using the grouping rule and the generated sample set identifier as metadata; invoking at least two differential expression analysis algorithms in parallel to perform differential gene expression analysis on the comparison queue; and combining the output results of each algorithm to generate the second entity feature containing a significance metric, wherein the significance metric includes at least a comprehensive corrected p-value and a comprehensive effect value.
[0032] In this embodiment, the grouping rules can be configured as follows: Group 1 (non-metastasis group M0): samples with the value of 'M_stage' in the clinical table 'clinical' are selected as 'M0'; Group 2 (metastasis group M1): samples with the value of 'M_stage' in the clinical table are selected as 'M1'.
[0033] It should be noted that the second entity feature is a data table, where each row uniquely corresponds to a gene, and includes at least the following fields: gene identifier, comprehensive effect value, comprehensive original p-value, and comprehensive corrected p-value; the grouping rules, the sample set selected according to the rules, and the timestamp of rule execution are structured and stored as immutable metadata of the comparison queue; when the process coordination and optimization module or a user requests to reproduce the analysis, the same comparison queue can be automatically and accurately reconstructed based on this metadata, thereby ensuring the complete reproducibility of the analysis process; the differential analysis algorithm is encapsulated as an independent containerized analysis service, the core of which is differential gene expression analysis, all using algorithms known and standard in the field; the parallel call refers to sending computation tasks containing the same comparison queue data to the analysis service simultaneously through an asynchronous task scheduler, and concurrently collecting the return results of each service, thereby avoiding the serial execution of the algorithm, significantly shortening the time to obtain all candidate results, and improving the overall system throughput efficiency.
[0034] In one possible implementation, for each gene being analyzed, its original p-value is extracted and aligned from the results returned by each differential analysis algorithm. With effect size Calculate and generate significance metrics for each gene, including: the original p-values corresponding to each gene. Perform fusion and calculate the comprehensive chi-square value of the gene. Based on this, the overall original p-value of the gene is obtained, and the specific formula is as follows: in, The p-value is represented by the original p-value calculated for this gene by the i-th differential analysis algorithm, where k represents the total number of differential analysis algorithms. The inverse variance weighted average method is used to calculate the effect values corresponding to this gene. Perform fusion and calculate the overall effect value of the gene. Specifically, it is expressed as: in, The p-value is represented as the weight, and its value can be the reciprocal of the variance of each effect value. Multiple tests are performed on the composite original p-values of all genes to obtain the composite corrected p-value for each gene.
[0035] In this embodiment, the effect value This represents a log2 folding change.
[0036] The filtering and sorting module is used to output a high-priority third entity feature based on the second entity feature, which is structured data containing a list of differentially expressed gene identifiers obtained from the comparison queue and at least one significance metric for each gene, through a preset weighted decision model and combined with various topological analyses.
[0037] Specifically, the screening and sorting module parses the second entity feature, which contains at least a list of identifiers for differentially expressed genes and related statistics; the gene identifier list is submitted to a protein-protein interaction database such as STRING through a pre-defined application programming interface to obtain network data, and internally constructs a graph data structure consisting of genes as a set of nodes, interactions as a set of edges, and STRING confidence scores as edge weights; various topology analyses are performed based on the constructed network graph, including at least two selected from the following algorithm set: maximum clique centrality, maximum neighborhood component, degree centrality, bottleneck, eccentricity, proximity centrality, radiality, betweenness centrality, stress, and clustering coefficient; the gene interaction network automatically constructed based on the second entity feature executes the calculations of the selected algorithms in parallel and outputs a topology score matrix with a uniform format; the weighted decision model processes the topology score matrix; according to the determined weight vector, the topology scores of each gene are weighted and summed to calculate a comprehensive robust score, and the genes are sorted according to this score to generate the third entity feature, i.e., a sorted gene list with a comprehensive score and individual item scores.
[0038] It should be noted that the weighted decision model is configured to support at least two weight configuration modes, including: a static configuration mode, which allows users to manually set static weight vectors for the various topology analysis algorithms participating in the fusion based on prior knowledge; a dynamic optimization mode, which automatically calls the entropy weight method to calculate and allocate weights based on the data characteristics corresponding to the second entity features, and calculates and allocates weights according to the dispersion of the output scores of each topology analysis algorithm; and a learning tuning mode, which receives historical verification data from the process coordination and optimization module, and adjusts the weight configuration through regression analysis to maximize the correlation between the fused ranking results and the effective target priority verified by external experiments.
[0039] In this embodiment, the weighted decision model calculates the comprehensive robustness score for each gene, specifically expressed as follows: in, Let be the standardized topological score of the b-th gene calculated using the j-th topological analysis algorithm. The weight coefficient is represented by the j-th topology analysis method in the weighted decision model; the differential gene list of gastric cancer M0 and M1 groups is received, the PPI network is automatically constructed, the scores of 11 internally integrated centrality algorithms are calculated in parallel, the weights are determined by the entropy weight method, and high-ranking genes such as ACTG2 are obtained after fusion, thereby efficiently locking potential key targets.
[0040] The multi-dimensional association parsing module is used to perform at least one of the following analyses on the third entity feature: association analysis with the category label, classification performance analysis of the data sample, and enrichment analysis based on a preset function set, to generate a multi-dimensional evaluation report of the third entity feature.
[0041] Specifically, the multi-dimensional association parsing module automatically receives the third entity feature identifier from the screening and sorting module through a predefined internal data interface, and automatically extracts the corresponding gene expression profile and clinical information dataset from the structured data pool generated by the data acquisition module based on the identifier. The multi-dimensional association parsing module is automatically triggered by the process coordination and optimization module after the screening and sorting module completes. After the multi-dimensional evaluation report is generated, it is automatically submitted to the knowledge base of the process coordination and optimization module for subsequent verification feedback and model optimization loops. The multi-dimensional evaluation report is a structured data object, and its automatic generation process includes: parsing and formatting the raw results output by each analysis process, integrating them with target identifiers and analysis parameter metadata according to a predefined template, and finally serializing them into a unified JSON document.
[0042] In this embodiment, integration based on a predefined template specifically refers to: maintaining a report template configuration file, which defines the root structure of the JSON document, the key names corresponding to each analysis result, and the storage path reference format of the chart file; when generating the report, the output of each analysis component is automatically filled into the corresponding position of the template.
[0043] In one possible implementation, the association analysis with the category label is configured to perform an automated prognostic value assessment pipeline, including: acquiring expression level data corresponding to the third entity feature and clinical label data including survival time and status; automatically calculating the optimal cutoff value for dividing into high and low expression groups, and grouping the samples accordingly; automatically invoking survival analysis models and tests to generate survival curves and significance p-values; automatically performing univariate and multivariate proportional hazards regression analyses to calculate the hazard ratio and confidence interval of the third entity feature as a risk factor; and automatically integrating the generated statistical indicators and charts as part of a multidimensional assessment report.
[0044] It should be noted that the prognostic value assessment pipeline is executed by calling the internally integrated survival analysis service component and risk regression analysis service component. The internally integrated survival analysis service component is constructed by encapsulating the Kaplan-Meier algorithm and the Log-Rank test algorithm into an executable unit with standardized input and output interfaces. The risk regression analysis service component is constructed by encapsulating the Cox proportional hazards model into a statistical service that receives expression vectors and survival data and returns hazard ratios and confidence intervals. These components are called through the APIs exposed by the process coordination and optimization module. The process coordination and optimization module is responsible for automatically scheduling the execution of these components and data transmission according to a predefined order.
[0045] In this embodiment, the logic of the process coordination and optimization module to automatically schedule tasks according to a predefined order is based on a definable workflow description file, which defines the execution order, dependencies, and service component identifiers corresponding to each task, such as prognostic analysis, classification performance analysis, and enrichment analysis.
[0046] In one possible implementation, the classification performance analysis of the data samples is configured to: explicitly divide the samples into positive and negative sets according to the category labels; automatically perform receiver operating characteristic (ROC) curve analysis to calculate the classification performance of the third entity feature on the samples at different expression thresholds and generate ROC curves; and automatically calculate the area under the curve, the optimal classification threshold, and its corresponding sensitivity and specificity as part of a multidimensional evaluation report.
[0047] It should be noted that the ROC curve analysis is achieved by calling an embedded, configurable classification performance evaluation library. This library receives feature representation vectors and binary classification labels, and automatically completes threshold traversal, performance calculation, and chart rendering.
[0048] In one possible implementation, the enrichment analysis based on a preset function set is configured as follows: automatically generating a gene ranking list for enrichment analysis based on the association between the third entity feature and related phenotypes; connecting to a local functional annotation knowledge base that can be updated periodically; invoking hypergeometric tests and gene set enrichment analysis algorithms in parallel to screen biological pathways that are significantly related to the gene list from the local functional annotation knowledge base; and automatically generating textual hypotheses describing the potential biological mechanisms of the third entity feature based on the significance ranking of the enrichment results and the gene list, as part of a multidimensional evaluation report.
[0049] It should be noted that the hypergeometric test and GSEA algorithm invoked in parallel are deployed in the system backend as independent computing services; the local functional annotation knowledge base is stored in the form of a relational database or graph database and exchanges data with these computing services through a dedicated query interface.
[0050] In this embodiment, the multi-dimensional association parsing module internally includes a task parsing and execution engine, which executes the following after the module starts: reads configuration instructions from the process coordination and optimization module, or parses out the list of analysis tasks that need to be performed on the third entity feature according to predefined rules; automatically extracts and converts data from the acquired gene expression profile and clinical information dataset according to the analysis task list, into a data format that meets the input requirements of each analysis service component, including but not limited to preparing survival time, status, and expression vectors for prognostic analysis; preparing binary labels and expression vectors for classification performance analysis; sequentially or in parallel calling the service components corresponding to each analysis task in the task list according to the workflow description file or built-in order, including but not limited to survival analysis service components, classification performance evaluation libraries, enrichment analysis calculation services, etc., and passing in the adapted data as input parameters; and listening to and collecting the result data returned by each called service component and delivering it to the report generation stage.
[0051] The process coordination and optimization module allows users to perform agile configuration based on templates to quickly respond to analysis needs, schedules the execution order and data flow between the data acquisition module, the difference queue management module, the filtering and sorting module, and the multi-dimensional correlation analysis module, and optimizes the parameters of the weighted decision model in the filtering and sorting module based on the comparison between the external verification report and the multi-dimensional evaluation report.
[0052] Specifically, the process coordination and optimization module includes at least a workflow engine, a data dependency manager, a process definition language processor, a continuous learning optimizer, and a knowledge base. The workflow engine receives analysis processes defined in the form of a directed acyclic graph (DAG), where nodes in the DAG correspond to data acquisition, difference analysis, network filtering, or multidimensional parsing tasks, and edges correspond to data dependencies between tasks. The workflow engine parses the DAG, determines the execution order of tasks based on dependencies, and schedules the execution of each task. The data dependency manager registers the data entities and corresponding metadata output by each module, where the metadata includes at least a data identifier, data type, generation task identifier, generation timestamp, and version number. When downstream tasks are executed, the data dependency manager automatically matches and provides corresponding versions of data entities from registered data entities based on the input data pattern declared by the downstream tasks; the process definition language processor parses the complete analysis process based on the declarative language definition, which includes data source identifiers, task sequences, task parameters, and inter-task dependencies; the process definition language processor converts the process definition into task scheduling instructions executable by the workflow engine; the continuous learning optimizer performs the following steps to optimize the weighted decision model in the screening and sorting module: receiving an external validation report, which contains experimental validation conclusions or unique findings on candidate targets previously output by the system. Establish clinical validation data; the experimental validation conclusions include, but are not limited to, gene expression levels validated by qRT-PCR and biological behavioral impacts validated by cell function experiments; based on the experimental validation conclusions in the external validation reports, assign validation priority scores to the candidate targets, thereby constructing a validation priority ranking; wherein, the validation priority scores are set according to a positive correlation between the clarity of the validation conclusions and the strength of experimental evidence; align the external validation reports with the multidimensional evaluation reports of the corresponding candidate targets generated by the system, and based on the aligned data, quantitatively evaluate the accuracy of the previous ranking results of the weighted decision model, generating a model performance deviation signal; wherein, the model performance deviation signal is obtained through... The correlation loss function between model ranking and validation priority is calculated. Using the model performance deviation signal as the training target, the set of weight parameters used to fuse multiple topology analysis scores in the weighted decision model is dynamically adjusted using the gradient descent algorithm. The set of weight parameters is an n-dimensional vector, the dimension of which is the number of topology analysis methods, and each weight corresponds to the fusion weight of one topology analysis method. The optimized set of weight parameters is updated to the weighted decision model. The knowledge base stores the optimized weight parameters, corresponding validation cases, and optimization process metadata. When the system executes a new analysis task, the knowledge base provides initial weight parameters for the weighted decision model and provides historical references for process optimization.
[0053] It should be noted that the model performance deviation signal is obtained by comparing the correlation loss function between model ranking and validation priority; in this embodiment, a negative Spearman rank correlation coefficient is used as the loss value.
[0054] As attached Figure 2 The method shown is an artificial intelligence-based prognostic assessment method for gastric cancer, which specifically includes: S1: Preprocess the acquired first state data and corresponding category labels to generate the first entity feature; S2: Divide the first entity feature into samples to generate a comparison queue, and process the comparison queue using a difference analysis algorithm to generate the second entity feature; S3: Based on the second entity feature, a high-priority third entity feature is output through a preset weighted decision model and combined with multiple topological analyses; S4: Perform at least one of the following analyses on the third entity feature: correlation analysis with the category label, classification performance analysis of the data sample, and enrichment analysis based on a preset function set, to generate a multidimensional evaluation report of the third entity feature; S5: Schedule the execution order and data flow between the data acquisition module, the difference queue management module, the filtering and sorting module, and the multi-dimensional correlation analysis module, and optimize the parameters of the weighted decision model in the filtering and sorting module based on the comparison between the external verification report and the multi-dimensional evaluation report.
[0055] Secondly: The accompanying drawings of the embodiments disclosed in this invention only involve the structures involved in the embodiments disclosed in this invention. Other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of this invention can be combined with each other. In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An artificial intelligence-based prognostic assessment system for gastric cancer, characterized in that, include: Data acquisition module: used to preprocess the acquired first state data and corresponding category labels to generate the first entity feature; The difference queue management module is used to divide the first entity feature into samples to generate a comparison queue, and to process the comparison queue through a difference analysis algorithm to generate the second entity feature. Filtering and sorting module: Based on the second entity features, it outputs high-priority third entity features through a preset weighted decision model and combined with various topological analyses; Multi-dimensional association analysis module: used to perform at least one of the following analyses on the third entity feature: association analysis with the category label, classification performance analysis of data samples, and enrichment analysis based on a preset function set, to generate a multi-dimensional evaluation report of the third entity feature; The process coordination and optimization module supports users in making agile configurations based on templates to quickly respond to analysis needs. It schedules the execution order and data flow between the data acquisition module, the difference queue management module, the filtering and sorting module, and the multi-dimensional correlation analysis module. Based on the comparison between the external verification report and the multi-dimensional evaluation report, it optimizes the parameters of the weighted decision model in the filtering and sorting module.
2. The gastric cancer prognostic assessment system based on artificial intelligence according to claim 1, characterized in that: The data acquisition module, in acquiring the first state data, specifically includes: In response to the user-specified data source identifier and acquisition task, the preset data source adapter corresponding to the target biomedical database is invoked; Through the data source adapter, at least two heterogeneous raw data sets are simultaneously obtained from the database in an automated manner. The raw data sets include at least: omics data containing quantitative gene expression values, and clinical information containing clinical attributes and status indicators of samples. The acquired omics data and the clinical information are logically linked through a common sample identifier to form the first state data with a clearly defined structure.
3. The gastric cancer prognostic assessment system based on artificial intelligence according to claim 1, characterized in that: The difference queue management module obtains the second entity feature, specifically including: Configurable grouping rules defined based on clinical information in the first entity features; The grouping rules are executed to filter out samples that meet the rule conditions from the first entity features, forming at least two sample sets, i.e., comparison queues; The grouping rules and the generated sample set identifier are used as metadata; At least two differential analysis algorithms are invoked in parallel to perform differential gene expression analysis on the comparison queue. Generate the second entity feature, which includes a saliency metric.
4. The gastric cancer prognostic assessment system based on artificial intelligence according to claim 3, characterized in that: The difference queue management module, in the second entity feature, includes at least the comprehensive corrected p-value and the comprehensive effect value as significant metrics, specifically including: The original p values Perform fusion and calculate the comprehensive chi-square value. The specific formula is as follows: in, denoted as the original p-value calculated for the gene by the i-th differential analysis algorithm, and k represents the total number of differential analysis algorithms; The inverse variance weighted average method was used to calculate the effect values. To perform fusion and calculate the combined effect value Specifically, it is expressed as: in, It is represented as a weight, and its value can be the reciprocal of the variance of each effect value; Multiple tests were performed on the original p-values to obtain the overall corrected p-values.
5. The gastric cancer prognostic assessment system based on artificial intelligence according to claim 1, characterized in that: The filtering and sorting module has a preset weighted decision model configured to support at least two weight configuration modes, including: In static configuration mode, static weight vectors are manually set for the various topology analysis algorithms involved in the fusion based on prior knowledge. In dynamic optimization mode, the entropy weight method is automatically invoked to calculate and allocate weights based on the dispersion of the output scores of each topology analysis algorithm. The learning optimization model adjusts the weight configuration through regression analysis to maximize the correlation between the fused ranking results and the effective target priorities verified by external experiments.
6. The gastric cancer prognostic assessment system based on artificial intelligence according to claim 1, characterized in that: The filtering and sorting module performs various topology analyses, including at least two selected from the following set of algorithms: Maximum clique centrality, maximum neighborhood component, degree centrality, bottleneck, eccentricity, proximity centrality, radiality, betweenness centrality, stress, and clustering coefficient; Furthermore, the gene interaction network automatically constructed based on the second entity features executes the selected algorithm in parallel and outputs a topology score matrix with a uniform format.
7. The gastric cancer prognostic assessment system based on artificial intelligence according to claim 1, characterized in that: The multi-dimensional association analysis module, in its association analysis with the category labels, is specifically configured to execute an automated prognostic value assessment pipeline, including: Obtain expression level data corresponding to the third entity feature and clinical label data including survival time and status; Automatically calculate the optimal cutoff value for dividing the high and low expression groups; Automatically invokes the survival analysis model to generate survival curves and significance P-values; Automatically perform univariate and multivariate proportional hazards regression analysis to calculate the hazard ratio and confidence interval of the third entity feature as a risk factor; The statistical indicators and charts generated automatically.
8. The gastric cancer prognostic assessment system based on artificial intelligence according to claim 1, characterized in that: The multi-dimensional association parsing module, for classifying data samples, is specifically configured as follows: Based on the category labels, the samples are clearly divided into positive and negative sets; Calculate the classification performance of the third entity feature on samples under different expression thresholds, and generate ROC curves; Automatically calculate the area under the curve, the optimal classification threshold, and the corresponding sensitivity and specificity.
9. The gastric cancer prognostic assessment system based on artificial intelligence according to claim 1, characterized in that: The multi-dimensional association parsing module, based on enrichment analysis of a preset function set, is specifically configured as follows: Automatically generate a gene sequence list for enrichment analysis; Connect to a local feature annotation knowledge base that can be regularly synchronized and updated; Select biological pathways that are significantly related to the gene list from the local functional annotation knowledge base; Automatically generate textual hypotheses describing the underlying biological mechanisms of the features of the third entity.
10. A method for prognostic assessment of gastric cancer based on artificial intelligence, comprising a system for prognostic assessment of gastric cancer based on artificial intelligence according to any one of claims 1-9, characterized in that, include: S1: Preprocess the acquired first state data and corresponding category labels to generate the first entity feature; S2: Divide the first entity feature into samples to generate a comparison queue, and process the comparison queue using a difference analysis algorithm to generate the second entity feature; S3: Based on the second entity feature, a high-priority third entity feature is output through a preset weighted decision model and combined with multiple topological analyses; S4: Perform at least one of the following analyses on the third entity feature: correlation analysis with the category label, classification performance analysis of the data sample, and enrichment analysis based on a preset function set, to generate a multidimensional evaluation report of the third entity feature; S5: Schedule the execution order and data flow between the data acquisition module, the difference queue management module, the filtering and sorting module, and the multi-dimensional correlation analysis module, and optimize the parameters of the weighted decision model in the filtering and sorting module based on the comparison between the external verification report and the multi-dimensional evaluation report.