Medical oncology specimen collection management system

By constructing blockchain and multi-omics databases, and combining adaptive quantile regression forest algorithm and graph structure spatiotemporal feature extraction, the quality and data integration problems in specimen collection in oncology were solved, realizing digital management of the entire specimen process and support for precision diagnosis and treatment.

CN121034575APending Publication Date: 2025-11-28NANTONG TUMOR HOSPITAL

Patent Information

Application Number
CN202511544975.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Current techniques for collecting oncology specimens suffer from problems such as insufficient sensitivity, heterogeneity, lack of standardization, uneven distribution of resources, difficulty in data integration, and ethical controversies, which limit the accuracy and efficiency of diagnosis and treatment.

Method used

Blockchain technology is used to store information about the entire process of specimens. Combined with adaptive quantile regression forest algorithm and graph structure spatiotemporal feature extraction, a multi-omics database and knowledge graph are constructed to realize dynamic early warning and anomaly scoring of specimen data. Data mining and sharing are carried out through graph neural network to ensure specimen quality and data privacy.

Benefits of technology

This has enabled a complete upgrade of the specimen collection and application chain, ensuring the integrity, traceability, and transparency of specimen information, improving the data-driven nature and accuracy of diagnosis and treatment, and promoting cross-platform data sharing and scientific research development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034575A_ABST
    Figure CN121034575A_ABST
Patent Text Reader

Abstract

The invention discloses a medical oncology specimen collection and management system, and relates to the technical field of specimen management, and the system comprises a collection and detection unit which is used for collecting biopsy specimens and corresponding specimen data; the tracking quality control unit is used for storing full-process information of the biopsy specimens by utilizing a block chain and dynamically evaluating and tracing quality information of the biopsy specimens; the integrated learning unit is used for constructing a multi-omics database, generating a tumor multi-omics knowledge graph and carrying out dynamic knowledge updating through data mining; and the decision management unit is used for sharing resources and monitoring authorized use records of the biopsy specimens. By integrating the block chain, federal learning and process reconstruction, full-chain upgrading of medical oncology specimens from'collection 'to'application' is realized, tumor diagnosis and treatment are promoted to step forward from'experience-driven 'to'data-driven', digitization, standardization and traceable management of biopsy specimen full-process information are realized, and the efficiency of medical oncology diagnosis and treatment is improved. And data support and decision support are provided for clinical practice and scientific research.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of specimen management, in particular to a tumor internal medicine specimen collection and management system. BACKGROUND

[0002] Tumor internal medicine is a discipline in clinical medicine that is responsible for the diagnosis and treatment of malignant tumors. Its core is to control cancer progression through non-surgical systemic treatment methods such as chemotherapy, targeted therapy, and immunotherapy. Tumor internal medicine doctors need to use imaging, pathology, and molecular detection techniques to determine the type and stage of tumors (such as TNM staging), and develop treatment plans based on individual characteristics of patients (such as gene mutations and PD-L1 expression levels), such as using osimertinib targeted therapy for lung cancer patients with EGFR mutation positive. In addition, tumor internal medicine also needs to closely cooperate with other departments, such as tumor surgery responsible for surgical resection, radiation oncology for radiotherapy, and through multidisciplinary consultation (MDT) to optimize diagnosis and treatment strategies, and manage treatment side effects (such as bone marrow suppression and cancer pain) to improve the quality of life of patients.

[0003] Tumor internal medicine specimens refer to biological samples obtained during diagnosis and treatment for analysis of tumor biological characteristics, covering tissues, fluids, bone marrow, and other types. Tissue specimens include biopsy or surgical resection of tumor tissue (such as lung biopsy to determine lung cancer type), fluid specimens involve circulating tumor cells (CTC) in blood, circulating tumor DNA (ctDNA), and cancer cells in pleural effusion and ascites, and bone marrow specimens are used for diagnosis of leukemia and other hematological tumors. These specimens provide key evidence for tumor typing (such as distinguishing HER2 status in breast cancer), treatment plan selection (such as using immunotherapy for patients with high PD-L1 expression), and drug resistance monitoring (such as detecting T790M mutation in patients with EGFR mutation resistance) through pathological examination (such as immunohistochemistry), molecular detection (such as gene sequencing), or liquid biopsy technology, and are the cornerstone of precision medicine.

[0004] Tumor internal medicine specimens must be collected in a standardized manner, as they directly determine the accuracy of diagnosis and effectiveness of treatment. The molecular characteristics of specimens (such as ALK fusion and microsatellite instability) are the direct basis for developing individualized plans, and dynamic monitoring of ctDNA can also detect early relapse or drug resistance mutations. In addition, specimen banks provide resources for scientific research and promote the development of new drugs. If specimens are not collected adequately or do not meet quality standards, it may lead to misdiagnosis, blind treatment, and even delay in diagnosis and treatment of patients.

[0005] Current technologies encompass molecular detection, liquid biopsy, and information technology tools, but they suffer from numerous technical shortcomings. These include insufficient sensitivity of liquid biopsy, heterogeneity of tissue specimens (single-point biopsy may miss drug-resistant clones), lack of standardization (differences in fixed times and storage conditions among different institutions lead to poor data comparability), uneven resource distribution (primary hospitals lack NGS testing capabilities, specimens are prone to degradation during long-distance transportation), difficulties in data integration (multi-omics data are scattered, lacking a unified analysis platform), and ethical controversies (the scope of informed consent for the secondary use of research samples is ambiguous). Summary of the Invention

[0006] Therefore, it is necessary to provide a specimen collection and management system for oncology to address the aforementioned technical problems.

[0007] This invention provides a specimen collection and management system for medical oncology, comprising: The collection and detection unit is used to collect biopsy specimens and corresponding specimen data; The traceability quality control unit is used to store biopsy specimen information throughout the entire process using blockchain and to dynamically evaluate and trace the quality information of biopsy specimens. The integrated learning unit is used to build a multi-omics database, generate a tumor multi-omics knowledge graph, and dynamically update knowledge through data mining; The decision-making management unit is used for resource sharing and monitoring the authorized use records of biopsy specimens; The quality control unit includes: a monitoring and early warning module and an evaluation and decision-making module; The monitoring and early warning module is used to dynamically optimize the early warning threshold in real time and identify abnormal information in the specimen data. This includes: predicting different quantiles based on the adaptive quantile regression forest algorithm and adjusting the early warning threshold of each specimen data in real time; constructing a graph structure using specimen data and specimen event nodes, and calculating the weighted sum of reconstruction error and prediction error by extracting the spatiotemporal features of the graph structure to obtain an anomaly score. The evaluation and decision-making module is used to perform image and biochemical index analysis on biopsy specimens, evaluate the specimen quality, and generate specimen processing suggestions.

[0008] Furthermore, the data collection and detection unit, the tracking and quality control unit, the integrated learning module, and the decision management unit are sequentially connected; The quality control unit also includes: an anomaly tracing module and a record reporting module; Among them, the anomaly traceability module is used to assign a unique digital identity to biopsy specimens, and record the entire process of data from collection, transportation, storage to testing on the blockchain, recording the entire chain of biopsy specimen operation logs; The report module is used to generate visualized quality control reports with multi-dimensional query and statistical analysis. Furthermore, the monitoring and early warning module, the assessment and decision-making module, the anomaly tracing module, and the recording and reporting module are connected sequentially.

[0009] Furthermore, a graph structure is constructed using the specimen data and specimen event nodes. By extracting the spatiotemporal features of the graph structure, the weighted sum of the reconstruction error and prediction error is calculated to obtain the anomaly score, which includes: Based on the abnormality score of the specimen data, the corresponding abnormality level is matched according to the preset grading rules, and different abnormality warning mechanisms are triggered to generate abnormal information of the biopsy specimen.

[0010] Furthermore, based on the adaptive quantile regression forest algorithm to predict different quantiles, the early warning thresholds for various sample data are adjusted in real time, including: The sample data is split into time series data and discrete time data. The time series data is segmented by a sliding window and periodic encoding is added. The time difference features are calculated for the discrete time data. Based on the historical distribution of the sample data, the target value is predicted at different quantiles by modeling the quantile regression forest algorithm, and the prediction results of all trees are averaged to obtain the quantile prediction value. By using quantile regression forest to predict different quantiles of the target variable, a dynamic early warning threshold is set. When new sample data arrives, the early warning threshold is automatically adjusted according to the new data distribution.

[0011] Furthermore, the formula for calculating the anomaly score is as follows: ; In the formula, Score t Indicates the current t Anomaly scoring at time steps; λ Indicates the weighting coefficient; x t Indicates the current t The numerical values ​​of the specimen data at each time step; Indicates the current t Predictions from the spatiotemporal graph attention network at each time step; Q α ( x t ) indicates the current t The warning threshold for the time step.

[0012] Furthermore, image and biochemical index analysis is performed on the biopsy specimens to assess their quality and generate specimen processing recommendations, including: The specimen data were divided into biopsy image data and physiological index data. The biopsy image data was subjected to image standardization processing, and the physiological index data was subjected to structure processing. Segment tumor cells in biopsy image data, obtain image features by outputting heatmaps, and calculate the proportion of tumor cells to total pixels and the necrotic area; A graph structure of physiological indicator data is constructed, and physiological features are extracted through cross-attention fusion. The degradation probability of quality risk prediction is then calculated. Image features and physiological features are mapped to a common latent space. The distribution parameters are output through a Bayesian neural network to calculate the quality confidence score. By combining the proportion of tumor cells, the output degradation probability, and the confidence score results, the quality of the current biopsy specimen is evaluated, and corresponding specimen processing suggestions are matched.

[0013] Furthermore, the integrated learning unit includes: a multi-source integration module, a joint analysis module, a knowledge graph module, and an update and mining module; Among them, the multi-source integration module is used to integrate and unify specimen data from different platforms; The joint analysis module is used to model multi-omics associations using graph neural networks and introduce federated learning shared models to generate multi-omics databases; The atlas construction module is used to construct a tumor multi-omics knowledge graph and identify potential associations between biopsy specimens and pathological information; Update the mining module to capture the latest literature for information mining and link key information to the knowledge graph; Furthermore, the multi-source integration module, joint analysis module, knowledge graph module, and update mining module are sequentially connected.

[0014] Furthermore, graph neural networks are used to model multi-omics associations, and federated learning sharing models are introduced to generate multi-omics databases, including: Standardize specimen data from different platforms, unify the names and identifiers of each type, and use cross-omics collaborative matrix decomposition to predict missing values ​​in the specimen data; Define nodes and edges in a heterogeneous graph structure, dynamically adjust edge weights, train a graph convolutional network locally on each platform, and extract low-dimensional embedding vectors of each omics node through a multi-layer graph neural network. The graph convolutional network parameters are shared across platforms. The central server uses a federated averaging algorithm to aggregate local models and then distributes the aggregated global model to each platform for the next round of training. Using the trained global model, low-dimensional embedding representations are generated for each node, and the low-dimensional embeddings and the predicted associations between nodes are stored in a multi-omics database.

[0015] Furthermore, constructing a multi-omics knowledge graph of tumors to identify potential associations between biopsy specimens and pathological information includes: Collect multi-omics data, extract key features of each omics data, convert them into graph structure representations, extract entities and relationships, and construct a knowledge network. Different types of entities and relationships are integrated into a knowledge graph to construct a multi-omics knowledge graph, and the entities and relationships in the knowledge graph are stored in a graph database; Spatiotemporal correlation modeling of different omics data and pathological information is performed to identify potential multi-omics associations and to mine potential patterns in knowledge graphs to identify potential associations between biopsy specimens and pathological information.

[0016] Furthermore, the decision management unit includes: an informed authorization module, a shared interaction module, a usage record module, and an early warning response module; The informed authorization module is used to dynamically authorize the use of specimens through smart terminals, desensitize specimen data, and associate it with the authorization approval number. The shared interaction module is used to establish cross-regional resource sharing, conduct virtual multidisciplinary consultations and patient communication and interaction through terminal communication, and transcribe consultation and interaction content in real time. The recording module is used to monitor the authorization time, scope, and usage records of biopsy specimens in real time using blockchain. The early warning and response module is used to issue real-time warnings and freeze unauthorized biopsy specimens. Furthermore, the informed authorization module, the shared interaction module, the usage record module, and the early warning response module are kept connected in sequence.

[0017] The beneficial effects of this invention are as follows: 1. By integrating blockchain, federated learning, and process refactoring (standardization, automation, and collaboration), the entire chain of oncology specimens from "collection" to "application" has been upgraded, driving oncology diagnosis and treatment from "experience-driven" to "data-driven." This has enabled the digitalization, standardization, and traceability management of biopsy specimen information throughout the entire process. Utilizing adaptive quantile regression forest algorithms and graph structure spatiotemporal feature extraction technology, dynamic early warning and anomaly scoring of specimen data are achieved, effectively detecting and responding to specimen quality issues. Simultaneously, image and biochemical indicator analysis generates processing suggestions to ensure specimen quality meets clinical standards. Furthermore, the construction and dynamic updating of multi-omics databases and knowledge graphs not only enhance the understanding of tumor biological mechanisms but also promote cross-platform data sharing and precise diagnostic and treatment decisions, providing strong data support and decision support for clinical practice and scientific research.

[0018] 2. This invention ensures that the entire lifecycle data of biopsy specimens (from collection, transportation, storage to testing) can be accurately recorded and traced, ensuring the integrity, traceability, and transparency of specimen information; by assigning a unique digital identity to each specimen and using blockchain technology to put the data on the chain, the immutability and efficient management of the data throughout the process can be achieved; by using graph neural networks (GNN) and federated learning technologies to model, mine associations, and generate multi-omics databases for multi-omics data, it is possible to efficiently extract potential associations from data from different platforms and different omics types while maintaining data privacy. Attached Figure Description

[0019] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a system block diagram of an oncology specimen collection and management system according to an embodiment of the present invention.

[0020] The reference numerals are: 1. Collection and detection unit; 2. Tracking and quality control unit; 3. Integration and learning unit; 4. Decision management unit. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0022] Please see Figure 1 A specimen collection and management system for oncology includes: a collection and testing unit 1, a tracking and quality control unit 2, an integrated learning unit 3, and a decision management unit 4; and the collection and testing unit 1, the tracking and quality control unit 2, the integrated learning unit 3, and the decision management unit 4 are sequentially connected.

[0023] Collection and detection unit 1 is used to collect biopsy specimens and corresponding specimen data.

[0024] In this invention, AI-guided precise sampling can be utilized. AI algorithms based on imaging data (CT / MRI) plan the puncture path in real time, avoiding necrotic areas (such as in lung nodule biopsies) to ensure the acquisition of samples with high tumor cell content. During endoscopy or surgery, augmented reality (AR) navigation is used to mark tumor boundaries, guiding precise resection (such as in breast-conserving surgery for breast cancer) to obtain a biopsy specimen. Simultaneously, the intelligent sampling device automatically records the collection time, location, and temperature, and synchronizes this data to a blockchain platform (tamper-proof). The biopsy specimen utilizes microfluidic chips and nanomaterial enrichment technology (such as increasing CTC capture efficiency to over 90%), addressing the issue of low sensitivity in liquid biopsies.

[0025] Specifically, biopsy specimens include tissue biopsy specimens and liquid biopsy specimens. Tissue biopsy specimens include tumor tissue sections, needle biopsies, etc., and are usually collected through endoscopy, needle aspiration, or surgery. These specimens are used for pathological examination, gene analysis, molecular testing, etc. Liquid biopsy specimens are bodily fluid specimens such as blood, urine, and saliva, mainly used to detect cell-free DNA (cfDNA) or circulating tumor cells (CTCs) in plasma for early detection and dynamic monitoring of tumor progression.

[0026] The specimen data includes the following aspects: 1. Basic Information: Patient information (name, age, gender, medical history, etc.); Specimen information (specimen type, collection location, collection method, collection time, etc.); Clinical diagnostic information (diagnosis, tumor type, stage, subtype, etc.).

[0027] 2. Testing Information: Detection methods (such as pathological examination, genomic analysis, molecular marker detection, etc.); Test results (such as gene mutations, protein expression, pathology reports, etc.).

[0028] 3. Specimen quality information: Specimen processing and storage information (temperature, humidity, storage time, etc.); Quality checks on the collected samples (e.g., whether they meet the collection standards, whether there is any contamination, etc.).

[0029] In addition, integrated electronic recording systems (such as electronic health record systems and LIMS systems) are introduced for the input and storage of specimen information. All specimen information can be automatically entered into the system by scanning QR codes or RFID tags, ensuring data accuracy and timeliness. Quality control checkpoints are set up at each stage of specimen collection, processing, and transportation. For example, monitoring whether the temperature and humidity of specimen storage meet standards, and whether temperature control during specimen transportation meets standards. All specimen quality information is recorded in a database. Each specimen has a unique digital identity (via QR code or RFID tag), and all information on collection, processing, transportation, and storage is recorded in real time, ensuring full traceability of the specimen. When quality problems occur, they can be traced back to the specific collection, storage, or transportation stage.

[0030] The tracking quality control unit 2 is used to store the entire process information of biopsy specimens using blockchain and to dynamically evaluate and trace the quality information of biopsy specimens.

[0031] In the description of this invention, the tracking quality control unit 2 includes: a monitoring and early warning module (not shown in the figure), an evaluation and decision module (not shown in the figure), an anomaly tracing module (not shown in the figure), and a recording and reporting module (not shown in the figure); and the monitoring and early warning module, the evaluation and decision module, the anomaly tracing module, and the recording and reporting module are connected in sequence.

[0032] The monitoring and early warning module is used to dynamically optimize the early warning threshold in real time and identify abnormal information in the specimen data.

[0033] In the description of this invention, the real-time dynamic optimization of the early warning threshold to identify abnormal information in the specimen data includes: S201. Based on the adaptive quantile regression forest algorithm, different quantiles are predicted, and the warning thresholds of various specimen data are adjusted in real time.

[0034] In the description of this invention, the method of predicting different quantiles based on the adaptive quantile regression forest algorithm and adjusting the warning thresholds of various sample data in real time includes: S2011. The sample data is split into time series data and discrete time data. The time series data is segmented by a sliding window and periodic coding is added. The time difference feature is calculated for the discrete time data.

[0035] Specifically, collect specimen data that includes both time-series and discrete event data. For example, this includes specimen temperature, humidity, storage time, and events during processing. Sliding window segmentation is used on the time-series data to capture short-term fluctuations in historical data, and time periodicity encoding (e.g., daily, weekly, seasonal factors) is added to extract time features. Time difference features (e.g., event intervals, processing delays) are calculated for the discrete event data.

[0036] S2012. Based on the historical distribution of the sample data, the target value is predicted at different quantiles by modeling the quantile regression forest algorithm and averaging the prediction results of all trees to obtain the quantile prediction value.

[0037] Quantile Regression Forest (QRF) is an extension of Random Forest that not only predicts the expected value of the target variable (mean regression), but also predicts the value of the target variable at a specific quantile (such as 0.5 or 0.95). Its main objective is to predict the response value of a sample based on a given quantile (e.g., median, upper quartile).

[0038] Random Forest (RF) consists of multiple decision trees, each trained by sampling a subset of the training data using a bootstrap sampling method. For each tree, the calculation is performed for each data point. Corresponding response y Quantile prediction. For a given quantile τ (e.g., ... =0.9 is used for upper limit warning), QRF calculates the quantile estimate of the tree node.

[0039] S2013. By using quantile regression forest to predict different quantiles of the target variable, a dynamic early warning threshold is set. When new sample data arrives, the early warning threshold is automatically adjusted according to the new data distribution.

[0040] Specifically, QRF models are used to predict different quantiles of the target variable (such as the upper quartile, lower quartile, etc.) to set dynamic warning thresholds. For example, for temperature data, if the goal is to monitor abnormal sample temperatures, an upper threshold based on QRF prediction can be set. ; In the formula, Indicates the upper limit threshold; Indicates the predicted quantile; This indicates the actual set coefficient (e.g., 1.5 times the standard deviation). It represents the standard deviation.

[0041] Whenever new sample data arrives, the QRF model automatically adjusts the threshold based on the new data distribution. By continuously learning from data changes, the system can update the warning threshold in real time, avoiding false alarms or missed alarms caused by fixed thresholds in traditional methods.

[0042] S202. Construct a graph structure using specimen data and specimen event nodes. Extract the spatiotemporal features of the graph structure and calculate the weighted sum of reconstruction error and prediction error to obtain anomaly score.

[0043] In the description of this invention, the formula for calculating the anomaly score is: ; In the formula, Score t Indicates the current t Anomaly scoring at time steps; λ Indicates the weighting coefficient; x t Indicates the current t The numerical values ​​of the specimen data at each time step; Indicates the current t Predictions from the spatiotemporal graph attention network at each time step; Q α ( x t ) indicates the current t The warning threshold for the time step.

[0044] S203. Based on the abnormal score value of the specimen data, match the corresponding abnormal level according to the preset grading rules, trigger different abnormal warning mechanisms, and generate abnormal information of the biopsy specimen.

[0045] Specifically, different anomaly levels can be defined based on anomaly scores, and corresponding early warning mechanisms can be triggered. For example, the sample data can be divided into different anomaly levels based on the magnitude of the anomaly score. Low risk: The anomaly score is less than the standard threshold, and the specimen quality is normal.

[0046] Medium risk: Abnormal scores are within the standard threshold, which may indicate minor issues and require attention.

[0047] High risk: Abnormal scores exceeding the standard threshold will trigger an immediate warning, which may affect experimental results or quality.

[0048] If the specimen's anomaly score reaches the medium or high risk level, the system will trigger an alert and notify the operator to conduct further inspections or take measures.

[0049] The evaluation and decision-making module is used to perform image and biochemical index analysis on biopsy specimens, evaluate the specimen quality, and generate specimen processing suggestions.

[0050] In the description of this invention, image and biochemical index analysis of biopsy specimens is performed to assess the specimen quality and generate specimen processing suggestions, including: S211. Divide the specimen data into biopsy image data and physiological indicator data. Perform image standardization processing on the biopsy image data and structure processing on the physiological indicator data.

[0051] Specifically, the specimen data are divided into two main categories: Biopsy image data: includes image data such as tumor tissue slices and microscopic images.

[0052] Physiological index data: including physiological parameters (such as blood indicators, gene expression data, pathological scores, etc.).

[0053] Image data standardization is a process used to eliminate the influence of factors such as image quality and lighting. Common standardization methods include: 1. Gray-level normalization: scaling the pixel values ​​of an image to make them fall within a uniform range. 2. Denoising: using filters (such as Gaussian filters, mean filters, etc.) to remove noise from the image and improve image quality.

[0054] Physiological indicator data should be structured to ensure that the data can be used for further analysis. For example, raw unstructured text information (such as medical records and diagnostic reports) can be transformed into numerical or categorical features, and then standardized and normalized.

[0055] S212. Segment the tumor cells in the biopsy image data, obtain image features by outputting a heatmap, and calculate the proportion of tumor cells in the total pixels and the necrotic area.

[0056] Specifically, image segmentation algorithms (such as U-Net, Mask R-CNN, etc.) are used to segment the tumor region in the biopsy image and extract the region of tumor cells. This step will help clarify the distribution and proportion of tumor cells in the image.

[0057] Heatmap generation techniques (such as Grad-CAM and LIME) are used to highlight the features of tumor cells in images, aiding in further analysis of key regions. The proportion of tumor cells in the overall image is then calculated, and the pixel ratio of tumor cells is used to assess the density and biological characteristics of the tumor in the specimen. Finally, image analysis techniques are used to identify and calculate necrotic areas in the tumor image; these areas are often associated with treatment response or prognosis.

[0058] S213. Construct a graph structure for physiological indicator data, extract physiological features through cross-attention fusion, and calculate the degradation probability for quality risk prediction.

[0059] Specifically, physiological indicator data (such as gene expression levels, cell activity indicators, etc.) are transformed into a graph structure. Each physiological indicator (or indicator category) is considered a node in the graph, and the relationships between nodes (such as correlation, causality, etc.) are represented by edges. Nodes represent the features of physiological data contained in each node (such as expression values, clinical data). Edges represent the edges of the graph constructed based on the correlation or biological relationships of the data.

[0060] To fuse image and physiological features, a cross-attention mechanism is employed. This method learns the interaction between the two modalities to extract useful information from the image and physiological data. Through cross-attention, the model can dynamically adjust weights between the two modalities, enhancing the correlation between image and physiological features. Based on the physiological features in the graph structure, the probability of specimen quality degradation is calculated, assessing the overall quality risk of the specimen. For example, mutations or expression levels of certain genes may be related to specimen degradation and processing quality.

[0061] S214. Map image features and physiological features to a common latent space, output distribution parameters through a Bayesian neural network, calculate quality confidence, combine the proportion of tumor cells, output degradation probability and confidence results to evaluate the quality of the current biopsy specimen, and match corresponding specimen processing suggestions.

[0062] Specifically, the features of image and physiological data are mapped into a common latent space. Deep learning models (such as joint embedding networks and Bayesian neural networks) are used to fuse features from different modalities into a shared space for comprehensive evaluation of specimen quality. Bayesian neural networks (BNNs) output distribution parameters for tumor cell percentage, degradation probability, and quality confidence. Bayesian networks can model uncertainty, outputting a probability distribution rather than a single predicted value. This helps quantify the uncertainty of the model and provides more information for subsequent decision-making.

[0063] Based on the output parameters of the Bayesian neural network, the quality confidence score of the specimen is calculated. For example, a comprehensive score is calculated based on factors such as the proportion of tumor cells, the proportion of necrotic areas, and the probability of degradation to obtain the final quality confidence score. Finally, the decision rule for specimen processing recommendations can be determined using the following formula: ; In the formula, R tumor Indicates the percentage of tumor cells; P degrade Indicates the probability of degradation. C quality This indicates the confidence level of the quality.

[0064] The anomaly traceability module is used to assign a unique digital identity to biopsy specimens, and record the entire process of data from collection, transportation, storage to testing on the blockchain, recording the entire operation log of the biopsy specimen.

[0065] Specifically, each collected specimen is assigned a unique digital identity (such as a QR code or RFID tag). This identity includes basic information about the specimen (such as patient ID, collection time, collection location, specimen type, etc.) to ensure that the specimen is not confused with other specimens throughout its lifecycle. The specimen's unique digital identity is then linked to all relevant data (such as collection information, transportation records, storage conditions, test results, etc.) to ensure that each data point can be traced back to a specific specimen.

[0066] Key operational data for each specimen (such as collection time, collection location, storage temperature, transportation records, and quality control inspection results) are recorded in real time on the blockchain. Blockchain technology ensures the immutability, traceability, and transparency of this data. Utilizing the decentralized and encrypted characteristics of blockchain, it ensures that specimen data cannot be maliciously tampered with at any stage, while also guaranteeing data transparency, with all operations publicly searchable.

[0067] The entire chain of operation logs is recorded in the following aspects: 1. Collection Stage: Record the time, location, and method of specimen collection (e.g., puncture biopsy, endoscopic biopsy), as well as relevant personnel information. Specimen collection information is automatically synchronized to the system and uploaded to the blockchain via real-time scanning of QR codes or RFID tags.

[0068] 2. Transportation Stage: Records the transportation process of specimens from the collection point to the laboratory or storage location, including transportation time, transportation method (such as cold chain transportation), and temperature and humidity conditions. During transportation, the system monitors and records environmental data such as temperature and humidity in real time to ensure the quality of the specimens during transportation.

[0069] 3. Storage Stage: Record the storage conditions of specimens in the laboratory or storage facility (such as storage temperature and humidity), and monitor the storage environment in real time using sensors. When abnormal storage conditions occur, the system will automatically record and send an alert.

[0070] 4. Testing Phase: Record the testing process of the specimen in the laboratory, including information on experimental equipment, test results, and operator information. All experimental results and test data are also recorded in real time and uploaded to the blockchain to ensure the reliability of the results.

[0071] The record and report module is used to generate visualized quality control reports with multi-dimensional query and statistical analysis.

[0072] Integrated learning unit 3 is used to build a multi-omics database, generate a tumor multi-omics knowledge graph, and dynamically update knowledge through data mining.

[0073] In the description of this invention, the integrated learning unit 3 includes: a multi-source integration module (not shown in the figure), a joint analysis module (not shown in the figure), a knowledge graph module (not shown in the figure), and an update mining module (not shown in the figure); and the multi-source integration module, the joint analysis module, the knowledge graph module, and the update mining module are connected in sequence.

[0074] The multi-source integration module is used to integrate and unify specimen data from different platforms.

[0075] The joint analysis module is used to model multi-omics associations using graph neural networks and introduce federated learning shared models to generate multi-omics databases.

[0076] In the description of this invention, the use of graph neural networks to model multi-omics associations and the introduction of federated learning sharing models to generate multi-omics databases include: S301. Standardize the specimen data from different platforms, unify the names and identifiers of each type, and use cross-omics collaborative matrix decomposition to predict missing values ​​in the specimen data.

[0077] Specifically, specimen data collected from different platforms (such as hospitals, research institutions, and laboratories) often differ in data format, naming standards, and units. To unify these data, standardization processing is necessary. Specific procedures include: 1. Unified naming and identification: Unify the identification of the same elements (such as genes, proteins, etc.) in different platforms and different omics data (such as gene expression, protein data, etc.) to ensure consistency across all platforms.

[0078] 2. Data conversion and unit unification: Convert units (such as concentration, time, etc.) used on different platforms to ensure data consistency and comparability.

[0079] For missing value prediction, cross-omics collaborative matrix factorization is used to predict missing values ​​in the sample data. Matrix factorization represents each omics data point as a product of low-rank matrices, and missing values ​​are filled using known data.

[0080] S302. Define the nodes and edges of the heterogeneous graph structure, dynamically adjust the edge weights, train the graph convolutional network locally on each platform, and extract the low-dimensional embedding vectors of each omics node through a multi-layer graph neural network.

[0081] Specifically, due to the heterogeneity of multi-omics data (such as different data types like genomics, transcriptomics, and proteomics), it is necessary to construct different nodes in a heterogeneous graph for each data type. Each data type corresponds to a node type, and nodes of different types are connected by edges, with the weights of the edges reflecting the correlation between nodes.

[0082] Nodes represent different types of specimen data, such as gene nodes, protein nodes, patient nodes, etc. The weights of the edges are dynamically adjusted based on the correlation or biological relationship of the data, reflecting the strength of the relationship between different data types.

[0083] Each platform's local data is trained using a Graph Convolutional Network (GCN). Through multiple layers of GCN, node features are passed around and aggregated, ultimately yielding a low-dimensional embedding representation of the nodes.

[0084] S303: All platforms share graph convolutional network parameters. The central server uses a federated averaging algorithm to aggregate local models and then distributes the aggregated global model to each platform for the next round of training.

[0085] Specifically, each platform (such as different hospitals, laboratories, or research institutions) trains its own graph convolutional network model based on local data. Each platform performs multiple training epochs on its own data and updates its local model parameters. Furthermore, the platforms do not share data, but rather share the trained graph convolutional network parameters (such as model weights). This approach ensures data privacy while allowing for joint training using data from multiple platforms.

[0086] In addition, the central server uses the Federated Avg algorithm from federated learning to aggregate the local model parameters from each platform. The specific steps are as follows: 1. Assume the model parameters for each platform are as follows: w k ,in k This indicates the platform number, and the central server performs a weighted average based on the model parameters and sample size of each platform. 2. The aggregated global model will be returned to each platform for the next round of local training.

[0087] S304. Using the trained global model, generate low-dimensional embedding representations for each node, and store the low-dimensional embeddings and the predicted associations between nodes in a multi-omics database.

[0088] Specifically, a trained global graph convolutional network model is used to infer information from sample data across various platforms, generating low-dimensional embeddings for each node. These embeddings effectively capture latent features in the sample data, representing the similarities and differences between data points. The generated low-dimensional embeddings and predicted relationships between nodes are stored in a multi-omics database. These embeddings can support subsequent joint analysis, querying, and inference. For example, graph databases (such as Neo4j) can be used to store node information, and deep data analysis can be performed based on graph querying and inference techniques.

[0089] Based on stored multi-omics data and embedded representations, further analysis can be performed to reveal potential associations between different omics data. Through reasoning using graph structures and graph neural networks, deep-seated connections between tumor-related genes, pathological features, and clinical data can be discovered, providing data support for medicine.

[0090] The atlas construction module is used to build a tumor multi-omics knowledge graph and identify potential associations between biopsy specimens and pathological information.

[0091] In the description of this invention, constructing a tumor multi-omics knowledge graph to identify potential associations between biopsy specimens and pathological information includes: S311. Collect multi-omics data, extract key features of each omics data, convert them into graph structure representations, extract entities and relationships, and construct a knowledge network.

[0092] Specifically, this involves collecting multi-omics tumor data from various sources, including genomic data, transcriptomic data, proteomic data, clinical data, and pathological information. This data can come from clinical trials, laboratory testing, imaging analysis, and other sources.

[0093] Key features were extracted from each type of omics data, including the following aspects: Genomic data: Extracting gene mutations, copy number variations (CNVs), SNPs (single nucleotide polymorphisms), etc.

[0094] Transcriptome data: Extracting gene expression levels, RNA splicing information, etc.

[0095] Proteomic data: Extracting protein expression levels, post-translational modifications, etc.

[0096] Clinical data: Extracting basic patient information, tumor type, pathological stage, etc.

[0097] Pathological data: Extracted tissue section images, pathology reports, and specimen information.

[0098] Transform the features of each group of learnings into a graph structure representation: Entities from different omics (such as genes, proteins, tumor types, patients, and pathological results) are represented as nodes in the graph. Relationships between entities (such as gene-protein interactions, the association between gene mutations and tumor types, and the relationship between patients and pathological information) are represented as edges, and the weights of the edges can be dynamically adjusted based on the strength or relevance of the relationship.

[0099] S312. Integrate different types of entities and relationships into a knowledge graph, construct a multi-omics knowledge graph, and store the entities and relationships in the knowledge graph in a graph database.

[0100] Specifically, this involves integrating different omics data (such as gene, protein, and pathological information) and their relationships into a unified knowledge graph. Examples include: the relationship between genes and tumor types, and the association between gene mutations and certain tumor types; the relationship between protein expression and clinical response, with the expression levels of certain proteins potentially correlated with patient treatment response; and the relationship between tumor types and pathological types, and the association between tumor pathological classifications and their clinical characteristics.

[0101] Based on the extracted entities and relationships, a multi-level, multi-omics knowledge graph is constructed, with each level representing different types of entities and the relationships between them. The knowledge graph describes complex biological and clinical data through nodes and edges.

[0102] Ultimately, graph databases (such as Neo4j and JanusGraph) are used to store the constructed knowledge graph, facilitating efficient querying, reasoning, and analysis. Graph databases can handle complex relational data and support multi-dimensional query and reasoning operations.

[0103] S313. Perform spatiotemporal correlation modeling on different omics data and pathological information, identify potential multi-omics associations, and mine potential patterns in the knowledge graph to identify potential associations between biopsy specimens and pathological information.

[0104] Specifically, there are often dynamic spatiotemporal correlations between multi-omics data and pathological information. For example, changes in gene expression in tumor cells at different stages, or the different patterns of correlation between pathological results and gene mutations as treatment progresses, can all be identified by using graph neural networks (GNNs) or other deep learning methods to model the spatiotemporal characteristics of nodes in a graph structure. Specifically, based on the patient's treatment history, changes in gene expression, and the pathological progression of the tumor, potential spatiotemporal correlations between specimen data and pathological information can be identified.

[0105] In multi-omics knowledge graphs, potential biological patterns or disease mechanisms can be uncovered through graph reasoning, path search, and other methods. For example, specific mutations in certain genes may lead to changes in protein expression, thereby affecting tumor progression or response to treatment.

[0106] The mining module has been updated to capture the latest literature for information mining and link key information to the knowledge graph.

[0107] Decision management unit 4 is used for resource sharing and monitoring the authorized use records of biopsy specimens.

[0108] In the description of this invention, the decision management unit 4 includes: an informed authorization module (not shown in the figure), a sharing interaction module (not shown in the figure), a usage record module (not shown in the figure), and an early warning response module (not shown in the figure); and the informed authorization module, the sharing interaction module, the usage record module, and the early warning response module are connected in sequence.

[0109] The informed authorization module is used to dynamically authorize the use of specimens through smart terminals, de-identify specimen data, and associate it with the authorization approval number.

[0110] Specifically, the informed consent module enables dynamic authorization of specimen usage through smart terminals (such as mobile devices and computer terminals). By setting permissions, authorized personnel can determine the uses of specimen data (such as scientific research, clinical research, commercialization, etc.) and conduct appropriate review and confirmation during authorization. Before specimen data is authorized for use, the system automatically performs data anonymization to ensure that the data does not contain sensitive personally identifiable information (such as patient names, ID numbers, etc.). The anonymized data can be securely used for analysis or sharing.

[0111] The shared interaction module is used to establish cross-regional resource sharing, conduct virtual multidisciplinary consultations and patient communication through terminal communication, and transcribe consultation and interaction content in real time.

[0112] Specifically, the shared interaction module provides a cross-regional data sharing and resource collaboration platform for medical and research institutions in different regions. Through cloud computing and network communication technologies, it enables real-time data transmission and sharing between different platforms, overcoming geographical and technological limitations. This module supports virtual multidisciplinary consultations via terminal devices, bringing together experts from various disciplines for remote diagnosis and discussion. Doctors and researchers can participate in consultations remotely and provide corresponding treatment suggestions based on specimen data.

[0113] The usage record module is used to monitor the authorization time, scope, and usage records of biopsy specimens in real time using blockchain.

[0114] The early warning and response module is used to issue real-time warnings and freeze unauthorized biopsy specimens.

[0115] Specifically, the early warning response module is responsible for monitoring specimen usage in real time, ensuring that only authorized specimens can be used. When the system detects unauthorized specimen use, it triggers an immediate early warning mechanism. Once unauthorized specimen use is detected, the system automatically freezes the usage permissions for the relevant specimens and notifies the administrator and relevant personnel to ensure that the violation cannot continue.

[0116] In summary, by leveraging the technical solutions described above, and integrating blockchain, federated learning, and process refactoring (standardization, automation, and collaboration), the entire chain of oncology specimens from "collection" to "application" has been upgraded, propelling oncology diagnosis and treatment from "experience-driven" to "data-driven." This enables the digitalization, standardization, and traceability management of biopsy specimen information throughout the entire process. By utilizing adaptive quantile regression forest algorithms and graph structure spatiotemporal feature extraction technology, dynamic early warning and anomaly scoring of specimen data are achieved, effectively detecting and responding to specimen quality issues. Simultaneously, image and biochemical indicator analysis generates processing suggestions to ensure specimen quality meets clinical standards. Furthermore, the construction and dynamic updating of multi-omics databases and knowledge graphs not only enhance the understanding of tumor biological mechanisms but also promote cross-platform data sharing and precise diagnostic and treatment decisions, providing strong data support and decision support for clinical practice and scientific research. This invention ensures that biopsy specimen data throughout its entire lifecycle (from collection, transportation, storage to testing) can be accurately recorded and traced, guaranteeing the integrity, traceability, and transparency of specimen information. By assigning each specimen a unique digital identity and using blockchain technology to upload the data, the invention achieves tamper-proof and efficient management of the data throughout the entire process. By using graph neural networks (GNN) and federated learning technologies to model, mine associations, and generate multi-omics databases from multi-omics data, the invention can efficiently extract potential associations from data from different platforms and different omics types while maintaining data privacy.

[0117] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

Claims

1. A specimen collection and management system for oncology, characterized in that, include: The collection and detection unit is used to collect biopsy specimens and corresponding specimen data; The traceability quality control unit is used to store the entire process information of biopsy specimens using blockchain and to dynamically evaluate and trace the quality information of biopsy specimens. The integrated learning unit is used to build a multi-omics database, generate a tumor multi-omics knowledge graph, and dynamically update knowledge through data mining; The decision-making management unit is used for resource sharing and monitoring the authorized use records of biopsy specimens; The tracking quality control unit includes: a monitoring and early warning module and an evaluation and decision-making module; The monitoring and early warning module is used to dynamically optimize the early warning threshold in real time and identify abnormal information in the specimen data. This includes: predicting different quantiles based on the adaptive quantile regression forest algorithm and adjusting the early warning threshold of each specimen data in real time; constructing a graph structure using specimen data and specimen event nodes, and calculating the weighted sum of reconstruction error and prediction error by extracting the spatiotemporal features of the graph structure to obtain an anomaly score. The evaluation and decision-making module is used to perform image and biochemical index analysis on biopsy specimens, evaluate the specimen quality of biopsy specimens, and generate specimen processing suggestions.

2. The oncology specimen collection and management system according to claim 1, characterized in that, The collection and detection unit, the tracking and quality control unit, the integrated learning module, and the decision management unit are connected in sequence. The tracking and quality control unit also includes: an anomaly tracing module and a recording and reporting module; The anomaly tracing module is used to assign a unique digital identity to the biopsy specimen, record the entire process of data from collection, transportation, storage to testing on the blockchain, and record the entire chain of biopsy specimen operation logs. The record reporting module is used to generate a visualized quality control report with multi-dimensional query and statistical analysis. Furthermore, the monitoring and early warning module, the evaluation and decision-making module, the anomaly tracing module, and the recording and reporting module are connected in sequence.

3. The oncology specimen collection and management system according to claim 1, characterized in that, The process of constructing a graph structure using specimen data and specimen event nodes, extracting the spatiotemporal features of the graph structure, calculating the weighted sum of reconstruction error and prediction error, and obtaining an anomaly score includes: Based on the abnormality score of the specimen data, the corresponding abnormality level is matched according to the preset grading rules, and different abnormality warning mechanisms are triggered to generate abnormal information of the biopsy specimen.

4. The oncology specimen collection and management system according to claim 1, characterized in that, The early warning thresholds for predicting different quantiles based on the adaptive quantile regression forest algorithm and adjusting various sample data in real time include: The sample data is split into time series data and discrete time data. The time series data is segmented by a sliding window and periodic encoding is added. The time difference features are calculated for the discrete time data. Based on the historical distribution of the sample data, the target value is predicted at different quantiles by modeling the quantile regression forest algorithm, and the prediction results of all trees are averaged to obtain the quantile prediction value. By using quantile regression forest to predict different quantiles of the target variable, a dynamic early warning threshold is set. When new sample data arrives, the early warning threshold is automatically adjusted according to the new data distribution.

5. The oncology specimen collection and management system according to claim 1, characterized in that, The formula for calculating the anomaly score is as follows: ; In the formula, Score t Indicates the current t Anomaly scoring at time steps; λ Indicates the weighting coefficient; x t Indicates the current t The numerical values ​​of the specimen data at each time step; Indicates the current t Predictions from the spatiotemporal graph attention network at each time step; Q α ( x t ) indicates the current t The warning threshold for the time step.

6. The oncology specimen collection and management system according to claim 1, characterized in that, The process of analyzing images and biochemical indicators of biopsy specimens, evaluating specimen quality, and generating specimen processing suggestions includes: The specimen data were divided into biopsy image data and physiological index data. The biopsy image data was subjected to image standardization processing, and the physiological index data was subjected to structure processing. Segment tumor cells in biopsy image data, obtain image features by outputting heatmaps, and calculate the proportion of tumor cells to total pixels and the necrotic area; A graph structure of physiological indicator data is constructed, and physiological features are extracted through cross-attention fusion. The degradation probability of quality risk prediction is then calculated. Image features and physiological features are mapped to a common latent space. The distribution parameters are output through a Bayesian neural network to calculate the quality confidence score. By combining the proportion of tumor cells, the output degradation probability, and the confidence score results, the quality of the current biopsy specimen is evaluated, and corresponding specimen processing suggestions are matched.

7. The oncology specimen collection and management system according to claim 1, characterized in that, The integrated learning unit includes: a multi-source integration module, a joint analysis module, a knowledge graph module, and an update and mining module; The multi-source integration module is used to integrate and unify specimen data from different platforms. The joint analysis module is used to model multi-omics associations using graph neural networks and to introduce a federated learning sharing model to generate a multi-omics database. The atlas construction module is used to construct a tumor multi-omics knowledge atlas to identify potential associations between biopsy specimens and pathological information; The update mining module is used to capture the latest literature for information mining and link key information to the knowledge graph; Furthermore, the multi-source integration module, the joint analysis module, the knowledge graph module, and the update mining module are sequentially connected.

8. The oncology specimen collection and management system according to claim 7, characterized in that, The method of using graph neural networks to model multi-omics associations and introducing federated learning sharing models to generate multi-omics databases includes: Standardize specimen data from different platforms, unify the names and identifiers of each type, and use cross-omics collaborative matrix decomposition to predict missing values ​​in the specimen data; Define nodes and edges in a heterogeneous graph structure, dynamically adjust edge weights, train a graph convolutional network locally on each platform, and extract low-dimensional embedding vectors of each omics node through a multi-layer graph neural network. The graph convolutional network parameters are shared across platforms. The central server uses a federated averaging algorithm to aggregate local models and then distributes the aggregated global model to each platform for the next round of training. Using the trained global model, low-dimensional embedding representations are generated for each node, and the low-dimensional embeddings and the predicted associations between nodes are stored in a multi-omics database.

9. The oncology specimen collection and management system according to claim 8, characterized in that, The construction of a tumor multi-omics knowledge graph to identify potential associations between biopsy specimens and pathological information includes: Collect multi-omics data, extract key features of each omics data, convert them into graph structure representations, extract entities and relationships, and construct a knowledge network. Different types of entities and relationships are integrated into a knowledge graph to construct a multi-omics knowledge graph, and the entities and relationships in the knowledge graph are stored in a graph database; Spatiotemporal correlation modeling of different omics data and pathological information is performed to identify potential multi-omics associations and to mine potential patterns in knowledge graphs to identify potential associations between biopsy specimens and pathological information.

10. The oncology specimen collection and management system according to claim 1, characterized in that, The decision management unit includes: an informed authorization module, a sharing and interaction module, a usage record module, and an early warning response module; The informed authorization module is used to dynamically authorize the use of specimens through smart terminals, desensitize specimen data, and associate it with authorization approval numbers. The shared interaction module is used to establish cross-regional resource sharing, conduct virtual multidisciplinary consultations and patient communication and interaction through terminal communication, and transcribe consultation and interaction content in real time. The usage record module is used to monitor the authorization time, authorization scope, and authorized usage records of biopsy specimens in real time using blockchain. The early warning response module is used to issue real-time warnings and freeze unauthorized biopsy specimens. Furthermore, the informed authorization module, the shared interaction module, the usage record module, and the early warning response module are connected sequentially.

Citation Information

Patent Citations

  • Fatigue detecting method, device, apparatus and readable storage medium

    CN109620269A

  • Hospital quality monitoring data analysis and fine management system and method

    CN117038025A

  • Sugar net recognition system based on block chain and federal learning

    CN119445234A

  • Video anomaly detection method based on optical flow reconstruction and variation prediction of enhanced memory

    CN119919848A

  • Deep learning-based lung cancer early screening system

    CN119920445A

Cited By

  • Block chain technology-based scarce human body specimen circulation management method

    CN121788008A