Lymphoma-oriented multi-dimensional special disease data construction and management system and method

By constructing a multi-layered data processing pipeline and intelligent information extraction technology, the issues of data integration, standardization, and privacy security in lymphoma disease data management have been resolved, achieving efficient data management and precision medicine support.

CN121662406APending Publication Date: 2026-03-13SICHUAN ACADEMY OF MEDICAL SCI SICHUAN PROVINCIAL PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Lymphoma-specific data management faces challenges such as difficulty in data integration and standardization, insufficient value mining of unstructured data, low level of data application, and issues related to data quality and privacy. Existing systems struggle to achieve efficient management throughout the entire lifecycle.

Method used

We construct a multi-dimensional disease-specific data construction and management system, which adopts a multi-layered data processing pipeline, including data collection, desensitization, cleaning, standardization, structuring and intelligent processing. It combines OCR and NLP technologies to extract key information and provides advanced application services through disease prediction models.

Benefits of technology

It enables standardized and structured management of multi-source heterogeneous lymphoma data, improves data quality and utilization, supports precise searching, rich statistical analysis and disease prediction, and ensures patient privacy and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121662406A_ABST
    Figure CN121662406A_ABST
Patent Text Reader

Abstract

The invention discloses a lymphoma-oriented multi-dimensional special disease data construction and management system and method. The system comprises an application server, a data governance server and a database server cluster, the application server collects multi-source heterogeneous data from each information system of a hospital through technologies such as ETL; the data governance server executes a multi-layer data processing assembly line, and sequentially completes data desensitization cleaning, unified patient main index management, unstructured information extraction based on optimized OCR and NLP engines, and standardized mapping according to a special disease standard data set; the database cluster adopts a standard database to support online query, and the data warehouse supports deep analysis; and the application server provides scientific research cooperation and disease prediction services based on the double libraries. According to the method, whole-process standardization, structured treatment and intelligent application of the special lymphoma disease data are realized, the problems of difficulty in integration of multi-source data and insufficient utilization of unstructured texts are effectively solved, and the ability of data to support clinical scientific research and precise diagnosis and treatment is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical information technology, specifically to a multi-dimensional disease-specific data construction and management system and method for lymphoma. Background Technology

[0002] Lymphoma, a highly heterogeneous hematologic malignancy, involves multiple stages in its diagnosis and treatment, including imaging examinations, laboratory tests, pathological diagnosis, drug therapy, radiotherapy, and stem cell transplantation, generating massive amounts of multi-source, heterogeneous clinical data. This data is scattered across multiple heterogeneous systems such as Hospital Information System (HIS), Electronic Medical Record System (EMR), Laboratory Information System (LIS), and Picture Archiving and Communication System (PACS), resulting in problems such as inconsistent data standards, heterogeneous formats, and information silos.

[0003] Currently, existing technologies face the following main technical bottlenecks in lymphoma-specific data management: 1. Significant challenges in data integration and standardization: Different medical information systems employ different data standards and coding systems, such as ICD-10 and SNOMED, ​​making direct data integration and utilization difficult. The lack of standardized datasets designed specifically for lymphoma specialty results in a lack of unified standards for data collection and management.

[0004] 2. Insufficient Value Mining of Unstructured Data: A large amount of key medical information, such as imaging examination reports (CT, MRI, PET-CT), pathology reports, etc., exists in the form of natural language text and is unstructured data. Traditional methods struggle to automatically and accurately extract key medical entities (such as tumor involvement sites, SUVmax values, etc.) from these reports, severely limiting the in-depth utilization of the data.

[0005] 3. Low level of data application: Existing data management systems focus on data storage and simple querying, lacking advanced functions (such as multi-condition combination retrieval and statistical analysis) and intelligent applications (such as disease progression prediction) for clinical research, and thus cannot effectively support precision medicine and clinical research.

[0006] 4. Data quality and privacy security: The data cleaning and de-identification processes are not perfect, resulting in problems such as data inconsistency, duplication, and missing data. At the same time, the protection of patient privacy faces challenges.

[0007] Therefore, there is an urgent need for a technical solution that can systematically solve the above problems and realize the full life cycle management of lymphoma disease data from collection and treatment to application. Summary of the Invention

[0008] In view of the shortcomings of the existing technology, the purpose of this invention is to provide a multi-dimensional disease-specific data construction and management system and method for lymphoma, so as to solve the problems mentioned in the background technology.

[0009] To achieve the above objectives, a specific embodiment of the present invention provides a multi-dimensional disease-specific data construction and management system for lymphoma, including at least one application server, a data governance server, and a database server cluster. The application server is equipped with a data acquisition module configured to collect multi-source heterogeneous raw data of lymphoma patients from hospital information systems, electronic medical record systems, laboratory systems, radiology systems, pathology systems, ultrasound systems, electrocardiogram systems, and surgical systems using database synchronization and ETL technologies, and temporarily store the data in the raw database. The data governance server is connected to the application server and configured to execute a multi-layer data processing pipeline, which includes: a data access and preliminary processing layer: storing data from the raw database into an intermediate database after data anonymization and data cleaning, wherein data cleaning includes at least one of format conversion, calculation, and dictionary mapping; a patient identity integration layer: performing unified patient master index management based on the data in the intermediate database; and an unstructured data intelligent processing layer: performing text recognition on medical report images through its integrated and deployed optical character recognition engine optimized for medical report images, which is adapted to common printing patterns in medical documents through a training set. The system handles text, handwritten text, and low-resolution noise interference. It also utilizes a natural language processing engine trained on a lymphoma-specific corpus, to perform named entity recognition and relation extraction on the identified text and raw unstructured text data. The named entity recognition and relation extraction types are predefined medical entities and relations strongly related to lymphoma diagnosis and treatment. This allows for the structured extraction of predefined key medical information from lymphoma-related medical imaging and pathology reports, addressing the specific technical problem of inaccurate and incomplete key information extraction from multi-source heterogeneous reports. A disease-specific standardization layer maps and standardizes all data according to a predefined lymphoma-specific standard dataset, outputting standardized data to a standard database. A database server cluster stores the processed data, including: the standard database, which stores standardized detailed data supporting online queries; a data warehouse connected to the standard database, used for deep processing and data aggregation of the data in the standard database to support upper-layer analytical applications; and an application server configured to provide data application service interfaces to users based on the standard database and data warehouse.

[0010] According to an embodiment of this application, a multi-dimensional disease data construction and management system and method for lymphoma is proposed to solve the technical problems of difficulty in integrating multi-source heterogeneous data of lymphoma, low utilization rate of unstructured data, uneven data quality and lack of advanced intelligent applications, so as to realize the standardized, structured and intelligent management of lymphoma disease data.

[0011] In addition, the multi-dimensional disease-specific data construction and management system and method for lymphoma proposed in this application may also have the following additional technical features: In one embodiment of this application, the predefined lymphoma-specific standard dataset includes multiple data modules, which at least include: patient demographic information, diagnosis and treatment overview, medical history and medical history, physical examination, CT scan, MRI scan, PET-CT scan, lymph node ultrasound scan, other imaging examinations, laboratory tests, MICM scan, pathological examination, treatment information, efficacy evaluation, adverse events, radiotherapy, hematopoietic stem cell transplantation, clinical trials, and follow-up information.

[0012] In one embodiment of this application, the field definitions of the standard dataset conform to at least one of the standards OMOP, CDASH, SNOMED, ​​ICD-10, and ICD-9-CM-3.

[0013] In one embodiment of this application, the application service interface is specifically used to support a scientific research collaboration platform and a disease prediction model; the scientific research collaboration platform provides functions including: scientific research project management, precise patient search, and medical statistical analysis, wherein the scientific research project management includes multi-center permission setting function; the disease prediction model is trained based on structured data in the data warehouse and outputs predictive indicators to assist clinical decision-making.

[0014] In one embodiment of this application, the patient precision search supports multiple retrieval modes, including: keyword retrieval, used for fuzzy matching based on any input keywords; advanced retrieval, used for combined queries based on data element fields and logical operators; and condition tree retrieval, used to construct complex retrieval condition trees by combining "AND", "OR", and "exclude" logical relationships through a graphical interface.

[0015] In one embodiment of this application, the analytical methods provided by the medical statistical analysis include at least one of frequency analysis, cross-chi-square analysis, categorical summary analysis, one-sample t-test, homogeneity of variance test, cluster analysis, principal component analysis, multi-factor ANOVA, hierarchical cluster analysis, and curve regression analysis.

[0016] A method for constructing and managing multidimensional disease-specific data for lymphoma includes the following steps: S1: Through database synchronization technology and ETL technology, collect multi-source heterogeneous raw data of lymphoma patients from hospital information systems, electronic medical record systems, laboratory systems, radiology systems, pathology systems, ultrasound systems, electrocardiogram systems and surgical systems; S2: The original data is temporarily stored in the original database, and data desensitization and data cleaning are performed, wherein data cleaning includes at least one of format conversion, calculation and dictionary mapping; S3: Based on the cleaned data, perform unified patient master index management to associate all scattered data of the same patient; S4: The optical character recognition engine optimized for medical report images performs text recognition on the medical report images, and the natural language processing engine trained on the lymphoma disease corpus performs named entity recognition and relation extraction on the recognized text and the original unstructured text data to extract predefined key medical information that is strongly related to the diagnosis and treatment of lymphoma. S5: Map and standardize all data according to the predefined standard dataset for lymphoma, obtain standardized data and store it in the standard database, and further synchronize the data to the data warehouse for in-depth processing and aggregation; S6: Based on the standard database and data warehouse, and through the data application service interface, respond to user requests and provide data query, data analysis, and calculation services based on disease prediction models.

[0017] In one embodiment of this application, the key medical information extracted in step S4 includes examination findings, examination conclusions, tumor involvement sites, sites of abnormally increased FDG uptake, and abnormally increased FDG uptake values. The key information corresponds to the imaging examination module and pathological examination module in the predefined lymphoma-specific standard dataset.

[0018] In one embodiment of this application, the step S6 of responding to a user request specifically includes: Responding to complex search criteria built on condition trees, it accurately filters out patient cohort data that meet the criteria from standard databases and data warehouses; Based on the patient cohort data, medical statistical analysis algorithms and disease prediction models are run to generate statistical analysis reports and predictive indicators.

[0019] The advantages of this invention compared to existing technologies are: (1) By constructing a standard dataset covering the entire diagnosis and treatment cycle of lymphoma (including multiple modules such as MICM examination and efficacy evaluation), and adopting a multi-layer data processing pipeline, the standardization and structuring of multi-source heterogeneous data were realized, which significantly improved the data quality.

[0020] (2) By adopting OCR and NLP technologies optimized for medical scenarios, key medical information can be automatically and accurately extracted from unstructured texts such as image reports and pathology reports, which greatly releases the value of data.

[0021] (3) The system supports precise patient search (keyword, advanced, conditional tree search), rich medical statistical analysis and AI-based disease prediction models, which strongly support clinical research and precision diagnosis and treatment.

[0022] (4) Through measures such as data anonymization and unified EMPI management, patient privacy and security were ensured, and medical data compliance requirements were met.

[0023] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a system architecture and data flow diagram of a multi-dimensional disease-specific data construction and management system and method for lymphoma according to one embodiment of the present invention; Figure 2 This is a detailed data governance pipeline diagram of a multi-dimensional disease-specific data construction and management system and method for lymphoma according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating the precise patient search process of a multi-dimensional disease-specific data construction and management system and method for lymphoma, as described in one embodiment of the present invention. Figure 4 This is a flowchart illustrating the medical statistical analysis and prediction application of a multi-dimensional disease-specific data construction and management system and method for lymphoma in one embodiment of the present invention. Figure 5 This is a sequence diagram of unstructured data intelligent processing for a multi-dimensional disease-specific data construction and management system and method for lymphoma, as described in one embodiment of the present invention. Figure 6 This is a flowchart illustrating the disease data mapping and standardization process of a multi-dimensional disease data construction and management system and method for lymphoma in one embodiment of the present invention. Figure 7 This is a flowchart illustrating the construction and application of a disease prediction model for a multi-dimensional disease-specific data construction and management system and method for lymphoma, as described in one embodiment of the present invention. Figure 8 This is a schematic diagram of the architecture of a multi-dimensional disease-specific data construction and management system for lymphoma according to one embodiment of the present invention. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] like Figures 1 to 8 As shown in the figure, the multi-dimensional disease data construction and management system and method for lymphoma according to an embodiment of the present invention consists of three core parts: an application server, a data governance server, and a database server cluster. Each component establishes a stable communication connection through the TCP / IP protocol, and a distributed deployment mode is adopted to ensure the efficiency and reliability of data processing.

[0028] Example 1: System Hardware Architecture and Deployment Environment This system employs a distributed architecture, with the hardware environment comprising multiple servers, specifically an application server cluster, a data governance server cluster, and a database server cluster. The application servers consist of at least two Dell PowerEdge R740 servers forming a load-balanced cluster, each configured with an Intel Xeon Silver 4210 processor, 64GB of memory, and a 2TB SAS hard drive, used to run data acquisition modules and application service interfaces. The data governance servers utilize Huawei FusionServer 2288H V5 servers, configured with two Intel Xeon Gold 5218 processors, 128GB of memory, and a 4TB NVMe SSD, used to run various compute-intensive tasks in the data governance pipeline. The database server cluster uses three IBM Db2 PureScale nodes, each configured with 256GB of memory and 16TB of SSD storage, to build a highly available database environment.

[0029] The software environment includes: CentOS 7.6 operating system, Oracle 19c and Greenplum 6.0 database management system, Apache NiFi 1.12 ETL tool, Spark NLP 3.0 natural language processing framework, and Java Spring Boot 2.5 application service framework. All servers are connected via 10 Gigabit Ethernet, deployed in the hospital's internal data center, and isolated from external networks by firewalls to ensure data security.

[0030] Example 2: Specific Implementation of the Data Acquisition Module The data acquisition module is deployed on the application server and uses various technologies to acquire data from multiple sources. For relational database data sources such as HIS and EMR systems, a real-time incremental acquisition method based on JDBC is used, and data changes are captured through database log parsing technology (CDC). For imaging systems such as PACS, the DICOM standard protocol is used for image data acquisition, while related text report data is acquired through the HL7 protocol.

[0031] The specific data collection process includes: first, configuring data source connection parameters, including IP address, port, database type, authentication information, etc.; then defining the data collection task scheduling strategy, using timed full collection for data with low change frequency such as patient basic information, and using real-time incremental collection for data with high change frequency such as test results; finally, setting data quality check rules to monitor and implement retry mechanisms for network anomalies and data format errors during the data collection process.

[0032] The data acquisition module also implements a breakpoint resume function. When the network is interrupted or the system fails, it can record the breakpoint position and resume acquisition from the breakpoint after recovery, ensuring data integrity. The acquired raw data is first stored in the temporary storage area of ​​the raw database and stamped with a timestamp and data source identifier.

[0033] Example 3: Detailed Workflow of the Data Governance Pipeline The multi-layered data processing pipeline executed by the data governance server specifically includes the following four layers: 3.1 Data Access and Preliminary Processing Layer This layer primarily handles data anonymization and basic cleaning. Data anonymization employs a hash-based pseudonymization process, irreversibly encrypting direct identifiers such as patient names, ID numbers, and phone numbers to generate unified pseudonym identifiers. Simultaneously, quasi-identifiers such as addresses are generalized, for example, by generalizing detailed addresses to the district / county level.

[0034] Data cleaning includes three sub-steps: format conversion to unify the display format of date, time, numerical, and other data; calculation processing to perform derivative calculations on the raw data, such as calculating age based on birth date and BMI based on height and weight; and dictionary mapping to map the local codes used within the hospital to a standard terminology system, such as mapping internal drug codes to ATC standard codes.

[0035] 3.2 Patient Identity Integration Layer This layer implements patient master index management using the EMPI algorithm. It employs a combination of rule-based matching and machine learning-based fuzzy matching. First, deterministic matching is performed, using unique identifiers such as ID card number and medical insurance card number for precise matching. For patients lacking unique identifiers, a probabilistic matching algorithm is used to calculate similarity scores for attributes such as name, gender, date of birth, and phone number. Records exceeding a set threshold after comprehensive weighting are considered to be from the same patient.

[0036] The matching algorithm uses Jaro-Winkler distance to calculate name similarity and absolute date difference to calculate date similarity. The final similarity score is a weighted sum of the similarities of each attribute. A globally unique patient master index is generated for each unique patient, and a mapping relationship is established between the patient master index and the local patient identifier in each business system.

[0037] 3.3 Intelligent Processing Layer for Unstructured Data This layer is specifically designed for processing unstructured text data in medical reports. The OCR engine optimized for medical report images adopts a deep learning-based CRNN-CTC architecture. During training, it uses a dataset of 100,000 medical report images with various noise levels, including stamp occlusions, handwritten annotations, and low-resolution images, enabling the engine to effectively resist these noise interferences.

[0038] The natural language processing engine employs a pre-trained model based on the BERT architecture and undergoes domain-adaptive training on 500,000 specially collected lymphoma-related medical records. The entity types defined in the named entity recognition task include: body parts, disease names, examination items, drug names, examination values, time information, and other key entities for lymphoma diagnosis and treatment. The relation extraction task defines relation types including: location-lesion relationships, examination-value relationships, and drug-dosage relationships.

[0039] The specific processing flow is as follows: First, the scanned report image is converted into text using an OCR engine. Then, word segmentation, part-of-speech tagging, named entity recognition, and relation extraction are performed using an NLP engine. Finally, the extracted structured information is output in a predefined JSON format.

[0040] 3.4 Standardization Layer for Specific Diseases This layer maps the processed data to a lymphoma-specific standard dataset. This dataset contains 19 main modules and 662 fields, each with clearly defined data type, value range, and terminology standards. The mapping process uses a rule-based transformation engine to convert the source data fields into target standard fields according to preset mapping rules.

[0041] For standard terminology, a code mapping method is used to establish a mapping table between local terms and standard terms. For numerical data, units are standardized and data precision is normalized. All mapping operations are logged in an audit log to ensure the traceability of data processing.

[0042] Example 4: Database Cluster Architecture Design The database server cluster adopts a layered storage architecture, including two layers: a standard database and a data warehouse.

[0043] The standard database uses Oracle 19c relational database to store standardized detailed data that supports online business queries. The database table structure is strictly designed according to the standard dataset for lymphoma, with each module corresponding to a main table, and related sub-modules linked through foreign keys. B-tree indexes are created for commonly used query conditions such as patient ID, examination date, and diagnosis name, and inverted indexes are created for full-text search requirements.

[0044] The data warehouse uses the Greenplum distributed database to support complex analytical queries. The data model employs a dimensional modeling approach, constructing a patient-centric fact table and dimension tables around dimensions such as time, diagnosis, and treatment. Data is synchronized from a standard database through regular ETL jobs, and preprocessing operations such as data aggregation and summarization are performed in the data warehouse to improve analytical query performance.

[0045] Example 5: Implementation of Application Service Interface Functionality The application server provides data application service interfaces through a RESTful API, which uses the OAuth 2.0 protocol for authentication and access control. The research collaboration platform implements the following core functions: 5.1 Research Project Management: Provides project creation, editing, and deletion functions, and supports multi-center permission settings. Project administrators can set project participants and their permission roles, such as read-only permissions, data export permissions, and patient inclusion / exclusion permissions. The system records project operation logs to meet research audit requirements.

[0046] 5.2 Precise Patient Search: Provides three search modes.

[0047] Keyword search supports natural language input, such as diffuse large B-cell lymphoma with an SUVmax greater than 10. The system automatically parses the query intent and converts it into background query conditions.

[0048] Advanced search provides a visual query condition building interface, allowing users to construct query conditions by selecting fields, operators, and inputting values ​​from dropdown menus.

[0049] Condition tree retrieval supports complex combinations of logical conditions. Users can drag and drop to build query condition trees that include AND, OR, and exclusion logic, such as a diagnosis name equal to diffuse large B-cell lymphoma and an IPI score greater than or equal to 3, or having received CAR-T therapy but excluding those over 70 years of age.

[0050] 5.3 Medical Statistical Analysis: Integrating multiple statistical analysis methods, it provides a complete toolchain from descriptive statistics to advanced analysis. Frequency analysis is used for the distribution statistics of categorical variables; cross-chi-square analysis is used for the association test between two categorical variables; one-sample t-test is used for comparing continuous variables with the population mean; survival analysis supports Kaplan-Meier curves and Cox regression models; machine learning analysis provides unsupervised learning methods such as cluster analysis and principal component analysis.

[0051] Example 6: Construction and Application of Disease Prediction Models The disease prediction model is trained on structured data from a data warehouse. First, feature engineering is performed to extract predictive features from patient demographic information, medical records, and test results. These features include static features such as age, gender, and underlying diseases, as well as dynamic features such as trends in test indicators and treatment response.

[0052] The model employs the Gradient Boosting Decision Tree (GBDT) algorithm, with parameter tuning using five-fold cross-validation. After training, the model provides services via an API interface, taking patient feature data as input and outputting indicators such as disease progression risk score and survival prediction. The model is periodically retrained with new data to maintain predictive performance.

[0053] Example 7: Complete Workflow Example Using a lymphoma clinical study as an example, the complete workflow of the system is illustrated: Step 1: After logging into the system, researchers create a new research project, name it "Analysis of Prognostic Factors in Diffuse Large B-cell Lymphoma," and configure multi-center collaboration permissions.

[0054] Step 2: Construct patient inclusion criteria using the conditional tree search function: diagnosis of diffuse large B-cell lymphoma, age 18-75 years, completion of at least 3 cycles of chemotherapy, and availability of baseline PET-CT scan results. The system returns a queue of eligible patients within seconds.

[0055] Step 3: Conduct a data quality assessment on the included patient cohort, examine the distribution of missing values, and impute missing values ​​for key variables.

[0056] Step 4: Perform univariate analysis to compare survival differences among different IPI score groups, and generate Kaplan-Meier survival curves and log-rank test p-values.

[0057] Step 5: Perform multivariate Cox regression analysis, including covariates such as age, stage, LDH level, and ECOG score, and calculate the hazard ratio (HR) and confidence interval.

[0058] Step Six: Using a disease prediction model, input new patient characteristics to obtain individualized prognostic risk scores and treatment recommendations.

[0059] Step 7: Export the analysis results, including statistical tables, charts, and model parameters, and generate a research report.

[0060] The entire process is completed on a unified platform, eliminating the need to switch between different systems and significantly improving research efficiency.

[0061] Example 8: System Monitoring and Maintenance The system establishes a comprehensive monitoring framework, including infrastructure monitoring, application performance monitoring, and business data monitoring. Infrastructure monitoring tracks server CPU, memory, and disk usage, and sets threshold alerts. Application performance monitoring records metrics such as API response time and concurrent user count. Business data monitoring tracks data quality metrics, such as data integrity, timeliness, and consistency.

[0062] The system generates weekly monitoring reports, including data access volume, data processing success rate, and user activity statistics. Regular system backups and recovery drills are performed to ensure business continuity. System performance optimizations are conducted quarterly, including database index rebuilding and query plan optimization.

[0063] The technical solutions described in the embodiments of this application construct a standard dataset for lymphoma covering the entire process of diagnosis and treatment, and innovatively adopt a multi-layered data governance pipeline that integrates intelligent information extraction technology. This achieves deep integration and high-quality structuring of multi-source heterogeneous data. Based on this, the application services provided, such as precise retrieval, multi-dimensional analysis, and prediction models, effectively solve the key technical bottlenecks of traditional medical data systems in data integration, utilization of unstructured text, and advanced analysis applications. This provides a powerful full-lifecycle data support platform for clinical research and precision diagnosis and treatment of lymphoma.

[0064] Obviously, the above-described embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention also intends to include these modifications and variations.

Claims

1. A multi-dimensional disease-specific data construction and management system for lymphoma, characterized in that, It includes at least one application server, data governance server, and database server cluster; The application server is equipped with a data acquisition module, which is configured to collect raw data of multi-source heterogeneous lymphoma patients from hospital information systems, electronic medical record systems, laboratory systems, radiology systems, pathology systems, ultrasound systems, electrocardiogram systems, and surgical systems through database synchronization technology and ETL technology, and temporarily store the data in the raw database. The data governance server is connected to the application server and configured to execute a multi-layer data processing pipeline, which includes: Data access and preliminary processing layer: After data anonymization and data cleaning, the data in the original database is stored in the intermediate database, where data cleaning includes at least one of format conversion, calculation and dictionary mapping; Patient identity integration layer: Based on data from the intermediate database, it performs unified patient master index management; The unstructured data intelligent processing layer: It performs text recognition on medical report images through its integrated optical character recognition engine optimized for medical report images. This engine is adapted to common medical document interference such as stamps, handwriting, and low-resolution noise through a training set. It also calls its integrated natural language processing engine, trained on a lymphoma-specific corpus, to perform named entity recognition and relation extraction on the recognized text and the original unstructured text data. The entity types for named entity recognition and the relation types for relation extraction are predefined medical entities and relations strongly related to lymphoma diagnosis and treatment. This allows for the structured extraction of predefined key medical information from lymphoma-related medical imaging and pathology reports, addressing the specific technical problem of inaccurate and incomplete key information extraction in multi-source heterogeneous reports. Disease-specific standardization layer: All data are mapped and standardized according to a predefined lymphoma disease-specific standard dataset, and the standardized data is output to the standard database. The database server cluster, used to store the processed data, includes: The standard database is used to store standardized detailed data that supports online querying; A data warehouse, connected to the standard database, is used for in-depth processing and data aggregation of the data in the standard database to support upper-level analytical applications; The application server is also configured to provide data application service interfaces to users based on the standard database and data warehouse.

2. The multi-dimensional disease-specific data construction and management system for lymphoma according to claim 1, characterized in that, The predefined lymphoma-specific standard dataset contains multiple data modules, which include at least: patient demographic information, diagnosis and treatment overview, medical history and medical history, physical examination, CT scan, MRI scan, PET-CT scan, lymph node ultrasound scan, other imaging examinations, laboratory tests, MICM scan, pathological examination, treatment information, efficacy evaluation, adverse events, radiotherapy, hematopoietic stem cell transplantation, clinical trials, and follow-up information.

3. The multi-dimensional disease-specific data construction and management system for lymphoma according to claim 2, characterized in that, The field definitions of the standard dataset conform to at least one of the following standards: OMOP, CDASH, SNOMED, ​​ICD-10, and ICD-9-CM-3.

4. The multi-dimensional disease-specific data construction and management system for lymphoma according to claim 1, characterized in that, The application service interface is specifically used to support scientific research collaboration platforms and disease prediction models; The research collaboration platform provides functions including: research project management, precise patient search, and medical statistical analysis, among which research project management includes multi-center permission setting function; The disease prediction model is trained based on structured data in the data warehouse and outputs predictive indicators to assist clinical decision-making.

5. A multi-dimensional disease-specific data construction and management system for lymphoma according to claim 4, characterized in that, The precise patient search supports multiple search modes, including: Keyword search is used to perform fuzzy matching based on any input keywords; Advanced search is used for combined queries based on data element fields and logical operators; Condition tree retrieval is used to construct complex retrieval condition trees by combining "AND", "OR" and "exclude" logical relationships through a graphical interface.

6. The multi-dimensional disease-specific data construction and management system for lymphoma according to claim 4, characterized in that, The analytical methods provided by the medical statistical analysis include at least one of the following: frequency analysis, cross-chi-square analysis, categorical summary analysis, one-sample t-test, homogeneity of variance test, cluster analysis, principal component analysis, multivariate analysis of variance, hierarchical cluster analysis, and curve regression analysis.

7. A method for constructing and managing multi-dimensional disease-specific data for lymphoma, characterized in that, Includes the following steps: S1: Through database synchronization technology and ETL technology, collect multi-source heterogeneous raw data of lymphoma patients from hospital information systems, electronic medical record systems, laboratory systems, radiology systems, pathology systems, ultrasound systems, electrocardiogram systems and surgical systems; S2: The original data is temporarily stored in the original database, and data desensitization and data cleaning are performed, wherein data cleaning includes at least one of format conversion, calculation and dictionary mapping; S3: Based on the cleaned data, perform unified patient master index management to associate all scattered data of the same patient; S4: The optical character recognition engine optimized for medical report images performs text recognition on the medical report images, and the natural language processing engine trained on the lymphoma disease corpus performs named entity recognition and relation extraction on the recognized text and the original unstructured text data to extract predefined key medical information that is strongly related to the diagnosis and treatment of lymphoma. S5: Map and standardize all data according to the predefined standard dataset for lymphoma, obtain standardized data and store it in the standard database, and further synchronize the data to the data warehouse for in-depth processing and aggregation; S6: Based on the standard database and data warehouse, and through the data application service interface, respond to user requests and provide data query, data analysis, and calculation services based on disease prediction models.

8. The method for constructing and managing multi-dimensional disease-specific data for lymphoma according to claim 7, characterized in that, The key medical information extracted in step S4 includes examination findings, examination conclusions, tumor involvement sites, sites of abnormally increased FDG uptake, and abnormally increased FDG uptake values. These key information correspond to the imaging examination module and pathological examination module in the predefined lymphoma-specific standard dataset.

9. A method for constructing and managing multi-dimensional disease-specific data for lymphoma according to claim 7, characterized in that, The response to the user request in step S6 specifically includes: Responding to complex search criteria built on condition trees, it accurately filters out patient cohort data that meet the criteria from standard databases and data warehouses; Based on the patient cohort data, medical statistical analysis algorithms and disease prediction models are run to generate statistical analysis reports and predictive indicators.