Medical health big data probing system

By integrating large language model technology with medical expertise, an intelligent medical data exploration system was built, which solved the problems of high professional knowledge requirements, lack of intelligent assistance capabilities, and insufficient privacy protection in the utilization of medical and health big data, and achieved efficient and secure data analysis and report generation.

CN120954601APending Publication Date: 2025-11-14SHANGHAI LINGZAI TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511058525.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

The current utilization of medical and health big data suffers from problems such as high requirements for professional knowledge, lack of intelligent assistance capabilities, fragmented exploration process, and insufficient privacy protection.

Method used

By integrating large language model technology with medical expertise, an intelligent medical data exploration system is constructed, including modules for requirements understanding, compliance review, code generation, query execution, and report generation, ensuring data security and privacy protection.

Benefits of technology

It enables automatic conversion of natural language requirements, compliance review, data analysis, and security protection, lowering the professional threshold, improving exploration efficiency, and ensuring data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954601A_ABST
    Figure CN120954601A_ABST
Patent Text Reader

Abstract

The invention relates to a medical health big data exploration system, and the system comprises a demand understanding module which is used for converting a data exploration demand into a structured exploration task; the compliance review module is used for performing compliance review on the probing task; the code generation module is used for generating a query code according to the probing task and the target database characteristics; the query execution module is used for querying query result data corresponding to the query code from a target database according to the query code; the query result analysis module is used for performing multi-dimensional analysis on the query result data to obtain a multi-dimensional analysis result; a report generation module; and the result presentation module is used for presenting the exploration report in an interactive mode in a user interface and supporting iterative improvement of the report by providing a feedback mechanism, so that the problems of data isomerism, difficulty in quality evaluation, low exploration efficiency, high professional threshold, insufficient privacy protection and the like in the existing medical data exploration process are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical big data analysis, and in particular to a medical and health big data exploration system. Background Technology

[0002] With the rapid development of medical informatization, medical institutions and health management systems have accumulated massive amounts of clinical data, health records, and medical management information. These big data resources contain enormous value, but the current utilization of medical and health big data still faces the following problems:

[0003] Professional knowledge requirements: Understanding and analyzing medical data requires professional medical knowledge, while data processing requires technical skills. This cross-disciplinary knowledge requirement raises the threshold for data utilization.

[0004] Lack of intelligent assistance capabilities: Traditional tools cannot understand medical needs expressed in natural language and cannot automatically convert them into effective data queries, requiring manual, tedious needs conversion and code writing;

[0005] Weak privacy protection mechanisms: Most tools treat privacy protection as an external requirement rather than a built-in function, lacking automated mechanisms for identifying and protecting sensitive information, which increases the risk of data breaches.

[0006] Fragmented exploration process: Existing methods typically handle the processes of requirement understanding, code generation, and result analysis in a separate manner, lacking an end-to-end integrated solution, resulting in an inefficient and error-prone exploration process. Summary of the Invention

[0007] (a) Technical problems to be solved

[0008] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a medical and health big data exploration system, which solves the technical problems such as the lack of intelligent assistance capabilities in the prior art.

[0009] (II) Technical Solution

[0010] To achieve the above objectives, the main technical solutions adopted by the present invention include:

[0011] This invention provides a medical and health big data exploration system, comprising:

[0012] The demand understanding module is used to obtain the data exploration needs input by users through natural language, and to call up relevant medical guidelines, medical literature and historical experience databases to automatically recommend the calculation logic and judgment criteria of various indicators, and to transform the data exploration needs into structured exploration tasks in real time.

[0013] The compliance review module is used to conduct compliance reviews of exploration tasks based on a multi-level compliance review mechanism, and to identify and filter high-risk exploration requests.

[0014] The code generation module is used to generate query code based on the exploration task and the characteristics of the target database, provided that compliance review has been passed.

[0015] The query execution module is used to retrieve query result data corresponding to the query code from the target database based on the query code.

[0016] The query results analysis module is used to perform multi-dimensional analysis on the query results data and obtain multi-dimensional analysis results;

[0017] The report generation module is used to generate hierarchical investigation reports based on multi-dimensional analysis results;

[0018] The compliance review module is also used to conduct compliance and security reviews of the investigation reports to ensure that the content is compliant and does not contain sensitive information;

[0019] The results presentation module is used to present the investigation report in an interactive manner in the user interface, provided that compliance and security reviews have been passed, and to support iterative improvement of the report by providing a feedback mechanism.

[0020] Therefore, by using the above technical solutions, the embodiments of this application can understand the exploration needs expressed in natural language, automatically conduct compliance reviews, generate efficient query codes, perform data analysis, and generate professional reports, while ensuring data security and privacy protection. This can solve the problems of data heterogeneity, difficulty in quality assessment, low exploration efficiency, high professional threshold, and insufficient privacy protection in the existing medical data exploration process.

[0021] In one possible embodiment, data exploration requirements include at least one of the following: exploration purpose and background, definition of exploration object, statistical indicators and grouping dimensions, data quality exploration requirements, and definition and calculation logic of exploration indicators.

[0022] In one possible embodiment, the target definition includes at least one of the following: the subject of the investigation, the disease type, the time range of the target, demographic characteristics, clinical characteristics, treatment-related characteristics, and other screening criteria.

[0023] In one possible embodiment, the statistical indicators in the statistical indicators and grouping dimensions include frequency statistics and central tendency statistics; wherein, frequency statistics cover the total number and proportion; and central tendency statistics include the mean and median.

[0024] In one possible embodiment, the statistical indicators and grouping dimensions include at least one of time-dimension grouping, spatial-dimension grouping, and clinical-dimension grouping.

[0025] In one possible embodiment, the compliance review module is specifically used to perform deep matching between the exploration task and the rule base to identify and filter high-risk exploration requests; wherein, the rule base includes at least one of the following: core data protection rules, sensitive information identification and protection rules, identity recognition and re-identification risk prevention and control rules, compliance and authorization control rules, system security and performance protection rules, and statistical quality and bias control rules.

[0026] In one possible embodiment, the sensitive information identification and protection rules include at least one of the following: rules for exposure to medically sensitive areas, rules for protection of biometric information, and rules for special protection of vulnerable groups.

[0027] In one possible embodiment, the identity recognition and re-identification risk control rules include at least one of the following: direct identifier exposure rules, quasi-identifier combination risk rules, spatiotemporal trajectory association risk rules, and social relationship network exposure rules.

[0028] In one possible embodiment, the compliance and authorization control rules include at least one of the following: hierarchical authorization verification rules, ethical review compliance rules, and legal and regulatory adaptation rules.

[0029] In one possible embodiment, the system security and performance protection rules include at least one of the following: query complexity and resource consumption rules and abnormal behavior detection rules.

[0030] To make the above-mentioned objectives, features and advantages to be achieved by the embodiments of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0031] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This illustration shows a structural diagram of a medical and health big data exploration system provided in an embodiment of this application;

[0033] Figure 2 This document illustrates a flowchart of a medical and health big data exploration method provided in an embodiment of this application.

[0034] Figure 3 This document illustrates a flowchart of a method for a user to input their data exploration needs via natural language, as provided in an embodiment of this application.

[0035] Figure 4 The diagram shows the structure of a key information extraction model provided in an embodiment of this application. Detailed Implementation

[0036] To better explain and facilitate understanding of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0037] Breakthroughs in large language model technology have offered new possibilities for solving related problems in existing solutions. Large language models possess the ability to understand natural language, process complex instructions, and generate specialized code. Combined with knowledge from the medical field, they hold promise for enabling more intelligent and efficient medical data exploration systems. Particularly in areas such as requirements understanding, content review, code generation, and result interpretation, large language models can significantly lower the barrier to entry, improve exploration efficiency, and provide strong support for the value mining of medical big data.

[0038] Based on this, this application provides a medical and health big data exploration system. By integrating large language model technology with medical expertise, an intelligent and automated medical data exploration platform is constructed. This platform can understand exploration needs expressed in natural language, automatically conduct compliance reviews, generate efficient query codes, perform data analysis, and generate professional reports, while ensuring data security and privacy protection. This solves the problems existing in the current medical data exploration process, such as data heterogeneity, difficulty in quality assessment, low exploration efficiency, high professional threshold, and insufficient privacy protection.

[0039] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present invention can be understood more clearly and thoroughly, and that the scope of the present invention can be fully conveyed to those skilled in the art.

[0040] Please see Figure 1 , Figure 1 A schematic diagram of the structure of a medical and health big data exploration system provided in an embodiment of this application is shown. Figure 1 As shown, the medical and health big data exploration system includes a user interaction layer, an application logic layer, a data processing layer, and an infrastructure layer.

[0041] The user interaction layer provides an intuitive web interface for users and includes the following components:

[0042] Intelligent dialogue component: A conversational interface based on a large language model, supporting natural language interaction;

[0043] Exploration Task Management Interface: Used to create, view, and manage exploration tasks;

[0044] The results presentation module features a report display interface that supports interactive functions such as dynamic filtering and drill-down analysis. For example, after passing compliance and security reviews, the investigation report can be presented interactively within the user interface, and a feedback mechanism can be provided to support iterative improvement of the report.

[0045] User feedback collection component: Collects user feedback on the exploration results.

[0046] In addition, the application logic layer is responsible for implementing the core functions of the system and includes the following modules:

[0047] Demand Understanding Module: It acquires the data exploration needs input by users through natural language, and calls the knowledge fusion module to call related medical guidelines, medical literature and historical experience bases, automatically recommends the calculation logic and judgment criteria of various indicators, and transforms the data exploration needs into structured exploration tasks in real time.

[0048] Knowledge fusion module: Integrates information from multiple sources, including medical knowledge bases and database structure information;

[0049] Compliance review module: Conducts compliance reviews of exploration tasks based on a multi-level compliance review mechanism, and identifies and filters high-risk exploration requests;

[0050] Code generation module: After passing compliance review, it generates query code based on the exploration task and the characteristics of the target database;

[0051] Query execution module: Based on the query code, retrieve the query result data corresponding to the query code from the target database;

[0052] Query Result Analysis Module: Performs multi-dimensional analysis on query result data to obtain multi-dimensional analysis results;

[0053] Report generation module: Generates tiered investigation reports based on multi-dimensional analysis results;

[0054] Compliance review module: Conducts compliance and security reviews of investigation reports to ensure that the content is compliant and does not contain sensitive information.

[0055] In addition, the data processing layer is responsible for data storage, access, and processing, and includes the following components:

[0056] Query execution engine: Securely executes the generated query code;

[0057] Data access adapter: Supports connections to different types of data sources;

[0058] Intermediate result cache: Stores intermediate query results;

[0059] Knowledge base management: Managing resources such as medical knowledge bases and data dictionaries;

[0060] Compliance review rule base management: Manages and stores compliance review rules.

[0061] In addition, the infrastructure layer provides the basic environment required for the system to operate, including the following components:

[0062] Computing resource management: Allocating and monitoring the computing resources used by the system;

[0063] Security and access control: Ensuring system and data security;

[0064] Logs and monitoring: Record system operation logs and monitor them;

[0065] Large Language Model Service: Provides reasoning capabilities for large language models.

[0066] In addition, to facilitate understanding of the application logic layer, the following provides a detailed description of the content involved in the application logic layer.

[0067] Specifically, such as Figure 2 As shown, Figure 2 A flowchart illustrating a medical and health big data exploration method provided in an embodiment of this application is shown. Figure 2 As shown, the method for exploring big data in healthcare includes:

[0068] In step S210, the user inputs their data to explore needs through natural language, which is then collected and clarified by the needs understanding module and transformed into a structured task.

[0069] Step S220: Conduct risk assessment and control of requirements through the compliance audit module.

[0070] In step S230, if the review is approved, the code generation module will intelligently generate efficient query code, which will then be securely executed by the query execution module to obtain the data.

[0071] In step S240, the obtained raw results will be processed by the query result analysis module and the report generation module to form a preliminary intelligent exploration report.

[0072] Step S250: The investigation report is examined through the compliance review module to ensure that the content is compliant and does not contain sensitive information.

[0073] In step S260, the review approval report is displayed interactively through the results presentation module, and users can provide feedback, which the system then iteratively improves.

[0074] Therefore, by utilizing the aforementioned technical solutions, this application can continuously optimize the accuracy and efficiency of the exploration. The entire process fully leverages the various technical solutions proposed in this application, aiming to provide medical researchers with an efficient, safe, and easy-to-use data exploration platform.

[0075] Furthermore, for the requirements understanding module, a complete exploration requirement may include at least one of the following: exploration purpose and background, definition of the exploration object, statistical indicators and grouping dimensions, data quality exploration requirements, and definition and calculation logic of exploration indicators.

[0076] In addition, the purpose and background of the investigation may include at least one of the following: the subject of the investigation, the type of disease, the time range of the target subjects, demographic characteristics, clinical characteristics, treatment-related characteristics, and other screening criteria.

[0077] The subject of the investigation may include the user's investigation purpose (e.g., clinical research, drug development, real-world research, market insights, etc.), the expected application scenarios of the investigation results, and the follow-up plans after the investigation.

[0078] The disease type can include the subject of the investigation (e.g., a patient group or a medical event), the disease type (e.g., a single disease, such as type 2 diabetes or non-small cell lung cancer, or a disease category, such as cardiovascular disease or diabetes), the time range of the target population (e.g., the time period of consultation, diagnosis, or treatment), demographic characteristics (e.g., age range, gender, and geographical distribution), clinical characteristics (e.g., diagnostic criteria, disease stage, severity), treatment-related characteristics (e.g., treatment method, medication, type of surgery), and other screening criteria (e.g., insurance type, department visited).

[0079] For statistical indicators and grouping dimensions, the statistical indicator exploration phase includes only two categories: frequency statistics and central tendency statistics. Frequency statistics cover totals and proportions (e.g., total number of patients, diagnosis rate); central tendency statistics include mean and median (e.g., average treatment cost, median survival time). Furthermore, the option to group statistics is available. If grouping is required, it can be performed according to at least one of the following dimensions: time, space, and clinical. Time-based grouping refers to displaying data distribution at different time granularities such as year, quarter, and month; spatial-based grouping refers to grouping by administrative division (province, city, district / county), medical institution level, specific hospital, department, etc.; clinical-based grouping refers to grouping by disease ICD coding system, disease severity classification, treatment plan category, etc., supporting aggregation by professional dimensions such as drug category, treatment pathway, and examination type, combined with patient demographic characteristics (age group, gender, occupation, etc.). Moreover, when multiple groupings exist, it can be determined whether cross-analysis is needed between the multiple grouping dimensions.

[0080] For data quality exploration needs, the specific list of data items that need to be understood is required. For these data items, users may want to know their completeness (missing rate), accuracy (outlier ratio, conformity with standard range), etc. (such as laboratory test results, imaging test results, medication status, follow-up data, etc.).

[0081] Regarding the definition and calculation logic of the exploration indicators, the judgment logic and calculation method for the screening and grouping conditions involved in the exploration process (for example, the specific standard for achieving blood glucose target is fasting blood glucose <7.0mmol / L, etc.).

[0082] To facilitate understanding of the specific process of the requirements understanding module, the following description uses specific examples.

[0083] Specifically, such as Figure 3 As shown, Figure 3 A flowchart illustrating a method for a user to probe their data needs by inputting natural language, according to an embodiment of this application, is shown. Figure 3 As shown, the method includes:

[0084] Step S310: Input of requirements and preliminary analysis.

[0085] Specifically, users submit data exploration requests via natural language (e.g., "I want to know about patients diagnosed with type 2 diabetes between 2018 and 2020 in a specified database for clinical research"). The system uses a large language model to parse the input and extract key information points, such as the disease domain ("type 2 diabetes"), time range ("2018 to 2020"), research subjects ("confirmed patients"), and exploration purpose ("clinical research").

[0086] It should be understood that the specific model structure of the large model mentioned in step S310 can be set according to actual needs, and the embodiments of this application are not limited thereto.

[0087] For example, such as Figure 4 As shown, Figure 4 A model structure diagram of a key information extraction model provided in an embodiment of this application is shown. For example... Figure 4 As shown, the key information extraction model includes a pre-trained embedding module for text vectorization representation, a context understanding module for establishing global semantic relationships, a sequence labeling layer for entity boundary recognition, a key information classification layer for information point existence detection, an information point aggregation module for cross-entity relationship modeling, a dynamic template matching module for structured information generation, and an information integration module. Specifically, the context understanding module is connected to the pre-trained embedding module and the sequence labeling layer, and the sequence labeling layer is also connected to the key information classification layer. The information point aggregation module is connected to the context understanding module, the sequence labeling layer, and the key information classification layer, respectively. The dynamic template matching module is connected to the context understanding module, and the information integration module is connected to the sequence labeling layer, the key information classification layer, the information point aggregation module, and the dynamic template matching module, respectively.

[0088] Furthermore, the pre-trained embedding module is used to segment the input text to obtain a token sequence, and can generate token embedding representations through a pre-trained BERT model, and can also embed the token representations into a bidirectional LSTM to obtain a context-aware token representation vector sequence;

[0089] This context understanding module is used to perform average pooling on the sequence of context-aware token representation vectors to obtain a global context vector, and can use a self-attention mechanism (or multi-head attention) to allow the token representation of each token to interact with the global context vector to obtain a globally context-enhanced token representation;

[0090] This sequence labeling layer is used to input the globally context-enhanced token representation into the CRF layer to predict the label of each token, so as to obtain the predicted label sequence;

[0091] This key information classification layer is used to input the global context vector into a multilayer perceptron (MLP) to obtain the existence probability of each key information point. Key information points can refer to information about relevant elements such as disease domain, time frame, research subject, and investigation purpose that need to be extracted from medical texts.

[0092] This information point aggregation module first extracts entities for each key information point type (disease, time range, research subject, investigation purpose) based on the predicted label sequence output by the sequence labeling layer, and collects the entity representation for each extracted entity. For example, for each entity, the representation vector of the entity is obtained by averaging the globally context-enhanced token representation vectors of all its contained tokens. Then, a node is constructed for each key information point type (disease, time range, research subject, investigation purpose), and the node feature is the product of the average value of all entities under that type and their corresponding existence probabilities (or a zero vector if there are no entities). For example, the disease node feature = disease existence probability × (average value of all disease entity vectors). Furthermore, based on the four nodes (disease node, time range node, research subject node, investigation purpose node), a fully connected graph (nodes include: disease, time range, research subject, investigation purpose) is constructed, and graph convolutional networks (GCN) or graph attention networks (GAT) can be used for message passing and node updates. For example, during message passing, nodes aggregate information from their neighbors. The updated node representation includes not only its own initial features but also information from related nodes in the graph (e.g., a disease node receives information from time nodes, study subject nodes, and investigation target nodes). Furthermore, the updated node representation will be used to adjust the representation of the sequence labeling layer (through an attention mechanism) or to correct classification results.

[0093] In other words, the information point aggregation module adjusts the representation of the sequence labeling layer (by adjusting its input representation through correction signals, thereby affecting the sequence labeling of the next round) and the output of the key information point classification layer (by correcting probability values). Thus, the information point aggregation module focuses on using graph networks for relational reasoning and error correction, and feeds the correction results back to the sequence labeling layer and the key information classification layer.

[0094] This dynamic template matching module defines a learnable query vector for each key information point type (e.g., a query vector q_time for "time range"). The similarity between the query vector and each globally context-enhanced token representation is then calculated. An attention weight for each similarity is obtained using softmax, and a weighted sum based on the similarity and attention weights yields the key information point representation. Finally, this key information point representation can be used for classification (presence / absence) or to directly generate text for the time range.

[0095] In other words, the dynamic template matching module focuses attention based on the original enhanced representation, avoiding the direct use of graph node representations that may contain correction noise, thus maintaining the purity of the extraction.

[0096] The information integration module integrates the predicted label sequence output by the sequence labeling layer, the existence probability of each key information point output by the key information classification layer, the relationship-enhanced node representation output by the information point aggregation module, and the key information point representation output by the dynamic template matching module. It also determines whether there are conflicts in the output information of each module. If there are conflicts, they are processed according to the preset conflict resolution rules, and the key information point extraction results are output after the conflict is resolved. If there are no conflicts, the key information point extraction results are directly output.

[0097] It should be understood that the specific rules of the preset conflict resolution rules can be set according to actual needs, and the embodiments of this application are not limited thereto.

[0098] For example, when there are two representations, "2018-20" and "2018-2020", the latter should be used.

[0099] Step S320: Verification and supplementation guidance of key information integrity.

[0100] Specifically, the system assesses the completeness and clarity of the identified key information. If the information is insufficient or ambiguous (e.g., the user only enters "I want to see diabetes data"), the system will proactively guide the user to supplement or clarify relevant details through multiple rounds of dialogue and clarifying questions. For example, the system might ask: "Do you mean type 1 diabetes or type 2 diabetes? Which time period are you interested in? What specific information about the patient do you want to know, such as age distribution, complication status, or treatment plan?" etc., until sufficiently clear information is obtained.

[0101] In step S330, the system asks the user whether they need AI to generate a specific exploration plan.

[0102] If the user needs AI to generate a specific exploration plan, then proceed to step S340; if the user does not need AI to generate a specific exploration plan, then proceed to step S350.

[0103] In step S340, the system will combine the key information points provided by the user and deeply invoke the knowledge fusion module. This knowledge fusion module can integrate multi-source information such as medical knowledge graphs, database structure information, exploration plan templates, research plan libraries, and medical coding mapping tables.

[0104] Specifically, the system will refer to commonly used analytical dimensions for similar research objectives in the exploration protocol template, combine disease association information from the medical knowledge graph (such as common complications of diabetes and related examination indicators), and existing high-quality research designs in the research protocol library to intelligently recommend one or more structured exploration protocols. Furthermore, the protocol may include suggested population characteristics, key indicators, grouping methods, etc. Users can modify and confirm the recommended protocols.

[0105] Step S350: Structured requirements construction and iterative confirmation.

[0106] Specifically, whether generated by AI or input by the user, the system transforms the exploration plan into structured exploration requirements. During this process, for the exploration indicators proposed by the user or recommended by AI, the system automatically recommends the calculation logic and industry-recognized judgment standards for each indicator (e.g., normal ranges of blood lipids for different age groups, standards for TNM staging of tumors, etc.) based on medical guidelines, literature, and historical exploration experience in the knowledge fusion module. It also converts natural language terms (e.g., "type 2 diabetes") into actual codes recognizable by the database (e.g., "E11") using the knowledge fusion module. Subsequently, the system performs the following checks using a large language model:

[0107] Reasonableness of the purpose and background of the investigation: Check whether the user is clear about the application scenario and follow-up plan of the investigation results, and whether the investigation object and investigation method are reasonable;

[0108] Reasonableness and consistency of requirement logic: For example, check whether the time range is reasonable and whether there are contradictions between the screening conditions;

[0109] Data source alignment and feasibility verification: Confirm that all data elements involved in the exploration requirements (such as specific examination items and drug names) can be found in the corresponding data tables and fields in the database structure information of the knowledge fusion module, and assess the feasibility of data acquisition;

[0110] Completeness and clarity of current structured requirements: Determine whether all necessary information for effective exploration has been included;

[0111] Based on the inspection results, the system will take the following measures:

[0112] If improvements are needed, the system will display the current exploration requirements to the user and ask targeted questions to guide the user in making improvements. The user's answers will be used to update the exploration requirements and may trigger a return to the information transformation and solution adjustment in step S340 above.

[0113] If the check passes, the system will display the final structured exploration requirements to the user and request the user's confirmation.

[0114] Step S360: Requirement confirmation and workflow.

[0115] Specifically, if the user confirms the final requirements, the structured exploration requirements will be sent to the subsequent review and execution process;

[0116] If the user has suggestions for modification to the final requirements, the system will wait for the user to input the modification content, and may return to step S340 or step S350 for corresponding adjustments depending on the modification.

[0117] Furthermore, after the requirements understanding module completes its execution, the compliance review module automatically reviews and controls the risks of the resulting structured investigation tasks. The core processes include: refined rule matching and multi-dimensional preliminary assessment, weighted risk comprehensive judgment and dynamic threshold adjustment, intelligent differentiated processing and guidance strategies, and transparent audit and user appeal mechanisms.

[0118] Specifically, for the aforementioned refined rule matching and multi-dimensional preliminary assessment, the system performs deep matching between structured investigation requirements and a dynamically updated, more comprehensive rule base. This rule base not only contains preset rules but can also learn and evolve from historical audit cases through machine learning. Specifically, the rule base mainly includes at least one of the following categories: core data protection rules, sensitive information identification and protection rules, identity recognition and re-identification risk control rules, compliance and authorization management rules, system security and performance protection rules, and statistical quality and bias control rules.

[0119] Furthermore, the core rules for data protection include the data minimization principle and the usage restriction rule. The data minimization principle means assessing whether the requested data fields exceed the minimum scope required for the research purpose, and conducting strict review based on business necessity and proportionality principles. The usage restriction rule means ensuring that data use is strictly limited to the research purpose applied for, preventing out-of-scope use or secondary exploitation.

[0120] Furthermore, the rules for identifying and protecting sensitive information include at least one of the following: rules for exposure in medically sensitive areas, rules for the protection of biometric information, and rules for special protection of vulnerable groups. Specifically, rules for exposure in medically sensitive areas refer to establishing a tiered protection mechanism for specific disease categories (mental illnesses, infectious diseases, genetic diseases, rare diseases), special drugs (narcotic drugs, psychotropic drugs, experimental drugs), and high-risk medical procedures; rules for the protection of biometric information refer to covering unique and unalterable biometric identifiers such as genetic information, biological sample data, and imaging characteristics; and rules for special protection of vulnerable groups refer to establishing stricter protection standards for special groups such as minors, the elderly, patients with mental disorders, and critically ill patients.

[0121] Furthermore, the risk prevention and control rules for identity recognition and re-identification include at least one of the following: direct identifier exposure rules, quasi-identifier combination risk rules, spatiotemporal trajectory association risk rules, and social relationship network exposure rules. Specifically, the direct identifier exposure rule refers to detecting the direct query risk of strong identifiers such as complete ID card numbers, medical insurance card numbers, hospitalization numbers, contact information, and detailed addresses; the quasi-identifier combination risk rule refers to assessing the query risk of combinations of quasi-identifiers such as age, gender, region, occupation, and disease, with particular attention to the re-identification possibility of small sample groups; the spatiotemporal trajectory association risk rule refers to identifying the exposure risk of individual behavioral trajectories that may be formed based on information such as visit time, location, and department; and the social relationship network exposure rule refers to detecting query patterns that may indirectly identify individuals through family members, colleagues, community relationships, etc.

[0122] Furthermore, compliance and authorization control rules include at least one of the following: tiered authorization verification rules, ethical review compliance rules, and legal and regulatory compliance rules. Tiered authorization verification rules involve fine-grained matching based on data sensitivity levels and user permission levels to achieve field-level access control; ethical review compliance rules refer to compliance verification based on the scope of approval by the institution's ethics committee, research agreement constraints, and informed consent restrictions; and legal and regulatory compliance rules refer to dynamically tracking the latest requirements of relevant regulations.

[0123] Furthermore, the system security and performance protection rules include at least one of the following: query complexity and resource consumption rules, and abnormal behavior detection rules. Query complexity and resource consumption rules assess the computational complexity of queries, data volume, and concurrency impact to prevent system overload and resource abuse. Abnormal behavior detection rules identify suspicious behavior patterns such as batch queries, frequent queries, and queries during abnormal time periods.

[0124] Furthermore, statistical quality and bias control rules include at least one of the following: sample representativeness assessment rules, statistical inference validity rules, and data quality assurance rules. Specifically, sample representativeness assessment rules examine whether query conditions may lead to sample selection bias and affect the representativeness of research results; statistical inference validity rules assess whether the query design conforms to statistical principles, avoiding inappropriate causal inferences or association analyses; and data quality assurance rules identify query conditions that may affect data integrity, accuracy, and consistency.

[0125] Some of these procedures are broadly applicable to all business scenarios, while others are only applicable to specific scenarios. For different exploration needs, the system selects a relevant subset of rules from the complete rule base and performs subsequent matching and risk assessment only on this subset.

[0126] Furthermore, for each matching rule, the large language model (combined with the knowledge fusion module) will play a key role, specifically:

[0127] Semantic understanding and contextual analysis: gaining a deeper understanding of the true intent behind the exploration needs, rather than just keyword matching, such as identifying users' attempts to obtain sensitive information through roundabout methods;

[0128] Risk level quantification: Combining the severity of the rule, the sensitivity of the data itself, and the specific context of the exploration needs, a more refined risk score (e.g., 0-1 points) is given for each matching rule, and a natural language explanation of the scoring basis is provided.

[0129] Potential risk discovery: Based on extensive medical and privacy protection knowledge, identify emerging risk patterns and attack methods that are not yet covered in the rule base.

[0130] Furthermore, for the aforementioned weighted risk assessment and dynamic threshold adjustment, the system references the scores of all rules (S_i, whose value ranges from [0,1]) and comprehensively considers the following factors:

[0131] The base weight (W_i) for rule settings: A base weight coefficient is set based on the sensitivity of the rule itself, with a value range of [0.2, 1].

[0132] Business Scenario Weight (B): Distinguishes the risk tolerance of different application scenarios such as scientific research, clinical decision support, and public health monitoring. The value range is [0.7, 1.2], with market insight < scientific research < clinical < public health.

[0133] User trust adjustment weight (U): dynamically adjusted based on the user's historical compliance record and the institution's reputation level, with a value range of [0.8, 1.2]. Low values ​​are assigned to high-trust users, and high values ​​are assigned to low-trust users.

[0134] External environment adjustment weight (E): Combines external factors such as special periods (e.g., public health emergencies), changes in regulatory policies, social attention, and media sensitivity, with a value range of [0,1]. Changes in the external environment are positive.

[0135] Total number of triggering rules (N): The total number of triggering rules is taken into consideration as one of the factors in calculating the risk of multiple rules superimposed, and is used to measure the additional risk that the rule superposition effect may bring.

[0136] Based on the above, the overall risk score is calculated using the formula:

[0137] Basic risk=Σ(S_i×W_i) / Σ(W_i)

[0138] The rule superposition factor = 1 + 0.2 × ln(1 + N)

[0139] Overall risk score = Basic risk × B × U × Rule superposition factor + Basic risk × E × 0.5.

[0140] After obtaining the overall risk score, the system will classify the risk levels based on dynamic risk thresholds. These thresholds are not fixed but can be adjusted according to the institution's risk tolerance, updates to laws and regulations, and new risk patterns learned by the system.

[0141] Furthermore, risk levels can be further refined, such as: extremely low risk, low risk, low-to-medium risk, medium risk, medium-to-high risk, high risk, extremely high risk, etc.

[0142] Furthermore, the aforementioned weighted risk assessment and dynamic threshold adjustment include:

[0143] Very low / low risk: Fast track, which may only require automated logging, and users will be unaware of it or only receive slight prompts;

[0144] Low to Medium Risk: The system not only alerts users to risks but also uses a large language model to generate specific and actionable modification suggestions, such as: "The age range you queried is too precise, which may lead to individual identification; we suggest broadening it to the xx age range" or "We suggest aggregating the xx field to reduce risk." Users can choose to accept the suggestions and modify them with one click, or adjust them manually.

[0145] Medium-high / high risk: In addition to explicit risk warnings and modification suggestions, this may trigger a second approval process, requiring designated personnel to review the application or supplementary medical ethics review materials. The system will provide the approver with a complete risk assessment report and a risk interpretation generated by a large language model;

[0146] Extremely high risk: Execution is prohibited in principle. The system will explain the reasons in detail and log the attempt. In certain strictly regulated scenarios, such requests may trigger security alerts.

[0147] Furthermore, the aforementioned transparent auditing and user complaint mechanism includes:

[0148] Provide detailed and traceable audit logs to record the submission of each investigation request, the review process, risk scoring, decision-making basis, and final processing results;

[0149] Establish a user appeal channel. If users disagree with the review results or system suggestions, they can submit an appeal. The appeal processing process and results will be recorded and used to improve the review model.

[0150] Furthermore, after the compliance review module completes its execution, the code generation module can use text2SQL to generate query code based on the approved exploration task requirements. Specifically:

[0151] Choose an appropriate query language (such as MySQL, Postgres query language, etc.) based on the type and structure of the target database;

[0152] By combining a data dictionary and an encoding mapping table, medical terms are converted into corresponding database fields and codes;

[0153] Construct efficient query statements, including appropriate table joins, conditional filtering, and aggregate functions;

[0154] Optimize the performance of the generated query code, such as by adding appropriate index hints and splitting complex queries.

[0155] Furthermore, after the code generation module finishes execution, the generated query code is safely executed through the query execution module. Specifically:

[0156] Run the query code in an isolated, secure execution environment;

[0157] Real-time monitoring and querying of execution status, including resource usage and execution progress;

[0158] Large query tasks are processed in segments, and intermediate results are cached.

[0159] Automatically handle exceptions during query execution.

[0160] Furthermore, after the query execution module completes its execution, the query result analysis module performs in-depth analysis and preprocessing of the query result data to prepare for report generation. Specifically:

[0161] Blurring of small sample data ensures data privacy.

[0162] Based on the grouping dimensions and statistical indicators, perform descriptive statistical analysis (such as calculating the mean, median, frequency, proportion, etc.);

[0163] Differential privacy technology can be used to add a suitable amount of noise to the statistical results (optional);

[0164] Based on the data quality exploration requirements, the completeness (missing rate), accuracy, and consistency of core data items are quantitatively evaluated.

[0165] Perform statistical and in-depth analysis on the query results to identify key patterns and anomalies.

[0166] Furthermore, after the query results analysis module completes its execution, the report generation module automatically generates an exploration report based on the multi-dimensional analysis results. Specifically:

[0167] Automatically select the appropriate visualization method based on the characteristics of the data;

[0168] Automatically generate charts, tables, and text to present data clearly and accurately;

[0169] Generate a hierarchical report structure based on the report template, including an abstract, detailed analysis, and technical appendices;

[0170] Highlighting key findings and providing expert explanations;

[0171] Utilize large language models for content polishing and refinement to ensure the report's professionalism and readability.

[0172] Furthermore, after the report generation module completes its execution, the compliance review module conducts a compliance and security review of the generated report content. Specifically:

[0173] Reports with extremely low risk or very low risk are published directly after review.

[0174] Reports of low to medium risk and medium risk are published after modifying or partially masking some of the report content using a large language model.

[0175] Reports classified as medium-high risk, high risk, or extremely high risk need to be revised in terms of data exploration requirements and resubmitted.

[0176] Furthermore, after the compliance review module conducts a compliance and security review of the generated report content, the results presentation module presents the investigation report through a web interface, specifically:

[0177] It provides interactive report display and supports features such as dynamic filtering;

[0178] Collect user feedback on the investigation results to support iterative improvement of the report;

[0179] Allow users to submit simple revision requests, such as adjusting specific statistical definitions or display methods;

[0180] Supports exporting reports in multiple formats (PDF, Excel, HTML, etc.).

[0181] Therefore, by leveraging the above technical solutions, this application embodiment integrates large language model technology with medical expertise to construct an intelligent and automated medical data exploration platform. This platform can understand exploration needs expressed in natural language, automatically conduct compliance reviews, generate efficient query codes, perform data analysis, and generate professional reports, while ensuring data security and privacy protection. This solves the problems existing in current medical data exploration processes, such as data heterogeneity, difficulty in quality assessment, low exploration efficiency, high professional threshold, and insufficient privacy protection.

[0182] Furthermore, this application also possesses the following technical advantages:

[0183] Lowering the barrier to entry: Through a natural language interactive interface, even medical workers without professional data analysis skills can efficiently complete data exploration tasks;

[0184] Improve exploration efficiency: Automated code and report generation processes reduce exploration tasks that traditionally take several days to just hours or less;

[0185] Enhanced data quality assessment: The system's built-in data quality assessment function can comprehensively evaluate the completeness, accuracy, and consistency of data, helping users quickly determine whether the data meets research needs;

[0186] Strengthen privacy protection: A multi-layered security audit mechanism ensures that the investigation process and results meet privacy protection requirements and reduces the risk of data leakage;

[0187] Enhanced depth of investigation: By combining intelligent analysis capabilities with medical expertise, it can discover data patterns and correlations that are difficult to identify using traditional methods;

[0188] Support for decision optimization: The generated professional reports provide data-driven references for research design, ensuring that research questions match available data.

[0189] It should be understood that the above-mentioned medical and health big data exploration system is merely exemplary, and those skilled in the art can make various modifications based on the above system, and the modified solutions also fall within the protection scope of this application.

[0190] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0191] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions.

[0192] It should be noted that any reference numerals placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In claims that enumerate several means, several of these means may be embodied by the same hardware. The use of the terms first, second, third, etc., is merely for convenience of expression and does not indicate any order. These terms can be understood as part of the component names.

[0193] Furthermore, it should be noted that in the description of this specification, the terms "one embodiment," "some embodiments," "embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0194] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the claims should be interpreted to include both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0195] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, then this invention should also include these modifications and variations.

Claims

1. A medical and health big data exploration system, characterized in that, include: The demand understanding module is used to obtain the data exploration demand input by the user through natural language, and to call the relevant medical guidelines, medical literature and historical experience database to automatically recommend the calculation logic and judgment criteria of various indicators, and to transform the data exploration demand into a structured exploration task in real time. The compliance review module is used to conduct compliance reviews of the exploration tasks according to a multi-level compliance review mechanism, and to identify and filter high-risk exploration requests; The code generation module is used to generate query code based on the exploration task and the characteristics of the target database, provided that the compliance review has been passed. The query execution module is used to query the query result data corresponding to the query code from the target database according to the query code; The query result analysis module is used to perform multi-dimensional analysis on the query result data to obtain multi-dimensional analysis results; The report generation module is used to generate a hierarchical exploration report based on the multi-dimensional analysis results. The compliance review module is also used to conduct compliance and security reviews of the investigation report to ensure that the content is compliant and does not contain sensitive information; The results presentation module is used to present the investigation report in an interactive manner in the user interface, provided that compliance and security reviews have been passed, and to support iterative improvement of the report by providing a feedback mechanism.

2. The medical and health big data exploration system according to claim 1, characterized in that, The data exploration requirements include at least one of the following: exploration purpose and background, definition of exploration object, statistical indicators and grouping dimensions, data quality exploration requirements, and definition and calculation logic of exploration indicators.

3. The medical and health big data exploration system according to claim 2, characterized in that, The definition of the target of investigation includes at least one of the following: the subject of investigation, the type of disease, the time range of the target, demographic characteristics, clinical characteristics, treatment-related characteristics, and other screening criteria.

4. The medical and health big data exploration system according to claim 2, characterized in that, The statistical indicators in the statistical indicators and grouping dimensions include frequency statistics and central tendency statistics; wherein, the frequency statistics cover the total number and proportion; the central tendency statistics include the mean and median.

5. The medical and health big data exploration system according to claim 4, characterized in that, The statistical indicators and grouping dimensions include at least one of the following: time dimension grouping, spatial dimension grouping, and clinical dimension grouping.

6. The medical and health big data exploration system according to claim 1, characterized in that, The compliance review module is specifically used to perform deep matching between the exploration task and the rule base to identify and filter high-risk exploration requests; wherein, the rule base includes at least one of the following: core data protection rules, sensitive information identification and protection rules, identity recognition and re-identification risk prevention and control rules, compliance and authorization control rules, system security and performance protection rules, and statistical quality and bias control rules.

7. The medical and health big data exploration system according to claim 6, characterized in that, The rules for identifying and protecting sensitive information include at least one of the following: rules for exposure to sensitive medical fields, rules for protecting biometric information, and rules for special protection of vulnerable groups.

8. The medical and health big data exploration system according to claim 6, characterized in that, The identity recognition and re-identification risk prevention and control rules include at least one of the following: direct identifier exposure rules, quasi-identifier combination risk rules, spatiotemporal trajectory association risk rules, and social relationship network exposure rules.

9. The medical and health big data exploration system according to claim 6, characterized in that, The compliance and authorization control rules include at least one of the following: hierarchical authorization verification rules, ethical review compliance rules, and legal and regulatory adaptation rules.

10. The medical and health big data exploration system according to claim 6, characterized in that, The system security and performance protection rules include at least one of the following: query complexity and resource consumption rules and abnormal behavior detection rules.

Citation Information

Cited By

  • Online medical intelligent medical guide system based on text graph embedding

    CN121483569A

  • Medical scientific research data space-oriented data discovery and privacy calculation cooperation method, equipment, medium and product

    CN122065347A

  • Method, device, medium and product for data discovery and privacy computing oriented to medical scientific research data space

    CN122065347B