Intelligent medical project value evaluation method based on medical health big data

By constructing a structured graph structure and causal driving factors based on the BERT-BiLSTM-CRF model and an improved density clustering algorithm, the scientificity and accuracy of existing medical project value assessment methods are solved, and refined evaluation and resource allocation of medical projects in different patient subgroups are achieved.

CN120708912AInactive Publication Date: 2025-09-26JINAN BEISEN ELECTRONIC INFORMATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510903857.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-09-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing medical project value assessment methods are deficient in scientificity, accuracy, and interpretability, and are unable to meet the needs of refined clinical management. In particular, when faced with complex multi-project interactions or project reconfiguration needs during disease evolution, there is a lack of in-depth modeling of patient characteristics, etiological drivers, and diagnosis and treatment complexity.

Method used

The BERT-BiLSTM-CRF joint model is used to transform unstructured medical record data into a structured graph structure. Patient feature vectors are constructed through graph embedding representation. Combined with an improved density clustering algorithm, multi-scale clustering is performed to construct etiology driving factors and evaluate the value of medical projects in different patient subgroups.

Benefits of technology

It achieves more accurate patient clustering and medical project value assessment, breaks through the bottleneck of traditional methods in modeling continuity of patient subgroups, provides a stable and explainable group basis, and provides a more forward-looking decision-making basis for medical insurance payment optimization and medical resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708912A_ABST
    Figure CN120708912A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of value assessment, and discloses an intelligent medical project value assessment method based on medical health big data, and the method comprises the steps: carrying out the graph embedding representation of a patient node in a structured medical record graph structure, and obtaining a patient feature vector with graph semantics; an improved density clustering algorithm is adopted to carry out multi-scale grouping on the patient feature vectors; and performing medical complexity modeling and medical income modeling on the medical project to obtain the medical complexity and medical income of the medical project in the corresponding patient subgroup, and evaluating the comprehensive value of the medical project. According to the method, unstructured medical record data is converted into a structured graph structure, a patient feature vector embedded in the graph is extracted, patient subgroup identification is realized by adopting an improved density clustering method in combination with a disease cause driving factor, complexity and income dual modeling is performed on diagnosis and treatment items of different patient subgroups, the comprehensive value of each medical item is evaluated, and the diagnosis and treatment efficiency is improved. And the medical resource configuration efficiency and the project precision evaluation level are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of value assessment, and in particular to a method for value assessment of smart medical projects based on medical and health big data. Background Art

[0002] In recent years, with the continuous integration of big data, artificial intelligence, and medical information systems, intelligent medical project management and value assessment have become important directions for promoting the optimal allocation of medical resources, medical insurance payment decisions, and clinical pathway management. Medical services refer to the health-promoting services provided by health professionals in accordance with professional standards to protect lives, diagnose and treat diseases, and the pharmaceuticals, medical devices, transportation, and hospital accommodations provided to support these services. This concept is widely used in daily life, as well as in policies, regulations, and national development strategies. A clear definition of medical services is necessary to standardize the services provided by medical institutions, define the scope of medical services and various types of life and health insurance, manage the doctor-patient relationship, and develop social health programs. Traditionally, medical services primarily provide diagnosis, treatment, prevention, and rehabilitation within hospitals, encompassing the entire process from patient consultation and establishment of a service relationship with the hospital to the completion of treatment, discharge, or death. While strengthening in-hospital medical services, modern medical services also emphasize extending social medical services beyond hospitals, including post-discharge follow-up, home care beds, public health education, disease screenings, social medical assistance and targeted support, and medical outreach to rural areas.

[0003] Across the spectrum of healthcare services, the resource consumption and clinical benefits of medical programs (such as examinations, treatments, and medications) exhibit significant heterogeneity. This heterogeneity is not only related to the attributes of the programs themselves, but also closely linked to multiple factors, including the patient's etiology, the complexity of the diagnosis and treatment pathway, and the presence of complications. Therefore, quantifying the value of medical programs across diverse patient contexts at the population level and establishing an interpretable and quantifiable program evaluation mechanism are key challenges in intelligent medical evaluation systems.

[0004] Current research primarily focuses on evaluating the average cost-effectiveness of programs based on static statistical indicators such as program frequency, cost, and medical insurance utilization, or simple regression models. While these approaches are easy to implement, they suffer from a significant "group average assumption" bias, ignoring patient variability in treatment needs, complexity, and response, making them inadequate for meeting the needs of refined clinical management. In existing research, CN120013316A provides a method, system, device, and media for analyzing medical service quality. These methods involve acquiring medical service evaluation data from various medical institutions, filtering variable data from the data, performing integrity processing and standardization on the variable data to ensure that the processed variable data meets integrity and standardization requirements, performing principal component analysis on the processed variable data to generate analysis results, and interpreting and analyzing the analysis results using a large AI model to provide corresponding optimization decision recommendations. This approach still focuses on path rule mining or service evaluation data mining, lacks in-depth modeling of patient characteristics, etiological drivers, and treatment complexity, and struggles to effectively respond to complex multi-program interactions or program reconfiguration needs during disease progression.

[0005] Therefore, the existing medical project value assessment framework has much room for improvement in terms of scientificity, accuracy and explainability. There is an urgent need for a targeted, clustered and multi-dimensional value assessment method for medical projects to support the implementation of key decision-making scenarios such as medical insurance payment optimization, clinical pathway standardization and precise allocation of medical resources. Summary of the Invention

[0006] In view of this, the present invention provides a smart medical project value assessment method based on medical health big data, which converts unstructured medical record data into a structured graph structure, wherein the structured graph structure is composed of patient nodes and medical entity nodes, and the medical entity types in the medical entity nodes include symptoms, previous diseases, examination indicator types, and diagnosis and treatment information. Based on the perception weight between the patient node and the medical entity node, the semantic association degree between the patient node and the medical entity node is measured, and in the embedding fusion process, a graph embedding representation based on semantic perception + structural enhancement is constructed, and medical entity nodes with larger perception weights are selected for fusion processing of semantic vector representation. The fusion processing result is used as the patient feature vector of the patient node, instead of extracting the patient medical record information from the patient node, and the patient medical record information is directly used as the patient feature vector, so that the patient node has more comprehensive feature semantics, realizing the process of constructing a graph embedding-oriented patient feature representation based on structured medical record data, forming a patient feature vector with the global semantics of all medical entity nodes and node relationships, and improving the accuracy and effectiveness of subsequent patient clustering;

[0007] During the clustering algorithm process, the neighborhood radius and distance representing density are optimized, and the mutation points are identified as boundary candidates in combination with the co-occurrence frequency of symptoms in the etiological driving factors. Differential-driven boundary screening is achieved based on the dynamic boundary threshold to improve boundary positioning accuracy and medical consistency. This method not only solves the problems of traditional density clustering algorithms being sensitive to global density and fuzzy boundaries, but also breaks through the bottleneck of existing methods in modeling the continuity of patient subgroups, providing a more stable and interpretable group basis for subsequent medical project modeling, and then conducting targeted medical benefit modeling and medical complexity modeling in different patient subgroups. From the comprehensive evaluation of the medical benefits and complexity of medical projects, the value of medical projects in patient subgroups is measured in a targeted manner, emphasizing basic stable effectiveness, and paying attention to the peak potential and stability of the benefits of medical projects in patient subgroups, to discover high-yield but risk-controlled medical projects.

[0008] To achieve the above objectives, the present invention provides a method for evaluating the value of smart medical projects based on medical and health big data, comprising the following steps:

[0009] S1: Obtain unstructured medical record data, use the BERT-BiLSTM-CRF joint model to extract structured semantic information, and construct the structured semantic information into structured medical record data;

[0010] S2: Map the structured medical record data into a graph structure to obtain a structured medical record graph structure. Perform graph embedding representation of the patient nodes in the structured medical record graph structure combined with multivariate relationships to obtain a patient feature vector with graph semantics.

[0011] S3: Construct etiological driving factors and use an improved density clustering algorithm to perform multi-scale clustering on patient feature vectors to obtain patients with similar etiological symptoms and form patient subgroups;

[0012] S4: Extract diagnosis and treatment information of patient subgroups from structured medical record data, and perform medical complexity modeling and medical benefit modeling on the medical items in the diagnosis and treatment information;

[0013] S5: Evaluate the comprehensive value of medical programs based on the medical complexity and medical benefits of medical programs in different patient subgroups.

[0014] As a further improvement method of the present invention:

[0015] Optionally, the unstructured medical record data is the patient's medical record text data, and the BERT-BiLSTM-CRF joint model is used to extract structured semantic information from the unstructured medical record data, including:

[0016] Use the BERT language model to contextually encode unstructured medical record data and extract semantic vector representations at the word and sentence level;

[0017] The semantic vector representation is input into the BiLSTM network to capture the dependencies between medical terms. The resulting semantic vector representation and the labels of the words and phrases associated with the semantic vector representation are obtained. The label types include patient ID, symptoms, previous diseases, examination index results, diagnosis and treatment information, timestamp information, and other types.

[0018] The CRF layer is used to jointly decode and optimize the labels of all semantic vector representations to obtain the final labels of each semantic vector representation and the associated words and sentences. The semantic vector representations and the associated words and sentences labeled with patient ID, symptoms, previous diseases, examination index results, diagnosis and treatment information, and timestamp information are used as structured medical record data. The diagnosis and treatment information includes diagnosis results and medical items. The medical items are the patient's treatment items, consisting of examination, treatment, and medication items.

[0019] Specifically, the medical items are composed of one or more item categories of examination, treatment, and medication items.

[0020] Compared with traditional medical record structuring methods based on rules or shallow models, this invention adopts a pre-trained language model combined with sequence modeling and structured decoding, which can effectively deal with the problems of semantic ambiguity, synonymous expressions and long text dependencies in medical texts. It uses the BiLSTM network to further enhance the sequence modeling capability. The extracted structured information has clear semantic boundaries and can be directly used in graph structure construction, patient representation modeling, and subsequent clustering and evaluation steps, realizing the key transformation from original medical record text to quantitative analysis.

[0021] Optionally, extract the patient ID and patient medical record information from the structured medical record data as a patient node, the patient medical record information is a concatenation result of semantic vector representations of symptoms, past diseases, examination index results, and diagnosis and treatment information associated with the patient ID, extract the medical entities and semantic vector representations of the medical entities from the structured medical record data to form a medical entity node, and the types of the medical entities include symptoms, past diseases, examination index types, and diagnosis and treatment information;

[0022] Constructing patient nodes and medical entity nodes into a structured medical record graph structure, wherein the structured medical record graph structure is composed of nodes and relationships between nodes, wherein the nodes include patient nodes and medical entity nodes, and the relationships between the nodes are the perception weights between the patient nodes and the medical entity nodes;

[0023] The perception weight calculation formula is:

[0024] ;

[0025] ;

[0026] ;

[0027] ;

[0028] in, Represents the patient node c and the medical entity node The perceptual weight between represents the node association coefficient, Represents the medical numerical relationship between the patient node c and the medical entity node node, represents the time decay coefficient, represents the coefficient weight, , and set They are 0.6, 0.2, and 0.2 respectively;

[0029] The higher the perception weight, the higher the semantic association between the patient node and the medical entity node, which is manifested in that the patient node has the symptoms or previous diseases or examination indicator types or diagnosis and treatment information corresponding to the medical entity node. The perception weight is composed of the node association coefficient, the medical numerical relationship and the time decay coefficient.

[0030] The node association coefficient is composed of the frequency of occurrence of the medical entity node in the patient node and the semantic similarity between the medical entity node and the patient node, which represents the importance of the medical entity node to the patient node;

[0031] The medical numerical relationship is the numerical relationship between the patient node and the medical entity node. The larger the medical numerical relationship, the more obvious abnormality the patient node has in the corresponding medical entity node, which is manifested by more severe symptoms or previous diseases, and more abnormal examination indicator results.

[0032] The time decay coefficient is the timestamp decay result of the medical entity in the medical entity node appearing in the patient medical record information of the patient node. The time decay coefficient is used to give a time weight penalty to obsolete medical entities to achieve dynamic modeling and avoid excessive influence of obsolete medical entities on the construction process of the patient feature vector. If the medical entity in the medical entity node does not appear in the patient medical record information of the patient node, the time decay coefficient between the patient node and the medical entity node is set to 0;

[0033] Specifically, the normalized representation of numerical indicators such as the intensity of the symptoms, the intensity of the previous diseases, and the deviation value between the examination index results and the normal results is the medical numerical relationship between the patient node and the medical entity node. The medical numerical relationship between the patient node and the diagnosis and treatment information is 1. If the patient node does not have the symptoms, previous diseases, or examination index types associated with the medical entity node, the medical numerical relationship between the patient node and the medical entity node is 0.

[0034] represents the cosine similarity between the semantic vector representations in the patient node c and the medical entity node node, Indicates the statistical importance of the medical entity node node in the patient node c, represents the frequency of occurrence of the medical entity associated with the medical entity node node in the structured medical record data of the patient associated with the patient node c. N represents the total number of patients. Indicates the number of patients who have a relationship with the medical entity node. represents the cosine similarity weight, Indicates the statistical importance weight, set is 0.6, is 0.4;

[0035] Specifically, in the calculation process of node association coefficient, semantic similarity and frequency are integrated to reflect the structural contribution of medical entities to patients. By combining contextual semantics and statistical strength, high-frequency and low-correlation problems (such as commonly found symptoms but not related to the current patient) are effectively handled, the weight of specific medical entities is increased, and common entities (such as "fever") are suppressed.

[0036] represents the attenuation control coefficient, Indicates the current timestamp, Indicates the timestamp of the medical entity in medical entity node node appearing in the patient medical record information of patient node c; the attenuation control coefficient is set to 0.2.

[0037] Optionally, a graph embedding representation is performed on the patient nodes in the structured medical record graph structure in combination with multivariate relationships to obtain a patient feature vector with graph semantics, including:

[0038] Based on the perception weight between the patient node and the medical entity node, the 10 medical entity nodes with the highest perception weight with the patient node are selected as the medical entity node set of the patient node, and the medical entity node set of the patient node is embedded and fused to obtain the patient feature vector of the patient node.

[0039] Specifically, the embedding fusion formula of the medical entity node set of the patient node c is:

[0040] ;

[0041] ;

[0042] in, represents the patient feature vector of patient node c, W represents the weight matrix, represents the set of medical entity nodes of patient node c, , Represents a collection of medical entity nodes Any medical entity node in Represents a medical entity node The semantic vector representation in Represents a set All weighted semantic vector representations in are aggregated. Represents the activation function, the selected activation function is the tanh function;

[0043] Represent patient node c and medical entity node respectively The node correlation coefficient, medical numerical relationship and time decay coefficient between them;

[0044] In the embedding fusion process, a graph embedding representation based on semantic perception + structural enhancement is constructed, and medical entity nodes with larger perception weights are selected for fusion processing of semantic vector representation. The fusion processing results are used as the patient feature vector of the patient node, instead of extracting the patient medical record information from the patient node. The patient medical record information is directly used as the patient feature vector, so that the patient node has more comprehensive feature semantics. This realizes the process of constructing a graph embedding-oriented patient feature representation based on structured medical record data, forming a patient feature vector with the global semantics of all medical entity nodes and node relationships, thereby improving the accuracy and effectiveness of subsequent patient clustering.

[0045] Optionally, based on structured medical record data, construct etiology drivers, including:

[0046] Extracting a symptom set of each patient from the structured medical record data and initially constructing a symptom matrix, wherein the symptom matrix is ​​in the form of a matrix with M rows and M columns, where M represents the total number of symptoms of all patients;

[0047] For any two different symptoms, the number of symptom sets in which these two symptoms appear is counted as the co-occurrence frequency of the two symptoms, and the co-occurrence frequency is mapped to the corresponding position of the symptom matrix. The matrix element in the mth row and eth column in the symptom matrix represents the co-occurrence frequency of the mth symptom and the eth symptom. , the symptom matrix after all matrix elements are mapped and filled is used as the cause driving factor.

[0048] Optionally, in combination with the etiological driving factors, an improved density clustering algorithm is used to perform multi-scale clustering on the patient feature vectors, including:

[0049] Obtain the patient feature vectors of all patient nodes, calculate the Mahalanobis distance between any two patient feature vectors, use the Mahalanobis distance to determine the adaptive neighborhood radius of the patient feature vector, and extract the patient feature vectors whose Mahalanobis distance is smaller than the adaptive neighborhood radius as neighbor vectors;

[0050] Specifically, the patient feature vector The adaptive neighborhood radius calculation formula is:

[0051] ;

[0052] in, Represents the patient feature vector The adaptive neighborhood radius of represents the initial neighborhood radius, represents the mean Mahalanobis distance between all patient feature vectors, represents the distance control factor, Represents the patient feature vector With the initial neighborhood radius The standard deviation of the distance between other patient feature vectors within the range, setting the distance control factor is 0.1;

[0053] Specifically, the adaptive neighborhood radius calculation formula is used to dynamically adjust the local neighborhood radius of each patient's feature vector, effectively adapting to local density differences and improving clustering stability. Density perception is combined with Mahalanobis distance, and the adaptive neighborhood radius is used as a normalization factor to make the distance measurement result adaptable to local density. If the patient's feature vector is located in a dense area, the adaptive neighborhood radius is small and the Mahalanobis distance is amplified, emphasizing the difference between the patient and the sparse area. Otherwise, the Mahalanobis distance is compressed to avoid the misclustering of marginal patients.

[0054] Based on the Mahalanobis distance between the neighbor vector and the patient feature vector, the core distance of the patient feature vector is generated, the core distance is used to generate the improved reachable distance between the patient feature vectors, and the patient feature vectors are sorted based on the improved reachable distance to construct an improved reachable distance curve;

[0055] Specifically, the patient feature vector The core distance is the patient feature vector The distance between the patient and the MinPtsth nearest feature vector; MinPts is a preset neighbor vector ranking, and MinPts is set to 5;

[0056] The reachable distance calculation method is improved based on the adaptive neighborhood radius, and the patient feature vector To patient feature vector The improved reachable distance is:

[0057] ;

[0058] in, Represents the patient feature vector To patient feature vector Improved reach distance, Represents the patient feature vector The core distance, Represents the patient feature vector and patient feature vector The Mahalanobis distance between

[0059] Specifically, the reachable distance is the patient feature vector Can the patient feature vector be used? The density reaches, if is smaller, indicating that the patient's feature vector and In areas of the same density, an improved reachable distance is defined, and an adaptive neighborhood radius is introduced to normalize the Mahalanobis distance. In areas with higher density, the neighborhood radius is smaller, and the normalized distance is larger. In areas with lower density, the neighborhood radius is larger, and the normalized distance is smaller. This helps distinguish high-density from low-density areas, enhances the robustness of boundary detection in complex feature spaces, and helps identify borderline cases, patients with ambiguous diagnoses, and patients with atypical presentations.

[0060] All patient feature vectors are sorted in the order of access with the smallest improved reachable distance using the OPTICS algorithm, the improved reachable distance from each patient feature vector to the previous patient feature vector after sorting is recorded, and the improved reachable distances are sequentially plotted as curves to form an improved reachable distance curve. The improved reachable distance curve is a two-dimensional curve graph, the abscissa of the improved reachable distance curve is the sorted patient feature vector, and the ordinate is the improved reachable distance from the patient feature vector corresponding to the abscissa to the previous patient feature vector after sorting;

[0061] In combination with the causal driving factors, the cluster boundaries in the improved reachable distance curve are identified, and the patient feature vectors in the improved reachable distance curve are clustered, where the cluster boundary is the horizontal coordinate in the improved reachable distance curve, and the patient feature vectors at the cluster boundary position and between the cluster boundaries, and the patients associated with the patient feature vectors belong to the same patient subgroup.

[0062] Specifically, the second-order difference sequence of the improved reachable distance curve is calculated, and the horizontal coordinate position of the local maximum point of the second-order difference sequence in the improved reachable distance curve is used as the candidate cluster boundary. Based on the etiology driving factor, the sum of the co-occurrence frequencies of the two patient symptoms associated with the patient feature vectors on both sides of the candidate cluster boundary in the etiology driving factor is identified. If the sum of the co-occurrence frequencies is lower than the current dynamic boundary threshold, the candidate cluster boundary is retained, and the retained candidate cluster boundary is used as the cluster boundary in the improved reachable distance curve. The patient feature vectors in the improved reachable distance curve are clustered. The calculation formula of the dynamic boundary threshold is:

[0063] ;

[0064] in, The dynamic boundary threshold representing the candidate cluster boundary, represents the initial boundary threshold, represents the adjustment coefficient, represents the sum of the maximum co-occurrence frequencies among all the candidate cluster boundaries, and P represents the sum of the co-occurrence frequencies of the candidate cluster boundaries; the adjustment parameter is set to 0.3;

[0065] Assume that the patient feature vectors on both sides of the candidate cluster boundary are H(1) and H(2) respectively, the patient corresponding to the patient feature vector H(1) has the first symptom and the second symptom, and the patient corresponding to the patient feature vector H(2) has the third symptom and the fourth symptom. Then the sum of the co-occurrence frequencies of the candidate cluster boundary in the etiology driver is: the co-occurrence frequency of the first symptom with the third and fourth symptoms in the etiology driver + the co-occurrence frequency of the second symptom with the third and fourth symptoms;

[0066] By combining the co-occurrence frequency of symptoms within the etiological drivers, clinical semantic mutation points are identified as boundary candidates. A dynamic boundary threshold is then used to implement differential-driven boundary screening, improving boundary positioning accuracy and medical consistency. This approach not only addresses the global density sensitivity and fuzzy boundary issues of traditional density clustering algorithms, but also overcomes the bottleneck of existing methods in modeling continuity across patient subgroups, providing a more stable and interpretable population foundation for subsequent medical project modeling.

[0067] Optionally, extracting the diagnosis and treatment information of the patient subgroup from the structured medical record data, and performing medical complexity modeling and medical benefit modeling on the medical items in the diagnosis and treatment information, including:

[0068] Extract non-overlapping medical items in the patient subgroups, extract the complexity parameters of the medical items and the treatment benefits in the patient subgroups, and calculate the medical complexity and medical benefits of the medical items in the patient subgroups, where the complexity parameters include the number of project categories included in the medical items and the time required for the medical items, and the treatment benefits include the 90-day disease recurrence rate after the implementation of the medical items, the current treatment achievement ratio of the medical items, and patient satisfaction.

[0069] Specifically, the calculation formula for the medical complexity is:

[0070] ;

[0071] in, Indicates the medical complexity of the medical project, Indicates the number of project categories included in the medical project. The medical project consists of one or more project categories among examination, treatment, and medication projects. Indicates the maximum number of project categories, is 3, Indicates the medical project duration of the medical project, Indicates the preset maximum time, set For 10 days, Are all complexity control coefficients, set ;

[0072] The calculation formula for the medical benefits is:

[0073] ;

[0074] in, represents the medical benefits of the medical project in the patient subgroup, It represents the mean 90-day disease recurrence rate, the mean current treatment outcome ratio and the mean patient satisfaction of the medical project in the process of implementing the medical project in the patient subgroup. Expressed full satisfaction, Both represent revenue control parameters, set In order , the satisfaction rating range is 0-10 points, and the current treatment achievement ratio of the medical project ranges from 0-1;

[0075] Optionally, the comprehensive value of the medical project is evaluated based on the medical benefits of the medical project in different patient subgroups, including:

[0076] The medical benefits of the medical project in each patient subgroup are obtained to form a medical benefit sequence of the medical project. The medical benefit sequence is a sequence composed of the medical benefits of the medical project in each patient subgroup. The sequence mean, sequence maximum and sequence standard deviation of the medical benefit sequence are calculated to evaluate the comprehensive value of the medical project. The greater the medical benefit created by the medical project per unit medical complexity, the greater the comprehensive value of the medical project.

[0077] As a preferred algorithm of the present invention, the evaluation formula of the comprehensive value is:

[0078] ;

[0079] in, Indicates the comprehensive value of medical projects, Indicates the medical complexity of the medical project, Indicates control parameters, setting is 0.001, Represents the sequence mean, sequence maximum, and sequence standard deviation of the medical revenue sequence of the medical project. Indicates the value weight coefficient, set They are 0.5 and 0.5 respectively.

[0080] Specifically, the first part of the comprehensive value assessment formula measures the average cost-effectiveness of medical projects in patient subgroups, emphasizing basic stable effectiveness; the second part focuses on the peak profit potential and stability of medical projects in patient subgroups, and is particularly suitable for discovering medical projects with high returns but controllable risks.

[0081] In order to solve the above problem, the present invention provides an electronic device, comprising:

[0082] a memory storing at least one instruction;

[0083] Communication interfaces to enable electronic equipment to communicate; and

[0084] The processor executes the instructions stored in the memory to implement the above-mentioned smart medical project value assessment method.

[0085] In order to solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is executed by a processor in an electronic device to implement the above-mentioned smart medical project value assessment method.

[0086] Compared with the existing technology, this invention proposes a smart medical project value assessment method based on medical and health big data, which has the following advantages:

[0087] First, this application introduces a multi-factor perception weight mechanism to dynamically calculate the semantic connection strength between patients and medical entities (such as diagnosis, drugs, examinations) in the structured medical record graph structure, and forms a patient feature representation based on the graph embedding fusion function. It weights and models multiple factors such as structural enhancement, semantic similarity, and time decay, and for the first time converts the complex dynamic relationship between nodes in the structured medical record graph structure into an explicit learnable vector; specifically, in response to the problems of diverse structures and incomparable edge weights in structured medical record graphs, multi-factor weighting mechanisms such as importance frequency and time decay are introduced for quantification, and semantic similarity is introduced to avoid the problem that simple structural embedding cannot accurately reflect the semantics of nodes.

[0088] At the same time, this application introduces a two-dimensional composite indicator model integrating mean, extreme value, and standard deviation into the value assessment of medical projects, effectively overcoming the problems of crude assessment and insufficient risk perception caused by traditional value calculations relying solely on a single statistic (such as the mean or ratio). By taking patient subgroups as units, and taking into account the global trends, extreme performance, and fluctuations of medical complexity and medical benefits, a comprehensive value function with an expectation-volatility-limit multi-layer trade-off mechanism is constructed to evaluate the value of medical projects. The first part measures the average cost-effectiveness of medical projects in patient subgroups, emphasizing basic stable performance; the second part focuses on the peak benefit potential and stability of medical projects in patient subgroups. It is particularly suitable for identifying medical projects with high benefits but controllable risks. By incorporating complexity and standard deviation terms for benefits, the model has risk adjustment capabilities, providing more forward-looking and strategic decision-making basis for medical insurance and hospitals. In addition, the model introduces a value weight coefficient for self-adjustment, allowing flexible adjustment of evaluation preferences according to policy orientations (such as efficacy priority, cost control priority), significantly enhancing the configurability and adaptability of the method. It not only improves the richness of evaluation dimensions, but also ensures the stability and sensitivity to differences of the results. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] Figure 1 A flowchart of a method for evaluating the value of an intelligent medical project based on medical and health big data is provided in accordance with one embodiment of the present invention.

[0090] Figure 2 A schematic diagram of a structured medical record diagram provided by one embodiment of the present invention.

[0091] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0092] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0093] The embodiment of the present application provides a method for evaluating the value of smart medical projects based on medical and health big data. The execution subject of the smart medical project value evaluation method includes but is not limited to at least one of the electronic devices such as a server and a terminal that can be configured to execute the method provided by the embodiment of the present application. In other words, the smart medical project value evaluation method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc.

[0094] Reference Figure 1 , embodiment 1 of the present invention is:

[0095] A method for evaluating the value of smart medical projects based on medical and health big data includes the following steps:

[0096] S1: Obtain unstructured medical record data, use the BERT-BiLSTM-CRF joint model to extract structured semantic information, and construct the structured semantic information into structured medical record data.

[0097] The unstructured medical record data is the patient's medical record text data. The BERT-BiLSTM-CRF joint model is used to extract the structured semantic information of the unstructured medical record data, including:

[0098] Use the BERT language model to contextually encode unstructured medical record data and extract semantic vector representations at the word and sentence level;

[0099] The semantic vector representation is input into the BiLSTM network to capture the dependencies between medical terms. The resulting semantic vector representation and the labels of the words and phrases associated with the semantic vector representation are obtained. The label types include patient ID, symptoms, previous diseases, examination index results, diagnosis and treatment information, timestamp information, and other types.

[0100] The CRF layer is used to jointly decode and optimize the labels of all semantic vector representations to obtain the final labels of each semantic vector representation and the associated words and sentences. The semantic vector representations and the associated words and sentences labeled with patient ID, symptoms, previous diseases, examination index results, diagnosis and treatment information, and timestamp information are used as structured medical record data. The diagnosis and treatment information includes diagnosis results and medical items. The medical items are the patient's treatment items, consisting of examination, treatment, and medication items.

[0101] Specifically, the medical items are composed of one or more item categories of examination, treatment, and medication items.

[0102] Compared with traditional medical record structuring methods based on rules or shallow models, this invention adopts a pre-trained language model combined with sequence modeling and structured decoding, which can effectively deal with the problems of semantic ambiguity, synonymous expressions and long text dependencies in medical texts. It uses the BiLSTM network to further enhance the sequence modeling capability. The extracted structured information has clear semantic boundaries and can be directly used in graph structure construction, patient representation modeling, and subsequent clustering and evaluation steps, realizing the key transformation from original medical record text to quantitative analysis.

[0103] S2: Map the structured medical record data into a graph structure to obtain a structured medical record graph structure. Perform a graph embedding representation of the patient nodes in the structured medical record graph structure combined with multivariate relationships to obtain a patient feature vector with graph semantics.

[0104] Extracting the patient ID and patient medical record information from the structured medical record data as a patient node, wherein the patient medical record information is a concatenation result of semantic vector representations of symptoms, previous diseases, examination index results, and diagnosis and treatment information associated with the patient ID; extracting medical entities and semantic vector representations of medical entities from the structured medical record data to form medical entity nodes, wherein the types of medical entities include symptoms, previous diseases, examination index types, and diagnosis and treatment information;

[0105] Constructing patient nodes and medical entity nodes into a structured medical record graph structure, wherein the structured medical record graph structure is composed of nodes and relationships between nodes, wherein the nodes include patient nodes and medical entity nodes, and the relationships between the nodes are the perception weights between the patient nodes and the medical entity nodes;

[0106] like Figure 2 Taking the structure diagram of a structured medical record graph as an example, the graph includes patient nodes and medical entity nodes. The patient nodes include patient nodes of three patients, namely patient node 1, patient node 2, and patient node 3. The types of medical entities in the medical entity nodes include symptoms, previous diseases, examination indicator types, and diagnosis and treatment information. There are lines between patient nodes and medical entity nodes, and the lines represent the perception weights between the patient nodes and the medical entity nodes.

[0107] The perception weight calculation formula is:

[0108] ;

[0109] ;

[0110] ;

[0111] ;

[0112] in, Represents the patient node c and the medical entity node The perceptual weight between represents the node association coefficient, Represents the medical numerical relationship between the patient node c and the medical entity node node, represents the time decay coefficient, represents the coefficient weight, , and set They are 0.6, 0.2, and 0.2 respectively;

[0113] The higher the perception weight, the higher the semantic association between the patient node and the medical entity node, which is manifested in that the patient node has the symptoms or previous diseases or examination indicator types or diagnosis and treatment information corresponding to the medical entity node. The perception weight is composed of the node association coefficient, the medical numerical relationship and the time decay coefficient.

[0114] The node association coefficient is composed of the frequency of occurrence of the medical entity node in the patient node and the semantic similarity between the medical entity node and the patient node, which represents the importance of the medical entity node to the patient node;

[0115] The medical numerical relationship is the numerical relationship between the patient node and the medical entity node. The larger the medical numerical relationship, the more obvious abnormality the patient node has in the corresponding medical entity node, which is manifested by more severe symptoms or previous diseases, and more abnormal examination indicator results.

[0116] The time decay coefficient is the timestamp decay result of the medical entity in the medical entity node appearing in the patient medical record information of the patient node. The time decay coefficient is used to give a time weight penalty to obsolete medical entities to achieve dynamic modeling and avoid excessive influence of obsolete medical entities on the construction process of the patient feature vector. If the medical entity in the medical entity node does not appear in the patient medical record information of the patient node, the time decay coefficient between the patient node and the medical entity node is set to 0;

[0117] Specifically, the normalized representation of numerical indicators such as the intensity of the symptoms, the intensity of the previous diseases, and the deviation value between the examination index results and the normal results is the medical numerical relationship between the patient node and the medical entity node. The medical numerical relationship between the patient node and the diagnosis and treatment information is 1. If the patient node does not have the symptoms, previous diseases, or examination index types associated with the medical entity node, the medical numerical relationship between the patient node and the medical entity node is 0.

[0118] represents the cosine similarity between the semantic vector representations in the patient node c and the medical entity node node, Indicates the statistical importance of the medical entity node node in the patient node c,

[0119] represents the frequency of occurrence of the medical entity associated with the medical entity node node in the structured medical record data of the patient associated with the patient node c. N represents the total number of patients. Indicates the number of patients who have a relationship with the medical entity node. represents the cosine similarity weight, Indicates the statistical importance weight, set is 0.6, is 0.4;

[0120] Specifically, in the calculation process of node association coefficient, semantic similarity and frequency are integrated to reflect the structural contribution of medical entities to patients. By combining contextual semantics and statistical strength, high-frequency and low-correlation problems (such as commonly found symptoms but not related to the current patient) are effectively handled, the weight of specific medical entities is increased, and common entities (such as "fever") are suppressed.

[0121] represents the attenuation control coefficient, Indicates the current timestamp, Indicates the timestamp of the medical entity in medical entity node node appearing in the patient medical record information of patient node c; the attenuation control coefficient is set to 0.2.

[0122] The patient nodes in the structured medical record graph are represented by graph embedding combined with multivariate relationships to obtain a patient feature vector with graph semantics, including:

[0123] Based on the perception weight between the patient node and the medical entity node, the 10 medical entity nodes with the highest perception weight with the patient node are selected as the medical entity node set of the patient node, and the medical entity node set of the patient node is embedded and fused to obtain the patient feature vector of the patient node.

[0124] Specifically, the embedding fusion formula of the medical entity node set of the patient node c is:

[0125] ;

[0126] ;

[0127] in, represents the patient feature vector of patient node c, W represents the weight matrix, represents the set of medical entity nodes of patient node c, , Represents a collection of medical entity nodes Any medical entity node in Represents a medical entity node The semantic vector representation in Represents a set All weighted semantic vector representations in are aggregated. Represents the activation function, the selected activation function is the tanh function;

[0128] Represent patient node c and medical entity node respectively The node correlation coefficient, medical numerical relationship and time decay coefficient between them;

[0129] In the embedding fusion process, a graph embedding representation based on semantic perception + structural enhancement is constructed, and medical entity nodes with larger perception weights are selected for fusion processing of semantic vector representation. The fusion processing results are used as the patient feature vector of the patient node, instead of extracting the patient medical record information from the patient node. The patient medical record information is directly used as the patient feature vector, so that the patient node has more comprehensive feature semantics. This realizes the process of constructing a graph embedding-oriented patient feature representation based on structured medical record data, forming a patient feature vector with the global semantics of all medical entity nodes and node relationships, thereby improving the accuracy and effectiveness of subsequent patient clustering.

[0130] S3: Construct etiological driving factors and use an improved density clustering algorithm to perform multi-scale clustering on patient feature vectors to obtain patients with similar etiological symptoms and form patient subgroups.

[0131] Based on structured medical record data, we construct etiology drivers, including:

[0132] Extracting a symptom set of each patient from the structured medical record data and initially constructing a symptom matrix, wherein the symptom matrix is ​​in the form of a matrix with M rows and M columns, where M represents the total number of symptoms of all patients;

[0133] For any two different symptoms, the number of symptom sets in which these two symptoms appear is counted as the co-occurrence frequency of the two symptoms, and the co-occurrence frequency is mapped to the corresponding position of the symptom matrix. The matrix element in the mth row and eth column in the symptom matrix represents the co-occurrence frequency of the mth symptom and the eth symptom. , the symptom matrix after all matrix elements are mapped and filled is used as the cause driving factor.

[0134] Combined with the above-mentioned etiological driving factors, an improved density clustering algorithm is used to perform multi-scale clustering of patient feature vectors, including:

[0135] Obtain the patient feature vectors of all patient nodes, calculate the Mahalanobis distance between any two patient feature vectors, use the Mahalanobis distance to determine the adaptive neighborhood radius of the patient feature vector, and extract the patient feature vectors whose Mahalanobis distance is smaller than the adaptive neighborhood radius as neighbor vectors;

[0136] Specifically, the patient feature vector The adaptive neighborhood radius calculation formula is:

[0137] ;

[0138] in, Represents the patient feature vector The adaptive neighborhood radius of represents the initial neighborhood radius, represents the mean Mahalanobis distance between all patient feature vectors, represents the distance control factor, Represents the patient feature vector With the initial neighborhood radius The standard deviation of the distance between other patient feature vectors within the range, setting the distance control factor is 0.1;

[0139] Specifically, the adaptive neighborhood radius calculation formula is used to dynamically adjust the local neighborhood radius of each patient's feature vector, effectively adapting to local density differences and improving clustering stability. Density perception is combined with Mahalanobis distance, and the adaptive neighborhood radius is used as a normalization factor to make the distance measurement result adaptable to local density. If the patient's feature vector is located in a dense area, the adaptive neighborhood radius is small and the Mahalanobis distance is amplified, emphasizing the difference between the patient and the sparse area. Otherwise, the Mahalanobis distance is compressed to avoid the misclustering of marginal patients.

[0140] Based on the Mahalanobis distance between the neighbor vector and the patient feature vector, the core distance of the patient feature vector is generated, the core distance is used to generate the improved reachable distance between the patient feature vectors, and the patient feature vectors are sorted based on the improved reachable distance to construct an improved reachable distance curve;

[0141] It should be noted that the patient feature vector The core distance is the patient feature vector The distance between the patient and the MinPtsth nearest feature vector; MinPts is a preset neighbor vector ranking, and MinPts is set to 5;

[0142] The reachable distance calculation method is improved based on the adaptive neighborhood radius, and the patient feature vector To patient feature vector The improved reachable distance is:

[0143] ;

[0144] in, Represents the patient feature vector To patient feature vector Improved reach distance, Represents the patient feature vector The core distance, Represents the patient feature vector and patient feature vector The Mahalanobis distance between

[0145] It should be noted that the reachable distance is the patient feature vector Can the patient feature vector be used? The density reaches, if is smaller, indicating that the patient's feature vector and In areas of the same density, an improved reachable distance is defined, and an adaptive neighborhood radius is introduced to normalize the Mahalanobis distance. In areas with higher density, the neighborhood radius is smaller, and the normalized distance is larger. In areas with lower density, the neighborhood radius is larger, and the normalized distance is smaller. This helps distinguish high-density from low-density areas, enhances the robustness of boundary detection in complex feature spaces, and helps identify borderline cases, patients with ambiguous diagnoses, and patients with atypical presentations.

[0146] All patient feature vectors are sorted in the order of access with the smallest improved reachable distance using the OPTICS algorithm, the improved reachable distance from each patient feature vector to the previous patient feature vector after sorting is recorded, and the improved reachable distances are sequentially plotted as curves to form an improved reachable distance curve. The improved reachable distance curve is a two-dimensional curve graph, the abscissa of the improved reachable distance curve is the sorted patient feature vector, and the ordinate is the improved reachable distance from the patient feature vector corresponding to the abscissa to the previous patient feature vector after sorting;

[0147] In combination with the causal driving factors, the cluster boundaries in the improved reachable distance curve are identified, and the patient feature vectors in the improved reachable distance curve are clustered, where the cluster boundary is the horizontal coordinate in the improved reachable distance curve, and the patient feature vectors at the cluster boundary position and between the cluster boundaries, and the patients associated with the patient feature vectors belong to the same patient subgroup.

[0148] It should be noted that the second-order difference sequence of the improved reachable distance curve is calculated, and the horizontal coordinate position of the local maximum point of the second-order difference sequence in the improved reachable distance curve is used as the candidate cluster boundary. Based on the etiology driving factor, the sum of the co-occurrence frequencies of the two patient symptoms associated with the patient feature vectors on both sides of the candidate cluster boundary in the etiology driving factor is identified. If the sum of the co-occurrence frequencies is lower than the current dynamic boundary threshold, the candidate cluster boundary is retained, and the retained candidate cluster boundary is used as the cluster boundary in the improved reachable distance curve. The patient feature vectors in the improved reachable distance curve are clustered. The calculation formula of the dynamic boundary threshold is:

[0149] ;

[0150] in, The dynamic boundary threshold representing the candidate cluster boundary, represents the initial boundary threshold, represents the adjustment coefficient, represents the sum of the maximum co-occurrence frequencies among all the candidate cluster boundaries, and P represents the sum of the co-occurrence frequencies of the candidate cluster boundaries; the adjustment parameter is set to 0.3;

[0151] Assume that the patient feature vectors on both sides of the candidate cluster boundary are H(1) and H(2) respectively, the patient corresponding to the patient feature vector H(1) has the first symptom and the second symptom, and the patient corresponding to the patient feature vector H(2) has the third symptom and the fourth symptom. Then the sum of the co-occurrence frequencies of the candidate cluster boundary in the etiology driver is: the co-occurrence frequency of the first symptom with the third and fourth symptoms in the etiology driver + the co-occurrence frequency of the second symptom with the third and fourth symptoms;

[0152] By combining the co-occurrence frequency of symptoms within the etiological drivers, clinical semantic mutation points are identified as boundary candidates. A dynamic boundary threshold is then used to implement differential-driven boundary screening, improving boundary positioning accuracy and medical consistency. This approach not only addresses the global density sensitivity and fuzzy boundary issues of traditional density clustering algorithms, but also overcomes the bottleneck of existing methods in modeling continuity across patient subgroups, providing a more stable and interpretable population foundation for subsequent medical project modeling.

[0153] S4: Extract diagnosis and treatment information of patient subgroups from structured medical record data, and perform medical complexity modeling and medical benefit modeling on the medical items in the diagnosis and treatment information.

[0154] Extracting the diagnosis and treatment information of the patient subgroup from the structured medical record data, and performing medical complexity modeling and medical benefit modeling on the medical items in the diagnosis and treatment information, including:

[0155] Extract non-overlapping medical items in the patient subgroups, extract the complexity parameters of the medical items and the treatment benefits in the patient subgroups, and calculate the medical complexity and medical benefits of the medical items in the patient subgroups, where the complexity parameters include the number of project categories included in the medical items and the time required for the medical items, and the treatment benefits include the 90-day disease recurrence rate after the implementation of the medical items, the current treatment achievement ratio of the medical items, and patient satisfaction.

[0156] Specifically, the calculation formula for the medical complexity is:

[0157] ;

[0158] in, Indicates the medical complexity of the medical project, Indicates the number of project categories included in the medical project. The medical project consists of one or more project categories among examination, treatment, and medication projects. Indicates the maximum number of project categories, is 3, Indicates the medical project duration of the medical project, Indicates the preset maximum time, set For 10 days, Are all complexity control coefficients, set ;

[0159] The calculation formula for the medical benefits is:

[0160] ;

[0161] in, represents the medical benefits of the medical project in the patient subgroup, It represents the mean 90-day disease recurrence rate, the mean current treatment outcome ratio and the mean patient satisfaction of the medical project in the process of implementing the medical project in the patient subgroup. Expressed full satisfaction, Both represent revenue control parameters, set In order The satisfaction rating range is 0-10, and the current treatment outcome ratio of the medical project ranges from 0-1.

[0162] S5: Evaluate the comprehensive value of medical programs based on the medical complexity and medical benefits of medical programs in different patient subgroups.

[0163] The medical benefits of the medical project in each patient subgroup are obtained to form a medical benefit sequence of the medical project. The medical benefit sequence is a sequence composed of the medical benefits of the medical project in each patient subgroup. The sequence mean, sequence maximum and sequence standard deviation of the medical benefit sequence are calculated to evaluate the comprehensive value of the medical project. The greater the medical benefit created by the medical project per unit medical complexity, the greater the comprehensive value of the medical project.

[0164] As a preferred algorithm of the present invention, the evaluation formula of the comprehensive value is:

[0165] ;

[0166] in, Indicates the comprehensive value of medical projects, Indicates the medical complexity of the medical project, Indicates control parameters, setting is 0.001, Represents the sequence mean, sequence maximum, and sequence standard deviation of the medical revenue sequence of the medical project. Indicates the value weight coefficient, set They are 0.5 and 0.5 respectively.

[0167] Specifically, the first part of the comprehensive value assessment formula measures the average cost-effectiveness of medical projects in patient subgroups, emphasizing basic stable effectiveness; the second part focuses on the peak profit potential and stability of medical projects in patient subgroups, and is particularly suitable for discovering medical projects with high returns but controllable risks.

[0168] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.

[0169] It should be noted that the serial numbers of the above-mentioned embodiments of the present invention are for descriptive purposes only and do not represent the advantages or disadvantages of the embodiments. In addition, the terms "including", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "including a ..." does not exclude the presence of other identical elements in the process, device, article or method comprising the element.

[0170] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0171] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A method for evaluating the value of smart medical projects based on medical and health big data, characterized by: The method comprises: S1: Obtain unstructured medical record data, use the BERT-BiLSTM-CRF joint model to extract structured semantic information, and construct the structured semantic information into structured medical record data; S2: Map the structured medical record data into a graph structure to obtain a structured medical record graph structure. Perform graph embedding representation of the patient nodes in the structured medical record graph structure combined with multivariate relationships to obtain a patient feature vector with graph semantics. S3: Construct etiological driving factors and use an improved density clustering algorithm to perform multi-scale clustering on patient feature vectors to obtain patients with similar etiological symptoms and form patient subgroups; S4: Extract diagnosis and treatment information of patient subgroups from structured medical record data, and perform medical complexity modeling and medical benefit modeling on the medical items in the diagnosis and treatment information; S5: Evaluate the comprehensive value of medical programs based on the medical complexity and medical benefits of medical programs in different patient subgroups.

2. The method for evaluating the value of smart medical projects based on medical and health big data according to claim 1, wherein: The unstructured medical record data is the patient's medical record text data. The BERT-BiLSTM-CRF joint model is used to extract the structured semantic information of the unstructured medical record data, including: Use the BERT language model to contextually encode unstructured medical record data and extract semantic vector representations at the word and sentence level; The semantic vector representation is input into the BiLSTM network to capture the dependencies between medical terms. The resulting semantic vector representation and the labels of the words and phrases associated with the semantic vector representation are obtained. The label types include patient ID, symptoms, previous diseases, examination index results, diagnosis and treatment information, timestamp information, and other types. The CRF layer is used to jointly decode and optimize the labels of all semantic vector representations to obtain the final labels of each semantic vector representation and the associated words and sentences. The semantic vector representations and the associated words and sentences labeled with patient ID, symptoms, previous diseases, examination index results, diagnosis and treatment information, and timestamp information are used as structured medical record data. The diagnosis and treatment information includes diagnosis results and medical items. The medical items are the patient's treatment items, consisting of examination, treatment, and medication items.

3. The method for evaluating the value of smart medical projects based on medical and health big data according to claim 2, characterized in that: Extracting the patient ID and patient medical record information from the structured medical record data as a patient node, wherein the patient medical record information is a concatenation result of semantic vector representations of symptoms, previous diseases, examination index results, and diagnosis and treatment information associated with the patient ID; extracting medical entities and semantic vector representations of medical entities from the structured medical record data to form medical entity nodes, wherein the types of medical entities include symptoms, previous diseases, examination index types, and diagnosis and treatment information; Constructing patient nodes and medical entity nodes into a structured medical record graph structure, wherein the structured medical record graph structure is composed of nodes and relationships between nodes, wherein the nodes include patient nodes and medical entity nodes, and the relationships between the nodes are the perception weights between the patient nodes and the medical entity nodes; The perception weight calculation formula is: ; ; ; ; in, Represents the patient node c and the medical entity node The perceptual weight between represents the node association coefficient, Represents the medical numerical relationship between the patient node c and the medical entity node node, represents the time decay coefficient, represents the coefficient weight, , and set They are 0.6, 0.2, and 0.2 respectively; The higher the perception weight, the higher the semantic association between the patient node and the medical entity node, which is manifested in that the patient node has the symptoms or previous diseases or examination indicator types or diagnosis and treatment information corresponding to the medical entity node. The perception weight is composed of the node association coefficient, the medical numerical relationship and the time decay coefficient. The node association coefficient is composed of the frequency of occurrence of the medical entity node in the patient node and the semantic similarity between the medical entity node and the patient node; The medical numerical relationship is the numerical relationship between the patient node and the medical entity node. The larger the medical numerical relationship, the more obvious anomalies the patient node has in the corresponding medical entity node. The time decay coefficient is the decay result of the timestamp of the medical entity in the medical entity node in the patient's medical record information. The time decay coefficient is used to give a time weight penalty to outdated medical entities to achieve dynamic modeling and avoid excessive influence of outdated medical entities on the construction process of patient feature vectors. represents the cosine similarity between the semantic vector representations in the patient node c and the medical entity node node, Indicates the statistical importance of the medical entity node node in the patient node c, represents the frequency of occurrence of the medical entity associated with the medical entity node node in the structured medical record data of the patient associated with the patient node c. N represents the total number of patients. Indicates the number of patients who have a relationship with the medical entity node. represents the cosine similarity weight, Indicates the statistical importance weight, set is 0.6, is 0.4; represents the attenuation control coefficient, Indicates the current timestamp, Indicates the timestamp of the medical entity in medical entity node node appearing in the patient medical record information of patient node c; the attenuation control coefficient is set to 0.

2.

4. The method for evaluating the value of smart medical projects based on medical and health big data according to claim 3, wherein: The patient nodes in the structured medical record graph are represented by graph embedding combined with multivariate relationships to obtain a patient feature vector with graph semantics, including: Based on the perception weight between the patient node and the medical entity node, the 10 medical entity nodes with the highest perception weight with the patient node are selected as the medical entity node set of the patient node, and the medical entity node set of the patient node is embedded and fused to obtain the patient feature vector of the patient node.

5. The method for evaluating the value of smart medical projects based on medical and health big data according to claim 1, wherein: Based on structured medical record data, we construct etiology drivers, including: Extracting a symptom set of each patient from the structured medical record data and initially constructing a symptom matrix, wherein the symptom matrix is ​​in the form of a matrix with M rows and M columns, where M represents the total number of symptoms of all patients; For any two different symptoms, the number of symptom sets in which these two symptoms appear is counted as the co-occurrence frequency of the two symptoms, and the co-occurrence frequency is mapped to the corresponding position of the symptom matrix. The matrix element in the mth row and eth column in the symptom matrix represents the co-occurrence frequency of the mth symptom and the eth symptom. , the symptom matrix after all matrix elements are mapped and filled is used as the cause driving factor.

6. The method for evaluating the value of smart medical projects based on medical and health big data according to claim 5, characterized in that: Combined with the above-mentioned etiological driving factors, an improved density clustering algorithm is used to perform multi-scale clustering of patient feature vectors, including: Obtain the patient feature vectors of all patient nodes, calculate the Mahalanobis distance between any two patient feature vectors, use the Mahalanobis distance to determine the adaptive neighborhood radius of the patient feature vector, and extract the patient feature vectors whose Mahalanobis distance is smaller than the adaptive neighborhood radius as neighbor vectors; Based on the Mahalanobis distance between the neighbor vector and the patient feature vector, the core distance of the patient feature vector is generated, the core distance is used to generate the improved reachable distance between the patient feature vectors, and the patient feature vectors are sorted based on the improved reachable distance to construct an improved reachable distance curve; The reachable distance calculation method is improved based on the adaptive neighborhood radius, and the patient feature vector To patient feature vector The improved reachable distance is: ; in, Represents the patient feature vector To patient feature vector Improved reach distance, Represents the patient feature vector The core distance, Represents the patient feature vector and patient feature vector The Mahalanobis distance between All patient feature vectors are sorted in the order of access with the smallest improved reachable distance using the OPTICS algorithm, the improved reachable distance from each patient feature vector to the previous patient feature vector after sorting is recorded, and the improved reachable distances are sequentially plotted as curves to form an improved reachable distance curve. The improved reachable distance curve is a two-dimensional curve graph, the abscissa of the improved reachable distance curve is the sorted patient feature vector, and the ordinate is the improved reachable distance from the patient feature vector corresponding to the abscissa to the previous patient feature vector after sorting; In combination with the causal driving factors, the cluster boundaries in the improved reachable distance curve are identified, and the patient feature vectors in the improved reachable distance curve are clustered, where the cluster boundary is the horizontal coordinate in the improved reachable distance curve, and the patient feature vectors at the cluster boundary position and between the cluster boundaries, and the patients associated with the patient feature vectors belong to the same patient subgroup.

7. The method for evaluating the value of smart medical projects based on medical and health big data according to claim 1, wherein: Extracting the diagnosis and treatment information of the patient subgroup from the structured medical record data, and performing medical complexity modeling and medical benefit modeling on the medical items in the diagnosis and treatment information, including: Extract non-overlapping medical items in the patient subgroups, extract the complexity parameters of the medical items and the treatment benefits in the patient subgroups, and calculate the medical complexity and medical benefits of the medical items in the patient subgroups, where the complexity parameters include the number of project categories included in the medical items and the time required for the medical items, and the treatment benefits include the 90-day disease recurrence rate after the implementation of the medical items, the current treatment achievement ratio of the medical items, and patient satisfaction.

8. The method for evaluating the value of smart medical projects based on medical and health big data according to claim 7, characterized in that: Evaluate the comprehensive value of the medical program based on its benefits across different patient subgroups, including: The medical benefits of the medical project in each patient subgroup are obtained to form a medical benefit sequence of the medical project. The medical benefit sequence is a sequence composed of the medical benefits of the medical project in each patient subgroup. The sequence mean, sequence maximum and sequence standard deviation of the medical benefit sequence are calculated to evaluate the comprehensive value of the medical project.