Big data-based medical treatment in different places data analysis method and system

By constructing a temporal medical technology network and a spatiotemporal graph neural network model, the limitations of prediction and evaluation in cross-regional medical treatment data analysis are overcome. This enables multi-dimensional assessment of the hospital's technological influence and feasible resource optimization, supporting hierarchical diagnosis and treatment and resource planning.

CN121641316BActive Publication Date: 2026-05-19BEIJING CHUANGZHI HEYU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING CHUANGZHI HEYU TECH CO LTD
Filing Date
2025-12-25
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies for analyzing cross-regional medical treatment data have several drawbacks, including an inability to predict future development trends, a limited range of evaluation dimensions, and difficulty in translating them into feasible hierarchical medical treatment pathways and resource optimization schemes.

Method used

By acquiring and integrating cross-regional medical treatment data with multi-source external data, a time-series medical technology network is constructed. A spatiotemporal graph neural network model is used to predict the technological influence of hospitals, and a hierarchical diagnosis and treatment recommendation path and a regional medical resource optimization allocation scheme are generated.

Benefits of technology

It enables multi-dimensional and dynamic evaluation of the hospital's technological influence, provides forward-looking predictive results, supports differentiated patient recommendation pathways and precise resource allocation, and forms a complete closed loop from data analysis to management decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121641316B_ABST
    Figure CN121641316B_ABST
Patent Text Reader

Abstract

The application discloses a medical treatment data analysis method and system based on big data, relates to the technical field of big data analysis, and solves the technical problems that it is difficult to capture the dynamic changes of hospitals in the medical technology network, the evaluation index is single, and it is also difficult to effectively predict the technical influence of the hospital in a specific period in the future based on the medical treatment behavior data. The application comprises the following steps: firstly, acquiring and fusing medical treatment data and multi-source external data, and performing pretreatment to obtain a standardized fusion data set; then, constructing a time sequence medical technology network and calculating the dynamic technical influence center degree of a hospital node; constructing a space-time graph neural network model, and predicting the technical influence score of each hospital node in the future; then, selecting a first medical service node based on the technical influence score; and finally, generating a hierarchical diagnosis and treatment recommendation path and a regional medical resource optimization configuration scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of big data analysis technology, specifically relating to a method and system for analyzing cross-regional medical treatment data based on big data. Background Technology

[0002] With the continuous improvement of my country's medical security system and the continuous expansion of the scale of cross-regional medical treatment, scientific analysis of cross-regional medical treatment data is of great significance for optimizing the allocation of medical resources, guiding the orderly flow of patients, and implementing the hierarchical medical system.

[0003] While existing technologies have enabled cross-regional medical data analysis to some extent, they still have significant limitations. First, multi-source data is usually statistically analyzed after the fact, making it impossible to predict future trends and thus leading to delayed decision-making. Second, the evaluation dimensions are singular, making it difficult to comprehensively consider a hospital's dynamic influence in cross-regional medical technology networks, its attractiveness to complex and critical illnesses, and the reach of its technology. Third, most existing studies remain at the level of hospital rankings, making it difficult to effectively translate the analysis results into implementable hierarchical medical pathways and precise resource optimization and allocation schemes.

[0004] Therefore, there is an urgent need in this field for an intelligent data analysis method that can dynamically, proactively, and multidimensionally assess the technological influence of hospitals and directly apply the analysis results to the optimal allocation of medical resources. Summary of the Invention

[0005] This invention aims to solve at least one of the technical problems existing in the prior art; to this end, this invention proposes a method and system for analyzing cross-regional medical treatment data based on big data, to solve the following technical problems:

[0006] While existing technologies have enabled cross-regional medical data analysis to some extent, they still have significant limitations. First, multi-source data is usually statistically analyzed after the fact, making it impossible to predict future trends and thus leading to delayed decision-making. Second, the evaluation dimensions are singular, making it difficult to comprehensively consider a hospital's dynamic influence in cross-regional medical technology networks, its attractiveness to complex and critical illnesses, and the reach of its technology. Third, most existing studies remain at the level of hospital rankings, making it difficult to effectively translate the analysis results into implementable hierarchical medical pathways and precise resource optimization and allocation schemes.

[0007] To address the aforementioned problems, the first aspect of this invention provides a method for analyzing cross-regional medical treatment data based on big data, comprising the following steps:

[0008] S1: Acquire and merge cross-regional medical treatment data and multi-source external data, preprocess the cross-regional medical treatment data and multi-source external data to obtain a standardized fused dataset;

[0009] S2: Based on the standardized fusion dataset, filter out medical records of specific diseases in different locations, construct a time-series medical technology network for specific diseases, and calculate the dynamic technical influence centrality of each hospital node in each time slice.

[0010] S3: Construct a spatiotemporal graph neural network model, and input the time-series medical technology network into the spatiotemporal graph neural network model. The time-series medical technology network includes hospital nodes, edge weights and dynamic technology influence centrality. The network is input into the spatiotemporal graph neural network model to output the predicted technology influence score of each hospital node in a future preset time slice.

[0011] S4: Based on the predicted technical influence score of each hospital node, a score threshold is set, and hospital nodes with predicted technical influence scores higher than the score threshold are selected as the first medical service nodes.

[0012] S5: Based on the first medical service node, generate a hierarchical diagnosis and treatment recommendation path and a regional medical resource optimization allocation scheme for specific diseases. The hierarchical diagnosis and treatment recommendation path generates priority medical flow directions for patients with different risk levels based on the spatial distance between the patient's place of origin and the first medical service node and the predicted technology influence score. The regional medical resource optimization allocation scheme identifies areas with high predicted technology influence scores but low local medical resource density and marks them as key areas that urgently need technical support and resource allocation.

[0013] Preferably, in step S1, the preprocessing includes outlier removal, missing value imputation, and data standardization. Specifically, missing value imputation employs a weighted mean imputation method based on the incidence rate of the disease.

[0014] The dimensions of the missing data to be filled are determined, including the number of out-of-town medical visits, the average length of hospital stay, and the average out-of-pocket cost per visit. Non-missing medical data is then filtered out. The The data must meet the following criteria: the data belongs to the same disease category as the missing data, and the per capita disposable income of the source region is at the same level as that of the source region of the missing data;

[0015] Get each Disease incidence rate in corresponding regions Calculate the imputation values ​​for missing data, specifically:

[0016]

[0017] in, Fill in the missing data with values. This represents the number of samples with non-missing data. An index identifier for a single, non-missing medical data sample that meets the screening criteria.

[0018] Preferably, in step S2, screening out-of-town medical records for specific diseases includes the following steps:

[0019] Construct a multi-dimensional core technology feature profile of the target disease, the profile including a diagnostic dimension composed of core disease diagnosis codes and a treatment dimension composed of key treatment operation codes;

[0020] From the standardized fusion dataset, records whose main diagnostic codes belong to the diagnostic dimension or whose main operational codes belong to the treatment dimension are selected as the initial set of selected records;

[0021] Based on the pre-trained BERT model, the diagnostic text summary and operation name of each record in the initial screening record set are semantically embedded, and the cosine similarity between the summary and operation name and the core technology feature vector of the target disease is calculated as the probability of technology relevance.

[0022] A probability threshold is set to include records with a probability of technical relevance higher than the threshold in the final screening record set. Each initially screened record is assigned a final record weight that is positively correlated with the probability of technical relevance, so as to form a weighted set of out-of-town medical treatment records for specific diseases.

[0023] Preferably, in step S2, constructing a time-series medical technology network for a specific disease includes the following steps:

[0024] The time-series medical technology network uses hospitals in various regions as network nodes, and the number of patients with specific diseases flowing from the source region to the target hospital within a unit time slice is used as the edge weight between nodes.

[0025] The edge weights between nodes are specifically as follows:

[0026]

[0027]

[0028] in, For time slices From the source region Flowing to target hospitals The sum of weighted records of patients with specific diseases. The total number of medical records for a specific disease. For summation indexing, representing a single medical record, For the first The final record weight of each record. For the source region per capita disposable income The average per capita disposable income nationwide. Target hospital The technical difficulty adjustment factor Target hospital The proportion of complex and difficult cases of this specific disease received. This represents the percentage of complex and difficult cases of this specific disease nationwide.

[0029] Preferably, in step S2, calculating the dynamic technical influence centrality of each hospital node in each time slice includes the following steps:

[0030] Centrality of input Specifically: ,in, For time slices Inside, all pointing to the target hospital The sum of the weights of the incoming edges;

[0031] Proximity Centrality Specifically: ,in, Target hospital From the network to other hospitals The shortest weighted path length;

[0032] Authority centrality is calculated iteratively based on the HITS algorithm to obtain the time slice of each node. Authority Centrality ;

[0033] The weights corresponding to the input centrality, proximity centrality, and authority centrality are determined based on the entropy weight method, and then weighted and added together to obtain the dynamic technical influence centrality.

[0034] Preferably, the determination of weights based on the entropy weight method includes the following steps:

[0035] Entropy weighting is used for each time slice. The weights are calculated dynamically, specifically as follows:

[0036] For time slices This forms a matrix consisting of the three centrality index values ​​of all hospital nodes, and the information entropy of each index is calculated. ,in, For nodes In indicators The proportion of the surface;

[0037] The weight of each indicator is calculated as follows: ,in, For the first The weights corresponding to each indicator.

[0038] Preferably, in step S3, constructing the spatiotemporal graph neural network model includes the following steps:

[0039] The spatiotemporal graph neural network model includes a graph attention network layer, a gated recurrent unit layer, and a feature interaction layer;

[0040] The graph attention network layer takes the dynamic technical influence centrality of hospital nodes and the edge weights between nodes as core inputs. After concatenation with the basic features of the nodes, attention coefficients are generated through the LeakyReLU activation function. A dynamic attention head number rule is introduced, and the results of the multi-head attention mechanism are concatenated to output the spatial feature vector of each hospital node in the corresponding time slice.

[0041] The gated recurrent unit layer assigns corresponding time decay weights to the spatial feature vectors of different time slices according to the time slice type of the time-series medical technology network, integrates the time decay weights into the calculation of the update gate and reset gate of GRU, and outputs the time feature vector of each hospital node.

[0042] The feature interaction layer uses element-wise multiplication to fuse the spatial feature vector and the temporal feature vector to obtain a spatiotemporal fusion feature vector. After linear transformation, it is processed by the ReLU activation function to obtain the prediction technology impact score of each hospital node in the future preset time slice.

[0043] Preferably, step S4 includes the following steps:

[0044] The scoring threshold for the first medical service node is set to the upper quartile of the prediction technology influence scores of all hospital nodes.

[0045] The predicted technical influence score of each hospital node is compared with the score threshold. Nodes with scores higher than the score threshold are directly included in the candidate pool. For the candidate first medical service node, if the number of similar high-level hospitals in its area exceeds the preset density threshold, it is determined to be a resource redundant node and is removed from the first medical service node list to obtain the final first medical service node list.

[0046] Preferably, step S5 includes the following steps:

[0047] Based on the geographical location and predicted technical influence score of the first medical service node, as well as the spatial distribution of the patient's origin, a recommendation network is constructed with the patient's origin and the first medical service node as vertices and potential medical flow as edges. A comprehensive recommendation index is calculated for each edge in the network. Based on the index, a direct recommendation path guided by technical influence is generated for high-risk patients, and K shortest recommendation paths that balance spatial distance and technical influence are generated for medium- and low-risk patients, forming a hierarchical diagnosis and treatment recommendation map.

[0048] Meanwhile, the specific method for optimizing the allocation of regional medical resources is as follows: each first medical service node is mapped to its respective administrative division, the medical resource density of each administrative division is calculated, and through spatial overlay analysis, the first medical service node whose predicted technical influence score is higher than the first preset threshold and whose administrative division medical resource density is lower than the second preset threshold is identified. The area where the node is located is marked as a key area that urgently needs technical support and resource allocation, and a resource allocation recommendation report is output.

[0049] A second aspect of the present invention provides a big data-based cross-regional medical treatment data analysis system, comprising the following modules:

[0050] Multi-source data fusion and preprocessing module: acquires and fuses out-of-town medical treatment data and multi-source external data, preprocesses the out-of-town medical treatment data and multi-source external data to obtain a standardized fusion dataset;

[0051] Time-series medical technology network construction module: Based on the standardized fusion dataset, filter out out-of-town medical treatment records for specific diseases, construct a time-series medical technology network for specific diseases, and calculate the dynamic technology influence centrality of each hospital node in each time slice;

[0052] Spatiotemporal prediction module for technological influence: Construct a spatiotemporal graph neural network model, input the time-series medical technology network into the spatiotemporal graph neural network model, wherein the time-series medical technology network includes hospital nodes, edge weights and dynamic technological influence centrality, and input into the spatiotemporal graph neural network model to output the predicted technological influence score of each hospital node in a future preset time slice;

[0053] Intelligent selection module for medical service nodes: Based on the predicted technical influence score of each hospital node, a score threshold is set, and hospital nodes with predicted technical influence scores higher than the score threshold are selected as the first medical service nodes;

[0054] The hierarchical diagnosis and treatment and resource optimization decision support module generates a hierarchical diagnosis and treatment recommendation path and a regional medical resource optimization allocation scheme based on the first medical service node. The hierarchical diagnosis and treatment recommendation path generates priority medical flow directions for patients with different risk levels based on the spatial distance between the patient's place of origin and the first medical service node and the predicted technology influence score. The regional medical resource optimization allocation scheme identifies areas with high predicted technology influence scores but low local medical resource density and marks them as key areas that urgently need technical support and resource allocation.

[0055] The beneficial effects of this invention are:

[0056] This invention utilizes a pre-trained language model to perform semantic understanding of medical records and calculate the probability of technical relevance, which greatly improves the accuracy of data screening for specific diseases, reduces interference from noisy data, and lays a high-quality data foundation for core network construction and influence calculation.

[0057] This invention constructs a temporal medical technology network by integrating multi-source data such as patient flow, regional economic level, and hospital technical difficulty. It innovatively proposes a dynamic composite index that integrates three centralities: entry weight, proximity, and authority, to calculate the dynamic technical influence centrality of hospitals. This achieves a multi-dimensional, dynamic, and precise quantitative assessment of hospital technical influence. Finally, the dynamic technical influence centrality of hospital nodes and the edge weights between nodes are used as core inputs to construct a spatiotemporal graph neural network model that combines graph attention mechanisms, gated recurrent units, and feature interaction layers. This model achieves high-precision prediction of future hospital technical influence. It can simultaneously capture the spatial dependencies and temporal evolution patterns of the medical technology network, making the prediction results forward-looking and reliable, providing a basis for dynamic resource planning.

[0058] This invention generates differentiated recommended paths for patients with different risk levels through network analysis algorithms, and identifies high-score, low-density key areas through spatial overlay analysis. This transforms complex predictive data into clear and actionable decision-making tools, directly serving patient medical guidance and precise allocation of government resources. It forms a complete closed loop from data analysis to management decision-making, significantly enhancing the practical value of the method. Attached Figure Description

[0059] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0060] Figure 2 This is a schematic diagram of the module flow of the present invention. Detailed Implementation

[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] Please see Figure 1 As shown, this invention is a data analysis method for cross-regional medical treatment based on big data, including the following steps:

[0063] S1: Acquire and merge cross-regional medical treatment data and multi-source external data, preprocess the cross-regional medical treatment data and multi-source external data to obtain a standardized fused dataset;

[0064] S2: Based on the standardized fusion dataset, filter out medical records of specific diseases in different locations, construct a time-series medical technology network for specific diseases, and calculate the dynamic technical influence centrality of each hospital node in each time slice.

[0065] S3: Construct a spatiotemporal graph neural network model, and input the time-series medical technology network into the spatiotemporal graph neural network model. The time-series medical technology network includes hospital nodes, edge weights and dynamic technology influence centrality. The network is input into the spatiotemporal graph neural network model to output the predicted technology influence score of each hospital node in a future preset time slice.

[0066] S4: Based on the predicted technical influence score of each hospital node, a score threshold is set, and hospital nodes with predicted technical influence scores higher than the score threshold are selected as the first medical service nodes.

[0067] S5: Based on the first medical service node, generate a hierarchical diagnosis and treatment recommendation path and a regional medical resource optimization allocation scheme for specific diseases. The hierarchical diagnosis and treatment recommendation path generates priority medical flow directions for patients with different risk levels based on the spatial distance between the patient's place of origin and the first medical service node and the predicted technology influence score. The regional medical resource optimization allocation scheme identifies areas with high predicted technology influence scores but low local medical resource density and marks them as key areas that urgently need technical support and resource allocation.

[0068] Specifically, based on the dataset from step S1, firstly, through core diagnostic / operational encoding combined with semantic filtering using a pre-trained BERT model, cross-regional medical records for specific diseases are accurately screened. Then, using hospitals as nodes and weighted cross-regional patient flows as directed edges, a temporal medical technology network for specific diseases is constructed. Finally, the dynamic technical influence centrality of each hospital node in each time slice is calculated. This centrality is the weighted sum of entry centrality, proximity centrality, and authority centrality, with the weights dynamically determined by the entropy weight method. A spatiotemporal graph neural network model integrating graph attention network, gated recurrent units, and feature interaction layers is constructed. The dynamic technical influence centrality of hospital nodes and the edge weights between nodes are used as core inputs and fed into the spatiotemporal graph neural network model for training. The force layer captures the spatial dependencies between hospitals, the gated recurrent unit layer learns the temporal evolution of the technical influence, and the feature interaction layer integrates spatiotemporal features, ultimately outputting the predicted technical influence score of each hospital node within a future preset time slice; receiving the predicted technical influence score output in step S3, the score threshold is set to the upper quartile of the scores of all hospital nodes, initially screening nodes above this threshold as candidates, and then checking the regional medical resource density of the candidate nodes. If there are too many similar high-level hospitals in the region, they are judged as resource redundancy and removed, ultimately forming the first medical service node list; based on the first medical service node list obtained in step S4, the following two parallel processes are executed: generating a hierarchical diagnosis and treatment recommendation path and generating a regional medical resource optimization allocation scheme.

[0069] In one embodiment of the present invention, step S1 includes preprocessing such as outlier removal, missing value imputation, and data standardization. Specifically, missing value imputation employs a weighted mean imputation method based on the incidence rate of the disease.

[0070] The dimensions of the missing data to be filled are determined, including the number of out-of-town medical visits, the average length of hospital stay, and the average out-of-pocket cost per visit. Non-missing medical data is then filtered out. The The data must meet the following criteria: the data belongs to the same disease category as the missing data, and the per capita disposable income of the source region is at the same level as that of the source region of the missing data;

[0071] Get each Disease incidence rate in corresponding regions Calculate the imputation values ​​for missing data, specifically:

[0072]

[0073] in, Fill in the missing data with values. This represents the number of samples with non-missing data. An index identifier for a single, non-missing medical data sample that meets the screening criteria.

[0074] Specifically, the per capita disposable income of the source region and the per capita disposable income of the source region with missing data are at the same level. By adopting the latest "National Residents' Per Capita Disposable Income Five-Division Grouping Data" released by the National Bureau of Statistics, all surveyed households across the country are arranged from low to high per capita disposable income and divided into five groups: low-income group, lower-middle-income group, middle-income group, upper-middle-income group, and high-income group. To simplify the model, the above five groups are combined into three preset levels, where the low-income group corresponds to the low-income group and the lower-middle-income group, the middle-income group corresponds to the middle-income group, and the high-income group corresponds to the upper-middle-income group and the high-income group.

[0075] In one embodiment of the present invention, step S2, screening out-of-town medical treatment records for specific diseases, includes the following steps:

[0076] Construct a multi-dimensional core technology feature profile of the target disease, the profile including a diagnostic dimension composed of core disease diagnosis codes and a treatment dimension composed of key treatment operation codes;

[0077] From the standardized fusion dataset, records whose main diagnostic codes belong to the diagnostic dimension or whose main operational codes belong to the treatment dimension are selected as the initial set of selected records;

[0078] Based on the pre-trained BERT model, the diagnostic text summary and operation name of each record in the initial screening record set are semantically embedded, and the cosine similarity between the summary and operation name and the core technology feature vector of the target disease is calculated as the probability of technology relevance.

[0079] A probability threshold is set to include records with a probability of technical relevance higher than the threshold in the final screening record set. Each initially screened record is assigned a final record weight that is positively correlated with the probability of technical relevance, so as to form a weighted set of out-of-town medical treatment records for specific diseases.

[0080] Specifically, firstly, a multi-dimensional core technical feature profile of the target disease is defined jointly by clinical experts and data scientists. This profile serves as a standard template for determining whether a medical record falls within the core treatment scope of the disease, and includes two dimensions: 1. Diagnostic dimension: composed of a series of core disease diagnostic codes (ICD-10). For example, for lung cancer, this set includes codes for all lung cancer subtypes such as C34.0 (malignant tumor of the main bronchus) and C34.1 (upper lobe, bronchus, or lung); 2. Treatment dimension: composed of a series of key treatment operation codes (ICD-9-CM-3). These codes represent the characteristic and complex features of the disease. For diagnostic surgeries, such as lung cancer, the set might include: 32.29: Thoracoscopic lobectomy (representing a core technology of minimally invasive surgery), 32.41: Bronchoscopic radiofrequency ablation (representing a core technology of interventional therapy); From the standardized fusion dataset, a preliminary screening based on coding rules is performed. The screening logic is as follows: A record is included in the initial screening record set if it meets either of the following conditions: its main diagnostic code belongs to the aforementioned "diagnostic dimension" set and its main operational code belongs to the aforementioned "treatment dimension" set. This aims to quickly narrow down the scope and exclude completely irrelevant cases; The diagnostic text of each record in the initial screening record set is then... The abstract (e.g., malignant tumor of the upper lobe of the right lung) and the operation name (e.g., thoracoscopic lobectomy) are concatenated to form a short text. A BERT model pre-trained on a medical corpus is used to convert this short text into a high-dimensional, machine-readable numerical vector (i.e., a semantic embedding vector). All diagnostic and treatment keywords in the core technology feature profile of the target disease are also converted into semantic embedding vectors using the same BERT model, and their average values ​​are calculated to form a unified core technology feature vector for the target disease. The cosine similarity between the semantic embedding vector of each record and the core technology feature vector is calculated and defined as the similarity between that record and the target disease feature vector. The probability threshold for the technical relevance of the core technology of the disease is determined by the model's performance on the validation set. Specifically, on the labeled validation set, a similarity value that results in the highest F1 score is selected as the probability threshold by plotting a PR curve. Records with a technical relevance probability higher than the probability threshold are included in the final screening record set. Each record in the final screening record set is assigned a final record weight, which is positively correlated with the technical relevance probability. Specifically, the relationship is: final record weight = technical relevance probability. This means that the higher the semantic similarity of a record to the core technology, the greater its weight in subsequent network construction and influence calculation.

[0081] In one embodiment of the present invention, step S2, constructing a time-series medical technology network for a specific disease, includes the following steps:

[0082] The time-series medical technology network uses hospitals in various regions as network nodes, and the number of patients with specific diseases flowing from the source region to the target hospital within a unit time slice is used as the edge weight between nodes.

[0083] The edge weights between nodes are specifically as follows:

[0084]

[0085]

[0086] in, For time slices From the source region Flowing to target hospitals The sum of weighted records of patients with specific diseases. The total number of medical records for a specific disease. For summation indexing, representing a single medical record, For the first The final record weight of each record. For the source region per capita disposable income The average per capita disposable income nationwide. Target hospital The technical difficulty adjustment factor Target hospital The proportion of complex and difficult cases of this specific disease received. This represents the percentage of complex and difficult cases of this specific disease nationwide.

[0087] In one embodiment of the present invention, step S2, calculating the dynamic technical influence centrality of each hospital node in each time slice, includes the following steps:

[0088] Centrality of input Specifically: ,in, For time slices Inside, all pointing to the target hospital The sum of the weights of the incoming edges;

[0089] Proximity Centrality Specifically: ,in, Target hospital From the network to other hospitals The shortest weighted path length;

[0090] Authority centrality is calculated iteratively based on the HITS algorithm to obtain the time slice of each node. Authority Centrality ;

[0091] The weights corresponding to the input centrality, proximity centrality, and authority centrality are determined based on the entropy weight method, and then weighted and added together to obtain the dynamic technical influence centrality.

[0092] Specifically, the weighted centrality is used to measure the weighted centrality in a time slice. Inside, the target hospital The sum of a hospital's ability to directly acquire patient resources from outside the network (i.e., all source regions) reflects its direct technological attractiveness. A higher value indicates that the hospital directly attracts more weighted out-of-town patients within a specific time period, and its market appeal is stronger; proximity centrality is used to measure the target hospital To all other hospitals in the network The distance reflects the hospital's pivotal role in the network and the efficiency of information / resource flow. A node that is close to other hospitals is often located at the center of the network. The higher the value, the more central the hospital is in the medical technology network and the higher the efficiency of indirect connections with other hospitals. Authority centrality, based on the HITS algorithm, is used to identify nodes in the network that are not only heavily connected but also connected to other important hospitals. It reflects the reputation of the hospital that is transmitted through the network structure and recognized by the algorithm.

[0093] In one embodiment of the present invention, determining the weights based on the entropy weight method includes the following steps:

[0094] Entropy weighting is used for each time slice. The weights are calculated dynamically, specifically as follows:

[0095] For time slices This forms a matrix consisting of the three centrality index values ​​of all hospital nodes, and the information entropy of each index is calculated. ,in, For nodes In indicators The proportion of the surface;

[0096] The weight of each indicator is calculated as follows: ,in, For the first The weights corresponding to each indicator.

[0097] In one embodiment of the present invention, step S3, constructing a spatiotemporal graph neural network model, includes the following steps:

[0098] The spatiotemporal graph neural network model includes a graph attention network layer, a gated recurrent unit layer, and a feature interaction layer;

[0099] The graph attention network layer takes the dynamic technical influence centrality of hospital nodes and the edge weights between nodes as core inputs. After concatenation with the basic features of the nodes, attention coefficients are generated through the LeakyReLU activation function. A dynamic attention head number rule is introduced, and the results of the multi-head attention mechanism are concatenated to output the spatial feature vector of each hospital node in the corresponding time slice.

[0100] The gated recurrent unit layer assigns corresponding time decay weights to the spatial feature vectors of different time slices according to the time slice type of the time-series medical technology network, integrates the time decay weights into the calculation of the update gate and reset gate of GRU, and outputs the time feature vector of each hospital node.

[0101] The feature interaction layer uses element-wise multiplication to fuse the spatial feature vector and the temporal feature vector to obtain a spatiotemporal fusion feature vector. After linear transformation, it is processed by the ReLU activation function to obtain the prediction technology impact score of each hospital node in the future preset time slice.

[0102] Specifically, considering the differences in patient mobility characteristics among different diseases, this layer introduces a dynamic attention head rule: for sensory diseases (where patients move over a wide range and have many associated nodes), 5 attention heads are used (to comprehensively capture spatial relationships between multiple regions); for rare diseases (where patients mostly flow to core hospitals), 3 attention heads are used (to focus on the relationships between core nodes).

[0103] In the gated recurrent unit layer, based on the time slice type of the time-series medical technology network (determined by the disease's treatment cycle; infectious diseases use weekly time slices, and rare diseases use monthly time slices), time decay weights are assigned to the spatial feature vectors of different time slices. Specifically, weekly time slices are assigned a time decay weight of 0.98 (due to faster data updates); monthly time slices are assigned a time decay weight of 0.95 (due to higher data stability, thus strengthening the contribution of recent data to prediction and reducing interference from long-term data). The time decay weights are integrated into the update and reset gates of the GRU for calculation. In the inputs of the update gate (which controls the proportion of historical information retention) and the reset gate (which controls the proportion of historical information forgetting), the hidden state of the previous moment (reflecting historical time characteristics) is multiplied by the time decay weight, then concatenated with the spatial feature vector of the current time slice, and finally the gating value is obtained through the Sigmoid activation function. The final output is the time feature vector of each hospital node.

[0104] The constructed time-series medical technology network for specific diseases (including hospital nodes, edge weights, and dynamic technology influence centrality) is divided into time slices to obtain training and testing sets. The training set is input into the graph attention network layer, which outputs spatial feature vectors for each time slice through medical attribute-weighted attention and disease-adapted attention head number. The spatial feature vectors are input into the gated recurrent unit layer, which, combined with time decay weights, outputs temporal feature vectors. The spatiotemporal feature vectors are fused, linearly transformed, and activated through the feature interaction layer to output the predicted technology influence score. The testing set is input into the trained model to output the predicted technology influence score of each hospital node in the future preset time slice, providing data support for subsequent medical service node identification and medical flow prediction.

[0105] In one embodiment of the present invention, step S4 includes the following steps:

[0106] The scoring threshold for the first medical service node is set to the upper quartile of the prediction technology influence scores of all hospital nodes.

[0107] The predicted technical influence score of each hospital node is compared with the score threshold. Nodes with scores higher than the score threshold are directly included in the candidate pool. For the candidate first medical service node, if the number of similar high-level hospitals in its area exceeds the preset density threshold, it is determined to be a resource redundant node and is removed from the first medical service node list to obtain the final first medical service node list.

[0108] Specifically, a regional resource density check is introduced to eliminate redundant nodes and prevent the emergence of too many similar high-level hospitals in a certain area, which would lead to resource consumption and layout imbalance. Specifically, provincial-level administrative divisions or national-level urban agglomerations (Yangtze River Delta, Guangdong-Hong Kong-Macao Greater Bay Area) are usually used as the basic check units. The preset density threshold is pre-set by medical and health planning experts based on factors such as regional population and development goals. For example, it can be set that the number of candidate first medical service nodes in the same area should not exceed 3. For each node in the candidate pool, the system automatically identifies the check area in which it is located. Then, it counts the number of all nodes in the candidate pool in that area. The judgment rule is: if the number exceeds the preset density threshold, the system determines that there is resource redundancy in the area. In this case, the system will retain the top N nodes with the highest predicted technical influence scores in the area (N equals the density threshold), and remove all candidate nodes ranked after Nth from the candidate pool. After optimizing the candidate pool through the above steps, the remaining hospital nodes in the pool constitute the final list of first medical service nodes.

[0109] In one embodiment of the present invention, step S5 includes the following steps:

[0110] Based on the geographical location and predicted technical influence score of the first medical service node, as well as the spatial distribution of the patient's origin, a recommendation network is constructed with the patient's origin and the first medical service node as vertices and potential medical flow as edges. A comprehensive recommendation index is calculated for each edge in the network. Based on the index, a direct recommendation path guided by technical influence is generated for high-risk patients, and K shortest recommendation paths that balance spatial distance and technical influence are generated for medium- and low-risk patients, forming a hierarchical diagnosis and treatment recommendation map.

[0111] Meanwhile, the specific method for optimizing the allocation of regional medical resources is as follows: each first medical service node is mapped to its respective administrative division, the medical resource density of each administrative division is calculated, and through spatial overlay analysis, the first medical service node whose predicted technical influence score is higher than the first preset threshold and whose administrative division medical resource density is lower than the second preset threshold is identified. The area where the node is located is marked as a key area that urgently needs technical support and resource allocation, and a resource allocation recommendation report is output.

[0112] Specifically, a medical treatment recommendation network is constructed as follows: The patient's origin (usually at the prefecture-level administrative region level) and the first medical service node (hospital) are both used as vertices of the network. A directed edge is established between each patient's origin and each first medical service node, pointing from the patient's origin to the hospital, representing a potential medical treatment flow. A comprehensive recommendation index is calculated for each edge in the recommendation network. This index represents the balance between the target hospital's technological attractiveness and spatial accessibility. The comprehensive recommendation index is calculated as follows: ,in, The overall recommendation index is as follows: To predict the score of technological influence, To achieve the maximum score, To standardize spatial distance, and For the corresponding weighting coefficients, where, for high-risk paths, a weighting coefficient is set. > The values ​​are 0.7 and 0.3 respectively; for low- to medium-risk routes, the values ​​are set as follows: The values ​​are all 0.5; differentiated paths are generated for patients with different risk levels, specifically: 1. High-risk patient path: guided by technological influence, for each high-risk patient (based on disease diagnosis codes, a predefined database of difficult and severe diseases is defined. This database is determined by clinical experts according to the "List of Difficult and Complex Diseases" issued by the National Health Commission and clinical consensus, including the subtypes with the worst prognosis and most complex treatment in a specific disease. If the patient's primary diagnosis code belongs to this database, he / she is automatically marked as a high-risk patient), the path with the highest comprehensive recommendation index is directly recommended, that is, it connects to the hospital with the strongest technical strength in the country to form a direct recommendation path; 2. Medium and low-risk patient path: balancing technological influence and spatial distance, the K shortest path algorithm is used. Here, "shortest" is defined as having the highest comprehensive recommendation index. The system calculates K paths with the highest recommendation index for each low-to-medium risk patient (the patient's disease belongs to the target disease category but does not meet the above high-risk criteria). All the above path results are visualized and rendered on an electronic map to form a hierarchical medical treatment recommendation map. This map can highlight the corresponding priority recommended medical flow based on different patient origins and risk levels. In generating a regional medical resource optimization allocation scheme, each first medical service node is mapped to its respective provincial administrative division, and the number density (e.g., number of nodes per ten million people) or spatial density (e.g., number of nodes per unit area) of first medical service nodes in each province is calculated. A first preset threshold is used to screen hospitals with high technical influence, for example, set as the median of the predicted technical influence score of all first medical service nodes. A second preset threshold is used to screen low resource density areas, for example, set as the lower quartile of the national provincial regional medical resource density.

[0113] Please see Figure 2 As shown, this invention is a big data-based cross-regional medical treatment data analysis system, which includes the following modules:

[0114] Multi-source data fusion and preprocessing module: acquires and fuses out-of-town medical treatment data and multi-source external data, preprocesses the out-of-town medical treatment data and multi-source external data to obtain a standardized fusion dataset;

[0115] Time-series medical technology network construction module: Based on the standardized fusion dataset, filter out out-of-town medical treatment records for specific diseases, construct a time-series medical technology network for specific diseases, and calculate the dynamic technology influence centrality of each hospital node in each time slice;

[0116] Spatiotemporal prediction module for technological influence: Construct a spatiotemporal graph neural network model, input the time-series medical technology network into the spatiotemporal graph neural network model, wherein the time-series medical technology network includes hospital nodes, edge weights and dynamic technological influence centrality, and input into the spatiotemporal graph neural network model to output the predicted technological influence score of each hospital node in a future preset time slice;

[0117] Intelligent selection module for medical service nodes: Based on the predicted technical influence score of each hospital node, a score threshold is set, and hospital nodes with predicted technical influence scores higher than the score threshold are selected as the first medical service nodes;

[0118] The hierarchical diagnosis and treatment and resource optimization decision support module generates a hierarchical diagnosis and treatment recommendation path and a regional medical resource optimization allocation scheme based on the first medical service node. The hierarchical diagnosis and treatment recommendation path generates priority medical flow directions for patients with different risk levels based on the spatial distance between the patient's place of origin and the first medical service node and the predicted technology influence score. The regional medical resource optimization allocation scheme identifies areas with high predicted technology influence scores but low local medical resource density and marks them as key areas that urgently need technical support and resource allocation.

[0119] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.

Claims

1. A data analysis method for cross-regional medical treatment based on big data, characterized in that, Includes the following steps: S1: Acquire and merge cross-regional medical treatment data and multi-source external data, preprocess the cross-regional medical treatment data and multi-source external data to obtain a standardized fused dataset; S2: Based on the standardized fusion dataset, filter out medical records of specific diseases in different locations, construct a time-series medical technology network for specific diseases, and calculate the dynamic technical influence centrality of each hospital node in each time slice. S3: Construct a spatiotemporal graph neural network model and input the time-series medical technology network into the spatiotemporal graph neural network model. The time-series medical technology network includes hospital nodes, edge weights, and dynamic technology influence centrality to output the predicted technology influence score of each hospital node in a future preset time slice. S4: Based on the predicted technical influence score of each hospital node, a score threshold is set, and hospital nodes with predicted technical influence scores higher than the score threshold are selected as the first medical service nodes. S5: Based on the first medical service node, generate a hierarchical diagnosis and treatment recommendation path and a regional medical resource optimization allocation scheme for specific diseases. The hierarchical diagnosis and treatment recommendation path generates priority medical flow for patients of different risk levels based on the spatial distance between the patient's place of origin and the first medical service node and the predicted technology influence score. The regional medical resource optimization allocation scheme identifies areas with high predicted technology influence scores but low local medical resource density and marks them as key areas that urgently need technical support and resource allocation. In step S2, constructing a time-series medical technology network for a specific disease includes the following steps: The time-series medical technology network uses hospitals in various regions as network nodes, and the number of patients with specific diseases flowing from the source region to the target hospital within a unit time slice is used as the edge weight between nodes. The edge weights between nodes are specifically as follows: in, For time slices From the source region Flowing to target hospitals The sum of weighted records of patients with specific diseases. The total number of medical records for a specific disease. For summation indexing, representing a single medical record, For the first The final record weight of each record. For the source region per capita disposable income The average per capita disposable income nationwide. Target hospital The technical difficulty adjustment factor Target hospital The proportion of complex and difficult cases of this specific disease received. This represents the percentage of complex and difficult cases of this specific disease nationwide. In step S2, the dynamic technical influence centrality of each hospital node in each time slice is calculated, including the following steps: Centrality of input Specifically: ,in, For time slices Inside, all pointing to the target hospital The sum of the weights of the incoming edges; Proximity Centrality Specifically: ,in, Target hospital From the network to other hospitals The shortest weighted path length; Authority centrality is calculated iteratively based on the HITS algorithm to obtain the time slice of each node. Authority Centrality ; The weights corresponding to the input centrality, proximity centrality, and authority centrality are determined based on the entropy weight method, and then weighted and added together to obtain the dynamic technical influence centrality.

2. The method for analyzing cross-regional medical treatment data based on big data according to claim 1, characterized in that, In step S1, preprocessing includes outlier removal, missing value imputation, and data standardization. Specifically, missing value imputation employs a weighted mean imputation method based on disease incidence rates. The dimensions of the missing data to be filled are determined, including the number of out-of-town medical visits, the average length of hospital stay, and the average out-of-pocket cost per visit. Non-missing medical data is then filtered out. The The data must meet the following criteria: the data belongs to the same disease category as the missing data, and the per capita disposable income of the source region is at the same level as that of the source region of the missing data; Get each Disease incidence rate in corresponding regions Calculate the imputation values ​​for missing data, specifically: in, Fill in the missing data with values. This represents the number of samples with non-missing data. An index identifier for a single, non-missing medical data sample that meets the screening criteria.

3. The method for analyzing cross-regional medical treatment data based on big data according to claim 1, characterized in that, In step S2, screening out-of-town medical records for specific diseases includes the following steps: Construct a multi-dimensional core technology feature profile of the target disease, the profile including a diagnostic dimension composed of core disease diagnosis codes and a treatment dimension composed of key treatment operation codes; From the standardized fusion dataset, records whose main diagnostic codes belong to the diagnostic dimension or whose main operational codes belong to the treatment dimension are selected as the initial set of selected records; Based on the pre-trained BERT model, the diagnostic text summary and operation name of each record in the initial screening record set are semantically embedded, and the cosine similarity between the summary and operation name and the core technology feature vector of the target disease is calculated as the probability of technology relevance. A probability threshold is set to include records with a probability of technical relevance higher than the threshold in the final screening record set. Each initially screened record is assigned a final record weight that is positively correlated with the probability of technical relevance, so as to form a weighted set of out-of-town medical treatment records for specific diseases.

4. The method for analyzing cross-regional medical treatment data based on big data according to claim 1, characterized in that, The method of determining weights based on entropy weight includes the following steps: Entropy weighting is used for each time slice. The weights are calculated dynamically, specifically as follows: For time slices This forms a matrix consisting of the three centrality index values ​​of all hospital nodes, and the information entropy of each index is calculated. ,in, For nodes In indicators The proportion of the surface; The weight of each indicator is calculated as follows: ,in, For the first The weights corresponding to each indicator.

5. The method for analyzing cross-regional medical treatment data based on big data according to claim 1, characterized in that, In step S3, constructing the spatiotemporal graph neural network model includes the following steps: The spatiotemporal graph neural network model includes a graph attention network layer, a gated recurrent unit layer, and a feature interaction layer; The graph attention network layer takes the dynamic technical influence centrality of hospital nodes and the edge weights between nodes as core inputs. After concatenation with the basic features of the nodes, attention coefficients are generated through the LeakyReLU activation function. A dynamic attention head number rule is introduced, and the results of the multi-head attention mechanism are concatenated to output the spatial feature vector of each hospital node in the corresponding time slice. The gated recurrent unit layer assigns corresponding time decay weights to the spatial feature vectors of different time slices according to the time slice type of the time-series medical technology network, integrates the time decay weights into the calculation of the update gate and reset gate of GRU, and outputs the time feature vector of each hospital node. The feature interaction layer uses element-wise multiplication to fuse the spatial feature vector and the temporal feature vector to obtain a spatiotemporal fusion feature vector. After linear transformation, it is processed by the ReLU activation function to obtain the prediction technology impact score of each hospital node in the future preset time slice.

6. The method for analyzing cross-regional medical treatment data based on big data according to claim 1, characterized in that, Step S4 includes the following steps: The scoring threshold for the first medical service node is set to the upper quartile of the prediction technology influence scores of all hospital nodes. The predicted technical influence score of each hospital node is compared with the score threshold. Nodes with scores higher than the score threshold are directly included in the candidate pool. For the candidate first medical service node, if the number of similar high-level hospitals in its area exceeds the preset density threshold, it is determined to be a resource redundant node and is removed from the first medical service node list to obtain the final first medical service node list.

7. The method for analyzing cross-regional medical treatment data based on big data according to claim 1, characterized in that, Step S5 includes the following steps: Based on the geographical location and predicted technical influence score of the first medical service node, as well as the spatial distribution of the patient's origin, a recommendation network is constructed with the patient's origin and the first medical service node as vertices and potential medical flow as edges. A comprehensive recommendation index is calculated for each edge in the network. Based on the index, a direct recommendation path guided by technical influence is generated for high-risk patients, and K shortest recommendation paths that balance spatial distance and technical influence are generated for medium- and low-risk patients, forming a hierarchical diagnosis and treatment recommendation map. Meanwhile, the specific method for optimizing the allocation of regional medical resources is as follows: each first medical service node is mapped to its respective administrative division, the medical resource density of each administrative division is calculated, and through spatial overlay analysis, the first medical service node whose predicted technical influence score is higher than the first preset threshold and whose administrative division medical resource density is lower than the second preset threshold is identified. The area where the node is located is marked as a key area that urgently needs technical support and resource allocation, and a resource allocation recommendation report is output.

8. A big data-based cross-regional medical treatment data analysis system, used to implement the big data-based cross-regional medical treatment data analysis method according to any one of claims 1-7, characterized in that, Includes the following modules: Multi-source data fusion and preprocessing module: acquires and fuses out-of-town medical treatment data and multi-source external data, preprocesses the out-of-town medical treatment data and multi-source external data to obtain a standardized fusion dataset; Time-series medical technology network construction module: Based on the standardized fusion dataset, filter out out-of-town medical treatment records for specific diseases, construct a time-series medical technology network for specific diseases, and calculate the dynamic technology influence centrality of each hospital node in each time slice; Spatiotemporal prediction module for technological influence: Construct a spatiotemporal graph neural network model and input the time-series medical technology network into the spatiotemporal graph neural network model. The time-series medical technology network includes hospital nodes, edge weights, and dynamic technological influence centrality to output the predicted technological influence score of each hospital node in a future preset time slice. Intelligent selection module for medical service nodes: Based on the predicted technical influence score of each hospital node, a score threshold is set, and hospital nodes with predicted technical influence scores higher than the score threshold are selected as the first medical service nodes; The hierarchical diagnosis and treatment and resource optimization decision support module generates a hierarchical diagnosis and treatment recommendation path and a regional medical resource optimization allocation scheme based on the first medical service node. The hierarchical diagnosis and treatment recommendation path generates priority medical flow directions for patients with different risk levels based on the spatial distance between the patient's place of origin and the first medical service node and the predicted technology influence score. The regional medical resource optimization allocation scheme identifies areas with high predicted technology influence scores but low local medical resource density and marks them as key areas that urgently need technical support and resource allocation.