Medical institution name matching method, system and device and storage medium
By combining a large language model with geo-fencing, the problem of inconsistent naming in matching medical institution names is solved, and automated matching with high precision and high recall rate is achieved, which has adaptive capabilities and reduces manual intervention.
Patent Information
- Application Number
- CN202511179814.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-08-22
AI Technical Summary
Existing technologies have data silos caused by inconsistent naming in medical institution name matching, resulting in low precision and recall, lack of semantic relationship understanding and context awareness, and poor model adaptability.
A pre-trained large language model is used for semantic analysis, combined with geo-fencing and confidence fusion models, the search radius is dynamically adjusted, weighted similarity and structured arbitration results are used for multi-level matching decisions, and the model is optimized through human-computer collaborative feedback.
It achieves high-precision and high-recall matching of medical institution names, can identify multiple semantic relationships, has adaptive capabilities, and reduces the workload of manual review.
Smart Images

Figure CN120687595A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a medical institution name matching method, system, device and storage medium. Background Art
[0002] In the development of modern medical information technology, data interconnection and interoperability are the cornerstones for achieving intelligent medical insurance cost control, public health monitoring, clinical research, and improving medical service efficiency. Medical data sources are diverse, encompassing electronic health record (EHR) systems, hospital information systems (HIS), regional health information platforms, medical insurance settlement databases, and commercial map point of interest (POI) databases.
[0003] However, a key technical bottleneck in integrating and analyzing these heterogeneous data sources is the inconsistent naming of the same medical institution entity across different databases. For example, a hospital might have multiple names, including its official full name, a commonly used abbreviation, a local nickname, and names that include campuses or departments. This naming diversity makes it impossible to effectively link data records referring to the same entity, creating "data silos" and severely hindering the realization of the value of medical big data.
[0004] Existing technology uses fuzzy matching methods based on string similarity algorithms. A common implementation involves preprocessing the two name strings to be matched (e.g., removing punctuation and converting them to equal case). Then, a similarity score is calculated using algorithms such as Levenshtein distance, Jaro-Winkler distance, or N-gram. Finally, a comparison is made with a preset threshold (e.g., 0.85) to determine whether the two names match. However, existing solutions suffer from the following inherent drawbacks: Semantic confusion and low accuracy: Existing methods are usually based on character matching and cannot understand the semantic connotation of words. Therefore, they are prone to mistakenly identify names such as "Z City First People's Hospital" and "Z Province First People's Hospital" as matches, which have highly similar character strings but belong to different entities, resulting in low matching accuracy.
[0005] Ignoring aliases leads to low recall: When the literal difference between the official full name and the commonly used abbreviation is huge, the string similarity score is extremely low, resulting in a low recall rate.
[0006] Lack of contextual awareness: Traditional methods isolate key contextual information such as geographic location. Two organizations with similar names that are geographically far apart are virtually impossible to be the same entity, but existing technologies cannot leverage this prior knowledge to aid in identification.
[0007] Lack of relationship recognition capability: This method can usually only provide a binary conclusion of "match" or "mismatch" and cannot identify more fine-grained hierarchical relationships such as "main hospital-branch hospital" or "institution-subordinate department".
[0008] The model is static and has poor adaptability: its matching logic and parameters are usually fixed, lacking the ability to learn from new data and self-optimize, and is difficult to adapt to emerging naming methods and complex scenarios. Summary of the Invention
[0009] In response to the above problems, the present invention provides a medical institution name matching method, system and storage medium, aiming to overcome the technical defects of the existing technology that rely solely on surface string similarity, resulting in low precision and recall rates, inability to understand semantic relationships, lack of contextual information utilization, and static model and poor adaptability.
[0010] According to a first aspect of an embodiment of the present disclosure, a method for matching medical institution names is provided, the method comprising the following steps: Receiving a source medical institution name string, performing semantic component analysis on the name string using a pre-trained large language model to identify an identification name representing the entity's core identity and at least one common name describing the entity's attributes; and assigning different preset weights to the identification name and the common name; Obtaining the geographic coordinates of the source medical institution, and dynamically determining a search radius based on at least one entity attribute of the source medical institution, constructing a geo-fence with the geographic coordinates as the center and the search radius, and searching within the geo-fence to obtain a set of candidate medical institutions; Calculate the weighted similarity score between the source medical institution and any candidate medical institution in the candidate medical institution set. When the weighted similarity score falls within the preset fuzzy interval, determine the semantic relationship between the source medical institution and the candidate medical institution, and output a structured arbitration result including the relationship type and relationship confidence. A confidence fusion model is constructed to fuse the weighted similarity scores and structured arbitration results, calculate the final matching confidence, and make matching decisions based on the final matching confidence.
[0011] In some embodiments, the weight assigned to a recognized name is higher than the weight assigned to a common name.
[0012] In some embodiments, the method further includes: when the final matching confidence does not reach a preset decision threshold, manually reviewing the source medical institution and the candidate medical institution, and retraining the confidence fusion model and / or the large language model based on the matching results obtained by the manual review.
[0013] In some embodiments, the confidence fusion model is a Bayesian inference model, which uses the weighted similarity score as the prior probability and the structured arbitration result as the observation evidence to calculate the posterior probability as the final matching confidence.
[0014] In some embodiments, the confidence fusion model is a supervised learning classifier, and the weighted similarity score and the structured arbitration result are input into a pre-trained supervised learning classifier to output a final matching probability.
[0015] In some embodiments, a pre-trained large language model or a medical field knowledge graph is used to determine the semantic relationship between the source medical institution and the candidate medical institutions.
[0016] According to a second aspect of an embodiment of the present disclosure, a medical institution name matching system is provided, the system comprising: A name structuring and weighting module is configured to receive a source medical institution name string, perform semantic component analysis on the name string using a pre-trained large language model, identify an identification name representing the entity's core identity, and at least one common name describing the entity's attributes; and assign different preset weights to the identification name and the common name; a dynamic geo-fence candidate set generation module, configured to obtain the geographic coordinates of a source medical institution, and dynamically determine a search radius based on at least one entity attribute of the source medical institution, construct a geo-fence with the geographic coordinates as the center and the search radius, and search within the geo-fence to obtain a set of candidate medical institutions; The multi-level matching and semantic relationship determination module is used to calculate the weighted similarity score between the source medical institution and any candidate medical institution in the candidate medical institution set. When the weighted similarity score falls within a preset fuzzy interval, the semantic relationship between the source medical institution and the candidate medical institution is determined, and a structured arbitration result including the relationship type and relationship confidence is output; The confidence fusion and decision module is used to build a confidence fusion model, fuse the weighted similarity score and the structured arbitration result, calculate the final matching confidence, and make matching decisions based on the final matching confidence.
[0017] In some embodiments, the system also includes a human-computer collaborative feedback and optimization module, which is used to manually review the source medical institution and the candidate medical institution when the final matching confidence does not reach a preset decision threshold, and retrain the confidence fusion model and / or large language model based on the matching results obtained by manual review.
[0018] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned medical institution name matching method when executing the program.
[0019] According to a fourth aspect of an embodiment of the present disclosure, a non-temporary computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the steps of the above-mentioned medical institution name matching method are implemented.
[0020] The embodiments of the present disclosure provide a medical institution name matching method, system, and storage medium, which are designed to achieve high-precision, high-recall, semantically aware, and adaptive automated matching of medical institution names. The beneficial effects include: Accuracy is significantly improved: By deeply integrating the semantic weight of the name with geographic spatial information, mismatches and missed matches caused by "similar names but distant geographical locations" or "large differences in names but close geographical locations and consistent core words" are effectively avoided, greatly improving the matching precision and recall rate.
[0021] Achieve semantic relationship understanding: When the weighted similarity score falls into the preset fuzzy interval, the semantic relationship between the source medical institution and the candidate medical institution is determined, and a structured arbitration result including the relationship type and relationship confidence is output. It can go beyond the binary matching of "yes / no" and identify various types of semantic relationships such as "same entity", "alias", "branch / campus", etc., providing richer entity relationship dimension information for downstream data applications.
[0022] Robustness and adaptability: Dynamic geofencing can intelligently adapt to the distribution characteristics of institutions at different levels; the Bayesian reasoning framework provides a unified probabilistic interpretation framework for the integration of multi-source evidence, and provides a clear expansion path for integrating more dimensions of evidence in the future to achieve model optimization.
[0023] High degree of automation, cost reduction and efficiency improvement: Internalizing complex expert judgment logic into an end-to-end automated process greatly reduces the workload of manual review and intervention, and significantly reduces the time and labor costs of medical data governance.
[0024] Possessing the ability to continuously learn and self-optimize: Through a collaborative human-machine feedback loop, the system can submit ambiguous cases where the model has difficulty making decisions for manual review, and feed the review results back to the model as high-quality training data, allowing it to continuously evolve during continued use, dynamically adapt to new naming methods and complex scenarios, and achieve long-term self-optimization of the model.
[0025] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0027] Figure 1 This is a flow chart of a medical institution name matching method according to an embodiment of the present invention; Figure 2 1 is a flowchart of a method for name structuring and semantic weighting according to an embodiment of the present invention; Figure 3 is a flowchart of a method for generating a dynamic geo-fence candidate set in an embodiment of the present invention; Figure 4 1 is a flow chart of a method for multi-level matching and intelligent arbitration and confidence fusion in an embodiment of the present invention; Figure 5 1 is a flow chart of a method for feedback and optimization of human-machine collaboration in an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of the medical institution name matching system in an embodiment of the present invention; Figure 7 It is a schematic diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.
[0029] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0030] The present invention provides a medical institution name matching method, system, electronic device, and computer-readable storage medium that integrate semantic weights, dynamic geofencing, and a probabilistic model. The method aims to achieve high-precision, high-recall, semantically aware, and adaptive automated matching of medical institution names. The following are specific embodiments: A method for matching medical institution names, comprising the following steps: S100, name structuring and semantic weighting step: receiving a source medical institution name string, performing semantic component analysis on the name string using a pre-trained large language model to identify an identification name representing the entity's core identity and at least one common name describing the entity's attributes; and assigning different preset weights to the identification name and the common name, preferably, the weight assigned to the identification name is higher than the weight assigned to the common name; S200, a step of generating a dynamic geo-fence candidate set: obtaining the geographic coordinates of a source medical institution, and dynamically determining a search radius based on at least one entity attribute (preset level) of the source medical institution, constructing a dynamic geo-fence, constructing a geo-fence with the geographic coordinates as the center and the search radius, and searching within the geo-fence to obtain a set of candidate medical institutions; S300, multi-level matching and intelligent arbitration step: Calculate a weighted similarity score between the source medical institution and any candidate medical institution in the candidate medical institution set. When the weighted similarity score falls within a preset fuzzy interval, trigger the intelligent arbitration module to determine the semantic relationship between the source medical institution and the candidate medical institution, and output a structured arbitration result including the relationship type and relationship confidence. S400, confidence fusion and decision-making step: Construct a confidence fusion model to fuse the weighted similarity score and the structured arbitration result, calculate the final matching confidence, and make a matching decision based on the final matching confidence.
[0031] In some preferred embodiments, Figure 1 As shown, the method also includes S500, a human-machine collaborative feedback and optimization step: when the final matching confidence does not reach a preset decision threshold, the source medical institution and the candidate medical institution are manually reviewed, and the confidence fusion model and / or the large language model are retrained based on the matching results obtained by the manual review.
[0032] In some preferred embodiments, the confidence fusion model is a Bayesian inference model, which uses the weighted similarity score as the prior probability and the structured arbitration result as the observation evidence to calculate the posterior probability as the final matching confidence.
[0033] In some preferred embodiments, the confidence fusion model is a supervised learning classifier, and the weighted similarity score and the structured arbitration result are input into a pre-trained supervised learning classifier to output a final matching probability.
[0034] In some preferred embodiments, a pre-trained large language model or a medical field knowledge graph is used to determine the semantic relationship between the source medical institution and the candidate medical institutions.
[0035] In a specific embodiment, the following steps may be included: Step S100: Name structuring and semantic weighting: Figure 2 As shown, this step aims to convert the unstructured name string into a structured data object containing semantic and importance information, laying the foundation for subsequent weighted matching. In an exemplary embodiment: 301. The system receives a source medical institution name string to be matched, such as "YY Hospital East Branch"; 302 and 303. The processor in the system calls a pre-trained language model, which can be a general-purpose large language model (for example, but not limited to the GPT series, LLaMA series), or a specialized model that has been fine-tuned in the medical field (for example, but not limited to PULSE, BioBERT, etc.). Through carefully designed prompt engineering, the model is instructed to perform the semantic component parsing task. For example, the prompt can be designed as: "Please parse the following medical institution name into three parts: 'identification name' (a word representing the core identity of the entity), 'common name' (a word describing the type, campus, department, etc.), and 'administrative division' (a geographical location word). Input: 'YY Hospital East Campus'". After receiving the instruction, the model analyzes and outputs a structured result: 304. Identify the name part: This refers to the core words in the name that are most distinctive and can identify the entity. For example, for "YY Hospital East Branch", the model should recognize "YY"; 305. Common name part: refers to common words that describe the type, level, and nature of the institution. For "YY Hospital East Branch", the model should recognize "hospital" and "East Branch"; Administrative division part: refers to the geographical location information explicitly included in the name. In this example, it is empty, but for "Z City H District People's Hospital", it is "Z City H District"; 306. Next, the system assigns different preset weights to the different parsed parts. For example, a higher weight is assigned to the "identification name part" (for example, the weight value is in the range of [0.7, 0.9]), and a lower weight is assigned to the "common name part" (for example, the weight value is in the range of [0.1, 0.3]). These weights reflect the importance of different words in determining entity identity. The weight values can be obtained by training and learning using a grid search or gradient optimization method on a labeled verification dataset, and the optimization goal is to maximize the F1 score (F1-Score) of the final matching task; 307. Finally, a structured object is generated, such as {“identification name”: “YY”, “common name”: [“hospital”, “East Hospital”], “weight”: {“identification name”: 0.8, “common name”: 0.2}}.
[0036] Step S200: Dynamic geo-fence candidate set generation: Figure 3 As shown in Figure 3, this step uses geospatial proximity as strong prior knowledge to greatly narrow the matching scope and filter out the vast majority of geographically irrelevant candidates, thereby improving matching efficiency and accuracy.
[0037] 401. The system obtains the geographic coordinates of the source medical institution. The acquisition methods may include: direct extraction: if the source data record already contains the latitude and longitude fields, use them directly; address resolution: if the source data only has a text address (such as "No. 1 Shuaifuyuan, H District, Z City"), call the geographic information service API (such as the API of Gaode Map and Baidu Map) to convert the text address into accurate latitude and longitude coordinates; keyword search: if the source data only has a name, call the location search API of the geographic information service, search with the name as the keyword, and obtain the most likely geographic coordinates; then, the system builds a dynamic geographic fence. The "dynamic" here is reflected in the fact that the search radius is not a fixed value, but is adjusted according to one or more preset attributes of the source medical institution. These attributes may include: 403 and 405, Institutional Level: Determined by keywords in the name (e.g., "Central Hospital," "Grade 3A," and "University Affiliated" are considered high-level) or associated external knowledge base information. For institutions determined to be high-level, use a larger search radius (e.g., 500 to 2000 meters); 402 and 404: For institutions determined to be low-level (such as "XX Community Health Service Station"), use a smaller search radius (for example, 100 meters to 300 meters).
[0038] Geographical characteristics: For example, the population density or medical resource density of the area where the geographic coordinates are located. In densely populated urban areas, a smaller radius can be used; in sparsely populated suburbs or rural areas, a larger radius can be used; The specific value of the search radius may be derived based on statistical analysis of a large amount of geographic information point of interest (POI) data. For example, the search radius may be set to a minimum radius that covers 99% of known branches or associated facilities of the same medical entity, and may be adjusted based on the region type. 406. Finally, the system uses the acquired geographic coordinates as the center and performs a circular area search with a dynamically determined radius, and selects all candidate POIs of the type "health care services" from a standardized geographic information database to form a highly geographically relevant candidate set.
[0039] Step S300: Multi-level matching and intelligent arbitration: Figure 4 As shown in the figure, this step is the core matching decision process, which adopts a multi-level sequential strategy from certainty to fuzziness and from simple to complex, and innovatively introduces an intelligent arbitration mechanism to handle difficult situations.
[0040] 501. For the source organization and each candidate POI in the candidate set, the system performs the following determinations in sequence: 502. Exact match determination: First, perform the simplest string complete consistency comparison. If the official name of the candidate POI is exactly the same as the source institution name string, it is directly determined to be a match, the match confidence is set to the highest value of 1.0, and the subsequent matching process for the candidate POI is terminated; 504. Weighted fuzzy matching determination: If the exact match fails, weighted fuzzy matching is initiated. This algorithm (which can be implemented as a weighted Jaro-Winkler algorithm or a token-based weighted Jaccard similarity algorithm, for example) uses the structured names and weights obtained in S100 when calculating similarity. For example, when comparing the source institution "YY Hospital East Branch" with the candidate institution "NN Academy of Medical Sciences YY Hospital": (1) The "identifying name" part ("YY") of the two matches exactly, contributing a score of 1.0 * weight 0.8 = 0.8; (2) The “common name” part of the two ([“hospital”, “East Hospital”] vs. [“hospital”]) partially matches, and the score is contributed by 0.5 * weight 0.2 = 0.1 according to the matching degree; (3) The final weighted similarity score is 0.8 + 0.1 = 0.9; 505. Intelligent Arbitration Triggering and Execution: The system presets a fuzzy interval, such as [0.70, 0.95]. The upper and lower thresholds of this interval are determined by analyzing the precision-recall curve on the test dataset to determine the optimal operating point to balance the computational overhead of invoking the intelligent arbitration module with the final model performance. 506. When the calculated weighted similarity score falls within this range, the intelligent arbitration module is triggered; In a preferred embodiment, the semantic relationship determination module is a pre-trained large language model (LLM). The system interacts with the LLM through a carefully designed prompt, setting an expert role for it and providing contextual information. For example: Prompt: "You are a medical data governance expert. Based on the following information, please determine whether the 'source name' and 'candidate name' point to the same core medical entity. Select the most appropriate relationship from ['same entity', 'alias', 'subordinate institution', 'unrelated entity'] and give a confidence level between 0 and 1. Please return the result in JSON format. Source name: 'YY Hospital East Branch', Candidate name: 'NN Academy of Medical Sciences YY Hospital'"; 507. LLM reasoning based on its huge knowledge base may return the following structured arbitration result: {“relation”: “subordinate organization”, “confidence”: 0.98}.
[0041] It should be noted that the large language model described in this invention is preferably used as a semantic reasoning component, and its output, the structured arbitration results, serve as intermediate evidence and are integrated with other evidence such as weighted similarity within a unified fusion framework. This approach leverages the reasoning capabilities of the LLM while avoiding the risks associated with its potential instability as the sole decision maker, thereby ensuring the robustness and explainability of the final decision.
[0042] Step S400: Decision based on confidence fusion model: Figure 4 As shown in the logical flow, this step aims to integrate evidence from different information sources (weighted text similarity, LLM expert judgment) in a unified and well-interpretable probabilistic framework, thereby deriving a final matching confidence that is more reliable than any single piece of evidence.
[0043] like Figure 4 In the logical flow shown, 508, the present invention constructs a confidence fusion model for fusing the weighted similarity score and the structured arbitration result to calculate the final matching confidence. In a preferred embodiment, the confidence fusion model is a Bayesian reasoning model. Specifically, the Bayesian formula is applied: , The meanings of each part are as follows: Prior probability P(Match): represents the initial belief about the possibility of two names matching before obtaining the strong evidence of intelligent arbitration. This value can be obtained from the weighted similarity score calculated in step S300 It is obtained by mapping through a preset mapping function (such as Sigmoid function), that is, , where k and b are pre-calibrated mapping function parameters.
[0044] Observational evidence: refers to the structured arbitration result output by the semantic relationship determination module in S300, for example, {“relation”: “subordinate organization”, “confidence”: 0.98}.
[0045] Likelihood (P) (Evidence | Match): This value represents the probability of the intelligent arbitration module's current judgment, assuming that the two names are indeed strongly correlated (matched). This value can be derived directly or by weighting the confidence returned in the arbitration result.
[0046] Posterior probability P(Match | Evidence): is the final quantity to be solved, representing the updated belief that the two names match after obtaining the expert opinion (evidence) of the intelligent arbitration.
[0047] 509. After calculating the posterior probability, in practical applications, it is necessary to normalize the posterior probabilities of all possible results (such as "match" and "mismatch") so that their sum is 1. This normalized posterior probability is used as the final match confidence.
[0048] Finally, the system compares this final match confidence with a final decision threshold (e.g., 0.90). If the confidence is higher than or equal to the threshold, the system automatically makes the final decision. If the confidence is lower than the threshold, the next step is triggered.
[0049] Step S500: Feedback and optimization of human-machine collaboration: Figure 5 As shown in Figure 2, this step aims to handle difficult cases that the automated model cannot be sure of, and to enable continuous learning of the system by introducing expert knowledge.
[0050] 601. Triggering and Presenting: When the final matching confidence calculated in step S400 is lower than the decision threshold, the system automatically triggers the manual review process and pushes the matching pairs to be reviewed and related auxiliary information to a graphical user interface (GUI). Figure 5 As shown in the interactive interface in , the interface clearly displays information such as 603, the source institution name, 604, the candidate institution name, the geographical location, 605, the weighted similarity score, and the suggestions given by the intelligent arbitration module to assist experts in making judgments.
[0051] 606. Manual Review and Adjudication: The review expert makes the final decision on the relationship between the two institutions based on the interface information and their own knowledge.
[0052] 607 and 608, Model Update and Optimization: After the system receives the authoritative matching results given by the experts, it uses them as a high-quality, labeled training sample. Figure 5 As shown in the data flow in , after accumulating a sufficient number of manual review samples, the system can periodically start the model retraining task to fine-tune the semantic analysis model in step S100 or retrain the confidence fusion model in step S400.
[0053] Through this closed-loop feedback mechanism, the method of the present invention can learn from uncertainty and achieve continuous self-iteration and optimization of the model.
[0054] For some steps of the above specific embodiment, there are the following alternatives: Alternative to step S100 (name structuring and semantic weighting): This step can also be implemented using a rule-based and dictionary-based approach. The system pre-builds a "common name dictionary" (containing terms such as "hospital," "center," "outpatient department," "branch," and "campus") and an "administrative division dictionary" (containing the names of provinces, cities, and districts nationwide). Using a string matching and elimination algorithm, the words found in the dictionary are removed from the original name, and the remaining part is considered the recognized name. This approach is simple to implement and has low computational cost.
[0055] An alternative to step S300 (multi-level matching and intelligent arbitration): The semantic relationship determination module can also be a medical knowledge graph. When arbitration is triggered, the system queries the graph to see if there are predefined relationships between the entity nodes representing two names, such as sameAs, alias, or subOrganizationOf, and uses this as the arbitration result. This solution provides extremely high accuracy.
[0056] As an alternative to step S400 (decision-making based on the confidence fusion model): The confidence fusion module can also be a supervised learning classifier (such as logistic regression or gradient boosting tree (XGBoost)). The weighted string score, geographic distance, relationship type (one-hot encoded) output by the LLM, and confidence are used as features and fed into a pre-trained classifier, which then directly outputs the final match probability. While this approach requires a large amount of labeled data for training, it may achieve higher performance.
[0057] Another embodiment is used to illustrate a medical institution name matching system, such as Figure 6 As shown, the system includes: The name structuring and weighting module 10 is configured to receive a source medical institution name string, perform semantic component analysis on the name string using a pre-trained large language model, identify an identification name representing the entity's core identity, and at least one common name describing the entity's attributes; and assign different preset weights to the identification name and the common name; a dynamic geo-fence candidate set generation module 20, configured to obtain the geographic coordinates of a source medical institution, and dynamically determine a search radius based on at least one entity attribute of the source medical institution, construct a geo-fence with the geographic coordinates as the center and the search radius, and search within the geo-fence to obtain a set of candidate medical institutions; The multi-level matching and semantic relationship determination module 30 is used to calculate the weighted similarity score between the source medical institution and any candidate medical institution in the candidate medical institution set. When the weighted similarity score falls within a preset fuzzy interval, the semantic relationship between the source medical institution and the candidate medical institution is determined, and a structured arbitration result including the relationship type and relationship confidence is output; The confidence fusion and decision module 40 is used to build a confidence fusion model for fusing the weighted similarity score and the structured arbitration result, calculating the final matching confidence, and making a matching decision based on the final matching confidence.
[0058] In a preferred embodiment, the system also includes a human-computer collaborative feedback and optimization module 50, which is used to manually review the source medical institution and the candidate medical institution when the final matching confidence does not reach the preset decision threshold, and retrain the confidence fusion model and / or the large language model based on the matching results obtained by manual review.
[0059] In addition to the above modules, the system may also include other components, such as Figure 6 As shown, its specific implementation is the same as the above-mentioned method embodiment. In addition, some components are irrelevant to the content of the embodiment of the present disclosure, so their illustration and description are omitted here.
[0060] The other specific working processes of the medical institution name matching system refer to the description of the above-mentioned medical institution name matching method embodiment and will not be repeated here.
[0061] Another embodiment is used to illustrate that the system of the present invention can also be used with the help of Figure 7 The architecture of the computing device shown is implemented. Figure 7 The architecture of the computing device is shown in FIG. Figure 7 As shown, a computer system 710, a system bus 730, one or more CPUs 740, an input / output 720, a memory 750, etc. The memory 750 can store various data or files used for computer processing and / or communication, as well as program instructions executed by the CPU, including the medical institution name matching method of the embodiment. Figure 7 The architecture shown is only exemplary and may be adjusted based on actual needs when implementing different devices. Figure 7 One or more components in the system. The memory 750, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, such as the program instructions / modules corresponding to the medical institution name matching method in the embodiment of the present invention (for example, the name structuring and weighting module 10, the dynamic geo-fence candidate set generation module 20, the multi-level matching and semantic relationship determination module 30, the confidence fusion and decision module 40 and the human-computer collaborative feedback and optimization module 50 in the medical institution name matching system). One or more CPUs 740 execute various functional applications and data processing of the system of the present invention by running the software programs, instructions and modules stored in the memory 750, that is, to implement the above-mentioned medical institution name matching method, which includes the following steps: Receiving a source medical institution name string, performing semantic component analysis on the name string using a pre-trained large language model to identify an identification name representing the entity's core identity and at least one common name describing the entity's attributes; and assigning different preset weights to the identification name and the common name; Obtaining the geographic coordinates of the source medical institution, and dynamically determining a search radius based on at least one entity attribute of the source medical institution, constructing a geo-fence with the geographic coordinates as the center and the search radius, and searching within the geo-fence to obtain a set of candidate medical institutions; Calculate the weighted similarity score between the source medical institution and any candidate medical institution in the candidate medical institution set. When the weighted similarity score falls within the preset fuzzy interval, determine the semantic relationship between the source medical institution and the candidate medical institution, and output a structured arbitration result including the relationship type and relationship confidence. A confidence fusion model is constructed to fuse the weighted similarity scores and structured arbitration results, calculate the final matching confidence, and make matching decisions based on the final matching confidence.
[0062] Of course, the processor of the server provided by the embodiment of the present invention is not limited to executing the method operations described above, but can also execute related operations in the medical institution name matching method provided by any embodiment of the present invention.
[0063] The memory 750 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data created based on the use of the terminal, etc. Furthermore, the memory 750 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some instances, the memory 750 may further include memory remotely located relative to one or more CPUs 740, and these remote memories may be connected to the device via a network. Examples of the aforementioned networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0064] The input / output 720 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The input / output 720 may also include a display device such as a display screen.
[0065] Embodiments of the present invention also provide a non-transitory computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the medical institution name matching method described in the above embodiments. The computer-readable storage medium of the embodiments of the present invention may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0066] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0067] The program code contained on the storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0068] In addition, other specific working processes of a non-transitory computer-readable storage medium refer to the description of the above-mentioned medical institution name matching method embodiment and are not repeated here.
[0069] In this document, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a step or method that comprises a series of elements includes not only those elements, but also includes other elements not expressly listed, or also includes elements inherent to such step or method.
[0070] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A medical institution name matching method, characterized in that: The method comprises the following steps: Receiving a source medical institution name string, performing semantic component analysis on the name string using a pre-trained large language model to identify an identification name representing the entity's core identity and at least one common name describing the entity's attributes; and assigning different preset weights to the identification name and the common name; Obtaining the geographic coordinates of the source medical institution, and dynamically determining a search radius based on at least one entity attribute of the source medical institution, constructing a geo-fence with the geographic coordinates as the center and the search radius, and searching within the geo-fence to obtain a set of candidate medical institutions; Calculate the weighted similarity score between the source medical institution and any candidate medical institution in the candidate medical institution set. When the weighted similarity score falls within the preset fuzzy interval, determine the semantic relationship between the source medical institution and the candidate medical institution, and output a structured arbitration result including the relationship type and relationship confidence. A confidence fusion model is constructed to fuse the weighted similarity scores and structured arbitration results, calculate the final matching confidence, and make matching decisions based on the final matching confidence.
2. The medical institution name matching method according to claim 1, characterized in that: The weight assigned to distinguished names is higher than the weight assigned to common names.
3. The medical institution name matching method according to claim 1, characterized in that: The method also includes: when the final matching confidence does not reach a preset decision threshold, manually reviewing the source medical institution and the candidate medical institution, and retraining the confidence fusion model and / or the large language model based on the matching results obtained by the manual review.
4. The medical institution name matching method according to claim 1, characterized in that: The confidence fusion model is a Bayesian inference model, which uses the weighted similarity score as the prior probability and the structured arbitration result as the observation evidence to calculate the posterior probability as the final matching confidence.
5. The medical institution name matching method according to claim 1, characterized in that: The confidence fusion model is a supervised learning classifier, which inputs the weighted similarity score and the structured arbitration result into a pre-trained supervised learning classifier and outputs the final matching probability.
6. The medical institution name matching method according to claim 1, characterized in that: Use pre-trained large language models or medical field knowledge graphs to determine the semantic relationship between the source medical institution and the candidate medical institutions.
7. A medical institution name matching system, characterized in that: The system comprises: A name structuring and weighting module is configured to receive a source medical institution name string, perform semantic component analysis on the name string using a pre-trained large language model, identify an identification name representing the entity's core identity, and at least one common name describing the entity's attributes; and assign different preset weights to the identification name and the common name; a dynamic geo-fence candidate set generation module, configured to obtain the geographic coordinates of a source medical institution, and dynamically determine a search radius based on at least one entity attribute of the source medical institution, construct a geo-fence with the geographic coordinates as the center and the search radius, and search within the geo-fence to obtain a set of candidate medical institutions; The multi-level matching and semantic relationship determination module is used to calculate the weighted similarity score between the source medical institution and any candidate medical institution in the candidate medical institution set. When the weighted similarity score falls within a preset fuzzy interval, the semantic relationship between the source medical institution and the candidate medical institution is determined, and a structured arbitration result including the relationship type and relationship confidence is output; The confidence fusion and decision module is used to build a confidence fusion model, fuse the weighted similarity score and the structured arbitration result, calculate the final matching confidence, and make matching decisions based on the final matching confidence.
8. The medical institution name matching system according to claim 7, characterized in that: The system also includes a human-machine collaborative feedback and optimization module, which is used to manually review the source medical institution and the candidate medical institution when the final matching confidence does not reach a preset decision threshold, and retrain the confidence fusion model and / or the large language model based on the matching results obtained by the manual review.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the medical institution name matching method according to any one of claims 1 to 6 are implemented.
10. A non-transitory computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instructions are executed by the processor, the steps of the medical institution name matching method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Address matching method and device
CN111950280A
Online geo-fencing address matching method
CN118939674A
Methods and systems for list filtering based on known entity matching
US20130204880A1