Method and system for generating settlement list based on medical data

By introducing medical insurance settlement rules constraints in the feature vector space, using medical ontology knowledge base and multi-dimensional indexing technology, the problem that coding in the existing technology does not meet the medical insurance settlement standards is solved, and the medical semantic rationality of the encoding and the compatibility of the settlement standards are achieved.

CN120068802AActive Publication Date: 2025-05-30SHENZHEN COMBIT INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510545442.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The prior art does not integrate the medical insurance settlement rules as constraints into the feature space construction during the encoding conversion process, resulting in the generated encodings that may be reasonable in medical semantics but do not meet the local medical insurance settlement standards, causing settlement disputes.

Method used

By introducing medical insurance settlement rules constraints into the feature vector space, using the medical ontology knowledge base to establish a mapping relationship for surgical coding, and combining multi-dimensional indexing technology to build a feature vector space, so that the coding meets the requirements of medical semantic similarity and medical insurance settlement rules at the same time.

Benefits of technology

Ensure that the generated operational coding is not only reasonable in medical semantics, but also in compliance with the standards and specifications of medical insurance settlement, and solves the problem of inconsistency between coding and settlement standards from the source, improving the accuracy of coding matching and the consistency of medical semantics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068802A_ABST
    Figure CN120068802A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for generating a settlement list based on medical data, and relates to the field of data processing, and the method comprises the steps: obtaining structured medical data and unstructured medical data; generating standardized data according to the structured medical data; generating a standardized text according to the unstructured medical data; fusing the standardized data and the standardized text; constructing a feature vector space according to the fused feature vectors; establishing a mapping relation of operation codes by using the medical ontology knowledge base; performing entity recognition on the fused feature vectors; and according to the entity identification result and the mapping relationship with the operation code, generating an operation and a corresponding operation code, and generating a medical insurance settlement list. In order to solve the problem that in the prior art, due to the fact that coding standards are not uniform, generated codes are possibly reasonable but do not conform to settlement standards, the codes meet the requirements of medical semantic similarity and medical insurance settlement rules at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and particularly to a method and system for generating a settlement list based on medical data. Background Art

[0002] With the continuous improvement of the medical security system, the role of the medical insurance settlement list in medical insurance payment has become increasingly important. As a key link between medical institutions and medical insurance departments, it is directly related to the economic benefits of medical institutions and the reasonable use of medical insurance funds. In the medical insurance settlement list, surgical and procedure codes and disease diagnosis codes are two of the most critical elements. The accuracy of surgical and procedure codes directly affects the score calculation of cases. Under-coding (insufficient coding) will result in medical institutions not being able to obtain due compensation, while over-coding (excessive coding) may cause unreasonable expenditure of medical insurance funds and even lead to compliance risks. To ensure that the medical insurance settlement list can objectively reflect the consumption of medical resources, medical institutions need to comprehensively analyze unstructured data including patient admission and discharge records, surgical operation records, etc., as well as structured data such as medical service price items, medication and consumable usage, and convert these heterogeneous data into standardized surgical and procedure codes. This process places high requirements on both data processing technology and medical expertise.

[0003] There are differences in medical insurance policies and coding standards in different regions. The surgical coding systems used within medical institutions (such as ICD-9-CM-3, ICD-10-PCS, etc.) often do not exactly correspond to the coding standards required for medical insurance settlement. Existing technologies do not incorporate medical insurance settlement rules as constraints into the construction of the feature space during the coding conversion process, resulting in generated codes that may be reasonable in medical semantics but do not conform to local medical insurance settlement standards, ultimately leading to settlement disputes.

[0004] In addition, existing technologies usually use flat coding mappings or simple classification trees, lacking the expression of the internal hierarchical relationships of surgical operations and unable to reflect multi-dimensional characteristics such as surgical systems, intervention methods, and operation methods. As a result, similar surgeries may obtain extremely different codes, which is not conducive to settlement analysis and auditing. Summary of the Invention

[0005] Aiming at the problem in the prior art that the lack of unified coding standards leads to generated codes that may be reasonable but do not conform to the settlement standards, this application provides a method and system for generating a settlement list based on medical data. By introducing medical insurance settlement rule constraints into the feature vector space, the generated codes can meet the requirements of both medical semantic similarity and medical insurance settlement rules.

[0006] The objectives of this application are achieved through the following technical solutions.

[0007] One aspect of the present application provides a method for generating a settlement list based on medical data, including: S1, obtaining structured medical data and unstructured medical data; S2, preprocessing the structured medical data to generate normalized data; performing text standardization processing on the unstructured medical data to generate standardized text; S3, fusing the normalized data and the standardized text to obtain a fused feature vector; constructing a feature vector space based on the fused feature vector; S4, establishing a mapping relationship of surgical codes according to the feature vector space by using a medical ontology knowledge base; S5, performing entity recognition on the fused feature vector to obtain an entity recognition result; S6, generating a surgery and corresponding operation code according to the entity recognition result and the mapping relationship with the surgical code; S7, generating a medical insurance settlement list according to the surgery and corresponding operation code.

[0008] Further, the structured medical data includes medical service prices, medication and consumable usage; the unstructured medical data includes patient admission and discharge records and surgical operation records.

[0009] Further, in S3, fusing the normalized data and the standardized text to obtain a fused feature vector includes: extracting quantitative features of the normalized data, where the quantitative features include medical service price features, drug usage features, and consumable consumption features; extracting text features of the standardized text, where the text features include semantic features and surgical term features; aligning the quantitative features and the text features to generate an initial fused feature; using an attention mechanism to perform weighted fusion on the initial fused feature to obtain a fused feature vector.

[0010] Further, constructing a feature vector space based on the fused feature vector includes: performing clustering analysis on the fused feature vector by using a knowledge-enhanced hierarchical clustering algorithm to obtain a feature vector category distribution; the knowledge-enhanced hierarchical clustering algorithm uses the synonymous relationship, hyponymy relationship, and co-occurrence relationship of surgical terms as clustering constraint conditions; calculating the semantic similarity between different surgical feature vectors based on the feature vector category distribution to construct a semantic similarity matrix; converting medical insurance settlement rules into spatial constraint conditions, where the medical insurance settlement rules include surgical operation type restrictions; partitioning the feature space through multi-dimensional indexing according to the semantic similarity matrix and the spatial constraint conditions to construct a feature vector space.

[0011] Among them, the synonymous relationship: This refers to the relationship where different medical terms or surgical descriptions, although having different forms of expression, actually refer to the same medical concept or surgical procedure. In medical terms, the synonymous relationship is very common because the same surgical procedure may have different ways of expression. For example, "laparoscopic cholecystectomy" = "laparoscopic removal of the gallbladder" = "LC"; "appendectomy" = "removal of the appendix" = "appendix surgery"; "PCI" = "percutaneous coronary intervention" = "coronary artery stenting".

[0012] The hierarchical relationship: This refers to the hierarchical relationship between medical concepts or surgical procedures, where one concept is a more specific form (subordinate relationship) or a more general form (superordinate relationship) of another concept. For example, "gastrectomy" is a subordinate concept of "digestive system surgery", while "digestive system surgery" is a superordinate concept of "gastrectomy". For example, superordinate: "digestive system surgery" → subordinate: "gastrointestinal surgery" → subordinate: "gastrectomy" → subordinate: "distal subtotal gastrectomy"; superordinate: "joint surgery" → subordinate: "knee joint surgery" → subordinate: "knee joint replacement" → subordinate: "total knee joint replacement"; superordinate: "vascular intervention surgery" → subordinate: "coronary artery intervention" → subordinate: "coronary artery stenting".

[0013] The co-occurrence relationship: This refers to the relationship between surgical procedures that often occur simultaneously or are performed in combination in medical practice. Some surgeries are usually carried out together clinically, and there is a statistical correlation or clinical combination necessity between them. For example, the co-occurrence relationship between "cholecystectomy" and "choledochoscopy" (co-occurrence rate is about 65%); the co-occurrence relationship between "hysterectomy" and "bilateral oophorectomy" (co-occurrence rate is about 43%); the co-occurrence relationship between "coronary artery bypass grafting" and "aortic valve replacement" (under specific indications).

[0014] The surgical operation type constraint: This refers to the reimbursement conditions, restrictions, or special regulations of medical insurance policies for specific types of surgeries. It includes requirements for indications, frequency limitations, combination limitations, etc. for certain surgeries, and these limitations will affect surgical coding and the final medical insurance settlement. For example, type constraint: the mutually exclusive coding constraint between minimally invasive surgery and open surgery; frequency limitation: "intraocular lens implantation" is limited to one medical insurance payment per eye per life cycle; combination limitation: "simple modified radical mastectomy" and "sentinel lymph node biopsy" cannot be coded and reimbursed simultaneously; level limitation: "transcatheter aortic valve replacement" is only coded and used in tertiary hospitals.

[0015] Further, to calculate the semantic similarity between different surgical feature vectors, the following formula is used: , where and are two surgical feature vectors, respectively representing the feature representations obtained in step S41 corresponding to different surgeries; and are and corresponding surgical terms, referring to the standardized surgical names in the medical knowledge graph; is the cosine similarity of the feature vectors; ; where is the term and in the medical knowledge graph, the number of shortest path edges, w is the path distance weight factor, used to adjust the influence degree of the path distance on the similarity, and the value range for different types of surgeries is [0.5, 2]; , where is the term and in the medical knowledge graph, the nearest common ancestor node, represents the distance from node T to the root node of the knowledge graph, and the value range is [0, 1]. The larger the value, the closer the taxonomic relationship between the two terms; α, β, and γ are weight coefficients, respectively controlling the proportion of the feature vector similarity, the path distance similarity, and the taxonomic similarity in the total similarity calculation, and α + β + γ = 1. The value range of α is [0.2, 0.4], the value range of β is [0.3, 0.5], and the value range of γ is [0.2, 0.4], which are dynamically adjusted according to different medical scenarios.

[0016] Further, according to the semantic similarity matrix and the spatial constraint conditions, through multi-dimensional indexing, the feature space is partitioned to construct a feature vector space, including: converting the surgical operation type restrictions in the medical insurance settlement rules into a binary relation matrix R, where indicates that operation i is compatible with operation j, indicates that operation i is mutually exclusive with operation j; constructing a directed acyclic graph according to the hierarchical relationship of the surgical operation types, where V is the set of operation types and E is the set of hierarchical constraint edges; constructing a three-level tree index structure according to the binary relation matrix R and the directed acyclic graph , where the first level is divided by the surgical system, the second level is divided by the intervention method, and the third level is divided by the operation method; assigning multi-dimensional space coordinates to the nodes in the three-level tree index structure to form an index structure containing spatial position information; mapping the semantic similarity matrix to the index structure containing spatial position information to obtain the feature vector space.

[0017] Among them, the existing technologies often only focus on medical semantic similarity and ignore the constraints of medical insurance settlement rules, resulting in that although the coding is medically reasonable, it does not meet the settlement standards. This application transforms the medical insurance settlement rules into spatial constraint conditions, introduces a binary relation matrix R to represent operation compatibility, and imposes settlement rule constraints when constructing the feature vector space, so that the generated coding satisfies both medical semantic similarity and medical insurance settlement rules from the design stage, fundamentally solving the problem of inconsistent coding standards.

[0018] Further, the surgical system represents the surgical area division based on human anatomy classification, including the circulatory system, digestive system, nervous system, skeletal muscle system, and urogenital system; the technical path for implementing the interventional surgery; the operation methods include but are not limited to resection, suture, repair, reconstruction, replacement, implantation, ablation, and decompression.

[0019] Further, multi-dimensional space coordinates are assigned to the nodes in the three-level tree index structure to form an index structure containing spatial position information, including: setting the number of dimensions n of the feature space; constructing a standard n-dimensional Hilbert space filling curve H; using the binary relation matrix R to deform and adjust the n-dimensional Hilbert space filling curve H to obtain a curve ; where the curve segments corresponding to the operations marked as compatible in the matrix R are adjacent, and the curve segments corresponding to the operations marked as mutually exclusive are separated; according to the hierarchical relationship of the three-level tree index structure, a spatial region is assigned to each node along the curve Allocate spatial regions.

[0020] Among them, the n-dimensional Hilbert space filling curve is a special continuous fractal curve that can recursively traverse and fill each point in the n-dimensional space while maintaining the relative proximity of neighboring points in space and sequence, thus mapping the high-dimensional space to a one-dimensional continuous path. In this application, multi-dimensional medical features (such as anatomical parts, operation methods, surgical difficulty, etc.) are uniformly mapped to the continuous curve; a unique curve coordinate is assigned to each coding node to support range query and nearest neighbor search; a spatially locally sensitive index structure is constructed to improve the feature retrieval speed and accuracy. For the discrete points in the n-dimensional space, the m-order Hilbert curve H can be defined by the following recursive function: : , where m represents the order of the curve, n represents the spatial dimension, and the function maps the one-dimensional index to the n-dimensional space coordinates.

[0021] Deformation adjustment is a space-filling curve topology optimization technology based on domain knowledge constraints. By applying the compatibility constraints in the binary relation matrix R to the standard Hilbert curve, the shape and density of the curve in a specific area are changed, so that the curve segments of compatible operations are close to each other while the curve segments of mutually exclusive operations are far from each other. In this application, the curve segments corresponding to the coding pairs of compatible operations (such as "cholecystectomy" and "choledochoscopy") are pulled closer; the curve segments corresponding to the coding pairs of mutually exclusive operations (such as "open surgery" and "laparoscopic surgery") are pulled farther apart; the medical insurance policy restriction conditions are transformed into the boundary constraints of curve deformation; the clinical knowledge of common surgical combinations is incorporated into the curve shape optimization goal. Specifically, the deformation adjustment can be expressed as finding a mapping function F such that: ; where F minimizes the objective function: ; d represents the spatial distance, and f represents the conversion function from the compatibility value to the ideal distance. Through this deformed n-dimensional Hilbert curve , a feature space index structure with medical semantic perception is constructed, so that the three-level tree index not only contains hierarchical relationships, but also contains complex medical knowledge such as compatibility and mutual exclusivity.

[0022] Furthermore, according to the hierarchical relationship of the three-level tree index structure, spatial regions are allocated for each node along the curve , including: according to the classification of surgical systems, the main trajectory of the curve is divided into multiple non-overlapping convex spatial region blocks, and each region block is assigned to a first-level node; within the spatial region block corresponding to each first-level node, according to the classification of intervention methods, along the curve the sub-trajectory is further divided into multiple continuous region segments, and each region segment is assigned to a second-level node; within the region segment corresponding to each second-level node, according to the classification of operation methods, along the curve the end trajectory, spatial coordinate points are assigned to each third-level node; among them, the main trajectory is the path of the Hilbert space-filling curve at the highest level, which is used to divide the largest granularity of spatial region blocks and corresponds to the first-level nodes (surgical system classification) in the three-level tree structure. For example, surgeries of the circulatory system (such as coronary artery bypass grafting, valve replacement) are mapped to large convex regions at the starting segment of the curve; surgeries of the digestive system (such as gastrectomy, colectomy) are mapped to continuous region blocks in the middle segment of the curve; surgeries of the orthopedic system (such as joint replacement, spinal fusion) are mapped to the corresponding regions in the latter segment of the curve; special transition regions are set at the intersections of the main trajectories for rare combined surgeries between different systems. The main trajectory can be expressed as a first-level piecewise function: , where is the division point, corresponding to the boundary of different surgical systems.

[0023] The sub-trajectory is the secondary path branch of the Hilbert curve within the regional block divided by each main trajectory, used to further divide the regional segment, corresponding to the second-level node (intervention method classification) in the three-level tree structure. For example, within the digestive system regional block, laparoscopic surgery is mapped to a continuous regional segment; open surgery is mapped to another adjacent but independent regional segment; endoscopic operation is mapped to a third regional segment; hybrid intervention methods (such as laparoscopic-assisted mini-incision surgery) are mapped to the transition area at the boundary of the intervention method. For the i-th regional block of the main trajectory, its sub-trajectory can be expressed as: wherein, represents the j-th sub-region division point within the i-th main region.

[0024] The end trajectory is the terminal path direction of the Hilbert curve at the finest granularity, used to allocate specific spatial coordinate points within the regional segment divided by the sub-trajectory, corresponding to the third-level node (operation method classification) in the three-level tree structure. For example, within the "laparoscopic digestive system" regional segment, operations of the "resection" type are mapped to a set of adjacent coordinate points; operations of the "anastomosis" type are mapped to another set of adjacent coordinate points; operations of the "biopsy" type are mapped to the third coordinate point area; composite operations (such as "resection + reconstruction") are mapped to the exact positions between the corresponding operation coordinate points. For the j-th sub-region within the i-th main region, its end trajectory can be expressed as a discrete point set: where represents the precise positioning point of the k-th operation method within the j-th sub-region of the i-th main region. In this application, this method of allocating the Hilbert curve with a multi-level trajectory structure creates an encoding index system that not only maintains the consistency of medical semantics but also has efficient spatial retrieval capabilities by directly mapping the hierarchical structure (system - intervention method - operation method) of the medical surgery classification system to different granularity trajectories of the space-filling curve. This structure not only supports precise coding queries but also enables similar surgery retrieval based on spatial proximity relationships and spatial constraint verification of medical insurance settlement rules.

[0025] Furthermore, in S6, according to the entity recognition result and the mapping relationship with the surgery code, generate the surgery and the corresponding operation code, including: performing encoding mapping on the entity recognition result according to the mapping relationship of the surgery code to obtain the preliminary code; according to the preliminary code, determining the category code, system code, and method code of each surgery according to the hierarchical division of the three-level tree index structure to form the hierarchical code; using the binary relation matrix R to verify whether the hierarchical code meets the compatibility constraints of the matrix R; using the positional relationship in the eigenvector space to allocate the final operation code for the hierarchical code combinations that pass the verification; Another aspect of the present application also provides a system for generating a settlement list based on medical data, which is used to execute a method for generating a settlement list based on medical data according to the present application.

[0026] Compared with the prior art, the advantages of the present application are as follows: The medical insurance settlement rules are not incorporated as constraint conditions into the construction of the feature space, resulting in the generated codes being possibly reasonable but not meeting the settlement standards. The present application converts the surgical operation type restrictions in the medical insurance settlement rules into a binary relation matrix R, and imposes these constraint conditions during the construction of the feature vector space, so that the organizational structure of the coding space itself contains the constraint information of the settlement rules, thus ensuring that the generated operation codes are not only reasonable in medical semantics but also meet the standards and specifications of medical insurance settlement, and solving the problem of inconsistent coding and settlement standards from the source.

[0027] In the prior art, when calculating the surgical similarity, a simple bag-of-words model or a basic TF-IDF vector cosine similarity is generally used. However, these methods only focus on the surface features of the text and ignore the synonymy, hyponymy, and co-occurrence relationships existing between terms, affecting the subsequent accurate matching of the codes. The present application designs a composite similarity calculation formula , comprehensively considering the cosine similarity of the feature vectors, the path distance similarity in the medical knowledge graph, and the taxonomic similarity of the nearest common ancestor nodes, and at the same time taking the semantic relationship of the surgical terms as the constraint condition for hierarchical clustering, realizing the accurate measurement of the surgical similarity in the medical professional field and improving the accuracy of code matching and the consistency of medical semantics.

[0028] In the prior art, a simple hash mapping or a linear index structure is generally used, which has the defects of scattered coding and difficulty in obtaining similar codes for similar operations. The present application constructs an n-dimensional Hilbert space filling curve and adjusts it according to the binary relation matrix R, making the curve segments of compatible operations adjacent and the curve segments of mutually exclusive operations separated, ensuring the continuity and locality of the coding space, improving the coding query and processing efficiency, and reducing the system calculation burden. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The present application will be further described in the form of exemplary embodiments, and these exemplary embodiments will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where: Figure 1 is an exemplary flowchart of a method for generating a settlement list based on medical data according to some embodiments of the present application; Figure 2 is an exemplary flowchart of generating a fused feature vector according to some embodiments of the present application; Figure 3An exemplary flowchart of constructing a feature vector space according to some embodiments of the present application; Figure 4 An exemplary flowchart of allocating spatial coordinate points according to some embodiments of the present application. Detailed implementation manners

[0030] The methods and systems provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0031] As Figure 1 shown, obtain structured medical data and unstructured medical data; preprocess the structured medical data to generate normalized data; perform text standardization processing on the unstructured medical data to generate standardized text; fuse the normalized data and the standardized text to obtain a fused feature vector; construct a feature vector space according to the fused feature vector; establish a mapping relationship of surgical coding by using a medical ontology knowledge base according to the feature vector space; perform entity recognition on the fused feature vector to obtain an entity recognition result; generate a surgery and corresponding operation code according to the entity recognition result and the mapping relationship with the surgical coding; generate a medical insurance settlement list according to the surgery and corresponding operation code.

[0032] S1. Obtain structured medical data and unstructured medical data. Among them, for obtaining structured medical data, for obtaining medical service price data: through the data interface of the hospital information system (HIS), obtain the detailed charge list data during the patient's hospitalization, including fields such as charge item code, item name, unit price, quantity, billing unit, executing department, and expense category. Adopt a combination mode of regular batch extraction and real-time push to ensure the integrity and timeliness of the data.

[0033] For obtaining medication information: extract the patient's medication records from the hospital pharmacy management system, including information such as generic name of the drug, trade name, specification, dosage form, usage and dosage, administration route, medication time, medication frequency, and drug code. For special medications (such as antibiotics, anesthetic drugs, psychotropic drugs, etc.), additionally mark their medication categories for subsequent processing.

[0034] For obtaining the usage situation of medical consumables: through the interface of the operating room management system and the material management system, extract the medical consumable information used during the patient's surgery, including consumable name, model, specification, quantity, usage time, usage site, consumable classification, etc. For high-value consumables, additionally record their batch numbers and barcode information to establish a traceability link.

[0035] Obtain unstructured data, and obtain patient admission and discharge records: Extract documents such as patient admission records, discharge summaries, and progress notes from the Electronic Medical Record System (EMR). These documents are usually stored in XML, HTML, or PDF formats. Implement data extraction through custom interfaces or HL7 standard interfaces, and retain the original structure and format information of the documents.

[0036] Obtain surgical operation records: Extract documents such as operation record sheets, anesthesia record sheets, and surgical nursing records from the surgical information system. These documents contain key information such as the surgical method, operation site, description of the surgical process, operation time, and anesthesia method. For multimedia materials (such as surgical images), extract their metadata information as an auxiliary reference.

[0037] S2. Preprocess structured medical data to generate normalized data; perform text standardization processing on unstructured medical data to generate standardized text. Among them, for preprocessing structured medical data, data cleaning: Detect and process outliers, missing values, and duplicate values in structured data. Identify outliers using statistical methods (such as the 3σ rule), and correct or mark them according to business rules; for missing values, adopt different filling strategies according to the data type, such as mean filling, mode filling, or derivation from previous and subsequent values; perform merging or deletion processing on duplicate values. Data standardization: Convert structured data in different systems into a unified data format and coding standard. Map medical service price items to the national medical insurance fee directory; uniformly map drug information to the National Essential Medicines List and the Drug Coding Rules; map medical consumables to the Rules for the Unique Identification System for Medical Devices. Temporal alignment: Align structured data from different sources along the time axis, and establish the temporal relationship between medication, use of consumables, and medical services, providing a basis for subsequent correlation analysis. Adopt the sliding time window technique to handle the problem of inconsistent timestamp precision. Numerical normalization: Uniformly convert numerical values with inconsistent measurement units, such as converting medication doses in different units to standard units; perform normalization processing on amount information to eliminate the impact of price changes in different periods.

[0038] For preprocessing unstructured medical data, document parsing: According to the characteristics of unstructured documents in different formats, adopt corresponding parsing techniques to extract text content. Use DOM or SAX parsers to extract text from XML / HTML documents; use professional PDF parsing tools to extract text content from PDF documents; retain the chapter structure information of the documents for subsequent refined processing.

[0039] Text tokenization and normalization: For the characteristics of medical texts, a medical professional tokenization system is used for Chinese word segmentation to identify professional vocabulary such as medical terms, drug names, and anatomical parts; text normalization processing is carried out, including unifying full-width / half-width characters, removing special symbols, unifying case, and traditional / simplified Chinese conversion; medical abbreviations and shorthands are expanded, such as "BP" expanded to "blood pressure" and "WBC" expanded to "white blood cell count"; numerical expressions and units are identified and standardized, such as "2mg / kg" standardized to "2 milligrams per kilogram"; Medical term normalization: Map synonymous medical terms to a standard medical term set. A medical ontology library (such as SNOMED CT, ICD-10, etc.) is used for term matching to convert non-standard expressions in the text into normalized expressions, such as unifying "cesarean section", "abdominal cesarean section", "C-section", etc. into the standard term "cesarean operation".

[0040] Named entity recognition: Use named entity recognition technology in the medical field to identify and label key entities such as operation names, anatomical parts, operation methods, and diagnosis names from the text, and establish relationships between entities, such as "operation - part", "diagnosis - operation" relationships, etc. BiLSTM-CRF or a pre-trained language model in the medical field is used to achieve high-precision entity recognition. Text structuring: Based on the identified medical entities and relationships, convert unstructured text into a semi-structured representation form, and extract key facts and attributes, such as attributes of an operation like "operation method", "operation part", "operation purpose", "instruments used", etc., to form a structured event description.

[0041] As Figure 2 shown in S3, fuse the normalized data and standardized text to obtain a fused feature vector, including: First, extract quantitative features. The medical service price feature can be extracted by constructing a service item vector, including item frequency, cumulative cost, average cost per time, and standard deviation; introduce temporal features to record the time distribution and phased characteristics of item execution; calculate the proportion distribution of various medical services to reflect the resource consumption structure.

[0042] Extract drug usage features, construct a drug usage vector according to pharmacological classification, including the number of drug types, days of medication, and total dose; extract the temporal pattern of drug use to identify the corresponding relationship between the drug use procedure and the treatment stage; calculate the usage intensity index of special drugs (such as antibiotics, anesthetic drugs).

[0043] Extract consumable consumption features, construct a consumable consumption vector, including the usage of high-value consumables and the consumption of conventional consumables; extract the combined pattern of consumable use to identify the set of consumables related to a specific operation; calculate the value distribution of various consumables as an indirect indicator of the surgical complexity.

[0044] Then extract text features. The semantic features can be extracted by using a pre-trained language model in the medical field to extract the context semantic representation of the text, using the attention mechanism to extract the semantic vectors of key sentences in the surgical description, and calculating the density and complexity of medical terms in the text to reflect the technical requirements of the surgery.

[0045] Extract term features, identify key terms such as surgical names, operation sites, operation methods, etc. in the text; construct a term co-occurrence matrix to capture the combination patterns between terms; extract the hierarchical relationship features of terms to reflect the classification information of surgeries.

[0046] Feature fusion: Establish the time correspondence relationship between structured data and text descriptions, perform spatial alignment according to the medical service execution department and the text record department, construct an alignment matrix, and record the corresponding intensity of data from different sources; splice the aligned quantitative features and text features to form an initial feature vector, apply principal component analysis (PCA) to reduce the feature dimension, eliminate redundant information, and perform feature normalization processing to eliminate the dimensional influence of different types of features. Design a multi-head attention mechanism, calculate the importance weights of different feature components, dynamically adjust the attention weight distribution according to the current medical scenario, and obtain the final fused feature vector through weighted summation. This vector comprehensively expresses the multi-dimensional characteristics of the patient's medical service.

[0047] As Figure 3 shown, construct a feature vector space. First, design clustering constraints, convert the synonymous relationships in the medical term ontology into hard constraints of "must be clustered together", convert the hyponymy relationships into structural constraints of "hierarchical attribution", and convert the co-occurrence relationships into soft constraints of "similarity enhancement"; adopt the agglomerative hierarchical clustering algorithm. Initially, regard each surgical feature vector as an independent category, incorporate knowledge constraints into each merger decision, preferentially merge categories that meet the constraint conditions, use the modified Ward's minimum variance method to calculate the distance between categories, reduce the influence of outliers, and generate a hierarchical tree structure to reflect the classification pedigree of surgical features. Through knowledge-enhanced hierarchical clustering, the problem that traditional clustering methods ignore prior knowledge in the medical field is effectively solved, and the formed category distribution is more in line with medical professional cognition.

[0048] Specifically, on the one hand, traditional vector similarity calculations (such as TF-IDF, Word2Vec, etc.) only focus on lexical or statistical features and cannot capture the complex semantic relationship network between medical terms, resulting in terms that are superficially similar but have very different medical meanings being misclassified. On the other hand, there is a lack of a technical path to organically combine medical ontology knowledge with numerical calculations, and it is impossible to use existing medical classification systems and term relationship networks to enhance similarity calculations. This application designs a composite similarity calculation formula: .

[0049] Vector cosine similarity : Calculate the cosine angle between two surgical feature vectors, which represents the similarity of traditional feature vectors and measures the directional consistency of two surgical feature vectors in a high-dimensional space. Its mathematical basis is the cosine formula of the included angle in the vector space model, with high computational efficiency and easy parallel processing. By adaptively adjusting the feature dimension weights, this component can highlight the contribution of key features to similarity and reduce the interference of noise features.

[0050] Path distance similarity: Utilize the concept association network in the medical knowledge graph to calculate the shortest path between terms. The path distance similarity formula: ; Convert the concept connections in the medical knowledge graph into a similarity metric. Encode the knowledge of term relationships scattered in medical literature and guidelines into a graph structure, enabling the calculation process to consider the semantic connection strength between terms. The introduced dynamic path weight w parameter can precisely control the influence degree of path distance in different surgical fields, solving the problem of inconsistent term connection densities in different medical subfields.

[0051] Taxonomic similarity: Based on the depth of the most recent common ancestor node, evaluate the genetic relationship of terms in the classification system. The taxonomic similarity formula: ; Evaluate the systematic relationship of terms in the classification tree. Quantify the hierarchical classification information into a similarity metric, which is particularly suitable for fields such as medical surgery with a strict classification system. By calculating the ratio of the depth of the common ancestor to the depth of the term itself, it effectively balances the comparison benchmarks of terms at different levels and solves the asymmetry problem in traditional hierarchical similarity calculations.

[0052] In summary, SimCos captures the similarity relationship at the feature representation level; SimPath captures the semantic connection relationship between terms; captures the classification hierarchical relationship between terms. Incorporate the structured relationship information in the medical knowledge graph into the similarity calculation, making up for the knowledge gap of the pure vector model. Through dynamic weight adjustment, adaptively balance the importance of different semantic relationships to meet the similarity evaluation needs of different types of surgeries.

[0053] Dynamically adjust the α, β, and γ weight parameters according to different medical scenarios; for routine surgeries with high standardization, increase the weight of vector similarity (α); for new and complex surgeries, increase the weight of knowledge graph path similarity (β); for multi-system joint surgeries, increase the weight of taxonomic similarity (γ). The composite similarity calculation method solves the problem that traditional similarity calculations ignore the special semantic relationships between medical terms. By integrating vector space similarity and knowledge graph relationships, it achieves an accurate measurement of medical surgery similarity.

[0054] Specifically, there is a fundamental technical contradiction in the field of medical insurance settlement list generation: the coding needs to meet both the medical semantic rationality and the compliance with medical insurance policies at the same time. The existing medical insurance settlement rules usually exist in text form, lacking an effective mechanism to convert these rules into computable constraint conditions, making it difficult for the system to automatically apply these rules. Secondly, in the traditional vector space, surgeries with similar semantics are mapped to adjacent positions, but this arrangement of semantic similarity often conflicts with the constraints of medical insurance policies, resulting in non-compliant coding. Finally, the medical coding space is highly complex, and it is difficult for traditional index structures to ensure both query efficiency and semantic fidelity at the same time, especially when multiple constraint conditions need to be considered simultaneously.

[0055] Therefore, in this application, first, the systematic conversion of the surgical operation type restrictions in the medical insurance settlement rules into a binary relation matrix R is carried out. Specifically, the surgical operation restriction rules are extracted from medical insurance policy documents (such as "Administrative Measures for Medical Insurance Diagnosis and Treatment Items", "Surgical Operation Specifications", etc.); then, an n×n binary relation matrix R (n is the total number of surgical operation types) is constructed, where indicates that operation i and operation j are compatible and can be coded simultaneously; indicates that operation i and operation j are mutually exclusive and should not be coded simultaneously.

[0056] For the mutually exclusive relationships clearly stipulated in the policy (such as "resection and repair should not be coded simultaneously at the same surgical site"), the corresponding matrix elements are directly set to 0; for surgeries with sequential dependence relationships in clinical practice but not clearly stipulated in the policy (such as "vascular resection" and "vascular reconstruction"), their co-occurrence patterns are identified by analyzing clinical data, and the corresponding matrix elements are set to 1.

[0057] For fuzzy relationships, a fuzzy compatibility degree [0, 1] is introduced to reflect the strength of policy constraints. For example, for operation pairs that "can be coded together depending on the situation", it may be set that R(i, j)=0.7, indicating medium-high compatibility but further clinical judgment is required. This fuzzy value is obtained through the scoring of medical insurance experts and the analysis of historical settlement success rates. This application converts the medical insurance policies expressed in natural language into precise mathematical representations, realizing a qualitative change from "rule description" to "computational constraint". Secondly, by introducing the fuzzy compatibility degree in the [0, 1] interval, the problem of coexistence of "strong constraints" and "weak suggestions" in medical insurance policies is solved, enabling the system to handle policy guidance of different intensities. Finally, the matrix form is convenient for efficient computer processing, supports large-scale matrix operations and parallel computing, providing a basis for subsequent space construction.

[0058] Then, based on the medical ontology and the medical insurance classification standard, a directed acyclic graph G=(V, E) of surgical operation types is constructed. Specifically, the ontology engineering and knowledge graph technology are applied to construct the hierarchical relationship graph. First, the concept hierarchy of surgical operations is extracted from the standard medical terminology systems (such as SNOMED CT, ICD-10-PCS); then, combined with the classification system in the medical insurance coding manual, the hierarchical relationship reflecting the local medical insurance policy is constructed.

[0059] In the graph G, the node V represents the operation type, from the high-level concept (such as "surgical operations on the circulatory system") to the specific operation (such as "coronary artery bypass grafting"); the edge E represents the hierarchical constraint relationship, including the "is-a" relationship (such as "coronary artery bypass grafting is a type of coronary artery surgery") and the composite relationship (such as "operation combinations that need to be used jointly").

[0060] The conflicts and loops in the hierarchical relationship are detected and eliminated through graph algorithms. For example, if an operation is reclassified after the update of the medical insurance policy, which may lead to conflicts in the hierarchical relationship, the system will detect such conflicts through transitive closure calculation and remind the administrator to make adjustments. On the one hand, this application represents the operation type and the hierarchical constraint through the node V and the edge E respectively, and transforms the complex classification system into a computable graph structure. By using the concepts of reachability and transitive closure in graph theory, the automatic derivation of "the upper-level rules constrain the lower-level operations" in the medical insurance policy is realized, greatly reducing the redundant storage of rules. On the other hand, the loops and contradictions are detected through graph algorithms to ensure the consistency of the coding rules and avoid conflicts and contradictions in policy interpretation.

[0061] A three-level tree-shaped index structure is constructed, and the hierarchical clustering and index optimization technology are used to construct the index structure. First, according to the hierarchical relationship in the directed acyclic graph G, the boundaries of the three-level classification are determined; then, the operation compatibility between the nodes at each level is analyzed to optimize the branch structure of the index tree. The first level is divided according to the surgical system, including the circulatory system, digestive system, nervous system, etc., reflecting the classification of surgical areas based on human anatomy. The system boundaries are automatically extracted by parsing the chapter structures of medical textbooks and surgical specifications, and verified by domain experts.

[0062] The second level is divided according to the intervention method, such as open surgery, endoscopic surgery, interventional therapy, etc., reflecting the technical path of surgical implementation. The classification of intervention methods is automatically identified by analyzing the operation approach, instruments used, and technical characteristics in the surgical description.

[0063] The third level is divided according to the operation method, including resection, suture, repair, reconstruction, replacement, implantation, ablation, and decompression, etc., reflecting the specific surgical behavior. The operation method is determined by semantic analysis of the surgical records to identify the core verb and the operation object.

[0064] Specifically, for common multi-system joint surgeries (such as "thoracoabdominal combined surgery"), cross-system index links are established so that they can be accessed from multiple first-level nodes. Each level captures different dimensional characteristics of the surgery, from anatomical location to technical path to specific operations, achieving a comprehensive representation of the multi-dimensional characteristics of the surgery. Using the hierarchical pruning ability of the tree structure, irrelevant branches can be quickly excluded during the query process, significantly improving the retrieval efficiency of large-scale coding libraries.

[0065] Next, multi-dimensional space coordinates are assigned to the nodes in the three-level tree-shaped index structure to form an index structure containing spatial location information. First, the optimal dimension is determined through principal component analysis (PCA) and feature importance evaluation. First, feature extraction is performed on the existing surgical dataset to obtain initial high-dimensional features (usually including 100 - 200 features); then PCA analysis is applied to select the first n principal components whose cumulative explained variance reaches more than 95%; finally, in combination with the evaluation of experts in the field of medical insurance coding, the final number of dimensions is determined. For example, surgeries of the circulatory system may require a higher dimension (n = 12) to express their complexity, while relatively simple surgeries of the skin system may only require a lower dimension (n = 8).

[0066] Then a Hilbert curve is generated using a recursive construction method. First, the n-dimensional space is divided into equal hypercubes; then the access order of these hypercubes is determined so that adjacent accessed hypercubes are also adjacent in space; the same partitioning and sorting rules are recursively applied to each hypercube until a preset precision level is reached (usually 16 - 20 levels, which can represent approximately points). The Gray code technique is used to optimize the direction change of the curve to ensure a smooth transition at the turning points of the curve. To meet the special requirements of medical coding, based on the standard Hilbert curve, the rotation matrix in curve generation is adjusted to have a higher resolution in the dimensions related to the surgical system.

[0067] Then the graph embedding algorithm and the force-directed model are applied to achieve curve deformation. First, the binary relation matrix R is converted into a weighted graph, where an attractive force is added between node pairs of and a repulsive force is added between node pairs of ; then, under the constraint of maintaining the continuity of the curve, the positions of the points on the curve are adjusted through an iterative optimization algorithm (such as the spring-charge model); finally, an order-preserving transformation is applied to maintain the local properties of the curve, ensuring that the adjusted curve H' is still a valid space-filling curve.

[0068] For relationships with fuzzy compatibility, elastic coefficient adjustment is adopted. The higher the compatibility, the greater the spring constant and the stronger the attractive force. For example, when two surgical operations (such as "intestinal resection" and "intestinal anastomosis") are marked as highly compatible in the medical insurance rules ( When it is (such as "coronary artery bypass grafting" and "percutaneous coronary intervention"), the corresponding curve segments will be pulled closer by a stronger attraction; while for mutually exclusive operations (such as "coronary artery bypass grafting" and "percutaneous coronary intervention"), they should not be coded simultaneously during the same hospitalization of the same patient, and the corresponding curve segments will be pushed apart by the repulsive force.

[0069] Map the pre-computed semantic similarity matrix into an index structure containing spatial location information to form the final feature vector space. Specifically, use multi-objective optimization and topology-preserving mapping techniques to achieve the mapping. First, construct the semantic similarity matrix of surgical operations, and its elements represent the semantic similarity degree between surgeries i and j; then, combine this similarity constraint with the aforementioned spatial coordinate assignment, and adjust the spatial coordinate distribution through an optimization algorithm. The specific implementation uses an optimization method that combines gradient descent and simulated annealing. The objective function is designed as: ; where, is the Euclidean distance between points i and j in space, is the semantic similarity, is the binary relation value, and are the weight coefficients for balancing semantic similarity and policy compatibility, and α and β are scaling factors.

[0070] Through such optimization, the system achieves a balance of three objectives: 1) surgeries with similar semantics are close in spatial location; 2) surgeries compatible with medical insurance rules are adjacent in spatial location; 3) surgeries mutually exclusive under medical insurance rules are separated in spatial location. Finally, use spatial indexing techniques such as locality-sensitive hashing (LSH) and R-tree to establish an efficient query mechanism for the constructed feature vector space, support k-nearest neighbor search and range query, and significantly improve the efficiency of coding matching.

[0071] As Figure 4 shown, according to the hierarchical relationship of the three-level tree-shaped index structure, allocate spatial regions for each level of nodes along the deformed curve The first-level spatial region division: Based on the importance and complexity of the surgical system, determine the spatial allocation ratio of each system. For example, the circulatory system accounts for 20%, the digestive system accounts for 18%, the nervous system accounts for 15%, etc. Then divide the curve H' along the main trajectory into continuous segments with corresponding ratios, and convert the set of spatial points covered by each segment of the curve into a convex spatial region block through the convex hull algorithm. To ensure the coherence of the regions, apply the minimum spanning tree algorithm to connect the region subsets that may be separated due to curve bending.

[0072] The second-level spatial region division: Within each first-level spatial region block, determine the spatial allocation ratio according to the medical resource consumption and technical complexity of different intervention methods in this system. For example, within the circulatory system, open surgery may account for 40%, interventional surgery may account for 35%, and minimally invasive surgery may account for 25%. Then within this region block, along The sub-trajectory is further subdivided into multiple consecutive regional segments. An interval tree data structure is applied to record the boundary coordinates of these regional segments for subsequent rapid positioning.

[0073] Third-level spatial point allocation: Within each second-level regional segment, according to the fine classification of the operation method, along the end trajectory, specific spatial coordinate points are allocated to each third-level node. For common operations (such as "resection"), more coordinate points are allocated; for rare operations (such as a specific "reconstruction"), only a small number of coordinate points may be allocated. To improve spatial utilization, a density adaptive algorithm is applied to increase the sampling density in key areas. This multi-level spatial allocation strategy not only maintains the spatial locality of the Hilbert curve but also reflects the inherent hierarchical relationship of medical surgery classification through hierarchical division, enabling the spatial structure itself to contain medical knowledge, thereby improving the accuracy of subsequent coding and the efficiency of retrieval.

[0074] S4. According to the feature vector space, establish a mapping relationship for surgical coding using the medical ontology knowledge base; in this embodiment, UMLS (Unified Medical Language System) and SNOMED CT are used as the basic medical ontologies, integrating the Chinese Classification Code of Clinical Surgery and Operations (CCC) and the medical insurance payment coding standard; construct an ontology knowledge base containing 179,000 medical concepts and 347,000 synonyms, covering 98.6% of common surgical terms; specifically, establish a mapping relationship for "Sub total gastrectomy" in the knowledge base with the ICD-9-CM code "43.89" and the CCC code "G3301003"; establish "anchor vectors" for 5,800 standard surgical codes in the feature vector space, and each anchor represents the standard position of a standard code; for each anchor, collect 3 - 8 typical medical record samples, extract features and calculate the average vector as the representation of this code; specifically, for the code "LC1201" of "laparoscopic cholecystectomy", collect 27 surgical records from 3 Class-A tertiary hospitals, and after extracting features, obtain an anchor vector of dimension 10 [0.78, 0.65, 0.12, 0.03, 0.92, 0.45, 0.32, 0.19, 0.56, 0.28].

[0075] According to the complexity and ambiguity of different surgical categories, set a dynamic similarity threshold matrix T. For surgeries with clearly defined boundaries (such as fracture reduction), set a higher threshold (T = 0.85); for surgeries with large variations in description (such as various repair surgeries), set a lower threshold (T = 0.65); specifically, the threshold matrix for digestive system surgeries shows that the mapping threshold for "gastrectomy" surgeries is 0.78, while the mapping threshold for "intestinal repair" surgeries is 0.67.

[0076] Construct a multi-level mapping path of "surgical description → feature vector → anchor vector → standard code". For 43 common variants of surgical descriptions, establish a synonym mapping table to achieve the normalization of different expressions. Specifically, map 13 expressions such as "partial hepatectomy", "hepatic segmentectomy", and "wedge resection of liver" to the standard expression of "partial hepatectomy" to improve the mapping accuracy. For the fuzzy region in the feature space (where multiple coding anchors are close), construct a weighted decision tree for secondary judgment. Use 364 clinical discriminative features as decision tree nodes, and determine the optimal branching path according to historical data. Specifically, for surgeries falling into the fuzzy region between "subtotal gastrectomy" and "total gastrectomy", the system uses 5 key features such as "whether to retain the cardia" and "percentage of resection range" for decision tree judgment.

[0077] S5. Perform entity recognition on the fused feature vector to obtain the entity recognition result. In this embodiment, a hybrid model combining BiLSTM-CRF (Bidirectional Long Short-Term Memory - Conditional Random Field) and attention mechanism is adopted. The model is trained on 12,800 annotated medical records, and the F1 score reaches 0.917. Specifically, for the surgical description of "The patient underwent laparotomy + partial resection of gastric antrum + Billroth II gastrojejunostomy", three surgical entities of "laparotomy", "partial resection of gastric antrum", and "Billroth II gastrojejunostomy" are recognized.

[0078] Use a customized medical entity recognition model to accurately identify anatomical location information. Construct a location relationship map containing 4,200 anatomical sites to support the derivation of the upper and lower levels of the sites. Specifically, identify "right upper lobe of the lung" as the precise anatomical location from "wedge resection of the right upper lobe of the lung" and establish a part-whole relationship with "lung".

[0079] Perform operation method recognition based on the surgical verb ontology library, which contains 486 standard surgical verbs. Use dependency syntactic analysis to extract the relationship between the verb and the receptor, and construct an operation-object pair. Specifically, identify the operation method as "puncture biopsy" and the operation object as "renal parenchyma" from "percutaneous puncture biopsy of renal parenchyma".

[0080] Adopt a hybrid model of rules and statistics to recognize surgical path information. Establish a dictionary containing 89 common interventional paths, covering main surgical types such as open, minimally invasive, and endoscopic. Specifically, identify "transcatheter" as the interventional path from "transcatheter aortic valve replacement" and classify it into the category of "interventional therapy".

[0081] Identify key modifiers affecting coding, such as "bilateral", "multi-site", "repeated", etc. Construct a modifier influence matrix to quantify the influence degree of various modifiers on the final coding. Specifically, identify the "bilateral" modifier in "bilateral inguinal hernia repair", and automatically adjust it to a surgery that requires two independent codings during the coding process.

[0082] S6. According to the entity recognition result and the mapping relationship with the surgical coding, generate the surgery and the corresponding operation coding. In this embodiment, the recognized surgical entity is converted into a feature vector, and the nearest k anchor vectors (usually k = 5) are searched in the feature space; the weighted K-nearest neighbor algorithm is used to calculate the probability score of each candidate coding according to the distance. Specifically, for the recognized "laparoscopic cholecystectomy" entity, after the system generates the feature vector, the nearest 5 anchor codings are found: LC1201 (distance 0.17), LC1202 (distance 0.22), LC1203 (distance 0.38), LC1205 (distance 0.45), KC1201 (distance 0.57), and LC1201 is initially selected as the best match. The core problem solved in this stage is the problem of semantic mapping uncertainty, that is, how to accurately map the unstructured surgical entity to the structured coding system. Based on the Vector Space Model and the nearest neighbor retrieval theory, this application embeds the surgical entity and the standard coding into a unified high-dimensional feature space at the same time. In this space, concepts with similar semantics are also close to each other geometrically, thus transforming the semantic matching problem into a quantifiable distance calculation problem.

[0083] Decompose the preliminary coding according to the three-level tree index structure, and extract the category code, system code and method code; for the uncertain part in the preliminary coding, use Bayesian network reasoning for supplementation; specifically, decompose LC1201 into L (laparoscopic category), C (digestive system), 12 (gallbladder), 01 (resection) to form a complete hierarchical coding. The core problem solved in this stage is the problem of multi-level semantic integration, that is, how to decompose the overall coding into components that conform to the medical ontology hierarchy. Based on Hierarchical Representation Learning and the tree structure traversal algorithm, this application splits the coding into components reflecting different semantic levels and verifies its effectiveness in the medical ontology.

[0084] Use the binary relation matrix R to check the compatibility between multiple surgical codes; for incompatible code combinations, use a rule-based priority strategy to resolve conflicts; specifically, when a patient undergoes both "cholecystectomy" and "choledocholithotomy", the system detects that the compatibility value of these two surgeries in the matrix R is 0.9 (highly compatible), and both codes are retained; while when detecting the combination of "open cholecystectomy" and "laparoscopic cholecystectomy", its compatibility value in the matrix R is 0.1 (almost mutually exclusive), and the system retains the more complex "open cholecystectomy" code through clinical rules. The core problem solved in this stage is the multi-code coordination consistency problem, that is, how to ensure that multiple surgical codes meet the compatibility constraints of medical insurance policies. This application is based on the Constraint Satisfaction Theory and matrix algebra, represents medical insurance rules as a binary relation matrix, and verifies the compliance of code combinations through matrix operations.

[0085] Specifically, in the binary relation matrix R, represents the compatibility degree between code i and code j: : completely compatible; : completely mutually exclusive; : partially compatible, the larger the value, the higher the compatibility degree. For the code set , its overall compatibility is calculated through matrix multiplication: ), when , it indicates that there are mutually exclusive codes in the combination. Based on the graph coloring theory, the mutually exclusive code problem is transformed into a graph coloring problem, and the greedy algorithm is applied to solve the conflict: MaxIndependent Set(Graph(C, E = {mutually exclusive relationship})), where the independent set represents the largest code subset that can be retained simultaneously.

[0086] Based on the positional relationship in the eigenvector space, assign the final code to the verified hierarchical code combination; use the medical insurance payment standard database to conduct the final verification of code compliance; specifically, for the combination of "laparoscopic cholecystectomy + common bile duct exploration and lithotomy", the system finally assigns the codes LC1201 + LC0702 and automatically adds the necessary code association markers.

[0087] S7, generate a medical insurance settlement list according to the surgery and the corresponding operation codes. Specifically, build a code - settlement item conversion engine to convert surgical codes into medical insurance settlement items; apply a regional policy adapter to adjust settlement items according to different regional medical insurance policies; implement hospital-level difference processing to adjust the scope of claimable settlement items according to the hospital level and qualifications. Build a settlement list generator to organize settlement items in a standard format; apply expense classification and summary to summarize expenses by categories such as drug expenses, surgical expenses, and material expenses; implement the division of out-of-pocket and medical insurance payments to clearly divide the patient's out-of-pocket part and the medical insurance payment part.

Claims

1. A method for generating a settlement list based on medical data, characterized in that: include: S1, obtain structured medical data and unstructured medical data; S2, preprocessing structured medical data to generate standardized data; Perform text standardization on unstructured medical data to generate standardized text; S3, fusing the normalized data and the standardized text to obtain a fused feature vector; constructing a feature vector space based on the fused feature vector; S4, based on the feature vector space, the mapping relationship of surgical codes is established using the medical ontology knowledge base; S5, performing entity recognition on the fused feature vector to obtain an entity recognition result; S6, generating the surgery and the corresponding operation code according to the entity recognition result and the mapping relationship with the surgery code; S7, generate a medical insurance settlement list based on the surgery and the corresponding operation code.

2. The method for generating a settlement list based on medical data according to claim 1, characterized in that: The structured medical data includes medical service prices, medication and consumables usage; The unstructured medical data includes patient admission and discharge records and surgical operation records.

3. The method for generating a settlement list based on medical data according to claim 2, characterized in that: S3, obtain the fused feature vector, including: Extracting quantitative features of the normalized data, wherein the quantitative features include medical service price features, drug use features, and consumables consumption features; Extracting text features of the standardized text, wherein the text features include semantic features and surgical term features; Align quantitative features and text features to generate initial fusion features; The initial fusion features are weightedly fused using the attention mechanism to obtain the fused feature vector.

4. The method for generating a settlement list based on medical data according to claim 3, characterized in that: Construct a feature vector space based on the fused feature vectors, including: A knowledge-enhanced hierarchical clustering algorithm is used to perform cluster analysis on the fused feature vectors to obtain the feature vector category distribution; the knowledge-enhanced hierarchical clustering algorithm uses the synonymy, hyponymy and concurrency relationships of surgical terms as clustering constraints; Based on the feature vector category distribution, the semantic similarity between different surgical feature vectors is calculated and a semantic similarity matrix is ​​constructed; Converting medical insurance settlement rules into spatial constraints, wherein the medical insurance settlement rules include restrictions on surgical operation types; According to the semantic similarity matrix and spatial constraints, the feature space is divided through multi-dimensional indexing to construct the feature vector space.

5. The method for generating a settlement list based on medical data according to claim 4, characterized in that: The semantic similarity between different surgical feature vectors is calculated using the following formula: in, and are two surgical feature vectors; and for and Corresponding surgical terms; is the cosine similarity of the feature vector; in, For term and The number of shortest path edges in the medical knowledge graph, w is the path distance weight factor; in, For term and The nearest common ancestor node in the medical knowledge graph, and Respectively represent nodes and The distance to the root node of the knowledge graph; α, β, and γ are weight coefficients.

6. The method for generating a settlement list based on medical data according to claim 4, characterized in that: Construct the feature vector space, including: The surgical operation type restrictions in the medical insurance settlement rules are converted into a binary relationship matrix R, where Indicates that operation i is compatible with operation j, Indicates that operation i and operation j are mutually exclusive; Constructing a directed acyclic graph based on the hierarchical relationship of surgical operation types , where V is the set of operation types and E is the set of hierarchical constraint edges; According to the binary relationship matrix R and the directed acyclic graph ,A three-level tree index structure is constructed, where the first level is divided by surgical system, the second level is divided by interventional method, and the third level is divided by operation method; Assign multi-dimensional spatial coordinates to the nodes in the three-level tree index structure to form an index structure containing spatial location information; The semantic similarity matrix is ​​mapped to an index structure containing spatial location information to obtain a feature vector space.

7. The method for generating a settlement list based on medical data according to claim 6, characterized in that: The surgical system refers to the division of surgical areas based on the human anatomy classification, including the circulatory system, digestive system, nervous system, musculoskeletal system and urogenital system; The technical path for implementing the interventional surgery; The methods of operation include excision, suturing, repair, reconstruction, replacement, implantation, ablation and decompression.

8. The method for generating a settlement list based on medical data according to claim 6, characterized in that: An index structure containing spatial location information is formed, including: Set the number of dimensions n of the feature space; Construct a standard n-dimensional Hilbert space-filling curve H; The n-dimensional Hilbert space filling curve H is deformed and adjusted using the binary relationship matrix R to obtain the curve ; Among them, the curve segments corresponding to the operations marked as compatible in the matrix R are adjacent, and the curve segments corresponding to the operations marked as mutually exclusive are separated; According to the hierarchical relationship of the three-level tree index structure, for each node along the curve Allocate space area.

9. The method for generating a settlement list based on medical data according to claim 8, characterized in that: For each node along the curve Allocate space areas, including: According to the classification of surgical systems, the curve The main trajectory of is divided into multiple non-overlapping convex space blocks, each of which is assigned to a first-level node; In the spatial area block corresponding to each first-level node, according to the classification of intervention methods, along the curve The sub-trajectory of is further divided into multiple continuous area segments, each of which is assigned to a second-level node; In the area segment corresponding to each second-level node, according to the classification of the operation method, along the curve The terminal trajectory assigns spatial coordinate points to each third-level node.

10. A system for generating a settlement list based on medical data, characterized in that: include: At least one processing unit; used to execute instructions to implement the method for generating a settlement list based on medical data as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Medical record diagnosis and operation ICD coding method based on mapping knowledge domain

    CN119230090A

  • Medical ICD coding classification method based on graph neural network

    CN119474981A

  • Medical text big data intelligent labeling and knowledge graph construction method and system

    CN119851968A

  • Medical providers knowledge base and interaction website

    US20130096937A1