A method and system for generating a settlement list based on medical data

By introducing medical insurance settlement rules constraints in the feature vector space, combining medical terminology relationships and multi-dimensional index structures, the problem of inconsistent coding standards is solved, and the consistency and efficient matching of coding in medical semantics and medical insurance settlement rules is achieved.

CN120068802BActive Publication Date: 2025-07-18SHENZHEN COMBIT INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510545442.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-18
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

When generating medical insurance settlement lists in the prior art, the encoding standards are not uniform, resulting in the generated encodings that may be reasonable but do not meet the settlement standards, and cannot meet the requirements of medical semantic similarity and medical insurance settlement rules at the same time, and there is a problem of inconsistent encodings.

Method used

By introducing medical insurance settlement rules constraints in the feature vector space, using knowledge-enhanced hierarchical clustering algorithm and multi-dimensional index structure, a feature vector space is constructed, combining the synonym relationship, upper and lower relationship and concurrent relationship of medical terms, and using compound similarity calculation formulas and binary relationship matrix R to ensure that the encoding meets both medical semantic similarity and medical insurance settlement rules.

Benefits of technology

The generated encoding is realized not only reasonable in medical semantics, but also complies with the medical insurance settlement standards, improves the accuracy and processing efficiency of coding matching, reduces the system's computing burden, and solves the problem of inconsistency between coding and settlement standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068802B_ABST
    Figure CN120068802B_ABST
Patent Text Reader

Abstract

The present application discloses a method and system for generating a settlement list based on medical data, relating to the field of data processing, including: obtaining structured medical data and unstructured medical data; generating normalized data according to the structured medical data; generating standardized text according to the unstructured medical data; fusing the normalized data and the standardized text; constructing a feature vector space according to the fused feature vectors; establishing a mapping relationship of surgical codes by using a medical ontology knowledge base; performing entity recognition on the fused feature vectors; generating surgical operations and corresponding operation codes according to the entity recognition results and the mapping relationship with the surgical codes, and generating a medical insurance settlement list. Aiming at the problem in the prior art that the generated codes may be reasonable but do not meet the settlement standards due to the non-uniform coding standards, the present application enables the codes to meet the requirements of both medical semantic similarity and medical insurance settlement rules.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and particularly to a method and system for generating a settlement list based on medical data. Background Art

[0002] With the continuous improvement of the medical security system, the role of the medical insurance settlement list in medical insurance payment has become increasingly important. As a key link between medical institutions and medical insurance departments, it is directly related to the economic benefits of medical institutions and the rational use of medical insurance funds. In the medical insurance settlement list, the operation and procedure codes and disease diagnosis codes are two of the most critical elements. The accuracy of the operation and procedure codes directly affects the score calculation of cases. Under-coding (insufficient coding) will cause medical institutions to fail to obtain due compensation, while over-coding (excessive coding) may result in unreasonable expenditure of medical insurance funds and even trigger compliance risks. To ensure that the medical insurance settlement list can objectively reflect the consumption of medical resources, medical institutions need to comprehensively analyze unstructured data including patient admission and discharge records, operation records, etc., as well as structured data such as medical service price items, medication and consumable usage, and convert this heterogeneous data into standardized operation and procedure codes. This process poses relatively high requirements for both data processing technology and medical expertise.

[0003] There are differences in medical insurance policies and coding standards in different regions. The operation coding systems used within medical institutions (such as ICD-9-CM-3, ICD-10-PCS, etc.) often do not fully correspond to the coding standards required for medical insurance settlement. Existing technologies do not incorporate medical insurance settlement rules as constraint conditions into the construction of the feature space during the coding conversion process, resulting in generated codes that may be reasonable in medical semantics but do not conform to local medical insurance settlement standards, ultimately leading to settlement disputes.

[0004] In addition, existing technologies usually use flat coding mappings or simple classification trees, lacking the expression of the internal hierarchical relationship of operations, and unable to reflect multi-dimensional characteristics such as surgical systems, intervention methods, and operation methods. As a result, similar operations may obtain extremely different codes, which is not conducive to settlement analysis and auditing. Summary of the Invention

[0005] In view of the problem in the prior art that the generated codes may be reasonable but do not conform to the settlement standards due to inconsistent coding standards, this application provides a method and system for generating a settlement list based on medical data. By introducing medical insurance settlement rule constraints into the feature vector space, the generated codes can meet the requirements of both medical semantic similarity and medical insurance settlement rules.

[0006] The objectives of this application are achieved through the following technical solutions.

[0007] One aspect of the present application provides a method for generating a settlement list based on medical data, including: S1, obtaining structured medical data and unstructured medical data; S2, preprocessing the structured medical data to generate normalized data; performing text standardization processing on the unstructured medical data to generate standardized text; S3, fusing the normalized data and the standardized text to obtain a fused feature vector; constructing a feature vector space according to the fused feature vector; S4, establishing a mapping relationship of surgical codes based on the feature vector space by using a medical ontology knowledge base; S5, performing entity recognition on the fused feature vector to obtain an entity recognition result; S6, generating a surgery and corresponding operation code according to the entity recognition result and the mapping relationship with the surgical code; S7, generating a medical insurance settlement list according to the surgery and corresponding operation code.

[0008] Further, the structured medical data includes medical service prices, medication and consumable usage; the unstructured medical data includes patient admission and discharge records and surgical operation records.

[0009] Further, S3, fusing the normalized data and the standardized text to obtain a fused feature vector, includes: extracting quantitative features of the normalized data, where the quantitative features include medical service price features, drug usage features, and consumable consumption features; extracting text features of the standardized text, where the text features include semantic features and surgical term features; aligning the quantitative features and the text features to generate an initial fused feature; using an attention mechanism to perform weighted fusion on the initial fused feature to obtain a fused feature vector.

[0010] Further, constructing a feature vector space according to the fused feature vector includes: performing clustering analysis on the fused feature vector by using a knowledge-enhanced hierarchical clustering algorithm to obtain a feature vector category distribution; the knowledge-enhanced hierarchical clustering algorithm uses the synonym relationship, hyponymy relationship, and co-occurrence relationship of surgical terms as clustering constraint conditions; calculating the semantic similarity between different surgical feature vectors based on the feature vector category distribution to construct a semantic similarity matrix; converting the medical insurance settlement rules into spatial constraint conditions, where the medical insurance settlement rules include surgical operation type restrictions; partitioning the feature space through multi-dimensional indexing according to the semantic similarity matrix and the spatial constraint conditions to construct a feature vector space.

[0011] Among them, the synonymous relationship: This refers to the relationship where different medical terms or surgical descriptions, although having different forms of expression, actually refer to the same medical concept or surgical operation. In medical terms, the synonymous relationship is very common because the same surgery may have different ways of expression. For example, "laparoscopic cholecystectomy" = "laparoscopic gallbladder removal" = "LC"; "appendectomy" = "appendix removal" = "appendix operation"; "PCI" = "percutaneous coronary intervention" = "coronary artery stent implantation".

[0012] The hierarchical relationship: This refers to the hierarchical relationship between medical concepts or surgical operations, where one concept is a more specific form (subordinate relationship) or a more general form (superordinate relationship) of another concept. For example, "gastrectomy" is a subordinate concept of "digestive system surgery", while "digestive system surgery" is a superordinate concept of "gastrectomy". For example, superordinate: "digestive system surgery" → subordinate: "gastrointestinal surgery" → subordinate: "gastrectomy" → subordinate: "distal subtotal gastrectomy"; superordinate: "joint surgery" → subordinate: "knee joint surgery" → subordinate: "knee joint replacement" → subordinate: "total knee joint replacement"; superordinate: "vascular intervention surgery" → subordinate: "coronary artery intervention" → subordinate: "coronary artery stent implantation".

[0013] The co-occurrence relationship: This refers to the relationship between surgical operations that often occur simultaneously or are performed in combination in medical practice. Some surgeries are usually performed together clinically, and there is a statistical correlation or clinical combination necessity between them. For example, the co-occurrence relationship between "cholecystectomy" and "choledochoscopy" (co-occurrence rate is about 65%); the co-occurrence relationship between "hysterectomy" and "bilateral oophorectomy" (co-occurrence rate is about 43%); the co-occurrence relationship between "coronary artery bypass grafting" and "aortic valve replacement" (under specific indications).

[0014] Surgical operation type constraint: This refers to the reimbursement conditions, restrictions, or special regulations of medical insurance policies for specific types of surgeries. It includes requirements for indications of certain surgeries, frequency restrictions, combination restrictions, etc. These restrictions will affect surgical coding and the final medical insurance settlement. For example, type constraint: mutually exclusive coding restrictions between minimally invasive surgery and open surgery; frequency restriction: the medical insurance payment for "intraocular lens implantation" is limited to once per eye per life cycle; combination restriction: "simple modified radical mastectomy" and "sentinel lymph node biopsy" cannot be coded and reimbursed simultaneously; level restriction: "transcatheter aortic valve replacement" can only be coded and used in tertiary hospitals.

[0015] Furthermore, the semantic similarity between different surgical feature vectors is calculated using the following formula:

[0016] Sim(V1, V2) = α × Sim cos (V1, V2) + β × Sim path (T1, T2) + γ × Sim lca (T1, T2), where V1 and V2 are two surgical feature vectors, respectively representing the feature representations corresponding to different surgeries obtained in step S41; T1 and T2 are the surgical terms corresponding to V1 and V2, referring to the standardized surgical names in the medical knowledge graph; Sim cos (V1, V2) is the cosine similarity of the feature vectors; where ShortestPath(T1, T2) is the number of the shortest path edges between terms T1 and T2 in the medical knowledge graph, and w is the path distance weight factor used to adjust the influence degree of the path distance on the similarity, with the value range [0.5, 2] for different types of surgeries; where lca(T1, T2) is the nearest common ancestor node of terms T1 and T2 in the medical knowledge graph, depth(T) represents the distance from node T to the root node of the knowledge graph, with the value range [0, 1], and the larger the value, the closer the taxonomic relationship between the two terms; α, β, and γ are weight coefficients, respectively controlling the proportions of the feature vector similarity, the path distance similarity, and the taxonomic similarity in the calculation of the total similarity, and α + β + γ = 1, with the value range of α being [0.2, 0.4], the value range of β being [0.3, 0.5], and the value range of γ being [0.2, 0.4], which are dynamically adjusted according to different medical scenarios.

[0017] Furthermore, according to the semantic similarity matrix and the spatial constraint conditions, the feature space is partitioned through multi-dimensional indexing to construct a feature vector space, including: converting the surgical operation type restrictions in the medical insurance settlement rules into a binary relation matrix R, where R(i, j) = 1 indicates that operation i is compatible with operation j, and R(i, j) = 0 indicates that operation i is mutually exclusive with operation j; constructing a directed acyclic graph G = (V, E) according to the hierarchical relationship of the surgical operation types, where V is the set of operation types and E is the set of hierarchical constraint edges; constructing a three-level tree index structure according to the binary relation matrix R and the directed acyclic graph G = (V, E), where the first level is divided by the surgical system, the second level is divided by the intervention method, and the third level is divided by the operation method; allocating multi-dimensional space coordinates to the nodes in the three-level tree index structure to form an index structure containing spatial position information; mapping the semantic similarity matrix to the index structure containing spatial position information to obtain the feature vector space.

[0018] Among them, the existing technologies often only focus on medical semantic similarity and ignore the constraints of medical insurance settlement rules, resulting in that although the coding is medically reasonable, it does not meet the settlement standards. This application transforms the medical insurance settlement rules into spatial constraint conditions, introduces a binary relation matrix R to represent operation compatibility, and imposes settlement rule constraints when constructing the feature vector space, so that the generated coding satisfies both medical semantic similarity and medical insurance settlement rules from the design stage, fundamentally solving the problem of inconsistent coding standards.

[0019] Furthermore, the surgical system represents the surgical area division based on human anatomy classification, including the circulatory system, digestive system, nervous system, musculoskeletal system, and urogenital system; the technical path for implementing the interventional surgery; the operation methods include but are not limited to resection, suture, repair, reconstruction, replacement, implantation, ablation, and decompression.

[0020] Furthermore, multi-dimensional spatial coordinates are assigned to the nodes in the three-level tree-shaped index structure to form an index structure containing spatial position information, including: setting the number of dimensions n of the feature space; constructing a standard n-dimensional Hilbert space-filling curve H; using the binary relation matrix R to deform and adjust the n-dimensional Hilbert space-filling curve H to obtain the curve H'; among them, the curve segments corresponding to the operations marked as compatible in the matrix R are adjacent, and the curve segments corresponding to the operations marked as mutually exclusive are separated; according to the hierarchical relationship of the three-level tree-shaped index structure, spatial regions are assigned to each node along the curve H'.

[0021] Among them, the n-dimensional Hilbert space-filling curve is a special continuous fractal curve that can recursively traverse and fill each point in the n-dimensional space while maintaining the relative proximity of neighboring points in space and sequence, thus mapping the high-dimensional space to a one-dimensional continuous path. In this application, multi-dimensional medical features (such as anatomical parts, operation methods, surgical difficulty, etc.) are uniformly mapped to the continuous curve; unique curve coordinates are assigned to each coding node to support range queries and nearest neighbor searches; a spatially locally sensitive index structure is constructed to improve the feature retrieval speed and accuracy. For 2 m discrete points in the n-dimensional space, the m-order Hilbert curve H can be defined by the following recursive function: H m : [0, 2 m ^n - 1] → [0, 1] n , where m represents the order of the curve, n represents the spatial dimension, and the function H m maps the one-dimensional index to the n-dimensional space coordinates.

[0022] Deformation adjustment is a space-filling curve topology optimization technique based on domain knowledge constraints. By applying the compatibility constraints in the binary relation matrix R to the standard Hilbert curve, it changes the shape and density of the curve in specific regions, making the curve segments of compatible operations approach each other while those of mutually exclusive operations move away from each other. In this application, the curve segments corresponding to the compatible operation coding pairs (such as "cholecystectomy" and "choledochoscopy") are pulled closer; the curve segments corresponding to the mutually exclusive operation coding pairs (such as "open surgery" and "laparoscopic surgery") are pulled farther apart; the medical insurance policy restriction conditions are transformed into boundary constraints for curve deformation; and the clinical knowledge of common surgical combinations is incorporated into the curve shape optimization objective. Specifically, deformation adjustment can be expressed as finding a mapping function F such that: H' = F(H, R); where F minimizes the objective function: d represents the spatial distance, and f represents the conversion function from the compatibility value to the ideal distance. Through this deformed n-dimensional Hilbert curve H', a feature space index structure with medical semantic perception is constructed, enabling the three-level tree index to contain not only hierarchical relationships but also complex medical knowledge such as compatibility and mutual exclusivity.

[0023] Furthermore, according to the hierarchical relationship of the three-level tree index structure, spatial regions are assigned to each node along the curve H', including: dividing the main trajectory of the curve H' into multiple non-overlapping convex spatial region blocks according to the classification of the surgical system, and each region block is assigned to a first-level node; within the spatial region block corresponding to each first-level node, according to the classification of the intervention method, the sub-trajectory along the curve H' is further divided into multiple continuous region segments, and each region segment is assigned to a second-level node; within the region segment corresponding to each second-level node, according to the classification of the operation method, along the end trajectory of the curve H', spatial coordinate points are assigned to each third-level node; where the main trajectory is the path of the Hilbert space-filling curve H' at the highest level, used to divide the largest-grained spatial region blocks, corresponding to the first-level nodes (surgical system classification) in the three-level tree structure. For example, cardiovascular surgeries (such as coronary artery bypass grafting, valve replacement) are mapped to the large convex region at the starting segment of the curve; digestive system surgeries (such as gastrectomy, colectomy) are mapped to the continuous region blocks in the middle segment of the curve; orthopedic system surgeries (such as joint replacement, spinal fusion) are mapped to the corresponding regions in the latter segment of the curve; special transition regions are set at the intersections of the main trajectories for rare combined surgeries between different systems. The main trajectory can be expressed as a first-level piecewise function of H': H′ main ={H′[0,t1),H′[t1,t2),.....,H′[t m-1 ,1]}, where t1,t2,......,t m-1 are the division points, corresponding to the boundaries of different surgical systems.

[0024] A sub-trajectory is a secondary path branch of the Hilbert curve H' within the regional blocks divided by each main trajectory, used to further divide the regional segments, corresponding to the second-level nodes (interventional method classification) in the three-level tree structure. For example, within the regional block of the digestive system, laparoscopic surgery is mapped to a continuous regional segment; open surgery is mapped to another adjacent but independent regional segment; endoscopic procedures are mapped to a third regional segment; hybrid interventional methods (such as laparoscopic-assisted mini-incision surgery) are mapped to the transitional area at the boundary of the interventional methods. For the i-th regional block of the main trajectory, its sub-trajectory can be expressed as:

[0025] H′ sub,i ={H′[t i-1 +δ i,0 ,t i-1 +δ i,1 ),H′[t i-1 +δ i,1 ,t i-1 +δ i,2 ),.....,H′[t i-1 +δ i,k-1 ,t i},where δ i,j represents the j-th sub-region division point within the i-th main region.

[0026] The terminal trajectory is the terminal path direction of the Hilbert curve H' at the finest granularity, used to allocate specific spatial coordinate points within the regional segments divided by the sub-trajectories, corresponding to the third-level nodes (operation method classification) in the three-level tree structure. For example, within the "laparoscopic digestive system" regional segment, operations of the "resection" type are mapped to a group of adjacent coordinate points; operations of the "anastomosis" type are mapped to another group of adjacent coordinate points; operations of the "biopsy" type are mapped to the third coordinate point region; composite operations (such as "resection + reconstruction") are mapped to the exact positions between the corresponding operation coordinate points. For the j-th sub-region within the i-th main region, its terminal trajectory can be expressed as a discrete point set:

[0027] H′ end,i ={H′(t i-1 +δ i,j-1 +ε i,j,0 ), H′(t i-1 +δ i,j-1 +ε i,j,1 ),......, H′(t i-1 +δ i,j-1 +ε i,j,l-1 )}, where ε i,j,kIt represents the precise positioning point of the k-th operation method in the j-th sub-region of the i-th main region. In this application, this Hilbert curve allocation method for a multi-level trajectory structure creates a coding index system that not only maintains medical semantic consistency but also has efficient spatial retrieval capabilities by directly mapping the hierarchical structure of the medical surgery classification system (system - intervention method - operation method) to different granularity trajectories of the space-filling curve. This structure not only supports precise coding queries but also enables similarity surgery retrieval based on spatial proximity relationships and spatial constraint verification of medical insurance settlement rules.

[0028] Further, in S6, according to the entity recognition result and the mapping relationship with the surgery code, generate the surgery and the corresponding operation code, including: perform coding mapping on the entity recognition result according to the mapping relationship of the surgery code to obtain the preliminary code; according to the preliminary code, determine the category code, system code, and method code of each surgery according to the hierarchical division of the three-level tree index structure to form a hierarchical code; use the binary relation matrix R to verify whether the hierarchical code meets the compatibility constraints of the matrix R; use the positional relationship in the eigenvector space to allocate the final operation code to the hierarchical code combination that passes the verification;

[0029] Another aspect of this application also provides a system for generating a settlement list based on medical data, which is used to execute a method for generating a settlement list based on medical data in this application.

[0030] Compared with the prior art, the advantages of this application are as follows:

[0031] The medical insurance settlement rules are not incorporated as constraint conditions into the construction of the feature space, resulting in the generated codes being possibly reasonable but not meeting the settlement standards. In this application, by converting the surgical operation type restrictions in the medical insurance settlement rules into the binary relation matrix R and imposing these constraint conditions during the construction of the eigenvector space, the organizational structure of the coding space itself contains the constraint information of the settlement rules, thus ensuring that the generated operation codes are not only reasonable in medical semantics but also meet the standards and specifications of medical insurance settlement, fundamentally solving the problem of inconsistent coding and settlement standards.

[0032] In the prior art, when calculating surgical similarity, generally a simple bag-of-words model or a basic TF-IDF vector cosine similarity is used. However, these methods only focus on the surface features of the text and ignore the synonymous, hyponymous, and co-occurrence relationships existing between terms, affecting the subsequent precise matching of the codes. In this application, by designing a composite similarity calculation formula

[0033] Sim(V1, V2) = α × Sim cos (V1, V2) + β × Sim path (T1, T2) + γ × Sim lca(T1, T2), by comprehensively considering the cosine similarity of feature vectors, the path distance similarity in the medical knowledge graph, and the taxonomic similarity of the nearest common ancestor nodes, and taking the semantic relationship of surgical terms as the constraint condition for hierarchical clustering, the accurate measurement of surgical similarity in the medical professional field is realized, and the accuracy of coding matching and the consistency of medical semantics are improved.

[0034] In the prior art, simple hash mapping or linear index structures are generally used, which have the defects of scattered coding and difficulty in obtaining similar codes for similar operations. In this application, an n-dimensional Hilbert space filling curve is constructed and deformed and adjusted according to the binary relation matrix R, so that the curve segments of compatible operations are adjacent and the curve segments of mutually exclusive operations are separated, ensuring the continuity and locality of the coding space, improving the coding query and processing efficiency, and reducing the system computing burden. Brief Description of the Drawings

[0035] This application will be further described in the form of exemplary embodiments, and these exemplary embodiments will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where:

[0036] Figure 1 is an exemplary flowchart of a method for generating a settlement list based on medical data according to some embodiments of this application;

[0037] Figure 2 is an exemplary flowchart of generating a fused feature vector according to some embodiments of this application;

[0038] Figure 3 is an exemplary flowchart of constructing a feature vector space according to some embodiments of this application;

[0039] Figure 4 is an exemplary flowchart of allocating space coordinate points according to some embodiments of this application. Detailed Embodiments

[0040] The methods and systems provided in the embodiments of this application will be described in detail below with reference to the drawings.

[0041] As Figure 1As shown, structured medical data and unstructured medical data are obtained; structured medical data are preprocessed to generate standardized data; unstructured medical data is subjected to text standardization processing to generate standardized text; standardized data and standardized text are fused to obtain fused feature vectors; feature vector space is constructed according to the fused feature vectors; based on the feature vector space, a mapping relationship of surgical codes is established using a medical ontology knowledge base; entity recognition is performed on the fused feature vectors to obtain entity recognition results; based on the entity recognition results and the mapping relationship with the surgical codes, the surgical operations and corresponding operation codes are generated; based on the surgical operations and corresponding operation codes, a medical insurance settlement list is generated.

[0042] S1, obtain structured medical data and unstructured medical data. Among them, obtain structured medical data, medical service price data acquisition: through the data interface of the hospital information system (HIS), obtain the detailed data of the patient's charge list during hospitalization, including charge item code, item name, unit price, quantity, billing unit, execution department, cost category and other fields. A combination of scheduled batch extraction and real-time push is adopted to ensure the integrity and timeliness of the data.

[0043] Medication information acquisition: Extract patient medication records from the hospital pharmacy management system, including generic name, trade name, specifications, dosage form, usage and dosage, route of administration, medication time, medication frequency, drug code, etc. For special medications (such as antibiotics, anesthetics, psychotropic drugs, etc.), additionally mark their medication categories for subsequent processing.

[0044] Obtaining the usage of consumables: Through the interface between the operating room management system and the material management system, extract the information of medical consumables used during the patient's operation, including the name, model, specification, quantity, usage time, usage site, consumable classification, etc. For high-value consumables, additionally record their batch number and barcode information to establish a traceability link.

[0045] Acquisition of unstructured data, patient admission and discharge records: Extract patient admission records, discharge summaries, medical records and other documents from the electronic medical record system (EMR), which are usually stored in XML, HTML or PDF format. Data extraction is achieved through custom interfaces or HL7 standard interfaces, retaining the original structure and format information of the document.

[0046] Surgical operation record acquisition: Extract surgical records, anesthesia records, surgical nursing records and other documents from the surgical information system. These documents contain key information such as surgical method, operation site, surgical process description, surgical time, anesthesia method, etc. For multimedia materials (such as surgical images), extract their metadata information as auxiliary reference.

[0047] S2. Preprocess the structured medical data to generate normalized data; perform text standardization on the unstructured medical data to generate standardized text. Among them, when preprocessing the structured medical data, data cleaning: Detect and process outliers, missing values, and duplicate values in the structured data. Identify outliers using statistical methods (such as the 3σ rule), and correct or mark them according to business rules; for missing values, adopt different filling strategies according to the data type, such as mean filling, mode filling, or derivation from previous and subsequent values; perform merging or deletion processing on duplicate values. Data standardization: Convert the structured data in different systems into a unified data format and coding standard. Map medical service price items to the national medical insurance fee catalog; uniformly map drug information to the National Essential Medicines List and the Drug Coding Rules; map medical consumables to the Rules for the Unique Identification System for Medical Devices. Temporal alignment: Align the structured data from different sources along the time axis, establish the temporal relationship between medication use, consumable use, and medical services, and provide a basis for subsequent correlation analysis. Adopt the sliding time window technique to handle the problem of inconsistent timestamp precision. Numerical normalization: Uniformly convert numerical values with inconsistent measurement units, such as converting medication doses in different units to standard units; perform normalization on amount information to eliminate the impact of price changes in different periods.

[0048] When preprocessing the unstructured medical data, document parsing: According to the characteristics of unstructured documents in different formats, adopt corresponding parsing techniques to extract text content. Use DOM or SAX parsers to extract text from XML / HTML documents; use professional PDF parsing tools to extract text content from PDF documents; retain the chapter structure information of the document for subsequent refined processing.

[0049] Text tokenization and standardization: Aiming at the characteristics of medical texts, adopt a medical professional tokenization system for Chinese tokenization to identify professional vocabulary such as medical terms, drug names, and anatomical parts; perform text standardization processing, including unifying full-width / half-width characters, removing special symbols, unifying case, and simplifying and traditional Chinese character conversion; expand medical abbreviations and shorthands, such as expanding "BP" to "blood pressure" and "WBC" to "white blood cell count"; identify and standardize numerical expressions and units, such as standardizing "2mg / kg" to "2 milligrams per kilogram"; Medical term normalization: Map synonymous medical terms to a standard medical term set. Adopt a medical ontology library (such as SNOMED CT, ICD-10, etc.) for term matching, and convert non-standard expressions in the text into normalized expressions, such as unifying "cesarean section", "abdominal cesarean section", "C-section", etc. into the standard term "cesarean section operation".

[0050] Named Entity Recognition: Using named entity recognition technology in the medical field, identify and label key entities such as surgical names, anatomical locations, operation methods, and diagnosis names from the text, and establish relationships between entities, such as "surgery - location" and "diagnosis - surgery" relationships. Implement high-precision entity recognition using BiLSTM-CRF or pre-trained language models in the medical field. Text Structuring: Based on the identified medical entities and relationships, convert unstructured text into a semi-structured representation form, extract key facts and attributes, such as "operation method", "operation location", "operation purpose", "instruments used" and other attributes of the surgery, to form a structured event description.

[0051] As Figure 2 shown in S3, fuse the normalized data and standardized text to obtain the fused feature vector, including: First, extract quantitative features. The medical service price features can be extracted by constructing service item vectors, including item frequency, cumulative cost, average cost per time, and standard deviation; introduce time series features to record the time distribution and phased features of item execution; calculate the proportion distribution of various medical services to reflect the resource consumption structure.

[0052] Extract drug use features, construct drug use vectors according to pharmacological classification, including the number of drug types, days of medication, and total dose; extract the medication time series pattern to identify the corresponding relationship between the medication procedure and the treatment stage; calculate the use intensity indicators of special drugs (such as antibiotics and anesthetic drugs).

[0053] Extract consumable consumption features, construct consumable consumption vectors, including the use of high-value consumables and the consumption of regular consumables; extract the combined pattern of consumable use to identify the set of consumables related to specific surgeries; calculate the value distribution of various consumables as an indirect indicator of surgical complexity.

[0054] Then extract text features. The semantic features can be extracted by using pre-trained language models in the medical field to extract the context semantic representation of the text, using the attention mechanism to extract the key sentence semantic vectors in the surgical description, and calculating the density and complexity of professional terms in the text to reflect the technical requirements of the surgery.

[0055] Extract term features, identify key terms such as surgical names, operation locations, and operation methods in the text; construct a term co-occurrence matrix to capture the combined pattern between terms; extract the hierarchical relationship features of terms to reflect the classification information of surgeries.

[0056] Feature fusion is performed to establish the time correspondence between structured data and text descriptions. Spatial alignment is carried out according to the medical service execution department and the text record department to construct an alignment matrix and record the corresponding strength of data from different sources. The aligned quantitative features and text features are concatenated to form an initial feature vector. Principal Component Analysis (PCA) is applied to reduce the feature dimension and eliminate redundant information, and feature normalization is used to eliminate the influence of the dimension of different types of features. A multi-head attention mechanism is designed to calculate the importance weights of different feature components, and the attention weight distribution is dynamically adjusted according to the current medical scenario. The final fused feature vector is obtained through weighted summation, which comprehensively expresses the multi-dimensional characteristics of the patient's medical service.

[0057] As Figure 3 shown, a feature vector space is constructed. First, clustering constraints are designed. The synonymous relationships in the medical term ontology are transformed into hard constraints of "must be clustered together", the hyponymy relationships are transformed into structural constraints of "hierarchical attribution", and the co-occurrence relationships are transformed into soft constraints of "similarity enhancement". The agglomerative hierarchical clustering algorithm is used. Initially, each surgical feature vector is regarded as an independent category. In each merging decision, knowledge constraints are incorporated, and the categories that meet the constraint conditions are preferentially merged. The modified Ward's minimum variance method is used to calculate the distance between classes to reduce the influence of outliers, and a hierarchical tree structure is generated to reflect the classification pedigree of surgical features. Through knowledge-enhanced hierarchical clustering, the problem that traditional clustering methods ignore prior knowledge in the medical field is effectively solved, and the formed category distribution is more in line with medical professional cognition.

[0058] Specifically, on the one hand, traditional vector similarity calculations (such as TF-IDF, Word2Vec, etc.) only focus on lexical or statistical features and cannot capture the complex semantic relationship network among medical terms, resulting in the misclassification of terms that are superficially similar but have very different medical meanings. On the other hand, there is a lack of a technical path to organically combine medical ontology knowledge with numerical calculations, and the existing medical classification system and term relationship network cannot be used to enhance similarity calculation. This application designs a composite similarity calculation formula:

[0059] Sim(V1, V2) = α × Sim cos (V1, V2) + β × Sim path (T1, T2) + γ × Sim lca (T1, x x).

[0060] The vector cosine similarity Sim cos(V1, V2): Calculate the cosine angle of two surgical feature vectors, which represents the similarity of traditional feature vectors and measures the direction consistency of two surgical feature vectors in a high-dimensional space. Its mathematical basis is the cosine formula of the included angle in the vector space model, with high computational efficiency and easy parallel processing. By adaptively adjusting the feature dimension weights, this component can highlight the contribution of key features to similarity and reduce the interference of noise features.

[0061] Path distance similarity: Utilize the concept association network in the medical knowledge graph to calculate the shortest path between terms. The path distance similarity formula: Convert the concept connections in the medical knowledge graph into similarity metrics. Encode the knowledge of term relationships scattered in medical literature and guidelines into a graph structure, enabling the calculation process to consider the semantic connection strength between terms. The introduced dynamic path weight w parameter can precisely regulate the influence degree of path distance in different surgical fields, solving the problem of uneven term connection density in different medical sub-domains.

[0062] Taxonomic similarity: Evaluate the genetic relationship of terms in the classification system based on the depth of the nearest common ancestor node. The taxonomic similarity formula: Evaluate the systematic relationship of terms in the classification tree. Quantify the hierarchical classification information into a similarity index, which is particularly suitable for fields with a strict classification system such as medical surgeries. By calculating the ratio of the depth of the common ancestor to the depth of the term itself, it effectively balances the comparison benchmarks of terms at different levels and solves the asymmetry problem in traditional hierarchical similarity calculations.

[0063] In summary, SimCos captures the similarity relationship at the feature representation level; SimPath captures the semantic connection relationship between terms;

[0064] Sim lca (T1, T2) captures the classification hierarchical relationship between terms. Incorporate the structured relationship information in the medical knowledge graph into the similarity calculation, making up for the knowledge gap of the pure vector model. Through dynamic weight adjustment, it adaptively balances the importance of different semantic relationships to meet the similarity evaluation requirements of different types of surgeries.

[0065] Dynamically adjust the α, β, γ weight parameters according to different medical scenarios; for routine surgeries with high standardization, increase the weight of vector similarity (α); for new and complex surgeries, increase the weight of knowledge graph path similarity (β); for multi-system joint surgeries, increase the weight of taxonomic similarity (γ). The composite similarity calculation method solves the problem that traditional similarity calculations ignore the special semantic relationships between medical terms. By integrating vector space similarity and knowledge graph relationships, it realizes the accurate measurement of medical surgery similarity.

[0066] Specifically, there is a fundamental technical contradiction in the field of medical insurance settlement list generation: the coding needs to satisfy both the medical semantic rationality and the compliance with medical insurance policies at the same time. The existing medical insurance settlement rules usually exist in the form of text, lacking an effective mechanism to convert these rules into computable constraint conditions, making it difficult for the system to automatically apply these rules. Secondly, in the traditional vector space, surgeries with similar semantics are mapped to adjacent positions, but this arrangement of semantic similarity often conflicts with the constraints of medical insurance policies, resulting in non-compliant coding. Finally, the medical coding space is highly complex, and traditional indexing structures are difficult to ensure both query efficiency and semantic fidelity, especially when multiple constraint conditions need to be considered simultaneously.

[0067] Therefore, in this application, first, the surgical operation type restrictions in the medical insurance settlement rules are systematically converted into a binary relation matrix R. Specifically, extract the surgical operation restriction rules from medical insurance policy documents (such as "Administrative Measures for Medical Insurance Diagnosis and Treatment Items", "Surgical Operation Specifications", etc.); then, construct an n×n binary relation matrix R (n is the total number of surgical operation types), where R(i,j)=1 indicates that operation i is compatible with operation j and can be coded simultaneously; R(i,j)=0 indicates that operation i and operation j are mutually exclusive and should not be coded simultaneously.

[0068] For the mutually exclusive relationships clearly stipulated in the policy (such as "resection and repair should not be coded simultaneously at the same surgical site"), directly set the corresponding matrix elements to 0; for surgeries with a sequential dependence relationship in clinical practice but not clearly stipulated in the policy (such as "vascular resection" and "vascular reconstruction"), identify their co-occurrence patterns by analyzing clinical data and set the corresponding matrix elements to 1.

[0069] For fuzzy relationships, introduce a fuzzy compatibility degree [0, 1] to reflect the strength of policy constraints. For example, for operation pairs that "can be combined coded depending on the situation", it may be set that R(i,j)=0.7, indicating medium to high compatibility but further clinical judgment is required. This fuzzy value is obtained through medical insurance expert scoring and analysis of historical settlement success rates. This application converts the medical insurance policies expressed in natural language into an accurate mathematical representation, achieving a qualitative change from "rule description" to "computational constraint". Secondly, by introducing the fuzzy compatibility degree in the [0, 1] interval, the problem of coexistence of "strong constraints" and "weak suggestions" in medical insurance policies is solved, enabling the system to handle policy guidance of different intensities. Finally, the matrix form is convenient for efficient computer processing, supports large-scale matrix operations and parallel computing, providing a basis for subsequent space construction.

[0070] Then, based on the medical ontology and the medical insurance classification standard, a directed acyclic graph G=(V, E) of surgical operation types is constructed. Specifically, ontology engineering and knowledge graph technology are applied to construct a hierarchical relationship graph. First, the concept hierarchy of surgical operations is extracted from standard medical terminology systems (such as SNOMED CT, ICD-10-PCS); then, combined with the classification system in the medical insurance coding manual, a hierarchical relationship reflecting local medical insurance policies is constructed.

[0071] In graph G, node V represents the operation type, from high-level concepts (such as "surgical operations on the circulatory system") to specific operations (such as "coronary artery bypass grafting"); edge E represents the hierarchical constraint relationship, including "is-a" relationships (such as "coronary artery bypass is a type of coronary artery surgery") and composite relationships (such as "operation combinations that need to be used jointly").

[0072] Conflicts and loops in the hierarchical relationship are detected and eliminated through graph algorithms. For example, if an operation is reclassified after the update of the medical insurance policy, which may lead to conflicts in the hierarchical relationship, the system will detect such conflicts through transitive closure calculation and remind the administrator to make adjustments. On the one hand, this application represents the operation type and hierarchical constraints through node V and edge E respectively, converting the complex classification system into a computable graph structure. Using the concepts of reachability and transitive closure in graph theory, the automatic derivation of "upper-level rules constraining lower-level operations" in medical insurance policies is realized, greatly reducing the redundant storage of rules. On the other hand, loops and contradictions are detected through graph algorithms to ensure the consistency of coding rules and avoid conflicts and contradictions in policy interpretation.

[0073] A three-level tree index structure is constructed, and hierarchical clustering and index optimization techniques are used to construct the index structure. First, according to the hierarchical relationship in the directed acyclic graph G, the boundaries of the three-level classification are determined; then, the operation compatibility between nodes at each level is analyzed to optimize the branch structure of the index tree. The first level is divided according to the surgical system, including the circulatory system, digestive system, nervous system, etc., reflecting the classification of surgical areas based on human anatomy. The system boundaries are automatically extracted by parsing the chapter structures of medical textbooks and surgical specifications and verified by domain experts.

[0074] The second level is divided according to the intervention method, such as open surgery, endoscopic surgery, interventional therapy, etc., reflecting the technical path of surgical implementation. The classification of intervention methods is automatically identified by analyzing the operation route, used instruments, and technical characteristics in the surgical description.

[0075] The third level is divided according to the operation method, including resection, suture, repair, reconstruction, replacement, implantation, ablation, and decompression, etc., reflecting the specific surgical behavior. The operation method is determined by semantic analysis of the surgical record to identify the core verb and the operation object.

[0076] Specifically, for common multi-system joint surgeries (such as "thoracoabdominal combined surgery"), cross-system index links are established, enabling access from multiple primary nodes. Each level captures different dimensional characteristics of the surgery, from anatomical location to technical path to specific operations, achieving a comprehensive representation of the multi-dimensional characteristics of the surgery. Using the hierarchical pruning ability of the tree structure, irrelevant branches can be quickly excluded during the query process, significantly improving the retrieval efficiency of large-scale coding libraries.

[0077] Next, multi-dimensional space coordinates are assigned to the nodes in the three-level tree-shaped index structure to form an index structure containing spatial location information. First, the optimal dimension is determined through principal component analysis (PCA) and feature importance evaluation. First, feature extraction is performed on the existing surgical dataset to obtain initial high-dimensional features (usually containing 100 - 200 features); then PCA analysis is applied to select the top n principal components with a cumulative explained variance of over 95%; finally, in combination with the evaluation of experts in the field of medical insurance coding, the final number of dimensions is determined. For example, surgeries of the circulatory system may require a higher dimension (n = 12) to express their complexity, while relatively simple surgeries of the skin system may only require a lower dimension (n = 8).

[0078] Then, a Hilbert curve is generated using a recursive construction method. First, the n-dimensional space is divided into 2 n equal hypercubes; then the access order of these hypercubes is determined so that adjacent hypercubes are also adjacent in space; the same division and sorting rules are recursively applied to each hypercube until a preset precision level is reached (usually 16 - 20 levels, which can represent approximately 2 20 points). The Gray code technique is used to optimize the direction change of the curve to ensure a smooth transition at the turning points of the curve. To meet the special requirements of medical coding, based on the standard Hilbert curve, the rotation matrix in curve generation is adjusted to have a higher resolution in the dimensions related to the surgical system.

[0079] Then, the graph embedding algorithm and the force-directed model are applied to achieve curve deformation. First, the binary relation matrix R is converted into a weighted graph, where an attractive force is added between node pairs with R(i,j) = 1, and a repulsive force is added between node pairs with R(i,j) = 0; then, under the constraint of maintaining curve continuity, the positions of points on the curve are adjusted through an iterative optimization algorithm (such as the spring-charge model); finally, an order-preserving transformation is applied to maintain the local properties of the curve, ensuring that the adjusted curve H' is still a valid space-filling curve.

[0080] For relationships with fuzzy compatibility, elastic coefficient adjustment is adopted. The higher the compatibility, the larger the spring constant and the stronger the attraction. For example, when two surgical operations (such as "intestinal resection" and "intestinal anastomosis") are marked as highly compatible (R(i,j) = 0.9) in the medical insurance rules, the corresponding curve segments will be pulled closer by a stronger attraction; while for mutually exclusive operations (such as "coronary artery bypass grafting" and "percutaneous coronary intervention"), they usually should not be coded simultaneously for the same patient during the same hospitalization, and the corresponding curve segments will be pushed apart by the repulsive force.

[0081] Map the pre-computed semantic similarity matrix into an index structure containing spatial location information to form the final feature vector space. Specifically, multi-objective optimization and topology-preserving mapping techniques are used to achieve the mapping. First, construct the semantic similarity matrix of surgical operations, where the element S(i,j) represents the semantic similarity degree between operations i and j; then, combine this similarity constraint with the aforementioned spatial coordinate assignment, and adjust the spatial coordinate distribution through an optimization algorithm. The specific implementation uses an optimization method that combines gradient descent and simulated annealing. The objective function is designed as:

[0082] where dist(i,j) is the Euclidean distance between points i and j in space, S(i,j) is the semantic similarity, R(i,j) is the binary relation value, λ1 and λ2 are weight coefficients for balancing semantic similarity and policy compatibility, and α and β are scaling factors.

[0083] Through such optimization, the system achieves a balance of three goals: 1) Surgical operations with similar semantics are close in spatial position; 2) Surgical operations that are compatible with medical insurance rules are adjacent in spatial position; 3) Surgical operations that are mutually exclusive in medical insurance rules are separated in spatial position. Finally, use spatial indexing techniques such as locality-sensitive hashing (LSH) and R-trees to establish an efficient query mechanism for the constructed feature vector space, supporting k-nearest neighbor search and range queries, significantly improving the efficiency of coding matching.

[0084] As Figure 4 shown, according to the hierarchical relationship of the three-level tree-shaped index structure, allocate spatial regions for each level of nodes along the deformed curve H'. First-level spatial region division: Based on the importance and complexity of the surgical system, determine the spatial allocation ratio of each system. For example, the circulatory system accounts for 20%, the digestive system accounts for 18%, the nervous system accounts for 15%, etc. Then divide the curve H' along the main trajectory into continuous segments with corresponding ratios, and convert the set of spatial points covered by each segment of the curve into a convex spatial region block through the convex hull algorithm. To ensure the coherence of the regions, apply the minimum spanning tree algorithm to connect the subset of regions that may be separated due to curve bending.

[0085] Second-level spatial region division: Within each first-level spatial region block, determine the spatial allocation ratio according to the medical resource consumption and technical complexity of different intervention methods in the system. For example, within the circulatory system, open surgery may account for 40%, interventional surgery for 35%, and minimally invasive surgery for 25%. Then, further divide it into multiple consecutive region segments along the sub-trajectory of H' within this region block. Use the interval tree data structure to record the boundary coordinates of these region segments for subsequent rapid positioning.

[0086] Third-level spatial point allocation: Within each second-level region segment, allocate specific spatial coordinate points for each third-level node along the end trajectory of H' according to the fine classification of the operation method. For common operations (such as "resection"), allocate more coordinate points; for rare operations (such as a specific "reconstruction"), only allocate a small number of coordinate points. To improve spatial utilization, apply the density adaptive algorithm to increase the sampling density in key regions. This multi-level spatial allocation strategy not only maintains the spatial locality of the Hilbert curve but also reflects the inherent hierarchical relationship of medical surgery classification through hierarchical division, making the spatial structure itself contain medical knowledge, thereby improving the accuracy of subsequent coding and the efficiency of retrieval.

[0087] S4. According to the feature vector space, establish the mapping relationship of surgical coding using the medical ontology knowledge base; in this embodiment, use UMLS (Unified Medical Language System) and SNOMED CT as the basic medical ontologies, integrate the Chinese Classification Code of Clinical Surgery and Procedures (CCC) and the medical insurance payment coding standard; construct an ontology knowledge base containing 179,000 medical concepts and 347,000 synonyms, covering 98.6% of common surgical terms; specifically, establish the mapping relationship between "Sub total gastrectomy" and the ICD-9-CM code "43.89" and the CCC code "G3301003" in the knowledge base; establish "anchor vectors" for 5,800 standard surgical codes in the feature vector space, and each anchor represents the standard position of a standard code; for each anchor, collect 3 - 8 typical medical record samples, extract features and calculate the average vector as the representation of this code; specifically, for the code "LC1201" of "Laparoscopic cholecystectomy", collect 27 surgical records from 3 tertiary hospitals, and obtain the anchor vector [0.78, 0.65, 0.12, 0.03, 0.92, 0.45, 0.32, 0.19, 0.56, 0.28] with a dimension of 10 after extracting features.

[0088] According to the complexity and ambiguity of different surgical categories, a dynamic similarity threshold matrix T is set. For surgeries that can be clearly defined (such as fracture reduction), a higher threshold is set (T = 0.85); for surgeries with large variations in descriptions (such as various repair surgeries), a lower threshold is set (T = 0.65); specifically, the threshold matrix for digestive system surgeries shows that the mapping threshold for "gastrectomy" surgeries is 0.78, while the mapping threshold for "intestinal repair" surgeries is 0.67.

[0089] Construct a multi-level mapping path of "surgical description → feature vector → anchor vector → standard code". For 43 common variants of surgical descriptions, a synonym mapping table is established to achieve the normalization of different expressions; specifically, 13 expressions such as "partial hepatectomy", "hepatic segmentectomy", and "wedge resection of the liver" are uniformly mapped to the standard expression of "partial hepatectomy" to improve the mapping accuracy. For fuzzy regions in the feature space (where multiple coding anchors are close in distance), a weighted decision tree is constructed for secondary judgment; 364 clinical discriminative features are used as decision tree nodes, and the optimal branching path is determined based on historical data; specifically, for surgeries falling into the fuzzy region between "subtotal gastrectomy" and "total gastrectomy", the system uses 5 key features such as "whether to retain the cardia" and "percentage of resection range" for decision tree judgment.

[0090] S5. Perform entity recognition on the fused feature vector to obtain the entity recognition result; in this embodiment, a hybrid model combining BiLSTM-CRF (Bidirectional Long Short-Term Memory - Conditional Random Field) and attention mechanism is used; the model is trained on 12,800 annotated medical records, and the F1 score reaches 0.917; specifically, for the surgical description of "The patient underwent laparotomy + partial resection of the gastric antrum + Billroth II gastrojejunostomy", three surgical entities of "laparotomy", "partial resection of the gastric antrum", and "Billroth II gastrojejunostomy" are recognized.

[0091] Use a customized medical entity recognition model to accurately identify anatomical location information; construct a location relationship map containing 4,200 anatomical parts to support the derivation of the upper and lower levels of the parts; specifically, "right upper lung lobe" is identified as the precise anatomical location from "wedge resection of the right upper lung lobe", and a part-whole relationship is established with "lung".

[0092] Perform operation method recognition based on the surgical verb ontology library, which contains 486 standard surgical verbs; use dependency syntax analysis to extract the relationship between the verb and the receptor, and construct an operation-object pair; specifically, the operation method is identified as "percutaneous biopsy" and the operation object is "renal parenchyma" from "percutaneous biopsy of the renal parenchyma".

[0093] Adopt a rule and statistics hybrid model to identify surgical path information; establish a dictionary containing 89 common interventional paths, covering major surgical types such as open surgery, minimally invasive surgery, and endoscopy; specifically, identify "transcatheter" as an interventional path from "transcatheter aortic valve replacement" and classify it into the "interventional treatment" category.

[0094] Identify key modifiers that affect coding, such as "bilateral", "multi-site", "repeated", etc.; construct a modifier influence matrix to quantify the influence degree of various modifiers on the final coding; specifically, identify the "bilateral" modifier in "bilateral inguinal hernia repair" and automatically adjust it to a surgery that requires two independent codings during the coding process.

[0095] S6. According to the entity recognition result and the mapping relationship with the surgical coding, generate the surgery and corresponding operation coding. In this embodiment, convert the recognized surgical entity into a feature vector, and find the nearest k anchor vectors (usually k = 5) in the feature space; use the weighted K-nearest neighbor algorithm to calculate the probability score of each candidate coding according to the distance; specifically, for the recognized "laparoscopic cholecystectomy" entity, after the system generates the feature vector, find the nearest 5 anchor codings: LC1201 (distance 0.17), LC1202 (distance 0.22), LC1203 (distance 0.38), LC1205 (distance 0.45), KC1201 (distance 0.57), and initially select LC1201 as the best match. The core problem solved in this stage is the semantic mapping uncertainty problem, that is, how to accurately map the unstructured surgical entity to the structured coding system. This application is based on the Vector Space Model and the nearest neighbor retrieval theory, and embeds the surgical entity and the standard coding into the unified high-dimensional feature space at the same time. In this space, concepts with similar semantics are also close to each other geometrically, thus transforming the semantic matching problem into a quantifiable distance calculation problem.

[0096] Decompose the preliminary coding according to the three-level tree index structure, and extract the category code, system code, and method code; for the uncertain part in the preliminary coding, use Bayesian network reasoning for supplementation; specifically, decompose LC1201 into L (laparoscopic category), C (digestive system), 12 (gallbladder), 01 (resection) to form a complete hierarchical coding. The core problem solved in this stage is the multi-level semantic integration problem, that is, how to decompose the overall coding into components that conform to the medical ontology hierarchy. This application is based on Hierarchical Representation Learning and the tree structure traversal algorithm, splits the coding into components reflecting different semantic levels, and verifies its effectiveness in the medical ontology.

[0097] Use the binary relation matrix R to check the compatibility between multiple surgical codes; for incompatible code combinations, use a rule-based priority strategy to resolve conflicts; specifically, when a patient undergoes both "cholecystectomy" and "choledocholithotomy", the system detects that the compatibility value of these two surgeries in the matrix R is 0.9 (highly compatible), and both codes are retained; while when detecting the combination of "open cholecystectomy" and "laparoscopic cholecystectomy", its compatibility value in the matrix R is 0.1 (almost mutually exclusive), and the system retains the more complex "open cholecystectomy" code through clinical rules. The core problem solved in this stage is the multi-code coordination consistency problem, that is, how to ensure that multiple surgical codes meet the compatibility constraints of medical insurance policies. This application is based on the Constraint Satisfaction Theory and matrix algebra, represents medical insurance rules as a binary relation matrix, and verifies the compliance of code combinations through matrix operations.

[0098] Specifically, in the binary relation matrix R, R(i, j) represents the compatibility degree between code i and code j: R(i, j) = 1: completely compatible; R(i, j) = 0: completely mutually exclusive; R(i, j) ∈ (0, 1): partially compatible, the larger the value, the higher the compatibility degree. For the code set C = {c1, c2,....., c n}, its overall compatibility is calculated through matrix multiplication: When Compatibility(C) = 0, it means that there are mutually exclusive codes in the combination. Based on the graph coloring theory, the mutually exclusive code problem is transformed into a graph coloring problem, and the greedy algorithm is applied to solve the conflict: MaxIndependent Set(Graph(C, E = {mutually exclusive relationship})), where the independent set represents the largest code subset that can be retained simultaneously.

[0099] Based on the positional relationship in the eigenvector space, assign the final code to the verified hierarchical code combination; use the medical insurance payment standard database to perform the final verification of code compliance; specifically, for the combination of "laparoscopic cholecystectomy + common bile duct exploration and lithotomy", the system finally assigns the codes LC1201 + LC0702 and automatically adds the necessary code association markers.

[0100] S7. Generate a medical insurance settlement list according to the surgery and the corresponding operation codes. Specifically, build a code-settlement item conversion engine to convert surgical codes into medical insurance settlement items; apply a regional policy adapter to adjust settlement items according to different regional medical insurance policies; implement hospital-level difference processing to adjust the scope of settlement items according to the hospital level and qualifications. Build a settlement list generator to organize settlement items in a standard format; apply expense classification and summary to summarize expenses by categories such as drug expenses, surgical expenses, and material expenses; implement the division of out-of-pocket and medical insurance payments to clearly divide the patient's out-of-pocket part and the medical insurance payment part.

Claims

1. A method for generating a settlement list based on medical data, characterized in that, Including: S1. Obtain structured medical data and unstructured medical data; S2. Preprocess the structured medical data to generate normalized data; Perform text standardization processing on the unstructured medical data to generate standardized text; S3. Integrate the normalized data and the standardized text to obtain a fused feature vector; construct a feature vector space based on the fused feature vector; S4. Establish a mapping relationship of surgical coding using the medical ontology knowledge base according to the feature vector space; S5. Perform entity recognition on the fused feature vector to obtain an entity recognition result; S6. Generate a surgery and corresponding operation code according to the entity recognition result and the mapping relationship with the surgical coding; S7. Generate a medical insurance settlement list according to the surgery and corresponding operation code; Constructing a feature vector space includes: converting the surgical operation type restrictions in the medical insurance settlement rules into a binary relation matrix R, where R(i,j)=1 indicates that operation i is compatible with operation j, and R(i,j)=0 indicates that operation i is mutually exclusive with operation j; constructing a directed acyclic graph G=(V,E) according to the hierarchical relationship of the surgical operation types, where V is the set of operation types and E is the set of hierarchical constraint edges; constructing a three-level tree index structure according to the binary relation matrix R and the directed acyclic graph G=(V,E), where the first level is divided by surgical system, the second level is divided by intervention method, and the third level is divided by operation method; assign multi-dimensional space coordinates to the nodes in the three-level tree index structure to form an index structure containing spatial position information; map the semantic similarity matrix to the index structure containing spatial position information to obtain a feature vector space; Forming an index structure containing spatial position information includes: setting the number of dimensions n of the feature space; constructing a standard n-dimensional Hilbert space filling curve H; using the binary relation matrix R to deform and adjust the n-dimensional Hilbert space filling curve H to obtain a curve H'; where, the curve segments corresponding to the operations marked as compatible in the matrix R are adjacent, and the curve segments corresponding to the operations marked as mutually exclusive are separated; according to the hierarchical relationship of the three-level tree index structure, assign spatial regions to each node along the curve H'; Assigning spatial regions to each node along the curve H' includes: dividing the main trajectory of the curve H' into multiple non-overlapping convex spatial region blocks according to the classification of the surgical system, and each region block is assigned to a first-level node; within the spatial region block corresponding to each first-level node, further divide it into multiple continuous region segments along the sub-trajectory of the curve H' according to the classification of the intervention method, and each region segment is assigned to a second-level node; within the region segment corresponding to each second-level node, assign spatial coordinate points to each third-level node along the end trajectory of the curve H'.

2. The method for generating a settlement list based on medical data according to claim 1, wherein: The structured medical data includes medical service prices, medication and consumable usage; The unstructured medical data includes patient admission and discharge records and surgical operation records.

3. The method for generating a settlement list based on medical data according to claim 2, wherein: S3. Obtain the fused feature vector, including: Extract the quantitative features of the normalized data, where the quantitative features include medical service price features, drug usage features, and consumable consumption features; Extract the text features of the standardized text, where the text features include semantic features and surgical term features; Align the quantitative features and text features to generate an initial fused feature; Use the attention mechanism to perform weighted fusion on the initial fused feature to obtain the fused feature vector.

4. The method for generating a settlement list based on medical data according to claim 3, wherein: Construct a feature vector space according to the fused feature vector, including: Use a knowledge-enhanced hierarchical clustering algorithm to perform clustering analysis on the fused feature vector to obtain the feature vector category distribution; the knowledge-enhanced hierarchical clustering algorithm uses the synonymy relationship, hyponymy relationship, and co-occurrence relationship of surgical terms as clustering constraint conditions; Based on the feature vector category distribution, calculate the semantic similarity between different surgical feature vectors and construct a semantic similarity matrix; Convert the medical insurance settlement rules into spatial constraint conditions, where the medical insurance settlement rules include surgical operation type restrictions; According to the semantic similarity matrix and spatial constraint conditions, perform feature space partitioning through multi-dimensional indexing to construct a feature vector space.

5. The method for generating a settlement list based on medical data according to claim 4, wherein: Calculate the semantic similarity between different surgical feature vectors using the following formula: Sim(V1,V2) = α × Sim cos (V1,V2) + β × Sim path (T1,T2) + γ × Sim lca (T1,T2) where V1 and V2 are two surgical feature vectors; T1 and T2 are the surgical terms corresponding to V1 and V2; Sim cos (V1, V2) is the cosine similarity of the feature vectors; where ShortestPath(T1,T2) is the number of edges of the shortest path between terms T1 and T2 in the medical knowledge graph, and w is the path distance weight factor; where lca(T1,T2) is the nearest common ancestor node of terms T1 and T2 in the medical knowledge graph, and depth(T1) and depth(T2) respectively represent the distances from nodes T1 and T2 to the root node of the knowledge graph; α, β, and γ are weight coefficients.

6. The method for generating a settlement list based on medical data according to claim 1, wherein: The surgical system represents the surgical area division based on human anatomy classification, including the circulatory system, digestive system, nervous system, skeletal muscle system, and urogenital system; The intervention methods include open surgery and endoscopic surgery; The operation methods include resection, suture, repair, reconstruction, replacement, implantation, ablation, and decompression.

7. A system for generating a settlement list based on medical data, wherein it includes: At least one processing unit; used to execute instructions to implement the method for generating a settlement list based on medical data according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Medical record diagnosis and operation ICD coding method based on mapping knowledge domain

    CN119230090A