Clinical data processing method and device based on large language model and storage medium

Through a method based on a large language model, high-precision structured conversion of spinal disease consultation texts was achieved, solving the problem of insufficient utilization of spinal disease consultation data in existing technologies, improving the recognition rate of anatomical terms and consultation efficiency, and meeting the standards for direct clinical use.

CN120674093APending Publication Date: 2025-09-19PEKING UNION MEDICAL COLLEGE HOSPITAL +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510741554.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the field of spinal diseases, existing technologies make it difficult to effectively utilize unstructured medical consultation data, the accuracy of professional medical terminology recognition is insufficient, and clinical decisions lack structured data support, resulting in low data entry efficiency and accuracy, high error rate in symptom description conversion, and frequent errors in the timing logic of treatment plan records.

Method used

A large language model-based approach, including BERT preprocessing, Transformer architecture feature extraction, grammatical constraint decoding, term importance weight calculation, and cross-field consistency checking, is used to generate structured documents that meet preset standards.

Benefits of technology

The accuracy of structured conversion of spinal disease consultation texts has been improved, the accuracy of spinal anatomical term recognition has increased by 47%, the accuracy of symptom duration extraction has increased to 89%, the error rate of associating treatment plans with anatomical positions has been reduced to 6%, the consultation efficiency has increased by 62%, and the medical record completeness score has increased from 2.8/5 to 4.5/5.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120674093A_ABST
    Figure CN120674093A_ABST
Patent Text Reader

Abstract

The invention relates to clinical data processing based on a large language model, in particular to a structural processing method for clinical text data of spinal diseases, and belongs to the cross field of artificial intelligence and medical science. The method comprises the following steps: performing preprocessing and feature extraction on the clinical text data of the spinal diseases based on the large language model, and then performing grammar constraint decoding; and calculating a duplication probability in combination with term importance weights and spinal disease clinical correlation, and finally performing consistency check to generate structured target data. According to the method and the device provided by the invention, the term recognition accuracy and inquiry efficiency of spinal anatomy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the intersection of artificial intelligence and medicine, and in particular relates to a clinical data processing method, device and storage medium based on a large language model. Background Art

[0002] Currently, the clinical diagnosis and treatment of spinal diseases faces three major technical bottlenecks: the difficulty in effectively utilizing unstructured medical consultation data, the insufficient accuracy in recognizing professional medical terminology, and the lack of structured data support for clinical decision-making. Traditional electronic medical record systems rely on manual data entry, which is inefficient and inaccurate. For example, medical staff take an average of 15 minutes to complete the entry of a single medical record, of which 60% is spent on processing the structured conversion of symptom descriptions. Approximately 18% of clinical entities are entered incorrectly due to differences in expression, and 30% of treatment plan records contain temporal logic errors. Existing natural language processing technology performs poorly in the field of spinal diseases. Test data shows that the F1 value of the general NER model in spinal anatomical term recognition is only 0.63, the accuracy of symptom duration extraction is less than 55%, and the error rate in associating treatment plans with anatomical locations is 22%.

[0003] Specific challenges in clinical practice include: spinal disease symptom descriptions have significant spatial characteristics (such as "radiating pain in the L4 nerve root area"), making it difficult for existing models to accurately parse the correspondence between anatomical location and symptoms; treatment follow-up records contain complex time series information, making traditional methods unable to effectively model the evolution of symptoms; and medical terminology has a large number of variant expressions (such as "ACDF" and "anterior cervical discectomy and fusion"), making entity normalization difficult. These technical shortcomings of existing technologies make it difficult to meet the accuracy requirements of clinical diagnosis and treatment. Summary of the Invention

[0004] In view of the above-mentioned deficiencies in the existing technology, the purpose of the invention is to provide a method and device for structuring clinical text data of spinal diseases based on a large language model, so as to achieve high-precision structured conversion of spinal disease consultation texts and improve the accuracy of spinal anatomical term recognition.

[0005] The first aspect of the present invention provides a method for structured processing of spinal disease clinical text data based on a large language model, comprising:

[0006] Preprocessing of spinal disease clinical text data based on the BERT large language model;

[0007] The Transformer architecture based on the large language model performs feature extraction on the preprocessed data;

[0008] Perform grammatical constraint decoding on clinical text data of spinal diseases after feature extraction, and generate a legal token set for each decoding step based on a dynamic vocabulary;

[0009] The replication probability was calculated based on the term importance weights and clinical relevance to spinal diseases;

[0010] Performing cross-field consistency check on the spinal disease clinical text data;

[0011] Transform consistency-checked spinal disease clinical text data into target formats to generate documents that meet pre-set standards.

[0012] A second aspect of the present invention provides a clinical data processing device based on a large language model, comprising:

[0013] Data preprocessing layer, feature encoding layer, grammatical constraint decoding layer, pointer generation layer, multimodal verification layer and output conversion layer;

[0014] The data preprocessing layer is configured to preprocess the spinal disease clinical text data based on the BERT large language model;

[0015] The feature encoding layer is configured to extract features from the preprocessed data based on the Transformer architecture of the large language model;

[0016] The grammatical constraint decoding layer is configured to perform grammatical constraint decoding on the clinical text data of spinal diseases after feature extraction, and generate a legal token set for each decoding step according to a dynamic vocabulary;

[0017] The pointer generation layer is configured to calculate replication probability based on term importance weights and clinical relevance of spinal diseases;

[0018] The multimodal verification layer is configured to perform cross-field consistency check on the spinal disease clinical text data;

[0019] The output conversion layer is configured to convert the spinal disease clinical text data that has passed the consistency check into a target format to generate a document that meets preset standards.

[0020] A third aspect of the present invention provides a device including a memory, a processor, and a user interface;

[0021] The memory is used to store computer programs;

[0022] The user interface is used to interact with the user;

[0023] The processor is used to read the computer program in the memory, and when the processor executes the computer program, it implements the above-mentioned clinical data processing method based on the large language model.

[0024] In a fourth aspect of the present invention, a processor-readable storage medium is proposed, wherein the processor-readable storage medium stores a computer program, and when the processor executes the computer program, the above-mentioned clinical data processing method based on the large language model is implemented.

[0025] The beneficial effects of the present invention are as follows:

[0026] The method and device described in the present invention achieve high-precision structured conversion of spinal disease consultation texts, improving the accuracy of spinal anatomical term recognition and consultation efficiency. In real-world applications in a tertiary-level Class A hospital, the F1 value of spinal anatomical term recognition reached 0.92, a 47% improvement over the baseline model; the accuracy of symptom duration extraction increased to 89%; and the error rate of associating treatment plans with anatomical locations decreased to 6%. The pass rate of structured medical records increased from 68% of traditional methods to 94%, meeting the standards for direct clinical use. Consultation efficiency has been improved, with the processing time for a single consultation record shortened from an average of 15 minutes for manual entry to 45 seconds; the medical record completeness score increased from 2.8 / 5 to 4.5 / 5; and the time series accuracy of follow-up records increased by 62%. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings are only for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. Throughout the drawings, the same reference numerals represent the same components. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings.

[0028] Figure 1 A hardware structure block diagram of a computing device for implementing the method of the present invention;

[0029] Figure 2 Schematic diagram of the flow of a clinical data processing method based on a large language model according to an embodiment of the present invention;

[0030] Figure 3 This is a flow chart of a medical terminology standardization method according to an embodiment of the present invention;

[0031] Figure 4 Schematic diagram of the flow of a comprehensive similarity calculation method according to an embodiment of the present invention;

[0032] Figure 5 A schematic diagram of a method for calculating replication probability according to an embodiment of the present invention;

[0033] Figure 6 Schematic diagram of a clinical data processing device based on a large language model according to an embodiment of the present invention;

[0034] Figure 7 Schematic diagram of another clinical data processing device based on a large language model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0035] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are part of the embodiments of the present invention, rather than all of the embodiments. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work should fall within the scope of protection of the present invention.

[0036] Furthermore, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts disclosed in the present invention.

[0037] In the description of the present invention, it should be noted that, unless otherwise expressly specified and limited, the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second" and "third" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance. The terms "installed", "connected" and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be a communication between the internal parts of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0038] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of methods and systems consistent with certain aspects of the present invention, as detailed in the appended claims.

[0039] In order to better explain the present invention, some terms are first explained.

[0040] 1. BERT: Bidirectional Encoder Representations from Transformers, a pre-trained language model based on the Transformer architecture;

[0041] 2. Transformer: Transformer is a neural network architecture based on the self-attention mechanism. By introducing the self-attention mechanism, it allows the model to dynamically focus on other elements in the sequence while processing each element in the sequence, thereby more efficiently capturing long-range dependencies.

[0042] 3. In this invention, the terms "terms" and "terms" are medical terms. Words and terms have the same meaning and are medical terms.

[0043] The embodiments of the present invention propose a clinical data processing method, device and storage medium based on a large language model, which solves the problems of low recognition accuracy and low efficiency when applying clinical data in existing methods.

[0044] According to this embodiment, an embodiment of a clinical data processing method based on a large language model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0045] The method embodiment provided in this embodiment can be executed in a mobile terminal, a computer terminal, a server or a similar computing device. Figure 1 The hardware structure block diagram of a computing device for a clinical data processing method based on a large language model is shown. Figure 1 As shown, the computing device may include one or more processors (the processor may include but is not limited to a microprocessor MCU or a programmable logic device FPGA, etc.), a memory for storing data, a transmission device for communication functions, and an input / output interface. The memory, transmission device, and input / output interface are connected to the processor via a bus. In addition, it may also include: a display, a keyboard, and a cursor control device connected to the input / output interface. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0046] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computing device. As described in the embodiments of the present disclosure, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0047] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the clinical data processing based on the large language model in the embodiment of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned clinical data processing method based on the large language model. The memory may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computing device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0048] The transmission device is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by a communications provider of the computing device. In one embodiment, the transmission device includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0049] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computing device.

[0050] It should be noted that, in some optional embodiments, the above Figure 1 The computing device shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the computing devices described above.

[0051] In the above operating environment, according to the first aspect of this embodiment, a clinical data processing method based on a large language model is provided. Figure 2 A schematic diagram showing the process of the method is shown in FIG. Figure 2 As shown, the method includes:

[0052] S201. Preprocessing of spinal disease clinical text data based on the BERT large language model;

[0053] In the present invention, this step may also be referred to as a text preprocessing stage, which performs medical terminology standardization.

[0054] S202, extracting features from the preprocessed data using the Transformer architecture based on the large language model;

[0055] In the present invention, this step can also be called the feature extraction stage, which integrates linguistic features and clinical features.

[0056] S203, performing grammatical constraint decoding on the clinical text data of spinal diseases after feature extraction, and generating a legal token set for each decoding step according to a dynamic vocabulary;

[0057] In the present invention, this step may also be referred to as the grammatical constraint decoding stage, which implements dynamic vocabulary control and dynamically generates a legal token set for each decoding step.

[0058] S204. Calculate the replication probability based on the term importance weight and the clinical relevance of spinal diseases;

[0059] In the present invention, this step may also be referred to as a pointer generation phase.

[0060] S205, performing a cross-field consistency check on the spinal disease clinical text data;

[0061] In the present invention, this step may also be referred to as a multimodal verification phase, which performs cross-field consistency checks, including value range verification.

[0062] S206. Convert the spinal disease clinical text data that has undergone consistency check into a target format to generate a document that meets preset standards.

[0063] In the present invention, this step can also be referred to as the output conversion stage, where an XML document conforming to the CDA standard is generated and converted into a target format through XSLT, thereby achieving structured processing of spinal disease clinical text data.

[0064] The clinical data processing method based on a large language model of the present invention processes clinical data through a large language model and then verifies its consistency to generate structured clinical data, thereby improving recognition accuracy and consultation efficiency when using these structured clinical data. In particular, in the clinical field of spinal diseases, after the method of the present invention structures the clinical text of spinal diseases based on a large language model, in real-world applications in tertiary-level A hospitals, the F1 value of spinal anatomical term recognition reached 0.92, an improvement of 47% over the baseline model; the accuracy of symptom duration extraction increased to 89%; and the error rate of associating treatment plans with anatomical locations decreased to 6%. The pass rate of structured medical records increased from 68% of traditional methods to 94%, meeting the standards for direct clinical use. The consultation efficiency was improved, and the processing time for a single consultation record was shortened from an average of 15 minutes for manual entry to 45 seconds; the medical record completeness score increased from 2.8 / 5 to 4.5 / 5; and the time series accuracy of follow-up records increased by 62%. Significant technical effects were achieved.

[0065] Optionally, in the above S201, the spinal disease clinical text data is preprocessed based on the BERT large language model to achieve medical terminology standardization and clinical entity annotation. Among them, the method of medical terminology standardization is as follows Figure 3 Shown, including:

[0066] S301, screening candidate terms;

[0067] In this step, a candidate set is obtained through double filtering, which includes: semantic-level term matching filtering and literal-level string matching.

[0068] As an optional example, the double filtering candidate set C(u) is implemented by the following formula: C(u) = {v∈T std |ClinicalED(u,v)≤θ ed and BERTScore(u,v)≥θ sem};

[0069] ClinicalED(u, v) represents the clinical edit distance, which indicates the similarity between medical terms u and v, both of which are medical terms. ed represents the character difference threshold, such as θ ed =2 means that 2 characters are allowed to differ. BERTScore(u,v) represents the semantic similarity between medical terms u and v calculated based on the BERT model. θ sem represents the semantic similarity threshold, such as θ sem =0.7 means the semantic similarity must be greater than or equal to 0.7. stdStandard Edit Distance (SEDD) is a numerical value that measures the minimum number of edit operations between two strings. The smaller the value, the more similar the two strings are. In this invention, it serves as a rapid screening tool to quickly and preliminarily determine which candidate medical terms are similar to the standard term, thereby improving the efficiency and accuracy of subsequent standardization processing.

[0070] Optionally, the editing operation includes one or a combination of the following:

[0071] Insertion of characters, deletion of characters, replacement of characters.

[0072] For example, there are two words: string a: "intervertebral disc"; string b: "intervertebral disc". Converting from "intervertebral disc" to "intervertebral disc" only requires one operation, that is, replacing "disc" with "pelvis", so: T std =1.

[0073] If the two words are more different, for example: string a: "intervertebral disc herniation"; string b: "intervertebral disc herniation". The conversion requires: "basin" → "basin" (replacement), "protrusion" → "disc" (replacement), a total of 2 operations: T std =2.

[0074] S302, performing anatomical space consistency check on the candidate terms;

[0075] Preferably, the anatomical space consistency check can be performed according to the following formula:

[0076] Valid(v|u)=I(AnatomyDist(u.loc, v.loc)≤δ);

[0077] Here, u and v are both medical terms. u.loc represents the anatomical location corresponding to the medical term u, and v.loc represents the anatomical location corresponding to the medical term v. For example, if the medical term u = "L4 / 5 disc herniation," then u.loc = "L4 / 5 disc" is the specific location. If the medical term v = "L5 nerve root compression," then v.loc = "L5 nerve root" is the specific location. Here, both terms u and v clearly correspond to specific anatomical locations (such as spinal segment number, nerve root location, etc.).

[0078] Valid(v|u) represents the result of the anatomical space consistency check. The result is 1 if the check passes, and 0 if the check fails.

[0079] AnatomyDist(u.loc, v.loc) represents the spatial distance or degree of difference in anatomical positions between the anatomical positions corresponding to terms u and v. Taking the spinal anatomy as an example, if the two terms describe the same anatomical region (such as "L4 intervertebral disc" and "L4 intervertebral disc"), then: AnatomyDist(u.loc, v.loc) = 0. If the terms describe adjacent segments or very close positions (such as "L4 intervertebral disc" and "L5 intervertebral disc"), then: AnatomyDist(u.loc, v.loc) = the first difference value (the smaller the position, the smaller the value), such as 1 or other defined units; if the term positions are far apart (such as cervical C2 intervertebral disc and lumbar L5 intervertebral disc): AnatomyDist(u.loc, v.loc) = the second difference value (the farther the position is, the larger the value). In this example, the first difference value is smaller than the second difference value. Therefore, during the medical terminology standardization process, this function can help the system determine whether the two terms are related or close in spatial anatomical structure.

[0080] I() represents an indicator function, which returns 1 if the condition is true, otherwise it returns 0;

[0081] δ is the position deviation threshold of the spinal segment. For example, δ=3 means that the position deviation of 3 spinal segments is allowed.

[0082] S303, calculating the comprehensive similarity of the candidate terms and ranking them;

[0083] Optionally, the sorting can be done from high to low similarity or from low to high similarity.

[0084] Optionally, the calculation of comprehensive similarity is as follows Figure 4 As shown, including:

[0085] S401, screening candidate terms using the first formula;

[0086] In this step, the first formula is used to quickly screen candidate terms to achieve literal string matching. The completed processing includes:

[0087] Spelling correction (e.g., "intervertebral pelvis" → "intervertebral disc"); word order change (e.g., "pain radiating to lower limbs" → "radiating pain in lower limbs"); and affix difference (e.g., "postoperative day 3" → "postoperative day 3"). In this step, literal string matching is used for high computational efficiency.

[0088] Among them, the first formula is:

[0089] ClinicalED(a,b)=min(ED(a,b),1+min a′∈Syn(a) ED(a′, b));

[0090] ClinicalED(a, b) is the clinical edit distance, which represents the similarity between medical terms a and b; a and b are both medical terms; ED(a, b) is the standard edit distance, which represents the minimum number of operations required to convert string a to b; Syn(a) represents the set of synonyms of medical term a; min a′∈Syn(a) ED(a′, b) represents the value with the smallest edit distance to b among all synonyms of a; a′ represents the synonym of a;

[0091] S402: directly use character string mapping for terms with similarity greater than a preset similarity in the results of the first formula screening, and perform deep semantic matching on the remaining terms using the second formula;

[0092] S403, performing anatomical space constraint check on the matching result of the second formula;

[0093] In this step, semantic-level term matching is achieved, and the processing mainly includes:

[0094] Synonym mapping (e.g., “ACDF” → “anterior cervical discectomy and fusion”);

[0095] abbreviation expansion (e.g., “LBP” → “lower back pain”);

[0096] Dialect standardization (e.g. “lumbar disc herniation” → “lumbar disc herniation”).

[0097] Among them, the second formula is:

[0098]

[0099] Among them, TermSim(u, v) represents the similarity between medical terms u and v; w1 and w2 are weight parameters; BERTScore(u, v) represents the semantic similarity of medical terms calculated based on the BERT model; AnatomyDist(u, v) represents the spatial distance of anatomical positions; SemDiff(u, v) represents the semantic difference between terms; k is the adjustment factor; u and v are both medical terms.

[0100] S304: Mark two medical terms whose comprehensive similarity is higher than a preset threshold as the same word.

[0101] Optionally, in this step, the preset threshold is set using an adaptive threshold strategy. Specifically, the preset threshold can be determined according to the following formula:

[0102] θ dynamic is the preset threshold, which is determined by the following formula:

[0103] θ dynamic =0.5+0.2·log(1+Freq(u));

[0104] Where Freq(u) represents the usage frequency of medical term u.

[0105] Optionally, the purpose of clinical entity annotation is to identify medically significant entities from clinical texts, such as diseases, symptoms, drugs, tests, etc. The steps of clinical entity annotation may include: data preparation, which involves collecting clinical text data on spinal diseases (such as electronic medical records and medical reports), performing word segmentation, and removing noise; annotation standard development, which involves developing annotation guidelines based on task requirements to clarify which text segments should be annotated as clinical entities and the categories of the entities; and annotation execution, which involves professional medical personnel or trained annotators performing entity annotation on clinical texts on spinal diseases according to the annotation guidelines.

[0106] Optionally, in S202 above, the extracting features from the preprocessed data using the Transformer architecture based on the large language model includes:

[0107] Modeling symptom descriptions through three attention mechanisms;

[0108] Optimize symptom descriptions through clinical knowledge;

[0109] Among them, the three attention mechanisms are:

[0110] The local window attention mechanism is used to process the detailed description of symptoms and handle the local dependencies in the detailed description of symptoms, such as the association between time and intensity in "lasting for 3 days and worsening at night";

[0111] A global sparse attention mechanism is used to capture key clinical entities, such as identifying core diagnoses such as "L4 / 5 intervertebral disc herniation" from long text.

[0112] The cross-modal attention mechanism is used to associate imaging test results with symptom descriptions, such as the imaging report describing "T2 high signal" in MRI with symptom descriptions such as "lower limb numbness".

[0113] The following table compares the three attention mechanisms mentioned above:

[0114] Table 1: Comparison of three attention mechanisms

[0115]

[0116] Among them, the query matrix Q of the local window attention mechanism uses text features, and the query matrix Q of the cross-modal attention mechanism uses image features.

[0117] Optionally, the key-value projection parameters of the three attention mechanisms are trained independently.

[0118] Optionally, the attention mechanism is implemented by the following formula:

[0119]

[0120] Among them, A hier is the hierarchical attention weight matrix; Q represents the query matrix, which represents the information that needs to be paid attention to at present; K local The key matrix representing the local context; K global A key matrix representing the global context; represents the matrix concatenation operation, combining local and global features; d represents the square root of the feature dimension; Softmax() is a normalization function used to convert the attention score into a probability distribution; T1 represents the transpose operation of the matrix.

[0121] Optionally, refine the symptom description according to the following formula:

[0122] H enh =H base +LayerNorm(W k ReLU(W c ·E kb ));

[0123] Among them, H enh represents the feature representation matrix after knowledge enhancement; H base Represents the feature matrix output by the basic encoder; LayerNorm represents the layer normalization operation; W k and W c are all trainable parameter matrices; ReLU() represents the rectified linear unit activation function, which is used to introduce nonlinear transformation; E kb Represents the clinical knowledge graph embedding matrix, which is used to contain the domain knowledge of spinal diseases.

[0124] Optionally, in S203 of this embodiment, generating a legal token set for each decoding step according to the dynamic vocabulary includes:

[0125] S203-1. Use the third formula to perform initial decoding to obtain a basic structure. When a key field is detected, use the fourth formula to perform decoding.

[0126] S203-2. Take the intersection of the decoding result of the third formula and the decoding result of the fourth company as the legal token set.

[0127] In step S203-1 of the present invention, during the initial decoding phase, the third formula is used to ensure the basic structure. Then, when a key field (e.g., a diagnosis conclusion) is detected, the process switches to the refined control phase and decodes using the fourth formula. Finally, in step S203-2, the intersection of the initial decoding phase and the refined control phase is used as the legal token set.

[0128] Optionally, the third formula is:

[0129]

[0130] v is the spinal medicine vocabulary; t is the decoding step number; V t is the token set obtained by decoding in step t;

[0131] V is the complete vocabulary; CFGCheck() is the context-free grammar check function; S1:t-1 represents the token sequence generated by the previous t-1 step decoding; Represents the matrix concatenation operation; G spine Special grammar for spinal diseases;

[0132] Optionally, the fourth formula is:

[0133] L t =Lookahead(y1:t-1, G, 3);

[0134] L t is the set of tokens obtained by decoding in the t-th step; y1:t-1 represents the token sequence generated by decoding in the previous t-1 step; G represents the set of grammatical rules that define the output structure; 3 represents the number of look-ahead steps, which is used to look forward 3 steps for grammatical analysis; Lookahead() represents the look-ahead analysis function, which is used to predict legal subsequent tokens based on the generated sequence and grammatical rules.

[0135] Optionally, in the above S204, the replication probability is calculated based on the term importance weight and the clinical relevance of spinal disease, such as Figure 5 Shown, including:

[0136] S501: Determine a generation mode according to a fifth formula, where the generation mode includes generation or copying;

[0137] Optionally, the generation probability is calculated according to the fifth formula, and if the generation probability is greater than a preset generation probability threshold, the generation mode is generation; otherwise, the generation mode is copy;

[0138] The fifth formula is:

[0139]

[0140] ρ gen To generate probability; Sigmoid() represents the S-type activation function, which maps the value to the interval [0, 1] as the probability value; Represents the parameter vector for generating probability calculation; tanh represents the hyperbolic tangent activation function, which maps the value to the interval [-1, 1]; W g Represents the weight matrix for generating probability calculation; ht Represents the current decoding state vector; c t represents the context vector; t is the decoding step number.

[0141] S502: If the generation mode is generation, the vocabulary probability is calculated according to the sixth formula; if the generation mode is copy, the input term is selected according to the seventh formula;

[0142] Optionally, the sixth formula is:

[0143] P vocab (w t )=softmax(W vocab h t +b vocab );

[0144] Among them, P vocab (w t ) means that when decoding in step t, word w is selected from the vocabulary t The probability distribution of w t The word chosen for step t; this represents the probability that the output of the current decoding step is generated from the vocabulary. vocab is the parameter matrix learned during the training process, which is used to transform the hidden state h of the decoder t Mapped to the space of each word in the vocabulary to calculate the probability of each word being selected. t Represents the current decoding state vector, which contains all the information that the model has integrated in the previous historical information (text context, decoded words) in the current decoding step. It is usually obtained through Transformer or other recurrent neural network structures. vocab Represents a bias term, which is used to adjust the probability distribution of the model output; it is used to adjust the probability distribution of the model output, which is usually also obtained through training. Softmax() represents a normalization function.

[0145] In the present invention, the sixth formula is used to determine the probability of selecting a word from the entire vocabulary as the output when decoding a certain step. First, use the current hidden state h of the model t The original score of each word in the vocabulary is calculated and then converted into a probability value between 0 and 1 through the softmax() function.

[0146] Optionally, the seventh formula is:

[0147]

[0148] i and j indicate the term numbers; represents the importance weight of the i-th term; q represents the query vector, which represents the current decoding state; k irepresents the key vector representation of the i-th term; T represents the temperature parameter used to control the smoothness of the distribution; τ represents the set of all candidate terms; exp() represents the natural exponential function; ∑j∈τ represents the summation of all candidate terms for normalization.

[0149] S503: Perform probability normalization according to the eighth formula.

[0150] Optionally, the eighth formula is:

[0151]

[0152] P norm (w t ) means that after normalization, the word w is selected in step t t The final probability value, w t The word selected in step t; the normalized result ensures that the sum of the probabilities of all words is 1, which is a property that the probability distribution must satisfy; S(w t ) represents the word w t The raw score at the current step is usually generated by the previous module (for example, the model's predicted score for the current word); T is the temperature parameter, which is used to control the smoothness of the output probability distribution; when T is small (for example, <1), the output probability distribution tends to be steeper and more certain, and the model is more inclined to select the word with the highest score (that is, the probability distribution is sharper); when T is large (for example, >1), the output probability distribution becomes flatter, and the model generates more even probabilities, reducing the probability of a single word being selected with excessive certainty. e() represents the natural exponential function, which is used to ensure that any real number becomes a positive number; represents the sum of the indexed scores of all selected words, and j represents the word number. It is used to normalize the entire probability to ensure that the sum of the probabilities of all candidate words is 1.

[0153] The eighth formula normalizes the probability of selecting each candidate word. First, each candidate word's score is exponentially transformed to ensure that all probabilities are positive. Then, the sum of the exponentially transformed scores of all candidate words is used for normalization to obtain an appropriate probability distribution.

[0154] Steps S501, S502 and S503 of the present invention together constitute a complete probability generation chain of the present invention.

[0155] Optionally, in S205 of the present invention, performing a cross-field consistency check on the spinal disease clinical text data includes:

[0156] Grammatical verification, clinical logic verification and medical fact verification;

[0157] The syntax check is used to ensure that the output complies with predetermined specifications, the clinical logic is used to verify the rationality of the numerical values, and the medical fact check is used to compare with the latest clinical guidelines.

[0158] Optional validation scoring methods for syntax checking, clinical logic checking, and medical fact checking are:

[0159] First, perform clinical logic dimension verification and convert the verification results into a score value of 0 or 1;

[0160] Determine a comprehensive scoring value according to the scoring value;

[0161] The clinical logic dimension verification includes:

[0162] Clinical logic dimension verification is performed using the following formula:

[0163] RangeCheck(f1,f)=I(u f -2σ f ≤f1≤u f +2σ f );

[0164] RangeCheck(f1, f) indicates the result of range checking on the value f1 in field f; f1 indicates the value to be verified; f indicates the field identifier; u f represents the expected mean of field f; σ f Represents the standard deviation of field f; I() represents an indicator function, which returns 1 if the condition is true, otherwise it returns 0.

[0165] The clinical logic dimension verification in this step, for example, when scoring pain, if the pain u f =5,σ f =2, then [1, 9] is allowed, indicating that pain levels 1 to 9 are allowed.

[0166] For example, in a blood test, the u of blood calcium f =2.25mmol / L, σ f =0.25, the allowable laboratory test value range is [1.75, 2.75].

[0167] Determining a comprehensive score value according to the score value includes:

[0168]

[0169] ValidScore represents the comprehensive score value; represents the product of the five evaluation dimensions, k represents the evaluation dimension number; η k represents the weight of the kth dimension; f k(R) represents the scoring function of the kth dimension on the structured result R; exp(-λ·ConflictCount) represents the conflict penalty term; λ represents the conflict penalty coefficient; ConflictCount represents the number of conflicts detected; R represents the structured result generated by the system.

[0170] Optional, the five evaluation dimensions are: syntax verification, clinical logic, medical facts (knowledge graph verification), timing consistency and data integrity.

[0171] Optionally, the clinical data processing method based on a large language model of the present invention may further include a verification optimization phase, using a cache mechanism to accelerate grammar verification.

[0172] The cache mechanism is implemented through the following formula:

[0173]

[0174] Among them, CacheHitRate represents the cache hit rate; MissCount represents the number of queries that miss the cache; TotalQuery represents the total number of queries; e -βt represents the decay factor over time; β represents the decay coefficient; t' represents the system operation time.

[0175] The clinical data processing method based on a large language model of the present invention achieves high-precision structured conversion of spinal disease consultation texts, improving the recognition rate of spinal anatomical terms and consultation efficiency. In real-world applications in tertiary-level A-class hospitals, the F1 value of spinal anatomical term recognition reached 0.92, an improvement of 47% compared to the baseline model; the accuracy of symptom duration extraction increased to 89%; and the error rate of associating treatment plans with anatomical positions decreased to 6%. The pass rate of structured medical records increased from 68% of traditional methods to 94%, meeting the standards for direct clinical use. The consultation efficiency was improved, and the processing time for a single consultation record was shortened from an average of 15 minutes for manual entry to 45 seconds; the medical record completeness score increased from 2.8 / 5 to 4.5 / 5; and the time series accuracy of follow-up records increased by 62%.

[0176] Optionally, the method of the present invention also includes a security mechanism, specifically including: data transmission using AES-256 encryption, access control using the RBAC model, and operation logs using blockchain for notarization.

[0177] Optionally, the method of the present invention also includes an exception handling mechanism. When an unregistered term is detected, a three-level processing strategy is initiated: first, a clinical synonym database is queried, then a literal similarity matching is used, and finally, the query is submitted to the manual review queue. Incomplete sentences are handled using bidirectional context prediction. The prediction formula can be:

[0178] CompletionScore(c|p,s)=λ1·P left (c|p)+(1-λ1)P right (c|s);

[0179] CompletionScore(c|p,s) represents the score of the completion content c under the conditions of the previous content p and the next content s; c is the content to be completed; p is the previous context; s is the next context; λ1 represents the previous context prediction weight; P left (c|p) represents the probability of predicting the completion content c based on the previous content p; (1-λ1) represents the prediction weight of the subsequent content; P right (c|s) represents the probability of predicting the completion content c based on the subsequent content s.

[0180] Optionally, the method of the present invention further includes a dialect processing module to achieve regional adaptation. As an optional example, dialect processing is performed according to the following formula:

[0181] DialectAdapt(t1)=argmax s∈S Prob(s|r)·Sim(t1,s);

[0182] DialectAdapt(t1) represents the standardization result of dialect t1; t1 is the input dialect term; argmax represents the parameter value that maximizes the objective function; s represents the candidate standard term; S represents the set of standard terms; Prob(s|r) represents the probability of using the standard term s under the condition of region r; Sim(t1, s) represents the similarity between the dialect term t1 and the standard term s.

[0183] Optionally, the method of the present invention further includes a new disease subtype expansion interface:

[0184] ExtendAPI= <T new , R new , M map >

[0185] ExtendAPI means extended interface definition; new A collection of terms representing newly added disease subtypes; R new Represents the relationship set of newly added disease subtypes; M map Represents the old and new term mapping matrix.

[0186] Optionally, the method of the present invention further includes multi-language support, which is achieved through Unicode standardization and can be expressed as:

[0187] MultilangualTerm(t2)=Normalize(Transliterate(t2));

[0188] MultilangualTerm(t2) represents the multilingual normalization result of term t2; t2 is the input term; Transliterate() represents the transliteration function, which is used to convert non-Latin characters into Latin character representation; Normalize() represents the Unicode normalization function, which processes character variants and combining characters.

[0189] Optionally, the method of the present invention further includes mobile terminal optimization, using model quantization technology, which can be expressed as:

[0190]

[0191] Quantize(W, b) represents the result of b-bit quantization of weight W; W represents the model weight matrix; b represents the number of quantization bits; Round() represents the rounding function; μ represents the mean of weight W; σ represents the standard deviation of weight W; 2 b-1 Indicates the scaling factor for b-bit quantization.

[0192] Optionally, in this invention, a three-stage training strategy can be used for training large models: 1. Basic pre-training: using medical literature data to train language comprehension capabilities; 2. Domain adaptation training: fine-tuning model parameters on spinal disease text; 3. Task fine-tuning: using labeled medical consultation data to optimize end-to-end performance. The optimization objective function for training is a multi-task loss, which can be expressed as:

[0193] L=α·L gen +β·L ptr +γ·L valid ;

[0194] L represents the total loss function; α represents the loss weight of the generated task; L gen represents the generation task loss; β represents the pointer task loss weight; L ptr represents the pointer network task loss; γ represents the verification task loss weight; L valid represents the verification task loss.

[0195] The learning rate scheduling adopts the cosine annealing strategy:

[0196]

[0197] η n represents the learning rate of the nth step; η min represents the minimum learning rate; η max represents the maximum learning rate; n represents the current number of learning steps; N represents the total number of training steps; Angle representation showing the progress of training; represents the cosine annealing factor.

[0198] Optionally, the method of the present invention further comprises a clinical effectiveness assessment, wherein the clinical effectiveness is assessed using the following formula:

[0199]

[0200] ClinicalScore represents the total clinically effective score; Metric i Represents the score of the i-th evaluation indicator; i is the evaluation indicator number; τ i is the threshold of the i-th evaluation indicator; I() is the indicator function, which returns 1 if the condition is true, otherwise it returns 0.

[0201] Optionally, the method of the present invention further includes a model update method, which adopts a rolling upgrade strategy. The rolling update window can be expressed as:

[0202] UpdateWindow=Max(T avg +3σ, T min );

[0203] UpdateWindow represents the update time window; T avg represents the average request processing time; σ represents the standard deviation of the request sorting time; T min Indicates the minimum update window time.

[0204] Optionally, the method of the present invention further includes a term base update method to achieve incremental update. The incremental synchronization priority score can be expressed as:

[0205]

[0206] SyncDelta represents the incremental synchronization priority score; t2 represents the newly added term; T new Indicates a set of newly added terms; Priority(t2) indicates the business priority of term t2; Urgency(t2) indicates the urgency of term t2; Priority(t2)·Urgency(t2) indicates the comprehensive synchronization priority of term t2.

[0207] The clinical data processing method based on the large language model of the present invention mainly includes the following core steps:

[0208] (1) Term Similarity Calculation: BERTScore(u, v) in the formula TermSim(u, v) directly relies on the semantic encoding capabilities of the pre-trained large language model. This step uses the parameterized text representation technology of large models such as BERT to calculate the cosine similarity of terms in a high-dimensional semantic space, capturing deep semantic associations that cannot be quantified by traditional edit distance. For example, in spinal anatomy term matching, BERT's context-awareness can distinguish the semantic equivalence of "C5 nerve root compression" and "C5 nerve involvement."

[0209] (2) Feature encoding layer: The hybrid Transformer module inherits the core components of the large language model: multi-head self-attention mechanism: by calculating the attention weight, long-distance dependency modeling is achieved; position encoding: using learnable dynamic position embedding instead of the fixed sine encoding of the traditional Transformer. Residual connection and layer normalization: key technologies to maintain the stability of large model training. E in the knowledge enhancement module kb The matrix is ​​obtained by fine-tuning the medical knowledge graph, continuing the parameter expansion method of the large language model (LLM).

[0210] (3) Pointer generation: Generation probability ρ gen The calculation of h is the same probability mixing mechanism as the large model, and the ratio of vocabulary generation and term replication is controlled by Sigmoid() gate. t The hidden state representation is inherited from the decoder architecture of the large model, and the parameter initialization of the tanh activation function follows the Xavier strategy of LLM.

[0211] (4) Semantic understanding optimization: Clinical knowledge enhancement formula H enh It is a variant of the large model adapter technology, which realizes knowledge injection by superimposing domain-specific parameters on the pre-trained base.

[0212] The following is a specific example with reference to specific examples.

[0213] Example 1: Acute attack of lumbar disc herniation

[0214] The input text is: "A 45-year-old male developed sudden severe low back pain with radiating pain in the left lower limb for 3 days after lifting heavy objects. The VAS score was 8 points. The straight leg raising test was positive at 30 degrees on the left and negative on the right. He had a history of L4 / 5 intervertebral disc herniation."

[0215] Using our large language model-based clinical data processing method, the time expression "3 days" is normalized to ISO 8601 format; the symptom description is split into two atomic entities: "low back pain" and "radiating pain in the left lower limb"; the anatomical location "L4 / 5" is associated with the symptoms; and a time series model is constructed based on past medical history and current symptoms. The output structure includes four modules: chief complaint, present medical history, physical examination, and preliminary diagnosis. The system automatically generates a differential diagnosis suggestion: spinal stenosis and cauda equina syndrome need to be ruled out.

[0216] The structured output of the pain assessment part is:

[0217]

[0218] Among them: PainAssessment represents pain assessment structured data; Location represents the pain location; Radiation represents the pain radiation area; Severity represents the severity of pain; Scale represents the pain rating scale used; Score represents the pain score value; AggravatingFactors represents the factors that aggravate pain; TemporalPattern represents the temporal pattern of pain.

[0219] Example 2: Postoperative follow-up of cervical spondylosis

[0220] Input text: "Six weeks after ACDF, the patient showed significant relief of neck pain, with the VAS dropping from 7 points before surgery to 2 points. However, a new decrease in right hand grip strength occurred, and the JOA score improved by 15 points."

[0221] The clinical data processing method based on a large language model in this invention automatically identifies "ACDF" as "anterior cervical discectomy and fusion," establishes a pre- and post-operative symptom comparison model, and converts changes in the JOA score into a treatment efficacy assessment. The output structure includes surgical information, symptom changes, and functional assessment. The system triggers a clinical alert: an electromyography (EMG) is recommended to rule out C5 radiculopathy.

[0222] The treatment effect evaluation formula is:

[0223]

[0224] ImprovementRate represents the treatment improvement rate, which indicates the improvement ratio relative to the preoperative status; CurrentScore represents the current score; MinScore represents the minimum value of the scoring scale; and PreopScore represents the preoperative score.

[0225] Example 3: Clinical Integration Solution

[0226] The clinical data processing method based on the large language model of the present invention is connected with the hospital information system to realize a two-way data flow: basic information of patients is obtained from the HIS system, and structured medical record data is returned after being processed by the method of the present invention.

[0227] The present invention's clinical data processing method based on a large language model (AI) utilizes an artificial intelligence (AI) large language model to process clinical data. In this method, AI analyzes medical data and automatically verifies the reliability of the results. First, the large language model reads and analyzes the patient's clinical data (such as medical records and test results) to obtain an analytical conclusion. Then, relevant data supporting this conclusion is collected, such as the test values ​​used to support the conclusion. Next, the accuracy and reliability of this supporting data are verified. First, the data source is checked for accuracy (e.g., the condition of the equipment used to collect this data and its completeness); second, the data is checked for timeliness and stability (e.g., whether the data is up-to-date and whether historical records are consistent). This dual verification process assigns a "reliability score" to the AI's analysis conclusion. Finally, based on this "reliability score," the AI's conclusion is evaluated and feedback is provided. A high score indicates that the AI's conclusion is credible; a low score indicates a problem or requires further verification. This invention can improve the accuracy and reliability of AI-powered clinical data analysis, enabling medical professionals to more confidently utilize AI results in practical applications, thereby improving diagnostic and research efficiency.

[0228] Another specific embodiment of the present invention discloses a clinical data processing device based on a large language model, such as Figure 6 As shown, including:

[0229] Data preprocessing layer 610 , feature encoding layer 620 , grammatical constraint decoding layer 630 , pointer generation layer 640 , multimodal verification layer 650 and output conversion layer 660 .

[0230] The data preprocessing layer 610 is configured to preprocess the spinal disease clinical text data based on the BERT large language model;

[0231] The feature encoding layer 620 is configured to extract features from the preprocessed data based on the Transformer architecture of the large language model;

[0232] The grammatical constraint decoding layer 630 is configured to perform grammatical constraint decoding on the spinal disease clinical text data after feature extraction, and generate a legal token set for each decoding step according to a dynamic vocabulary;

[0233] The pointer generation layer 640 is configured to calculate the replication probability based on the term importance weight and the clinical relevance of spinal diseases;

[0234] The multimodal verification layer 650 is configured to perform cross-field consistency check on the spinal disease clinical text data;

[0235] The output conversion layer 660 is configured to convert the spinal disease clinical text data that has undergone consistency check into a target format to generate a document that meets preset standards.

[0236] Optionally, the data pre-processing layer 610 is further configured to:

[0237] Perform medical terminology standardization and clinical entity annotation;

[0238] The implementation of medical terminology standardization includes:

[0239] Screen candidate terms;

[0240] Performing anatomical spatial consistency checking on the candidate terms;

[0241] Calculating the comprehensive similarity of the candidate terms and ranking them;

[0242] Two medical terms with a comprehensive similarity higher than a preset threshold are marked as the same word.

[0243] Optionally, the comprehensive similarity is determined according to the following method:

[0244] Use the first formula to filter candidate terms;

[0245] The terms with similarity greater than the preset similarity in the first formula screening results are directly mapped with strings, and the remaining terms are deep semantically matched using the second formula;

[0246] performing anatomical space constraint checking on the matching result of the second formula;

[0247] Among them, the first formula is:

[0248]

[0249] ClinicalED(a, b) is the clinical edit distance, which represents the similarity between medical terms a and b; a and b are both medical terms; ED(a, b) is the standard edit distance, which represents the minimum number of operations required to convert string a to b; Syn(a) represents the set of synonyms of medical term a; min a′∈Syn(a) ED(a′, b) represents the value with the smallest edit distance to b among all synonyms of a; a′ represents the synonym of a;

[0250] The second formula is:

[0251]

[0252] TermSim(u, v) represents the similarity between medical terms u and v; w1 and w2 are weight parameters; BERTScore(u, v) represents the semantic similarity of medical terms calculated based on the BERT model; AnatomyDist(u, v) represents the spatial distance of anatomical locations; SemDiff(u, v) represents the semantic difference between terms; k is the adjustment factor; u and v are both medical terms.

[0253] Optionally, the feature encoding layer 620 is further configured to:

[0254] Modeling symptom descriptions through three attention mechanisms;

[0255] Optimize symptom descriptions through clinical knowledge;

[0256] Among them, the three attention mechanisms are:

[0257] Local window attention mechanism for processing symptom details;

[0258] Global sparse attention mechanism to capture key clinical entities;

[0259] Cross-modal attention mechanism for associating imaging findings with symptom descriptions.

[0260] Optionally, the attention mechanism is implemented by the following formula:

[0261]

[0262] Among them, A hier is the hierarchical attention weight matrix; Q represents the query matrix, which represents the information that needs to be paid attention to at present; K local The key matrix representing the local context; K global A key matrix representing the global context; represents the matrix concatenation operation, combining local and global features; d represents the square root of the feature dimension; T1 represents the transposition operation of the matrix; Softmax() is a normalization function used to convert the attention score into a probability distribution;

[0263] Among them, the query matrix Q of the local window attention mechanism uses text features, and the query matrix Q of the cross-modal attention mechanism uses image features.

[0264] Optionally, refine the symptom description according to the following formula:

[0265] H enh =H base +LayerNorm(W k ReLU(W c ·E kb));

[0266] Among them, H enh represents the feature representation matrix after knowledge enhancement; H base Represents the feature matrix output by the basic encoder; LayerNorm represents the layer normalization operation; W k and W c are all trainable parameter matrices; ReLU() represents the rectified linear unit activation function, which is used to introduce nonlinear transformation; E kb Represents the clinical knowledge graph embedding matrix, which is used to contain the domain knowledge of spinal diseases.

[0267] Optionally, the syntax constraint decoding layer 630 is further configured to:

[0268] Use the third formula to perform initial decoding to obtain the basic structure, and when the key field is detected, use the fourth formula to perform decoding;

[0269] The intersection of the decoding result of the third formula and the decoding result of the fourth formula is used as the legal token set;

[0270] Wherein, the third formula is:

[0271]

[0272] v is the spinal medicine vocabulary; t is the decoding step number; V t is the set of tokens obtained by decoding in step t; V is the complete vocabulary; CFGCheck() is the context-free grammar check function; S1:t-1 represents the token sequence generated by decoding in the previous step t-1; Represents the matrix concatenation operation; G spine Special grammar for spinal diseases;

[0273] The fourth formula is:

[0274] L t =Lookahead(y1:t-1, G, 3);

[0275] L t is the set of tokens obtained by decoding in the t-th step; y1:t-1 represents the token sequence generated by decoding in the previous t-1 step; G represents the set of grammatical rules that define the output structure; 3 represents the number of look-ahead steps, which is used to look forward 3 steps for grammatical analysis; Lookahead() represents the look-ahead analysis function, which is used to predict legal subsequent tokens based on the generated sequence and grammatical rules.

[0276] Optionally, the pointer generation layer 640 is further configured to:

[0277] Determine a generation mode according to a fifth formula, wherein the generation mode includes generation or replication;

[0278] If the generation mode is generation, the vocabulary probability is calculated according to the sixth formula, and if the generation mode is copy, the input term is selected according to the seventh formula;

[0279] Probability normalization is performed according to the eighth formula;

[0280] Wherein, determining the generation mode according to the fifth formula includes:

[0281] Calculating the generation probability according to the fifth formula, if the generation probability is greater than a preset generation probability threshold, the generation mode is generation; otherwise, the generation mode is copy;

[0282] The fifth formula is:

[0283]

[0284] ρ gen To generate probability; Sigmoid() represents the S-type activation function, which maps the value to the interval [0, 1] as the probability value; Represents the parameter vector for generating probability calculation; tanh represents the hyperbolic tangent activation function, which maps the value to the interval [-1, 1]; W g Represents the weight matrix for generating probability calculation; h t Represents the current decoding state vector; c t represents the context vector; t is the decoding step number;

[0285] Optionally, the sixth formula is:

[0286] P vocab (w t )=softmax(W vocab h t +b vocab );

[0287] Among them, P vocab (w t ) means that when decoding in step t, word w is selected from the vocabulary t This represents the probability that the output of the current decoding step is generated from the vocabulary. vocab is the parameter matrix learned during the training process, which is used to transform the hidden state h of the decoder t Mapped to the space of each word in the vocabulary to calculate the probability of each word being selected. t Represents the current decoding state vector, which contains all the information that the model has integrated in the previous historical information (text context, decoded words) in the current decoding step. It is usually obtained through Transformer or other recurrent neural network structures. vocabRepresents a bias term, which is used to adjust the probability distribution of model output; It is used to adjust the probability distribution of model output and is usually also obtained through training and learning.

[0288] The seventh formula is:

[0289]

[0290] i and j indicate the term numbers; represents the importance weight of the i-th term; q represents the query vector, which represents the current decoding state; k i represents the key vector representation of the i-th term; T represents the temperature parameter used to control the smoothness of the distribution; τ represents the set of all candidate terms; exp() represents the natural exponential function; ∑ j∈τ represents the sum of all candidate terms for normalization;

[0291] Optionally, the eighth formula is:

[0292]

[0293] P norm (w t ) means that after normalization, the word w is selected in step t t The final probability value, w t The word selected in step t; the normalized result ensures that the sum of the probabilities of all words is 1, which is a property that the probability distribution must satisfy; S(w t ) represents the word w t The raw score at the current step is usually generated by the previous module (for example, the model's predicted score for the current word); T is the temperature parameter, which is used to control the smoothness of the output probability distribution; when T is small (for example, <1), the output probability distribution tends to be steeper and more certain, and the model is more inclined to select the word with the highest score (that is, the probability distribution is sharper); when T is large (for example, >1), the output probability distribution becomes flatter, and the model generates more even probabilities, reducing the probability of a single word being selected with excessive certainty. e() represents the natural exponential function, which is used to ensure that any real number becomes a positive number; represents the sum of the indexed scores of all selected words, and j represents the word number. It is used to normalize the entire probability to ensure that the sum of the probabilities of all candidate words is 1.

[0294] The multimodal authentication layer 650 is also configured to:

[0295] Grammatical verification, clinical logic verification and medical fact verification;

[0296] The syntax check is used to ensure that the output complies with predetermined specifications, the clinical logic is used to verify the rationality of numerical values, and the medical fact check is used to compare with the latest clinical guidelines.

[0297] Optionally, the verification scoring method for the syntax check, clinical logic check and medical fact check is:

[0298] First, perform clinical logic dimension verification and convert the verification results into a score value of 0 or 1;

[0299] Determine a comprehensive scoring value based on the scoring value;

[0300] The clinical logic dimension verification includes:

[0301] The clinical logic dimension is verified by the following formula:

[0302] RangeCheck(f1,f)=I(u f -2σ f ≤f1≤u f +2σ f );

[0303] RangeCheck(f1, f) indicates the result of range checking on the value f1 in field f; f1 indicates the value to be verified; f indicates the field identifier; u f represents the expected mean of field f; σ f Indicates the standard deviation of the field f; I() represents an indicator function, which returns 1 if the condition is true, otherwise it returns 0;

[0304] Determining a comprehensive score value according to the score value includes:

[0305]

[0306] ValidScore represents the comprehensive score value; represents the product of the five evaluation dimensions, k represents the evaluation dimension number; η k represents the weight of the kth dimension; f k (R) represents the scoring function of the kth dimension on the structured result R; exp(-λ·ConflictCount) represents the conflict penalty term; λ represents the conflict penalty coefficient; ConflictCount represents the number of conflicts detected; R represents the structured result generated by the system.

[0307] Optional, such as Figure 6 The illustrated apparatus 600 further includes an optimized verification layer configured to accelerate grammar verification using a caching mechanism;

[0308] The cache mechanism is implemented through the following formula:

[0309]

[0310] Among them, CacheHitRate represents the cache hit rate; MissCount represents the number of queries that miss the cache; TotalQuery represents the total number of queries; e -βt represents the decay factor over time; β represents the decay coefficient; t' represents the system operation time.

[0311] It should be noted that the device provided in this embodiment belongs to the same inventive concept as the above-mentioned method embodiment. The device of this embodiment can implement all the method steps of the above-mentioned method embodiment, solve the same technical problems, and obtain the same technical effects. The similarities will not be repeated here.

[0312] Another specific embodiment of the present invention discloses an electronic device, such as Figure 7 As shown, it includes: a memory 702 and one or more processors 701.

[0313] One or more application programs are stored in the memory 702, and the one or more application programs are suitable for being executed by the one or more processors 701 to implement:

[0314] Preprocessing of spinal disease clinical text data based on the BERT large language model;

[0315] The Transformer architecture based on the large language model performs feature extraction on the preprocessed data;

[0316] Perform grammatical constraint decoding on clinical text data of spinal diseases after feature extraction, and generate a legal token set for each decoding step based on a dynamic vocabulary;

[0317] The replication probability was calculated based on the term importance weights and clinical relevance to spinal diseases;

[0318] Performing cross-field consistency check on the spinal disease clinical text data;

[0319] Transform consistency-checked spinal disease clinical text data into target formats to generate documents that meet pre-set standards.

[0320] like Figure 7 As shown, the electronic device includes: a processor 701 and a memory 702. The processor 701 and the memory 702 are connected, for example, via a bus interface.

[0321] The structure of the electronic device does not constitute a limitation to the embodiments of the present invention.

[0322] Processor 701 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the present disclosure. Processor 701 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0323] The bus interface may include a path to transfer information between the above components. The bus interface may be a PCI bus or an EISA bus. The bus interface may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0324] The memory 702 may be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or an EEPROM, a CD-ROM or other optical disk storage, an optical disc storage (including a compact disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.

[0325] It should be noted that the device provided in this embodiment belongs to the same inventive concept as the above-mentioned method embodiment. The device of this embodiment can implement all the method steps of the above-mentioned method embodiment, solve the same technical problems, and obtain the same technical effects. The similarities will not be repeated here.

[0326] Another specific embodiment of the present invention discloses a computer-readable storage medium having a computer program stored thereon, which can be loaded and executed by a processor to perform the clinical data processing method based on a large language as described in the first aspect.

[0327] The applicant of the present invention has made a detailed explanation and description of the implementation examples of the present invention in conjunction with the drawings in the specification. However, those skilled in the art should understand that the above implementation examples are only preferred implementation plans of the present invention, and the detailed description is only to help readers better understand the spirit of the present invention, and is not a limitation on the scope of protection of the present invention. On the contrary, any improvements or modifications based on the inventive spirit of the present invention should fall within the scope of protection of the present invention.

[0328] Finally, it should be noted that the above embodiments are merely illustrative of the technical solutions of the embodiments of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they may still modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein with equivalents; and such modifications or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or replacements that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be covered by the scope of protection of the present invention.

Claims

1. A clinical data processing method based on a large language model, characterized in that: include: Preprocessing of spinal disease clinical text data based on the BERT large language model; The Transformer architecture based on the large language model performs feature extraction on the preprocessed data; Perform grammatical constraint decoding on clinical text data of spinal diseases after feature extraction, and generate a legal token set for each decoding step based on a dynamic vocabulary; The replication probability was calculated based on the term importance weights and clinical relevance to spinal diseases; Performing cross-field consistency check on the spinal disease clinical text data; Convert clinical text data of spinal diseases that have been checked for consistency into the target format to generate documents that meet the preset standards; The preprocessing of spinal disease clinical text data includes: Perform medical terminology standardization and clinical entity annotation; The implementation of medical terminology standardization includes: Screen candidate terms; Performing anatomical spatial consistency checking on the candidate terms; Calculating the comprehensive similarity of the candidate terms and ranking them; Two medical terms with a comprehensive similarity higher than a preset threshold are marked as the same word.

2. The method according to claim 1, characterized in that The comprehensive similarity is determined according to the following method: Use the first formula to filter candidate terms; The terms with similarity greater than the preset similarity in the first formula screening results are directly mapped with strings, and the remaining terms are deep semantically matched using the second formula; performing anatomical space constraint checking on the matching result of the second formula; Among them, the first formula is: ClinicalED(a,b)=min(ED(a,b),1+min a′∈Syn(a) ED(a′,b)); ClinicalED(a,b) is the clinical edit distance, which represents the similarity between medical terms a and b; a and b are both medical terms; ED(a,b) is the standard edit distance, which represents the minimum number of operations required to convert string a to b; Syn(a) represents the set of synonyms of medical term a; min a′∈Syn(a) ED(a′,b) represents the value with the smallest edit distance to b among all synonyms of a; a′ represents the synonym of a; The second formula is: TermSim(u,v) represents the similarity between medical terms u and v; w1 and w2 are weight parameters; BERTScore(u,v) represents the semantic similarity of medical terms calculated based on the BERT model; AnatomyDist(u,v) represents the spatial distance of the calculated anatomical position; SemDiff(u,v) represents the semantic difference between terms; k is the adjustment factor; u and v are both medical terms.

3. The method according to claim 1, characterized in that The Transformer architecture based on the large language model performs feature extraction on the preprocessed data, including: Modeling symptom descriptions through three attention mechanisms; Optimize symptom descriptions through clinical knowledge; Among them, the three attention mechanisms are: Local window attention mechanism for processing symptom details; Global sparse attention mechanism to capture key clinical entities; Cross-modal attention mechanism for associating imaging findings with symptom descriptions; The attention mechanism is implemented by the following formula: Among them, A hier is the hierarchical attention weight matrix; Q represents the query matrix, which represents the information that needs to be paid attention to at present; K local The key matrix representing the local context; K global A key matrix representing the global context; represents the matrix concatenation operation, combining local and global features; d represents the square root of the feature dimension; T1 represents the transposition operation of the matrix; Softmax() is a normalization function used to convert the attention score into a probability distribution; Among them, the query matrix Q of the local window attention mechanism uses text features, and the query matrix Q of the cross-modal attention mechanism uses image features.

4. The method according to claim 3, characterized in that The symptom description is optimized according to the following formula: H enh =H base +LayerNorm(W k ·ReLU(W c ·E kb )); Among them, H enh represents the feature representation matrix after knowledge enhancement; H base Represents the feature matrix output by the basic encoder; LayerNorm represents the layer normalization operation; W k and W c are all trainable parameter matrices; ReLU() represents the rectified linear unit activation function, which is used to introduce nonlinear transformation; E kb Represents the clinical knowledge graph embedding matrix, which is used to contain the domain knowledge of spinal diseases.

5. The method according to claim 1, wherein The generation of a legal token set for each decoding step according to the dynamic vocabulary includes: Use the third formula to perform initial decoding to obtain the basic structure, and when the key field is detected, use the fourth formula to perform decoding; The intersection of the decoding result of the third formula and the decoding result of the fourth formula is used as the legal token set; Wherein, the third formula is: Where v is the spinal medicine vocabulary; t is the decoding step number; V t is the set of tokens obtained by decoding in step t; V is the complete vocabulary; CFGCheck() is the context-free grammar check function; S1:t-1 represents the token sequence generated by decoding in the previous step t-1; Represents the matrix concatenation operation; G spine Special grammar for spinal diseases; The fourth formula is: L t =Lookahead(y1:t-1,G,3); Among them, L t is the set of tokens obtained by decoding in step t; y1:t-1 represents the token sequence generated by decoding in the previous step t-1; G represents the set of grammatical rules that define the output structure; 3 represents the number of lookahead steps, which is used to look forward 3 steps for grammatical analysis; Lookahead() represents the lookahead analysis function, which is used to predict legal subsequent tokens based on the generated sequence and grammatical rules.

6. The method according to claim 1, characterized in that The calculation of replication probability based on term importance weights and clinical relevance to spinal diseases includes: Determine a generation mode according to a fifth formula, wherein the generation mode includes generation or replication; If the generation mode is generation, the vocabulary probability is calculated according to the sixth formula, and if the generation mode is copy, the input term is selected according to the seventh formula; Probability normalization is performed according to the eighth formula; Wherein, determining the generation mode according to the fifth formula includes: Calculating the generation probability according to the fifth formula, if the generation probability is greater than a preset generation probability threshold, the generation mode is generation; otherwise, the generation mode is copy; The fifth formula is: Among them, ρ gen To generate probability; Sigmoid() represents the S-type activation function, which maps the value to the [0,1] interval as the probability value; Represents the parameter vector for generating probability calculation; tanh represents the hyperbolic tangent activation function, which maps the value to the [-1,1] interval; W g Represents the weight matrix for generating probability calculation; h t Represents the current decoding state vector; c t represents the context vector; t is the decoding step number; The sixth formula is: P vocab (w t )=softmax(W vocab h t +b vocab ); Among them, P vocab (w t ) means that when decoding in step t, word w is selected from the vocabulary t The probability distribution of w t The word chosen for step t; W vocab is the parameter matrix learned during the training process; b vocab Represents the bias term, which is used to adjust the probability distribution of model output; softmax() represents the normalization function; The seventh formula is: Where i and j represent the term numbers; represents the importance weight of the i-th term; q represents the query vector, which represents the current decoding state; k i represents the key vector representation of the i-th term; T represents the temperature parameter used to control the smoothness of the distribution; τ represents the set of all candidate terms; exp() represents the natural exponential function; ∑ j∈τ represents the sum of all candidate terms for normalization; The eighth formula is: Among them, P norm (w t ) means that after normalization, the word w is selected in step t t The final probability value, w t The word selected in step t; S(w t ) represents the word w t The original score at the current step; T is the temperature parameter, which is used to control the smoothness of the output probability distribution; e () represents the natural exponential function, which is used to ensure that any real number becomes a positive number; It represents the sum of the indexed scores of all selected words, and j represents the number of the word.

7. The method according to claim 1, characterized in that The cross-field consistency check of the spinal disease clinical text data includes: Grammatical verification, clinical logic verification and medical fact verification; The syntax check is used to ensure that the output meets the predetermined specifications, the clinical logic is used to verify the rationality of the numerical value, and the medical fact check is used to compare with the latest clinical guidelines; The verification scoring method for the grammatical verification, clinical logic verification and medical fact verification is: First, perform clinical logic dimension verification and convert the verification results into a score value of 0 or 1; Determine a comprehensive scoring value according to the scoring value; The clinical logic dimension verification includes: Clinical logic dimension verification is performed using the following formula: RangeCheck(f1,f)=I(u f -2σ f ≤f1≤u f +2σ f ); Among them, RangeCheck(f1,f) represents the result of range checking on the value f1 in field f; f1 represents the value to be verified; f represents the field identifier; u f represents the expected mean of field f; σ f Indicates the standard deviation of the field f; I() represents an indicator function, which returns 1 if the condition is true, otherwise it returns 0; Determining a comprehensive score value according to the score value includes: Among them, ValidScore represents the comprehensive score value; represents the product of the five evaluation dimensions, k represents the evaluation dimension number; η k represents the weight of the kth dimension; f k (R) represents the scoring function of the kth dimension on the structured result R; exp(-λ·ConflictCount) represents the conflict penalty term; λ represents the conflict penalty coefficient; ConflictCount represents the number of conflicts detected; R represents the structured result generated by the system.

8. The method according to claim 1, characterized in that Also includes: Use cache mechanism to speed up grammar verification; The cache mechanism is implemented through the following formula: Among them, CacheHitRate represents the cache hit rate; MissCount represents the number of queries that miss the cache; TotalQuery represents the total number of queries; e -βt‘ represents the decay factor over time; β represents the decay coefficient; t' represents the system operation time.

9. A clinical data processing device based on a large language model, characterized in that: A method for structuring spinal disease clinical text data based on a large language model according to any one of claims 1 to 8 is provided, the device comprising: Data preprocessing layer, feature encoding layer, grammatical constraint decoding layer, pointer generation layer, multimodal verification layer and output conversion layer; The data preprocessing layer is configured to preprocess the spinal disease clinical text data based on the BERT large language model; The feature encoding layer is configured to extract features from the preprocessed data based on the Transformer architecture of the large language model; The grammatical constraint decoding layer is configured to perform grammatical constraint decoding on the clinical text data of spinal diseases after feature extraction, and generate a legal token set for each decoding step according to a dynamic vocabulary; The pointer generation layer is configured to calculate replication probability based on term importance weights and clinical relevance of spinal diseases; The multimodal verification layer is configured to perform cross-field consistency check on the spinal disease clinical text data; The output conversion layer is configured to convert the spinal disease clinical text data that has passed the consistency check into a target format to generate a document that meets preset standards.

10. A device for structuring spinal disease clinical text data based on a large language model, characterized in that: including a memory, a processor, and a user interface; The memory is used to store computer programs; The user interface is used to interact with the user; The processor is configured to read the computer program in the memory, and when the processor executes the computer program, implements the clinical data processing method based on a large language model as claimed in any one of claims 1 to 8.

Citation Information

Cited By

  • Medical image data management method and system based on large language model

    CN121148734A

  • Case quality evaluation method, computer device and storage medium

    CN122114746A