Intelligent stroke staging method, device and storage medium
By applying pre-training models in the electronic medical record system for data preprocessing and intelligent staging, the problem of staging of stroke diseases in the existing technology is solved, and efficient and accurate automated staging is achieved.
Patent Information
- Application Number
- CN202510090503.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The existing electronic medical record system relies on manual reading and summary of experienced clinicians in the staging of stroke diseases, resulting in a large amount of time and research funding.
By obtaining the patient's complaints, current medical history and diagnostic data from the electronic medical record system, using pre-trained disease classification model, time entity recognition model and symptom standardization matching model, data preprocessing, disease identification, symptom time extraction and standardization processing are automated, and intelligent staging is finally carried out according to the preset stroke disease staging rules.
It realizes automated staging of stroke diseases, reduces the workload and time of clinicians, reduces the consumption of research funds, and improves the accuracy and efficiency of staging results.
Smart Images

Figure CN119517386B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical information technology, and in particular to a method, device and storage medium for intelligent staging of stroke. Background Art
[0002] Stroke is an acute neurological deficit syndrome and a cerebral blood circulation disorder caused by cerebrovascular disease. It can be divided into cerebral infarction and cerebral hemorrhage.
[0003] With the development of information technology and the popularization of electronic medical record systems in my country, a large amount of clinical data related to stroke is generated every day, providing important real-world data for clinical research on stroke. The characteristics, treatment patterns and prognosis of different disease stages of stroke are significantly different. At present, all stroke patients are included in the electronic medical record system, but there is a lack of directly available disease staging methods. In the electronic medical record system data, accurate screening of stroke patients according to different disease stages is the basis and premise of real-world research, which directly affects the accuracy of research results. The current hospital electronic medical record system does not directly collect and count stroke disease staging information. If disease staging data is to be obtained, experienced clinicians are required to read the relevant data in the electronic medical record system and summarize the results. This process not only consumes the energy of clinicians, but is also very long and costs a lot of research funds. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide an intelligent stroke staging method, device and storage medium to solve the problem that the existing electronic medical record system for staging of stroke diseases relies on experienced clinicians to read relevant data in the electronic medical record system for manual experience summary, which is not only time-consuming and labor-intensive for clinicians to implement, but also consumes a lot of research funds.
[0005] According to a first aspect of an embodiment of the present invention, a method for intelligent staging of stroke is provided, the method comprising:
[0006] Obtain patient complaints, history of present illness, and diagnostic data from the electronic medical record system;
[0007] Preprocessing the chief complaint, current medical history, and diagnostic data;
[0008] Inputting the preprocessed chief complaint, current medical history and diagnosis data into a pretrained disease classification model, if the output of the pretrained disease classification model is cerebral hemorrhage or cerebral infarction;
[0009] Inputting the patient's chief complaint data into a pre-trained temporal entity recognition model, the pre-trained temporal entity recognition model outputs a symptom description and its corresponding onset time and symptom time category triplet;
[0010] Inputting the symptom description and its corresponding onset time and symptom time category triplet into a pre-trained symptom standardization matching model, and the pre-trained symptom standardization matching model outputs standardized symptom description, onset time and symptom time category according to a preset standard output format;
[0011] Matching and screening the standardized symptom descriptions in a preset stroke symptom standard library to obtain stroke symptoms and their onset time and symptom time categories;
[0012] If there are multiple onset times corresponding to the same symptom, the multiple onset times are screened according to the symptom time category and the preset stroke onset time rule to determine the onset time required for the stroke disease staging;
[0013] The stroke symptoms and their corresponding onset time are judged by using preset stroke disease staging rules to obtain the patient's stroke disease staging result.
[0014] Preferably,
[0015] The pre-processing of the chief complaint, current medical history and diagnosis data includes:
[0016] Data cleaning is performed on the chief complaint, current medical history, and diagnosis data to remove special characters and non-text content;
[0017] After data cleaning, the chief complaint, current medical history, and diagnosis data were segmented;
[0018] Stop words were removed from the chief complaint, current medical history, and diagnosis data after word segmentation.
[0019] Preferably,
[0020] The preset stroke disease staging rules include:
[0021] Stroke diseases are divided into ultra-early stage, acute stage, recovery stage and sequelae stage, and the corresponding onset time is set for ultra-early stage, acute stage, recovery stage and sequelae stage respectively;
[0022] The method of obtaining the patient's stroke disease staging results includes:
[0023] The patient's stroke symptoms and their corresponding time are matched with the onset time corresponding to each period of the set stroke disease to obtain the patient's stroke disease stage.
[0024] Preferably,
[0025] The training of the pre-trained disease classification model includes:
[0026] Build a customized dual-tower cross-attention stroke disease classification model framework based on the BioClinicalBERT pre-trained model, and set the loss function of the model framework;
[0027] obtaining chief complaints, history of present illness, and diagnosis data of a plurality of patients from the electronic medical record system;
[0028] Preprocessing the chief complaints, current medical histories, and diagnostic data of the multiple patients to obtain training data;
[0029] The training data is used as the input of the customized dual-tower cross-attention stroke disease classification model framework, and the customized dual-tower cross-attention stroke disease classification model framework is iteratively trained until the loss value of the loss function no longer decreases or reaches a preset number of iterations, thereby obtaining a pre-trained disease classification model.
[0030] Preferably,
[0031] The loss function of the customized dual-tower cross-attention stroke disease classification model framework adopts a smooth L1 loss function;
[0032] The expression of the smooth L1 loss function is: In the formula, represents the loss value; y represents the true label; represents the predicted value; Represents the balance parameter, which is used to control the balance between squared loss and absolute loss.
[0033] Preferably,
[0034] The training of the pre-trained temporal entity recognition model includes:
[0035] Build a temporal entity recognition model framework based on the ChatGLM large language model, set the loss function of the model framework, and use the cross entropy loss function;
[0036] obtaining chief complaint information of a plurality of patients from the electronic medical record system;
[0037] Preprocessing the chief complaint information of the plurality of patients to obtain training data;
[0038] The training data is used as the input of the temporal entity recognition model framework, and the temporal entity recognition model framework is iteratively trained until the loss value of the loss function no longer decreases or reaches a preset iteration round, thereby obtaining a pre-trained temporal entity recognition model.
[0039] According to a second aspect of an embodiment of the present invention, there is provided an intelligent stroke staging device, the device comprising:
[0040] Data acquisition module: used to obtain the patient's chief complaint, current medical history and diagnosis data from the electronic medical record system;
[0041] Preprocessing module: used for preprocessing the chief complaint, current medical history and diagnostic data;
[0042] Disease identification module: used to input the pre-processed chief complaint, current medical history and diagnosis data into the pre-trained disease classification model, if the output of the pre-trained disease classification model is cerebral hemorrhage or cerebral infarction;
[0043] Symptom time acquisition module: used to input the patient's main complaint data into a pre-trained time entity recognition model, and the pre-trained time entity recognition model outputs symptom descriptions and their corresponding onset times;
[0044] Symptom description standardization module: used to input the symptom description and its corresponding onset time into a pre-trained symptom standardization matching model, and the pre-trained symptom standardization matching model outputs standardized symptom description, onset time and symptom time category according to a preset standard output format;
[0045] Stroke symptom time acquisition module: used to match and screen the standardized symptom description in the preset stroke symptom standard library to obtain the stroke symptoms and their onset time;
[0046] Onset time screening module: if there are multiple onset times corresponding to the same onset symptom, the multiple onset times are screened according to the symptom time category and the preset stroke onset time rule to determine the onset time required for stroke disease staging;
[0047] Staging judgment module: used to judge the stroke symptoms and their corresponding onset time by using preset stroke disease staging rules to obtain the patient's stroke disease staging results.
[0048] According to a third aspect of an embodiment of the present invention, there is provided a storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a main controller, each step in the above method is implemented.
[0049] The technical solution provided by the embodiments of the present invention may have the following beneficial effects:
[0050] The present application obtains the patient's chief complaint, current medical history and discharge diagnosis information based on the existing electronic medical record system, and determines whether the patient is admitted to the hospital due to cerebral hemorrhage or cerebral infarction based on a pre-trained disease classification model. If so, the patient's chief complaint data is input into the pre-trained time entity recognition model to obtain the symptom description and its corresponding onset time, and then performs standardized description processing. The standardized symptom description is matched and screened in a preset stroke symptom standard library, and the symptoms of stroke onset and their corresponding time are obtained according to the preset stroke onset time rules; finally, the patient's stroke disease is staged according to the preset stroke disease staging rules; thereby solving the problem in the prior art that the staging of stroke disease in the electronic medical record system relies on experienced clinicians to read the relevant data in the electronic medical record system for manual experience summary, which is not only time-consuming and labor-intensive for clinicians to implement, but also consumes a lot of research funds.
[0051] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0053] Figure 1 is a flow chart of a stroke intelligent staging method according to an exemplary embodiment;
[0054] Figure 2 is a schematic diagram of the architecture of a stroke intelligent staging method according to another exemplary embodiment;
[0055] Figure 3 is a schematic diagram of a time value acquisition process framework according to another exemplary embodiment;
[0056] Figure 4 is a system schematic diagram of a stroke intelligent staging device according to another exemplary embodiment;
[0057] In the attached figure: 1-data acquisition module, 2-preprocessing module, 3-disease identification module, 4-symptom time acquisition module, 5-symptom description standardization module, 6-stroke symptom time acquisition module, 7-onset time screening module, 8-staging judgment module. DETAILED DESCRIPTION
[0058] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0059] Embodiment 1:
[0060] Figure 1 is a flow chart of a stroke intelligent staging method according to an exemplary embodiment. Figure 1 As shown, the method includes:
[0061] S1, obtain the patient's chief complaint, history of present illness, and diagnosis data from the electronic medical record system;
[0062] S2, preprocessing the chief complaint, current medical history, and diagnostic data;
[0063] S3, inputting the preprocessed chief complaint, current medical history and diagnosis data into a pretrained disease classification model, if the output of the pretrained disease classification model is cerebral hemorrhage or cerebral infarction;
[0064] S4, inputting the patient's chief complaint data into a pre-trained temporal entity recognition model, wherein the pre-trained temporal entity recognition model outputs a symptom description and its corresponding onset time and symptom time category triplet;
[0065] S5, inputting the symptom description and its corresponding onset time and symptom time category triplet into a pre-trained symptom standardization matching model, and the pre-trained symptom standardization matching model outputs standardized symptom description, onset time and symptom time category according to a preset standard output format;
[0066] S6, matching and screening the standardized symptom description in a preset stroke symptom standard library to obtain stroke symptoms and their onset time and symptom time category;
[0067] S7, if there are multiple onset times corresponding to the same symptom, the multiple onset times are screened according to the symptom time category and the preset stroke onset time rule to determine the onset time required for the stroke disease staging;
[0068] S8, judging the stroke symptoms and their corresponding onset time by using a preset stroke disease staging rule to obtain a stroke disease staging result for the patient;
[0069] It is understandable that if Figure 2As shown, the present application obtains the patient's chief complaint, current medical history and diagnostic data in the electronic medical record system, wherein the chief complaint refers to the first item in the inpatient medical record, which is used to describe the main symptoms (or signs) and duration of the patient's current visit. The first diagnosis can be generated based on the chief complaint, and the language of the chief complaint usually uses medical terms; the current medical history refers to the detailed circumstances of the occurrence, evolution, diagnosis and treatment of the patient's current disease, which is written in chronological order. The current medical history is the most important part of the medical history, including the initial symptoms of the current disease, the entire process from the beginning to the visit, that is, the occurrence, development, evolution and diagnosis and treatment process; diagnostic data refers to the conclusions drawn on the cause, location, nature, degree of damage, etc. of the patient's disease through medical examination and clinical analysis. Diagnostic data usually include pathological diagnosis, cytological diagnosis, clinical diagnosis and other types. These diagnostic data are the results of comprehensive analysis by doctors based on the patient's condition and examination results;
[0070] Preprocess the acquired chief complaint, current medical history and diagnosis data, including: data cleaning: remove useless information, such as special characters, non-text content, etc.; word segmentation: divide the text into words; stop word removal: remove common stop words, such as prepositions, indefinite articles, etc.;
[0071] The pre-processed patient complaints, current medical history, and diagnosis data are input into the pre-trained disease classification model, where the training of the disease classification model includes:
[0072] The framework of the disease classification model uses a customized dual-tower cross-attention stroke disease classification model framework based on the pre-trained model of BioClinicalBERT. The BioClinicalBERT model framework is a pre-trained language model designed specifically for the biomedical field. It is optimized and improved based on the BERT (Bidirectional Encoder Representations from Transformers) model to better meet the processing needs of biomedical texts. The BioClinicalBERT model framework is based on BERT and has been specifically pre-trained and optimized for the biomedical field. It is trained on a large amount of biomedical text data to learn specific language features and knowledge in the biomedical field, thereby performing well in biomedical-related natural language processing tasks. The BioClinicalBERT model framework mainly includes the following parts:
[0073] Pre-training data: Use a large amount of biomedical literature, papers, clinical records, etc. as training data;
[0074] Pre-training tasks: Uses similar pre-training tasks as BERT, including masked language model (MLM) and sentence-level continuity prediction tasks, but with specific adjustments and optimizations for biomedical text;
[0075] Model structure: The encoder part based on Transformer adopts a multi-layer bidirectional Transformer structure, which can fully understand the bidirectional contextual information in the text;
[0076] Fine-tuning process: Based on pre-training, fine-tuning is performed in combination with specific biomedical tasks to adapt to different downstream task requirements; the advantage of the BioClinicalBERT model framework is that it can accurately understand the professional terms and complex contexts in the biomedical field, provide high-quality semantic representation, and thus perform well in various biomedical NLP tasks.
[0077] The Huber loss is used as the loss function in the training process of the customized dual-tower cross-attention stroke disease classification model framework. The expression of the smooth L1 loss function is: In the formula, represents the loss value; y represents the true label; represents the predicted value; represents the balance parameter, which is used to control the balance between square loss and absolute loss;
[0078] It is worth mentioning that the smoothed L1 loss function is more advantageous than the mean square error (MSE) in dealing with outliers. When the error is large, the smoothed L1 loss function becomes a linear growth, similar to the mean absolute error (MAE), thereby reducing the impact of outliers on the model. Unlike MAE, the smoothed L1 loss function is smooth at the turning point, which makes it more stable and converges faster during the optimization process. By adjusting the parameter δ, a flexible balance can be made between MSE and MAE to make it suitable for different application scenarios. When the difference between the true value and the predicted value is less than δ, the smoothed L1 loss function uses MSE and can converge quickly. When the difference is greater than δ, MAE is used, which is insensitive to outliers. Compared with MSE, the smoothed L1 loss function can fall near the minimum value more accurately when the gradient decreases, and is more robust to outliers.
[0079] Then, a large amount of patient complaints, current medical history, and diagnosis data are obtained from the electronic medical record system. After the above preprocessing, the customized dual-tower cross-attention stroke disease classification model framework is iteratively trained until the loss value of the smoothed L1 loss function no longer decreases or reaches the preset iteration round, thereby obtaining a pre-trained disease classification model;
[0080] If the output of the pre-trained disease classification model is that the patient is admitted to the hospital due to cerebral hemorrhage or cerebral infarction, the patient's chief complaint information is input into the pre-trained temporal entity recognition model, wherein the training of the temporal entity recognition model includes:
[0081] A temporal entity recognition model framework based on the ChatGLM large language model is selected, and the cross entropy loss function is used as the loss function in the model framework training process;
[0082] Similarly, the chief complaint information of multiple patients is obtained from the electronic medical record system; the chief complaint information of multiple patients is preprocessed as above to obtain training data; the temporal entity recognition model framework based on the ChatGLM large language model is iteratively trained through the training data until the loss value of the cross entropy loss no longer decreases or reaches a preset iteration round;
[0083] It is worth mentioning that
[0084] The ChatGLM large language model framework is a generative language model based on the Transformer architecture, which is specially designed to handle natural language generation tasks in dialogue systems. The ChatGLM model relies on the Transformer architecture, has highly parallel computing capabilities, and can capture long-distance language dependencies. The ChatGLM model is built based on OpenAI's GPT model framework, using a large-scale pre-training data set to learn language patterns and the ability to generate text. It can understand context and generate coherent and natural responses, and is suitable for building dialogue systems, intelligent customer service, chatbots and other applications, providing a more interactive and humane dialogue experience. The core principle of the ChatGLM model lies in its highly parallel computing capabilities based on the Transformer architecture and its ability to capture long-distance language dependencies, which enables the model to perform well in handling complex natural language generation tasks and generate coherent and natural text responses.
[0085] The calculation method of the cross entropy loss function is simple and can be directly implemented through standard mathematical libraries. At the same time, its intuitiveness allows us to easily understand the difference between the model prediction and the actual situation. The cross entropy loss function has good mathematical properties, such as convexity and differentiability, which help to ensure the stability and effectiveness of the optimization process. It can handle multi-category classification problems well, and obtain the total loss by calculating the loss of each category separately and summing them up. Compared with other loss functions, the cross entropy loss function can converge to the optimal solution faster during the back propagation process because its gradient is proportional to the difference between the predicted value and the true value, thus avoiding the problem of gradient vanishing or exploding.
[0086] Relevant symptoms and their onset time are obtained through a pre-trained temporal entity recognition model, and the relevant symptoms and their onset time are standardized to obtain a standardized symptom description corresponding to the symptom. The output of the temporal entity recognition model is in a preset standard output format, including: standardized symptom description, onset time, and symptom time category, wherein the symptom time category includes: the symptom is the first onset, recurrence, or aggravation, for example: "{symptom, time, first onset}, {symptom, time, recurrence}, {symptom, time, aggravation}"; the standardization process can be achieved through a pre-trained symptom standardization model, that is, the description of the relevant symptoms is converted into a language described by special medical terms, so as to facilitate subsequent matching and screening;
[0087] The standardized symptoms are matched and screened in the pre-set stroke symptom standard library to obtain the stroke onset symptoms and their corresponding time. If there are multiple onset times corresponding to the same onset symptom, the multiple onset times are screened according to the symptom time category and the pre-set stroke onset time rules to determine the onset time required for the stroke disease staging; wherein, the pre-set stroke symptom standard library and onset time rules include:
[0088] As attached Figure 3As shown, first determine whether the time corresponding to the symptoms of stroke is one. If so, take a single time directly. If not, determine whether the chief complaint diagnosis contains: "acute, subacute, recovery period, old / sequelae". If it does, take the corresponding single time of "acute, subacute, recovery period, old / sequelae", where "acute" corresponds to less than or equal to 14 days, subacute corresponds to 15 days to 2 months, recovery period corresponds to 2 months to 6 months, and old / sequelae corresponds to more than 6 months; if it does not contain "acute, subacute, recovery period, old / sequelae" or contains but the corresponding time is more than one, determine whether there is "cavity*infarction" in the chief complaint diagnosis. If so, determine whether the diagnosis ranking is greater than or equal to 4 If yes, the time cannot be determined. If not, determine whether the diagnostic ranking is 2 or 3. If yes, take the maximum time of non-"cerebral hemorrhage", and determine whether it includes "acute, subacute, recovery period, old / sequelae". If not, the maximum time of non-"cerebral hemorrhage" is the time value. If it is included, continue to determine whether the maximum time of non-"cerebral hemorrhage" is within the corresponding range. If it is, determine the maximum time of non-"cerebral hemorrhage" as the time value. If not, the time value cannot be determined. If the chief complaint diagnosis does not include "cavitary infarction" or the diagnostic ranking is not 2 or 3; then determine whether it is a single symptom. If yes, determine whether the symptom time category is recurrence. If yes, take the recurrence time. At the same time, determine whether it contains "acute, subacute, recovery period, old / sequelae". If not, determine the relapse time as the time value. If included, continue to determine whether the relapse time is within the corresponding range. If it is, determine the relapse time as the time value. If not, the time cannot be determined. If the time symptom category is not relapse, determine whether the symptom time category is aggravation. If so, determine whether the onset time is less than or equal to 2 months. If so, take the onset time. If not, take the aggravation time, and continue to determine whether it contains "acute, subacute, recovery period, old / sequelae". If not, determine the onset time or aggravation time as the time value. If included, determine the onset time or aggravation time. Whether the time is within the corresponding range, if so, the onset time or aggravation time is determined as the time value, if not, the time value cannot be determined; if it is not a single symptom or the symptom time category is not aggravation, then determine whether there are two times, if so, then determine whether the larger time of the two times is less than or equal to 2 months, if so, then take the larger time, if not, then take the smaller time, and continue to determine whether it contains "acute, subacute, recovery period, old / sequelae", if not, then take the larger time or the smaller time as the time value, if included, then determine whether the larger time or the smaller time is within the corresponding range, if yes, then the larger time or the smaller time is the time value, if not, then the time cannot be determined;If there are not two times, determine whether the symptom time category is relapse. If so, take the relapse time and continue to determine whether it contains: "acute, subacute, recovery period, old / sequelae". If not, take the relapse time as the time value. If included, determine whether the relapse time is within the corresponding range. If so, take the relapse time as the time value. If not, the time value cannot be determined. If the symptom time category is not relapse, determine whether the symptom time category is aggravation. If so, continue to determine whether the minimum onset time is less than or equal to two months. If so, take the minimum onset time. If not, take the aggravation. Time, and then continue to determine whether it contains "acute, subacute, recovery period, obsolete / sequelae". If not, the minimum onset time or aggravation time is the time value. If it is included, continue to determine whether the minimum onset time or aggravation time is within the corresponding range. If it is, the minimum onset time or aggravation time is the time value. If not, the time value cannot be determined. If the symptom time category is not aggravation, determine whether it contains "acute, subacute, recovery period, obsolete / sequelae". If so, take the corresponding time range as the time value. If it does not contain "acute, subacute, recovery period, obsolete / sequelae", the time cannot be determined.
[0089] Finally, the patient's staging results are obtained according to the stroke staging rules defined by the research needs, the patient's stroke symptoms and their corresponding time. The example of the pre-set stroke staging rules here is: the stroke disease is divided into ultra-early stage (onset time is less than or equal to 6 hours), acute stage (onset time is 6 hours to 2 weeks), recovery stage (onset time is 2 weeks to 6 months), and sequelae stage (onset time is more than 6 months).
[0090] Most of the existing methods for staging other diseases based on electronic medical record system data extract time from the chief complaint text based on regular expressions to obtain the onset time of the disease, and then determine the disease stage according to the disease staging rules. First, a patient in the electronic medical record system often has multiple disease diagnoses. There is no one-to-one correspondence between the time extracted by this type of method and the diagnosed disease, and the onset time and disease stage of a specific disease cannot be determined. Secondly, this type of method is only responsible for extracting all the time information in the chief complaint, and cannot handle the correspondence problem between multiple symptoms and multiple times in the chief complaint text. The extraction of symptom corresponding time is inaccurate. Finally, if a symptom in the chief complaint contains multiple time information, the onset of a specific disease cannot be determined. Which one is the disease time? Even if we further use machine learning algorithms combined with pre-defined disease rules to match the diseases in the chief complaint text and their corresponding onset times, the disease symptoms are diverse and the descriptions are different. The pre-enumerated symptom rules are relatively fixed and rely on professionals to manually update the rule base. They lack generalization capabilities and are difficult to accurately extract the onset time and stage the disease. Based on this, this embodiment uses artificial intelligence technology to innovatively construct a three-stage model for stroke scenarios, "disease population identification-symptom standardization and time extraction and positioning-dynamic staging" model, which unifies population screening and disease staging to achieve integrated and automated intelligent and precise stroke disease staging.
[0091] In the three-stage model of this embodiment, the first stage uses the BioClinicalBERT pre-trained language model, combined with the two parts of text data from different input sources required by this patent, current medical history and diagnosis data, to construct a two-tower basic classification model based on cross-attention, which can dynamically consider the differences and connections between the target disease-stroke from the current medical history and diagnosis data, enhance the model's ability to identify people with stroke, and more accurately identify hospitalized patients due to stroke to lay a foundation for the next step of disease time extraction.
[0092] In the second stage, after the stroke population model in the first stage, accurate hospitalized patients due to stroke were screened out. Then, the ChatGLM large language model was used as the base model. Based on its powerful language comprehension ability, the main complaint text of such patients was input, and the triple recognition downstream task of the disease and its corresponding time was innovatively constructed. In this stage, different disease symptoms can be automatically identified and standardized, and then the corresponding time is output. According to whether the disease symptoms are recurring or occurring for the first time, the output standard format is defined as "{symptom, time, first occurrence}, {symptom, time, recurrence}, {symptom, time, aggravation}", etc. Then, according to the stroke symptom standard library pre-injected into the base large language model, the onset symptoms belonging to the stroke disease and their corresponding onset time are matched. If a symptom has multiple times, the onset time of the symptom is determined according to the onset time rule injected into the base large model;
[0093] In the third stage, the disease staging rules are determined according to the research scenario requirements. The researchers dynamically input the disease staging rules. Based on the symptoms of stroke and their corresponding times screened out in the second stage, combined with the staging rules, the staging of stroke disease is obtained, thereby completing the intelligent staging of stroke.
[0094] Embodiment 2:
[0095] Figure 4 is a system schematic diagram of a stroke intelligent staging device according to another exemplary embodiment, the device comprising:
[0096] Data acquisition module 1: used to obtain the patient's chief complaint, current medical history and diagnosis data from the electronic medical record system;
[0097] Preprocessing module 2: used for preprocessing the chief complaint, current medical history and diagnostic data;
[0098] Disease identification module 3: used to input the pre-processed chief complaint, current medical history and diagnosis data into the pre-trained disease classification model, if the output of the pre-trained disease classification model is cerebral hemorrhage or cerebral infarction;
[0099] Symptom time acquisition module 4: used to input the patient's main complaint data into a pre-trained time entity recognition model, and the pre-trained time entity recognition model outputs symptom descriptions and their corresponding onset times;
[0100] Symptom description standardization module 5: used to input the symptom description and its corresponding onset time into a pre-trained symptom standardization matching model, and the pre-trained symptom standardization matching model outputs standardized symptom description, onset time and symptom time category according to a preset standard output format;
[0101] Stroke symptom time acquisition module 6: used to match and screen the standardized symptom description in a preset stroke symptom standard library to obtain the stroke symptoms and their onset time;
[0102] Onset time screening module 7: for screening multiple onset times according to the symptom time category and the preset stroke onset time rule if there are multiple onset times corresponding to the same onset symptom, so as to determine the onset time required for stroke disease staging;
[0103] The staging judgment module 8 is used to judge the stroke symptoms and their corresponding onset time according to the preset stroke disease staging rules, and obtain the patient's stroke disease staging result.
[0104] Embodiment three:
[0105] This embodiment provides a storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a main controller, each step in the above method is implemented;
[0106] It is understandable that the storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0107] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.
[0108] It should be noted that, in the description of the present invention, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "plurality" refers to at least two.
[0109] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention belong.
[0110] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0111] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0112] In addition, each functional unit in each embodiment of the present invention may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0113] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0114] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0115] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. An intelligent stroke staging method, characterized in that: The method comprises: Obtain patient complaints, history of present illness, and diagnostic data from the electronic medical record system; Preprocessing the chief complaint, current medical history, and diagnostic data; Inputting the preprocessed chief complaint, current medical history and diagnosis data into a pretrained disease classification model, if the output of the pretrained disease classification model is cerebral hemorrhage or cerebral infarction; Inputting the patient's chief complaint data into a pre-trained temporal entity recognition model, the pre-trained temporal entity recognition model outputs a symptom description and its corresponding onset time and symptom time category triplet; Inputting the symptom description and its corresponding onset time and symptom time category triplet into a pre-trained symptom standardization matching model, and the pre-trained symptom standardization matching model outputs standardized symptom description, onset time and symptom time category according to a preset standard output format; Matching and screening the standardized symptom descriptions in a preset stroke symptom standard library to obtain stroke symptoms and their onset time and symptom time categories; If there are multiple onset times corresponding to the same symptom, the multiple onset times are screened according to the symptom time category and the preset stroke onset time rule to determine the onset time required for the stroke disease staging; The stroke symptoms and their corresponding onset time are judged by using the preset stroke disease staging rules to obtain the patient's stroke disease staging result.
2. The method according to claim 1, characterized in that The pre-processing of the chief complaint, current medical history and diagnosis data includes: Data cleaning is performed on the chief complaint, current medical history, and diagnosis data to remove special characters and non-text content; After data cleaning, the chief complaint, current medical history, and diagnosis data were segmented; Stop words were removed from the chief complaint, current medical history, and diagnosis data after word segmentation.
3. The method according to claim 2, characterized in that The preset stroke disease staging rules include: Stroke diseases are divided into ultra-early stage, acute stage, recovery stage and sequelae stage, and the corresponding onset time is set for ultra-early stage, acute stage, recovery stage and sequelae stage respectively; The method of obtaining the patient's stroke disease staging results includes: The patient's stroke symptoms and their corresponding time are matched with the onset time corresponding to each period of the stroke disease to obtain the patient's stroke disease stage.
4. The method according to claim 3, characterized in that: The training of the pre-trained disease classification model includes: Build a customized dual-tower cross-attention stroke disease classification model framework based on the BioClinicalBERT pre-trained model, and set the loss function of the model framework; obtaining chief complaints, present medical history, and diagnosis data of a plurality of patients from the electronic medical record system; Preprocessing the chief complaints, current medical histories, and diagnostic data of the multiple patients to obtain training data; The training data is used as the input of the customized dual-tower cross-attention stroke disease classification model framework, and the customized dual-tower cross-attention stroke disease classification model framework is iteratively trained until the loss value of the loss function no longer decreases or reaches a preset number of iterations, thereby obtaining a pre-trained disease classification model.
5. The method according to claim 4, characterized in that The loss function of the customized dual-tower cross-attention stroke disease classification model framework adopts a smooth L1 loss function; The expression of the smooth L1 loss function is: Where L1(y), y * ) represents the loss value; y represents the true label; y * Represents the predicted value; δ represents the balance parameter, which is used to control the balance between square loss and absolute loss.
6. The method according to claim 5, characterized in that The training of the pre-trained temporal entity recognition model includes: Build a temporal entity recognition model framework based on the ChatGLM large language model, set the loss function of the model framework, and use the cross entropy loss function; obtaining chief complaint information of a plurality of patients from the electronic medical record system; Preprocessing the chief complaint information of the plurality of patients to obtain training data; The training data is used as the input of the temporal entity recognition model framework, and the temporal entity recognition model framework is iteratively trained until the loss value of the loss function no longer decreases or reaches a preset iteration round, thereby obtaining a pre-trained temporal entity recognition model.
7. Intelligent stroke staging device, characterized in that: The device comprises: Data acquisition module: used to obtain the patient's chief complaint, current medical history and diagnosis data from the electronic medical record system; Preprocessing module: used for preprocessing the chief complaint, current medical history and diagnostic data; Disease identification module: used to input the pre-processed chief complaint, current medical history and diagnosis data into the pre-trained disease classification model, if the output of the pre-trained disease classification model is cerebral hemorrhage or cerebral infarction; Symptom time acquisition module: used for inputting the patient's main complaint data into a pre-trained time entity recognition model, and the pre-trained time entity recognition model outputs a symptom description and its corresponding onset time and symptom time category triplet; Symptom description standardization module: used for inputting the symptom description and its corresponding onset time and symptom time category triplet into a pre-trained symptom standardization matching model, and the pre-trained symptom standardization matching model outputs standardized symptom description, onset time and symptom time category according to a preset standard output format; A stroke symptom time acquisition module is used to match and screen the standardized symptom description in a preset stroke symptom standard library to obtain stroke symptoms and their onset time and symptom time category; Onset time screening module: if there are multiple onset times corresponding to the same onset symptom, the multiple onset times are screened according to the symptom time category and the preset stroke onset time rule to determine the onset time required for stroke disease staging; Staging judgment module: used to judge the stroke symptoms and their corresponding onset time by using preset stroke disease staging rules to obtain the patient's stroke disease staging results.
8. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the main controller, each step of the intelligent stroke staging method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Device and system for identifying early cerebral infarction and cerebral hemorrhage
CN115120223A
Method for identifying early cerebral infarction and cerebral hemorrhage
CN115137338A