Gastrointestinal bleeding risk assessment system based on deep learning

By extracting and analyzing keywords and non-keywords from the text data of patients with gastrointestinal bleeding, and combining aspirin dosing intervals and semantic features, the attention weights of the neural network model are adjusted, thus solving the problem of low accuracy in gastrointestinal bleeding risk assessment in existing technologies and achieving higher assessment precision.

CN120878239BActive Publication Date: 2025-12-09ORDNANCE IND HYGIENIC INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511367003.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2025-12-09
Estimated Expiration
2045-09-24

AI Technical Summary

Technical Problem

Existing technologies, when using deep learning methods to combine semantic information of patients with gastrointestinal bleeding for risk assessment, fail to effectively consider the changes between the semantic information of keywords and text data, resulting in low accuracy of risk assessment.

Method used

The data acquisition and preprocessing module extracts keywords and non-keywords from the text data. Combining the variation characteristics of the text data and the aspirin dosing interval, the formal risk feature value is determined. The semantic risk feature value is determined by word vector analysis and hierarchical clustering. The attention weight is used to adjust the output of the gastrointestinal bleeding risk assessment coefficient of the neural network model.

Benefits of technology

It improves the accuracy of gastrointestinal bleeding risk assessment, reduces noise interference, and enhances the precision of prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120878239B_ABST
    Figure CN120878239B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of medical data mining, in particular to a digestive tract hemorrhage risk assessment system based on deep learning. Firstly, non-structural text data is preprocessed to determine key words and non-key words; then, in combination with risk features represented by text data changes and the influence of aspirin on hemorrhage risk, form risk feature values are determined in the dimension of text form; then, in combination with semantic information of the text data and the text semantics of each risk link corresponding to the digestive tract risk, correlation analysis is carried out, semantic risk feature values are determined in the dimension of semantic information; finally, in combination with the form risk feature values and the semantic risk feature values, the attention weight of each text evaluation dimension is determined, and a neural network model is adjusted according to the attention weight, so that the accuracy of the output digestive tract hemorrhage risk assessment coefficient is higher.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical data mining, in particular to a gastrointestinal bleeding risk assessment system based on deep learning. BACKGROUND

[0002] In the prior art, when gastrointestinal bleeding risk assessment is performed by combining the semantic information of gastrointestinal bleeding patients with a deep learning method, a vector composed of the frequency of occurrence of extracted keywords related to gastrointestinal bleeding is directly input into a neural network model, and a gastrointestinal bleeding risk assessment coefficient is output. However, the prior art does not consider the semantic information of the keywords themselves and the gastrointestinal bleeding risk exhibited by the changes between text data when performing gastrointestinal bleeding risk assessment, resulting in low accuracy of the gastrointestinal bleeding risk assessment coefficient output by the prior art after directly inputting the vector composed of the frequency of occurrence of extracted keywords related to gastrointestinal bleeding into the neural network model. SUMMARY

[0003] In order to solve the technical problem of low accuracy of the gastrointestinal bleeding risk assessment coefficient output by the prior art when combining the semantic information of gastrointestinal bleeding patients with a deep learning method, the purpose of the present application is to provide a gastrointestinal bleeding risk assessment system based on deep learning, and the technical solution adopted is as follows:

[0004] The first aspect of the present application provides a gastrointestinal bleeding risk assessment system based on deep learning, comprising:

[0005] A data acquisition and preprocessing module is configured to acquire each text data of a gastrointestinal bleeding patient in each text evaluation dimension in a medical database, and extract keywords and all non-keyword tokens in the text data.

[0006] A risk feature value determination module is configured to determine a formal risk feature value of each text evaluation dimension according to the recording time interval between the current text data and the previous text data, the aspirin taking interval of the current text data, the number of non-keyword tokens in the sentence containing the keyword, and the change stability of various keywords in each text evaluation dimension; and determine a semantic risk feature value of each text evaluation dimension according to the correlation between the overall vector feature of the keyword vector in each text evaluation dimension and the preset reference vector of each risk link.

[0007] A gastrointestinal bleeding risk assessment module is configured to determine an attention weight of each text evaluation dimension according to the formal risk feature value and the semantic risk feature value, adjust a neural network model according to the attention weight, and output a gastrointestinal bleeding risk assessment coefficient of the gastrointestinal bleeding patient at the current time.

[0008] Further, the process of extracting the keywords and all non-keyword words in the text data comprises:

[0009] Each piece of text data is input into the BioBERT model for word segmentation and keyword extraction, and the corresponding keywords and non-keyword words of gastrointestinal bleeding are output; the non-keyword words are other words besides the keywords.

[0010] Further, the process of obtaining the formal risk feature value comprises:

[0011] Under each text evaluation dimension, each keyword in the current text data is sequentially taken as a target word.

[0012] According to the average number of non-keyword words in all sentences containing the target word in each piece of text data, the description word quantity of the target word in each piece of text data is determined; a set composed of all kinds of keywords in each piece of text data is taken as a corresponding keyword set; according to the coincidence of the keyword set and the relative reduction of the description word quantity between the current text data and the last piece of text data, the symptom description coefficient of the target word in the current text data is determined.

[0013] The time interval length between the recording time of the current text data and the recording time of the previous text data is negatively correlated to determine the time interval coefficient of the current text data.

[0014] According to the time interval between the recording time of the current text data and the latest aspirin medication date of the gastrointestinal bleeding patient, the risk development coefficient of the current text data is determined.

[0015] According to the product of the risk development coefficient, the time interval coefficient and the symptom description coefficient, the weighted structural risk feature value of the target word in the current text data is determined.

[0016] Under each text evaluation dimension, according to the average of the weighted structural risk feature values of all kinds of keywords in the current text data, the corresponding formal risk feature value is determined.

[0017] Further, the process of obtaining the symptom description coefficient comprises:

[0018] According to the difference between the description word quantity of the target word in the last piece of text data of the current text data and the description word quantity of the target word in the current text data, the description word reduction quantity of the target word in the current text data is determined.

[0019] determine a reference set according to an intersection between a keyword set of the current text data and a keyword set of previous text data corresponding to the keyword set of the current text data; and determine a keyword combination stability of the current text data according to a ratio between a number of keyword categories in the reference set and a number of keyword categories in the keyword set of the current text data;

[0020] positively map a product between the reduced number of descriptors and the keyword combination stability to determine a symptom description coefficient of a target word in the current text data.

[0021] Further, the obtaining process of the risk development coefficient comprises:

[0022] take each time when the patient with gastrointestinal bleeding takes aspirin as a taking time; and take a time interval length between a recording time of the current text data and each taking time as a corresponding determination time length;

[0023] when the determination time length is less than or equal to a preset first time threshold, take a normalized value of the determination time length as a reference development coefficient of the corresponding taking time;

[0024] when the determination time length is greater than the preset first time threshold and less than or equal to a preset second time threshold, take a preset highest risk feature value as the reference development coefficient of the corresponding taking time;

[0025] when the determination time length is greater than the preset second time threshold, take a negative correlation mapping value of the determination time length as the reference development coefficient of the corresponding taking time;

[0026] take a maximum value of the reference development coefficients of all the taking times as the risk development coefficient of the current text data.

[0027] Further, the obtaining process of the semantic risk feature value comprises:

[0028] under each text evaluation dimension, convert all the keywords in the current text data into word vectors by using word2vec to determine a word vector of each keyword;

[0029] determine a corresponding semantic similarity according to a cosine similarity between the word vector of each keyword and the word vector of each other keyword; and perform hierarchical clustering on the current text data by using a negative correlation mapping value of the semantic similarity between keywords as a clustering distance to obtain at least two keyword clustering clusters;

[0030] determine a center vector according to a mean vector of the word vectors of all the keywords in each keyword clustering cluster; and determine a risk score according to a mean value of all the semantic similarities between the keywords in each keyword clustering cluster.

[0031] According to the vector similarity distribution between the center vector and the preset reference word vector of each risk link, the risk association strength of each keyword clustering cluster is determined;

[0032] According to the product between the risk association strength and the risk score, the local risk characteristic value of each keyword clustering cluster is determined; and according to the mean value of the local risk characteristic values of all keyword clustering clusters, the semantic risk characteristic value of each text evaluation dimension is determined.

[0033] Further, the risk association strength acquisition process comprises:

[0034] The cosine similarity between the center vector and the preset reference word vector of each risk link is calculated to determine the link matching degree between each keyword clustering cluster and each risk link; and according to the maximum value of the link matching degrees between each keyword clustering cluster and all risk links, the corresponding risk association strength is determined.

[0035] Further, the attention weight acquisition process comprises:

[0036] According to the product between the formal risk characteristic value and the semantic risk characteristic value, the attention weight of each text evaluation dimension is determined.

[0037] Further, the gastrointestinal bleeding risk assessment coefficient acquisition process comprises:

[0038] In each text evaluation dimension, the occurrence frequency of each keyword in the current text data is taken as the corresponding frequency characteristic value; the frequency characteristic values of all keywords in the current text data are arranged in the order of occurrence of all keywords in all text data to determine the feature vector of the current text data in each text evaluation dimension;

[0039] The product between the feature vector of the current text data in each text evaluation dimension and the corresponding attention weight is taken as a weighted vector; and after the weighted vectors of all text evaluation dimensions are spliced, a joint feature vector is determined;

[0040] The joint feature vector is input into the trained multi-layer fully connected neural network to output a gastrointestinal bleeding risk assessment index.

[0041] Further, after outputting the gastrointestinal bleeding risk assessment coefficient of the gastrointestinal bleeding patient at the current time, the method further comprises:

[0042] When the gastrointestinal bleeding risk assessment index is greater than a preset risk threshold, a bleeding warning is issued.

[0043] When the digestive tract bleeding risk assessment is less than or equal to a preset risk threshold, no bleeding warning is issued.

[0044] In a second aspect, the present application provides a computer device, comprising a memory and a processor. The memory is configured to store computer program code, and the processor is configured to call and run the computer program code from the memory to execute the system as claimed in the first aspect or any embodiment of the first aspect of the present application.

[0045] In a third aspect, the present application provides a computer program product, comprising computer program code, which, when executed, performs the system as claimed in the first aspect or any embodiment of the first aspect of the present application.

[0046] In a fourth aspect, the present application provides a computer-readable storage medium, which stores computer program code, which, when executed, performs the system as claimed in the first aspect or any embodiment of the first aspect of the present application.

[0047] The present application has the following beneficial effects:

[0048] Firstly, the present application preprocesses the non-structured text data to determine the key words and non-key words therein; then, in combination with the risk features represented by the text data changes and the influence of aspirin on the bleeding risk, the formal risk feature values are determined in the dimension of the text form; then, in combination with the semantic information of the text data and the text semantics of each risk link corresponding to the digestive tract risk, the relevance analysis is performed, and the semantic risk feature values are determined in the dimension of the semantic information; finally, in combination with the formal risk feature values and the semantic risk feature values, the attention weights of each text evaluation dimension are determined, and the neural network model is adjusted according to the attention weights, so that the accuracy of the output digestive tract bleeding risk assessment coefficient is higher; the present application further improves the influence of the key modal data on the final prediction result, reduces the noise interference of the weak modal data, and further improves the accuracy of the prediction result. BRIEF DESCRIPTION OF DRAWINGS

[0049] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0050] Figure 1 The structural diagram of a digestive tract bleeding risk assessment system based on deep learning provided by an embodiment of the present application;

[0051] Figure 2 A computer device structure schematic diagram provided by an embodiment of the present application. DETAILED DESCRIPTION

[0052] In order to further illustrate the technical means and effects taken by the present application to achieve the predetermined object of the application, the following describes in detail the specific implementation, structure, features and effects of a digestive tract bleeding risk assessment system based on deep learning according to the present application, in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment, and the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form. In addition, the terms "first", "second" are only used for descriptive purposes, and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Therefore, the features with "first", "second" can be explicitly or implicitly included one or more features.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.

[0054] The specific scheme of the digestive tract bleeding risk assessment system based on deep learning provided by the present application is specifically described below in combination with the accompanying drawings.

[0055] The present application provides a digestive tract bleeding risk assessment system based on deep learning. Please refer to Figure 1 , which shows the structure diagram of a digestive tract bleeding risk assessment system based on deep learning provided by an embodiment of the present application. The system comprises a data acquisition and preprocessing module 101, a risk feature value determination module 102 and a digestive tract bleeding risk assessment module 103.

[0056] The data acquisition and preprocessing module 101 is used to acquire each text data of the digestive tract bleeding patient in each text evaluation dimension in the medical database; extract the keywords and all non-key words in the text data.

[0057] In a specific implementation manner of the embodiment of the present application, the text evaluation dimensions include clinical documents such as chief complaint, physical examination report, doctor's evaluation record, discharge record and medical record, and the implementer can adjust the text evaluation dimensions according to the specific implementation environment; in the medical database, each text data of each text evaluation dimension of the patient with gastrointestinal bleeding is collected at each record time and the corresponding record time is marked; the current text data in the embodiment of the present application is the text data whose record time is closest to the current time. Further, the embodiment of the present application determines the attention weight according to the risk characteristics embodied by the text data, so that the neural network model after the attention weight adjustment can improve the influence of the key modal data on the final prediction result, reduce the noise interference of the weak modal data, further improve the accuracy of the prediction result, and make the accuracy of the finally obtained gastrointestinal bleeding risk evaluation coefficient higher.

[0058] And the analysis of the text data usually needs to extract the keywords and various word segmentation in the text data based on the idea of natural language processing method. Preferably, in some possible implementation manners of the embodiment of the present application, the process of extracting the keywords and all non-key word segmentation in the text data includes:

[0059] Each text data is input into the BioBERT model for word segmentation and keyword extraction, and the corresponding keywords and non-key word segmentation of gastrointestinal bleeding are output; the non-key word segmentation is other word segmentation except the keywords.

[0060] Before inputting the text data into the BioBERT model, the BioBERT model needs to be fine-tuned, specifically: the historical text data used for training is taken as a fine-tuning dataset, and is divided into a training set and a validation set according to 7:3; the dataset label system is artificially annotated, including drug, pathological entity, organ part, pathological diagnosis opinion and other entities or diagnosis and treatment conclusions, and the entities are annotated in BIOES format; the dmis-lab / biobert-base-cased-v1.1 provided in HuggingFace Transformers is used as a pre-training model; a small amount of fine-tuning is performed on the basis model of BioBERT, and the specific fine-tuning settings are as follows: the task type is set to sequence labeling (Named Entity Recognition); the input format is set to [CLS]+token1+token2+…+[SEP]; the output label is the BIO label corresponding to each token; the loss function is defined as cross-entropy loss, the optimizer AdamW is set, the fine-tuning epoch is 3-5 rounds, and the batch size can be set to 16-32. The fine-tuned BioBERT model is used to fine-tune the training set, and the F1-score is monitored on the validation set to stop early, so as to realize the fine-tuning of the BioBERT model. Each piece of text data of the current patient is input into the fine-tuned BioBERT model, and the key words of the gastrointestinal bleeding are extracted, and the formatted JSON object of the key words is output, so as to determine the key words and all non-key words in each piece of text data.

[0061] In other possible implementations of the embodiments of the present application, the jieba segmentation method can also be used to segment each piece of text data, and the Latent Dirichlet Allocation (LDA) model can be used to extract key words, and the TF-IDF algorithm can also be used to extract key words, wherein the LDA model, the TF-IDF algorithm and the jieba segmentation method are technical means known to those skilled in the art, and will not be further limited and described.

[0062] The risk feature value determination module 102 is configured to determine a formal risk feature value of each text evaluation dimension according to a record time interval between the current text data and the last piece of text data, an aspirin taking interval of the current text data, a non-key word quantity change in a sentence containing the key words, and a change stability of various key words; and determine a semantic risk feature value of each text evaluation dimension according to an association between a word vector overall vector feature of the key words in each text evaluation dimension and a preset reference word vector of each risk link.

[0063] After extracting the keywords from the text data, before analyzing the semantic features of each keyword in the text data, it needs to be considered that the writing form or writing mode of the clinical document in the medical data itself contains rich information. When recording the text information of the patient with gastrointestinal bleeding, even if the terms such as "black stool" or "vomiting blood" are not directly mentioned, the writing behavior of the doctor will have systematic differences due to the severity of the disease. The difference is reflected in the surface features of the text: when facing a more serious case of gastrointestinal bleeding, the cognitive resources of the doctor are mainly concentrated on the rescue of the patient, which leads to a specific mode of writing of the text data: the length of the descriptive sentence is shortened (the continuous appearance of a certain physiological risk represented by the keyword indicates that this risk has not been cured during the treatment process, and in the subsequent record, the change compared with the last time is generally recorded instead of describing the risk represented by the keyword again), the text data recording interval is short (reflecting the rapid change of the disease).

[0064] In the medical text, the keywords do not exist alone. When multiple same keywords exist in two continuous text data, that is, the combination mode of the keywords is the same, it often indicates that multiple concurrent factors are linked, and the risk degree is greatly increased. Therefore, when evaluating the description features of the keywords, the combination co-occurrence features of the keywords should also be combined to further improve the accuracy of feature evaluation. In addition, considering that the risk of gastrointestinal bleeding caused by the time interval after taking aspirin is different, the non-semantic bleeding risk feature evaluation can be combined with the aspirin taking interval of the current text data. Therefore, the embodiment of the present application further determines the formal risk feature value of each text evaluation dimension according to the recording time interval between the current text data and the last text data, the aspirin taking interval of the current text data, the number of non-keyword word changes in the sentence where the keyword is located, and the change stability of various keywords under each text evaluation dimension. The greater the formal risk feature value is, the more significant the bleeding risk feature corresponding to the non-semantic feature is, and then the reference value of the data of the corresponding text evaluation dimension should be higher.

[0065] Preferably, in some possible implementation manners of the embodiment of the present application, the acquisition process of the formal risk feature value comprises:

[0066] Under each text evaluation dimension, each keyword in the current text data is sequentially taken as a target word;

[0067] The number of descriptive words for the target word in each text dataset is determined by the average number of non-keyword segments in all sentences containing the target word. The set of all types of keywords in each text dataset is taken as the corresponding keyword set. The symptom description coefficient of the target word in the current text dataset is determined based on the overlap of the keyword sets between the current and previous text datasets and the relative decrease in the number of descriptive words. In a specific implementation of this invention, the process of obtaining the symptom description coefficient includes:

[0068] Based on the difference between the number of descriptive words for the target word in the previous text data and the number of descriptive words for the target word in the current text data, determine the reduction in descriptive words for the target word in the current text data; based on the intersection between the keyword set of the current text data and the keyword set of its corresponding previous text data, determine the reference set; based on the ratio between the number of keyword types in the reference set and the number of keyword types in the keyword set of the current text data, determine the stability of the keyword combination in the current text data.

[0069] Firstly, when gastrointestinal bleeding is severe, doctors will focus their cognitive resources on patient resuscitation, resulting in shorter descriptive sentences. Therefore, the greater the reduction in descriptive words, that is, the greater the reduction in descriptive words for the target word in the current text data, the higher the risk of gastrointestinal bleeding. In addition, the more similar the keyword combinations between the current text data and the previous text data, that is, the closer the keyword set of the reference set is to the current text data, the more it corresponds to the situation where multiple concurrent factors lead to a greater increase in risk, and the higher the corresponding risk of gastrointestinal bleeding.

[0070] Therefore, a positive correlation mapping is further performed between the product of the reduction in descriptive words and the stability of keyword combinations to determine the symptom description coefficient of the target words in the current text data. This ensures that the larger the symptom description coefficient, the higher the risk of gastrointestinal bleeding reflected by the corresponding text evaluation dimension, and the more attention should be paid to the corresponding text evaluation dimension.

[0071] In one specific implementation of this invention, the process of obtaining the symptom description coefficient is expressed by the following formula: ;in, For the first Target words in the current text data of each text evaluation dimension Symptom description coefficient; For the first The target words in the previous text data of the current text data for each text evaluation dimension The number of descriptive words and target words in the current text data the difference between the number of the adjectives in the current text data and the number of the adjectives in the target text data, i.e., the target word the number of the adjectives in the current text data; the number of the adjectives in the current text data; the number of the adjectives in the intersection of the keyword set of the current text data and the keyword set of the previous text data in the first text evaluation dimension; the number of the adjectives in the keyword set of the current text data in the first text evaluation dimension; the number of the adjectives in the keyword set of the current text data in the first text evaluation dimension; the keyword combination stability of the current text data in the first text evaluation dimension; the keyword combination stability of the current text data in the first text evaluation dimension; the keyword combination stability of the current text data in the first text evaluation dimension; the function is normalized, and other normalization functions can be used instead, such as a linear normalization function, which will not be described further herein.

[0072] It should be noted that, in order to ensure that the calculation result is meaningful, when performing fractional operation, if the denominator is 0, a parameter adjustment factor greater than 0 is added to the denominator to prevent the denominator from being 0. The value of the parameter adjustment factor is set by the implementer according to the actual situation, and the present application is set to 0.1.

[0073] The time interval length between the recording time of the current text data and the recording time of the previous text data is negatively correlated to determine the time interval coefficient of the current text data. The smaller the time interval between the recording times of adjacent text data, the shorter the text data recording interval, the faster the disease changes. Therefore, the smaller the time interval coefficient of the current text data in the corresponding text evaluation dimension, the higher the risk of gastrointestinal bleeding at the current time, and the more attention needs to be paid to the corresponding text evaluation dimension.

[0074] In one specific implementation of the present application, the method of negatively correlating the time interval length between the recording time of the current text data and the recording time of the previous text data adopts: normalizing the time interval length between the recording time of the current text data and the recording time of the previous text data to determine the reference interval length; the difference between the real number 1 and the reference interval length is taken as the negative correlation mapping result, i.e., the time interval coefficient of the current text data.

[0075] Further considering the influence of aspirin medication, the risk development coefficient of the current text data is determined according to the time interval between the recording time of the current text data and the latest aspirin medication date of the gastrointestinal bleeding patient. In one specific implementation of the present application, the process of obtaining the risk development coefficient includes:

[0076] The moment of taking aspirin each time by the patient with gastrointestinal bleeding is taken as the taking moment; the length of the time interval between the recording moment of the current text data and each taking moment is taken as the corresponding determination time length; when the determination time length is less than or equal to a preset first time length threshold, the normalized value of the determination time length is taken as the reference development coefficient of the corresponding taking moment; when the determination time length is greater than the preset first time length threshold and less than or equal to a preset second time length threshold, a preset highest risk feature value is taken as the reference development coefficient of the corresponding taking moment; when the determination time length is greater than the preset second time length threshold, a negative correlation mapping value of the determination time length is taken as the reference development coefficient of the corresponding taking moment; and the maximum value of the reference development coefficients of all taking moments is taken as the risk development coefficient of the current text data.

[0077] In a specific implementation manner of the embodiment of the present application, the negative correlation mapping method of the determination time length adopts: the reference difference is determined by normalizing the difference between the determination time length and the preset second time length threshold; and the difference between the real number 1 and the reference difference is taken as the negative correlation mapping result of the determination time length. It should be noted that, except for special description, the normalization method in the embodiment of the present application adopts linear normalization, which will not be described further here.

[0078] In a specific implementation manner of the embodiment of the present application, the first time length threshold is set to 31 days, the second time length threshold is set to 90 days, and the preset highest risk feature value is set to 1. According to the paper corresponding to the prediction of non-variceal upper gastrointestinal bleeding of the patient taking enteric-coated aspirin by the aspirin risk score, the high incidence stage of gastrointestinal bleeding is within one year of taking aspirin, and the relative risk of upper gastrointestinal bleeding is the highest when taking aspirin for 31 to 90 days; the effect has time accumulation, and the bleeding risk is in the rising period but not at the peak within 0 to 30 days, and the risk decreases compared with the peak stage within 90 days to 1 year due to the adaptation of some patients. Therefore, according to the time period in which each determination time length is located, the maximum value is selected as the risk development coefficient, so that the greater the risk development coefficient, the higher the gastrointestinal bleeding risk reflected by the current text data in the corresponding text evaluation dimension, and the more attention needs to be paid to the corresponding text evaluation dimension.

[0079] Finally, according to the correlation, the weighted structural risk feature value of the target word in the current text data is determined according to the product of the risk development coefficient, the time interval coefficient and the symptom description coefficient. It should be noted that, in addition to the product, the implementer can also determine the weighted structural risk feature value according to the specific implementation environment, for example, normalizing the sum value between the risk development coefficient, the time interval coefficient and the symptom description coefficient to determine the weighted structural risk feature value; so that the greater the weighted structural risk feature value, the more attention needs to be paid to the text evaluation dimension in which the target word is located.

[0080] Since the current text data under each text evaluation dimension typically corresponds to multiple keywords, further, under each text evaluation dimension, the corresponding formal risk feature value is determined based on the mean of the weighted structural risk feature values ​​of all keywords in the current text data. This ensures that the larger the formal risk feature value, the higher the attention should be paid to the corresponding text evaluation dimension at the non-semantic level, and a greater attention weight should be given to the features of the text evaluation dimension to improve the accuracy of the gastrointestinal bleeding risk assessment coefficient output by the neural network model.

[0081] In one specific implementation of this invention, the process of obtaining the formal risk feature value is expressed by the following formula: ;in, For the first Formal risk feature values ​​for each text assessment dimension; For the first The number of keyword categories in the current text data for each text evaluation dimension; For the first In the current text data of the text evaluation dimension, the first... Symptom description coefficients for various keywords; For the first The time interval coefficient of the current text data under each text evaluation dimension; For the first Risk development coefficient of current text data under each text evaluation dimension; For the first In the current text data of the text evaluation dimension, the first... Weighted structural risk eigenvalues ​​for various keywords.

[0082] Further analysis requires considering the semantic features of each text assessment dimension. Stronger semantic associations among keywords within a single dimension indicate more consistent risk signals and a higher likelihood of corresponding gastrointestinal bleeding risk. For example, the strong association between deep ulcers and vascular exposure in PACS clearly points to severe mucosal damage. Furthermore, for the current text data of each assessment dimension, the more similar the semantic features of the corresponding keywords are to the semantic features of each risk factor representing bleeding risk, the more significant the bleeding risk characteristic of the corresponding text assessment data at the current moment, and the more attention should be paid to that text assessment dimension. Therefore, based on the correlation between the overall word vector features of keywords in each text assessment dimension and the preset benchmark word vectors of each risk factor, the semantic risk feature value of each text assessment dimension is determined. A higher semantic risk feature value indicates a higher gastrointestinal bleeding risk reflected by the corresponding text assessment dimension, and thus, a greater focus on that text assessment dimension at the semantic analysis level.

[0083] Preferably, in some possible implementation manners of the embodiments of the present application, the process of obtaining the semantic risk feature value comprises:

[0084] Under each text evaluation dimension, all keywords in the current text data are converted into word vector form by word2vec, and the word vector of each keyword is determined; it should be noted that word2vec is a word vector conversion means known to those skilled in the art, and will not be further limited and described here.

[0085] According to the cosine similarity between the word vector of each keyword and the word vector of each other keyword, the corresponding semantic similarity is determined; in the current text data, the negative correlation mapping value of the semantic similarity between keywords is taken as the clustering distance for hierarchical clustering, and at least two keyword clustering clusters are obtained; in one specific implementation manner of the embodiments of the present application, the reciprocal of the semantic similarity is taken as the clustering distance for hierarchical clustering. It should be noted that hierarchical clustering is a technical means known to those skilled in the art, and will not be further described here. By hierarchical clustering, keywords with similar word vectors, i.e., similar semantics, are clustered into the same keyword clustering cluster; under each text evaluation dimension, the higher the semantic similarity between the keywords in the corresponding keyword clustering cluster, the more it conforms to the feature that the semantic correlation between keywords in a single dimension is stronger, and the higher the risk of digestive tract bleeding reflected by the corresponding text evaluation dimension, so the average of all semantic similarities between all keywords in each keyword clustering cluster is further determined to determine the corresponding risk score; so that the greater the risk score, the more attention needs to be paid to the corresponding text evaluation dimension.

[0086] Further, the average vector of the word vectors of all keywords in each keyword clustering cluster is determined to determine the corresponding center vector; according to the vector similarity distribution between the center vector and the preset reference word vector of each risk link, the risk association strength of each keyword clustering cluster is determined; the process of obtaining the risk association strength comprises: calculating the cosine similarity between the center vector and the preset reference word vector of each risk link, and determining the link matching degree between each keyword clustering cluster and each risk link.

[0087] The central vector represents the semantic features of each keyword cluster as a whole, and each risk link can directly or indirectly reflect the state of gastrointestinal bleeding. In a specific implementation of this invention, the types of risk links include four links: mucosal injury, drug-driven, symptom warning, and protective effect. Among them, the keywords of the mucosal injury link include "ulcer", "erosion", and "vascular exposure"; the keywords of the drug-driven link include "NSAIDs combination" and "high-dose aspirin"; the keywords of the symptom warning link include "melena", "hematemesis", and "anemia"; and the keywords of the protective effect include "regularity". "PPI" and "Hp eradication"; and in this embodiment of the invention, the preset benchmark word vector for each risk link is the average vector of all word vectors corresponding to all keywords in each risk link. The preset benchmark word vector obtained by the average vector represents the overall semantic features of each risk link. Therefore, for each central vector, the smaller the cosine similarity between it and a certain preset benchmark word vector, the more the semantic features of the overall keywords of the keyword cluster of the central vector are consistent with the semantic features of the keywords of the risk link corresponding to the preset benchmark word vector. That is, the more likely the corresponding keyword cluster is to correspond to the risk features of gastrointestinal bleeding.

[0088] Therefore, the risk association strength is further determined based on the maximum value of the link matching degree between each keyword cluster and all risk links. The greater the risk association strength, the more likely the corresponding keyword cluster is to correspond to the state of gastrointestinal bleeding, and the higher the risk of gastrointestinal bleeding reflected. Therefore, the corresponding text evaluation dimension needs to be paid more attention to.

[0089] Finally, based on the correlation, the local risk feature value of each keyword cluster is determined by multiplying the risk association strength with the risk score; the semantic risk feature value of each text evaluation dimension is determined by the mean of the local risk feature values ​​of all keyword clusters; the larger the obtained semantic risk feature value, the higher the risk of gastrointestinal bleeding reflected by the corresponding text evaluation dimension, and the more attention needs to be paid to the corresponding text evaluation dimension at the semantic analysis level.

[0090] In one specific implementation of this invention, the process of obtaining semantic risk feature values ​​includes: ;in, For the first Semantic risk feature values ​​for each text evaluation dimension; For the first The number of keyword clusters in the current text data for each text evaluation dimension; For the first In the current text data of the text evaluation dimension, the first... a mean value of all semantic similarities between all keywords in a keyword clustering cluster, i.e. a corresponding risk score; is the maximum value of the link matching degree between the i-th keyword clustering cluster in the current text data of the j-th text evaluation dimension and all risk links, i.e. the corresponding risk association strength; is the maximum value of the link matching degree between the i-th keyword clustering cluster in the current text data of the j-th text evaluation dimension and all risk links, i.e. the corresponding risk association strength; is the maximum value of the link matching degree between the i-th keyword clustering cluster in the current text data of the j-th text evaluation dimension and all risk links, i.e. the corresponding risk association strength; is the maximum value of the link matching degree between the i-th keyword clustering cluster in the current text data of the j-th text evaluation dimension and all risk links, i.e. the corresponding risk association strength; is the maximum value of the link matching degree between the i-th keyword clustering cluster in the current text data of the j-th text evaluation dimension and all risk links, i.e. the corresponding risk association strength; is the maximum value of the link matching degree between the i-th keyword clustering cluster in the current text data of the j-th text evaluation dimension and all risk links, i.e. the corresponding risk association strength.

[0091] The gastrointestinal bleeding risk assessment module 103 is configured to determine the attention weight of each text evaluation dimension according to the formal risk feature value and the semantic risk feature value, and adjust the neural network model according to the attention weight to output the gastrointestinal bleeding risk assessment coefficient of the gastrointestinal bleeding patient at the current time.

[0092] Finally, based on the dimension attention demand respectively represented by the formal risk feature value and the semantic risk feature value of each text evaluation dimension, the attention weight of each text evaluation dimension is determined according to the formal risk feature value and the semantic risk feature value, so that when the neural network model is adjusted according to the attention weight of each text evaluation dimension, the text evaluation dimension with a larger attention weight has a greater influence on the output result.

[0093] Preferably, in a specific implementation manner of the embodiment of the present application, the attention weight of each text evaluation dimension is determined according to the product between the formal risk feature value and the semantic risk feature value, so that the greater the attention weight is, the greater the influence of the corresponding text evaluation dimension on the output of the neural network model is. The process of obtaining the attention weight is represented by a formula as follows: ; wherein, is the attention weight of the j-th text evaluation dimension; is the attention weight of the j-th text evaluation dimension; is the semantic risk feature value of the j-th text evaluation dimension; is the semantic risk feature value of the j-th text evaluation dimension; is the formal risk feature value of the j-th text evaluation dimension. is the formal risk feature value of the j-th text evaluation dimension.

[0094] Finally, the gastrointestinal bleeding risk assessment coefficient of the gastrointestinal bleeding patient at the current time is output by adjusting the neural network model according to the attention weight of each text evaluation dimension. Preferably, in some possible implementation manners of the embodiment of the present application, the process of obtaining the gastrointestinal bleeding risk assessment coefficient includes:

[0095] In each text evaluation dimension, the frequency of occurrence of each keyword in the current text data is taken as the corresponding frequency feature value; the frequency feature values of all keywords in the current text data are arranged in the order of occurrence of all keywords in all text data to determine the feature vector of the current text data in each text evaluation dimension; wherein the frequency feature value of the keyword not present in the current text data is set to 0 and exists in the corresponding feature vector. The product of the feature vector of the current text data in each text evaluation dimension and the corresponding attention weight is taken as the weighted vector; that is, each element in the feature vector is weighted by the attention weight to determine the corresponding weighted vector; thereby improving the output result of each text evaluation dimension.

[0096] Then the weighted vectors of all text evaluation dimensions are spliced to determine the joint feature vector; the joint feature vector is input into the trained multi-layer fully connected neural network to output the digestive tract risk evaluation index. In one specific implementation manner of the embodiment of the present application, the multi-layer fully connected neural network adopts a three-layer fully connected neural network; wherein the first layer is fully connected+BatchNorm+ReLU+Dropout; the second layer is fully connected+BatchNorm+ReLU+Dropout; the third layer is fully connected+Sigmoid output; the input dimension is the fusion feature dimension, and 1 prediction probability value is output; the digestive tract bleeding results of the data are manually annotated according to the multi-modal data of the historical digestive tract bleeding patients; the joint feature vectors of these multi-modal data are extracted and divided into a training set and a validation set according to 7:3; a binary classification cross-entropy is used as a loss function, an AdamW optimizer is used, and an end-to-end joint training is performed to realize the training of the prediction neural network; then the joint feature vector of the current moment of the digestive tract bleeding patient in the embodiment of the present application is input into the trained three-layer fully connected neural network to output the digestive tract bleeding risk probability, and the digestive tract bleeding risk probability is taken as the digestive tract bleeding risk evaluation coefficient of the current moment of the digestive tract bleeding patient.

[0097] In one specific implementation manner of the embodiment of the present application, after outputting the digestive tract bleeding risk evaluation coefficient of the current moment of the digestive tract bleeding patient, it further includes: when the digestive tract bleeding risk evaluation index is greater than a preset risk threshold, it is indicated that the digestive tract bleeding risk at the current moment is high, and a bleeding warning is issued; when the digestive tract bleeding risk evaluation is less than or equal to the preset risk threshold, no bleeding warning is issued; wherein the preset risk threshold is set to 0.75, which can be adjusted according to the specific implementation environment, so that the digestive tract bleeding risk evaluation index can be responded according to the size, and the evaluation result is more specific.

[0098] To sum up, the deep learning-based digestive tract bleeding risk assessment system first preprocesses the non-structured text data, determines the key words and non-key words in the text data; then, in combination with the risk features represented by the text data changes and the influence of aspirin on the bleeding risk, determines the formal risk feature values in the text form dimension; then, in combination with the semantic information of the text data and the text semantics of each risk link corresponding to the digestive tract risk, performs relevance analysis, and determines the semantic risk feature values in the semantic information dimension; finally, in combination with the formal risk feature values and the semantic risk feature values, determines the attention weights of each text evaluation dimension, and adjusts the neural network model according to the attention weights, so that the accuracy of the output digestive tract bleeding risk assessment coefficient is higher; the influence of the key modal data on the final prediction result is further improved, the noise interference of the weak modal data is reduced, and the accuracy of the prediction result is further improved.

[0099] The embodiment of the present application also provides a computer device, please refer to Figure 2 which shows a computer device structure schematic diagram provided by an embodiment of the present application, the computer device includes memory 201, processor 202 and computer program 203 stored in the memory 201 and running on the processor 202, wherein, when the processor 202 executes the computer program 203, the computer device can execute any one of the foregoing deep learning-based digestive tract bleeding risk assessment system.

[0100] The embodiment of the present application also provides a computer program product, when the computer program product runs on the computer device, so that the computer device can execute any one of the foregoing deep learning-based digestive tract bleeding risk assessment system.

[0101] The embodiment of the present application also provides a computer readable storage medium, the computer readable storage medium has computer program code stored therein, when the computer program code runs on the computer device, so that the computer device can execute any one of the foregoing deep learning-based digestive tract bleeding risk assessment system.

[0102] In the embodiments provided in the present application, it should be understood that the computer device, computer program product and computer readable storage medium provided are all used to execute the corresponding system provided in the foregoing, and therefore the beneficial effects that can be achieved are referable to the beneficial effects of the system provided in the foregoing, which will not be described herein again.

[0103] It should be noted that: the sequence of the above-mentioned embodiments is only for description, and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or can be advantageous.

[0104] The various embodiments described in this specification are presented by way of example, and each embodiment is not inherently more important than any other embodiment. To the extent that any embodiment is directed to a distinct, independently applicable inventive concept, it is to be understood that the inventive concept(s) can be realized in a multitude of alternative ways. Each embodiment is presented for the purpose of illustrating one or more aspects of the inventive concept(s), and the application should not be construed as requiring that the inventive concept(s) be limited to only those embodiments.

Claims

1.A deep learning-based digestive tract bleeding risk assessment system, characterized by, The system comprises: a data acquisition preprocessing module, configured to acquire each piece of text data of a patient with gastrointestinal bleeding in each text evaluation dimension in a medical database; extract keywords and all non-key words in the text data; a risk feature value determination module, configured to sequentially take each keyword in the current text data as a target word under each text evaluation dimension; determine the number of descriptive words of the target word in each piece of text data according to the mean value of the number of non-key words in all sentences containing the target word in each piece of text data; take a set composed of all kinds of keywords in each piece of text data as a corresponding keyword set; and determine a symptom description coefficient of the target word in the current text data according to the coincidence of the keyword set and the relative reduction of the number of descriptive words between the current text data and the previous piece of text data; determine a time interval coefficient of the current text data by negatively correlating the length of the time interval between the recording time of the current text data and the recording time of the previous text data; determine a risk development coefficient of the current text data according to the time interval between the recording time of the current text data and the date of the latest aspirin medication of the patient with gastrointestinal bleeding; determine a weighted structural risk feature value of the target word in the current text data according to the product of the risk development coefficient, the time interval coefficient and the symptom description coefficient; determine a corresponding formal risk feature value according to the mean value of the weighted structural risk feature values of all kinds of keywords in the current text data under each text evaluation dimension; convert all keywords in the current text data into word vectors by word2vec to determine the word vector of each keyword under each text evaluation dimension; determine a corresponding semantic similarity according to the cosine similarity between the word vector of each keyword and the word vector of each other keyword; and perform hierarchical clustering on the current text data by taking the negatively correlated mapping value of the semantic similarity between keywords as a clustering distance to obtain at least two keyword clustering clusters; determine a corresponding center vector according to the mean vector of the word vectors of all keywords in each keyword clustering cluster; and determine a corresponding risk score according to the mean value of all semantic similarities between all keywords in each keyword clustering cluster; determine the risk association strength of each keyword clustering cluster according to the vector similarity distribution between the center vector and the preset reference word vector of each risk link; determine the local risk feature value of each keyword clustering cluster according to the product of the risk association strength and the risk score; and determine the semantic risk feature value of each text evaluation dimension according to the mean value of the local risk feature values of all keyword clustering clusters; a gastrointestinal bleeding risk assessment module, configured to determine an attention weight of each text evaluation dimension according to the formal risk feature value and the semantic risk feature value; adjust a neural network model according to the attention weight to output a gastrointestinal bleeding risk assessment coefficient of the patient with gastrointestinal bleeding at the current time. 2.The digestive tract bleeding risk assessment system based on deep learning according to claim 1, wherein, The process of extracting the keywords and all non-key words in the text data comprises: The text data is input into the BioBERT model for word segmentation and keyword extraction, and the corresponding keywords of the gastrointestinal bleeding and non-key words are output. 3.The digestive tract bleeding risk assessment system based on deep learning according to claim 1, wherein, The symptom description coefficient acquisition process includes: According to the difference between the description word quantity of the target word in the last text data of the current text data and the description word quantity of the target word in the current text data, the description word reduction quantity of the target word in the current text data is determined. According to the intersection between the keyword set of the current text data and the keyword set of the corresponding last text data, a reference set is determined; and according to the ratio between the number of keyword categories in the reference set and the number of keyword categories in the keyword set of the current text data, the keyword combination stability of the current text data is determined. The product of the description word reduction quantity and the keyword combination stability is positively correlated, and the symptom description coefficient of the target word in the current text data is determined. 4.The digestive tract bleeding risk assessment system based on deep learning according to claim 1, wherein, The risk development coefficient acquisition process includes: The time interval length between the recording time of the current text data and each taking time is taken as the corresponding judgment time length; When the judgment time length is less than or equal to the preset first time threshold, the normalized value of the judgment time length is taken as the reference development coefficient of the corresponding taking time; When the judgment time length is greater than the preset first time threshold and less than or equal to the preset second time threshold, the preset highest risk feature value is taken as the reference development coefficient of the corresponding taking time; When the judgment time length is greater than the preset second time threshold, the negative correlation mapping value of the judgment time length is taken as the reference development coefficient of the corresponding taking time; The maximum value of the reference development coefficients of all taking times is taken as the risk development coefficient of the current text data. 5.The digestive tract bleeding risk assessment system based on deep learning according to claim 1, wherein, The risk association strength acquisition process includes: The cosine similarity between the center vector and each preset reference word vector of each risk link is calculated to determine the link matching degree between each keyword clustering cluster and each risk link; and the maximum value of the link matching degrees between each keyword clustering cluster and all risk links is determined as the corresponding risk association strength. 6.The digestive tract bleeding risk assessment system based on deep learning according to claim 1, wherein, The attention weight acquisition process includes: The product of the formal risk feature value and the semantic risk feature value is determined as the attention weight of each text evaluation dimension. 7.The digestive tract bleeding risk assessment system based on deep learning according to claim 1, wherein, The gastrointestinal bleeding risk assessment coefficient acquisition process includes: In each text evaluation dimension, the occurrence frequency of each keyword in the current text data is taken as the corresponding frequency feature value; and the frequency feature values of all keywords in the current text data are arranged in the order of occurrence of all keywords in all text data to determine the feature vector of the current text data in each text evaluation dimension; The product of the feature vector of the current text data in each text evaluation dimension and the corresponding attention weight is taken as a weighted vector; and the joint feature vector is determined by concatenating the weighted vectors of all text evaluation dimensions. Input the joint feature vector into the trained multi-layer fully connected neural network, and output a digestive tract risk assessment index. 8.The digestive tract bleeding risk assessment system based on deep learning according to claim 1, wherein, The output digestive tract bleeding risk assessment coefficient of the patient at the current time also includes: When the digestive tract bleeding risk assessment index is greater than the preset risk threshold, a bleeding warning is issued; When the digestive tract bleeding risk assessment index is less than or equal to the preset risk threshold, no bleeding warning is issued.

Citation Information

Patent Citations

  • Risk assessment system based on medical record, electronic equipment and storage medium

    CN120727296A