A disease prediction method and system based on multi-granularity feature fusion

By adopting a multi-grained feature fusion method in disease prediction, combining word features, conceptual features, conceptual relationship features and attribute-value features, and training parallel adaptive convolutional neural network models, the problem of difficulty in making full use of medical data semantic information in the existing technology is solved, and the performance of disease prediction models is improved.

CN112331332BActive Publication Date: 2025-05-30GUANGZHOU ANYUE INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011095993.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-14
Publication Date
2025-05-30
Estimated Expiration
2040-10-14

AI Technical Summary

Technical Problem

The prior art is difficult to fully utilize the semantic information in medical data in disease prediction, resulting in insufficient performance of the prediction model.

Method used

A disease prediction method based on multi-grained feature fusion is adopted, and a parallel adaptive convolutional neural network model is trained to classify disease types by obtaining and fusing multiple granularity features (word features, conceptual features, conceptual relationship features and attribute-value features).

Benefits of technology

Through the fusion of multi-grained feature, semantic information in medical text can be more comprehensively understood, and the performance and accuracy of disease prediction models can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112331332B_ABST
    Figure CN112331332B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a disease prediction method and system based on multi-granularity feature fusion, including: obtaining fusion features based on a disease to be predicted; inputting the fusion features into a trained disease prediction model to obtain a classification result of the disease type; wherein, the disease prediction model is obtained by training with fusion features of multiple diseases based on a parallel adaptive convolutional neural network model. By adopting the prediction method of multi-granularity feature fusion, the embodiment of the present invention not only uses fine-grained word and concept features, but also uses larger-granularity concept relationships and attribute-value features to fully understand the semantic information in medical texts, thereby improving the performance of the model in disease prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a disease prediction method and system based on multi-granularity feature fusion. Background Art

[0002] Disease prediction is to automatically classify diseases into different categories using existing semantic analysis techniques, which can help doctors or patients quickly understand the current disease state of the patient, and schedule and coordinate key medical resources according to the prediction of possible intervention means.

[0003] So far, the construction methods of prediction models are mainly divided into two categories: hypothesis-driven methods and data-driven methods. The former starts from the hypotheses proposed by clinical experts based on observations and clinical experience, then finds facts from medical data, and uses deductive reasoning to verify the authenticity of the hypotheses. The prediction model is derived from a set of verified hypotheses. Generally speaking, hypothesis-driven methods cannot fully utilize the valuable information contained in medical data. Data-driven methods use fully labeled medical data sets to train machine learning models to achieve disease prediction. Traditional machine learning models require domain experts to specify clinical features in a special way, and the success of the final prediction model largely depends on the complex supervision of manually designed feature selection. For example, Senthilkmar Mohan et al. proposed a linear hybrid random forest model for heart disease prediction in the article "Effective Heart Disease Prediction Using Hybrid Machine Learning Techniques" published in 2019. Deep learning can reduce the complexity of traditional machine learning feature selection and automatically learn deeper features from data, and has now become the main method for prediction models. And disease prediction methods based on deep learning usually use words or concept vectors as the main feature expressions of medical texts. For example, the article "Augmenting Embedding with Domain Knowledge for Oral Disease Diagnosis Prediction" published by Guangkai Li, Songmao Zhang et al. in SmartCom 2018 learns concepts related to symptoms and diagnoses from domain ontologies and uses neural networks to learn the concept features in electronic medical records to construct an oral disease prediction model. However, only considering words or concept vectors is likely to cause insufficient extraction of semantic information contained in medical texts due to their too small feature granularity, and cannot provide correct medical decisions. Summary of the Invention

[0004] An embodiment of the present invention provides a disease prediction method and system based on multi-granularity feature fusion to solve the defects existing in the prior art.

[0005] In a first aspect, an embodiment of the present invention provides a disease prediction method based on multi-granularity feature fusion, including:

[0006] Obtain the fusion features based on the disease to be predicted;

[0007] Input the fusion features into the trained disease prediction model to obtain the classification result of the disease type; wherein, the disease prediction model is obtained by training with the fusion features of multiple diseases based on a parallel adaptive convolutional neural network model.

[0008] Further, the disease prediction model is obtained through the following steps:

[0009] Obtain the text to be processed, and obtain the preprocessed text after preprocessing the text to be processed;

[0010] Extract features from the preprocessed text to obtain the extracted features;

[0011] Fuse the extracted features based on multi-granularity features to obtain the fusion features of multiple diseases;

[0012] Obtain a parallel adaptive convolutional neural network model, and input the fusion features of multiple diseases into the parallel adaptive convolutional neural network model for training to obtain the disease prediction model.

[0013] Further, the step of obtaining the text to be processed and obtaining the preprocessed text after preprocessing the text to be processed specifically includes:

[0014] Manually annotate the medical text data according to the target category to be predicted, and then load the domain ontology to obtain the text to be processed;

[0015] Segment the text to be processed into Chinese character strings according to punctuation marks, numbers, and space symbols, and remove stop words to obtain the preprocessed text.

[0016] Further, the step of extracting features from the preprocessed text to obtain the extracted features specifically includes:

[0017] Extract features from the preprocessed text through concept feature extraction, word feature extraction, concept relationship feature extraction, and attribute and value feature extraction to obtain the extracted features.

[0018] Further, the feature extraction of the preprocessed text through concept feature extraction, word feature extraction, concept relationship feature extraction, and attribute and value feature extraction to obtain the extracted features specifically includes:

[0019] Map the preprocessed text to the domain ontology to obtain text data, segment the text data into semantic sets through the maximum matching method, use the word2vec model to transform the concept self-feature type and concept type features that can find matching ones in the domain ontology into vector forms, and extract concept features by combining the concept self-feature type and the concept type features;

[0020] Use the word2vec model to transform the concept self-feature type and concept type features that cannot find matching ones in the domain ontology into vector forms, and extract word features;

[0021] Combine the word features, position features, and negation word features to extract the relationship trigger words between concepts, and combine the concept features to represent the concept features and the relationship trigger words as concept relationship features;

[0022] Further represent the concept features as disease and time results including numerical types, and detection and examination results including the numerical types and category types, to obtain attribute and value features.

[0023] Further, the fusion of the extracted features based on multi-granularity features to obtain the fusion features of multiple diseases specifically includes:

[0024] For categories with large differences in prediction targets, directly splice the extracted features into vectors, or for categories with high similarity in prediction targets, use a weight-based feature fusion method to fuse the extracted features to obtain the fusion features of multiple diseases.

[0025] Further, obtain a parallel adaptive convolutional neural network model, input the fusion features of multiple diseases into the parallel adaptive convolutional neural network model for training to obtain the disease prediction model, specifically including:

[0026] According to the differences in the concept relationship features and the attribute and value features, segment the sentence into different parts to extract the semantic information contained in the sentence;

[0027] Fuse the semantic information with the concept features and the word features to train the parallel adaptive convolutional neural network model, and use dropout operation in the convolutional layer and zero padding to maintain the validity of the sentence to obtain the disease prediction model.

[0028] Second aspect, the embodiments of the present invention further provide a disease prediction system based on multi-granularity feature fusion, including:

[0029] An acquisition module, configured to acquire fusion features based on a disease to be predicted;

[0030] A processing module, configured to input the fusion features into a trained disease prediction model to obtain a classification result of the disease type; wherein, the disease prediction model is obtained by training with fusion features of multiple diseases based on a parallel adaptive convolutional neural network model.

[0031] Third aspect, the embodiments of the present invention further provide an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any one of the above-mentioned disease prediction methods based on multi-granularity feature fusion are implemented.

[0032] Fourth aspect, the embodiments of the present invention further provide a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned disease prediction methods based on multi-granularity feature fusion are implemented.

[0033] The disease prediction method and system based on multi-granularity feature fusion provided by the embodiments of the present invention adopt a prediction method of multi-granularity feature fusion, not only using fine-grained word and concept features, but also using larger-granularity concept relationships and attribute-value features to fully understand the semantic information in medical texts, thereby improving the performance of the model for disease prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 is a schematic flowchart of a disease prediction method based on multi-granularity feature fusion provided by an embodiment of the present invention;

[0036] Figure 2 is a schematic diagram of the decomposition of the process modules provided by an embodiment of the present invention:

[0037] Figure 3 is a schematic structural diagram of a disease prediction system based on multi-granularity feature fusion provided by an embodiment of the present invention;

[0038] Figure 4It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0040] In view of the problems existing in the prior art, an embodiment of the present invention proposes a disease prediction method based on multi-granularity feature fusion. This method extracts features of different granularities and fuses them based on an existing medical ontology and an annotated corpus to train a disease prediction model. The trained model can provide a category corresponding to the prediction target and can be used in applications related to disease prediction, such as disease type prediction or disease severity level prediction.

[0041] Figure 1 It is a schematic flowchart of a disease prediction method based on multi-granularity feature fusion provided by an embodiment of the present invention. As Figure 1 shown, it includes:

[0042] S1. Obtain fusion features based on the disease to be predicted;

[0043] S2. Input the fusion features into the trained disease prediction model to obtain a classification result of the disease type. Among them, the disease prediction model is obtained by training with fusion features of multiple diseases based on a parallel adaptive convolutional neural network model.

[0044] Specifically, fusion features related to the disease to be predicted are obtained through certain technical means, and the fusion features are input into a pre-trained disease prediction model to obtain a final classification result of the disease type. Among them, the disease prediction model is based on a parallel adaptive convolutional neural network and is obtained by training with fusion features of multiple diseases.

[0045] By adopting the prediction method of multi-granularity feature fusion in the embodiment of the present invention, not only fine-grained word and concept features are adopted, but also larger-granularity concept relationships and attribute-value features are adopted to fully understand the semantic information in medical texts and improve the performance of the model for disease prediction.

[0046] Based on the above embodiment, the disease prediction model is obtained through the following steps:

[0047] Obtain the text to be processed, and obtain the preprocessed text after preprocessing the text to be processed;

[0048] Perform feature extraction on the preprocessed text to obtain extracted features;

[0049] Fuse the extracted features based on multi-granularity features to obtain the fused features of the multiple diseases;

[0050] Obtain a parallel adaptive convolutional neural network model, and input the fused features of the multiple diseases into the parallel adaptive convolutional neural network model for training to obtain the disease prediction model.

[0051] Specifically, as Figure 2 shown, when training the disease prediction model, first perform data preprocessing 1 on the domain ontology to obtain the preprocessed text, then pass the preprocessed text through feature extraction 2 including concept feature 21, word feature 22, concept relationship feature 23, and attribute-value feature 24 to obtain the extracted features, and then fuse the extracted features based on multi-granularity features 3, including directly performing vector splicing 31 or a fusion method 32 based on feature weights, to obtain the fused features of the multiple diseases. Further, based on the obtained parallel adaptive convolutional neural network model, use the fused features of the multiple diseases to train the model to obtain the trained disease prediction model 4, and finally use the trained model for disease type classification 5.

[0052] Based on any of the above embodiments, the obtaining of the text to be processed and the preprocessing of the text to be processed to obtain the preprocessed text specifically include:

[0053] Manually annotate the medical text data according to the target category to be predicted, and then load the domain ontology to obtain the text to be processed;

[0054] Segment the text to be processed into Chinese character strings according to punctuation marks, numbers, and space symbols, and remove stop words to obtain the preprocessed text.

[0055] Specifically, manually annotate the medical text data according to the target category to be predicted, and then load the domain ontology; segment the text to be processed into Chinese character strings according to punctuation marks, numbers, and space symbols, and remove stop words to obtain the preprocessed text.

[0056] Based on any of the above embodiments, the performing of feature extraction on the preprocessed text to obtain the extracted features specifically includes:

[0057] Perform feature extraction on the preprocessed text through concept feature extraction, word feature extraction, concept relationship feature extraction, and attribute and value feature extraction to obtain the extracted features.

[0058] Among them, the feature extraction of the preprocessed text through concept feature extraction, word feature extraction, concept relationship feature extraction, and attribute and value feature extraction to obtain the extracted features specifically includes:

[0059] Map the preprocessed text to the domain ontology to obtain text data, segment the text data into semantic sets by the maximum matching method, use the word2vec model to convert the concept self-feature type and concept type feature that can find matching ones in the domain ontology into vector forms, and extract concept features by combining the concept self-feature type and the concept type feature;

[0060] Use the word2vec model to convert the concept self-feature type and concept type feature that cannot find matching ones in the domain ontology into vector forms and extract word features;

[0061] Combine the word features, position features, and negation word features to extract the relationship trigger words between concepts, and combine the concept features to represent the concept features and the relationship trigger words as concept relationship features;

[0062] Further represent the concept features as disease and time results including numerical types, and detection and inspection results including the numerical types and category types to obtain attribute and value features.

[0063] Specifically, it is divided into four steps: concept feature extraction, word feature extraction, concept relationship feature extraction, and attribute-value feature extraction.

[0064] Concept features include concept self-features and concept type features. First, map the preprocessed text to the domain ontology, and segment the text data into semantic sets {Y 1 ,…Y n} ∈ D, where D is the text data and contains a concept set {C 1 ,…C n} ∈ Y that can find matching ones in the domain ontology, and there are corresponding concept types {C 1type ,…C Ntype}. Secondly, use the word2vec model to convert the concepts and concept types into d-dimensional vector forms. Finally, extract concept features by combining the concept self-feature and the concept type feature, denoted as where c i is the concept self-feature belonging to the concept set {C 1 ,…C N}, c itype is the type feature of concept c i belonging to {C 1type ,…C Ntype}, It is a vector splicing operation.

[0065] The word feature refers to the semantics for which no matching concept can be found in the domain ontology, denoted as {W 1 ,…W n} ∈ Y. Similarly, word2vec is used to convert words into d-dimensional vector forms, denoted as w = {w 1 ,…w n}.

[0066] By combining the aforementioned word features, position features, and negation word features, relationship trigger words between concepts are extracted. And by combining the aforementioned concept features, the concept relationship features are represented in triple form, denoted as p i = (e i , r i , e o ), p = {p 1 …p n}, p i ∈ p, where e i and e o represent concept features, and r i represents the relationship trigger word between concepts. There is {s 1 …s i …s n} ∈ D, where s i consists of m semantics s i = {w 1 …p i …q o …w m}, where e i and e o represent the concepts included in s i . {w 1 …w m} are the word features in sentence s i . Each word has two relative distances between concept features e i and e o , denoted as Since negation words can change the meaning of words, negation word features are extracted by loading negation word points, denoted as {n 1 …n m} ∈ w, where w represents the set of word features. Finally, the relationship trigger word between concepts can be expressed by the formula as where concept features e i and e o , and relationship trigger word r i are in the same spatial dimension, expressed as

[0067] The attribute-value features include two categories: disease-time and test-result. An attribute refers to a conceptual feature. The values in disease-time only include numerical types, and the values in test-result include numerical types and categorical types. For numerical types, both the numerical value and its corresponding unit symbol are considered. For example, for a numerical value V i it has its corresponding unit symbol U i , and the updated numerical type is calculated as where u i represents the vector form of the unit symbol. The disease-time feature is represented as t i =(e o , v m ). For the numerical types in test-result, the index level features need to be extracted. For example, if there is a concept C={C 1 , C 2 , …, C n}, its corresponding numerical values v={v 1 , v 2 , …, v n} and index levels L={L 1 , L 2 , L 3}, the value of the test result can be represented in a triple form z i =(e i , v i , l i ), where e o and e i are conceptual features, {v m , v i}∈v, and l i is the index level vector. Categorical types have no unit symbols and are usually composed of strings. For example: negative, positive, etc. Therefore, the negation word features need to be extracted to accurately express the semantics contained in the text. For categorical types without negation words, their categorical vectors are directly extracted. For categorical types with negation words, combining the categorical features and negation word features, it can be represented as where b m is the categorical feature and n m is the negation word vector. Therefore, the categorical type of test-result can be represented as k i =(e m , g m ), where e m represents the conceptual feature. t i , z i , k i ∈q i , q={q 1 , …, q n}, q i ∈q, where q represents the set of attribute-value features.

[0068] Based on any of the above embodiments, fusing the extracted features based on multi-granularity features to obtain the fused features of multiple diseases specifically includes:

[0069] For categories with large differences in prediction targets, directly splice the extracted features into vectors; or for categories with high similarity in prediction targets, use a weight-based feature fusion method to fuse the extracted features to obtain the fused features of multiple diseases.

[0070] Specifically, different feature fusion methods are adopted according to different prediction targets. For categories with large differences in prediction targets, the extracted features can be directly spliced into vectors; for categories with high similarity in prediction targets, a weight-based feature fusion method is adopted, and the specific description is as follows:

[0071] Directly splicing the extracted features into vectors, the formula can be expressed as:

[0072]

[0073] Among them, e i represents the concept feature, w i represents the word feature, p i represents the concept relationship feature, q i represents the attribute-value feature.

[0074] The weight-based feature fusion method, the formula can be expressed as:

[0075] First, different weights are respectively set for each type of feature according to its importance in this type of feature. For example, 4 weights are set, and the calculation formula can be expressed as:

[0076]

[0077]

[0078]

[0079]

[0080] Among them, e i represents the concept feature, w i represents the word feature, p i represents the concept relationship feature, q i represents the attribute-value feature. α i ∈[0,1], and

[0081] Secondly, calculate the weight-based eigenvalue by combining the weights and feature vectors obtained in the above formula.

[0082]

[0083] Among them, CE i represents the weight-based concept feature, WE i represents the weight-based word feature, RE i represents the weight-based concept relationship feature, VE i represents the weight-based attribute-value feature.

[0084] According to the above content, the weight-based concept feature, word feature, concept relationship feature, and attribute-value feature are fused as the input of the parallel adaptive convolutional neural network to train the disease prediction model.

[0085]

[0086] Based on any of the above embodiments, obtaining the parallel adaptive convolutional neural network model, inputting the fusion features of the multiple diseases into the parallel adaptive convolutional neural network model for training to obtain the disease prediction model specifically includes:

[0087] Segment the sentence into different parts according to the differences between the concept relationship feature and the attribute and value features to extract the semantic information contained in the sentence;

[0088] Fuse the semantic information with the concept feature and the word feature to train the parallel adaptive convolutional neural network model, and adopt the dropout operation in the convolutional layer and adopt zero padding to maintain the validity of the sentence to obtain the disease prediction model.

[0089] Specifically, a parallel adaptive convolutional neural network is used to train the disease prediction model, and the specific formula is as follows:

[0090] Convolutional layer: There is a sentence s i ={w 1 ,w 2 ,…,w m}, where w j is the j-th word vector in the sentence s i . h is the length of the convolutional kernel, indicating that it contains h words. The convolution operation for the j-th word is:

[0091] c j =f(k·w i:i+h-1 +b)

[0092] where is the matrix of the convolutional kernel, b is the bias, w i:i+h-1represents the combination of word vectors from the $i$-th to the $i + h - 1$-th, $f(·)$ represents a non-linear activation function, usually ReLU, $c$ j represents a feature map after convolution operation, sentence $s$ i The feature map of is expressed as: Suppose there are $l$ convolution kernels of length $h$, $1 < i < l$, the feature map is expressed as:

[0093]

[0094] Parallel adaptive pooling layer: First, the sentence is segmented into different parts according to the concept relationship and attribute-value features, and two types of features are learned in parallel.

[0095] For the concept relationship feature, the sentence $c$ is divided into three parts according to the position of the concept pair j as $[c$ j1 , $c$ j2 , $c$ j3 . Secondly, the most important information in the sentence is obtained by calculating the maximum value of each part, and the calculation formula is as follows: Finally, all the feature maps after convolution operation are concatenated to obtain the feature vector $b$ of sentence $s$ i $ = ReLU(v)$. sp

[0096] For the attribute-value feature, the sentence $c$ is divided into two parts according to the position of the concept j as $[c$ j1 , $c$ j2 . Secondly, the information of the value most relevant to the concept relationship in the sentence is obtained by calculating the maximum value of each part, and the calculation formula is as follows: All the feature maps after convolution operation are concatenated to obtain the feature vector $b$ of sentence $s$ i $ = ReLU(u)$. Finally, the final sentence feature vector is obtained by combining the sentence feature vectors of the concept relationship and the attribute value sq

[0097]

[0098] Finally, the extracted concept relationships, attribute-features and concepts, and word features are combined, and the result is put into the classification layer of the parallel adaptive convolutional neural network, and the classification result of the final disease type is generated through the softmax classifier. Based on different feature fusion methods, the result formula generated by the classifier is as follows:

[0099] (1) Directly perform vector concatenation on the extracted features:

[0100] $O = softmax(W$ o $h$ i ​​+b s )

[0101] r s = argmax(O)

[0102] where e i is the concept feature, w i is the word feature, p i is the concept relationship feature, q i is the attribute - value feature, b s is the feature vector of sentence s i W o is the weight, O ∈ [1, n] indicates there are n relationship types, r s is the final relationship category label.

[0103] (2) Feature fusion method based on weights:

[0104]

[0105] D = softmax(W o f i +b s )

[0106] r s = argmax(D)

[0107] where CE i represents the concept feature based on weights, WE i represents the word feature based on weights, RE i represents the concept relationship feature based on weights, VE i represents the attribute - value feature based on weights. b s is the feature vector of sentence s i W o is the weight, D ∈ [1, n] indicates there are n relationship types, r s is the final relationship category label.

[0108] The disease prediction system based on multi - granularity feature fusion provided by the embodiments of the present invention will be described below. The disease prediction system based on multi - granularity feature fusion described below can be correspondingly referred to the disease prediction method based on multi - granularity feature fusion described above.

[0109] Figure 3 is a schematic structural diagram of a disease prediction system based on multi - granularity feature fusion provided by the embodiments of the present invention. As Figure 3 shown, it includes: an acquisition module 310 and a processing module 320; where:

[0110] The acquisition module 310 is used to acquire the fusion features based on the disease to be predicted; the processing module 320 is used to input the fusion features into the trained disease prediction model to obtain the classification result of the disease type; wherein, the disease prediction model is obtained by training with the fusion features of multiple diseases based on a parallel adaptive convolutional neural network model.

[0111] In the embodiment of the present invention, by adopting the prediction method of multi-granularity feature fusion, not only fine-grained word and concept features are adopted, but also larger-granularity concept relationships and attribute-value features are adopted to fully understand the semantic information in medical texts and improve the performance of the model for disease prediction.

[0112] Figure 4 An example of the entity structure diagram of an electronic device is shown as Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete communication with each other through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute the disease prediction method based on multi-granularity feature fusion. The method includes: acquiring the fusion features based on the disease to be predicted; inputting the fusion features into the trained disease prediction model to obtain the classification result of the disease type; wherein, the disease prediction model is obtained by training with the fusion features of multiple diseases based on a parallel adaptive convolutional neural network model.

[0113] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software function units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks and other various media that can store program codes.

[0114] On the other hand, an embodiment of the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the disease prediction method based on multi-granularity feature fusion provided by each of the above method embodiments. The method includes: obtaining fusion features based on a disease to be predicted; inputting the fusion features into a trained disease prediction model to obtain a classification result of the disease type; wherein, the disease prediction model is obtained by training with fusion features of multiple diseases based on a parallel adaptive convolutional neural network model.

[0115] In another aspect, an embodiment of the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the disease prediction method based on multi-granularity feature fusion provided by each of the above embodiments. The method includes: obtaining fusion features based on a disease to be predicted; inputting the fusion features into a trained disease prediction model to obtain a classification result of the disease type; wherein, the disease prediction model is obtained by training with fusion features of multiple diseases based on a parallel adaptive convolutional neural network model.

[0116] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0117] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, also by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A disease prediction system based on multi-granularity feature fusion, characterized in that, it includes: An acquisition module for acquiring fusion features based on the disease to be predicted; A processing module for inputting the fusion features into a trained disease prediction model to obtain a classification result of the disease type; wherein, the disease prediction model is obtained by training with fusion features of multiple diseases based on a parallel adaptive convolutional neural network model; The disease prediction model is obtained through the following steps: Acquire the text to be processed, and obtain the preprocessed text after preprocessing the text to be processed; Extract features from the preprocessed text to obtain extracted features; Fuse the extracted features based on multi-granularity features to obtain the fusion features of multiple diseases; Acquire a parallel adaptive convolutional neural network model, input the fusion features of multiple diseases into the parallel adaptive convolutional neural network model for training to obtain the disease prediction model; The extracting features from the preprocessed text to obtain extracted features specifically includes: Extracting features from the preprocessed text through concept feature extraction, word feature extraction, concept relationship feature extraction, and attribute and value feature extraction to obtain the extracted features; The extracting features from the preprocessed text through concept feature extraction, word feature extraction, concept relationship feature extraction, and attribute and value feature extraction to obtain the extracted features specifically includes: Map the preprocessed text to the domain ontology to obtain text data, segment the text data into semantic sets by the maximum matching method, use the word2vec model to convert the concept self-feature type and concept type feature that can find matching concepts in the domain ontology into vector forms, and extract concept features by combining the concept self-feature type and the concept type feature; Use the word2vec model to convert the concept self-feature type and concept type feature that cannot find matching concepts in the domain ontology into vector forms and extract word features; Extract the relationship trigger words between concepts by combining the word features, location features, and negation word features, and combine the concept features to represent the concept features and the relationship trigger words as concept relationship features. The concept relationship features are represented in the form of a triple, denoted as p i =(e i ,r i ,e o ), where e i and e o represent concept features, r i represents the relationship trigger word between concepts, and there is {s 1 …s i …s n} ∈ D, where D is the text data, and s i consists of m semantics s i ={w 1 …p i …q o …w m}, where {w 1 …w m} are the word features in the sentence s i . Each word has two relative distances with respect to the concept features e i and e o , denoted as The negation word feature is denoted as {n 1 …n m} ∈ w, where w represents the set of word features. The relationship trigger word between concepts can be expressed by the formula as Further represent the concept features as disease and time results including numerical types, and detection and inspection results including the numerical types and category types to obtain attribute and value features; The acquiring a parallel adaptive convolutional neural network model, inputting the fusion features of multiple diseases into the parallel adaptive convolutional neural network model for training to obtain the disease prediction model specifically includes: Segment the sentence into different parts according to the differences of the concept relationship features and the attribute and value features to extract the semantic information contained in the sentence; Fuse the semantic information with the concept features and the word features to train the parallel adaptive convolutional neural network model, and adopt dropout operation in the convolutional layer and use zero padding to maintain the validity of the sentence to obtain the disease prediction model.

2. The disease prediction system based on multi-granularity feature fusion according to claim 1, characterized in that, The acquiring the text to be processed, and obtaining the preprocessed text after preprocessing the text to be processed specifically includes: Manually annotate the medical text data according to the target category to be predicted, and then load the domain ontology to obtain the text to be processed; Segment the text to be processed into Chinese character strings according to punctuation marks, numbers, and space symbols, and remove stop words to obtain the preprocessed text.

3. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor executes the program, the following steps are implemented: Obtain the fusion features based on the disease to be predicted; Input the fusion features into the trained disease prediction model to obtain the classification result of the disease type; wherein, the disease prediction model is obtained by training with the fusion features of multiple diseases based on a parallel adaptive convolutional neural network model; The disease prediction model is obtained through the following steps: Obtain the text to be processed, and obtain the preprocessed text after preprocessing the text to be processed; Extract features from the preprocessed text to obtain the extracted features; Fuse the extracted features based on multi-granularity features to obtain the fusion features of multiple diseases; Obtain a parallel adaptive convolutional neural network model, input the fusion features of multiple diseases into the parallel adaptive convolutional neural network model for training to obtain the disease prediction model; The step of extracting features from the preprocessed text to obtain the extracted features specifically includes: Extract features from the preprocessed text through concept feature extraction, word feature extraction, concept relationship feature extraction, and attribute and value feature extraction to obtain the extracted features; The step of extracting features from the preprocessed text through concept feature extraction, word feature extraction, concept relationship feature extraction, and attribute and value feature extraction to obtain the extracted features specifically includes: Map the preprocessed text to the domain ontology to obtain text data, segment the text data into semantic sets by the maximum matching method, use the word2vec model to transform the concept self-feature type and concept type feature that can find matching concepts in the domain ontology into vector form, and extract concept features by combining the concept self-feature type and the concept type feature; Use the word2vec model to transform the concept self-feature type and concept type feature that cannot find matching concepts in the domain ontology into vector form, and extract word features; Extract the relationship trigger words between concepts by combining the word features, position features, and negation word features, and represent the concept features and the relationship trigger words as concept relationship features in the form of triples, denoted as p i =(e i ,r i ,e o ), where e i and e o represent concept features, r i represents the relationship trigger word between concepts, and {s 1 …s i …s n} ∈ D, where D is the text data, and s i consists of m semantics s i ={w 1 …p i …q o …w m}, where {w 1 …w m} are the word features in the sentence s i , and each word has two relative distances with respect to the concept features e i and e o , denoted as The negation word feature is denoted as {n 1 …n m} ∈ w, w represents the set of word features, and the relationship trigger word between concepts can be represented by the formula Further represent the concept features as the disease and time results including numerical types, and the detection and inspection results including the numerical types and category types to obtain the attribute and value features; The step of obtaining a parallel adaptive convolutional neural network model, inputting the fusion features of multiple diseases into the parallel adaptive convolutional neural network model for training to obtain the disease prediction model specifically includes: Segment the sentence into different parts according to the differences between the concept relationship features and the attribute and value features to extract the semantic information contained in the sentence; Fuse the semantic information with the concept features and the word features to train the parallel adaptive convolutional neural network model, and perform dropout operations in the convolutional layer and use zero padding to maintain the validity of the sentence, thereby obtaining the disease prediction model.

4. A non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, the following steps are implemented: Obtain the fusion features based on the disease to be predicted; Input the fusion features into the trained disease prediction model to obtain the classification result of the disease type; wherein, the disease prediction model is obtained by training with the fusion features of multiple diseases based on a parallel adaptive convolutional neural network model; The disease prediction model is obtained through the following steps: Obtain the text to be processed, and obtain the preprocessed text after preprocessing the text to be processed; Extract features from the preprocessed text to obtain the extracted features; Fuse the extracted features based on multi-granularity features to obtain the fusion features of multiple diseases; Obtain a parallel adaptive convolutional neural network model, input the fusion features of multiple diseases into the parallel adaptive convolutional neural network model for training to obtain the disease prediction model; The step of extracting features from the preprocessed text to obtain the extracted features specifically includes: Extract features from the preprocessed text through concept feature extraction, word feature extraction, concept relationship feature extraction, and attribute and value feature extraction to obtain the extracted features; The step of extracting features from the preprocessed text through concept feature extraction, word feature extraction, concept relationship feature extraction, and attribute and value feature extraction to obtain the extracted features specifically includes: Map the preprocessed text to the domain ontology to obtain text data, segment the text data into semantic sets by the maximum matching method, use the word2vec model to convert the concept self-feature type and concept type features that can find matching concepts in the domain ontology into vector forms, and extract concept features by combining the concept self-feature type and the concept type features; Use the word2vec model to convert the concept self-feature type and concept type features that cannot find matching concepts in the domain ontology into vector forms to extract word features; Extract the relationship trigger words between concepts by combining the word features, location features, and negation word features, and combine the concept features to represent the concept features and the relationship trigger words as concept relationship features. The concept relationship features are represented in the form of a triple, denoted as p i =(e i ,r i ,e o ), where e i and e o represent concept features, r i represents the relationship trigger word between concepts, and {s 1 …s i …s n} ∈ D, where D is the text data, and s i consists of m semantics s i ={w 1 …p i …q o …w m}, where {w 1 …w m} are the word features in the sentence s i . Each word has two relative distances with respect to the concept features e i and e o , denoted as The negation word features are denoted as {n 1 …n m} ∈ w, where w represents the set of word features. The relationship trigger word between concepts can be expressed by the formula Further represent the concept features as the disease and time results including numerical types, and the detection and inspection results including the numerical types and category types to obtain the attribute and value features; The step of obtaining a parallel adaptive convolutional neural network model, inputting the fusion features of multiple diseases into the parallel adaptive convolutional neural network model for training to obtain the disease prediction model specifically includes: Segment the sentence into different parts according to the differences between the concept relationship features and the attribute and value features to extract the semantic information contained in the sentence; Fuse the semantic information with the concept features and the word features to train the parallel adaptive convolutional neural network model, perform dropout operations in the convolutional layer, and use zero padding to maintain the validity of the sentence, thereby obtaining the disease prediction model.

Citation Information

Patent Citations

  • Convolutional neural network-based medical analysis assistant system

    CN109192299A

  • Named entity recognition method based on feature fusion

    CN109800437A

  • Diagnosis and treatment scheme prediction method and device

    CN110297908A