Chronic disease data analysis method and system based on natural language processing and integrated training
Through natural language processing and integrated training methods, chronic disease instances and optimized analysis models are solved, and the traditional method is insufficient in analyzing unstructured medical texts and processing chronic disease data complexity, achieving more efficient and accurate chronic disease data analysis.
Patent Information
- Application Number
- CN202510479534.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Traditional chronic disease data analysis methods are difficult to fully tap potential information in unstructured medical texts, the model accuracy and generalization capabilities are insufficient, and a single model is difficult to comprehensively analyze the diversity and complexity of chronic disease data.
Using a method based on natural language processing and integration training, chronic disease instances are generated through entity recognition and relationship extraction, and the initial analysis model is obtained using the integrated training framework to compare learning, and a chronic disease recognition model that has been trained is obtained through multi-model integration optimization.
Effective analysis of chronic disease data is achieved, the accuracy and generalization ability of the model are improved, and the diversity and complexity of chronic disease data can be handled more comprehensively.
Smart Images

Figure CN120015352A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a chronic disease data analysis method and system based on natural language processing and integrated training. Background Art
[0002] With the development of medical information technology, a large amount of unstructured medical text data contains rich chronic disease information, but it is difficult to use directly. Traditional chronic disease data analysis methods have limitations when processing such complex data, such as the inability to fully explore the potential information in the text, and insufficient model accuracy and generalization ability. At the same time, a single model is difficult to fully analyze the diversity and complexity of chronic disease data. Summary of the invention
[0003] The object of the present invention is to provide a chronic disease data analysis method and system based on natural language processing and integrated training.
[0004] In a first aspect, an embodiment of the present invention provides a chronic disease data analysis method based on natural language processing and integrated training, comprising: Use natural language processing technology to perform entity recognition and relationship extraction on unstructured medical texts to generate the first chronic disease instance without labeled chronic disease target values; Acquire a pre-trained initial analysis model, wherein the pre-trained initial analysis model is obtained by comparative learning based on the first chronic disease instance through an integrated training framework, wherein the integrated training framework integrates feature representations output by multiple heterogeneous models; Acquire a second chronic disease instance corresponding to the disease analysis indicator, where the second chronic disease instance is a chronic disease instance with a chronic disease target value marked; Based on the second chronic disease instance, the pre-trained initial analysis model is optimized by multi-model integration to obtain a trained chronic disease recognition model corresponding to the disease analysis index; The chronic disease symptoms to be analyzed are obtained, and the chronic disease symptoms to be analyzed are identified based on the trained chronic disease recognition model to obtain the chronic disease data analysis results of the chronic disease symptoms to be analyzed.
[0005] In a possible implementation, the method further includes: Obtaining a first chronic disease instance, wherein the first chronic disease instance includes a plurality of first chronic disease symptom instances; Determine a first chronic disease symptom instance topology of each first chronic disease symptom instance according to the symptom description information of each first chronic disease symptom instance, and determine a comparative learning feature encoding of each first chronic disease symptom instance according to the topology of each first chronic disease symptom instance; Performing recognition processing on each first chronic disease symptom instance topology based on a preset basic model to obtain a first feature vector instance corresponding to each first chronic disease symptom instance topology; Identify each first feature vector instance based on the trained feature recognition component to obtain a first symptom inference result corresponding to each first chronic disease symptom instance; The preset basic model is integrated and optimized for training based on the comparative learning feature coding of each first chronic disease symptom instance and the first symptom inference result to obtain a trained initial analysis model.
[0006] In a possible implementation, determining the comparative learning feature encoding of each first chronic disease symptom instance according to the topology of each first chronic disease symptom instance includes: Acquire multiple preset symptom topology patterns, wherein each symptom topology pattern corresponds to a symptom combination structure with clinical diagnostic significance; Associating each first chronic disease symptom instance topology with the multiple symptom topology patterns to obtain an association coefficient parameter of each first chronic disease symptom instance topology; According to the correlation coefficient parameter corresponding to the topology of each first chronic disease symptom instance, the comparative learning feature coding of each first chronic disease symptom instance is determined.
[0007] In a possible implementation, associating each first chronic disease symptom instance topology with the multiple symptom topology patterns to obtain an association coefficient parameter of each first chronic disease symptom instance topology includes: Associating the target first chronic disease symptom instance topology with each symptom topology pattern to obtain a correlation coefficient corresponding to each symptom topology pattern; The symptom topology pattern whose correlation coefficient exceeds the correlation coefficient threshold is used as the target symptom topology pattern corresponding to the target first chronic disease symptom instance topology; The target symptom topology pattern corresponding to the target first chronic disease symptom instance topology is used as the correlation coefficient parameter of the target first chronic disease symptom instance topology.
[0008] In a possible implementation, determining the comparative learning feature coding of each first chronic disease symptom instance according to the correlation coefficient parameter corresponding to the topology of each first chronic disease symptom instance includes: Obtaining the core feature code of the target symptom topology pattern corresponding to the target first chronic disease symptom instance topology; Determining, based on the core feature codes, control feature codes in the plurality of symptom topology patterns that are not associated with the target chronic disease symptom instance topology; The core feature code and the control feature code are subjected to feature coupling processing to obtain a comparative learning feature code of the target first chronic disease symptom instance topology.
[0009] In a possible implementation, the preset basic model includes a plurality of physical sign interaction networks and at least one physical sign fusion network, and the identification processing of each first chronic disease symptom instance topology based on the preset basic model to obtain a first feature vector instance corresponding to each first chronic disease symptom instance topology includes: According to the symptom description information of the target chronic disease symptom instance, obtaining the physical sign feature vector of each symptom entity in the topology of the target first chronic disease symptom instance, the associated physical sign entity of each symptom entity, and the association strength vector of the pathological association relationship between each symptom entity and the associated physical sign entity; Performing a feature conversion operation on the physical sign feature vector of each symptom entity to obtain a basic physical sign vector of each symptom entity; Performing feature fusion on the basic sign vector of the associated sign entity and the associated strength vector to obtain a fusion vector; Acquire a first weight tensor of a first feature aggregation module for interactive aggregation processing, perform a linear transformation operation on the first weight tensor and the fusion vector, and obtain a first transformation coefficient tensor; Processing the first transformation coefficient tensor based on a preset first nonlinear mapping function to obtain a first-order interaction aggregation vector of the symptom entity; Performing sign aggregation processing on the input sign vector of each symptom entity and the first-order interactive aggregation vector to obtain the first-order sign vector of each symptom entity; Perform interactive aggregation processing on the preceding order sign vector and the association strength vector of the associated sign entity based on the target sign interaction network to obtain the target order interactive aggregation vector of each symptom entity; Performing sign aggregation processing on the preceding order sign vector and the target order interactive aggregation vector of each symptom entity to obtain the target order sign vector of each symptom entity; Taking the target-order sign vector of each symptom entity as the feature vector of each symptom entity; The feature vector of each symptom entity is processed based on the physical sign fusion network to obtain a first feature vector instance corresponding to the target first chronic disease symptom instance topology.
[0010] In a possible implementation, performing sign aggregation processing on the input sign vector of each symptom entity and the first-order interaction aggregation vector to obtain the first-order sign vector of each symptom entity includes: Obtaining a second weight tensor of a second feature aggregation module for performing vital sign aggregation processing; Performing a linear transformation operation on the second weight tensor and the first-order interaction aggregation vector to obtain a second transformation coefficient tensor; The input sign vector of each symptom entity and the second transformation coefficient tensor are processed based on a preset second nonlinear mapping function to obtain the first-order sign vector of each symptom entity.
[0011] In a possible implementation, the step of obtaining, based on the symptom description information of the target chronic disease symptom instance, the sign feature vector of each symptom entity in the topology of the target first chronic disease symptom instance, the associated sign entity of each symptom entity, and the association strength vector of the pathological association relationship between each symptom entity and the associated sign entity includes: Acquire clinical indicator data of each sub-symptom instance in the target first chronic disease symptom instance according to the symptom description information of the target first chronic disease symptom instance; Performing feature coupling processing on the clinical indicator data of each sub-symptom instance to obtain a physical sign feature vector of each symptom entity in the target first chronic disease symptom instance topology; Taking the control sub-symptom having a clinical association path with each sub-symptom instance as the associated sub-symptom of each sub-symptom instance; Acquire clinical indicator data of a clinical association path between each sub-symptom instance and the associated sub-symptom; The clinical indicator data of the clinical association pathway is subjected to feature coupling processing to obtain an association strength vector of the pathological association relationship between each symptom entity and the associated sign entity.
[0012] In a possible implementation, the optimizing the pre-trained initial analysis model based on the second chronic disease example to obtain the trained chronic disease recognition model corresponding to the disease analysis index includes: Obtain a second chronic disease symptom instance topology for each second chronic disease symptom instance in the second chronic disease instance; Based on the pre-trained initial analysis model, each second chronic disease symptom instance topology is identified and processed to obtain a second feature vector instance corresponding to each second chronic disease symptom instance topology; Identify each second feature vector instance based on the trained feature recognition component to obtain a second symptom inference result corresponding to each second chronic disease symptom instance; The pre-trained initial analysis model is integrated and optimized according to the chronic disease target value of each second chronic disease symptom instance and the second symptom inference result to obtain a trained chronic disease recognition model corresponding to the disease analysis index.
[0013] In a second aspect, an embodiment of the present invention provides a server system, including a server, wherein the server is used to execute the method described in the first aspect.
[0014] Compared with the prior art, the beneficial effects provided by the present invention include: adopting a chronic disease data analysis method and system based on natural language processing and integrated training disclosed by the present invention, generating a first chronic disease instance without annotated chronic disease target value from unstructured medical text by using natural language processing technology, and obtaining a pre-trained initial analysis model based on this through comparative learning through an integrated training framework. Then obtain a second chronic disease instance with annotated chronic disease target value, optimize the multi-model integration of the initial analysis model, and obtain a trained chronic disease recognition model. Finally, use the model to identify the symptoms of the chronic disease to be analyzed, obtain the analysis results of the chronic disease data, and realize effective analysis of the chronic disease data. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative work.
[0016] Figure 1 A schematic diagram of the steps of a method for analyzing chronic disease data based on natural language processing and integrated training provided in an embodiment of the present invention; Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0018] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.
[0019] In order to solve the technical problems in the aforementioned background technology, Figure 1A flow chart of a chronic disease data analysis method based on natural language processing and integrated training provided in an embodiment of the present disclosure is provided below. The chronic disease data analysis method based on natural language processing and integrated training is introduced in detail.
[0020] Step S201, performing entity recognition and relationship extraction on unstructured medical text by natural language processing technology to generate a first chronic disease instance without annotated chronic disease target value; Step S202, obtaining an initial analysis model that has been trained in advance, wherein the initial analysis model that has been trained in advance is obtained by comparative learning based on the first chronic disease instance through an integrated training framework, wherein the integrated training framework integrates feature representations output by multiple heterogeneous models; Step S203, obtaining a second chronic disease instance corresponding to the disease analysis indicator, where the second chronic disease instance is a chronic disease instance with a chronic disease target value marked; Step S204, performing multi-model integrated optimization on the pre-trained initial analysis model based on the second chronic disease instance to obtain a trained chronic disease recognition model corresponding to the disease analysis index; Step S205, obtaining the chronic disease symptoms to be analyzed, and performing recognition processing on the chronic disease symptoms to be analyzed based on the trained chronic disease recognition model to obtain the chronic disease data analysis results of the chronic disease symptoms to be analyzed.
[0021] In an embodiment of the present invention, for example, a large number of patients' medical records are stored in the electronic medical record system of a hospital. Most of these medical records exist in the form of unstructured text, such as documents scanned and entered by doctors after handwriting, or descriptions of the condition freely entered by doctors in the system. After receiving these unstructured medical texts, the server uses the entity recognition algorithm in natural language processing technology to identify various entities related to chronic diseases from the text, such as "hypertension", "diabetes", "blood sugar level", "blood pressure value" and other disease names and related physical signs. At the same time, through the relationship extraction algorithm, the relationship between these entities is sorted out, such as "patient-suffering-hypertension" and "hypertension-related-to-blood pressure value". After such processing, the server integrates the extracted and sorted information to generate the first chronic disease instance without the target value of the chronic disease labeled. For example, a medical record describes that "the patient has recently felt dizzy, and the blood pressure is 160 / 100mmHg, and has a history of hypertension for many years." The server identifies entities such as "dizziness", "hypertension", "blood pressure value 160 / 100mmHg", and relationships such as "patient-has-history of hypertension" and "hypertension-association-blood pressure value", forming a first chronic disease instance, but at this time, this instance has no chronic disease target value annotation such as whether the diagnosis standard of hypertension is met. The server has previously used a large number of similar first chronic disease instances to perform preliminary training through an integrated training framework. In this process, the server is like a learned researcher, integrating multiple different types (heterogeneous) models, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) in deep learning, and decision tree models in traditional machine learning. These models each extract and learn features from the first chronic disease instance from different perspectives. For example, a CNN model may be better at capturing local feature patterns in a text, while an RNN model has a better grasp of the sequence information of the text. The server integrates the feature representations output by these heterogeneous models and then performs comparative learning. In contrastive learning, the server compares the features between different instances to find out the difference features between similar instances and different instances. For example, by comparing the first chronic disease instances generated by the medical records of different hypertensive patients, it is found that those instances with similar blood pressure values and similar symptoms have certain common features, while the features of the instances of hypotensive patients are significantly different. After such a training process, the server obtains an initial analysis model that has been trained in advance, which has the ability to preliminarily analyze information related to chronic diseases. The hospital's professional doctor team, based on clinical diagnostic criteria and past experience, has detailedly annotated some chronic disease instances to form second chronic disease instances. These instances not only contain information extracted from medical records similar to the first chronic disease instance, but also have clear chronic disease target values. For example, for the second chronic disease instance corresponding to the hypertension disease analysis indicator, it is marked whether the patient is diagnosed with hypertension (yes / no) and the grade of hypertension (first, second, third, etc.).The server obtains these annotated second chronic disease instances from the hospital's database, providing a key basis for the subsequent optimization of the initial analysis model. After the server obtains the second chronic disease instance, it uses the integrated training framework again. First, the relevant information of each second chronic disease instance is converted into a format that the model can process, such as generating the second chronic disease symptom instance topology (similar to the processing method of the first chronic disease symptom instance topology). Then, the pre-trained initial analysis model is used to identify and process each second chronic disease symptom instance topology, obtain the second feature vector instance corresponding to each topology, and let the initial analysis model perform a preliminary analysis of these instances with standard answers. Next, each second feature vector instance is identified based on the feature recognition component that has been trained, and the second symptom inference result corresponding to each second chronic disease symptom instance is obtained, that is, the final judgment of the initial analysis model on the instance. After that, the server performs integrated optimization training on the initial analysis model based on the chronic disease target value of each second chronic disease symptom instance (that is, the standard answer annotated by the doctor) and the second symptom inference result given by the model. For example, if the initial analysis model determines that a hypertension instance is first-level hypertension, but the doctor annotates it as second-level hypertension, the server will adjust the parameters of the model to allow the model to more accurately identify similar situations. After such an optimization training process for a large number of second chronic disease instances, the server finally obtained a trained chronic disease recognition model for the disease analysis indicators (such as hypertension diagnosis and classification), which became more accurate and reliable in identifying related chronic disease symptoms. When a new patient comes to the hospital for treatment, the doctor records the patient's chronic disease-related symptom description, forms chronic disease symptom information to be analyzed, and uploads it to the server. For example, the new patient describes "I am often thirsty and weak recently, and the blood sugar value is 11mmol / L after physical examination". In the early stage, multiple heterogeneous models (such as physical sign interaction network and physical sign fusion network) have been used to integrate and train the first chronic disease instance to obtain an initial analysis model with preliminary analysis capabilities. On this basis, the second chronic disease instance with the chronic disease target value marked is obtained, which is like a "standard answer" set. During multi-model integrated optimization, the second chronic disease instance is first converted into a model-processable format (such as generating a symptom instance topology), processed by the initial analysis model to obtain the second feature vector instance, and then the second symptom inference result is obtained through the feature recognition component. Subsequently, the parameters of the initial analysis model are adjusted according to the difference between the chronic disease target value and the inference result. For example, in the case of hypertension, if the model's judgment of the classification does not match the doctor's annotation, the parameters are adjusted. Through a large number of such optimization trainings, the advantages of multiple models are integrated to improve the model's recognition accuracy of chronic disease symptoms, and finally obtain a trained, more accurate and reliable chronic disease recognition model. After the server obtains these chronic disease symptoms to be analyzed, it immediately calls the trained chronic disease recognition model. The model recognizes and processes the chronic disease symptoms to be analyzed.During this process, the model will analyze factors such as the relationship between symptoms and similarities with past instances. Ultimately, the server obtains the chronic disease data analysis results of the chronic disease symptoms to be analyzed, such as determining that the patient may have diabetes, and gives relevant risk assessments or further examination recommendations, providing strong support for the doctor's diagnosis.
[0022] In the embodiments of the present invention, the following implementation modes are also provided.
[0023] Obtaining a first chronic disease instance, wherein the first chronic disease instance includes a plurality of first chronic disease symptom instances; Determine a first chronic disease symptom instance topology of each first chronic disease symptom instance according to the symptom description information of each first chronic disease symptom instance, and determine a comparative learning feature encoding of each first chronic disease symptom instance according to the topology of each first chronic disease symptom instance; Performing recognition processing on each first chronic disease symptom instance topology based on a preset basic model to obtain a first feature vector instance corresponding to each first chronic disease symptom instance topology; Identify each first feature vector instance based on the trained feature recognition component to obtain a first symptom inference result corresponding to each first chronic disease symptom instance; The preset basic model is integrated and optimized for training based on the comparative learning feature coding of each first chronic disease symptom instance and the first symptom inference result to obtain a trained initial analysis model.
[0024] In an embodiment of the present invention, illustratively, when processing chronic disease data, the server first obtains a first chronic disease instance. These instances are derived from the results generated by entity recognition and relationship extraction of unstructured medical texts by natural language processing technology, and each first chronic disease instance contains multiple first chronic disease symptom instances. For example, in the data of a hospital, a first chronic disease instance may cover multiple symptoms of a patient with hypertension, such as dizziness, headache, palpitations, etc., which constitute different first chronic disease symptom instances. Then, the server determines the first chronic disease symptom instance topology according to the symptom description information of each first chronic disease symptom instance. Taking the symptom instance of "dizziness" as an example, the server will analyze the relevant information, such as the frequency of dizziness, duration, whether it is accompanied by other symptoms, etc., and construct this information into a topological structure to show the relationship between the symptom and its related factors. After that, the comparative learning feature encoding is determined according to this topology. The server first obtains multiple pre-set symptom topology patterns, which are symptom combination structures with clinical diagnostic significance summarized by medical experts based on a large amount of clinical experience. The server associates the "dizziness" symptom instance topology with these patterns. For example, it is found that the dizziness symptom is highly correlated with the pattern of "frequent dizziness and accompanied by elevated blood pressure", thereby obtaining the correlation coefficient parameter, and then determining the comparative learning feature encoding, which highlights the correlation characteristics of the symptom with the specific pattern. Then, the server identifies and processes each first chronic disease symptom instance topology based on the preset basic model. The preset basic model includes multiple sign interaction networks and at least one sign fusion network. Taking the "headache" symptom instance topology as an example, the server obtains the sign feature vector of each symptom entity (such as headache intensity, onset site, etc.), the associated sign entity (such as whether it is related to fatigue), and the association strength vector of the pathological association relationship between them according to the description of the disease. Through a series of operations, such as feature conversion of the sign feature vector, fusion of the basic sign vector and the association strength vector of the associated sign entity, etc., the first feature vector instance corresponding to the symptom instance topology is finally obtained, and this vector instance comprehensively reflects the multi-faceted characteristics of the symptom. After that, the server identifies each first feature vector instance based on the feature recognition component that has completed the training. This component acts like a professional diagnostic assistant, analyzing the first feature vector instance. For example, for the first feature vector instance corresponding to the "palpitations" symptom, it identifies and gives the first symptom inference result to determine the disease tendency or severity that the palpitations symptom may correspond to. Finally, the server integrates and optimizes the preset basic model based on the comparative learning feature encoding and the first symptom inference result of each first chronic disease symptom instance.For example, for the symptom of "chest tightness", if the comparative learning feature encoding shows that it is related to a typical symptom pattern, but the first symptom inference result deviates from expectations, the server will adjust the parameters of the preset basic model. After repeated optimization of a large number of first chronic disease symptom instances, the server will finally obtain the initial analysis model that has completed training, making its analysis of chronic disease symptoms more accurate.
[0025] In the embodiment of the present invention, the determining of the comparative learning feature encoding of each first chronic disease symptom instance according to the topology of each first chronic disease symptom instance can be implemented through the following examples.
[0026] Acquire multiple preset symptom topology patterns, wherein each symptom topology pattern corresponds to a symptom combination structure with clinical diagnostic significance; Associating each first chronic disease symptom instance topology with the multiple symptom topology patterns to obtain an association coefficient parameter of each first chronic disease symptom instance topology; According to the correlation coefficient parameter corresponding to the topology of each first chronic disease symptom instance, the comparative learning feature coding of each first chronic disease symptom instance is determined.
[0027] In an embodiment of the present invention, exemplarily, when the server processes the topology of the first chronic disease symptom instance to determine the comparative learning feature coding, it first obtains a plurality of pre-set symptom topology patterns. These patterns are formulated by experts in the medical field in combination with long-term clinical practice and research results, and each pattern corresponds to a symptom combination structure with clinical diagnostic significance. For example, in the symptom topology patterns related to cardiovascular diseases, there may be a combination pattern of "chest pain + palpitations + aggravation after exercise", which is of great significance for judging cardiovascular diseases such as coronary heart disease; there is also a pattern of "dizziness + abnormally high blood pressure + headache", which is closely related to hypertensive emergencies.
[0028] Next, the server associates each first chronic disease symptom instance topology with these pre-set multiple symptom topology patterns, thereby obtaining the correlation coefficient parameters of each first chronic disease symptom instance topology. Suppose there is currently a first chronic disease symptom instance topology described as "a patient often feels palpitations, and the symptoms worsen after fatigue, accompanied by mild chest pain." The server compares this topology with each symptom topology pattern. When associated with the "chest pain + palpitations + aggravation after exercise" pattern, because "palpitations" are similar to "palpitations", "aggravated after fatigue" and "aggravated after exercise" are related, and "mild chest pain" also meets some of the characteristics of this pattern, a higher correlation coefficient is calculated by the algorithm; when compared with other irrelevant patterns (such as symptom topology patterns for respiratory diseases), the correlation coefficient is lower. In this way, the server finds the most relevant symptom topology pattern for the first chronic disease symptom instance topology and determines its correlation coefficient parameters.
[0029] Finally, the server determines the comparative learning feature coding of each first chronic disease symptom instance according to the correlation coefficient parameter corresponding to each first chronic disease symptom instance topology. Continuing with the above example, after the server obtains the correlation coefficient parameter associated with the "chest pain + palpitations + aggravation after exercise" mode, it first extracts the core feature coding of the mode, that is, the coding representing the key symptom features of "chest pain, palpitations, and aggravation after exercise". Then, the control feature coding that is not associated with the topology of the target chronic disease symptom instance in multiple symptom topology patterns is determined, such as the feature coding of other respiratory disease symptom topology patterns. The server performs feature coupling processing on the core feature coding and the control feature coding, and combines the coding representing the key features of the relevant pattern with the coding representing the features of the irrelevant pattern with a specific algorithm, thereby generating the comparative learning feature coding of the first chronic disease symptom instance. This coding not only highlights the correlation between the symptom instance and the specific clinical diagnosis pattern, but also strengthens its unique characteristics by comparing with the irrelevant pattern, providing more targeted information for subsequent analysis and diagnosis.
[0030] In the embodiment of the present invention, associating each first chronic disease symptom instance topology with the multiple symptom topology patterns to obtain the association coefficient parameter of each first chronic disease symptom instance topology can be implemented through the following examples.
[0031] Associating the target first chronic disease symptom instance topology with each symptom topology pattern to obtain a correlation coefficient corresponding to each symptom topology pattern; The symptom topology pattern whose correlation coefficient exceeds the correlation coefficient threshold is used as the target symptom topology pattern corresponding to the target first chronic disease symptom instance topology; The target symptom topology pattern corresponding to the target first chronic disease symptom instance topology is used as the correlation coefficient parameter of the target first chronic disease symptom instance topology.
[0032] In an embodiment of the present invention, illustratively, when the server processes the association between the first chronic disease symptom instance topology and the symptom topology pattern, it takes a specific target first chronic disease symptom instance topology as an example to carry out the operation. Assume that this target first chronic disease symptom instance topology describes the symptoms of a patient: "The patient has frequently felt dizzy recently, accompanied by tinnitus, and the dizziness worsens when standing up." First, the server associates this target first chronic disease symptom instance topology with each pre-set symptom topology pattern to obtain the correlation coefficient corresponding to each symptom topology pattern. For example, among the many symptom topology patterns, there is a "dizziness + increased blood pressure + headache" pattern for hypertension, a "tinnitus + hearing loss + ear pain" pattern for ear diseases, and a "dizziness + fatigue + sudden drop in blood pressure when standing up" pattern for postural hypotension. The server uses a specific algorithm to analyze the matching degree of the symptoms in the target instance topology with the symptoms in each pattern, the similarity of the logical relationship between the symptoms, and other factors to calculate the correlation coefficient. For the "dizziness + high blood pressure + headache" pattern, since the target instance has the symptom of "dizziness" but does not mention high blood pressure and headache, the calculated correlation coefficient is relatively low; for the "tinnitus + hearing loss + ear pain" pattern, although there is a symptom of "tinnitus", there is a lack of descriptions of hearing loss and ear pain, and the correlation coefficient is not high; and for the "dizziness when standing + fatigue + sudden drop in blood pressure" pattern, "dizziness when standing" is a complete match, and "dizziness" and "fatigue" often occur together in some cases of postural hypotension, so the correlation coefficient corresponding to this pattern is relatively high. Next, the server will set a correlation coefficient threshold, which is based on medical experience and a large amount of data statistics, and is used to filter out the pattern that is most relevant to the target instance topology. The server uses the symptom topology pattern with a correlation coefficient exceeding the correlation coefficient threshold as the target symptom topology pattern corresponding to the target first chronic disease symptom instance topology. Assuming that the correlation coefficient threshold is set to 0.6, in the above calculation, the correlation coefficient of the "dizziness + fatigue + sudden drop in blood pressure when standing" mode reaches 0.7, which exceeds the threshold, while the correlation coefficients of other modes have not reached it, then the "dizziness + fatigue + sudden drop in blood pressure when standing" mode is determined to be the target symptom topology mode. Finally, the server uses the target symptom topology mode corresponding to the target first chronic disease symptom instance topology as the correlation coefficient parameter of the target first chronic disease symptom instance topology. This is because the target symptom topology mode best represents the characteristic tendency of the current target first chronic disease symptom instance topology. Subsequent operations such as comparative learning feature encoding based on this correlation coefficient parameter (i.e., the target symptom topology mode) can provide more targeted and accurate information for analyzing the target first chronic disease symptom instance, which helps to more accurately judge the patient's condition and possible chronic diseases.
[0033] In the embodiment of the present invention, determining the comparative learning feature coding of each first chronic disease symptom instance according to the correlation coefficient parameter corresponding to the topology of each first chronic disease symptom instance can be implemented through the following example.
[0034] Obtaining the core feature code of the target symptom topology pattern corresponding to the target first chronic disease symptom instance topology; Determining, according to the core feature code, a control feature code in the plurality of symptom topology patterns that is not associated with the target chronic disease symptom instance topology; The core feature code and the control feature code are subjected to feature coupling processing to obtain a comparative learning feature code of the target first chronic disease symptom instance topology.
[0035] In the embodiment of the present invention, illustratively, taking the server processing a specific first chronic disease symptom instance topology as an example, assuming that the target first chronic disease symptom instance topology is described as "the patient often feels chest tightness, aggravated after activity, and occasionally accompanied by palpitations", the server has determined that its corresponding target symptom topology pattern is "chest tightness + aggravated after activity + palpitations", which has a high correlation with coronary heart disease. First, the server obtains the core feature coding of the target symptom topology pattern corresponding to the target first chronic disease symptom instance topology. For the target symptom topology pattern of "chest tightness + aggravated after activity + palpitations", its core feature coding is a set of code information that can represent these three key symptoms and their mutual relationship. According to the pre-set coding rules, the server converts "chest tightness", "aggravated after activity" and "palpitations" into specific feature codes respectively, and combines them into core feature codes according to their logical relationship in the pattern. For example, "chest tightness" corresponds to code A, "aggravated after activity" corresponds to code B, and "palpitations" corresponds to code C. The core feature coding may be arranged in the order of ABC, and has metadata indicating the association relationship between them. Next, the server determines the control feature codes that are not associated with the target chronic disease symptom instance topology in multiple symptom topology patterns based on the core feature codes. The server traverses all pre-set symptom topology patterns to find patterns that are significantly different from the current target symptom topology pattern. For example, the "cough + sputum + dyspnea" symptom topology pattern for respiratory diseases has a low correlation with the current target pattern related to coronary heart disease. The server also converts "cough", "sputum" and "dyspnea" into feature codes according to the coding rules, assuming that they are D, E, and F respectively, to form the control feature code DEF and related associated metadata. Finally, the server performs feature coupling processing on the core feature code and the control feature code to obtain the comparative learning feature code of the target first chronic disease symptom instance topology. The server uses a specific algorithm to fuse the core feature code ABC and the control feature code DEF. For example, through a weighted fusion method, each code in the core feature code is given a higher weight, and the code in the control feature code is given a lower weight, and then combined in a certain order and mathematically operated to generate a new code. The assumed operation rule is to perform weighted addition of the code values at the corresponding positions to obtain a new set of values, which are then normalized to ultimately form the comparative learning feature encoding of the topology of the target first chronic disease symptom instance. This encoding not only highlights the core features that are closely related to the target symptom topology pattern, but also reflects the difference from other unrelated patterns through the comparative feature encoding, so that subsequent analysis can more accurately distinguish the symptom instance from other situations, and the auxiliary server can perform more accurate chronic disease analysis and diagnosis based on this encoding.
[0036] In an embodiment of the present invention, the preset basic model includes multiple vital sign interaction networks and at least one vital sign fusion network. The recognition and processing of each first chronic disease symptom instance topology based on the preset basic model to obtain the first feature vector instance corresponding to each first chronic disease symptom instance topology can be implemented through the following example.
[0037] According to the symptom description information of the target chronic disease symptom instance, obtaining the physical sign feature vector of each symptom entity in the topology of the target first chronic disease symptom instance, the associated physical sign entity of each symptom entity, and the association strength vector of the pathological association relationship between each symptom entity and the associated physical sign entity; Performing a feature conversion operation on the physical sign feature vector of each symptom entity to obtain a basic physical sign vector of each symptom entity; Based on the multiple sign interaction networks, the basic sign vector of each symptom entity, the basic sign vector of the associated sign entity and the associated strength vector are aggregated to obtain a feature vector of each symptom entity; The feature vector of each symptom entity is processed based on the physical sign fusion network to obtain a first feature vector instance corresponding to the target first chronic disease symptom instance topology.
[0038] In the embodiment of the present invention, for example, it is assumed that the server is processing a chronic disease symptom instance of a patient, and the target chronic disease symptom instance is described as "the patient has suffered from hypertension for a long time, and has recently developed dizziness symptoms, and the blood pressure value rises significantly when dizzy, accompanied by occasional palpitations". The server identifies and processes the topology of the target first chronic disease symptom instance based on the preset basic model to obtain the corresponding first feature vector instance. For the symptom entity "dizziness", the server extracts relevant information from the disease description information. Through the analysis and learning of a large amount of clinical data, the server knows that dizziness may be related to blood pressure values. Therefore, "blood pressure value" is the associated sign entity of the "dizziness" symptom entity. From the description, it can be seen that blood pressure rises significantly when dizzy. The server quantifies this degree of association into an association strength vector based on a specific algorithm, such as [0.8] (the higher the value, the stronger the association). At the same time, for "dizziness" itself, the server converts the characteristics of dizziness (such as the frequency and degree of dizziness) into a sign feature vector based on clinical knowledge and data, assuming that it is [0.6, 0.4], which represent the quantified values of the frequency and degree of dizziness respectively. For the "palpitations" symptom entity, the server analyzes and finds that it may be associated with signs related to heart function. In this case, the associated sign entity can be "heart rate". Since the description only mentions occasional palpitations, the server determines its association strength vector with heart rate, for example, [0.5]. The sign feature vector of "palpitations" itself is quantified as [0.3, 0.7] based on the frequency and intensity of palpitations. "Hypertension" as a symptom entity, its associated sign entity can be "blood pressure control status", and the association strength vector is assumed to be [0.9], because hypertension is closely related to blood pressure control. The sign feature vector of "hypertension" can be quantified as [0.7, 0.5] based on the course of hypertension, blood pressure fluctuation range, etc. For the sign feature vector of "dizziness" [0.6, 0.4], the server uses a preset conversion algorithm, which can be a linear transformation containing a weight matrix. After calculation, the basic sign vector is obtained, which is assumed to become [0.5, 0.6]. This conversion process is to convert the original sign feature vector into a form that is more suitable for subsequent processing and highlight the key features related to disease diagnosis. The physical sign feature vector [0.3, 0.7] of "palpitations" undergoes the same feature conversion operation to obtain a basic physical sign vector, for example, [0.4, 0.8]. The physical sign feature vector [0.7, 0.5] of "hypertension" is converted to obtain a basic physical sign vector, assuming it is [0.8, 0.4]. Taking "dizziness" as an example, multiple physical sign interaction networks will consider the basic physical sign vector [0.5, 0.6] of "dizziness", the basic physical sign vector corresponding to "blood pressure value" (associated physical sign entity) (assuming that the blood pressure value is quantified as [0.7, 0.3] based on some measurement data) and the association strength vector [0.8]. The physical sign interaction network uses a specific aggregation algorithm, which can be a weighted sum combined with a nonlinear transformation.After multiplying each element of the basic sign vector of "dizziness" by the association strength vector, the corresponding elements of the basic sign vector of "blood pressure value" are weighted and added, and then processed by nonlinear function to obtain the feature vector of the "dizziness" symptom entity, which is assumed to be [0.65, 0.55]. For "palpitations", the basic sign vector [0.4, 0.8], the basic sign vector of "heart rate" (associated sign entity) (assumed to be [0.6, 0.2]) and the association strength vector [0.5] are combined, and the feature vector of the "palpitations" symptom entity is obtained through aggregation processing of the sign interaction network, such as [0.45, 0.75]. "Hypertension" is combined with its own basic sign vector [0.8, 0.4], the basic sign vector of "blood pressure control" (associated sign entity) (assumed to be [0.9, 0.1]) and the association strength vector [0.9], and the feature vector is obtained through aggregation processing, which is assumed to be [0.88, 0.35]. The server inputs the feature vectors of the three symptom entities "dizziness" [0.65, 0.55], "palpitations" [0.45, 0.75], and "hypertension" [0.88, 0.35] into the sign fusion network. The sign fusion network will comprehensively consider the relationship between these feature vectors and fuse the features of each symptom entity through a series of operations, such as matrix multiplication and convolution operations. Finally, a unified vector is output, that is, the first feature vector instance corresponding to the topology of the target first chronic disease symptom instance, which is assumed to be [0.6, 0.5, 0.7, 0.4]. This first feature vector instance comprehensively reflects the characteristics of each symptom entity and their mutual relationship in the topology of the target first chronic disease symptom instance, providing a key data basis for subsequent symptom inference and disease analysis.
[0039] In an embodiment of the present invention, the basic sign vector of each symptom entity, the basic sign vector of the associated sign entity and the associated strength vector are aggregated based on the multiple sign interaction networks to obtain the feature vector of each symptom entity, which can be implemented through the following examples.
[0040] Based on the first symptom interaction network, the basic symptom vector and the association strength vector of the associated symptom entity are interactively aggregated to obtain the first-order interactive aggregation vector of each symptom entity; Performing sign aggregation processing on the input sign vector of each symptom entity and the first-order interactive aggregation vector to obtain the first-order sign vector of each symptom entity; Perform interactive aggregation processing on the preceding order sign vector and the association strength vector of the associated sign entity based on the target sign interaction network to obtain the target order interactive aggregation vector of each symptom entity; Performing sign aggregation processing on the preceding order sign vector and the target order interactive aggregation vector of each symptom entity to obtain the target order sign vector of each symptom entity; The target-order sign vector of each symptom entity is used as the feature vector of each symptom entity.
[0041] In the embodiment of the present invention, for example, it is assumed that the server is processing a target first chronic disease symptom instance topology described as "the patient has recently developed headache symptoms, the degree of headache worsens with the increase of blood pressure, accompanied by neck stiffness, and neck stiffness is related to the frequency of headache attacks." For the "headache" symptom entity, its associated sign entity is "blood pressure." Assume that the basic sign vector of "blood pressure" obtained through the previous steps is [0.7, 0.3] (representing the quantitative characteristics of systolic and diastolic blood pressure, respectively), and the correlation strength vector between headache and blood pressure is [0.8] (indicating the degree of correlation between the degree of headache and the increase of blood pressure). The first sign interaction network will be based on a preset algorithm, such as multiplying the correlation strength vector with each element of the "blood pressure" basic sign vector to obtain [0.56, 0.24], which is the first-order interaction aggregation vector of the "headache" symptom entity. This vector reflects the initial impact characteristics of blood pressure on headache based on the correlation strength. For the "stiff neck" symptom entity, its associated sign entity can be regarded as "headache attack frequency." Assume that the basic sign vector of "headache frequency" is [0.6] (simply quantified as a numerical value of the frequency of attacks), and the correlation strength vector between neck stiffness and headache frequency is [0.7]. The first sign interaction network multiplies the two through the same or similar algorithm to obtain [0.42]. This is the first-order interaction aggregation vector of the "stiff neck" symptom entity, reflecting the initial impact of headache frequency on neck stiffness. Taking "headache" as an example, its input sign vector (that is, the basic sign vector) is [0.5, 0.6] (representing the quantitative characteristics of the degree and duration of headache, respectively). The server obtains a specific weight tensor for sign aggregation processing, assuming it is [0.4, 0.6]. The first-order interaction aggregation vector [0.56, 0.24] is linearly transformed with the weight tensor, that is, the corresponding elements are multiplied and then added, and [0.56*0.4+0.24*0.6]=[0.368] is obtained. Then, the input sign vector [0.5, 0.6] and the transformation coefficient tensor [0.368] are processed based on a pre-set nonlinear mapping function (such as the Sigmoid function). The elements of the input sign vector are combined with the transformation coefficient tensor and calculated by the Sigmoid function to obtain the first-order sign vector of the "headache" symptom entity, which is assumed to be [0.65, 0.7]. This first-order sign vector combines the characteristics of headache itself and the initial impact of blood pressure on it. For "stiff neck", its input sign vector is [0.4, 0.7] (representing the quantitative characteristics of the degree and range of stiff neck, respectively), and the weight tensor used for sign aggregation processing is assumed to be [0.3, 0.7]. The first-order interaction aggregation vector [0.42] is linearly transformed with the weight tensor to obtain [0.42*0.3+0.42*0.7]=[0.42]. After being processed by the nonlinear mapping function, the first-order sign vector of the "stiff neck" symptom entity is obtained, such as [0.55, 0.8], which combines the characteristics of stiff neck itself and the initial impact of headache frequency on it.Still taking "headache" as an example, the preceding order sign vector here is the first order sign vector [0.65, 0.7] just obtained, the associated sign entity is still "blood pressure", and the association strength vector is still [0.8]. The target sign interaction network uses a more complex algorithm, such as weighted convolution operation on the preceding order sign vector and the association strength vector, assuming that [0.52, 0.56] is obtained, which is the target order interaction aggregation vector of the "headache" symptom entity. This vector further explores the deeper association characteristics between blood pressure and headache based on the preceding order. For "stiff neck", the preceding order sign vector is [0.55, 0.8], and the associated sign entity is "headache frequency", and the association strength vector is [0.7]. After similar complex processing of the target sign interaction network, it is assumed that [0.44, 0.56] is obtained as the target order interaction aggregation vector of the "stiff neck" symptom entity, which reflects the further association characteristics between headache frequency and stiff neck based on the preceding order. For "headache", the previous order sign vector [0.65, 0.7] and the target order interaction aggregation vector [0.52, 0.56] are processed for sign aggregation. The server obtains the appropriate weight tensor again, assuming it is [0.5, 0.5]. The two vectors and the weight tensor are linearly transformed and nonlinearly mapped (similar to the previous steps) to obtain the target order sign vector of the "headache" symptom entity, assuming it is [0.6, 0.63]. This target order sign vector more comprehensively integrates the characteristics of headache itself and the multi-level impact of blood pressure on it. For "stiff neck", the previous order sign vector [0.55, 0.8] and the target order interaction aggregation vector [0.44, 0.56] are linearly transformed and nonlinearly mapped under the corresponding weight tensor (assuming it is [0.4, 0.6]) to obtain the target order sign vector of the "stiff neck" symptom entity, such as [0.5, 0.7], which integrates the characteristics of stiff neck itself and the multi-faceted impact of headache attack frequency on it. Finally, the feature vector of the "headache" symptom entity is the target-order sign vector [0.6, 0.63] obtained above, and the feature vector of the "stiff neck" symptom entity is [0.5, 0.7]. These feature vectors comprehensively and deeply reflect the complex relationship between each symptom entity and its associated sign entities, providing key feature data for the subsequent comprehensive analysis of the topology of the entire target first chronic disease symptom instance.
[0042] In an embodiment of the present invention, the interactive aggregation processing of the basic sign vector and the associated strength vector of the associated sign entity based on the first sign interaction network to obtain the first-order interactive aggregation vector of each symptom entity can be implemented through the following examples.
[0043] Performing feature fusion on the basic sign vector of the associated sign entity and the associated strength vector to obtain a fusion vector; Acquire a first weight tensor of a first feature aggregation module for interactive aggregation processing, perform a linear transformation operation on the first weight tensor and the fusion vector, and obtain a first transformation coefficient tensor; The first transformation coefficient tensor is processed based on a preset first nonlinear mapping function to obtain a first-order interaction aggregation vector of the symptom entity.
[0044] In an embodiment of the present invention, for example, it is assumed that the server is processing the chronic disease symptom data of a patient, and the symptom description of the patient is "recent joint pain symptoms, the degree of pain is closely related to the amount of activity, accompanied by slight swelling, and the degree of swelling is related to the frequency of pain". For the "joint pain" symptom entity, its associated sign entity is "activity". Assume that the basic sign vector of "activity" obtained in the previous step is [0.6, 0.4], which respectively represent the quantitative values of daily activity time and activity intensity. The correlation strength vector between joint pain and activity is [0.8], which indicates the close correlation between the degree of pain and activity. The server performs feature fusion on the correlation strength vector [0.8] and the basic sign vector [0.6, 0.4] of "activity". A simple fusion method can be to multiply the correlation strength vector with each element of the basic sign vector to obtain the fusion vector [0.48, 0.32]. This fusion vector preliminarily reflects the characteristics of the influence of activity on joint pain based on the correlation strength. For the "swelling" symptom entity, its associated sign entity is "pain frequency". Assume that the basic sign vector of "pain frequency" is [0.7], which is simply quantified as the frequency value of pain attacks. The correlation strength vector of swelling and pain frequency is [0.7]. In the same way, the server multiplies the two to obtain the fusion vector [0.49], which reflects the initial impact of pain frequency on swelling based on the correlation strength. Taking "joint pain" as an example, the server obtains the first weight tensor of the first feature aggregation module for interactive aggregation processing, which is assumed to be [0.5, 0.5]. The first weight tensor [0.5, 0.5] and the fusion vector [0.48, 0.32] are linearly transformed. The linear transformation operation can be the multiplication and addition of the corresponding elements, that is, (0.5*0.48+0.5*0.32)=0.4. The first transformation coefficient tensor obtained is [0.4]. This coefficient tensor comprehensively considers the information of weights and fusion vectors, preparing for subsequent nonlinear processing. For the "swelling" symptom entity, it is assumed that the first weight tensor is [0.6]. It is linearly transformed with the fusion vector [0.49], that is, 0.6*0.49=0.294. The first transformation coefficient tensor obtained is [0.294], which combines the weight and the information after the swelling and pain frequencies are fused. For "joint pain", the server is based on a pre-set first nonlinear mapping function, such as the Sigmoid function. Substitute the first transformation coefficient tensor [0.4] into the Sigmoid function for processing. For example, f(0.4)≈0.6 is calculated. The first-order interaction aggregation vector of the "joint pain" symptom entity is [0.6]. This first-order interaction aggregation vector further adjusts the coefficients obtained by the previous linear transformation through nonlinear mapping, so that it can better reflect the complex correlation characteristics between joint pain and activity. For "swelling", the first transformation coefficient tensor [0.294] is also processed using the Sigmoid function.Calculate f(0.294)≈0.573. The first-order interaction aggregation vector of the "swelling" symptom entity is [0.573]. This vector comprehensively reflects the nonlinearly adjusted correlation characteristics between swelling and pain frequency, and provides a key intermediate result for subsequent more in-depth sign aggregation processing. Through such a series of operations, the server generates a first-order interaction aggregation vector for each symptom entity. These vectors play a connecting role in the entire sign analysis process, gradually exploring and integrating the complex relationship between symptoms and related signs, and providing strong support for accurate analysis of chronic disease symptoms.
[0045] In the embodiment of the present invention, the input sign vector of each symptom entity and the first-order interaction aggregation vector are subjected to sign aggregation processing to obtain the first-order sign vector of each symptom entity, which can be implemented through the following examples.
[0046] Obtaining a second weight tensor of a second feature aggregation module for performing vital sign aggregation processing; Performing a linear transformation operation on the second weight tensor and the first-order interaction aggregation vector to obtain a second transformation coefficient tensor; The input sign vector of each symptom entity and the second transformation coefficient tensor are processed based on a preset second nonlinear mapping function to obtain the first-order sign vector of each symptom entity.
[0047] In an embodiment of the present invention, exemplarily, continuing with the above-mentioned case of "the patient has recently developed joint pain symptoms, the degree of pain is closely related to the amount of activity, and is accompanied by slight swelling, the degree of swelling is related to the frequency of pain" as an example, the server then performs sign aggregation processing on the input sign vector and the first-order interactive aggregation vector of each symptom entity to obtain the first-order sign vector of each symptom entity. For the "joint pain" symptom entity, the server obtains the second weight tensor of the second feature aggregation module for sign aggregation processing according to the system preset or the parameters obtained based on a large amount of data training. Assuming that this second weight tensor is [0.3, 0.7], this weight tensor reflects the importance distribution of different feature information in the process of sign aggregation. For example, the weight setting here means that in the aggregation process, some features related to pain (corresponding to the part with a weight of 0.7) may be more valued, while another part of the features (corresponding to the part with a weight of 0.3) may be relatively lightly considered. This is based on the experience and algorithm setting of the symptom analysis of joint pain. For the "swelling" symptom entity, it is assumed that the second weight tensor obtained is [0.4, 0.6]. Similarly, this weight distribution reflects the difference in the degree of importance attached to different related features when processing swelling sign aggregation. It may focus more on features related to the degree of swelling (weight 0.6), while paying relatively less attention to other related features (weight 0.4). Taking "joint pain" as an example, the first-order interaction aggregation vector of "joint pain" obtained before is [0.6] (assuming it is obtained after the previous steps). The second weight tensor [0.3,0.7] is linearly transformed with the first-order interaction aggregation vector [0.6]. Since the first-order interaction aggregation vector has only one value, it can be multiplied and added with each element of the weight tensor (similar to the simplified form of matrix multiplication), that is, (0.3*0.6+0.7*0.6)=0.6. In this way, the second transformation coefficient tensor [0.6] is obtained. This second transformation coefficient tensor integrates the importance information represented by the first-order interaction aggregation vector and the weight tensor, providing a basis for subsequent further processing. For the "swelling" symptom entity, its first-order interaction aggregation vector is assumed to be [0.573]. The second weight tensor [0.4, 0.6] is linearly transformed with it, that is, (0.4*0.573+0.6*0.573)=0.573. The obtained second transformation coefficient tensor is [0.573], which combines the swelling first-order interaction aggregation vector and the feature importance information represented by the weight tensor. For "joint pain", assume that its input sign vector is [0.5, 0.8], which represents the quantitative values of the degree and duration of joint pain respectively. The server is based on a pre-set second nonlinear mapping function, such as the ReLU function (RectifiedLinearUnit, f(x)=max(0,x)). Each element of the input sign vector is combined with the second transformation coefficient tensor [0.6].First, the first element 0.5 and 0.6 of the input sign vector are operated (for example, added) to obtain 0.5+0.6=1.1, and then processed by the ReLU function, f(1.1)=1.1. The same operation is performed on the second element 0.8, 0.8+0.6=1.4, f(1.4)=1.4. In this way, the first-order sign vector of the "joint pain" symptom entity is obtained as [1.1,1.4]. This first-order sign vector not only contains the input sign information of the joint pain itself, but also integrates the first-order interactive aggregation information related to the activity amount, and is adjusted through nonlinear mapping to more comprehensively and accurately reflect the comprehensive sign characteristics of joint pain. For the "swelling" symptom entity, its input sign vector is assumed to be [0.4,0.7], representing the range of swelling and the quantitative value of hardness, respectively. Combine the input sign vector with the second transformation coefficient tensor [0.573]. For the first element 0.4+0.573=0.973, after ReLU function processing, f(0.973)=0.973. For the second element 0.7+0.573=1.273, f(1.273)=1.273. The first-order sign vector of the "swelling" symptom entity is [0.973,1.273]. This first-order sign vector combines the original signs of swelling, the first-order interactive aggregation information related to pain frequency, and the adjustment after nonlinear mapping, providing a more representative feature vector for further analysis of swelling symptoms. Through this series of operations, the server generates a first-order sign vector for each symptom entity. These vectors further refine and integrate the sign information in the process of chronic disease symptom analysis, which helps to analyze the patient's condition more accurately.
[0048] In an embodiment of the present invention, the method of obtaining the sign feature vector of each symptom entity in the topology of the target first chronic disease symptom instance, the associated sign entity of each symptom entity, and the association strength vector of the pathological association relationship between each symptom entity and the associated sign entity based on the disease description information of the target chronic disease symptom instance can be implemented through the following examples.
[0049] Acquire clinical indicator data of each sub-symptom instance in the target first chronic disease symptom instance according to the symptom description information of the target first chronic disease symptom instance; Performing feature coupling processing on the clinical indicator data of each sub-symptom instance to obtain a physical sign feature vector of each symptom entity in the target first chronic disease symptom instance topology; Taking the control sub-symptom having a clinical association path with each sub-symptom instance as the associated sub-symptom of each sub-symptom instance; Acquire clinical indicator data of a clinical association path between each sub-symptom instance and the associated sub-symptom; The clinical indicator data of the clinical association pathway is subjected to feature coupling processing to obtain an association strength vector of the pathological association relationship between each symptom entity and the associated sign entity.
[0050] In the embodiment of the present invention, for example, it is assumed that the server is processing a patient's chronic disease symptom description: "The patient has suffered from diabetes for many years, and recently developed polydipsia and polyphagia symptoms, with large blood sugar fluctuations and weight loss. There is a certain correlation between weight loss and blood sugar fluctuations." For the sub-symptom instance of "polydipsia", the server obtains clinical indicator data from the symptom description information and related medical records. For example, it is found that the patient's daily water intake reaches 3000ml. This specific value is an important clinical indicator data of the "polydipsia" sub-symptom instance. For the "polyphagia" sub-symptom instance, it is obtained that the patient's recent meal intake has increased by 50% compared with the previous one. This food intake change ratio is the clinical indicator data of "polyphagia". For "large blood sugar fluctuations", the server obtains that the patient's fasting blood sugar value fluctuates between 6-10mmol / L in the past week. This group of blood sugar value ranges is the clinical indicator data of the sub-symptom instance. Regarding "weight loss", it is learned that the patient has lost 5kg in weight in the past month. This 5kg weight change is the clinical indicator data of the "weight loss" sub-symptom instance. For "excessive drinking", the server converts the daily water intake of 3000ml into a feature vector based on the preset feature coupling algorithm. Assume that the algorithm compares the water intake with the normal range, and combines the common water intake characteristics of diabetic patients to generate a physical sign feature vector [0.8, 0.2]. Among them, 0.8 may indicate the degree of exceeding the normal water intake, and 0.2 indicates the degree of conformity with the typical polydipsia characteristics of diabetes. For "excessive eating", the data of increasing the amount of food per meal by 50% is converted into a physical sign feature vector through the feature coupling algorithm, such as [0.7, 0.3]. 0.7 reflects the increase in food intake, and 0.3 reflects the closeness of the association with the polyphagia symptoms of diabetes. The blood sugar value range of "large blood sugar fluctuations" is 6-10mmol / L. After feature coupling processing, the physical sign feature vector [0.9, 0.1] may be obtained. 0.9 highlights the degree of blood sugar fluctuations beyond the normal range, and 0.1 indicates the consistency with the blood sugar fluctuation characteristics of diabetes. The 5kg weight change of "weight loss" is generated through the feature coupling algorithm to generate a sign feature vector [0.6, 0.4], where 0.6 represents the weight loss rate and 0.4 represents the degree of fit with the symptoms of diabetes weight loss. For the "polydipsia" sub-symptom example, based on medical knowledge and a large amount of case data, the server identifies that "large blood sugar fluctuations" have a clinical association path with it, because the increased blood sugar in diabetic patients stimulates the thirst center and causes polydipsia. Therefore, "large blood sugar fluctuations" are the associated sub-symptoms of "polydipsia". "Poor eating" also has a clinical association path with "large blood sugar fluctuations". In the state of high blood sugar, the body's cells cannot get enough energy, which stimulates the appetite center and causes polydipsia. Therefore, "large blood sugar fluctuations" are the associated sub-symptoms of "polydipsia"."Weight loss" and "large blood sugar fluctuations" also have clinical association paths. Large blood sugar fluctuations lead to metabolic disorders in the body, increased decomposition of fat and protein, and thus cause weight loss, so "large blood sugar fluctuations" are the associated sub-symptoms of "weight loss". For the clinical association path of "excessive drinking" and "large blood sugar fluctuations", the server obtains data on the significant increase in water intake when the patient's blood sugar rises. For example, when blood sugar rises from 7mmol / L to 9mmol / L, the amount of water intake increases from 2500ml to 3500ml. This set of corresponding change data of blood sugar value and water intake is the clinical indicator data of their clinical association path. In the clinical association path of "excessive eating" and "large blood sugar fluctuations", the server obtains data on the increase in food intake when the patient's blood sugar rises. For example, when blood sugar rises from 6mmol / L to 8mmol / L, the amount of food per meal increases from 200g to 300g. This set of data on changes in blood sugar value and food intake is the clinical indicator data of this clinical association path. In terms of the clinical association path of "weight loss" and "large blood sugar fluctuations", data on the speed of weight loss during blood sugar fluctuations are obtained. For example, within half a month when blood sugar fluctuates between 8-10mmol / L, the weight decreases by 3kg. This set of data on blood sugar fluctuation range and weight loss is the clinical indicator data of the clinical association path. For the clinical association path data of "drinking more" and "large blood sugar fluctuation", the server processes it through the feature coupling algorithm. Taking into account the relationship between the increase in blood sugar and the increase in water intake, an association strength vector [0.7] is generated. This value indicates a strong degree of association between drinking more and blood sugar fluctuation. After feature coupling processing, the clinical association path data of "eating more" and "large blood sugar fluctuation" generates an association strength vector [0.6], which reflects a certain degree of association between eating more and blood sugar fluctuation. After feature coupling processing, the clinical association path data of "weight loss" and "large blood sugar fluctuation" obtains an association strength vector [0.8], which indicates that the degree of association between weight loss and blood sugar fluctuation is high. Through these steps, the server comprehensively and meticulously obtains the physical sign feature vector of each symptom entity in the target first chronic disease symptom instance topology, the associated physical sign entities, and the association strength vector of the pathological association relationship between them, providing a key data foundation for subsequent analysis and processing based on the preset basic model.
[0051] In an embodiment of the present invention, the model optimization of the pre-trained initial analysis model based on the second chronic disease instance to obtain the trained chronic disease recognition model corresponding to the disease analysis index can be implemented through the following examples.
[0052] Obtain a second chronic disease symptom instance topology for each second chronic disease symptom instance in the second chronic disease instance; Based on the pre-trained initial analysis model, each second chronic disease symptom instance topology is identified and processed to obtain a second feature vector instance corresponding to each second chronic disease symptom instance topology; Identify each second feature vector instance based on the trained feature recognition component to obtain a second symptom inference result corresponding to each second chronic disease symptom instance; The pre-trained initial analysis model is integrated and optimized according to the chronic disease target value of each second chronic disease symptom instance and the second symptom inference result to obtain a trained chronic disease recognition model corresponding to the disease analysis index.
[0053] In an embodiment of the present invention, illustratively, it is assumed that the server is processing data analysis of chronic diseases related to diabetes, and has obtained a batch of second chronic disease instances, which have been annotated with chronic disease target values by professional doctors based on clinical diagnostic standards and experience. Taking a specific second chronic disease instance as an example, the instance is described as "the patient has a history of diabetes for many years, has recently developed symptoms of blurred vision, and the blood sugar level continues to be between 10-12mmol / L, the glycated hemoglobin index is 8%, and the body feels obvious fatigue." The server constructs a second chronic disease symptom instance topology based on the disease description information. For the "blurred vision" symptom, the server analyzes related factors, such as the frequency and degree of blurred vision, whether there is a temporal correlation with changes in blood sugar levels, etc., and constructs this information into a topological structure. For "blood sugar levels continue to be between 10-12mmol / L", the blood sugar level range, fluctuations, and the relationship with other symptoms are recorded to form a topological structure. Similarly, similar processing is performed on "Glycated hemoglobin index is 8%" and "physical fatigue is obvious", and the topological structure of each second chronic disease symptom instance is constructed respectively. These topological structures show the relationship between each symptom and its related factors, forming the second chronic disease symptom instance topology. The server uses the initial analysis model that has been trained in advance to identify and process each second chronic disease symptom instance topology constructed above. For example, for the symptom instance topology of "blurred vision", the initial analysis model analyzes each factor in the topology based on its internal algorithm and parameters, such as quantifying the degree of blurred vision into a numerical value, combining its association with the change of blood sugar value and other information, and through a series of calculations and conversions, generates a corresponding feature vector, assuming it is [0.7, 0.3, 0.1]. This vector comprehensively reflects the characteristics of the "blurred vision" symptom and its related factors, that is, the second feature vector instance corresponding to the "blurred vision" symptom instance topology. Similarly, for the symptom instance topology of "blood sugar value continues to be between 10-12mmol / L", the model analyzes the characteristics of blood sugar value range, fluctuation, etc., and generates the corresponding second feature vector instance, such as [0.8, 0.2]. The symptom instance topologies of "glycosylated hemoglobin index is 8%" and "obvious physical fatigue" are processed similarly to obtain the corresponding second feature vector instances. The server uses the feature recognition component that has been trained to identify each second feature vector instance. For example, for the second feature vector instance [0.7, 0.3, 0.1] of "blurred vision", the feature recognition component determines the degree of correlation between the symptom represented by the feature vector and diabetic retinopathy based on the knowledge and patterns learned in its training, and gives the second symptom inference result, such as "the blurred vision symptom is moderately correlated with diabetic retinopathy."For the second feature vector instance [0.8, 0.2] of "blood sugar value is continuously between 10-12mmol / L", the feature recognition component determines the degree of influence of the current blood sugar value on the control of diabetes, and obtains the second symptom inference result of "poor blood sugar control". The second feature vector instances corresponding to "glycosylated hemoglobin index is 8%" and "obvious physical fatigue" are also identified, and the corresponding second symptom inference results are given respectively. It is known that each second chronic disease symptom instance has a chronic disease target value annotated by a professional doctor. For example, for the "blurred vision" symptom in the above example, the chronic disease target value annotated by the doctor may be "high risk of diabetic retinopathy". However, the second symptom inference result given by the model is "the blurred vision symptom is moderately related to diabetic retinopathy", which is different from the two. The server adjusts the parameters of the initial analysis model based on this difference. For example, in the algorithm module related to blurred vision and diabetic retinopathy in the model, the relevant weight parameters are adjusted so that the model can more accurately infer the result consistent with the doctor's annotation when it encounters a similar symptom instance topology next time. For other symptoms in the second chronic disease instance, such as "blood sugar level is continuously between 10-12mmol / L", "glycosylated hemoglobin index is 8%", "obvious physical fatigue", etc., similar parameter adjustments are made to the initial analysis model according to the difference between the chronic disease target value and the second symptom inference result. Through such an optimization training process for a large number of second chronic disease instances, the server gradually optimizes the initial analysis model and finally obtains a chronic disease recognition model that has been trained for diabetes disease analysis indicators. This model becomes more accurate and reliable in identifying diabetes-related chronic disease symptoms and analyzing the condition.
[0054] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned chronic disease data analysis method based on natural language processing and integrated training. Figure 2 As shown, Figure 2 The block diagram of the computer device 100 provided in the embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112 and a communication unit 113. To achieve data transmission or interaction, the memory 111, the processor 112 and the communication unit 113 are directly or indirectly electrically connected to each other. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0055] For illustrative purposes, the foregoing description is made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise form disclosed. Numerous modifications and variations are possible in accordance with the above teachings. These embodiments are selected and described in order to best illustrate the principles of the present disclosure and its practical application, so that those skilled in the art can best utilize the present disclosure and utilize various embodiments with different modifications to suit the intended specific application.
Claims
1. A chronic disease data analysis method based on natural language processing and ensemble training, characterized in that: include: Use natural language processing technology to perform entity recognition and relationship extraction on unstructured medical texts to generate the first chronic disease instance without labeled chronic disease target values; Acquire a pre-trained initial analysis model, wherein the pre-trained initial analysis model is obtained by comparative learning based on the first chronic disease instance through an integrated training framework, wherein the integrated training framework integrates feature representations output by multiple heterogeneous models; Acquire a second chronic disease instance corresponding to the disease analysis indicator, where the second chronic disease instance is a chronic disease instance with a chronic disease target value marked; Based on the second chronic disease instance, the pre-trained initial analysis model is optimized by multi-model integration to obtain a trained chronic disease recognition model corresponding to the disease analysis index; The chronic disease symptoms to be analyzed are obtained, and the chronic disease symptoms to be analyzed are identified based on the trained chronic disease recognition model to obtain the chronic disease data analysis results of the chronic disease symptoms to be analyzed.
2. The method according to claim 1, characterized in that The method further comprises: Obtaining a first chronic disease instance, wherein the first chronic disease instance includes a plurality of first chronic disease symptom instances; Determine a first chronic disease symptom instance topology of each first chronic disease symptom instance according to the symptom description information of each first chronic disease symptom instance, and determine a comparative learning feature encoding of each first chronic disease symptom instance according to the topology of each first chronic disease symptom instance; Performing recognition processing on each first chronic disease symptom instance topology based on a preset basic model to obtain a first feature vector instance corresponding to each first chronic disease symptom instance topology; Identify each first feature vector instance based on the trained feature recognition component to obtain a first symptom inference result corresponding to each first chronic disease symptom instance; The preset basic model is integrated and optimized for training based on the comparative learning feature coding of each first chronic disease symptom instance and the first symptom inference result to obtain a trained initial analysis model.
3. The method according to claim 2, characterized in that The step of determining the comparative learning feature encoding of each first chronic disease symptom instance according to the topology of each first chronic disease symptom instance includes: Acquire multiple preset symptom topology patterns, wherein each symptom topology pattern corresponds to a symptom combination structure with clinical diagnostic significance; Associating each first chronic disease symptom instance topology with the multiple symptom topology patterns to obtain an association coefficient parameter of each first chronic disease symptom instance topology; According to the correlation coefficient parameter corresponding to the topology of each first chronic disease symptom instance, the comparative learning feature coding of each first chronic disease symptom instance is determined.
4. The method according to claim 3, characterized in that The step of associating each first chronic disease symptom instance topology with the multiple symptom topology patterns to obtain an association coefficient parameter of each first chronic disease symptom instance topology includes: Associating the target first chronic disease symptom instance topology with each symptom topology pattern to obtain a correlation coefficient corresponding to each symptom topology pattern; The symptom topology pattern whose correlation coefficient exceeds the correlation coefficient threshold is used as the target symptom topology pattern corresponding to the target first chronic disease symptom instance topology; The target symptom topology pattern corresponding to the target first chronic disease symptom instance topology is used as the correlation coefficient parameter of the target first chronic disease symptom instance topology.
5. The method according to claim 4, characterized in that The determining, according to the correlation coefficient parameter corresponding to the topology of each first chronic disease symptom instance, the comparative learning feature coding of each first chronic disease symptom instance comprises: Obtaining the core feature code of the target symptom topology pattern corresponding to the target first chronic disease symptom instance topology; Determining, based on the core feature codes, control feature codes in the plurality of symptom topology patterns that are not associated with the target chronic disease symptom instance topology; The core feature code and the control feature code are subjected to feature coupling processing to obtain a comparative learning feature code of the target first chronic disease symptom instance topology.
6. The method according to claim 2, characterized in that The preset basic model includes a plurality of physical sign interaction networks and at least one physical sign fusion network, and the recognition processing of each first chronic disease symptom instance topology based on the preset basic model to obtain a first feature vector instance corresponding to each first chronic disease symptom instance topology includes: According to the symptom description information of the target chronic disease symptom instance, obtaining the physical sign feature vector of each symptom entity in the topology of the target first chronic disease symptom instance, the associated physical sign entity of each symptom entity, and the association strength vector of the pathological association relationship between each symptom entity and the associated physical sign entity; Performing a feature conversion operation on the physical sign feature vector of each symptom entity to obtain a basic physical sign vector of each symptom entity; Performing feature fusion on the basic sign vector of the associated sign entity and the associated strength vector to obtain a fusion vector; Acquire a first weight tensor of a first feature aggregation module for interactive aggregation processing, perform a linear transformation operation on the first weight tensor and the fusion vector, and obtain a first transformation coefficient tensor; Processing the first transformation coefficient tensor based on a preset first nonlinear mapping function to obtain a first-order interaction aggregation vector of the symptom entity; Performing sign aggregation processing on the input sign vector of each symptom entity and the first-order interactive aggregation vector to obtain the first-order sign vector of each symptom entity; Perform interactive aggregation processing on the preceding order sign vector and the association strength vector of the associated sign entity based on the target sign interaction network to obtain the target order interactive aggregation vector of each symptom entity; Performing sign aggregation processing on the preceding order sign vector and the target order interactive aggregation vector of each symptom entity to obtain the target order sign vector of each symptom entity; Taking the target-order sign vector of each symptom entity as the feature vector of each symptom entity; The feature vector of each symptom entity is processed based on the physical sign fusion network to obtain a first feature vector instance corresponding to the target first chronic disease symptom instance topology.
7. The method according to claim 6, characterized in that The step of performing sign aggregation processing on the input sign vector of each symptom entity and the first-order interaction aggregation vector to obtain the first-order sign vector of each symptom entity includes: Obtaining a second weight tensor of a second feature aggregation module for performing vital sign aggregation processing; Performing a linear transformation operation on the second weight tensor and the first-order interaction aggregation vector to obtain a second transformation coefficient tensor; The input sign vector of each symptom entity and the second transformation coefficient tensor are processed based on a preset second nonlinear mapping function to obtain the first-order sign vector of each symptom entity.
8. The method according to claim 6, characterized in that The step of obtaining, according to the symptom description information of the target chronic disease symptom instance, the sign feature vector of each symptom entity in the topology of the target first chronic disease symptom instance, the associated sign entity of each symptom entity, and the association strength vector of the pathological association relationship between each symptom entity and the associated sign entity, comprises: Acquire clinical indicator data of each sub-symptom instance in the target first chronic disease symptom instance according to the symptom description information of the target first chronic disease symptom instance; Performing feature coupling processing on the clinical indicator data of each sub-symptom instance to obtain a physical sign feature vector of each symptom entity in the target first chronic disease symptom instance topology; Taking the control sub-symptom having a clinical association path with each sub-symptom instance as the associated sub-symptom of each sub-symptom instance; Acquire clinical indicator data of a clinical association path between each sub-symptom instance and the associated sub-symptom; The clinical indicator data of the clinical association pathway is subjected to feature coupling processing to obtain an association strength vector of the pathological association relationship between each symptom entity and the associated sign entity.
9. The method according to claim 1, characterized in that: The step of optimizing the pre-trained initial analysis model based on the second chronic disease instance to obtain a trained chronic disease recognition model corresponding to the disease analysis indicator includes: Obtain a second chronic disease symptom instance topology for each second chronic disease symptom instance in the second chronic disease instance; Based on the pre-trained initial analysis model, each second chronic disease symptom instance topology is identified and processed to obtain a second feature vector instance corresponding to each second chronic disease symptom instance topology; Identify each second feature vector instance based on the trained feature recognition component to obtain a second symptom inference result corresponding to each second chronic disease symptom instance; The pre-trained initial analysis model is integrated and optimized according to the chronic disease target value of each second chronic disease symptom instance and the second symptom inference result to obtain a trained chronic disease recognition model corresponding to the disease analysis index.
10. A server system, characterized in that: The method comprises a server, wherein the server is used to execute the method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Chronic disease client service system and service method
CN107731288A
Medical data processing and system based on migration learning
CN108520780A
Chronic disease course stage analysis system and device and storage medium
CN112037876A
Chronic disease data analysis method and system based on natural language processing and integrated training
CN112287665A
System and method to generate a summary template for summarizing structured medical reports
EP4495805A1
Cited By
Chronic disease medical data acquisition, analysis and management method and cloud platform
CN120853870A
A chronic disease medical data collection, analysis and management method and cloud platform
CN120853870B