Medical diagnosis system through clinical interviews
The medical diagnostic system uses AI to build country-specific disease models for precise diagnosis by integrating demographic and health criteria, addressing symptom conveyance challenges and enhancing telemedicine accuracy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- KANG YOUNG HOON
- Filing Date
- 2026-01-14
- Publication Date
- 2026-07-23
AI Technical Summary
Patients face challenges in accurately conveying symptoms during short consultations, leading to difficulties for doctors in diagnosis, and existing telemedicine systems lack accuracy due to one-sided questionnaires without direct specialist interaction.
A medical diagnostic system using artificial intelligence constructs country-specific disease diagnosis learning models based on prevalence data, incorporating demographic, regional, temporal, and health status criteria, and utilizes a big data server to optimize diagnostic models through feedback loops.
Enables accurate and efficient predictive diagnosis by systematically collecting and analyzing medical history information, allowing quick and precise disease identification even in remote areas.
Smart Images

Figure KR2026000826_23072026_PF_FP_ABST
Abstract
Description
Medical diagnostic system through medical history taking
[0001] The present invention relates to a medical diagnostic system through medical history taking that receives symptoms through medical history taking and provides a predictive diagnosis for a specific disease using artificial intelligence through a diagnostic learning model built by learning from diagnostic information on various diseases.
[0002] Nowadays, patients visiting hospitals often find it difficult to accurately convey their symptoms during a short consultation of about 5 to 10 minutes. Furthermore, it is even more challenging for doctors to accurately identify the condition when the patient lacks expressive ability or medical knowledge. Additionally, there is a problem of significant time loss, as patients tend to prefer large hospitals and a proper appointment system is not in place, leading them to wait 2 to 3 hours for a short consultation of about 5 to 10 minutes.
[0003] Meanwhile, with the recent development of the Internet, telemedicine services utilizing it have been developed, and medical services such as diagnosis, treatment, and consultation are being actively provided online.
[0004] However, since the medical history was not based on a direct conversation with a specialist but rather on one-sided responses to pre-written questionnaires asking about the patient's symptoms, there were inevitably limitations to the accuracy of the diagnosis.
[0005] To address these issues, online medical questionnaire technology is emerging that enables more accurate and efficient questionnaire creation by utilizing artificial intelligence to analyze user input and generate customized questionnaires.
[0006] Furthermore, to improve the accuracy of online medical questionnaires, it is necessary to construct country-specific disease diagnosis learning models based on prevalence data according to various national demographic, regional, temporal, or health status criteria, and to provide predicted diagnostic information using artificial intelligence.
[0007] The objective of the present invention is to provide a medical diagnostic system through medical history taking that receives systematically and objectively prepared medical history information and additional symptoms based on the patient's personal constitution or personality, constructs a national disease diagnosis learning model based on prevalence data according to national demographic, regional, temporal, or health status criteria, and provides predictive diagnostic information for specific diseases using artificial intelligence.
[0008] A medical diagnosis system through medical history taking according to one embodiment of the present invention may include: a big data server that collects and stores medical learning data in real time from multiple network servers in multiple countries and learns the medical learning data to build a country-specific disease diagnosis learning model based on the prevalence rate of each disease in each country; a user terminal that communicates with the big data server via a wired or wireless network to provide a medical diagnosis service through medical history taking and has a dedicated program installed to provide feedback on actual diagnosis results received by a user from a specialist; and a diagnosis information providing server that receives pre-prepared medical history questions and additional medical history questions personally added by the user through the dedicated program installed on the user terminal using the country-specific disease diagnosis learning model built by the big data server, infers predictive diagnosis information corresponding to the medical history questions using artificial intelligence, and provides the inferred predictive diagnosis information back to the dedicated program installed on the user terminal.
[0009] The above big data server can receive feedback on the above predictive diagnostic information and the above actual diagnostic results, compare them, and retrain and optimize the above country-specific disease diagnosis learning model by adding the above additional medical history items.
[0010] According to one embodiment, the big data server collects the medical learning data in real time from a public medical database provided by a domestic or overseas government or public institution, a medical institution database storing medical history data and diagnostic data from a hospital or medical institution, a medical academic database storing medical academic paper data, or a medical data sharing platform providing medical data, and the country-specific disease diagnosis learning model can be constructed by extracting prevalence data based on country-specific demographic criteria, regional criteria, time criteria, or health status criteria for a specific disease from the medical learning data, and learning the medical learning data based on the prevalence data.
[0011] Here, the aforementioned demographic criteria are classified by gender, age distribution, or socioeconomic status of the nationals of the country concerned; the aforementioned socioeconomic status is classified by income level, education level, or family composition type of the nationals of the country concerned; the aforementioned income level is classified based on the national's Gross National Income (GNI) per capita; the aforementioned education level is classified based on the average highest level of education attained by the nationals of the country concerned; the aforementioned family composition type is classified by referring to the household composition survey or the Population and Housing Census of the nationals of the country concerned; the aforementioned regional criteria are classified by population density, economic criteria, housing type criteria, or environmental criteria according to the administrative divisions of the country concerned; the aforementioned administrative division criteria are classified into provinces, cities, counties, districts, neighborhoods, towns, or townships; the aforementioned population density criteria are classified by the number of residents per area (persons / km²) residing in the aforementioned administrative divisions; the aforementioned economic criteria are classified according to the average income of residents residing in the aforementioned administrative divisions; and the aforementioned housing type criteria are classified according to apartment complexes, residential areas, or commercial areas existing in the aforementioned administrative divisions. Classification, and the above environmental standards are classified according to the industrial area, residential area, tourist area, mountainous area, or coastal area corresponding to the above administrative district standards.
[0012] The above time criteria are classified according to the passage of time by year or season for residents corresponding to each of the above regional criteria, and the above health status criteria are classified according to criteria by underlying disease, lifestyle, or genetic factor; the above criteria by underlying disease are classified according to the prevalence rate of patients with specific underlying diseases in the relevant country, the above criteria by lifestyle are classified according to the prevalence rate based on lifestyle habits regarding smoking, drinking, sleep duration, or exercise duration, and the above criteria by genetic factor can be classified according to the prevalence rate based on the family history of patients in the relevant country.
[0013] When training the above-mentioned big data server for the country-specific disease diagnosis learning model, the server may use a feature vector in which the prevalence data corresponding to each of the above characteristics, which are classified by each characteristic including the patient's age, nationality, gender, place of residence, duration of residence, family history, occupation, status, income level, lifestyle habits, or presence or absence of underlying diseases extracted from the above-mentioned medical learning data, is assigned a weight as shown in the following mathematical formula 1.
[0014] [Mathematical Formula 1]
[0015]
[0016] (Here, P d is the prevalence of specific disease d for each characteristic, X i is the characteristic vector of patient i)
[0017] According to one embodiment, when the big data server retrains the country-specific disease diagnosis learning model, it compares the actual diagnosis result with the predicted diagnosis information derived by the diagnosis information providing server by extracting the medical questionnaire items and additional medical questionnaire items entered from the user terminal or the medical learning data, and evaluates the performance of the country-specific disease diagnosis learning model according to the accuracy according to Equation 2 below using a confusion matrix as shown in the following table.
[0018] Depending on the accuracy threshold mentioned above, existing questions among the above questions may be deleted or retained, and additional questions may be added to the above questions.
[0019]
[0020] [Mathematical Formula 2]
[0021]
[0022] (Here, TP is True Positive, TN is True Negative, FP is False Positive, and FN is False Negative.)
[0023]
[0024] When the above-mentioned big data server retrains the above-mentioned country-specific disease diagnosis learning model, it can determine the Sasang constitution type based on the MBTI (Myers-Briggs Type Indicator) type, weight, height, metabolic capacity, body temperature, digestive capacity, personality determined by the MBTI type, sleep duration, exercise duration, or underlying diseases from the above-mentioned medical history information and additional medical history information entered from the above-mentioned user terminal.
[0025] According to one embodiment, data regarding the MBTI type or the Sasang constitution type is collected, and the prevalence data is subdivided according to the MBTI type or the Sasang constitution type to analyze disease occurrence patterns based on an individual's personality type or constitution, thereby allowing the country-specific disease diagnosis learning model to be retrained and optimized.
[0026] According to the present invention, through systematic and objective medical history taking, patients residing in remote areas where it is difficult to receive high-quality medical services can be diagnosed quickly and accurately for the type of disease, allowing specialists to quickly predict the disease based on this data and even manage the effectiveness of treatment, thereby enabling the provision of accurate medical services.
[0027] In addition, according to the present invention, by receiving systematically and objectively prepared medical history and additional symptom information based on the patient's personal constitution or personality, and by constructing a national disease diagnosis learning model based on prevalence data according to national demographic criteria, regional criteria, time criteria, or health status criteria, and utilizing artificial intelligence, accurate predictive diagnosis information for a specific disease can be provided.
[0028] FIG. 1 is a diagram illustrating the configuration of a medical diagnostic system through medical history taking according to one embodiment of the present invention.
[0029] Figure 2 is a diagram of the questionnaire items displayed in the dedicated program of the user terminal in the present invention.
[0030] FIG. 3 is a diagram showing the derivation of predictive diagnostic information in response to a user's medical history questions in a medical diagnostic system through medical history questions according to one embodiment of the present invention.
[0031] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0032] Hereinafter, various embodiments of the present invention are described with reference to the accompanying drawings. However, this is not intended to limit the present invention to specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives of the embodiments of the present invention.
[0033] Embodiments of the present invention are described below with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present invention may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present invention in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.
[0034] Additionally, terms such as “…part,” “…unit,” and “module” described in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware, software, or a combination of hardware and software.
[0035] Furthermore, throughout the specification, when a part is described as being "connected" to another part, this includes not only cases where they are "directly connected," but also cases where they are "electrically connected" with other components in between.
[0036] Furthermore, when a part is said to "include" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but rather may include additional components, and it should be understood as not excluding in advance the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0037] The embodiments can be implemented in various forms of products such as personal computers, laptop computers, tablet computers, smartphones, televisions, smart home appliances, intelligent automobiles, kiosks, and wearable devices.
[0038] In the embodiment, the Artificial Intelligence (AI) system is a computer system that implements human-level intelligence, and unlike existing rule-based smart systems, it is a system in which the machine learns and makes decisions on its own.
[0039] As artificial intelligence systems improve their recognition rates and become capable of understanding user preferences more accurately with continued use, existing rule-based smart systems are gradually being replaced by deep learning-based AI systems.
[0040] Artificial intelligence technology consists of machine learning and component technologies utilizing machine learning. Machine learning is an algorithmic technology that autonomously classifies and learns the characteristics of input data, while component technologies are technologies that mimic the cognitive and judgmental functions of the human brain by utilizing machine learning algorithms such as deep learning, and are comprised of technological fields such as linguistic understanding, visual understanding, reasoning / prediction, knowledge representation, and motion control.
[0041] Here, linguistic understanding refers to technologies that recognize, apply, and process human language and text, including natural language processing, machine translation, dialogue systems, question answering, and speech recognition / synthesis. Visual understanding refers to technologies that perceive and process objects in a manner similar to human vision, including object recognition, object tracking, image search, person recognition, scene understanding, spatial understanding, and image enhancement. Inference and prediction refer to technologies that logically reason and predict by judging information, including knowledge / probability-based inference, optimization prediction, preference-based planning, and recommendation. Knowledge representation refers to technologies that automatically process human experiential information into knowledge data, including knowledge construction (data generation / classification) and knowledge management (data utilization). Motion control refers to technologies that control the autonomous driving of vehicles and the movement of robots, including motion control (navigation, collision, driving) and manipulation control (behavior control).
[0042] Generally, to apply machine learning algorithms to real-world situations, training is performed using a trial-and-error method due to the inherent characteristics of machine learning methodologies. In particular, deep learning requires hundreds of thousands of iterations. Since it is impossible to execute this in a real physical external environment, training is instead performed through simulations by virtually recreating the actual physical environment on a computer.
[0043] FIG. 1 is a diagram illustrating the configuration of a medical diagnosis system (1) through medical questioning according to an embodiment of the present invention. FIG. 2 is a diagram of medical questioning items displayed in a dedicated program of a user terminal (300) in the present invention. FIG. 3 is a diagram showing the derivation of predictive diagnosis information (40) in response to the user's medical questioning items in the medical diagnosis system (1) through medical questioning according to an embodiment of the present invention.
[0044] Referring to FIG. 1, a medical diagnosis system (1) through medical history taking according to one embodiment of the present invention may include a big data server (100), a diagnosis information providing server (200), and a user terminal (300).
[0045] The big data server (100) can collect and store medical learning data in real time from multiple network servers in multiple countries, and learn from the said medical learning data to build a country-specific disease diagnosis learning model (50) based on the prevalence rate of each disease in each country.
[0046] Prevalence rates for specific diseases vary by country. Major factors include the mode of transmission of infectious diseases, population density in each country, the quality of healthcare systems, vaccination rates, social health behaviors, and climatic and environmental factors.
[0047] For example, prevalence rates can vary depending on how infectious diseases spread, and diseases may spread more rapidly in densely populated countries. In countries with high-quality healthcare systems, infectious diseases can be detected and treated quickly, which can lead to lower prevalence rates. Additionally, in countries with high vaccination rates, immunity against infectious diseases is enhanced, which can result in lower prevalence rates.
[0048] In addition to this, social health behaviors and climatic and environmental factors can also influence prevalence rates. For example, in countries with good social health behaviors, the likelihood of maintaining healthy lifestyles and taking preventive measures against infectious diseases increases, which can lead to lower prevalence rates.
[0049] For accurate diagnosis, by constructing country-specific disease diagnosis learning models based on different prevalence rates across nations, various differences such as genetic, environmental, and lifestyle factors are reflected, enabling the provision of customized prevention and diagnostic information tailored to users in each country.
[0050] The user terminal (300) communicates with the big data server (100) via a wired or wireless network and provides medical diagnostic services through medical questioning, and may be equipped with a dedicated program to provide feedback on the actual diagnostic results (30) received by the user from a specialist. Here, the user terminal (300) may include a smartphone or a kiosk.
[0051] The aforementioned dedicated program can provide strong user authentication and security features to protect personal information and medical data, and can provide personalized medical questionnaires based on the user's past medical records, lifestyle habits, current symptoms, etc.
[0052] It includes a feature that allows users to consult with a doctor remotely when necessary, enabling them to receive additional medical advice based on their medical history. Furthermore, it supports various languages, allowing users worldwide to use it conveniently.
[0053] The above dedicated program may include a function that can continuously improve the performance of the dedicated program by providing feedback to the big data server (100) on the actual diagnosis results (30) received by the user from a specialist.
[0054] Referring to FIGS. 2 and 3, the dedicated program installed on the user terminal (300) can receive pre-prepared medical questionnaire items (10), additional medical questionnaire items (20) for newly adding symptoms that the user wants to personally report, and actual diagnosis results (30) received by the user from a specialist, and transmit them to the diagnosis information provision server (200).
[0055] The diagnostic information providing server (200) receives the medical history information (10) and additional medical history information (20) from the user terminal (300), and can infer predictive diagnostic information (40) corresponding to the medical history information using artificial intelligence through the country-specific disease diagnosis learning model (50) of the learned big data server (100).
[0056] Additionally, the diagnostic information providing server (200) can provide the inferred predicted diagnostic information (40) from the diagnostic information providing server (200) back to the dedicated program installed on the user terminal (300). The user can feed back the actual diagnostic result (30) received from a specialist back to the diagnostic information providing server (200) through the dedicated program.
[0057] The diagnostic information provision server (200) can use the country-specific disease diagnosis learning model (50) built in the big data server (100) to receive pre-prepared medical questionnaire items (10) and additional medical questionnaire items (20) that the user personally adds, through the dedicated program installed on the user terminal (300), and use artificial intelligence to provide predictive diagnostic information (40) corresponding to the medical questionnaire items to the dedicated program installed on the user terminal (300).
[0058] The above big data server (100) can receive feedback from the above prediction diagnosis information (40) and the above actual diagnosis result (30) from the above user terminal (300), compare them, and retrain the above country-specific disease diagnosis learning model (50) by adding the above additional medical history items to optimize it.
[0059] The big data server (100) can collect the medical learning data in real time from a public medical database (DB) provided by a government or public institution in Korea or abroad, a medical institution database in which medical questionnaire data and diagnostic data of a hospital or medical institution are stored, a medical academic database in which medical academic paper data is stored, or a medical data sharing platform that provides medical data.
[0060] Domestic public medical databases include the Health Insurance Review & Assessment Service (HIRA), which provides various medical data such as treatment history, prescription history, and health checkup results; the National Health Insurance Service (NHIS), which includes health insurance claims data, health checkup data, and mortality information; the National Cancer Center, which provides diagnosis and treatment data and survival rate data for cancer patients; and the National Medical Center, which includes diagnosis and treatment data for various diseases.
[0061] Overseas public medical databases include the National Institutes of Health (NIH) in the United States, which provides research data and clinical trial data on various diseases; the National Health Service (NHS) in the United Kingdom, which includes patient medical records, prescription history, and health checkup results; the European Union’s European Health Data Space (EHDS), which integrates and provides medical data from various countries within Europe; and the OMOP CDM (Observational Medical Outcomes Partnership Common Data Model), a standardized data model managed by the international non-profit organization OHDSI and used in various health and medical research.
[0062] In addition, medical data provided by data science platforms such as Kaggle can be utilized; Kaggle offers a variety of medical datasets, through which training data can be obtained.
[0063] The UCI Machine Learning Repository also provides various medical data, which can be utilized to build training data.
[0064] The country-specific disease diagnosis learning model (50) is intended to build a large-scale big database of medical learning data collected from these databases and to provide an artificial intelligence algorithm trained by machine learning to analyze the data.
[0065] To implement this, an artificial neural network can be used, which is implemented by software and can be composed of neurons connected by a network.
[0066] An artificial neural network has the function of connecting multiple neurons, which are units of computational processing, to receive the outputs of other connected neurons with weights assigned according to appropriate connection strengths, and to produce an output through its own activation function. One or more neurons form a layer, and this layer is divided into three types: an input layer, a hidden layer, and an output layer. The input of an artificial neural network is transmitted to the neurons in the input layer, and the output of the neurons in the input layer is transmitted to the neurons in the hidden layer.
[0067] Finally, the output of the neuron in the hidden layer is transmitted to the neuron in the output layer, and the output of the neuron in the output layer becomes the output of the artificial neural network. There is one input layer and one output layer, but there may be no hidden layer or one or more.
[0068] The present invention may be composed of prevalence data extracted from medical learning data, such as domestic or overseas public medical databases, based on national demographic, regional, temporal, or health status criteria for a specific disease, a plurality of input layers for pre-prepared medical history questions, and an output layer for a specific disease corresponding to the medical history questions. Optimal results can be derived using such an artificial neural network.
[0069] That is, the above-mentioned country-specific disease diagnosis learning model (50) can be constructed by extracting prevalence data based on country-specific demographic criteria, regional criteria, time criteria, or health status criteria for a specific disease from the above-mentioned medical learning data, and learning the above-mentioned medical learning data based on the prevalence data.
[0070] Here, the aforementioned demographic criteria can be classified by gender, age distribution, or socioeconomic status of the citizens of the country concerned.
[0071] The aforementioned socioeconomic status can be classified by the income level, education level, or family composition of the citizens of the country concerned.
[0072] The above income levels can be classified based on the country's Gross National Income per capita (GNI per capita).
[0073] The above educational levels can be classified based on the average highest level of education attained by the citizens of the country.
[0074] The above family member types can be classified by referring to the household composition survey or population and housing census of the respective country.
[0075] The above regional criteria can be classified into population density criteria, economic criteria, housing type criteria, or environmental criteria according to the administrative division criteria of the country concerned.
[0076] The above administrative division criteria can be classified into provinces, provinces, cities, counties, districts, neighborhoods, towns, or townships.
[0077] The above population density criteria can be classified as the number of residents per area (persons / km²) residing in the above administrative district criteria.
[0078] The above economic criteria can be classified according to the average income of residents living in the above administrative district criteria.
[0079] The above housing type criteria can be classified according to apartment complexes, residential areas, or commercial areas existing in the above administrative district criteria.
[0080] The above environmental standards can be classified according to industrial areas, residential areas, tourist areas, mountainous areas, or coastal areas to which the above administrative district standards apply.
[0081] The above time criteria can be classified according to the passage of time by year or season for residents corresponding to each of the above regional criteria.
[0082] The above health status criteria can be classified by underlying diseases, lifestyle habits, or genetic factors.
[0083] The above criteria by underlying disease can be classified according to the prevalence of patients with specific underlying diseases in the respective country.
[0084] The above criteria by lifestyle can be classified according to the prevalence rates associated with lifestyle habits regarding smoking, drinking, sleep duration, or exercise duration.
[0085] The above criteria by genetic factor can be classified according to the prevalence rate based on the family history of patients in the respective country.
[0086] Specifically, looking at demographic criteria, it is possible to indicate the prevalence of diseases in specific age groups. For example, hypertension may show a higher prevalence among the elderly.
[0087] It can indicate differences in disease prevalence rates between men and women. For example, the smoking rate among men is higher than that among women.
[0088] Disease prevalence rates can be indicated according to socioeconomic status based on income level, education level, etc. For example, the prevalence of obesity may be high among low-income groups.
[0089] Next, looking at regional criteria, the prevalence of specific diseases can vary by country and between urban and rural areas depending on vaccination rates and the level of the healthcare system. For example, the prevalence of respiratory diseases caused by air pollution may be high in urban areas.
[0090] Next, looking at it from a temporal perspective, the prevalence of a disease at a specific point in time can be represented. For example, the prevalence of hypertension can be investigated as of January 1, 2024. Also, the prevalence of a disease can be represented over a specific period. For example, the prevalence of influenza over the entire year of 2024 can be investigated.
[0091] Next, when examined based on health status, the prevalence of chronic underlying diseases may be higher than that of the general population.
[0092] To give a specific example, looking at prevalence rates based on demographic criteria, the prevalence of hypertension among the population aged 65 and older in Korea is approximately 60%. Additionally, the smoking rate for Korean men is 36.7%, while that for women is 6.6%. In terms of socioeconomic status, the prevalence of obesity is high among low-income groups, with the prevalence among low-income groups in the United States being approximately 35%.
[0093] Looking at prevalence rates based on regional criteria, the prevalence of respiratory diseases caused by air pollution is high in urban areas, with the prevalence of respiratory diseases in Beijing, China being approximately 20%. Additionally, by country, the prevalence of diabetes in Korea is approximately 10%, while in the United States it is approximately 8%.
[0094] Looking at prevalence rates based on timeframes, the prevalence of hypertension in Korea is approximately 30% as of January 1, 2024. Additionally, the prevalence of influenza for the entire year of 2023 is about 5%. The prevalence of acute respiratory diseases can vary depending on the specific season. For example, the prevalence of winter influenza is approximately 10%.
[0095] Looking at prevalence rates based on health status, the prevalence of chronic diseases is higher among people with chronic diseases than in the general population, with the prevalence of hypertension among people with chronic diseases being about 70%.
[0096]
[0097] Meanwhile, when the big data server (100) learns the country-specific disease diagnosis learning model (50), it may use a characteristic vector in which the prevalence data corresponding to each characteristic, which is classified by each characteristic including the patient's age, nationality, gender, place of residence, period of residence, family history, occupation, status, income level, lifestyle habits, or presence or absence of underlying diseases extracted from the medical learning data, is assigned a weight as shown in the following mathematical formula 1.
[0098]
[0099] [Mathematical Formula 1]
[0100]
[0101] (Here, P d is the prevalence of specific disease d for each characteristic, X i is the characteristic vector of patient i)
[0102]
[0103] For example, the prevalence rates according to each characteristic learned from medical learning data are as follows.
[0104] · Nationality and Age: Prevalence of diabetes in Koreans aged 50 and older is approximately 20%
[0105] · Gender: Prevalence of diabetes is approximately 12% in men and 8% in women
[0106] · Family history: Prevalence of diabetes is approximately 15% with a family history and 7% without.
[0107] · Lifestyle: Diabetes prevalence of approximately 14% in smokers, and approximately 8% in non-smokers
[0108] · Current symptoms: Diabetes prevalence is approximately 18% among those who feel fatigued, and approximately 6% among those who do not.
[0109] · Region: Diabetes prevalence among urban residents is approximately 11%, and among rural residents is approximately 9%
[0110] · Income level: Diabetes prevalence rate of approximately 9% for high-income groups and 11% for low-income groups
[0111] · Presence of chronic disease: Diabetes prevalence is approximately 20% with chronic disease, and approximately 5% without.
[0112] · Sleep duration: The prevalence of diabetes is approximately 12% for those sleeping less than 7 hours, and approximately 8% for those sleeping 7 hours or more.
[0113] · Exercise time: Diabetes prevalence is approximately 7% for those exercising 3 or more times a week, and approximately 13% for those exercising less than 3 times a week.
[0114]
[0115] In this case, the feature vector may include various characteristics and prevalence data for each user. It is said that the feature vector may include the following characteristics.
[0116] 1. Nationality: Patient's nationality
[0117] 2. Age: Patient's age
[0118] 3. Gender: Patient's gender (Male=1, Female=0)
[0119] 4. Family history: Presence or absence of family history of diabetes (Yes=1, No=0)
[0120] 5. Lifestyle: Smoking status (Smoking=1, Non-smoking=0)
[0121] 6. Current symptoms: Fatigue, thirst, frequent urination, etc. (Symptoms present=1, Absent=0)
[0122] 7. Region: Urban or Rural (Urban=1, Rural=0)
[0123] 8. Income Level: High-income group=1, Low-income group=0
[0124] 9. Presence of chronic disease: Present=1, Absent=0
[0125] 10. Sleep Time: Sleep Time (in hours)
[0126] 11. Exercise Time: Exercise time per week (in hours)
[0127]
[0128] For example, assuming it includes the following characteristics (where 1 applies and 0 does not),
[0129] 1. Nationality: South Korea
[0130] 2. Age: 50
[0131] 3. Gender: Male (1)
[0132] 4. Family history: Yes (1)
[0133] 5. Lifestyle: Smoking (1)
[0134] 6. Current symptoms: Fatigue (1)
[0135] 7. Region: City (1)
[0136] 8. Income level: High-income group (1)
[0137] 9. Presence of chronic disease: Yes (1)
[0138] 10. Sleep duration: 6 hours
[0139] 11. Exercise time: 2 hours
[0140]
[0141] At this time, a feature vector can be constructed by assigning the prevalence rate corresponding to each characteristic as a weight.
[0142]
[0143] 1. Prevalence Data:
[0144] Page=0.20
[0145] Pgender=0.12
[0146] Pfamily=0.15
[0147] Phabit=0.14
[0148] Psymptom=0.18
[0149] Pregion=0.11
[0150] Pincome=0.09
[0151] Pchronic=0.20
[0152] Psleep=0.12
[0153] Pexercise=0.13
[0154] 2. Patient data: Xi represents the feature vector of patient i.
[0155] 3. Weighting: Weights are assigned by multiplying each characteristic Xi by the prevalence rate P corresponding to it.
[0156] Wi = [10,0.12,0.15,0.14,0.18,0.11,0.09,0.20,0.72,0.26]
[0157]
[0158] Through this process, feature vectors for learning can be constructed by assigning different prevalence rates as weights according to each characteristic. This enables the derivation of more accurate diagnostic results.
[0159]
[0160] Meanwhile, when the big data server (100) retrains the country-specific disease diagnosis learning model (50), it compares the predicted diagnosis information (40) derived by the diagnosis information providing server (200) by extracting the medical questionnaire items and additional medical questionnaire items entered from the user terminal (300) or the medical learning data with the actual diagnosed disease, and can use a confusion matrix as shown in Table 1 below.
[0161] The confusion matrix is a tool used to evaluate the performance of classification models. By representing the relationship between predicted and actual results in a tabular form, it allows for the calculation of the model's accuracy, precision, recall, and other metrics. The confusion matrix consists of the following four elements.
[0162] 1. True Positive (TP): When the model correctly predicts data that is actually positive as positive.
[0163] 2. False Positive (FP): Cases where the model incorrectly predicts data that is actually negative as positive.
[0164] 3. True Negative (TN): Cases where the model correctly predicts data that is actually negative as negative.
[0165] 4. False Negative (FN): Cases where the model incorrectly predicts data that is actually positive as negative.
[0166]
[0167] [Table 1]
[0168]
[0169]
[0170] Using this confusion matrix, the performance of the country-specific disease diagnosis learning model (50) can be evaluated according to the accuracy according to the following mathematical formula 2. However, the performance evaluation of the country-specific disease diagnosis learning model (50) can be performed when a sufficient sample group (N=1000) or more has been collected.
[0171]
[0172] [Mathematical Formula 2]
[0173]
[0174] (Here, TP is True Positive, TN is True Negative, FP is False Positive, and FN is False Negative.)
[0175]
[0176] Depending on the accuracy threshold mentioned above, existing questions among the above questions may be deleted or retained, and additional questions may be added to the above questions.
[0177]
[0178] For example, assume that the Random Forest algorithm is selected for the country-specific disease diagnosis learning model (50). Random Forest is a machine learning algorithm that uses an ensemble learning method, and performs predictions by combining multiple decision trees, so that each tree is learned independently and the results are combined to make a final prediction.
[0179] As follows, the predicted diagnosis is: if it is diabetes,
[0180] The relationship between the predicted diagnosis result and the actual diagnosis result can be represented using a confusion matrix such as the following Table 2.
[0181] [Table 2]
[0182]
[0183] In this case, assuming the following [Table 3] represents the predicted and actual diagnostic results for 5 patients as an example,
[0184] [Table 3]
[0185]
[0186]
[0187] The confusion matrix can be calculated as shown in the following [Table 4].
[0188] [Table 4]
[0189]
[0190]
[0191] At this time, the accuracy can be calculated using the following mathematical formula 3, in accordance with the aforementioned mathematical formula 2.
[0192] [Mathematical Formula 3]
[0193]
[0194] (Here, TP is True Positive, TN is True Negative, FP is False Positive, and FN is False Negative.)
[0195]
[0196] Here, depending on the threshold of the above accuracy, existing questions among the above questions may be deleted or retained, and additional questions may be added to the above questions.
[0197] For example, if the accuracy is 0.5 or higher, the existing questions in the above-mentioned questions can be maintained, and additional questions can be added to the above-mentioned questions.
[0198] If the accuracy is less than 0.5, the existing questions in the above questionnaire may be deleted, and the additional questions may not be added to the above questionnaire.
[0199]
[0200] Additionally, when a small number of users input questionnaires through the user terminal (300), they may input different or contradictory questionnaires multiple times within a relatively short period, provide duplicate responses multiple times, or input questionnaires without any sincerity.
[0201] In this case, in order to properly evaluate the country-specific disease diagnosis learning model (50), a process for accurate individual identification and collection of uncontaminated data is required.
[0202] To this end, all data inputs entered from the user terminal (300) can be carefully selected and authenticated to filter out contradictory questions or duplicate responses. That is, contradictory or duplicate responses can be filtered out by utilizing a dataset that corresponds to the accurate context in various situations.
[0203] In addition, the vulnerability of the country-specific disease diagnosis learning model (50) can be tested and repeatedly improved by simulating hostile scenarios that hostile users may take. Furthermore, the dataset containing old medical history can be periodically reviewed and updated.
[0204]
[0205] Meanwhile, when the above big data server (100) retrains the above country-specific disease diagnosis learning model (50), it can determine and collect Sasang constitution based on MBTI (Myers-Briggs Type Indicator) type, weight, height, metabolic capacity, body temperature, digestive capacity, personality determined by MBTI type, sleep time, exercise time, or underlying disease from the above medical history information and above additional medical history information entered from the above user terminal (300).
[0206] In this way, data regarding the above MBTI type or the above Sasang constitution can be collected, and the prevalence data can be subdivided according to the above MBTI type or the above Sasang constitution to analyze the disease occurrence pattern according to an individual's personality type or constitution, thereby retraining and optimizing the above country-specific disease diagnosis learning model (50).
[0207] Some research results are revealing a correlation between specific MBTI types and diseases. For example, according to ADHD research, it has been reported that N (Intuitive) and P (Perceiving) types are more prevalent among patients with ADHD.
[0208] Furthermore, studies investigating the correlation between specific personality disorders and MBTI types have reported that, for example, Borderline Personality Disorder is more prevalent in N (Intuitive) and J (Judging) types. In a study examining the correlation between mood disorders and MBTI types, F (Feeling) types were found to be more vulnerable to mood disorders.
[0209] Similarly, there may be a significant relationship between Sasang constitution and disease prevalence. Sasang constitution is a method of traditional Korean medicine that classifies human constitutions into four types: Taeyang-in, Soyang-in, Taeum-in, and Soeum-in. Each constitution possesses unique physical and psychological characteristics, and accordingly, susceptibility to specific diseases may vary.
[0210] For example, according to some research results, Taeum-in is more susceptible to diseases such as hypertension, diabetes, metabolic syndrome, stroke, non-alcoholic fatty liver disease, and obstructive sleep apnea; Soyang-in is more susceptible to diseases such as digestive disorders, allergies, and asthma; and Soeum-in is more susceptible to diseases such as digestive disorders, chronic fatigue, and depression.
[0211] Meanwhile, the user terminal (300) forming part of the present invention may be a medical examination kiosk (310) installed in a hospital. FIG. 4 is a perspective view of a medical examination kiosk (310) included in the user terminal (300) forming part of the present invention. Referring to FIG. 4, the medical examination kiosk (310) is a device that helps patients easily complete a medical examination when visiting a hospital. The medical examination kiosk (310) can provide convenience to both patients and medical staff by allowing a non-face-to-face medical examination to be conducted before a face-to-face consultation with a specialist.
[0212] The medical history kiosk (310) of the present invention may additionally be equipped with a blood pressure monitor (315). The blood pressure monitor (315) allows the user to easily measure blood pressure directly and transmit the data to the hospital system in real time along with the medical history input data. By doing so, the accuracy and reliability of the medical history data can be ensured.
[0213] Kiosks equipped with blood pressure monitors have already been installed in some medical institutions. However, conventional kiosks with blood pressure monitors had a problem in that the monitors were fixed, requiring patients to assume an uncomfortable posture to measure their blood pressure.
[0214] The medical examination kiosk (310) of the present invention may provide a moving part (500) that allows the blood pressure monitor (315) to move left and right to a certain extent, provided that the blood pressure monitor (315) is not fixed to the kiosk body but is provided at the bottom of the blood pressure monitor (315) and the kiosk body so that the user can comfortably measure blood pressure.
[0215] The moving part (500) is provided with spherical joint parts (511) at both ends and a cushioning part (513) at the center.
[0216] A spring (515) is provided within the above buffer (513), and the spring (515) can provide elastic restoring force when the spherical joint rods (514) on both sides are compressed.
[0217] Each moving part (500) provided at the lower corner of the blood pressure monitor (315) can move while cushioning when the user measures blood pressure, so the user can comfortably measure blood pressure.
[0218] Even if the length of the connecting fixing rod (10) is limited, the connecting fixing rod (10) can be extended using this connecting fixing rod extension device (500) to allow it to be deformed while reflecting the movement of the ground, so that the condition of the ground can be accurately monitored.
[0219] The embodiments described above may be implemented as hardware components, software components, and / or combinations of hardware and software components. For example, the devices, methods, and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.
[0220] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.
[0221] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
[0222] Although the embodiments have been described above with reference to the limited drawings, those skilled in the art can apply various technical modifications and variations based on the above. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0223] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.
Claims
1. A big data server that collects and stores medical learning data in real time from multiple network servers in multiple countries, and learns from the said medical learning data to build a country-specific disease diagnosis learning model based on the prevalence rate of each disease in each country; A user terminal that communicates with the big data server via a wired or wireless network, provides medical diagnostic services through medical history taking, and has a dedicated program installed to provide feedback on actual diagnostic results received by the user from a specialist; and A diagnostic information providing server that uses the country-specific disease diagnosis learning model built on the big data server to receive pre-prepared medical questionnaire items and additional medical questionnaire items personally added by the user through the dedicated program installed on the user terminal, infers predictive diagnostic information corresponding to the medical questionnaire items using artificial intelligence, and provides the inferred predictive diagnostic information back to the dedicated program installed on the user terminal; Includes, A medical diagnosis system through medical history taking, characterized in that the above-mentioned big data server receives feedback on the above-mentioned predictive diagnosis information and the above-mentioned actual diagnosis results, compares them, and retrains and optimizes the above-mentioned country-specific disease diagnosis learning model by adding the above-mentioned additional medical history items.
2. In Paragraph 1, The above big data server is, Collecting the above medical learning data in real time from public medical databases provided by domestic or overseas governments or public institutions, medical institution databases storing medical history and diagnostic data from hospitals or medical institutions, medical academic databases storing medical academic paper data, or medical data sharing platforms providing medical data, and The above country-specific disease diagnosis learning model is, A medical diagnosis system through medical history taking, characterized by extracting prevalence data based on national demographic criteria, regional criteria, time criteria, or health status criteria for a specific disease from the medical learning data, and constructing the medical learning data based on the prevalence data.
3. In Paragraph 2, The above demographic criteria are, It is classified according to the gender, age distribution, or socioeconomic status of the nationals of the country concerned, and the aforementioned socioeconomic status is classified by the income level, education level, or family composition type of the nationals of the country concerned, the aforementioned income level is classified based on the country's Gross National Income (GNI per capita), the aforementioned education level is classified based on the average highest level of education of the nationals of the country concerned, and the aforementioned family composition type is classified by referring to the household composition survey or the Population and Housing Census of the nationals of the country concerned. The above regional criteria are, It is classified according to population density criteria, economic criteria, housing type criteria, or environmental criteria based on the administrative division criteria of the country concerned, wherein the administrative division criteria are classified into provinces, provinces, cities, counties, districts, neighborhoods, towns, or townships; the population density criteria are classified by the number of residents per area (persons / km²) residing in the administrative division criteria; the economic criteria are classified according to the average income of residents residing in the administrative division criteria; the housing type criteria are classified according to apartment complexes, residential areas, or commercial areas existing in the administrative division criteria; and the environmental criteria are classified according to industrial areas, residential areas, tourist areas, mountainous areas, or coastal areas corresponding to the administrative division criteria. The above time standard is, Classify residents corresponding to each of the above regional criteria according to the passage of time by year or season, and A medical diagnostic system based on medical history taking, characterized in that the above-mentioned health status criteria are classified according to underlying diseases, lifestyle habits, or genetic factors, the above-mentioned criteria for underlying diseases are classified according to the prevalence rate of patients with specific underlying diseases in the country, the above-mentioned criteria for lifestyle habits are classified according to the prevalence rate of lifestyle habits regarding smoking, drinking, sleep duration, or exercise duration, and the above-mentioned criteria for genetic factors are classified according to the prevalence rate of patients' family history in the country.
4. In Paragraph 2, The above big data server is, When training the above country-specific disease diagnosis learning model, A medical diagnosis system through medical history taking, characterized by using a feature vector in which the prevalence data corresponding to each characteristic extracted from the medical learning data, including the patient's age, nationality, gender, place of residence, duration of residence, family history, occupation, status, income level, lifestyle habits, or presence or absence of underlying diseases, is assigned a weight as shown in the following mathematical formula 1. [Mathematical Formula 1] (Here, P d is the prevalence of specific disease d for each characteristic, X i is the characteristic vector of patient i) 5. In Paragraph 4, The above big data server is, When retraining the above country-specific disease diagnosis learning models, The predicted diagnosis information derived by the diagnostic information providing server by extracting the medical history questions and additional questions entered from the user terminal or the medical learning data, and the actual diagnosis result are compared, and the performance of the country-specific disease diagnosis learning model is evaluated according to the accuracy in Equation 2 below using a confusion matrix as shown in the following table. Depending on the accuracy threshold above, delete or retain existing questionnaire items among the above questionnaire items, and add the above additional questionnaire items to the above questionnaire items, [Mathematical Formula 2] (Here, TP is True Positive, TN is True Negative, FP is False Positive, and FN is False Negative.) The above big data server is, When retraining the above country-specific disease diagnosis learning models, By determining the Sasang constitutional type based on the MBTI (Myers-Briggs Type Indicator) type, weight, height, metabolic rate, body temperature, digestive rate, personality determined by the MBTI type, sleep duration, exercise duration, or underlying diseases from the above-mentioned medical questionnaire items and additional medical questionnaire items entered at the above-mentioned user terminal, A medical diagnostic system through medical history taking, characterized by collecting data on the above MBTI type or the above Sasang constitution type, subdividing the above prevalence data according to the above MBTI type or the above Sasang constitution type to analyze disease occurrence patterns according to an individual's personality type or constitution, and retraining and optimizing the above country-specific disease diagnosis learning model.