Cancer risk prediction device and method

JP2026127021APending Publication Date: 2026-08-05SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION +4
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
Filing Date
2025-10-17
Publication Date
2026-08-05

AI Technical Summary

Benefits of technology

【0012】 上述した本発明の課題解決手段によれば、個別化された癌発生リスク予測情報を提供し、現在の食習慣タイプを分析して健康的な食習慣への改善を推奨し、生活習慣改善のための具体的な推奨事項を提示する。また、ハイリスク群には適切な検診を推奨することで検診の精度を高め、不必要な検診によるリスク(例:偽陽性や大腸内視鏡などの後続検査による副作用)も減少させることができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026127021000001_ABST
    Figure 2026127021000001_ABST
Patent Text Reader

Abstract

We predict the risk of developing cancer using a cancer risk prediction model that is suited to the characteristics of Koreans. [Solution] The cancer risk prediction device according to the present invention comprises a memory storing a risk prediction program and a processor that executes the risk prediction program. The risk prediction program applies questionnaire responses regarding multiple causative factors to a causative rate calculation model to calculate the incidence rate of breast cancer and outputs a breast cancer risk level based on the breast cancer incidence rate. The causative rate calculation model uses the multiple causative factors for a healthy cohort and calculates the incidence rate of breast cancer using a mathematical formula that shows the relationship between the multiple factors and breast cancer occurrence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a cancer occurrence risk prediction device and method for predicting the risk of cancer occurrence.

Background Art

[0002] South Korea has a high incidence of various cancers globally. Cancer has a high treatment effect if detected early, but the mortality rate tends to increase rapidly as the detection is delayed. Therefore, in South Korea, the importance of early intervention has been emphasized to reduce the burden of cancer, and the need for cancer prevention and improvement of screening has emerged. In addition, due to the rapid progress of aging, it is expected that the consumption of human and economic resources due to cancer will further increase. Although early detection and prognosis improvement of cancer have been promoted through various cancer screening programs, these programs alone have limitations in preventing the occurrence of cancer itself.

[0003] Cancer prevention is broadly classified into primary prevention (identifying high-risk groups and removing risk factors), secondary prevention (early detection and diagnosis), and tertiary prevention (preventing recurrence and improving prognosis). In order to accurately predict the cancer occurrence risk, a prediction model based on reliable data is essential. Currently, cancer prediction models used at home and abroad are mainly developed based on research targeting foreigners. Since there are differences in genes and lifestyles between Koreans and foreigners, there is a high possibility of inaccurate prediction for Koreans.

[0004] Moreover, the majority of existing studies are based on patient-control group studies, are vulnerable to recall bias, and only consider variables that have already been clarified by statistical methods, so the model fitness is often low. Furthermore, although there are studies using machine learning techniques, medical knowledge is not fully reflected, and there is a risk of including inappropriate variables. Therefore, for accurate cancer prediction, a new approach that utilizes a prospective Korean cohort and considers both statistical models and machine learning models was necessary.

[0005] Therefore, this invention proposes a method for predicting cancer risk suitable for Koreans by constructing a predictive model based on urban cohort HEXA data, which is the largest prospective cohort in Japan consisting of 170,000 people.

[0006] On the other hand, among the various factors contributing to cancer development, dietary factors are an important and modifiable risk factor. Dietary factors influence the development of systemic diseases associated with contact with intestinal epithelial tissue and nutritional components, and are closely related to cancer development. While the importance of adjusting dietary habits for cancer prevention is widely known, existing dietary guidelines apply the same guidelines to everyone, only recommending increases or decreases in specific micronutrients or food groups, making them difficult to adapt to individual dietary habits. Furthermore, these dietary guidelines are often based on studies targeting Westerners and are therefore not suitable for Koreans.

[0007] Therefore, in this invention, based on food intake frequency data from 170,000 Koreans, a machine learning model and an unsupervised clustering algorithm were used to extract and interpret multiple Korean dietary habit types that best explain the characteristics of Korean eating habits. Furthermore, dietary habit questionnaire data from new users was collected and classified into predefined Korean dietary habit types. Based on the dietary habit type of the user, other risk factors, and the cancer risk calculated by a cancer prediction model, personalized cancer prevention solutions were proposed, aiming to achieve more effective cancer prevention. [Overview of the project] [Problems that the invention aims to solve]

[0008] The technical objective of this invention is to provide a cancer risk prediction device and method that predicts the risk of cancer development using a cancer risk prediction model suited to the characteristics of Koreans, in order to solve the aforementioned problems.

[0009] However, the technical challenges that this embodiment aims to address are not limited to those described above, and other technical challenges may exist. [Means for solving the problem]

[0010] As a technical means for solving the above-mentioned technical problems, a cancer risk prediction device according to one embodiment of the present invention includes a memory storing a risk prediction program and a processor that executes the risk prediction program. The risk prediction program inputs the user's responses to a pre-set questionnaire into a dietary habit type classification model to determine the user's dietary habit type, and inputs at least one user-specific characteristic factor, including the dietary habit type, into a cancer risk prediction model to calculate the user's cancer risk. The dietary habit type classification model is generated based on data regarding the names of foods consumed and the frequency of consumption of each food for multiple users, and the cancer risk prediction model includes a calculation formula that calculates the cancer risk based on at least one user-specific characteristic factor, including the dietary habit type.

[0011] Furthermore, a cancer risk prediction method according to another embodiment of the present invention includes the steps of determining the user's dietary habit type by inputting the user's responses to a pre-set questionnaire into a dietary habit type classification model, and calculating the user's cancer risk by inputting at least one user-specific characteristic factor, including the dietary habit type, into a cancer risk prediction model, wherein the dietary habit type classification model is generated based on data regarding the names of foods consumed and the frequency of consumption of each food for multiple users, and the cancer risk prediction model includes a calculation formula for calculating the cancer risk based on at least one user-specific characteristic factor, including the dietary habit type. [Effects of the Invention]

[0012] According to the problem-solving means of the present invention described above, personalized cancer risk prediction information is provided, current dietary habits are analyzed and improvements to healthy eating habits are recommended, and specific recommendations for improving lifestyle habits are presented. Furthermore, by recommending appropriate screenings to high-risk groups, the accuracy of screenings can be improved, and risks from unnecessary screenings (e.g., false positives and side effects from subsequent examinations such as colonoscopy) can also be reduced.

[0013] Furthermore, according to the present invention, after a user has undergone an assessment of their cancer risk, they can specifically improve their dietary and lifestyle habits through a customized solution. For example, if a user is advised to reduce their intake of white rice and processed meats and increase their intake of nuts and vegetables, these changes in dietary habits can effectively reduce the risk of metabolic diseases and cancer. This goes beyond mere theoretical recommendations; by providing a concrete and practical action plan that users can actually implement in their daily lives, it can improve the user's health in the long term and bring economic benefits such as reduced medical expenses. [Brief explanation of the drawing]

[0014] [Figure 1] This is a schematic block diagram showing a cancer risk prediction device according to one embodiment of the present invention. [Figure 2] This is a block diagram showing the detailed configuration of a risk prediction program according to one embodiment of the present invention. [Figure 3] This figure shows the detailed configuration of a risk prediction program according to one embodiment of the present invention. [Figure 4] This figure shows the process of constructing a dietary habit type classification model according to one embodiment of the present invention. [Figure 5] This figure shows the process of classifying user-specific eating habits in the execution process of an eating habit type classification model according to one embodiment of the present invention. [Figure 6] This is a flowchart illustrating the operation method of a cancer risk prediction device according to one embodiment of the present invention. [Figure 7] This flowchart illustrates a method for providing a cancer prevention solution using a cancer risk prediction device based on one embodiment of the present invention. [Figure 8] This flowchart shows a method for providing a dietary habit solution, which is part of a cancer prevention solution for a cancer risk prediction device according to one embodiment of the present invention. [Modes for carrying out the invention]

[0015] The present invention will now be described in detail with reference to the accompanying drawings. However, the present invention can be implemented in various forms and is not limited to the embodiments described herein. Furthermore, the accompanying drawings are provided to facilitate understanding of the embodiments disclosed herein and do not limit the technical ideas disclosed herein. In order to clearly illustrate the present invention in the drawings, parts unrelated to the description have been omitted, and the size, shape, and form of each component shown in the drawings can be varied. Throughout the specification, identical or similar parts are denoted by the same or similar reference numerals.

[0016] The suffixes "module" and "part" used in the following description of components are added or mixed together solely for the convenience of drafting the specification and do not inherently have a distinct meaning or role. Furthermore, in describing the embodiments disclosed herein, detailed explanations of related well-known technologies have been omitted if it was deemed that such detailed explanations would obscure the gist of the embodiments disclosed herein.

[0017] Throughout this specification, when a part is described as being “connected (joined, in contact with, or joined)” to another part, this includes not only “directly connected (joined, in contact with, or joined)” but also “indirectly connected (joined, in contact with, or joined)” through other members. Furthermore, when a part is described as “containing (having or comprising)” a component, unless otherwise stated, this does not exclude other components, but rather means that it may further “contain (have or comprise)” other components.

[0018] In this specification, ordinal terms such as "first" and "second" are used solely to distinguish one component from another and do not restrict the order or relationship of the components. For example, the first component of the present invention may also be called the second component, and similarly, the second component may also be called the first component.

[0019] In this specification, the "~ part" includes a unit realized by hardware, a unit realized by software, or a unit realized using both. Also, one unit may be realized using two or more hardware components, and two or more units may be realized by one hardware component. On the other hand, the "~ part" is not meant to be limited to software or hardware. The "~ part" may be configured on an addressable storage medium and may also be configured to be executed by one or more processors. Therefore, as an example, the "~ part" includes components such as software components, object-oriented software components, class components, and task components, and processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functions provided in the components and the "~ part" may be integrated into fewer components or "~ parts", or separated into additional components or "~ parts". Also, the components and the "~ part" may be implemented to be executed by one or more CPUs within a device.

[0020] The network means a connection structure capable of information exchange between each node such as a terminal and a server, and includes a local area network (LAN), a wide area communication network (WAN), the Internet (WWW), wired and wireless data communication networks, telephone networks, wired and wireless television communication networks, etc. As an example of a wireless data communication network, 3G, 4G, 5G, 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), WIMAX (Worldwide Interoperability for Microwave Access), Wi-Fi, Bluetooth (registered trademark) communication, infrared communication, ultrasonic communication, visible light communication (VLC), LiFi (Light Fidelity), etc. can be mentioned, but it is not limited to these.

[0021] The user terminal can be implemented as a computer or mobile device capable of connecting to the cancer risk prediction device via a network. Here, the computer includes, for example, a laptop, desktop, or other device equipped with a web browser, and the mobile terminal includes, for example, all types of handheld wireless communication devices, such as various smartphones and tablet PCs, as long as portability and mobility are guaranteed.

[0022] Figure 1 is a schematic block diagram showing a cancer risk prediction device according to one embodiment of the present invention.

[0023] Referring to Figure 1, a cancer risk prediction device (100) according to one embodiment of the present invention will be described. The cancer risk prediction device (100) inputs questionnaire responses entered by the user into a dietary habit type classification model to determine the dietary habit type, and inputs at least one user-specific characteristic factor, including this dietary habit type, into the cancer risk prediction model to calculate the user's cancer risk. Such a cancer risk prediction device (100) can be implemented in the form of a computing device, or it can be implemented in the form of a server that provides cancer risk prediction services. When the cancer risk prediction device (100) is implemented as a server, it can operate in cloud computing service models such as SaaS (Software as a Service), PaaS (Platform as a Service), or IaaS (Infrastructure as a Service). Furthermore, the localized content provision device (100) can be built in the form of a private cloud, public cloud, or hybrid cloud.

[0024] The cancer risk prediction device (100) includes memory (110) and a processor (120), and may further include a communication module (130) and a database (140).

[0025] The memory (110) stores the risk prediction program, and the memory (110) should be interpreted as encompassing a non-volatile memory device that maintains recorded information even without a power supply, and a volatile memory device that requires power to maintain recorded information. The risk prediction program inputs the user's responses to a pre-set questionnaire into a dietary habit type classification model to determine the user's dietary habit type, and inputs at least one user-specific characteristic factor, including the dietary habit type, into a cancer risk prediction model to calculate the user's cancer risk. At this time, the dietary habit type classification model is generated based on data regarding the names of foods consumed by multiple users and the frequency of consumption of each food (number of times consumed per unit period, amount consumed per time, etc.). The cancer risk prediction model also includes a calculation formula that calculates the cancer risk based on at least one user-specific characteristic factor, including the dietary habit type.

[0026] The memory (110) can perform the function of temporarily or permanently storing data processed by the processor (120). The memory (110) may include, but is not limited to, a volatile storage device that requires power to maintain the stored information, a magnetic storage medium, or a flash storage medium.

[0027] The processor (120) executes a risk prediction program stored in memory (110) to determine the user's dietary habit type based on a dietary habit type classification model and calculates the user's cancer risk based on a cancer risk prediction model. The processor (120) also executes the risk prediction program and provides the function of controlling the hardware of the cancer risk prediction device (100) in accordance with the execution of the program. In other words, by executing the program, the processor (120) can perform necessary hardware control functions such as file system, memory allocation, network, basic libraries, timers, device control (display, media, input devices, 3D, etc.), and other utilities.

[0028] A processor (120) can mean a data processing device embedded in hardware, which has physically structured circuitry to perform functions expressed by code or instructions contained within a program, for example. Examples of such data processing devices embedded in hardware include microprocessors, central processing units (CPUs), processor cores, multiprocessors, ASICs (application-specific integrated circuits), FPGAs (field programmable gate arrays), etc., but the scope of the present invention is not limited to these.

[0029] For reference, the components shown in Figure 1 according to the embodiment of the present invention refer to software, or hardware components such as FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit), which perform predetermined roles.

[0030] However, “components” are not limited to software or hardware; each component may be configured on an addressable storage medium and may be configured to run by one or more processors.

[0031] Therefore, as an example, the components include software components, object-oriented software components, class components, task components, and processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, variables, etc. The components and the functions they provide may be integrated into fewer components or separated into additional components.

[0032] The communication module (130) may include a device that includes hardware and software necessary to connect with other network devices by wire or wireless connection to send and receive signals such as control signals or data signals in order to perform signal data communication with external devices.

[0033] The database (140) may store various data necessary for the risk prediction program to operate. For example, it may store user-specific questionnaire response data and information on the user's dietary habit type determined by a dietary habit type classification model. In addition to the user's dietary habit type, it may also manage information on demographic factors, lifestyle factors, disease history factors, family history factors, weight factors, anthropometric indicator factors, and biomarker factors as user-specific characteristic factors, as well as various detailed information for determining them.

[0034] Figure 2 is a block diagram showing the detailed configuration of a risk prediction program according to one embodiment of the present invention.

[0035] The risk prediction program (200) can broadly include a dietary habit type classification model (210), a cancer incidence risk prediction model (220), and a personalized cancer prevention solution provision unit (230).

[0036] The dietary habit type classification model (210) is generated based on data on the names of foods consumed and the frequency of consumption of each food from multiple users. When a user's questionnaire responses are entered, the model determines the user's dietary habit type.

[0037] Figure 3 shows the detailed configuration of a risk prediction program according to one embodiment of the present invention, Figure 4 shows the construction process of a dietary habit type classification model according to one embodiment of the present invention, and Figure 5 shows the process of classifying user-specific dietary habit types in the execution process of the dietary habit type classification model according to one embodiment of the present invention.

[0038] Referring to Figure 3, the urban cohort that forms the basis of the present invention is constructed by conducting a questionnaire survey with various factors on 170,000 Korean subjects during the construction process, and building a dietary habit type classification model based on the questionnaire response data. In this case, the questionnaire survey may include demographic factors, anthropometric measurements, biometric measurements, lifestyle factors, medical history, family history, and dietary factors for each user. More specifically, questionnaire responses regarding demographic factors may include information such as each user's age, sex, education level, marital status, or income level. Questionnaire responses regarding anthropometric measurements may include information on the user's height and weight, and information on changes in weight. Questionnaire responses regarding biometric measurements may include various biomarker measurement information such as the user's blood pressure, blood glucose level, and electrocardiogram. Questionnaire responses regarding lifestyle factors may include information on the user's exercise habits (type of exercise, frequency, intensity, etc.), drinking habits (frequency of drinking, average intake, etc.), smoking habits (frequency of smoking, amount smoked, etc.), and sleep habits (sleep quality, sleep duration, etc.). Furthermore, responses to the medical history questionnaire may include information about the user's various diseases, such as whether or not they have blood pressure-related diseases, blood glucose-related diseases, kidney diseases, or metabolic syndrome. Responses to the family history questionnaire may include information about whether or not the user's parents, siblings, or other family members have cancer-related diseases. Information about dietary factors may include information about the frequency and amount of each food consumed by the user. In particular, in this invention, questions are asked about the frequency and amount of intake of 106 pre-selected individual foods during the questionnaire survey, and each user's eating habit type is classified based on their responses. The number of individual foods is just an example, and the number of foods can be designed to differ depending on the model design policy. It is also possible to reduce the original 106 types of foods to 35 food groups using various quantitative and qualitative criteria, and then input these 35 food groups into the autoencoder model.

[0039] The eating habit type classification model (210) can be constructed through a data dimensionality reduction process using a machine learning model and a clustering process using an unsupervised clustering algorithm.

[0040] In particular, Figure 4 is based on an autoencoder model, which performs dimensionality reduction on the original data containing information on L pre-selected individual food items. Here, L represents the number of individual food items, and as in the example above, L can be set to 106, but this can be changed depending on the embodiment.

[0041] Conventional research extracted eating habit types by comparing the distance to a group based on user questionnaire data. However, applying this method directly makes it difficult to extract appropriate clusters due to the large dimensionality of the data. In this invention, to solve this problem, a machine learning model is used to compress multiple food data into an appropriate dimensionality and condense the information.

[0042] As a machine learning model, an autoencoder model composed of a deep neural network (DNN) can be used. An autoencoder consists of two components: an encoder and a decoder. The encoder compresses the original data, and the decoder reconstructs the compressed data. The encoder is used to extract eating habit types, and the decoder is used to train the encoder.

[0043] Autoencoder models are trained by compressing the original data with an encoder and then restoring it with a decoder, minimizing the difference between the original and restored data. Specifically, by repeatedly training the model using a loss function to minimize the difference between the original and restored data, the intrinsic parameters of the encoder and decoder are learned. After training is complete, the encoder becomes generalized to effectively compress new data. This allows high-dimensional (L-dimensional) data to be compressed into multiple lower-dimensional (M-dimensional, where M is a smaller value than L) data for use. While this invention presents an autoencoder as a machine learning model for dimensionality reduction, other methods such as attention models and contrast learning can also be used.

[0044] Next, a clustering algorithm is applied to the M-dimensional compressed data via the encoder to extract K clusters of eating habit types. The clustering process is described in detail with reference to Figure 5.

[0045] Referring to Figure 5, the process by which the eating habit type classification model (210) classifies the user's eating habit type is explained. When the user inputs their answers to the dietary questionnaire, these are input into the eating habit type classification model (210), the probability of belonging to each eating habit type is calculated, and the type with the highest calculated probability is determined as the user's eating habit type. To achieve this, a classifier can be added that classifies the original data into eating habit types using multiple eating habit types constructed by a clustering process.

[0046] As a clustering algorithm, the K-means algorithm, a commonly used unsupervised learning algorithm, can be used. The K-means algorithm divides data into k clusters by calculating the distance between each data point to find the k cluster centers and assigning each data point to the nearest cluster. Since the appropriate value of k varies depending on the data, the silhouette coefficient is used to find the most suitable k. The silhouette coefficient is an index that evaluates clustering by comparing the difference in distance between the cluster to which each data point is assigned and the other clusters; a value closer to 1 indicates good clustering, and a value closer to -1 indicates poor clustering. Even if the silhouette coefficient of one cluster is high, optimal clustering is evaluated as having a balanced distribution when considering the average silhouette coefficient of the clusters as a whole.

[0047] For example, if we set k to 3 and distinguish between male and female data so that three eating habit types are formed for each gender, the following eating habit types may be derived.

[0048] Male eating habit type 1: A diet that primarily consists of carbohydrates, especially rice, and consumes very little of other foods. Male eating habit type 2: A diet that primarily consists of wheat flour products such as noodles, bread, pizza, and hamburgers. Male eating habit type 3: A diet that involves consuming large amounts of pickles, processed meats, and salted seafood, resulting in a high sodium intake.

[0049] Female eating habit type 1: A diet that involves consuming a lot of pizza / hamburgers, and a lot of desserts such as cakes, cookies, dairy products, and beverages. Female eating habit type 2: A diet that consists mainly of carbohydrates, especially rice, and does not consume much of other foods. Female Eating Habit Type 3: A diet that involves consuming more bread, mochi (rice cakes), and noodles than rice, and a wide variety of foods including soybeans, vegetables, meat, and seafood. High fat intake.

[0050] Thus, the dietary habit type classification model (210) may include a classifier that uses multiple dietary habit types generated by applying a machine learning model to data collected through questionnaire responses during a cohort study to reduce the data's dimensionality, and then applying a clustering algorithm to the dimensionality-reduced data.

[0051] The classifier is also constructed based on an artificial neural network layer and can be trained using the training data used in the construction process of the aforementioned eating habit type classification model (210). As shown in the diagram, the raw data representing the user's questionnaire responses is input into an encoder, which is then input into a dimensionality-reduced compressed classifier. The classifier also utilizes multiple eating habit types generated by a clustering algorithm as a classification system.

[0052] With this configuration, the dietary habit type classification model (210) can determine the user's dietary habit type from the user's questionnaire responses related to diet. The determined dietary habit type is then input into the cancer risk prediction model and used to calculate the cancer risk. In some embodiments, in addition to information about the user's dietary habit type, information about the probability of being classified into the corresponding dietary habit type can also be transmitted to the cancer risk prediction model. On the other hand, for the multiple dietary habit types extracted as described above, the cancer risk for each type can be determined by survival analysis, and the clusters can be ranked from the lowest to the highest risk for each type of cancer.

[0053] Referring again to Figures 2 and 3, the cancer risk prediction model (220) calculates the user's cancer risk based on the dietary habit type and other user-specific characteristics determined by the dietary habit type classification model (210). In this case, the other user-specific characteristics may include at least one of the following: demographic factors, lifestyle factors, disease history factors, family history factors, weight factors, anthropometric indicator factors, and biomarker factors.

[0054] The cancer risk prediction model (220) is based on survival analysis. Survival analysis is a concept developed in statistics and refers to a method for predicting the period from the start of observation until a specific event occurs. The cancer risk prediction model (220) in this invention aims to predict the time from the start of observation until cancer is diagnosed.

[0055] However, since not everyone develops cancer and observations must be based on a limited timeframe, it is common practice to calculate a risk rate (hazard) and use it instead of predicting a specific time. The Cox proportional hazards model (Cox, DR; Oakes, D.(1984). Analysis of Survival Data. New York: Chapman & Hall. ISBN 978-0412244902) can be used to calculate such risks. In this case, the time-dependent risk for each observed user is calculated based on user-specific characteristics, as shown in the following formula.

[0056]

number

[0057] In Equation 1, h(t) is the risk of developing cancer. h0(t) is a value set identically for each individual and changes continuously over time. β i x is a coefficient set for each feature factor, i These are values ​​that represent each characteristic factor.

[0058] Alternatively, we can use equation 2, which is obtained by applying a logarithmic function to equation 1.

[0059]

number

[0060] In this case, the values ​​of each characteristic factor can be provided in a quantifiable form. The values ​​of each characteristic factor can be represented as a vector and quantified using one-hot encoding. For example, if eating habit types are classified into n types, a vector consisting of n variables (x1, x2, ..., x n By representing it as ) and applying one-hot encoding, where each variable is either 0 or 1, the corresponding eating habit type can be directly indicated.

[0061] Furthermore, in this invention, since the risk ratio between individuals is determined, it is not necessarily required to determine the value of h0(t). For example, if the risk rate of user A is determined to be h_A(t)=h0(t)×exp(1.0) and the risk rate of user B is determined to be h_B(t)=h0(t)×exp(2.0), the risk ratio between user A and user B can be calculated without determining h0(t), as h_A(t):h_B(t)=h0(t)×exp(1.0):h0(t)×exp(2.0)=exp(1.0):exp(2.0)=2.72:7.39.

[0062] In this invention, the β value that best represents the relative risk of cancer was calculated based on the Cox proportional hazards model described above. Furthermore, individual Cox proportional hazards models can be constructed for each of the multiple cancer types. Since the degree of influence of risk factors differs for each cancer type, a Cox proportional hazards model corresponding to each cancer type can be designed in a manner in which the β value of cancer incidence risk is set differently for each cancer type.

[0063] We developed a cancer risk prediction model using risk factors influencing cancer risk within urban cohort data, cancer risk factor data revealed in previous studies, and the proportional hazards model, a type of statistical analysis method described above. Using this cancer risk prediction model, we calculated and presented the cancer risk of a hypothetical subject, who provided lifestyle factors, physical measurements, family history, and medical history, as a hazard ratio.

[0064] Referring again to Figure 2, the Personalized Cancer Prevention Solution Provider (230) provides various cancer prevention solutions, such as family history solutions, lifestyle solutions, and dietary solutions, for cancer types with a high risk of developing cancer, as determined by the cancer risk prediction model (220).

[0065] Figure 6 is a flowchart showing the operation method of a cancer risk prediction device according to one embodiment of the present invention, Figure 7 is a flowchart showing a method for providing a cancer prevention solution using a cancer risk prediction device according to one embodiment of the present invention, and Figure 8 is a flowchart showing a method for providing a dietary habit solution among the cancer prevention solutions of a cancer risk prediction device according to one embodiment of the present invention.

[0066] First, the cancer risk prediction device (100) inputs the user's questionnaire responses into a dietary habit type classification model (210) to determine the dietary habit type for each user (S110). As mentioned above, the dietary habit type classification model is generated based on data on the names of foods consumed and the frequency of consumption of each food for multiple users, and is constructed using data compression processing and clustering algorithms with machine learning models.

[0067] Next, the user-specific characteristics, including the determined dietary habit type, are input into the cancer risk prediction model to calculate the cancer risk (S120). At this time, the user-specific characteristics may include, in addition to the dietary habit type, at least one of the following: demographic factors, lifestyle factors, disease history factors, family history factors, anthropometric indicator factors, and biomarker factors, which are determined based on the user's questionnaire responses. These user-specific characteristics are input into the cancer risk prediction model based on the aforementioned formula 1 or formula 2 to calculate the cancer risk. At this time, the cancer risk prediction model can determine and output one of several risk stages according to the calculated cancer risk value. For example, the risk stages can be divided into five stages, and the corresponding risk stage is determined and output according to the interval to which the calculated risk value belongs.

[0068] On the other hand, when the risk levels are divided into five risk levels (the fifth being the most dangerous), the cancer type with the highest risk of developing cancer can be assigned to cancer type A, and cancer types belonging to lower risk levels (for example, risk level 3 or higher) can be assigned to cancer type B (S122 in Figure 7).

[0069] On the other hand, the cancer risk prediction device (100) of the present invention may further include, in addition to the step of outputting a cancer risk, a step (S130) of providing a cancer prevention solution corresponding to that risk.

[0070] Referring to Figure 7, the stage of providing cancer prevention solutions (S130) may include the stage of providing family history solutions (S132), the stage of providing lifestyle solutions (S134), and the stage of providing dietary solutions (S136).

[0071] To further elaborate on the stage of providing the family history solution (S132), if there is a family history of a specific cancer type among the cancer types for which the cancer risk was calculated in the previous stage (S120), that cancer type can be assigned to cancer type C, and recommendations for genetic testing or health checkup centers for cancer type C, or screening items necessary for diagnosing cancer type C can be output. If there are no cancer types with a family history, this procedure is terminated. The recommendations based on the family history solution are based on pre-configured data, and the system databases the types of genetic testing recommended for each cancer type, a list of health checkup centers that perform the relevant tests, or a list of health screening items essential for diagnosing each cancer type, and outputs recommendations corresponding to cancer type C. For example, if there is a family history of lung cancer, genetic testing for lung cancer diagnosis, health checkup centers that perform the relevant tests, or screening items necessary for diagnosing lung cancer may be output as recommended information. For example, the following family history solution may be output.

[0072] "Mr. Hong Gil-dong has a family history of colorectal cancer, which may increase his risk of developing colorectal cancer due to genetic factors. Increased cancer risk due to family history requires more meticulous management for prevention and early detection. In such cases, it is essential to accurately understand your genetic risks through genetic testing and to thoroughly implement regular health checkups and preventive measures. Below is a list of recommended institutions where you can receive genetic testing and health checkups."

[0073] Next, in the stage of providing lifestyle solutions (S134), lifestyle solutions are provided based on the content of cancer type A and cancer type B determined in the previous stage (S122). Such lifestyle solutions can be selected and output from among multiple lifestyle solutions recommended for each cancer type that are appropriate to the user's information. The specific content of the lifestyle solutions may be pre-prepared as text messages in a database, etc., encouraging avoidance of lifestyle habits that increase the risk for each cancer type, while simultaneously recommending the maintenance of lifestyle habits that reduce the risk. In this case, visual materials such as body shape simulation diagrams and signal light colors (good / average / bad, etc.) indicating risk signals for each item may be used in the output.

[0074] Furthermore, in the stage of providing lifestyle solutions (S134), characteristic factors influencing cancer type A and cancer type B can be identified, and lifestyle solutions corresponding to characteristic factors common to cancer types A and B can be provided. In particular, various messages encouraging improvement of lifestyle habits related to each characteristic factor can be combined and output. For example, the following lifestyle solutions may be output.

[0075] "Hong Gil-dong's Body Mass Index (BMI) is 26.12 kg / m²." 2This indicates you are overweight. Since this may increase your risk of colorectal cancer, it is advisable to pay attention to weight management. Fortunately, your blood sugar, cholesterol, and blood pressure are well managed, but please note that your triglyceride levels are slightly high. While your regular exercise is very commendable, alcohol consumption and smoking increase the risk of various cancers, including colorectal cancer, so reducing alcohol intake or quitting smoking would be beneficial for maintaining your health.

[0076] Next, in the step of providing a dietary habit solution (S136), the user's dietary habit type is determined, and a dietary habit solution is output to reach a dietary habit type that can reduce the risk of developing cancer.

[0077] Referring to Figure 8, the specific details are explained by classifying the user's eating habit type using the aforementioned eating habit type classification model (210), and then sorting the classification results based on the probability of being classified into each type (S140).

[0078] Based on the sorting results of the eating habit types, the eating habit type with the highest probability is determined as eating habit type A, and the similarities and differences between type A and the actual eating habits based on the user's questionnaire responses can be provided as content constituting the eating habit solution (S142). In this case, since the user's questionnaire responses include whether or not they consume multiple foods (e.g., 106 items) and the frequency of consumption, the similarities between eating habit type A and the user's actual questionnaire results are determined as whether or not they consume the same foods and the frequency of consumption, and the differences can be output in the form of a list of foods that type A people consume but the user does not, or a list of foods that type A people do not consume but the user does.

[0079] Next, from the results of sorting the dietary habit types, the dietary habit type B with the lowest incidence of cancer type A, determined in the previous step (S122), is determined from among the types ranked lower than or equal to the threshold rank (S144). At this time, the threshold rank can be set to, for example, the 3rd, 4th, or 5th rank. Since type B is similar to the user's actual dietary habit type and can reduce the risk of developing cancer type A, it is presented to the user as a recommended dietary habit type.

[0080] Next, after determining that dietary habit type B is the dietary habit type to recommend to the user, the system identifies the similarities and differences between dietary habit type B and the user's actual dietary habits based on their questionnaire responses, and then outputs a dietary habit recommendation solution based on this (S142). Specifically, for foods common to dietary habit type B, consumption is recommended as before, and for foods that differ from dietary habit type B, consumption is also recommended for those consumed by people with type B, while conversely, for foods that people with type B do not consume, the system advises against consumption.

[0081] For example, the following dietary habit solutions may be output.

[0082] "Your total calorie intake is high, and you consume more rice, noodles, processed meats, and chicken than average. In contrast, you consume significantly less nuts, milk, yogurt, and fruit than average. You also consume slightly more coffee than average and slightly less vegetables than average."

[0083] "Your eating habit type is 'Type A (tentative name).' People belonging to this type are characterized by high intake of rice, noodles, and meat, and a high total calorie intake. While they consume a lot of carbohydrates and meat, they tend to consume relatively little vegetables, seafood, and dairy products. People belonging to this type tend to have a high risk of metabolic diseases such as hypertension and diabetes, as well as a high risk of cancer. Your eating habits are similar to this type, as you consume a lot of rice, noodles, processed meat and chicken, and little nuts and yogurt."

[0084] "Your eating habits share similar characteristics with 'Type B (tentative name)'. People belonging to this type have a high total calorie intake and a large intake of carbohydrates, but they also consume a wide variety of foods, including nuts, legumes, vegetables, seafood, dairy products, and fruits. People belonging to this type tend to have a low long-term cancer risk (ranked N). Your eating habits share the common characteristics of high intake of noodles and bread, and a high total calorie intake."

[0085] Therefore, to prevent cancer, it is desirable to improve your dietary habits to belong to "Type B" rather than "Type A," which you currently belong to. For example, try reducing your intake of rice and processed meats. Also, try increasing your intake of nuts, yogurt, vegetables, and fruits.

[0086] "It is known that consuming large amounts of processed meats such as ham and sausage increases the risk of cancer. Instead, try replacing them with unprocessed, fresh meats. Processed meats may contain compounds that are potentially carcinogenic."

[0087] Thus, the present invention makes it possible to predict a user's risk of developing cancer based on characteristic factors, including the user's eating habits. Furthermore, it can provide multiple solutions that reduce the user's risk of developing cancer.

[0088] The present invention can also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules executed by a computer. Computer-readable media means any available medium accessible by a computer, and includes all volatile and non-volatile media, removable and non-removable media. Computer-readable media may also include computer storage media. Computer storage media includes all volatile and non-volatile, removable and non-removable media implemented by any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data.

[0089] Furthermore, although the methods and systems of the present invention are described in relation to specific embodiments, some or all of their components or operations can be implemented using a computer system having a general-purpose hardware architecture.

[0090] Those skilled in the art will understand that the present invention can be readily modified into other specific forms without altering the technical idea or essential features based on the above description. Therefore, the embodiments described above should be understood to be illustrative and not restrictive in all respects. The scope of the present invention is defined by the claims set forth below, and their meaning and scope, as well as all modifications or variations derived from equivalent concepts, should be construed as being within the scope of the present invention.

[0091] The scope of this application should be indicated by the claims set forth below, rather than by the detailed description above, and all changes or modifications derived from the meaning and scope of the claims, as well as equivalent concepts thereto, should be interpreted as being included within the scope of this application.

Claims

1. A device that predicts the risk of developing cancer, The memory where the risk prediction program is stored, The system includes a processor that executes the aforementioned risk prediction program, The aforementioned risk prediction program, The user's responses to a pre-set questionnaire are input into a dietary habit type classification model to determine the user's dietary habit type. At least one user-specific characteristic factor, including the aforementioned dietary habit type, is input into the cancer risk prediction model to calculate the user's cancer risk. The aforementioned dietary habit type classification model was generated based on data regarding the names of foods consumed and the frequency of consumption of each food item by multiple users. The cancer risk prediction device includes a cancer risk prediction model that includes a calculation formula for calculating the cancer risk based on at least one user-specific characteristic factor, including the type of eating habits.

2. A cancer risk prediction device according to claim 1, The responses to the aforementioned questionnaire include information to confirm at least one of the following factors: dietary factors, demographic factors, lifestyle factors, disease history factors, family history factors, anthropometric indicator factors, and biomarker factors. The aforementioned user-specific characteristic factors are: A cancer risk prediction device that includes at least one of the demographic factors, lifestyle factors, disease history factors, family history factors, anthropometric index factors, and biomarker factors.

3. A cancer risk prediction device according to claim 1, The aforementioned dietary habit type classification model is, Data on the names of foods consumed and the frequency of consumption of each food item from multiple users are clustered to generate multiple classification clusters. A cancer risk prediction device in which the dietary habit type is determined as one of the multiple classification clusters by inputting the user's questionnaire responses into the dietary habit type classification model.

4. A cancer risk prediction device according to claim 3, The aforementioned dietary habit type classification model performs clustering based on data obtained by inputting data on the names of foods consumed and the frequency of consumption of each food by multiple users into a machine learning model to reduce dimensionality before performing the clustering, thereby providing a cancer risk prediction device.

5. A cancer risk prediction device according to claim 1, The above calculation formula is given by formula 1 below: [Math 1] The aforementioned h(t) is the risk of developing cancer, and the aforementioned h 0 A cancer risk prediction device in which (t) is a constant, β is a user-specific characteristic factor, and x is the weight set for the user-specific characteristic factor.

6. A cancer risk prediction device according to claim 5, The cancer risk prediction model includes multiple calculation formulas with different settings for multiple cancer types, and the weights are set for each cancer type, thereby providing a cancer risk prediction device.

7. A cancer risk prediction device according to claim 1, The aforementioned cancer risk prediction model is a cancer risk prediction device that calculates the cancer risk for each cancer type using multiple calculation formulas with different settings for each cancer type.

8. A cancer risk prediction device according to claim 1, The aforementioned cancer risk prediction model is a cancer risk prediction device that determines and outputs one of several risk stages divided into predetermined intervals, according to the calculated cancer risk value.

9. A cancer risk prediction device according to claim 1, The cancer risk prediction device is characterized in that the risk prediction program assigns cancer types with a family history to cancer type C among the cancer types for which the cancer risk has been calculated, and provides a cancer prevention solution in the form of outputting genetic testing, recommendations from health checkup centers, or screening items for the diagnosis of cancer type C.

10. A cancer risk prediction device according to claim 1, The risk prediction program is characterized by assigning the cancer with the highest risk of developing cancer to cancer type A, and assigning cancer types with a lower risk than cancer type A but with a risk of a predetermined level or higher to cancer type B, thereby providing a cancer risk prediction device.

11. A cancer risk prediction device according to claim 10, The risk prediction program provides lifestyle solutions that address characteristic factors that commonly affect cancer type A and cancer type B. The aforementioned lifestyle solution is a cancer risk prediction device characterized by including wording that encourages improvement of lifestyle habits related to the aforementioned characteristic factors.

12. A cancer risk prediction device according to claim 10, The risk prediction program is characterized by determining the second dietary habit type with the lowest incidence of cancer type A from among the dietary habit types that have a lower probability than the first dietary habit type with the highest probability, as classified by the dietary habit type classification model, and providing a dietary habit solution that recommends the second dietary habit type to the user.

13. A cancer risk prediction device according to claim 12, The risk prediction program calculates the similarities and differences between the second eating habit type and the actual eating habits based on the user's questionnaire responses. The text recommends the consumption of foods that meet the aforementioned common characteristics, For foods that fall under the aforementioned differences and that the user has not consumed, the text should include wording recommending their consumption. A cancer risk prediction device characterized by providing a dietary habit solution that includes wording encouraging individuals to refrain from consuming foods that fall under the aforementioned differences but are not consumed by those with the second dietary habit type.

14. A method for predicting the risk of developing cancer, (a) A step of inputting the user's responses to a pre-set questionnaire into a dietary habit type classification model to determine the user's dietary habit type; and (b) The process includes inputting at least one user-specific characteristic factor, including the dietary habit type, into a cancer risk prediction model and calculating the user's cancer risk, The aforementioned dietary habit type classification model was generated based on data regarding the names of foods consumed and the frequency of consumption of each food item by multiple users. A method for predicting the risk of cancer, characterized in that the cancer risk prediction model includes a calculation formula that calculates the risk of cancer based on at least one user-specific characteristic factor, including the type of eating habits.

15. A method for predicting the risk of cancer development according to claim 14, The responses to the aforementioned questionnaire include information to confirm at least one of the following factors: dietary factors, demographic factors, lifestyle factors, medical history factors, family history factors, anthropometric indicator factors, and biomarker factors. A method for predicting cancer risk, characterized in that the user-specific characteristic factors include at least one of the demographic factors, lifestyle factors, medical history factors, family history factors, anthropometric index factors, and biomarker factors.

16. A method for predicting the risk of cancer development according to claim 14, The aforementioned dietary habit type classification model generates multiple classification clusters by clustering data on the names of foods consumed and the frequency of consumption of each food for multiple users. The method for predicting cancer risk is characterized in that step (a) above determines the type of eating habit as one of the multiple classification clusters by inputting the user's questionnaire responses into the eating habit type classification model.

17. A method for predicting the risk of cancer development according to claim 16, A method for predicting cancer risk, characterized in that the dietary habit type classification model performs clustering based on data obtained by inputting data on the names of foods consumed and the frequency of consumption of each food by multiple users into a machine learning model and performing dimensionality reduction on that data.

18. A method for predicting the risk of cancer development according to claim 14, The aforementioned calculation formula is given by the following equation 1: [Math 2] The aforementioned h(t) is the risk of developing cancer, and the aforementioned h 0 A method for predicting the risk of cancer development, characterized in that (t) is a constant, β is a user-specific characteristic factor, and x is the weight set for the characteristic factor.

19. A method for predicting the risk of developing cancer according to claim 18, The cancer risk prediction model is characterized in that it includes different calculation formulas for each of several cancer types, and the weights are set for each cancer type.

20. A method for predicting the risk of cancer development according to claim 14, The above step (b) is, A method for predicting cancer risk, which includes a step of calculating the cancer risk for each cancer type using different calculation formulas for multiple cancer types.

21. A method for predicting the risk of cancer development according to claim 14, The above step (b) is, A method for predicting cancer risk, comprising the step of determining and outputting one of several risk stages divided into predetermined intervals, based on the calculated cancer risk value.

22. A method for predicting the risk of cancer development according to claim 14, A method for predicting cancer incidence risk, further comprising the step of assigning cancer types with a family history among those for which cancer incidence risk has been calculated as cancer type C, and providing a cancer prevention solution in the form of outputting genetic testing for cancer type C, recommendations from health checkup centers, or screening items necessary for diagnosing cancer type C.

23. A method for predicting the risk of cancer development according to claim 14, A method for predicting cancer incidence risk, further comprising the steps of assigning the cancer with the highest risk of developing cancer as cancer type A, and assigning cancer types with a lower risk than cancer type A but with a risk of a predetermined stage or higher as cancer type B.

24. A method for predicting the risk of developing cancer according to claim 23, The process further includes providing lifestyle solutions based on characteristic factors that commonly affect cancer type A and cancer type B, A method for predicting cancer risk, characterized in that the lifestyle solution includes wording that encourages improvement of lifestyle habits related to the characteristic factors.

25. A method for predicting the risk of developing cancer according to claim 23, A method for predicting cancer risk, further comprising the steps of determining a second dietary habit type with the lowest incidence of cancer type A from among multiple dietary habit types classified by the aforementioned dietary habit type classification model, with a lower probability ranking than the first dietary habit type which is classified with the highest probability, and providing a dietary habit solution that recommends the second dietary habit type to the user.

26. A method for predicting the risk of developing cancer according to claim 25, The process of providing the aforementioned dietary habit solution is: The similarities and differences between the aforementioned second eating habit type and the actual eating habits based on the user's questionnaire responses are calculated. Statements recommending the consumption of foods that share common characteristics, Regarding the foods that fall under the category of differences, the text recommends consumption of foods that the user has not consumed, and A method for predicting cancer risk, comprising the step of providing a dietary habit solution that includes wording encouraging people to refrain from consuming foods that do not consume among those foods that fall under the category of differences, specifically those foods of the second dietary habit type.