CKD queue data-based causal relationship identification method and device
By preprocessing the CKD queue data and optimizing the feature correlation matrix, using the CVD event prediction model to identify causal relationships, the problem of difficulty in identifying interaction effects of multiple risk factors in the CKD queue data in the prior art is solved, and the accuracy of CKD secondary CVD risk prediction is improved.
Patent Information
- Application Number
- CN202510197020.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art is difficult to effectively identify the interaction effects between multiple risk factors in CKD queue data, resulting in inaccurate prediction of risk of CKD secondary CVD.
By preprocessing the CKD queue data, the feature vector and feature association matrix are constructed, and the CVD event prediction model of the adaptive one-dimensional convolution kernel and feedforward neural network are trained, and the feature association matrix is optimized to identify causal relationships.
It realizes effective extraction of the interaction effects between multiple risk factors in the CKD cohort data, improves the accuracy of risk prediction of CKD secondary CVD, and provides a theoretical basis for clinical practice.
Smart Images

Figure CN120089372A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical data processing, and particularly to a method, device, storage medium and electronic device for identifying causal relationships based on CKD cohort data. Background Art
[0002] The Global Burden of Disease (GBD) Chronic Kidney Disease (CKD) Collaboration's study on the burden of CKD in 195 countries and regions around the world from 1990 to 2017 shows that in 2017, there were 697.5 million CKD patients globally. Among them, there were 132.3 million in China, accounting for nearly 1 / 5. In particular, the number of patients with Disability Adjusted Life Year (DALY) caused by CKD reached 35.8 million, and the DALYs related to Cardiovascular Disease (CVD) were as high as 25.3 million. The number of CVD deaths caused by CKD was also as high as 1.36 million. Based on this, to further reduce the CVD disease burden caused by CKD, specific CVD risk factors should be identified as early as possible to guide clinical integration of existing and emerging therapies to improve CVD prognosis.
[0003] In recent years, continuous research reports have confirmed that the α-Klotho protein (Kl) produced by the kidney, as a mandatory receptor for fibroblast growth factor 23 (FGF23) secreted by the bone, constitutes the renal-bone endocrine axis. By regulating multiple feedback loops among multiple organs such as the kidney, intestine, bone, and parathyroid gland, it precisely regulates the calcium and phosphorus fluxes in the body. In recent years, continuous research reports have confirmed that during the CKD process, the reduction of Kl and the excess of FGF23 will accelerate the progression of CKD and the risk of CVD by amplifying traditional and non-traditional risk factors. However, most of the research processes of this view are based on traditional statistical analysis methods. The methods based on traditional statistical analysis not only ignore the complex interaction between Kl and FGF23, but also ignore the complex interaction between these two and other factors. Summary of the Invention
[0004] The embodiments of this application provide a method, device, storage medium and electronic device for identifying causal relationships based on CKD cohort data, which can effectively extract the interaction effects between multiple risk factors and provide a theoretical basis for predicting the risk of CKD secondary to CVD.
[0005] The embodiments of this application provide a method for identifying causal relationships based on CKD cohort data, including: Obtain the CKD cohort corpus and preprocess the CKD cohort corpus; Construct a first feature vector based on the preprocessed CKD cohort corpus; Construct an initial feature correlation matrix and initialize the initial feature correlation matrix to obtain a first feature correlation matrix; Input the first feature vector and the first feature correlation matrix into the CVD event prediction model for training to obtain the classification result of the CKD cohort corpus and the optimized feature correlation matrix.
[0006] As a further improvement of the present invention, in the above causal relationship recognition method based on CKD cohort data, wherein preprocessing the CKD cohort corpus includes: Delete the invalid feature columns in the CKD cohort corpus; Process the numerical feature outliers in the CKD cohort corpus after deleting the invalid feature columns; Process the symbolic feature outliers in the CKD cohort corpus after deleting the invalid feature columns.
[0007] As a further improvement of the present invention, in the above causal relationship recognition method based on CKD cohort data, wherein constructing the first feature vector based on the preprocessed CKD cohort corpus includes: Normalize the preprocessed CKD cohort corpus to obtain a first feature vector.
[0008] As a further improvement of the present invention, in the above causal relationship recognition method based on CKD cohort data, wherein the CVD event prediction model includes an adaptive one-dimensional convolutional kernel and a feedforward neural network; The step of inputting the first feature vector and the first feature correlation matrix into the CVD event prediction model for training to obtain the classification result of the CKD cohort corpus and the optimized feature correlation matrix includes: Sort the first feature correlation matrix to obtain a second feature correlation matrix; Perform a dot product operation on the second feature correlation matrix and the first feature vector to obtain a second feature vector; Concatenate the first feature vector and the second feature vector to obtain a third feature vector; Perform a dot product operation on the third feature vector and the adaptive one-dimensional convolutional kernel to obtain a fourth feature vector; Input the fourth feature vector into the feedforward neural network to obtain the classification result of the CKD cohort corpus and the optimized feature correlation matrix.
[0009] As a further improvement of the present invention, in the above-mentioned causal relationship recognition method based on CKD queue data, the training process of the CVD event prediction model includes: Obtain a CKD data set, where the CKD data set includes CKD queue corpus and corresponding true classification labels; Preprocess the CKD queue corpus, construct a first feature vector, and construct a first feature correlation matrix; Input the first feature vector and the first feature correlation matrix into the CVD event prediction model for processing to obtain the classification result of the CKD queue corpus; Calculate the squared loss function based on the classification result of the CKD queue corpus and the corresponding true classification label, and iteratively train the CVD event prediction model based on the squared loss function to obtain a trained CVD event prediction model and an optimized feature correlation matrix.
[0010] As a further improvement of the present invention, in the above-mentioned causal relationship recognition method based on CKD queue data, the processing of numerical feature outliers in the CKD queue corpus after deleting invalid feature columns includes: Detect missing values and illegal values in the numerical features of the CKD queue corpus, and perform median or mean filling processing on the missing values and illegal values; Construct a normal range table for numerical features, and sequentially determine whether each numerical feature is out of bounds according to the normal range table; Fill the numerical features with out-of-bounds situations with maximum, minimum, or median values.
[0011] As a further improvement of the present invention, in the above-mentioned causal relationship recognition method based on CKD queue data, the processing of symbolic feature outliers in the CKD queue corpus after deleting invalid feature columns includes: Detect missing values and illegal values in the symbolic features of the CKD queue corpus, and process them according to the principle of the most recent occurrence or the principle of the most frequent occurrence; Replace the symbols in the symbolic features.
[0012] The embodiment of the present application also provides a causal relationship recognition device based on CKD queue data, including: An acquisition and preprocessing module, configured to acquire a CKD queue corpus and preprocess the CKD queue corpus; A first generation module, configured to construct a first feature vector based on the preprocessed CKD queue corpus; A second generation module, configured to construct an initial feature correlation matrix and initialize the initial feature correlation matrix to obtain a first feature correlation matrix; An identification and optimization module, configured to input the first feature vector and the first feature correlation matrix into a CVD event prediction model for training processing, so as to obtain a classification result of the CKD cohort corpus and an optimized feature correlation matrix.
[0013] An embodiment of the present application further provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are suitable for being loaded by a processor to execute any one of the above-mentioned causal relationship identification methods based on CKD cohort data.
[0014] An embodiment of the present application further provides an electronic device, including a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used for the steps in any one of the above-mentioned causal relationship identification methods based on CKD cohort data.
[0015] The causal relationship identification method, device, storage medium and electronic device based on CKD cohort data provided by the present application. The present application preprocesses the CKD cohort corpus, constructs a first feature vector based on the preprocessed CKD cohort corpus, and processes the constructed first feature correlation matrix and the first feature vector through a CVD event prediction model to obtain an optimized feature correlation matrix. The present application performs a series of preprocessing on the CKD cohort corpus to ensure the data integrity of the CKD cohort corpus. The present application describes the causal relationship between CKD cohort corpora through a feature correlation matrix, can effectively extract the interactive effects between multiple risk factors, and provides a theoretical basis for the risk prediction of CKD secondary to CVD. Description of the Drawings
[0016] The following combines the drawings and describes the technical solutions and other beneficial effects of the present application in detail through the specific embodiments of the present application, which will be obvious.
[0017] Figure 1 It is a flowchart of the causal relationship identification method based on CKD cohort data provided by an embodiment of the present application.
[0018] Figure 2 It is a schematic structural diagram of the causal relationship identification device based on CKD cohort data provided by an embodiment of the present application.
[0019] Figure 3 It is a schematic structural diagram of the electronic device provided by an embodiment of the present application. Detailed Embodiments
[0020] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0021] Traditional statistical analysis methods cannot effectively extract the interaction effects between multiple risk factors. For example, they cannot consider the combined effects of FGF23 and hyperphosphatemia on CVD. Although most modern machine learning methods can model the interaction effects between multiple risk factors, they cannot effectively integrate some existing medical causal knowledge. For example, for a certain patient, when their serum creatinine level increases, eGFR will necessarily decrease. To combine the advantages of both, a small number of studies have designed corresponding machine learning models from the perspective of probabilistic graphical models, setting artificially defined constraints between different risk factors. While ensuring the causal relationships between a few risk factors, the machine learning model comprehensively considers the interaction effects between multiple risk factors and estimates the probability of cardiovascular events occurring under different conditions.
[0022] Based on this, it is not difficult to see that although there are already a few studies that can artificially set constraints between different risk factors, the prerequisite for doing so is to know all relevant medical causal knowledge and make constraints. At the same time, when facing some relatively complex tasks, it is often impossible to know all the relevant medical causal knowledge required.
[0023] To solve the above problems, the embodiments of the present application provide a causal relationship recognition method, device, storage medium, and electronic device based on CKD cohort data. A causal relationship recognition device based on CKD cohort data provided by the embodiments of the present application can be integrated in an electronic device, and the electronic device can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a microprocessing box, or other devices, etc.
[0024] Please refer to Figure 1 , Figure 1 which is a flowchart of the causal relationship recognition method based on CKD cohort data provided by the embodiments of the present application. It is applied to an electronic device, and the causal relationship recognition method based on CKD cohort data includes the following steps: S1. Obtain the CKD cohort corpus and preprocess the CKD cohort corpus.
[0025] In one embodiment, the preprocessing of the CKD cohort corpus in step S1 may specifically include the following steps: S11, delete invalid feature columns in the CKD cohort corpus.
[0026] Specifically, the CKD cohort corpus contains hundreds of feature columns, among which sensitive feature columns (such as 'name' and 'ID number') and useless feature columns (such as 'follow-up time' and 'number') should be deleted.
[0027] S12, processing the numerical feature outliers in the CKD cohort corpus after deleting the invalid feature columns.
[0028] Specifically, step S12 includes the following steps: S121, detect missing values and illegal values in the numeric features in the CKD cohort corpus, and perform median or mean filling processing on the missing values and illegal values; S122, constructing a normal range table of digital features, and judging whether each digital feature is out of bounds in turn according to the normal range table; S123, fill in the numerical features that are out of bounds with the maximum, minimum or median value.
[0029] Specifically, we first detect missing values and illegal values in digital features, and fill them in according to the median or mean for different situations. Secondly, we build a normal range table for all digital features based on professional medical knowledge. Finally, we determine whether each digital feature is out of bounds. For out-of-bounds situations, we use the maximum, minimum or median to fill in.
[0030] S13, processing symbolic feature outliers in the CKD cohort corpus after deleting invalid feature columns.
[0031] Specifically, step S13 includes the following steps: S131, detect missing values and illegal values in symbolic features in the CKD cohort corpus, and process them according to the most recent occurrence principle or the most frequent occurrence principle; S132, replacing the symbol in the symbolic feature.
[0032] Specifically, missing values and illegal values in symbolic features are first detected and processed according to the most recent occurrence principle or the most frequent occurrence principle. Secondly, the symbols in symbolic features are processed uniformly, for example, "yes" and "no" are replaced with "yes" and "no".
[0033] In view of the problems that the existing CKD electronic case cohort data often have human filling errors and omissions in actual clinical work, an automated verification program and a filling program are constructed to ensure the integrity of the data. In view of the imbalance of numerical data in the CKD electronic case cohort data, different numerical features are balanced through feature normalization technology. In view of the non-uniformity of symbolic features in the CKD electronic case cohort data, unified processing is carried out through the principle of the most recent occurrence or the principle of the most frequent occurrence.
[0034] S2. Construct a first feature vector based on the preprocessed CKD cohort corpus.
[0035] Specifically, the CKD cohort corpus data after preprocessing is standardized to obtain a first feature vector, and the size of the first feature vector is: the number of features × 1.
[0036] S3. Construct an initial feature correlation matrix and initialize the initial feature correlation matrix to obtain a first feature correlation matrix.
[0037] Set an initial feature correlation matrix (this matrix is a real symmetric matrix and the diagonal elements are 0), and the size of the initial feature correlation matrix is: the number of features × the number of features. Use the random initialization function to initialize the initial feature correlation matrix to obtain a first feature correlation matrix.
[0038] S4. Input the first feature vector and the first feature correlation matrix into the CVD event prediction model for training to obtain the classification result of the CKD cohort corpus and the optimized feature correlation matrix.
[0039] Specifically, the CVD event prediction model includes an adaptive one-dimensional convolution kernel and a feedforward neural network. Step S4 includes the following steps: S41. Sort the first feature correlation matrix to obtain a second feature correlation matrix.
[0040] Specifically, taking rows or columns as units, sort the first feature correlation matrix, retain the top two elements with the highest values, and set the remaining element values to 0 to obtain a second feature correlation matrix, and its size is: the number of features × the number of features.
[0041] S42. Perform a dot product operation on the second feature correlation matrix and the first feature vector to obtain a second feature vector.
[0042] Among them, the size of the second feature vector is: the number of features × 1.
[0043] S43. Concatenate the first feature vector and the second feature correlation matrix to obtain a third feature vector.
[0044] Among them, the size of the third feature vector is: the number of features × 2.
[0045] S44. Perform a dot product operation on the third eigenvector and the adaptive one-dimensional convolutional kernel to obtain a fourth eigenvector.
[0046] Among them, the size of the third eigenvector is: the number of features × 1.
[0047] S45. Input the fourth eigenvector into the feedforward neural network to obtain the classification result of the CKD queue corpus and the optimized feature correlation matrix.
[0048] The feature correlation matrix is a real symmetric matrix, and the values inside represent whether there is a causal relationship between features. The causal relationship between CKD queue corpora can be represented by the optimized feature correlation matrix.
[0049] Furthermore, the training, validation, and testing processes of the CVD event prediction model include: A1. Obtain a CKD dataset, where the CKD dataset includes CKD queue corpora and corresponding true classification labels.
[0050] A2. Divide the CKD queue corpora into a training set, a validation set, and a test set according to the ratio of 6:2:2.
[0051] A3. Preprocess the CKD queue corpora, construct a first eigenvector, and construct a first feature correlation matrix.
[0052] A4. Input the first eigenvector and the first feature correlation matrix into the CVD event prediction model for processing to obtain the classification result of the CKD queue corpora.
[0053] A5. Calculate the squared loss function based on the classification result of the CKD queue corpora and the corresponding true classification labels, and iteratively train the CVD event prediction model based on the squared loss function to obtain a trained CVD event prediction model and an optimized feature correlation matrix.
[0054] A6. Validate the CVD event prediction model through the validation set and test the CVD event prediction model through the test set.
[0055] In this application, the CKD queue corpora are preprocessed, a first eigenvector is constructed based on the preprocessed CKD queue corpora, and the constructed first feature correlation matrix and the first eigenvector are processed by the CVD event prediction model to obtain an optimized feature correlation matrix. Through a series of preprocessing of the CKD queue corpora, this application ensures the data integrity of the CKD queue corpora. By describing the causal relationship between CKD queue corpora through the feature correlation matrix, this application can effectively extract the interactive effects between multiple risk factors, providing a theoretical basis for predicting the risk of CKD secondary to CVD.
[0056] According to the method described in the above embodiments, this embodiment will be further described from the perspective of a causal relationship recognition device based on CKD queue data. The causal relationship recognition device based on CKD queue data can be specifically implemented as an independent entity or integrated in an electronic device, which can be a terminal, a server, or other devices. Among them, the terminal can include a tablet computer, a laptop computer, a personal computer (PC), a microprocessing box, or other devices, etc.
[0057] Please refer to Figure 2 , Figure 2 which specifically describes the causal relationship recognition device based on CKD queue data provided in the embodiments of the present application. Applied in an electronic device, the causal relationship recognition device based on CKD queue data may include: An acquisition and preprocessing module, configured to acquire CKD queue corpus and preprocess the CKD queue corpus; A first generation module, configured to construct a first feature vector based on the preprocessed CKD queue corpus; A second generation module, configured to construct an initial feature association matrix and initialize the initial feature association matrix to obtain a first feature association matrix; An identification and optimization module, configured to input the first feature vector and the first feature association matrix into a CVD event prediction model for training processing to obtain a classification result of the CKD queue corpus and an optimized feature association matrix.
[0058] Specifically in implementation, the above-mentioned various modules and / or units can be implemented as independent entities, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of the above-mentioned various modules and / or units, reference can be made to the foregoing method embodiments, and the specific beneficial effects that can be achieved can also be referred to the beneficial effects in the foregoing method embodiments, which will not be elaborated herein.
[0059] In addition, the embodiments of the present application further provide an electronic device, which can be a computer, a tablet computer, or other devices. The electronic device can implement the steps in any of the embodiments of the causal relationship recognition method based on CKD queue data provided in the embodiments of the present application. Therefore, the beneficial effects that can be achieved by any of the causal relationship recognition methods based on CKD queue data provided in the embodiments of the present invention can be realized. For details, please refer to the foregoing embodiments, which will not be elaborated herein.
[0060] Figure 3The specific structural block diagram of the electronic device provided by the embodiment of the present invention is shown. This electronic device can be used to implement the causal relationship identification method based on CKD queue data provided in the above embodiment. The electronic device 500 can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a microprocessing box, or other devices, etc.
[0061] The RF circuit 510 is used to receive and send electromagnetic waves, realize the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices. The RF circuit 510 can include various existing circuit elements for performing these functions. For example, antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity module (SIM) cards, memories, and so on. The RF circuit 510 can communicate with various networks such as the Internet, enterprise intranets, wireless networks, or communicate with other devices through a wireless network. The above wireless network can include a cellular phone network, a wireless local area network, or a metropolitan area network. The above wireless network can use various communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging, and short messages, and any other suitable communication protocols, and even can include those protocols that have not been developed yet.
[0062] The memory 520 can be used to store software programs and modules, such as the corresponding program instructions / modules in the above embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, that is, realizes functions such as taking pictures with the front camera, processing the captured images, and switching the display colors of the display content on the display screen. The memory 520 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 520 may further include a memory remotely disposed relative to the processor 580, and these remote memories can be connected to the electronic device 500 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0063] The input unit 530 can be used to receive input digital or character information, and generate a keyboard and a mouse related to user settings and function controls. The display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, and these graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 540 may include a display panel 541. Optionally, the display panel 541 can be configured in the form of an LCD (Liquid Crystal Display) or an OLED (Organic Light-Emitting Diode).
[0064] The audio circuit 560, the speaker 561, and the microphone 562 can provide an audio interface between the user and the electronic device 500. The audio circuit 560 can transmit the electrical signal converted from the received audio data to the speaker 561, and the speaker 561 converts it into a sound signal for output; on the other hand, the microphone 562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 560 and then converted into audio data. After the audio data is output to the processor 580 for processing, it is sent to another terminal, for example, through the RF circuit 510, or the audio data is output to the memory 520 for further processing. The audio circuit 560 may also include an earphone jack to provide communication between the peripheral earphone and the electronic device 500.
[0065] The electronic device 500 can help the user receive requests, send information, etc. through the transmission module 570 (such as a Wi-Fi module), and it provides the user with wireless broadband Internet access. Although the transmission module 570 is shown in the figure, it can be understood that it does not belong to the essential components of the electronic device 500 and can be omitted completely as needed without changing the essence of the invention.
[0066] The processor 580 is the control center of the electronic device 500, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 520, and by calling the data stored in the memory 520, it performs various functions of the electronic device 500 and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 580 may include one or more processing cores; in some embodiments, the processor 580 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 580 either.
[0067] The electronic device 500 further includes a power supply 590 (such as a battery) for supplying power to each component. In some embodiments, the power supply can be logically connected to the processor 580 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 590 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0068] Although not shown, the electronic device 500 further includes a camera (such as a front camera, a rear camera), a Bluetooth module, etc., which will not be elaborated here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal further includes a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations: Obtain the CKD queue corpus and preprocess the CKD queue corpus; Construct a first feature vector based on the preprocessed CKD queue corpus; Construct an initial feature association matrix and initialize the initial feature association matrix to obtain a first feature association matrix; Input the first feature vector and the first feature association matrix into the CVD event prediction model for training processing to obtain the classification result of the CKD queue corpus and the optimized feature association matrix.
[0069] In specific implementation, the above-mentioned each module can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of the above-mentioned each module, reference can be made to the method embodiments described above, which will not be elaborated here.
[0070] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. For this purpose, an embodiment of the present invention provides a storage medium, which stores multiple instructions that can be loaded by a processor to execute the steps of any one of the embodiments of the causal relationship recognition method based on CKD queue data provided by the embodiments of the present invention.
[0071] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0072] Since the instructions stored in the storage medium can execute the steps in any one of the embodiments of the causal relationship recognition method based on CKD queue data provided by the embodiments of the present invention, the beneficial effects that can be achieved by any of the causal relationship recognition methods based on CKD queue data provided by the embodiments of the present invention can be realized. For details, please refer to the previous embodiments and will not be repeated here.
[0073] The above has introduced in detail a causal relationship recognition method, device, storage medium, and electronic device based on CKD queue data provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A causal relationship identification method based on CKD cohort data, characterized in that: The method comprises: Acquire CKD cohort corpus, and preprocess the CKD cohort corpus; Construct the first feature vector based on the preprocessed CKD cohort corpus; Constructing an initial feature association matrix, and initializing the initial feature association matrix to obtain a first feature association matrix; The first feature vector and the first feature association matrix are input into a CVD event prediction model for training processing to obtain a classification result of the CKD cohort corpus and an optimized feature association matrix.
2. The causal relationship identification method based on CKD cohort data according to claim 1, characterized in that: The CKD cohort corpus is preprocessed, including: Deleting invalid feature columns in the CKD cohort corpus; Handle outliers of numeric features in the CKD cohort corpus after deleting invalid feature columns; Handling symbolic feature outliers in the CKD cohort corpus after removing invalid feature columns.
3. The causal relationship identification method based on CKD cohort data according to claim 1, characterized in that: The constructing a first feature vector based on the preprocessed CKD cohort corpus includes: The preprocessed CKD cohort corpus was standardized to obtain the first feature vector.
4. The causal relationship identification method based on CKD cohort data according to claim 1, characterized in that: The CVD event prediction model includes an adaptive one-dimensional convolution kernel and a feedforward neural network; The step of inputting the first feature vector and the first feature association matrix into a CVD event prediction model for training processing to obtain the classification result of the CKD cohort corpus and the optimized feature association matrix includes: Sorting the first feature association matrix to obtain a second feature association matrix; Performing a dot product operation on the second characteristic association matrix and the first characteristic vector to obtain a second characteristic vector; Concatenate the first eigenvector and the second eigenvector to obtain a third eigenvector; Performing a dot product operation on the third eigenvector and the adaptive one-dimensional convolution kernel to obtain a fourth eigenvector; The fourth feature vector is input into the feedforward neural network to obtain the classification result of the CKD cohort corpus and the optimized feature association matrix.
5. The causal relationship identification method based on CKD cohort data according to claim 1, characterized in that: The training process of the CVD event prediction model includes: Obtain a CKD dataset, wherein the CKD dataset includes a CKD cohort corpus and corresponding true classification labels; Preprocess the CKD cohort corpus, construct the first feature vector, and construct the first feature association matrix; Inputting the first feature vector and the first feature association matrix into a CVD event prediction model for processing to obtain a classification result of the CKD cohort corpus; The square loss function is calculated based on the classification results of the CKD cohort corpus and the corresponding true classification labels, and the CVD event prediction model is iteratively trained based on the square loss function to obtain a trained CVD event prediction model and an optimized feature association matrix.
6. The causal relationship identification method based on CKD cohort data according to claim 2, characterized in that: The processing of the numerical feature outliers in the CKD cohort corpus after deleting the invalid feature columns includes: Detect missing values and illegal values in the numeric features in the CKD cohort corpus, and fill the missing values and illegal values with the median or mean; Constructing a normal range table of digital features, and judging whether each digital feature is out of bounds in turn according to the normal range table; For numerical features that are out of bounds, use the maximum, minimum, or median number for filling.
7. The causal relationship identification method based on CKD cohort data according to claim 2, characterized in that: The processing of symbolic feature outliers in the CKD cohort corpus after deleting invalid feature columns includes: Detect missing values and illegal values in symbolic features in the CKD cohort corpus and process them according to the most recent occurrence principle or the most frequent occurrence principle; Replace the symbols in symbolic features.
8. A causal relationship identification device based on CKD cohort data, characterized in that: include: An acquisition and preprocessing module, used for acquiring CKD cohort corpus and preprocessing the CKD cohort corpus; A first generating module, configured to construct a first feature vector based on the preprocessed CKD cohort corpus; The second generating module is used to construct an initial feature association matrix and initialize the initial feature association matrix to obtain a first feature association matrix; The identification and optimization module is used to input the first feature vector and the first feature association matrix into the CVD event prediction model for training processing to obtain the classification result of the CKD cohort corpus and the optimized feature association matrix.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the causal relationship identification method based on CKD cohort data as described in any one of claims 1 to 7.
10. An electronic device, characterized in that: It includes a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to execute the steps in the causal relationship identification method based on CKD cohort data as described in any one of claims 1 to 7.