Blood vessel data completion method, device, equipment, medium and product

By encoding and decoding the missing categorical variables in the vascular data and using the vascular data completion model to achieve interpolation, the problem that the existing technology cannot complete categorical variables is solved, and the completion accuracy of vascular data is improved.

CN120045859APending Publication Date: 2025-05-27SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510510900.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing vascular data completion technology cannot achieve the completion of categorical variables, and can only process numerical variables.

Method used

By encoding the classified variables with missing data, input them into the trained vascular data completion model, interpolation and decoding, the corresponding interpolation data is obtained, thereby achieving the completion of the classified variables.

Benefits of technology

Automatic completion of classified variables in vascular data is achieved, and the accuracy and reliability of vascular data completion is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045859A_ABST
    Figure CN120045859A_ABST
Patent Text Reader

Abstract

The invention discloses a blood vessel data completion method, device and equipment, a medium and a product. The invention relates to the technical field of big data. The method comprises the steps that blood vessel data with data missing are acquired, and the blood vessel data with the data missing comprise classification variables with the data missing; for the classification variables with data missing, encoding the classification variables with data missing to obtain encoded classification variables, inputting the encoded classification variables into the trained blood vessel data completion model to obtain interpolation classification variables to be decoded, and decoding the interpolation classification variables to be decoded to obtain interpolation classification variables to be decoded. Interpolation data corresponding to the classification variables with data missing are obtained; and determining a classification variable of data completion based on the classification variable with data missing and the interpolation data corresponding to the classification variable with data missing. According to the scheme, automatic completion of the blood vessel data is realized through coding, the blood vessel data completion model and decoding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of big data technology, and in particular to a vascular data completion method, device, equipment, medium and product. Background Art

[0002] In cardiovascular and cerebrovascular disease risk assessment, due to technical limitations or disease progression, vascular data may be partially missing or unavailable, affecting the accuracy and reliability of risk assessment.

[0003] In current vascular data completion technology, attention mechanisms and Transformer models are used to selectively focus on relevant variables and complete missing data in vascular datasets.

[0004] In the process of implementing the present disclosure, there are at least the following technical problems in the prior art: the existing vascular data completion technology solution can only complete the numerical variables, but cannot complete the categorical variables. Summary of the invention

[0005] The present disclosure provides a vascular data completion method, device, equipment, medium and product, which realize the completion of classification variables in vascular data.

[0006] According to one aspect of the present disclosure, a method for completing blood vessel data is provided, comprising:

[0007] Acquiring blood vessel data with missing data, wherein the blood vessel data with missing data includes a categorical variable with missing data;

[0008] For the categorical variable with missing data, encode the categorical variable with missing data to obtain the encoded categorical variable, input the encoded categorical variable into the trained vascular data completion model to obtain an interpolation categorical variable to be decoded, and decode the interpolation categorical variable to be decoded to obtain interpolation data corresponding to the categorical variable with missing data;

[0009] The categorical variable with data completion is determined based on the categorical variable with missing data and the interpolation data corresponding to the categorical variable with missing data.

[0010] According to another aspect of the present disclosure, a blood vessel data completion device is provided, comprising:

[0011] A data-missing vascular data acquisition module, used to acquire data-missing vascular data, wherein the data-missing vascular data includes categorical variables with missing data;

[0012] a module for determining interpolation data corresponding to categorical variables, for the categorical variables with missing data, encoding the categorical variables with missing data to obtain the encoded categorical variables, inputting the encoded categorical variables into the trained vascular data completion model to obtain interpolation categorical variables to be decoded, decoding the interpolation categorical variables to be decoded, and obtaining interpolation data corresponding to the categorical variables with missing data;

[0013] The data-completing categorical variable determination module is used to determine the data-completing categorical variable based on the categorical variable with missing data and the interpolation data corresponding to the categorical variable with missing data.

[0014] According to another aspect of the present disclosure, an electronic device is provided, the electronic device comprising:

[0015] at least one processor;

[0016] and a memory communicatively coupled to the at least one processor;

[0017] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the blood vessel data completion method described in any embodiment of the present disclosure.

[0018] According to another aspect of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the blood vessel data completion method described in any embodiment of the present disclosure when executed.

[0019] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the blood vessel data completion method as described in any one of the embodiments of the present disclosure is implemented.

[0020] The technical solution of the embodiment of the present disclosure is to obtain vascular data with missing data, wherein the vascular data with missing data includes categorical variables with missing data; for categorical variables with missing data, encode the categorical variables with missing data to obtain the encoded categorical variables, input the encoded categorical variables into the trained vascular data completion model to obtain the interpolation categorical variables to be decoded, decode the interpolation categorical variables to be decoded to obtain the interpolation data corresponding to the categorical variables with missing data; determine the categorical variables for data completion based on the categorical variables with missing data and the interpolation data corresponding to the categorical variables with missing data. The above solution realizes automatic completion of vascular data through encoding, vascular data completion model and decoding.

[0021] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 is a flow chart of a blood vessel data completion method provided according to an embodiment of the present disclosure;

[0024] Figure 2 is a flow chart of another blood vessel data completion method provided according to an embodiment of the present disclosure;

[0025] Figure 3 is a schematic diagram of categorical variable encoding provided according to an embodiment of the present disclosure;

[0026] Figure 4 is a flow chart of another blood vessel data completion method provided according to an embodiment of the present disclosure;

[0027] Figure 5 is a flow chart of a training process of a vascular data completion model provided according to an embodiment of the present disclosure;

[0028] Figure 6 is a flow chart of another blood vessel data completion method provided according to an embodiment of the present disclosure;

[0029] Figure 7 is a structural schematic diagram of a blood vessel data completion device provided according to an embodiment of the present disclosure;

[0030] Figure 8 It is a structural schematic diagram of an electronic device for implementing the blood vessel data completion method of an embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] In order to enable those skilled in the art to better understand the disclosed solution, the technical solution in the disclosed embodiment will be clearly and completely described below in conjunction with the drawings in the disclosed embodiment. Obviously, the described embodiment is only a part of the disclosed embodiment, not all of the embodiments. Based on the embodiments in the disclosed embodiment, all other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the disclosed embodiment.

[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. The acquisition, storage, use, processing, etc. of data in the technical solution of the present disclosure comply with the relevant provisions of national laws and regulations.

[0033] The technical field and related terms of the embodiments of the present disclosure are briefly described below.

[0034] Vascular data: It can be cardiovascular data around the surgical period, including preoperative, intraoperative and postoperative stages. Its monitoring indicators may include but are not limited to blood pressure, blood sugar and blood lipids.

[0035] Numerical variables: The variable value is quantitative, expressed as a numerical value, can be obtained through measurement, and often has a unit of measurement, such as height (cm), weight (kg), and blood pressure (mmHg).

[0036] Categorical variables: variables used to describe the attributes or characteristics of things. For example, "gender" is a categorical variable, and its variable value is "male" or "female", and "drinking status" is also a categorical variable, and its variable value is "drinking", "drinking a little" or "drinking a lot".

[0037] Figure 1 This is a flow chart of a blood vessel data completion method provided by an embodiment of the present disclosure. This embodiment is applicable to the case of automatic blood vessel data completion. The method can be executed by a blood vessel data completion device. The blood vessel data completion device can be implemented in the form of hardware and / or software. The blood vessel data completion device can be configured in electronic devices such as terminals and servers. Figure 1 As shown, the method includes:

[0038] S110 . Acquire blood vessel data with missing data, wherein the blood vessel data with missing data includes categorical variables with missing data.

[0039] The vascular data with missing data refers to incomplete vascular data, which may include categorical variables and numerical variables. For example, the vascular data may include numerical variables that are missing part of the blood pressure data and categorical variables that are missing part of the smoking and drinking status.

[0040] Exemplarily, the vascular data with missing data can be read from a preset storage path of the electronic device, or can be obtained from other devices or the cloud that are communicatively connected to the electronic device, and no specific limitation is made here.

[0041] S120. For the categorical variable with missing data, encode the categorical variable with missing data to obtain the encoded categorical variable, input the encoded categorical variable into the trained vascular data completion model to obtain the interpolation categorical variable to be decoded, decode the interpolation categorical variable to be decoded, and obtain the interpolation data corresponding to the categorical variable with missing data.

[0042] Among them, encoding is the process of converting categorical variables with missing data from discrete variables to other forms or formats to facilitate the reading and use of the vascular data completion model.

[0043] Specifically, the categorical variables with missing data can be converted into coded categorical variables in the form of binary or embedded vectors, etc. For example, assuming that the categorical variable is drinking status, the drinking status can include light, moderate and heavy drinking, light drinking can be encoded as [0, 0, 1], moderate drinking can be encoded as [0, 1, 0], and heavy drinking can be encoded as [1, 0, 0].

[0044] In the disclosed embodiment, the vascular data completion model refers to an artificial intelligence model used to predict interpolation data, where the interpolation data is data missing from the original vascular data. In other words, the vascular data can be completed by predicting the interpolation data.

[0045] The vascular data completion model can be a prediction model based on a conditional diffusion model or other neural network-based construction. The training steps of the vascular data completion model can include: obtaining a vascular training data set, which can include numerical variables, categorical variables, and missing values, and then training the initial neural network based on the vascular training data set to obtain the vascular data completion model.

[0046] Decoding is the process of converting an imputed categorical variable from another form or format into a discrete variable for subsequent vascular disease risk assessment. An imputed categorical variable refers to data to be decoded that can be decoded into imputed data.

[0047] Specifically, the interpolated categorical variables in the form of binary or embedded vectors can be converted into discrete variables. For example, assuming that the interpolated categorical variable is drinking status, the drinking status can include light, moderate and heavy, [0,0,1] can be decoded as light, [0,1,0] can be decoded as moderate, and [0,0,1] can be decoded as heavy.

[0048] S130. Determine a categorical variable with data completion based on the categorical variable with missing data and the interpolation data corresponding to the categorical variable with missing data.

[0049] The categorical variables of data completion refer to the complete categorical variables in the vascular data.

[0050] Specifically, the interpolated data can be spliced ​​or inserted into the categorical variable with missing data to obtain a categorical variable with complete data.

[0051] On the basis of the above embodiments, optionally, the vascular data with missing data also includes numerical variables with missing data; accordingly, after obtaining the vascular data with missing data, it also includes: for the numerical variables with missing data, inputting the numerical variables with missing data into a trained vascular data completion model to obtain interpolated data corresponding to the numerical variables with missing data.

[0052] It should be noted that numerical variables can be directly input into the vascular data completion model without encoding and decoding operations.

[0053] Exemplarily, assuming that the numerical variable is blood pressure data with missing data, the blood pressure data with missing data can be input into the vascular data completion model, and the vascular data completion model restores the blood pressure data with missing data. The vascular data completion model can output the missing blood pressure data, that is, output the interpolated data.

[0054] The technical solution of the embodiment of the present disclosure is to obtain vascular data with missing data, wherein the vascular data with missing data includes categorical variables with missing data; for categorical variables with missing data, encode the categorical variables with missing data to obtain the encoded categorical variables, input the encoded categorical variables into the trained vascular data completion model to obtain the interpolation categorical variables to be decoded, decode the interpolation categorical variables to be decoded to obtain the interpolation data corresponding to the categorical variables with missing data; determine the categorical variables for data completion based on the categorical variables with missing data and the interpolation data corresponding to the categorical variables with missing data. The above solution realizes automatic completion of vascular data through encoding, vascular data completion model and decoding.

[0055] Figure 2 This is a flow chart of another vascular data completion method provided by an embodiment of the present disclosure. The method of this embodiment can be combined with the various optional solutions in the vascular data completion method provided in the above embodiments. Based on the above embodiments, this embodiment further refines the step of encoding the categorical variables with missing data.

[0056] like Figure 2 As shown, the method includes:

[0057] S210. Acquire blood vessel data with missing data, wherein the blood vessel data with missing data includes categorical variables with missing data.

[0058] S220. For the categorical variable with missing data, encode the categorical variable with missing data to obtain the encoded categorical variable, wherein the encoding of the categorical variable with missing data includes one or more of the following steps: one-hot encoding the categorical variable with missing data; analog bit encoding the categorical variable with missing data; embedded encoding the categorical variable with missing data.

[0059] S230, inputting the encoded categorical variables into the trained vascular data completion model to obtain interpolation categorical variables to be decoded, decoding the interpolation categorical variables to be decoded, and obtaining interpolation data corresponding to the categorical variables with missing data.

[0060] S240: Determine a categorical variable with data completion based on the categorical variable with missing data and the interpolation data corresponding to the categorical variable with missing data.

[0061] Among them, one-hot encoding can create a new binary feature for each possible value of a discrete variable, where only one feature can be activated (marked as 1) at any given time, and all other features are marked as 0. Analog bits encoding is used to convert discrete variables into binary bit data, where each category is assigned a unique fixed-length binary code. Embedding encoding is used to convert all discrete variables into embedding vectors of the same length.

[0062] For example, Figure 3 A schematic diagram of a categorical variable encoding provided for an embodiment of the present disclosure. Assuming that the categorical variable is 3, when the encoding method is one-hot encoding, the categorical variable 3 can be converted into binary bit data 100, and the binary bit data 100 indicates that the third category is selected, represented by 1, and the other two categories are not selected, represented by 0. When the encoding method is analog bit encoding, the categorical variable 3 can be converted into binary bit data 11, and the binary bit data 11 indicates that the third category is selected. Compared with one-hot encoding, fewer bits can be used to represent the category. When the encoding method is embedded encoding, the categorical variable 3 can be converted into E3, and E3 is the embedded representation of the categorical variable 3.

[0063] On the basis of the above embodiments, optionally, the interpolation categorical variable to be decoded is decoded to obtain the interpolation data corresponding to the categorical variable with missing data, including: when the encoding method is one-hot encoding, the category corresponding to the index with the largest value in the interpolation categorical variable to be decoded is used as the interpolation data corresponding to the categorical variable with missing data; when the encoding method is analog bit encoding, for any element in the interpolation categorical variable to be decoded, if the element is greater than zero, the element is updated to 1, and if the element is not greater than zero, the element is updated to -1; the categorical variable after the element is updated is mapped to obtain the interpolation data corresponding to the categorical variable with missing data; when the encoding method is embedded encoding, the Euclidean distance between the interpolation categorical variable to be decoded and the embedded vector corresponding to each category is determined, and the category with the smallest Euclidean distance is used as the interpolation data corresponding to the categorical variable with missing data.

[0064] For example, assuming that the interpolation categorical variable is drinking status, which may include mild, moderate, and severe drinking, the interpolation categorical variable may be [1,0,0], 1 being the largest element in the interpolation categorical variable, 1 corresponding to index 3, 3 corresponding to category severe, that is, the interpolation data obtained by decoding is severe drinking. In the case where the encoding method is analog bit encoding, the interpolation categorical variable may be [2,1]. According to the above threshold judgment update rule, [2,1] may be updated to [1,1], and then [1,1] may be decoded back to the severe category according to the predefined mapping relationship, and severe is used as the interpolation data, wherein the predefined mapping relationship includes: [0,1] corresponding to mild, [1,0] corresponding to moderate, and [1,1] corresponding to severe. When the encoding method is embedded encoding, the interpolation classification variable can be E3, the embedding vector corresponding to the mild condition can be e1, the embedding vector corresponding to the moderate condition can be e2, and the embedding vector corresponding to the severe condition can be e3. The Euclidean distances between E3 and e1, e2, and e3 are calculated respectively. If the Euclidean distance between E3 and e3 is the smallest, the severe category corresponding to e3 is used as the interpolation data.

[0065] It should be noted that, when the encoding method is embedded encoding, the numerical variable can be embedded encoded so that the length of the numerical variable is the same as that of the categorical variable. Then, the numerical variable after embedding encoding can be input into the trained vascular data completion model to obtain the numerical variable to be decoded, and then each element in the numerical variable after embedding encoding is divided by the element at the corresponding position in the numerical variable to be decoded, and then the average value of the division results of each element is calculated, and the average value is used as the interpolation data.

[0066] The technical solution of the disclosed embodiment can enhance the adaptability of the vascular data completion model to different types of data through three categorical variable encoding technologies: one-hot encoding, analog bit encoding, and embedded encoding. This enables the vascular data completion model to not only process traditional numerical data, but also effectively process categorical data, providing a comprehensive and flexible missing value interpolation strategy for vascular data completion.

[0067] Figure 4 This is a flow chart of another vascular data completion method provided by an embodiment of the present disclosure. The method of this embodiment can be combined with the various optional solutions in the vascular data completion method provided in the above embodiments. Based on the above embodiments, this embodiment further refines the training process of the vascular data completion model.

[0068] like Figure 4 As shown, the method includes:

[0069] S310 , obtaining blood vessel training data, wherein the blood vessel training data includes observed data and missing data, and the observed data includes conditional observed values ​​and interpolation targets.

[0070] S320: Add noise to the interpolation target to obtain the interpolation target after the noise is added.

[0071] S330, inputting the interpolation target after adding noise and the conditional observation value into the conditional diffusion model to be trained to obtain predicted noise.

[0072] S340 , determining a model loss based on the predicted noise and the real noise, and updating model parameters of the conditional diffusion model based on the model loss until a training stop condition is met, thereby obtaining a vascular data completion model.

[0073] Among them, the observed data refers to the actual vascular sample data collected from the patient during the perioperative period. Missing data refers to the vascular sample data that was not collected during the perioperative period. In the embodiment of the present disclosure, the observed data can be divided into conditional observations and interpolation targets. The conditional observations refer to the data observed under specific conditions or constraints, and the interpolation targets are the data in the observed data other than the conditional observations.

[0074] For example, Figure 5 Flow chart of a training process of a vascular data completion model provided by an embodiment of the present disclosure. Figure 5As shown in the figure, the green area in the rectangle represents the observed data, the white area in the rectangle represents the missing data, the blue area represents the conditional observed value, the red diagonal area represents the interpolation target, and the green grid area represents the interpolation target after adding noise. The interpolation target and the conditional observed value after adding noise are input into the conditional diffusion model to be trained to obtain the predicted noise, thereby minimizing the loss between the predicted noise and the true noise, and updating the model parameters of the conditional diffusion model until the training stop condition is met to obtain the vascular data completion model.

[0075] S350: Acquire blood vessel data with missing data, wherein the blood vessel data with missing data includes categorical variables with missing data.

[0076] S360. For the categorical variable with missing data, encode the categorical variable with missing data to obtain the encoded categorical variable, input the encoded categorical variable into the trained vascular data completion model to obtain the interpolation categorical variable to be decoded, decode the interpolation categorical variable to be decoded, and obtain the interpolation data corresponding to the categorical variable with missing data.

[0077] S370: Determine a categorical variable for data completion based on the categorical variable with missing data and the interpolation data corresponding to the categorical variable with missing data.

[0078] Exemplarily, the present disclosure compares the performance of the vascular data completion model based on the conditional diffusion model with the existing data completion methods: mean / mode, multiple interpolation based on chain equations (MICE(linear)), multiple interpolation based on random forest (MissForest) and generative adversarial interpolation network (GAIN). The performance comparison uses two evaluation indicators: root mean square error (RMSE) and error rate. Diaberes represents the vascular data of diabetic patients, and COVID-19 represents the vascular data of patients with the new coronavirus. The root mean square error is used to evaluate the interpolation accuracy of numerical variables, and the error rate is used to evaluate the interpolation accuracy of categorical variables. One-hot represents unique hot encoding, Analog bits represents analog bit encoding, and FT represents embedded encoding.

[0079] Table 1

[0080]

[0081] The comparison results are shown in Table 1. The experimental results show that one-hot encoding, simulated bit encoding, and embedded encoding all perform well in processing categorical variables. The conditional diffusion model can better understand and predict the data relationships in vascular data, thereby improving the quality and accuracy of interpolated data.

[0082] The technical solution of the disclosed embodiment obtains a vascular data completion model by training a conditional diffusion model, thereby effectively improving the accuracy of the interpolated data.

[0083] Figure 6 This is a flow chart of another vascular data completion method provided by an embodiment of the present disclosure. The method of this embodiment is a preferred example of the above embodiment. The numerical variable includes at least one of blood sugar, blood lipids and blood pressure; the categorical variable includes at least one of smoking status, drinking status and exercise status. Figure 6 As shown, the method includes:

[0084] Step 1: Collect multimodal data.

[0085] Specifically, numerical variables such as blood sugar, blood lipids and blood pressure can be collected through blood sugar collection equipment, blood lipid collection equipment, blood pressure collection equipment, etc. Categorical variables such as smoking status, drinking status and exercise status can be input through input devices such as keyboard, mouse or touch screen. It should be noted that the above collection or input may be incomplete or the input data may be incomplete.

[0086] Step 2: Preprocess the multimodal data.

[0087] Specifically, preprocessing operations such as outlier filtering and normalization can be performed on the multimodal data to obtain standard vascular data with missing data.

[0088] Step 3: Input the vascular data with missing data into the vascular data completion model to obtain interpolated data.

[0089] Specifically, for categorical variables with missing data, the categorical variables with missing data are encoded to obtain encoded categorical variables, the encoded categorical variables are input into the trained vascular data completion model to obtain interpolated categorical variables to be decoded, the interpolated categorical variables to be decoded are decoded to obtain interpolated data corresponding to the categorical variables with missing data. For numerical variables with missing data, the numerical variables with missing data are input into the trained vascular data completion model to obtain interpolated data corresponding to the numerical variables with missing data.

[0090] Step 4: Complete the vascular data based on the interpolated data.

[0091] Specifically, the interpolation data is inserted into the blood vessel data with missing data to obtain the completed blood vessel data.

[0092] The technical solution of the embodiment of the present disclosure realizes automatic completion of vascular data through encoding, vascular data completion model and decoding.

[0093] Figure 7 FIG. 1 is a schematic diagram of a blood vessel data completion device provided by an embodiment of the present disclosure. Figure 7 As shown, the device comprises:

[0094] A data-missing vascular data acquisition module 510, used to acquire data-missing vascular data, wherein the data-missing vascular data includes categorical variables with missing data;

[0095] The interpolation data determination module 520 corresponding to the categorical variable is used for encoding the categorical variable with missing data to obtain the encoded categorical variable, inputting the encoded categorical variable into the trained vascular data completion model to obtain the interpolation categorical variable to be decoded, decoding the interpolation categorical variable to be decoded, and obtaining the interpolation data corresponding to the categorical variable with missing data;

[0096] The data-completed categorical variable determination module 530 is used to determine the data-completed categorical variable based on the categorical variable with missing data and the interpolated data corresponding to the categorical variable with missing data.

[0097] The technical solution of the embodiment of the present disclosure is to obtain vascular data with missing data, wherein the vascular data with missing data includes categorical variables with missing data; for categorical variables with missing data, encode the categorical variables with missing data to obtain the encoded categorical variables, input the encoded categorical variables into the trained vascular data completion model to obtain the interpolation categorical variables to be decoded, decode the interpolation categorical variables to be decoded to obtain the interpolation data corresponding to the categorical variables with missing data; determine the categorical variables for data completion based on the categorical variables with missing data and the interpolation data corresponding to the categorical variables with missing data. The above solution realizes automatic completion of vascular data through encoding, vascular data completion model and decoding.

[0098] Based on any optional technical solution in the embodiments of the present disclosure, optionally, encoding the categorical variable with missing data includes one or more of the following steps:

[0099] Perform one-hot encoding on the categorical variables with missing data;

[0100] Performing simulated bit coding on the categorical variables with missing data;

[0101] Embedding coding is performed on the categorical variables with missing data.

[0102] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the interpolation data determination module 520 corresponding to the categorical variable includes:

[0103] A decoding unit corresponding to one-hot encoding, used for, when the encoding mode is one-hot encoding, taking the category corresponding to the index with the largest value in the interpolation categorical variable to be decoded as the interpolation data corresponding to the categorical variable with missing data;

[0104] The decoding unit corresponding to the analog bit coding is used for, when the coding mode is the analog bit coding, for any element in the interpolation classification variable to be decoded, if the element is greater than zero, then updating the element to 1, if the element is not greater than zero, then updating the element to -1; mapping the classification variable after the element is updated to obtain the interpolation data corresponding to the classification variable with missing data;

[0105] The decoding unit corresponding to the embedded coding is used to determine the Euclidean distance between the interpolation categorical variable to be decoded and the embedded vector corresponding to each category when the encoding method is embedded coding, and take the category with the smallest Euclidean distance as the interpolation data corresponding to the categorical variable with missing data.

[0106] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the training process of the blood vessel data completion model includes:

[0107] Acquire blood vessel training data, wherein the blood vessel training data includes observed data and missing data, and the observed data includes conditional observed values ​​and interpolation targets;

[0108] adding noise to the interpolation target to obtain an interpolation target after the noise is added;

[0109] Inputting the interpolation target after adding noise and the conditional observation value into the conditional diffusion model to be trained to obtain predicted noise;

[0110] A model loss is determined based on the predicted noise and the real noise, and model parameters of the conditional diffusion model are updated based on the model loss until a training stop condition is met, thereby obtaining a vascular data completion model.

[0111] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the blood vessel data with missing data further includes a numerical variable with missing data;

[0112] Accordingly, the vascular data completion device includes:

[0113] The numerical variable interpolation module is used to input the numerical variable with missing data into the trained vascular data completion model to obtain interpolation data corresponding to the numerical variable with missing data.

[0114] Based on any optional technical solution in the embodiments of the present disclosure, optionally, the numerical variable includes at least one of heart rate, blood pressure, body temperature and respiratory rate; the categorical variable includes at least one of smoking status, drinking status and exercise status.

[0115] The vascular data completion device provided in the embodiments of the present disclosure can execute the vascular data completion method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0116] Figure 8 A schematic diagram of an electronic device 10 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0117] like Figure 8 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The I / O interface 15 is also connected to the bus 14.

[0118] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0119] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a vascular data completion method, which includes:

[0120] Acquiring blood vessel data with missing data, wherein the blood vessel data with missing data includes a categorical variable with missing data;

[0121] For the categorical variable with missing data, encode the categorical variable with missing data to obtain the encoded categorical variable, input the encoded categorical variable into the trained vascular data completion model to obtain an interpolation categorical variable to be decoded, and decode the interpolation categorical variable to be decoded to obtain interpolation data corresponding to the categorical variable with missing data;

[0122] The categorical variable with data completion is determined based on the categorical variable with missing data and the interpolation data corresponding to the categorical variable with missing data.

[0123] In some embodiments, the vascular data completion method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the vascular data completion method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the vascular data completion method in any other appropriate manner (e.g., by means of firmware).

[0124] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0125] Computer programs for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0126] In the context of the present disclosure, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device, or equipment. A computer-readable storage medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0127] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0128] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0129] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0130] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of this disclosure can be achieved, and this document is not limited here.

[0131] The embodiments of the present disclosure also provide a computer program product, including a computer program, which, when executed by a processor, implements the blood vessel data completion method provided in any embodiment of the present disclosure.

[0132] In the process of implementation, the computer program product can be written in one or more programming languages ​​or a combination thereof to perform the computer program code for the disclosed operation, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., using an Internet service provider to connect through the Internet).

[0133] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication unit 19, or installed from the storage unit 18, or installed from the ROM 12. When the computer program is executed by the processor 11, the above-mentioned functions defined in the method of the embodiment of the present invention are executed.

[0134] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A blood vessel data completion method, characterized in that: include: Acquiring blood vessel data with missing data, wherein the blood vessel data with missing data includes a categorical variable with missing data; For the categorical variable with missing data, encode the categorical variable with missing data to obtain the encoded categorical variable, input the encoded categorical variable into the trained vascular data completion model to obtain an interpolation categorical variable to be decoded, and decode the interpolation categorical variable to be decoded to obtain interpolation data corresponding to the categorical variable with missing data; The categorical variable with data completion is determined based on the categorical variable with missing data and the interpolation data corresponding to the categorical variable with missing data.

2. The method according to claim 1, characterized in that The encoding of the categorical variable with missing data comprises one or more of the following steps: Perform one-hot encoding on the categorical variables with missing data; Performing simulated bit coding on the categorical variables with missing data; Embedding coding is performed on the categorical variables with missing data.

3. The method according to claim 2, characterized in that The decoding of the interpolation categorical variable to be decoded to obtain the interpolation data corresponding to the categorical variable with missing data includes: When the encoding method is one-hot encoding, the category corresponding to the index with the largest value in the interpolation categorical variable to be decoded is used as the interpolation data corresponding to the categorical variable with missing data; In the case where the encoding method is analog bit encoding, for any element in the interpolation categorical variable to be decoded, if the element is greater than zero, the element is updated to 1, and if the element is not greater than zero, the element is updated to -1; the categorical variable after the element is updated is mapped to obtain the interpolation data corresponding to the categorical variable with missing data; When the encoding method is embedded encoding, the Euclidean distance between the interpolation categorical variable to be decoded and the embedded vector corresponding to each category is determined, and the category with the smallest Euclidean distance is used as the interpolation data corresponding to the categorical variable with missing data.

4. The method according to claim 1, characterized in that: The training process of the vascular data completion model includes: Acquire blood vessel training data, wherein the blood vessel training data includes observed data and missing data, and the observed data includes conditional observed values ​​and interpolation targets; adding noise to the interpolation target to obtain an interpolation target after the noise is added; Inputting the interpolation target after adding noise and the conditional observation value into the conditional diffusion model to be trained to obtain predicted noise; A model loss is determined based on the predicted noise and the real noise, and model parameters of the conditional diffusion model are updated based on the model loss until a training stop condition is met, thereby obtaining a vascular data completion model.

5. The method according to claim 1, characterized in that The blood vessel data with missing data also includes numerical variables with missing data; Accordingly, after obtaining the blood vessel data with missing data, the following steps are also included: For the numerical variables with missing data, the numerical variables with missing data are input into the trained blood vessel data completion model to obtain interpolation data corresponding to the numerical variables with missing data.

6. The method according to claim 5, characterized in that The numerical variables include at least one of blood pressure, blood lipids and blood sugar; the categorical variables include at least one of smoking status, drinking status and exercise status.

7. A blood vessel data completion device, characterized in that: include: A data-missing vascular data acquisition module, used to acquire data-missing vascular data, wherein the data-missing vascular data includes categorical variables with missing data; a module for determining interpolation data corresponding to categorical variables, for the categorical variables with missing data, encoding the categorical variables with missing data to obtain the encoded categorical variables, inputting the encoded categorical variables into the trained vascular data completion model to obtain interpolation categorical variables to be decoded, decoding the interpolation categorical variables to be decoded, and obtaining interpolation data corresponding to the categorical variables with missing data; The data-completing categorical variable determination module is used to determine the data-completing categorical variable based on the categorical variable with missing data and the interpolation data corresponding to the categorical variable with missing data.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the blood vessel data completion method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the blood vessel data completion method according to any one of claims 1 to 6 when executed.

10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the blood vessel data completion method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Traffic flow completion and prediction method

    CN110555018A

  • Self-supervised trajectory completion method in general scene

    CN118550907A

  • Method, device and equipment for complementing missing data based on support vector machine and storage medium

    CN118733980A

  • Point cloud completion method and system based on self-supervised conditional diffusion model

    CN119599904A